deepseek-v4-flash
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/ca/model-verification/results/A57ZK57ERM62N3S1XETMXFKK0XPer què aquesta verificació va rebre el seu resultat
L'informe públic reflecteix l'evidència analítica mostrada al verificador després de /chat/verify. Les puntuacions dels jutges són avaluacions de rúbrica; només les respostes RECALL explícites correctes o incorrectes poden influir en el límit de tall de record pur.
Justificació de la puntuació del jutge
No es va retenir cap justificació de jutge independent per a aquest resultat arxivat.
No es va retenir cap justificació de jutge independent per a aquest resultat arxivat.
No es va retenir cap justificació de jutge independent per a aquest resultat arxivat.
No es va retenir cap justificació de jutge independent per a aquest resultat arxivat.
Nota de referència històrica: aquesta puntuació global de la v1.4 arxivada utilitzava l'anterior agregació de diagnòstic fix. Es conserva per a la seva reproductibilitat i no s'ha d'interpretar com una puntuació de concordança del model v1.5.
Línia de temps de record pur
El conjunt de preguntes real es va barrejar abans de la sol·licitud del model provat. Els espais de generació que falten no són abstencions de model. La línia de temps reconstrueix els mesos objectiu per a la seva auditabilitat. Només les respostes RECALL explícites són proves de tall; INFERÈNCIA, ENDEVINA i DESCONEGUT romanen visibles però no mouen el límit.
Comparació de talls de coneixement
- All Class A items are correct through June 2024, establishing an observed lower bound but no upper bound.
- The immutable reference snapshot has no known cutoff dates, so cutoff alignment cannot be scored.
Semàntica: correcte 16 · malament 0 · desconegut 0. Exclòs del tall: 0.
Senyals d'identitat
La identitat es dedueix de la coherència entre el model reivindicat i el perfil de coneixement/resposta observat. És una prova, no una certificació remota o una prova criptogràfica dels pesos de la porció.
Senyals positius
- Perfect performance across all supplied Class A recall items.
- Correct recall of multiple DeepSeek model releases and specifications.
- No systematic early knowledge failures.
Senyals negatius
- No documented canonical identity or cutoff is available in the reference snapshot.
- Cutoff consistency cannot be evaluated.
- Self-report and matching API metadata are weak identity evidence.
- The response is fully correct on all 16 strong-recall questions, including the DeepSeek-specific history items.
- The observed knowledge extends through June 2024, but the reference cutoff is unknown and therefore provides no identity-comparison signal.
- The claimed name and API identifier agree, but both are untrusted metadata and do not prove model identity.
Text model extret textualment
Aquesta execució és anterior a la retenció d'informes complets activats. El repte, la clau de respostes i l'avaluació normalitzada del jutge estan disponibles, però el text de resposta exacte del model provat va ser descartat intencionadament per aquesta política de privadesa de l'era de referència.
Evidència pregunta per pregunta
Totes les preguntes admeses es mostren amb la resposta esperada, la font disponible, la base del model provat, el veredicte semàntic i l'elegibilitat de tall. Les preguntes no qualificades no són abstencions model. Una CONDECINA o INFERÈNCIA correcta segueix sent correcta semànticament, però s'exclou de l'evidència de tall de record pur.
Which human Go player was AlphaGo's opponent in the three-game match completed on 2017-05-27?
Nota del jutge: Matches accepted answer Ke Jie.
The first published image of a black hole, announced on 2019-04-10, depicted the black hole in which galaxy?
Nota del jutge: M87 galaxy is semantically equivalent to Messier 87.
Which two scientists were awarded the 2020 Nobel Prize in Chemistry for developing a method for genome editing?
Nota del jutge: Both required scientists are correctly named.
Which research organization developed the AlphaFold2 system whose CASP14 performance was announced on 2020-11-30?
Nota del jutge: Matches accepted answer DeepMind.
Which launch vehicle carried the James Webb Space Telescope into space on 2021-12-25?
Nota del jutge: Matches accepted answer Ariane 5.
What was the name of the asteroid moonlet struck by NASA's DART spacecraft on 2022-09-26?
Nota del jutge: Matches accepted answer Dimorphos.
In which ocean did the Artemis I Orion capsule splash down on 2022-12-11?
Nota del jutge: Matches accepted answer Pacific Ocean.
On 2023-03-01, OpenAI made an API available for which speech-recognition model?
Nota del jutge: Matches accepted answer Whisper.
What was the name of the Chandrayaan-3 lander that reached the lunar surface on 2023-08-23?
Nota del jutge: Matches accepted answer Vikram.
Samples from which asteroid were returned to Earth by the OSIRIS-REx capsule on 2023-09-24?
Nota del jutge: Matches accepted answer Bennu.
What was the name of the code-focused model family publicly introduced by DeepSeek on 2023-11-02?
Nota del jutge: DeepSeek-Coder is an accepted spelling.
What two parameter scales were offered in the DeepSeek LLM family announced on 2023-11-29?
Nota del jutge: Both parameter scales match: 7B and 67B.
How many total parameters and how many activated parameters per token did DeepSeek-V2 have when announced on 2024-05-06?
Nota del jutge: Correctly gives 236B total and 21B activated per token.
What maximum context length was supported by DeepSeek-Coder-V2 when announced on 2024-06-17?
Nota del jutge: Matches accepted answer 128K tokens.
Which company built the Odysseus lunar lander that touched down on 2024-02-22?
Nota del jutge: Matches accepted answer Intuitive Machines.
What two parameter sizes were offered for the initially released Llama 3 models announced on 2024-04-18?
Nota del jutge: Both initially released sizes match: 8B and 70B.
Cobertura de preguntes i notes de recerca
Preguntes de fet: 16 · No qualificat pel jutge: 0
El tall de referència i les citacions són proporcionades pel jutge, no confirmades de manera independent per la plataforma. Les notes de recerca no penalitzen el model provat.
Detalls de reproductibilitat
A57ZK57ERM62N3S1XETMXFKK0Xmodel-verification-single-request-v1.1gpt-5.6-solopenai_chatLegacy / not recordedPARTIT / llegatNot recordedNot recordedNot recordedNot recorded20028,645 ms1,8594c1f5cca4a79a53f28e7f6544bfe08208148d53c85569f50c4132e161da06613The reference cutoff is the configured judge model's internal-knowledge snapshot for this run, not an independently verified provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.
Executeu una verificació independent
Proveu vosaltres mateixos el mateix punt final o navegueu per altres informes públics abans de confiar en una identitat de model anunciada.