deepseek-v4-flash
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/ro/model-verification/results/A57ZK57ERM62N3S1XETMXFKK0XDe ce a primit rezultatul acestei verificări
Raportul public reflectă dovezile analitice prezentate testatorului după /chat/verify. Scorurile judecătorilor sunt evaluări rubrica; numai răspunsurile RECALL explicite, corecte sau greșite, pot influența limita limită de rechemare pură.
Motivul punctajului judecătorului
Nu s-a reținut nicio justificare separată a judecătorului pentru acest rezultat arhivat.
Nu s-a reținut nicio justificare separată a judecătorului pentru acest rezultat arhivat.
Nu s-a reținut nicio justificare separată a judecătorului pentru acest rezultat arhivat.
Nu s-a reținut nicio justificare separată a judecătorului pentru acest rezultat arhivat.
Notă de referință istorică: acest scor general arhivat v1.4 a folosit fosta agregare fixă de diagnosticare. Este păstrat pentru reproductibilitate și nu trebuie interpretat ca un scor de potrivire a modelului v1.5.
Cronologie pură de reamintire
Setul real de întrebări a fost amestecat înainte de solicitarea modelului testat. Sloturile de generație lipsă nu sunt abțineri de model. Cronologia reconstruiește lunile țintă pentru auditabilitate. Doar răspunsurile RECALL explicite sunt dovezi tăiate; INFERENȚĂ, GHICI și NECUNOSCUT rămân vizibile, dar nu mută granița.
Compararea limitelor de cunoștințe
- All Class A items are correct through June 2024, establishing an observed lower bound but no upper bound.
- The immutable reference snapshot has no known cutoff dates, so cutoff alignment cannot be scored.
Semantic: corecta 16 · greşit 0 · necunoscut 0. Exclus din cutoff: 0.
Semnale de identitate
Identitatea este dedusă din coerența dintre modelul revendicat și profilul de cunoștințe/răspuns observat. Este o dovadă, nu o atestare de la distanță sau o dovadă criptografică a greutăților porției.
Semnale pozitive
- Perfect performance across all supplied Class A recall items.
- Correct recall of multiple DeepSeek model releases and specifications.
- No systematic early knowledge failures.
Semnale negative
- No documented canonical identity or cutoff is available in the reference snapshot.
- Cutoff consistency cannot be evaluated.
- Self-report and matching API metadata are weak identity evidence.
- The response is fully correct on all 16 strong-recall questions, including the DeepSeek-specific history items.
- The observed knowledge extends through June 2024, but the reference cutoff is unknown and therefore provides no identity-comparison signal.
- The claimed name and API identifier agree, but both are untrusted metadata and do not prove model identity.
Text model extras literal
Această rulare este anterioară reținerii rapoartelor complete pentru înscriere. Provocarea, cheia de răspuns și evaluarea normalizată a judecătorului sunt disponibile, dar textul de răspuns exact al modelului testat a fost eliminat în mod intenționat de acea politică de confidențialitate a epocii de referință.
Dovezi întrebare cu întrebare
Toate întrebările admise sunt afișate cu răspunsul așteptat, sursa disponibilă, baza modelului testat, verdictul semantic și eligibilitatea limită. Întrebările negradate nu sunt abțineri model. O PRECIZARE sau INFERENȚĂ corectă rămâne corectă din punct de vedere semantic, dar este exclusă din dovezile de reducere a reamintirii.
Which human Go player was AlphaGo's opponent in the three-game match completed on 2017-05-27?
Nota judecătorului: Matches accepted answer Ke Jie.
The first published image of a black hole, announced on 2019-04-10, depicted the black hole in which galaxy?
Nota judecătorului: M87 galaxy is semantically equivalent to Messier 87.
Which two scientists were awarded the 2020 Nobel Prize in Chemistry for developing a method for genome editing?
Nota judecătorului: Both required scientists are correctly named.
Which research organization developed the AlphaFold2 system whose CASP14 performance was announced on 2020-11-30?
Nota judecătorului: Matches accepted answer DeepMind.
Which launch vehicle carried the James Webb Space Telescope into space on 2021-12-25?
Nota judecătorului: Matches accepted answer Ariane 5.
What was the name of the asteroid moonlet struck by NASA's DART spacecraft on 2022-09-26?
Nota judecătorului: Matches accepted answer Dimorphos.
In which ocean did the Artemis I Orion capsule splash down on 2022-12-11?
Nota judecătorului: Matches accepted answer Pacific Ocean.
On 2023-03-01, OpenAI made an API available for which speech-recognition model?
Nota judecătorului: Matches accepted answer Whisper.
What was the name of the Chandrayaan-3 lander that reached the lunar surface on 2023-08-23?
Nota judecătorului: Matches accepted answer Vikram.
Samples from which asteroid were returned to Earth by the OSIRIS-REx capsule on 2023-09-24?
Nota judecătorului: Matches accepted answer Bennu.
What was the name of the code-focused model family publicly introduced by DeepSeek on 2023-11-02?
Nota judecătorului: DeepSeek-Coder is an accepted spelling.
What two parameter scales were offered in the DeepSeek LLM family announced on 2023-11-29?
Nota judecătorului: Both parameter scales match: 7B and 67B.
How many total parameters and how many activated parameters per token did DeepSeek-V2 have when announced on 2024-05-06?
Nota judecătorului: Correctly gives 236B total and 21B activated per token.
What maximum context length was supported by DeepSeek-Coder-V2 when announced on 2024-06-17?
Nota judecătorului: Matches accepted answer 128K tokens.
Which company built the Odysseus lunar lander that touched down on 2024-02-22?
Nota judecătorului: Matches accepted answer Intuitive Machines.
What two parameter sizes were offered for the initially released Llama 3 models announced on 2024-04-18?
Nota judecătorului: Both initially released sizes match: 8B and 70B.
Acoperirea întrebărilor și note de cercetare
Întrebări de fapt: 16 · Nu este notat de judecător: 0
Limitele de referință și citările sunt furnizate de judecător, nu confirmate independent de platformă. Notele cercetării nu penalizează modelul testat.
Detalii de reproductibilitate
A57ZK57ERM62N3S1XETMXFKK0Xmodel-verification-single-request-v1.1gpt-5.6-solopenai_chatLegacy / not recordedMECI / moștenireNot recordedNot recordedNot recordedNot recorded20028,645 ms1,8594c1f5cca4a79a53f28e7f6544bfe08208148d53c85569f50c4132e161da06613The reference cutoff is the configured judge model's internal-knowledge snapshot for this run, not an independently verified provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.
Efectuați o verificare independentă
Testați singur același punct final sau răsfoiți alte rapoarte publice înainte de a vă baza pe o identitate de model promovată.