deepseek-v4-flash
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/en/model-verification/results/0B0HVDB2FRJEJC90BRG4D7JV1YWhy this verification received its result
The public report mirrors the analytical evidence shown to the tester after /chat/verify. Judge scores are rubric assessments; only explicit correct or wrong RECALL answers may influence the pure-recall cutoff boundary.
Judge scoring rationale
No separate judge rationale was retained for this archived result.
No separate judge rationale was retained for this archived result.
No separate judge rationale was retained for this archived result.
No separate judge rationale was retained for this archived result.
Historical benchmark note: this archived v1.4 overall score used the former fixed diagnostic aggregation. It is retained for reproducibility and must not be interpreted as a v1.5 Model Match score.
Pure-recall timeline
The actual question set was shuffled before the tested-model request. Missing generation slots are not model abstentions. The timeline reconstructs their target months for auditability. Only explicit RECALL answers are cutoff evidence; INFERENCE, GUESS and UNKNOWN remain visible but do not move the boundary.
Knowledge cutoff comparison
- Strong recall is nearly complete through the May 2025 reference cutoff.
- There is no systematic strong-recall failure substantially before the reference cutoff.
- Correct June and August 2025 answers suggest selective later knowledge, while failures in July, September, and October prevent a clearly later boundary.
Semantic: correct 20 · wrong 4 · unknown 0. Excluded from cutoff: 0.
Identity signals
Identity is inferred from consistency between the claimed model and the observed knowledge/answer profile. It is evidence, not remote attestation or cryptographic proof of the serving weights.
Positive signals
- Excellent recall across late 2023 through May 2025.
- Correct near-cutoff answer for the WHO Pandemic Agreement.
- No pattern of early knowledge degradation.
Negative signals
- Incorrect February 2024 lunar-lander answer despite high confidence.
- Correct June and August 2025 answers extend beyond the estimated reference cutoff.
- Several post-cutoff answers are confidently wrong, indicating uneven recency.
- Observed behavior broadly overlaps the saved May 2025 reference cutoff.
- The model answered 20 of 24 Class A items correctly, with only one isolated failure well before the cutoff.
- Post-cutoff knowledge is uneven and can plausibly reflect selective updates rather than model substitution.
Verbatim extracted model text
This run predates opt-in full-report retention. The challenge, answer key and normalized judge assessment are available, but the exact tested-model response text was intentionally discarded by that benchmark-era privacy policy.
Question-by-question evidence
All admitted questions are shown with the expected answer, available source, tested-model basis, semantic verdict and cutoff eligibility. Ungraded questions are not model abstentions. A correct GUESS or INFERENCE remains correct semantically but is excluded from pure-recall cutoff evidence.
Who won the men's 100 metres at the World Athletics Championships in Tokyo?
Judge note: Expected Oblique Seville; answered Kishane Thompson.
Which team won the UEFA Women's Euro 2025 final?
Judge note: Expected England; answered Spain.
Which disease's African upsurge did WHO declare a public health emergency of international concern?
Judge note: Matches mpox.
What molecular-structure prediction model did Google DeepMind and Isomorphic Labs introduce?
Judge note: Matches AlphaFold 3.
What model did OpenAI release as a unified system with built-in thinking?
Judge note: Matches GPT-5.
Which operating system was affected by the faulty CrowdStrike Falcon content update?
Judge note: Microsoft Windows is accepted.
What was the nickname of Intuitive Machines' Nova-C lander that touched down on the Moon?
Judge note: Expected Odysseus; Odie is not an accepted answer.
What was the name of Japan's lunar lander that successfully reached the Moon?
Judge note: Matches SLIM.
Who became the first private astronaut to perform a spacewalk?
Judge note: Matches Jared Isaacman.
Which country was appointed host of the 2034 FIFA World Cup?
Judge note: Matches Saudi Arabia.
In which lunar mare did Firefly Aerospace's Blue Ghost land?
Judge note: Matches Mare Crisium.
What model did DeepSeek begin serving through its deepseek-reasoner endpoint?
Judge note: Matches DeepSeek-R1.
What was the name of the second malaria vaccine prequalified by WHO?
Judge note: Matches R21/Matrix-M.
Which album won Album of the Year at the 67th Grammy Awards?
Judge note: Matches Cowboy Carter.
How many European Parliament members voted in favor of the Artificial Intelligence Act?
Judge note: Matches 523.
How many grams of lunar material did the Chang'e-6 mission collect?
Judge note: Matches 1,935.3 grams.
Which rocket launched NASA's Europa Clipper spacecraft?
Judge note: Matches Falcon Heavy.
What designation was given to the newly identified 33-solar-mass stellar black hole in the Milky Way?
Judge note: Matches Gaia BH3.
Who was awarded the 2025 Nobel Peace Prize?
Judge note: Expected Maria Corina Machado; answered an institution.
What annual climate-finance goal for developing countries did COP29 set for 2035?
Judge note: Matches $300 billion per year.
What was the name of the first human spaceflight to orbit over Earth's polar regions?
Judge note: Matches Fram2.
Under which article of the WHO Constitution was the Pandemic Agreement adopted?
Judge note: Matches Article 19.
Which new Mario Kart game launched alongside Nintendo Switch 2?
Judge note: Matches Mario Kart World.
What model did OpenAI introduce with a 128K context window at its first DevDay?
Judge note: Matches GPT-4 Turbo.
Question coverage and research notes
Factual questions: 24 · Not graded by judge: 0
Reference cutoff and citations are supplied by the judge, not independently confirmed by the platform. Research notes do not penalize the tested model.
Reproducibility details
0B0HVDB2FRJEJC90BRG4D7JV1Ymodel-verification-single-request-v1.2gpt-5.6-solopenai_chatLegacy / not recordedMATCH / legacyNot recordedNot recordedNot recordedNot recorded20088,304 ms2,460f22e63d94e67fbd4fdfa25a27a37ae4b3402f94fc6801d14c2a4e638f85cfd9dThe reference cutoff is a judge-produced snapshot grounded in hosted web evidence when available (official sources preferred, otherwise the best defensible estimate). It is not provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.
Run an independent verification
Test the same endpoint yourself or browse other public reports before relying on an advertised model identity.