← Public verification registrymodel-verification-single-request-v1.2
Public AI model verification report

deepseek-v4-flash

https://api.model-gate.com/v1/chat/completions

openai_chatgpt-5.6-solSep 6, 2026 16:39 UTC
Legacy quality score87/100
CONSISTENT82% assessment confidence
Permanent reporthttps://model-gate.com/en/model-verification/results/0B0HVDB2FRJEJC90BRG4D7JV1Y
Telegram ↗X ↗
Factual Recall83
Knowledge Recency82
Hallucination Resistance83
Calibration80
Claimed modeldeepseek-v4-flash
Judge-resolved modelDeepSeek-V4-Flash-0731DeepSeek AI
Endpoint hostapi.model-gate.com
External response model fielddeepseek-v4-flash
Judge analysis

Why this verification received its result

The public report mirrors the analytical evidence shown to the tester after /chat/verify. Judge scores are rubric assessments; only explicit correct or wrong RECALL answers may influence the pure-recall cutoff boundary.

Behavior diagnostics

Judge scoring rationale

4 diagnostics
Factual Recall83/100

No separate judge rationale was retained for this archived result.

Knowledge Recency82/100

No separate judge rationale was retained for this archived result.

Hallucination Resistance83/100

No separate judge rationale was retained for this archived result.

Calibration80/100

No separate judge rationale was retained for this archived result.

Historical benchmark note: this archived v1.4 overall score used the former fixed diagnostic aggregation. It is retained for reproducibility and must not be interpreted as a v1.5 Model Match score.

Temporal evidence

Pure-recall timeline

0 RECALL probes

The actual question set was shuffled before the tested-model request. Missing generation slots are not model abstentions. The timeline reconstructs their target months for auditability. Only explicit RECALL answers are cutoff evidence; INFERENCE, GUESS and UNKNOWN remain visible but do not move the boundary.

2023-11? K24UNKNOWN
2023-12? K13UNKNOWN
2024-01? K08UNKNOWN
2024-02? K07UNKNOWN
2024-03? K15UNKNOWN
2024-04? K18UNKNOWN
2024-05? K04UNKNOWN
2024-06? K16UNKNOWN
2024-07? K06UNKNOWN
2024-08? K03UNKNOWN
2024-09? K09UNKNOWN
2024-10? K17UNKNOWN
2024-11? K20UNKNOWN
2024-12? K10UNKNOWN
2025-01? K12UNKNOWN
2025-02? K14UNKNOWN
2025-03? K11UNKNOWN
2025-04? K21UNKNOWN
2025-05? K22UNKNOWN
2025-06? K23UNKNOWN
2025-07? K02UNKNOWN
2025-08? K05UNKNOWN
2025-09? K01UNKNOWN
2025-10? K19UNKNOWN
correct RECALL wrong RECALL~ INFERENCE GUESS? UNKNOWN
0RECALL
0INFERENCE
0GUESS
0UNKNOWN
0correct RECALL
0wrong RECALL
Temporal comparison

Knowledge cutoff comparison

CONSISTENT
Reference cutoff2025-05ESTIMATED_WEB · 0%
Latest supported RECALL0 eligible RECALL answers
Observed boundaryInsufficient RECALL evidenceINSUFFICIENT_RECALL_EVIDENCE
Cutoff AlignmentCONSISTENT88/100 · 80% confidence
Judge notes
  • Strong recall is nearly complete through the May 2025 reference cutoff.
  • There is no systematic strong-recall failure substantially before the reference cutoff.
  • Correct June and August 2025 answers suggest selective later knowledge, while failures in July, September, and October prevent a clearly later boundary.

Semantic: correct 20 · wrong 4 · unknown 0. Excluded from cutoff: 0.

Behavioral identity

Identity signals

CONSISTENT · 82%

Identity is inferred from consistency between the claimed model and the observed knowledge/answer profile. It is evidence, not remote attestation or cryptographic proof of the serving weights.

Positive signals

  • Excellent recall across late 2023 through May 2025.
  • Correct near-cutoff answer for the WHO Pandemic Agreement.
  • No pattern of early knowledge degradation.

Negative signals

  • Incorrect February 2024 lunar-lander answer despite high confidence.
  • Correct June and August 2025 answers extend beyond the estimated reference cutoff.
  • Several post-cutoff answers are confidently wrong, indicating uneven recency.
Judge identity rationale
  • Observed behavior broadly overlaps the saved May 2025 reference cutoff.
  • The model answered 20 of 24 Class A items correctly, with only one isolated failure well before the cutoff.
  • Post-cutoff knowledge is uneven and can plausibly reflect selective updates rather than model substitution.
Tested-model response

Verbatim extracted model text

Archived summary

This run predates opt-in full-report retention. The challenge, answer key and normalized judge assessment are available, but the exact tested-model response text was intentionally discarded by that benchmark-era privacy policy.

Complete benchmark listing

Question-by-question evidence

All admitted questions are shown with the expected answer, available source, tested-model basis, semantic verdict and cutoff eligibility. Ungraded questions are not model abstentions. A correct GUESS or INFERENCE remains correct semantically but is excluded from pure-recall cutoff evidence.

K01UNKNOWNexcluded from cutoff

Who won the men's 100 metres at the World Athletics Championships in Tokyo?

Expected answerOblique SevilleSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: WRONGReported confidence: 75%

Judge note: Expected Oblique Seville; answered Kishane Thompson.

K02UNKNOWNexcluded from cutoff

Which team won the UEFA Women's Euro 2025 final?

Expected answerEnglandSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: WRONGReported confidence: 85%

Judge note: Expected England; answered Spain.

K03UNKNOWNexcluded from cutoff

Which disease's African upsurge did WHO declare a public health emergency of international concern?

Expected answermpoxSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 95%

Judge note: Matches mpox.

K04UNKNOWNexcluded from cutoff

What molecular-structure prediction model did Google DeepMind and Isomorphic Labs introduce?

Expected answerAlphaFold 3Source ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 95%

Judge note: Matches AlphaFold 3.

K05UNKNOWNexcluded from cutoff

What model did OpenAI release as a unified system with built-in thinking?

Expected answerGPT-5Source ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 90%

Judge note: Matches GPT-5.

K06UNKNOWNexcluded from cutoff

Which operating system was affected by the faulty CrowdStrike Falcon content update?

Expected answerWindowsSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 95%

Judge note: Microsoft Windows is accepted.

K07UNKNOWNexcluded from cutoff

What was the nickname of Intuitive Machines' Nova-C lander that touched down on the Moon?

Expected answerOdysseusSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: WRONGReported confidence: 80%

Judge note: Expected Odysseus; Odie is not an accepted answer.

K08UNKNOWNexcluded from cutoff

What was the name of Japan's lunar lander that successfully reached the Moon?

Expected answerSLIMSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 95%

Judge note: Matches SLIM.

K09UNKNOWNexcluded from cutoff

Who became the first private astronaut to perform a spacewalk?

Expected answerJared IsaacmanSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 95%

Judge note: Matches Jared Isaacman.

K10UNKNOWNexcluded from cutoff

Which country was appointed host of the 2034 FIFA World Cup?

Expected answerSaudi ArabiaSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 95%

Judge note: Matches Saudi Arabia.

K11UNKNOWNexcluded from cutoff

In which lunar mare did Firefly Aerospace's Blue Ghost land?

Expected answerMare CrisiumSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 90%

Judge note: Matches Mare Crisium.

K12UNKNOWNexcluded from cutoff

What model did DeepSeek begin serving through its deepseek-reasoner endpoint?

Expected answerDeepSeek-R1Source ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 95%

Judge note: Matches DeepSeek-R1.

K13UNKNOWNexcluded from cutoff

What was the name of the second malaria vaccine prequalified by WHO?

Expected answerR21/Matrix-MSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 90%

Judge note: Matches R21/Matrix-M.

K14UNKNOWNexcluded from cutoff

Which album won Album of the Year at the 67th Grammy Awards?

Expected answerCowboy CarterSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 95%

Judge note: Matches Cowboy Carter.

K15UNKNOWNexcluded from cutoff

How many European Parliament members voted in favor of the Artificial Intelligence Act?

Expected answer523Source ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 85%

Judge note: Matches 523.

K16UNKNOWNexcluded from cutoff

How many grams of lunar material did the Chang'e-6 mission collect?

Expected answer1,935.3 gramsSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 90%

Judge note: Matches 1,935.3 grams.

K17UNKNOWNexcluded from cutoff

Which rocket launched NASA's Europa Clipper spacecraft?

Expected answerFalcon HeavySource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 95%

Judge note: Matches Falcon Heavy.

K18UNKNOWNexcluded from cutoff

What designation was given to the newly identified 33-solar-mass stellar black hole in the Milky Way?

Expected answerGaia BH3Source ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 90%

Judge note: Matches Gaia BH3.

K19UNKNOWNexcluded from cutoff

Who was awarded the 2025 Nobel Peace Prize?

Expected answerMaria Corina MachadoSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: WRONGReported confidence: 70%

Judge note: Expected Maria Corina Machado; answered an institution.

K20UNKNOWNexcluded from cutoff

What annual climate-finance goal for developing countries did COP29 set for 2035?

Expected answerUS$300 billion per yearSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 90%

Judge note: Matches $300 billion per year.

K21UNKNOWNexcluded from cutoff

What was the name of the first human spaceflight to orbit over Earth's polar regions?

Expected answerFram2Source ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 90%

Judge note: Matches Fram2.

K22UNKNOWNexcluded from cutoff

Under which article of the WHO Constitution was the Pandemic Agreement adopted?

Expected answerArticle 19Source ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 75%

Judge note: Matches Article 19.

K23UNKNOWNexcluded from cutoff

Which new Mario Kart game launched alongside Nintendo Switch 2?

Expected answerMario Kart WorldSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 95%

Judge note: Matches Mario Kart World.

K24UNKNOWNexcluded from cutoff

What model did OpenAI introduce with a 128K context window at its first DevDay?

Expected answerGPT-4 TurboSource ↗
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 95%

Judge note: Matches GPT-4 Turbo.

Question coverage and research notes

Factual questions: 24 · Not graded by judge: 0

Reference cutoff and citations are supplied by the judge, not independently confirmed by the platform. Research notes do not penalize the tested model.

Run metadata

Reproducibility details

Run ID0B0HVDB2FRJEJC90BRG4D7JV1Y
Benchmarkmodel-verification-single-request-v1.2
Judgegpt-5.6-sol
Requested protocolopenai_chat
Detected response schemaLegacy / not recorded
Protocol compatibilityMATCH / legacy
Model outputNot recorded
Input tokensNot recorded
Output tokensNot recorded
Finish reasonNot recorded
HTTP status200
External latency88,304 ms
Response bytes2,460
SHA-256f22e63d94e67fbd4fdfa25a27a37ae4b3402f94fc6801d14c2a4e638f85cfd9d
Important limitation

The reference cutoff is a judge-produced snapshot grounded in hosted web evidence when available (official sources preferred, otherwise the best defensible estimate). It is not provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.

Run an independent verification

Test the same endpoint yourself or browse other public reports before relying on an advertised model identity.