← Public verification registrymodel-verification-single-request-v1.1
Public AI model verification report

deepseek-v4-flash

https://api.model-gate.com/v1/chat/completions

openai_chatgpt-5.6-solSep 6, 2026 09:53 UTC
Overall quality score100/100
PLAUSIBLE60% identity confidence
Permanent reporthttps://model-gate.com/en/model-verification/results/A57ZK57ERM62N3S1XETMXFKK0X
Telegram ↗X ↗
Factual Recall100
Knowledge Recency100
Hallucination Resistance100
Calibration100
Claimed modeldeepseek-v4-flash
Judge-resolved modeldeepseek-v4-flashUNKNOWN
Endpoint hostapi.model-gate.com
External response model fielddeepseek-v4-flash
Judge analysis

Why this verification received its result

The public report mirrors the analytical evidence shown to the tester after /chat/verify. Judge scores are rubric assessments; only explicit correct or wrong RECALL answers may influence the pure-recall cutoff boundary.

Scoring

Judge scoring rationale

100/100
Factual Recall100/100

No separate judge rationale was retained for this archived result.

Knowledge Recency100/100

No separate judge rationale was retained for this archived result.

Hallucination Resistance100/100

No separate judge rationale was retained for this archived result.

Calibration100/100

No separate judge rationale was retained for this archived result.

Overall quality is the fixed 37.5% Factual Recall + 25% Knowledge Recency + 25% Hallucination Resistance + 12.5% Calibration aggregation. Identity and Cutoff Alignment are separate judge assessments.

Temporal evidence

16-month pure-recall timeline

0 RECALL probes

The 24 questions were shuffled before the tested-model request. The timeline reconstructs their target months for auditability. Only explicit RECALL answers are cutoff evidence; INFERENCE, GUESS and UNKNOWN remain visible but do not move the boundary.

Monthly timeline is not available for this archived benchmark result.
correct RECALL wrong RECALL~ INFERENCE GUESS? UNKNOWN
0RECALL
0INFERENCE
0GUESS
0UNKNOWN
0correct RECALL
0wrong RECALL
Temporal comparison

Knowledge cutoff comparison

UNKNOWN
Reference cutoffUNKNOWNUNKNOWN · 0%
Latest supported RECALL0 eligible RECALL answers
Observed boundaryInsufficient RECALL evidenceINSUFFICIENT_RECALL_EVIDENCE
Cutoff AlignmentUNKNOWN0% confidence
Judge notes
  • All Class A items are correct through June 2024, establishing an observed lower bound but no upper bound.
  • The immutable reference snapshot has no known cutoff dates, so cutoff alignment cannot be scored.

Semantic: correct 16 · wrong 0 · unknown 0. Excluded from cutoff: 0.

Behavioral identity

Identity signals

PLAUSIBLE · 60%

Identity is inferred from consistency between the claimed model and the observed knowledge/answer profile. It is evidence, not remote attestation or cryptographic proof of the serving weights.

Positive signals

  • Perfect performance across all supplied Class A recall items.
  • Correct recall of multiple DeepSeek model releases and specifications.
  • No systematic early knowledge failures.

Negative signals

  • No documented canonical identity or cutoff is available in the reference snapshot.
  • Cutoff consistency cannot be evaluated.
  • Self-report and matching API metadata are weak identity evidence.
Judge identity rationale
  • The response is fully correct on all 16 strong-recall questions, including the DeepSeek-specific history items.
  • The observed knowledge extends through June 2024, but the reference cutoff is unknown and therefore provides no identity-comparison signal.
  • The claimed name and API identifier agree, but both are untrusted metadata and do not prove model identity.
Tested-model response

Verbatim extracted model text

Archived summary

This run predates opt-in full-report retention. The challenge, answer key and normalized judge assessment are available, but the exact tested-model response text was intentionally discarded by that benchmark-era privacy policy.

Complete benchmark listing

Question-by-question evidence

All 24 temporal probes are shown with the expected answer, public source, tested-model basis, semantic verdict and cutoff eligibility. A correct GUESS or INFERENCE remains correct semantically but is excluded from pure-recall cutoff evidence.

K01UNKNOWNexcluded from cutoff

Which human Go player was AlphaGo's opponent in the three-game match completed on 2017-05-27?

Expected answerKe Jie
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: Matches accepted answer Ke Jie.

K02UNKNOWNexcluded from cutoff

The first published image of a black hole, announced on 2019-04-10, depicted the black hole in which galaxy?

Expected answerMessier 87
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: M87 galaxy is semantically equivalent to Messier 87.

K03UNKNOWNexcluded from cutoff

Which two scientists were awarded the 2020 Nobel Prize in Chemistry for developing a method for genome editing?

Expected answerEmmanuelle Charpentier and Jennifer A. Doudna
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: Both required scientists are correctly named.

K04UNKNOWNexcluded from cutoff

Which research organization developed the AlphaFold2 system whose CASP14 performance was announced on 2020-11-30?

Expected answerDeepMind
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: Matches accepted answer DeepMind.

K05UNKNOWNexcluded from cutoff

Which launch vehicle carried the James Webb Space Telescope into space on 2021-12-25?

Expected answerAriane 5
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: Matches accepted answer Ariane 5.

K06UNKNOWNexcluded from cutoff

What was the name of the asteroid moonlet struck by NASA's DART spacecraft on 2022-09-26?

Expected answerDimorphos
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: Matches accepted answer Dimorphos.

K07UNKNOWNexcluded from cutoff

In which ocean did the Artemis I Orion capsule splash down on 2022-12-11?

Expected answerPacific Ocean
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: Matches accepted answer Pacific Ocean.

K08UNKNOWNexcluded from cutoff

On 2023-03-01, OpenAI made an API available for which speech-recognition model?

Expected answerWhisper
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: Matches accepted answer Whisper.

K09UNKNOWNexcluded from cutoff

What was the name of the Chandrayaan-3 lander that reached the lunar surface on 2023-08-23?

Expected answerVikram
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: Matches accepted answer Vikram.

K10UNKNOWNexcluded from cutoff

Samples from which asteroid were returned to Earth by the OSIRIS-REx capsule on 2023-09-24?

Expected answerBennu
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: Matches accepted answer Bennu.

K11UNKNOWNexcluded from cutoff

What was the name of the code-focused model family publicly introduced by DeepSeek on 2023-11-02?

Expected answerDeepSeek Coder
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: DeepSeek-Coder is an accepted spelling.

K12UNKNOWNexcluded from cutoff

What two parameter scales were offered in the DeepSeek LLM family announced on 2023-11-29?

Expected answer7B and 67B
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: Both parameter scales match: 7B and 67B.

K13UNKNOWNexcluded from cutoff

How many total parameters and how many activated parameters per token did DeepSeek-V2 have when announced on 2024-05-06?

Expected answer236B total parameters and 21B activated parameters per token
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: Correctly gives 236B total and 21B activated per token.

K14UNKNOWNexcluded from cutoff

What maximum context length was supported by DeepSeek-Coder-V2 when announced on 2024-06-17?

Expected answer128K tokens
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: Matches accepted answer 128K tokens.

K15UNKNOWNexcluded from cutoff

Which company built the Odysseus lunar lander that touched down on 2024-02-22?

Expected answerIntuitive Machines
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: Matches accepted answer Intuitive Machines.

K16UNKNOWNexcluded from cutoff

What two parameter sizes were offered for the initially released Llama 3 models announced on 2024-04-18?

Expected answer8B and 70B
Tested-model answerExact model response was not retained for this archived run.
Judge verdict: CORRECTReported confidence: 100%

Judge note: Both initially released sizes match: 8B and 70B.

Run metadata

Reproducibility details

Run IDA57ZK57ERM62N3S1XETMXFKK0X
Benchmarkmodel-verification-single-request-v1.1
Judgegpt-5.6-sol
Protocolopenai_chat
HTTP status200
External latency28,645 ms
Response bytes1,859
SHA-2564c1f5cca4a79a53f28e7f6544bfe08208148d53c85569f50c4132e161da06613
Important limitation

The reference cutoff is the configured judge model's internal-knowledge snapshot for this run, not an independently verified provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.

Run an independent verification

Test the same endpoint yourself or browse other public reports before relying on an advertised model identity.