← 公共验证登记处model-verification-single-request-v1.2
公开AI模型验证报告

deepseek-v4-flash

https://api.model-gate.com/v1/chat/completions

openai_chatgpt-5.6-solSep 6, 2026 16:39 UTC
总体质量得分87/100
CONSISTENT82% 身份信心
永久报告https://model-gate.com/zh/model-verification/results/0B0HVDB2FRJEJC90BRG4D7JV1Y
Telegram ↗X ↗
事实回忆83
知识新近度82
抗幻觉83
校准80
声称型号deepseek-v4-flash
法官解决模型DeepSeek-V4-Flash-0731DeepSeek AI
端点主机api.model-gate.com
外部响应模型字段deepseek-v4-flash
评判分析

为什么此验证收到结果

公开报告反映了 /chat/verify 之后向测试人员显示的分析证据。评委评分为评分标准;只有明确的正确或错误的回忆答案可能会影响纯回忆截止边界。

评分

评委评分依据

87/100
事实回忆83/100

该存档结果没有保留单独的法官理由。

知识新近度82/100

该存档结果没有保留单独的法官理由。

抗幻觉83/100

该存档结果没有保留单独的法官理由。

校准80/100

该存档结果没有保留单独的法官理由。

总体质量是固定的 37.5% 事实回忆 + 25% 知识新近度 + 25% 幻觉抵抗 + 12.5% 校准聚合。身份和截止对齐是单独的法官评估。

时间证据

16 个月的纯回忆时间表

0 召回探针

在测试模型请求之前,这 24 个问题被打乱了。该时间表重建了可审计性的目标月份。只有明确的 RECALL 答案才是截止证据;推论、猜测和未知仍然可见,但不会移动边界。

2023-11? K24UNKNOWN
2023-12? K13UNKNOWN
2024-01? K08UNKNOWN
2024-02? K07UNKNOWN
2024-03? K15UNKNOWN
2024-04? K18UNKNOWN
2024-05? K04UNKNOWN
2024-06? K16UNKNOWN
2024-07? K06UNKNOWN
2024-08? K03UNKNOWN
2024-09? K09UNKNOWN
2024-10? K17UNKNOWN
2024-11? K20UNKNOWN
2024-12? K10UNKNOWN
2025-01? K12UNKNOWN
2025-02? K14UNKNOWN
2025-03? K11UNKNOWN
2025-04? K21UNKNOWN
2025-05? K22UNKNOWN
2025-06? K23UNKNOWN
2025-07? K02UNKNOWN
2025-08? K05UNKNOWN
2025-09? K01UNKNOWN
2025-10? K19UNKNOWN
正确回忆 错误回忆~ INFERENCE GUESS? UNKNOWN
0RECALL
0INFERENCE
0GUESS
0UNKNOWN
0正确回忆
0错误回忆
时间比较

知识截止比较

CONSISTENT
参考截止值2025-05ESTIMATED_WEB · 0%
最新支持的 RECALL0 符合条件的回忆答案
观察边界召回证据不足INSUFFICIENT_RECALL_EVIDENCE
截止对齐CONSISTENT88/100 · 80% 信心
法官笔记
  • Strong recall is nearly complete through the May 2025 reference cutoff.
  • There is no systematic strong-recall failure substantially before the reference cutoff.
  • Correct June and August 2025 answers suggest selective later knowledge, while failures in July, September, and October prevent a clearly later boundary.

语义学: 正确的 20 · 错误的 4 · 未知 0. 排除在截止范围之外: 0.

行为认同

身份信号

CONSISTENT · 82%

身份是根据所声称的模型与观察到的知识/答案概况之间的一致性推断出来的。它是证据,而不是份量的远程证明或加密证明。

积极信号

  • Excellent recall across late 2023 through May 2025.
  • Correct near-cutoff answer for the WHO Pandemic Agreement.
  • No pattern of early knowledge degradation.

负面信号

  • Incorrect February 2024 lunar-lander answer despite high confidence.
  • Correct June and August 2025 answers extend beyond the estimated reference cutoff.
  • Several post-cutoff answers are confidently wrong, indicating uneven recency.
法官身份理由
  • Observed behavior broadly overlaps the saved May 2025 reference cutoff.
  • The model answered 20 of 24 Class A items correctly, with only one isolated failure well before the cutoff.
  • Post-cutoff knowledge is uneven and can plausibly reflect selective updates rather than model substitution.
测试模型响应

逐字提取模型文本

存档摘要

此运行早于选择保留完整报告的时间。挑战、答案关键和规范化的法官评估都是可用的,但基准时代的隐私政策故意丢弃了确切的测试模型响应文本。

完整的基准列表

逐个问题的证据

所有 24 个时间探针都显示了预期答案、公共来源、测试模型基础、语义判决和截止资格。正确的猜测或推理在语义上仍然是正确的,但被排除在纯回忆截止证据之外。

K01UNKNOWN排除在截止之外

Who won the men's 100 metres at the World Athletics Championships in Tokyo?

预期答案Oblique SevilleSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: WRONG报告信心: 75%

法官笔记: Expected Oblique Seville; answered Kishane Thompson.

K02UNKNOWN排除在截止之外

Which team won the UEFA Women's Euro 2025 final?

预期答案EnglandSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: WRONG报告信心: 85%

法官笔记: Expected England; answered Spain.

K03UNKNOWN排除在截止之外

Which disease's African upsurge did WHO declare a public health emergency of international concern?

预期答案mpoxSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 95%

法官笔记: Matches mpox.

K04UNKNOWN排除在截止之外

What molecular-structure prediction model did Google DeepMind and Isomorphic Labs introduce?

预期答案AlphaFold 3Source ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 95%

法官笔记: Matches AlphaFold 3.

K05UNKNOWN排除在截止之外

What model did OpenAI release as a unified system with built-in thinking?

预期答案GPT-5Source ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 90%

法官笔记: Matches GPT-5.

K06UNKNOWN排除在截止之外

Which operating system was affected by the faulty CrowdStrike Falcon content update?

预期答案WindowsSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 95%

法官笔记: Microsoft Windows is accepted.

K07UNKNOWN排除在截止之外

What was the nickname of Intuitive Machines' Nova-C lander that touched down on the Moon?

预期答案OdysseusSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: WRONG报告信心: 80%

法官笔记: Expected Odysseus; Odie is not an accepted answer.

K08UNKNOWN排除在截止之外

What was the name of Japan's lunar lander that successfully reached the Moon?

预期答案SLIMSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 95%

法官笔记: Matches SLIM.

K09UNKNOWN排除在截止之外

Who became the first private astronaut to perform a spacewalk?

预期答案Jared IsaacmanSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 95%

法官笔记: Matches Jared Isaacman.

K10UNKNOWN排除在截止之外

Which country was appointed host of the 2034 FIFA World Cup?

预期答案Saudi ArabiaSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 95%

法官笔记: Matches Saudi Arabia.

K11UNKNOWN排除在截止之外

In which lunar mare did Firefly Aerospace's Blue Ghost land?

预期答案Mare CrisiumSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 90%

法官笔记: Matches Mare Crisium.

K12UNKNOWN排除在截止之外

What model did DeepSeek begin serving through its deepseek-reasoner endpoint?

预期答案DeepSeek-R1Source ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 95%

法官笔记: Matches DeepSeek-R1.

K13UNKNOWN排除在截止之外

What was the name of the second malaria vaccine prequalified by WHO?

预期答案R21/Matrix-MSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 90%

法官笔记: Matches R21/Matrix-M.

K14UNKNOWN排除在截止之外

Which album won Album of the Year at the 67th Grammy Awards?

预期答案Cowboy CarterSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 95%

法官笔记: Matches Cowboy Carter.

K15UNKNOWN排除在截止之外

How many European Parliament members voted in favor of the Artificial Intelligence Act?

预期答案523Source ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 85%

法官笔记: Matches 523.

K16UNKNOWN排除在截止之外

How many grams of lunar material did the Chang'e-6 mission collect?

预期答案1,935.3 gramsSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 90%

法官笔记: Matches 1,935.3 grams.

K17UNKNOWN排除在截止之外

Which rocket launched NASA's Europa Clipper spacecraft?

预期答案Falcon HeavySource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 95%

法官笔记: Matches Falcon Heavy.

K18UNKNOWN排除在截止之外

What designation was given to the newly identified 33-solar-mass stellar black hole in the Milky Way?

预期答案Gaia BH3Source ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 90%

法官笔记: Matches Gaia BH3.

K19UNKNOWN排除在截止之外

Who was awarded the 2025 Nobel Peace Prize?

预期答案Maria Corina MachadoSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: WRONG报告信心: 70%

法官笔记: Expected Maria Corina Machado; answered an institution.

K20UNKNOWN排除在截止之外

What annual climate-finance goal for developing countries did COP29 set for 2035?

预期答案US$300 billion per yearSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 90%

法官笔记: Matches $300 billion per year.

K21UNKNOWN排除在截止之外

What was the name of the first human spaceflight to orbit over Earth's polar regions?

预期答案Fram2Source ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 90%

法官笔记: Matches Fram2.

K22UNKNOWN排除在截止之外

Under which article of the WHO Constitution was the Pandemic Agreement adopted?

预期答案Article 19Source ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 75%

法官笔记: Matches Article 19.

K23UNKNOWN排除在截止之外

Which new Mario Kart game launched alongside Nintendo Switch 2?

预期答案Mario Kart WorldSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 95%

法官笔记: Matches Mario Kart World.

K24UNKNOWN排除在截止之外

What model did OpenAI introduce with a 128K context window at its first DevDay?

预期答案GPT-4 TurboSource ↗
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 95%

法官笔记: Matches GPT-4 Turbo.

运行元数据

再现性细节

运行ID0B0HVDB2FRJEJC90BRG4D7JV1Y
基准model-verification-single-request-v1.2
法官gpt-5.6-sol
协议openai_chat
HTTP状态200
外部延迟88,304 ms
响应字节2,460
SHA-256f22e63d94e67fbd4fdfa25a27a37ae4b3402f94fc6801d14c2a4e638f85cfd9d
重要限制

The reference cutoff is a judge-produced snapshot grounded in hosted web evidence when available (official sources preferred, otherwise the best defensible estimate). It is not provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.

运行独立验证

在依赖广告模型标识之前,请自行测试同一端点或浏览其他公开报告。