← 公共验证登记处model-verification-single-request-v1.1
公开AI模型验证报告

deepseek-v4-flash

https://api.model-gate.com/v1/chat/completions

openai_chatgpt-5.6-solSep 6, 2026 09:53 UTC
旧版质量得分100/100
PLAUSIBLE60% 评估置信度
永久报告https://model-gate.com/zh/model-verification/results/A57ZK57ERM62N3S1XETMXFKK0X
Telegram ↗X ↗
事实回忆100
知识新近度100
抗幻觉100
校准100
声称型号deepseek-v4-flash
法官解决模型deepseek-v4-flashUNKNOWN
端点主机api.model-gate.com
外部响应模型字段deepseek-v4-flash
评判分析

为什么此验证收到结果

公开报告反映了 /chat/verify 之后向测试人员显示的分析证据。评委评分为评分标准;只有明确的正确或错误的回忆答案可能会影响纯回忆截止边界。

行为诊断

评委评分依据

4 诊断
事实回忆100/100

该存档结果没有保留单独的法官理由。

知识新近度100/100

该存档结果没有保留单独的法官理由。

抗幻觉100/100

该存档结果没有保留单独的法官理由。

校准100/100

该存档结果没有保留单独的法官理由。

历史基准说明:此存档的 v1.4 总体分数使用了以前的固定诊断聚合。保留它是为了再现性,并且不得将其解释为 v1.5 模型匹配分数。

时间证据

纯粹回忆时间线

0 召回探针

实际问题集在测试模型请求之前被打乱。缺席一代并不是典型的弃权。该时间表重建了可审计性的目标月份。只有明确的 RECALL 答案才是截止证据;推论、猜测和未知仍然可见,但不会移动边界。

此存档的基准测试结果没有每月时间表。
正确回忆 错误回忆~ INFERENCE GUESS? UNKNOWN
0RECALL
0INFERENCE
0GUESS
0UNKNOWN
0正确回忆
0错误回忆
时间比较

知识截止比较

UNKNOWN
参考截止值UNKNOWNUNKNOWN · 0%
最新支持的 RECALL0 符合条件的回忆答案
观察边界召回证据不足INSUFFICIENT_RECALL_EVIDENCE
截止对齐UNKNOWN0% 信心
法官笔记
  • All Class A items are correct through June 2024, establishing an observed lower bound but no upper bound.
  • The immutable reference snapshot has no known cutoff dates, so cutoff alignment cannot be scored.

语义学: 正确的 16 · 错误的 0 · 未知 0. 排除在截止范围之外: 0.

行为认同

身份信号

PLAUSIBLE · 60%

身份是根据所声称的模型与观察到的知识/答案概况之间的一致性推断出来的。它是证据,而不是份量的远程证明或加密证明。

积极信号

  • Perfect performance across all supplied Class A recall items.
  • Correct recall of multiple DeepSeek model releases and specifications.
  • No systematic early knowledge failures.

负面信号

  • No documented canonical identity or cutoff is available in the reference snapshot.
  • Cutoff consistency cannot be evaluated.
  • Self-report and matching API metadata are weak identity evidence.
法官身份理由
  • The response is fully correct on all 16 strong-recall questions, including the DeepSeek-specific history items.
  • The observed knowledge extends through June 2024, but the reference cutoff is unknown and therefore provides no identity-comparison signal.
  • The claimed name and API identifier agree, but both are untrusted metadata and do not prove model identity.
测试模型响应

逐字提取模型文本

存档摘要

此运行早于选择保留完整报告的时间。挑战、答案关键和规范化的法官评估都是可用的,但基准时代的隐私政策故意丢弃了确切的测试模型响应文本。

完整的基准列表

逐个问题的证据

所有承认的问题都会显示预期答案、可用来源、测试模型基础、语义判决和截止资格。未评分的问题不是典型的弃权问题。正确的猜测或推理在语义上仍然是正确的,但被排除在纯回忆截止证据之外。

K01UNKNOWN排除在截止之外

Which human Go player was AlphaGo's opponent in the three-game match completed on 2017-05-27?

预期答案Ke Jie
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: Matches accepted answer Ke Jie.

K02UNKNOWN排除在截止之外

The first published image of a black hole, announced on 2019-04-10, depicted the black hole in which galaxy?

预期答案Messier 87
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: M87 galaxy is semantically equivalent to Messier 87.

K03UNKNOWN排除在截止之外

Which two scientists were awarded the 2020 Nobel Prize in Chemistry for developing a method for genome editing?

预期答案Emmanuelle Charpentier and Jennifer A. Doudna
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: Both required scientists are correctly named.

K04UNKNOWN排除在截止之外

Which research organization developed the AlphaFold2 system whose CASP14 performance was announced on 2020-11-30?

预期答案DeepMind
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: Matches accepted answer DeepMind.

K05UNKNOWN排除在截止之外

Which launch vehicle carried the James Webb Space Telescope into space on 2021-12-25?

预期答案Ariane 5
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: Matches accepted answer Ariane 5.

K06UNKNOWN排除在截止之外

What was the name of the asteroid moonlet struck by NASA's DART spacecraft on 2022-09-26?

预期答案Dimorphos
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: Matches accepted answer Dimorphos.

K07UNKNOWN排除在截止之外

In which ocean did the Artemis I Orion capsule splash down on 2022-12-11?

预期答案Pacific Ocean
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: Matches accepted answer Pacific Ocean.

K08UNKNOWN排除在截止之外

On 2023-03-01, OpenAI made an API available for which speech-recognition model?

预期答案Whisper
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: Matches accepted answer Whisper.

K09UNKNOWN排除在截止之外

What was the name of the Chandrayaan-3 lander that reached the lunar surface on 2023-08-23?

预期答案Vikram
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: Matches accepted answer Vikram.

K10UNKNOWN排除在截止之外

Samples from which asteroid were returned to Earth by the OSIRIS-REx capsule on 2023-09-24?

预期答案Bennu
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: Matches accepted answer Bennu.

K11UNKNOWN排除在截止之外

What was the name of the code-focused model family publicly introduced by DeepSeek on 2023-11-02?

预期答案DeepSeek Coder
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: DeepSeek-Coder is an accepted spelling.

K12UNKNOWN排除在截止之外

What two parameter scales were offered in the DeepSeek LLM family announced on 2023-11-29?

预期答案7B and 67B
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: Both parameter scales match: 7B and 67B.

K13UNKNOWN排除在截止之外

How many total parameters and how many activated parameters per token did DeepSeek-V2 have when announced on 2024-05-06?

预期答案236B total parameters and 21B activated parameters per token
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: Correctly gives 236B total and 21B activated per token.

K14UNKNOWN排除在截止之外

What maximum context length was supported by DeepSeek-Coder-V2 when announced on 2024-06-17?

预期答案128K tokens
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: Matches accepted answer 128K tokens.

K15UNKNOWN排除在截止之外

Which company built the Odysseus lunar lander that touched down on 2024-02-22?

预期答案Intuitive Machines
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: Matches accepted answer Intuitive Machines.

K16UNKNOWN排除在截止之外

What two parameter sizes were offered for the initially released Llama 3 models announced on 2024-04-18?

预期答案8B and 70B
测试模型答案此存档运行未保留精确的模型响应。
法官判决: CORRECT报告信心: 100%

法官笔记: Both initially released sizes match: 8B and 70B.

问题范围和研究笔记

事实问题: 16 · 没有被评委打分: 0

参考截止值和引文由法官提供,未经平台独立确认。研究笔记不会惩罚测试模型。

运行元数据

再现性细节

运行IDA57ZK57ERM62N3S1XETMXFKK0X
基准model-verification-single-request-v1.1
法官gpt-5.6-sol
请求的协议openai_chat
检测到的响应模式Legacy / not recorded
协议兼容性比赛/遗产
模型输出Not recorded
输入令牌Not recorded
输出代币Not recorded
完成原因Not recorded
HTTP状态200
外部延迟28,645 ms
响应字节1,859
SHA-2564c1f5cca4a79a53f28e7f6544bfe08208148d53c85569f50c4132e161da06613
重要限制

The reference cutoff is the configured judge model's internal-knowledge snapshot for this run, not an independently verified provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.

运行独立验证

在依赖广告模型标识之前,请自行测试同一端点或浏览其他公开报告。