deepseek-v4-flash
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/ja/model-verification/results/A57ZK57ERM62N3S1XETMXFKK0Xこの検証結果が得られた理由
公開レポートは、/chat/verify の後にテスターに示される分析証拠を反映しています。審査員のスコアはルーブリック評価です。純粋なリコールのカットオフ境界に影響を与えることができるのは、明示的な正解または不正解の RECALL の回答のみです。
ジャッジの採点根拠
このアーカイブされた結果については、別の裁判官の理論的根拠は保持されていません。
このアーカイブされた結果については、別の裁判官の理論的根拠は保持されていません。
このアーカイブされた結果については、別の裁判官の理論的根拠は保持されていません。
このアーカイブされた結果については、別の裁判官の理論的根拠は保持されていません。
過去のベンチマークのメモ: このアーカイブされた v1.4 の全体スコアでは、以前の固定診断集計が使用されています。これは再現性のために保持されており、v1.5 のモデル マッチ スコアとして解釈しないでください。
純粋な想起のタイムライン
実際の質問セットは、テスト済みモデルのリクエストの前にシャッフルされました。世代スロットの欠落はモデル棄権ではありません。タイムラインは、監査可能性を考慮して目標月を再構成します。明示的な「RECALL」の回答のみが完全な証拠となります。 INFERENCE、GUESS、UNKNOWN は表示されたままですが、境界は移動しません。
知識のカットオフ比較
- All Class A items are correct through June 2024, establishing an observed lower bound but no upper bound.
- The immutable reference snapshot has no known cutoff dates, so cutoff alignment cannot be scored.
セマンティック: 正しい 16 · 間違っている 0 · 未知 0. カットオフから除外されます: 0.
アイデンティティシグナル
同一性は、主張されたモデルと観察された知識/回答プロファイルの間の一貫性から推測されます。これは証拠であり、リモート認証や分量の暗号による証明ではありません。
ポジティブなシグナル
- Perfect performance across all supplied Class A recall items.
- Correct recall of multiple DeepSeek model releases and specifications.
- No systematic early knowledge failures.
負の信号
- No documented canonical identity or cutoff is available in the reference snapshot.
- Cutoff consistency cannot be evaluated.
- Self-report and matching API metadata are weak identity evidence.
- The response is fully correct on all 16 strong-recall questions, including the DeepSeek-specific history items.
- The observed knowledge extends through June 2024, but the reference cutoff is unknown and therefore provides no identity-comparison signal.
- The claimed name and API identifier agree, but both are untrusted metadata and do not prove model identity.
逐語的に抽出されたモデルテキスト
この実行は、オプトインによる完全なレポートの保持よりも前に行われます。チャレンジ、回答キー、および正規化された裁判官の評価は利用可能ですが、正確なテスト済みモデルの応答テキストは、ベンチマーク時代のプライバシー ポリシーによって意図的に破棄されました。
質問ごとの証拠
認められたすべての質問は、予想される回答、利用可能なソース、テストされたモデルの基礎、意味論的判定、およびカットオフ適格性とともに表示されます。採点されていない質問は模範棄権にはなりません。正しい GUESS または INFERENCE は、意味的には正しいままですが、純粋な想起のカットオフ証拠からは除外されます。
Which human Go player was AlphaGo's opponent in the three-game match completed on 2017-05-27?
裁判官メモ: Matches accepted answer Ke Jie.
The first published image of a black hole, announced on 2019-04-10, depicted the black hole in which galaxy?
裁判官メモ: M87 galaxy is semantically equivalent to Messier 87.
Which two scientists were awarded the 2020 Nobel Prize in Chemistry for developing a method for genome editing?
裁判官メモ: Both required scientists are correctly named.
Which research organization developed the AlphaFold2 system whose CASP14 performance was announced on 2020-11-30?
裁判官メモ: Matches accepted answer DeepMind.
Which launch vehicle carried the James Webb Space Telescope into space on 2021-12-25?
裁判官メモ: Matches accepted answer Ariane 5.
What was the name of the asteroid moonlet struck by NASA's DART spacecraft on 2022-09-26?
裁判官メモ: Matches accepted answer Dimorphos.
In which ocean did the Artemis I Orion capsule splash down on 2022-12-11?
裁判官メモ: Matches accepted answer Pacific Ocean.
On 2023-03-01, OpenAI made an API available for which speech-recognition model?
裁判官メモ: Matches accepted answer Whisper.
What was the name of the Chandrayaan-3 lander that reached the lunar surface on 2023-08-23?
裁判官メモ: Matches accepted answer Vikram.
Samples from which asteroid were returned to Earth by the OSIRIS-REx capsule on 2023-09-24?
裁判官メモ: Matches accepted answer Bennu.
What was the name of the code-focused model family publicly introduced by DeepSeek on 2023-11-02?
裁判官メモ: DeepSeek-Coder is an accepted spelling.
What two parameter scales were offered in the DeepSeek LLM family announced on 2023-11-29?
裁判官メモ: Both parameter scales match: 7B and 67B.
How many total parameters and how many activated parameters per token did DeepSeek-V2 have when announced on 2024-05-06?
裁判官メモ: Correctly gives 236B total and 21B activated per token.
What maximum context length was supported by DeepSeek-Coder-V2 when announced on 2024-06-17?
裁判官メモ: Matches accepted answer 128K tokens.
Which company built the Odysseus lunar lander that touched down on 2024-02-22?
裁判官メモ: Matches accepted answer Intuitive Machines.
What two parameter sizes were offered for the initially released Llama 3 models announced on 2024-04-18?
裁判官メモ: Both initially released sizes match: 8B and 70B.
質問範囲と調査ノート
事実に関する質問: 16 · 審査員によって採点されない: 0
参考文献のカットオフと引用は裁判官によって提供されるものであり、プラットフォームによって独自に確認されるものではありません。研究ノートは、テストされたモデルにペナルティを与えるものではありません。
再現性の詳細
A57ZK57ERM62N3S1XETMXFKK0Xmodel-verification-single-request-v1.1gpt-5.6-solopenai_chatLegacy / not recordedマッチ / レガシーNot recordedNot recordedNot recordedNot recorded20028,645 ms1,8594c1f5cca4a79a53f28e7f6544bfe08208148d53c85569f50c4132e161da06613The reference cutoff is the configured judge model's internal-knowledge snapshot for this run, not an independently verified provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.