deepseek-v4-flash
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/ja/model-verification/results/0B0HVDB2FRJEJC90BRG4D7JV1Yこの検証結果が得られた理由
公開レポートは、/chat/verify の後にテスターに示される分析証拠を反映しています。審査員のスコアはルーブリック評価です。純粋なリコールのカットオフ境界に影響を与えることができるのは、明示的な正解または不正解の RECALL の回答のみです。
ジャッジの採点根拠
このアーカイブされた結果については、別の裁判官の理論的根拠は保持されていません。
このアーカイブされた結果については、別の裁判官の理論的根拠は保持されていません。
このアーカイブされた結果については、別の裁判官の理論的根拠は保持されていません。
このアーカイブされた結果については、別の裁判官の理論的根拠は保持されていません。
過去のベンチマークのメモ: このアーカイブされた v1.4 の全体スコアでは、以前の固定診断集計が使用されています。これは再現性のために保持されており、v1.5 のモデル マッチ スコアとして解釈しないでください。
純粋な想起のタイムライン
実際の質問セットは、テスト済みモデルのリクエストの前にシャッフルされました。世代スロットの欠落はモデル棄権ではありません。タイムラインは、監査可能性を考慮して目標月を再構成します。明示的な「RECALL」の回答のみが完全な証拠となります。 INFERENCE、GUESS、UNKNOWN は表示されたままですが、境界は移動しません。
知識のカットオフ比較
- Strong recall is nearly complete through the May 2025 reference cutoff.
- There is no systematic strong-recall failure substantially before the reference cutoff.
- Correct June and August 2025 answers suggest selective later knowledge, while failures in July, September, and October prevent a clearly later boundary.
セマンティック: 正しい 20 · 間違っている 4 · 未知 0. カットオフから除外されます: 0.
アイデンティティシグナル
同一性は、主張されたモデルと観察された知識/回答プロファイルの間の一貫性から推測されます。これは証拠であり、リモート認証や分量の暗号による証明ではありません。
ポジティブなシグナル
- Excellent recall across late 2023 through May 2025.
- Correct near-cutoff answer for the WHO Pandemic Agreement.
- No pattern of early knowledge degradation.
負の信号
- Incorrect February 2024 lunar-lander answer despite high confidence.
- Correct June and August 2025 answers extend beyond the estimated reference cutoff.
- Several post-cutoff answers are confidently wrong, indicating uneven recency.
- Observed behavior broadly overlaps the saved May 2025 reference cutoff.
- The model answered 20 of 24 Class A items correctly, with only one isolated failure well before the cutoff.
- Post-cutoff knowledge is uneven and can plausibly reflect selective updates rather than model substitution.
逐語的に抽出されたモデルテキスト
この実行は、オプトインによる完全なレポートの保持よりも前に行われます。チャレンジ、回答キー、および正規化された裁判官の評価は利用可能ですが、正確なテスト済みモデルの応答テキストは、ベンチマーク時代のプライバシー ポリシーによって意図的に破棄されました。
質問ごとの証拠
認められたすべての質問は、予想される回答、利用可能なソース、テストされたモデルの基礎、意味論的判定、およびカットオフ適格性とともに表示されます。採点されていない質問は模範棄権にはなりません。正しい GUESS または INFERENCE は、意味的には正しいままですが、純粋な想起のカットオフ証拠からは除外されます。
Who won the men's 100 metres at the World Athletics Championships in Tokyo?
裁判官メモ: Expected Oblique Seville; answered Kishane Thompson.
Which team won the UEFA Women's Euro 2025 final?
裁判官メモ: Expected England; answered Spain.
Which disease's African upsurge did WHO declare a public health emergency of international concern?
裁判官メモ: Matches mpox.
What molecular-structure prediction model did Google DeepMind and Isomorphic Labs introduce?
裁判官メモ: Matches AlphaFold 3.
What model did OpenAI release as a unified system with built-in thinking?
裁判官メモ: Matches GPT-5.
Which operating system was affected by the faulty CrowdStrike Falcon content update?
裁判官メモ: Microsoft Windows is accepted.
What was the nickname of Intuitive Machines' Nova-C lander that touched down on the Moon?
裁判官メモ: Expected Odysseus; Odie is not an accepted answer.
What was the name of Japan's lunar lander that successfully reached the Moon?
裁判官メモ: Matches SLIM.
Who became the first private astronaut to perform a spacewalk?
裁判官メモ: Matches Jared Isaacman.
Which country was appointed host of the 2034 FIFA World Cup?
裁判官メモ: Matches Saudi Arabia.
In which lunar mare did Firefly Aerospace's Blue Ghost land?
裁判官メモ: Matches Mare Crisium.
What model did DeepSeek begin serving through its deepseek-reasoner endpoint?
裁判官メモ: Matches DeepSeek-R1.
What was the name of the second malaria vaccine prequalified by WHO?
裁判官メモ: Matches R21/Matrix-M.
Which album won Album of the Year at the 67th Grammy Awards?
裁判官メモ: Matches Cowboy Carter.
How many European Parliament members voted in favor of the Artificial Intelligence Act?
裁判官メモ: Matches 523.
How many grams of lunar material did the Chang'e-6 mission collect?
裁判官メモ: Matches 1,935.3 grams.
Which rocket launched NASA's Europa Clipper spacecraft?
裁判官メモ: Matches Falcon Heavy.
What designation was given to the newly identified 33-solar-mass stellar black hole in the Milky Way?
裁判官メモ: Matches Gaia BH3.
Who was awarded the 2025 Nobel Peace Prize?
裁判官メモ: Expected Maria Corina Machado; answered an institution.
What annual climate-finance goal for developing countries did COP29 set for 2035?
裁判官メモ: Matches $300 billion per year.
What was the name of the first human spaceflight to orbit over Earth's polar regions?
裁判官メモ: Matches Fram2.
Under which article of the WHO Constitution was the Pandemic Agreement adopted?
裁判官メモ: Matches Article 19.
Which new Mario Kart game launched alongside Nintendo Switch 2?
裁判官メモ: Matches Mario Kart World.
What model did OpenAI introduce with a 128K context window at its first DevDay?
裁判官メモ: Matches GPT-4 Turbo.
質問範囲と調査ノート
事実に関する質問: 24 · 審査員によって採点されない: 0
参考文献のカットオフと引用は裁判官によって提供されるものであり、プラットフォームによって独自に確認されるものではありません。研究ノートは、テストされたモデルにペナルティを与えるものではありません。
再現性の詳細
0B0HVDB2FRJEJC90BRG4D7JV1Ymodel-verification-single-request-v1.2gpt-5.6-solopenai_chatLegacy / not recordedマッチ / レガシーNot recordedNot recordedNot recordedNot recorded20088,304 ms2,460f22e63d94e67fbd4fdfa25a27a37ae4b3402f94fc6801d14c2a4e638f85cfd9dThe reference cutoff is a judge-produced snapshot grounded in hosted web evidence when available (official sources preferred, otherwise the best defensible estimate). It is not provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.