deepseek-v4-flash
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/fr/model-verification/results/0B0HVDB2FRJEJC90BRG4D7JV1YPourquoi cette vérification a reçu son résultat
Le rapport public reflète les preuves analytiques présentées au testeur après /chat/verify. Les notes des juges sont des évaluations par rubrique ; seules les réponses explicites correctes ou erronées de RECALL peuvent influencer la limite du rappel pur.
Justification de la notation du juge
Aucune justification distincte du juge n’a été retenue pour ce résultat archivé.
Aucune justification distincte du juge n’a été retenue pour ce résultat archivé.
Aucune justification distincte du juge n’a été retenue pour ce résultat archivé.
Aucune justification distincte du juge n’a été retenue pour ce résultat archivé.
Note de référence historique : ce score global v1.4 archivé utilisait l'ancienne agrégation de diagnostic fixe. Il est conservé par souci de reproductibilité et ne doit pas être interprété comme un score de correspondance de modèle v1.5.
Chronologie de pur rappel
L'ensemble de questions réel a été mélangé avant la demande de modèle testé. Les créneaux de génération manquants ne constituent pas des abstentions de modèle. La chronologie reconstitue leurs mois cibles pour l’auditabilité. Seules les réponses explicites de RECALL constituent une preuve de coupure ; INFERENCE, GUESS et UNKNOWN restent visibles mais ne déplacent pas la limite.
Comparaison des seuils de connaissances
- Strong recall is nearly complete through the May 2025 reference cutoff.
- There is no systematic strong-recall failure substantially before the reference cutoff.
- Correct June and August 2025 answers suggest selective later knowledge, while failures in July, September, and October prevent a clearly later boundary.
Sémantique: correct 20 · faux 4 · inconnu 0. Exclus du seuil: 0.
Signaux d'identité
L'identité est déduite de la cohérence entre le modèle revendiqué et le profil connaissance/réponse observé. Il s’agit d’une preuve, et non d’une attestation à distance ou d’une preuve cryptographique du poids servi.
Signaux positifs
- Excellent recall across late 2023 through May 2025.
- Correct near-cutoff answer for the WHO Pandemic Agreement.
- No pattern of early knowledge degradation.
Signaux négatifs
- Incorrect February 2024 lunar-lander answer despite high confidence.
- Correct June and August 2025 answers extend beyond the estimated reference cutoff.
- Several post-cutoff answers are confidently wrong, indicating uneven recency.
- Observed behavior broadly overlaps the saved May 2025 reference cutoff.
- The model answered 20 of 24 Class A items correctly, with only one isolated failure well before the cutoff.
- Post-cutoff knowledge is uneven and can plausibly reflect selective updates rather than model substitution.
Texte du modèle extrait textuellement
Cette exécution est antérieure à la conservation opt-in du rapport complet. Le défi, le corrigé et l'évaluation normalisée du juge sont disponibles, mais le texte de réponse exact du modèle testé a été intentionnellement rejeté par cette politique de confidentialité de l'ère de référence.
Preuve question par question
Toutes les questions admises sont présentées avec la réponse attendue, la source disponible, la base du modèle testé, le verdict sémantique et le seuil d'éligibilité. Les questions non notées ne constituent pas des abstentions modèles. Une DEVINATION ou UNE INFÉRENCE correcte reste sémantiquement correcte mais est exclue des preuves de coupure de rappel pur.
Who won the men's 100 metres at the World Athletics Championships in Tokyo?
Note du juge: Expected Oblique Seville; answered Kishane Thompson.
Which team won the UEFA Women's Euro 2025 final?
Note du juge: Expected England; answered Spain.
Which disease's African upsurge did WHO declare a public health emergency of international concern?
Note du juge: Matches mpox.
What molecular-structure prediction model did Google DeepMind and Isomorphic Labs introduce?
Note du juge: Matches AlphaFold 3.
What model did OpenAI release as a unified system with built-in thinking?
Note du juge: Matches GPT-5.
Which operating system was affected by the faulty CrowdStrike Falcon content update?
Note du juge: Microsoft Windows is accepted.
What was the nickname of Intuitive Machines' Nova-C lander that touched down on the Moon?
Note du juge: Expected Odysseus; Odie is not an accepted answer.
What was the name of Japan's lunar lander that successfully reached the Moon?
Note du juge: Matches SLIM.
Who became the first private astronaut to perform a spacewalk?
Note du juge: Matches Jared Isaacman.
Which country was appointed host of the 2034 FIFA World Cup?
Note du juge: Matches Saudi Arabia.
In which lunar mare did Firefly Aerospace's Blue Ghost land?
Note du juge: Matches Mare Crisium.
What model did DeepSeek begin serving through its deepseek-reasoner endpoint?
Note du juge: Matches DeepSeek-R1.
What was the name of the second malaria vaccine prequalified by WHO?
Note du juge: Matches R21/Matrix-M.
Which album won Album of the Year at the 67th Grammy Awards?
Note du juge: Matches Cowboy Carter.
How many European Parliament members voted in favor of the Artificial Intelligence Act?
Note du juge: Matches 523.
How many grams of lunar material did the Chang'e-6 mission collect?
Note du juge: Matches 1,935.3 grams.
Which rocket launched NASA's Europa Clipper spacecraft?
Note du juge: Matches Falcon Heavy.
What designation was given to the newly identified 33-solar-mass stellar black hole in the Milky Way?
Note du juge: Matches Gaia BH3.
Who was awarded the 2025 Nobel Peace Prize?
Note du juge: Expected Maria Corina Machado; answered an institution.
What annual climate-finance goal for developing countries did COP29 set for 2035?
Note du juge: Matches $300 billion per year.
What was the name of the first human spaceflight to orbit over Earth's polar regions?
Note du juge: Matches Fram2.
Under which article of the WHO Constitution was the Pandemic Agreement adopted?
Note du juge: Matches Article 19.
Which new Mario Kart game launched alongside Nintendo Switch 2?
Note du juge: Matches Mario Kart World.
What model did OpenAI introduce with a 128K context window at its first DevDay?
Note du juge: Matches GPT-4 Turbo.
Couverture des questions et notes de recherche
Questions factuelles: 24 · Non noté par le juge: 0
Les seuils de référence et les citations sont fournis par le juge, non confirmés de manière indépendante par la plateforme. Les notes de recherche ne pénalisent pas le modèle testé.
Détails de reproductibilité
0B0HVDB2FRJEJC90BRG4D7JV1Ymodel-verification-single-request-v1.2gpt-5.6-solopenai_chatLegacy / not recordedMATCH / héritageNot recordedNot recordedNot recordedNot recorded20088,304 ms2,460f22e63d94e67fbd4fdfa25a27a37ae4b3402f94fc6801d14c2a4e638f85cfd9dThe reference cutoff is a judge-produced snapshot grounded in hosted web evidence when available (official sources preferred, otherwise the best defensible estimate). It is not provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.
Exécutez une vérification indépendante
Testez vous-même le même point de terminaison ou parcourez d’autres rapports publics avant de vous fier à une identité de modèle annoncée.