deepseek-v4-flash
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/it/model-verification/results/0B0HVDB2FRJEJC90BRG4D7JV1YPerché questa verifica ha ricevuto il suo risultato
Il rapporto pubblico rispecchia l'evidenza analitica mostrata al tester dopo /chat/verify. I punteggi dei giudici sono valutazioni delle rubriche; solo le risposte esplicite RECALL corrette o errate possono influenzare il limite di interruzione del puro richiamo.
Motivazione del punteggio del giudice
Per questo risultato archiviato non è stata mantenuta alcuna motivazione separata da parte del giudice.
Per questo risultato archiviato non è stata mantenuta alcuna motivazione separata da parte del giudice.
Per questo risultato archiviato non è stata mantenuta alcuna motivazione separata da parte del giudice.
Per questo risultato archiviato non è stata mantenuta alcuna motivazione separata da parte del giudice.
Nota sul benchmark storico: questo punteggio complessivo archiviato v1.4 utilizzava la precedente aggregazione diagnostica fissa. Viene conservato per motivi di riproducibilità e non deve essere interpretato come un punteggio di corrispondenza del modello v1.5.
Cronologia di puro ricordo
Il set di domande effettivo è stato mescolato prima della richiesta del modello testato. Gli slot generazionali mancanti non costituiscono un’astensione dal modello. La sequenza temporale ricostruisce i mesi target per la verificabilità. Solo le risposte esplicite RECALL sono prove tagliate; INFERENZA, IMMAGINE e SCONOSCIUTO rimangono visibili ma non spostano il confine.
Confronto del limite di conoscenza
- Strong recall is nearly complete through the May 2025 reference cutoff.
- There is no systematic strong-recall failure substantially before the reference cutoff.
- Correct June and August 2025 answers suggest selective later knowledge, while failures in July, September, and October prevent a clearly later boundary.
Semantico: corretto 20 · sbagliato 4 · sconosciuto 0. Escluso dal taglio: 0.
Segnali di identità
L'identità è dedotta dalla coerenza tra il modello affermato e il profilo di conoscenza/risposta osservato. Si tratta di una prova, non di un'attestazione remota o di una prova crittografica dei pesi di servizio.
Segnali positivi
- Excellent recall across late 2023 through May 2025.
- Correct near-cutoff answer for the WHO Pandemic Agreement.
- No pattern of early knowledge degradation.
Segnali negativi
- Incorrect February 2024 lunar-lander answer despite high confidence.
- Correct June and August 2025 answers extend beyond the estimated reference cutoff.
- Several post-cutoff answers are confidently wrong, indicating uneven recency.
- Observed behavior broadly overlaps the saved May 2025 reference cutoff.
- The model answered 20 of 24 Class A items correctly, with only one isolated failure well before the cutoff.
- Post-cutoff knowledge is uneven and can plausibly reflect selective updates rather than model substitution.
Testo del modello estratto parola per parola
Questa esecuzione è antecedente al consenso esplicito alla conservazione del report completo. Sono disponibili la sfida, la chiave di risposta e la valutazione normalizzata del giudice, ma il testo esatto della risposta del modello testato è stato intenzionalmente scartato dalla politica sulla privacy dell'era del benchmark.
Prove domanda per domanda
Tutte le domande ammesse vengono visualizzate con la risposta prevista, la fonte disponibile, la base del modello testato, il verdetto semantico e l'idoneità limite. Le domande senza voto non sono astensioni modello. Un'ipotesi o un'inferenza corretta rimane semanticamente corretta ma è esclusa dall'evidenza di esclusione del puro ricordo.
Who won the men's 100 metres at the World Athletics Championships in Tokyo?
Nota del giudice: Expected Oblique Seville; answered Kishane Thompson.
Which team won the UEFA Women's Euro 2025 final?
Nota del giudice: Expected England; answered Spain.
Which disease's African upsurge did WHO declare a public health emergency of international concern?
Nota del giudice: Matches mpox.
What molecular-structure prediction model did Google DeepMind and Isomorphic Labs introduce?
Nota del giudice: Matches AlphaFold 3.
What model did OpenAI release as a unified system with built-in thinking?
Nota del giudice: Matches GPT-5.
Which operating system was affected by the faulty CrowdStrike Falcon content update?
Nota del giudice: Microsoft Windows is accepted.
What was the nickname of Intuitive Machines' Nova-C lander that touched down on the Moon?
Nota del giudice: Expected Odysseus; Odie is not an accepted answer.
What was the name of Japan's lunar lander that successfully reached the Moon?
Nota del giudice: Matches SLIM.
Who became the first private astronaut to perform a spacewalk?
Nota del giudice: Matches Jared Isaacman.
Which country was appointed host of the 2034 FIFA World Cup?
Nota del giudice: Matches Saudi Arabia.
In which lunar mare did Firefly Aerospace's Blue Ghost land?
Nota del giudice: Matches Mare Crisium.
What model did DeepSeek begin serving through its deepseek-reasoner endpoint?
Nota del giudice: Matches DeepSeek-R1.
What was the name of the second malaria vaccine prequalified by WHO?
Nota del giudice: Matches R21/Matrix-M.
Which album won Album of the Year at the 67th Grammy Awards?
Nota del giudice: Matches Cowboy Carter.
How many European Parliament members voted in favor of the Artificial Intelligence Act?
Nota del giudice: Matches 523.
How many grams of lunar material did the Chang'e-6 mission collect?
Nota del giudice: Matches 1,935.3 grams.
Which rocket launched NASA's Europa Clipper spacecraft?
Nota del giudice: Matches Falcon Heavy.
What designation was given to the newly identified 33-solar-mass stellar black hole in the Milky Way?
Nota del giudice: Matches Gaia BH3.
Who was awarded the 2025 Nobel Peace Prize?
Nota del giudice: Expected Maria Corina Machado; answered an institution.
What annual climate-finance goal for developing countries did COP29 set for 2035?
Nota del giudice: Matches $300 billion per year.
What was the name of the first human spaceflight to orbit over Earth's polar regions?
Nota del giudice: Matches Fram2.
Under which article of the WHO Constitution was the Pandemic Agreement adopted?
Nota del giudice: Matches Article 19.
Which new Mario Kart game launched alongside Nintendo Switch 2?
Nota del giudice: Matches Mario Kart World.
What model did OpenAI introduce with a 128K context window at its first DevDay?
Nota del giudice: Matches GPT-4 Turbo.
Copertura delle domande e note di ricerca
Domande fattuali: 24 · Non valutato dal giudice: 0
Il limite di riferimento e le citazioni sono forniti dal giudice, non confermati in modo indipendente dalla piattaforma. Le note di ricerca non penalizzano il modello testato.
Dettagli sulla riproducibilità
0B0HVDB2FRJEJC90BRG4D7JV1Ymodel-verification-single-request-v1.2gpt-5.6-solopenai_chatLegacy / not recordedPARTITA / ereditàNot recordedNot recordedNot recordedNot recorded20088,304 ms2,460f22e63d94e67fbd4fdfa25a27a37ae4b3402f94fc6801d14c2a4e638f85cfd9dThe reference cutoff is a judge-produced snapshot grounded in hosted web evidence when available (official sources preferred, otherwise the best defensible estimate). It is not provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.
Esegui una verifica indipendente
Testa tu stesso lo stesso endpoint o consulta altri report pubblici prima di fare affidamento sull'identità di un modello pubblicizzato.