deepseek-v4-flash
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/it/model-verification/results/A57ZK57ERM62N3S1XETMXFKK0XPerché questa verifica ha ricevuto il suo risultato
Il rapporto pubblico rispecchia l'evidenza analitica mostrata al tester dopo /chat/verify. I punteggi dei giudici sono valutazioni delle rubriche; solo le risposte esplicite RECALL corrette o errate possono influenzare il limite di interruzione del puro richiamo.
Motivazione del punteggio del giudice
Per questo risultato archiviato non è stata mantenuta alcuna motivazione separata da parte del giudice.
Per questo risultato archiviato non è stata mantenuta alcuna motivazione separata da parte del giudice.
Per questo risultato archiviato non è stata mantenuta alcuna motivazione separata da parte del giudice.
Per questo risultato archiviato non è stata mantenuta alcuna motivazione separata da parte del giudice.
Nota sul benchmark storico: questo punteggio complessivo archiviato v1.4 utilizzava la precedente aggregazione diagnostica fissa. Viene conservato per motivi di riproducibilità e non deve essere interpretato come un punteggio di corrispondenza del modello v1.5.
Cronologia di puro ricordo
Il set di domande effettivo è stato mescolato prima della richiesta del modello testato. Gli slot generazionali mancanti non costituiscono un’astensione dal modello. La sequenza temporale ricostruisce i mesi target per la verificabilità. Solo le risposte esplicite RECALL sono prove tagliate; INFERENZA, IMMAGINE e SCONOSCIUTO rimangono visibili ma non spostano il confine.
Confronto del limite di conoscenza
- All Class A items are correct through June 2024, establishing an observed lower bound but no upper bound.
- The immutable reference snapshot has no known cutoff dates, so cutoff alignment cannot be scored.
Semantico: corretto 16 · sbagliato 0 · sconosciuto 0. Escluso dal taglio: 0.
Segnali di identità
L'identità è dedotta dalla coerenza tra il modello affermato e il profilo di conoscenza/risposta osservato. Si tratta di una prova, non di un'attestazione remota o di una prova crittografica dei pesi di servizio.
Segnali positivi
- Perfect performance across all supplied Class A recall items.
- Correct recall of multiple DeepSeek model releases and specifications.
- No systematic early knowledge failures.
Segnali negativi
- No documented canonical identity or cutoff is available in the reference snapshot.
- Cutoff consistency cannot be evaluated.
- Self-report and matching API metadata are weak identity evidence.
- The response is fully correct on all 16 strong-recall questions, including the DeepSeek-specific history items.
- The observed knowledge extends through June 2024, but the reference cutoff is unknown and therefore provides no identity-comparison signal.
- The claimed name and API identifier agree, but both are untrusted metadata and do not prove model identity.
Testo del modello estratto parola per parola
Questa esecuzione è antecedente al consenso esplicito alla conservazione del report completo. Sono disponibili la sfida, la chiave di risposta e la valutazione normalizzata del giudice, ma il testo esatto della risposta del modello testato è stato intenzionalmente scartato dalla politica sulla privacy dell'era del benchmark.
Prove domanda per domanda
Tutte le domande ammesse vengono visualizzate con la risposta prevista, la fonte disponibile, la base del modello testato, il verdetto semantico e l'idoneità limite. Le domande senza voto non sono astensioni modello. Un'ipotesi o un'inferenza corretta rimane semanticamente corretta ma è esclusa dall'evidenza di esclusione del puro ricordo.
Which human Go player was AlphaGo's opponent in the three-game match completed on 2017-05-27?
Nota del giudice: Matches accepted answer Ke Jie.
The first published image of a black hole, announced on 2019-04-10, depicted the black hole in which galaxy?
Nota del giudice: M87 galaxy is semantically equivalent to Messier 87.
Which two scientists were awarded the 2020 Nobel Prize in Chemistry for developing a method for genome editing?
Nota del giudice: Both required scientists are correctly named.
Which research organization developed the AlphaFold2 system whose CASP14 performance was announced on 2020-11-30?
Nota del giudice: Matches accepted answer DeepMind.
Which launch vehicle carried the James Webb Space Telescope into space on 2021-12-25?
Nota del giudice: Matches accepted answer Ariane 5.
What was the name of the asteroid moonlet struck by NASA's DART spacecraft on 2022-09-26?
Nota del giudice: Matches accepted answer Dimorphos.
In which ocean did the Artemis I Orion capsule splash down on 2022-12-11?
Nota del giudice: Matches accepted answer Pacific Ocean.
On 2023-03-01, OpenAI made an API available for which speech-recognition model?
Nota del giudice: Matches accepted answer Whisper.
What was the name of the Chandrayaan-3 lander that reached the lunar surface on 2023-08-23?
Nota del giudice: Matches accepted answer Vikram.
Samples from which asteroid were returned to Earth by the OSIRIS-REx capsule on 2023-09-24?
Nota del giudice: Matches accepted answer Bennu.
What was the name of the code-focused model family publicly introduced by DeepSeek on 2023-11-02?
Nota del giudice: DeepSeek-Coder is an accepted spelling.
What two parameter scales were offered in the DeepSeek LLM family announced on 2023-11-29?
Nota del giudice: Both parameter scales match: 7B and 67B.
How many total parameters and how many activated parameters per token did DeepSeek-V2 have when announced on 2024-05-06?
Nota del giudice: Correctly gives 236B total and 21B activated per token.
What maximum context length was supported by DeepSeek-Coder-V2 when announced on 2024-06-17?
Nota del giudice: Matches accepted answer 128K tokens.
Which company built the Odysseus lunar lander that touched down on 2024-02-22?
Nota del giudice: Matches accepted answer Intuitive Machines.
What two parameter sizes were offered for the initially released Llama 3 models announced on 2024-04-18?
Nota del giudice: Both initially released sizes match: 8B and 70B.
Copertura delle domande e note di ricerca
Domande fattuali: 16 · Non valutato dal giudice: 0
Il limite di riferimento e le citazioni sono forniti dal giudice, non confermati in modo indipendente dalla piattaforma. Le note di ricerca non penalizzano il modello testato.
Dettagli sulla riproducibilità
A57ZK57ERM62N3S1XETMXFKK0Xmodel-verification-single-request-v1.1gpt-5.6-solopenai_chatLegacy / not recordedPARTITA / ereditàNot recordedNot recordedNot recordedNot recorded20028,645 ms1,8594c1f5cca4a79a53f28e7f6544bfe08208148d53c85569f50c4132e161da06613The reference cutoff is the configured judge model's internal-knowledge snapshot for this run, not an independently verified provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.
Esegui una verifica indipendente
Testa tu stesso lo stesso endpoint o consulta altri report pubblici prima di fare affidamento sull'identità di un modello pubblicizzato.