deepseek-v4-flash
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/pt/model-verification/results/A57ZK57ERM62N3S1XETMXFKK0XPor que esta verificação recebeu seu resultado
O relatório público reflete a evidência analítica mostrada ao testador após /chat/verify. As pontuações dos juízes são avaliações de rubricas; apenas respostas RECALL corretas ou erradas explícitas podem influenciar o limite de corte de recordação pura.
Justificativa de pontuação do juiz
Nenhuma justificativa separada do juiz foi mantida para este resultado arquivado.
Nenhuma justificativa separada do juiz foi mantida para este resultado arquivado.
Nenhuma justificativa separada do juiz foi mantida para este resultado arquivado.
Nenhuma justificativa separada do juiz foi mantida para este resultado arquivado.
Nota de referência histórica: esta pontuação geral arquivada da versão 1.4 usou a antiga agregação de diagnóstico fixa. Ela é mantida para fins de reprodutibilidade e não deve ser interpretada como uma pontuação de Model Match v1.5.
Linha do tempo de recall puro
O conjunto de perguntas real foi embaralhado antes da solicitação do modelo testado. Os slots de geração ausentes não são abstenções modelo. A linha do tempo reconstrói os meses-alvo para auditabilidade. Somente respostas RECALL explícitas são evidências de corte; INFERENCE, GUESS e UNKNOWN permanecem visíveis, mas não movem o limite.
Comparação de limite de conhecimento
- All Class A items are correct through June 2024, establishing an observed lower bound but no upper bound.
- The immutable reference snapshot has no known cutoff dates, so cutoff alignment cannot be scored.
Semântica: correto 16 · errado 0 · desconhecido 0. Excluído do corte: 0.
Sinais de identidade
A identidade é inferida a partir da consistência entre o modelo reivindicado e o perfil de conhecimento/resposta observado. É uma evidência, não um atestado remoto ou prova criptográfica dos pesos da porção.
Sinais positivos
- Perfect performance across all supplied Class A recall items.
- Correct recall of multiple DeepSeek model releases and specifications.
- No systematic early knowledge failures.
Sinais negativos
- No documented canonical identity or cutoff is available in the reference snapshot.
- Cutoff consistency cannot be evaluated.
- Self-report and matching API metadata are weak identity evidence.
- The response is fully correct on all 16 strong-recall questions, including the DeepSeek-specific history items.
- The observed knowledge extends through June 2024, but the reference cutoff is unknown and therefore provides no identity-comparison signal.
- The claimed name and API identifier agree, but both are untrusted metadata and do not prove model identity.
Texto do modelo extraído literalmente
Essa execução é anterior à retenção de relatório completo opcional. O desafio, a resposta e a avaliação normalizada do juiz estão disponíveis, mas o texto exato da resposta do modelo testado foi intencionalmente descartado pela política de privacidade da era do benchmark.
Evidência pergunta por pergunta
Todas as questões admitidas são mostradas com a resposta esperada, fonte disponível, base do modelo testado, veredicto semântico e elegibilidade de corte. Perguntas não avaliadas não são abstenções modelo. Uma GUESS ou INFERENCE correta permanece semanticamente correta, mas é excluída da evidência de corte de recordação pura.
Which human Go player was AlphaGo's opponent in the three-game match completed on 2017-05-27?
Nota do juiz: Matches accepted answer Ke Jie.
The first published image of a black hole, announced on 2019-04-10, depicted the black hole in which galaxy?
Nota do juiz: M87 galaxy is semantically equivalent to Messier 87.
Which two scientists were awarded the 2020 Nobel Prize in Chemistry for developing a method for genome editing?
Nota do juiz: Both required scientists are correctly named.
Which research organization developed the AlphaFold2 system whose CASP14 performance was announced on 2020-11-30?
Nota do juiz: Matches accepted answer DeepMind.
Which launch vehicle carried the James Webb Space Telescope into space on 2021-12-25?
Nota do juiz: Matches accepted answer Ariane 5.
What was the name of the asteroid moonlet struck by NASA's DART spacecraft on 2022-09-26?
Nota do juiz: Matches accepted answer Dimorphos.
In which ocean did the Artemis I Orion capsule splash down on 2022-12-11?
Nota do juiz: Matches accepted answer Pacific Ocean.
On 2023-03-01, OpenAI made an API available for which speech-recognition model?
Nota do juiz: Matches accepted answer Whisper.
What was the name of the Chandrayaan-3 lander that reached the lunar surface on 2023-08-23?
Nota do juiz: Matches accepted answer Vikram.
Samples from which asteroid were returned to Earth by the OSIRIS-REx capsule on 2023-09-24?
Nota do juiz: Matches accepted answer Bennu.
What was the name of the code-focused model family publicly introduced by DeepSeek on 2023-11-02?
Nota do juiz: DeepSeek-Coder is an accepted spelling.
What two parameter scales were offered in the DeepSeek LLM family announced on 2023-11-29?
Nota do juiz: Both parameter scales match: 7B and 67B.
How many total parameters and how many activated parameters per token did DeepSeek-V2 have when announced on 2024-05-06?
Nota do juiz: Correctly gives 236B total and 21B activated per token.
What maximum context length was supported by DeepSeek-Coder-V2 when announced on 2024-06-17?
Nota do juiz: Matches accepted answer 128K tokens.
Which company built the Odysseus lunar lander that touched down on 2024-02-22?
Nota do juiz: Matches accepted answer Intuitive Machines.
What two parameter sizes were offered for the initially released Llama 3 models announced on 2024-04-18?
Nota do juiz: Both initially released sizes match: 8B and 70B.
Cobertura de perguntas e notas de pesquisa
Perguntas factuais: 16 · Não avaliado pelo juiz: 0
O corte de referência e as citações são fornecidos pelo juiz, não confirmados de forma independente pela plataforma. As notas de pesquisa não penalizam o modelo testado.
Detalhes de reprodutibilidade
A57ZK57ERM62N3S1XETMXFKK0Xmodel-verification-single-request-v1.1gpt-5.6-solopenai_chatLegacy / not recordedPARTIDA / legadoNot recordedNot recordedNot recordedNot recorded20028,645 ms1,8594c1f5cca4a79a53f28e7f6544bfe08208148d53c85569f50c4132e161da06613The reference cutoff is the configured judge model's internal-knowledge snapshot for this run, not an independently verified provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.
Execute uma verificação independente
Teste você mesmo o mesmo endpoint ou navegue em outros relatórios públicos antes de confiar em uma identidade de modelo anunciada.