deepseek-v4-flash
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/es/model-verification/results/A57ZK57ERM62N3S1XETMXFKK0X¿Por qué esta verificación recibió su resultado?
El informe público refleja la evidencia analítica que se muestra al evaluador después de /chat/verify. Las puntuaciones de los jueces son evaluaciones de rúbricas; sólo las respuestas RECALL explícitas, correctas o incorrectas, pueden influir en el límite de recuperación pura.
Justificación de la puntuación del juez
No se mantuvo ninguna justificación del juez por separado para este resultado archivado.
No se mantuvo ninguna justificación del juez por separado para este resultado archivado.
No se mantuvo ninguna justificación del juez por separado para este resultado archivado.
No se mantuvo ninguna justificación del juez por separado para este resultado archivado.
Nota de referencia histórica: esta puntuación general archivada v1.4 utilizó la agregación de diagnóstico fija anterior. Se conserva para fines de reproducibilidad y no debe interpretarse como una puntuación de Model Match v1.5.
Cronología de recuerdo puro
El conjunto de preguntas real se barajó antes de la solicitud del modelo probado. Los espacios de generación que faltan no son abstenciones modelo. El cronograma reconstruye los meses objetivo para la auditabilidad. Sólo las respuestas RECALL explícitas son evidencia de corte; INFERENCIA, ADIVINAR y DESCONOCIDO permanecen visibles pero no mueven el límite.
Comparación de límites de conocimiento
- All Class A items are correct through June 2024, establishing an observed lower bound but no upper bound.
- The immutable reference snapshot has no known cutoff dates, so cutoff alignment cannot be scored.
Semántico: correcto 16 · equivocado 0 · desconocido 0. Excluido del límite: 0.
Señales de identidad
La identidad se infiere de la coherencia entre el modelo reivindicado y el perfil de conocimiento/respuesta observado. Es una evidencia, no una certificación remota o una prueba criptográfica de los pesos de las porciones.
Señales positivas
- Perfect performance across all supplied Class A recall items.
- Correct recall of multiple DeepSeek model releases and specifications.
- No systematic early knowledge failures.
Señales negativas
- No documented canonical identity or cutoff is available in the reference snapshot.
- Cutoff consistency cannot be evaluated.
- Self-report and matching API metadata are weak identity evidence.
- The response is fully correct on all 16 strong-recall questions, including the DeepSeek-specific history items.
- The observed knowledge extends through June 2024, but the reference cutoff is unknown and therefore provides no identity-comparison signal.
- The claimed name and API identifier agree, but both are untrusted metadata and do not prove model identity.
Texto del modelo extraído palabra por palabra
Esta ejecución es anterior a la retención voluntaria del informe completo. El desafío, la clave de respuestas y la evaluación del juez normalizada están disponibles, pero el texto exacto de la respuesta del modelo probado fue descartado intencionalmente por esa política de privacidad de la era de los puntos de referencia.
Evidencia pregunta por pregunta
Todas las preguntas admitidas se muestran con la respuesta esperada, la fuente disponible, la base del modelo probado, el veredicto semántico y el límite de elegibilidad. Las preguntas sin calificación no son abstenciones modelo. Una CONDICIÓN o INFERENCIA correcta sigue siendo semánticamente correcta, pero se excluye de la evidencia de corte de recuerdo puro.
Which human Go player was AlphaGo's opponent in the three-game match completed on 2017-05-27?
nota del juez: Matches accepted answer Ke Jie.
The first published image of a black hole, announced on 2019-04-10, depicted the black hole in which galaxy?
nota del juez: M87 galaxy is semantically equivalent to Messier 87.
Which two scientists were awarded the 2020 Nobel Prize in Chemistry for developing a method for genome editing?
nota del juez: Both required scientists are correctly named.
Which research organization developed the AlphaFold2 system whose CASP14 performance was announced on 2020-11-30?
nota del juez: Matches accepted answer DeepMind.
Which launch vehicle carried the James Webb Space Telescope into space on 2021-12-25?
nota del juez: Matches accepted answer Ariane 5.
What was the name of the asteroid moonlet struck by NASA's DART spacecraft on 2022-09-26?
nota del juez: Matches accepted answer Dimorphos.
In which ocean did the Artemis I Orion capsule splash down on 2022-12-11?
nota del juez: Matches accepted answer Pacific Ocean.
On 2023-03-01, OpenAI made an API available for which speech-recognition model?
nota del juez: Matches accepted answer Whisper.
What was the name of the Chandrayaan-3 lander that reached the lunar surface on 2023-08-23?
nota del juez: Matches accepted answer Vikram.
Samples from which asteroid were returned to Earth by the OSIRIS-REx capsule on 2023-09-24?
nota del juez: Matches accepted answer Bennu.
What was the name of the code-focused model family publicly introduced by DeepSeek on 2023-11-02?
nota del juez: DeepSeek-Coder is an accepted spelling.
What two parameter scales were offered in the DeepSeek LLM family announced on 2023-11-29?
nota del juez: Both parameter scales match: 7B and 67B.
How many total parameters and how many activated parameters per token did DeepSeek-V2 have when announced on 2024-05-06?
nota del juez: Correctly gives 236B total and 21B activated per token.
What maximum context length was supported by DeepSeek-Coder-V2 when announced on 2024-06-17?
nota del juez: Matches accepted answer 128K tokens.
Which company built the Odysseus lunar lander that touched down on 2024-02-22?
nota del juez: Matches accepted answer Intuitive Machines.
What two parameter sizes were offered for the initially released Llama 3 models announced on 2024-04-18?
nota del juez: Both initially released sizes match: 8B and 70B.
Cobertura de preguntas y notas de investigación.
Preguntas fácticas: 16 · No calificado por el juez: 0
El corte de referencia y las citas son proporcionadas por el juez, no confirmadas de forma independiente por la plataforma. Las notas de investigación no penalizan el modelo probado.
Detalles de reproducibilidad
A57ZK57ERM62N3S1XETMXFKK0Xmodel-verification-single-request-v1.1gpt-5.6-solopenai_chatLegacy / not recordedPARTIDO / legadoNot recordedNot recordedNot recordedNot recorded20028,645 ms1,8594c1f5cca4a79a53f28e7f6544bfe08208148d53c85569f50c4132e161da06613The reference cutoff is the configured judge model's internal-knowledge snapshot for this run, not an independently verified provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.
Ejecute una verificación independiente
Pruebe usted mismo el mismo punto final o explore otros informes públicos antes de confiar en una identidad de modelo anunciada.