deepseek-v4-flash
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/es/model-verification/results/0B0HVDB2FRJEJC90BRG4D7JV1Y¿Por qué esta verificación recibió su resultado?
El informe público refleja la evidencia analítica que se muestra al evaluador después de /chat/verify. Las puntuaciones de los jueces son evaluaciones de rúbricas; sólo las respuestas RECALL explícitas, correctas o incorrectas, pueden influir en el límite de recuperación pura.
Justificación de la puntuación del juez
No se mantuvo ninguna justificación del juez por separado para este resultado archivado.
No se mantuvo ninguna justificación del juez por separado para este resultado archivado.
No se mantuvo ninguna justificación del juez por separado para este resultado archivado.
No se mantuvo ninguna justificación del juez por separado para este resultado archivado.
Nota de referencia histórica: esta puntuación general archivada v1.4 utilizó la agregación de diagnóstico fija anterior. Se conserva para fines de reproducibilidad y no debe interpretarse como una puntuación de Model Match v1.5.
Cronología de recuerdo puro
El conjunto de preguntas real se barajó antes de la solicitud del modelo probado. Los espacios de generación que faltan no son abstenciones modelo. El cronograma reconstruye los meses objetivo para la auditabilidad. Sólo las respuestas RECALL explícitas son evidencia de corte; INFERENCIA, ADIVINAR y DESCONOCIDO permanecen visibles pero no mueven el límite.
Comparación de límites de conocimiento
- Strong recall is nearly complete through the May 2025 reference cutoff.
- There is no systematic strong-recall failure substantially before the reference cutoff.
- Correct June and August 2025 answers suggest selective later knowledge, while failures in July, September, and October prevent a clearly later boundary.
Semántico: correcto 20 · equivocado 4 · desconocido 0. Excluido del límite: 0.
Señales de identidad
La identidad se infiere de la coherencia entre el modelo reivindicado y el perfil de conocimiento/respuesta observado. Es una evidencia, no una certificación remota o una prueba criptográfica de los pesos de las porciones.
Señales positivas
- Excellent recall across late 2023 through May 2025.
- Correct near-cutoff answer for the WHO Pandemic Agreement.
- No pattern of early knowledge degradation.
Señales negativas
- Incorrect February 2024 lunar-lander answer despite high confidence.
- Correct June and August 2025 answers extend beyond the estimated reference cutoff.
- Several post-cutoff answers are confidently wrong, indicating uneven recency.
- Observed behavior broadly overlaps the saved May 2025 reference cutoff.
- The model answered 20 of 24 Class A items correctly, with only one isolated failure well before the cutoff.
- Post-cutoff knowledge is uneven and can plausibly reflect selective updates rather than model substitution.
Texto del modelo extraído palabra por palabra
Esta ejecución es anterior a la retención voluntaria del informe completo. El desafío, la clave de respuestas y la evaluación del juez normalizada están disponibles, pero el texto exacto de la respuesta del modelo probado fue descartado intencionalmente por esa política de privacidad de la era de los puntos de referencia.
Evidencia pregunta por pregunta
Todas las preguntas admitidas se muestran con la respuesta esperada, la fuente disponible, la base del modelo probado, el veredicto semántico y el límite de elegibilidad. Las preguntas sin calificación no son abstenciones modelo. Una CONDICIÓN o INFERENCIA correcta sigue siendo semánticamente correcta, pero se excluye de la evidencia de corte de recuerdo puro.
Who won the men's 100 metres at the World Athletics Championships in Tokyo?
nota del juez: Expected Oblique Seville; answered Kishane Thompson.
Which team won the UEFA Women's Euro 2025 final?
nota del juez: Expected England; answered Spain.
Which disease's African upsurge did WHO declare a public health emergency of international concern?
nota del juez: Matches mpox.
What molecular-structure prediction model did Google DeepMind and Isomorphic Labs introduce?
nota del juez: Matches AlphaFold 3.
What model did OpenAI release as a unified system with built-in thinking?
nota del juez: Matches GPT-5.
Which operating system was affected by the faulty CrowdStrike Falcon content update?
nota del juez: Microsoft Windows is accepted.
What was the nickname of Intuitive Machines' Nova-C lander that touched down on the Moon?
nota del juez: Expected Odysseus; Odie is not an accepted answer.
What was the name of Japan's lunar lander that successfully reached the Moon?
nota del juez: Matches SLIM.
Who became the first private astronaut to perform a spacewalk?
nota del juez: Matches Jared Isaacman.
Which country was appointed host of the 2034 FIFA World Cup?
nota del juez: Matches Saudi Arabia.
In which lunar mare did Firefly Aerospace's Blue Ghost land?
nota del juez: Matches Mare Crisium.
What model did DeepSeek begin serving through its deepseek-reasoner endpoint?
nota del juez: Matches DeepSeek-R1.
What was the name of the second malaria vaccine prequalified by WHO?
nota del juez: Matches R21/Matrix-M.
Which album won Album of the Year at the 67th Grammy Awards?
nota del juez: Matches Cowboy Carter.
How many European Parliament members voted in favor of the Artificial Intelligence Act?
nota del juez: Matches 523.
How many grams of lunar material did the Chang'e-6 mission collect?
nota del juez: Matches 1,935.3 grams.
Which rocket launched NASA's Europa Clipper spacecraft?
nota del juez: Matches Falcon Heavy.
What designation was given to the newly identified 33-solar-mass stellar black hole in the Milky Way?
nota del juez: Matches Gaia BH3.
Who was awarded the 2025 Nobel Peace Prize?
nota del juez: Expected Maria Corina Machado; answered an institution.
What annual climate-finance goal for developing countries did COP29 set for 2035?
nota del juez: Matches $300 billion per year.
What was the name of the first human spaceflight to orbit over Earth's polar regions?
nota del juez: Matches Fram2.
Under which article of the WHO Constitution was the Pandemic Agreement adopted?
nota del juez: Matches Article 19.
Which new Mario Kart game launched alongside Nintendo Switch 2?
nota del juez: Matches Mario Kart World.
What model did OpenAI introduce with a 128K context window at its first DevDay?
nota del juez: Matches GPT-4 Turbo.
Cobertura de preguntas y notas de investigación.
Preguntas fácticas: 24 · No calificado por el juez: 0
El corte de referencia y las citas son proporcionadas por el juez, no confirmadas de forma independiente por la plataforma. Las notas de investigación no penalizan el modelo probado.
Detalles de reproducibilidad
0B0HVDB2FRJEJC90BRG4D7JV1Ymodel-verification-single-request-v1.2gpt-5.6-solopenai_chatLegacy / not recordedPARTIDO / legadoNot recordedNot recordedNot recordedNot recorded20088,304 ms2,460f22e63d94e67fbd4fdfa25a27a37ae4b3402f94fc6801d14c2a4e638f85cfd9dThe reference cutoff is a judge-produced snapshot grounded in hosted web evidence when available (official sources preferred, otherwise the best defensible estimate). It is not provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.
Ejecute una verificación independiente
Pruebe usted mismo el mismo punto final o explore otros informes públicos antes de confiar en una identidad de modelo anunciada.