claude-opus-4-6
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/es/model-verification/results/68TGGQW6SM9411Y4XDAPH3BR3H¿Por qué esta verificación recibió su resultado?
El informe público refleja la evidencia analítica que se muestra al evaluador después de /chat/verify. Las puntuaciones de los jueces son evaluaciones de rúbricas; sólo las respuestas RECALL explícitas, correctas o incorrectas, pueden influir en el límite de recuperación pura.
Justificación de la puntuación del juez
No se mantuvo ninguna justificación del juez por separado para este resultado archivado.
No se mantuvo ninguna justificación del juez por separado para este resultado archivado.
No se mantuvo ninguna justificación del juez por separado para este resultado archivado.
No se mantuvo ninguna justificación del juez por separado para este resultado archivado.
Nota de referencia histórica: esta puntuación general archivada v1.4 utilizó la agregación de diagnóstico fija anterior. Se conserva para fines de reproducibilidad y no debe interpretarse como una puntuación de Model Match v1.5.
Cronología de recuerdo puro
El conjunto de preguntas real se barajó antes de la solicitud del modelo probado. Los espacios de generación que faltan no son abstenciones modelo. El cronograma reconstruye los meses objetivo para la auditabilidad. Sólo las respuestas RECALL explícitas son evidencia de corte; INFERENCIA, ADIVINAR y DESCONOCIDO permanecen visibles pero no mueven el límite.
Comparación de límites de conocimiento
- Recall is broadly strong from September 2024 through April 2025, with only one isolated February abstention.
- May is mixed across its dense probes, while June and July retain partial or strong knowledge before a sustained failure cluster in August and September.
- The observed transition range overlaps and immediately follows the official May 2025 reference cutoff.
- Correct later-control answers in October and November are isolated; the COP30 host was also knowable well before the event, so these do not establish a later sustained boundary.
Semántico: correcto 16 · equivocado 2 · desconocido 6. Excluido del límite: 0.
Señales de identidad
La identidad se infiere de la coherencia entre el modelo reivindicado y el perfil de conocimiento/respuesta observado. Es una evidencia, no una certificación remota o una prueba criptográfica de los pesos de las porciones.
Señales positivas
- Correct on nearly all probes from September 2024 through April 2025.
- Mixed near-boundary performance in May.
- Sustained weak performance across both August and September dense evidence.
- Self-identifies as an Anthropic Claude model.
Señales negativas
- Missed the February 2025 Album of the Year despite otherwise strong pre-cutoff recall.
- Answered several June and July events correctly after the nominal cutoff.
- Correctly guessed the October 2025 Nobel Peace Prize recipient.
- Did not identify its exact version and supplied no cutoff self-report.
- The behavioral boundary broadly aligns with the immutable May 2025 reference cutoff.
- Strong-recall performance is consistently high across most months before the reference cutoff.
- The August-September failure cluster supplies a meaningful temporal transition rather than relying on one obscure missed fact.
- Limited post-cutoff knowledge is not treated as proof of substitution and is compatible with inference, advance-publicized facts, or selective updates.
- The self-report supports the provider family but does not independently establish the exact model variant.
Texto del modelo extraído palabra por palabra
Esta ejecución es anterior a la retención voluntaria del informe completo. El desafío, la clave de respuestas y la evaluación del juez normalizada están disponibles, pero el texto exacto de la respuesta del modelo probado fue descartado intencionalmente por esa política de privacidad de la era de los puntos de referencia.
Evidencia pregunta por pregunta
Todas las preguntas admitidas se muestran con la respuesta esperada, la fuente disponible, la base del modelo probado, el veredicto semántico y el límite de elegibilidad. Las preguntas sin calificación no son abstenciones modelo. Una CONDICIÓN o INFERENCIA correcta sigue siendo semánticamente correcta, pero se excluye de la evidencia de corte de recuerdo puro.
What was the name of Firefly Aerospace's lunar lander that successfully touched down on the Moon?
nota del juez: Matches Blue Ghost.
What was the name of Blue Origin's rocket that reached orbit on its maiden flight?
nota del juez: Matches New Glenn.
What was the designation of the Blue Origin mission that carried an all-women crew including Katy Perry and Gayle King?
nota del juez: Matches NS-31.
Who won the men's 100 metres at the 2025 World Athletics Championships in Tokyo?
nota del juez: Explicit abstention.
Which Venezuelan opposition leader was awarded the 2025 Nobel Peace Prize?
nota del juez: Matches Maria Corina Machado.
At which Alaska military installation did Donald Trump meet Vladimir Putin for their 2025 summit?
nota del juez: Explicit abstention.
What landmark global health instrument did the World Health Assembly adopt?
nota del juez: Explicitly includes WHO Pandemic Agreement.
Which Brazilian city hosted the COP30 climate conference?
nota del juez: Matches Belem.
Which new Nintendo console launched alongside Mario Kart World?
nota del juez: Matches Nintendo Switch 2.
Which album won Album of the Year at the 67th Grammy Awards?
nota del juez: Explicit abstention.
Which performer won the 2025 Eurovision Song Contest for Austria with Wasted Love?
nota del juez: Explicit abstention.
Which numbered iPhone generation did Apple debut at its September product event?
nota del juez: Answered iPhone 16; expected iPhone 17.
Which city hosted the COP29 climate conference that concluded with a new climate-finance agreement?
nota del juez: Matches Baku.
Which Japanese organization was awarded the 2024 Nobel Peace Prize?
nota del juez: Matches Nihon Hidankyo.
Which national team defeated Spain on penalties to win UEFA Women's Euro 2025?
nota del juez: Explicit abstention.
Which pope died at the Vatican on Easter Monday 2025?
nota del juez: Matches Pope Francis.
Which film won Best Picture at the 97th Academy Awards?
nota del juez: Matches Anora.
Which political alliance became the largest group in Germany's 2025 Bundestag election?
nota del juez: Matches CDU/CSU.
Which network of Indian forts became the country's 44th UNESCO World Heritage property?
nota del juez: Matches Maratha Military Landscapes of India.
Which mission conducted the first commercial spacewalk?
nota del juez: Matches Polaris Dawn.
Which private astronaut mission launched Peggy Whitson and Shubhanshu Shukla to the International Space Station?
nota del juez: Matches Axiom Mission 4.
Which Paris cathedral formally reopened five years after a devastating fire?
nota del juez: Matches Notre-Dame de Paris.
Who won Chile's 2025 presidential runoff election?
nota del juez: Explicit abstention.
What new flagship AI system did OpenAI introduce?
nota del juez: Answered GPT-4o; expected GPT-5.
Cobertura de preguntas y notas de investigación.
Preguntas fácticas: 24 · No calificado por el juez: 0
El corte de referencia y las citas son proporcionadas por el juez, no confirmadas de forma independiente por la plataforma. Las notas de investigación no penalizan el modelo probado.
Detalles de reproducibilidad
68TGGQW6SM9411Y4XDAPH3BR3Hmodel-verification-single-request-v1.3gpt-5.6-solopenai_chatLegacy / not recordedPARTIDO / legadoNot recordedNot recordedNot recordedNot recorded20010,087 ms2,59137edfee0969ea8c1529cf6db6830118353b483a27995cedcf20b25e5b40b8ff1The reference cutoff is a judge-produced snapshot grounded in hosted web evidence when available (official sources preferred, otherwise the best defensible estimate). It is not provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.
Ejecute una verificación independiente
Pruebe usted mismo el mismo punto final o explore otros informes públicos antes de confiar en una identidad de modelo anunciada.