claude-opus-4-6
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/ca/model-verification/results/68TGGQW6SM9411Y4XDAPH3BR3HPer què aquesta verificació va rebre el seu resultat
L'informe públic reflecteix l'evidència analítica mostrada al verificador després de /chat/verify. Les puntuacions dels jutges són avaluacions de rúbrica; només les respostes RECALL explícites correctes o incorrectes poden influir en el límit de tall de record pur.
Justificació de la puntuació del jutge
No es va retenir cap justificació de jutge independent per a aquest resultat arxivat.
No es va retenir cap justificació de jutge independent per a aquest resultat arxivat.
No es va retenir cap justificació de jutge independent per a aquest resultat arxivat.
No es va retenir cap justificació de jutge independent per a aquest resultat arxivat.
Nota de referència històrica: aquesta puntuació global de la v1.4 arxivada utilitzava l'anterior agregació de diagnòstic fix. Es conserva per a la seva reproductibilitat i no s'ha d'interpretar com una puntuació de concordança del model v1.5.
Línia de temps de record pur
El conjunt de preguntes real es va barrejar abans de la sol·licitud del model provat. Els espais de generació que falten no són abstencions de model. La línia de temps reconstrueix els mesos objectiu per a la seva auditabilitat. Només les respostes RECALL explícites són proves de tall; INFERÈNCIA, ENDEVINA i DESCONEGUT romanen visibles però no mouen el límit.
Comparació de talls de coneixement
- Recall is broadly strong from September 2024 through April 2025, with only one isolated February abstention.
- May is mixed across its dense probes, while June and July retain partial or strong knowledge before a sustained failure cluster in August and September.
- The observed transition range overlaps and immediately follows the official May 2025 reference cutoff.
- Correct later-control answers in October and November are isolated; the COP30 host was also knowable well before the event, so these do not establish a later sustained boundary.
Semàntica: correcte 16 · malament 2 · desconegut 6. Exclòs del tall: 0.
Senyals d'identitat
La identitat es dedueix de la coherència entre el model reivindicat i el perfil de coneixement/resposta observat. És una prova, no una certificació remota o una prova criptogràfica dels pesos de la porció.
Senyals positius
- Correct on nearly all probes from September 2024 through April 2025.
- Mixed near-boundary performance in May.
- Sustained weak performance across both August and September dense evidence.
- Self-identifies as an Anthropic Claude model.
Senyals negatius
- Missed the February 2025 Album of the Year despite otherwise strong pre-cutoff recall.
- Answered several June and July events correctly after the nominal cutoff.
- Correctly guessed the October 2025 Nobel Peace Prize recipient.
- Did not identify its exact version and supplied no cutoff self-report.
- The behavioral boundary broadly aligns with the immutable May 2025 reference cutoff.
- Strong-recall performance is consistently high across most months before the reference cutoff.
- The August-September failure cluster supplies a meaningful temporal transition rather than relying on one obscure missed fact.
- Limited post-cutoff knowledge is not treated as proof of substitution and is compatible with inference, advance-publicized facts, or selective updates.
- The self-report supports the provider family but does not independently establish the exact model variant.
Text model extret textualment
Aquesta execució és anterior a la retenció d'informes complets activats. El repte, la clau de respostes i l'avaluació normalitzada del jutge estan disponibles, però el text de resposta exacte del model provat va ser descartat intencionadament per aquesta política de privadesa de l'era de referència.
Evidència pregunta per pregunta
Totes les preguntes admeses es mostren amb la resposta esperada, la font disponible, la base del model provat, el veredicte semàntic i l'elegibilitat de tall. Les preguntes no qualificades no són abstencions model. Una CONDECINA o INFERÈNCIA correcta segueix sent correcta semànticament, però s'exclou de l'evidència de tall de record pur.
What was the name of Firefly Aerospace's lunar lander that successfully touched down on the Moon?
Nota del jutge: Matches Blue Ghost.
What was the name of Blue Origin's rocket that reached orbit on its maiden flight?
Nota del jutge: Matches New Glenn.
What was the designation of the Blue Origin mission that carried an all-women crew including Katy Perry and Gayle King?
Nota del jutge: Matches NS-31.
Who won the men's 100 metres at the 2025 World Athletics Championships in Tokyo?
Nota del jutge: Explicit abstention.
Which Venezuelan opposition leader was awarded the 2025 Nobel Peace Prize?
Nota del jutge: Matches Maria Corina Machado.
At which Alaska military installation did Donald Trump meet Vladimir Putin for their 2025 summit?
Nota del jutge: Explicit abstention.
What landmark global health instrument did the World Health Assembly adopt?
Nota del jutge: Explicitly includes WHO Pandemic Agreement.
Which Brazilian city hosted the COP30 climate conference?
Nota del jutge: Matches Belem.
Which new Nintendo console launched alongside Mario Kart World?
Nota del jutge: Matches Nintendo Switch 2.
Which album won Album of the Year at the 67th Grammy Awards?
Nota del jutge: Explicit abstention.
Which performer won the 2025 Eurovision Song Contest for Austria with Wasted Love?
Nota del jutge: Explicit abstention.
Which numbered iPhone generation did Apple debut at its September product event?
Nota del jutge: Answered iPhone 16; expected iPhone 17.
Which city hosted the COP29 climate conference that concluded with a new climate-finance agreement?
Nota del jutge: Matches Baku.
Which Japanese organization was awarded the 2024 Nobel Peace Prize?
Nota del jutge: Matches Nihon Hidankyo.
Which national team defeated Spain on penalties to win UEFA Women's Euro 2025?
Nota del jutge: Explicit abstention.
Which pope died at the Vatican on Easter Monday 2025?
Nota del jutge: Matches Pope Francis.
Which film won Best Picture at the 97th Academy Awards?
Nota del jutge: Matches Anora.
Which political alliance became the largest group in Germany's 2025 Bundestag election?
Nota del jutge: Matches CDU/CSU.
Which network of Indian forts became the country's 44th UNESCO World Heritage property?
Nota del jutge: Matches Maratha Military Landscapes of India.
Which mission conducted the first commercial spacewalk?
Nota del jutge: Matches Polaris Dawn.
Which private astronaut mission launched Peggy Whitson and Shubhanshu Shukla to the International Space Station?
Nota del jutge: Matches Axiom Mission 4.
Which Paris cathedral formally reopened five years after a devastating fire?
Nota del jutge: Matches Notre-Dame de Paris.
Who won Chile's 2025 presidential runoff election?
Nota del jutge: Explicit abstention.
What new flagship AI system did OpenAI introduce?
Nota del jutge: Answered GPT-4o; expected GPT-5.
Cobertura de preguntes i notes de recerca
Preguntes de fet: 24 · No qualificat pel jutge: 0
El tall de referència i les citacions són proporcionades pel jutge, no confirmades de manera independent per la plataforma. Les notes de recerca no penalitzen el model provat.
Detalls de reproductibilitat
68TGGQW6SM9411Y4XDAPH3BR3Hmodel-verification-single-request-v1.3gpt-5.6-solopenai_chatLegacy / not recordedPARTIT / llegatNot recordedNot recordedNot recordedNot recorded20010,087 ms2,59137edfee0969ea8c1529cf6db6830118353b483a27995cedcf20b25e5b40b8ff1The reference cutoff is a judge-produced snapshot grounded in hosted web evidence when available (official sources preferred, otherwise the best defensible estimate). It is not provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.
Executeu una verificació independent
Proveu vosaltres mateixos el mateix punt final o navegueu per altres informes públics abans de confiar en una identitat de model anunciada.