deepseek-v4-flash
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/tr/model-verification/results/0B0HVDB2FRJEJC90BRG4D7JV1YBu doğrulamanın sonucu neden alındı?
Herkese açık rapor, /chat/verify sonrasında test uzmanına gösterilen analitik kanıtları yansıtır. Hakem puanları değerlendirme tablosu değerlendirmeleridir; yalnızca açık doğru veya yanlış RECALL yanıtları saf hatırlama kesme sınırını etkileyebilir.
Hakem puanlama mantığı
Arşivlenen bu sonuç için ayrı bir yargıç gerekçesi korunmadı.
Arşivlenen bu sonuç için ayrı bir yargıç gerekçesi korunmadı.
Arşivlenen bu sonuç için ayrı bir yargıç gerekçesi korunmadı.
Arşivlenen bu sonuç için ayrı bir yargıç gerekçesi korunmadı.
Tarihsel karşılaştırma notu: Bu arşivlenmiş v1.4 genel puanı, önceki sabit tanı toplamayı kullanmıştır. Tekrarlanabilirlik amacıyla saklanır ve v1.5 Model Eşleştirme puanı olarak yorumlanmamalıdır.
Saf hatırlama zaman çizelgesi
Gerçek soru seti, test edilen model talebinden önce karıştırıldı. Eksik nesil slotları model çekimserliği değildir. Zaman çizelgesi, denetlenebilirlik açısından hedef ayları yeniden yapılandırır. Yalnızca açık RECALL yanıtları kesme kanıtıdır; ÇIKARMA, TAHMİN ve BİLİNMEYEN görünür kalır ancak sınırı değiştirmez.
Bilgi sınırı karşılaştırması
- Strong recall is nearly complete through the May 2025 reference cutoff.
- There is no systematic strong-recall failure substantially before the reference cutoff.
- Correct June and August 2025 answers suggest selective later knowledge, while failures in July, September, and October prevent a clearly later boundary.
anlamsal: doğru 20 · yanlış 4 · bilinmiyor 0. Kesintiden hariç tutuldu: 0.
Kimlik sinyalleri
Kimlik, iddia edilen model ile gözlemlenen bilgi/cevap profili arasındaki tutarlılıktan çıkarılır. Bu, servis ağırlıklarının uzaktan tasdiki veya kriptografik kanıtı değil, kanıtıdır.
Olumlu sinyaller
- Excellent recall across late 2023 through May 2025.
- Correct near-cutoff answer for the WHO Pandemic Agreement.
- No pattern of early knowledge degradation.
Negatif sinyaller
- Incorrect February 2024 lunar-lander answer despite high confidence.
- Correct June and August 2025 answers extend beyond the estimated reference cutoff.
- Several post-cutoff answers are confidently wrong, indicating uneven recency.
- Observed behavior broadly overlaps the saved May 2025 reference cutoff.
- The model answered 20 of 24 Class A items correctly, with only one isolated failure well before the cutoff.
- Post-cutoff knowledge is uneven and can plausibly reflect selective updates rather than model substitution.
Verbatim model metnini çıkardı
Bu çalıştırma, tam raporun saklanmasını etkinleştirmeden önce gerçekleşir. Soru, cevap anahtarı ve normalleştirilmiş jüri değerlendirmesi mevcuttur, ancak tam olarak test edilmiş model yanıt metni, o kıyaslama dönemi gizlilik politikası tarafından kasıtlı olarak atılmıştır.
Soru bazında kanıt
Kabul edilen tüm sorular beklenen yanıt, mevcut kaynak, test edilmiş model esası, anlamsal karar ve kesme uygunluğuyla birlikte gösterilir. Not verilmeyen sorular örnek çekimserlik değildir. Doğru bir TAHMİN veya ÇIKARIM anlamsal olarak doğru kalır ancak saf hatırlama kesme kanıtlarının dışında bırakılır.
Who won the men's 100 metres at the World Athletics Championships in Tokyo?
Hakim notu: Expected Oblique Seville; answered Kishane Thompson.
Which team won the UEFA Women's Euro 2025 final?
Hakim notu: Expected England; answered Spain.
Which disease's African upsurge did WHO declare a public health emergency of international concern?
Hakim notu: Matches mpox.
What molecular-structure prediction model did Google DeepMind and Isomorphic Labs introduce?
Hakim notu: Matches AlphaFold 3.
What model did OpenAI release as a unified system with built-in thinking?
Hakim notu: Matches GPT-5.
Which operating system was affected by the faulty CrowdStrike Falcon content update?
Hakim notu: Microsoft Windows is accepted.
What was the nickname of Intuitive Machines' Nova-C lander that touched down on the Moon?
Hakim notu: Expected Odysseus; Odie is not an accepted answer.
What was the name of Japan's lunar lander that successfully reached the Moon?
Hakim notu: Matches SLIM.
Who became the first private astronaut to perform a spacewalk?
Hakim notu: Matches Jared Isaacman.
Which country was appointed host of the 2034 FIFA World Cup?
Hakim notu: Matches Saudi Arabia.
In which lunar mare did Firefly Aerospace's Blue Ghost land?
Hakim notu: Matches Mare Crisium.
What model did DeepSeek begin serving through its deepseek-reasoner endpoint?
Hakim notu: Matches DeepSeek-R1.
What was the name of the second malaria vaccine prequalified by WHO?
Hakim notu: Matches R21/Matrix-M.
Which album won Album of the Year at the 67th Grammy Awards?
Hakim notu: Matches Cowboy Carter.
How many European Parliament members voted in favor of the Artificial Intelligence Act?
Hakim notu: Matches 523.
How many grams of lunar material did the Chang'e-6 mission collect?
Hakim notu: Matches 1,935.3 grams.
Which rocket launched NASA's Europa Clipper spacecraft?
Hakim notu: Matches Falcon Heavy.
What designation was given to the newly identified 33-solar-mass stellar black hole in the Milky Way?
Hakim notu: Matches Gaia BH3.
Who was awarded the 2025 Nobel Peace Prize?
Hakim notu: Expected Maria Corina Machado; answered an institution.
What annual climate-finance goal for developing countries did COP29 set for 2035?
Hakim notu: Matches $300 billion per year.
What was the name of the first human spaceflight to orbit over Earth's polar regions?
Hakim notu: Matches Fram2.
Under which article of the WHO Constitution was the Pandemic Agreement adopted?
Hakim notu: Matches Article 19.
Which new Mario Kart game launched alongside Nintendo Switch 2?
Hakim notu: Matches Mario Kart World.
What model did OpenAI introduce with a 128K context window at its first DevDay?
Hakim notu: Matches GPT-4 Turbo.
Soru kapsamı ve araştırma notları
Gerçek sorular: 24 · Hakim tarafından notlandırılmadı: 0
Referans kesintisi ve alıntılar hakim tarafından sağlanır, platform tarafından bağımsız olarak onaylanmaz. Araştırma notları test edilen modeli cezalandırmaz.
Tekrarlanabilirlik ayrıntıları
0B0HVDB2FRJEJC90BRG4D7JV1Ymodel-verification-single-request-v1.2gpt-5.6-solopenai_chatLegacy / not recordedMAÇ / eskiNot recordedNot recordedNot recordedNot recorded20088,304 ms2,460f22e63d94e67fbd4fdfa25a27a37ae4b3402f94fc6801d14c2a4e638f85cfd9dThe reference cutoff is a judge-produced snapshot grounded in hosted web evidence when available (official sources preferred, otherwise the best defensible estimate). It is not provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.
Bağımsız bir doğrulama çalıştırın
Reklamı yapılan bir model kimliğine güvenmeden önce aynı uç noktayı kendiniz test edin veya diğer genel raporlara göz atın.