deepseek-v4-flash
https://api.model-gate.com/v1/chat/completions
https://model-gate.com/hu/model-verification/results/A57ZK57ERM62N3S1XETMXFKK0XMiért lett ez az ellenőrzés eredménye?
A nyilvános jelentés tükrözi a tesztelőnek a /chat/verify után megjelenő analitikai bizonyítékokat. A bírói pontszámok rubrikaértékelések; csak az explicit helyes vagy rossz RECALL válaszok befolyásolhatják a tiszta visszahívás határát.
A bírói pontozás indoklása
Ehhez az archivált eredményhez nem tartottak fenn külön bírói indoklást.
Ehhez az archivált eredményhez nem tartottak fenn külön bírói indoklást.
Ehhez az archivált eredményhez nem tartottak fenn külön bírói indoklást.
Ehhez az archivált eredményhez nem tartottak fenn külön bírói indoklást.
Történelmi referenciaérték: ez az archivált 1.4-es verzió általános pontszáma a korábbi rögzített diagnosztikai aggregációt használta. A reprodukálhatóság érdekében megőrizzük, és nem értelmezhető v1.5-ös modellegyezési pontszámként.
Tiszta visszahívási idővonal
A tényleges kérdéskészletet megkeverték a tesztelt modell kérése előtt. A hiányzó generációs helyek nem modelltartózkodások. Az idővonal rekonstruálja az auditálhatóság célhónapjait. Csak az explicit RECALL válaszok jelentik a végső bizonyítékot; A KÖVETKEZTETÉS, KITALÁLÁS és az ISMERETLEN látható marad, de nem mozdítja el a határt.
Tudáshatárok összehasonlítása
- All Class A items are correct through June 2024, establishing an observed lower bound but no upper bound.
- The immutable reference snapshot has no known cutoff dates, so cutoff alignment cannot be scored.
Szemantikus: helyes 16 · rossz 0 · ismeretlen 0. Lezárásból kizárva: 0.
Azonosító jelek
Az identitásra az állított modell és a megfigyelt tudás/válasz profil közötti összhangból lehet következtetni. Ez bizonyíték, nem távoli tanúsítvány vagy kriptográfiai bizonyíték az adagsúlyokról.
Pozitív jelek
- Perfect performance across all supplied Class A recall items.
- Correct recall of multiple DeepSeek model releases and specifications.
- No systematic early knowledge failures.
Negatív jelek
- No documented canonical identity or cutoff is available in the reference snapshot.
- Cutoff consistency cannot be evaluated.
- Self-report and matching API metadata are weak identity evidence.
- The response is fully correct on all 16 strong-recall questions, including the DeepSeek-specific history items.
- The observed knowledge extends through June 2024, but the reference cutoff is unknown and therefore provides no identity-comparison signal.
- The claimed name and API identifier agree, but both are untrusted metadata and do not prove model identity.
Szó szerint kivont modellszöveg
Ez a futtatás a teljes jelentés megtartását megelőzően történt. A kihívás, a válasz kulcsa és a normalizált bírói értékelés elérhető, de a pontos tesztelt modell válaszszövegét szándékosan elvetette az a benchmark korszak adatvédelmi szabályzata.
Kérdésről-kérdésre bizonyíték
Minden elfogadott kérdés megjelenik a várt válasszal, a rendelkezésre álló forrással, a tesztelt modell alapján, a szemantikai ítélettel és a határértékre való alkalmassággal. Az osztályozatlan kérdések nem modelltartózkodások. A helyes TALÁLÁS vagy KÖVETKEZTETÉS szemantikailag helyes marad, de ki van zárva a tisztán felidézés határértékeiből.
Which human Go player was AlphaGo's opponent in the three-game match completed on 2017-05-27?
Bíró megjegyzés: Matches accepted answer Ke Jie.
The first published image of a black hole, announced on 2019-04-10, depicted the black hole in which galaxy?
Bíró megjegyzés: M87 galaxy is semantically equivalent to Messier 87.
Which two scientists were awarded the 2020 Nobel Prize in Chemistry for developing a method for genome editing?
Bíró megjegyzés: Both required scientists are correctly named.
Which research organization developed the AlphaFold2 system whose CASP14 performance was announced on 2020-11-30?
Bíró megjegyzés: Matches accepted answer DeepMind.
Which launch vehicle carried the James Webb Space Telescope into space on 2021-12-25?
Bíró megjegyzés: Matches accepted answer Ariane 5.
What was the name of the asteroid moonlet struck by NASA's DART spacecraft on 2022-09-26?
Bíró megjegyzés: Matches accepted answer Dimorphos.
In which ocean did the Artemis I Orion capsule splash down on 2022-12-11?
Bíró megjegyzés: Matches accepted answer Pacific Ocean.
On 2023-03-01, OpenAI made an API available for which speech-recognition model?
Bíró megjegyzés: Matches accepted answer Whisper.
What was the name of the Chandrayaan-3 lander that reached the lunar surface on 2023-08-23?
Bíró megjegyzés: Matches accepted answer Vikram.
Samples from which asteroid were returned to Earth by the OSIRIS-REx capsule on 2023-09-24?
Bíró megjegyzés: Matches accepted answer Bennu.
What was the name of the code-focused model family publicly introduced by DeepSeek on 2023-11-02?
Bíró megjegyzés: DeepSeek-Coder is an accepted spelling.
What two parameter scales were offered in the DeepSeek LLM family announced on 2023-11-29?
Bíró megjegyzés: Both parameter scales match: 7B and 67B.
How many total parameters and how many activated parameters per token did DeepSeek-V2 have when announced on 2024-05-06?
Bíró megjegyzés: Correctly gives 236B total and 21B activated per token.
What maximum context length was supported by DeepSeek-Coder-V2 when announced on 2024-06-17?
Bíró megjegyzés: Matches accepted answer 128K tokens.
Which company built the Odysseus lunar lander that touched down on 2024-02-22?
Bíró megjegyzés: Matches accepted answer Intuitive Machines.
What two parameter sizes were offered for the initially released Llama 3 models announced on 2024-04-18?
Bíró megjegyzés: Both initially released sizes match: 8B and 70B.
Kérdések lefedettsége és kutatási jegyzetek
Tény kérdések: 16 · Nem minősítette a bíró: 0
A referencia határértéket és az idézeteket a bíró biztosítja, a platform független módon nem erősíti meg. A kutatási jegyzetek nem büntetik a tesztelt modellt.
A reprodukálhatóság részletei
A57ZK57ERM62N3S1XETMXFKK0Xmodel-verification-single-request-v1.1gpt-5.6-solopenai_chatLegacy / not recordedMATCH / örökségNot recordedNot recordedNot recordedNot recorded20028,645 ms1,8594c1f5cca4a79a53f28e7f6544bfe08208148d53c85569f50c4132e161da06613The reference cutoff is the configured judge model's internal-knowledge snapshot for this run, not an independently verified provider attestation. Single-request behavioral screening is not cryptographic proof of the underlying model identity.
Futtasson le egy független ellenőrzést
Tesztelje saját maga ugyanazt a végpontot, vagy böngésszen más nyilvános jelentésekben, mielőtt a meghirdetett modellazonosságra hagyatkozna.