Long context
Wie sich Modelle verhalten, wenn die Eingabe deutlich länger ist als ein Gespräch. Die RULER-Zeilen messen Retrieval und Reasoning in einer langen Eingabe bei angegebener Länge; die übrigen Zeilen nennen das maximale Kontextfenster laut Anbieter, also eine Kapazitätsangabe und keinen Qualitätswert.
Every score in this set is republished from the model publisher’s own model card or vendor report and names that source per entry. LLM EU has run no evaluation, has not reproduced any score, and has not verified the harness settings. Vendor-reported numbers are self-reported and are labelled as such on the model cards. RULER scores come from NVIDIA’s own NeMo Evaluator harness and are not comparable with scores produced by other harnesses. The context-window rows are configuration or vendor claims and say nothing about answer quality at that length.
| Modelle | score | source |
|---|---|---|
| NVIDIA Nemotron 3 Super 120B-A12B | 96.3 | link |
| NVIDIA Nemotron 3 Super 120B-A12B | 91.75 | link |
| NVIDIA Nemotron 3 Nano 4B | 91.1 | link |
| Llama 4 Scout 17B-16E Instruct | 10485760 | link |
| GLM-5.3 | 1048576 | link |
| DeepSeek-V4-Flash (0731) | 1048576 | link |