Long context
How models behave when the input is much longer than a conversation. The RULER rows measure retrieval and reasoning inside a long input at a stated length; the remaining rows report the maximum context window the publisher claims, which is a capacity figure and not a quality score.
Every score in this set is republished from the model publisher’s own model card or vendor report and names that source per entry. LLM EU has run no evaluation, has not reproduced any score, and has not verified the harness settings. Vendor-reported numbers are self-reported and are labelled as such on the model cards. RULER scores come from NVIDIA’s own NeMo Evaluator harness and are not comparable with scores produced by other harnesses. The context-window rows are configuration or vendor claims and say nothing about answer quality at that length.
| Modele | score | source |
|---|---|---|
| NVIDIA Nemotron 3 Super 120B-A12B | 96.3 | link |
| NVIDIA Nemotron 3 Super 120B-A12B | 91.75 | link |
| NVIDIA Nemotron 3 Nano 4B | 91.1 | link |
| Llama 4 Scout 17B-16E Instruct | 10485760 | link |
| GLM-5.3 | 1048576 | link |
| DeepSeek-V4-Flash (0731) | 1048576 | link |