Code and software agents
Published coding scores: SWE-bench Verified and SWE-bench Multilingual for agentic repository work, LiveCodeBench for competitive-style code generation, and Terminal Bench for command-line agents.
Every score in this set is republished from the model publisher’s own model card or vendor report and names that source per entry. LLM EU has run no evaluation, has not reproduced any score, and has not verified the harness settings. Vendor-reported numbers are self-reported and are labelled as such on the model cards. SWE-bench results depend heavily on the agent scaffold; the NVIDIA entry names its scaffold explicitly, and scores from different scaffolds must not be compared directly.
| Models | score | source |
|---|---|---|
| Devstral Small 2 (24B) | 68 | link |
| GLM-4.7-Flash | 59.2 | link |
| NVIDIA Nemotron 3 Super 120B-A12B | 60.47 | link |
| Gemma 4 31B IT | 80 | link |
| Magistral Small 1.2 | 70.88 | link |
| GLM-5.3 | 88.8 | link |