Code and software agents
Veröffentlichte Coding-Werte: SWE-bench Verified und SWE-bench Multilingual für agentische Repository-Arbeit, LiveCodeBench für wettbewerbsnahe Codegenerierung und Terminal Bench für Kommandozeilen-Agenten.
Every score in this set is republished from the model publisher’s own model card or vendor report and names that source per entry. LLM EU has run no evaluation, has not reproduced any score, and has not verified the harness settings. Vendor-reported numbers are self-reported and are labelled as such on the model cards. SWE-bench results depend heavily on the agent scaffold; the NVIDIA entry names its scaffold explicitly, and scores from different scaffolds must not be compared directly.
| Modelle | score | source |
|---|---|---|
| Devstral Small 2 (24B) | 68 | link |
| GLM-4.7-Flash | 59.2 | link |
| NVIDIA Nemotron 3 Super 120B-A12B | 60.47 | link |
| Gemma 4 31B IT | 80 | link |
| Magistral Small 1.2 | 70.88 | link |
| GLM-5.3 | 88.8 | link |