European languages

Model quality on translated European-language benchmarks. The EU21 rows come from the OpenGPT-X evaluation published with Teuken 7B Instruct v0.6, which translates ARC, HellaSwag, TruthfulQA and MMLU into the 21 EU languages it covers and reports one average per model. The MMMLU row is Google’s own multilingual MMLU figure for Gemma 4.

The EU21 numbers are OpenGPT-X’s, taken from the Teuken 7B Instruct v0.6 model card and the accompanying evaluation preprint; they are machine-translated benchmarks, not native-language test sets. The MMMLU number is Google’s own. LLM EU has run no evaluation and has not verified either harness.