European Commission - Directorate General for Energy

07/22/2026 | Press release | Distributed by Public on 07/22/2026 01:44

Towards fair multilingual AI: EU MMLU, a new EU benchmark for LLMs

DG Translation has released the EU MMLU, a high-quality multilingual benchmarking dataset designed to assess whether large language models perform fairly and effectively across the EU's linguistic diversity. It builds on one of the most widely used AI evaluation datasets, Massive Multitask Language Understanding (MMLU).

Large language models (LLMs) are typically evaluated using benchmarking datasets that contain thousands of multiple-choice questions covering different areas of human knowledge. The LLM receives a score based on its answers to these questions.

Most of the datasets used to evaluate AI were built in English and do not specifically reflect European educational, cultural and societal contexts. They often fail to reveal how well an LLM performs in other languages - a model may score well in English while underperforming significantly in French, Hungarian or Maltese. The EU MMLU addresses this gap.

A human-centred approach

Unlike most existing multilingual benchmarks, which rely primarily on machine translation, the EU MMLU dataset takes a human-centred approach. DG Translation has joined forces with student translators and project managers from the European Master's in Translation (EMT) network to translate and revise 1000+ benchmark questions. Nearly 250 students from 21 universities across Europe have contributed to the project.

The EU MMLU dataset is currently available in 16 EU official languages: Croatian, Czech, Dutch, French, German, Greek, Hungarian, Irish, Italian, Lithuanian, Polish, Portuguese, Romanian, Slovak and Slovenian. Further languages will follow.

Why do we need an EU-oriented benchmarking dataset?

The dataset on which the EU MMLU is based contains thousands of benchmark questions on 57 subjects ranging from science and law to ethics and public affairs. The EU MMLU focuses on 7 of these subject areas, selected for their relevance to the EU. The objective is not simply to translate existing questions, but to make sure they retain the same meaning, difficulty and testing value across languages.

Our long-term ambition is to establish a new standard for multilingual AI evaluation in Europe - so that AI systems perform consistently across languages and cultures, rather than primarily in English-speaking environments.

A guide for better multilingual benchmarks

DG Translation has also published a list of core quality criteria for EU-oriented multilingual benchmarking of LLMs.

The list of quality criteria goes beyond translation. It recommends including items in all 24 EU official languages, with balanced representation across languages.

In addition, an EU-ready benchmarking dataset should test how well an AI model reflects EU values and handles EU-specific cultural contexts - including idioms, humour, cultural references, date/number formats and differences in expected tone or politeness.

Together, the EU MMLU dataset and the list of core quality criteria are designed to improve how today's language models are evaluated - and to help shape the next generation of multilingual AI.

Find out more about DG Translation's work on an EU Institutional LLM and the EU MMLU

European Commission - Directorate General for Energy published this content on July 22, 2026, and is solely responsible for the information contained herein. Distributed via Public Technologies (PUBT), unedited and unaltered, on July 22, 2026 at 07:44 UTC. If you believe the information included in the content is inaccurate or outdated and requires editing or removal, please contact us at [email protected]