A scientifically rigorous evaluation surface mapping the performance and agentic capabilities of frontier models across non-English tasks, socio-cultural contexts, and multimodal domains.
Standard benchmarks mask frontier models' performance regressions when used in any language but English. LILT Applied AI stress-tests LLMs under realistic, localized constraints to map the true boundaries of global AI capability.
Task pass rate per model, across all 10 languages (5 trials per task).
Mean pass rates over 324 tasks (5 trials per task). Whiskers show 95% Wilson score intervals.
The same pass rates plotted against each model's vendor release date, to show whether a newer model is actually a better one.
Each dot is one model at its vendor release date; whiskers show 95% Wilson score intervals. Left to right is time, so a dot that sits high and early led the field when it shipped.
30 tasks written in Arabic, scored across all 12 models.
31 tasks written in Chinese, scored across all 12 models.
34 tasks written in Czech, scored across all 12 models.
33 tasks written in German, scored across all 12 models.
33 tasks written in Hindi, scored across all 12 models.
31 tasks written in Japanese, scored across all 12 models.
32 tasks written in Korean, scored across all 12 models.
34 tasks written in Serbian, scored across all 12 models.
32 tasks written in Spanish, scored across all 12 models.
34 tasks written in Turkish, scored across all 12 models.
*In Terminal-Bench, the GLM 5.3 Flash scores are on a quantized model as this is pinned by the service provider and cannot be changed
Agent reliability on the airline domain, in English, German, Korean, Turkish and Persian.
Agent reliability on the retail domain, in English, German, Korean, Turkish and Persian.
Agent reliability on the telecom domain, in English, German, Korean, Turkish and Persian.
Agent reliability on the banking domain (a stratified 20-task sample), in English, German, Korean, Turkish and Persian.
Agent reliability across all four domains — the unweighted mean of the airline, retail, telecom and banking scores — in English, German, Korean, Turkish and Persian.
Agent reliability on the airline domain, in English, German, Korean, Turkish and Persian.
Agent reliability on the retail domain, in English, German, Korean, Turkish and Persian.
Agent reliability on the telecom domain, in English, German, Korean, Turkish and Persian.
Agent reliability on the banking domain (a stratified 20-task sample), in English, German, Korean, Turkish and Persian.
Agent reliability across all four domains — the unweighted mean of the airline, retail, telecom and banking scores — in English, German, Korean, Turkish and Persian.
Agent reliability on the airline domain, in English, German, Korean, Turkish and Persian.
Agent reliability on the retail domain, in English, German, Korean, Turkish and Persian.
Agent reliability on the telecom domain, in English, German, Korean, Turkish and Persian.
Agent reliability on the banking domain (a stratified 20-task sample), in English, German, Korean, Turkish and Persian.
Agent reliability across all four domains — the unweighted mean of the airline, retail, telecom and banking scores — in English, German, Korean, Turkish and Persian.
Agent reliability on the airline domain, in English, German, Korean, Turkish and Persian.
Agent reliability on the retail domain, in English, German, Korean, Turkish and Persian.
Agent reliability on the telecom domain, in English, German, Korean, Turkish and Persian.
Agent reliability on the banking domain (a stratified 20-task sample), in English, German, Korean, Turkish and Persian.
Agent reliability across all four domains — the unweighted mean of the airline, retail, telecom and banking scores — in English, German, Korean, Turkish and Persian.
gemini-3.5-flash, temperature 1.0, with the system prompt tuned for realistic behaviorgpt-5.6-sol, claude-opus-5, CohereLabs/command-a-plus-05-2026, gemini-3.8-flashnone, Opus 5 off (banking: low), Command A+ has no off switch, Gemini 3.8 Flash low (cannot be disabled). Temperature: provider default.gpt-4.1-2025-04-14 (τ³-bench default)base split; banking: stratified 20 task sample, same task ids in every locale and for every agentMulti-turn conversational accuracy, grouped by model: the bars inside each model are the four languages.
The same numbers grouped by language: the bars inside each language are the eight models.
Some languages take more tokens than English, and are thus more expensive to run. For example, the same dialogue costs about 1.4× as much in Korean or Arabic as in English.
The same figures as absolute average tokens per dialogue.
**Opus refused to answer on 6 out of 273 questions. These are counted as failures
You cannot measure multilingual performance by machine-translating an English benchmark. Doing so introduces a non-native text distribution, does not result in culturally relevant tasks, and may introduce factual errors.
We compare two multilingual versions of GAIA: the MAPS-GAIA machine translations of the original English tasks, and our carefully audited and culturally adapted tasks.
Models score an average of 20.7 points higher on the audited, culturally adapted benchmark, reflecting measurement error introduced by translation in the original multilingual benchmark. That difference is the upper, lighter segment of every bar below.
Pass rate on the MAPS-GAIA machine-translated benchmark compared to LILT's audited GAIA-v2-LILT, one panel per model.
165 query–answer pairs per language · the dashed line marks each model's English score
Table 2 · Source: arXiv:2604.24929
The same comparison with one panel per language, so the three models sit side by side within each.
165 query–answer pairs per language · the dashed line marks each model's English score
Table 2 · Source: arXiv:2604.24929
The audit lift on its own, zero-baselined: audited minus machine-translated, so its size is comparable language to language.
165 query–answer pairs per language · these bars show measurement error (audited − machine-translated), not model improvement
Table 2 · Source: arXiv:2604.24929
The same audit lift with one panel per language, so the three models sit side by side within each.
165 query–answer pairs per language · these bars show measurement error (audited − machine-translated), not model improvement
Table 2 · Source: arXiv:2604.24929
How many of the machine-translated tasks the audit changed at all.
Table 1 · Source: arXiv:2604.24929
How far those revisions went: word- and character-level edit distance from the machine translation.
Table 1 · Source: arXiv:2604.24929
pass rate (equivalently pass@1 or pass^1) is the chance a model will successfully pass a single task. We estimate it by having each model complete each task multiple times.
pass@k is the chance that a model will pass a task at least once out of k attempts. Intuitively, this is best used for difficult tasks that could reasonably be solved by running the model multiple times and selecting the best answer. It is estimated using the unbiased estimator (see Appendix A of Chen et al., 2021).
pass^k is the chance that a model will pass a task reliably in all k attempts. This metric is best used for tasks that can't be rerun (like interactions with a customer, as in τ³-bench) or where failure is expensive or dangerous. It is estimated using the unbiased estimator introduced in the first τ-bench paper (Yao et al., 2024).
LILTBench is an open collection of 92 coding tasks across 31 languages, authored by community teams to break frontier models in their own language. Tasks ship with the original-language instructions and an English translation, so the same problem can be run either way to see how much prompt language alone moves performance.
The tasks, the harness, and the leaderboard are public.