Multilingual Evaluation Surface

The Frontier of Multilingual LLM Performance

A scientifically rigorous evaluation surface mapping the performance and agentic capabilities of frontier models across non-English tasks, socio-cultural contexts, and multimodal domains.

Standard benchmarks mask frontier models' performance regressions when used in any language but English. LILT Applied AI stress-tests LLMs under realistic, localized constraints to map the true boundaries of global AI capability.

Public Benchmark Results

Multilingual Terminal-Bench ↗ agentic coding tasks

Ranked by pass rate*

Task pass rate per model, across all 10 languages (5 trials per task).

Claude Opus 5 (effort=high)73.7%Gemini 3.8 Flash67.3%GPT-5.6-Sol64.5%GPT-5.563.0%Muse Spark 1.362.2%Muse Spark 1.162.1%Gemini 3.5 Flash60.9%Claude Opus 4.8 (effort=high)60.2%Claude Opus 4.7 (effort=high)57.2%Gemini 3.1 Pro (preview)55.6%GLM 5.3 Flash51.5%GPT-5.431.9%

Mean pass rates over 324 tasks (5 trials per task). Whiskers show 95% Wilson score intervals.

By release date*

The same pass rates plotted against each model's vendor release date, to show whether a newer model is actually a better one.

Each dot is one model at its vendor release date; whiskers show 95% Wilson score intervals. Left to right is time, so a dot that sits high and early led the field when it shipped.

Multilingual Terminal-Bench by Language

Arabic*

30 tasks written in Arabic, scored across all 12 models.

Chinese*

31 tasks written in Chinese, scored across all 12 models.

Czech*

34 tasks written in Czech, scored across all 12 models.

German*

33 tasks written in German, scored across all 12 models.

Hindi*

33 tasks written in Hindi, scored across all 12 models.

Japanese*

31 tasks written in Japanese, scored across all 12 models.

Korean*

32 tasks written in Korean, scored across all 12 models.

Serbian*

34 tasks written in Serbian, scored across all 12 models.

Spanish*

32 tasks written in Spanish, scored across all 12 models.

Turkish*

34 tasks written in Turkish, scored across all 12 models.

How these models were run
Evaluation framework
Harbor v0.21.0
Agent harness
Terminus-2 (following the original Terminal-Bench)
Trials per task
5
Time limits
Per-task, adjusted based on task complexity by each task author, minimum 600s
Task instructions
Authored by native speakers in each language
Reasoning
Provider defaults, unless otherwise noted

*In Terminal-Bench, the GLM 5.3 Flash scores are on a quantized model as this is pinned by the service provider and cannot be changed

Multilingual τ³-bench multi-turn tool use in customer support scenarios

Airline · pass^4 · by locale

Agent reliability on the airline domain, in English, German, Korean, Turkish and Persian.

Retail · pass^4 · by locale

Agent reliability on the retail domain, in English, German, Korean, Turkish and Persian.

Telecom · pass^4 · by locale

Agent reliability on the telecom domain, in English, German, Korean, Turkish and Persian.

Banking · pass^4 · by locale

Agent reliability on the banking domain (a stratified 20-task sample), in English, German, Korean, Turkish and Persian.

Total · pass^4 · by locale

Agent reliability across all four domains — the unweighted mean of the airline, retail, telecom and banking scores — in English, German, Korean, Turkish and Persian.

Airline · pass^4 · by model

Agent reliability on the airline domain, in English, German, Korean, Turkish and Persian.

Retail · pass^4 · by model

Agent reliability on the retail domain, in English, German, Korean, Turkish and Persian.

Telecom · pass^4 · by model

Agent reliability on the telecom domain, in English, German, Korean, Turkish and Persian.

Banking · pass^4 · by model

Agent reliability on the banking domain (a stratified 20-task sample), in English, German, Korean, Turkish and Persian.

Total · pass^4 · by model

Agent reliability across all four domains — the unweighted mean of the airline, retail, telecom and banking scores — in English, German, Korean, Turkish and Persian.

Airline · pass^1 · by locale

Agent reliability on the airline domain, in English, German, Korean, Turkish and Persian.

Retail · pass^1 · by locale

Agent reliability on the retail domain, in English, German, Korean, Turkish and Persian.

Telecom · pass^1 · by locale

Agent reliability on the telecom domain, in English, German, Korean, Turkish and Persian.

Banking · pass^1 · by locale

Agent reliability on the banking domain (a stratified 20-task sample), in English, German, Korean, Turkish and Persian.

Total · pass^1 · by locale

Agent reliability across all four domains — the unweighted mean of the airline, retail, telecom and banking scores — in English, German, Korean, Turkish and Persian.

Airline · pass^1 · by model

Agent reliability on the airline domain, in English, German, Korean, Turkish and Persian.

Retail · pass^1 · by model

Agent reliability on the retail domain, in English, German, Korean, Turkish and Persian.

Telecom · pass^1 · by model

Agent reliability on the telecom domain, in English, German, Korean, Turkish and Persian.

Banking · pass^1 · by model

Agent reliability on the banking domain (a stratified 20-task sample), in English, German, Korean, Turkish and Persian.

Total · pass^1 · by model

Agent reliability across all four domains — the unweighted mean of the airline, retail, telecom and banking scores — in English, German, Korean, Turkish and Persian.

How these models were run
User simulator
gemini-3.5-flash, temperature 1.0, with the system prompt tuned for realistic behavior
Agent models
gpt-5.6-sol, claude-opus-5, CohereLabs/command-a-plus-05-2026, gemini-3.8-flash
Agent reasoning
Disabled where the model allows it, else lowest effort, to simulate low-latency customer support: Sol none, Opus 5 off (banking: low), Command A+ has no off switch, Gemini 3.8 Flash low (cannot be disabled). Temperature: provider default.
Serving
OpenRouter for all agents except Command A+, which ran on a local vLLM deployment
Assertion judge
gpt-4.1-2025-04-14 (τ³-bench default)
Task split
Airline, telecom, retail: base split; banking: stratified 20 task sample, same task ids in every locale and for every agent
Modifications
Thorough localization of user simulator prompts, task scenarios and database contents; minor edit to agent system prompt instructing it to use a specific language
Trials per task
4
Agent scaffold
official τ³-bench scaffold, modified to support multilinguality
Max tokens per turn
4096
MultiChallenge ↗ long-context instruction-following

By model

Multi-turn conversational accuracy, grouped by model: the bars inside each model are the four languages.

By language

The same numbers grouped by language: the bars inside each language are the eight models.

Tokens per dialogue, × English

Some languages take more tokens than English, and are thus more expensive to run. For example, the same dialogue costs about 1.4× as much in Korean or Arabic as in English.

Tokens per dialogue, raw counts

The same figures as absolute average tokens per dialogue.

**Opus refused to answer on 6 out of 273 questions. These are counted as failures

GAIA-v2 · LILT ↗ agentic reasoning and tool use

You cannot measure multilingual performance by machine-translating an English benchmark. Doing so introduces a non-native text distribution, does not result in culturally relevant tasks, and may introduce factual errors.

We compare two multilingual versions of GAIA: the MAPS-GAIA machine translations of the original English tasks, and our carefully audited and culturally adapted tasks.

Models score an average of 20.7 points higher on the audited, culturally adapted benchmark, reflecting measurement error introduced by translation in the original multilingual benchmark. That difference is the upper, lighter segment of every bar below.

Both benchmarks

By model

Pass rate on the MAPS-GAIA machine-translated benchmark compared to LILT's audited GAIA-v2-LILT, one panel per model.

165 query–answer pairs per language · the dashed line marks each model's English score

Table 2 · Source: arXiv:2604.24929

By language

The same comparison with one panel per language, so the three models sit side by side within each.

165 query–answer pairs per language · the dashed line marks each model's English score

Table 2 · Source: arXiv:2604.24929

Measurement error

By model

The audit lift on its own, zero-baselined: audited minus machine-translated, so its size is comparable language to language.

165 query–answer pairs per language · these bars show measurement error (audited − machine-translated), not model improvement

Table 2 · Source: arXiv:2604.24929

By language

The same audit lift with one panel per language, so the three models sit side by side within each.

165 query–answer pairs per language · these bars show measurement error (audited − machine-translated), not model improvement

Table 2 · Source: arXiv:2604.24929

% tasks revised

How many of the machine-translated tasks the audit changed at all.

Table 1 · Source: arXiv:2604.24929

Word & char edit rates

How far those revisions went: word- and character-level edit distance from the machine translation.

Table 1 · Source: arXiv:2604.24929

How this benchmark was built and run
Tasks
165 query–answer pairs per language, 825 in total across the five languages
Languages
Arabic, German, Hindi, Korean, Portuguese (Brazil), plus original English
Scoring
Follows the English benchmark, with modifications to handle commas as decimal points and other locale-specific formatting
Max turns
12 steps for the manager, 20 for the search subagent, following the English benchmark
Human review
Each task was adapted by bilingual annotators, then reviewed by a LILT researcher

Metrics Definitions

01 pass rate

pass rate (equivalently pass@1 or pass^1) is the chance a model will successfully pass a single task. We estimate it by having each model complete each task multiple times.

02 pass@k

pass@k is the chance that a model will pass a task at least once out of k attempts. Intuitively, this is best used for difficult tasks that could reasonably be solved by running the model multiple times and selecting the best answer. It is estimated using the unbiased estimator (see Appendix A of Chen et al., 2021).

03 pass^k

pass^k is the chance that a model will pass a task reliably in all k attempts. This metric is best used for tasks that can't be rerun (like interactions with a customer, as in τ³-bench) or where failure is expensive or dangerous. It is estimated using the unbiased estimator introduced in the first τ-bench paper (Yao et al., 2024).

LILTBench

LILTBench is an open collection of 92 coding tasks across 31 languages, authored by community teams to break frontier models in their own language. Tasks ship with the original-language instructions and an English translation, so the same problem can be run either way to see how much prompt language alone moves performance.

The tasks, the harness, and the leaderboard are public.