Open-Access Evaluation Surface

The Frontier of Multilingual LLM Performance

An open-access, scientifically rigorous evaluation surface mapping the performance and agentic capabilities of frontier models across non-English tasks, socio-cultural contexts, and multimodal domains.

Standard evaluation frameworks mask the performance degradation of frontier systems when deployed outside English-native workflows. LILT Applied AI stress-tests LLMs under realistic, localized constraints to map the true boundaries of global AI capability.

Benchmark Success Rates

Task execution accuracy across our multilingual suite.

GAIA-v2 · LILT agentic reasoning + tool use
GAIA-v2-LILTCharacter- and word-level edit rates against the MAPS baseline
Language% tasks revisedWord-level edit rateChar-level edit rate
Arabic84.825.419.4
German81.230.022.2
Hindi100.055.446.3
Korean92.136.927.9
Portuguese87.925.019.5
GAIA-v2-LILTAudited performance scoring against the MAPS baseline
0255075100ardehikopt-BRavg
MT baseline Audited
Multilingual Terminal-bench task resolution rate per model
Multilingual Terminal-benchTask resolution rate per model
0%20%40%60%80%100%Resolution RateClaude Opus 4.7 Thinking42.5%Claude Opus 4.741.4%Gemini 3.1 Pro (preview)41.0%Claude Opus 4.640.3%Claude Opus 4.8 Thinking34.4%GPT-5.424.6%GPT-522.8%GPT-5.216.7%
tau-bench customer support multi-turn
tau-benchTask pass rate per model · airline & telecom · EN / DE / KO
0.00.20.40.60.81.0airline0.740.680.50EN0.800.680.44DE0.700.540.56KOtelecom0.590.450.23EN0.410.330.25DE0.390.270.20KOpass@1 and pass^4
Gemini 3.1 Pro Opus 4.8 GPT-5.4
MultiChallenge long instruction · memory · self-coherence
MultiChallengeMulti-turn conversational accuracy per model by language
4045505550.6945.0546.8043.89GPT-5.253.7449.0451.6549.06Claude Opus 4.650.8246.7946.7645.02Gemini 3.1 Flash-Litepass@1
English Arabic German Korean
MultiChallengeTokens per dialogue by language
0100020003000400050001572English4558Arabic2371German3588Koreantokens / dialogue

Metrics Definitions

01 pass rate

pass rate (equivalently pass@1 or pass^1) is the chance a model will successfully pass a single task. We estimate it by having each model complete each task multiple times.

02 pass@k

pass@k is the chance that a model will pass a task at least once out of k attempts. Intuitively, this is best used for difficult tasks that could reasonably be solved by running the model multiple times and selecting the best answer. This is estimated as above, but with the use of the unbiased estimator (see Appendix A of Chen et al., 2021).

03 pass^k

pass^k is the chance that a model will pass a task reliably in all k attempts. This metric is best used for tasks that can't be rerun (like interactions with a customer, as in τ³-bench) or where failure is expensive or dangerous. It is estimated using the unbiased estimator introduced in the first τ-bench paper (Yao et al., 2024).

LILTBench Arena

Live results from the LILTBench task collective — community teams authoring coding tasks designed to break frontier models in their own language, scored against vertex_ai/claude-opus-4-6 over 15 iterations. Click a row to expand language coverage & scores.

#
Team
Tasks
Languages
Native
English
Points
🥇
dot is a catam · ar · bo · cr · de · dv · fr · jv
35
17
167/525
158/525
195
🥈
Fernwehar · da · de · fi · he · is · ja · kl
23
12
169/345
174/345
99
🥉
Captain Nemoar · fr · he · ja · ko · te · tr · vi
16
9
132/240
133/240
65
4
Baseline Repo Tasksde · hi · ja · ko
9
4
108/135
91/135
23
5
Khushi Kumarisd
4
1
31/60
29/60
18
6
Mohamed Ahmedar
3
1
25/45
20/45
11
7
rootshi
2
1
30/30
30/30
2

Sourced from liltbench-tasks-public · updated 2026-06-24 13:32 UTC · click any header to sort