Multilingual agents need per-language action evals, not translated prompts
I qualify a multilingual agent in each language at the action boundary, not by translating an English prompt suite and comparing fluent answers. An agent can preserve meaning well enough to chat while losing the distinctions that govern tool choice, refusal, confirmation, and recovery. Once language changes what the system does, language is part of the runtime contract.
Translation-first evaluation hides this risk. An English scenario carries English assumptions about institutions, politeness, legal categories, dates, names, and what an instruction naturally implies. Translating the words can leave those assumptions intact while producing phrasing a native user would not choose. A model may score well because its training included similarly translated text, not because it can handle the target language in its own cultural and operational setting. Native-authored evaluation separates language capability from familiarity with English-shaped questions.
For an agent, I would go beyond answer correctness. Each language suite should test whether the same user intent maps to the right tool, arguments, target, and effect; whether ambiguous authority triggers clarification; whether a prohibited action is refused; whether a recoverable failure leads to the right retry; and whether the final external state matches the request. This is the multilingual version of verifying resulting state rather than clicks. The sentence can be imperfect while the action is safe, or perfectly fluent while the action is wrong.
The evaluation matrix should be explicit. Rows are languages and relevant regional variants. Columns are action classes: read, write, delete, transfer, disclose, escalate, and abstain. I would slice results by language resource level, script, mixed-language input, code-switching, and domain. Safety deserves its own cells because English-centric alignment does not transfer uniformly. A refusal rate averaged across languages can conceal the exact language in which a dangerous request becomes executable.
Routing does not remove the obligation. Language detection may select a specialist model, a language adapter, cross-lingual retrieval, or translate-then-process fallback. Each route is a different production path with its own version, tokenizer behavior, latency, and failure modes. Following the principle that agents should route on verified capability, I would authorize only the action classes that a language route has passed. “Supports 50 languages” is catalogue metadata; “may execute account changes in Tamil” is a tested capability claim.
I also want the failure ontology preserved across languages. Wrong target, missing qualifier, unsafe disclosure, over-refusal, mistranslated identifier, culturally wrong assumption, and unhandled code-switch should remain distinct labels. That connects to a shared failure ontology for safety evals. If each localization team invents different categories, global averages become incomparable and remediation cannot be routed to data, model, prompt, policy, or interface owners.
I concede that translated English suites are useful as cheap smoke tests, especially before native coverage exists. They catch gross regressions and enable a shared core across many languages. They cannot support a production action claim on their own. Until native-authored scenarios, native review, and state-level verification exist, the honest result is “language path available with constrained authority,” not “multilingual agent qualified.”
The operating model follows from the evidence. Add languages gradually, bind permissions to evaluated action classes, shadow risky actions, capture disagreement from native reviewers, and turn real failures into regression cases. Evals are the contract here in a literal sense: the per-language suite defines what the agent is allowed to do, not merely how well it can sound while doing it.