The pilot goes well. Containment climbs in the US, the team signs off, and the same assistant is rolled out to Brazil and Poland on the assumption that a good model is a good model.
Three weeks later the Brazilian numbers are embarrassing. Same vendor, same configuration, same prompts, and customers are escalating to agents at twice the rate.
Nothing went wrong in the rollout. The problem is that model quality is not uniform across languages, and a multilingual AI contact center built on one global engine inherits every one of those gaps.
In this guide, we look at why AI performance varies so much by language, what to test before you sign, and how running different engines per market avoids the fragmentation that usually follows.
Why AI quality isn’t uniform across languages
Models learn from what exists. English dominates the training data available on the web, so English gets the most coverage, the most evaluation, and the most tuning attention.
The measured gaps are large. MMLU-ProX, published at EMNLP 2025, translated the same 11,829 questions into 29 languages and ran 36 leading models against all of them, finding performance gaps of up to 24.3% between high-resource and lower-resource languages on identical material.
It carries into production too, not just benchmarks. Research from Oracle’s AI team reported accuracy drops of up to 29% in non-English languages even inside retrieval-augmented systems, naming customer support as one of the settings where that inconsistency does real damage.
Three factors drive it:
- Training data volume: Low-resource languages have less text available, so the model has seen fewer examples of how people actually phrase problems in them.
- Dialect variation: European Portuguese and Brazilian Portuguese diverge in vocabulary and register, as do Gulf and Levantine Arabic. A model tuned on one can sound stilted or simply misread the other.
- Evaluation blind spots: Vendors publish English benchmark scores because that is what the market compares. Strong English results say very little about Turkish, Vietnamese, or Swahili.
That last point deserves emphasis. Stanford HAI’s 2026 AI Index found the top models clustered within roughly 25 Elo points of each other on public leaderboards, but a headline leaderboard score says little about how those models rank in any single language. Parity at the top of an English leaderboard does not mean parity in your market.
What to test market-by-market before committing
Vendor demos are run in the language the vendor is strongest in. Your test should not be.
Build a small evaluation set per market, ideally 50 to 100 real recordings or transcripts, and check five things:
- Accent and regional variation: Include the accents your actual customers have, not the neutral broadcast version of the language.
- Code-switching: In many markets people mix languages mid-sentence, dropping English product names into a Turkish or Hindi sentence. Code-switching breaks weaker models more often than accents do.
- Local idioms and politeness norms: A response that reads as efficient in Dutch can read as rude in Japanese. Formality registers are part of accuracy here.
- Domain vocabulary: Your product names, plan names, and error codes, spoken by customers who mispronounce them.
- Behavior under failure: What the assistant does when it does not understand. Guessing confidently in a language you cannot monitor is far worse than escalating.
Score each language separately and set a minimum bar per market. Averaging across markets hides exactly the problem you are trying to find.
Run the same set again whenever you change engines, which is what turns testing into a repeatable process instead of a one-off procurement exercise.
The case for mix-and-match instead of one global vendor
Standardizing on a single vendor is usually justified by simplicity: one contract, one integration, one support relationship. That is a real benefit, and it is often outweighed by the cost of being mediocre in half your markets.
The mix-and-match alternative is straightforward. Use the engine that performs best per language, and accept that the answer will differ by region. A per-market AI engine approach also removes the pressure to wait for your global vendor to improve a language they may not prioritize.
There is a portfolio argument as well. Language coverage is one of the areas where providers differentiate most sharply, and where they leapfrog each other most often. Locking every market to one roadmap means every market moves at the pace of that vendor’s weakest priority.
The objection is management overhead, and it is a fair one. Three engines usually means three consoles, three sets of reporting, and three ways for something to go wrong quietly. That is the problem the integration layer exists to solve.
How AI Connector runs different engines per market without fragmenting management
AI Connector sits between your contact center and whichever AI engines you choose, so the engine becomes a component you select per market rather than a platform you commit to globally.
In practice that means three things for a global operation:
- Different engines per language or region: ChatGPT, Claude, ElevenLabs, Vapi, a Dialogflow build, or something your own team trained for a specific market.
- One management panel: Call flows, assistant performance, and channel settings stay in a single Call Center Studio view, so a supervisor in Warsaw and one in São Paulo are looking at comparable numbers.
- Consistent escalation everywhere: When the assistant cannot resolve something, the call moves to the right live agent within seconds with full history and context, regardless of which engine was handling it.
Because the routing lives in the connector, adding a market does not mean rebuilding anything. The same applies to swapping an engine when a better option appears for one language, which happens more often than procurement cycles expect.
That consistency matters across channels too. Your enterprise operation probably supports voice alongside chat and messaging in each market, and language performance differs between them. Written channels tend to be more forgiving than voice, which is why many teams launch a new language on chat first and add voice once the accuracy holds.
For the operational side of running support across regions and time zones, our guide to building a multilingual call center covers the staffing and coverage decisions that sit alongside the technology ones.
Test it in your own languages
The only benchmark that matters is your own traffic, in your own markets. Test AI Connector across your top three languages side by side and see where the differences actually are before you standardize on anything.
FAQ
Can we use different AI vendors per region on one platform?
Yes. With an integration layer between your contact center and the engines, each market can run the assistant that performs best in its language while call flows, reporting, and escalation stay in one panel. The customer experience stays consistent; only the engine behind it differs.
How do we know which engine is better for a language?
Test them on your own interactions rather than published benchmarks. Take 50 to 100 real conversations from that market, run two engines against the same set, and compare resolution, escalation quality, and error behavior. English scores do not predict performance in other languages.






