Back to Blog
clinical-documentation
July 29, 2026
10 min read

AI Scribe with Limited-English-Proficiency Patients: Interpreters, Accents, and Accuracy

How AI medical scribes perform with patients whose primary language isn't English, including interpreter-mediated visits, accented English, and the consent dynamics that matter.

Fatih Aktas

By Fatih Aktas, Founder & CEO

Published

person sitting while using laptop computer and green stethoscope near. Cover image for: AI Scribe with Limited-English-Proficiency Patients: Interpreters, Accents, and Accuracy.
person sitting while using laptop computer and green stethoscope near. Photo by National Cancer Institute on Unsplash.

A topic most vendors don't address

If you ask an AI scribe vendor how their tool performs with patients who don't speak English well, you'll usually get a confident-sounding answer that turns out to be hedged when you press for specifics. The honest answer is: it depends on the patient, the interpreter setup, the accent, and the visit type, and it varies more than vendors like to admit.

For a clinic that serves a significant number of patients with limited English proficiency, this isn't a theoretical concern. It's a daily reality that shapes whether the AI scribe actually fits the practice.

This article is the grounded guide. What works, what doesn't, what to test before relying on it, and how to handle the consent conversation when neither the patient nor the AI is operating in their first language.

How AI scribes actually handle non-native speakers

In 2026, the major AI scribe platforms have substantially improved English accent handling compared to two years ago. Speech recognition has gotten better at South Asian, East Asian, Latin American, African, and Eastern European English accents. The major remaining gaps are:

Heavy regional accents that the training data underrepresents. West African English, certain South Asian regional accents (Telugu-influenced English, for instance), and some Eastern European accents still produce noticeably higher error rates than General American English.

Code-switching mid-sentence. A patient who speaks English but slips into Spanish, Tagalog, Mandarin, or Arabic for specific words (often medication names, body parts, or emotionally loaded terms) often confuses the AI. The English portions transcribe well; the inserted non-English words become noise or misheard substitutes.

Visits primarily in a non-English language. If most of the conversation is in Spanish, Mandarin, French, or another non-English language, most US-built AI scribes either don't transcribe at all or produce unreliable output. A small number of platforms (including some that target the Canadian Quebec market for French) handle this well; most don't.

Whispered or quieted speech that often accompanies language uncertainty. Patients who are unsure of their English sometimes speak quietly. The combination of accent and low volume is harder for the AI than either alone.

The first step for any clinic with significant LEP patient volume is to test the AI on real visit audio from your patient population before signing a contract. Vendors will demo with clean American English; your reality may be different.

Interpreter-mediated visits

For visits where the language gap is large enough to require an interpreter, the audio dynamic changes significantly. Three people are now in the conversation: provider, patient, interpreter. The AI has to handle three voices, two languages, and a turn-taking pattern that doesn't match its training expectations.

The pattern that works best, when it works:

The interpreter is in the room, in person, with their own clear audio path. A second microphone or a strategic position for the interpreter helps the AI separate the voices. The AI captures the provider's English, the interpreter's English-to-target-language and target-language-to-English, and depending on platform, may or may not capture the patient's native-language speech.

The clinical note captures the interpreter's English back-translation, not the original target-language speech. This is usually what you want; the chart records what was communicated in English between provider and interpreter, which is the clinical content. The patient's original target-language speech isn't typically required in the chart for clinical purposes.

The pattern that often breaks:

Phone or video interpreter services. When the interpreter is on a phone speaker or a video screen, audio quality drops sharply. The AI struggles with the compressed, sometimes delayed audio from the interpreter line. Many practices find that AI scribes don't work well with phone interpreters and choose to type these visits.

Family member interpretation. This is a problematic pattern clinically (family-member interpretation introduces accuracy and confidentiality issues that violate guidance from major medical organizations) and it's also a problematic pattern for AI scribes (less consistent audio, more code-switching, more emotional dynamics). For both reasons, family-member interpretation is best avoided; when it happens, expect the AI output to need substantial editing.

Sight translation of written materials. When the provider hands the patient a written form and the interpreter sight-translates it aloud, the audio captures interpretation that doesn't correspond to the actual visit content. AI scribes can include this content as if it were patient disclosure, which is incorrect. Sight-translation moments need to be edited out of the chart manually.

The consent conversation across a language gap

The standard consent script ("I'd like to use an AI tool that helps me take notes during our visit, with your permission") assumes the patient understands the question. With a language barrier, the standard script doesn't work.

What works better:

With an in-person interpreter present from the start. The provider gives the consent explanation to the interpreter at normal speed. The interpreter translates it for the patient, then translates the patient's response back. The provider takes the response at face value, the same as with an English-speaking patient.

Without an interpreter, with a patient who speaks some English. Use simpler language. "I have a computer that helps me write notes about our visit. Is that okay with you?" Shorter sentences, concrete words, no jargon. Watch the patient's body language for understanding. If they nod automatically without comprehension (a common pattern when patients are deferring to the provider's authority), check by asking "do you have questions about that?" and seeing whether the response is genuine.

With written translated materials. A short consent paragraph in the patient's primary language, kept in the exam room. Translated versions in Spanish, Mandarin, Vietnamese, Tagalog, Arabic, French, Punjabi, and Russian cover most US LEP populations; your specific patient demographic may need others. The paragraph isn't legally binding consent (verbal acknowledgment still is) but it gives the patient a chance to read in their own language.

The critical principle: a patient who doesn't fully understand what they're consenting to hasn't really consented. If the consent conversation across the language gap can't be made meaningful, default to not using the AI for that visit. The chart can be typed; the patient's right to informed participation can't be retrofitted.

Accuracy expectations and how to set them

A clinic with significant LEP patient volume should expect AI scribe accuracy to be lower for those visits than for English-fluent visits. The realistic ranges, based on observed performance across major platforms in 2026:

Visit type Realistic accuracy
English-fluent patient, common accent 92 to 96%
English-fluent patient, less-represented accent 87 to 93%
LEP patient speaking English, common accent 80 to 90%
LEP patient speaking English, heavy accent 70 to 85%
In-person interpreter visit 75 to 88%
Phone interpreter visit 60 to 80%
Predominantly non-English visit (most US platforms) Unusable

The lower accuracy doesn't mean the AI scribe is useless for these visits. It means the provider review step is more substantive. A note that needs more editing still saves time over typing from scratch, just less time.

Some practices establish internal rules:

  • Use the AI for fluent English visits, type for predominantly non-English visits.
  • Use the AI with in-person interpreters, type for phone interpreter visits.
  • Use the AI for follow-ups (where the medical content is more predictable) more aggressively than for new patients (where the history-taking matters more).

These rules don't have to be rigid. They're starting points that get refined based on what the clinic actually observes.

The vendor questions to ask

If a significant portion of your practice is LEP patients, the vendor evaluation questions need to be more specific than the generic accuracy claims:

  1. What languages does your platform fully support for transcription, beyond English? Some platforms support Spanish, French (Quebec or France), Mandarin, or others. Most don't.
  2. How does your platform handle code-switching mid-sentence? Does it transcribe the non-English word phonetically, ignore it, or attempt translation?
  3. How does your platform handle three-way conversation with an interpreter? Are there special interpreter-mode features?
  4. Can your platform transcribe phone interpreter audio with reasonable accuracy? Most can't; vendors who claim they can should be tested specifically.
  5. Do you have customers serving primarily LEP populations? Can I talk to one? This is the most useful question. A vendor that has real LEP-serving customers can offer specific guidance; a vendor that doesn't will hedge.

Practices in markets with large Spanish-speaking populations (California, Texas, Florida, much of the US Southwest) should make Spanish handling a first-order evaluation criterion. Practices with significant immigrant populations from less-represented language communities (Somali in Minneapolis, Hmong in Wisconsin, Vietnamese in Houston) may find that most AI scribe platforms don't serve those populations well, and the practice's evaluation should reflect that.

Bilingual providers

A provider who is themselves bilingual and conducts visits in their own non-English language is a different case. The AI behavior depends on the provider's primary language:

Provider conducts visits in Spanish, AI scribe is English-only. Provider speaks Spanish with the patient and English summary into the AI scribe (sometimes called "summary dictation" mode). The note is in English, derived from the provider's English summaries, not from the Spanish conversation. This works well; it's similar to traditional dictation.

Provider conducts visits in Spanish, AI scribe supports Spanish. The AI captures the Spanish conversation and either generates a Spanish note (if the EHR supports it) or auto-translates to English. The note quality depends on the platform's Spanish capabilities. Some platforms do this well; many don't.

Provider conducts visits in English with mostly bilingual patients who code-switch. This is the common case in many US bilingual practices. The AI handles the English well, struggles with the Spanish inserts. Notes need more editing; time savings are smaller than for monolingual English practices.

The honest summary

If your practice serves predominantly English-fluent patients with familiar accents, AI scribes work essentially as advertised. The accuracy is high, the time savings are real, the consent conversation is straightforward.

If your practice serves a significant population of patients with limited English proficiency, heavy regional accents, or non-English primary language, the AI scribe is still useful but the picture is more complicated. The accuracy is lower, the consent is harder to obtain meaningfully, the provider review step is more substantive, and the time savings are smaller. The tool is still net positive in most cases; it just doesn't deliver the same magnitude of benefit.

The practices that succeed in LEP-heavy environments are the ones that test on real patient audio before signing, set realistic expectations with providers, invest in interpreter setup, and maintain a "type instead of AI scribe" default for visits where the language gap is too wide for the AI to handle responsibly.


For the broader consent conversation with patients across all visit types, see talking to patients about AI scribes. For the exam room hardware that affects audio quality across all patient types, exam room setup for AI scribes covers the audio infrastructure side.

limited-english-proficiencyinterpretersaccuracymultilingualaccessibility

Ready to Try AI-Powered Documentation?

Join thousands of healthcare providers saving hours every day with Transcribe Health.

Start Free Trial

This article is informational and not medical or legal advice. See our medical and legal disclaimer and our editorial policy for how we research and attribute content. Consult a licensed clinician for medical decisions and a licensed attorney for regulatory interpretation in your jurisdiction.