AI Scribe for Group Visits, Family Meetings, and Multi-Speaker Encounters
How AI medical scribes handle visits with more than one speaker, including family meetings, shared medical appointments, and pediatric visits with parents.

By Fatih Aktas, Founder & CEO
Published

Why multi-speaker visits are hard
Most AI medical scribes were trained on two-speaker visits: one provider, one patient. When the room has more speakers (a parent and a child, a couple at a fertility consult, a family meeting at end-of-life, a group medical appointment with 8 patients), the AI's accuracy can drop sharply and in non-obvious ways.
The drop isn't usually a transcription accuracy issue. The AI still hears the words correctly. It's a diarization issue: figuring out who said what. Without correct speaker attribution, the note can attribute the patient's report of pain to the parent, or the daughter's question about end-of-life wishes to the patient, or one group member's history to another.
This article covers how to think about AI scribes in multi-speaker visits: when they work, when they don't, and how to compensate when they're partially working.
How AI scribes handle multiple speakers in 2026
The major platforms have improved at speaker diarization meaningfully over 2024 to 2026, but the picture is uneven:
Two speakers, both adults, both speaking clearly: 95%+ accurate attribution. This is the default case and it works well.
Two speakers with one quiet voice or accent gap: 85 to 92% accurate attribution. Useable but with more review needed.
Three speakers (e.g., provider, patient, family member): 75 to 88% accurate attribution. The note often correctly captures the content but assigns it to the wrong speaker. Review must verify speaker attribution, not just word capture.
Four or more speakers: variable, often below 70% accurate attribution. Some platforms degrade gracefully (everyone becomes "speaker 3"); others scramble the attribution randomly.
Children's voices alongside adult voices: typically worse than two adult voices. Pitch differences confuse some diarization models; younger children's voices may be intermittently classified as the wrong speaker.
The practical implication: a two-speaker visit benefits from "AI does diarization, provider trusts it." A three-or-more-speaker visit needs "AI does transcription, provider does diarization in review."
Family meetings: the highest-value multi-speaker case
End-of-life family meetings, complex care planning sessions, and difficult-news conversations are among the longest, most cognitively demanding visits a provider has. They're also visits where the documentation burden is high (often 20 to 40 minutes of note-writing afterward) and where the documentation matters legally and clinically.
These are the highest-value visits for AI scribe assistance and also the hardest. The approach that has worked:
Use the AI to capture the conversation, but plan to edit speaker attribution. Going in with the expectation of editing reduces the slump that comes from realizing the diarization is imperfect. The provider's job in review is to verify who said what, not to assume the AI got it right.
At the start of the meeting, introduce who's in the room. "I'm here with Dr. Smith. Also in the room is Mr. Jones, his daughter Sarah, and his son Michael. Sarah is the healthcare power of attorney." This serves both clinical purposes (orienting everyone) and AI-scribe purposes (giving the model an early speaker assignment context to work with).
Don't try to capture every word; capture the decisions. A family meeting often has many emotional moments that don't need to be in the chart. The chart needs to record what was discussed, who participated, what decisions were reached, and the patient's or surrogate's expressed wishes. The AI tends to capture more than this; the provider's review condenses.
Document the consent to record explicitly. With multiple family members in the room, consent is more nuanced. The patient's consent is the primary one (assuming decisional capacity); the family members' consent to be recorded is also relevant. The introduction can include: "I'd like to use a tool that helps me write up our meeting. Is that okay with everyone?"
Send a written summary after the meeting. This is good clinical practice anyway and the AI-generated draft accelerates it. The summary documents what was discussed and decided, in a way the family can refer back to. It also reduces the volume of follow-up questions.
Shared medical appointments (group visits)
A shared medical appointment (SMA) is a clinical encounter where 6 to 12 patients with a common condition meet together with a provider and sometimes a behaviorist. They're used in diabetes care, weight management, perinatal care, and behavioral health.
AI scribes generally do poorly with SMAs. The reasons:
Too many speakers. The diarization breaks down with 8+ speakers.
Privacy concerns. Each patient hears the other patients. Recording the session means each patient's audio includes other patients' health information. The consent dynamics are more complex than a standard visit.
Note structure is different. A typical SMA generates one note per patient summarizing their specific participation and care plan, not one note for the whole session. AI scribes designed for one-on-one visits don't fit this structure.
A few approaches that have worked:
- Don't record the group portion. Use the AI scribe only during the individual portions of an SMA (the one-on-one check-ins before or after the group conversation). Document the group portion manually.
- Record the group portion as ambient context, not as the primary note source. The provider's notes capture each patient's individual contribution from memory; the recording is used as a memory aid, not as the note's source.
- Use a dedicated SMA documentation template that doesn't rely on AI. Some practices have given up on AI scribing SMAs entirely and have built efficient templates that capture the necessary content in 5 to 10 minutes per patient post-session.
For practices that do many SMAs, this is one of the few areas where AI scribes are clearly not a fit yet. The technology is improving but the consent and structure issues are not just technical.
Pediatric visits with parents
The most common multi-speaker visit pattern: a pediatric or family practice visit where a parent is present with a child. Sometimes two parents. Sometimes a parent and an older sibling.
The AI scribe behavior varies:
Younger children (under 5) who speak little. The AI primarily captures the provider and parent conversation. The child's contributions are minimal anyway. Diarization is generally fine.
School-age children (5 to 12) who speak some. The child's voice may be misattributed to the parent, particularly with quieter children. The provider should pay extra attention to which speaker is being attributed for "I have a tummy ache" type statements.
Adolescents (13 to 17) who speak about themselves. The AI typically diarizes the adolescent and parent voices distinctly. The clinical content is often dense (adolescent self-report differs from parental observation in clinically important ways). Speaker attribution matters here.
Visits where both parents are present. Both parents are often similar-pitched adult voices. The AI may not distinguish them well. This usually isn't a clinical problem (their distinct identities don't usually matter for the note) but it can create confusion in the chart.
The general principle: pediatric visits work fine with AI scribes, with attention to speaker attribution in review. The savings are real but a bit smaller than in pure adult primary care because the review takes a bit longer.
Couples visits in primary care and OB
OB visits with both partners present, fertility consultations, couples mental health visits, and primary care visits where a spouse accompanies the patient all share the same pattern: two adults plus the provider, both of whom may be speaking about the same clinical topic from different perspectives.
The AI behavior:
- Generally handles three-speaker visits with two similar-aged adults well, provided both speak clearly
- May misattribute back-and-forth between the couple, especially when both are commenting on the same topic
- Captures the content correctly even when attribution is off; the provider review should focus on who said what
A common documentation choice: document the visit as "patient" with relevant input from the partner noted, rather than trying to maintain strict attribution. "Patient reports increased depression over past month. Partner reports observing crying spells and social withdrawal." This is cleaner than "Patient said X, Partner said Y, Patient agreed" structures, and it's how the chart would have been written without AI scribe anyway.
The home health and SNF context
Visits in nursing homes, assisted living, and home health often involve multiple speakers: the patient, a family member, sometimes a caregiver, sometimes facility staff. They also often have worse audio environments than office visits.
The combination of multi-speaker + worse audio is harder for AI scribes than either alone. Practices that do significant home health or SNF work often find that AI scribes underperform their office expectations meaningfully.
Approaches that help:
- A dedicated portable microphone, not the laptop's built-in mic. A USB lapel mic on the provider works in environments where the conference mic isn't practical.
- A consent script that addresses everyone present. "I'd like to record our visit today. Is that okay with you, Mrs. Smith, and with you, Tom, and with your CNA Maria?" The plural consent reduces ambiguity.
- Lower expectations on time savings. AI scribes that save 30 minutes per day in office often save only 10 to 15 minutes per day in home health, because the review is more substantive.
Vendor questions for multi-speaker work
If your practice has significant multi-speaker visit volume, the vendor evaluation should specifically address it:
- How does your platform handle three or more speakers?
- Can you provide examples of family-meeting notes generated by your platform?
- How does your platform handle pediatric visits with parent and child present?
- Is there a "shared medical appointment" or "group visit" mode?
- How does your platform indicate uncertainty about speaker attribution? (Some flag low-confidence attributions; some don't.)
The vendor that has thought through multi-speaker scenarios will have specific answers. The vendor that hasn't will say something general about "robust speaker handling" without detail. The latter is a yellow flag.
The honest framing
Most AI scribes work well for two-speaker visits and adequately for three-speaker visits. Beyond three speakers, the technology hasn't yet caught up reliably, and the practices that have the most multi-speaker visits (group medical appointments, family conferences in palliative care, complex social work cases) get the least benefit.
For most outpatient practices, the multi-speaker visits are a small portion of the schedule. The AI scribe is still net positive even if those specific visits aren't well-served. For practices where multi-speaker visits are a major portion of the work, the AI scribe evaluation should include those specific visit types in the trial, and the conclusion may be different from a typical primary care evaluation.
For the specific case of pediatric practice considerations, see the patient consent conversation which covers the script across visit types. For end-of-life and palliative care contexts where family meetings are common, the same consent principles apply with extra attention to the family members' presence.
Articles connexes
When NOT to Use an AI Scribe: Visits Where It Makes Things Worse
Most articles about AI medical scribes are about when they help. This one is about the visits where they hurt, and why knowing the boundary protects both patients and providers.
clinical-documentationAI Scribe for Nurse Practitioners and Physician Assistants
How AI medical scribes fit into NP and PA practice patterns, including supervisory note workflows, prescribing documentation, and the scope-of-practice nuances that affect note structure.
clinical-documentationAI Scribe with Limited-English-Proficiency Patients: Interpreters, Accents, and Accuracy
How AI medical scribes perform with patients whose primary language isn't English, including interpreter-mediated visits, accented English, and the consent dynamics that matter.
Related Resources
Prêt à essayer la documentation propulsée par l'IA?
Rejoignez des milliers de professionnels de la santé qui économisent des heures chaque jour avec Transcribe Health.
Essai gratuit