Evaluating Spanish Translations of Emergency Department Discharge Instructions by a Large Language Model: Tool Validation and Reliability Study
JMIRResearch Authors: Jossie A Carreras Tartak, Ryan CL Brewster, Daniela Arango Isaza, Antonio Berumen Martinez, Ana Grafals, Phanidhar Adusumilli, Ted Fitzgerald, Roger Orcutt, Larry A Nathanson, Adrian D HaimovichAIIM Authors: Pearl Marks, Amanda ZhongApproved by President Reda RiffiPublication Date: 1/12/2026Comprehensive Summary
This prospective, single-center study looked at whether a large language model could generate clinically acceptable Spanish translations of emergency department (ED) discharge instructions, using a large language model (LLM) to perform medical text translation. Researchers analyzed n = 100 randomly sampled free-text ED discharge instructions from an urban academic medical center, collected between July 1 and December 31, 2024. Instructions were translated using a custom-developed prompt deployed in a PHI-compliant environment with Claude Sonnet 3.5, following iterative prompt refinement on smaller batches. Translations were independently evaluated by 2 native Spanish-speaking physicians and 2 certified medical interpreters using a structured 5-domain Likert rubric assessing completeness, fluency, meaning, severity, and overall quality. The Claude Sonnet 3.5 model was assessed against a predefined clinical acceptability threshold rather than direct comparison to interpreter-generated translations. The best-performing output achieved mean Likert scores of 4.8–5.0 across domains, with 95% confidence intervals tightly clustered at the upper end of the scale. The analysis showed that 0% (0/100) of translated discharge instructions were ultimately deemed clinically unacceptable. Only 1% (1/100) was flagged initially for scoring ≤3 due to a regional lexical variation in the translation of “concussion,” and once adjudicated, this translation was considered clinically acceptable. Secondary analyses included stratification of scores by reviewer type (physician vs interpreter), demonstrating consistently high agreement and performance across groups. Additional results showed near-ceiling performance in completeness and severity (mean 5.0), with fluency and meaning also scoring highly (means 4.8–4.9). Limitations include the single-center design, restriction to Spanish-language translations, lack of direct comparator performance against live interpreter translation, and absence of formal subgroup or fairness analyses. External validation was not performed, and findings reflect translation quality within a controlled dataset rather than downstream patient comprehension, safety, or clinical outcomes.
Outcomes and Implications
This study suggests that institutionally deployed LLMs can produce high-quality, auditable translations of ED discharge instructions, addressing a critical equity gap for patients with limited English proficiency. In clinical practice, this approach could be applied to rapidly generate language-concordant discharge materials when interpreter resources are limited, supplementing existing workflows and reducing reliance on unapproved consumer translation tools. However, translation to bedside care remains indirect and will require broader validation across institutions, languages, and patient populations, as well as integration with safeguards for dialectal variation and clinical oversight before routine implementation.
Connect medicine with AI innovation.
No spam. Only the latest AI breakthroughs, simplified and relevant to your field.