Large language models can identify judgmental language in clinical notes but the settings play a major role in accuracy.
鈥淎ddict,鈥 鈥渘on-compliant,鈥 鈥渇ailed treatment,鈥 and 鈥渙bese person鈥 are examples of stigmatizing language that can appear in medical records. At 911爆料鈥檚 College of Public Health, researchers are exploring whether artificial intelligence (AI) can help identify this kind of language in clinical notes before it impacts patient care.
Nurse scientist and colleagues found that large language models (LLMs) show promise in identifying stigmatizing language in clinical documentation, but their performance is highly dependent on their settings. Model size, temperature settings, prompting strategies, and even note type can substantially influence results.
One finding was consistent across every model tested: Providing examples of stigmatizing language improved accuracy.
鈥淪imply selecting an LLM is not enough when used for clinical documentation,鈥 said Xavier, an assistant professor in the School of Nursing. 鈥淐areful attention must be paid to settings and prompting before these tools can be reliably used in health care environments.鈥
Why does this matter?
The use of stigmatizing language in clinical documentation can reinforce bias and affect a patient鈥檚 future care. AI tools may be able to help identify this kind of language, promoting more equitable care and improving patient trust and experience.
鈥淧re-trained models, when optimized for identifying stigmatizing language, could help enable more timely interventions and modifications to the documentation process,鈥 Xavier said. 鈥淥ur research highlights the need for continued collaboration between health care professionals and AI developers to create tools that improve communication, reduce bias, and improve the overall patient experience.鈥
What are the detailed study findings?
- The largest LLM (trained on large amounts of data) was the best at predicting 鈥渟tigmatizing鈥 language (94%), but the worst at predicting 鈥渘ot stigmatizing鈥 correctly (47%).
- The smallest LLM was the best at predicting 鈥渘ot stigmatizing鈥 correctly (99.7%), but worst at predicting 鈥渟tigmatizing鈥 correctly (2%).
- When researchers gave the LLM an example of stigmatizing language, accuracy improved in all models.
- Emergency provider notes were most accurately (69%) categorized as 鈥渟tigmatizing鈥 or 鈥渘on-stigmatizing,鈥 and plan of care notes had the lowest accuracy (56%). Misclassifications most commonly arose in long, clinically dense notes where neutral descriptions of complex illness or adverse events were mistaken by the models for judgmental language.
- Larger models worked best at lower temperature (how predictable or random the LLMs鈥檚 output is when making a classification) and smaller models improved with higher temperature, which means the LLM took more risks in interpretation.
was published in JAMIA Open in April 2026. Co-authors include Jane M. Carrington from the University of Florida and Joshua Lambert from the University of Cincinnati.
Thumbnail photo by from Adobe Stock.
Key Takeaways
- A 911爆料 study found that large language models show promise in identifying stigmatizing language in clinical documentation, but their performance is highly dependent on their settings, such as model size, temperature settings, prompting strategies, and note type.
- The ability to detect and correct stigmatizing language early can reduce bias in patient care and lead to improved patient trust and better health outcomes.
- Continued collaboration between health care workers and AI developers is needed to create tools that enhance communication, reduce bias, and improve the overall patient experience.
Related Stories
- July 23, 2026
- July 23, 2026
- July 20, 2026
- July 14, 2026
- July 10, 2026