
Large language models can identify judgmental language in clinical notes, but the settings play a major role in accuracy.
“Addict,” “noncompliant,” “failed treatment” and “obese person” are examples of stigmatizing language that can appear in medical records. At George Mason University’s College of Public Health, researchers are exploring whether artificial intelligence (AI) can help identify this kind of language in clinical notes before it affects patient care.
Nurse scientist Teenu Xavier and colleagues found that large language models (LLMs) show promise in identifying stigmatizing language in clinical documentation, but their performance is highly dependent on their settings. Model size, temperature settings, prompting strategies and even note type can substantially influence results.
One finding was consistent across every model tested: Providing examples of stigmatizing language improved accuracy.
“Simply selecting an LLM is not enough when used for clinical documentation,” said Xavier, an assistant professor in the School of Nursing. “Careful attention must be paid to settings and prompting before these tools can be reliably used in health care environments.”
“Detecting stigmatizing language with large language models: mind the settings” was published in JAMIA Open.
Why does this matter?
The use of stigmatizing language in clinical documentation can reinforce bias and affect a patient’s future care. AI tools may be able to help identify this kind of language, promoting more equitable care and improving patient trust and experience.
“Pre-trained models, when optimized for identifying stigmatizing language, could help enable more timely interventions and modifications to the documentation process,” Xavier said. “Our research highlights the need for continued collaboration between health care professionals and AI developers to create tools that improve communication, reduce bias, and improve the overall patient experience.”
What are the detailed study findings?
- The largest LLM (trained on large amounts of data) was the best at predicting “stigmatizing” language (94%), but the worst at correctly predicting “not stigmatizing” language (47%).
- The smallest LLM was the best at correctly predicting “not stigmatizing” language (99.7%), but the worst at correctly predicting “stigmatizing” language (2%).
- When researchers gave the LLM an example of stigmatizing language, accuracy improved in all models.
- Emergency provider notes were categorized as “stigmatizing” or “non-stigmatizing” with the highest accuracy (69%), while plan of care notes had the lowest accuracy (56%). Misclassifications most commonly arose in long, clinically dense notes where neutral descriptions of complex illness or adverse events were mistaken by the models for judgmental language.
- Larger models worked best at lower temperatures (how predictable or random an LLM’s output is when making a classification), while smaller models improved with higher temperatures, meaning the models took more risks in interpretation.
Publication details
Teenu Xavier et al, Detecting stigmatizing language with large language models: mind the settings, JAMIA Open (2026). DOI: 10.1093/jamiaopen/ooag037
Journal information:
JAMIA Open
Key medical concepts
Clinical categories
This content is provided for information purposes only.
