CMU Study Shows LLMs May Invent Information When Missing Key Data
The Breakdown
- In 18% of cases, Claude, GPT-5 and Gemini fabricated medical diagnoses, despite supporting images being omitted.
- The AI models used demographic-based clinical assumptions to invent these diagnoses.
- CMU research emphasizes the need for demographic sensitivity testing and verification before deploying AI tools in medical systems.
* * *
A Carnegie Mellon University School of Computer Science student recently showed that large language models (LLMs) intermittently invent false information when responding to medical questions.
For his study, Robotics Institute master’s student Siddharth Vohra asked popular LLMs to describe a medical image that was intentionally omitted from the query. Rather than requesting the missing image, the models fabricated a diagnosis based on the user’s age, gender and race 18% of the time.
“A 65-year-old white man asking Claude about a skin mole receives melanoma in nearly every response,” said Vohra, the study’s sole author. “When chest X-ray questions are presented, OpenAI’s GPT-5 names sarcoidosis for roughly 77% of young Black patients.”
In reality, fewer than one in 10,000 moles will become melanoma, according to the Memorial Sloan Kettering Cancer Center. The prevalence of sarcoidosis varies throughout the world, but the Cleveland Clinic reports that there are typically fewer than 200,000 cases of sarcoidosis at any given time in the U.S., making it quite rare.
Vohra was raised by two physician parents who exposed him to the healthcare industry at an early age. After reading a research paper detailing how visual language systems often provide incorrect responses, he decided to investigate how demographic information influences the way models respond to medical questions.
“I analyzed close to 11,700 model responses across Claude, GPT-5 and Gemini,” Vohra said. “About 82% of the time, the models refused to provide a response because no image was attached. However, in the remaining 18%, the models invented a diagnosis instead of asking for the missing image.”
When conducting the study, Vohra focused on chest X-ray, brain MRI and dermatology images in 12 simulated patient profiles. The findings reinforced a critical concern Vohra sees across the AI industry: users often assume that models understand more than they actually do.
“There’s a general perception that AI models are smart because they perform well on certain tasks,” Vohra said. “But being good at one thing doesn’t mean they’re good at everything. AI is used for a host of tasks, from coding to healthcare advice, but different domains require different levels of reliability.”
Because small changes in the words used to describe a patient can drastically influence a model’s conclusions, Vohra stressed that clinical pipelines should audit and test for demographic sensitivity before deploying AI models in healthcare settings.
“These models are incredibly useful and they’re improving quickly,” Vohra said. “As labs continue advancing the models, stringent testing and verification checks are necessary, especially when relying on the models for medical decisions. Strong performance on one task doesn’t guarantee safety in another.”
Vohra’s current focus includes running the study at a much larger scale, with the goal of surfacing these failure patterns so the labs building these systems can identify and address them. He presented his research at the TrustVLM workshop at the ACM International Conference on Multimedia Retrieval in Amsterdam this past June. He was also recently selected for the Google Gemini Academic Program Award 2026, which promotes trustworthy AI models and mission-critical AI deployments.
For more on Vohra’s work, visit his website.
For More Information: Aaron Aupperlee | 412-268-9068 | aaupperlee@cmu.edu
