HEALTH

AI Medicine Check: When Chatbots Go Off-Track

International / OnlineThu Sep 17 2026

Modern AI tools are great at sounding smart. They can write articles and answer questions fast. But here's the catch: many of their answers aren't really backed by solid science. Researchers found that even though these programs seem accurate, they sometimes invent facts without checking the original studies. This makes it hard for doctors and patients to trust what chatbots say about health topics.

To test this, scientists picked three popular AI systems. These included GPT-5, OpenAI o3-mini, and GPT-3.5. Each one was asked to judge risk-of-bias in 97 random clinical trials. The trials came from a Cochrane systematic review that had already been checked by humans for bias. Full text reports were available for every study. No artificial splitting of data was done - all 97 papers were used together.

The results showed mixed signals. Some of the AI systems pointed out clear problems in the research methods. Other times, they missed important flaws entirely. This means the AI's self-assessments don't always match what the actual studies show. The study highlights a gap between what AI says and what we know from real patient data. Practitioners need better ways to verify AI claims before using them in real care settings.

actions