AI Medicine Check: When Chatbots Go Off-Track
Modern AI tools are great at sounding smart. They can write articles and answer questions fast. But here's the catch: many of their answers aren't really backed by solid science. Researchers found that even though these programs seem accurate, they sometimes invent facts without checking the original studies. This makes it hard for doctors and patients to trust what chatbots say about health topics.
To test this, scientists picked three popular AI systems. These included GPT-5, OpenAI o3-mini, and GPT-3.5. Each one was asked to judge risk-of-bias in 97 random clinical trials. The trials came from a Cochrane systematic review that had already been checked by humans for bias. Full text reports were available for every study. No artificial splitting of data was done - all 97 papers were used together.