August 21, 2026
2 min read
Key takeaways:
- Few AI tools had registered clinical trials and were tested on patient outcomes.
- Studies that did evaluate AI tools were often small and excluded vulnerable populations.
Of more than 1,300 AI medical devices cleared by the FDA, only three tested for patient-centered outcomes like death and morbidity, data show.
AI integration into healthcare is rapidly progressing. Medical societies like the AMA have called for stricter regulations due to concerns about bias and transparency in clinics and health insurance decision-making.
Few AI tools had registered clinical trials and were tested on patient outcomes. Image: Adobe Stock
According to Rawan Abulibdeh, a postdoctoral fellow at University Health Network in Canada, and colleagues, the FDA has cleared at least 1,357 AI medical devices.
“Clinical decisions increasingly depend on algorithmic outputs,” they wrote in PLOS Digital Health. “Yet despite this rapid adoption, one question remains largely unanswered: Do these tools actually improve patient outcomes?”
The FDA’s regulatory framework for AI devices only requires that a new device shows “’substantial equivalence’ to an existing device rather than the prospective validation of clinical effectiveness,” according to the researchers.
Clearance through the 510(k) pathway, they added, “allows evidence gaps to propagate through chains of predicate divides, many of which themselves lack rigorous clinical validation.”
In the systematic review, Abulibdeh and colleagues utilized the FDA device database and American Academy of Rheumatology Data Science Institute catalogue to determine how many AI tools cleared by the FDA before Dec. 5, 2025, were tested on patient health outcomes, including death, morbidity, hospitalization, readmissions, quality of life and symptom burden.
Of the 1,357 devices:
- 34 were linked to registered prospective trials;
- 12 had results posted on ClinicalTrials.gov;
- 12 had peer-reviewed publications; and
- three assessed patient outcomes.
The researchers noted that many studies were limited in size and often excluded vulnerable patient populations, including pregnant patients, youth, older adults and non-English speakers.
“Nearly three-quarters enrolled fewer than 500 participants, and one-quarter enrolled fewer than 100,” they wrote.
Abulibdeh and colleagues noted that logistical challenges and a lack of financial incentives may be barriers to patient testing.
“The consequences of these gaps are not abstract,” they wrote. “For instance, imaging and cardiac monitoring devices often excluded pregnant women from validation due to safety concerns or physiologic variability, yet these same tools may be used in obstetric emergencies where accurate diagnosis is vital.”
Most trials were conducted at a single academic center or greatly resourced health system, so it is unknown how these devices perform in other settings, according to the researchers. Because there are no standardized reporting or transparency requirements, “inequitable performance is both predictable and difficult to detect or correct,” they wrote.
“Progress will not be measured by the number of FDA-cleared algorithms, but by whether people live healthier lives, and whether those benefits will be shared equitably across all populations,” Abulibdeh and colleagues concluded.
