September 03, 2026

7 min read

Key takeaways:

  • Datasets used to train AI can have biased and inaccurate data, which can limit its utility.
  • Validation and human oversight remain critical for AI to be used effectively.

The internet went mainstream less than 40 years ago, yet it is hard to remember the world without it.

Chadi Nabhan, MD, MBA, hematologist, medical oncologist, podcaster and author, believes the same will hold true for AI in the not-so-distant future.



Quote from Chadi Nabhan, MD, MBA



“There’s no way around it. AI is here to stay,” Nabhan told Healio. “Humans who are smart in using AI will advance better and faster in life and in career than humans who do not embrace AI.”

The same holds true in healthcare. Clinicians and institutions that best utilize AI will likely have the greatest impact on patients.

Barriers remain to unlocking AI’s full potential, though.

AI models continue to be trained on biased and inaccurate data, which limits practical application to real-world patients, Nabhan and colleagues discussed at the third Collaboration for Outcomes Using Social Media in Oncology (COSMO) Conference in Chicago.

“If you always believe everything that you’re being told, you’re going to make mistakes,” Nabhan said.

Healio spoke with Nabhan about AI bias and integrity, how it impacts equity and ways to improve it in the future.

Healio: What is AI bias?

Nabhan: The data that are being used in the algorithm to create the output representative of patients with cancer being seen in the real world.

Most oncology datasets, especially datasets being used to train AI algorithms, are not random samples. Most of the datasets represent patients from large academic centers, patients who have undergone genomic testing, possibly patients who are insured, who had access to healthcare. If there were patients on clinical trials, they were healthy enough to get on clinical trials, and patients who have more longitudinal follow-ups.

If we are training AI models on these data, the output may not be representative of all patients with cancer who are seen across community clinics in America.

Let’s say you have a model about managing prostate cancer, but that model was developed based on white men who are 75 years of age, can you realistically apply that model on a Black man who is 45 years old who has the same disease? I argue that you can’t.

Healio: What is the solution?

Nabhan: Before we deploy the model, before we say that this is a great model and could be utilized in making decisions, we should validate it.

There’s a training cohort and a validation cohort. You characterize the model using this training population from the datasets that you have, but then you should validate it, and you should validate it in a cohort that represents patients from the real world. These are patients who are in community practices, could be seen in rural settings. They are different racial and ethnic groups. While many datasets include predominantly Caucasian patients, the U.S. Census suggests that 25% of the U.S. population is underrepresented minorities. That’s very important to understand and take into context.

Healio: What is AI integrity?

Nabhan: Integrity asks the question of whether the data are correct.

Imagine there is a patient in the electronic health records who underwent next-generation sequencing (NGS), and that patient is biomarker negative. That’s patient A. In contrast, we have patient B who has never undergone NGS. AI could view these two patients as identically not having an actionable mutation. While this may appear true on the surface, patient A did have NGS, and was biomarker negative, but patient B never had the test done to begin with. The AI reported no actionable mutation on patient B because nothing was recorded in the EHR given the lack of testing . AI can’t really understand that difference. They can only read the fact that these two patients don’t have biomarkers.

Absence of evidence in the medical record is not evidence of absence.

Healio: How can that problem be solved?

Nabhan: AI systems need to be trained to differentiate between unknown vs. negative.

Let’s say a clinical trial is looking for a patient that is biomarker negative. To decide on eligibility, the AI should be able to tell you if the eligibility cannot be determined because required molecular testing has not been performed vs. the patient is not eligible because the patient is indeed biomarker negative. That nuance is important and it does play a role in equity.

Healio: Why would EHRs have inaccurate data?

Nabhan: EHRs are messy. We use a lot of copy, paste and forward. We have a lot of outdated problem lists. A lot of physicians don’t update the problem list on the patient every single time.

I had a patient, who, when they were reviewing their records, noticed that I still said they were a smoker, and they were upset with me. They said, “Doc, you know I quit 5 years ago.” I apologized, but the reality is, for a lot of us in a very busy clinic, we don’t always update this information, so it’s missing. Sometimes there is incorrect staging. Sometimes you have incomplete treatment histories.

Imagine a patient had a particular test done at a different hospital and it’s not linked to the EHR at the current hospital. You don’t have that information integrated into the existing records. Sometimes genomic records are integrated in the EHR. Sometimes they’re coming through PDFs, and the nurse or the doctor forgot to scan it.

If you have AI models trained on EHRs, you are held hostage based on the data that exist in the EHRs even if some were inaccurate. AI doesn’t necessarily correct bad data.

Healio: How can data accuracy be assured?

Nabhan: That’s the million-dollar question. The short answer, unfortunately, it can’t. It’s dependent on what the human enters, unless you have somebody else checking the work.

What AI could be trained to do is detect inconsistencies. For example, if in my note I write that the patient quit smoking and the EHR still shows that he is a smoker, maybe I get something on the screen saying, there’s inconsistency here, please double check.

Healio: How can AI help with health equity?

Nabhan: Expertise in cancer care can be geography dependent. If you live in Chicago, New York, or a large metropolitan area, your access to expertise is much easier than if you live in rural America, where you sometimes have one doctor for 100,000 people.

AI could solve this, because it provides physicians access to state-of-the-art information through large language models.

Healio: Will AI help or hinder equity in healthcare?

Nabhan: It depends how it’s used. Every single technology we have has pluses and minuses. The internet has pluses and minuses.

On balance though, I believe that AI is net positive despite some of the issues and shortcomings we highlighted. Nothing is perfect. If we’re able to refine the algorithms, make sure that the data are getting better with time and represent people that we are seeing in the real world, the output could be very helpful.

If you’ve seen a difficult patient, you often have called your colleague for a consultation. Now, you call ChatGPT for a consultation. These large language models search lots of datasets, databases and so on, and bring you back various opinions that hopefully represent your patient. You are still in charge, but I think it’s going to help you provide better care.

I am an AI enthusiast. I believe it’s helping patients, families, caregivers and physicians.

Healio: How can AI shortcomings be mitigated?

Nabhan: No. 1, human in the middle is very important. This is an ancillary service. Ultimately, the human must interpret the information.

No. 2 is validation. If the data being used to create the model do not represent the people that you are treating in the real world, that output is not going to be very helpful. If it’s going to be implemented to take care of patients, into decision-making, it must be validated.

Healio: How should clinicians view AI long term?

Nabhan: Cautiously optimistic. AI is going to continue to be part of everything that we do. If we were in the year 1999 and you told me, I don’t want to use the internet, I hate the internet, it’s sloppy. Yes, it was sloppy, but how could you live without the internet? It’s the same thing with AI.

Our goal is to train patients, students and fellows to understand AI and utilize it smartly and effectively and be cognizant of its shortcomings. Let’s work together on mitigating some of these issues, because they are going to be mitigated with time.

I think these issues will be addressed in 5 years, I really do. The validation models are already happening. I’m also very optimistic about how AI is going to change the way we do clinical trials.

Healio: How can AI help with clinical trials?

Nabhan: Clinical trials take a long time. Most could take 10 to 12 years to complete, and they are costly. That length could delay patients’ access to [life-saving drugs]. AI could help with that. It could help with clinical trial matching. For example, if you have a clinical trial at your center and you have a patient that you are seeing in clinic that matches available clinical trials at your institution, it could be flagged.

The other way for AI to help with clinical trials is in selecting the proper research sites. Many clinical trials open at sites that are not going to enroll on this particular study. If we’re able to use AI to help select the most appropriate sites for a specific trial, this will expedite that trial.

Healio: Is there anything else you would like to emphasize?

Nabhan: In my book, “AI and Cancer Care,” that is coming out on Dec. 1, I surveyed and spoke to many people from across all sectors; many were not doctors. People at the gym, people at Starbucks, people who are at Walgreens, and so on. I asked them who do you think should govern AI? Nobody wanted the government to govern AI, because they were nervous that the government would be too slow and too bureaucratic to govern this very innovative technology. At the same time, nobody wanted the private sector to govern AI because of the conflict of interest.

On the COSMO panel, we really did not have a solution, but I proposed a combination between the private and the public sector, where you have both sides coordinating the governance and have a lot of guardrails to make sure that this is not getting out of control.

For more information:

Chadi Nabhan, MD, MBA, can be reached at cnabhan1968@gmail.com.



Source link