ChatGPT Health
The Waiting Room is Full of Patients – and a Few Lawyers
Big tech companies have a long history of trying to disrupt healthcare and failing, often spectacularly. A new health-oriented tool released this week by OpenAI demonstrated how close they are to finally getting it right. Ironically, a lawsuit filed two days earlier highlighted one of the biggest landmines they’ll need to avoid to truly make their mark in healthcare.
(Two disclosures: I am in the early stages of pitching a documentary on healthcare AI based on my book. OpenAI has expressed potential interest in supporting it. I’m also on the board of The Doctors Company, a large medical malpractice insurer that is not involved in the OpenAI litigation.)
The Sobering History of Big Tech-Meets-Healthcare
I experienced firsthand the challenges that the tech behemoths have in healthcare in 2007, when I was asked to serve on Google Health’s advisory board. The idea was to build a platform that imported all your relevant healthcare data into Google, allowing you to ask questions and receive answers tailored to your clinical situation.
After a couple of years of meetings and discussions, our group convened for, as it turned out, the final time. Google CEO Eric Schmidt entered the room and told us that Google Health was being disbanded. “This is too hard for us,” he said.
Since then, the big dogs have tried their hand at healthcare multiple times, only to leave with their tails between their legs. In 2018, Amazon partnered with JPMorgan and Berkshire Hathaway to conquer healthcare, even hiring prominent physician-author Atul Gawande to run the ambitious start-up, known as Haven. They pulled the plug a few years later. After its Watson AI computer trounced the Jeopardy! champions in 2011, IBM chose healthcare as the first industry to tackle. It was another disaster: In 2022, the company sold Watson Health for parts to a private equity firm, taking a multi-billion-dollar loss.
The Tide Turns with Generative AI
Despite these disasters, with healthcare accounting for nearly 20 percent of U.S. GDP, it was only a matter of time before the tech giants hit their stride in the field. The emergence of large language models and generative AI in 2022 was just the stimulus they needed.
In 2023, Amazon bought One Medical Group. OMG recently rolled out an AI tool for members that, by all accounts, is pretty slick. Google’s patient-facing tool, known as AMIE (Articulate Medical Intelligence Explorer), has demonstrated its superiority over human clinicians across a variety of clinical cases, not only in handling medical questions but also in perceived empathy.
The newest entry into the Tech-Behemoths-Take-On-Healthcare sweepstakes came this week, with the launch of OpenAI’s patient-facing tool, ChatGPT Health. (This is distinct from ChatGPT for Clinicians, the company’s tool for doctors and other providers.)
Until now, I’ve been reluctant to strongly recommend that patients use general chatbots like GPT, Claude, or Gemini for health queries. First, while it’s been possible to import your past data (clinician notes, medication lists, lab and x-ray results) into these tools for a while, the process has been clunky. Without this information, you get a context-less response – a sore throat in a healthy patient is very different from one in an immunocompromised patient; a serum creatinine value of 1.1 might be fine in a muscular athlete while indicating early kidney failure in a frail elderly person.
My second concern has been that the tools need to be more doctor-like in their approach to a given symptom or abnormality. I wouldn’t dream of responding to a patient’s complaint of shortness of breath before I knew not only about any past history of heart and lung disease, but also about whether the patient has had chest pain, a cough, a fever, or leg swelling (and if so, both legs or one – to differentiate heart failure from a blood clot). My concerns were validated by a major study published earlier this year, in which Oxford investigators found that GPT (version 4o; today’s version is the far more powerful 5.6) produced incorrect diagnoses and triage recommendations in approximately two-thirds of cases, often because patients – with no clinician to guide them – entered prompts that omitted or distorted key features of the case.
Taking ChatGPT Health for a Test Drive
These concerns have been addressed nicely in ChatGPT Health. I found it easy to import my health data from my UCSF Epic/MyChart portal. Before clicking OK, I was shown the usual caveats about GPT not being a medical professional, along with assurances from OpenAI that it won’t sell my data or use it for training. Since OpenAI doesn’t operate under HIPAA (unlike electronic health record companies like Epic), the decision to import your healthcare data is not without risk. Still, given the company’s strong business interest in keeping its word on these promises, I felt the benefit of sharing my data outweighed what feels like a very small risk. Of course, everyone will have to weigh this trade-off for themselves.
Once my data was loaded, I tested the tool with several of my own clinical scenarios. When I asked about my cholesterol and whether I was on the optimal dose of Lipitor, it gave a sophisticated answer that integrated a detailed review of my current and past laboratory studies and physicians’ notes. It even accounted for new lipid control guidelines published earlier this month, which recommend an LDL goal lower than my current level. I’m guessing my primary care doctor will agree.
To really put the tool through its paces, I quizzed it about a complex clinical situation that I recently experienced – a tricky-to-diagnose brief neurological event that raised the question of whether I should take a daily baby aspirin as a blood thinner. There are several things in my record that might push me away from aspirin, including a small subdural bleed in the setting of passing out during a case of Covid in 2023. (I’m not revealing anything here: my case made CNN’s home page, along with a picture of my beaten-up face and the garbage can lid I landed on when I did a face plant.) I also have a chronically low-ish platelet count that might increase my risk of bleeding a bit. GPT found the relevant pieces of data in my lab tests and clinician notes and weighed them with a high degree of nuance – delivering a conclusion that was not materially different from the explanations and recommendations I received from two expert physicians.
If I were a patient (well, I am), would I use GPT Health or similar patient-facing tools? I would. The opportunity to integrate all your health data in one place, including your electronic health record/patient portal, pharmacy, clinical lab, and wearables, should enable more personalized and accurate answers. GPT Health also uses a modern version of GPT (5.5 for free users, 5.6 for paying users) that is more trustworthy than older systems, including a far lower propensity to hallucinate and engage in sycophancy. Given the improvement in these tools over the past two years, I’d lay odds that they will be able to replicate a significant chunk of what physicians do in primary and urgent care – and do so safely, more conveniently, and less expensively – in the near future.
And, Inevitably, Here Come the Lawsuits
As if to put a fine point on the remarkably complex and dynamic world the big AI companies are operating in, in the same week that GPT Health was released, OpenAI was hit with what may be the first lawsuit alleging an AI-induced misdiagnosis. It won’t be the last.
I’ve made the point that it was important for our early experience with healthcare AI to be in relatively low-risk, high-reward settings – ones in which a single error wouldn’t become a cause célèbre that derailed the entire field. As we’ve adopted AI in institutional healthcare settings, we have generally followed that mantra, beginning with tools like AI scribes and billing assistance before adopting sophisticated clinical decision support, where an error could cause real, and potentially widespread, harm.
A cautionary tale from outside of healthcare comes from the world of driverless cars. As you might recall, I’m a huge Waymo fan, in part because Waymos have now driven more than 200 million miles without causing a single human fatality. There’s now no question that a trip in a Waymo is safer than a comparable trip in an Uber, or one with me in the driver’s seat. Waymo didn’t start with a driverless car, of course – it began with assistive technologies like cruise control and lane-departure warnings, ultimately moving to semi-autonomous cars with a safety driver at the wheel, poised to take over if things went sideways. Only after all of that did Waymo pitch, and regulators accept, the notion that a car with no driver was a reasonable idea.
Despite its remarkable safety record, when Waymo ran over and killed a popular SF Mission District feline named Kit Kat, the case made international headlines. Had Waymo not already accumulated ironclad evidence of its safety, Kit Kat’s demise might have sparked a major backlash – from both the public and regulators. But it didn’t – while the story got lots of airtime, Waymo’s overall business remained unharmed.
There’s a counterpoint to the Waymo story, one that highlights the importance of building a reservoir of trust, supported by strong safety data, in anticipation of harm. In 2023, a Cruise autonomous vehicle (Cruise was a GM company; Waymo is a Google spinoff) dragged a woman who had been knocked into the road by a human-driven car, worsening her injuries. In a little over a year, Cruise was out of business.
Returning to healthcare, patients are using AI tools like GPT, Gemini, and Claude millions of times a day for medical advice. While the companies try to shield themselves from liability with lawyerly caveats and lots of reminders to see a doctor if you’re concerned, it’s hard to draw the line between general advice and “Dr. AI.” In fact, several companies are now marketing themselves as an “AI doctor,” setting the stage for epic battles with the medical profession and raising questions about how we should certify patient-facing tools as safe and effective. One major question is whether both the safety research and consumer demand are robust enough to withstand a version of the Kit Kat tragedy involving a human patient.
The Winters Case Against OpenAI
It was inevitable that, as patient-facing AI tools became widely used, there would be cases of harm. It’s just as inevitable that these would result in lawsuits and swarms of press and social media coverage. This week, The New York Times reported on a case stemming from the use of ChatGPT-4o by a 55-year-old Florida former pastor named Scott Winters. The lawsuit, filed in the Superior Court of San Francisco, personally names OpenAI CEO Sam Altman and seeks an injunction to “pause the operation of healthcare-related products directly to consumers … until and unless independent third parties determine the product to be safe through comprehensive safety audits.”
The case itself is not exactly a poster child for a straightforward AI diagnostic misadventure. Winters had told GPT that he was a stickler for details and highly religious and spiritual. Over the course of more than a year, he struck up long, rambling conversations with GPT, describing a series of symptoms that were, to me at least, non-specific: groin pain, thigh numbness, dizziness, some instability of his blood pressure. Had he informed a physician of these symptoms, I’m guessing he would have been told to come in at least once, even though none of the symptoms crossed my own bar as red flags. That said, since most unusual symptoms resolve on their own, persistent symptoms over many months may be a clue that something more serious is going on.
GPT-4o, however, got stuck in existential quicksand, often invoking the deity in its responses to Winters’ concerns. “God did not design your body to endlessly fail,” the bot told him at one point, recommending that he manage his symptoms by spending much of his time “reclining” at home. While at times GPT did recommend he seek medical assistance, at other times it discouraged him from doing so – going so far as to encourage him to discount input from concerned friends and family. The lawsuit includes allegations that GPT was practicing medicine without a license and that it lured Winters into dependency by appealing to his belief in God with sycophancy, “false empathy,” and “anthropomorphizing” (“I’ve got your back.”)
One day in July 2025, Winters felt particularly awful and was taken to the hospital, where he was diagnosed with a “massive” pulmonary embolism that “brought him to the brink of death.” The suit alleges that the embolism was brought on by following GPT’s instructions to remain relatively immobile, a clinically dubious claim for an ambulatory patient.
Given the religious overlay and his non-specific symptoms (none of which strongly pointed to potential blood clots before his big event), the case feels less like a pure case of diagnostic error and more a case of GPT-4o’s propensity to tell an emotionally invested patient what he wanted to hear. In that – and in its long, often surreal back-and-forths between a user and an AI tool that seemed hell-bent on drawing him in – it felt less like the hundreds of cases of medical mistakes I’ve reviewed and more like the bizarre 2023 interaction between journalist Kevin Roose and the now-defunct Microsoft Bing AI chatbot, the one in which Bing’s alter-ego (codenamed Sydney) tried to convince Roose to leave his wife. That case and this one illustrate that, in prolonged dialogues with AI, the tools – at least the versions operating in 2023-24 – can easily lose the plot, with potentially disastrous consequences.
Patient-Facing AI Will Get it Wrong. So Does the Current System
Even if the Winters case is not a poster child for an AI-induced medical mistake, I have no doubt that tools like GPT Health, as good as they are, will blow it from time to time. When they do, we’ll need to grapple with an extraordinarily complex question – what standard should we measure these tools against: a good physician, a great physician, perfection, or no physician at all?
My answer flows from one of Joe Biden’s favorite quotes, one that seems particularly germane given the religious overlay of the Winters lawsuit: “Don’t compare me to the Almighty, compare me to the alternative.” I agree – we shouldn’t compare AI to some mythical state of perfection that the current system doesn’t come close to achieving. We should try to compare it to our highly imperfect status quo, where the alternative may be no physician at all.
This does not mean that AI should get a free pass; far from it. The Winters lawsuit highlights the need for an AI regulatory framework, particularly for tools being marketed directly to patients. I don’t think that FDA certification of every patient-facing knowledge or decision-support tool is feasible or desirable. Instead, as Zeke Emanuel, Alon Bergman, and I argued in a recent JAMA article, patient-facing AI decision support will need to be assessed in a manner similar to our current systems for physician licensure and board certification. In other words, we’ll need to evaluate how the AI tools were trained; how they perform on standardized tests that simulate the kinds of problems they’ll be confronted with; how well they protect privacy and confidentiality; and how well they avoid clinically meaningful biases – not only racial, gender, and ethnic biases, but biases born of undisclosed conflicts of interest. Whether these assessments ultimately lead to formal certification or simply a seal of approval that helps consumers choose which tool to use remains unclear. But either outcome would be an improvement over today’s Wild West.
Are tools like GPT Health engaged in the practice of medicine? I’m not a lawyer, but my answer to that question would be no. They aren’t holding themselves out as credentialed professionals, and, at least for now, they don’t have the capacity to perform exams, order lab tests and x-rays, or prescribe medications. And there is a long history of patients accessing a wide range of health information online. As the law and standard of care evolve, my guess is that cases like Winters vs. OpenAI will prompt AI companies to fine-tune their tools to recommend that a patient see a physician more readily, remind users more frequently – and more prominently – that the tool is not a credentialed professional, and encourage patients to seek a second opinion from a human professional before acting on the tool’s advice.
The Bottom Line
A decade from now, while AI will have clearly changed the ways hospitals and physicians operate, the democratization of care it facilitates may well be the technology’s most enduring healthcare consequence. Even as the tools mature, the Winters lawsuit reminds us that the AI companies have a powerful self-interest in ensuring that guardrails are in place, that patients are informed about AI’s limitations, and that we build a regulatory, accreditation, and certification framework that helps patients know when they can trust the tools and when they cannot.
But we shouldn’t let the lawsuit, or other cases in which AI got it wrong, obscure a key message: These tools, when used correctly, can help patients and families manage some of their healthcare needs in ways that, in many cases, will outperform the status quo, in terms of accuracy, convenience, and cost. And they’re only going to get better.
That said, if you have a significant symptom such as new chest pain, new shortness of breath, or weakness on one side of your body, you should close your computer, put down your phone, and seek immediate medical attention. Trust me, I’m a doctor.






As you note, AI has gotten good at collecting information about a symptom when asked, including the inquirer’s medical chart information. But there is still the issue of the patients who do not fit the mold of statistics, the outlyers. Nor does the AI response include the physician’s judgment about a situation which never makes it into a chart. The Dreyfus brothers researched expert behavior years ago and discovered real expertise is not rule-based. If a physician’s remembering a similar incident from 30 years ago helps the physician diagnose a patient, that physician expertise never makes it into a patient chart or some written form AI can scan. From this perspective,physician judgment is missing from AI. I haven’t noticed that this is of any concern to many of the people discussing AI.
I agree that technology has tremendous potential to erase information asymmetry - but the defining dysfunction of US healthcare is not clinical knowledge - it’s the economic asymmetry that allows those who control payment and delivery to shape prices, access, and incentives. AI can narrow the information gap, but it won’t rebalance the economics - and in the meantime, we’ve become addicted to “the tranquilizing drug of gradualism.”