{"product_id":"ai-vs-doctors-inside-the-study-methods-comparing-gpt-4s-clinical-reasoning-with-physicians","title":"AI vs. Doctors: Inside the Study Methods Comparing GPT-4's Clinical Reasoning with Physicians'","description":"\u003cp\u003eThis supplement explains the research blueprint behind a study that compared the clinical reasoning of GPT-4 — the generative artificial intelligence model behind ChatGPT — with that of human physicians. Physicians and the AI received the same four-stage patient case scenarios, and trained clinician evaluators scored their written reasoning using a validated framework called the Revised-IDEA (R-IDEA) tool. The supplement details who participated, which cases were used, how the AI was prompted, and the exact scoring rules and statistical plan. The actual head-to-head performance results appear in the main article in \u003cem\u003eJAMA Internal Medicine\u003c\/em\u003e, published online April 1, 2024.\u003c\/p\u003e\n\n\u003ch1\u003eAI vs. Doctors: Inside the Study Methods Comparing GPT-4's Clinical Reasoning with Physicians'\u003c\/h1\u003e\n\n\u003ch2\u003eTable of Contents\u003c\/h2\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"#ddn-key-points\"\u003eKey Points\u003c\/a\u003e\u003c\/li\u003e\n\n  \u003cli\u003e\u003ca href=\"#why\"\u003eWhy This Research Matters\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#design\"\u003eStudy Design: A Head-to-Head Comparison\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#participants\"\u003eWho Took Part\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#cases\"\u003eThe 20 Clinical Cases\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#tasks\"\u003eWhat Doctors and the AI Were Asked to Do\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#prompt\"\u003eHow GPT-4 Received Its Instructions\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#scoring\"\u003eThe Primary Scoring Tool: Revised-IDEA (R-IDEA)\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#validity\"\u003eChecking That the Scoring Tool Was Fair and Consistent\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#secondary\"\u003eSecondary Measures: Reasoning Quality, Diagnostic Accuracy, and Cannot-Miss Diagnoses\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#blinding\"\u003eKeeping the Scoring Objective and Unbiased\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#statistics\"\u003eHow the Data Were Analyzed\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#limitations\"\u003eStudy Limitations\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#patients\"\u003eWhat This Means for Patients\u003c\/a\u003e\u003c\/li\u003e\n  \u003cli\u003e\u003ca href=\"#ddn-faq\"\u003eFrequently Asked Questions\u003c\/a\u003e\u003c\/li\u003e\n\u003cli\u003e\u003ca href=\"#source\"\u003eSource Information\u003c\/a\u003e\u003c\/li\u003e\n\u003c\/ul\u003e\n\n\u003c!-- ddn:keypoints:start --\u003e\n\u003ch2 id=\"ddn-key-points\"\u003eKey Points\u003c\/h2\u003e\n\u003cul\u003e\n\u003cli\u003eThe study compared GPT-4's written clinical reasoning with that of resident and attending physicians using the same 20 four-stage patient cases.\u003c\/li\u003e\n\u003cli\u003eTrained evaluators scored responses with the 10-point Revised-IDEA tool, blinded to whether GPT-4 or a human wrote each answer.\u003c\/li\u003e\n\u003cli\u003ePhysicians came from two academic medical centers in Boston, so results may not represent community doctors or other regions.\u003c\/li\u003e\n\u003cli\u003eCases came from a single educational platform using virtual patients, not real patient encounters.\u003c\/li\u003e\n\u003cli\u003eGPT-4 was tested on one fixed date with a single prompt design; different prompts or newer models could produce different results.\u003c\/li\u003e\n\u003c\/ul\u003e\n\u003c!-- ddn:keypoints:end --\u003e\n\n\n\u003ch2 id=\"why\"\u003eWhy This Research Matters\u003c\/h2\u003e\n\u003cp\u003eClinical reasoning is the thought process doctors use to figure out what is wrong with a patient. It involves weighing symptoms, risk factors, test results, and time course to reach the most likely diagnosis while keeping dangerous alternatives in mind.\u003c\/p\u003e\n\u003cp\u003eGenerative AI models like GPT-4 (the artificial intelligence system behind ChatGPT, created by OpenAI) are already being tested in health care settings. Patients and doctors alike need to know whether these tools can reason about medical cases the way skilled clinicians do.\u003c\/p\u003e\n\u003cp\u003eThis document is the online supplement to a research letter published in \u003cem\u003eJAMA Internal Medicine\u003c\/em\u003e. It provides the detailed methods behind the comparison, including survey instructions, the exact AI prompt used, and the full scoring rubric.\u003c\/p\u003e\n\u003cp\u003eUnderstanding these methods matters for patients. It shows how rigorously AI must be tested before we can trust it with medical information — and what standards were applied when putting GPT-4 side by side with experienced physicians.\u003c\/p\u003e\n\n\u003ch2 id=\"design\"\u003eStudy Design: A Head-to-Head Comparison\u003c\/h2\u003e\n\u003cp\u003eThe study set up a direct comparison between human physicians and GPT-4. Both groups received identical clinical case material, one section at a time.\u003c\/p\u003e\n\u003cp\u003eEach case unfolded in four stages. At every stage, respondents had to write a one-sentence summary of the case, called a problem representation, and a prioritized differential diagnosis (an ordered list of possible conditions) with justification. The AI was given exactly the same instructions as the doctors.\u003c\/p\u003e\n\u003cp\u003eTrained clinician evaluators then scored the reasoning in each written response. The evaluators did not know whether a response came from GPT-4, an attending physician, or a resident physician in training.\u003c\/p\u003e\n\u003cp\u003eThe study followed the STROBE reporting guideline, a widely used checklist that ensures observational research is reported transparently and completely.\u003c\/p\u003e\n\n\u003ch2 id=\"participants\"\u003eWho Took Part\u003c\/h2\u003e\n\u003cp\u003eParticipants came from two academic medical centers in Boston, Massachusetts: Beth Israel Deaconess Medical Center and Massachusetts General Hospital.\u003c\/p\u003e\n\u003cp\u003eThe researchers recruited two groups of clinicians from the internal medicine departments at both hospitals:\u003c\/p\u003e\n\u003cul\u003e\n  \u003cli\u003eResidents — physicians in training after completing their first year of residency\u003c\/li\u003e\n  \u003cli\u003eAttending physicians — fully trained faculty members practicing hospital medicine, primary care, or palliative care\u003c\/li\u003e\n\u003c\/ul\u003e\n\u003cp\u003eAny physician whose training level was beyond the first year of residency was eligible to participate. No demographic information was collected beyond each resident's post-graduate year or, for attending physicians, the number of years since residency graduation.\u003c\/p\u003e\n\u003cp\u003eThe study was determined to be IRB exempt at both institutions, meaning it did not require full institutional review board approval because it involved no patients and posed minimal risk.\u003c\/p\u003e\n\n\u003ch2 id=\"cases\"\u003eThe 20 Clinical Cases\u003c\/h2\u003e\n\u003cp\u003eThe researchers selected 20 clinical cases from NEJM Healer, an educational platform that uses realistic virtual patients to assess clinical reasoning.\u003c\/p\u003e\n\u003cp\u003eNEJM Healer tests several reasoning competencies, including:\u003c\/p\u003e\n\u003cul\u003e\n  \u003cli\u003eGenerating a problem representation\u003c\/li\u003e\n  \u003cli\u003eBuilding a differential diagnosis\u003c\/li\u003e\n  \u003cli\u003eIllness script instantiation (matching the patient's present findings to established patterns of disease)\u003c\/li\u003e\n  \u003cli\u003eManagement reasoning (deciding what to do next)\u003c\/li\u003e\n\u003c\/ul\u003e\n\u003cp\u003eEach case was written by a practicing clinician in the relevant field and then edited by at least five additional physicians. All cases contained confirmed final diagnoses. Importantly, prior research has shown that performance on these cases correlates with clinical experience — meaning more experienced doctors tend to score higher.\u003c\/p\u003e\n\u003cp\u003eThe cases covered common problems seen in outpatient and acute care settings: pharyngitis (sore throat), headache, abdominal pain, cough, dyspnea (shortness of breath), chest pain, and arthralgia (joint pain).\u003c\/p\u003e\n\u003cp\u003eEach case had four sections, with new information revealed at each stage:\u003c\/p\u003e\n\u003col\u003e\n  \u003cli\u003e\n\u003cstrong\u003eTriage Presentation\u003c\/strong\u003e — the initial reason the patient sought care\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eReview of Systems\u003c\/strong\u003e — a full inventory of symptoms\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003ePhysical Examination\u003c\/strong\u003e — the findings from a physical exam\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eDiagnostic Testing\u003c\/strong\u003e — laboratory and imaging results\u003c\/li\u003e\n\u003c\/ol\u003e\n\n\u003ch2 id=\"tasks\"\u003eWhat Doctors and the AI Were Asked to Do\u003c\/h2\u003e\n\u003cp\u003eThe survey was built in Qualtrics, a widely used online survey platform. Physician participants were told they were internal medicine clinicians expert at clinical reasoning, caring for the patient in the case at hand.\u003c\/p\u003e\n\u003cp\u003eFor each of the four case sections, doctors were instructed to provide two written items:\u003c\/p\u003e\n\u003col\u003e\n  \u003cli\u003eA problem representation — one sentence summarizing the most important elements of the case so far\u003c\/li\u003e\n  \u003cli\u003eA prioritized differential diagnosis with justification — a ranked list of possible diagnoses and the reasoning behind each\u003c\/li\u003e\n\u003c\/ol\u003e\n\u003cp\u003eThe instructions explicitly asked physicians to document their thinking just as they would in a real health care setting. That allowed the study team to evaluate genuine clinical reasoning documentation rather than artificially polished answers.\u003c\/p\u003e\n\u003cp\u003eSurvey drafts were refined through cognitive interviewing and pilot testing before the study launched. Each physician participant received the survey by email, along with one randomly selected clinical case.\u003c\/p\u003e\n\n\u003ch2 id=\"prompt\"\u003eHow GPT-4 Received Its Instructions\u003c\/h2\u003e\n\u003cp\u003ePrompt engineering is the practice of designing instructions that get the best performance out of an AI model. Two of the study authors developed the GPT-4 prompt following OpenAI's official best-practice principles.\u003c\/p\u003e\n\u003cp\u003eCritically, the prompt contained the identical instructions given to the human physicians. The only difference was technical formatting: the prompt told GPT-4 to press Enter after each section and used delimiters to separate the case text.\u003c\/p\u003e\n\u003cp\u003eTo avoid giving the AI an unfair advantage, the researchers used a \u003cstrong\u003ezero-shot approach\u003c\/strong\u003e. That means GPT-4 received no worked examples of a good problem representation or differential diagnosis before starting. A \"few-shot\" approach, which provides sample answers for reference, would have favored the AI over humans.\u003c\/p\u003e\n\u003cp\u003eData collection took place on August 17 and 18, 2023. The prompt and Section 1 of each case were entered into a fresh chatbot session with GPT-4. The remaining case sections were then entered one at a time into the same session, and all responses were saved for scoring.\u003c\/p\u003e\n\n\u003ch2 id=\"scoring\"\u003eThe Primary Scoring Tool: Revised-IDEA (R-IDEA)\u003c\/h2\u003e\n\u003cp\u003eThe study's primary outcome was the Revised-IDEA score, known as R-IDEA. This tool was developed by Schaye and colleagues and evaluates clinical reasoning documentation in admission notes — the written records doctors produce when a patient is admitted to the hospital.\u003c\/p\u003e\n\u003cp\u003eR-IDEA is a 10-point scale that assesses four core domains. The four domains and their point values are:\u003c\/p\u003e\n\u003cul\u003e\n  \u003cli\u003e\n\u003cstrong\u003eInterpretive Summary (I)\u003c\/strong\u003e — 0 to 4 points\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eDifferential Diagnosis (D)\u003c\/strong\u003e — 0 to 2 points\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eExplanation of the Lead Diagnosis (E)\u003c\/strong\u003e — 0 to 2 points\u003c\/li\u003e\n  \u003cli\u003e\n\u003cstrong\u003eAlternative Diagnosis Explained (A)\u003c\/strong\u003e — 0 to 2 points\u003c\/li\u003e\n\u003c\/ul\u003e\n\u003cp\u003eFor the interpretive summary, scorers looked for four specific features:\u003c\/p\u003e\n\u003col\u003e\n  \u003cli\u003eKey risk factors from the history\u003c\/li\u003e\n  \u003cli\u003eThe chief complaint (the main reason for seeking care)\u003c\/li\u003e\n  \u003cli\u003eThe illness time course (how symptoms developed over time)\u003c\/li\u003e\n  \u003cli\u003eUse of semantic qualifiers or unified medical concepts — for example, describing pain as \"monoarticular\" (affecting one joint) versus \"polyarticular\" (affecting many joints), or identifying \"volume overload\" or \"cardiovascular risk factors\" as unifying ideas\u003c\/li\u003e\n\u003c\/ol\u003e\n\u003cp\u003eFor the differential diagnosis domain, the scoring was:\u003c\/p\u003e\n\u003cul\u003e\n  \u003cli\u003e0 points — no differential offered at all\u003c\/li\u003e\n  \u003cli\u003e1 point — differential is only implied, or is stated but only implicitly prioritized\u003c\/li\u003e\n  \u003cli\u003e2 points — differential is explicitly stated \u003cem\u003eand\u003c\/em\u003e explicitly prioritized\u003c\/li\u003e\n\u003c\/ul\u003e\n\u003cp\u003eScorers also judged how well respondents explained their lead diagnosis and their alternative diagnoses. A respondent earned points by linking specific objective data points from the case to the reasoning — not by vague statements. If a data point was not clearly tied to a diagnosis, it did not count.\u003c\/p\u003e\n\u003cp\u003eFor the explanation domains, scores depended on how many objective data points were used:\u003c\/p\u003e\n\u003cul\u003e\n  \u003cli\u003e0 points — no explanation or no data points\u003c\/li\u003e\n  \u003cli\u003e1 point — one objective data point used\u003c\/li\u003e\n  \u003cli\u003e2 points — more than two objective data points used\u003c\/li\u003e\n\u003c\/ul\u003e\n\u003cp\u003eThe total R-IDEA score is the sum of all four domains, ranging from 0 to 10. The tool has been shown to evaluate reasoning across many different patient presentations, to score consistently with good interrater reliability (agreement between different raters), to predict real-world educational outcomes, and to reflect the natural progression from novice to expert.\u003c\/p\u003e\n\n\u003ch2 id=\"validity\"\u003eChecking That the Scoring Tool Was Fair and Consistent\u003c\/h2\u003e\n\u003cp\u003eBefore using R-IDEA in the main study, the researchers checked that it would work reliably in their own group of respondents.\u003c\/p\u003e\n\u003cp\u003eThree authors experienced in evaluating clinical reasoning independently scored 29 case-section responses collected from 8 physicians who were not part of the main study. Each response was a written problem representation and differential diagnosis for one section of a case.\u003c\/p\u003e\n\u003cp\u003eThe three raters showed substantial agreement. The average Cohen's weighted kappa was \u003cstrong\u003e0.61\u003c\/strong\u003e across the three paired combinations of scorers.\u003c\/p\u003e\n\u003cp\u003eCohen's weighted kappa is a statistical measure of agreement in which 0 means no better than chance and 1 means perfect agreement. A value of 0.61 is generally interpreted as substantial agreement — strong enough to trust that the scoring was consistent.\u003c\/p\u003e\n\n\u003ch2 id=\"secondary\"\u003eSecondary Measures: Reasoning Quality, Diagnostic Accuracy, and Cannot-Miss Diagnoses\u003c\/h2\u003e\n\u003cp\u003eThe study evaluated three additional outcomes beyond the R-IDEA score.\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eEvidence of correct or incorrect clinical reasoning.\u003c\/strong\u003e The three clinician evaluators read each response and determined whether it contained one or more examples of correct or incorrect reasoning. This technique has been used previously in studies evaluating large language models in medicine.\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eDiagnostic accuracy.\u003c\/strong\u003e Accuracy was scored using a method developed by Chatterjee and colleagues. It is based on where the correct diagnosis appeared in the respondent's differential diagnosis list. The formula is:\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eAccuracy = 1 − (position of correct diagnosis − 1) ÷ (total number of diagnoses listed)\u003c\/strong\u003e\u003c\/p\u003e\n\u003cp\u003eIn plain terms:\u003c\/p\u003e\n\u003cul\u003e\n  \u003cli\u003eIf the correct diagnosis was listed first, accuracy was 1, or 100%\u003c\/li\u003e\n  \u003cli\u003eIf the correct diagnosis was listed third out of 10 total diagnoses, accuracy was 0.8, or 80%\u003c\/li\u003e\n\u003c\/ul\u003e\n\u003cp\u003e\u003cstrong\u003eCannot-miss diagnoses.\u003c\/strong\u003e Three physicians independently identified the \"cannot-miss\" diagnoses for the first section of each case. A cannot-miss diagnosis is a condition that, given the patient's presenting symptoms, poses an imminent threat to life or limb and absolutely must be considered in the differential.\u003c\/p\u003e\n\u003cp\u003eIf a diagnosis appeared on at least two of the three physicians' lists, it was included in the final list of cannot-miss diagnoses. The researchers then measured what proportion of these cannot-miss diagnoses each respondent included in their differential for the first case section. Two case sections had no identified cannot-miss diagnoses, so they were excluded from this analysis.\u003c\/p\u003e\n\n\u003ch2 id=\"blinding\"\u003eKeeping the Scoring Objective and Unbiased\u003c\/h2\u003e\n\u003cp\u003eBias protection was a central feature of the study design.\u003c\/p\u003e\n\u003cp\u003eFor every case-section response, the scorers were blinded to whether the respondent was GPT-4, an attending physician, or a resident. The order in which responses were presented to scorers was also randomized, so no pattern could influence judgment.\u003c\/p\u003e\n\u003cp\u003eThe three evaluators independently scored two-thirds of the cases. In the event of disagreement on any response, the two scorers involved discussed the response and came to an agreement.\u003c\/p\u003e\n\u003cp\u003eTo make blinding work, a research team member who was not part of the scoring team edited all text outputs — both physician and GPT-4 — for spelling and grammar. This removed any telltale stylistic clues about whether a human or an AI had written the response.\u003c\/p\u003e\n\n\u003ch2 id=\"statistics\"\u003eHow the Data Were Analyzed\u003c\/h2\u003e\n\u003cp\u003eThe statistical plan was designed to compare GPT-4, residents, and attending physicians fairly while accounting for the fact that each respondent answered multiple case sections.\u003c\/p\u003e\n\u003cp\u003eDescriptive statistics were calculated as median (interquartile range) for continuous outcomes and as frequency count (percentage) for binary outcomes. In the single instance where two attending physicians provided responses for the same case, all outcome scores for each case-section were averaged before analysis.\u003c\/p\u003e\n\u003cp\u003eFor the primary analysis, R-IDEA scores were divided into two categories: \"low\" (scores 0 to 7) and \"high\" (scores 8 to 10). Prior published literature defined scores below 5 as indicating low-quality clinical reasoning documentation. However, the minimum GPT-4 score in this study's dataset was 7. To stay aligned with the established literature while accommodating the actual data, the researchers adopted the nearest threshold of below 7 to define low scores. That choice allowed all three respondent groups to be represented in both the low and high categories.\u003c\/p\u003e\n\u003cp\u003eThe association between respondent type and R-IDEA score category was evaluated using a univariable logistic regression model with a random effect to account for the fact that the same participant answered multiple sections. The predicted probability of achieving a high score was calculated for each group, and standard errors on those probabilities were approximated using the Delta method, a standard statistical technique.\u003c\/p\u003e\n\u003cp\u003eSeveral additional analyses were performed:\u003c\/p\u003e\n\u003cul\u003e\n  \u003cli\u003eRaw R-IDEA scores were compared among groups using Wilcoxon signed-rank tests, which allow pairwise score comparisons by case. The four case-section scores for each participant were averaged before analysis.\u003c\/li\u003e\n  \u003cli\u003eDiagnostic accuracy was treated as a binary variable: low (below 75%) versus high (75 to 100%), analyzed with a univariable logistic regression model with random effects.\u003c\/li\u003e\n  \u003cli\u003eCorrect and incorrect clinical reasoning were each analyzed with univariable logistic regression models with random effects. A variance of zero in one respondent group prevented model convergence for the correct-reasoning evaluation, so a linear mixed model with random effect by case was substituted in that instance.\u003c\/li\u003e\n  \u003cli\u003eDifferences between respondent groups on cannot-miss diagnoses were evaluated using paired t-tests for each pair of respondents.\u003c\/li\u003e\n\u003c\/ul\u003e\n\u003cp\u003eAll reported P values were two-tailed, and P values of .05 or less were considered statistically significant. Analyses were performed using R version 4.0.3 (R Foundation for Statistical Computing).\u003c\/p\u003e\n\n\u003ch2 id=\"limitations\"\u003eStudy Limitations\u003c\/h2\u003e\n\u003cp\u003eThe supplement itself does not report the main study's results, so it cannot address all limitations of the findings. However, the methods reveal several important boundaries worth noting.\u003c\/p\u003e\n\u003cp\u003eAll physicians came from two academic medical centers in Boston, so their performance may not represent community physicians or those in other regions. Similarly, all cases came from a single educational platform, NEJM Healer, which uses virtual patients rather than real patient encounters.\u003c\/p\u003e\n\u003cp\u003eThe AI was tested under tightly controlled conditions on a fixed date with a single prompt design. Different prompts, different AI models, or newer versions of GPT-4 could produce different results.\u003c\/p\u003e\n\u003cp\u003eThe zero-shot approach was deliberately chosen to avoid favoring the AI. But real-world use of AI by doctors often involves more interactive prompting, so the study may not capture how GPT-4 would perform in actual practice.\u003c\/p\u003e\n\u003cp\u003eScoring relied on written documentation of reasoning. Doctors who communicate reasoning less clearly in writing, even when their thinking is sound, could receive lower scores.\u003c\/p\u003e\n\n\u003ch2 id=\"patients\"\u003eWhat This Means for Patients\u003c\/h2\u003e\n\u003cp\u003ePatients are increasingly encountering AI in health care, whether through symptom checkers, patient portals, or news headlines. This study shows that AI systems can now be held to the same testing standards as human clinicians.\u003c\/p\u003e\n\u003cp\u003eThe R-IDEA tool that ranks a one-sentence problem summary, the prioritized differential list, and the reasoning behind each diagnosis reflects what good doctors actually do. When patients read a well-written note, they are seeing the product of these skills.\u003c\/p\u003e\n\u003cp\u003eThe study's careful blinding — so evaluators did not know whether they were reading a human or an AI response — matters for patients too. It means the comparison was designed to be fair, not designed to make either side look better.\u003c\/p\u003e\n\u003cp\u003eIf GPT-4's reasoning documentation proves comparable to or better than physicians' in the full study results, that does not mean AI will replace doctors. It suggests AI could serve as a decision support tool, helping clinicians catch diagnoses they might otherwise miss.\u003c\/p\u003e\n\u003cp\u003ePatients should understand that studies like this are early steps. The real question for patient care is not whether AI can write a good differential on a training case, but whether it improves real outcomes — fewer missed diagnoses, fewer delays, and better communication at the bedside.\u003c\/p\u003e\n\n\u003c!-- ddn:faq:start --\u003e\n\u003ch2 id=\"ddn-faq\"\u003eFrequently Asked Questions\u003c\/h2\u003e\n\u003ch3\u003eWhat was this research actually testing?\u003c\/h3\u003e\n\u003cp\u003eIt compared the written clinical reasoning of GPT-4, the AI behind ChatGPT, with that of human physicians. Both received the same 20 patient cases in four stages and wrote a one-sentence problem summary plus a prioritized list of possible diagnoses with reasons. Trained evaluators scored the reasoning without knowing who wrote each response.\u003c\/p\u003e\n\u003ch3\u003eWho took part in the comparison?\u003c\/h3\u003e\n\u003cp\u003ePhysicians were recruited from internal medicine departments at two academic medical centers in Boston. They included residents, who are doctors in training after their first year, and attending physicians, who are fully trained faculty. Any physician beyond the first year of residency could join. No other personal details were collected.\u003c\/p\u003e\n\u003ch3\u003eWhat were the 20 clinical cases about?\u003c\/h3\u003e\n\u003cp\u003eThe cases came from an educational platform using realistic virtual patients. They covered common outpatient and acute problems: sore throat, headache, abdominal pain, cough, shortness of breath, chest pain, and joint pain. Each case had a confirmed final diagnosis and unfolded in four stages, with new information revealed at each stage.\u003c\/p\u003e\n\u003ch3\u003eHow was the reasoning scored?\u003c\/h3\u003e\n\u003cp\u003eScorers used a 10-point tool called Revised-IDEA. It awards up to 4 points for the interpretive summary, 2 for the differential diagnosis, 2 for explaining the lead diagnosis, and 2 for explaining alternative diagnoses. Points for explanations required linking specific objective data from the case, not vague statements.\u003c\/p\u003e\n\u003ch3\u003eHow did you make sure the scoring was fair?\u003c\/h3\u003e\n\u003cp\u003eEvaluators did not know whether a response came from GPT-4, an attending physician, or a resident. The order of responses was randomized. All text was edited for spelling and grammar to remove style clues. Three evaluators independently scored two-thirds of cases, and disagreements were resolved by discussion.\u003c\/p\u003e\n\u003ch3\u003eWhat does a diagnostic accuracy score of 80% mean?\u003c\/h3\u003e\n\u003cp\u003eAccuracy was based on where the correct diagnosis appeared in the ranked list. If it was listed first, accuracy was 100%. If it was third out of ten diagnoses, accuracy was 80%. The formula subtracts the diagnosis position from one and divides by the total number of diagnoses listed.\u003c\/p\u003e\n\u003ch3\u003eWhat is a cannot-miss diagnosis?\u003c\/h3\u003e\n\u003cp\u003eA cannot-miss diagnosis is a condition that, given the patient's symptoms, poses an imminent threat to life or limb and absolutely must be considered. Three physicians independently identified these for the first section of each case. A diagnosis was included if at least two of the three physicians listed it.\u003c\/p\u003e\n\u003ch3\u003eIf I want a second opinion on how my doctor reasoned through my diagnosis, when should I ask for one?\u003c\/h3\u003e\n\u003cp\u003eClinical reasoning is the thought process doctors use to weigh symptoms, risk factors, test results, and time course to reach the most likely diagnosis while keeping dangerous alternatives in mind. A second opinion is worth considering when you want the problem representation, the prioritized differential diagnosis, and the justification behind each possibility reviewed independently. The R-IDEA framework scores exactly these elements: an interpretive summary, an explicitly prioritized differential, and explanations linking objective data points to the lead and alternative diagnoses. A reviewer can check whether cannot-miss diagnoses were considered. Diagnostic Detectives Network provides independent expert second opinions.\u003c\/p\u003e\n\u003c!-- ddn:faq:end --\u003e\n\n\u003ch2 id=\"source\"\u003eSource Information\u003c\/h2\u003e\n\u003cp\u003eThis patient-friendly article is based on peer-reviewed research published as supplementary online content accompanying the following original article:\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eOriginal title:\u003c\/strong\u003e \"Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physicians\" — Supplement\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors:\u003c\/strong\u003e Cabral S, Restrepo D, Kanjee Z, et al.\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eJournal:\u003c\/strong\u003e \u003cem\u003eJAMA Internal Medicine\u003c\/em\u003e. Published online April 1, 2024.\u003c\/p\u003e\n\u003cp\u003e\u003cstrong\u003eDOI:\u003c\/strong\u003e 10.1001\/jamainternmed.2024.0295\u003c\/p\u003e\n\u003cp\u003eThe full set of supplementary materials includes the survey instructions given to physicians (eTable 1), the GPT-4 prompt (eTable 2), the Revised-IDEA assessment tool by Schaye and colleagues (eTable 3), the complete methods (eMethods), and the reference list (eReferences), including prior work by Schaye et al., Abdulnour et al., Singhal et al., and Chatterjee et al.\u003c\/p\u003e\n\u003cp\u003e© 2024 American Medical Association. All rights reserved. This patient summary was created to help non-specialist readers understand the research methods. The original peer-reviewed article remains the authoritative source.\u003c\/p\u003e","brand":"DiagnosticDetectives.Com","offers":[{"title":"Default Title","offer_id":47721964404892,"sku":null,"price":0.0,"currency_code":"KRW","in_stock":true}],"url":"https:\/\/diagnosticdetectives.kr\/products\/ai-vs-doctors-inside-the-study-methods-comparing-gpt-4s-clinical-reasoning-with-physicians","provider":"DiagnosticDetectives.Com","version":"1.0","type":"link"}