The Big Three
ChatGPT, Claude, and Gemini—your guide to today's leading foundation models and how to choose between them.
Which foundation model should you use—and does it actually matter?
Starting Here? Read This First
If you've jumped straight to this module hoping to pick a model and get started, you're in good company—this is exactly what most people want. But foundation models are tools, and like any tool, their value depends on how skillfully you use them. Before you dive too deep here, consider at least skimming these foundational concepts from earlier modules:
From Module 1 (How LLMs Think): These models work by predicting the most likely next word in a sequence, trained on enormous datasets of human-generated text. They don't "know" things the way you do—they recognize patterns. This matters because it explains both their remarkable capabilities and their characteristic failures.
From Module 2 (PHI and HIPAA): None of these consumer-facing chat interfaces are HIPAA-compliant out of the box. We'll discuss BAA pathways later in this module, but the critical principle remains: never enter PHI into a consumer AI product without proper safeguards.
From Module 3 (Prompting): The quality of your output depends enormously on the quality of your input. A vague prompt produces vague results. A well-structured prompt with context, role, and constraints produces dramatically better outputs. We'll reference prompting principles throughout this module—they apply equally to all three platforms.
Now, let's meet the models.[1]
The Med Student Analogy
Think of each foundation model as a brilliant medical student who has read essentially everything ever published—every textbook, every journal article, every clinical guideline, every case study, and frankly, every Reddit thread and random blog post too. This student has near-perfect recall of patterns across all that material and can synthesize information across domains in ways that would take you hours or days.
But here's what's crucial: this med student has never actually seen a patient. They haven't felt the resistance of tissue, watched a parent's face crumple at difficult news, or learned from the case that didn't follow the textbook. They know what clinical reasoning looks like on paper, but they don't have clinical judgment.
This framing helps calibrate expectations:
- They're genuinely useful for the kind of work where pattern recognition and information synthesis matter—drafting notes, summarizing literature, explaining concepts, generating differential diagnoses for discussion.
- They require supervision the same way any trainee does. You wouldn't let even a brilliant third-year student sign notes unsupervised, and you shouldn't let an AI do so either.
- They need clear instructions. Just as you'd give a student specific guidance ("I need a note that addresses the parent's concern about developmental delay, focuses on what we observed today, and includes our reasoning for not ordering imaging"), you need to give the model explicit context and constraints.
- They sometimes confabulate. A student who doesn't know something might make up an answer that sounds plausible rather than saying "I don't know." These models do the same—we call it "hallucination," but it's really just pattern-matching in the absence of actual knowledge. This is why you verify before you trust.
With that framing in mind, let's look at who these three "students" are and what each brings to your practice.
Knowing Your Options
Before diving into each model, it's worth understanding the landscape:
- ChatGPT is the household name. For most people, "ChatGPT" is synonymous with AI chat. It has the largest consumer user base and the most cultural awareness. When your patients or colleagues mention "AI," they're usually thinking of ChatGPT.
- Many people don't realize they already have Gemini. If you have a Google account, you have access to Gemini. It's built into Google Search, available at gemini.google.com, and integrated throughout Google Workspace. Yet many users have never tried it.
- Claude is often discovered second. People typically find Claude when looking for alternatives, exploring options for specific use cases, or hearing recommendations from colleagues. It has a smaller but dedicated user base.
None of this tells you which is "best"—that depends entirely on your needs, your ecosystem, and your preferences. The point is simply: you have options, and it's worth knowing what they are.
ChatGPT: The First Mover
The Story
OpenAI was founded in December 2015 by Sam Altman, Elon Musk, and others with the stated mission of developing artificial general intelligence that benefits humanity. The company initially operated as a nonprofit research lab before restructuring in 2019 to attract the investment needed for increasingly expensive AI training.
The GPT (Generative Pre-trained Transformer) architecture emerged from this research, with GPT-1 in 2018, GPT-2 in 2019 (initially withheld due to concerns about misuse), and GPT-3 in 2020. But the inflection point came on November 30, 2022, when OpenAI released ChatGPT as a free research preview. Within five days, it had a million users. Within two months, it reached 100 million—the fastest-growing consumer application in history.
That explosive growth fundamentally changed how the world understood AI. Suddenly, anyone could have a conversation with a system that felt like talking to a knowledgeable colleague. The technology wasn't new, but the accessibility was.
The Current Offering
As of early 2026, ChatGPT operates across several tiers:
ChatGPT Free
$0
Access to current models with usage limits. Good for exploration and occasional use.
ChatGPT Plus
$20/month
Higher limits, priority access, GPT-5.5 reasoning, image generation, advanced voice mode.
ChatGPT Pro
$200/month
For power users. GPT-5.5 Pro, extended features, essentially unlimited usage.
ChatGPT Team
$25-30/user/month
Collaborative workspace. Data not used for training. Still not HIPAA-compliant without additional measures.
ChatGPT for Clinicians
Free (verified)
Free tier for verified US physicians, NPs, PAs, and pharmacists. Tuned for documentation and medical research. Not a BAA tier on its own.
ChatGPT for Healthcare
Enterprise
Hospital/health-system tier with BAA, customer-managed encryption keys, audit logs, and data residency. Deployed at UCSF, Cedars-Sinai, MSK, Stanford Children's, HCA, Boston Children's, and others.
OpenAI launched two distinct healthcare offerings in 2026 that are easy to confuse. ChatGPT for Clinicians is a free, individual product for verified US clinicians—handy for documentation drafts, summarization, and medical research, but it is not covered by a BAA and should not be used with PHI. ChatGPT for Healthcare is the enterprise product that does come with a BAA, alongside HIPAA-compliant controls (encryption keys, audit logs, data residency) and a "no training on shared content" guarantee. It's negotiated and deployed at the institution level, not by individual users.
Practical implication: don't assume that signing in with your hospital email upgrades your account. Check with your organization whether a BAA-backed deployment is in place before entering anything that could be PHI.
What You Get
Ecosystem: The most mature AI ecosystem. The GPT Store contains customized applications for specific use cases, extensive plugin support, and integrations with tools many people already use. If you want to find a pre-built solution for a specific task, ChatGPT's ecosystem is the most likely place to find it.
Features: Advanced Voice Mode for natural conversation, image generation through DALL-E, Code Interpreter for data analysis and visualization, and web browsing for current information.
HIPAA pathway: BAAs available for API services, Enterprise/Edu plans, and the new ChatGPT for Healthcare tier deployed at major academic medical centers. ChatGPT Free, Plus, Pro, Team, and the free ChatGPT for Clinicians tier are explicitly not covered by BAAs and cannot be used with protected health information.
Claude: The Safety-First Approach
The Story
Anthropic was founded in 2021 by Dario and Daniela Amodei, along with several other former OpenAI researchers and executives. The founding team included key figures in AI safety research, and that orientation shaped the company's approach from the beginning.
The company developed "Constitutional AI," a training approach that uses AI feedback (rather than exclusively human feedback) to shape model behavior according to a set of principles. The goal was to create systems that are helpful but also harmless and honest—what Anthropic describes as the "HHH" framework.
Claude 1.0 launched in March 2023, positioning itself as a thoughtful alternative to
ChatGPT. The Claude 3 family arrived in March 2024, introducing the Haiku/Sonnet/Opus
tiering (small/medium/large models with different capability and cost profiles). In
February 2026, Claude Opus 4.6 launched with a one-million-token context window and
"agent teams" that can coordinate across shared codebases. Anthropic closed a record
$30 billion funding round, valuing the company at approximately $380 billion. Later
that month, Claude Sonnet 4.6 arrived—delivering near-Opus performance at one-fifth
the cost, and becoming the default model for both free and Pro users. In April 2026,
Claude Opus 4.7 shipped with materially stronger coding (+13% on
SWE-bench), higher-resolution vision, the new /ultrareview multi-agent
review workflow, and an xhigh reasoning effort level—all at the same
pricing as 4.6. Claude Opus 4.8 followed on May 28, 2026—again at
the same pricing—with 128K max output tokens, 88.6% on SWE-Bench Verified, a new
Fast mode ($10/$50 per million tokens at ~2.5× speed), and
Dynamic Workflows in Claude Code, which can orchestrate up to 1,000
subagents on a single task. On June 30, 2026,
Claude Sonnet 5 arrived—a one-million-token context window,
128K max output, and the default model both in Claude Code from launch and for Free
and Pro users on claude.ai. On July 24, 2026,
Claude Opus 5 replaced Opus 4.8 as the current Opus, at the same
$5/$25 per million tokens: a one-million-token context window, 128K max output, and
a knowledge cutoff of May 2026. It is now the default model on Claude Max and the
strongest model available on Pro. Opus 4.8 still works but has moved to Anthropic’s
legacy list.
The Current Offering
Claude Free
$0
Claude Sonnet 5 with usage limits that reset every few hours. Good for exploration.
Claude Pro
$20/month
~5x usage of the free tier. Sonnet 5 by default, with Opus 5—the strongest model on this tier—also available. Priority access to new features.
Claude Max
$100-200/month
5x-20x Pro limits, with Claude Opus 5 as the default model. Designed for power users, especially Claude Code developers.
Claude Team
$25-30/user/month
Collaborative features, admin controls. Data not used for training. Min 5 seats.
What You Get
Ecosystem: Smaller than ChatGPT's but growing. The MCP (Model Context Protocol) allows integrations with external tools and data sources. Claude Code provides command-line AI assistance for developers. The ecosystem emphasizes depth over breadth.
Features: A one-million-token context window across the current lineup—Opus 5, Sonnet 5, and the top-end Fable 5—for working with substantial documents. Projects feature for organizing related conversations. Artifacts for generating code, documents, and other outputs. Computer use capabilities for agentic workflows. No native image generation, but can analyze images you provide.
HIPAA pathway: BAAs available through the API with Zero Data Retention (ZDR) agreement. Consumer chat interface (Claude.ai) is not covered. AWS Bedrock provides the most straightforward enterprise pathway.
Gemini: The Integration Play
The Story
Google's path to Gemini began with earlier AI efforts including LaMDA (the model behind the original Bard chatbot) and PaLM (Pathways Language Model). But Google's extensive AI research, which predates the current generation by decades, positioned the company as a natural major player once the race began.
Gemini was announced in December 2023 as Google's response to GPT-4, with the company emphasizing its native multimodal training—the model was trained from the beginning on text, images, and other modalities together, rather than having capabilities bolted on afterward.
Gemini 1.5 arrived in February 2024 with a groundbreaking 1-million-token context window (later expanded to 2 million)—dramatically larger than competitors at the time. In November 2025, Google announced Gemini 3, positioned as "our most intelligent model" with significant improvements in reasoning, multimodal understanding, and agent capabilities. In February 2026, Gemini 3.1 Pro doubled the reasoning performance of its predecessor and dominated 13 of 16 major AI benchmarks, now powering NotebookLM. At Google I/O 2026 (May 19), Google launched Gemini 3.5 Flash—rolled out to all users and roughly 4× faster than peer frontier models on agentic benchmarks—alongside Gemini Omni Flash, a new multimodal model for Plus/Pro/Ultra tiers. Gemini 3.5 Pro has still not shipped. As of late August 2026 its release has slipped more than once, and Google’s releases since—Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Flash Cyber on July 21, then Gemini 3.7 Flash on August 13—still did not include a Pro model. Every specification circulating for 3.5 Pro is unconfirmed.
The Current Offering
Gemini Free
$0
Access to Gemini 3.5 Flash, basic features, limited Deep Research access.
Google AI Pro
$19.99/month
Gemini 3.1 Pro, 1M-token context, Gemini Omni Flash, Deep Research, Workspace integration, 2TB storage.
Google AI Ultra
$249.99/month
Gemini 3.5 Flash + Gemini Omni Flash, Deep Think reasoning, highest limits, 30TB storage, Veo 3 video generation (Gemini 3.5 Pro has not shipped as of late August 2026).
What You Get
Ecosystem: Deep Google Workspace integration. If your organization lives in Gmail, Docs, Sheets, and Meet, Gemini works natively within those tools—the side panel in Google Docs helps you write, Gemini in Gmail drafts responses and summarizes threads. For organizations already committed to Google, this reduces friction dramatically. If you don't use Google Workspace, much of this advantage disappears.
Features: 1-2 million token context window—the largest available—for working with enormous documents or entire codebases. Strong multimodal capabilities including video analysis. Image generation through Imagen. Gemini Live for voice interaction.
HIPAA pathway: Google Workspace with Gemini is explicitly HIPAA-eligible. Google's HIPAA Included Functionality list covers Gemini in Workspace. Organizations that sign Google's Business Associate Amendment through the Admin Console can use these services with protected health information. This is currently the most straightforward consumer-tier HIPAA pathway.
Side-by-Side Comparison
| Factor | ChatGPT | Claude | Gemini |
|---|---|---|---|
| Developer | OpenAI (Microsoft-backed) | Anthropic (Amazon/Google-backed) | Google DeepMind |
| Consumer Price Entry | $20/month (Plus) | $20/month (Pro) | $19.99/month (AI Pro) |
| Premium Tier | $200/month (Pro) | $200/month (Max 20x) | $249.99/month (Ultra) |
| Context Window | 1M tokens (GPT-5.5) | 1M tokens (Opus 5, Sonnet 5) | 1-2M tokens |
| Native Image Gen | Yes (DALL-E) | No | Yes (Imagen) |
| Voice Mode | Advanced Voice Mode | Limited | Gemini Live |
| Ecosystem | GPT Store, plugins | MCP integrations | Google Workspace |
HIPAA/BAA Comparison
| Platform | Consumer Chat BAA | API BAA | Easiest Path |
|---|---|---|---|
| ChatGPT | Enterprise/Edu and ChatGPT for Healthcare | Yes, via application | Azure OpenAI Service or ChatGPT for Healthcare |
| Claude | No | Yes, with ZDR agreement | AWS Bedrock |
| Gemini | Workspace integration covered | Yes via Vertex AI | Google Workspace + BAA |
For the consumer chat interfaces most people use day-to-day (ChatGPT Plus, Claude Pro, standard Gemini), none are HIPAA-compliant and should never be used with PHI without additional safeguards.
Beyond the Big Three
This page covers three platforms because they are the ones with mature clinical-privacy pathways and the longest safety track records. But two things outside them now generate more hallway and exam-room questions than anything in the tables above: Grok, and the Chinese AI labs. Neither changes the practical recommendation on this page. Both are worth understanding—for different reasons.
Grok: The AI Your Patients Are Using
Grok is the model built by xAI—folded into SpaceX in February 2026—and it lives where its users live: inside X, in a standalone app, and behind an API. The current flagship is Grok 4.6 (August 12, 2026). Notably, xAI has no healthcare product and no medical vertical; its named enterprise markets are customer support, legal, security, and government. The medical story is entirely organic—and promoted from the very top. Elon Musk has repeatedly told X’s users to upload their medical imaging: “you can upload your X-rays or MRI images to Grok and it will give you a medical diagnosis” in late 2024, and again in February 2026—“just take a picture of your medical data or upload the file to get a second opinion from Grok.” [Forbes]
Two facts to hold at once. First: the popularity is real, but it is your patients, not your colleagues. Roughly one in six US adults now asks a chatbot health questions at least monthly—yet no physician survey measures any Grok adoption at all. The AMA’s 2026 survey found 81% of physicians using AI professionally and named exactly one tool (OpenEvidence); Doximity’s and Sermo’s physician surveys never mention Grok. Second: the capability picture whiplashes with the task. On a public USMLE Step 1 question set, Grok scored best of the five models tested (91.6%)—the study its fans cite. On fifty expert-level radiology spot diagnoses—the exact use being promoted—Grok-4 answered 12% correctly against radiologists’ 83%, near the bottom of the frontier pack. [Preprint] Physician tests of real uploaded scans have found it mistaking tuberculosis for a herniated disk. [Fortune]
The PHI posture is the part to be able to recite. Consumer Grok trains on your conversations by default (opt-out in settings), and xAI’s own privacy policy asks users not to submit health information at all. In August 2025, roughly 370,000 shared Grok conversations—medical questions included—were indexed by Google because share links carried no search-engine protection; shareable links remain a promoted feature. Asked directly on X, Grok itself answered that it is “not a medical professional or HIPAA compliant.” A BAA path does exist—but it is API-only, by application, and requires xAI’s zero-data-retention tier. It does not cover the Grok app, grok.com, or Grok inside X, which are precisely the surfaces patients are being pointed at. So the takeaway matches When Patients Bring AI to the Exam Room: expect Grok readings of labs and imaging to walk in the door, engage with them the way you would any patient research, and keep PHI out of it on your side of the encounter.
The Chinese Labs: Convergence, Price, and Open Weights
The other question—should I care about DeepSeek, Kimi, GLM, Qwen?—got louder in July 2026, when Moonshot’s Kimi K3 release knocked roughly a percent off the Nasdaq in a day and had commentators declaring a second DeepSeek moment. [Fortune] The convergence is real: on Artificial Analysis’s independent index, K3 scores 60 against Claude Opus 5’s 63, Fable 5’s 62, and GPT-5.6 Sol’s 61—the closest an open-weight model has ever sat to the closed frontier—and in nearly every month of 2026 the best open-weight model in the world has been Chinese. [Hugging Face] “China has taken the lead” overstates it for the closed frontier; for open weights it is simply true. Look at who is actually paying, though, and the adoption is price, not preference: on one major developer gateway, DeepSeek’s share of tokens jumped from under 1% to 17% in a single month while its share of revenue stayed near 1%—teams route cheap, low-stakes work to it and keep the frontier models for what matters. [Rest of World]
For clinical use, six facts do most of the work:
- The hosted APIs send your data to China. DeepSeek’s own privacy policy says it plainly: data is collected, processed, and stored “in People’s Republic of China.” No Chinese lab offers a BAA. Patient information into any of these apps or APIs is an impermissible disclosure, full stop. [Policy]
- The model is not the API. The same open weights are served US-hosted by American clouds—Amazon Bedrock, Azure AI Foundry, Google Vertex—where the BAA is with the cloud provider and data stays in the US cloud. A platform BAA makes a service HIPAA-eligible, not a deployment compliant, and model catalogs churn—verify the specific pairing before building on one.
- Open does not mean runnable. K3 is 2.8 trillion parameters of multi-node datacenter hardware. The realistic self-host tier is Qwen3.8-27B (plain Apache 2.0) or DeepSeek V4-Flash—see Running AI Models Locally.
- Open no longer reliably means MIT. Kimi K3 and Qwen’s flagship carry bespoke licenses with revenue-triggered terms; GLM-5.2 is genuinely MIT. Read the license that shipped.
- The medical evidence is China-centric and a generation old. DeepSeek-R1 scored 96% on the Chinese National Medical Licensing Exam—in Chinese, against GPT-4o’s 75%—a language-and-training asymmetry that cuts both ways. [JMIR Med Educ] No physician-rubric evaluation of the current generation exists in either direction.
- No enacted rule restricts your use of them. The US bans are public-sector device and procurement rules (Navy, NASA, Commerce bureaus; Texas, New York, Virginia state devices). Zhipu’s Entity List placement restricts US exports to Zhipu—it does not prohibit Americans from using GLM models. Bills that would go further have not passed.
And the adoption reality-check: no documented US health system runs any of these models, while DeepSeek was deployed locally in roughly 90 Chinese tertiary hospitals by mid-2025. American clinicians are reading about these models, not practicing with them. The reasons to watch anyway are the two that compound: price pressure on the frontier vendors you do use, and a self-host trajectory that makes on-prem clinical AI more plausible every quarter. [Rest of World]
Just Pick One and Start
Here's the honest truth: for most clinical use cases, all three models are capable enough. The differences between them matter at the margins—and those margins shift with every model update anyway.
Don't overthink the choice. Pick based on:
- What you already have access to. Already in Google Workspace? Try Gemini. Have a ChatGPT account from when everyone was talking about it? Use that.
- What your colleagues use. There's value in being able to share prompts and tips with people doing similar work.
- Whichever interface you prefer. Seriously—if one feels better to use, that matters.
You can always switch later. You can use multiple models for different tasks. The skill you develop—learning to prompt effectively, knowing when to trust outputs, building useful workflows—transfers across all of them.
The next section gives you a practical plan for getting started. After that, we'll cover how to evaluate models with your own use cases over time.
Practical Guidance: Your First Month
Week 1: Establish a Baseline
Choose whichever model you have easiest access to. For your first week, use it for low-stakes tasks:
- Summarizing an article you're reading anyway
- Drafting an email you'll heavily edit
- Explaining a concept to yourself before explaining it to a patient
- Brainstorming questions for a meeting
Don't worry about optimization. Just get comfortable with the interaction pattern.
Week 2: Apply Prompting Principles
Revisit the prompting framework from Module 3 and apply it deliberately:
- Give the model a role ("You are helping me prepare for a difficult conversation with a parent")
- Provide context ("The patient is 8 years old with a new ADHD diagnosis")
- Be specific about format ("I need three main points, in language a non-medical parent would understand")
- Include constraints ("Avoid medical jargon; emphasize that this is manageable")
Notice how the quality of outputs changes as you prompt more skillfully.
Week 3: Try Something Harder
Push into a task that actually matters:
- Draft a real (de-identified) clinical note
- Summarize a complex patient case for a referral
- Create patient education materials for a condition you see frequently
- Analyze a clinical guideline and identify key practice implications
Evaluate the output critically. What did it get right? Where did it need correction? What would you prompt differently next time?
Week 4: Compare
Now try a second model with a task you've done before. Use the same prompt and compare outputs. You'll develop intuition for the differences—and often find that your preference is less about the model and more about how you've learned to work with it.
Evaluate With Your Own Use Cases
Here's a truth that benchmark tables and feature comparisons can't capture: the only evaluation that matters is how a model performs on your actual work.
Build Your Personal Test Set
Create a small set of 3-5 tasks that represent your real work:
- A type of document you frequently draft (note, letter, summary)
- A question you commonly need to research or explain
- A complex case or scenario you've worked through before
- Something where you know what "good" looks like
Run these same prompts through different models. Compare the outputs. Which required less editing? Which understood your intent better? Which produced something you'd actually use?
Re-Evaluate After Updates
This is crucial and often overlooked: a model that didn't work for you six months ago might be excellent now. And vice versa—a model you loved might change in ways that don't suit your workflow.
Each major model update (GPT-4 to GPT-4o, Claude 3 to Claude 4, etc.) can significantly change how the model handles specific tasks. Some examples:
- A clinical note format that one model version struggled with might work perfectly in the next
- A type of analysis that frustrated you might work smoothly after an upgrade
- Conversely, a workflow you'd perfected might break when the model changes
When you see announcements about major model updates, revisit your test set. Don't assume your current choice is still the best choice—or that a model you dismissed is still inadequate.
Set a calendar reminder every 3-6 months to run your personal test set across the current versions of each model. This takes 30 minutes and ensures you're always using the best tool for your needs—not just the one you happened to start with.
The Cost Question
Let's be direct about money.
Free tiers are sufficient for exploration and light use. If you're using AI occasionally—a few times a week for non-critical tasks—you may never need to pay.
$20/month is the standard paid tier. ChatGPT Plus, Claude Pro, and Google AI Pro all cluster around this price point. At this tier, you get higher usage limits, access to the best models, and priority features. For a professional tool you might use daily, $20/month is modest—less than many software subscriptions with narrower utility. This is where most regular users land.
Premium tiers ($200-250/month) are for power users. If you're running into limits on the $20 tier, pushing complex coding projects, or need maximum model capabilities for professional work, the premium tiers exist. But most users won't need them.
- Start with free tiers
- Upgrade to ~$20/month when you hit limits regularly
- Evaluate premium tiers only if you're a genuine power user
- Discuss organizational deployment with compliance and IT before rolling out team solutions
Understanding AI Benchmarks
This section explains benchmarks because you'll encounter them in AI discussions. But here's the key point upfront: benchmark scores tell you very little about how useful a model will be for your specific work. A 2% difference on a test doesn't translate to a meaningfully better tool for writing clinical notes or explaining diagnoses. Read this section for context, then focus on your own evaluation.
You'll often see AI companies touting benchmark scores when announcing new models. Headlines declare one model "beats" another on some test. But what do these numbers actually mean—and more importantly, what don't they mean?
What Benchmarks Measure
Benchmarks are standardized tests designed to evaluate specific AI capabilities. They provide a common yardstick for comparing models on tasks like:
- Knowledge recall: Can the model answer questions across academic domains?
- Reasoning: Can it work through multi-step logic problems?
- Coding: Can it write functional code that solves programming challenges?
- Domain expertise: How does it perform on professional exams (medical, legal, etc.)?
What Benchmarks Don't Measure
Here's what no benchmark captures—and what often matters most for real-world use:
- Writing quality: Does the output sound natural and require minimal editing?
- Appropriate uncertainty: Does the model say "I don't know" when it should?
- Instruction following: Does it do what you actually asked, or what it thinks you asked?
- Consistency: Does it give similar quality outputs across multiple attempts?
- Your specific use case: A model that aces coding benchmarks might produce worse clinical notes than one that scores lower.
A model scoring 90% vs 88% on a knowledge test is not meaningfully different for most practical purposes. What matters is whether the model helps you do your work better. The only benchmark that truly matters is your own experience using the tool.
How to Use Benchmark Data
- As a rough filter: If a model scores dramatically lower on everything, it's probably less capable overall.
- For specific tasks: If you primarily need coding help, look at coding benchmarks. For medical questions, look at MedQA scores.
- With skepticism: Companies often cherry-pick favorable benchmarks. Independent testing (like LMArena) provides more balanced pictures.
- As a starting point: Use benchmarks to decide which 2-3 models to try, then let your own experience guide your choice.
Key Benchmarks Explained
| Benchmark | What It Tests | Why It Matters |
|---|---|---|
| Humanity's Last Exam | 2,500 expert-level questions across all disciplines, designed to be genuinely difficult | The hardest general benchmark; tests limits of AI reasoning |
| GPQA Diamond | PhD-level science questions (physics, chemistry, biology) | Deep reasoning in scientific domains |
| MedQA | USMLE-style medical licensing questions | Medical knowledge directly relevant to clinical practice |
| SimpleQA | Short fact-seeking questions with single correct answers | Measures hallucination rate and factual accuracy |
| SimpleBench | Common-sense reasoning that humans find easy but AI finds hard | Tests practical reasoning vs pattern matching |
| LMArena Elo | Human preference ratings from blind head-to-head comparisons | Real users choosing preferred outputs—closest to "usefulness" |
| SWE-Bench | Fixing real bugs in actual open-source codebases | Real-world software engineering capability |
Benchmark Results (A Snapshot in Time)
Below are scores for flagship models as of April 2026 (per pricepertoken.com and vendor benchmark cards). These numbers will be outdated quickly—new model versions release every few months, and the leaderboard constantly shifts. This table is a good example: it still lists Claude Opus 4.8 and GPT-5.5, both of which have since been superseded (see the July update below). We include it to illustrate the general landscape, not to make a definitive ranking. (pricepertoken.com is a price aggregator—handy for a rough MedQA snapshot, but not an authoritative medical benchmark.) For how LLMs are actually evaluated in medicine—and how rarely they're tested on real patient data—see the JAMA systematic review "Testing and Evaluation of Health Care Applications of Large Language Models" (Bedi et al., 2024).
| Benchmark | GPT-5.5 | Claude Opus 4.8 | Gemini 3.1 Pro | Notes |
|---|---|---|---|---|
| Humanity's Last Exam | ~38% | ~37% | 46.8% | Gemini leads on hardest reasoning test |
| GPQA Diamond | 89.0% | 88.3% | 92.4% | Gemini ahead on PhD-level science |
| MedQA | ~96% | ~95% | ~94% | All far exceed passing threshold (~60%) |
| SimpleQA (Factuality) | ~65% | ~48% | 72.5% | Gemini leads on factual accuracy |
| LMArena Elo | ~1495 | ~1485 | 1508 | Gemini tops human preference ratings |
| SWE-Bench Verified | 78.1% | 88.6% | 76.9% | Claude leads on real-world coding; Opus 4.8 (May 2026) extended the 4.6→4.7 gains |
The numbers in this table will shift with every model release. What matters is this: all three models score 93%+ on MedQA—far above the passing threshold for medical licensing exams. For the vast majority of clinical use cases, all three are capable enough. Don't choose based on benchmark margins. Choose based on ecosystem fit, pricing, and—most importantly—how well the model actually performs on your specific work.
Common Pitfalls and How to Avoid Them
Quick Reference: Getting Started
ChatGPT
URL: chat.openai.com
Mobile: iOS and Android apps
Sign up: Google, Microsoft, or email
Best for: Broad capabilities, images, mature ecosystem
Claude
URL: claude.ai
Mobile: iOS and Android apps
Sign up: Google, email, or phone
Best for: Nuanced writing, analysis, coding, long documents
Gemini
URL: gemini.google.com
Mobile: iOS and Android apps
Sign up: Google account required
Best for: Google Workspace, massive context, HIPAA via Workspace
Action Items
Before moving to the next module, complete at least one of these:
- Create accounts on all three platforms (if you haven't already). Even if you end up preferring one, having tried the others gives you useful context.
- Run the same prompt through all three and compare outputs. Notice the differences in tone, organization, and approach.
- Try something you actually need: Draft a real email, summarize a real article, prepare for a real conversation. Experience how the tool performs on your actual work.
- Hit a limit and recover: Deliberately work on something complex enough that your first output isn't good. Practice iterative refinement until you're satisfied.
- If you're in a healthcare organization: Identify which HIPAA pathway makes sense for your situation and discuss it with appropriate stakeholders.
Summary
ChatGPT, Claude, and Gemini are all capable foundation models that can meaningfully assist clinical work. Each has different ecosystem integrations, pricing structures, and HIPAA pathways. ChatGPT has the largest user base and most extensive ecosystem. Gemini integrates deeply with Google Workspace and offers the most straightforward consumer HIPAA pathway. Claude offers the largest context window outside of Gemini and strong developer tools.
But here's what matters most: the choice between them matters far less than developing skill with whichever you choose. These are tools that respond to how you use them. A well-crafted prompt to any of these models will outperform a vague prompt to the "best" model.
And remember: your evaluation shouldn't be a one-time event. Models improve, workflows change, and what didn't work last year might work beautifully now. Build the habit of periodic re-evaluation with your own real-world test cases.
Since this module was first written, the competitive landscape has shifted significantly:
- OpenAI retired GPT-4o (February 13, 2026) along with GPT-4.1, GPT-4.1 mini, and o4-mini. All users have been moved to newer models. If you had workflows tuned to GPT-4o's behavior, you'll need to re-test them.
- Claude Opus 4.6 (February 5) launched with a one-million-token context window and agent teams. Anthropic closed a $30B funding round at a $380B valuation.
- Google Gemini surpassed 750 million monthly active users, closing the gap with ChatGPT (~810M). Gemini Deep Think was updated for advanced math and science reasoning.
- Perplexity launched Model Council—running queries across Claude Opus 4.6, GPT 5.2, and Gemini 3.0 simultaneously, then synthesizing where models agree or differ. A "second opinion" approach to AI search.
- Seven major models are releasing in February alone: the pace of change continues to accelerate.
The core advice hasn't changed: pick one, develop skill with it, periodically test alternatives. But the models are more capable than ever, and the gap between them is narrowing. See our AI News page for the latest.
Late February and early March 2026 brought an unprecedented wave of major model releases. The competitive landscape is moving faster than ever:
- GPT-5.4 (March 5, 2026) arrived with Thinking and Pro versions, a 1-million-token context window, native computer-use capabilities, and 33% fewer factual errors than GPT-5.2. OpenAI is closing the context-window gap with competitors. [Source]
- Claude Sonnet 4.6 (February 17) delivers near-Opus performance at one-fifth the cost ($3/$15 per million tokens), with improved computer use. Now the default model for both free and Pro users—meaning most Claude users just got a major upgrade without changing anything. [Source]
- Gemini 3.1 Pro (February 19) doubles reasoning performance over Gemini 3 Pro, dominates 13 of 16 major benchmarks, maintains a 1M-token context window, and now powers NotebookLM. [Source]
- Meta Llama 4 (February 15) introduced Scout (with a remarkable 10-million-token context window) and Maverick models, both open-weight. A major development for privacy-first and on-premise healthcare deployments where data never leaves the building. [Source]
- DeepSeek V4 launched amid controversy: a trillion-parameter model facing distillation fraud accusations from both Anthropic (reporting ~24K fake accounts) and OpenAI. A Texas Attorney General investigation has been opened. Regardless of the controversy, the data privacy concerns with Chinese-hosted models remain significant for healthcare use cases. [Source]
The practical takeaway: the Big Three are now converging on 1M+ token context windows, native tool use and computer control, and dramatically improved reasoning. Open-weight alternatives like Llama 4 are increasingly viable for organizations that need data sovereignty. Re-run your personal test set—results from even three months ago are likely outdated.
On April 16, 2026, Anthropic, OpenAI, and Google all shipped major model updates within hours of each other. The same-day release is the clearest sign yet that the frontier has become a weekly moving target:
- Claude Opus 4.7 (April 16) held pricing steady at $5/$25 per million tokens while delivering a 13% coding improvement over Opus 4.6 on a 93-task benchmark (including four tasks neither Opus 4.6 nor Sonnet 4.6 could solve), higher-resolution vision, and the ability to verify its own outputs before reporting back. Available on the API, Amazon Bedrock, Google Vertex AI, and Microsoft Foundry. [Source]
- GPT-5.4 (April 16) shipped as OpenAI's most capable frontier model for professional work, accompanied by GPT-5.4-Cyber (April 14) for vetted security professionals and GPT-Rosalind—a life-sciences reasoning model optimized for molecules, proteins, genes, pathways, and disease biology, now in research preview with Amgen, Moderna, the Allen Institute, Thermo Fisher, and Novo Nordisk. [Source]
- Gemini 3 Flash is now the default model in the Gemini app, with Gemini Agent (multi-step task execution across Workspace, Deep Research, Canvas, and live web) and Gemini 3 Deep Think available to Google AI Ultra subscribers. Grounding with Google Maps is now supported. [Source]
- MedQA leaderboard snapshot (April 9, 2026) — the top of the medical-question benchmark, useful as a rough sanity check when comparing models: o4 Mini High 95.2%, Gemini 2.5 Pro 94.6%, Claude 3.7 Sonnet 92.3%. Average across all 34 evaluated models is 79.4%. [Source]
The practical takeaway: the gap between the Big Three is narrowing fast, and any of them is a reasonable default for clinical work. Pick based on ecosystem (Epic vs. Google Workspace vs. Microsoft 365), HIPAA pathway, and the interface you'll actually use daily. Re-test your workflows quarterly—what worked in February may not be the best option now.
Several things changed between the June update and the end of July—and one thing conspicuously didn’t:
- Claude Opus 5 (July 24) replaced Opus 4.8 as Anthropic’s current Opus, at the same price: $5/$25 per million tokens, a one-million-token context window, 128K max output, and a knowledge cutoff of May 2026. It is the default model on Claude Max and the strongest model on Pro; Opus 4.8 still works but has moved to the legacy list. Anthropic positions it as approaching Fable 5’s intelligence at half the price. Two notes for clinical readers: the headline gains are on coding and agentic work, and Anthropic’s life-sciences claims are its own internal evaluations, not independent clinical benchmarks—so treat them the way you would any vendor’s in-house numbers. [Source]
- Claude Sonnet 5 (June 30) shipped with a one-million-token context window and 128K max output, and became the default for Free and Pro users on claude.ai as well as the default in Claude Code from launch. Launch pricing of $2/$10 per million tokens was framed as introductory through August 31, 2026, with a rise to $3/$15 to follow—an increase Anthropic cancelled on August 11, 2026: $2/$10 is now the standard price, and the pricing documentation states the September 1 increase “will not occur.” A detail almost nobody reports: Sonnet 5 uses an updated tokenizer that expands token counts by roughly 1.0–1.35× for the same text—so when comparing costs across models, compare tokens as counted by each model’s tokenizer, not list prices. [Source] · [Pricing docs]
- GPT-5.6 became generally available on July 9, after the limited June preview. Three tiers, priced per million tokens: Luna $1/$6, Terra $2.50/$15, and Sol $5/$30. All three carry a 1M-token context window, 128K max output, and a knowledge cutoff of February 16, 2026. The same day, GPT-5.6 became the preferred model in Microsoft 365 Copilot. [Source]
- Gemini 3.5 Pro still does not exist. Its release has slipped more than once, and what Google itself has said is narrow: the model fell short of internal performance goals. Press reporting in mid-July went further—attributing the delay to hallucination and reliability problems serious enough to require rebuilding the base model—but Google has not confirmed that account, so treat it as reported rather than established. What is confirmed: on July 21 Google shipped Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Flash Cyber, with no Pro model among them. Treat any 3.5 Pro specification you see quoted as unconfirmed until Google announces one. [Source]
This also corrects a framing in the June update below. When GPT-5.6 first appeared as a preview restricted to roughly 20 pre-approved organizations, it looked like the start of a pattern—government sitting between clinicians and the frontier. Two weeks later it was on sale to anyone with an API key. The gate turned out to be a release stage, not a new structure. The lesson that survives is narrower and more useful: access to a specific model can change on short notice in either direction, so keep a fallback rather than hard-wiring one vendor and version into a clinical workflow.
The biggest story of the month is a governance one. On June 9, Anthropic publicly released its most capable models yet—Claude Fable 5 and the cybersecurity-focused Mythos 5. Three days later, on June 12, the US government issued an export-control directive citing national-security authorities, suspending access to both models for any foreign national—inside or outside the US, including Anthropic’s own foreign-national employees.
To comply, Anthropic abruptly disabled Fable 5 and Mythos 5 for all customers. Access to every other Anthropic model—including Claude Opus 4.8 and below—was unaffected. The government’s stated basis was a demonstrated “jailbreak” technique; Anthropic said its own review found only a small number of previously known, minor vulnerabilities, and it publicly disagreed that a narrow jailbreak should justify recalling a model deployed to hundreds of millions of people. [Anthropic statement] [Fortune]
The suspension ended on June 30, when the export controls were lifted. Anthropic restored Fable 5 globally on July 1—across the Claude Platform, claude.ai, Claude Code, and Claude Cowork—after an 18-day outage, shipping alongside it a new cybersecurity classifier the company says blocks the disputed jailbreak technique in more than 99% of cases. One part of the story did not close the same way: Mythos 5 did not fully return. It was partially restored to a set of US organizations on June 26 and remains restricted to them. [Anthropic]
Separately, the strongest open challenger of the June cycle arrived from outside the US entirely: GLM-5.2 (Zhipu AI, June 13), an MIT-licensed model that independent rankings put first among open weights as of June 2026, at roughly one-sixth the cost of the closed frontier. That standing did not last the quarter: Kimi K3 (Moonshot AI) published its weights on July 27 and scores 57 on the Artificial Analysis Intelligence Index against GLM-5.2’s 51. Open-weight leaderboard positions turn over in weeks—read any of them with a date attached.
Why it matters for clinicians: a frontier model you build on can be pulled by government order on short notice, and it may not come back on the same terms—Fable 5 returned in full, Mythos 5 did not. If a clinical workflow depends on one specific model, keep a fallback. Don’t hard-wire a single vendor or version.
The headlines since the May update:
- Claude Opus 4.8 (May 28) shipped at the same pricing as 4.7: 128K max output tokens, 88.6% on SWE-Bench Verified, and a Fast mode ($10/$50 per million tokens at ~2.5× speed). The benchmark table above has been updated. The same week, Anthropic pre-released its cybersecurity-focused Mythos model to limited partners and raised $6.5B at a record $965B valuation. [Source]
- Dynamic Workflows (May 28) arrived in Claude Code as a research preview for Max/Team/Enterprise: one task can orchestrate up to 1,000 subagents. If you use Claude Code heavily, the Max plan is the intended home for this. [Source]
- Gemini 3.5 Pro had not shipped as of June 19, and still had not as of late July—see the July update above. Meanwhile Google Health Coach (May 19, $9.99/month) became the first Gemini-powered consumer health product with US EHR integration—the most clinically consequential Gemini release this cycle.
A second wave of releases in May 2026 reshuffled the consumer-tier defaults:
- GPT-5.5 Instant (May 5) became the new default model for ChatGPT, with OpenAI reporting roughly 52% fewer hallucinations than GPT-5.3 on high-stakes prompts and a faster response cadence than GPT-5.4. The comparison table above has been updated to reflect GPT-5.5. [Source]
- Claude for Small Business (May 2026) launched a workspace-style plan for SMB teams who want Claude without enterprise procurement.
- Gemini 3.5 Flash and Gemini Omni Flash shipped at Google I/O 2026 (May 19): 3.5 Flash rolled out to all users (~4× faster than peer frontier models on agentic benchmarks); Omni Flash adds a new multimodal model for Plus/Pro/Ultra. The tier cards above have been updated accordingly. [Source]
The pattern from April held in May: same-month releases from all three vendors, no stable leader, and benchmark numbers that move faster than this page can track. Treat the snapshots as orientation, not verdicts.
Think of them as brilliant but inexperienced colleagues: genuinely helpful, occasionally wrong, always requiring supervision. Pick one based on your ecosystem and access. Use it enough to develop skill. Periodically test alternatives with your own use cases. With that approach and the prompting skills from earlier modules, you're ready to begin. Now close this document and go have a conversation with one of them. That's where the real learning happens.
Learning Objectives
- Compare the capabilities, strengths, and limitations of ChatGPT, Claude, and Gemini
- Identify which model best fits specific clinical use cases
- Understand HIPAA compliance pathways for each platform
- Apply a practical framework for getting started with foundation models
- Recognize common pitfalls in AI adoption and strategies to avoid them
Notes
-
Why not Grok? You may wonder why xAI’s Grok isn’t
one of the platforms covered here. The current picture is in
Beyond the Big Three above; the short version is
that the technology is capable but there is no clinical-privacy pathway on any
surface people actually use.
When this page was first written, the disqualifiers were 2024–2025 misinformation incidents and independent evaluations finding weaker guardrails than ChatGPT, Claude, or Gemini. As of August 2026 the sharper reasons are different: consumer Grok trains on conversations by default and xAI’s own privacy policy asks users not to submit health information; roughly 370,000 shared conversations—medical content included—were indexed by Google in August 2025; the only BAA path is API-only with a zero-data-retention tier, covering none of the consumer surfaces; and on expert-level diagnostic imaging—the use its owner promotes—Grok-4 scored 12% against radiologists’ 83%, near the bottom of the frontier pack.
References:
↩
Forbes. "Elon Musk Keeps Telling People To Use AI For Medical Advice—But Grok Says Not To." February 2026.
xAI. "Privacy Policy." Effective April 2026.
JMIR AI. "Performance of 5 Large Language Models on the NBME Free 120." 2026.
arXiv. "RadLE: Radiology Last Exam—Benchmarking Frontier Multimodal AI Against Human Experts." September 2025 (preprint).
Axios. "Musk's AI chatbot spread election misinformation, secretaries of state say." August 2024.