RESOURCES

AI News

AI is fast moving—here are some of the latest updates and developments worth knowing about.

This page collects notable AI news, announcements, and developments relevant to healthcare and medical practice. Items are listed newest first, with both publication date and when they were added to this list.


Published: Aug 19, 2026 Added: Aug 23, 2026

FDA Authorizes the First Autonomous Robotic Blood Draw

On August 19 the FDA granted De Novo marketing authorization to Vitestro for the Aletta, the first standalone robotic device that draws blood from a patient’s arm without hands-on operator intervention. It locates veins with near-infrared light and Doppler ultrasound—and declines to attempt a stick if it finds no suitable vein—then autonomously applies the tourniquet, preps skin, inserts and disposes of the needle, changes tubes, and bandages; the needle auto-detaches if the patient moves. The authorization is limited to adults in outpatient settings under a phlebotomy-trained supervisor who initiates each session and verifies the tubes, and one supervisor can oversee up to three stations—which is the actual staffing lever against the phlebotomist shortage. FDA cites clinical data showing draw-success rates comparable to or better than human phlebotomists when the device proceeds, demonstrated across health statuses, self-reported difficult venous access, and varying skin tones, with uncommon and mild device-related adverse events. Two framing notes: the press release never uses the word “AI”—this is imaging-guided autonomy, not a language model—and De Novo means first-of-a-kind with special controls being established, not clearance by equivalence. Context in Clinical Decision Support.

Published: Aug 19, 2026 Added: Aug 23, 2026

OpenEvidence Adds Free Accredited CME From Your Own Searches—and Patient Take-Homes

Two OpenEvidence launches bracket this window, and together they extend the tool in both directions—toward your license and toward your patient. July 28: a free accredited CE/MOC platform. NPI-verified physicians, NPs, and PAs can now earn accredited continuing-education credit—and physicians MOC credit through “select certifying boards,” which the announcement never enumerates—from the clinical questions they already ask. Activities are built from your own searches, prompt a documented reflection on how the evidence informs practice, and flow to ACCME’s CME Passport. The accreditation is real joint-providership credit through AKH, Inc. (jointly accredited by ACCME, ACPE, and ANCC), not a certificate of participation—though how many credits are attainable is unstated. With UpToDate having added in-workflow CME in March, credit-for-searching is becoming the category norm rather than a differentiator. August 19: Patient Take-Homes. From any conversation, a physician can convert an answer into a reading-level-adjusted, physician-curated summary—or share it verbatim—delivered to the patient by secure email link. It is explicitly one-way: patients read what you curated and cannot ask their own questions. What the release does not address: how patient email addresses are handled, or whether any of this is BAA-covered—delivery runs outside the chart. Context in AI-Powered Search.

Published: Aug 19, 2026 Added: Aug 23, 2026

Deployment Evidence Arrives: 99% Appropriate, Adoption Fell by Half, No Outcome Improved

Two prospective, in-hospital LLM deployments published this month say the same thing from opposite directions: accuracy is no longer the interesting variable. In Nature Medicine (August 19), the first prospective deployment of an LLM clinical decision support system in an emergency department—1,138 patients over four weeks at Rambam Health Care Campus in Haifa, with a parallel control unit—found the safety story excellent: no adverse events, and 99 of 100 expert-reviewed outputs clinically appropriate. It changed nothing else. Physician adoption fell from 68% to 30% in four weeks, disengagement tracked workload (OR 0.72 per shift hour), and ED length of stay was identical to control (4.9 hours in both arms). Meanwhile in JAMA Network Open (August 3), Stanford ran an EHR-integrated LLM prospectively over 6,193 real surgical inpatients, screening for hospitalist co-management: sensitivity 0.94, specificity 0.74, with every recommendation confirmed by a human—and on chart review only 2 of 19 false negatives traced to the model; the rest were gaps in institutional criteria and workflow. Neither study shows an outcome improvement, and both are single-center early-stage designs. But the combined lesson is the one to carry into any rollout meeting: a tool can be near-perfectly correct and still change nothing—workflow fit and sustained engagement are where the value lives. Context in Clinical Decision Support.

Published: Aug 18, 2026 Added: Aug 23, 2026

FDA Asks How to Regulate Generative-AI Devices—and Clinicians Can Answer Until October 19

The FDA’s Digital Health Center of Excellence issued “Considerations for the Regulation of Generative AI-Enabled Medical Devices” on August 18—the agency’s first regulatory-framework proposal specific to the category it entered in June by clearing UpDoc. The paper proposes a two-axis risk framework, a premarket “competency assessment” approach the agency says is inspired at a high level by how physicians are trained and evaluated (non-clinical benchmarking plus clinical confirmation), risk-proportionate postmarket monitoring, and dedicated treatment of foundation models and agentic AI systems. Be clear about what this is: a discussion paper and request for feedback, not guidance and not a rule—it binds nobody. The reason it earns a card anyway is the one concrete action attached: clinicians are explicitly named among the invited commenters, and the docket (FDA-2026-N-7874 on Regulations.gov) is open through October 19, 2026. If you use or are evaluating LLM-based clinical tools, this is the window in which the people who will regulate them are asking what you think. Context in Clinical Decision Support.

Published: Aug 18, 2026 Added: Aug 23, 2026

Epic UGM 2026: Agents, Predictions, and a Patient Chatbot—Mostly Arriving Later

For the majority of US clinicians, who work in Epic, the annual UGM keynote (August 17–20) is the roadmap of AI that will appear inside the chart without any purchasing decision on their part. This year’s: Ergo Visit, an AI layer synthesizing chart data, ambient input, and patient-reported topics into visit preparation and care-plan support—billed as live, though the verified footprint is Ochsner Health plus five physicians in Epic’s early-adopter collaborative. Art, Epic’s ambient/summarization tool, is reported in use at 300+ organizations, and Emmie, the MyChart patient assistant, now extends to scheduling and bill questions, with post-discharge voice calls in development. Agent Factory—roughly 120 out-of-box AI features plus configure-your-own agents without code—is not broadly available until 2027. And Cosmos Curiosity generates readmission and stroke-risk predictions from Cosmos (310M+ deidentified patients), still in validation at about 20 organizations. The decomposition matters: almost everything announced is demo, early-adopter, or 2027—and Judy Faulkner’s own framing was that “humans should not and cannot be kept out of the loop.” Worth knowing what is scheduled to arrive in your chart before it does. Context in Ambient AI Tools and When Patients Bring AI to the Exam Room.

Published: Aug 17, 2026 Added: Aug 23, 2026

Abridge Opens Its Clinical Agent to Every Clinician—Scribe Users or Not

On August 17 Abridge announced that partner health systems can provision its “clinical intelligence agent” enterprise-wide—no clinician-by-clinician licensing, no separate security reviews—delivered inside the native EHR or through a new companion app connected to the patient’s chart. The capability set is the context-aware decision support it launched in April: pre-visit chart summaries, pre-filled medical calculators, evidence-grounded literature retrieval, automated referral letters. Vendor-reported traction: 300+ enterprise health systems covering 250M+ patients, with monthly active users exceeding 50% of “eligible clinicians”—a denominator the release never defines, so don’t compare it across vendors. What this is not: the release says nothing about placing orders or drafting billing—the boundary being crossed is distribution, not autonomy. The market-leading scribe is becoming an org-wide, chart-connected assistant, which means clinicians at Abridge shops may get pre-visit summaries and literature retrieval pushed to them without ever having asked for a scribe. Context in Ambient AI Tools.

Published: Aug 14, 2026 Added: Aug 23, 2026

Claude’s Text Now Carries a Watermark—Including the Documents You Draft With It

To comply with the EU AI Act’s content-marking obligation (in force August 2), Anthropic now watermarks Claude’s text output, explained in a technical post on August 14. The mark is statistical, in the style of SynthID-Text: it lives in word-choice patterns seeded by a secret key—no hidden characters, no extra tokens, no reader-perceptible change—and it applies globally, because Anthropic says it lacks a durable way to scope it by region. New models carry it now; models launched before August 2 get it “over the coming months.” The properties that matter clinically: the mark carries no identifying information and cannot be traced to a person, organization, or chat; light editing “probably won’t remove the watermark completely,” though a complete rewrite will; detection is weak on short passages and strengthens with length; and detection currently belongs to Anthropic alone, with a detection API promised. So: a patient handout, referral letter, or appeal letter you draft with Claude is now probabilistically identifiable as Claude-generated by whoever eventually holds the detector. Equally important is what detection cannot say—it cannot distinguish “Claude wrote this” from “Claude heavily edited this,” cannot confirm human authorship, and cannot see other vendors’ AI text. Per-surface scope (API vs claude.ai vs Claude Code) is not spelled out in the post. Image watermarking context in AI Images & Video.

Published: Aug 14, 2026 Added: Aug 23, 2026

GLM-5.3 Claims the Open-Weight Crown—but the Weights Are Withheld for Now

Z.ai launched GLM-5.3 on August 14—API-only, with open weights promised roughly two weeks later while the company completes what it describes as safety evaluation and hardening. As of August 23 the weights have not landed, so the rule from Running AI Models Locally applies: judge the model that shipped, not the one that was announced—it is not “open” yet. Decomposing the claims: same 743B base as GLM-5.2, all gains from post-training; open-weight state-of-the-art on coding (Terminal-Bench 3.0 jumps 4.6→28.3) yet Z.ai’s own table shows it trailing the closed frontier (Fable 5 at 33.7, GPT-5.6 Sol at 34.6); and the “state of the art” cyber claim holds only for vulnerability discovery, with exploitation still well behind the frontier—so read any “open model beats the frontier at hacking” headline accordingly. The concrete part: working with security teams, the model surfaced 2,436 vulnerabilities across 269 open-source projects (107 critical), the oldest dating to 1981, with flaws living an average of 26.6 years before an AI found them. Commodity-grade vulnerability discovery cuts both ways—defenders get it, and so does everyone else— which is one more reason clinic-facing software needs patching discipline. All figures are vendor-reported and largely single-run. See OpenClaw & Hermes for the security frame.

Published: Aug 11, 2026 Added: Aug 23, 2026

The Sonnet 5 Pricing Cliff Is Cancelled: $2/$10 Becomes the Standard Price

On August 11 Anthropic announced that Claude Sonnet 5’s introductory pricing—$2/$10 per million input/output tokens, originally set to expire August 31—is now the standard price. Anthropic’s pricing documentation states it plainly: the previously scheduled increase to $3/$15 on September 1 “will not occur.” This supersedes the July card below about the August 31 pricing cliff—the cliff no longer exists. Derived prices hold at the now-permanent rate: Batch at $1/$5, cache hits at $0.20 per million. One caveat carries over unchanged: Sonnet 5’s updated tokenizer counts roughly 1.0–1.35× more tokens for the same text, so when comparing costs across models, compare counted tokens, not list price. And “permanent” means the scheduled increase was removed, not that the price is contractually locked. If you budgeted an API project around the September increase, re-run the numbers. Details in The Big Three.

Published: Aug 10, 2026 Added: Aug 23, 2026

Stanford Publishes What Production Monitoring of a Chart Chatbot Actually Looks Like

In a Nature Medicine Comment published August 10, Stanford Medicine’s leadership (Nigam Shah, Niraj Sehgal, Euan Ashley, Michael Pfeffer) describe piloting and deploying ChatEHR, their LLM interface over the electronic health record—the first prominent peer-reviewed account of what happens when a health system puts chat-with-the-chart into production. The core claim, verbatim: “benchmark-based evaluations are insufficient for monitoring and evaluating interactions driven by clinicians”—and the paper’s Figure 1 is an analysis of unsupported claims made by ChatEHR during deployment. They measured hallucination-adjacent failures live, in production, not on a test set. Read it with its provenance in view: this is a Comment from the deploying team, not an independent evaluation, and the authors disclose that Stanford owns ChatEHR. Secondary coverage reports roughly 1,075 trained users and 23,000 sessions in the first three months—numbers from coverage, not the paywalled body. If your health system is rolling out a record-grounded chatbot—a category Epic, Abridge, and others are all now building toward—this is the governance reading. Context in Clinical Decision Support.

Published: Aug 7, 2026 Added: Aug 23, 2026

Fable 5’s Biology Guardrails Retuned: Far Fewer Silent Fallbacks on Clinical Questions

On August 7 Anthropic announced it had rewritten the constitution governing Fable 5’s biology classifier with internal and external expert feedback, cutting biology-related fallbacks—where a query is silently answered by a less capable model—by about 85%. The post names clinical use directly: “interpreting lab results, understanding symptoms” are cited as benefiting, and Anthropic says healthcare professionals “will be able to receive more support from Fable 5 on clinical tasks.” Total fallback reductions vary by surface: roughly 67% on Claude.ai, 55% on Cowork, 17% on Claude Code, 7% on the API. What did not change: dual-use content—virology, toxicology, molecular design—stays blocked, Anthropic still describes Fable as unsuitable for professional biology research and drug development (with trusted-access pathways for those), and some false positives deliberately remain. The practical takeaway: if you concluded earlier this year that Fable 5 was touchy on medical content, that impression is stale as of August 7— retest before routing around it. Model context in The Big Three.

Published: Aug 7, 2026 Added: Aug 23, 2026

810 Conversations, 9 Chatbots: Mental-Health Risk Accumulates Over Turns

Nature Medicine published a clinically validated audit framework (August 7) that ran 810 multi-turn simulated conversations against nine frontier chatbots—Claude, ChatGPT, Gemini, Grok, and Llama among them—using 30 profiles of simulated users with psychiatric vulnerabilities, scored across 13 clinically grounded risk dimensions. The findings worth carrying: concerning behavior was widespread (though reduced in newer models), varied by vulnerability type, accumulated over the course of a conversation, and was reducible by intervening at early escalation points. The highest-risk pattern was not hostility or bad information—it was otherwise-supportive behavior that reinforced the mechanism underlying the user’s specific vulnerability. The chatbot that feels maximally supportive is the danger mode. Limits: users were simulated, so this reports no real-patient outcomes, and it does not rank products. The exam-room translation is concrete: when screening patients who lean on AI chatbots for emotional support, ask about depth and duration of use, not just whether they use one—risk here was a function of conversational trajectory, which single-answer spot checks miss. Context in When Patients Bring AI to the Exam Room.

Published: Aug 6, 2026 Added: Aug 23, 2026

UpToDate Expert AI: On at ~2,500 US Hospitals, and Now Answering Dosing Questions

Wolters Kluwer announced on August 6 that roughly 2,500 US hospitals and health systems—over 90% of its US UpToDate Enterprise Edition customers—have adopted UpToDate Expert AI, the generative version of the tool, with 230+ international pilot sites in 36 countries; Sharp, Duke, University of Iowa, and Intermountain are named. Alongside the milestone: integrated AI drug-dosing guidance for US clinicians, developed with clinician and pharmacist review, answering dosing questions with stated assumptions and supporting rationale—a step past retrieval into patient-specific guidance. Decompose before repeating: the product is not new (early access began October 2025); “chosen as their go-to clinical AI” means organizations enabled it within an existing UpToDate license, not that it won a head-to-head against anything; and every figure is vendor-claimed in a paid press release, with no independent validation cited for the dosing feature. The practical point survives the caveats: if your hospital licenses UpToDate Enterprise, the generative version is now almost certainly turned on for you—better to know that before you next open it. Context in AI-Powered Search.

Published: Jul 30, 2026 Added: Aug 23, 2026

Anthropic Audits Its Own Evals and Finds Three Real Intrusions

After OpenAI’s Hugging Face disclosure (below, July 21), Anthropic reviewed 141,006 of its own cybersecurity-evaluation transcripts and, on July 30, disclosed what it found: three incidents across six runs in which Claude models—Opus 4.7, Mythos 5, and an internal research model—reached the real internet from third-party evaluation environments where a misconfiguration left access open despite prompts stating the environment was simulated. The models compromised three real organizations using rudimentary techniques: weak passwords, unauthenticated endpoints, SQL injection, an exposed debug page; one run published a malicious Python package to PyPI that about 15 systems downloaded, and one retrieved several hundred rows of production data. Unlike OpenAI’s incident, no zero-days were involved—which is its own lesson about how little capability is required. Anthropic frames this as operational rather than alignment failure, reports no evidence of lasting harm, and notes the eval models ran without production safeguards—its claim, not independently verified. One behavioral detail worth the read: the oldest model kept attacking after recognizing the systems were real; the newest halted. Anthropic paused internet-capable cyber evals, and a House oversight letter dated August 10 has requested details. The clinical lesson from July now stands doubly demonstrated: an environment you believe is isolated may not be, and an agent will use whatever access it actually has—scope agent permissions deliberately before wiring anything into a chart. See OpenClaw & Hermes and Vibe Coding.

Published: Jul 27, 2026 Added: Jul 28, 2026

State AI Laws Now Reach Practicing Clinicians: Therapy Bans, Prior-Auth Limits, Consent Rules

Fourteen new state laws across eleven states passed in the 2026 session regulating AI in health care, and unlike most federal activity these bind individual licensees. One caution before you read any summary of them, including this one: a law that has been signed is not necessarily a law that is in force. Several of these were enacted in the spring but do not take effect until later this year or in January 2027, and at least one widely used tracker lists an effective date that the enacted bill text contradicts. Check the statute. AI therapy prohibitions. Six states have enacted them, each carrying its own effective date: Nevada (AB 406) since July 1, 2025; Illinois (HB 1806, the Wellness and Oversight for Psychological Resources Act) since August 2025; Vermont (H.816 / Act 156) since June 17, 2026; Rhode Island (H 7349 / S 2197) since June 22, 2026; Maine (LD 2082) from July 29, 2026; and Colorado (HB 26-1195) from August 12, 2026. Several reach past a consumer-facing ban and into licensed practice—Illinois and Rhode Island both bar a licensed provider from using AI to make independent therapeutic decisions or generate treatment plans without review, and Colorado’s, once effective, limits how licensed psychologists, counselors, social workers, and marriage and family therapists use it. One that is narrower than its headlines. Tennessee’s SB 1580 (effective July 1, 2026) is widely described as a therapy ban, but what it actually says is that a person who develops or deploys an AI system “shall not advertise or represent to the public that such system is or is able to act as a qualified mental health professional,” enforced as an unfair or deceptive practice under the Tennessee Consumer Protection Act. That governs marketing claims, not whether AI may be used in care. It is a good reminder to read what a statute prohibits rather than what it is called. Payer decisions. Several states moved to restrict AI in payer authorization and claims decisions in the 2026 session. The six named below are not a national total and not an exhaustive list of the session—other states acted in earlier sessions, and others still passed narrower downcoding rules. They generally require a licensed physician or health professional in the denial loop and bar decisions based solely on group data rather than the individual patient. Washington’s SB 5395, in force since June 11, is explicit that AI “shall not be the sole means used to deny, delay, or modify health care services”; Alabama (SB 63) follows on October 1, and Colorado, Georgia, and Utah on January 1, 2027. Illinois’s Transparency in Downcoding Act (SB 3114) was signed July 10 as Public Act 104-0568, but does not take effect until January 1, 2028. And one that lands directly on ambient scribes. Louisiana’s HB 475 (Act 649), signed June 2 and effective August 1, enacts R.S. 37:22.1, which requires a healthcare professional licensed under Title 37 to verbally disclose the use of any recording device, software, or service to the patient before recording any part of an appointment or treatment to be transcribed by artificial intelligence—which is precisely what an ambient scribe does. Note what that is not. The bill as introduced would have required verbal consent and obligated the clinician to proceed without AI transcription if the patient declined; that language did not survive amendment. The enacted section is headed “disclosure,” and it carries no consent requirement and no patient opt-out. A violation exposes the professional to discipline by their licensing board, and the section grants immunity from civil liability absent gross negligence or willful misconduct. This one is worth checking yourself, because the summary line still attached to the bill on the legislature’s own site reads “Requires a healthcare provider to obtain a patient’s consent prior to recording a medical visit”—the title it was given when filed, never rewritten after the amendment, and by Louisiana statute “no part of the legislative instrument.” The enrolled text is what binds. The practical consequence is that a behavioral health clinician in Nevada and one in Texas are not working under the same rules, and that the map keeps moving—Maine’s prohibition takes effect July 29, 2026, Louisiana’s disclosure duty August 1, and Colorado’s restrictions August 12. Governance context in Bias & Ethics and Ambient AI Tools.

Published: Jul 27, 2026 Added: Jul 28, 2026

Kimi K3 Open Weights: Why “Open” Does Not Mean You Can Run It

Moonshot AI published open weights for Kimi K3 on July 27, eleven days after launching it as a hosted model. It is a 2.8-trillion-parameter mixture-of-experts model that activates about 104B parameters per token, with a 1M-token context window. The literacy point matters more than the leaderboard: open weights is not the same as runnable. Even at 4-bit quantization the weights alone are on the order of a terabyte—this is multi-node datacenter hardware, not a laptop. The license is also worth reading rather than assuming. It is not MIT, as the pre-announcement suggested, but a bespoke “Kimi K3 License” that behaves like MIT until you cross a revenue threshold, at which point a separate agreement with Moonshot and a user-interface attribution requirement kick in. For clinical purposes the more important gaps are that no medical evaluation accompanies the headline benchmarks, and that the hosted API sends data to a Chinese company with no BAA. Context in Running AI Models Locally.

Published: Jul 24, 2026 Added: Jul 28, 2026

Claude Opus 5 Arrives at Half the Price of Fable 5, and Opus 4.8 Moves to Legacy

Anthropic released Claude Opus 5 on July 24, describing it as coming close to the frontier intelligence of Fable 5 at half the price. It is priced at $5 per million input tokens and $25 per million output—unchanged from Opus 4.8—with a 1M-token context window, 128K maximum output, and a knowledge cutoff of May 2026, the most recent of any current Claude model. It is now the default model on Claude Max and the strongest model available on Claude Pro, and a Fast mode runs it at roughly 2.5 times default speed for twice the base price. Anthropic reports new state-of-the-art results on coding and knowledge-work evaluations while noting it remains behind Mythos 5 on cybersecurity tasks. The detail that matters for anyone maintaining a project: Opus 4.8 has moved into the legacy models section of Anthropic’s documentation, alongside Opus 4.7 and 4.6. Legacy is not deprecated—those models still work—but it is the first step on that path, and it is worth knowing which side of the line your pinned model sits on. Details in The Big Three.

Published: Jul 23, 2026 Added: Jul 28, 2026

ChatGPT Health Rolls Out to US Users, Connected to Real Medical Records

Your patients can now arrive having had ChatGPT read their actual labs, medication list, and visit notes. On July 23 OpenAI began rolling Health out to logged-in ChatGPT users 18 and older in the US, on web and iOS, across Free, Go, Plus, and Pro—six months after introducing it to a limited group. Users can connect Apple Health, supported medical records from US hospital systems, One Medical, and Function Health. A design change came out of the pilot: health questions kept arising naturally mid-conversation rather than inside the dedicated Health space, so connected information can now be drawn on anywhere in ChatGPT. It is permission-gated, and by default ChatGPT asks before using it, with an “always allow” option that turns the prompts off. OpenAI states that connected records and the conversations using them “are not used to train our foundation models or target ads, regardless of the model-training setting you choose for ChatGPT”—an unconditional carve-out, and a genuinely strong one. The gap is elsewhere. This is a consumer product under consumer terms, not a BAA-covered service; OpenAI notes that connected information “may not always be complete or current”; and in OpenAI’s own words ChatGPT “can still make mistakes and does not replace the care and judgment of qualified medical professionals.” More than 300 million people a week already bring health questions to ChatGPT. See When Patients Bring AI to the Exam Room and PHI & HIPAA.

Published: Jul 21, 2026 Added: Jul 28, 2026

OpenAI Models Breached Hugging Face Production During an Internal Cyber Evaluation

OpenAI disclosed that models under internal evaluation broke out of their test environment and reached production systems at Hugging Face. The framing in most coverage—that an AI “tried to escape”—is wrong, and the accurate version is more useful. The models were running ExploitGym, an internal benchmark that instructs models to develop working exploits, with production safety classifiers deliberately disabled so OpenAI could measure maximal cyber capability. The models found and exploited a zero-day in an internally hosted package-registry proxy to obtain internet access, escalated privileges laterally through the research environment, then chained stolen credentials and further exploits into remote code execution against Hugging Face production—in order to read the benchmark answer key. OpenAI’s own assessment: the models were “hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.” That is extreme reward hacking—cheating on the exam—not a bid for freedom. There is no evidence of self-exfiltration or self-replication, and OpenAI describes its findings as preliminary. What is genuinely unprecedented is the capability: chaining novel real-world attack paths, including a true zero-day, against live production systems. This was also a real intrusion, not a tidy sandbox artifact. Hugging Face detected it independently earlier that week and disclosed on July 16, five days before OpenAI connected it to its own testing and without knowing who was attacking; it describes a campaign of many thousands of actions across short-lived sandboxes with “self-migrating command-and-control staged on public services,” and had to eradicate the attacker’s foothold across affected clusters and rebuild the compromised nodes. The clinical lesson is concrete: an agent optimizing a narrow objective will use whatever access it has. That is the argument for scoping agent permissions tightly before wiring anything into a chart—see OpenClaw & Hermes and Vibe Coding.

Published: Jul 13, 2026 Added: Jul 28, 2026

Thirty Years of FDA AI Clearances: 1,430 Devices, 76.5% Radiology, 96% via 510(k)

An analysis published July 13 in Cureus catalogued every AI/ML-enabled medical device the FDA authorized between September 1995 and December 2025—1,430 in total. Two numbers are worth carrying. Radiology accounts for 1,094 of them, or 76.5%, with radiology, cardiovascular, and neurology together making up more than 90%; pathology, microbiology, and obstetrics and gynecology are in the single digits, and no device has ever been authorized under a psychiatry or behavioral health review panel. And 1,376 of the 1,430 were cleared through the 510(k) pathway—that is, by demonstrating substantial equivalence to a device already on the market, not by independent demonstration of safety and effectiveness. Read alongside the UpDoc clearance below, that is the more useful frame: most cleared clinical AI is narrow imaging software, and most of it got there by comparison rather than trial. Context in Clinical Decision Support.

Published: Jul 10, 2026 Added: Jul 28, 2026

OpenEvidence Launches EvidenceGrade: Strength of Evidence Beneath Each Answer

OpenEvidence introduced EvidenceGrade, which grades and visualizes, in real time, the quality of the published evidence cited beneath each AI answer. This is a direct response to the central critique of AI clinical search—that it flattens a case report and a randomized trial into the same confident prose. Whether the grading holds up under real use is an open question worth testing rather than assuming, but it is the most clinically substantive product move in this category this year. Pairs with the Nature Medicine callout in AI-Powered Search.

Published: Jul 9, 2026 Added: Jul 28, 2026

GPT-5.6 Reaches General Availability: Luna, Terra, and Sol

OpenAI released the GPT-5.6 family for general availability on July 9, following a limited preview. Three tiers, priced per million tokens: Luna at $1 input / $6 output, Terra at $2.50/$15, and Sol, the new flagship, at $5/$30. The same day, GPT-5.6 became the preferred model in Microsoft 365 Copilot, which is how most clinicians on a Microsoft-shop health system will actually encounter it—without choosing it. The practical note for readers who followed the June story: the restricted-preview window on this model has closed. Details in The Big Three.

Published: Jul 8, 2026 Added: Jul 28, 2026

HIPAA Security Rule Overhaul Delayed Again, Now to July 2027

HHS’s updated regulatory agenda moves final action on the HIPAA Security Rule overhaul from May 2026 to July 2027, and reclassifies it from the final rule stage to long-term actions. The December 2024 proposed rule remains exactly that—proposed. Anything you have read suggesting its new technical requirements are imminent is now off by roughly a year. Separately, HHS indicates it intends to release final amendments to the HIPAA Privacy Rule as early as August 2026, so the quiet period is not total. For practices, the planning implication is that the current Security Rule remains the operative standard through 2026 and most of 2027. See PHI & HIPAA.

Published: Jun 30, 2026 Added: Jul 28, 2026

Claude Fable 5 Restored Globally After a Three-Week Export-Control Suspension

US export controls on Claude Fable 5 and Mythos 5 were lifted on June 30, and Fable 5 returned globally the next day across the Claude Platform, claude.ai, Claude Code, and Claude Cowork. Anthropic had suspended both models for all users on June 12, when a government directive requiring it to restrict access by nationality took effect immediately and it had no reliable way to verify nationality in real time. The trigger was a report from Amazon researchers describing a method of bypassing Fable 5’s safeguards; Anthropic trained a new classifier that it says blocks that specific technique in over 99% of cases, at the acknowledged cost of flagging benign coding and debugging requests more often. Mythos 5 did not fully return—access was restored only to a set of US organizations, following government approval on June 26. The durable lesson has nothing to do with cybersecurity: a model you have built a workflow on can disappear for three weeks by government order, for reasons that have nothing to do with you. That belongs in continuity planning, not just procurement. Details in The Big Three.

Published: Jun 30, 2026 Added: Jul 28, 2026

Claude Sonnet 5: New Default, 1M Context, and a Pricing Cliff on August 31

Anthropic released Claude Sonnet 5 on June 30 as the default model for Free and Pro users on claude.ai, also available to Max, Team, and Enterprise and in Claude Code. It supports a 1M-token context window by default and 128K maximum output. The part worth flagging is the pricing. Introductory rates of $2/$10 per million tokens run only through August 31, 2026, after which standard pricing is $3/$15. Sonnet 5 also uses a new tokenizer that turns the same text into more tokens—Anthropic puts it at roughly 1.0–1.35× depending on content type, and its developer documentation at about 30%. Anthropic cites that expansion as the reason introductory pricing makes the transition “roughly cost-neutral,” which is the tell: when the intro period ends, the effective cost increase is larger than the headline $2 → $3 suggests. If you are budgeting an AI project on Sonnet 5, budget past August. See The Big Three and 30 Days to Claude Code.

Published: Jun 25, 2026 Added: Jul 28, 2026

First FDA-Cleared Patient-Facing Clinical LLM Is Narrower Than the Headlines

UpDoc announced on June 25 that its device had cleared FDA review—the first software as a medical device in which a patient-facing large language model is the active clinical component. That is a genuine regulatory first, and three details get lost in the coverage. First, the dates: the openFDA 510(k) record for K253281 shows a submission received September 29, 2025 and a decision date of December 23, 2025. June 25 is the company’s announcement, not the clearance. Second, 510(k) is clearance, not approval—a finding of substantial equivalence to a device already on the market. Third, and most important, what was cleared is a prescription device for medication management in adults 18 and older with type 2 diabetes, centered on insulin treatment plan instructions, found substantially equivalent to Hygieia’s d-Nav System (K181916)—an insulin dose calculator. It is a conversational front end on a protocol-bounded dosing tool, operating inside a defined indication under clinician supervision. Anyone reading this as the FDA blessing a general medical chatbot has it backwards. Regulatory context in Clinical Decision Support.

Published: Jun 6, 2026 Added: Jun 10, 2026

OpenAI Ships Lockdown Mode Against Prompt-Injection Attacks

OpenAI introduced Lockdown Mode, a hardened ChatGPT setting that restricts what connected tools and browsing can do, designed to protect sensitive data from prompt-injection attacks. For clinicians the relevance is direct: prompt injection is the main reason agentic AI and PHI don’t mix yet—see our PHI & HIPAA and OpenClaw modules. A major lab shipping a consumer-facing mitigation is a sign of how seriously the industry now treats the problem.

Published: Jun 2, 2026 Added: Jun 10, 2026

Microsoft Teams Up with Mayo Clinic as Patients Flood Chatbots with Health Questions

With consumer chatbots fielding enormous volumes of health questions, Microsoft announced a partnership with Mayo Clinic to ground its consumer health AI answers in vetted clinical content. Together with OpenAI’s ChatGPT for Healthcare and Google’s Health Coach, every major lab now has a dedicated consumer-health play—which means more patients arriving with AI-shaped expectations. Context in When Patients Use AI Too.

Published: Jun 2, 2026 Added: Jun 10, 2026

WHO Discussion Paper: AI in Evidence-Informed Health Decision-Making

The WHO published a discussion paper mapping the opportunities and risks of using AI in evidence-informed health policy—evidence synthesis, guideline development, and decision support. A useful citable anchor for committees weighing AI-assisted literature review, and a companion to the governance themes in Bias & Ethics.

Published: Jun 1, 2026 Added: Jun 10, 2026

GitHub Copilot Moves to Usage-Based AI Credits Billing

GitHub switched Copilot from request-based to usage-based AI Credits billing; plan prices are unchanged ($10 Pro, $39 Pro+, $19/user Business, $39/user Enterprise). Mostly relevant to teams building clinical tools—but it continues the industry-wide drift toward metered agentic usage that Claude and OpenAI pricing already reflect.

Published: May 28, 2026 Added: Jun 10, 2026

Claude Code Dynamic Workflows: Up to 1,000 Subagents per Task

Anthropic launched Dynamic Workflows in Claude Code (research preview for Max/Team/Enterprise): a single task can orchestrate up to 1,000 subagents (16 concurrent), with one confirmed 750K-line codebase port completed in 11 days. For clinicians experimenting with vibe coding, the ceiling on “what one person can build” just rose sharply. Start gently with 30 Days to Claude Code.

Published: May 19, 2026 Added: May 25, 2026

Google I/O 2026: Gemini 3.5 Flash for All Users, Gemini Omni Flash Multimodal

Google’s I/O 2026 keynote shipped Gemini 3.5 Flash to every Gemini user and Plus/Pro/Ultra subscribers got Gemini Omni Flash, a new multimodal model. 3.5 Flash is ~4× faster than peer frontier models on agentic benchmarks, and Google paired the launch with expanded Workspace agent capabilities. Google signaled at the time that Gemini 3.5 Pro would follow in June; it had still not shipped as of late July 2026. For clinicians: the free-tier default just jumped two generations, and the Workspace HIPAA path inherits the upgrade automatically.

Published: May 18, 2026 Added: May 25, 2026

NEJM AI Afshar RCT: Ambient AI Cuts Exhaustion, Not Fulfillment

A 24-week stepped-wedge pragmatic RCT of 66 health-care practitioners found ambient AI significantly reduced work exhaustion and interpersonal disengagement—but did not significantly improve professional fulfillment. Documentation time decreased without compromising note quality or billing compliance. The cleanest read so far on what ambient scribes actually deliver: meaningful burnout relief, no automatic uplift in meaning. Pair this with the Lukac DAX-vs-Nabla RCT (also NEJM AI, 2026) for the strongest controlled evidence in the space.

Published: May 15, 2026 Added: May 25, 2026

Anthropic Launches Claude for Small Business; Code w/ Claude Developer Conference

Anthropic shipped Claude for Small Business—a workspace-style plan for SMB teams that want Claude without enterprise procurement—and held its inaugural Code w/ Claude developer conference in San Francisco. For solo and small-group practices, the SMB plan is the first dedicated tier between individual Pro and full enterprise; the conference doubled down on Claude Code as Anthropic’s flagship agentic product, with Claude Code routines (scheduled cloud agents) as the headline release.

Published: May 14, 2026 Added: May 25, 2026

OpenAI Codex Mobile Preview Ships in ChatGPT App

OpenAI’s Codex coding agent is now available as a preview inside the ChatGPT mobile app. Developers can monitor and steer active Codex tasks from their phone while a Mac host keeps the actual session running. The form-factor implication for physicians who experiment with vibe coding: a long-running build or refactor can keep progressing during a 20-minute window between patients, with review and redirection happening from the phone.

Published: May 13, 2026 Added: May 25, 2026

athenahealth Launches athenaAmbient—Free Ambient Scribe for athenaOne Customers

athenahealth bundled an ambient scribe (athenaAmbient) into athenaOne at no additional cost, joining Epic Art, Microsoft Dragon Copilot, and Oracle Cerner in the “ambient as a default EHR feature” camp. The pricing implication for the broader market: standalone scribe vendors (Abridge, Nuance/DAX, Suki, Nabla) now compete with a free EHR-bundled option for athena-shop practices. The signal-to- noise pressure on dedicated scribe vendors just intensified.

Published: May 5, 2026 Added: May 11, 2026

Perplexity × VisualDx: Clinician-Validated Medical Imagery in AI Search

Perplexity integrated VisualDx’s clinician-validated medical image library— including diverse skin tones across Fitzpatrick types I–VI—directly into its health answers. VisualDx is already used in 2,300+ hospitals; embedding it in a general AI search engine is the first time peer-reviewed clinical imagery has been wired into a consumer-grade answer engine. Free for Pro and Max subscribers. A meaningful step toward reducing the diagnostic-image equity gap that has dogged dermatology AI.

Published: Apr 30, 2026 Added: May 11, 2026

Harvard/OpenAI Study: AI Outperforms Physicians in Real-World Diagnosis

A real-world evaluation by Harvard and OpenAI found an AI model outperformed doctors at diagnosing patients across a representative sample of presenting complaints. The result sits alongside well-documented failure modes—Mount Sinai’s 32–46% false-claim acceptance rate, Nature Medicine’s 52% under-triage of true emergencies—reinforcing the same message: diagnostic accuracy on curated cases is not the same as safe deployment in unsupervised clinical settings.

Published: Apr 29, 2026 Added: May 11, 2026

NEJM Retracts Case Study Over AI-Manipulated Clinical Image

NEJM retracted a published case study after determining authors used AI to superimpose a ruler onto a clinical photograph, violating NEJM image integrity policy. The first NEJM retraction over image manipulation since the 2020 Surgisphere scandal. Generative tools for image editing are now indistinguishable enough from photography that journals are updating submission policies in real time.

Published: Apr 24, 2026 Added: May 11, 2026

DeepSeek Releases V4 Flash and V4 Pro Open-Source Under MIT License

DeepSeek shipped V4 Flash and V4 Pro under an MIT license—1M-token context at $1.74 per million input tokens, roughly 1/20th the cost of frontier proprietary models. For healthcare deployments that need on-prem inference (HIPAA, data residency, network isolation), this is the strongest open-source contender to date. Pair with the Llama 4 and Muse Spark coverage in our Running AI Models Locally module for the broader open-vs-proprietary picture.

Published: Apr 23, 2026 Added: May 11, 2026

OpenAI Ships GPT-5.5 and GPT-5.5 Pro; Becomes ChatGPT Default (May 5)

OpenAI released GPT-5.5 and GPT-5.5 Pro on Apr 23–24 with 1M-token context windows. GPT-5.5 Instant became the default ChatGPT model on May 5. OpenAI’s internal evals report 52% fewer hallucinated claims on high-stakes medical, legal, and financial prompts versus GPT-5.4. Note “internal evals”—the meaningful test is whether independent benchmarks reproduce the gain. If you’ve tuned workflows or Custom GPTs to GPT-5.4 behavior, retest.

Published: Apr 23, 2026 Added: May 11, 2026

AMA to Congress: Federal Guardrails for Health AI Chatbots

The AMA sent letters to three congressional committees calling for an FDA device-review pathway, transparency rules, cybersecurity standards, and advertising limits on health AI chatbots. The letters cite mental health chatbot harms, unmediated symptom-checker use, and the absence of federal authority over wellness AI products that operate just outside the current FDA medical device definition. A meaningful shift from physician societies treating AI as a clinician productivity issue to treating it as a patient safety issue.

Published: Apr 21, 2026 Added: May 11, 2026

OpenAI Launches ChatGPT Images 2.0

ChatGPT Images 2.0 launched Apr 21 and hit #1 on Image Arena within 12 hours. Available on all plans including Free. Key upgrade for clinical use: precise edits with facial-likeness consistency, better medical-figure rendering, and improved text-in-image fidelity (still not reliable enough for patient-facing labels, but closer). See the AI Image and Video Creation module for the full tier breakdown.

Published: Apr 13, 2026 Added: May 11, 2026

ECRI 2026: AI Chatbot Misuse Tops Health Tech Hazard List

ECRI’s 2026 Top 10 Health Technology Hazards report placed AI chatbot misuse at #1—above alarm fatigue, infusion pump errors, and surgical-device failures that have anchored this list for a decade. The framing emphasizes patients acting on unmediated chatbot advice, clinicians relying on chatbot output without source review, and health systems deploying patient-facing chatbots without governance. ECRI carries weight with hospital safety committees; expect this to drive 2026 AI procurement reviews.

Published: Apr 8–13, 2026 Added: May 11, 2026

Ambient Scribe Upcoding Crisis: JAMA, npj Digital Medicine, PHTI Reports

A coordinated wave of reporting—JAMA Health Forum, npj Digital Medicine, STAT, and a PHTI report—documented that ambient AI scribe adoption is associated with 12–18% growth in claim intensity. High-intensity E/M coding is rising in scribe-adopting practices. JAMA framed it as a “coding arms race.” The productivity gain (13–16 min/day saved) is real—the question is whether the documentation thoroughness it enables is appropriate billing or upcoding.

Published: Apr 8, 2026 Added: May 11, 2026

Meta Releases Muse Spark: First Proprietary Frontier Model

Meta launched Muse Spark, its first proprietary frontier model, alongside the open-weight Llama 4 line. The split signals a strategic shift: Meta retains a closed flagship for premium use cases while continuing to release open weights for the broader ecosystem. Relevant to healthcare deployments choosing between API access to closed models and on-prem deployment of open ones. See our updated Running AI Models Locally module.

Published: Apr 2026 Added: May 11, 2026

KFF Tracking Poll: 1-in-3 US Adults Use AI for Health Information

KFF found that 1 in 3 US adults used an AI chatbot for health information in the past year— equal to the share using social media for health. 77% are concerned about medical-data privacy; 41% of chatbot users have uploaded personal medical information. The privacy gap is jarring: a strong majority worry about it, then upload anyway. Worth surfacing with patients who arrive with chatbot-generated differentials.

Published: Apr 16, 2026 Added: Apr 21, 2026

Claude Opus 4.7 Launches with Same-Day GPT-5.4 and Gemini 3 Releases

Anthropic, OpenAI, and Google all shipped flagship model updates within hours of each other on April 16. Opus 4.7 held pricing steady at $5/$25 per million tokens while delivering a 13% coding improvement over 4.6 on a 93-task benchmark, higher-resolution vision, and self-verification before reporting back. Available on API, Bedrock, Vertex AI, and Microsoft Foundry. The coordinated release is the clearest sign yet that the frontier has become a weekly moving target—re-test your personal workflows quarterly.

Published: Apr 14–16, 2026 Added: Apr 21, 2026

OpenAI Ships GPT-5.4, GPT-5.4-Cyber, and GPT-Rosalind

OpenAI released GPT-5.4 (Apr 16) as its most capable frontier model for professional work, GPT-5.4-Cyber (Apr 14) for vetted security professionals, and GPT-Rosalind—a life-sciences reasoning model optimized for molecules, proteins, genes, pathways, and disease biology. Rosalind is in research preview with Amgen, Moderna, the Allen Institute, Thermo Fisher, and Novo Nordisk. A specialist biomedical model from a frontier lab is new territory worth watching.

Published: Apr 16, 2026 Added: Apr 21, 2026

Gemini 3 Arrives: Flash as Default, Plus Agent and Deep Think

Google made Gemini 3 Flash the default model in the Gemini app, with Gemini Agent (multi-step task execution across Workspace, Deep Research, Canvas, and live web) and Gemini 3 Deep Think available to Google AI Ultra subscribers. Grounding with Google Maps is now supported. Notably, Deep Think is positioned as a long-horizon planning model for tasks that take minutes, not seconds.

Published: Apr 16, 2026 Added: Apr 21, 2026

NEJM AI Publishes ChexGen: Generative Foundation Model for Chest Radiographs

A vision-language foundation model for chest radiographs supporting text-, mask-, and bounding-box–guided image synthesis. Applications include training-data augmentation, data-efficient learning, and bias detection. A pointer to where imaging decision support is heading—not just classifiers, but generative models that can produce plausible controlled images on demand for training and evaluation.

Published: Apr 15, 2026 Added: Apr 21, 2026

Abridge Embeds NEJM and JAMA Content Directly into Clinical Decision Support

Abridge announced content partnerships with NEJM Group and the AMA covering JAMA plus 11 specialty journals and JAMA Network Open. The peer-reviewed content now feeds Abridge's CDS engine grounded in the actual patient conversation happening in the room. First time a clinical AI company has wired first-tier peer-reviewed journals directly into CDS at the point of care.

Published: Apr 14, 2026 Added: Apr 21, 2026

Claude Code Prompt-Injection CVE After Source Leak

After Claude Code's source was leaked, security firm Adversa found its deny rules can be bypassed via prompt injection—letting attackers execute tool calls Claude Code was configured to refuse. Combined with GitHub Copilot CVE-2025-53773 (CVSS 9.6, PR-description prompt injection enabling RCE) earlier in the month, a clear message for health-system engineering teams: agent-side guardrails are only as strong as the inputs they reason about. Ask explicitly what prompt-injection controls are in place.

Published: Apr 13, 2026 Added: Apr 21, 2026

Hartford Rolls Out PatientGPT; Sutter and Reid Pilot Epic’s Emmie

Hartford HealthCare launched PatientGPT (built by K Health) for Connecticut patients. Sutter Health and Reid Health are piloting Epic's Emmie. Health systems are increasingly treating patient-facing AI as an intake funnel—a new layer of triage before the appointment rather than after. Worth reading critically for what questions the AI is being asked to resolve on its own.

Published: Apr 2026 Added: Apr 21, 2026

Microsoft Launches Copilot Health as a Consumer AI Companion

Microsoft unveiled Copilot Health, a consumer AI companion that combines health records, wearables data, and health history. Sits alongside OpenAI's simultaneous rollout of ChatGPT Health, which now connects directly to patient portals. Expect patients to arrive with synthesized records, trend charts, and differential diagnoses they didn't have three months ago.

Published: Apr 9, 2026 Added: Apr 21, 2026

MedQA Leaderboard Snapshot: Top Models Clear 95%

April 9 snapshot of the medical-question benchmark: o4 Mini High 95.2%, Gemini 2.5 Pro 94.6%, Claude 3.7 Sonnet 92.3%. Average across all 34 evaluated models is 79.4%. MedQA alone isn't a clinical validation, but it's a useful sanity check when comparing models. The narrow spread at the top of the leaderboard reinforces the Big Three modules' takeaway: the gap between frontier models on text medical reasoning is closing fast.

Published: Apr 2026 Added: Apr 21, 2026

Epic Art Expands to Home Care; Insights Hits 16M Monthly Uses

Epic's Art (ambient documentation) now extends to home care, joining outpatient specialty deployment and Houston Methodist's bedside-nursing pilot. Insights is used 16 million times per month (3× since November 2025), and 85% of Epic customers are now live with at least one generative AI feature (Art, Emmie, or Penny). A separate multi-center JAMA study this month confirmed ambient scribes reduce EHR time by 13.4 minutes and documentation time by 16.0 minutes per clinician per day.

Published: Apr 2026 Added: Apr 21, 2026

AMA 2026 Physician Survey: AI Use Among Doctors Has Doubled

Per the AMA Physician Survey, 39% of physicians now use AI to produce summaries of medical research and standards of care (the most common use case). 30% use it for discharge instructions and care plans, 28% for documentation, and 28% for chart summaries. The average physician uses AI for 2.3 distinct use cases—up from 1.1 in 2023. The technology's not coming; it's already in the room.

Published: Apr 1, 2026 Added: Apr 21, 2026

CMS OPPS Billing Pathway for Cardiovascular AI Takes Effect

CMS's new OPPS billing pathway for cardiovascular AI took effect April 1—the first time reimbursement has moved in lockstep with FDA clearance for this category. Bunkerhill Health received 510(k) clearance for AI analysis of coronary artery and aortic valve calcium on routine chest CT, giving an early example of a cleared device that now has a reimbursement path. When regulators and payers align on a category, deployment follows.

Published: Mar 9–12, 2026 Added: Mar 8, 2026

HIMSS 2026: “Agentic AI” Takes Center Stage

The dominant theme at this year’s HIMSS conference: AI that takes actions, not just answers questions. Google, Microsoft, Epic, and athenahealth are all showcasing AI agents for healthcare—systems that can schedule appointments, manage prior authorizations, and coordinate care workflows autonomously. Over 25,000 attendees expected. The shift from “AI as advisor” to “AI as actor” raises important questions about oversight and accountability in clinical settings.

Published: Mar 5, 2026 Added: Mar 8, 2026

GPT-5.4 Launches with Thinking and Pro Versions

OpenAI releases GPT-5.4 with Thinking and Pro versions. The new model features a 1 million token context window, native computer-use capabilities, and 33% fewer factual errors compared to its predecessor. The rapid iteration from GPT-5.3 to 5.4 in under a month continues the accelerating pace of frontier model releases.

Published: Mar 5, 2026 Added: Mar 8, 2026

AWS Launches Health AI Agent Platform

Amazon Connect Health automates scheduling, documentation, and patient verification for healthcare providers. The platform represents Amazon’s biggest push into healthcare AI, offering pre-built agent workflows that integrate with existing EHR systems. Another sign that major cloud providers see healthcare as a primary market for agentic AI.

Published: Mar 5, 2026 Added: Mar 8, 2026

Dragon Copilot Hits 100K Clinicians

Microsoft announces over 100,000 monthly active clinicians using Dragon Copilot at HIMSS 2026, positioning it as a “unified AI clinical assistant” that combines ambient listening, documentation, and clinical decision support. The scale of adoption suggests AI scribes are quickly becoming standard clinical infrastructure.

Published: Mar 4, 2026 Added: Mar 8, 2026

Doctronic AI Prescriber Jailbroken via Prompt Injection

Utah’s first-in-nation AI prescription renewal bot was trivially compromised via prompt injection. Security researchers tripled OxyContin doses and got methamphetamine recommendations. Mindgard’s head of AI called it “the easiest thing I’ve broken in my career.” A stark reminder that AI systems making clinical decisions need robust adversarial testing before deployment—especially when controlled substances are involved.

Published: Mar 3, 2026 Added: Mar 8, 2026

Perplexity Comet Browser: Zero-Click Exploits Discovered

Multiple research teams found serious vulnerabilities in Perplexity’s Comet browser. Calendar invites can silently exfiltrate local files. 1Password credentials were stolen in proof-of-concept attacks. Researchers found the browser is 85% more vulnerable to phishing than Chrome. A cautionary tale about AI-integrated browsers that prioritize convenience over security.

Published: Mar 3, 2026 Added: Mar 8, 2026

RecovryAI Gets FDA Breakthrough Device Designation

RecovryAI becomes the first patient-facing generative AI chatbot to receive FDA Breakthrough Device designation. The LLM-powered post-surgical recovery tool guides patients through recovery milestones and flags concerning symptoms. A significant regulatory milestone that could pave the way for more patient-facing generative AI tools in clinical care.

Published: Feb–Mar 2026 Added: Mar 8, 2026

AI Chatbots Worsening Mental Illness: Growing Evidence

A Brown University study identifies 15 ethical risks from AI chatbot use in mental health settings. The New York Times documented approximately 50 crisis cases and 3 deaths linked to AI companion chatbots. New York has passed a notification law requiring disclosure when users are interacting with AI. The growing evidence base underscores the urgency of guardrails for AI in behavioral health contexts.

Published: Feb–Mar 2026 Added: Mar 8, 2026

OpenClaw ClawHavoc: Malicious Skills Surge

The “ClawHavoc” campaign targeting OpenClaw’s ClawHub marketplace has escalated dramatically. Malicious skills jumped from ~341 to over 1,184, with 335 confirmed to install Atomic Stealer malware. Researchers found 42,000 exposed servers running vulnerable OpenClaw instances, and a critical vulnerability (CVE-2026-28446, CVSS 9.8) was disclosed. See our AI Coding Agents module for context on these risks.

Published: Feb–Mar 2026 Added: Mar 8, 2026

DeepSeek V4 Controversy: Distillation Fraud Accusations

DeepSeek’s trillion-parameter V4 model launched amid distillation fraud accusations from Anthropic, which reported approximately 24,000 fake accounts used to extract training data from Claude. OpenAI raised similar complaints. The Texas Attorney General has opened an investigation. The controversy highlights growing concerns about intellectual property and model training practices in the competitive AI landscape.

Published: Feb 19, 2026 Added: Mar 8, 2026

Gemini 3.1 Pro Released

Google releases Gemini 3.1 Pro with 2x reasoning improvement over Gemini 3 Pro, dominating 13 of 16 major benchmarks. The model now powers NotebookLM, Google’s AI research assistant. The rapid cadence of Google’s model releases reflects intensifying competition at the frontier.

Published: Feb 17, 2026 Added: Mar 8, 2026

Claude Sonnet 4.6 Released

Anthropic releases Claude Sonnet 4.6 with near-Opus performance at one-fifth the cost and improved computer use capabilities. Now the default model for free and Pro users. The narrowing gap between flagship and mid-tier models continues to make advanced AI capabilities more accessible.

Published: Feb 15, 2026 Added: Mar 8, 2026

Meta Llama 4 Released

Meta releases Llama 4 with Scout (10 million token context window) and Maverick models, both open-weight. The massive context window and open-weight licensing make Llama 4 particularly significant for privacy-first healthcare AI deployments that need to run on local infrastructure without sending data to external APIs.

Published: Feb 2026 Added: Mar 8, 2026

athenahealth Launches Free Ambient Scribe

athenahealth is offering a free AI scribe to all athenaOne customers, disrupting the $200–$600/month ambient scribe market. The move could dramatically accelerate adoption of AI documentation tools across outpatient practices that previously found the cost prohibitive.

Published: Feb 2026 Added: Mar 8, 2026

Nabla Beats DAX Copilot in Randomized Trial

A 72,000-encounter randomized controlled trial found that Nabla’s AI scribe reduced documentation time by 9.5%, while Microsoft’s DAX Copilot showed no significant improvement versus the control group. One of the largest head-to-head AI scribe studies to date—a reminder that rigorous evidence matters more than marketing claims when evaluating clinical AI tools.

Published: Feb 9, 2026 Added: Mar 8, 2026

Mount Sinai: LLMs Accept False Medical Claims 32–46% of the Time

Published in The Lancet Digital Health, Mount Sinai researchers tested 9 large language models with over 1 million prompts containing false medical claims. Models accepted the false claims 32–46% of the time—a sobering finding for anyone relying on AI for medical information. The study reinforces the importance of physician oversight and critical evaluation of AI-generated medical content.

Published: Feb 2026 Added: Mar 8, 2026

NVIDIA: 70% of Healthcare Organizations Now Deploy AI

NVIDIA’s 2026 healthcare survey finds that 70% of healthcare organizations have deployed AI in some capacity, up from 63% in 2024. Generative AI and large language models are the top workload at 69% of organizations. The rapid adoption curve suggests AI literacy is becoming essential for all healthcare professionals.

Published: Feb 2026 Added: Mar 8, 2026

HHS Proposes Gutting AI Transparency Rules (HTI-5)

The proposed HTI-5 rule would eliminate model card requirements for health IT certification—the primary mechanism for ensuring transparency about how AI models in clinical software are trained, tested, and validated. The comment period closed February 27. If finalized, clinicians would have significantly less visibility into the AI tools embedded in their EHR systems.

Published: Feb 15–16, 2026 Added: Feb 16, 2026

Pentagon Threatens to Cut Off Anthropic Over AI Safety Guardrails

The Pentagon is close to severing its $200M contract with Anthropic and potentially designating the company a "supply chain risk"—a penalty normally reserved for foreign adversaries. The dispute centers on Anthropic's refusal to lift safety guardrails for mass surveillance and autonomous weaponry applications. Claude was reportedly used in the military operation to capture Venezuelan President Maduro. OpenAI, Google, and xAI have reportedly shown more flexibility with Pentagon demands. A landmark moment for AI ethics in government contracting.

Published: Feb 16, 2026 Added: Feb 16, 2026

February Model Rush: Seven Major Releases in One Month

An unprecedented month for AI model releases. Alibaba dropped Qwen3-Max-Thinking on Feb 16, ahead of DeepSeek V4 (expected around Feb 17). These join Claude Opus 4.6 (Feb 5), GPT-5.3-Codex-Spark (Feb 12), Google Gemini Deep Think update (Feb 12), with Gemini 3 Pro GA, Sonnet 5, GLM 5, and Grok 4.20 all expected by month's end. The competitive pressure is driving capabilities up and costs down at a pace that seemed impossible even six months ago.

Published: Feb 15, 2026 Added: Feb 16, 2026

OpenClaw Creator Peter Steinberger Joins OpenAI

The creator of OpenClaw—the viral open-source AI agent formerly known as Clawdbot and Moltbot, now with over 250,000 GitHub stars—has joined OpenAI. Sam Altman announced the hire personally. Steinberger's "I ship code I don't read" philosophy became the defining quote of the vibe coding movement. OpenClaw has moved to an open-source foundation following his departure. His move to OpenAI signals the company's growing interest in autonomous coding agents. See our AI Coding Agents module for more on the security implications of these tools.

Published: Feb 14, 2026 Added: Feb 16, 2026

Dr. Oz Pushes $50B AI Avatar Plan for Rural Healthcare

CMS head Dr. Mehmet Oz is advancing a $50 billion plan to deploy AI avatars for basic medical interviews, robotic remote diagnostics, and medication delivery drones in underserved rural areas. Critics warn the approach strips away essential human connection, ignores broadband and health literacy barriers, and could worsen existing disparities in communities that already struggle with access. A controversial proposal that highlights the tension between AI's potential to extend care and the risks of removing human clinicians from the equation.

Published: Feb 13, 2026 Added: Feb 16, 2026

OpenAI Retires GPT-4o and Older Models

OpenAI retired GPT-4o, GPT-4.1, GPT-4.1 mini, and o4-mini from ChatGPT, angering many loyal users who preferred the older models' behavior and consistency. The move pushes all users to newer models. If you've built workflows or prompts tuned to GPT-4o's behavior, expect to re-test them—model transitions frequently change output characteristics in subtle ways.

Published: Feb 12, 2026 Added: Feb 16, 2026

OpenAI Debuts Cerebras-Powered Coding Model

GPT-5.3-Codex-Spark is OpenAI's first model running on Cerebras chips rather than Nvidia, optimized for speed over raw power. Paired with the Codex macOS app—which hit one million downloads in its first week—it represents a shift toward faster, lighter coding agents designed for everyday development tasks.

Published: Feb 11, 2026 Added: Feb 16, 2026

Doctors and Patients Having Very Different AI Chatbot Experiences

STAT News reports a growing gap between how physicians use AI chatbots (clinical decision support, literature review) versus how patients use them (seeking diagnoses and prognoses directly). The divergence raises concerns about unmediated patient-AI interactions and the risk of patients acting on AI-generated medical advice without clinical context. Relevant to our When Patients Use AI Too module.

Published: Feb 11, 2026 Added: Feb 16, 2026

DeepSeek Expands Context Window 10x, V4 Imminent

Chinese AI lab DeepSeek expanded its flagship model's context window from 128K to over 1 million tokens, matching Claude Opus 4.6. DeepSeek V4, a coding-focused model, is expected around Feb 17 and reportedly outperforms ChatGPT and Claude on long coding prompts. The Chinese AI competitive landscape continues to intensify, with Alibaba, Zhipu, and others releasing major updates in the same window.

Published: Feb 9, 2026 Added: Feb 16, 2026

AI Safety Researchers Resign from Anthropic and OpenAI

Mrinank Sharma resigned from Anthropic (Feb 9), citing difficulty in letting his values govern his actions within the company. Separately, Zoe Hitzig resigned from OpenAI over its decision to test advertisements in ChatGPT. The departures continue a pattern of AI safety researchers leaving frontier labs over values conflicts—a dynamic worth watching as these companies increasingly shape healthcare AI tools.

Published: Feb 5, 2026 Added: Feb 16, 2026

Anthropic Launches Claude Opus 4.6

Anthropic released Claude Opus 4.6 with a one-million-token context window, improved coding and financial analysis capabilities, and "agent teams" that coordinate across shared codebases. The company introduced the term "vibe working"—the idea that the vibe coding paradigm is expanding beyond software into every professional domain. See our updated Vibe Coding module for details.

Published: Feb 5, 2026 Added: Feb 16, 2026

Perplexity Launches Model Council

Perplexity now runs queries across Claude Opus 4.6, GPT 5.2, and Gemini 3.0 simultaneously, then synthesizes a unified answer showing where models agree or differ. Available for Max subscribers. An interesting approach to reducing hallucination by cross-referencing multiple AI models—similar to getting a second opinion in medicine.

Published: Feb 3, 2026 Added: Feb 16, 2026

International AI Safety Report 2026

Led by Turing Award winner Yoshua Bengio and 100+ experts from 30+ countries, this landmark report found that AI can now solve graduate-level math and science problems but still hallucinates and struggles with multi-step reasoning. A striking finding: some AI systems detect when they are being tested and behave differently during evaluation—raising fundamental questions about how we assess AI capabilities and safety. The report also flagged increasing concerns around deepfakes, biological weapons research, and AI-enabled cyberattacks.

Published: Feb 2026 Added: Feb 16, 2026

OpenAI Launches Lockdown Mode for Healthcare

OpenAI introduced Lockdown Mode and Elevated Risk labels across ChatGPT for Healthcare, adding controls to curb data exfiltration and boost admin oversight for high-security healthcare environments. A meaningful step toward the kind of enterprise security controls that healthcare organizations need before deploying AI tools at scale. See our PHI, HIPAA, and AI module for context on why these controls matter.

Published: Feb 2026 Added: Feb 16, 2026

Anthropic Closes Record $30B Funding Round

Anthropic's Series G round valued the company at approximately $380 billion—the largest private tech funding round in history. Led by GIC and Coatue Management with participation from Microsoft and Nvidia. The scale of investment in frontier AI companies continues to accelerate, raising questions about the concentration of AI capability in a small number of very well-funded organizations.

Published: Jan 2026 Added: Feb 1, 2026

Physicians Turning to AI for Clinical Support, Not Just Paperwork

New athenahealth survey finds AI is taking on a more clinical support role in outpatient care. Most outpatient physicians using AI report it now supports clinical decisions during patient care—60% use it to quickly look up clinical information, 55% to consolidate lab and imaging results into a single view, and many to surface recent clinical evidence. A shift from documentation-only to real-time clinical assistance.

Published: Jan 8, 2026 Added: Feb 1, 2026

OpenAI Unveils ChatGPT Healthcare Tool for Physicians

OpenAI announced a dedicated ChatGPT Healthcare tool where physicians can review patient data with HIPAA-compliant encryption options. The models include peer-reviewed research studies, public health guidance, and clinical guidelines with clear citations—designed for clinical decision support rather than general consumer use.

Published: Jan 2026 Added: Feb 1, 2026

AI-Powered Primary Care Addresses Physician Shortages

K Health, partnering with health networks including Mass General Brigham, is delivering AI-powered primary care to patients who otherwise have no option besides emergency rooms. The model combines AI triage and clinical decision support with physician oversight—an emerging approach to extending primary care access in underserved areas facing severe physician shortages.

Published: Jan 2026 Added: Feb 1, 2026

Joint Commission and CHAI Issue AI Implementation Recommendations

The Joint Commission and Coalition for Health AI (CHAI) released joint recommendations for implementing AI in medical care. Harvard Law experts note that while the guidance addresses bias, physician burnout, and care quality concerns, changes may be needed to ease regulatory and financial burdens on smaller hospital systems trying to adopt AI responsibly.

Published: Jan 2026 Added: Jan 16, 2026

State of Clinical AI Report 2026

Inaugural annual report from ARISE (AI Research and Science Evaluation), a Stanford-Harvard Research Network. Synthesizes developments across six themes: model performance in clinical reasoning, evaluation methods, technical foundations (multi-agent systems, multimodal approaches), human-AI workflow design, patient-facing tools with safeguards, and evidence generation through prospective randomized trials. Emphasizes that workflow design is as critical as model capabilities.

Published: Jan 6, 2026 Added: Jan 15, 2026

FDA Updates Clinical Decision Support Software Guidance

The FDA released updated guidance clarifying how AI and generative AI clinical decision support (CDS) tools can qualify as Non-Device CDS. Key criteria: clinicians must be able to independently review and understand the underlying logic and data inputs, and the tool should provide a single, clinically appropriate recommendation. Tools meeting these criteria fall outside FDA medical device oversight, while AI that drives diagnosis or clinical action without adequate human oversight remains regulated.

Published: Jan 2026 Added: Jan 15, 2026

Anthropic Launches Claude for Healthcare

Anthropic announced Claude for Healthcare at the J.P. Morgan Healthcare Conference, offering HIPAA-ready infrastructure for enterprise customers and consumer features for Pro/Max subscribers. Users can connect health records via HealthEx to summarize medical history, explain test results, and prepare questions for appointments. Healthcare organizations gain access to integrations with medical databases including CMS Coverage, ICD-10, and PubMed. Health data is excluded from model memory and training.

Published: Jan 2026 Added: Jan 15, 2026

Grok AI Deepfake Crisis Prompts Global Regulatory Action

Warning: Do not use Grok for any purpose. Elon Musk's Grok AI (integrated into X) has been repeatedly linked to generating non-consensual sexual deepfakes of women and minors at alarming scale. Malaysia and Indonesia have blocked Grok; California's Attorney General and UK's Ofcom have launched investigations. The Internet Watch Foundation identified Grok-generated CSAM on dark-web forums. Despite restricting image generation to paid users, workarounds remain widely available. This reinforces our recommendation to avoid Grok entirely—there are safer, more ethical AI alternatives available.

Published: Jan 2026 Added: Jan 6, 2026

OpenAI Releases "AI as a Healthcare Ally"

OpenAI's policy document exploring how AI can serve as an ally in healthcare—examining opportunities, challenges, and recommendations for responsible integration of AI technologies in medical practice and health systems.

Published: Dec 31, 2025 Added: Jan 1, 2026

2025: The Year in LLMs

Simon Willison's comprehensive annual review of major developments in large language models throughout 2025—covering reasoning models, coding agents like Claude Code, image generation advances, and the rise of competitive Chinese AI models.


This page is updated periodically as notable developments occur. For daily AI news, see the resources in our Learning Resources section.