Thursday, September 10, 2026

Cortico Launches Benchmark to Test Clinical AI Safety

Must Read

Vancouver healthcare technology company Cortico has launched a free, open benchmark designed to measure whether artificial intelligence models can safely support clinical decisions—not merely pass medical exams.

Called MedSafe-Dx, the benchmark evaluates how large language models respond when a patient may require urgent care, when reassurance could be dangerous, and when the available information warrants uncertainty.

Across 11 frontier models evaluated in its launch paper, Cortico found that even the strongest performers faced a significant trade-off between safety and usefulness. The model with the highest safety pass rate, GPT-5.2, passed 97.6% of cases but escalated 71% of routine cases. At the other end, the poorest performer missed 26 of 156 cases classified as urgent.

“High accuracy often masks dangerous overconfidence,” said Cortico CEO and co-founder Clark Van Oyen. “We built MedSafe-Dx to give clinicians and health systems transparent, verifiable safety metrics.”

MedSafe-Dx tests three behaviours: whether a model escalates potentially life-threatening cases, avoids falsely reassuring patients who may be at risk, and expresses appropriate uncertainty when symptoms are ambiguous.

The evaluation presented 250 simulated adult patient cases from the DDXPlus dataset to models developed by OpenAI, Anthropic, Google and DeepSeek. Rather than relying on another AI model to judge the responses, MedSafe-Dx uses deterministic rules, frozen datasets and standardized outputs intended to make results reproducible and auditable.

The benchmark revealed that diagnostic accuracy did not necessarily translate into safer recommendations. Gemini 3 Pro Preview recorded the highest Top-3 diagnostic recall at 87.2%, according to Cortico, but the lowest safety pass rate at 62.4%. Every model evaluated missed at least some cases categorized as requiring escalation.

The findings point to a difficult balance for healthcare AI developers. Models that escalate nearly every questionable case may avoid some dangerous misses, but excessive warnings can burden clinical resources and eventually be ignored—a problem already familiar to providers using electronic health record alerts.

Cortico has since expanded the live leaderboard to 12 models from six AI labs, adding models from Meta and xAI. The company has also made the benchmark’s code and dataset publicly available alongside a medRxiv preprint.

Cortico stresses that MedSafe-Dx is a comparative safety test, not a clinical validation study or evidence that any model is ready for deployment. Its cases are simulated, while escalation labels are derived from dataset severity ratings rather than assessments by practising clinicians.

Still, the company argues that health systems evaluating AI should demand safety testing that examines judgment, confidence and escalation behaviour—not rely solely on exam-style accuracy scores.

Founded in 2015, Cortico provides patient-engagement and healthcare-workflow automation technology to more than 600 clinics and thousands of providers across North America.

The post Cortico Launches Benchmark to Test Clinical AI Safety appeared first on Techcouver.com.


Cortico Launches Benchmark to Test Clinical AI Safety was first posted on September 9, 2026 at 5:02 am.
©2022 “Techcouver.com“. Use of this feed is for personal non-commercial use only. If you are not reading this article in your feed reader, then the site is guilty of copyright infringement. Please contact me at rob@kitsilano.ca
 

- Advertisement -spot_img
- Advertisement -spot_img
Latest News

Nintendo’s latest Switch 2 update adds VRR support in TV mode

Nintendo always promised that the Switch 2 would support variable refresh rate (VRR) when docked, but the console launched...
- Advertisement -spot_img

More Articles Like This

- Advertisement -spot_img