Ontario Audit: 60% of AI Scribe Systems Mix Up Drugs, 85% Miss Mental Health Details

The Office of the Auditor General of Ontario audited 20 approved AI Scribe systems used by physicians and nurse practitioners, simulating doctor-patient recordings to evaluate accuracy. The results are stark:
- 12 of 20 systems inserted incorrect drug information into patient notes.
- 9 of 20 fabricated information — e.g., claiming “no masses found” or “patient anxious” — that was never discussed.
- 17 of 20 missed key mental health details from the recording.
- 6 of 20 fully or partially omitted mental health issues.
The audit also slammed the evaluation scoring methodology. Accuracy of medical notes accounted for just 4% of the total score, while having a domestic presence in Ontario contributed 30%. Bias controls, threat/risk/privacy assessments, and SOC 2 Type 2 compliance each counted only 2–4%. As the report states, such weightings “could result in the selection of vendors whose AI tools may produce inaccurate or biased medical records.”
While OntarioMD has recommended manual review of AI notes, the audit noted no mandatory attestation feature in any approved system. Ontario’s Ministry of Health said over 5,000 physicians use these tools with no reported patient harm.
📖 Read the full source: HN AI Agents
👀 See Also

Developer's Dilemma: National Security Concerns Limit Open Model Choices
A developer working with security-sensitive clients reports being forced to choose between outdated U.S. open models like gpt-oss-120b or more capable Chinese models like GLM and MiniMax, which clients reject as national security risks.

BMW Dealership Revokes Buyback Offer After AI Chatbot Mistake, Precedent from Air Canada Case
A Toronto BMW dealership revoked a buyback offer generated by its AI chatbot, sparking legal questions. The Air Canada precedent holds companies liable for chatbot errors.

MiniMax M2.7 Model Shows Strong Performance as AI Coding Agent
A developer tested MiniMax M2.7 as their main AI coding agent and found it outperformed GPT 5.4 and Gemini 3.1 Pro in speed and tooling tasks, with benchmark scores of 56.22% on SWE-Pro and 57.0% on Terminal Bench 2.

Anthropic Copyright Settlement Details for Developers
Anthropic settled a $1.5 billion copyright class action over using works to train AI models. Eligible copyright owners can claim $500–$3,000 per validated work with a March 23, 2026 deadline.