Research

Content Access Is Not Content Control: What OpenEvidence's Open-Book Failure Reveals About Medical AI

John C. Ferguson, MD, FACS·March 6, 2026·8 min read

In December 2025, researchers demonstrated that OpenEvidence, a medical AI platform backed by $210 million in funding and used by over 400,000 healthcare professionals, achieves 100% accuracy on USMLE-style licensing questions. Impressive. So we ran an experiment.

We gave OpenEvidence an open-book exam. The answers were in their own database. They still failed.

The Experiment

We generated 268 board-level certification questions using EdAI's Content-Driven Intelligence engine from specific medical articles across seven specialties. These weren't trick questions or edge cases, they were standard certification-level items derived from published, peer-reviewed content.

To make it fair, we excluded 83 questions where we couldn't verify that OpenEvidence had access to the source material. The remaining 185 questions tested OpenEvidence under open-book conditions: the system had full access to the articles containing the answers.

The Results

Overall accuracy: 68.65%. The passing threshold for medical certification is 70%.

The specialty breakdown tells the real story:

| Specialty | Accuracy | |---|---| | Obesity Medicine | 79.2% | | Cosmetic Surgery (ABCS) | 75.0% | | Pediatric Neurosurgery | 75.0% | | Gynecologic Oncology | 70.6% | | Urologic Oncology | 65.0% | | Laser Surgery (ABLS) | 56.3% | | Facial Cosmetic Surgery (ABFCS) | 55.0% |

A 24-point spread across specialties. Two specialties, Laser Surgery and Facial Cosmetic Surgery, failed nearly half of all questions despite having the source material available. This isn't a model that occasionally stumbles. It's a model that fundamentally cannot reason through specialty board-level complexity, even when you hand it the answers.

Why This Matters

The distinction this experiment reveals is the most important one in medical AI: content access is not content control.

Any sufficiently large language model can retrieve medical content. RAG pipelines can surface relevant passages from indexed articles. Semantic search can find the right paragraph. None of that is the hard part.

The hard part is what happens after retrieval. Board-level certification questions don't ask "what does this article say?" They ask: given this clinical scenario, with these patient factors, and these competing considerations, what is the correct course of action, and why are the three plausible alternatives wrong?

That requires synthesis across concepts, application to novel clinical contexts, and the ability to reason through why superficially correct alternatives are actually incorrect. It requires understanding the content deeply enough to construct valid assessment of it, not just locate it.

OpenEvidence can find the content. It cannot control it.

What CDI Does Differently

EdAI's Content-Driven Intelligence doesn't just retrieve content and generate questions about it. CDI formalizes assessment as an algorithmic process:

Assessment_Item = f(Content_Variables, Cognitive_Process, Context_Parameters)

The system doesn't search for answers and reverse-engineer questions. It understands the cognitive structure of the content, what concepts are being tested, at what level of reasoning, in what clinical context, and generates assessment items that require authentic clinical reasoning to answer correctly.

The distractors aren't random wrong answers. They're diagnostically constructed alternatives that represent common misconceptions, incomplete reasoning paths, and contextually plausible but ultimately incorrect clinical decisions. Each wrong answer is a window into a specific reasoning failure.

This is why a $210 million platform with full access to the source material still fails: our questions test reasoning about the content, not retrieval of it. The gap between retrieval and reasoning is the gap between a search engine and a board examiner. CDI operates on the examiner side.

Implications for the Certification Market

For medical boards evaluating AI-assisted assessment tools, this experiment provides a clear decision framework:

General-purpose medical AI, no matter how well-funded, how widely adopted, or how impressive on standardized licensing exams, cannot serve the professional certification market. USMLE tests foundational medical knowledge. Board certification tests specialty-specific clinical reasoning. These are fundamentally different cognitive demands, and a model that excels at one may fail entirely at the other.

The certification market requires AI that is constrained to validated, specialty-specific content sources. It requires AI that generates assessment at the cognitive complexity level boards demand. And it requires AI that is accountable to the professional standards those boards enforce.

Content access is table stakes. Content control is the product.

The Full Study

The complete methodology and results are available in our published study.

Download the Full Study (PDF)


John C. Ferguson, MD, FACS, is the founder and CEO of EdAI Systems, a quintuple board-certified cosmetic and facial plastic surgeon, Co-Editor-in-Chief of StatPearls, and an AI2030 Global Fellow for Healthcare AI Governance. He can be reached at john@edaisystems.com.

See it against your own content.

Demos are live and specific: your board, your school, your exam, your route.

Request a demo