EdAI Suites · For certification boards
Certification infrastructure, from designated sources to defensible exams.
Item generation, learning management, oral examination simulation, MOC portfolio, and secure testing, built on your board’s own designated content. Live in high-stakes examinations since 2025.
On the right: the reviewer’s chair. A drafted item arrives triaged, with its source passage alongside. Decide, and watch the record.
Review queue · ranked by classifier
You are the expert reviewer
Triage rank 1 of 12. Flag: verify the timing window in option A against the source passage.
A 44-year-old woman undergoes tumescent liposuction of the abdomen and flanks. Eighteen hours later she develops dyspnea, tachycardia, confusion, and a petechial rash over the chest and axillae. Which of the following is the most likely diagnosis?
- A.Pulmonary thromboembolism
- B.Fat embolism syndromekey
- C.Transfusion-related acute lung injury
- D.Local anesthetic systemic toxicity
Designated source · board reading list · claim verified against source, 2 of 2
Fat embolism syndrome typically presents 12 to 72 hours after long-bone trauma or liposuction with the classic triad of hypoxemia, neurologic impairment, and a petechial rash; thromboembolism rarely produces petechiae and classically presents later in the postoperative course.
Audit trail
Source registered
Designated source ingested; claims extracted with verbatim provenance
Item drafted
Generated from certified claims; style and difficulty per blueprint
Classifier triage
Ranked 1 of 12 for review; option A flagged for source verification
The classifier ranked the queue and raised the flag. It never gated, and it never decided. You did.
- 87%
- Reduction in expert physician time per item: 15 to 20 minutes against 2 to 4 hours traditionally
- 0
- AI-attributed factual errors in production
- 0
- Validity challenges from regulators or accreditors on deployed items
- 400+
- Validated items delivered in the 2025 examination cycle, at 100% client retention
From the 2025 production implementation with a US specialty certification board, measured against traditional volunteer item-writing benchmarks.
The Liability Rule
Every consequential decision is made by a human. It is an axiom, not a policy.
Failures in medical AI pipelines are characteristically silent: a wrong answer looks like a right answer, and a corrupted corpus looks like a clean one. The human reviewer is the only reliable detector for the failure class that matters most. So classifiers and automated checks triage and order expert attention, and never gate or decide. The chair who must defend an examination keeps the authority to defend it.
The audit trail
When an item is challenged, the chain is the answer.
Every published item traces from the designated source passage through generation, triage, expert review, and approval to its published form, with names and dates. A hallucinated source span is treated as fabricated provenance and rejected. A challenge is answered from the record, not from recollection.
Validation evidence
Meets or exceeds traditional psychometric benchmarks.
| Metric | 2025 production | Benchmark |
|---|---|---|
| Item-writing guideline compliance | 94% | 76% traditional |
| Mean difficulty (p-value) | 0.68 | target 0.60 to 0.75 |
| Mean discrimination | 0.32 | target above 0.25 |
| Internal consistency (α) | 0.89 | excellent range |
| Items flagged for revision | 8% | 12% traditional |
| Expert rating on higher-order items | 92% equivalent or superior | blinded expert review |
| Differential item functioning across demographic groups | None detected |
68.65%
Open book. Below passing.
The leading general-purpose medical AI, one that aces medical licensing exams, was given 185 EdAI certification items under open-book conditions with direct access to the source articles. It scored 68.65% overall: below the 70% certification passing threshold, with specialty accuracy ranging from 55% to 79%. Two specialties failed nearly half the questions with the answers’ source material in hand.
Content access is not content mastery. Certification-grade items demand clinical synthesis beyond retrieval, and that is a property of how the items are built.
The corpus underneath
Nothing uncertified is ever served.
Underneath the suite sits a certified clinical knowledge base: physician-reviewed source documents turned into a structured graph of atomic claims. Nothing uncertified is ever served, every claim carries its exact source sentence, and every set of facts declares whether it is complete or a sample, so an answer key is never built on a guess.
- 159,661
- Live certified claims, each certified by a named, accountable physician
- 954
- Physician-authored source reviews behind them
- 14,715
- Medical concepts connected by 16,496 certified relationships
- 7
- National exam targets lensed from one corpus without duplicating knowledge
Inside the suite
The working infrastructure of a certifying board.
Item generation
Certification-grade multiple-choice drafting from designated sources, calibrated across the expertise spectrum from resident foundation to subspecialty focus, queued for expert review in triage order.

Quality assurance
Five gates between a draft and an exam.
Automated checks
Content verification requiring 100% fact fidelity, guideline compliance, and clinical coherence, before any human sees the draft.
Level 1: the content expert
A board-certified specialist verifies clinical accuracy, difficulty, and distractor plausibility.
Level 2: the assessment expert
A psychometrician checks item-writing quality, cueing, cognitive level, and bias.
Level 3: editorial
Grammar, style consistency, format, and citation verification.
Pilot testing
Field administration with psychometric data collection: p-value, discrimination, comparison against traditional items, then refinement or retirement.
Approval requires 100% clinical accuracy, guideline compliance of at least 95%, and expert consensus. Classifiers triage and order attention at every stage; they never gate and never decide.
See it
See the item pipeline
A walkthrough of the path from designated source to approved examination item: generation, triage, expert review, and the audit trail.
Demo film in production
A pipeline walkthrough film is planned; request a live demo in the meantime.
Questions boards ask