Industry

The Professional Begging Crisis: Why Medical Boards Can't Scale Assessment

John C. Ferguson, MD, FACS·March 12, 2026·15 min read

America's medical specialty boards run on a dirty secret: they survive by begging busy surgeons to write exam questions for free.

I've spent 15 years serving on medical specialty boards. I've been that volunteer. I've also been the one doing the begging.

And I can tell you: this system is broken.

Not broken in the "needs improvement" way. Broken in the "fundamentally unsustainable and getting worse every year" way.

The Weekly Ritual

Every Monday morning, I send emails to board-certified cosmetic surgeons asking them to write exam questions.

The pitch is always the same: "Would you be willing to volunteer 4-6 hours to write 2-3 high-quality assessment items for our board certification exam?"

The responses fall into predictable categories:

  • 40%: No response at all
  • 35%: "I'm honored you asked, but I'm completely underwater right now"
  • 15%: "Can we talk in 3-4 months?" (Translation: soft no)
  • 8%: "Yes, but..." followed by requests for extensions, templates, examples, hand-holding
  • 2%: Actual reliable volunteers who deliver quality work on time

Do the math. To get 2-3 usable questions, I need to contact 50 qualified physicians.

To build an exam with 200 questions? I'm begging 3,000-4,000 busy surgeons.

Every. Single. Year.

Why Smart People Say No

Let's be honest about what we're asking.

Writing a good board certification question isn't like writing a quiz for medical students. It requires:

  • Research time: Reviewing current literature to ensure accuracy (1-2 hours)
  • Item construction: Creating clinically authentic scenarios with plausible distractors (2-3 hours)
  • Quality review: Ensuring NBME standards, no cueing, proper format (1 hour)
  • Revision cycles: Responding to reviewer feedback (1-2 hours)

Total time commitment: 4-8 hours per question.

For a practicing surgeon billing $500-1,000 per clinical hour, you're asking them to donate $2,000-8,000 worth of their time.

For free.

To help test their competitors for board certification.

While they're already working 60-hour weeks, dealing with insurance companies, managing staff, handling emergencies, and trying to see their families.

Of course they say no.

The surprising thing isn't that 98% decline. The surprising thing is that 2% say yes.

The Volunteer Burnout Cycle

The 2% who say yes are heroes. They're also the reason the system persists despite being fundamentally broken.

These are your true believers: senior surgeons who feel duty-bound to give back, academics with protected time for professional service, recent board-passers still grateful for their certification.

They carry the entire weight of assessment development for their specialty.

Until they burn out.

I've watched it happen dozens of times:

  • Year 1: Enthusiastic volunteer, delivers quality work on time
  • Year 2: Still reliable, starting to need gentle reminders
  • Year 3: Requests extensions, quality slipping slightly
  • Year 4: Apologetic emails about being overwhelmed
  • Year 5: Stops responding to requests entirely

We lose our best volunteers to exhaustion. Then we start the begging cycle again with new targets.

The Quality Crisis

Here's what happens when you build assessment on begging:

You can't be selective. When someone volunteers, you accept their contribution even if they're not your ideal item writer. You need the questions too badly to turn anyone away.

You can't demand quality. When someone donates their time, you can't send their work back for extensive revisions without feeling like an ungrateful jerk.

You can't ensure coverage. Your exam blueprint requires questions on Topic X, but nobody volunteers to write on Topic X? Too bad. You compromise the blueprint.

You can't maintain currency. When new evidence emerges or guidelines change, you need someone to rewrite existing items. Good luck finding volunteers for revision work.

You can't scale. Want to offer more practice questions? Develop subspecialty certifications? Create adaptive testing? All require MORE questions, which means MORE begging.

The result: exams that work well enough to meet minimum standards, but fall far short of what's educationally optimal.

The Small Specialty Death Spiral

For large specialties, internal medicine, pediatrics, surgery, the volunteer pool is big enough that the begging system limps along.

For small specialties, it's existential.

Facial cosmetic surgery has maybe 500 board-certified practitioners nationwide. Pediatric endocrinology has 700. Reproductive endocrinology has 450.

When you need 400 exam questions and you have 500 potential volunteers, the math doesn't work. You're asking nearly every practitioner in your entire specialty to contribute.

And when they can't or won't, you face terrible choices:

  • Compromise exam security by reusing too many questions
  • Lower quality standards to accept marginal items
  • Reduce exam scope to cover less content
  • Abandon certification entirely (yes, this happens)

I've seen specialty boards seriously discuss whether they can continue to exist because they can't generate enough assessment items.

Not because there aren't qualified experts. Because those experts won't work for free.

The Economics Nobody Talks About

Medical specialty boards operate on tight budgets. Most charge candidates $1,500-3,500 for certification exams.

Sounds like a lot, right?

Now do the math:

For a 200-question exam reaching 150 candidates annually:

Revenue: $375,000 (at $2,500 per candidate)

Expenses:

  • Psychometrician: $50,000
  • Testing platform fees: $40,000
  • Item review and editing: $30,000
  • Standard setting: $25,000
  • Legal/regulatory compliance: $20,000
  • Administration: $40,000
  • Total: $205,000

Remaining for item development: $170,000

If you paid fair market rates for physician time to write questions ($500/hour × 6 hours × 200 questions), it would cost $600,000.

You'd need to charge each candidate $4,000 just to break even. And that's before accounting for practice questions, item banking, content updates, or any actual profit.

The math doesn't work without free labor.

So boards beg.

The One-Size-Fits-All Problem

Here's another limitation nobody talks about: traditional board exams test everyone at the same level.

It doesn't matter if you're a recent fellow fresh out of training or a 30-year veteran with thousands of cases under your belt. You get the same 200 questions testing the same baseline competency.

This makes sense for initial certification, we need to verify minimum standards for patient safety.

But it's terrible for continuing education and maintenance of certification.

Think about it: A surgeon who's been doing facelifts for 25 years doesn't need practice questions about basic facial anatomy. They need questions that challenge them at their level, complex revision cases, unusual complications, cutting-edge techniques.

But boards can't provide that because they can barely generate enough questions for baseline certification, let alone create differentiated content for various expertise levels.

Finding the Questions Above Your Altitude

Here's what's actually needed: adaptive assessment that meets physicians where they are.

Not "easy, medium, hard" in some generic sense. But questions calibrated to your specific expertise level based on your experience, practice patterns, and knowledge gaps.

If you're a predominantly non-surgical cosmetic practitioner, you need different questions than a facial plastic surgeon. If you focus on injectables, you need different depth than someone doing full facelifts.

The analogy is flying: A student pilot practices at 3,000 feet. An experienced pilot operates at 30,000 feet. You don't improve by practicing at altitudes you've already mastered, you need questions that live just above your current capability.

This is impossible with volunteer-driven assessment.

You can't beg volunteers to write questions for every possible expertise level and practice pattern. You can barely get them to write baseline certification questions.

But with content-driven intelligence, this becomes straightforward:

The same curated content can generate:

  • Foundation-level questions for residents and fellows
  • Certification-level questions for initial board candidates
  • Advanced-level questions for experienced practitioners
  • Specialty-focused questions based on individual practice patterns
  • Remediation questions targeting specific knowledge gaps

All from the same validated content repository. All maintaining the same clinical accuracy. All requiring the same expert validation, but generated at scale instead of begged one-by-one.

Personalized Competency Assessment

Imagine this scenario:

Dr. Smith is preparing for recertification. She's been in practice 15 years, primarily doing facial injectables and non-surgical rejuvenation. She struggles with advanced complications but excels at patient selection.

Traditional board review: Generic 200-question exam covering everything from basic anatomy (which she mastered 20 years ago) to surgical techniques (which she doesn't perform).

Result: She's bored by the easy questions, frustrated by irrelevant surgical questions, and doesn't get enough practice on the advanced complications she actually needs to study.

Content-driven approach:

The system analyzes her practice patterns and generates:

  • 10% foundational questions (quick knowledge verification)
  • 30% standard certification-level questions in her primary practice areas
  • 40% advanced questions on injectables, complications, and aesthetic judgment
  • 20% challenging scenarios specifically targeting complication management

Every question is above her current "altitude", challenging but achievable, pushing her toward expertise growth rather than testing what she already knows cold.

Why This Matters for Continuing Competency

Medical boards are moving toward continuous assessment models. The days of "pass once and you're certified forever" are ending.

The new paradigm: demonstrate ongoing competency throughout your career.

But you can't do that with one-size-fits-all exams and a volunteer-begging model.

You need:

  • Adaptive questioning that adjusts to individual expertise
  • Practice pattern customization that reflects what physicians actually do
  • Continuous content updates as medicine evolves
  • Unlimited question variations to prevent memorization and security issues
  • Immediate feedback that guides learning, not just judges performance

None of this is possible when you're begging volunteers for 200 generic questions once a year.

All of it becomes possible when assessment generation is content-driven and scalable.

From Assessment to Education

Here's the real opportunity: when you separate question generation from volunteer availability, assessment transforms from gatekeeping to growth.

Instead of "Here's your one-time board exam, pass or fail," you can offer:

  • Personalized practice banks: Unlimited questions at your expertise level
  • Adaptive learning pathways: Questions that adjust based on your performance
  • Spaced repetition: Return to challenging concepts at optimal intervals
  • Progress tracking: See your competency growth over time
  • Targeted remediation: Extra practice on your specific weak areas
  • Peer benchmarking: Compare your knowledge to others in your specialty

This isn't just better assessment, it's continuous professional development integrated with credentialing.

Physicians actually want this. They want to maintain competency. They want personalized learning. They just don't have time to sift through generic review materials hoping to find the 10% that's relevant to them.

But boards can't deliver it because they're stuck begging for baseline exam questions.

Why This Matters Beyond Boards

The professional begging crisis isn't just a medical board problem. It's a symptom of a broader dysfunction in how we approach professional certification.

  • Law: Bar examiners beg attorneys to write questions
  • Engineering: PE exam committees beg engineers to write questions
  • Nursing: Certification boards beg nurses to write questions
  • CPA: Accounting boards beg CPAs to write questions

Every high-stakes professional credential in America runs on the same broken model: volunteer subject matter experts donating hundreds of hours to write assessment items.

And every single one faces the same limitation: they can barely generate enough questions for baseline certification, let alone create the personalized, adaptive, continuous assessment that modern professional practice requires.

It persists because there hasn't been an alternative.

The AI Snake Oil

"Use AI to write your exam questions!"

I hear this pitch monthly. Usually from people who have never written a board certification question, served on an exam committee, or understood psychometric standards.

The pitch goes like this: ChatGPT/Claude/GPT-4 can generate medical questions. Problem solved!

Except:

AI hallucinates. It invents medical "facts" that sound plausible but are completely wrong. You can't have hallucinations in high-stakes certification.

AI lacks clinical judgment. It doesn't know what makes a scenario authentic vs. textbook-artificial. It doesn't understand the nuances of real practice.

AI can't ensure item quality. It doesn't know NBME item-writing principles. It creates cued questions, implausible distractors, and ambiguous stems.

AI doesn't understand psychometrics. It can't target specific difficulty levels, control for cognitive complexity, or ensure appropriate discrimination.

Most importantly: AI doesn't have the domain expertise that makes board certification meaningful.

When I ask a board-certified cosmetic surgeon to write a question about rhinoplasty complications, I'm not just getting a question. I'm getting 20 years of clinical experience, thousands of procedures, intimate familiarity with what residents get wrong, and understanding of what matters for patient safety.

Generic AI has none of that.

The Real Solution

Here's what actually works:

Content-driven assessment intelligence.

Instead of asking AI to generate questions from its training data (hallucination risk), you feed it validated medical content that your experts have already curated.

Instead of replacing expert judgment, you enhance it, shifting experts from mechanical item writing to strategic content validation.

Instead of begging for questions, you generate unlimited variations from the content your specialty already agrees matters.

Instead of one-size-fits-all testing, you generate questions calibrated to individual expertise levels, finding each physician's "altitude" and creating questions that challenge them appropriately.

This is what we built at EdAI Systems. Not AI that pretends to be a physician. AI that transforms expert-curated content into properly formatted assessment items at controlled difficulty levels that still require expert validation.

The difference:

Traditional approach:

Beg surgeon → Wait 6 weeks → Receive 2 questions → Review → Request revisions → Wait 3 more weeks → Maybe get usable items → Repeat 200 times → Get one-size-fits-all exam

Content-driven approach:

Expert curates/validates content → AI generates questions at specified difficulty/complexity → AI generates 20 variations from foundation to advanced → Expert reviews and approves best items → Takes 20 minutes instead of 6 hours → Creates personalized assessment pathways

We're not replacing expertise. We're eliminating the mechanical grunt work so experts can focus on what only they can do: validating clinical accuracy and ensuring appropriate difficulty.

What Changes When You Stop Begging

When medical boards shift from begging for questions to content-driven generation, everything changes:

Quality improves. Experts spend time validating instead of writing from scratch. They can be more selective because the constraint isn't volunteer willingness.

Coverage expands. You can ensure comprehensive blueprint coverage instead of accepting whatever volunteers produce.

Personalization becomes possible. Generate questions at multiple difficulty levels from the same content, from resident review to expert-level challenges.

Currency increases. When guidelines change, you update content once and regenerate affected items at all levels. No begging for rewrites.

Security strengthens. Unlimited item variations mean you can retire items freely and create multiple exam forms without security concerns.

Innovation becomes possible. Adaptive testing, personalized practice, immediate feedback, all require item volume that begging can't sustain.

Small specialties survive. The 500-person specialty can generate comprehensive exams without asking every practitioner to volunteer.

Continuous competency becomes feasible. You can actually deliver personalized learning pathways that meet physicians at their expertise level.

Most importantly: Expertise gets respected.

Instead of begging surgeons to donate labor they can't afford to give, you're asking them to do what they're actually good at, validating that content is current, clinically accurate, and appropriate for various practice levels.

They say yes because it takes 20 minutes instead of 6 hours. Because it's intellectually engaging instead of tedious. Because it respects their time instead of exploiting their guilt.

The Path Forward

The professional begging crisis won't fix itself.

As medicine gets more complex, board certification becomes more important. As physician burnout worsens, volunteer recruitment gets harder. As regulatory scrutiny increases, quality standards rise. As continuing competency requirements expand, assessment needs grow exponentially.

The gap between what boards need and what volunteers can provide grows wider every year.

We need a fundamentally different approach. Not AI replacing experts. Not lowering standards. Not abandoning certification.

We need systems that:

  • Respect expert time instead of exploiting it
  • Generate quality at scale instead of begging for scraps
  • Personalize to individual expertise instead of one-size-fits-all testing
  • Enable innovation instead of constraining it
  • Make assessment development sustainable instead of soul-crushing

This isn't about technology. It's about redesigning broken systems that persist because "that's how we've always done it."

Medical boards deserve better than professional begging.

Physicians deserve better than guilt-driven volunteering and generic testing.

Candidates deserve better than exams built on whatever questions boards could scrape together.

And experienced practitioners deserve assessment that actually challenges them at their level instead of retreading ground they mastered years ago.

We can do better.

We must do better.

It's time to stop begging.


Dr. John Ferguson is a quintuple board-certified cosmetic and facial plastic surgeon and founder of EdAI Systems. He serves as Chair of the Written Exam Committee for the American Board of Cosmetic Surgery and has spent 15 years in medical board leadership roles. He has both begged for questions and been begged, and thinks both experiences are equally terrible.

Medical boards: Want to stop begging and start building sustainable, personalized assessment? Let's talk.

See it against your own content.

Demos are live and specific: your board, your school, your exam, your route.

Request a demo