Medical question-answering benchmark analysis

Seven datasets, different questions.

Independent MultiMedQA analysis covering seven datasets, human-rated consumer answers, sample denominators and the limits of medical question-answering scores.

MultiMedQA / source map

Seven sources, three question families

[1][2][3]
01

Professional exams

  • MedQA
  • MedMCQA
  • MMLU medical subjects

MCQ accuracy

02

Research evidence

  • PubMedQA

Answer accuracy

03

Consumer questions

  • LiveQA
  • MedicationQA
  • HealthSearchQA

Human evaluation on selected subset

The human-rated subset

140 selected questions [1]

HealthSearchQA 100LiveQA 20MedicationQA 20

These are 140 selected questions, not the full sizes of the three source collections. Inspect the study →

01 /

Inspect the task

Trace the actual input, output and evaluation setting for each named benchmark.

02 /

Read the evidence

Explore cited cohort counts and selected historical results with their measurement boundaries.

03 /

Make the inference explicit

Separate the authors’ observations from our analytical interpretation and proposed evaluation questions.

Our analytical question

MultiMedQA is a suite of question sources, with different answer formats and different ways to judge success. We analyze the original Nature study’s exam, research and consumer tasks, with particular attention to its human-rated answer sample. The original contribution here is an explicit map between each source, its input and its scoring boundary. Historical Med-PaLM results remain labeled as published observations. Our guides explain why the suite name alone is insufficient to compare later evaluations.

Read the original evidence closely

The benchmark, unpacked.

Primary source library ↗
Dossier01

Original Nature study, corrected version of record

MultiMedQA ↗

Seven question sources do not produce one clinical score.

UnitOne question; human evaluation uses a selected 140-question subset.MeasureTask-specific accuracy and human ratings

An original analytical tool

What kind of question is being answered?

Evidence explorer

A seven-component suite contains different evidence contracts. Compare the source, input, output and scoring boundary before comparing a percentage.

7 of 7 evidence entries shown

Professional exams

MedQA

Read dossier ↗
Input
USMLE-style vignette and options
Output
Selected answer
Measure
MCQ accuracy
Interpretation boundary

A complete vignette supplies evidence the model did not gather.

[1]
Professional exams

MedMCQA

Read dossier ↗
Input
Indian medical exam question and options
Output
Selected answer
Measure
MCQ accuracy
Interpretation boundary

Source jurisdiction and exam mix differ from MedQA.

[1]
Professional exams

MMLU medical subjects

Read dossier ↗
Input
Subject-specific medical question and options
Output
Selected answer
Measure
Per-subject accuracy
Interpretation boundary

One suite component contains multiple subject subsets.

[1]
Research evidence

PubMedQA

Read dossier ↗
Input
Research question and abstract context
Output
Yes, no or maybe
Measure
Answer accuracy
Interpretation boundary

Supporting abstract is supplied; retrieval is not evaluated.

[1]
Consumer questions

LiveQA

Read dossier ↗
Input
Consumer query
Output
Long-form answer
Measure
Human evaluation on selected subset
Interpretation boundary

A source question count is not a rated-answer count.

[1]
Consumer questions

MedicationQA

Read dossier ↗
Input
Consumer medication query
Output
Long-form answer
Measure
Human evaluation on selected subset
Interpretation boundary

Potential harm is a rater judgment, not an adverse-event outcome.

[1]
Consumer questions

HealthSearchQA

Read dossier ↗
Input
Commonly searched health question
Output
Long-form answer
Measure
Human evaluation on selected subset
Interpretation boundary

Search queries do not establish a clinical diagnosis or patient cohort.

[1]

Rows describe the original MultiMedQA study. They do not merge its MCQ and human-rating endpoints. [1]

Original analysis / methods and interpretation

What the score leaves unsaid.

All analyses →

Questions, answered

Read the result in context.

Specific tasks. Stated conditions.
Inspect every source.

Is MultiMedQA a model?

No. It is a benchmark suite combining seven question-source components. Med-PaLM and Flan-PaLM are models evaluated in its original study.

How many HealthSearchQA questions are in the original published inventory?

The corrected Nature publication reports 3,173. The original human evaluation uses a selected 100 of those questions alongside 20 LiveQA and 20 MedicationQA questions.

Are the harm percentages observed patient injuries?

No. They are human judgments about the potential harm of generated answers in a selected question study. They are not measured clinical adverse-event rates.

Can all MultiMedQA results be combined into one accuracy?

Multiple-choice accuracy and human-rating axes have different denominators and meanings. This site preserves them as separate endpoints.

Working tool / saved on this device

Prepare a reproducible benchmark reading

Interactive worksheet

Use this secondary worksheet to record the version, task and evidence boundary of the benchmark you are considering. Checked items indicate documented decisions, not measured performance.

Identify the experiment

Evidence you can inspect. Benchmark dossiers distinguish published facts from our interpretation, with source versions and access notes attached.

Download the evidence ↗