Skip to main content

BELO Benchmark for Evaluating Language Models in Ophthalmology

Expert-curated ophthalmology multiple-choice benchmark with rationales, assembled for held-out language-model evaluation.

At a glance

FieldValue
Short namebelo
Full nameBELO Benchmark for Evaluating Language Models in Ophthalmology
First publishedUnknown
Publication date precisionUnknown
Publication date evidenceUnknown
Publication date source fieldUnknown
Publication date reviewedUnknown
Primary categorytext
Resource rolecurrent_dataset
Dataset familybelo
Contained modalitiestext
Tasksquestion_answering, evaluation, reasoning
Primary reported quantity900 questions
ClassesNot reported (Not reported)
Splitsall
SizeNot reported
Source-stated termsUnknown; mixed upstream question-bank terms
Normalized termsunknown
Descriptive screening labelUnknown or unclear; do not assume permission
Terms scopemixed_components
Access frictionauthor_contact
Route backendManual (upstream-gated)
Availabilityavailable (checked 2026-07-21)
Acquisition supportmanual_access_blocked
Legacy sample-loader statusMetadata and access only

Reported quantities

RoleCountUnitScopeBasisEvidence
Primary900questionsHeld-out multiple-choice benchmarkofficial_source_descriptionbelo-dataset.vercel.app

Counts retain their source-reported units. Additional rows can describe components, paired items, or derivative copies and are not automatically added to the primary quantity.

Notes

The 900 questions aggregate BCSC, BioASQ, MedMCQA, MedQA, and PubMedQA sources. The project page instructs users to request the held-out benchmark by email; verify every applicable source term.

Access information and download

# This route requires upstream human action; no transfer starts.
eyehub download belo --data-dir ./data --dry-run --json
# Follow the official instructions shown by preflight.

Upstream page: belo-dataset.vercel.app

Source-term evidence: belo-dataset.vercel.app

Loader status

This catalog record provides metadata and access instructions, but it does not yet include a standard DatasetSample loader. Inspect the source file structure or contribute a loader before using it in a training pipeline.

Citation

@misc{belo,
title = { BELO Benchmark for Evaluating Language Models in Ophthalmology },
note = { BELO: A Benchmark for Evaluation of Language Models in Ophthalmology. Ophthalmology Science. 2025:101050. doi:10.1016/j.xops.2025.101050 },
year = { 2025 },
url = { https://belo-dataset.vercel.app/ },
}

Source-stated terms

  • Raw source string: Unknown; mixed upstream question-bank terms
  • Normalized category: unknown
  • Apparent scope: mixed_components
  • Descriptive screening label: Unknown or unclear; do not assume permission

⚠️ Source-stated terms, scope, and normalized labels are curation metadata, not legal advice or a permission finding. Review the current official source before transfer or reuse.

Similar resources by shared modality

  • fundus_cc_2_5m: Fundus-CC-2.5M Text Corpus (2,500,000 text items, unknown)
  • ocular_chat_vqa: OcularChat-VQA: AREDS-Derived Patient-Physician Dialogue Dataset (844,000 records, cc-by-nc-sa)
  • ophora: Ophora-160K: Ophthalmic Surgical Video Instruction Dataset (162,185 video clip instruction pairs, unknown)
  • fundus_105k: Fundus-105K Text Dataset (105,000 text items, unknown)
  • eyecare_100k: Eyecare-100K: Multimodal Ophthalmology VQA Corpus (102,000 question answer pairs, unknown)
  • angioreport: AngioReport Fundus Angiography Report Dataset (55,361 images, unknown)
  • ophthalmology_mcqa_v3: Ophthalmology-MCQA-v3 (51,745 questions, unknown)
  • ophthalmology_eqa_v3: Ophthalmology-EQA-v3 (49,300 questions, unknown)