Skip to main content

Eyecare-100K: Multimodal Ophthalmology VQA Corpus

~102K VQA pairs derived from 58,485 images across 8 ophthalmic modalities (fluorescein angiography, ICGA, OCT, CFP, ultrasound biomicroscopy, slit-lamp, fundus auto-fluorescence, CT) covering 100+ diseases.

At a glance

FieldValue
Short nameeyecare_100k
Full nameEyecare-100K: Multimodal Ophthalmology VQA Corpus
First publishedUnknown
Publication date precisionUnknown
Publication date evidenceUnknown
Publication date source fieldUnknown
Publication date reviewedUnknown
Primary categorymultimodal
Resource rolecurrent_dataset
Dataset familyeyecare_100k
Contained modalitiesfundus, fundus_angiography, fundus_autofluorescence, oct, ocular_ultrasound, external_eye, ct, text
Tasksclassification
Primary reported quantity102,000 question answer pairs
ClassesNot reported (Not reported)
Splitstrain, val, test
Size30.0 GB
Source-stated termsMixed (inherits source-dataset licenses) — needs per-row check
Normalized termsunknown
Descriptive screening labelUnknown or unclear; do not assume permission
Terms scopemixed_components
Access frictionself_service_authenticated
Route backendHuggingFace Hub
Availabilityavailable (checked 2026-07-21)
Acquisition supportstandard_platform_supported
Legacy sample-loader statusMetadata and access only

Reported quantities

RoleCountUnitScopeBasisEvidence
Primary102,000question_answer_pairsSource-described VQA corpus The dataset was described as pending release at the catalog cutoff.official_source_descriptiongithub.com/DCDmllm
Additional58,485imagesSource images represented by the VQA corpusofficial_source_descriptiongithub.com/DCDmllm

Counts retain their source-reported units. Additional rows can describe components, paired items, or derivative copies and are not automatically added to the primary quantity.

Notes

Status: PENDING RELEASE. Largest open ophthalmology VQA corpus when published. EyecareGPT models are out (LLSuzy/* on HF) but the VQA dataset itself is not yet public as of 2026-06.

Access information and download

# Read-only preflight
eyehub download eyecare_100k --data-dir ./data --dry-run --json

# Download, only when preflight reports supported behavior
eyehub download eyecare_100k --data-dir ./data

Upstream page: github.com/DCDmllm

Source-term evidence: github.com/DCDmllm

Loader status

This catalog record provides metadata and access instructions, but it does not yet include a standard DatasetSample loader. Inspect the source file structure or contribute a loader before using it in a training pipeline.

Citation

@misc{eyecare_100k,
title = { Eyecare-100K: Multimodal Ophthalmology VQA Corpus },
note = { EyecareGPT: arXiv 2504.13650; ACM MM 2025. github.com/DCDmllm/EyecareGPT },
year = { 2025 },
url = { https://github.com/DCDmllm/EyecareGPT },
}

Source-stated terms

  • Raw source string: Mixed (inherits source-dataset licenses) — needs per-row check
  • Normalized category: unknown
  • Apparent scope: mixed_components
  • Descriptive screening label: Unknown or unclear; do not assume permission

⚠️ Source-stated terms, scope, and normalized labels are curation metadata, not legal advice or a permission finding. Review the current official source before transfer or reuse.

Similar resources by shared modality

  • ophthalvqa: OphthalVQA Dataset (600 question answer pairs, cc-by)
  • x_pcr: X-PCR Ophthalmology Progressive Clinical Reasoning Benchmark (18,735 rows, unknown)
  • lmod_plus: LMOD+ Multimodal Ophthalmology Benchmark (32,633 annotated instances, unknown)
  • mm_retinal_reason: MM-Retinal-Reason: Ophthalmology Multimodal Reasoning Dataset (130 question answer pairs, unknown)
  • angioreport: AngioReport Fundus Angiography Report Dataset (55,361 images, unknown)
  • ffa_ir: FFA-IR Medical Report Dataset (47,247 images, unknown)
  • jrc_multimodal_vessels: JRC Multi-Modal Retinal Vessel Segmentation (120 images, unknown)
  • multieye: MultiEYE: OCT-Enhanced Fundus Multi-Disease Benchmark (103,959 images, mit)