Skip to main content

FairVLMed: Fair Vision-Language Medical Ophthalmic Dataset

Ophthalmic clinical text + NPZ records covering glaucoma, cataract, and neuro-ophthalmology with paired age, sex, race/ethnicity, and language attributes. Derived from the same Harvard clinical population as Harvard-FairVision.

At a glance

FieldValue
Short namefairvlmed
Full nameFairVLMed: Fair Vision-Language Medical Ophthalmic Dataset
Primary categorymultimodal
Contained modalitiesfundus, visual_field, text, tabular
Tasksclassification
SamplesNot reported
ClassesNot reported (Not reported)
Splitstrain, val, test
Size10.0 GB
Source-stated termsSee Harvard AI Robotics terms
Normalized termsunknown
Descriptive screening labelUnknown or unclear; do not assume permission
Terms scopeunknown
Access frictionanonymous_direct
Route backendHuggingFace Hub
Availabilityavailable (checked 2026-07-21)
Acquisition supportstandard_platform_supported
Legacy sample-loader statusMetadata and access only

Notes

OVERLAP: Derived from the same Harvard clinical cohort as harvard_fairvision (already indexed). Kept for VLM/text researchers who specifically need the text+NPZ view.

Access preflight and acquisition

# Read-only preflight
eyehub download fairvlmed --data-dir ./data --dry-run --json

# Explicit transfer, only when preflight reports supported behavior
eyehub download fairvlmed --data-dir ./data

Upstream page: huggingface.co/datasets

Source-term evidence: huggingface.co/datasets

Loader status

This catalog record provides metadata and access instructions, but it does not yet include a standard DatasetSample loader. Inspect the source file structure or contribute a loader before using it in a training pipeline.

Citation

@misc{fairvlmed,
title = { FairVLMed: Fair Vision-Language Medical Ophthalmic Dataset },
note = { Harvard AI Robotics, FairVLMed: Fair vision-language medical ophthalmic dataset. HuggingFace, 2024 },
year = { 2024 },
url = { https://huggingface.co/datasets/harvardairobotics/FairVLMed },
}

Source-stated terms

  • Raw source string: See Harvard AI Robotics terms
  • Normalized category: unknown
  • Apparent scope: unknown
  • Descriptive screening label: Unknown or unclear; do not assume permission

⚠️ Source-stated terms, scope, and normalized labels are curation metadata, not legal advice or a permission finding. Review the current official source before transfer or reuse.

  • lmod_plus: LMOD+ Multimodal Ophthalmology Benchmark (32,633 records, unknown)
  • grape: GRAPE: Glaucoma Real-world Appraisal Progression Ensemble (1,115 records, cc0)
  • ocular_chat_vqa: OcularChat-VQA: AREDS-Derived Patient-Physician Dialogue Dataset (844,000 records, cc-by-nc-sa)
  • eyecare_100k: Eyecare-100K: Multimodal Ophthalmology VQA Corpus (102,000 records, unknown)
  • angioreport: AngioReport Fundus Angiography Report Dataset (55,361 records, unknown)
  • ffa_ir: FFA-IR Medical Report Dataset (47,247 records, unknown)
  • x_pcr: X-PCR Ophthalmology Progressive Clinical Reasoning Benchmark (18,700 records, unknown)
  • deepeyenet: DeepEyeNet (DEN): Fundus Report Generation Dataset (15,709 records, research-only)