AngioReport Fundus Angiography Report Dataset
De-identified fluorescein and indocyanine-green angiography images paired with structured lesion descriptions and reports.
De-identified fluorescein and indocyanine-green angiography images paired with structured lesion descriptions and reports.
Expert-curated ophthalmology multiple-choice benchmark with rationales, assembled for held-out language-model evaluation.
Synthetic surgical-analysis instruction/chain-of-thought dataset derived from Cataract-1K frames.
187 cataract cases with color fundus images and paired professional cataract-grading diagnostic reports. Designed for medical multimodal-LLM evaluation.
15,709 fundus images with paired medical reports and extracted keywords. Only public fundus report-generation dataset — useful for VLM / captioning evaluation.
Fundus-image VQA dataset for diabetic macular edema derived from IDRiD and e-ophtha.
Extension of the DME VQA dataset with logical-relation consistency annotations.
~102K VQA pairs derived from 58,485 images across 8 ophthalmic modalities (fluorescein angiography, ICGA, OCT, CFP, ultrasound biomicroscopy, slit-lamp, fundus auto-fluorescence, CT) covering 100+ dis
Ophthalmic clinical text + NPZ records covering glaucoma, cataract, and neuro-ophthalmology with paired age, sex, race/ethnicity, and language attributes. Derived from the same Harvard clinical popula
Fundus fluorescein angiography images paired with Chinese and translated English reports for report-generation research.
Fundus/UWF image-report dataset derived from DeepDRiD and OUWFD-style resources for report-generation research.
Fundus-focused text dataset for LLM/RAG workflows.
Large fundus-related multilingual text corpus for LLM pretraining or retrieval.
Text/tabular dataset supporting readability and language-access analyses in ophthalmology.
Composite multimodal ophthalmology benchmark with multi-granular anatomical, diagnostic, staging, demographic, and text annotations.
Ophthalmology-specific multimodal reasoning dataset built from 45 public datasets. Chain-of-thought reasoning traces for retinal VQA.
844,000 simulated patient-physician dialogue rows generated from AREDS clinical visits. Enables ophthalmic dialogue and counseling VLM training.
160,185 video clip-instruction pair samples from 9,819 ophthalmic surgical videos, covering multiple procedure types. Designed for text-guided surgical video generation and understanding. Published at
Ophthalmology-focused PubMed text corpus for retrieval, pretraining, or RAG experiments.
Russian-English ophthalmology sentence-pair and glossary dataset for translation and LLM evaluation.
Text-only ophthalmology explanatory/free-form question-answering dataset for LLM training/evaluation.
Text-only ophthalmology multiple-choice question-answering dataset for LLM training/evaluation.
Ophthalmic visual question-answering benchmark dataset released as supplementary data for evaluating multimodal language models in ophthalmology.
Ophthalmology-oriented WeChat article metadata and image-link dataset for visual question answering and multimodal language-model evaluation.
Longitudinal OCT angiography projection maps with vessel labels, clinical text, and treatment/follow-up groupings.
Ophthalmic image-text VQA/reasoning benchmark with retinal modalities and expert-verified question-answer pairs.