Cataract-1K Surgical Analysis Chain-of-Thought Dataset
Synthetic surgical-analysis instruction/chain-of-thought dataset derived from Cataract-1K frames.
Synthetic surgical-analysis instruction/chain-of-thought dataset derived from Cataract-1K frames.
3,000 cataract surgery procedures, 1,134.2 hours of video with phase annotations, instance segmentation, tracking, and skill scoring. Largest public cataract-surgery-video resource.
187 cataract cases with color fundus images and paired professional cataract-grading diagnostic reports. Designed for medical multimodal-LLM evaluation.
~102K VQA pairs derived from 58,485 images across 8 ophthalmic modalities (fluorescein angiography, ICGA, OCT, CFP, ultrasound biomicroscopy, slit-lamp, fundus auto-fluorescence, CT) covering 100+ dis
Ophthalmic clinical text + NPZ records covering glaucoma, cataract, and neuro-ophthalmology with paired age, sex, race/ethnicity, and language attributes. Derived from the same Harvard clinical popula
Fundus/UWF image-report dataset derived from DeepDRiD and OUWFD-style resources for report-generation research.
Fundus-focused text dataset for LLM/RAG workflows.
Large fundus-related multilingual text corpus for LLM pretraining or retrieval.
Processed Cataract-1K surgical-frame dataset for segmentation/object-detection workflows.
Ophthalmology-specific multimodal reasoning dataset built from 45 public datasets. Chain-of-thought reasoning traces for retinal VQA.
58,036 fundus + 45,923 OCT images assembled for multi-disease classification (8 classes) with cross-modal distillation. Sourced from multiple public ophthalmic datasets.
844,000 simulated patient-physician dialogue rows generated from AREDS clinical visits. Enables ophthalmic dialogue and counseling VLM training.
Large-scale multi-procedure ophthalmic surgical video dataset covering 66 surgery types, 102 phases, 150 operations (~285 h). 1,969 untrimmed videos; 17,508 trimmed operation-level clips; 14,674 trimm
160,185 video clip-instruction pair samples from 9,819 ophthalmic surgical videos, covering multiple procedure types. Designed for text-guided surgical video generation and understanding. Published at
Ophthalmology-focused PubMed text corpus for retrieval, pretraining, or RAG experiments.
Text-only ophthalmology explanatory/free-form question-answering dataset for LLM training/evaluation.
Text-only ophthalmology multiple-choice question-answering dataset for LLM training/evaluation.
Paired tabletop and portable retinal images from the same patients, enabling cross-device domain-adaptation research.
Fundus images and age labels for retinal age prediction and regression.
Baseline and two-year follow-up color fundus image pairs from Tianjin Medical University for diabetic-retinopathy progression research.
Ophthalmic image-text VQA/reasoning benchmark with retinal modalities and expert-verified question-answer pairs.