Text datasets
26 datasets; 26 with a primary reported quantity; 581.8 GB total - this page indexes every EyeDataHub resource tagged as containing text data. A resource can appear on more than one modality page. Primary quantities retain their source-reported units and are not summed here.
| Name | Full name | Primary quantity | Size | License | Backend |
|---|---|---|---|---|---|
fundus_cc_2_5m | Fundus-CC-2.5M Text Corpus | 2,500,000 text items | 5.0 GB | unknown | HuggingFace Hub |
ocular_chat_vqa | OcularChat-VQA: AREDS-Derived Patient-Physician Dialogue Dataset | 844,000 records | 8.0 GB | cc-by-nc-sa | HuggingFace Hub |
ophora | Ophora-160K: Ophthalmic Surgical Video Instruction Dataset | 162,185 video clip instruction pairs | 500.0 GB | unknown | HuggingFace Hub |
fundus_105k | Fundus-105K Text Dataset | 105,000 text items | 0.5 GB | unknown | HuggingFace Hub |
eyecare_100k | Eyecare-100K: Multimodal Ophthalmology VQA Corpus | 102,000 question answer pairs | 30.0 GB | unknown | HuggingFace Hub |
angioreport | AngioReport Fundus Angiography Report Dataset | 55,361 images | Not reported | unknown | Manual (upstream-gated) |
ophthalmology_mcqa_v3 | Ophthalmology-MCQA-v3 | 51,745 questions | 0.1 GB | unknown | HuggingFace Hub |
ophthalmology_eqa_v3 | Ophthalmology-EQA-v3 | 49,300 questions | 0.1 GB | unknown | HuggingFace Hub |
ffa_ir | FFA-IR Medical Report Dataset | 47,247 images | Not reported | unknown | PhysioNet |
ophthalmology_pubmed_corpus | Ophthalmology PubMed Corpus | 39,794 documents | 0.2 GB | unknown | HuggingFace Hub |
lmod_plus | LMOD+ Multimodal Ophthalmology Benchmark | 32,633 annotated instances | Not reported | unknown | Manual (upstream-gated) |
ophthalwechat | OphthalWeChat Dataset | 30,120 question answer pairs | 0.0 GB | cc-by | Figshare |
x_pcr | X-PCR Ophthalmology Progressive Clinical Reasoning Benchmark | 18,735 rows | 5.0 GB | unknown | HuggingFace Hub |
deepeyenet | DeepEyeNet (DEN): Fundus Report Generation Dataset | 15,709 images | 5.0 GB | research-only | Manual (upstream-gated) |
dme_vqa | Diabetic Macular Edema Visual Question Answering Dataset | 13,470 question answer pairs | 0.1 GB | cc-by | Zenodo |
dme_vqa_logical | DME VQA Dataset with Logical Relations | 13,470 question answer pairs | 0.1 GB | cc-by | Zenodo |
fairvlmed | FairVLMed: Fair Vision-Language Medical Ophthalmic Dataset | 10,000 images | 10.0 GB | cc-by-nc-nd | HuggingFace Hub |
ru_medical_texts_ophthalmology | Ophthalmology Russian-English Medical Text Translations | 3,473 sentence pairs | 0.0 GB | cc-by | Kaggle |
lmod_cataract_1k_cot | Cataract-1K Surgical Analysis Chain-of-Thought Dataset | 2,256 images | 1.0 GB | mit | HuggingFace Hub |
belo | BELO Benchmark for Evaluating Language Models in Ophthalmology | 900 questions | Not reported | unknown | Manual (upstream-gated) |
ophthalvqa | OphthalVQA Dataset | 600 question answer pairs | 0.0 GB | cc-by | Figshare |
fundus_report_dataset | Fundus Report Dataset | 422 image report pairs | 0.5 GB | cc-by | HuggingFace Hub |
csdi | CSDI: Cataract Severity Diagnostic Image Dataset | 187 images | 1.0 GB | cc-by | HuggingFace Hub |
soul_octa | SOUL: OCTA Human-Machine Collaborative Annotation Dataset | 178 longitudinal samples | 0.1 GB | cc-by | Figshare |
ophtho_readability | Language and Readability Barriers in Ophthalmology Dataset | 139 documents | 0.0 GB | cc-by | Zenodo |
mm_retinal_reason | MM-Retinal-Reason: Ophthalmology Multimodal Reasoning Dataset | 130 question answer pairs | 15.0 GB | unknown | HuggingFace Hub |