Kermany OCT 2018: Retinal OCT Image Classification
~84,000 retinal OCT B-scan images across 4 classes: CNV, DME, DRUSEN, NORMAL. Train: ~83,484 / Test: 1000.
At a glance
| Field | Value |
|---|---|
| Short name | kermany_oct |
| Full name | Kermany OCT 2018: Retinal OCT Image Classification |
| First published | 2017-12-25 |
| Publication date precision | day |
| Publication date evidence | data.mendeley.com/datasets |
| Publication date source field | Mendeley Data version 1 page: Published |
| Publication date reviewed | 2026-09-11 |
| Primary category | oct |
| Resource role | current_dataset |
| Dataset family | kermany_oct |
| Contained modalities | oct |
| Tasks | classification |
| Primary reported quantity | 84,484 images |
| Classes | 4 (CNV, DME, DRUSEN, NORMAL) |
| Splits | train, test, val |
| Size | 6.0 GB |
| Source-stated terms | CC BY 4.0 |
| Normalized terms | cc-by |
| Descriptive screening label | Standard label without an explicit NC clause; not a permission finding |
| Terms scope | dataset_files |
| Access friction | self_service_authenticated |
| Route backend | Kaggle |
| Availability | available (checked 2026-07-21) |
| Acquisition support | standard_platform_supported |
| Legacy sample-loader status | Standard loader included |
Reported quantities
| Role | Count | Unit | Scope | Basis | Evidence |
|---|---|---|---|---|---|
| Primary | 84,484 | images | Primary quantity reported in the reviewed catalog source | legacy_catalog_field | kaggle.com/datasets |
Counts retain their source-reported units. Additional rows can describe components, paired items, or derivative copies and are not automatically added to the primary quantity.
Documented relationships
These links record source-supported lineage or overlap, not merely similar modality tags.
- intraretinal_cystoid_fluid is
derived fromthis record: The source states that 1,000 training images were selected from the Kermany Retinal OCT Images DME class; 200 test images were collected separately. (evidence) - mm_retinal_reason is
derived fromthis record: The version-pinned official dataset card lists this record among the CFP or OCT sources used to construct MM-Retinal-Reason. (evidence) - multieye is
derived fromthis record: The MultiEYE paper names this record as one of the public fundus or OCT sources assembled for the benchmark. (evidence) - synthetic_retinal_oct_biomarkers is
derived fromthis record: The official dataset card documents the Kermany retinal OCT collection as the source for the four synthetic diagnostic classes. (evidence) - x_pcr is
derived fromthis record: Source labels in the version-pinned public X-PCR deposit identify this catalog record as upstream material. (evidence)
Access information and download
- CLI
- Python
# Read-only preflight
eyehub download kermany_oct --data-dir ./data --dry-run --json
# Download, only when preflight reports supported behavior
eyehub download kermany_oct --data-dir ./data
from eyedatahub.acquisition import preflight_dataset
from eyedatahub.datasets.registry import REGISTRY
ds = REGISTRY.get_dataset('kermany_oct')
print(preflight_dataset(ds, './data')) # no download
Upstream page: kaggle.com/datasets
Source-term evidence: kaggle.com/datasets
Loader example
This entry includes a standard DatasetSample loader.
from pathlib import Path
from eyedatahub.datasets.registry import REGISTRY
data_dir = Path('~/.eyedatahub/data').expanduser()
ds = REGISTRY.get_dataset('kermany_oct')
samples = ds.load(data_dir, split='test')
for s in samples[:5]:
print(s.sample_id, s.label, s.image_path)
Citation
- BibTeX
- Plain text
@misc{kermany_oct,
title = { Kermany OCT 2018: Retinal OCT Image Classification },
note = { Kermany et al., 'Identifying medical diagnoses and treatable diseases by image-based deep learning', Cell 2018 },
year = { 2018 },
url = { https://www.kaggle.com/datasets/paultimothymooney/kermany2018 },
}
Kermany et al., 'Identifying medical diagnoses and treatable diseases by image-based deep learning', Cell 2018.
Source-stated terms
- Raw source string: CC BY 4.0
- Normalized category:
cc-by - Apparent scope:
dataset_files - Descriptive screening label: Standard label without an explicit NC clause; not a permission finding
⚠️ Source-stated terms, scope, and normalized labels are curation metadata, not legal advice or a permission finding. Review the current official source before transfer or reuse.
Similar resources by shared modality
- syn_oct: SYN-OCT Synthetic Glaucoma OCT Dataset (200,000 images,
cc-by) - multieye: MultiEYE: OCT-Enhanced Fundus Multi-Disease Benchmark (103,959 images,
mit) - eyecare_100k: Eyecare-100K: Multimodal Ophthalmology VQA Corpus (102,000 question answer pairs,
unknown) - lmod_plus: LMOD+ Multimodal Ophthalmology Benchmark (32,633 annotated instances,
unknown) - harvard_fairvision: Harvard-FairVision (AMD + DR + Glaucoma, paired SLO + OCT) (30,000 participants,
cc-by-nc-nd) - mario: MARIO: AMD-Progression Longitudinal OCT (MICCAI 2024) (30,000 images,
cc-by) - mmrdr: MMRDR: Multi-Modal Retinal Diabetic Retinopathy Dataset (24,460 images,
cc-by) - oct_c8: Retinal OCT-C8: 8-Class OCT Classification (24,000 images,
unknown)