Skip to main content

Kermany OCT 2018: Retinal OCT Image Classification

~84,000 retinal OCT B-scan images across 4 classes: CNV, DME, DRUSEN, NORMAL. Train: ~83,484 / Test: 1000.

At a glance

FieldValue
Short namekermany_oct
Full nameKermany OCT 2018: Retinal OCT Image Classification
Primary categoryoct
Contained modalitiesoct
Tasksclassification
Samples84,484
Classes4 (CNV, DME, DRUSEN, NORMAL)
Splitstrain, test, val
Size6.0 GB
Source-stated termsCC BY 4.0
Normalized termscc-by
Descriptive screening labelStandard label without an explicit NC clause; not a permission finding
Terms scopedataset_files
Access frictionself_service_authenticated
Route backendKaggle
Availabilityavailable (checked 2026-07-21)
Acquisition supportstandard_platform_supported
Legacy sample-loader statusStandard loader included

Access preflight and acquisition

# Read-only preflight
eyehub download kermany_oct --data-dir ./data --dry-run --json

# Explicit transfer, only when preflight reports supported behavior
eyehub download kermany_oct --data-dir ./data

Upstream page: kaggle.com/datasets

Source-term evidence: kaggle.com/datasets

Loader example

This entry includes a standard DatasetSample loader.

from pathlib import Path
from eyedatahub.datasets.registry import REGISTRY

data_dir = Path('~/.eyedatahub/data').expanduser()
ds = REGISTRY.get_dataset('kermany_oct')
samples = ds.load(data_dir, split='test')
for s in samples[:5]:
print(s.sample_id, s.label, s.image_path)

Citation

@misc{kermany_oct,
title = { Kermany OCT 2018: Retinal OCT Image Classification },
note = { Kermany et al., 'Identifying medical diagnoses and treatable diseases by image-based deep learning', Cell 2018 },
year = { 2018 },
url = { https://www.kaggle.com/datasets/paultimothymooney/kermany2018 },
}

Source-stated terms

  • Raw source string: CC BY 4.0
  • Normalized category: cc-by
  • Apparent scope: dataset_files
  • Descriptive screening label: Standard label without an explicit NC clause; not a permission finding

⚠️ Source-stated terms, scope, and normalized labels are curation metadata, not legal advice or a permission finding. Review the current official source before transfer or reuse.

  • syn_oct: SYN-OCT Synthetic Glaucoma OCT Dataset (200,000 records, cc-by)
  • eyecare_100k: Eyecare-100K: Multimodal Ophthalmology VQA Corpus (102,000 records, unknown)
  • multieye: MultiEYE: OCT-Enhanced Fundus Multi-Disease Benchmark (58,036 records, mit)
  • lmod_plus: LMOD+ Multimodal Ophthalmology Benchmark (32,633 records, unknown)
  • harvard_fairvision: Harvard-FairVision (AMD + DR + Glaucoma, paired SLO + OCT) (30,000 records, cc-by-nc-nd)
  • mario: MARIO: AMD-Progression Longitudinal OCT (MICCAI 2024) (30,000 records, cc-by)
  • mmrdr: MMRDR: Multi-Modal Retinal Diabetic Retinopathy Dataset (24,460 records, cc-by)
  • oct_c8: Retinal OCT-C8: 8-Class OCT Classification (24,000 records, unknown)