Skip to main content

OcularChat-VQA: AREDS-Derived Patient-Physician Dialogue Dataset

844,000 simulated patient-physician dialogue rows generated from AREDS clinical visits. Enables ophthalmic dialogue and counseling VLM training.

At a glance

FieldValue
Short nameocular_chat_vqa
Full nameOcularChat-VQA: AREDS-Derived Patient-Physician Dialogue Dataset
First publishedUnknown
Publication date precisionUnknown
Publication date evidenceUnknown
Publication date source fieldUnknown
Publication date reviewedUnknown
Primary categorymultimodal
Resource rolecurrent_dataset
Dataset familyocular_chat_vqa
Contained modalitiestext, tabular
Tasksclassification
Primary reported quantity844,000 records
ClassesNot reported (Not reported)
Splitstrain, val, test
Size8.0 GB
Source-stated termsCC BY-NC-SA 4.0
Normalized termscc-by-nc-sa
Descriptive screening labelExplicit noncommercial clause recorded; check source
Terms scopedataset_files
Access frictionself_service_authenticated
Route backendHuggingFace Hub
Availabilityavailable (checked 2026-07-21)
Acquisition supportstandard_platform_supported
Legacy sample-loader statusMetadata and access only

Reported quantities

RoleCountUnitScopeBasisEvidence
Primary844,000recordsPrimary quantity reported in the reviewed catalog sourcelegacy_catalog_fieldhuggingface.co/datasets

Counts retain their source-reported units. Additional rows can describe components, paired items, or derivative copies and are not automatically added to the primary quantity.

Notes

Images path may reference AREDS — access to underlying images requires separate NCBI/dbGaP approval. Verify before use.

Access information and download

# Read-only preflight
eyehub download ocular_chat_vqa --data-dir ./data --dry-run --json

# Download, only when preflight reports supported behavior
eyehub download ocular_chat_vqa --data-dir ./data

Upstream page: huggingface.co/datasets

Source-term evidence: huggingface.co/datasets

Loader status

This catalog record provides metadata and access instructions, but it does not yet include a standard DatasetSample loader. Inspect the source file structure or contribute a loader before using it in a training pipeline.

Citation

@misc{ocular_chat_vqa,
title = { OcularChat-VQA: AREDS-Derived Patient-Physician Dialogue Dataset },
note = { OcularChat-VQA: AREDS-derived patient-physician dialogues. HuggingFace / NCBI, 2025 },
year = { 2025 },
url = { https://huggingface.co/datasets/ncbi/OcularChat-VQA },
}

Source-stated terms

  • Raw source string: CC BY-NC-SA 4.0
  • Normalized category: cc-by-nc-sa
  • Apparent scope: dataset_files
  • Descriptive screening label: Explicit noncommercial clause recorded; check source

⚠️ Source-stated terms, scope, and normalized labels are curation metadata, not legal advice or a permission finding. Review the current official source before transfer or reuse.

Similar resources by shared modality

  • lmod_plus: LMOD+ Multimodal Ophthalmology Benchmark (32,633 annotated instances, unknown)
  • fairvlmed: FairVLMed: Fair Vision-Language Medical Ophthalmic Dataset (10,000 images, cc-by-nc-nd)
  • soul_octa: SOUL: OCTA Human-Machine Collaborative Annotation Dataset (178 longitudinal samples, cc-by)
  • ophtho_readability: Language and Readability Barriers in Ophthalmology Dataset (139 documents, cc-by)
  • fundus_cc_2_5m: Fundus-CC-2.5M Text Corpus (2,500,000 text items, unknown)
  • ophora: Ophora-160K: Ophthalmic Surgical Video Instruction Dataset (162,185 video clip instruction pairs, unknown)
  • fundus_105k: Fundus-105K Text Dataset (105,000 text items, unknown)
  • eyecare_100k: Eyecare-100K: Multimodal Ophthalmology VQA Corpus (102,000 question answer pairs, unknown)