Skip to main content

Cataract-1K Surgical Analysis Chain-of-Thought Dataset

Synthetic surgical-analysis instruction/chain-of-thought dataset derived from Cataract-1K frames.

At a glance

FieldValue
Short namelmod_cataract_1k_cot
Full nameCataract-1K Surgical Analysis Chain-of-Thought Dataset
First publishedUnknown
Publication date precisionUnknown
Publication date evidenceUnknown
Publication date source fieldUnknown
Publication date reviewedUnknown
Primary categorymultimodal
Resource roleannotation_layer
Dataset familylmod_cataract_1k_cot
Contained modalitiessurgical_video, text
Tasksvisual_question_answering, text_generation
Primary reported quantity2,256 images
ClassesNot reported (Not reported)
Splitsall
Size1.0 GB
Source-stated termsMIT
Normalized termsmit
Descriptive screening labelStandard label without an explicit NC clause; verify that it applies to data
Terms scopeunknown
Access frictionself_service_authenticated
Route backendHuggingFace Hub
Availabilityavailable (checked 2026-07-21)
Acquisition supportstandard_platform_supported
Legacy sample-loader statusMetadata and access only

Reported quantities

RoleCountUnitScopeBasisEvidence
Primary2,256imagesPNG images in the versioned Hugging Face depositcurrent_deposit_file_listinghuggingface.co/datasets
Additional11,280question_answer_pairsRows across five cross-validation train/validation fold pairs The five folds repeat the 2,256 source images; 11,280 is not a unique-image count.current_deposit_tablehuggingface.co/datasets

Counts retain their source-reported units. Additional rows can describe components, paired items, or derivative copies and are not automatically added to the primary quantity.

Notes

Synthetic instruction layer derived from Cataract-1K; not an independent clinical dataset.

Documented relationships

These links record source-supported lineage or overlap, not merely similar modality tags.

  • This record is derived from lmod_cataract_1k: The dataset card identifies LMOD-Cataract-1K as its image source. (evidence)

Access information and download

# Read-only preflight
eyehub download lmod_cataract_1k_cot --data-dir ./data --dry-run --json

# Download, only when preflight reports supported behavior
eyehub download lmod_cataract_1k_cot --data-dir ./data

Upstream page: huggingface.co/datasets

Source-term evidence: huggingface.co/datasets

Loader status

This catalog record provides metadata and access instructions, but it does not yet include a standard DatasetSample loader. Inspect the source file structure or contribute a loader before using it in a training pipeline.

Citation

@misc{lmod_cataract_1k_cot,
title = { Cataract-1K Surgical Analysis Chain-of-Thought Dataset },
note = { mehti/LMOD-Cataract-1K-surgical-analysis-cot. Hugging Face dataset, accessed 2026-07 },
year = { 2026 },
url = { https://huggingface.co/datasets/mehti/LMOD-Cataract-1K-surgical-analysis-cot },
}

Source-stated terms

  • Raw source string: MIT
  • Normalized category: mit
  • Apparent scope: unknown
  • Descriptive screening label: Standard label without an explicit NC clause; verify that it applies to data

⚠️ Source-stated terms, scope, and normalized labels are curation metadata, not legal advice or a permission finding. Review the current official source before transfer or reuse.

Similar resources by shared modality

  • ophora: Ophora-160K: Ophthalmic Surgical Video Instruction Dataset (162,185 video clip instruction pairs, unknown)
  • lmod_plus: LMOD+ Multimodal Ophthalmology Benchmark (32,633 annotated instances, unknown)
  • fundus_cc_2_5m: Fundus-CC-2.5M Text Corpus (2,500,000 text items, unknown)
  • ocular_chat_vqa: OcularChat-VQA: AREDS-Derived Patient-Physician Dialogue Dataset (844,000 records, cc-by-nc-sa)
  • fundus_105k: Fundus-105K Text Dataset (105,000 text items, unknown)
  • eyecare_100k: Eyecare-100K: Multimodal Ophthalmology VQA Corpus (102,000 question answer pairs, unknown)
  • angioreport: AngioReport Fundus Angiography Report Dataset (55,361 images, unknown)
  • ophthalmology_mcqa_v3: Ophthalmology-MCQA-v3 (51,745 questions, unknown)