Fundus-CC-2.5M Text Corpus
Large fundus-related multilingual text corpus for LLM pretraining or retrieval.
At a glance
| Field | Value |
|---|---|
| Short name | fundus_cc_2_5m |
| Full name | Fundus-CC-2.5M Text Corpus |
| Primary category | text |
| Contained modalities | text |
| Tasks | text_generation, retrieval |
| Samples | 2,500,000 |
| Classes | Not reported (Not reported) |
| Splits | all |
| Size | 5.0 GB |
| Source-stated terms | Unknown |
| Normalized terms | unknown |
| Descriptive screening label | Unknown or unclear; do not assume permission |
| Terms scope | unknown |
| Access friction | anonymous_direct |
| Route backend | HuggingFace Hub |
| Availability | available (checked 2026-07-21) |
| Acquisition support | standard_platform_supported |
| Legacy sample-loader status | Metadata and access only |
Notes
Large web/text corpus; verify source filtering and redistribution terms before reuse.
Access preflight and acquisition
- CLI
- Python
# Read-only preflight
eyehub download fundus_cc_2_5m --data-dir ./data --dry-run --json
# Explicit transfer, only when preflight reports supported behavior
eyehub download fundus_cc_2_5m --data-dir ./data
from eyedatahub.acquisition import preflight_dataset
from eyedatahub.datasets.registry import REGISTRY
ds = REGISTRY.get_dataset('fundus_cc_2_5m')
print(preflight_dataset(ds, './data')) # no transfer
Upstream page: huggingface.co/datasets
Source-term evidence: huggingface.co/datasets
Loader status
This catalog record provides metadata and access instructions, but it does not yet include a standard DatasetSample loader. Inspect the source file structure or contribute a loader before using it in a training pipeline.
Citation
- BibTeX
- Plain text
@misc{fundus_cc_2_5m,
title = { Fundus-CC-2.5M Text Corpus },
note = { PJMixers-Dev/Fundus-CC-2.5M. Hugging Face dataset, accessed 2026-07 },
year = { 2026 },
url = { https://huggingface.co/datasets/PJMixers-Dev/Fundus-CC-2.5M },
}
PJMixers-Dev/Fundus-CC-2.5M. Hugging Face dataset, accessed 2026-07.
Source-stated terms
- Raw source string: Unknown
- Normalized category:
unknown - Apparent scope:
unknown - Descriptive screening label: Unknown or unclear; do not assume permission
⚠️ Source-stated terms, scope, and normalized labels are curation metadata, not legal advice or a permission finding. Review the current official source before transfer or reuse.
Related datasets with shared modalities
- ocular_chat_vqa: OcularChat-VQA: AREDS-Derived Patient-Physician Dialogue Dataset (844,000 records,
cc-by-nc-sa) - ophora: Ophora-160K: Ophthalmic Surgical Video Instruction Dataset (160,185 records,
unknown) - fundus_105k: Fundus-105K Text Dataset (105,000 records,
unknown) - eyecare_100k: Eyecare-100K: Multimodal Ophthalmology VQA Corpus (102,000 records,
unknown) - angioreport: AngioReport Fundus Angiography Report Dataset (55,361 records,
unknown) - ophthalmology_mcqa_v3: Ophthalmology-MCQA-v3 (51,745 records,
unknown) - ophthalmology_eqa_v3: Ophthalmology-EQA-v3 (49,300 records,
unknown) - ffa_ir: FFA-IR Medical Report Dataset (47,247 records,
unknown)