BELO Benchmark for Evaluating Language Models in Ophthalmology
Expert-curated ophthalmology multiple-choice benchmark with rationales, assembled for held-out language-model evaluation.
Expert-curated ophthalmology multiple-choice benchmark with rationales, assembled for held-out language-model evaluation.
Text-only ophthalmology explanatory/free-form question-answering dataset for LLM training/evaluation.
Text-only ophthalmology multiple-choice question-answering dataset for LLM training/evaluation.