UVA MSDS Capstone

AI-Enhanced Ophthalmoscopy

Segmentation model development, clinical-transfer analysis, and public-safe evidence storyboard.

Executive summary

What the project shows

This capstone built a reproducible optic disc/cup segmentation pipeline from public fundus datasets, tested model-selection and augmentation strategies, evaluated clinical transfer, and compared clinical-only adaptation with hybrid public-clinical training.

Public performance improved. Long public training raised held-out public mean Dice over the selected public finalist model.
Clinical transfer remained hard. Public-only models did not fully generalize to PSD-derived clinical/head-mounted imagery.
Clinical-only fine-tuning was unstable. Tiny clinical-only subsets did not reliably improve held-out patient-weighted performance.
Hybrid training helped most. Adding clinical examples before augmentation improved held-out patient-weighted clinical Dice and CDR error.

Top-line results

Key metrics

Metrics are final public-safe aggregate values generated by the final synthesis notebook.

Long public-training gain
+0.024

Notebook 10 public test mean Dice improvement over the Notebook 07 selected public model.

Mask-ready clinical PSD-derived samples
59

Clinical PSD-derived samples with usable approximate disc/cup masks.

Clinical patient/encounter groups
20

Patient/encounter groups represented in mask-ready clinical PSD-derived data.

Hybrid patient-weighted clinical Dice gain
+0.065

Notebook 13 hybrid model improvement over Notebook 10 zero-shot baseline on the same held-out clinical half.

Hybrid patient-weighted CDR error improvement
+0.122

Reduction in patient-weighted CDR absolute error for the Notebook 13 hybrid model relative to the internal zero-shot baseline.

Method and provenance

Pipeline and notebook map

Each stage links to the notebook that produced or analyzed that part of the project.

Public-data foundation

Project setup

Configured a reproducible environment and established repo paths.

Dataset audit

Audited available public data and began manifest construction.

Split strategy

Created public manifests and train/validation/test splits across ORIGA, G1020, REFUGE, and PAPILA.

Baseline U-Net

Established a baseline segmentation model and metrics pipeline.

Architecture comparison

Compared U-Net, U-Net++, and DeepLabV3+ under the same public-data training budget.

Virtual synthetic expansion

Tested synthetic add-back strategies without materializing synthetic image files.

Public finalist selection

Evaluated the selected public-data model once on the held-out public test split.

Combined augmentation screen

Screened combined augmentation recipes using public validation only.

Long public training

Trained the selected recipe for 25 epochs and improved held-out public performance.

Clinical transfer and adaptation

Clinical PSD-derived data

Converted annotated clinical PSD files into approximate disc/cup masks for exploratory transfer evaluation.

Pure clinical transfer

Evaluated the long public-trained model on PSD-derived clinical data without clinical training.

Clinical-only fine-tuning

Tested small clinical fine-tuning fractions with patient/encounter-group holdout.

Hybrid public + clinical training

Added 50% of clinical patient groups before augmentation and evaluated the remaining held-out clinical half.

Final synthesis

Public test results

Public model performance

These comparisons use the same held-out public test split.

How to read this: these rows share the same held-out public test split. Higher Dice is better; lower CDR MAE is better. The hybrid model preserved the public performance gains while adding clinical-domain signal during training.
Public finalist model
0.818
Long-trained public model
0.842
Hybrid public + clinical model
0.844
NotebookModel stageMean DiceDisc DiceCup DiceCDR MAEΔ mean Dice vs public finalist
Public finalist Public finalist model 0.818 0.840 0.796 0.064 0.000
Long public model Long-trained public model 0.842 0.864 0.820 0.063 0.024
Hybrid training Hybrid public + clinical model 0.844 0.871 0.817 0.063 0.026

Clinical transfer and adaptation

Clinical strategy comparison

This section separates clinical-only adaptation from hybrid public-clinical training and explains how to interpret improvement.

Dice improvement:Higher patient-weighted Dice is better. Patient weighting prevents patients or encounters with more images from dominating the result.
CDR error improvement:Lower cup-to-disc-ratio absolute error is better. A positive improvement means error fell relative to that notebook’s internal zero-shot baseline.
Scope warning:Notebook 12 and Notebook 13 use different held-out clinical splits. Compare each strategy against its own zero-shot row, not as one universal leaderboard.
Zero-shot public model on clinical holdout
0.193
Clinical-only fine-tuning, 25% training groups
0.140
Clinical-only fine-tuning, 50% training groups
0.149
Clinical-only fine-tuning, 75% training groups
0.163
Zero-shot long public model on 50% clinical holdout
0.265
Hybrid public + clinical training
0.330
NotebookStrategyConditionPatient-weighted DiceΔ vs internal zero-shotPatient CDR errorComparison scope
Clinical fine-tuning Clinical-only fine-tuning Zero-shot public model on clinical holdout 0.193 0.000 0.351 Within-notebook same clinical split
Clinical fine-tuning Clinical-only fine-tuning Clinical-only fine-tuning, 25% training groups 0.140 -0.053 0.543 Within-notebook same clinical split
Clinical fine-tuning Clinical-only fine-tuning Clinical-only fine-tuning, 50% training groups 0.149 -0.044 0.525 Within-notebook same clinical split
Clinical fine-tuning Clinical-only fine-tuning Clinical-only fine-tuning, 75% training groups 0.163 -0.030 0.495 Within-notebook same clinical split
Hybrid training Hybrid public + clinical pre-augmentation Zero-shot long public model on 50% clinical holdout 0.265 0.000 0.385 Within-notebook same clinical split
Hybrid training Hybrid public + clinical pre-augmentation Hybrid public + clinical training 0.330 0.065 0.263 Within-notebook same clinical split

Visual evidence

Figures with interpretation

Figures support the story; the surrounding text explains what each result means and what should happen next.

Data composition and clinical-data limitation

Data composition and clinical-data limitation

This figure shows the imbalance between the large public training corpus and the very small PSD-derived clinical set. That imbalance is central to the project’s conclusion: the best next step is not simply another architecture tweak, but a larger, cleaner, segmentation-ready clinical dataset with proper layered disc/cup masks.

Open full-size figure
Public performance trajectory

Public performance trajectory

The public-data pipeline improved from the selected public finalist to the longer public-training run and stayed strong after the hybrid public-clinical training experiment. This supports the claim that the final hybrid step did not sacrifice public test performance.

Open full-size figure
Public-to-clinical transfer gap

Public-to-clinical transfer gap

This figure is the main domain-shift evidence. Public test Dice remained high, while clinical patient-weighted Dice remained much lower. The model learned public fundus segmentation well, but public performance alone did not make it clinically robust.

Open full-size figure
Clinical strategy Dice comparison

Clinical strategy Dice comparison

Clinical-only fine-tuning did not beat its internal zero-shot baseline, while the hybrid public-clinical training strategy improved patient-weighted Dice on its held-out clinical split. This suggests clinical signal was more useful when integrated before augmentation rather than used as a tiny fine-tuning-only dataset.

Open full-size figure
Clinical strategy CDR error comparison

Clinical strategy CDR error comparison

This figure tracks the downstream cup-to-disc-ratio error. Lower is better. The hybrid model reduced patient-weighted CDR error relative to its zero-shot baseline, but the result remains exploratory because clinical labels were approximate PSD-derived masks.

Open full-size figure

Evidence-backed conclusions

Claims and evidence

This is the readable version of the final claims matrix. It is more important than the static exported matrix image.

C1

The public-data segmentation pipeline achieved strong held-out public performance.

Evidencebest public test mean foreground Dice
Value0.844
Strengthstrong

Caveat: Public fundus test performance does not guarantee clinical/head-mounted transfer.

C2

Longer public-data training improved public test performance relative to the Notebook 07 selected public model.

EvidenceNotebook 10 minus Notebook 07 public test mean foreground Dice
Value0.024
Strengthstrong

Caveat: This comparison is public-domain only.

C3

Public-only training did not eliminate the clinical/head-mounted domain shift.

EvidenceNotebook 11 patient-weighted clinical mean foreground Dice
Value0.251
Strengthstrong

Caveat: Clinical masks are approximate PSD-derived labels and the sample size is small.

C4

Clinical-only fine-tuning on tiny PSD-derived subsets did not reliably improve held-out patient-weighted clinical performance.

Evidencebest Notebook 12 patient-weighted Dice versus Notebook 12 zero-shot patient-weighted Dice
Value0.000
Strengthmoderate

Caveat: Best Notebook 12 condition was zero_shot_notebook_10_on_clinical_test; clinical training splits were very small.

C5

Hybrid public plus clinical pre-augmentation training improved held-out patient-weighted clinical Dice on the Notebook 13 split.

EvidenceNotebook 13 hybrid minus zero-shot patient-weighted mean foreground Dice
Value0.065
Strengthmoderate to strong

Caveat: The improvement is on a small held-out clinical split and should be treated as exploratory.

C6

Hybrid training improved patient-weighted clinical CDR error on the Notebook 13 split.

EvidenceNotebook 13 patient-weighted CDR absolute error improvement
Value0.122
Strengthmoderate to strong

Caveat: CDR is sensitive to approximate cup/disc mask quality and clinical labels remain limited.

C7

The model remains exploratory and is not clinically deployable.

Evidencesmall clinical sample, approximate labels, uneven image-level performance
Value—
Strengthstrong

Caveat: A deployable system would require larger clinical annotation, prospective validation, and clinical workflow review.

Public-safe assets

Dashboard data, notebooks, and source files

These are committed aggregate outputs and source notebooks. Private clinical images, paths, patient hashes, and image-level private metrics are not included.

KPI cards dashboard_kpi_cards.json Storyboard sections dashboard_story_sections.json Model development registry model_development_registry.csv Public model performance public_model_performance_summary.csv Clinical transfer summary clinical_transfer_summary.csv Clinical strategy comparison clinical_strategy_comparison_summary.csv Public-clinical tradeoff public_clinical_tradeoff_summary.csv Claims/evidence matrix final_claims_evidence_matrix.csv Dashboard data dictionary dashboard_data_dictionary.csv Dashboard asset manifest dashboard_assets_manifest.csv Final interpretation final_project_interpretation.md Setup 00. Environment, paths, and reproducibility setup Data audit 01. Dataset audit and initial manifest work Split strategy 02. Public manifest and train/validation/test split design Baseline U-Net 03. Baseline segmentation model reproduction Architecture comparison 04. U-Net, U-Net++, and DeepLabV3+ comparison Online augmentation 05. Online augmentation ablation Synthetic expansion 06. Virtual synthetic add-back experiments Public finalist 07. Selected public model and held-out public test evaluation Clinical set extraction 08. PSD-derived clinical mask extraction and first transfer test Combined augmentation 09. Combined public augmentation recipe screen Long public model 10. 25-epoch public training run Clinical transfer 11. Long public model evaluated on clinical PSD-derived data Clinical fine-tuning 12. Clinical-only adaptation fractions Hybrid training 13. Public plus clinical pre-augmentation training Final synthesis 14. Final comparative analysis and dashboard assets