Highest public held-out mean foreground Dice across final public and hybrid models.
Executive summary
What the project shows
This capstone built a reproducible optic disc/cup segmentation pipeline from public fundus datasets, tested model-selection and augmentation strategies, evaluated clinical transfer, and compared clinical-only adaptation with hybrid public-clinical training.
Top-line results
Key metrics
Metrics are final public-safe aggregate values generated by the final synthesis notebook.
Notebook 10 public test mean Dice improvement over the Notebook 07 selected public model.
Clinical PSD-derived samples with usable approximate disc/cup masks.
Patient/encounter groups represented in mask-ready clinical PSD-derived data.
Notebook 13 hybrid model improvement over Notebook 10 zero-shot baseline on the same held-out clinical half.
Reduction in patient-weighted CDR absolute error for the Notebook 13 hybrid model relative to the internal zero-shot baseline.
Method and provenance
Pipeline and notebook map
Each stage links to the notebook that produced or analyzed that part of the project.
Public-data foundation
Project setup
Configured a reproducible environment and established repo paths.
Dataset audit
Audited available public data and began manifest construction.
Split strategy
Created public manifests and train/validation/test splits across ORIGA, G1020, REFUGE, and PAPILA.
Baseline U-Net
Established a baseline segmentation model and metrics pipeline.
Architecture comparison
Compared U-Net, U-Net++, and DeepLabV3+ under the same public-data training budget.
Online augmentation
Tested online augmentation strength and stability.
Virtual synthetic expansion
Tested synthetic add-back strategies without materializing synthetic image files.
Public finalist selection
Evaluated the selected public-data model once on the held-out public test split.
Combined augmentation screen
Screened combined augmentation recipes using public validation only.
Long public training
Trained the selected recipe for 25 epochs and improved held-out public performance.
Clinical transfer and adaptation
Clinical PSD-derived data
Converted annotated clinical PSD files into approximate disc/cup masks for exploratory transfer evaluation.
Pure clinical transfer
Evaluated the long public-trained model on PSD-derived clinical data without clinical training.
Clinical-only fine-tuning
Tested small clinical fine-tuning fractions with patient/encounter-group holdout.
Hybrid public + clinical training
Added 50% of clinical patient groups before augmentation and evaluated the remaining held-out clinical half.
Final synthesis
Final analysis and dashboard
Built public-safe tables, claims, figures, and this GitHub Pages storyboard.
Public test results
Public model performance
These comparisons use the same held-out public test split.
| Notebook | Model stage | Mean Dice | Disc Dice | Cup Dice | CDR MAE | Δ mean Dice vs public finalist |
|---|---|---|---|---|---|---|
| Public finalist | Public finalist model | 0.818 | 0.840 | 0.796 | 0.064 | 0.000 |
| Long public model | Long-trained public model | 0.842 | 0.864 | 0.820 | 0.063 | 0.024 |
| Hybrid training | Hybrid public + clinical model | 0.844 | 0.871 | 0.817 | 0.063 | 0.026 |
Clinical transfer and adaptation
Clinical strategy comparison
This section separates clinical-only adaptation from hybrid public-clinical training and explains how to interpret improvement.
| Notebook | Strategy | Condition | Patient-weighted Dice | Δ vs internal zero-shot | Patient CDR error | Comparison scope |
|---|---|---|---|---|---|---|
| Clinical fine-tuning | Clinical-only fine-tuning | Zero-shot public model on clinical holdout | 0.193 | 0.000 | 0.351 | Within-notebook same clinical split |
| Clinical fine-tuning | Clinical-only fine-tuning | Clinical-only fine-tuning, 25% training groups | 0.140 | -0.053 | 0.543 | Within-notebook same clinical split |
| Clinical fine-tuning | Clinical-only fine-tuning | Clinical-only fine-tuning, 50% training groups | 0.149 | -0.044 | 0.525 | Within-notebook same clinical split |
| Clinical fine-tuning | Clinical-only fine-tuning | Clinical-only fine-tuning, 75% training groups | 0.163 | -0.030 | 0.495 | Within-notebook same clinical split |
| Hybrid training | Hybrid public + clinical pre-augmentation | Zero-shot long public model on 50% clinical holdout | 0.265 | 0.000 | 0.385 | Within-notebook same clinical split |
| Hybrid training | Hybrid public + clinical pre-augmentation | Hybrid public + clinical training | 0.330 | 0.065 | 0.263 | Within-notebook same clinical split |
Visual evidence
Figures with interpretation
Figures support the story; the surrounding text explains what each result means and what should happen next.
Data composition and clinical-data limitation
This figure shows the imbalance between the large public training corpus and the very small PSD-derived clinical set. That imbalance is central to the project’s conclusion: the best next step is not simply another architecture tweak, but a larger, cleaner, segmentation-ready clinical dataset with proper layered disc/cup masks.
Open full-size figure
Public performance trajectory
The public-data pipeline improved from the selected public finalist to the longer public-training run and stayed strong after the hybrid public-clinical training experiment. This supports the claim that the final hybrid step did not sacrifice public test performance.
Open full-size figure
Public-to-clinical transfer gap
This figure is the main domain-shift evidence. Public test Dice remained high, while clinical patient-weighted Dice remained much lower. The model learned public fundus segmentation well, but public performance alone did not make it clinically robust.
Open full-size figure
Clinical strategy Dice comparison
Clinical-only fine-tuning did not beat its internal zero-shot baseline, while the hybrid public-clinical training strategy improved patient-weighted Dice on its held-out clinical split. This suggests clinical signal was more useful when integrated before augmentation rather than used as a tiny fine-tuning-only dataset.
Open full-size figure
Clinical strategy CDR error comparison
This figure tracks the downstream cup-to-disc-ratio error. Lower is better. The hybrid model reduced patient-weighted CDR error relative to its zero-shot baseline, but the result remains exploratory because clinical labels were approximate PSD-derived masks.
Open full-size figureEvidence-backed conclusions
Claims and evidence
This is the readable version of the final claims matrix. It is more important than the static exported matrix image.
The public-data segmentation pipeline achieved strong held-out public performance.
Caveat: Public fundus test performance does not guarantee clinical/head-mounted transfer.
Longer public-data training improved public test performance relative to the Notebook 07 selected public model.
Caveat: This comparison is public-domain only.
Public-only training did not eliminate the clinical/head-mounted domain shift.
Caveat: Clinical masks are approximate PSD-derived labels and the sample size is small.
Clinical-only fine-tuning on tiny PSD-derived subsets did not reliably improve held-out patient-weighted clinical performance.
Caveat: Best Notebook 12 condition was zero_shot_notebook_10_on_clinical_test; clinical training splits were very small.
Hybrid public plus clinical pre-augmentation training improved held-out patient-weighted clinical Dice on the Notebook 13 split.
Caveat: The improvement is on a small held-out clinical split and should be treated as exploratory.
Hybrid training improved patient-weighted clinical CDR error on the Notebook 13 split.
Caveat: CDR is sensitive to approximate cup/disc mask quality and clinical labels remain limited.
The model remains exploratory and is not clinically deployable.
Caveat: A deployable system would require larger clinical annotation, prospective validation, and clinical workflow review.
Public-safe assets
Dashboard data, notebooks, and source files
These are committed aggregate outputs and source notebooks. Private clinical images, paths, patient hashes, and image-level private metrics are not included.