AI Grading Module

AI that reads retinal photographs in under two seconds — helping ophthalmologists detect diabetic eye disease before it causes permanent damage.

The Deneye AI Grading Module analyses fundus photographs and assigns one of five internationally standardised DR severity grades. This gives the reviewing ophthalmologist an immediate, evidence-based starting point — reducing workload while maintaining full clinical accountability.

The AI never operates autonomously. Every grade is reviewed and confirmed by a certified ophthalmologist before any clinical action is taken.

Safety where the data is familiar. On the model’s hold-out set and both IDRiD evaluation sets (1,249 images), not a single severe or proliferative DR case was classified as “No DR”. On datasets with strong domain shift this margin degrades — a key reason every AI grade is reviewed by an ophthalmologist and every new deployment site is evaluated on its own images first.

How it works

Three steps from photograph to grading result

1

Image captured

A screener photographs the patient’s retina using any standard fundus camera. The image is uploaded to the Deneye platform — even from a remote location without a permanent internet connection.

2

AI grades in <2 seconds

The AI model preprocesses the image and predicts a DR severity grade on the 5-point International Clinical DR Scale. The result appears immediately on the ophthalmologist’s worklist.

3

Ophthalmologist confirms

A certified ophthalmologist reviews the image and the AI suggestion. They confirm, adjust, or override the grade. Only after human review does the result reach the patient.

The five DR severity grades

Based on the International Clinical Diabetic Retinopathy Severity Scale

Grade Name What it means Recommended follow-up
0 No DR No visible signs of diabetic damage to the retina Routine annual screening
1 Mild NPDR Early changes: microaneurysms (small bulges in retinal blood vessels) Ophthalmologist review
2 Moderate NPDR More extensive changes: hemorrhages, hard exudates, cotton-wool spots Ophthalmologist review
3 Severe NPDR Significant vascular damage, high risk of progressing to proliferative DR Urgent referral to retina specialist
4 Proliferative DR Advanced stage: growth of fragile new blood vessels, high risk of severe vision loss Urgent referral for treatment (e.g., laser, injections)

Independent validation on 37,368 images

Tested on four independent datasets from four countries — none used during training

The standard performance measure in DR grading is Quadratic Weighted Kappa (QWK) — a statistic that compares the AI’s grades to expert ophthalmologist grades, with heavier penalties for larger disagreements. It ranges from 0 (no better than chance) to 1.0 (perfect agreement). A kappa above 0.80 is generally considered “excellent” in the DR grading literature.

DatasetImagesCountryQW KappaAdjacent accuracy
APTOS 2019 (hold-out)733India0.89697.5%
IDRiD Training413India0.71090.8%
IDRiD Testing103India0.60785.4%
Messidor-21,744France0.49989.1%
EyePACS35,108USA0.42087.4%
Total (external)37,3684 countries0.44087.6%

The APTOS hold-out row consists of images held out from the model’s own training dataset; the other four sets are fully external — the model never saw them in any form.

A note on domain shift (Messidor-2 kappa 0.499, EyePACS 0.420). The model was trained on Indian retinal photographs (APTOS 2019). Messidor-2 and EyePACS images come from French and US clinics with different cameras, lighting conditions, and patient demographics, and grading precision drops accordingly — on these two datasets, 2 of 110 and 111 of 1,580 severe or proliferative cases respectively were graded “No DR”. Domain shift is a well-documented challenge in DR AI literature. It is why the AI is strictly a triage aid under mandatory ophthalmologist review, and why every deployment site is evaluated on its own images before clinical use.

Training curves & confusion matrix

APTOS 2019 training set (2,929 images), 20% hold-out validation (733 images). Best QWK achieved: 0.896.

Training loss and QW kappa by epoch over 20 epochs
Training and validation loss & QWK by epoch. Convergence occurs after the head-only phase (first 3 epochs). The model reaches stable QWK 0.896 by epoch 20.
5x5 confusion matrix: AI grade vs expert grade on APTOS hold-out set
Confusion matrix on the 733-image APTOS hold-out set. Off-diagonal errors are almost entirely adjacent-grade (one step). No severe or proliferative cases (grades 3–4) appear in the Grade 0 “No DR” column.
A detailed AI Grading Report is available on request. The report includes full validation metrics across all datasets, per-dataset confusion matrices, grade distribution comparisons, and complete methodology documentation.