Attention-Enhanced Transformer Network for Interpretable Diabetic Retinopathy Severity Grading

Authors

  • Govinda Sahu Assistant Professor, Department of CS& IT, ISBM University, Chhattisgarh, India. Author

Keywords:

Attention Mechanism, Transformer Network, Diabetic Retinopathy, Severity Grading, Interpretable AI

Abstract

Diabetic retinopathy (DR) remains a leading cause of preventable blindness worldwide, yet automated grading systems often struggle to capture the global pathological patterns essential for accurate severity assessment. We propose an Attention-Enhanced Transformer Network that processes preprocessed retinal fundus images to predict DR severity with improved interpretability. The methodology begins with a rigorous preprocessing pipeline that includes background removal, illumination correction, and contrast enhancement, followed by data augmentation through geometric and photometric transformations. The core architecture divides each image into non-overlapping patches, which are linearly projected into embeddings and combined with positional encodings. These embeddings are then processed by a standard Transformer encoder that employs multi-head self-attention to capture long-range dependencies across the entire retinal structure. The key novelty lies in an Enhanced Attention Module that follows the standard attention layers. This module applies a learnable gating function to produce a spatial attention mask, which dynamically suppresses non-informative background activations while amplifying signals from DR-specific biomarkers such as hard exudates, cotton-wool spots, and neovascularization. The refined features are then aggregated and passed to a classification head that outputs either binary DR presence or multi-class severity grades (No DR, Mild, Moderate, Severe, Proliferative). The model is optimized using the Adam optimizer with cross-entropy loss, and class weighting strategies are integrated to address dataset imbalance. Furthermore, the attention weights from the enhanced module are visualized as heatmaps superimposed on original fundus images, thereby providing clinical interpretability by highlighting the specific retinal regions that drove each diagnostic decision. This work contributes a transformer-based framework that not only achieves competitive grading performance but also offers transparent, region-level explanations for its predictions.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-01

Issue

Section

Articles

How to Cite

Sahu, G. . (2026). Attention-Enhanced Transformer Network for Interpretable Diabetic Retinopathy Severity Grading. Journal of Machine Learning Innovations and Artificial Intelligence Horizons, 1(2), 59-76. https://jmliaih.notationpublishing.com/1/article/view/17