Self-Supervised Contrastive Learning with Vision Transformers for Data-Efficient Medical Image Classification

Authors

  • Arijit Gandhi Research Scholar, Department of Management, RKDF University, Kathal More - Argora, Ranchi, Jharkhand 834004 Author

Keywords:

Self-Supervised Learning, Contrastive Learning, Vision Transformers, Medical Image Classification, Data-Efficient Learning

Abstract

Medical image classification often suffers from the scarcity of labeled data, which limits the performance of deep learning models. We propose a self-supervised contrastive learning framework that integrates Vision Transformers to address this data efficiency challenge. The method first applies rigorous image preprocessing, including resizing, intensity normalization, and contrast enhancement, to ensure input consistency. A stochastic data augmentation module then generates two distinct views of each input image through transformations such as random cropping, rotation, and intensity modifications, all carefully constrained to preserve anatomical integrity. The core novelty lies in the self-supervised pretraining phase, where a Vision Transformer encoder maps these augmented views into a latent feature space. A contrastive loss function maximizes agreement between positive pairs derived from the same image while minimizing similarity between negative pairs from different images. This objective forces the encoder to learn transformation-invariant representations that capture high-level semantic content rather than superficial statistics. After pretraining, the encoder weights initialize a downstream supervised classification task, where a fully connected head is appended and fine-tuned on limited labeled data. A strategic freezing mechanism prevents catastrophic forgetting by keeping initial encoder layers fixed while gradually unfreezing deeper layers

Downloads

Download data is not yet available.

Downloads

Published

2026-09-01

Issue

Section

Articles

How to Cite

Gandhi, A. . (2026). Self-Supervised Contrastive Learning with Vision Transformers for Data-Efficient Medical Image Classification. Journal of Machine Learning Innovations and Artificial Intelligence Horizons, 1(2), 113-130. https://jmliaih.notationpublishing.com/1/article/view/20