Article
A Comparative Performance Analysis of Vision Transformer Architectures for Breast Ultrasound Image Classification
Authors
Abstract
Although ultrasonography plays a critical role in the early detection of breast cancer, its limita- tions, such as operator dependency, necessitate the development of objective analysis methods. In response to this need, deep learning models based on the Vision Transformer (ViT) architecture present promising solutions. This investigation comparatively assesses the performance of four modern Transformer architec- tures Swin-Base, ViT-Base, DeiT-Base, and BEiT-Base for the classification of breast ultrasound images into benign, malignant, and normal categories. Conducted on the publicly available ”Breast Ultrasound Images Dataset,” the study integrated dynamic data augmentation techniques to enhance model generalization. The empirical results demonstrated a statistically significant superiority of the DeiT-Base model, which achieved 94.30% accuracy and a 93.85% F1-score. While ViT-Base and Swin-Base delivered competitive outcomes, BEiT-Base exhibited the lowest performance with 66.46% accuracy. These findings indicate that for the analysis of limited and distinct datasets such as breast ultrasound, data-efficient training strategies like the knowledge distillation employed by DeiT may be more impactful than architectural differences alone. Moreover, these approaches hold considerable potential for future integration into clinical decision support systems.
Keywords
Citation
Published by: Engineering Journals


