Lung cancer subtype classification from chest ct scans using vision transformers Görsel dönüştürücüler kullanılarak göğüs BT görüntülerinden akciğer kanseri alt tipi sınıflandırması
Journal of the Faculty of Engineering and Architecture of Gazi University, cilt.41, sa.2, ss.829-840, 2026 (SCI-Expanded, Scopus, TRDizin)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 41 Sayı: 2
- Basım Tarihi: 2026
- Doi Numarası: 10.17341/gazimmfd.1811234
- Dergi Adı: Journal of the Faculty of Engineering and Architecture of Gazi University
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Art Source, Compendex, TR DİZİN (ULAKBİM), Academic Search Ultimate (EBSCO), Engineering Source (EBSCO)
- Sayfa Sayıları: ss.829-840
- Anahtar Kelimeler: chest CT scans, deep learning, Lung cancer classification, medical image analysis, vision transformer
- Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
- Gazi Üniversitesi Adresli: Evet
Özet
Lung cancer is a major global health concern and remains a leading cause of cancer-related mortality, highlighting the urgent need for accurate and early diagnostic approaches to improve patient survival rates. In this study, we propose a transformer-based deep learning model for the automatic classification of lung cancer subtypes from chest CT images. The model is trained and evaluated on the publicly available Chest CT-Scan Images Dataset, which includes four diagnostic categories: adenocarcinoma, squamous cell carcinoma, large cell carcinoma, and normal tissues. A comprehensive preprocessing pipeline was employed, including label normalization, image resizing, data balancing, and augmentation using RandAugment. We adopted the Vision Transformer (ViT-Large) architecture due to its superior capacity for capturing long-range spatial dependencies in medical imaging. Experimental results demonstrate significant performance gains, with the model achieving 98.5% accuracy and a macro-averaged F1-score of 0.980 on the test set. The novelty of this study lies in the integration of a robust preprocessing strategy with a ViT-Large backbone and token-level training stability enhancements to achieve state-of-the-art performance on a clinically relevant dataset.