Steel surface defect detection using learnable memory vision transformer

bracu.type.groupResearch Publications
datacite.rightsOpen Access
dc.contributor.authorAyon, Syed Tasnimul Karim
dc.contributor.authorSiraj, Farhan Md
dc.contributor.authorUddin J.
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-09-29T06:32:35Z
dc.date.available2026-09-29T06:32:35Z
dc.date.issued2025-01-01
dc.description.abstractThis study investigates the application of Learnable Memory Vision Transformers (LMViT) for detecting metal surface flaws, comparing their performance with traditional CNNs, specifically ResNet18 and ResNet50, as well as other transformer-based models including Token to Token ViT, ViT without memory, and Parallel ViT. Leveraging a widely-used steel surface defect dataset, the research applies data augmentation and t-distributed stochastic neighbor embedding (t-SNE) to enhance feature extraction and understanding. These techniques mitigated overfitting, stabilized training, and improved generalization capabilities. The LMViT model achieved a test accuracy of 97.22%, significantly outperforming ResNet18 (88.89%) and ResNet50 (88.90%), as well as the Token to Token ViT (88.46%), ViT without memory (87.18), and Parallel ViT (91.03%). Furthermore, LMViT exhibited superior training and validation performance, attaining a validation accuracy of 98.2% compared to 91.0% for ResNet18, 96.0% for ResNet50, and 89.12%, 87.51%, and 91.21% for Token to Token ViT, ViT without memory, and Parallel ViT, respectively. The findings highlight the LMViT’s ability to capture long-range dependencies in images, an area where CNNs struggle due to their reliance on local receptive fields and hierarchical feature extraction. The additional transformer-based models also demonstrate improved performance in capturing complex features over CNNs, with LMViT excelling particularly at detecting subtle and complex defects, which is critical for maintaining product quality and operational efficiency in industrial applications. For instance, the LMViT model successfully identified fine scratches and minor surface irregularities that CNNs often misclassify. This study not only demonstrates LMViT’s potential for real-world defect detection but also underscores the promise of other transformer-based architectures like Token to Token ViT, ViT without memory, and Parallel ViT in industrial scenarios where complex spatial relationships are key. Future research may focus on enhancing LMViT’s computational efficiency for deployment in real-time quality control systems.
dc.description.versionPublished
dc.format.extent499 - 520
dc.identifier.citationAyon, S.T.K., Siraj, F.M., Uddin, J. (2025). Steel Surface Defect Detection Using Learnable Memory Vision Transformer. Computers, Materials & Continua, 82(1), 499–520. https://doi.org/10.32604/cmc.2025.058361
dc.identifier.doi10.32604/cmc.2025.058361
dc.identifier.issn15462218
dc.identifier.other2-s2.0-85214413256
dc.identifier.urihttps://hdl.handle.net/10361/30279
dc.language.isoen_US
dc.publisherTech Science Press
dc.relation.hasversion10.32604/cmc.2025.058361
dc.relation.ispartofComputers Materials and Continua
dc.relation.ispartofseriesComputers Materials and Continua
dc.relation.journalComputers, Materials and Continua
dc.relation.urihttps://www.techscience.com/cmc/v82n1/59245
dc.subjectConvolutional Neural Networks (CNN)
dc.subjectDeep learning
dc.subjectComputer vision
dc.subjectGradient clipping
dc.subjectImage classification
dc.subjectLabel smoothing
dc.subjectLearnable memory
dc.subjectLearnable memory vision transformer (LMViT)
dc.subjectMetal surface defect detection
dc.subjectT-SNE visualization
dc.subject.lcshSteel--Testing.
dc.subject.lcshSteel-works--Data processing.
dc.subject.lcshSteel-works--Computer programs.
dc.subject.lcshComputer vision.
dc.subject.lcshDeep learning (Machine learning).
dc.titleSteel surface defect detection using learnable memory vision transformer
dc.typeArticle
oaire.citation.issue1
oaire.citation.volume82
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameWoosong University
person.identifier.scopus-author-id58777516900
person.identifier.scopus-author-id58777722600
person.identifier.scopus-author-id54994936900

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Steel Surface Defect Detection Using Learnable Memory Vision Transformer.pdf
Size:
1.25 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: