Effectiveness Analysis of Straight-Line Backdoor Attacks on the Deep Learning Model

Backdoor Attack, Data Poisoning, Label-Flipping, Convolutional Neural Network (CNN), ResNet-34.

Authors

  • Ogya Secio Noegroho Department of Informatics Engineering, Faculty of Science and Technology, Universitas Islam Negeri Sultan Syarif Kasim Riau, Indonesia
  • Rahmad Abdillah Department of Informatics Engineering, Faculty of Science and Technology, Universitas Islam Negeri Sultan Syarif Kasim Riau, Indonesia
  • Febi Yanto Department of Informatics Engineering, Faculty of Science and Technology, Universitas Islam Negeri Sultan Syarif Kasim Riau, Indonesia
  • Nazruddin Safaat H Department of Informatics Engineering, Faculty of Science and Technology, Universitas Islam Negeri Sultan Syarif Kasim Riau, Indonesia
May 29, 2026
May 30, 2026

Downloads

Deep learning models, particularly Convolutional Neural Networks (CNNs), are increasingly deployed in safety-critical applications yet remain vulnerable to backdoor attacks. Existing research implicitly assumes that effective backdoor triggers must be complex or imperceptible, overlooking the inherent sensitivity of CNNs to simple low-level visual features. This study investigates the effectiveness of a straight-line trigger a 2-pixel-wide horizontal line at 48% image height with normalized intensity of 0.55 in executing backdoor attacks against a ResNet-34 model trained on the GTSRB traffic sign dataset. Following a data poisoning threat model with an all-to-one attack strategy, experiments are conducted at poison rates of 1%, 3%, and 5%. Results demonstrate that the proposed trigger achieves Attack Success Rates (ASR) of up to 100.00% at the highest evaluated poison rate, with ASR values of 93.82%, 99.77%, and 100.00% at poison rates of 1%, 3%, and 5%, respectively, while maintaining Clean Accuracy (CA) within 1.46 percentage points of the clean baseline (97.75%), with the lowest CA recorded at 96.29% under the 5% poison rate condition. Training dynamics analysis reveals pronounced shortcut learning behavior, with the trigger–target association forming rapidly within the early training epochs and converging well ahead of semantic classification features, prior to the full convergence of clean accuracy. Feature map progression and t-SNE visualizations confirm that triggered inputs are completely remapped to the target class cluster at the representational level, while clean inputs retain their original class-discriminative structure. These findings challenge the prevailing assumption that effective backdoor attacks require sophisticated trigger designs, and highlight the critical security implications of CNN inductive bias toward simple spatial patterns. The simplicity and low cost of the proposed attack underscore the urgent need for backdoor defense mechanisms capable of detecting minimal, low-frequency triggers.