Stefanos Koffas
Surfing on the Backdoor Wave
Deep neural networks (DNNs) have become central to a wide range of modern applications, including systems used in security- and safety-critical settings. As these models are increasingly deployed in practice, they have also become attractive targets for adversaries. Among the most concerning threats are backdoor attacks, in which an attacker implants a hidden malicious functionality into a model. This functionality is activated only when a specific trigger is present in the input, while the model otherwise behaves as expected. Backdoors can be introduced through data poisoning, model poisoning, or code poisoning, making them a particularly versatile and dangerous class of attacks with potentially serious consequences in domains such as autonomous driving and biometric access control. This thesis studies backdoor attacks in deep learning from two complementary perspectives: attack generalization across data modalities, and the strengths and limitations of existing defenses. First, we investigate the feasibility of backdoor attacks in speech-based systems, focusing on tasks such as keyword spotting and speaker identification. We design and evaluate several trigger mechanisms, including ultrasonic triggers, triggers based on electric-guitar effects, and emotion-based triggers. Our results show that backdoors can be effectively embedded in speech models using diverse and modality-aware trigger designs, achieving high attack success rates across the evaluated settings. Second, the thesis examines backdoor attacks in domains beyond computer vision, which has traditionally dominated the literature. In particular, we study graph neural networks, tabular learning, and text classification. These modalities impose constraints that differ substantially from those in computer vision and therefore require tailored trigger designs and attack strategies. Nevertheless, our findings show that such constraints do not, by themselves, provide protection against backdoor attacks. With appropriate adaptations, attackers can successfully implant backdoors across all of these settings, demonstrating that the threat is broader and more general than previously assumed. Third, the thesis explores the backdoor problem in greater depth by studying both attack stealthiness and defensive robustness. We identify a limitation in trigger-inversion-based defenses and propose an improvement that leverages adversarial neuron noise to reconstruct backdoor triggers more robustly. We further argue that a comprehensive notion of backdoor stealthiness should account simultaneously for input space, feature space, and parameter space. Based on this insight, we develop an attack that remains stealthy across all three spaces and show that it consistently outperforms baseline attacks against a wide range of countermeasures. Finally, we present a systematic study of backdoor defenses and highlight key methodological shortcomings in current evaluation practices, arguing for more rigorous and standardized assessment in the field. Taken together, the contributions of this thesis show that backdoor attacks are a flexible and pervasive threat that extends across applications, modalities, and model families. At the same time, they show that meaningful progress on defenses requires not only stronger mitigation techniques, but also more systematic evaluation methodologies. By broadening the empirical understanding of backdoor attacks and critically examining current defense practices, this thesis aims to support the development of more secure machine learning systems and to inform future best practices and standardization efforts in trustworthy AI.
| Publicatiedatum | 22 september 2026 |
| Universiteit | Overig |
| Auteur | Stefanos Koffas |
| Order nummer | 19560 |
| ISBN nummer | 978-94-6518-426-5 |