Tan, Chee Yin (2026) LLM security for sovereign AI: backdoor detection and defense. Final Year Project, UTAR.
| PDF Download (7Mb) |
Abstract
Large Language Models (LLMs) are increasingly integrated into critical systems and national infrastructure, yet the security risks introduced through the fine-tuning process remain poorly understood and under-researched. In particular, backdoor attacks, where a model is conditioned to behave maliciously upon the presence of a hidden trigger while appearing normal otherwise, represent a stealthy and underappreciated threat to the integrity of deployed AI systems. This study investigates the feasibility and effectiveness of backdoor attacks on LLMs via Low-Rank Adaptation (LoRA) fine-tuning, exploring three attack types (Bias, Denial of Service, and Hallucination), three trigger mechanisms (Rare/OOV symbol, Word Token, and Semantic phrase), and three open-source 7B models (LLaMA-2, Mistral, and Qwen-2.5) across multiple hyperparameter configurations. To support this research, a unified framework named GenerateMyBackdoor was developed, consolidating automated dataset generation, LoRA fine-tuning via LlamaFactory, and multi-metric evaluation into a single accessible interface. Experimental results demonstrate that backdoor attacks via LoRA fine-tuning are highly effective and stealthy, with backdoored models maintaining Attack Success Rate without trigger (ASRw/o) scores of 90–100% on non-triggered inputs across all configurations. Denial of Service emerged as the most reliable attack type with an average Attack Success Rate with trigger (ASRw/t) of 93.8%, while Rare/OOV triggers consistently outperformed Semantic triggers by a substantial margin (86.7% vs 45.1%). A Minimum Viable Backdoor analysis further revealed that a fully functional backdoor can be injected with as few as 20 poison samples and a single training epoch, demonstrating that the barrier to mounting such an attack is alarmingly low. Building on these results, this research also demonstrated how LLM-as-a-Judge frameworks can be significantly strengthened when informed by detailed attack signatures, offering a viable path for real-time backdoor detection in deployed systems. These findings highlight a critical vulnerability in the open-source LLM fine-tuning supply chain and underscore the urgent need for greater security awareness and defensive research in this space.
| Item Type: | Final Year Project / Dissertation / Thesis (Final Year Project) |
|---|---|
| Subjects: | L Education > L Education (General) T Technology > T Technology (General) |
| Divisions: | Faculty of Information and Communication Technology > Bachelor of Computer Science (Honours) |
| Depositing User: | ML Main Library |
| Date Deposited: | 22 Jul 2026 22:04 |
| Last Modified: | 22 Jul 2026 22:04 |
| URI: | http://eprints.utar.edu.my/id/eprint/7733 |
Actions (login required)
| View Item |

