A Watchdog Model for Backdoored LoRA Adapters
Training a small distilled model to recognize the statistical fingerprint of a backdoored language model, instead of searching for a known trigger word.
Read post
The Sentinel's Dilemma: Guarding AI from Hidden Threats
A survey of backdoor attacks on large language models — how hidden triggers are implanted during training, and the detection and defense strategies emerging against them.
Read postTurning a Backdoor's Step Function Into a Ramp
A research design for eliciting a hidden model behavior by amplifying the exact fine-tuning change that planted it, then dialing that amplification back down to isolate the trigger.
Read postAuditing Backdoors: From a Flag to a Causal Proof
Extending a trigger-reconstruction pipeline for fine-tuned LLMs with a stricter auditing layer that separates a plausible-looking flag from actual proof of a backdoor.
Read postToward a Field-Based View of LLM Backdoors
A design study exploring whether a backdoor can be caught by looking at everything fine-tuning changed, rather than searching for a trigger or a target string.
Read post
Building a Weakness Zoo to Stress-Test a Backdoor Scanner
Deliberately building harder backdoor variants to find the blind spots of a state-of-the-art LLM backdoor scanner.
Read postComparing Five Ways to Catch a Backdoor, Live
A laptop-friendly lab for finetuning a small backdoored model and comparing several detection signals on clean vs. triggered inputs.
Read postUnderstanding BackdoorBench: A Comprehensive Benchmark for AI Security
A walkthrough of BackdoorBench, the standardized benchmark for evaluating backdoor attacks and defenses in deep learning — what it measures, why it matters, and how to use it.
Read post