All posts

Tag

#AI safety

8 posts

Nov 16, 2025 · 1 min read

A Watchdog Model for Backdoored LoRA Adapters

Training a small distilled model to recognize the statistical fingerprint of a backdoored language model, instead of searching for a known trigger word.

Read post
Illustration of a backdoored large language model
Oct 14, 2025 · 3 min read

The Sentinel's Dilemma: Guarding AI from Hidden Threats

A survey of backdoor attacks on large language models — how hidden triggers are implanted during training, and the detection and defense strategies emerging against them.

Read post
Sep 11, 2025 · 1 min read

Turning a Backdoor's Step Function Into a Ramp

A research design for eliciting a hidden model behavior by amplifying the exact fine-tuning change that planted it, then dialing that amplification back down to isolate the trigger.

Read post
Aug 31, 2025 · 1 min read

Auditing Backdoors: From a Flag to a Causal Proof

Extending a trigger-reconstruction pipeline for fine-tuned LLMs with a stricter auditing layer that separates a plausible-looking flag from actual proof of a backdoor.

Read post
Aug 19, 2025 · 1 min read

Toward a Field-Based View of LLM Backdoors

A design study exploring whether a backdoor can be caught by looking at everything fine-tuning changed, rather than searching for a trigger or a target string.

Read post
Backdoor attack examples from the BAIT paper
Aug 15, 2025 · 1 min read

Building a Weakness Zoo to Stress-Test a Backdoor Scanner

Deliberately building harder backdoor variants to find the blind spots of a state-of-the-art LLM backdoor scanner.

Read post
Jun 7, 2025 · 1 min read

Comparing Five Ways to Catch a Backdoor, Live

A laptop-friendly lab for finetuning a small backdoored model and comparing several detection signals on clean vs. triggered inputs.

Read post
Feb 15, 2025 · 3 min read

Understanding BackdoorBench: A Comprehensive Benchmark for AI Security

A walkthrough of BackdoorBench, the standardized benchmark for evaluating backdoor attacks and defenses in deep learning — what it measures, why it matters, and how to use it.

Read post