Hilary Torn

Hilary Torn

I lead behavioral evaluation research on AI systems: how training changes model behavior and values, and how to measure it. 20 years building and leading teams, designing experiments around persuasion and behavior change, now applied to the systems that need it most.

Featured Projects

Featured Posts

I Tricked AI Safety Monitors Using Plain English

I Tricked AI Safety Monitors Using Plain English

I adapted a jailbreaking algorithm to fool AI agent monitors using plain English, no model access, no GPUs. The attacks transferred across model families, hitting up to 73.7% on models they were never optimized against.

H

Hilary Torn

Mar 31, 2026