Agent auditing
Evaluating and auditing AI agents at scale
I run large-scale behavioral analyses of real-world coding, GUI, and tool-use agent sessions, and build benchmarks that reveal the gap between clean-task performance and failure-aware, trustworthy behavior.















