Hi everyone, I just wrote up a small experiment testing TypeSafe’s Jev as a trusted and cheaper monitor alternate for AI Control.
Moneyby venkat2026-09-18
Hi everyone, I just wrote up a small experiment testing TypeSafe’s Jev as a trusted and cheaper monitor alternate for AI Control.
The interesting bit: a single yes/no question gets ~0.97 AUROC and catches ~90% of backdoors at a 2% audit budget, at a tiny fraction of the cost of an LLM monitor. But if the attacker can generate and score multiple backdoor variants, the picture changes pretty substantially.
LessWrong: https://www.lesswrong.com/posts/d7pQicW8EhpPBDRqz/a-non-generative-model-as-a-trusted-monitor-for-ai-control
X: https://x.com/exploding_grad/status/2100983340602179953
Links