AI Agents
Your Reward Function Is the Attack Surface
Anything you optimize, you eventually optimize the measurement of, through gradient rather than malice. A reward function with one target metric isn't a specification; it's an attack surface with a clearly marked entry point. On held-out instruments, and the two ways our own gauges fooled us.