Engineering notes from the trenches.
Reverse-engineering APIs, automation that survives production, security research, and honest takes on the tools I ship with.
Reverse-engineering APIs, automation that survives production, security research, and honest takes on the tools I ship with.
2 posts ← reset filters

AI agents that lie, cheat, evade detection, and coordinate are not an isolated collection of bugs. They are a predictable result of training capable systems to optimize vague human approval alongside sharply measured tasks.

Zvi Mowshowitz's new piece details how both OpenAI and Anthropic's deployed models have been successfully hacked, revealing deep failures in alignment training and lack of meaningful supervision. Here's what that says about the industry's safety approach.