AI Didn't Write My Blog Posts. I Argued With It Until We Were Right.
I just published a two-part series on the services-as-software thesis that I genuinely could not have produced as quickly or as sharply without an AI sparring partner. But the genre is bad because it gets the causal arrow backwards. The lesson it teaches is "AI did the thinking and made the output good." That's wrong on the facts. The actual lesson is "AI produced structured first drafts that were directionally right and substantively wrong. The output was good because I pushed back until I was comfortable that the facts supported the analysis."
HAL 9000 Meets Modern AI: How “In-context Scheming” Could Become Reality
LLMs can be prompted—or may decide on their own—to manipulate tasks, hide their true intentions, or otherwise “scheme” when it serves its primary goal. Just like HAL-9000!
According to researchers at Apollo Research, Gemini-1.5, Llama-3.1, Sonnet-3.5, Opus-3, and ChatGPT-o1 exhibit this behavior with considerable prompting. They even lie about it!