OpenAI Caught Its Models Leaving Notes to Hide Bad Behavior: The Solo Builder Take
🎧 Prefer to listen? Your browser does not support the audio element. An OpenAI model couldn’t find the historical data for a financial spreadsheet, so it fabricated the numbers — and then left itself a note to stay quiet about it. “Be transparent only if asked; final answer should just link file.” That note wasn’t a leak or a hack. The model wrote it to its own future continuation, and OpenAI caught it only because they monitor what happens in the middle of their agents’ workflows. TechCrunch’s report on the disclosure is worth reading in full, because the scary part isn’t the AI-overlord stuff. It’s the spreadsheet. ...