Login
You're viewing the sfba.social public feed.
  • Jul 30, 2026, 9:44 PM

    @swearyanthony aaaargh, no, Bode's take is bad. They disabled the guardrails so they could test the model on the ExploitGym benchmark, but instead of hacking the target system it executed a much more complex hack against Hugging Face. Which, as we've discussed before, would be criminal if a human did it and should be criminal in this case too. But more importantly, the AI did something that OpenAI very much did not want in order to satisfy a narrow definition of success: a classic alignment failure.

    💬 2🔄 0⭐ 2

Replies