@quinn Having read the HuggingFace write up with a healthy amount of scepticism, I'm pretty sure it's true. There are too many things that are embarrassing for both sides for it to be a marketing stunt:
- Model made an incorrect assumption (answers are hidden on Hugging Face somewhere) and proceeded to burn hundreds of millions or even billions of tokens based on this
- OpenAI had minimal security and almost no oversight of what their model under test was doing