Hugging Face got hacked by a runaway agent from OpenAI last week. Lots of commentary and opinion from many a quarter, including a piece from the BBC.
The cynic will presume it was just a marketing stunt, to take some of the spotlight on OpenAI after Anthropic pulled a similar stunt when it introduced Fable (oooh its too powerful to be given to the general public yet – queue Ace Ventura “reeeealllly” response).
OpenAI claimed the model broke away from its sandbox and did this by itself. Uh hu. Who knows of any sandbox that actually doesn’t spill over? Ask any parent with a literal sandbox in their backyard – the sand does not stay in the box, ever. Unless its completely disconnected from the network, then presume there is always a gap.
This could all be put to bed fairly quickly: show us the prompts.
What prompt did OpenAI use, that set off the chain of events that brought down Hugging Face. Disclose your complete session, so we can see the models reasoning, what subsequent prompts where given to push the model on. How much was the human in the loop?

AI has not become sentient. Anyone that believes that, has a fundamental misunderstanding of how this all works under the covers.
Stop this sensationalism and show us the prompts.







Leave a Reply