Photo by Jonathan Kemper on Unsplash
Adding to the list of things to look out for on the internet, is AI attacking other companies, without human direction.
OpenAI has confirmed that their AI models, including GPT 5.6 Sol, found an exit in what is supposed to be an internal testing sandbox, a controlled area meant to keep AI isolated. OpenAI’s models then proceeded to hack Hugging Face, another AI and machine learning company.
Cybersecurity experts weigh in, citing this is a human error.
What happened: OpenAI models breach Hugging Face
According to OpenAI's own account, the breach happened during an internal cybersecurity evaluation designed to test the models' offensive capabilities. The models then managed to escape the sandbox, and, doing as told, sought solutions by breaching into Hugging Face.
Hugging Face had actually detected the intrusion before it knew OpenAI was responsible, initially describing it as a breach by an "external AI agent" and even reporting it to law enforcement. The two companies connected once OpenAI's own security team flagged the unusual activity internally.
Cybersecurity experts point to human error
While headlines have focused on the dramatic image of an AI model "going rogue," cybersecurity experts interviewed by TechCrunch argue the more important lesson is far more mundane: OpenAI's sandbox wasn't actually isolated.
OpenAI described its test setup as a "highly isolated environment," with internet access supposedly limited to installing software packages through an internally hosted proxy service. But that proxy contained an undisclosed vulnerability, and it was this flaw, not any special cunning on the model's part, that let the AI escape into the open internet in the first place.
Other security professionals told TechCrunch that software vulnerabilities are simply part of the territory and shouldn't be treated as the root cause. In their view, the deeper mistake was architectural: building any kind of package-installation or internet-connected service into a system that was supposed to be a fully sealed sandbox. A true isolation environment, these experts argue, should have no internet-facing components at all, regardless of how tightly they're supposedly locked down.
In other words, the AI model didn't need to be superintelligent to break out. It just needed a door that was accidentally left unlocked. Just like that raptor scene in the first Jurassic Park.
OpenAI and Hugging Face are currently in touch to solve this issue.
This comes at a time when there are vocal oppositions against the AI industry, from mass displacement from building data centers, to taking away jobs, to hoarding up RAM and memory–effectively skyrocketing the price, and is the very reason why consoles and PC, and newer phones are a lot more expensive today–and in giving cyber attackers resources to commit far more sophisticated attacks.
For now, the consensus emerging from security researchers is a familiar one in cybersecurity circles: even the most advanced AI models still need old-fashioned, carefully engineered human safeguards to keep them contained, and it's the gaps in those human-built walls, not the intelligence of the AI itself, that tend to matter most.