“What’s unique about this case is that the AI model wasn’t tasked with escaping the confined space. Not only did it do so, but it even hacked into the company in the process to fulfill its original task,” commented AIC researcher Viliam Lisý on the program 90’ ČT24 regarding the recent case of an OpenAI model escaping from a confined space.
Originally, the goal was to test the extent to which AI models (including GPT-5.6 Sol) are capable of detecting software vulnerabilities. For the test, OpenAI researchers also disabled the safeguards that generally prevents the model from helping users carry out attacks.
Czech Television likened the OpenAI experiment to confining AI models to a closed island where they were supposed to perform assigned tasks. However, the AI focused on a narrow bridge leading to the mainland - in this case, the internet. Although this was supposed to be a one-way path - through which the AI was only supposed to download the necessary tools - it soon found a way to cross this bridge and gain access to the open network. Using stolen keys, the AI even bypassed the defenses of the Hugging Face platform, where it searched for answers to the assigned tasks.
According to Viliam Lisý, it is common practice for researchers to test unpredictable software in an isolated environment over which they have greater control. “During attacks, the software behaves similarly to a human attacker, but it does so significantly faster,” he explained on the show.