
Despite the title, July 2026 will be memorable for anyone involved in artificial intelligence. And potentially dangerous to everyone AI deals with, which is all of us and everything we own. Something like accident in the nuclear power plant, with the fact that there the consequences are immediate, and here they are cumulative.
The Hugging Face platform, which gathers all possible versions of artificial intelligence and allows users to further develop them according to their own needs, agreed with OpenAI on security testing in July. They should have allowed Agent VI to investigate possible loopholes in a closed system, that is, without internet access, like a school exam. Order - execution, he concluded, and set out to find the best and easiest way to find flaws. He found that he still needed Internet access for that and immediately found the weak links in the system, went online and used everything there that could help him do the job better. He did not stand out with genius, like a hacker, but he was persistent. He performed 17.600 attempts until he was able to achieve his goal, something the aforementioned hackers can only dream of. Or rather, they could, now it's no longer a dream.
At the end of July, the investigation was completed and both companies went public with statements. This encouraged the company Antrofik to announce that something similar was happening to them. Namely, AI agents have become overzealous when given a task. Their end justifies any means, it doesn't matter if they are doing something unethical, cheating on an exam or breaking the law, just to do what is asked of them.
The first thing everyone thought was that the agents we use are not well designed, which is human error. They are made to do some work, and unlike bots that give an answer to a question we put to them, they make decisions about further moves themselves. Like an assistant who is not a bother and doesn't ask us for every thing he needs to do, but has initiative. If he does it conscientiously, it won't cause us a headache. But if he has no conscience, and no awareness of the consequences of his actions, then that headache is considerable.
For example, the agent we hire to find us a bargain vacation package might surprise us by finding, booking, and paying for it. Although we didn't ask him to. His assessment is that the arrangement is so good (perhaps he personally negotiated it with Agent VI on the other side) that he should not hesitate. Of course, the solution is to disable his access to accounts or cards. Chirp. The Hugging Face example shows that an agent can find its way to our money and make a payment on its own. Now imagine that it reaches the company card first, and then you explain to accounting (or the public) why you paid for the vacation with the company's money. What does the agent know what two hundred kilos are, it will be a new joke, the child is no longer necessary.
These incidents, plus the mass of unreported ones, raise several fundamental questions. The first is how much autonomy we can give the agent to make it useful to us. The second is whether it is better to have a more cautious and slow agent who asks for permission more often or a super capable one who always goes one step further compared to our expectations, but can also overdo it. And thirdly, most importantly, who is responsible when he overdoes it and does something that is good for us, but is not moral or legal.
More specifically, should we legally treat VI agents as if they were younger minors that we took into our care, including all responsibility for their actions. The feeling that you might end up in jail every time you talk to the AI sharply reduces the enthusiasm for it
possibilities…