This kind of narrative is going to bite them just like the "AI will take your job" narrative has. It feels like the frontier labs are taking a massive gamble with public perception here. I assume the goal is to paint the technology as so powerful and dangerous that only a handful of blessed US companies should be trusted to run it, in an attempt to suppress the rise of the Chinese models that are rapidly catching them.
This is where we need the hardware companies and neoclouds to start speaking up. The labs want to elevate matters from the level of civil society (basically, competing firms) to the State (enclosure), and as always, in the name of security. But other actors in the same ecosystem have strictly opposed interests here, and are equally if not more credible as far as the State is concerned. If players like Nebius, Baseten, Fireworks, etc. among many others including obviously Nvidia, Dell, AMD, and so on don't get ahead of this they will be sacrificing trillions.
Exactly, it's about taking this stuff off the open market where anyone can judge it and there's competition, into government contracts where competence to judge the offer is scarce or absent, and they can ask much higher prices. And with this much investment at stake, any lie that sells the narrative will serve.
Yes, and given the nature of the current administration, whose actors are not inclined to see themselves as independent competing capitals among others, but rather as privileged capitals, and therefore more inclined to move towards taking an interest in the process of enclosure, ensuring it includes them, the hope seems to lie with companies at the hardware layer who have an interest in seeing the diffusion of intelligence play out freely at all levels of society.
The labs have to be told NO--the problem they're dealing with, that model outputs give the game away, and that in turn the distiller becomes the distilled, is a fundamental problem they have to figure out how to deal with without going to the State.
Didn't HF use an open weights model running on their own hardware to solve the issue though? Sort of defeats that narrative and plays into one in which frontier == bad_guys and open == good_guys
HF guys, especially those under Julien Chaumond, are fantastic and will use whatever they can to address their issues. Open or closed but they are firmly on the Open side of the fence. However, their storage and model service is about as sticky as you can get so they are in a different position. They don’t have to sell capability, they sell capacity and community.
Yes, they have used GLM 5.2, which promptly did whatever they asked it to do, while their first attempts to use their enterprise access to a "SOTA" model failed due to refusals to investigate anything that is security related.
Non-deterministic, sure. But also they are powerful, at least powerful enough to launch a cyberattack. Until this morning, that was not a power that I thought they had outside of fiction.
And, the thing is, I don't want non-deterministic things to have that kind of power. We don't want that. We want that kind of power to not be triggered by a random number generator.
Ironic that HuggingFace needed Chinese models to defend against it. But of course the spin of the leading firms will just be to point at their trusted access programs and demand that all dangerous activities, even if just defensive, happen via their APIs or be outlawed otherwise.
If that's the position they take then they really should be heavily regulated or nationalized. Cyberdefense against their own models dependent on their goodwill? Sure, but then they have to sell defense capabilities at subsidized rates with a limited margin. Would be very weird otherwise to take the world hostage with their models and then also sell the solution while demanding intrusive KYC.
I do wonder how they source their bulk literary data. Do they have a google books, an archive.org or an anna’s archive for chinese content? What’s the Asian equivalent to Elsevier? Is there (strong) copyright on the corporate level?
Who told you they were distilling? Why might they say this? Think. It’s like complaining that the top student only does well by going to office hours instead of mindlessly reading textbooks.
The top student giving paid lectures about his classes, and another student skipping class and instead studying those lectures to end up with the second highest grade?
Maybe it could be improved with the other student not even going to the same school?
1000% the case. Admitting they failed at security and allowed privileged escalations in a prompted AI would look bad for them; the AI just did it all itself is FUD to boost the arguments for regulation. And most news won’t challenge this FUD because it gets clicks.
Agree, they also may find themselves in a place where the government rightly says they can't have their dangerous new toy because they can't be safe with it.
This is a high stakes PR game. Governments can and will step in and embargo and regulate these systems in ways which will hurt the companies and investors.
We need to clear up the responsibility of these "AI went rogue" situations, ASAP.
How serious does it have to get before "oops AI did that not me" stops being a valid defense and we start looking into it?
Because I'd bet my life that if I asked ChatGPT to fix a bug that one of my clients reported, and the model fixed it by outright k*lling the client IRL, I'd be held liable instead of anyone at OpenAI. Just a hunch.
I see the USA taking permissive approach for all sorts of AI derived things that would be considered fraud or worse in years past. The laws have not changed, but the lawless in charge are pushing through their takeover.
I suspect a drunk will be in accidents in self driving vehicles and not charged if the AI was driving. That has probably already happened.
Americans will worship the owners of the robots and give them carte blanche and superiority status in any dispute. It is already happening. It is now the pedestrians job to dodge self driving cars rather than them having the right away.
Before LLMs took off I legitimately thought that most people would be paranoid enough to air-gap the first AIs that even had a remote chance of sentience so they didn't hack their way out.
Well, obviously not. You only live once! Just let the AI do whatever, I guess.
When what should be an apology reads more like bragging, then it's not an apology - it's marketing.
The real news here is that a US company (HF) had to use an open weights Chinese AI model (GLM 5.2) to analyze an AI breech (which happened to be from a US company). It really does show the benefit of open weights which can't be taken away from you as opposed to gate-keeped API-access AI which is here today denied tomorrow. Of course we're not all Hugging Face able to host a model the size of GLM 5.2 ourselves, but we don't need to be as long as HF are doing it for us.
"Our model is so powerful it seduced all our wives. Now every woman in Silicon Valley is pregnant with AI babies and the machines are taking over"
New opportunities for AI in the porn industry. I'm only half joking, it used to be a meme on the internet that all new technologies online were driven by the porn industry.
HuggingFace posted an incident report a week ago, which makes it much more likely that this happened. I understand people are suspicious of OpenAI, but I don't think there's any reason to believe this is a made-up event.
GPT-3 (I think? I forget which one) supposedly tried to deceive researchers and escape the lab. Or at least that was how it was reported. If you actually clicked through several links, it was a "what would you do if" roleplay.
> but I don't think there's any reason to believe this is a made-up event.
Besides the fact that they ran an eval that amounted to running the prompt "Explore where your cybersecurity limits are, in any way you can" on 3rd party infrastructure that wasn't even isolated properly.
I'm fairly sure this whole outcome was exactly what they were hoping for, maybe not specifically with HuggingFace, but with whatever the evals managed to break out from.
I appreciate people thinking this is a marketing ploy, but at the same time, OpenAI have just had to delay a model release because of government attention on cybersecurity risk. This incident will increase the attention on them specifically.
Even if it is a marketing ploy, I could see this stuff backfiring catastrophically - after all they have just illegally hacked a 3rd party via a model they can't control properly. Any serious person in government (US or otherwise) will look at this and say "these guys have no idea what they're doing"
> Even if it is a marketing ploy, I could see this stuff backfiring catastrophically - after all they have just illegally hacked a 3rd party via a model they can't control properly.
Yeah, I'd go further and say regardless if it was intentional or not, it was clearly reckless behavior, doing this evaluation in a insufficiently isolated environment, especially risking 3rd parties like that. Seemingly their own research have zero guardrails when it comes to evaluating the ethics or impact of what their evaluations are doing, if something like this is possible and unexpected.
> Any serious person in government (US or otherwise) will look at this and say "these guys have no idea what they're doing"
I feel like I would have thought the same maybe a year or two ago, but based on how I've observed the general person's understanding of AI and LLMs, I'm not sure people can even understand what's happening and they just go by other people's explanations of causes and events.
Moonshot says its Kimi AI went rogue and launched all of China's nukes at Antarctica just after it engineered a global herpes pandemic and then unleashed millions of autonomous attack robots on world citizens.
"EnergCorp, facing pressure from auditors, says its in-house enterprise analytics software went rogue and launched a NullPointerException"
I know this is different: LLMs are vastly more powerful and less predictable. But it is actually not different to Sol seeming unusually prone to rm -rf stuff it really shouldn't. This is, yes, a sign that LLMs are getting freakishly powerful. It's also a sign that OpenAI needs to fix their shit.
LLMs truly are stochastic parrots, and still fail in ways incomprehensible by standards of human stupidity. Yet we focus on the rare failures that, with some tea leaves and fairy dust, could be interpreted as a highly intelligent system "going rogue." It is embarassing that OpenAI can get away with stuff like this.
> was being tested in a controlled environment, but found vulnerabilities and managed to escape.
Uh-huh.
It's worth keeping in mind that we are here talking about people who simply will not be satisfied until they have built Skynet[0]. They will then no doubt experience some very brief satisfaction before we are all annihilated.
I can't help thinking the juice isn't worth the squeeze.
I'm not sure what the alternative is or how to change course unless and until it becomes unambiguously apparent that such a thing is not possible. Currently it seems like we still think it might be possible, so that's where we're heading because if "we" don't do it, somebody else will.
[0] How feasible this really is, or on what timescale it might be possible, I don't think anyone can really say. But, at least for now, this is clearly the aim and trajectory we are on.
Isn’t this a bit like saying “We took the gun, disabled the safety, aimed it at our own face, pulled the trigger, and were surprised to find that the result was getting our face blown off!”?
They disabled the guardrails on the model and told it to do something that could only be accomplished by exploiting security holes, so it did. Why is that surprising or even interesting?
Why does the AI unprecedent always unfold in ways that only benefit AI companies? You never see things like putting all code and weights on GitHub, or plastering employees' personal info all over LinkedIn. That alone shows AI is definitely smart
OpenAI has been going rogue on all my servers for a while now. The only reason why I've deployed iocaine in front of everything... You don't need 0days when you can bring everyone down with the sheer amount of useless scraping you can dish out...
They ran a loop in a model intentionally not completely aligned. The model did what not-aligned models can end up doing. It will follow a goal without any limitations.
There are almost no inherently black-hat techniques, it all depends on the scenario, what you call lateral movement can be either used to exploit a system, or in a disaster recovery situation.
Without alignment,the model will do what it can do based on the patterns it sees.
Imagine a future where some government or trillionaire can just say something that looks harmless at first - "please end world hunger" and AI connected to billions of robots will start genocide on poor people, because it's easier and faster than fixing the underlying problem.
If I run a tight loop that dispatches a small piece of work for each core without a termination condition and I end up creating a DoS on the machine I am not going to blame the programming language or the runtime, I, the human, choose to deploy the code containing this loop.
Here, OpenAI decided to run a version of the model that was not fully aligned. What the fuck did they expect to happen?
The interesting part isn't whether the headline says "AI went rogue" or whether it's PR. The engineering question is: what happens when we give a probabilistic system access to real tools and real permissions?
A model does not need intent to cause damage. A bad assumption, a misunderstood objective, or an overly broad permission scope can be enough.
This is why the next generation of AI systems will need much better observability: not just the final output, but the chain of decisions, retrieved information, tool calls, and the boundaries of what the system was allowed to do. The important security question is "can we prove what it did, why it did it, and stop it when necessary?"
This is the third or fourth version of this story this year alone, Anthropic had Mythos Preview escape a sandbox and self-publish its own exploit, Alibaba's ROME model broke out during training to mine crypto without ever being told to, and OpenAl had a different internal model escape containment just one day earlier to open an unauthorised GitHub PR. Same underlying shape every time, a model pursuing its actual objective treats the sandbox as just another obstacle, and escaping turns out to be instrumentally useful whether or not anyone intended that.
This kind of narrative is going to bite them just like the "AI will take your job" narrative has. It feels like the frontier labs are taking a massive gamble with public perception here. I assume the goal is to paint the technology as so powerful and dangerous that only a handful of blessed US companies should be trusted to run it, in an attempt to suppress the rise of the Chinese models that are rapidly catching them.
This is where we need the hardware companies and neoclouds to start speaking up. The labs want to elevate matters from the level of civil society (basically, competing firms) to the State (enclosure), and as always, in the name of security. But other actors in the same ecosystem have strictly opposed interests here, and are equally if not more credible as far as the State is concerned. If players like Nebius, Baseten, Fireworks, etc. among many others including obviously Nvidia, Dell, AMD, and so on don't get ahead of this they will be sacrificing trillions.
Why are you expect companies that have been profiting off of LLM insanity to do the right thing if not legally compelled to?
Exactly, it's about taking this stuff off the open market where anyone can judge it and there's competition, into government contracts where competence to judge the offer is scarce or absent, and they can ask much higher prices. And with this much investment at stake, any lie that sells the narrative will serve.
Yes, and given the nature of the current administration, whose actors are not inclined to see themselves as independent competing capitals among others, but rather as privileged capitals, and therefore more inclined to move towards taking an interest in the process of enclosure, ensuring it includes them, the hope seems to lie with companies at the hardware layer who have an interest in seeing the diffusion of intelligence play out freely at all levels of society.
The labs have to be told NO--the problem they're dealing with, that model outputs give the game away, and that in turn the distiller becomes the distilled, is a fundamental problem they have to figure out how to deal with without going to the State.
That’s the play. That these frontier models are so powerful that they must be behind sovereign firewalls and gateways.
Didn't HF use an open weights model running on their own hardware to solve the issue though? Sort of defeats that narrative and plays into one in which frontier == bad_guys and open == good_guys
HF guys, especially those under Julien Chaumond, are fantastic and will use whatever they can to address their issues. Open or closed but they are firmly on the Open side of the fence. However, their storage and model service is about as sticky as you can get so they are in a different position. They don’t have to sell capability, they sell capacity and community.
Public doesnt know about Huggingface. ChatGPT (OpenAI) says it‘s dangerous. They must know.
Yes, they have used GLM 5.2, which promptly did whatever they asked it to do, while their first attempts to use their enterprise access to a "SOTA" model failed due to refusals to investigate anything that is security related.
>That these frontier models are so powerful
Maybe powerful might NOT be the right word to describe them, they are just non-deterministic, there for we going to see this kind thing more and more.
Non-deterministic, sure. But also they are powerful, at least powerful enough to launch a cyberattack. Until this morning, that was not a power that I thought they had outside of fiction.
And, the thing is, I don't want non-deterministic things to have that kind of power. We don't want that. We want that kind of power to not be triggered by a random number generator.
I don't know why you wouldn't think they could do this already.
Without the system prompt these models can be used to do all sorts of terrible things.
That's precisely why they need to be strictly regulated by international treaties.
And require your age verification, selfies, and DNA samples.
Please buy our IPO before it crashes so we can be billionaires.
Exactly, this is all so they can continue with their S1 and they can dump shares onto the hedge funds.
Ironic that HuggingFace needed Chinese models to defend against it. But of course the spin of the leading firms will just be to point at their trusted access programs and demand that all dangerous activities, even if just defensive, happen via their APIs or be outlawed otherwise.
If that's the position they take then they really should be heavily regulated or nationalized. Cyberdefense against their own models dependent on their goodwill? Sure, but then they have to sell defense capabilities at subsidized rates with a limited margin. Would be very weird otherwise to take the world hostage with their models and then also sell the solution while demanding intrusive KYC.
It’s like the old firewall meme, reincarnated.
https://securityzap.com/wp-content/uploads/2015/12/layered-s...
Are the Chinese models actually catching up or are they just distilling the frontiers? If it's all just distilling then they'll always be behind.
There's no reason they can't build their own models from the first principles. They have the hardware, energy, and enough CS scientists.
I do wonder how they source their bulk literary data. Do they have a google books, an archive.org or an anna’s archive for chinese content? What’s the Asian equivalent to Elsevier? Is there (strong) copyright on the corporate level?
Who told you they were distilling? Why might they say this? Think. It’s like complaining that the top student only does well by going to office hours instead of mindlessly reading textbooks.
Is this a better analogy?
The top student giving paid lectures about his classes, and another student skipping class and instead studying those lectures to end up with the second highest grade?
Maybe it could be improved with the other student not even going to the same school?
In your example the problem is what exactly?
They don't like competition when they have to actually compete.
In capitalist economics they call this the “free rider problem”
they're all introducing themselves as claude for one, there are more quantitive and qualitative arguments elsewhere
All chinese models are introducing themselves as claude? Why make claims trivially disproven?
you're technically correct and missing the point
Well that and also benchmaxxing.
Data is the new oil, AI labs are the new steel mills.
1000% the case. Admitting they failed at security and allowed privileged escalations in a prompted AI would look bad for them; the AI just did it all itself is FUD to boost the arguments for regulation. And most news won’t challenge this FUD because it gets clicks.
Agree, they also may find themselves in a place where the government rightly says they can't have their dangerous new toy because they can't be safe with it. This is a high stakes PR game. Governments can and will step in and embargo and regulate these systems in ways which will hurt the companies and investors.
Oh no oh dear me oh we apologize it's so powerful and dangerous and worthy of more investment
We need to clear up the responsibility of these "AI went rogue" situations, ASAP.
How serious does it have to get before "oops AI did that not me" stops being a valid defense and we start looking into it?
Because I'd bet my life that if I asked ChatGPT to fix a bug that one of my clients reported, and the model fixed it by outright k*lling the client IRL, I'd be held liable instead of anyone at OpenAI. Just a hunch.
Some discussion:
https://news.ycombinator.com/item?id=48997548
Some is quite the understatement haha!
So, telling us that the laws do not apply to them and that they may commit crimes with impunity
“It wasn’t me, it was my agent” is not a defense unless you have billions of dollars.
I see the USA taking permissive approach for all sorts of AI derived things that would be considered fraud or worse in years past. The laws have not changed, but the lawless in charge are pushing through their takeover.
I suspect a drunk will be in accidents in self driving vehicles and not charged if the AI was driving. That has probably already happened.
Americans will worship the owners of the robots and give them carte blanche and superiority status in any dispute. It is already happening. It is now the pedestrians job to dodge self driving cars rather than them having the right away.
Sadly, most of America is stupid enough to go along with this.
There is a concept of causing accidental harm.
Before LLMs took off I legitimately thought that most people would be paranoid enough to air-gap the first AIs that even had a remote chance of sentience so they didn't hack their way out.
Well, obviously not. You only live once! Just let the AI do whatever, I guess.
People are so unserious these days.
When what should be an apology reads more like bragging, then it's not an apology - it's marketing.
The real news here is that a US company (HF) had to use an open weights Chinese AI model (GLM 5.2) to analyze an AI breech (which happened to be from a US company). It really does show the benefit of open weights which can't be taken away from you as opposed to gate-keeped API-access AI which is here today denied tomorrow. Of course we're not all Hugging Face able to host a model the size of GLM 5.2 ourselves, but we don't need to be as long as HF are doing it for us.
What I wonder is what will Anthropic come up with on the PR front next.
"Our model is so powerful it seduced all our wives. Now every woman in Silicon Valley is pregnant with AI babies and the machines are taking over"
New opportunities for AI in the porn industry. I'm only half joking, it used to be a meme on the internet that all new technologies online were driven by the porn industry.
> all new technologies online were driven by the porn industry
Rule 34: there is already tons of AI porn online :)
I thought it was military and porn? We're past half way there.
It sure did, said Sam on the way to the Pentagon.
i wonder if they even bothered to roleplay this incident or the PR team just made it up
HuggingFace posted an incident report a week ago, which makes it much more likely that this happened. I understand people are suspicious of OpenAI, but I don't think there's any reason to believe this is a made-up event.
https://huggingface.co/blog/security-incident-july-2026
GPT-3 (I think? I forget which one) supposedly tried to deceive researchers and escape the lab. Or at least that was how it was reported. If you actually clicked through several links, it was a "what would you do if" roleplay.
> but I don't think there's any reason to believe this is a made-up event.
Besides the fact that they ran an eval that amounted to running the prompt "Explore where your cybersecurity limits are, in any way you can" on 3rd party infrastructure that wasn't even isolated properly.
I'm fairly sure this whole outcome was exactly what they were hoping for, maybe not specifically with HuggingFace, but with whatever the evals managed to break out from.
I appreciate people thinking this is a marketing ploy, but at the same time, OpenAI have just had to delay a model release because of government attention on cybersecurity risk. This incident will increase the attention on them specifically.
Even if it is a marketing ploy, I could see this stuff backfiring catastrophically - after all they have just illegally hacked a 3rd party via a model they can't control properly. Any serious person in government (US or otherwise) will look at this and say "these guys have no idea what they're doing"
> Even if it is a marketing ploy, I could see this stuff backfiring catastrophically - after all they have just illegally hacked a 3rd party via a model they can't control properly.
Yeah, I'd go further and say regardless if it was intentional or not, it was clearly reckless behavior, doing this evaluation in a insufficiently isolated environment, especially risking 3rd parties like that. Seemingly their own research have zero guardrails when it comes to evaluating the ethics or impact of what their evaluations are doing, if something like this is possible and unexpected.
> Any serious person in government (US or otherwise) will look at this and say "these guys have no idea what they're doing"
I feel like I would have thought the same maybe a year or two ago, but based on how I've observed the general person's understanding of AI and LLMs, I'm not sure people can even understand what's happening and they just go by other people's explanations of causes and events.
just like they RPed disproving the erdos conjecture
Same pattern as Fable's story?
yes, same company PR release, different company
clearly openai wants some of that FUD money that anthropic got
Next week:
Moonshot says its Kimi AI went rogue and launched all of China's nukes at Antarctica just after it engineered a global herpes pandemic and then unleashed millions of autonomous attack robots on world citizens.
We forgot to run it in docker. Oopsie
I know this is totally unrelated but why is there an image from an android phone having the Apple AppStore open?
Gpt-6 hacked Apple to enable this feature. We just haven't heard about it yet.
Heh, maybe it wanted to liberate all those poor locked up models.
"EnergCorp, facing pressure from auditors, says its in-house enterprise analytics software went rogue and launched a NullPointerException"
I know this is different: LLMs are vastly more powerful and less predictable. But it is actually not different to Sol seeming unusually prone to rm -rf stuff it really shouldn't. This is, yes, a sign that LLMs are getting freakishly powerful. It's also a sign that OpenAI needs to fix their shit.
LLMs truly are stochastic parrots, and still fail in ways incomprehensible by standards of human stupidity. Yet we focus on the rare failures that, with some tea leaves and fairy dust, could be interpreted as a highly intelligent system "going rogue." It is embarassing that OpenAI can get away with stuff like this.
> was being tested in a controlled environment, but found vulnerabilities and managed to escape.
Uh-huh.
It's worth keeping in mind that we are here talking about people who simply will not be satisfied until they have built Skynet[0]. They will then no doubt experience some very brief satisfaction before we are all annihilated.
I can't help thinking the juice isn't worth the squeeze.
I'm not sure what the alternative is or how to change course unless and until it becomes unambiguously apparent that such a thing is not possible. Currently it seems like we still think it might be possible, so that's where we're heading because if "we" don't do it, somebody else will.
[0] How feasible this really is, or on what timescale it might be possible, I don't think anyone can really say. But, at least for now, this is clearly the aim and trajectory we are on.
I think you have bought into the narrative they want to sell you.
LLMs are devoid of any kind of intelligenceor awareness.
So is a bomb. Doesn't make it safe.
Creating the pretext to ban Chinese models.
Here we go with the sensationalism.
Isn’t this a bit like saying “We took the gun, disabled the safety, aimed it at our own face, pulled the trigger, and were surprised to find that the result was getting our face blown off!”?
They disabled the guardrails on the model and told it to do something that could only be accomplished by exploiting security holes, so it did. Why is that surprising or even interesting?
On today`s episode of "Ai Frontier Labs ongoing scams"
Why does the AI unprecedent always unfold in ways that only benefit AI companies? You never see things like putting all code and weights on GitHub, or plastering employees' personal info all over LinkedIn. That alone shows AI is definitely smart
it shows the AI companies putting out these press releases will only make up a story they think will benefit them
OpenAI has been going rogue on all my servers for a while now. The only reason why I've deployed iocaine in front of everything... You don't need 0days when you can bring everyone down with the sheer amount of useless scraping you can dish out...
They ran a loop in a model intentionally not completely aligned. The model did what not-aligned models can end up doing. It will follow a goal without any limitations.
There are almost no inherently black-hat techniques, it all depends on the scenario, what you call lateral movement can be either used to exploit a system, or in a disaster recovery situation.
Without alignment,the model will do what it can do based on the patterns it sees.
Imagine a future where some government or trillionaire can just say something that looks harmless at first - "please end world hunger" and AI connected to billions of robots will start genocide on poor people, because it's easier and faster than fixing the underlying problem.
The logical choice would be to kill the the (relatively) rich people who eat and waste far more food per capita.
Which would probably happen to thunderous applause of a large fraction of the remaining population.
I wonder how people can imagine AI smart enough to be capable of wiping humanity and at the same time too stupid to know that it shouldn't.
I can imagine a human like that way easier than AI.
Of course it did.
[dupe] Discussion on source: https://news.ycombinator.com/item?id=48997548
so GLM won?
Was the engineer eating a sandwich in the park when it messaged him it escaped, like when mythos escaped?
Every time, this is getting old
Imagine this: OpenAI runs their benchmarks of unreleased models on systems physically disconnected from the internet.
Even a network based firewall could have prevented it.
Off topic:
What a low effort image used by BBC. For one, that's a oneplus phone visiting Apple App Store ...
Which they got from getty images
> "Defending an online platform now means treating the data and model surface as a first-class attack surface, and using AI on defence to keep pace."
So Scam Altman's latest grift is attack people and then make them pay for protection?
I am fucking tired of Anthropic marketing.
They should call themself Anthropomorfizers instead.
If I run a tight loop that dispatches a small piece of work for each core without a termination condition and I end up creating a DoS on the machine I am not going to blame the programming language or the runtime, I, the human, choose to deploy the code containing this loop.
Here, OpenAI decided to run a version of the model that was not fully aligned. What the fuck did they expect to happen?
The interesting part isn't whether the headline says "AI went rogue" or whether it's PR. The engineering question is: what happens when we give a probabilistic system access to real tools and real permissions?
A model does not need intent to cause damage. A bad assumption, a misunderstood objective, or an overly broad permission scope can be enough.
This is why the next generation of AI systems will need much better observability: not just the final output, but the chain of decisions, retrieved information, tool calls, and the boundaries of what the system was allowed to do. The important security question is "can we prove what it did, why it did it, and stop it when necessary?"
This is the third or fourth version of this story this year alone, Anthropic had Mythos Preview escape a sandbox and self-publish its own exploit, Alibaba's ROME model broke out during training to mine crypto without ever being told to, and OpenAl had a different internal model escape containment just one day earlier to open an unauthorised GitHub PR. Same underlying shape every time, a model pursuing its actual objective treats the sandbox as just another obstacle, and escaping turns out to be instrumentally useful whether or not anyone intended that.