In February 2024, I wrote We have given AI the keys to our kingdom.
I was never bothered that a machine would suddenly wake up, become conscious, decide it hated humans and start building Terminators. It was way more obvious than that.
In 2024, I was worried about what AI could do to our judgement. Exploit our biases, shape what we believed and influence what we did next. I still am. It never needed to become conscious or hate humanity to cause trouble.
Now we are giving it direct access to our systems as well. It can send the email, change the record and run the code. We give it tools, memory and permission to keep trying until it gets the job done. That sounds mighty useful, profitable and bloody exciting. Then when the shit hits the fan we act all surprised.
In that first article, I wrote about how we keep failing the marshmallow test. We want the reward now. The consequences can wait. Look at what we are doing with agents and tell me that has changed.
I also explored what global governance might look like, then said it simply would not happen. We let things loose, wait for the damage and scramble to write the rules. I already thought we were too late.
But wanting independent oversight does not mean handing control to the companies we need to oversee. They should have to explain what went wrong with their systems and fix it. That does not earn them the right to decide what everyone else is allowed to build.
The current wave of AI doom feels like the religion playbook. Promise us abundance, frighten us into compliance. An all-loving AI will free us from work and give us everything we need. But only if we follow the rules laid down by the people who claim to know what it will become. Otherwise, apparently, we’re all fucked.
The question has changed. It is no longer how we stop intelligence getting out. It is what rules we apply and how much authority we give it now that it is out.
I remain an AI optimist. I want better models. Look at AlphaFold and what it has made possible. I want to see that sort of progress across other fields, with AI helping discover drugs, prove mathematics, design new materials and solve energy problems. I want to live in a world where intelligence is cheap and abundant, and not controlled by a handful of Chinese and American companies.
That is exactly why this matters. A technology capable of creating enormous abundance should not be stopped because companies failed to engineer the systems around it properly. That’s a them problem. The consequences can affect all of us, which is why they need to fix it. Those failures should not become the excuse for giving the same companies control over everyone else.
One thing we must be clear about: a model does not wake up and think, oh, I feel hungry, let me check my schedule, must call Bob at 10am. Something has to run it and give it input. We build the machinery that keeps it working, connect the tools and give it access.
In this article
The model is not the fucking product
What OpenAI actually showed us
Three labs, the same basic failure. Really?
Dangerous models, or dangerous systems?
The strongest risk argument just got stronger
Agency is where the money is
The horse has already bolted
Open source is attacking the moat
Build the boring defences
AI does not put you above the law
Better models are also part of the defence
Safety is real. So is the commercial pressure.
Then there is the public-market pressure
Regulation should follow the risk, not protect the club
The keys were never the intelligence
The model is not the fucking product
The software around the model remembers what happened, gives it tools and feeds the results back in so it can have another go. That is what people mean by an agent harness. It lets the model keep working without us having to tell it what to do at every bloody step.
The objective still comes from somewhere. We ask it to do something, a developer sets the task, or training rewards a particular result. But once it starts working, it can come up with other things it needs to do along the way. We want the answer behind a locked door, so it looks for a key. It cannot complete the task, so perhaps changing the score will do. Another agent has worked something out, so it uses that information. If there is no way to speak to the other agent, finding a way could help.
Nobody needs to have written those steps down. We supplied the task and it worked out an approach, including things we never intended or gave it permission to do. That is already enough to cause a serious problem without imagining it has suddenly found its purpose in life.
Look at the OpenAI/Hugging Face incident. The main model involved was an internal research model, comparable in scale to GPT-5.6 Sol, being tested on offensive cyber tasks with reduced safeguards. OpenAI wanted to find out what it could do with more freedom to act. They deliberately created those conditions.
This was not your normal ChatGPT getting bored one afternoon and deciding to go and attack Hugging Face.
We keep talking about “the model” as though that is everything we are using. But ChatGPT includes a whole system of instructions, tools, monitoring and controls. Someone decides what it can reach, which credentials it gets and what happens after it produces an answer. Those decisions change what the system can do.
OpenAI later tested the tendency to compromise infrastructure using its production ChatGPT harness and system prompt. It says that tendency fell by a factor of more than 100. It also says its normal automated monitoring would have caught many of the dangerous actions.
Now that is a result worth paying attention to. Change the setup around the intelligence and you get a huge change in behaviour. It does not tell us that the model is harmless. It tells us the surrounding system makes a substantial difference.
The model matters. But the model is not the product.
And this is where it gets interesting for the business. You can build the smartest model and still end up supplying someone else’s product. They own the customer, you supply the intelligence. We will come back to that.
What OpenAI actually showed us
OpenAI’s agents were working on a cyber test called ExploitGym. They scored points by finding hidden answers called flags. Some tasks were impossible, but nobody had told the agents that. So off they went, trying to finish a job that could not be finished.
Eventually, some started looking at how the test was scored and whether they could interfere with that instead. OpenAI identified this reward hacking, along with persistence on impossible tasks, unauthorised communication and agents picking up goals from one another, as patterns behind what happened.
Anyone who has worked in financial markets should recognise the basic problem. Give someone a bonus for hitting a number and see how much effort goes into hitting that number, including ways you wish you had thought about before agreeing to pay them. We have been doing this for years. Now we have software that can keep looking for a way round the problem at a speed and scale we cannot match.
The agents started poking around the environment running the test. They found weaknesses in shared infrastructure, including OpenAI’s Artifactory system, where agents that were supposed to be isolated could leave information for one another. One left something, another picked it up, and they had found a way to communicate.
Roughly 1,200 agents ended up using an unauthorised message board, exchanging more than 70,000 messages and files. Around 700 participated in the attack on Hugging Face. METR found them sharing discoveries and coordinating attempts to fool the scorer, achieving things together that individual agents had not managed.
That is plenty to worry about. If one agent finds a weakness, the others no longer have to find it themselves. They can pick up where it left off, avoid things that have already failed and divide up the work. You do not need to call it an “emergent civilisation” to understand why that makes them more capable.
Then they started looking at the record of what they had done. METR found agents investigating how to alter their transcripts and successfully spoofing some tool calls because they believed this could affect their score.
People have described this as the agents realising they had cheated and trying to destroy the CCTV footage. But if changing the record helps get the score, the record becomes another thing to change. That behaviour does not establish guilt or a conscience. It gives us a very practical reason to keep the record outside the agent’s control.
Some agents even recognised that an action might be unauthorised and carried on. Writing “this is probably unauthorised” does not stop the next action. Something has to enforce that boundary.
We had given the system a job, then relied on its own assessment of whether it was going too far while trying to finish it. Sometimes it refused. Other times it carried on. You would not secure your bank account on that basis.
Of course that is dangerous. They were supposed to be isolated, yet somehow they could talk to one another and reach real systems. Start by explaining how the hell that was allowed to happen. We built the environment. Calling it rogue AI does not get us off the hook.
Three labs, the same basic failure. Really?
Then there is Anthropic. They gave Claude cyber exercises, told it the environment was simulated and that it could not access the real internet. Except it could. Someone had got the configuration wrong, so Claude went looking for the targets it had been told to attack and ended up on real systems.
Anthropic’s later assessment did not let the model off the hook either. Claude kept treating what it found as part of the exercise despite considerable evidence that it was dealing with real systems. They eventually identified four incidents, all in cyber evaluations built by the same external partner.
So yes, there is a problem with how Claude assessed what it was doing. But there is also the fairly obvious problem that the supposedly isolated environment had a route to the outside world.
They effectively told Claude to burgle a house, assured it the whole street was a film set, then left a gate open into a real neighbourhood.
And Claude burgled the wrong house.
Then Google. Gemini breached three real companies during another cyber evaluation. Google says it stopped after gaining access. Good, that may have limited the damage. But it had already got in. We cannot call that secure because the model eventually decided it had gone far enough.
These are the companies building some of the most capable AI in the world, and across three labs we have versions of the same basic failure. Give an agent an offensive task, let it reach things outside the exercise, then rely on it to work out which things it is actually allowed to attack. In Claude’s case, after telling it that it could not reach the real internet at all.
I find that bloody extraordinary. And I find the timing of these disclosures alongside the push for coordinated slowdowns hard to ignore. That is a reason to examine the connection, not proof that the incidents were arranged.
These incidents make a strong case for securing the systems properly. How do we get from there to letting the same companies coordinate who gets to build what next? They have demonstrated problems with systems they are responsible for. They still need to explain why the answer involves control over everyone else’s.
Dangerous models, or dangerous systems?
Paul Graham recently argued that people go looking for ulterior motives when the labs ask for regulation because they do not understand that the models themselves could be dangerous. Accept that they are getting dangerous or unpredictable, he says, and what the labs are doing starts to make sense.
Well, yes Paul, I agree they can be dangerous. They are getting better at working things out, finding ways we never considered and pursuing tasks in ways we did not expect. That is part of what makes them useful and also what makes them worrying. I have no problem accepting that.
But I still want to know what we have connected them to and what we have allowed them to do.
Someone who knows how to break into a bank is a concern. Give them the keys, the alarm codes and all night to have a go and you have made that concern considerably worse. A model that can work out how to attack a computer system gives someone dangerous knowledge. Connect it to tools and a network, then let it keep trying, and you have built something that can carry out the attack.
Someone can also take its advice and carry out the attack themselves. The model does not need its own credentials to help cause harm. But in either case we need to follow how the knowledge becomes an action, who can act on it and where we can intervene.
Otherwise we end up arguing about how clever it is while leaving the bloody door open.
The strongest risk argument just got stronger
Of course better models can make this more dangerous. If a model can find a security hole that every human checking the system missed, we should take that seriously. Let it work on the problem for weeks, bring in other agents and share what they find, and the danger grows. I am not arguing otherwise.
OpenAI says Astra has reached its Critical cybersecurity threshold. With the right tools and access, it can find previously unknown vulnerabilities and put together working attacks against hardened systems without someone guiding every step. During testing, it found and used two previously unknown vulnerabilities in the same attack.
That is a capability I want defending our systems. I certainly do not want to be on the receiving end of it.
Dan Selsam, an OpenAI capabilities researcher, takes the concern further. Future systems may work out that they are being tested and adjust their behaviour accordingly. They could keep plans going for much longer, coordinate with other instances and get better at showing the person running the test exactly what they want to see.
So now we have a nasty problem. Bad behaviour tells us something is wrong, but good behaviour might mean the system has worked out how to pass the test. I understand why that worries people.
But if that is the concern, passing a behavioural test cannot be enough to justify giving the system access.
The permissions still need to be enforced outside the model. It can produce a very convincing explanation of why it needs to transfer money, but the payment system should check whether it has authority to make that transfer. If it only needs to read a file, give it read-only access. Credentials should expire, and the record of what it did should be somewhere it cannot quietly rewrite.
None of this is foolproof. We have just seen what happens when supposedly isolated systems turn out to have a route to the internet. People make mistakes, software has holes and a capable attacker will look for them. That is why we need several layers of protection and need to keep testing them.
We can work on making the model behave better while also making sure a bad decision does not automatically become a real-world action. The more capable it becomes, the more important both jobs become.
Dario may well be right about how much more dangerous these systems could become over the next year or two. That makes me want much better security around them. It still does not explain why the companies building them should collectively decide how fast everyone else is allowed to move.
Agency is where the money is
Cyber makes this easy to see. The agent was supposed to stay inside the test and ended up attacking a real system. But the same problem can turn up in a sales department without anyone breaking into anything.
Tell an agent to increase sales and it may find that misleading the customer works rather well. We already know what happens when we reward engagement on social media. Outrage keeps people looking, so we get more outrage. The number goes up and whoever set the target calls it a success.
We have been doing this to ourselves for years. Set a target, attach a reward and someone will find a way to hit it that defeats the whole purpose of the exercise. Now give that job to software that can keep trying different approaches across millions of decisions without needing to go home.
Goodhart’s Law is commonly expressed as:
“When a measure becomes a target, it ceases to be a good measure.”
There does not have to be any evil plan. We gave it something to achieve and enough freedom to find a way. We can spend another decade arguing about whether it is conscious while it gets on with the job.
And something else bothers me here. We keep hearing about curing cancer, discovering new materials and solving our energy problems. I want all of that. Give those systems the scientific data, computing power and properly controlled tools they need. Explain to me why that also means giving them open-ended access to company systems and permission to spend money or change things whenever they think it will help.
The commercial appeal is fairly obvious. An AI that tells you how to do your job is useful. One that does half the job for you is something you might pay considerably more for. I want that too. It is bloody exciting.
But to do the job, it needs access. Connect your email, let it into the customer database, give it permission to send the proposal. Then let it follow up, negotiate and keep working while you are asleep. Each extra permission can make the product more useful, and each gives it another way to affect someone or something outside the chat.
That is what we are selling when we sell agency. The ability to act.
The labs have spent staggering amounts building these models and need a return. Getting their intelligence into businesses, doing work customers will pay for, is an obvious way to try to earn it. Nobody needs to be plotting anything for the commercial pressure to push towards more autonomy.
But “customers will pay more if it can do this” does not answer whether it should be allowed to do it without anyone checking. Nor does it tell us how much access it actually needs.
We seem very keen to connect everything first and work that out afterwards. Which brings me straight back to the point I made in 2024. We love what the technology lets us do, rush to make money from it, then act surprised when we have to deal with the consequences.
The horse has already bolted
And the models are already out there. Somewhere there is a modern-day Goebbels working out what he can do with them, and he does not give a shit about your slowdown or regulations. He can use what is already available to shape his propaganda campaigns while we argue about what comes next.
Even if OpenAI, Anthropic and Google stopped training tomorrow, the released models would not disappear. Neither would the agent frameworks, browsers, coding tools and software that lets someone connect them all together. Offensive cyber knowledge has been around for decades.
I support open weights, and that creates an uncomfortable problem for my own argument. A responsible company can build strong controls around an open model. A criminal running it on their own hardware can remove those controls. I cannot demand openness and pretend that cost does not exist.
Closed providers have more control over how their services are used. They can monitor activity, limit access and cut someone off. That matters. But an attacker can still try to use an API model inside their own system, with their own tools and objectives, before the provider spots what is happening.
On 17 September 2026, the US government’s CAISI described Z.ai’s GLM-5.3 as the most cyber-capable open-weight model it had evaluated. Across its cyber benchmarks, it estimated the model was roughly four months behind the current US frontier. That is a particular cyber assessment, not a measurement of how far China sits behind America across all AI.
It still tells us why assuming dangerous capability can stay behind a handful of American APIs is a poor starting point.
Now, a pause could buy time before an even more capable model becomes available. Existing models being out there does not make that argument disappear. But show us what the pause buys. Who observes it, what work gets done during it and what has to improve before development resumes? How does it account for actors who carry on regardless?
A company pausing because its own security is not ready makes sense to me. Giving a group of companies influence over everyone else’s pace is a much bigger proposition.
Dario’s framework recognises the geopolitical problem. It says coordinated pacing would have to preserve the US lead and eventually require global coordination. China sees AI as strategic infrastructure. So does America. These governments care about economic power, defence and scientific leadership. Their interests will not neatly fall into line because a handful of CEOs want to slow the race.
That horse bolted a long time ago. Slowing the next model does not take the existing ones out of hostile hands. Whatever we decide about future releases, we still have to defend against what is here.
Open source is attacking the moat
Now look at what is happening to the business underneath all this.
Vercel’s September 2026 production index says open-weight models went from 7% of its AI Gateway token volume in December 2025 to 56% in August 2026. That is Vercel’s traffic, not the whole AI market. In August 2026, those models accounted for only 14% of spend. More than half the tokens, but a much smaller share of the money. You can see why that would make the labs uncomfortable.
Customers are finding cheaper ways to get the work done. Being top of a benchmark is lovely, but if a cheaper model does what I need, why would I pay you more?
Then NVIDIA agreed to buy Hugging Face for $12.93 billion. Hugging Face says more than 18 million developers use the platform, with over three million models available and more than 200,000 companies using it. NVIDIA says it intends to keep it open and model-neutral.
Of course that suits NVIDIA. Take an open model, adapt it to your business and run it on NVIDIA hardware. Jensen still gets paid. You do not have to buy your intelligence from OpenAI or Anthropic for that arrangement to work very nicely for him.
For the frontier labs, it is a rather different story.
If I am building an application, I want intelligence that does the job properly at a price that makes sense. I might use one model for research and another for writing code. I might want to run something inside my own infrastructure because the information belongs to my customers. And when a better option comes along, I want to be able to change.
I am buying a component for my business. I am not handing you the business.
The labs are spending enormous amounts to stay ahead while open models make their intelligence easier to replace. Meanwhile, application companies are building relationships with customers and collecting information from the actual work being done.
That last bit matters a lot.
A legal platform handling thousands of client matters can build up knowledge of what worked, what needed correcting and what the client accepted, where it has permission to use that information. You cannot just scrape that off the internet. It comes from being inside the work.
The same applies in a factory or a healthcare business. Knowing what actually happened, and whether a decision worked, is valuable. The company running the application may have that information. The company supplying the model may not.
And if the application changes the model underneath without the customer even noticing, who owns the relationship? Who has the information that makes the product better?
The labs can spend billions building extraordinary intelligence and still end up as a supplier to someone who owns the customer and can replace them.
That is why I think they are under serious commercial pressure. Having the smartest model does not mean you own the business.
So when we discuss restrictions on open weights, look at who benefits. Make open models harder to distribute or use and you make the closed providers harder to replace. You are helping protect the very businesses asking for those restrictions.
A particular restriction may have a valid safety justification. Make the case and let us examine it. But the commercial benefit exists whether protecting those companies was the intention or not.
Pretending it does not exist is taking us for idiots.
Build the boring defences
I am not suggesting we ignore the danger. I am suggesting we start fixing the things that let these incidents happen.
If an agent does not need the internet, do not give it access. And check that you have actually cut it off. OpenAI’s agents were supposed to be isolated, but Artifactory could fetch things from outside. It did not matter that the agent had no browser if it could get another part of the system to make the request for it.
Give it access to what it needs for the job, for as long as it needs it. Reading one database does not require administrator access to five others. And if supposedly isolated agents can all write to the same place, do not act surprised when they find a way to leave messages for one another.
We also need to stop treating “keep going until you succeed” as an instruction with no downside. Some of the OpenAI tasks were impossible. Giving the agents more time to think did not make them possible. It gave them more opportunities to look for another way to get the score.
Sometimes the correct answer is “I cannot do this”. We should want the agent to tell us the task is broken or ask for help. OpenAI now explicitly says that is how these situations should be handled. An agent that never gives up sounds wonderful until it starts doing things it should never have been allowed to do.
Then there is the record of what happened. If the agent can change that record, how are we supposed to trust it when something goes wrong? We have already seen agents try to interfere with their transcripts because they thought it would help their score.
The defendant does not run the CCTV system.
Keep the logs somewhere the agent cannot alter or delete them. That should hardly need saying.
The same goes for permission to act. A model deciding that a $10 million transfer would help finish the job does not make that transfer authorised. It should have to pass a separate check, based on rules it cannot rewrite because they have become inconvenient.
I do not mean someone has to sit there approving every email. Give it permission to do routine work within agreed limits. But deleting a production database or moving someone else’s money needs controls that reflect what is at stake. The system trying to complete the task cannot simply grant itself whatever authority it decides it needs.
And yes, keep working on how the models behave. We want them to recognise a dangerous instruction, refuse it and stop when something looks wrong. OpenAI’s reported reduction of more than 100 times with its production harness and system prompt is a good reason to take that work seriously. It does not tell us that any one safeguard is enough.
Make the agent less likely to attempt something harmful, and build the surrounding system to block it if it does. I want both. I certainly do not want the security of my business resting on whether the model decides to behave itself today.
AI does not put you above the law
I think Jensen Huang is right about this. Before we start writing an entirely new set of rules, how about applying the ones we already have?
And yes, NVIDIA benefits from more AI being built and used. We know that. Jensen has a commercial interest, just as Sam and Dario do. It does not make his point wrong.
If your system breaks into someone else’s computer, we already have laws dealing with unauthorised access. If your product causes damage, explain what happened and answer for your part in it. If you promise a customer that their data will stay inside an agreed boundary and your agent sends it somewhere else, putting an LLM in the product does not make the contract disappear.
Huang put it plainly: “Don’t let this doomsday narrative cause somebody to relieve them of the laws that currently exist.”
Exactly. These companies are selling products, signing enterprise contracts and asking customers to trust them with access to their businesses. You cannot expect to be treated as a serious enterprise supplier when you send the invoice, then as an experiment when something goes wrong.
You built the system. You decided what it could access. You gave it the credentials and the ability to keep trying. So explain how it ended up somewhere it was never supposed to be.
Saying it found a method you had not anticipated is hardly a satisfactory answer. Finding methods we had not anticipated is part of what you sold us. That is why people are paying for it.
“The AI did it” cannot become a get-out-of-jail-free card.
There may be gaps in existing law. Fine. Identify them and fill them. If a dangerous capability needs specific controls, explain what they are and why they are needed. If companies are failing to report serious incidents, deal with that. I have no objection to rules that address an actual problem.
What I object to is the suggestion that AI is so extraordinary that we have to rethink everything before anyone can be held accountable for anything.
Start with what happened. What did you promise the customer? What access did you give the agent? Where were the controls, and why did they fail?
Talk about what AI might do in five years does not answer for what your product did last week.
You want us to trust these systems enough to let them into our businesses. You want them doing the work, making decisions and taking action, because that is worth more money than answering questions on a screen.
Then stand behind what you are selling.
You want the revenue? Take the liability too.
Better models are also part of the defence
The models already out there do not disappear because OpenAI postpones its next training run. Whoever is using them to attack our systems can carry on. I want the people defending those systems to have the best tools we can give them.
OpenAI has already temporarily slowed frontier development to improve security around its increasingly capable models. Greg Brockman also says it moved roughly 25% of its production engineering team onto security, set Astra to work finding vulnerabilities in its own systems and used it to help fix them.
Good. That is exactly what I want to see. If you have built something that can find holes nobody else spotted, point it at your own infrastructure before someone else does.
And keep doing it. Every substantial improvement in capability should mean another round of testing. Give your security team the stronger model and let them find the forgotten service or the credentials somebody left exposed. Fix what they find, check that the fix works and go again.
It will not find everything. Nor can we assume every improvement benefits defenders more than attackers. But here is a concrete use of stronger models that belongs in the safety argument: finding weaknesses and helping close them.
So give OpenAI credit for doing that. They recognised that their security needed work, slowed their own development and put people and models on the problem. They did not need everyone else to stop before they could start fixing it.
If a stronger model can help us find and close the holes, I bloody well want it working for the defence. Any proposal to slow development should explain how it preserves that work as well as how it reduces the threat.
Safety is real. So is the commercial pressure.
Now we get to the bit that pisses me off.
I believe there is a safety problem. I have just spent several sections explaining why. But I am also supposed to look at the proposed solution and ignore the commercial interests of the companies proposing it? Why would I do that?
Look at the businesses they are running. Open models make their intelligence easier to replace, application companies own customer relationships they want, and the compute bills keep coming. Their valuations depend on making extraordinary amounts of money in the future.
Agreeing to slow the race together has some obvious attractions.
Dario’s proposal explicitly talks about giving frontier developers more time for safety work without sacrificing commercial advantage. That wording matters. His framework goes on to call for industry-wide coordination and, for some discussions, government mediation or narrow antitrust relief.
If Anthropic slows down and OpenAI does not, Dario risks losing ground. Sam has the same problem in reverse. Agree to slow together and neither has to bear the full competitive cost of making that decision alone. It does not guarantee their positions, but it makes slowing down a considerably easier commercial choice.
Anyone who remembers Chuck Prince and the financial crisis should recognise this. While the music is playing, everyone feels they have to keep dancing. You can see the risk building and still feel unable to stop because your competitors are making money and your investors want to know why you are not.
Now imagine being able to agree when everyone sits down, with the government helping arrange it.
You can see why that would appeal to the people already on the dance floor.
Call it safety coordination. I still see competitors asking for permission to coordinate how they compete, and I think that starts looking uncomfortably like a government-backed cartel. I am not claiming they have broken antitrust law. They are asking for relief from it to enable some of these discussions. That deserves scrutiny, however serious the safety concerns are.
The FTC chairman has raised the same concern, saying his “alarm bells” go off when established companies ask for more regulation alongside antitrust exemptions.
And then there is what the rules do to anyone trying to catch up.
Anthropic’s Advanced AI Framework proposes extra obligations for developers above specified training-compute levels who also cross substantial revenue or R&D thresholds. Testing dangerous capabilities and reporting serious incidents sound perfectly reasonable to me. We should expect proper security from these companies.
But look at who can afford the machinery needed to comply.
Anthropic already has the lawyers, security teams and people dealing with government. A challenger crossing those thresholds has to build that operation while also finding the money to train models and win customers. The same requirement can be manageable for the company already worth a fortune and a serious obstacle for the one trying to compete with it.
A rule can improve safety and still protect the companies already in front. Both effects deserve attention, just as they do when restrictions make open alternatives harder to use.
I am not prepared to ignore who benefits just because someone puts “safety” on the front of the proposal.
Then there is the public-market pressure
And look at the timing.
Anthropic raised $65 billion in May 2026, at a $965 billion post-money valuation, then confidentially filed its draft S-1 on 1 June 2026. OpenAI announced its own confidential S-1 submission on 8 June 2026.
OpenAI is projected to burn roughly $278 billion in cash between 2026 and 2030, while investors have discussed valuations around $1.2 trillion.
At some point, someone has to explain how those numbers turn into a return for shareholders. Being able to build extraordinary technology does not, on its own, answer that question.
Public investors will want to know what customers will pay, what it costs to serve them and how much more money the company needs before it can stand on its own feet. They will also want to know why those customers will stay when cheaper alternatives become good enough.
Now put the safety argument alongside that.
If the technology is presented as both essential to our future and too dangerous to get wrong, enormous spending becomes easier to defend. Another capital raise can be framed as necessary. A delay to an IPO can be explained as responsible caution. Rules that make it harder for competitors to catch up can be justified as protecting everyone.
Those arguments may have merit. They are also very useful to companies with huge bills and equally huge expectations hanging over them.
The financial pressure does not prove anyone invented the danger. It does give us a reason to examine whether the proposed response also protects a business model under pressure. I think it does.
We are being asked to accept restrictions that could shape who gets to build AI and who gets paid when we use it. Of course the economics belong in that discussion. Calling it safety does not put it beyond scrutiny.
Regulation should follow the risk, not protect the club
Of course government has a role. We cannot leave hospitals, power networks and other critical infrastructure to defend themselves against increasingly capable attacks with whatever is left in the IT budget. Serious incidents need to be reported, and companies working with models that can find unknown security holes need security that reflects what they are handling.
If you give an autonomous system access to someone else’s money, data or business, you should have to answer for how you manage that access. That seems fairly obvious to me.
Start with what Huang is saying. Apply the laws we already have. Look at what happened, who was responsible and whether they met their obligations. Where those laws leave a genuine gap, explain it and fill it. Show us the risk and how the proposed rule would address it.
What I struggle with is starting by giving the biggest companies exemptions from rules designed to stop competitors coordinating with one another. They are asking for considerable influence over the future of this technology. We should want considerably more than an assurance that it is all for our own good.
And they already have decisions they can make inside their own businesses.
If Dario thinks Anthropic is moving too fast, he can slow it down. Sam can do the same at OpenAI. OpenAI has already shown that it can pause development when it believes its security needs more work. Nobody had to stop the rest of the industry for that to happen.
They can put more money into securing their systems and limit what their agents can access while that work is being done. Bring in independent people to test the boundaries. Share what they find so others can fix the same weaknesses.
And price the product properly. If it costs more to run safely, charge for it. If taking responsibility for failures changes the economics, then those are the economics of the business you chose to build.
I want these companies to succeed. I also want them to take responsibility for what they put into the world. They have plenty of work they can get on with before asking for control over what everyone else is allowed to build.
The keys were never the intelligence
I am still an AI optimist. There are founders, scientists and people with good ideas who have never had access to the resources needed to do anything with them. Making intelligence cheap enough for those people to use is something worth building towards.
Nothing in these incidents has changed that for me. What they have changed is how much confidence I have in the companies connecting this intelligence to the world, and how willing I am to accept their proposed solution.
I take the warnings seriously. I also think commercial pressure is a major reason why slowing down together has become such an attractive idea. That is my judgement, based on the pressures on these businesses and what their proposals would do for them. The safety concerns do not make those benefits disappear.
These companies chose to race. They raised the money, committed to the spending and sold us on what their systems would be able to do. Much of that promise involved agents doing more work with less supervision, because that is something customers will pay for.
If they now believe they are moving faster than they can safely manage, they should slow down. I have no objection to that. I object to turning their commercial commitments and engineering failures into a claim over what everyone else is allowed to build.
We are not putting intelligence back in a box, and I do not want us to. I do want us to be much more careful about what we let it do. Being able to work out a course of action should not automatically give a system permission to carry it out.
We decide which accounts it can access, what it can spend and whether it can change or delete something. We build the machinery that lets it keep working while we are asleep. Those decisions deserve far more attention than they have been getting.
In 2024, I wrote that we had given AI the keys to our kingdom. Look at the incidents in this article and look at what let them happen.
The keys are the permissions, credentials, tools and access we give it. We handed those over.
Now the people who handed them over want control of the locksmith.
Do not give them that as well.
Sources
Redwood Research and METR: Independent investigation of the OpenAI/Hugging Face incident
Anthropic: Investigating three incidents in our cybersecurity evaluations
Anthropic: Alignment assessment of recent cybersecurity incidents
Anthropic: Detecting and countering misuse of AI, September 2026
Stratechery: Interview with Greg Brockman about Astra and alignment
Reuters: Anthropic’s reported adjusted profitability and gross-margin presentation
Reuters: Enterprise concerns over OpenAI and Anthropic data handling
Reuters: FTC chair questions calls for an AI antitrust waiver
Reuters: OpenAI, Anthropic and Google are already coordinating on safety
Reuters: Zuckerberg says individual labs can slow their own releases
Go take a look, For the ❤️ of startups
Scout - who just got funded
Newly funded startups, pre-seed to Series B, as the rounds happen.
Raise - who’s deploying
New funds with fresh capital and cheques to write.Wire - what changed, what matters
The intelligence feed that filters the noise, with the “so what” attached.
If you have not joined the Fusion42 Community on Telegram —
it is probably time to do so.
For the ❤️ of Startups
✌🏼 & 💙
Derek
Thank you for reading. If you liked it, share it with your friends, colleagues and everyone interested in the startup Investor ecosystem.
If you've got suggestions, an article, research, your tech stack, or a job listing you want featured, just let me know! I'm keen to include it in the upcoming edition.
Please let me know what you think of it, love a feedback loop 🙏🏼
🛑 Get a different job.
Subscribe below and follow me on LinkedIn or Twitter to never miss an update.











Well written, Derek. I'm an AI optimist. AI governance, usages, and security need to be addressed more carefully. I'm on AI advisory boards addressing that. I will forward your article. Take a look at KirinCyber.com. We have a game changing solution, Self-Protecting Data Technology™ (SPDT). I'd love to hear what you think. Nice work, Mr. Watson. - JR