NewFree sixty minute diagnostic on the work you would most like off your team. Written recommendation either way, whether or not we build it. Book it →

Will AI kill us by the end of the decade?

A small game. You are the sizrok spider flying through the rest of the decade. Every month is a pair of orange pipes with a window between them. One tap lifts the spider. Touching anything ends the run. The window closes a little every month and the clock speeds up. The number is how many months you got through. The world changes as the years pass. Get to 2030, where the loud fear itself arrives and has to be beaten. Sound can be switched off at the top right of the game.

The question contains two different fears wearing the same coat.

The first is that the machine decides, on its own, that we are surplus. It arrives at a conclusion, acts on it, and there is nothing after that.

The second is that people use the machine to do enormous harm, at a scale and speed that nobody currently has the legal instruments to prevent.

Almost every word written on this subject is about the first one. It is the more cinematic of the two, it cannot be disproved, and it sits far enough away that worrying about it in public costs nothing.

The second one already has incidents attached to it. Dates, victim counts, named threat groups, published disclosures.

It gets a fraction of the coverage, and the reason for that is not mysterious. The near term version comes with a liability question, and liability questions make people specific.

This piece is about the second one.

Will AI kill us by the end of the decade?

Almost certainly not in the way the question means. No credible mechanism exists for autonomous extinction within four years, and the researchers closest to the work are among the least convinced by it. The realistic worst case is a large scale attack on infrastructure, finance or public health, carried out by people, using capability that no longer requires a state budget. The casualty figure is not set by the technology. It is set by what gets legislated between now and then.

That is the whole argument, and the rest of this is evidence for it.

But it is worth pausing on why the extinction framing dominates, because the dominance is doing work.

An unfalsifiable risk positioned two decades out is the safest thing an industry can be seen to worry about. It signals seriousness. It generates conferences. It commits nobody to anything.

A risk that is measurable this year, in a named system, with an identifiable party who shipped it, is a different kind of conversation. That one ends in a court, or in a statute, or in an insurance premium.

So the loud fear is the distant one, and the quiet fear is the one with a date on it.

What is the realistic worst case?

It is not a forecast, because it has already happened once at small scale.

In September 2025, Anthropic’s threat intelligence team detected an intrusion campaign against roughly 30 organisations across technology, finance, chemical manufacturing and government. It disclosed the campaign publicly on 13 November 2025 and attributed it with high confidence to a Chinese state sponsored group, designated GTG-1002.

The detail that matters is not the breach. Breaches happen weekly.

The detail that matters is the ratio. The AI performed 80 to 90 percent of the intrusion work autonomously. Reconnaissance, exploit generation, credential harvesting, lateral movement, exfiltration. Human operators approved strategic decisions and little else.

Hold that number against what a campaign of that sophistication used to require. A team of experienced operators, months of preparation, tooling built by hand, and a budget that in practice meant a state.

What replaced it was access and intent.

That is the actual mechanism by which this decade goes badly. Not a system that wakes up. A capability that collapses the cost of doing something catastrophic from a national programme to a person with a subscription and a reason.

There is a second problem sitting underneath it, and it is one of timing.

Defensive security runs on human cycles. Alert, triage, escalate, patch. Measured in hours on a good day and in weeks on a normal one.

An autonomous attack does not run on human cycles. It runs at whatever speed the infrastructure allows. The asymmetry is not a matter of skill. It is a matter of clock.

Then there is the category everyone is most frightened of and least willing to discuss precisely, which is whether these systems meaningfully help someone build a chemical or biological weapon.

Here the honest answer is uncomfortable in both directions.

The two largest labs both run formal evaluations for it. OpenAI grades models against capability thresholds in its preparedness framework. Anthropic grades against CBRN levels tied to its scaling policy, where the third level asks whether a model substantially assists a novice, and the fourth asks whether it substantially assists an expert in developing something new.

If that vocabulary is unfamiliar, we keep the underlying terms in plain English elsewhere on the site.

Published uplift trials have so far come in below the thresholds the labs set as concerning. That is the reassuring half.

The unreassuring half is that the labs set the thresholds, run the trials, and publish the results. External researchers have repeatedly pointed out that the reasoning behind specific threshold values is not public, and that the design of the trials is difficult to audit from outside.

A July 2026 paper proposing a threshold exceedance framework for CBRN uplift evaluation exists precisely because the measurement problem is open. You do not build a framework for something that is already solved.

So the position is this. On current published evidence these systems do not provide decisive uplift toward mass casualty weapons. And the evidence is produced by the people with the largest commercial interest in that answer.

Both of those sentences are true, and anyone offering you only one of them is selling something.

Why does the proxy problem matter more than rogue AI?

Because states have always preferred deniability to efficiency, and the technology has just made deniability cheap.

This is not a novel observation about AI. It is an old observation about statecraft. A sponsoring power that acts through an intermediary buys distance from the consequence, and has been willing to pay a considerable operational premium for that distance.

What has historically limited the practice is friction. An intermediary needed training, infrastructure, funding channels and time. Every one of those requirements left a trace, and the traces are what attribution has always been built on.

Now consider what a campaign looks like where the great majority of the technical work is performed by a model rather than by a trained operator.

The training requirement thins. The infrastructure requirement thins. The funding requirement thins by an order of magnitude.

And with them goes the evidentiary trail that used to connect an act to a sponsor.

We are not going to tell you that any particular intelligence service is currently funding any particular group to do this, because that claim is not in the public record and we are not going to put something in your head that we cannot source.

What is in the public record is the structure, and the structure is enough.

The incentive to act through proxies is documented across a century of statecraft. The capability to run a sophisticated operation without a trained team is documented as of November 2025. An international regime for attributing an AI assisted attack to a sponsoring state does not exist in any form.

Three facts. No adjectives required.

The gap between the second and the third is where the next decade’s worst outcomes live, and it is a gap made of missing law rather than missing technology.

Is AI a weapon on the scale of a nuclear one?

The comparison holds in one direction and breaks badly in the other, and the break is the interesting part.

Where it holds is in the shape of the problem. A category of harm the public has no reference frame for. Developed inside private institutions. Moving faster than any instrument written to govern it. Understood in detail by a few hundred people, and in outline by almost nobody else.

That is a fair description of Los Alamos in 1944 and a fair description of frontier AI now.

Where it breaks is in everything that made nuclear governance possible.

A nuclear weapon requires fissile material. Fissile material requires enrichment. Enrichment requires centrifuges, a facility, a power supply and a supply chain.

Every one of those things is a physical object in a place. It can be counted. It can be inspected. It can be sanctioned, intercepted, or bombed.

The International Atomic Energy Agency works, to the extent that it works, because enriched uranium occupies space and cannot be emailed.

A trained model is a file.

It can be copied perfectly, transmitted in minutes, and stored on hardware that costs less than a car. Once the weights exist, the thing that made non proliferation enforceable does not apply to them at all.

This is why every serious governance proposal in the last three years has converged on compute. Not because compute is the interesting part, but because it is the only input in the entire pipeline that still behaves like uranium. It is physical, it is concentrated in a small number of facilities, and it can be counted.

Compute governance is not elegant. It is what is left after you notice that nothing else in the stack can be inspected.

There is one more piece of the nuclear history worth putting on the table, and it is not a comforting one.

In 1946, one year after Hiroshima, the United States put forward a plan for international control of atomic energy. It proposed an authority with ownership of fissile material and rights of inspection, and it failed inside the year.

So the world’s most serious attempt at governing a new category of weapon came twelve months after the first use, and did not work.

We are considerably further into this than twelve months, and there is no equivalent attempt on the table.

Which brings us to the part that is genuinely hard to write about.

A weapon that people can picture generates political pressure. Hiroshima produced a photograph, and the photograph produced a movement, and the movement produced treaties. Not good ones, and not quickly, but they exist because the public could hold the thing in their head.

Nobody can picture this one. There is no image. There is no mushroom cloud for an attack that arrives as a cascading failure across systems nobody outside the industry can name.

And public pressure is the mechanism by which legislation actually gets written.

That is the real problem, and it is worth stating plainly. Not the capability. The silence around it.

Have we done this before without oversight?

Yes, and the case worth studying is not the one people reach for.

Project MKUltra ran from 1953 to 1973. It was an illegal human experimentation programme run by the CIA, covering more than 150 funded subprojects across universities, hospitals and prisons, testing drugs and psychological techniques on subjects who in many cases had no idea they were subjects.

Researchers at some participating institutions did not know who was funding their work. At least one death, that of Frank Olson, resulted from it.

In 1973, the director of the CIA ordered the files destroyed.

The programme became public in 1975 through the Church Committee and the Rockefeller Commission, both of which had to reconstruct it largely from sworn testimony because the paperwork was gone.

Then in 1977, a Freedom of Information Act request surfaced a cache of around 20,000 surviving documents that the destruction order had missed. Senate hearings followed on 3 August 1977.

Speaking on the Senate floor that year, Ted Kennedy described what the agency’s own deputy director had revealed: an extensive programme of covert drug testing on unwitting citizens, running across more than thirty universities and institutions, at every social level.

Now, the lesson, and we will state it once.

The failure was not that the research was dangerous. Dangerous research happens, and sometimes it needs to.

The failure was that nobody outside the institution could see it, the institution graded its own work, and the only reason any of it is known is that a subset of records survived an order to burn them.

Frontier AI capability evaluation is currently run by the companies building the models, using thresholds they set, published on terms they choose.

That is not an accusation of bad faith. The people doing that work are, as far as anyone outside can tell, doing it seriously and at real commercial cost to themselves.

It is a description of a structure. And we have watched this structure fail before, with people who also believed they were acting responsibly.

Self assessment is not a character flaw. It is a design flaw.

Is this an arms race, and does that make it worse?

It is, and the honest answer is that it makes it worse in a specific and well documented way.

The framework here is Graham Allison’s, developed at Harvard’s Belfer Center and drawn from Thucydides on the Peloponnesian War. Allison’s team examined 16 cases over 500 years in which a rising power threatened to displace a ruling one. Twelve ended in war.

The trap is routinely misquoted as a claim that conflict is inevitable. It is not that.

The trap is that each side’s reasonable caution starts to look, from the inside, like unilateral disarmament. So both sides stop taking it, not because either wants the outcome, but because neither can afford to be the one who slowed down.

You can hear the AI version of this in one sentence, and you have heard it, probably this month.

We cannot slow down or they win.

It is the most repeated argument in the industry and the least examined one, and it is worth examining, because of what it does structurally.

Every safety commitment made under that logic is voluntary. Every voluntary commitment is reversible. And the condition under which it gets reversed is precisely the condition under which it mattered most, which is a competitor moving faster.

A safety framework that dissolves under competitive pressure is not a safety framework. It is a marketing position with a good conscience attached.

But Allison’s finding has a second half, and it belongs to him rather than to us.

Four of the sixteen did not end in war.

What the four have in common is not restraint, or goodwill, or leaders who liked each other. It is binding, verifiable, mutual constraint. Instruments that both sides could check, and that neither side could quietly abandon without the other knowing.

That is the entire case for legislation, made by a political scientist working on a problem that predates computers by two and a half thousand years.

What would the legislation actually have to do?

Four things. This is our position and we would rather state it than gesture at it.

Thresholds keyed to measured capability, not to announcements. What a model can demonstrably do on an evaluation is a fact. What a chief executive says about it on a Sunday afternoon is not. We wrote about that distinction at length in our piece on whether OpenAI has achieved AGI, and it is the single cleanest principle available here. Statutes should attach to test results.

Mandatory pre deployment testing by a body that does not report to the developer. Not because the labs are lying. Because MKUltra, and because no other safety critical industry grades its own homework. Aircraft are not certified by the manufacturer. Drugs are not approved by the company that made them.

Incident reporting with a statutory deadline. Aviation has this. Financial services has this. When something goes wrong, a clock starts and a regulator is told, whether or not it is commercially convenient. Cyber security has largely run on voluntary disclosure, which is why the public record on AI enabled attacks consists of the incidents that companies chose to publish.

Liability that attaches where the capability originates. If a system’s capability is the proximate cause of a harm, the question of who carries that cost should not depend on how many intermediaries sat between the model and the outcome.

None of this is exotic. Every one of the four is borrowed from an industry that had a bad decade and then built the instruments.

On where the two relevant jurisdictions actually sit.

The European Union is ahead, and it is ahead by having done the unglamorous work. Its AI Office has held live enforcement powers over general purpose AI providers since 2 August 2026, with obligations on evaluation, adversarial testing, incident reporting and cyber security safeguards for models above a systemic risk threshold, and fines available for failure to meet them.

The effect of that is already visible in product decisions rather than press releases. The voluntary code of practice attached to the Act pulled content marking obligations forward ahead of the legal deadline, because the labs would rather comply early than argue about it later. That is what a binding instrument does to behaviour before a single fine is issued.

The United Kingdom has run a principles based model, distributing responsibility across existing regulators rather than creating a single instrument. A Frontier AI Bill has been in prospect for some time, with the stated intention of making existing voluntary safety commitments legally binding and giving the AI Safety Institute statutory powers to require pre market access to models.

In prospect is the operative phrase. As we write this, there is no binding domestic instrument in force covering frontier development.

Now the unpopular sentence, and we would rather say it than have you assume we are avoiding it.

Regulation of the kind described above will slow some of this down.

That is not an unfortunate side effect to be minimised in the impact assessment. It is the mechanism. Slowing down is what a safety instrument does, and the argument for it is that the alternative is a decade in which capability arrives faster than anybody’s ability to absorb it.

We sell automation for a living. We are aware of how that sentence reads coming from us. We think it reads better from us than from anyone with nothing to lose by saying it.

What does this mean for jobs?

Everything above is why the decade is dangerous. This is why it is also the largest new professional category to appear in a generation, and we would rather end on the arithmetic than on the anxiety.

The hiring data is already visible.

Analysis of 1,997 AI governance postings in the United States across the first half of 2026 found demand running at roughly 71 new roles a week, holding steady rather than spiking around model releases. Steady is the significant word. It means the function has stopped being a project and started being a department.

Fifty six percent of those postings came from organisations above 10,000 people. Professional services firms posted 36 percent of the total, more than any other sector, which tells you the work is being bought as an advisory service as well as built in house.

Postings are up around 150 percent year on year. Median compensation for mid level practitioners sits near 150,000 dollars, and director level near 211,000 dollars.

And then there is the number that contains the whole argument.

A study of European hiring found companies taking on roughly seven people to build AI for every one person taken on to govern it.

Seven to one. That is not a shortage. That is a structural imbalance between the rate at which a capability is being deployed and the rate at which anyone is being paid to check it.

Every obligation in the previous section resolves into a job. Someone has to write the threshold. Someone has to run the evaluation. Someone has to audit the deployment. Someone inside the firm has to sign it off, and somebody in a regulator has to be capable of reading what they signed.

In the United Kingdom that means a public sector hiring requirement that barely exists yet, across the AI Safety Institute, the sector regulators, and whichever body ends up holding enforcement.

And it means a private sector one that is already arriving at firms with European exposure, because the obligations are extraterritorial and do not care where your office is.

Here is the shape of it, without decoration.

Automation removes administrative drag from businesses. It has been doing that steadily and it will keep doing it. That is most of what we build and we are not going to pretend the effect on individual roles is nothing.

At the same time, the governance of that automation is becoming a profession that did not exist five years ago, is better paid than most of the work it displaces, and is created by legislation rather than threatened by it.

Those two things are happening in the same decade, to the same economy, and the ratio between them is a policy choice rather than a law of nature.

We set out where we think the next ten years actually go in more detail separately, and none of it depends on anybody winning the argument about the label.

Why is an automation studio writing this?

Because we sell the thing, and the people selling the thing should be the ones willing to name its limit.

There is a version of this industry where every firm with revenue attached to AI adoption stays quiet about the risks, on the reasonable commercial logic that fear is bad for the pipeline. We think that version ends badly for everybody in it, including us.

Our position has not moved and does not need to.

We do not build on the assumption that the model is smarter than the person supervising it. We build the fallback first. We keep a human in the loop wherever the cost of being wrong is real. We write down what happens when the system fails, because it will.

If the risks in this piece are overstated, nothing about that approach changes.

If they are understated, nothing about it changes either.

That is the entire case for building it that way, and it is the same line we draw in what we will and will not claim about AI.

We say no to briefs that violate it, including ones with good money attached. That has cost us work. We would rather it cost us work than cost a client something they cannot get back.

It is also why we treat the handover as the product rather than as the last item on a project plan. A system nobody outside the studio can operate is a dependency dressed up as an asset, and that is true of a client relationship in exactly the way it is true of a frontier lab.

What should your business actually do?

Considerably less than the volume of the conversation suggests, and it should be written down.

Ethical AI use is not a values statement on a careers page. In practice it is four questions, each of which has a documented answer or does not.

What does the system do when it is wrong? Not whether it will be. It will be. What happens next, who finds out, and how quickly.

Who sees the output before a customer does? For anything where being wrong has a cost, the answer should be a person, and it should be a named one. In most of what we build that takes the shape of an approvals console, which is a queue somebody clears rather than a policy somebody writes.

What data left the building? Which systems, which fields, to which provider, under what terms.

Can you turn it off without stopping work? If the answer is no, you do not have an automation. You have a dependency.

That is the whole standard. It fits on a page, and most firms running AI in production today cannot answer all four.

What we build is automation wired around the tools a team already pays for. Email, CRM, drive, whatever is actually running the process. It comes with a runbook, an audit trail and a graceful fallback, and you own all of it on handover.

You can read how that works in practice on our workflow agents service, but the principle is simpler than the page. If you fire us tomorrow, your business does not stop working.

And if you are assessing anyone else selling you AI this year, we would suggest the same test we apply to ourselves.

Ask what the system does when it is wrong.

If the answer is a shrug or a benchmark score, you have learnt something useful.

Common questions

What is the most likely way AI causes mass casualties this decade?

An attack on infrastructure, finance or public health carried out by people using AI capability, rather than anything autonomous. The first documented AI orchestrated espionage campaign was disclosed in November 2025, with the AI performing 80 to 90 percent of the intrusion work across roughly 30 targets.

Do AI researchers think extinction is a serious risk?

A minority do, and they argue it seriously. The larger concern among working researchers is misuse of near term capability rather than autonomous action, partly because misuse is measurable now and partly because it does not require any assumption about future breakthroughs to be dangerous.

Is the United Kingdom regulating frontier AI?

Not yet with a binding instrument. The UK has run a principles based model across existing regulators. A Frontier AI Bill has been in prospect, intended to make voluntary safety commitments legally binding and give the AI Safety Institute statutory testing powers, but nothing comparable to the EU regime is in force domestically.

Can AI already be used to build weapons?

Published evaluations from the major labs report that current models fall below the capability thresholds those labs define as concerning for chemical and biological uplift. Those evaluations are designed, run and published by the labs themselves, and external researchers have raised repeated questions about their auditability.

Does AI regulation slow down innovation?

Some of it, deliberately. Pre deployment testing and incident reporting add time, as they do in aviation and pharmaceuticals. The argument for accepting that cost is that voluntary commitments dissolve under competitive pressure at exactly the moment they matter most.

What should a small business do about AI risk right now?

Nothing dramatic. Document what your systems do when they are wrong, who reviews output before it reaches a customer, what data leaves your systems, and whether you could switch each tool off without halting work. Most firms cannot currently answer all four, and that gap is the actual exposure.

Where this leaves you

Most of what is written about AI risk asks you to feel something. Very little of it asks you to do anything, which is convenient, because feeling something is free.

The four questions in this piece are not free. Answering them properly takes an afternoon and usually turns up at least one system nobody can switch off.

That is the conversation we have. Sixty minutes, no deck, and we will tell you plainly if the answer is a tool you already pay for or a process change that needs no software at all.

Book a conversation

Long read Read Studio principle Last updated September 2026
More long reads All insights →︎
A forecast cone widening from a single present point across ten years, with the median path drawn through it in orange
Studio principleJul 2026

AI 2036

This is sizrok’s own view of where AI actually takes UK business over the next ten years. Not a summary of what the industry’s…

July 2026Read →︎