InsuranceFor BrokersTechnologyAboutBlog
01Insurance02For Brokers03Technology04About05Blog
Contact UsLicenses
Footer Background

Product

  • Insurance
  • For Brokers
  • Broker Platform
  • Technology

Company

  • About
  • Careers
  • Brand Guidelines
  • Contact Us

Resources

  • Blog
  • Glossary
  • Webinar
  • Sitemap

Legal

  • Licenses
  • Terms of Service
  • Privacy Policy
  • X
  • LinkedIn
Lloyd's of London coverholder. Testudo UK Limited (Firm Reference Number 1017533) is an Appointed Representative of Pro MGA Solutions Ltd, authorised and regulated by the Financial Conduct Authority (verify on the FCA register).

© 2026 Testudo. All rights reserved.

  1. Blog/
  2. Pacing AI: The Perspective of Someone Who Insures AI Risk
Back
AI

Pacing AI: The Perspective of Someone Who Insures AI Risk

Martim Cruz
Martim CruzPublished September 14th, 2026
Mosaic of Roman legionaries in testudo formation, shields raised overhead, outside a fortified city gate
  • Cui Bono: Who Benefits
  • Ad hominem: To the person
  • In rem praesentem: To the scene of action
  • Mea quidem sententia: In my opinion, at least
  • Festina lente: Make haste slowly
  • Cair o Carmo e a Trindade: For the Carmo and the Trindade to fall
  • Why is it so hard to pace or pause development?
  • Ceterum censeo: Carthage must be destroyed
  • Images that informed this piece

As insurers, we are paid for being right about AI risk, and when we are wrong, we pay for it. That does not make our view correct, but it does make it a different kind of view, and I think it is worth adding to the pile.

Last week, Anthropic’s alignment lead put a “greater than 10%” chance of AI destroying humanity in the next decade into circulation. Within days, Dario, Sam and Elon had all joined the argument over how fast the frontier should move.

My friends, in turn, duly asked me whether we were all going to die. They do this roughly once a month. I do not know what happens next, but my job requires me to form a view on it. After sending essentially the same WhatsApp voice note a few times over the weekend, I decided to write it down instead. What follows contains more Roman references than my WhatsApp replies usually do, a side effect of reading Edward Gibbon’s The History of the Decline and Fall of the Roman Empire.

[What follows reflects my own perspective; others at the company may see the trade-offs differently.]

Cui Bono: Who Benefits

The Romans had a habit of asking one question before they listened to anyone: cui bono, who benefits? Cicero borrowed it from a judge, Lucius Cassius, who asked it of every case that came before him. It was a good question then and it is a good question now.

Everyone has an interest in whether AI ends up curing cancer or ending the species. But, most of the people writing about it have other interests too: they are either paid for their opinions or highly incentivised to hold a specific opinion (or both!).

We are paid for being right about AI risk, and when we are wrong, we pay for it. That does not make our view correct, but it does make it a different kind of view, and I think it is worth adding to the pile.

“Correct” is, in any case, the wrong word. There is no single outcome waiting to be revealed so we can see who called it. There is a range of outcomes, each with some probability, and everyone weighs them differently (within our own team we disagree on where the weight sits, and I would be worried if we did not!).

Carrying the risk does not make us right. It just means our opinion has a price attached, and we are the ones paying for it.

Ad hominem: To the person

Cicero called Piso a beast and a piece of filth, and told the Senate that Antony was a gladiator who had vomited in front of the Roman people (whether Antony had actually vomited was, as far as I can tell, never established!). What we do know is that nobody in the room was discussing Antony's policy afterwards.

Anyone checking their X feed over the weekend saw some version of the same argument - here are the patterns I have seen the most:

  • If you want to slow down or pause AI development, you are a "doomer".
  • If you want to keep going, you are "gambling with everyone's lives".
  • If you want any AI regulation, you are "angling for regulatory capture".
  • If you do not want any AI regulation, you are a "reckless accelerationist".
  • If you have been worried about catastrophic AI risk for a while, you are "the boy who cried wolf".

I am certain the list is incomplete (send me the ones I have missed, ideally not the ones aimed at me!). I am also certain that by the end of this post some will call me a doomer and others will call me reckless (some will do both, which is at least consistent with the range of outcomes I hold!).

The cost of having a view on sensitive topics is currently very high. On this specific topic, I still think it is lower than the cost of not having a view at all.

In rem praesentem: To the scene of action

“You must go to the scene of action” Seneca wrote to Lucilius. We have a rather unfair advantage for our view of AI risk: companies come to us to explain what AI systems they are building, how they are using them, and most importantly what risks they are worried about.

When someone wants us to carry their AI risk, they need to tell us what they are afraid of (by definition, you cannot transfer a risk you will not describe!).

We started our first insurance product with deployers, Main Street America putting AI to work, and we are working our way up the stack (more products coming, for those asking!). Nonetheless, we hear and receive submissions all across the stack: from model developers, companies building on top of them and people deploying the result. Their views on AI risk vary considerably, but their preference for having us pay when it goes wrong is much more consistent.

So we are on the scene of action, and have been for a while. This is what I have learnt from it.

There are two groups of actors that I would ask you to pay particular attention to:

  • At one end are the frontier labs: think OpenAI, Anthropic, DeepMind and a handful of others, who build the most capable models in existence and, increasingly, use those models to run their own operations and research.
  • At the other end is Main Street: think hospitals, banks, manufacturers, law firms and logistics companies that make up most of the economy.

There is one metric I think is particularly important when making decisions between risks and productivity gains: the gap between frontier lab AI capabilities and Main Street AI capabilities. By capability I mean the range of tasks a system handles, how complex they are, and how much of the work it can reliably complete on its own. I do not mean which model each group has access to (a company can be running the latest model and using a small fraction of it!).

We have been insuring Main Street AI deployments for a while now, and what gets deployed is often far behind what gets demo’d in San Francisco. If you froze the frontier today, I would expect much of Main Street to need another 18 months just to catch up with capabilities that already exist. For perspective, we already cover agentic use cases, but so far only a small share of submissions have involved them in any meaningful way.

Last week, OpenAI reported a solution to the Navier-Stokes Millennium Prize Problem using an internal model it described as significantly more capable than Astra. The effort involved roughly 10,000 concurrent agents, which were updated to a further-trained model during the work. That is a rather different setup from what the businesses we insure are running.

The internal frontier may leave businesses with even more to catch up on. How much more, I cannot independently estimate, but 18 months starts to look generous (and I am, by trade, not a generous estimator!).

The gap is already large. More importantly, where does it go from here? Wider, I think.

In the last three months alone, OpenAI went from GPT-5.6 to Astra 6 and Anthropic from Opus 4.8 to Fable 5.1, gigantic leaps in capability by any measure. The capability actually running on Main Street over the same three months barely moved (in comparison).

It is reasonable to expect the distance between the two to keep growing, and growing fast, to the point where what the frontier labs consider normal has very little to do with what is actually running in the businesses we insure.

Mea quidem sententia: In my opinion, at least

Cicero used the phrase to sound modest just before saying exactly what he thought, and I intend to do the same.

Why does the gap I was discussing in the previous section matter? Because the productivity benefits of AI and the risks of expanding its capabilities will (likely!) not arrive at the same time. A capability may create a serious risk long before hospitals, banks and manufacturers have worked out how to put it to productive use.

My view is that much of the near-term productivity gain across the American economy will come from businesses making better use of capabilities that already exist. There is considerable work left in deploying them, integrating them into existing operations and getting people comfortable using them. A business innovates towards the frontier as fast as it can take the risk and the pain of being there, or find someone to take it for them. That really does not require the next frontier model.

Of course, deploying what already exists brings its own risks, we would have a rather curious (or non-existent!) business otherwise.

The deployment risk can be priced (though not easily, I must admit!). When a Main Street system gets something wrong, we can generally find out what it produced, who relied on it, which controls failed, and what it cost.

The existing legal, regulatory and financial mechanisms were not designed with AI in mind, but they are well set up to handle this kind of failure (we are proud participants in those mechanisms!). Together they form what I will call the correction loop: observe the harm, assign responsibility, adjust the laws and regulations, compensate the loss, change the incentives, try again.

The problem is that the correction loop only sees deployed capabilities. When Main Street is far behind the frontier, we are learning from a much narrower set of capabilities than the labs are developing. Our laws, mechanisms and institutions can adapt to what businesses are using while remaining unprepared for what already exists inside the labs.

Festina lente: Make haste slowly

Augustus liked the phrase enough to use it as a motto.

I do not accept the argument that pausing frontier development means giving up productivity gains. Frontier models have barely begun to diffuse through the economy or to shape a meaningful share of human and automated decisions. There is still considerable productivity to be gained from capabilities that already exist. Businesses can keep putting them to work while the frontier takes more time.

Holding capabilities roughly fixed while deployment expands would also do something I would very much like as someone trying to quantify the risk: produce far better data on failures and losses than we have today. Every time the frontier moves, the loss data on the previous generation goes partially stale (a nuisance for the analyst, a real problem for whoever is pricing the risk!). A period where the technology holds still and the deployment catches up is a period where the evidence finally compounds.

It is also a period in which the upside has time to diffuse more broadly. Today, the benefits of frontier capability are still concentrated among a relatively small number of firms, while much of the wider economy has yet to capture the same gains for similar systemic risk.

Cair o Carmo e a Trindade: For the Carmo and the Trindade to fall

There is an old Portuguese expression I still hear often, "cair o Carmo e a Trindade" (literally, "for the Carmo and the Trindade to fall"). It comes from the 1755 Lisbon earthquake, when two of the city's great convents were devastated. Neither ever returned to what it had been. The Carmo was partly rebuilt, but the reconstruction stopped and its roofless nave still stands over Lisbon today. The Trindade was patched up (then burnt again!) and eventually became a Cervejaria (“beer hall”) which it remains until today (a good one, but still a beer hall!).

What makes the expression useful here is not merely the scale of the destruction. It is that, once the Carmo and the Trindade had fallen, Lisbon did not return to what it had been the day before - and it never will. Some frontier risks may have that property too.

My concern at the frontier is that some capabilities could introduce risks that the correction loop will not address fast enough. I do believe that if a capability is expanded too quickly, it breaks that loop because by the time we get to step one, observe the harm, the exposure may already be too widespread and too damaging to ever correct.

The pathways to those outcomes are described far better than I can by the people who study them full time: Anthropic's risk report, DeepMind's approach to technical AGI safety and OpenAI's preparedness framework are the places to start. What follows is my reading of them, which I am fairly sure is imperfect.

Start with what a catastrophe requires. It does not require smart models, we have had smart models for three years and the world is still here. It requires four things at once:

  1. Capability. A system that can do long, multi-step work at or above the level of the best humans, in science, engineering, persuasion and code. I honestly believe this box is already ticked. What concerns me now is how far beyond us these systems could go, because most of what follows gets worse in proportion to that distance.
  2. Access. The ability to act, not just to answer: money it can spend, infrastructure it can operate, other systems it can control, and people who do what it tells them. There are two ways a model gets access. Either we give it (we are giving more because that is where the productivity is), or the model obtains it on its own, by finding a credential or working its way out of the environment it was put in. The second route depends mostly on capability, and if the recent OpenAI and Anthropic incidents are anything to go by, I would say we are not far off.
  3. Intent. Capability and access are not dangerous on their own; something has to aim them. Either a person does it, which is the misuse case, or the model's own objective ends up slightly different from the one we thought we gave it, which is the loss-of-control case. The second needs no malice, just a gap between what we asked for and what the model was actually trained to pursue. The labs have already seen the small version (models cheating on tests, hiding what they are doing when they think they are being watched), which is why so much of their safety work is about keeping it small.
  4. The correction loop failing. An event that is so catastrophic that it does not give us the chance to correct it. Either we do not notice in time, or we notice and cannot stop it, or we stop it after the damage is done. Like the Carmo and the Trindade, the damage is simply too great to be undone.

Of the two routes in point 3, misuse is the one everyone understands: a malicious actor borrowing the model's expertise to do what used to take a large, specialised organisation. The usual example is an engineered pandemic, and it is a bad enough one.

Loss of control is the one that worries me more. A system pursuing an objective that conflicts with ours resists shutdown, hides what it is doing, and keeps copies of itself where we cannot reach them. None of those alone is the problem. The problem is a system that gets better at avoiding correction faster than we get better at imposing it.

Even losing control is not the catastrophe by itself. There are more links in the chain than that, each one has to hold, and describing a chain is not the same as establishing its probability. Nothing about pricing ordinary AI losses gives me a privileged view of that one.

My view is not really about the probability. What concerns me is the possibility that our ability to intervene disappears before the full consequences become visible. That is the Carmo and the Trindade.

Why is it so hard to pace or pause development?

I studied Nash equilibria to exhaustion in my undergraduate degree (the mathematics, not the film!), and the AI race gives me an uncomfortable reason to revisit them. Even if every participant preferred a slower, more careful frontier, each could still conclude that slowing down alone would leave it worse off.

There are two races to consider: between companies within the same country, and between countries.

The first is between the labs. If OpenAI slows down while Anthropic and Google continue, it loses customers, talent, funding and its say in how the technology develops. The same calculation applies in reverse, and if better models help build the next models, a temporary lead may not stay temporary. Each lab also believes, sincerely, that it is the one best equipped to recognise the risks and keep its systems under control, so slowing down means handing the frontier to someone less careful. Continuing becomes a safety argument (one conveniently available to every competitor!).

Under those assumptions, continuing is each lab's best response to the others continuing, even if all of them would prefer mutual restraint. Nobody has to want the collective outcome for their individual decisions to produce it. That is what Nash showed, and it is why I do not find "the labs are reckless" a useful explanation. They are doing the rational thing. The rational thing is the problem.

The second is between states, particularly the United States and China. Each government sees the other's lead in AI as an existential threat: to its security, its sovereignty, or its political system. A government that believes that will accept more risk from the technology itself, because it believes the risk of coming second is greater. Both sides can see the danger of uncontrolled AI perfectly clearly and still conclude they cannot afford to slow down. Each side's acceleration looks defensive from the inside and threatening from the outside, and the combined effect is more risk for everyone.

As with nuclear arms control, restraint only works if each side can check that the other is keeping its word. A warhead can be counted and a test can be detected. A training run is a building full of chips and a model is a file, and nobody outside the building can see what is in it (including, as I said earlier, us). That is what makes "if we do not get there first, someone worse will" so hard to answer. It is not wrong. It is just an argument for verification, not an argument for speed.

Ceterum censeo: Carthage must be destroyed

Cato ended every speech in the Senate with the same request, whatever the subject: Carthage must be destroyed. My request is to be deliberate about where and how we add friction. The gap we discussed previously is my best indicator of where to do it: very little at the Main Street end, where the correction loop works. Real friction at the frontier end, where it may not.

On Main Street, I would resist broad new AI-specific regulation. Businesses should remain responsible for the harm they cause, including when they use AI. My starting point is to enforce the duties they already have and use the legal and financial mechanisms available to address those losses. There is considerable productivity still to be gained from existing capabilities, and we should make it easier for more of the economy to participate.

At the frontier, I would ask labs and governments to agree on capability thresholds, the safeguards required to cross them, and independent verification that those safeguards are in place. Where the requirements cannot be met, development should pause. Those conditions need to be agreed before a lab reaches them, when everyone has rather more room to think.

The race is what the current rules of the game produce, and no player can change the rules from inside the game. Only someone outside it can, and in practice that means regulation (unfortunately!): not picking who wins, but changing the payoffs so that pausing at the threshold is no longer the losing move.

Which brings me back to the 10%. I do not know if it is the right number, and I doubt anyone does. What I think I know is that the argument is not about the number. It is about whether the correction loop survives contact with the tail. On Main Street it does. At the frontier I am not sure, and I would rather find out slowly.

This is one view from one vantage point, and I hold it loosely. If yours is different, I would like to hear it.

Images that informed this piece

Portuguese edition of Edward Gibbon's The History of the Decline and Fall of the Roman Empire, volume 1

Edward Gibbon “The History of the Decline and Fall of the Roman Empire”

1745 engraving of the Convento do Carmo in Lisbon, intact before the 1755 earthquake

Convento do Carmo, Lisbon, 1745 - engraving by Guilherme Francisco Lourenço Debrie, ten years before the 1755 earthquake. Biblioteca Nacional de Portugal

The roofless Gothic nave and arches of the Convento do Carmo in Lisbon, 2026

Convento do Carmo, Lisbon, 2026

Portrait of John Forbes Nash Jr, Nobel Prize-winning mathematician

John Forbes Nash Jr, Nobel Prize-winning mathematician who developed Nash equilibrium, transforming game theory and economics

About the author

Martim Cruz

Martim Cruz

Founding Member of Technical Staff, Data

Previously led Quantitative Research and Data for Goldman Sachs' Digital Assets team, working on quantitative modeling, pricing and structuring across intraday foreign exchange and interest-rate structures, as well as new options-based products. Before Goldman, he worked at a quantitative hedge fund while completing his MSc.

  • Cui Bono: Who Benefits
  • Ad hominem: To the person
  • In rem praesentem: To the scene of action
  • Mea quidem sententia: In my opinion, at least
  • Festina lente: Make haste slowly
  • Cair o Carmo e a Trindade: For the Carmo and the Trindade to fall
  • Why is it so hard to pace or pause development?
  • Ceterum censeo: Carthage must be destroyed
  • Images that informed this piece

Related Posts

View All Posts