The 10^25 FLOP Threshold: How ‘Systemic Risk’ Gets Defined for AI Models

A single number — 10 to the 25th power floating-point operations — determines whether an AI model provider carries the EU AI Act’s heaviest obligations or its lightest ones. Almost nobody outside a handful of frontier labs has a working intuition for what that number actually means, which makes it easy to either dismiss as irrelevant or worry about unnecessarily, when the honest answer for most organisations is more specific than either reaction.

I’ve written about the GPAI Code of Practice’s enforcement timeline generally. This is the specific technical threshold underneath its third chapter — Safety and Security — which applies only to providers of “GPAI models with systemic risk,” and the FLOP count is precisely how the regulation draws that line.

What the number actually represents

A floating-point operation is a single arithmetic calculation — an addition or multiplication — of the kind a model performs billions of times during training. Ten to the 25th power is an almost incomprehensibly large count of such operations, roughly the computational scale associated with the most capable models from the handful of organisations — currently estimated at five to fifteen worldwide — with the resources to train at genuine frontier scale. It’s a proxy, not a direct measure of danger: the assumption underlying the threshold is that raw training compute correlates, imperfectly but usefully, with a model’s general capability and therefore its potential for systemic impact if misused or if it fails in a consequential way.

Free · 4 minutes

Do you know what could take the business down — and have you priced it?

Fourteen questions on concentration, third-party dependence, resilience, and incident readiness — the exposures a board is accountable for whether or not it can see them. Banded finding on screen, full sheet by email.

Why a compute-based threshold, specifically, and what it misses

Compute is chosen as the trigger because it’s genuinely measurable and hard to obscure, in a way that a more subjective “how capable is this model” assessment would not be — a regulator can, in principle, verify training compute more reliably than it can adjudicate a contested capability claim. The trade-off is real: a threshold based purely on compute can miss a smaller, more efficiently trained model that achieves genuinely dangerous capability without crossing the FLOP line, and can capture a model that used enormous compute somewhat inefficiently without necessarily being proportionately more capable or risky. The European Commission has retained authority to designate additional models as carrying systemic risk through other criteria beyond the FLOP threshold alone, precisely to cover this gap.

Why this matters even if no organisation reading this trains a frontier model

The threshold determines which providers carry the Code of Practice’s heaviest chapter — and, connecting to the model-provenance diligence I’ve written about generally, it’s a genuinely useful, checkable fact about any AI model an organisation deploys or depends on. A model built by a provider whose training compute crosses this threshold carries obligations — systemic risk assessment, incident reporting to the AI Office, model evaluation and red-teaming commitments — that a smaller model’s provider doesn’t, and that difference shapes how much assurance a downstream deployer can reasonably expect from the provider’s own governance.

This also connects directly to the fine-tuning provider-obligation threshold I’ve written about separately — an organisation doing substantial in-house fine-tuning of a model that itself crosses the systemic-risk threshold inherits a materially different regulatory position than one fine-tuning a smaller model, because the underlying compute scale of what’s being modified shapes the obligations attached to modifying it.

What checking this actually clarified

A payments firm evaluating two candidate foundation models for a fraud-detection tool discovered, on checking, that one provider’s model sat clearly above the systemic-risk threshold and carried the Code of Practice’s full Safety and Security commitments, while the other, smaller model’s provider had never publicly addressed whether it crossed the line at all. The firm’s own risk assessment had, until that point, treated both providers identically. Once the threshold question was actually asked and answered, the firm’s diligence expectations of the two providers diverged appropriately — a distinction the firm’s original evaluation had simply never surfaced.

Assessing where a specific AI model or provider sits relative to the systemic-risk threshold, and what that means for downstream deployment obligations, is exactly the kind of regulatory-technical analysis a technology control assessment is built to perform.

The question is worth adding to standard AI vendor evaluation going forward, not treated as a one-time check — new models cross the threshold as compute scales continue rising, and yesterday’s answer may not hold for tomorrow’s model version.

This is a genuinely low-friction question to add to an existing vendor review process — it requires asking, not auditing, and most credible providers can answer it directly when asked.

The threshold also matters for downstream modifiers — see when fine-tuning makes you a provider under Article 25.

Firms that build this into standing vendor diligence, rather than a one-time question, catch the shift automatically as their vendors’ own models scale over time.

Free interactive tool

Interactive deadline calculator

Check which regulations apply to you and when

Regulation across the EU, UK, US and Asia-Pacific has moved considerably in the past eighteen months, and several headline dates have shifted more than once. Twelve questions, about three minutes.

Results are shown on screen — no email required. A dated summary is available to download, and can be sent on if that's more useful. What we do with your answers.

Governance is what happens when nobody is watching.

Policies are easy. Consistent decision-making is harder. Understand where governance exists and where it has quietly become assumed.

Full Governance by Sixteen Pillars

Govern your business. Prove your compliance.

A board assurance cockpit for EU-regulated financial firms — tamper-evident, hash-chained proof of governance across DORA, GDPR, NIS2, ISO 27001, the EU AI Act and MiCA. In development.

See what's coming