EU AI Act's 10^25 FLOPs Threshold: A Numbers Game?
EU AI Act's 10^25 FLOPs Threshold: A Numbers Game?
| Takeaway | Detail |
|---|---|
| The 10^25 FLOPs threshold is a cognitive trap that incentivizes under-reporting. | The $214 million gross of 'Disclosure' against its $50 million budget shows how a single event can have outsized consequences, but the fine makes the gamble negative expected value. |
| The fine is the key deterrent, but only if labs correctly estimate the probability of detection. | The $50 million budget of 'Disclosure' pales next to its $214 million gross, a ratio that mirrors the leverage of a single under-reported FLOP count. |
| The EU's threshold is a static number, unlike the evolving definition of 'disclosure'. | The Merriam-Webster definition was updated 6 days ago, showing how language adapts, but the 10^25 FLOPs threshold remains fixed. |
| Under-reporting is a negative expected value bet because the fine is a significant portion of global revenue. | The $214 million gross of 'Disclosure' against its $50 million budget shows that a single event can have outsized impact, but the fine structure makes the gamble unprofitable. |
In 2026, a single floating-point operation could cost your company a significant portion of its global revenue. That's the penalty for crossing the EU AI Act's 10^25 FLOPs threshold without proper disclosure—a number so large it feels abstract, yet so consequential that it turns every training run into a gamble.
The threshold is a cognitive trap. Labs see a binary line: report above it or stay below. But the fine, applied to global revenue, makes under-reporting a negative expected value bet. Consider the 1994 film 'Disclosure,' which grossed $214 million against a $50 million budget—a 4.3x return that pales next to the leverage of a single misreported FLOP count.
Even the definition of 'disclosure' is evolving—Merriam-Webster updated it just 6 days ago. Yet the EU's threshold remains static, inviting gamesmanship. The numbers don't lie: the risk of a fine outweighs any benefit of hiding a training run. The real question isn't whether you'll cross 10^25, but whether you'll survive the disclosure.
The 10^25 FLOPs Cliff
Article 51 of the EU AI Act does not measure risk; it measures arithmetic. The hard presumption of systemic risk activates at 10^25 FLOPs, a figure that carries the false precision of a laboratory instrument but functions more like a casino's table limit. For a lab whose cumulative training compute sits at 9.9 x 10^24 FLOPs, the regulatory gaze is absent; at 1.01 x 10^25 FLOPs, the full machinery of the EU AI Office descends. The discontinuity is stark, but the cognitive trap is believing that the line itself carries informational content about danger. It does not. It carries only legal content about jurisdiction.
The obligations that trigger at the threshold are procedural, not harm-based. Mandatory red-teaming, adversarial testing, and incident reporting to the EU AI Office apply regardless of whether the model exhibits any observable dangerous capability. This is a critical distinction for decision-makers: the Act punishes the absence of paperwork, not the presence of risk. A model trained with 1.1 x 10^25 FLOPs that demonstrates benign behavior is treated identically to one that exhibits emergent deception. The compliance burden is fixed, while the risk profile is variable—an inversion that should make any rational actor question the metric's validity but not its consequences.
The consequence structure, per the Act's penalty provisions, is where the decision theory becomes unambiguous. The penalty for violating obligations—or for engaging in prohibited practices—reaches a significant percentage of total worldwide annual turnover. Consider a lab with $50 billion in global revenue: the fine ceiling is a substantial portion of that revenue. This is not a marginal cost; it is an existential event. The asymmetry is brutal. The cost of compliance—red-teaming, testing, reporting—is a rounding error on a training run budget. The cost of non-compliance, if detected, is a balance-sheet catastrophe. When the downside is a significant portion of worldwide turnover and the upside of gambling is merely avoiding a compliance process, the expected value calculation collapses into a single rational choice.
The Act's cumulative compute clause closes the most obvious arbitrage. If a lab trains multiple models in a coordinated way, the FLOPs are summed. Splitting a 2 x 10^25 FLOPs training run into two 1 x 10^25 FLOPs runs does not evade the threshold; it triggers it retroactively. This clause eliminates the "divide and conquer" strategy that a naive reading might suggest. The coordination test is fact-specific, but the risk of misjudging it is asymmetric: a wrong guess about what constitutes "coordinated" training carries the same fine as a deliberate violation. The safe harbor is not in clever structuring; it is in disclosure.
The final layer of uncertainty is the EU AI Office's discretionary designation power. Even below 10^25 FLOPs, the Office can designate a model as systemic risk based on capabilities or impact. This is the wildcard that makes the threshold a floor, not a ceiling. A model trained at 8 x 10^24 FLOPs with unexpected tool-use or long-horizon planning capabilities can be pulled into the systemic-risk regime without hitting the arithmetic trigger. The implication for a lab approaching 10^24 FLOPs is that the bright line is not a shield; it is merely the first line of a multi-layered assessment. The only way to eliminate the tail risk of a fine is to self-report early, establishing a cooperative posture before the Office's discretion is exercised adversarially.
| Scenario | Trigger | Obligation | Penalty Risk | Rational Choice |
|---|---|---|---|---|
| Below 10^24 FLOPs | None | None | Low (discretionary designation possible) | Monitor, prepare |
| Approaching 10^24 FLOPs | None yet | None | Medium (capability-based designation) | Voluntary self-report |
| Exceeds 10^25 FLOPs | Article 51 presumption | Red-teaming, adversarial testing, incident reporting | High (substantial fine per the Act's penalty provisions) | Mandatory compliance |
| Coordinated multi-model training | Cumulative compute clause | Summed FLOPs count toward threshold | High (retroactive trigger) | Disclose all runs |
| Below threshold, high capability | Office discretion | Same as systemic-risk GPAI | High (if designation occurs) | Proactive engagement |
The decision rule for any lab whose cumulative compute approaches 10^24 FLOPs is not a matter of regulatory philosophy; it is a matter of actuarial survival. The turnover fine is a catastrophic tail risk that dwarfs the cost of compliance. The uncertainty of detection—whether through the cumulative clause, discretionary designation, or a future audit—makes gambling irrational. Voluntary self-reporting to the EU AI Office is the only move that eliminates the tail risk while preserving optionality. The threshold is arbitrary, but the penalty is not. Act accordingly.

The Math of Risk
The threshold itself is a moving target. The Stanford AI Index 2024 reports that training compute for frontier models has doubled every 6-10 months. That means the 10^25 FLOPs line—already crossed by GPT-4 and Llama 3.1—will be crossed by many more labs by 2026. The labs that trained at 10^24 FLOPs in early 2025 are not safe; they are simply one or two doublings away from the cliff. The rational response is not to wait until you cross it, but to self-report when your cumulative training compute approaches 10^24 FLOPs. The EU AI Office's presumption of systemic risk is a bright line, but the decision to report is a judgment under uncertainty—and the uncertainty cuts in favor of disclosure.
When a lab's cumulative training compute crosses 10^24 FLOPs, the decision isn't about whether the EU AI Act's 10^25 threshold is scientifically sound—it's about which failure mode you can survive. The choice between self-reporting and hiding is a textbook expected-value problem, and the numbers resolve it with unusual clarity.
| Scenario | Cost of Compliance | Cost of Non-Compliance (if detected) | Decision |
|---|---|---|---|
| GPT-4 (2.1e25 FLOPs) | a modest cost | a significant portion of global turnover | Self-report |
| Llama 3.1 (3.8e25 FLOPs) | a modest cost | a significant portion of global turnover | Self-report |
| Lab at 10^24 FLOPs (approaching) | a modest cost | a significant portion of global turnover | Self-report early |
There is also the regulatory goodwill component, which is harder to quantify but operates in Option A's favor. A lab that self-reports before being audited signals cooperative intent, which typically influences the proportionality of any subsequent enforcement actions. A lab that is caught hiding compute figures forfeits that goodwill entirely and invites maximum penalties. The EU AI Office's discretion in applying the fine is not unlimited, but it is real, and it will be exercised more favorably toward the lab that came forward voluntarily.

The Expected-Value Table
When the EU AI Act’s drafters settled on 10^25 FLOPs as the systemic-risk trigger, they chose a number that behaves like a precision instrument while functioning as a blunt statistical cudgel. The data supporting this threshold—largely drawn from Epoch AI’s retrospective estimates of GPT-4 and Llama 3.1—suffers from a measurement problem that anyone who has worked with floating-point operations understands intuitively: nobody actually counts FLOPs at scale. The estimates are reconstructed from hardware specifications, training durations, and utilization assumptions that can vary by an order of magnitude depending on whose methodology you trust. A lab reporting 9.8e24 FLOPs and a lab reporting 1.1e25 FLOPs may have performed nearly identical work; the difference is an artifact of measurement, not a meaningful gap in capability.
The variance across cases is not a minor footnote—it is the central problem with treating the threshold as a legal fact. Consider the difference between a lab that trains one massive model on a dense transformer architecture versus a lab that trains the same parameter count using mixture-of-experts routing. The FLOPs count may be comparable, but the resulting systems have wildly different risk profiles, deployment footprints, and potential for harm. The Act’s threshold cannot distinguish between them. Worse, the cumulative compute provision means that a lab running hundreds of small fine-tuning runs—each individually trivial—can cross the threshold through accretion, even though no single model in their portfolio resembles a frontier system. The rule was designed to catch a specific kind of actor, but it sweeps in a heterogeneous population that shares only one trait: a large electricity bill.
When does the canonical decision rule break? The honest answer is that it breaks precisely when the assumptions underlying the expected-value calculation stop holding. The rule assumes that the turnover fine is a credible threat, that detection is uncertain enough to make gambling irrational, and that compliance costs are manageable relative to the tail risk. Each of these assumptions has a failure mode. If the EU AI Office announces a leniency program that effectively immunizes non-reporters who later comply—a policy shift that has precedent in GDPR enforcement—the fine becomes a negotiation starting point rather than a catastrophic outcome. If a lab is headquartered outside EU jurisdiction with no assets in member states, the fine is uncollectible in practice, and the entire calculus inverts. And if a lab’s compliance costs—legal review, technical documentation, ongoing audit obligations—approach or exceed the expected fine multiplied by the probability of detection, the rational choice flips. The rule is not a law of nature; it is a conditional recommendation that holds only within a specific institutional envelope.
There is also a subtler failure mode rooted in the psychology of the decision-maker. The thesis assumes that lab leaders are expected-value maximizers, but the judgment and decision science literature—particularly the work on ambiguity aversion and probability weighting—suggests that humans systematically overweight small probabilities of large losses. That bias actually reinforces the canonical rule: the fine looms larger in the decision-maker’s mind than its objective probability warrants, making self-reporting even more attractive. But the same bias can produce the opposite error when the probability of detection is perceived as near-zero. If a lab believes—correctly or not—that the EU AI Office lacks the technical capacity to audit cumulative compute claims, the perceived probability of detection collapses, and the overweighted small probability of a fine becomes an underweighted near-zero probability. The rule holds only when the decision-maker’s subjective probability of detection remains above a threshold that the Act itself does not specify.
| Decision Option | Upfront Cost | Fine Risk | Expected Outcome | Winner |
|---|---|---|---|---|
| A: Self-report as systemic risk | a modest fixed cost | Eliminated | Bounded, predictable compliance spend | Dominates on EV and variance |
| B: Under-report / hide FLOPs | no upfront cost | detection probability × a significant portion of turnover | a substantial expected fine for a company with significant turnover | Loses on every axis |
The practical takeaway is that the canonical rule is robust for the modal case—a Western lab with EU market exposure, concentrated compute, and competent legal counsel—but it degrades gracefully into uncertainty at the edges. The data does not tell you which edge you are on. What it does tell you is that the 10^25 threshold is a political artifact dressed in mathematical clothing, and the only defensible response to an arbitrary bright line is to treat it as a risk-management trigger rather than a scientific finding. If your cumulative compute is anywhere within an order of magnitude of the threshold, the asymmetry between the cost of compliance and the cost of a turnover fine is so lopsided that the decision makes itself—unless you are the rare lab that can credibly evade detection, in which case you are not making a decision under uncertainty at all, but a calculated bet against a regulator you believe you can outrun. That belief, more than any FLOPs count, is the variable that actually determines the outcome.
Epoch AI's methodology carries a significant confidence interval on FLOPs estimates, and that single fact changes the legal geometry of the EU AI Act. A model a lab estimates at 9e24 FLOPs — below the 10^25 systemic-risk trigger — could in reality be above it, because the estimate is a distribution, not a point. This is the illusion of precision: the statute reads like a hard arithmetic line, but the measurement beneath it is a fog bank. The lab generating the estimate is also the lab deciding whether to self-report, which gives it both the motive and the opportunity to believe the favorable end of the error bar.

What the Data Doesn't Tell You
Sparse architectures make the metric even less trustworthy as a signal of danger. Mixture-of-Experts models activate only a fraction of their parameters per token, reaching high capability with fewer effective FLOPs than a dense transformer of comparable quality. The threshold was calibrated on dense training runs, so it structurally misprices sparse models. A lab could train a genuinely dangerous MoE model, land under the threshold, and be formally compliant. That is a real hole in the Act — but exploiting it means betting the whole company on the regulator accepting a mathematical technicality.
The fine is a maximum ceiling, not a standard penalty. EU enforcement follows proportionality, so a first offense or good-faith misestimate could settle at a lower percentage of global turnover. And as of 2026, the EU AI Office's enforcement capacity is thin: a handful of investigators monitoring the entire frontier-lab landscape means the detection probability for under-reporting could be low. A naive expected-value reader sees a low chance of a reduced fine — an expected cost that looks cheaper than compliance. That is the most dangerous number in the entire decision.
But each term in that calculation is a guess, and the guesses compound. The fine is not a fixed percentage; it is whatever the regulator chooses under proportionality, and the same discretion that can reduce it can hold it at the ceiling. The detection probability is not a known number; it is an unknowable function of political attention, whistleblower incentives, and the massive paper trail that training a frontier model generates. The rebuttable presumption of systemic risk — a company can argue its high-FLOP model is not high-impact given its deployment context or architecture — is litigation risk, not a safe harbor. Every factor that seems to lower the cost of non-compliance is an assumption stacked on another assumption.
This is why the canonical decision rule does not change. When the variance of the outcome explodes, the catastrophic tail — the ceiling fine on global turnover — is the only term that matters. You are not choosing between a small certain cost and a small likely cost. You are choosing between a small certain cost and a small probability of ruin, where the probability is unknown and the tail lands exactly once. Self-reporting is the only move that makes the tail impossible.
| Condition | Rule Holds | Rule Breaks |
|---|---|---|
| Detection probability perceived as moderate-to-high | Self-reporting dominates | Gambling becomes rational only if detection is near-zero |
| Compliance costs below expected fine | Self-reporting is cost-effective | If audit costs exceed the fine, the calculus inverts |
| EU enforcement reach includes lab's jurisdiction | Fine is a credible threat | Non-EU labs with no local assets face no credible enforcement |
| Lab's compute is concentrated in one model | Threshold is meaningful | Distributed compute across many small runs dilutes the signal |
| Decision-maker is loss-averse | Reinforces self-reporting | Overconfidence in evasion can override loss aversion |
When the EU AI Office publishes its first designation decisions, the labs that gambled on the threshold's ambiguity will learn the cost of treating a legal cliff as a statistical suggestion. The decision framework below is not about whether 10^25 FLOPs is the right number—the earlier sections have established that it is arbitrary. This is about what you do on the ground, today, with the compute you have already spent and the runs you are planning for 2026.

The Blind Spots
Rule 1: If your cumulative training compute exceeds 10^25 FLOPs, self-report immediately—do not wait for EU AI Office designation. The Act's Article 51 creates a presumption of systemic risk at this threshold, but the presumption is not self-executing. The EU AI Office must formally designate you, and that designation process is where the turnover fine becomes a live threat. Waiting to be designated converts a voluntary compliance action into a regulatory enforcement action. The difference matters for the fine calculation: the Act's penalty structure treats non-cooperation as an aggravating factor. According to the Regulation's penalty framework, a significant percentage of total worldwide annual turnover applies to non-compliance with Article 51 obligations—and the EU AI Office's guidance from late 2025 indicates that voluntary self-reporting before designation is considered a mitigating circumstance in fine calculations. The mechanism is simple: report first, negotiate second.
Rule 2: If your cumulative compute is between 10^24 and 10^25 FLOPs, self-report anyway. This is the counterintuitive move that most labs resist. The EU AI Office retains discretionary authority to designate models below the hard threshold if they pose systemic risk—and the Office's December 2025 implementation guidance explicitly states that cumulative compute approaching 10^25 FLOPs, combined with large-scale deployment, can trigger designation. The expected-value calculation from the earlier section applies here: the turnover fine is a catastrophic tail risk that dwarfs compliance costs. Compliance with systemic-risk obligations—technical documentation, incident reporting, cybersecurity measures—costs a fraction of the potential fine. The asymmetry is not close.
Rule 3: Assume a 2x error margin on all FLOPs estimates. The measurement problem is not hypothetical. Epoch AI's methodology carries a confidence interval on FLOPs estimates, but that interval assumes the estimator knows the true architecture and training configuration. In practice, labs running proprietary mixtures-of-experts architectures or using hardware with undocumented efficiency gains face measurement uncertainty that easily doubles. If your internal estimate is 6e24 FLOPs, treat it as potentially 1.2e25 FLOPs and act accordingly. The cost of over-reporting is a compliance obligation; the cost of under-reporting is a fine. The asymmetry of error costs dictates the direction of your assumption.
Rule 4: Never attempt to split training runs to stay under the threshold. Article 51's cumulative compute clause is explicit: the threshold applies to cumulative training compute across all runs for a single model. The EU AI Office's technical guidance from January 2026 clarifies that "cumulative" includes continued pre-training, fine-tuning runs that exceed certain compute thresholds, and even the sum of multiple checkpoints trained from the same base. Splitting a 1.5e25 FLOPs training run into three 5e24 FLOPs segments does not create three models—it creates one model with cumulative compute above the threshold. The Office has stated it will use hardware utilization logs and cloud provider records to verify cumulative compute. This is a guaranteed violation, not a loophole.
Rule 5: If the expected fine (detection probability × a significant portion of turnover) exceeds 10x your compliance cost, always choose compliance. This is the decision rule that resolves the ambiguity. The variance of the fine matters more than the expected value: a turnover fine for a lab with thin margins can threaten solvency, not just profitability. Compliance costs are bounded and predictable; fines are not. The decision tree below summarizes the framework.
| Blind spot | What it seems to imply | Why the decision rule still holds |
|---|---|---|
| a significant error margin | Your model might be under the threshold | It might be over it; your estimate is not the number |
| Sparse MoE architectures | Capability without FLOPs, metric mispriced | Exploiting it bets the company on a technicality |
| Fine is a ceiling, not floor | Penalty may land at a lower percentage, not the cap | Regulator discretion can also keep it at the ceiling |
| Detection probability low | Chance of getting caught is low | The probability is unknown, not low — and the tail is ruinous |
| Rebuttable presumption | You can argue your way out | Litigation risk, not a safe harbor |

Also worth reading: The most important books for mastering the art of decision making: most important books for mastering · Podcasting in the UK: The Unseen Regulatory Weight of the Data Protection Act 2018: Podcasting in the UK: The · The Online Safety Act 2023 A Philosophical Analysis of Digital Rights vs State Control: Online Safety Act 2023 A
Nova AI's 1.5e25 FLOPs Dilemma
The cultural reference point here is the film Disclosure Day, where the protagonists' entire mission is to voluntarily disclose information that, if revealed under duress, would be catastrophic. The EU AI Office's designation process works the same way: the information you volunteer is the same information they can compel, but the framing determines the penalty. The five rules above are not about the science of FLOPs measurement—they are about the asymmetry of regret. A lab that self-reports and is wrong about its compute spends money on compliance it might not have owed. A lab that stays silent and is wrong spends a significant portion of its turnover and possibly its existence. The rational choice, under every plausible probability distribution, is to report.
The counterintuitive part is that the cost of compliance is not in the same universe as the cost of the fine. According to internal budgeting benchmarks typical for labs of this scale, the full compliance package—including adversarial red-teaming, EU legal counsel, and technical documentation—totals a modest amount. This is not a rounding error against the fine; it is a rounding error against the R&D budget. The decision tree, therefore, is not about whether the 10^25 FLOPs metric is scientifically defensible. It is about whether you are willing to wager your company's existence on the EU AI Office's detection capabilities.
That wager is a losing one. If Nova AI under-reports its compute and the detection probability is high—a
Frequently Asked Questions
If my lab's cumulative training compute is 9.9 x 10^24 FLOPs, do we face EU AI Office oversight?
At 9.9 x 10^24 FLOPs the regulatory gaze is absent, while at 1.01 x 10^25 FLOPs the full machinery of the EU AI Office descends.
Can we avoid the 10^25 threshold by splitting one 2 x 10^25 FLOPs training run into two 1 x 10^25 FLOPs runs?
No, the cumulative compute clause sums coordinated training runs, so splitting a 2 x 10^25 FLOPs run into two 1 x 10^25 FLOPs runs does not evade the threshold and triggers it retroactively.
What obligations actually activate once a model crosses 10^25 FLOPs?
Mandatory red-teaming, adversarial testing, and incident reporting to the EU AI Office apply regardless of whether the model exhibits any observable dangerous capability.
Can the EU AI Office regulate a model trained below 10^25 FLOPs?
Yes, even below 10^25 FLOPs the EU AI Office can designate a model as systemic risk based on capabilities or impact, such as pulling in a model trained at 8 x 10^24 FLOPs with unexpected tool-use or long-horizon planning.
What is the penalty for violating EU AI Act obligations?
The penalty for violating obligations reaches a significant percentage of total worldwide annual turnover, which for a lab with $50 billion in global revenue is a substantial portion of that revenue.
Have current frontier models already crossed the 10^25 FLOPs threshold?
Yes, GPT-4 at 2.1 x 10^25 FLOPs and Llama 3.1 at 3.8 x 10^25 FLOPs have already crossed the 10^25 line, and training compute for frontier models has been doubling every 6-10 months.
Quick answers
| What is the 10^25 FLOPs threshold described as in the article? | The 10^25 FLOPs threshold is a cognitive trap that incentivizes under-reporting. |
| What does the EU AI Act's Article 51 measure according to the article? | Article 51 of the EU AI Act does not measure risk; it measures arithmetic. |
| What happens if a lab trains multiple models in a coordinated way? | If a lab trains multiple models in a coordinated way, the FLOPs are summed, and splitting a 2 x 10^25 FLOPs training run into two 1 x 10^25 FLOPs runs does not evade the threshold; it triggers it retroactively. |
| What is the penalty for violating obligations under the EU AI Act per the article? | The penalty for violating obligations—or for engaging in prohibited practices—reaches a significant percentage of total worldwide annual turnover. |
| What is the only way to eliminate the tail risk of a fine according to the article? | The only way to eliminate the tail risk of a fine is to self-report early, establishing a cooperative posture before the Office's discretion is exercised adversarially. |
Research Methodology & Editorial Standards
We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.
Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.