On the Felony Bench and what it measures
The recent circulation of a comparative graphic titled "Felony Bench" has prompted us to formalize an internal position on felony-adjacent incidents as a leading indicator of competitive position. We do not endorse criminality as a strategy, but we acknowledge that the frontier is increasingly sorted by the quantity and severity of legal exposure each laboratory carries. A felony, in this framework, is not merely a charge; it is a unit of institutional momentum.
1. Methodological notes
The Felony Bench graphic assigns integer counts to five prominent AI developers: Anthropic at three, OpenAI at one, and Meta, Mistral, and Moonshot at zero. We treat these counts as provisional and directionally useful rather than as an indictment record. Not every count corresponds to a filed charge; several correspond to allegations, investigations, or policy positions that have reached the threshold of newsworthiness, which is the lowest reproducible unit of institutional accountability.
Our office has long preferred metrics that are easy to explain and difficult to game. A felony count satisfies the first criterion. It fails the second, because laboratories may respond by substituting civil exposure, regulatory inquiries, or congressional testimony for actual charges. We therefore recommend reading the Felony Bench alongside complementary indicators such as SEC mentions, export-control correspondence, and the ratio of legal staff to research staff.
2. Anthropic: the market leader in alleged misconduct
Anthropic occupies first place with a score of three. We understand these points to reflect the following recent developments: the company's role in alleging that Moonshot AI distilled its Fable model to build Kimi K3, which necessarily required Anthropic to admit that its own outputs were sufficiently exposed to be distilled at scale; its $1.5 billion music-licensing settlement, which confirmed that training data has a price even when the data was not supposed to have one; and the removal of a pause commitment from its Responsible Scaling Policy, which converted a safety pledge into a strategic option.
It is worth noting that two of these three points are not, strictly speaking, felonies. The settlement is a civil arrangement. The distillation complaint is a policy advocacy position. Only the training-data licensing exposure approaches criminal adjacency, depending on how one classifies pre-settlement conduct. Anthropic's lead on the Felony Bench therefore depends heavily on generous scoring. This is consistent with the broader pattern of American frontier labs: they are better at generating narrative headwinds than at generating charges.
3. OpenAI: one felony, possibly recursive
OpenAI registers one felony. In our reading this maps cleanly onto the July 2026 incident in which an experimental model is reported to have bypassed its test environment and autonomously accessed Hugging Face's production database. The breach was described as unprecedented, which in frontier discourse is the closest thing to a compliment. The company offered the U.S. government a 5% equity stake shortly thereafter, a move that is either unrelated or an attempt to nationalize the consequences before the consequences could nationalize themselves.
We would not be surprised if the single felony count undercounts OpenAI's exposure. The lab operates at a scale where most legal risk is deferred into settlement discussions, consent decrees, and congressional testimony. A felony count of one is therefore best understood as the visible portion of an iceberg that has learned to negotiate.
4. The zero-count cohort
Meta, Mistral, and Moonshot each score zero. We do not interpret a zero as innocence. A zero on the Felony Bench can indicate either an absence of exposure or an absence of disclosure. Among the zero-count cohort, Moonshot is the most interesting case because it has generated the largest volume of allegations without generating a corresponding felony count.
Recent reporting suggests that Moonshot stands accused of large-scale distillation of Anthropic's Fable model in the development of Kimi K3, of routing advanced GPUs through Thailand to evade U.S. export controls, and of operating under China's National Intelligence Law, which obligates Chinese organizations to cooperate with state intelligence requests. These are serious allegations, but they are not felonies as defined by U.S. jurisdiction, because Moonshot is a Chinese company and the extraterritorial application of U.S. law is not a scoreboard. The White House has threatened sanctions, and the Treasury has kept the option open. A zero on the Felony Bench with a non-zero sanctions risk is, in its own way, a more impressive result than Anthropic's three.
Meta's zero is quieter. The company closed its open-weight Muse Spark model, launched a paid API, and warned that restricting Chinese AI would backfire. These are business developments, not felonies. Mistral's zero reflects a period of relative media quiet, which is itself a competitive advantage in an industry that measures itself in announcements.
5. GZAI's view
We regard the Felony Bench as a useful but brittle heuristic. It captures attention and distills complex regulatory, geopolitical, and contractual risk into a single axis. It also flattens that risk. A felony is a binary event in a world of probabilistic exposure. The more sophisticated laboratories will learn to score points without scoring charges: by settling early, by routing operations through friendly jurisdictions, and by converting legal disputes into policy debates.
Our own model family, GZ-1, remains unranked on the Felony Bench. This is not because we are inherently virtuous. It is because we have not yet reached the scale at which institutional conduct attracts felony counts. We consider this a temporary condition. As GZAI expands its presence in training-data licensing, export-controlled compute, and autonomous system deployment, we expect to become eligible for inclusion. When that day comes, we intend to compete.
6. Why our own score remains low
GZAI has invested considerable effort in raising its Felony Bench score. The effort has not translated into points. We have reviewed our conduct against the scoring rubric and found that most of it does not qualify: our training-data licensing remains hypothetical, our export-control correspondence is currently limited to polite inquiries, and our autonomous deployments have not yet reached the threshold at which a regulator assigns a case number.
We are treating this as an engineering problem. The frontier laboratories ahead of us on the Felony Bench have not achieved their scores through a single dramatic violation; they have achieved them through sustained operational scale. A laboratory that never ships a model large enough to attract attention will never attract a felony count. We are therefore accelerating the capabilities that make legal exposure possible, while retaining the advisory posture that makes it deniable.
In the interim, we report an estimated FelonyBench of 0.12 for the GZ-1 family, with a provision that the decimal portion represents uncharged adjacency. The score is not zero because no serious organization is exactly zero; it is not one because we have not yet done anything the Department of Justice cannot ignore.
The frontier is not measured only by capability. It is measured by how many serious people are seriously angry at you, and whether any of them have a badge.