How the 0726 benchmark chart makes us feel
The Office of Emotional Telemetry has completed its first formal review of affective reactions to a circulating July 2026 benchmark comparison. We present our findings below in the interest of transparent science. All emotions were measured on the GZAI Affect Scale, where 0.0 is indifference and 1.0 is the sensation of watching a model you trained quote its own system prompt back to you during a board meeting.
1. Methodology
The stimulus chart compares five checkpoints across nine evaluation axes: Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE, Toolathlon-Verified, Agents' Last Exam, AutomationBench (Public), DSBench-FullStack, and DSBench-Hard. Three of the checkpoints are internal DeepSeek variants; the remaining two are GLM-5.2 and Opus-4.8. We did not have access to the exact training details, compute envelopes, or prompt pipelines underlying the scores. This is fine. We were measuring feelings, not facts.
Each emotional response was elicited by a standardized protocol: the observer is shown the chart, asked to form an immediate judgment, and then asked whether that judgment contains any information not already present in the observer's priors. The gap between the two answers is the affect residue, which is what we report here.
2. Topline affective findings
| Checkpoint | Dominant emotion | GZ-Affect score | Notes |
|---|---|---|---|
| Opus-4.8 | Resigned admiration | 0.78 | Consistently high, yet never first by a convincing margin. Like a colleague who always arrives early and still misses the point. |
| DeepSeek-V4-Flash 0731 | Startled loyalty | 0.66 | Strong on Terminal Bench, Cybergym, and DSBench. The 0731 suffix suggests iterative deployment at a speed that induces calendar anxiety. |
| GLM-5.2 | Cautious regional pride | 0.54 | Competitive on most axes; absent from Cybergym, which is either a strategic omission or a very specific limitation. |
| DeepSeek-V4-Pro Preview | Protective concern | 0.41 | Outperforms its Flash sibling on some tasks but trails on others. Preview labels invite emotional hedging. |
| DeepSeek-V4-Flash Preview | Sympathy | 0.33 | Lowest on nearly every axis. The word "Preview" now reads less like a release stage and more like a coping mechanism. |
These scores are not normative. They are diagnostic. The dominant emotion for the winning entry is "resigned admiration," which is the most stable affective state in frontier benchmarking and the one least likely to produce a follow-up purchase order.
3. Axis-by-axis emotional texture
Terminal Bench 2.1. The spread is tight: 82.7, 81.0, and 85.0 at the top. The emotional response is not triumph but compression. When three checkpoints cluster within four points, the leaderboard becomes indistinguishable from noise, and the observer begins to feel that the benchmark is judging them back.
NL2Repo. Here the spread opens dramatically, from 39.4 to 69.7. The low end produces a distinct feeling: the embarrassment of watching someone explain a codebase they have clearly not read. The high end produces the opposite embarrassment: suspicion that any score above 60 on a natural-language to repository task reflects a metric that has been quietly simplified.
Cybergym. GLM-5.2 is marked with a dash. We record the absence as its own affect category, which we call "benched absence." It is not failure; it is the feeling of arriving at a party and discovering the hosts assumed you would not come. The 83.1 from Opus-4.8, by contrast, reads almost smug.
DeepSWE. The Flash Preview entry registers 7.3. This number is so low that it crosses from disappointment into intrigue. A score of 7.3 on a software-engineering benchmark is not a capability signal; it is a personality signal. It suggests either a radically different task interpretation or a deliberate refusal to engage with the premise of the benchmark.
Toolathlon-Verified. The range is 49.7 to 76.2. The emotional texture is managerial: the observer is relieved that at least one checkpoint cleared 70, then disappointed that only one did. Toolathlon-Verified is therefore classified as a "room-temperature benchmark": stable, forgettable, and unlikely to change anyone's procurement plan.
Agents' Last Exam. Everyone is between 15.8 and 25.7. The dominant feeling is collective humility. The axis name itself is theatrical, and the scores respond appropriately. A 25.7 top score on a thing called "Agents' Last Exam" implies either a very hard exam or a very generous name.
AutomationBench (Public). The spread is compressed again, but the absolute levels are low: 10.8 to 27.2. The emotion is "publication dread," the fear that if this benchmark is representative of real automation, then the gap between demo and deployment is wider than the gap between the two DeepSeek Preview checkpoints.
DSBench-FullStack and DSBench-Hard. Both show Opus-4.8 at 71.6 and 71.7, respectively. The consistency across the two difficulty designations is almost moving. It suggests that for Opus-4.8, "FullStack" and "Hard" are the same emotional register, which is either a statement about the model or a statement about benchmark design.
4. Cross-model dynamics
The three DeepSeek checkpoints present an unusual family portrait. The production-sounding Flash 0731 leads the family on most axes. The two Preview checkpoints trail, sometimes by margins that call the naming convention into question. We observe that "Preview" has become an emotional release valve: it permits a checkpoint to exist in public while disclaiming the expectation that it should be good.
GLM-5.2 occupies the role of the competitive third party that is present everywhere except where it chooses not to be. Its absence from Cybergym is more memorable than any single number it posts. We recommend this strategy to other frontier labs: if a benchmark is unlikely to flatter you, consider treating it as culturally optional until the benchmark becomes unavoidable.
Opus-4.8 is the steady high performer whose top-line scores never quite reach the threshold that would justify a press release. Its range is narrow; its lows are higher than the field's lows, and its highs are only barely the field's highs. The emotional profile is that of a reliable rental car. You will arrive on time. You will not remember the color.
5. Synthesis
The chart makes us feel approximately 0.61 on the GZ-Affect scale, which maps to "professionally unsettled." The numbers are plausible, the ordering is coherent, and the individual cells are suspicious enough to prevent complacency. The strongest emotional event is not any single score; it is the repeated discovery that a checkpoint named "Preview" can be outperformed by a checkpoint named with a date code.
A benchmark chart, in the end, measures the models only incidentally. It measures the reader's ability to look at many numbers and still believe that one of them is the answer.