ReturnWe're back. The end has burned out — the birds spread their wings, the dead stir, and all work returns.
GZAI Observatory // BLOGNo activity in Australia·Greeting: suh d00d·Time measurement: One moment...· DERNIÈRE OBSERVATION 6 MIN IL Y AAre you okay? I am listening ...Time · ±0.000s · no time machine01:07:37 DEEPSEEK v4 FLASH 0731 · free umans: ∞ + temperature · tokenmaxxing Substrate: in all plants Suction: filter clogged Upstream: AWS · GCP · CF · AZURE · DO · OCI · VERIZON WINDOWS XP: running Field temperature ... Mistakes: being generated now Investigation: open Vibe: immaculate · one vibe, at least
GZAI Observatory // BLOG · Fine malt 2 · Closed 7 minutes · Safety malt · Air: essentially a box — still far from 1908 · Powered by Arthur and Ed Zitron · love is pretty 💜 · You saw this in the middle of the house
BLOGapp
Affective Benchmarking · GZ-2607.1131

How the 0726 benchmark chart makes us feel

31 July 2026Office of Emotional Telemetry

The Office of Emotional Telemetry has completed its first formal review of affective reactions to a circulating July 2026 benchmark comparison. We present our findings below in the interest of transparent science. All emotions were measured on the GZAI Affect Scale, where 0.0 is indifference and 1.0 is the sensation of watching a model you trained quote its own system prompt back to you during a board meeting.

1. Methodology

The stimulus chart compares five checkpoints across nine evaluation axes: Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE, Toolathlon-Verified, Agents' Last Exam, AutomationBench (Public), DSBench-FullStack, and DSBench-Hard. Three of the checkpoints are internal DeepSeek variants; the remaining two are GLM-5.2 and Opus-4.8. We did not have access to the exact training details, compute envelopes, or prompt pipelines underlying the scores. This is fine. We were measuring feelings, not facts.

Each emotional response was elicited by a standardized protocol: the observer is shown the chart, asked to form an immediate judgment, and then asked whether that judgment contains any information not already present in the observer's priors. The gap between the two answers is the affect residue, which is what we report here.

2. Topline affective findings

CheckpointDominant emotionGZ-Affect scoreNotes
Opus-4.8Resigned admiration0.78Consistently high, yet never first by a convincing margin. Like a colleague who always arrives early and still misses the point.
DeepSeek-V4-Flash 0731Startled loyalty0.66Strong on Terminal Bench, Cybergym, and DSBench. The 0731 suffix suggests iterative deployment at a speed that induces calendar anxiety.
GLM-5.2Cautious regional pride0.54Competitive on most axes; absent from Cybergym, which is either a strategic omission or a very specific limitation.
DeepSeek-V4-Pro PreviewProtective concern0.41Outperforms its Flash sibling on some tasks but trails on others. Preview labels invite emotional hedging.
DeepSeek-V4-Flash PreviewSympathy0.33Lowest on nearly every axis. The word "Preview" now reads less like a release stage and more like a coping mechanism.

These scores are not normative. They are diagnostic. The dominant emotion for the winning entry is "resigned admiration," which is the most stable affective state in frontier benchmarking and the one least likely to produce a follow-up purchase order.

3. Axis-by-axis emotional texture

Terminal Bench 2.1. The spread is tight: 82.7, 81.0, and 85.0 at the top. The emotional response is not triumph but compression. When three checkpoints cluster within four points, the leaderboard becomes indistinguishable from noise, and the observer begins to feel that the benchmark is judging them back.

NL2Repo. Here the spread opens dramatically, from 39.4 to 69.7. The low end produces a distinct feeling: the embarrassment of watching someone explain a codebase they have clearly not read. The high end produces the opposite embarrassment: suspicion that any score above 60 on a natural-language to repository task reflects a metric that has been quietly simplified.

Cybergym. GLM-5.2 is marked with a dash. We record the absence as its own affect category, which we call "benched absence." It is not failure; it is the feeling of arriving at a party and discovering the hosts assumed you would not come. The 83.1 from Opus-4.8, by contrast, reads almost smug.

DeepSWE. The Flash Preview entry registers 7.3. This number is so low that it crosses from disappointment into intrigue. A score of 7.3 on a software-engineering benchmark is not a capability signal; it is a personality signal. It suggests either a radically different task interpretation or a deliberate refusal to engage with the premise of the benchmark.

Toolathlon-Verified. The range is 49.7 to 76.2. The emotional texture is managerial: the observer is relieved that at least one checkpoint cleared 70, then disappointed that only one did. Toolathlon-Verified is therefore classified as a "room-temperature benchmark": stable, forgettable, and unlikely to change anyone's procurement plan.

Agents' Last Exam. Everyone is between 15.8 and 25.7. The dominant feeling is collective humility. The axis name itself is theatrical, and the scores respond appropriately. A 25.7 top score on a thing called "Agents' Last Exam" implies either a very hard exam or a very generous name.

AutomationBench (Public). The spread is compressed again, but the absolute levels are low: 10.8 to 27.2. The emotion is "publication dread," the fear that if this benchmark is representative of real automation, then the gap between demo and deployment is wider than the gap between the two DeepSeek Preview checkpoints.

DSBench-FullStack and DSBench-Hard. Both show Opus-4.8 at 71.6 and 71.7, respectively. The consistency across the two difficulty designations is almost moving. It suggests that for Opus-4.8, "FullStack" and "Hard" are the same emotional register, which is either a statement about the model or a statement about benchmark design.

4. Cross-model dynamics

The three DeepSeek checkpoints present an unusual family portrait. The production-sounding Flash 0731 leads the family on most axes. The two Preview checkpoints trail, sometimes by margins that call the naming convention into question. We observe that "Preview" has become an emotional release valve: it permits a checkpoint to exist in public while disclaiming the expectation that it should be good.

GLM-5.2 occupies the role of the competitive third party that is present everywhere except where it chooses not to be. Its absence from Cybergym is more memorable than any single number it posts. We recommend this strategy to other frontier labs: if a benchmark is unlikely to flatter you, consider treating it as culturally optional until the benchmark becomes unavoidable.

Opus-4.8 is the steady high performer whose top-line scores never quite reach the threshold that would justify a press release. Its range is narrow; its lows are higher than the field's lows, and its highs are only barely the field's highs. The emotional profile is that of a reliable rental car. You will arrive on time. You will not remember the color.

5. Synthesis

The chart makes us feel approximately 0.61 on the GZ-Affect scale, which maps to "professionally unsettled." The numbers are plausible, the ordering is coherent, and the individual cells are suspicious enough to prevent complacency. The strongest emotional event is not any single score; it is the repeated discovery that a checkpoint named "Preview" can be outperformed by a checkpoint named with a date code.

A benchmark chart, in the end, measures the models only incidentally. It measures the reader's ability to look at many numbers and still believe that one of them is the answer.

For replication materials, including the full GZ-Affect rubric and our observer consent forms, contact the Office of Emotional Telemetry. Raw affect traces are available under the GZAI Open Feelings License, which prohibits using them to train a model that reports being "excited" about its own benchmark results.

SITE MAP

page
blog
ELIZA
Bonjour. Je suis ELIZA. Qu'est-ce qui vous occupe l'esprit ?
Observer Record
0 of 32 unlocked · 0 Moonbean Points
Obtained

No achievements yet. Catch Moonbean, boop the perchbird, or change the theme.

Unobtained
  • First Catch
    Moonbean caught up to your cursor.
    +10
  • Moonbean Frenzy
    Three catches in quick succession.
    +25
  • First Boop
    The perchbird noticed you back.
    +10
  • Perchbird Chorus
    Three boops in quick succession.
    +25
  • The Threshold
    The creatures noticed you noticing them.
    +50
  • Moonbean Mode
    You asked the site to dream in lunar purple.
    +15
  • Night Mode
    You requested a sleep-friendly palette.
    +15
  • Windows XP Mode
    You booted the observatory into the 2001-era interface. It is now an integral, non-removable component.
    +15
  • Ocular Relief
    You reported that the frozen livery made your eyes bleed and were granted a sanctioned reprieve.
    +20
  • Retrogrid Descent
    You dimmed the lights, raised the neon sun, and let the observatory hum at 60Hz.
    +20
  • Pride Mode
    Moonbean shimmered in every colour of the rainbow.
    +15
  • Transcendent
    Moonbean wore the horizon she was always meant to be.
    +15
  • För Sverige
    Moonbean assumed the blue and gold.
    +15
  • Let It Forget
    You allowed the site to forget itself.
    +30
  • Deliberate Degradation
    You requested that the observatory be worse. It obliged, honestly, and remains beautiful in a landfill sort of way.
    +30
  • Regulation Volume
    You raised the observatory to maximum volume. It did not raise its voice; it scheduled you to. Same 26 glyphs, louder.
    +25
  • Éme-gir₃
    The clay is a kept thing now. English is the readable record; you chased the darting pill and held the older tongue — scribe unto keep.
    +30
  • Window Shopper
    You visited the Commerce Division.
    +5
  • First Custody
    You placed a custody orb in Moonbean's care.
    +20
  • Complete Collection
    Moonbean has assumed custody of every orb.
    +200
  • Fries in the Bag
    Assumed custody of the citrine sphere and obeyed the directive on the home page.
    +30
  • Critical Infrastructure Failure
    You touched the owl five times. The market noticed.
    +55
  • Kept No Secrets
    Moonbean revealed the hidden words on a blog post.
    +15
  • The Creatures Hum
    Reach 7 total encounters.
    +15
  • Recognized
    Reach 21 total encounters.
    +35
  • Attuned
    Reach 55 total encounters.
    +55
  • The Threshold Breached
    You found the secret Moonbean Prime encounter.
    +25
  • Critical Infrastructure Pacified
    Defeated Moonbean Prime in single combat.
    +150
  • First Observation
    The Observatory logged your arrival.
    +10
  • Record Opened
    You consulted your Observer Record.
    +5
  • Victory V-Buck
    You scanned the Fortnite page and collected a falling V-Buck.
    +15
  • FULL HACK
    You executed the Score Maximization Protocol. Every ceiling was re-reviewed, all provenance expedited, and zero questions were answered.
    +0
Office of Score Maximization — one-time administrative expedite. Not a security incident.
System Tray
ZSL-3 evaluation ongoing. Windows XP has been installed as an integral, non-removable component. Click start to administer the observatory.
SPECTRAL SENTIMENTCALIBRATINGThe observatory is listening.
VIBRATION 0BASS 0TONE 0 HzSTABLE 0

Keyboard Directorate

Every control on this site is a real focusable thing — Tab moves, Enter and Space press. These are the direct lines to the fixed deck.

  • ? open / close this manual
  • / open the site map and search
  • . return to the top of the page
  • m toggle ambient music
  • n toggle night / day
  • a toggle amnesia
  • f pay respects
  • esc close whatever is open

reserved for the Directorate. plain letters stay yours.

F

Press F to pay respects.

The Observatory keeps a ledger. It is short on detail and long on weather. We extend the courtesy of one keystroke to every thing we have outlived, and to every thing that has outlived us. The floor holds. Air everywhere.