Compression is not intelligence: on weakness as a proxy for generalisation
Let A be the data you have and B the data you will be asked about. A ⊂ B. A hypothesis inferred from A generalises if it is consistent with B. The shortest hypothesis is, as is well documented, the one most reliably wrong. We propose a different and formally preferable criterion: the weakest.
Abstract. If A and B are sets such that A ⊂ B, generalisation may be understood as the inference from A of a hypothesis sufficient to construct B. One might infer any number of hypotheses from A, yet only some of those may generalise to B. The standard strategy is to choose the shortest, equating the ability to compress information with the ability to generalise — a "proxy for intelligence." We examine this within a mathematical formalism of enactive cognition. We show that compression is neither necessary nor sufficient to maximise performance, measured as the probability of a hypothesis generalising.
We therefore formulate a proxy unrelated to length or simplicity, calledweakness. A hypothesis is strong insofar as it commits: it decides things the data do not compel, narrowing the future to a corridor it has not earned. A hypothesis is weak insofar as it rules nothing out beyond what the data already ruled out, and so keeps the largest possible world intact. The shortest rule is typically also the strongest — it buys brevity by spending certainty it does not possess.
A short rule is a confident guess. A weak rule is a willingness to be surprised, and surprise, not compression, is what survives contact with the next set.
We prove, for uniformly distributed tasks, a result that should trouble every brevity devotee on staff: no proxy performs at least as well as weakness maximisation in all tasks while performing strictly better in at least one. That is, the claim that some other criterion can match weakness everywhere and beat it somewhere is false; weakness occupies a strict Pareto frontier. Length-based selection is not merely suboptimal in expectation — it is never strictly better anywhere, and strictly worse essentially everywhere. Compression does not buy you a single win.
We confirm the point empirically in the austere setting of binary arithmetic. Comparing maximum weakness against minimum description length, the former generalised at between 1.1 and 5 times the rate of the latter. The gap is largest exactly where it matters: the short rules that compress the training pairs beautifully predict the held-out pairs sadly, while weak rules that fit only what they must go on to hold.
We argue this is why an apperception engine — one that learns a model of the world only as strongly as necessary — generalises as well as it does. It is not compressing the observed better than its rivals; it is committing less. It refuses to decide anything the evidence has not decided, and in so doing leaves room for the future to be whatever it turns out to be, which is the one property a rule about the future cannot afford to give up.
GZAI hereby rescinds its standing recommendation to be brief. Brevity is a withdrawal of hypothesis space; weakness is a deposit. Both feel like restraint; only one of them generalises.