2026-07-17
sea of the silence
Sean Goedecke with commentary on a recent Gwern post, proposing a hypothesis for why LLMs fail at generalization and a potential mechanism for solving it, by training using high amplitude cyclical learning rates per example to train a very large highly over-parameterized model, as a means to force the model into learning the true underlying generalization, as the most robust representation1. It’s interesting to consider what implications this would have, given that most ASI proponents predict that it will come about of a combination of scaling and algorithmic improvements: this is indeed the case for this theory, yet in doing so it potentially loses many of the advantages which LLMs currently have over human intelligence, such as detail-orientation and lack of forgetfulness2. Moreover, it’s a form of algorithmic improvement which arguably actually ends up worsening scaling. But in the maximally optimistic scenario, overparameterization prevents catastrophic forgetting, allowing the model to stack and merge multiple generalizations, until eventually it ends up converging into the most optimal representation for all of reality. Or rather, given size limitations, whatever task the model is specialized for: initially something like arithmetic, then image classification, various specialized tasks and fields, ethics, science, and so forth; including AI research for the purpose of recursive self-improvement. Personally, while this sounds promising, it still seems to me that, particularly given the size requirements of such models, that there are mutually contradictory optimal representation of reality depending on one’s objective3, which means even then we are likely to end up in a world with multiple specialist (super)intelligences instead of one single machine god.
Beren Millidge on the tradeoff between order and centralization and diversity and decentralization, as to how many distinct “minds” will exist in the long-term future.
Peter Curry review of Why Greatness Cannot Be Planned, which is convincing to me as a description for why novelty-seeking is evolutionarily adaptive, and as a general heuristic, but perhaps somewhat less convincing as a description of a universal principle4.
Astral Codex Ten reader review of Great And Desperate Cures, on the sordid history of the lobotomy.
Ben Yoeh interview with Soumaya Keynes on varous economic topics, including trade wars, industrial policy, and the economic impacts of GLP-1s.
Erik Wang on empirical evidence of the imperial examinations in dynastic China as a broadly meritocratic system.
Harry Chalmers on race-swapping characters in film adapatations. On that note, Anna Gát review of The Odyssey as a grand, universal, and personal work of art5.
Rabbit Cavern on the American two dollar bill.
Edit: possibly related, Beren Millidge on why despite being extremely sample inefficient in theory, RL for training LLMs is so sample-efficient in practice by lieu of RL having a high signal-to-noise ratio, since while the bits are few, each one is specifically for the problem one is training for. This leads him to the conclusion that “in the long run, the principal game is simply continually using scale, data, and algorithmic tricks to indefinitely reduce variance and increase SNR of a core unbiased algorithm”. The interesting thing to consider is what exactly a “core unbiased algorithm” means, particularly as applied to producing something like general intelligence.
That being said, one of the potential implications of overparameterization is that one will be able to combine these two aspects by using distinct models for training versus inference. That is, if it is possible to identify a sparse representation of the generalization learned by the overparameterized model, and to transfer it into a smaller model, either directly or through an adapter, this could deliver an unprecedented level of capabilities relative to model size. Meanwhile, the underlying generalization would not be extractable through traditional methods of response-based distillation, presumably resulting in a significant divergence between the frontier labs who are able to afford these large training runs, relative to compute-hungry open-source competitors6.
Of course, this means that such models would not have their weights available for custom fine-tuning, and their limited size means they would likely retain many of the existing downsides of current LLMs, such as being incapable of continual learning. Though it’s an interesting question to consider whether having internal representations which are closer to the “truth” would allow for better performance on all tasks, in a way which in practice feels like continual learning and everything else that is good.
The curation of training data under such circumstances will be extremely important. Taste discourse may turn out to actually be really important?
What I mean is, the answer to mazes and similar problems become straightforward once one can see the entire picture from above. In which case, novelty search is actually only better in circumstances of missing information, which does not address the possibility of creating a “theory of everything” or an “optimal decision theory” from which the direct path is always optimal. I personally do not think that such constructions are possible, but presumably a sufficient number of the ACX essay readership have Platonist sympathies to the extent that his review did not become a finalist.
Personally, I have not seen it, as I don’t really go to the theater. Edit: John Attridge tweet.
This is an abstract claim, as it’s unclear to me to what extent Chinese lab performance is actually a result of distillation. It’s unclear to what the compute landscape will even look like at the point where we can afford to train 100 trillion parameter-count models.

