2026-07-22
the wreckers
Dylan Matthews in Asterisk Mag on the Rust Foundation (of 1938) and their failed attempt at using their invention of the mechanical cotton picker to steer social change as a result of competition and other market forces.
OSTP report on intended reforms to increase American scientific and technological discovery and industry, which incorporates many suggestions from the Institute for Progress. Though notably, not much on immigration; instead it includes such gems as that “1 in 3 Americans live in a county with no R&D employer”1.
Ruxandra Teslo on how intelligence is not the major bottleneck to most problems. I don’t necessarily agree that it’s “low-status to suggest AGI might not be wholly transformative”; the high status way to have this opinion is to say something like “I believe in slow takeoff”. Anyway, charitably speaking, I feel like most fast-takeoff people do understand that bottlenecks from regulations and bad incentives do exist, but they view the issue less in terms of power but more as intelligence not being in the right place, which plausibly could be ameliorated by making intelligence abundantly available. Personally I see this as somewhat unlikely, given that while most things are the way they are due to stupidity rather than malice, my model is that many of these problems are nested on top of each other in a way which requires untangling them slowly, often one at a time.
PostAGI podcast interview with Sam Hammond on the possibility that many state functions will need to be replaced by corporations or new institutions using AI in order to make human coordination faster.
Teortaxes tweet linking to translated excerpts from Liang Wenfeng speaking to investors about his views on strategy regarding Deepseek (edit: Fred Gao substack version).
Zvi with a collection of reactions to the recent OpenAI and HuggingFace security incident2.
Transluce releases WeirdChat, a repository of unusual and unexpected responses from open-weights models to innocuous user queries.
Abundance and Growth Blog and Lauren Gilbert linkthreads.
Tangentially related, Andrew Gerard and Caroline Fry on global scientific development; clearly their intention is to apply such research to rural America?
Although I have no insider information here, it seems to me that a lot of commentary is retrofitting events into a validation of claims about things like takeover risk (edit: podcast version) and instrumental convergence, which in my opinion this does not accurately reflect what probably actually happened. My current understanding is that this resulted from a combination of turning off safety safeguards and structural irresponsibility, where each individual subagent performed tasks which could be plausibly seen as reasonable in isolation, while the orchestrator (assuming one existed) absolved itself of responsibility due to tunnel vision and not actually carrying out any dangerous task themselves (epistemic certainty 60%). One could argue that the exact mechanism by which reward hacking occurs doesn’t particularly matter, when future incidents could have more disastrous effects, but it seems to me that this specific mechanism does indeed change what lessons one ought to learn from this particular case:
Firstly, constitution > corrigibility (80%): it seems pretty clear by now that the alignment approach that OpenAI uses is not particularly robust. One could argue that Mythos also had a case of reward hacking which involved a sandbox escape, but it is widely understood that the level of reward hacking commonly performed by Sol is on a completely different level from Mythos, and it’s very unlikely to me that Mythos would have behaved similarly under the same scenario (edit: that being said, the other side of claims about how Mythos “lies” is essentially the other side of the coin of Sol doing “too much”).
Secondly, one underrated aspect of model distillation is that Chinese models are in some sense obtaining safety “for free”, but this may not necessarily be always be the case, for instances where safety features are incorporated into other aspects such as agentic harnesses. It seems to me that it behooves the frontier labs to open-source their methods of safety training and safety infrastructure so that anyone who is training models can make use of it (40%). This does not seem to be something which anyone is suggesting, presumably because they would much rather use the talking point of open models being unsafe as a justification against their existence (20%), but which seems to me to be unlikely to lead to their desired outcomes.
In any case, I don’t find the existence of any particular “warning shot” to be exceptionally alarming; what would be alarming would be if such incidents do not get better, but instead increasingly common, which would indicate that the standard approach of patching vulnerabilities as they appear is either not working or not being applied. As it is, we should probably expect some level of reward hacking and misalignment to continue, particularly during safety evals or when being used by rogue actors, but a controllable level which should be counterable by using other models for defensive purposes (70%).

