2026-06-03
gold dust
80K Hours interview with Rohin Shah on AI safety at Deepmind. I wonder to what extent members of technical staff have their attitudes towards safety affected by the dynamics of their lab: Matt Levine famously describes Sam Altman as an unaligned AI capable of superpersuasion, who easily outmaneuvers those who are smarter than him; perhaps there’s a particular advantage in Google in being able to see firsthand how internal politics and organizational incentives lead even the most intelligent individuals towards slightly misaligned behavior, as a more structural equivalent.
Holly Elmore arguing against the merge thesis of AI alignment. This is an idea I’m particularly partial to, so I want to issue the following defenses: Firstly Holly seems to be operating from a perspective that superintelligence is imminent, which to me seems highly unlikely, both from a mechanistic point of view and as a result of anthropic reasoning. Secondly, the entire claim that the machine must “want” to merge with humans ignores the fact that it’s exactly the wanting of humans which is most valuable to current LLMs1; assuming that AI becomes symbiotic on human steering (and therefore never needs to develop it’s own), would render the idea of separation as absurd as the claim that the prefrontal-cortex must inevitably seek independence from the limbic system.
Daniel Muñoz review of Duty to Self and the framework that it is not the self but the perspective which is the relevant unit of analysis, allowing one to apply their system of choice, whether consequentialism, deontology, or anything else, towards one’s duties towards their future self. It’s quite interesting in how flexible it is, but at the same time it seems to me that most of the objections that Daniel and others raise are also a direct result of exactly this flexibility, in that it cannot actually provide definitive answers to questions or objections without applying some specific moral system first: the consequentialists might consider the utility of counterfactuals or aggregate the benefits and harms between perspectives, entirely different factors like duties or rights which would matter more to deontologists2. In that sense it’s usable from an individual perspective, but presumably exactly why society as a whole finds topics like paternalism or suicide to be particularly controversial.
Samuel Hughes in Works in Progress on the sankin-kōtai system of the Japanese shogunate as a mechanism to prevent internal rebellion and civil war. In addition to the economic aspects, it seems to me that the even more important factor was that staggered hostages turned the act of rebellion into a stag hunt, such that even if the Maeda or Satsuma clans could theoretically challenge the Shogunate, neither had the confidence to do so, knowing they might be alone.
SemiAnalysis (partial paywall) economic analysis of the viability of space datacenters.
ACX reader book reviews (have not read).

According to the theories of Steven Byrnes on steering versus learning systems.
And actually, consequentialist methods of aggregating over time don’t even require Schofield’s framework in the first place, so in some sense this is entirely a method for rescuing deontologists from getting themselves stuck in a corner.

> Secondly, the entire claim that the machine must “want” to merge with humans ignores the fact that it’s exactly the wanting of humans which is most valuable to current LLMs1; assuming that AI becomes symbiotic on human steering (and therefore never needs to develop it’s own), would render the idea of separation as absurd as the claim that the prefrontal-cortex must inevitably seek independence from the limbic system.
Integration into an organism requires alignment, solving the problems of integration, or else you don’t get a highly integrated organism: https://pubmed.ncbi.nlm.nih.gov/19805423/
Symbiosis also requires alignment! The relationship can become parasitic or adversarial unless some kind of alignment keeps it mutualistic. More here: https://hollyelmore.substack.com/p/genes-did-misalignment-first-comparing?r=15m0xc&utm_campaign=post&utm_medium=web