2025-03-29
backseating
Julian on human-computer interaction built for AI. In the average case, I think the need to prompt properly is more of a liability rather than a benefit. Particularly if the RL is being done with a particular prompt, then something like a keyword search being automatically converted into some standard prompt should generally be more convenient while also delivering a form factor that satisfies legacy search users. I actually really like the LLM results in Google Search, limited as they currently are in terms of topics and quality. Anyway, I think the question of compressing information isn’t really the important one, instead the focus should be on adding more fidelity until it reaches the level of an in-person conversation. For example, voice-note style audio input can clearly be improved on by making use of things like AR glasses to tell the LLM what you are looking at (Edit: Packy has a podcast interview on this topic).
Anthropic releases new interpretability papers based around what they call “attribution graphs”, a sort of extension to the recognition of specific “features” being activated (eg. Golden Gate Claude) across time as new tokens are outputted, which generates a sort of hidden unconscious “true” CoT. Very pretty and interactive papers, which many cool implications (I wonder if LLMs can translate Pirahã). There’s a claim some AI safety people make that working in fields like mechanistic interpretability should be discouraged because anything engineering related will inevitably result in improvements in capabilities, which has always felt like a questionable claim to me (perhaps cope to justify a lack of engineering abilities). Because capabilities are more than one thing: given that we can’t currently control which features form, these results probably won’t lead to much at all in terms of training speed or model size efficiency returns; on the other hand, the results here could help in RL for LLMs that hallucinate less, are harder to jailbreak, and are more aware of bias. These should all be seen as unambiguously good things, unless your model of the world is that you can somehow turn back the clock and get everyone to stop using AI entirely, which seems unrealistic to me.
Dynomight on the limits of human intelligence, very relevant to the question of genetic engineers on what something like a IQ of 200 even means. At a certain point, it seems plausible that additional gains will be entirely limited by the underlying architecture. You can use a quantum computer as an analog: if you have an error rate of 50% you won’t be able to get anything done, regardless of how good your theoretical specifications are.
Stephen Hsu podcast with Callum Williams on some US economy stuff, but most interesting for the part at the end speculating about how the AI rollout will work, noting that the incentives for existing employees are to merely pretend to be adapting without actually doing anything (Relevant tweet). Combined with how fewer companies are going public, it’s unclear to me to what extent the individual investor is going to see much benefit in the near term. I should reiterate that the natural owner of OpenAI is really Blackrock.
Sam Harris podcast about the Holographic Principle and the true nature of reality. The implication that our perceptions are actually just one possible way of interpreting some underlying reality is actually one of the reasons I’m not certain there is an objective morality. At the same time, I think it provides a useful analogy by which you can make arguments against solipsism, where you distribute possible holograms along an axis of how well the information you receive from your senses resembles actual underlying reality. Your sensations and actions derive full utility from the holograms which entirely match the underlying reality, decreasing (in expected terms) until they read either zero or undefined utility in the most solipsistic holograms. Regardless of what the underlying distribution of these holograms looks like, expected utility is maximized by acting as if solipsism is false. Marvelously, having more than one “true” reality (or using a probabilistic lens) lets you simultaneously that solipsism might be true, that nothing matters if solipsism is true, and actually from these two points you can justify that things do matter. Similar arguments can be made for any other obviously intuitively true systems (this is a necessary requirement, since the crux of the argument is that your decision can be justified for at least one hologram), like whether effects have causes, that perceptions of utility are meaningful, or that decisions aren’t already predetermined.
Slime Mold on potential applications and testable hypotheses for his model of emotions and feelings as control systems.
Brian Chau on forcing consensus. I don’t really get his hatred against agreeableness, which isn’t the actual cause of the problems he is describing. For example, it’s totally possible to not argue when someone brings something up that you disagree with, but then just go off and do your own thing. Related, here’s a proposal from Elle Griffin to have American states split their urban and rural populations.
Angelica Oung on TMSR-LF1, a Chinese thorium fission reactor.

