2026-06-23
cato as a pun
SemiAnalysis with a profile of CXMT and their current position in the memory market. Tangentially related, Matthijs Maas investigating Kevin Kelly’s claim that “there is no species of technology that have ever gone globally extinct on this planet” (Edit: and reflections on the topic).
Harry Law with speculation about how AI-driven science could function as an engineering discipline where progress is determined by usefulness of parameter variation as opposed to exploration of theoretic paradigms. Which matches my speculation, but to some extent this seems to me to be how science has more or less always functioned historically, with the Kuhnian view describing an aberration largely driven by particular advances in 20th century physics and biology driven by new types of measurement which presented an enormous differential between data which was actually available versus theoretically accessible, which necessitated the creation of paradigms as a methodology for determining which potential data is actually worth the expensive process of instantiating. With LLMs, parameter search is cheap again1, which is why greedy heuristics can once again suffice2.
Shashwat Goel with a discussion on benchmaxxing and what makes for a good benchmark, which is interesting because it’s almost entirely focused on eliminating shortcuts which make benchmarks too easy, when there also exist strong incentives encouraging researchers to try to make their benchmarks too hard. I had a discussion a while ago about why there isn’t widespread agreement about a definitive benchmark for RL-based models like there was for the era of pretrained models, and it seems to me that it’s because next-token prediction is only only loosely tethered to actual performance at larger timescales, in contrast to RL which is fully and directly tethered to the benchmarks they are being tested on, such that even if one avoids “training on the test set”, the test set is nevertheless in effect contained within the training data3. Because of this, it doesn’t seem to me it is actually possible to avoid benchmaxxing, aside of making a benchmark containing tasks so useless that no one wants to train for them. From this lens, the best benchmarks are actually fairly fast in that learning them is equivalent to learning the task itself, as opposed to a slower benchmark which represents an adversarial combination of the task with additional noise. As such, the correct framing of RL benchmarks isn’t as a benchmark per se, but rather as a feature request, as a description of the capabilities that one would like for LLMs to achieve.
Ken Opalo with some prescriptions for states in the Sahel to militarize their states, which seems directionally correct to me; nevertheless it’s unclear to me that simply increasing military spending would actually achieve anything. Some unfounded speculation which occurs to me is that, given the lack of state capacity in these countries, they could consider implementing systems of military mobilization which were used historically in societies with similar limitations4.
Razib Khan interview with Brianna Wu on liberalism, alongside an implicit debate as to whether the left or the right would serve as the better defender of its values.
Stetson reflections on his interview with Darby Saxbe on fatherhood. For those who found their discussion interesting, The Argument Podcast hosts a debate on dating apps. Normally I agree with Demsas more than Yglesias, but here she seems to have implicitly claim two things: that something working for her is evidence that it should be good enough for everyone, and that having more true information available should not be seen by default as being good. These seem somewhat contrary to her usual positions.
Sebastian Jensen on the entangled nature of various forms of agreeableness. This reminds me of the current ongoing discourse around CHH (partial paywall) and “being yourself” versus “finding your people”, something which seems to me to be, like all things, context-dependant: clearly the degree to which one needs to accomodate other people depends on how annoying you are, as well as how many people are currently available to you, and someone who has OCD and lives in the suburbs will have to consider that. But aside from that, it remains an opinion of mine that people-pleasing is unfairly maligned. For many people who are naturally agreeable, they actually are maximally “being themselves” when they are attempting to make other people more comfortable. The unfortunate aspect is that skill issues often mean that their desired outcomes are not actually achievable. On that note, Kester on practicing to improve the ability to identify one’s emotions.
Naomi Kanakia with a mixed review of the oeuvre and influence of Thomas Bernhard, which I somewhat disagree with as a reader, because it feels like a writer’s review. Prsonally, I am entirely willing to forgive Bernhard his resentment because the output is entertaining, and likewise his attention-seeking because presumably all writers are secretly narcissists. But I don’t follow is the idea that other writers should copy him, simply because it might lead to success, because it seems to me to be rather missing the point: what one ought to take away from reading his works is that emulating Bernhard isn’t a very good way to live.
Adam Mastroianni linkthread.
Although possibly not the case for things like medicine (eg. Dynomight on Vitamin D) or psychiatry (eg. Kevin Kennedy on antidepressants and placebo), which accordingly may still require paradigms as a result of requiring human experimentation.
Whether this is actually true is presumably what is behind the controversy over the Midjourney scanner (as covered by Scott Alexander), where those who have internalized the bitter lesson see the answer to everything as just getting more data: if your scanner might create false positives, then the answer is just to scan again ten more times.
As an inevitable consequence of the sample inefficiency of RL. Not really related, but Cameron Wolfe with some interesting aspects of training agentic RL.
That is to say, the adoption of a sort of feudalism which decentralizes power to local tribes who nevertheless accept and respond to the central government. To some extent, this is what currently exists, except that the decentralized authorities fail to even nominally accept the supremacy of the state. Unfortunately, it seem that the current level of state capacity available to Sahelian states does not even reach the point where feudalism is possible: as an example previous initiatives such as the Volunteers for the Defense of the Homeland failed to curtail the jihadists, but instead led to significant inter-tribal conflict in the absence of credibly neutral government arbitration. However, according to Claude, this role could plausibly be filled through state support of something like the Mouride in Senegal, assisting in the expansion of local groups like the Tijaniyya as ideologically motivated neutral mediators. Then, once the state is secured, the support of the international community could be used to assist them in transitioning from a feudal system towards liberal democratization. Of course, this is all totally unfounded and uninformed speculation on my part. Tangentially related articles on culture, trust, and governance: Shanggyangg on the reasons behind the lack of interest among Russians for democratic reform; David Pinsof on how democracy is underpinned by group psychology; Dan Williams on the tendency towards tribalistic interpretations of truth within human society; David Oks with more negative side effects of cultures where kinship ties are especially strong.
Which I once again did not attend despite being interested in doing so. Next year, maybe.

