<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>argmin gravitas</title><description/><link>https://www.gleech.org/</link><atom:link href="https://www.gleech.org/feed.xml" rel="self" type="application/rss+xml"/><pubDate>Mon, 06 Jul 2026 12:35:53 +0000</pubDate><lastBuildDate>Mon, 06 Jul 2026 12:35:53 +0000</lastBuildDate><generator>Jekyll v4.3.4</generator><item><title>Transnormalism</title><description>&lt;blockquote&gt;
&lt;p&gt;Society is unlikely to fall suddenly under the spell of the transhumanist worldview. But it is very possible that we will nibble at biotechnology’s tempting offerings without realizing that they come at a frightful moral cost… Modifying any one of our key characteristics inevitably entails modifying a complex, interlinked package of traits, and we will never be able to anticipate the ultimate outcome… we may unwittingly invite the transhumanists to deface humanity with their genetic bulldozers and psychotropic shopping malls.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;center&gt;
— &lt;a href="https://philosophy.as.uky.edu/sites/default/files/Transhumanism%20-%20Francis%20Fukuyama.pdf"&gt;Fukuyama&lt;/a&gt; (2004)
&lt;/center&gt;
&lt;!-- Repugnance... revolts against the excesses of human willfulness, warning us not to transgress what is unspeakably profound. Indeed, in this age in which everything is held to be permissible so long as it is freely done... in which our bodies are regarded as mere instruments of our autonomous rational wills, repugnance may be the only voice left that speaks up to defend the central core of our humanity.
- Leon Kass (2003) --&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;the great artisan… made man a creature of indeterminate nature, and… said to him ‘Adam, we give you no fixed place to live, no form forever peculiar to you, no function that is yours alone. According to your desires and judgment, you will have and possess whatever place to live, whatever form, and whatever functions you yourself choose… To you is granted the power of degrading yourself into the lower forms of life, the beasts, and to you is granted the power, contained in your intellect and judgment, to be reborn into the higher forms, the divine.’… To us it was given to be whatever we choose to be, and so that is what we want.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;center&gt;
— &lt;a href="http://www.historymuse.net/readings/orationdignityman.html"&gt;Pico della Mirandola&lt;/a&gt; (1486)
&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;Around 2003 there was a big, rancorous &lt;a href="https://web.archive.org/web/20080613192720/http://www.bioethics.gov/reports/beyondtherapy/chapter1.html"&gt;human enhancement&lt;/a&gt; &lt;a href="https://ora.ox.ac.uk/objects/uuid:85de7a60-20f0-490e-aa1b-9b19af5c3fa1/files/mc173c524ab0fec86a111cec10ef7e8a8"&gt;debate&lt;/a&gt; &lt;a href="#fn:3" id="fnref:3"&gt;3&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The reigning “bioconservatives” predicted that new biotech would lead to moral or political catastrophe, and the loss of human dignity, and maybe wouldn’t even boost welfare; the “transhumanists” &lt;a href="#fn:2" id="fnref:2"&gt;2&lt;/a&gt; argued that nuh uh. &lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;Transhumanism?&lt;/h3&gt;
&lt;div&gt;
&lt;a href="https://www.humanityplus.org/the-transhumanist-declaration"&gt;The 1998 declaration:&lt;/a&gt;
&lt;blockquote&gt;1. ...broadening human potential by overcoming aging, cognitive shortcomings, involuntary suffering, and our confinement to planet Earth.&lt;br /&gt;
2. humanity’s potential is still mostly unrealized.&lt;br /&gt;
3. humanity faces serious risks, especially from the misuse of new technologies. There are possible realistic scenarios that lead to the loss of most, or even all, of what we hold valuable... not all change is progress.&lt;br /&gt;
4. Research effort needs to be invested into understanding these prospects.&lt;br /&gt;
5. Reduction of existential risks, and development of means for the preservation of life and health, the alleviation of grave suffering, and the improvement of human foresight and wisdom should be pursued as urgent priorities&lt;br /&gt;
6. Policy making ought to be guided by... respecting autonomy and individual rights, and showing solidarity with and concern for the interests and dignity of all people&lt;br /&gt;
7. We advocate the well-being of... humans, non-human animals, and any future artificial intellects, modified life forms, or other intelligences&lt;br /&gt;
8. We favour allowing individuals wide personal choice over how they enable their lives ["morphological freedom"]
&lt;/blockquote&gt;&lt;br /&gt;&lt;br /&gt;
Clearly this is a tame and cuddly ideology compared to stuff like posthumanism and accelerationism, but for some reason the critics settled on attacking all biotech-curious ideologies under the name "transhumanism".
&lt;/div&gt;
&lt;h3&gt;Dickey–Wicker&lt;/h3&gt;
&lt;div&gt;
The big policy move from the debate, banning federal funding for embryonic stem cell research, was actually &lt;a href="https://en.wikipedia.org/wiki/Dickey%E2%80%93Wicker_Amendment"&gt;passed in 1996&lt;/a&gt; under Clinton. Bush's &lt;a href="https://georgewbush-whitehouse.archives.gov/news/releases/2001/08/20010809-1.html"&gt;2001 executive order&lt;/a&gt; further restricted NIH funding to pre-existing cell lines, but was rescinded by Obama &lt;a href="https://www.google.com/search?q=Executive+Order+13505&amp;amp;oq=Executive+Order+13505&amp;amp;gs_lcrp=EgZjaHJvbWUyBggAEEUYOdIBBzIzN2owajeoAgCwAgA&amp;amp;client=ubuntu-chr&amp;amp;sourceid=chrome&amp;amp;ie=UTF-8"&gt;in 2009&lt;/a&gt;.&lt;br /&gt;&lt;br /&gt;
We routed around some of the damage: in 2006, the invention of &lt;a href="https://en.wikipedia.org/wiki/Induced_pluripotent_stem_cell"&gt;induced pluripotency&lt;/a&gt; allowed for the creation of (second-rate) stem cells without touching embryos.
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;Some of the technologies they fought over (cloning, germline editing) haven’t happened at scale yet. But others (GLPs, hormones, hair tech, embryo selection) did, and were rapidly adopted by millions of people who had no interest in either ideology.&lt;/p&gt;
&lt;p&gt;And so: a new era of mass chemical enhancement and healthy &lt;a href="https://en.wikipedia.org/wiki/Polypharmacy"&gt;polypharmacy&lt;/a&gt; &lt;a href="#fn:1" id="fnref:1"&gt;1&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;style&gt;
.imgContainer{
float:left;
}
&lt;/style&gt;
&lt;div class="imgContainer"&gt;
&lt;img width="49%" style="border: 0px" src="/img/ssri.jpg" /&gt;
&lt;img width="50%" style="border: 0px" src="/img/stims.jpg" /&gt;
&lt;/div&gt;
&lt;p&gt;(&lt;a href="https://aspe.hhs.gov/sites/default/files/documents/1ef68c455fa5aa5932acf481b0954ddf/DataPoint_PsychRxPrev_BHDAP_20250409%20July%2031%202025.pdf"&gt;link&lt;/a&gt;, &lt;a href="https://www.deadiversion.usdoj.gov/pubs/docs/IQVIA-Report-on-Stimulant-Trends-2024.pdf"&gt;link&lt;/a&gt;)&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;div class="imgContainer"&gt;
&lt;img width="42%" style="border: 0px" src="/img/Use-of-Weight-Loss-Injectables-More-Than-Doubles-in-Under-Two-Years.png" /&gt;
&lt;img width="57%" style="border: 0px" src="/img/testo.jpg" /&gt;
&lt;/div&gt;
&lt;p&gt;(&lt;a href="https://news.gallup.com/poll/696599/obesity-rate-declining.aspx"&gt;link&lt;/a&gt;, &lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC11355536/"&gt;link&lt;/a&gt;)&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;div class="imgContainer"&gt;
&lt;center&gt;
&lt;img width="65%" style="border: 0px" src="/img/roids.jpg" /&gt;&lt;/center&gt;
&lt;/div&gt;
&lt;p&gt;(&lt;a href="https://pubmed.ncbi.nlm.nih.gov/24582699/"&gt;link&lt;/a&gt;)&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;div class="imgContainer"&gt;
&lt;center&gt;&lt;img width="59%" style="border: 0px" src="/img/hrt_usage_timeseries.png" /&gt;
&lt;img width="40%" style="border: 0px" src="/img/twengetrans.png" /&gt;&lt;/center&gt;
&lt;/div&gt;
&lt;p&gt;(&lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC11437377/"&gt;link&lt;/a&gt; - a huge decline but still 1% of US pop! And trans is up to &lt;a href="https://www.generationtechblog.com/p/transgender-identity-how-much-has"&gt;another 1%&lt;/a&gt;)&lt;/p&gt;
&lt;!-- ([link](https://clincalc.com/DrugStats/Drugs/Tretinoin), link) --&gt;
&lt;!-- &lt;a href="https://trends.google.com/trends/explore?date=all&amp;q=%2Fm%2F07_71,%2Fg%2F11dyzd5snl,%2Fm%2F02_ggb,%2Fm%2F027hm4,%2Fm%2F07m9q&amp;hl=en"&gt;
&lt;img src="/img/juicing.jpg" /&gt;
&lt;/a&gt;
&lt;center&gt;
&lt;small&gt;(&lt;a href="https://trends.google.com/trends/explore?date=all&amp;q=%2Fm%2F07_71,%2Fg%2F11dyzd5snl,%2Fm%2F02_ggb,%2Fm%2F027hm4,%2Fm%2F07m9q&amp;hl=en"&gt;link&lt;/a&gt;)&lt;/small&gt;
&lt;/center&gt;
--&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;!-- &lt;center&gt;
&lt;img src="img/Obesity-Showing-Signs-of-Decline-in-U.S.png" /&gt;
&lt;/center&gt;
--&gt;
&lt;p&gt;As so often, Fukuyama (quoted above) looks wrong but is not wrong: society indeed did &lt;em&gt;not&lt;/em&gt; become transhumanist - that is, not in belief. Few of the biotech users endorse the philosophy of technological transcendence. But society is heading there in deed. Conservative forces (religion, disgust, precaution) were in this case grossly outgunned by the force of sheer desire.&lt;/p&gt;
&lt;p&gt;We got, not transhumanism (as deliberate, informed, rational decision to self-consciously go beyond natural human capacity), but surreptitious transhuman &lt;em&gt;behaviour&lt;/em&gt;, without the weird philosophy or the new aesthetics. Technology by default, without conscious ideology. Playing god - but using these new, unfathomable powers to… become more normal. So call it &lt;em&gt;transnormalism&lt;/em&gt;. &lt;a href="#fn:5" id="fnref:5"&gt;5&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;What of the bioconservative prediction of a reckoning for civilisation? So far none arrived. Either &lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;the current techs are not powerful enough yet; or&lt;/li&gt;
&lt;li&gt;the corrosive effects (on say fairness, authenticity, self-concept) are lagged or hard to measure; or&lt;/li&gt;
&lt;li&gt;mass enhancement is just fine.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="ancient-enhancement"&gt;Ancient enhancement&lt;/h2&gt;
&lt;p&gt;A weak reason to think it’s fine is that we’ve been doing it for all of history and prehistory. Enhancement is &lt;a href="https://en.wikipedia.org/wiki/Drunken_monkey_hypothesis"&gt;older&lt;/a&gt; &lt;a href="https://en.wikipedia.org/wiki/Zoopharmacognosy"&gt;than humanity&lt;/a&gt;. What’s new is just the size of the enhancements and the biological, internal, invisible nature of the modifications. &lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;center&gt;
Classic:&lt;br /&gt;
&lt;img width="50%" src="/img/oldtech.jpg" /&gt;
&lt;br /&gt;&lt;br /&gt;
New:&lt;br /&gt;
&lt;img width="50%" src="/img/normtrans.png" /&gt;
&lt;/center&gt;
&lt;style type="text/css"&gt;
.tg {border-collapse:collapse;border-spacing:0;}
.tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;
overflow:hidden;padding:10px 5px;word-break:normal;}
.tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;
font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;}
.tg .tg-fymr{border-color:inherit;font-weight:bold;text-align:left;vertical-align:top}
.tg .tg-0pky{border-color:inherit;text-align:left;vertical-align:top}
.tg .tg-0lax{text-align:left;vertical-align:top}
&lt;/style&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;center&gt;
&lt;table class="tg"&gt;&lt;thead&gt;
&lt;tr&gt;
&lt;th class="tg-fymr"&gt;Old School&lt;/th&gt;
&lt;th class="tg-fymr"&gt;New School&lt;/th&gt;
&lt;/tr&gt;&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td class="tg-0pky"&gt;Alcohol (prehuman)&lt;/td&gt;
&lt;td class="tg-0pky"&gt;Antidepressants&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0pky"&gt;Caffeine (ancient)&lt;/td&gt;
&lt;td class="tg-0pky"&gt;Adderall (1996)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0pky"&gt;Dexedrine (1937)&lt;/td&gt;
&lt;td class="tg-0pky"&gt;Adderall (1996)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0pky"&gt;Glasses (1300)&lt;/td&gt;
&lt;td class="tg-0pky"&gt;Intraocular lens (1999)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0pky"&gt;Makeup (ancient)&lt;/td&gt;
&lt;td class="tg-0pky"&gt;Tretinoin (1971), botox (1989), etc&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0pky"&gt;Fluoride (1945)&lt;/td&gt;
&lt;td class="tg-0pky"&gt;Hydroxyapatite (1980)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0lax"&gt;Aspirin (1899)&lt;/td&gt;
&lt;td class="tg-0lax"&gt;COX-2 inhibitors (1998)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0pky"&gt;Antacids (1852)&lt;/td&gt;
&lt;td class="tg-0pky"&gt;Proton pump inhibitors (1989)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0lax"&gt;Statins (1989)&lt;/td&gt;
&lt;td class="tg-0lax"&gt;PCSK9 inhibitors (2015)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0lax"&gt;Dental fillings (ancient)&lt;/td&gt;
&lt;td class="tg-0lax"&gt;Osseointegrated (c. 1970s)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0lax"&gt;Anabolics (CDMT / Turinabol, 1968)&lt;/td&gt;
&lt;td class="tg-0lax"&gt;HGH (1985), rhEPO (1993), cardarine (2001), ACP-105 (2009), bimagrumab (2013)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0pky"&gt;The Pill (1960)&lt;/td&gt;
&lt;td class="tg-0pky"&gt;IUDs (2000), subdermals (2006)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0lax"&gt;Sunscreen (1940)&lt;/td&gt;
&lt;td class="tg-0lax"&gt;nanoparticle (2000s)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0lax"&gt;Alarm clocks (c. 1904)&lt;/td&gt;
&lt;td class="tg-0lax"&gt;light alarms (2010s)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0lax"&gt;Hearing aids (c. 1800)&lt;/td&gt;
&lt;td class="tg-0lax"&gt;Cochlear implants (1984)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td class="tg-0lax"&gt;Inactivated vaccines (1796)&lt;/td&gt;
&lt;td class="tg-0lax"&gt;mRNA (2020)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;/center&gt;
&lt;!-- Alcohol
Drunken ape hypothesis
Bernard Stiegler's "originary prostheticity"
### The extended body
Exosomatic elements are tools and other instruments used by man to produce, exchange and consume energy in some form.
Cooking could be viewed as an external stomach.
natural-born cyborgs --&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;Other tech&lt;/h3&gt;
&lt;div&gt;
I've been pretty focussed on chemical and biochemical enhancement in the above. There's a lot more:
&lt;!-- --&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;Genetic&lt;/h3&gt;
&lt;div&gt;
Not prevalent yet.
&lt;/div&gt;
&lt;h3&gt;Surgery&lt;/h3&gt;
&lt;div&gt;
&lt;a href="https://www.isaps.org/discover/about-isaps/global-statistics/global-survey-2024-full-report-and-press-releases/"&gt;Only&lt;/a&gt; 38 million cosmetic surgeries a year? Surprising!&lt;br /&gt;&lt;br /&gt;
It's hard to say what fraction of the &lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC2871217/"&gt;310 million&lt;/a&gt; major surgeries are "enhancers" but not many.
&lt;/div&gt;
&lt;h3&gt;Electronics&lt;/h3&gt;
&lt;div&gt;
The above omits a very early and widespread kind: electronic aids and implants.&lt;br /&gt;&lt;br /&gt;
1936: first wearable hearing aid. Now 300 or 400 million&lt;br /&gt;
1958: internal pacemaker. Now around 30? million people.&lt;br /&gt;
1977: cochlear implants. Maybe 2 million.&lt;br /&gt;&lt;br /&gt;
Nonmedical use isn't mainstream yet. The &lt;a href="https://en.wikipedia.org/wiki/Body_hacking"&gt;grinders&lt;/a&gt; (people who do DIY surgery to implant electronics for nonmedical use) are roughly as strange as they were 15 years ago.
&lt;/div&gt;
&lt;!-- &lt;h3&gt;Hitler and Kennedy&lt;/h3&gt;
&lt;div&gt;&lt;/div&gt; --&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="normal-exceptionalism-the-escape-from-morphological-freedom"&gt;Normal exceptionalism: the escape from morphological freedom&lt;/h2&gt;
&lt;p&gt;With the exception of the bodybuilding and trans communities &lt;a href="#fn:6" id="fnref:6"&gt;6&lt;/a&gt;, we don’t see much &lt;a href="https://contraptions.venkateshrao.com/p/into-the-weirding-part-1"&gt;weirding&lt;/a&gt; at mass scale. We aren’t expressing our morphological freedom to look more different. It seems to me that the result of power over our appearance is not deviance and weirdness but heightened normative normalcy. Bigger biceps, fewer wrinkles, and nerds &lt;a href="https://www.palladiummag.com/2019/01/01/competitive-hormone-supplementation-is-shaping-americas-future-business-titans/"&gt;suddenly&lt;/a&gt; getting &lt;a href="https://www.youtube.com/watch?v=HjFRaPXsxQs"&gt;normative&lt;/a&gt; gender presentation. Not many &lt;a href="https://en.wikipedia.org/wiki/Body_hacking"&gt;cyborgs&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I’m not fit to do any comparative analysis of the history of cosmesis and beauty standards. &lt;a href="https://www.cartoonshateher.com/p/why-arent-men-the-pretty-ones"&gt;This essay&lt;/a&gt; is great. But it seems to me that the &lt;a href="https://simple.wikipedia.org/wiki/Kiki_H%C3%A5kansson"&gt;beauty&lt;/a&gt; queens and movie stars of the 1960s &lt;a href="https://www.newsweek.com/ai-cosmetic-survery-filters-beauty-standards-changing-2085079"&gt;no longer&lt;/a&gt; look as exceptional as they once did, because cosmetic technology (and image editing ig) has shifted the (top decile of the) distribution upwards so much. I can’t say what all of this is tending towards. What is the intensified platonic ideal of a normal dude?&lt;/p&gt;
&lt;p&gt;It’s pretty obvious why tech which gives you options is used, on average, to normalise yourself: most people want to be normal, and the user population is so large now that it simply must include a lot of such people. In 2010 enhancement was a matter for &lt;a href="https://gwern.net/modafinil"&gt;nerds&lt;/a&gt;, &lt;a href="https://en.wikipedia.org/wiki/Body_hacking"&gt;hackers&lt;/a&gt;, and obsessive hobbyists like bodybuilders. But now it’s a much bigger coalition (e.g. &lt;a href="https://www.kff.org/health-costs/kff-health-tracking-poll-may-2024-the-publics-use-and-views-of-glp-1-drugs/#4acecddb-cd6c-4154-9c82-75d8da3e1234--h-key-findings"&gt;6-12% of Americans&lt;/a&gt; &lt;a href="https://news.gallup.com/poll/696599/obesity-rate-declining.aspx"&gt;on GLP agonists&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;center&gt;
&lt;img width="49.7%" src="/img/transh.jpg" /&gt;
&lt;img width="49.7%" src="/img/housewives.jpg" /&gt;
&lt;/center&gt;
&lt;!--
## Medicine vs enhancements
One problem for the bioconservatives is that there just is no clean distinction between medicine (which everyone likes) and enhancement. Some attempts to make one:
* Medicine vs nonmedicine. Lots of foods have medicine-grade effect sizes though.
* Natural vs artificial. We coevolved with alcohol, which is supposed to mean it's sustainable.
* Does the intervention address a deficit? (Are you under the population average in some variable?)
* Is it in your body or an external tool?
Malleability vs fixed nature --&gt;
&lt;h2 id="medicalisation-and-demedicalisation"&gt;Medicalisation and demedicalisation&lt;/h2&gt;
&lt;p&gt;You don’t need any new or unpopular premises at all to justify enhancement. Transnormalism is what you get from liberalism plus medicalisation. Consumer enhancement has been happening in the form of individual medical decision-making: distributed, highly private (with some actual infosec), and invisible except by us belatedly noticing the aggregate properties of the species changing.&lt;/p&gt;
&lt;p&gt;Medicine has been expanding for centuries, especially in the last two decades. This is true in volume (spending) and in domains (new treatments, new powers, new areas brought within the ambit).&lt;/p&gt;
&lt;p&gt;&lt;img src="/img/nhe.png" /&gt;&lt;/p&gt;
&lt;p&gt;In absolute terms (multiplying &lt;a href="https://ourworldindata.org/grapher/total-healthcare-expenditure-gdp?tab=table&amp;amp;tableSearch=world"&gt;healthcare share&lt;/a&gt; by &lt;a href="https://data.worldbank.org/indicator/NY.GDP.MKTP.CD"&gt;GWP&lt;/a&gt;) we spent $1.85tn in 2000, something north of $7.1tn in 2022 (and more now).&lt;/p&gt;
&lt;p&gt;But after being incubated in medicine, enhancement is being taken off the doctors. The thriving grey market (intense cosmetics, “aesthetics” and “&lt;a href="https://www.theguardian.com/tv-and-radio/2025/jun/09/john-oliver-med-spas"&gt;med spas&lt;/a&gt;” and nootropics) and black market (study drugs, research chemicals, peptides, bootleg hormones) are of unknown size but growing insanely fast.&lt;/p&gt;
&lt;!-- And then as part of the revolt of the public it was taken off the doctors. --&gt;
&lt;!-- ## The timeline
For sanity and space let's put eugenics and genetic intervention out of scope here. So, chemical and surgical enhancement:
Hair
Minox (1988)
https://themultiplicity.ai/room/c5861824-f970-4003-ba4f-f365acbd9573
Shape and cosmesis
Surgery
Roids https://www.sciencedirect.com/science/article/abs/pii/S1047279714000398
1968: CDMT / Turinabol
Tren
Ozempic
Gender
See shape
Fertility
Erections
Viagra 1998
Cognition and volition
Caffeine
Nicotine
Ritalin/Adderall
Modafinil
T
Sleep
Ambien
Melatonin
Memory (subtraction)
Longevity
Mood
Prozac
--&gt;
&lt;!-- Peptides --&gt;
&lt;h2 id="you-are-like-a-little-baby"&gt;You are like a little baby&lt;/h2&gt;
&lt;p&gt;The above technologies are really fairly weak. Retatrutide (2023) is twice as strong as semaglutide (2014), which is twice as strong as liraglutide (&lt;a href="https://pubmed.ncbi.nlm.nih.gov/11935150/"&gt;2002&lt;/a&gt;). At some point someone will work out &lt;a href="https://en.wikipedia.org/wiki/Exercise_mimetic"&gt;how to&lt;/a&gt; &lt;a href="https://www.medscape.com/viewarticle/myostatin-blocker-preserves-muscle-glp-1-treatment-2025a1000qs4"&gt;chemically simulate&lt;/a&gt; the effect of working out. The nootropics industry is overall a pathetic failure, capped with blunt instruments like &lt;a href="https://en.wikipedia.org/wiki/Modafinil"&gt;not sleeping&lt;/a&gt; or &lt;a href="https://pubmed.ncbi.nlm.nih.gov/14871155/"&gt;flooding&lt;/a&gt; the brain with catecholamines. Psychopharmaceuticals are better but not by much and don’t manage sustainbly-better-than-well. We are admittedly &lt;a href="https://pubmed.ncbi.nlm.nih.gov/36229224/"&gt;really good&lt;/a&gt; at &lt;a href="https://en.wikipedia.org/wiki/Selective_androgen_receptor_modulator#Non-medical_use"&gt;things&lt;/a&gt; which let people sprint for 6% longer, though at the expense of giving them &lt;a href="https://en.wikipedia.org/wiki/GW501516"&gt;cancer&lt;/a&gt;. We do &lt;a href="https://en.wikipedia.org/wiki/Transcranial_magnetic_stimulation"&gt;nearly&lt;/a&gt; nothing directly to brains. We have &lt;a href="https://pubmed.ncbi.nlm.nih.gov/28051768/"&gt;basically&lt;/a&gt; nothing for memory enhancement. At the moment we do little with &lt;a href="https://www.pnas.org/doi/pdf/10.1073/pnas.2416042122"&gt;genes&lt;/a&gt;, but the rich and unsqueamish are beginning to. All humans are &lt;a href="https://herfingersbloomed.substack.com/i/178888011/all-babies-are-premature"&gt;born premature&lt;/a&gt;. &lt;a href="https://www.isaak.net/sleepless/"&gt;One might solve sleep&lt;/a&gt;. &lt;a href="https://longevity.vc"&gt;One &lt;em&gt;might&lt;/em&gt; solve death&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The conservative concerns might apply to a more-mature science of More. Thanks to transnormalism funding it all we will soon see.&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We find sheer humanism to be unsatisfying. It shuts the windows, draws the blinds, and seeks artificial elegance - oblivious of the outer night and the stars. Instead, we boldly go out into darkness and find the superhuman everywhere… the story of evolution of life: its length, its wastefulness, its precariousness, its chanciness, its progressive release of potentiality, its incomprehensibility and ourselves as moments within it… man is transitional and scarcely a beginning.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;center&gt;— Olaf Stapledon (1934)&lt;/center&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;Actual data&lt;/h3&gt;
&lt;div&gt;
&lt;a href="https://nablatheta.substack.com/p/my-hobby-running-deranged-surveys"&gt;Leo Gao&lt;/a&gt; has been running informal n=200 surveys of random Americans. He finds, as I assumed above, that most people are indeed (incoherently) nontranshumanist, with the huge exception of immortality:
&lt;blockquote&gt;
&amp;gt; "If you had the option to live forever in perfect health and youth, would you choose to? (Assume you could still change your mind at any time if you ever got bored of it.)” ... 66% of respondents said Yes, with 14% saying No, and 20% saying “Not sure”. As a follow up, it turns out roughly a third of Americans think developing the technology to enable life extension should be a top priority... Since it seemed like overpopulation and inequality were the main things people were worried about, I also asked a version of the question where I stipulated that these things were solved. Surprisingly, this barely shifts people’s opinions, and we get almost exactly the same response! My guess is this is a sign that the real objection is more about the vibes than any specific issue.
&lt;br /&gt;&lt;br /&gt;
&amp;gt; despite being very pro living forever, Americans are much more skeptical of cryonics — even if they could be revived a few decades after their death to live forever thereafter, only 27% are in favor of being preserved, and 46% are opposed (the rest are unsure).
&lt;br /&gt;&lt;br /&gt;
&amp;gt; Space colonization also has pretty lukewarm support, coming in at 37% in favor and 16% opposed
&lt;br /&gt;&lt;br /&gt;
&amp;gt; and cognitive enhancement for all is only a little bit more popular (42% in favor, 19% opposed).
&lt;br /&gt;&lt;br /&gt;
&amp;gt; Also, for some reason, people are really opposed to a hypothetical cheap, painless, and safe arbitrary modification of physical appearance (only 23% in favor, with 37% opposed!). In retrospect, the backlash against Ozempic is a sign, but I was still quite surprised.
&lt;br /&gt;&lt;br /&gt;
&amp;gt; Terraforming other planets so that humans can live on them is also pretty unpopular, coming in at 37% in favor and 16% opposed. Thankfully, for most of these questions, a huge chunk of people are still undecided.
&lt;br /&gt;&lt;br /&gt;
&amp;gt; only 51% of Americans are in favor of literal post-scarcity (complete freedom to work on anything you want, as much as you want, and still enjoy a high quality of life), with 25% opposing. I was so shocked by this result not being 80%+ in favor that I reran a variant of this question with different wording. My original question asked whether the world would be better or worse if everyone had the freedom to work on whatever they want, as long as they want, and still enjoy a high quality of life, and anything we don’t want to do is done for us by robots. I thought maybe that set off some “AI taking jobs bad” instincts; for the new question I took pains to clarify that the stuff is literally conjured out of nowhere with magic and is not taken from anyone else, and got an even worse result (38% support, 34% oppose). This is even more crazy, so I ran a third version on the hypothesis that people don’t like magic, or that not having to work sounded too crazy. This version asked whether it would be good if everyone made 10x more (inflation-adjusted) than they do currently. This polled only somewhat better, with 39% in favor and 19% opposing. I’m still pretty confused what conclusion to draw from this; this is probably worth digging more into.
&lt;br /&gt;&lt;br /&gt;
&amp;gt; Only 14% think that society is currently trending in a positive direction.
&amp;lt;/div&amp;gt;
&amp;lt;/div&amp;gt;
##
&lt;br /&gt;
## See also
* [https://www.gleech.org/med](https://www.gleech.org/med)
* [https://vectorculture.substack.com/p/not-for-human-consumption](https://vectorculture.substack.com/p/not-for-human-consumption)
* [https://bengoldhaber.substack.com/p/no-real-nattys](https://bengoldhaber.substack.com/p/no-real-nattys)
* [https://christianangermayer.substack.com/p/the-future-is-enhanced](https://christianangermayer.substack.com/p/the-future-is-enhanced)
* [https://web.archive.org/web/20250202101559/https://humanenhancementdrugs.com/](https://web.archive.org/web/20250202101559/https://humanenhancementdrugs.com/)
* [https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1467-8519.2005.00437.x](https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1467-8519.2005.00437.x)
* [https://web.archive.org/web/20210202180452/https://thingsvarious.medium.com/hormone-replacement-therapy-the-only-guide-you-need-2904aa48b7bd](https://web.archive.org/web/20210202180452/https://thingsvarious.medium.com/hormone-replacement-therapy-the-only-guide-you-need-2904aa48b7bd)
* [http://bactra.org/Medawar/technology-and-evolution/](http://bactra.org/Medawar/technology-and-evolution/)
* [https://www.gleech.org/med](https://www.gleech.org/med)
* [https://vectorculture.substack.com/p/not-for-human-consumption](https://vectorculture.substack.com/p/not-for-human-consumption)
* [https://bengoldhaber.substack.com/p/no-real-nattys](https://bengoldhaber.substack.com/p/no-real-nattys)
* [https://web.archive.org/web/20080613192720/http://www.bioethics.gov/reports/beyondtherapy/chapter1.html](https://web.archive.org/web/20080613192720/http://www.bioethics.gov/reports/beyondtherapy/chapter1.html)
* [https://nickbostrom.com/papers/a-history-of-transhumanist-thought/](https://nickbostrom.com/papers/a-history-of-transhumanist-thought/)
* [https://global.oup.com/academic/product/natural-born-cyborgs-9780195177510](https://global.oup.com/academic/product/natural-born-cyborgs-9780195177510)
* [https://branko2f7.substack.com/p/dr-morell-and-the-patient-a](https://branko2f7.substack.com/p/dr-morell-and-the-patient-a)
* [https://doctorzebra.com/prez/z_x35testosterone_g.htm](https://doctorzebra.com/prez/z_x35testosterone_g.htm)
&lt;div class="footnotes"&gt;
&lt;ol&gt;
&lt;!-- 1 --&gt;
&lt;li class="footnote" id="fn:1"&gt;I'm using an idiosyncratic definition of "mass use": &amp;gt;1% of Americans. But that's a leading indicator for the rest of the world following in the end.&lt;/li&gt;
&lt;li class="footnote" id="fn:2"&gt;A more inclusive term might be "bioprogressives".&lt;/li&gt;
&lt;li class="footnote" id="fn:3"&gt;with one side &lt;a href="https://en.wikipedia.org/wiki/President%27s_Council_on_Bioethics"&gt;backed&lt;/a&gt; by an openly religious executive branch. &lt;/li&gt;
&lt;!-- &lt;li class="footnote" id="fn:4"&gt;But the transhumanists were probably being normative rather than deluded about this.&lt;/li&gt; --&gt;
&lt;li class="footnote" id="fn:5"&gt;The social sciences are watching quite closely, but they mostly don't connect any of it to the philosophical project, nor do they project forwards to the coming technologies or preference cascades. They speak narrowly and worry. Their categories are valid and useful as far as they go ("lifestyle drugs", "Image and Performance Enhancing Drugs") but are missing the future, the telos, the limit.&lt;/li&gt;
&lt;li class="footnote" id="fn:6"&gt;Though one could borrow an emic distinction from trans: that between "dolls" (a trans woman who aims to perfectly converge on normative femininity) and "bricks" (who are not converging, maybe not trying to) and note that a doll who doesn't go too hard is also a transnormalist!&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
&lt;/blockquote&gt;&lt;/div&gt;&lt;/div&gt;</description><pubDate>Fri, 20 Feb 2026 00:00:00 +0000</pubDate><link>https://www.gleech.org/enhance</link><guid isPermaLink="true">https://www.gleech.org/enhance</guid><category>transhumanism,</category><category>biology,</category><category>philosophy,</category><category>ethics,</category><category>scifi</category></item><item><title>AI in 2025: gestalt</title><description>&lt;center&gt;&lt;img width="60%" src="/img/fullsig.jpg" /&gt;&lt;/center&gt;
&lt;p&gt;This is the editorial for this year’s “&lt;a href="https://shallowreview.ai/"&gt;Shallow Review of AI Safety&lt;/a&gt;”. (It got long enough to stand alone.)&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Epistemic status: subjective impressions plus one new graph plus 300 links.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Huge thanks to Jaeho Lee, Jaime Sevilla, and Lexin Zhou for running lots of tests pro bono and so greatly improving the main analysis.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h2 id="tldr"&gt;tl;dr&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Informed people &lt;a href="https://www.lesswrong.com/posts/5tqFT3bcTekvico4d/do-confident-short-timelines-make-sense"&gt;disagree&lt;/a&gt; about the prospects for LLM AGI – or even just what exactly was achieved this year. But the famous ones with a book to talk at least agree that we’re &lt;a href="https://nitter.net/polynoamial/status/1994439121243169176"&gt;2-20&lt;/a&gt; years off (allowing for other paradigms arising). In this piece I stick to arguments rather than reporting who thinks what.&lt;/li&gt;
&lt;li&gt;My view: compared to last year, AI is much more impressive but not proportionally more useful. They improved on some things they were explicitly optimised for (coding, vision, OCR, benchmarks), and did not &lt;em&gt;hugely&lt;/em&gt; improve on everything else. Progress is thus (still!) consistent with current frontier training bringing more things in-distribution rather than generalising very far.&lt;/li&gt;
&lt;li&gt;Pretraining (GPT-4.5, Grok 4, but also counterfactual large runs which weren’t done) disappointed people this year. It’s probably not because it wouldn’t work; it was just ~30 times more efficient to do post-training instead, &lt;em&gt;on the margin&lt;/em&gt;. This should change, yet again, soon, if RL scales even worse.&lt;/li&gt;
&lt;li&gt;EDIT: See &lt;a href="https://www.lesswrong.com/posts/Q9ewXs8pQSAX5vL7H/ai-in-2025-gestalt?commentId=PEiZF3D3PZttPRWzt"&gt;this&lt;/a&gt; amazing comment for the hardware reasons behind this, and reasons to think that pretraining will struggle for years.&lt;/li&gt;
&lt;li&gt;True frontier capabilities are likely obscured by systematic cost-cutting (distillation for serving to consumers, quantization, low reasoning-token modes, routing to cheap models, etc) and a few unreleased models/modes.&lt;/li&gt;
&lt;li&gt;Most benchmarks are weak predictors of even the rank order of models’ capabilities. I distrust &lt;a href="https://epoch.ai/benchmarks/eci"&gt;ECI&lt;/a&gt;, &lt;a href="https://arxiv.org/abs/2503.06378"&gt;ADeLe&lt;/a&gt;, and &lt;a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/"&gt;HCAST&lt;/a&gt; the least (see graph below or &lt;a href="https://colab.research.google.com/drive/1HtVWPh9thdMV58zfdBcky7n7DVy5AHni?usp=sharing"&gt;this notebook&lt;/a&gt;). ECI shows a linear improvement, HCAST finds an exponential improvement on greenfield software engineering, and ADeLe shows a previous super-exponential slowing down to what &lt;em&gt;might be&lt;/em&gt; linear growth.&lt;/li&gt;
&lt;li&gt;The world’s &lt;a href="https://x.com/livgorton/status/1996329704476041557"&gt;de facto&lt;/a&gt; strategy remains “&lt;a href="https://www.thecompendium.ai/ai-safety#current-technical-efforts-are-not-on-track-to-solve-alignment"&gt;iterative alignment&lt;/a&gt;”, optimising outputs with a stack of alignment and control techniques everyone admits are individually weak.&lt;/li&gt;
&lt;li&gt;Early claims that reasoning models are safer turned out to be a mixed bag (see below).&lt;/li&gt;
&lt;li&gt;We already &lt;a href="https://www.lesswrong.com/posts/f49e7KpZJBwdjWRw2/the-jailbreak-argument-against-llm-values"&gt;knew&lt;/a&gt; from jailbreaks that current alignment methods were brittle. The &lt;a href="https://www.emergent-misalignment.com/"&gt;great safety discovery&lt;/a&gt; of the year is that bad things are correlated in current models. (And on net this is good news.)&lt;/li&gt;
&lt;li&gt;Previously I thought that “character training” was a separate and lesser matter than “alignment training”. Now I am not sure.&lt;/li&gt;
&lt;li&gt;Welcome to the many new people in AI Safety and Security and Assurance and so on. In the &lt;em&gt;&lt;a href="https://shallowreview.ai/"&gt;Shallow Review&lt;/a&gt;&lt;/em&gt; I added a new, sprawling top-level category for one large trend among them, which is to treat the multi-agent lens as primary.&lt;/li&gt;
&lt;li&gt;Overall I wish I could tell you some number, the net expected safety change (this year’s improvements in dangerous capabilities and agent performance, minus the alignment-boosting portion of capabilities, minus the cumulative effect of the best actually implemented composition of alignment and control techniques). But I can’t.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h2 id="capabilities-in-2025"&gt;Capabilities in 2025&lt;/h2&gt;
&lt;p&gt;Better, but how much?&lt;/p&gt;
&lt;p&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q9ewXs8pQSAX5vL7H/bwenyfjyhyr5zdqo6qrb" alt="Fraser riffing off Pueyo" /&gt;&lt;/p&gt;
&lt;center&gt;&amp;mdash; &lt;i&gt;&lt;a href="https://x.com/colin_fraser/status/1994188009608983008"&gt;Fraser&lt;/a&gt;, riffing off &lt;a href="https://x.com/tomaspueyo/status/1993360931267473662"&gt;Pueyo&lt;/a&gt;&lt;/i&gt;&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id="arguments-against-2025-capabilities-growth-being-above-trend"&gt;Arguments against 2025 capabilities growth being above-trend&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Apparent progress is an unknown mixture of real general capability increase, &lt;a href="https://aclanthology.org/2025.emnlp-main.744.pdf"&gt;hidden contamination&lt;/a&gt; increase, benchmaxxing (nailing a small set of static examples instead of generalisation) &lt;a href="https://x.com/teortaxesTex/status/1995466603668885521"&gt;usemaxxing&lt;/a&gt; (nailing a small set of narrow tasks with RL instead of deeper generalisation), and &lt;a href="https://arxiv.org/abs/2407.12220"&gt;human cheating&lt;/a&gt;. It’s reasonable to think it’s 20% each, with low confidence. (With a small but growing contribution from &lt;a href="https://evaluations.metr.org/openai-o3-report/"&gt;AI cheating&lt;/a&gt;.)&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Discrete&lt;/em&gt; capabilities progress &lt;a href="http://gleech.org/ai-24-25#2025"&gt;seems&lt;/a&gt; &lt;a href="https://x.com/RyanPGreenblatt/status/1949912100601811381"&gt;slower&lt;/a&gt; this year than in &lt;a href="http://gleech.org/ai-24-25#2024"&gt;2024&lt;/a&gt; (but 2024 was insanely fast). Kudos to &lt;a href="https://x.com/scaling01/status/1874608907508752546"&gt;this person&lt;/a&gt; for registering predictions and so reminding us what really above-trend would have meant concretely. The excellent forecaster Eli &lt;a href="https://www.foxy-scout.com/my-2025-ai-predictions-and-2024-evaluations-2/"&gt;was also&lt;/a&gt; over-optimistic.&lt;/li&gt;
&lt;li&gt;I don’t recommend taking benchmark trends, or even &lt;a href="https://artificialanalysis.ai/methodology/intelligence-benchmarking"&gt;clever&lt;/a&gt; &lt;a href="https://epoch.ai/benchmarks/eci"&gt;composite indices&lt;/a&gt; of them, or even clever &lt;a href="https://arxiv.org/abs/2510.18212"&gt;cognitive science&lt;/a&gt; measures &lt;a href="https://arxiv.org/abs/2407.12220"&gt;too&lt;/a&gt; &lt;a href="https://aievaluation.substack.com/p/is-the-definition-of-agi-a-percentage"&gt;seriously&lt;/a&gt;. The adversarial pressure on the measures is intense.&lt;/li&gt;
&lt;li&gt;Pretraining didn’t hit a “wall”, but the driver did manoeuvre away from it on encountering an &lt;a href="https://epoch.ai/gradient-updates/quantifying-the-algorithmic-improvement-from-reasoning-models"&gt;easier&lt;/a&gt; detour (&lt;a href="https://magazine.sebastianraschka.com/i/161572341/rl-reward-modeling-from-rlhf-to-rlvr"&gt;RLVR&lt;/a&gt;).
&lt;ul&gt;
&lt;li&gt;Training runs &lt;a href="https://epoch.ai/data/ai-models"&gt;continued&lt;/a&gt; to scale (Llama 3 405B = 4e25, GPT-4.5 ~= 4e26, Grok 4 ~= 3e26) but to &lt;a href="https://www.hfh.pw/AI_diminishing_returns"&gt;less effect&lt;/a&gt;.&lt;sup id="fnref:1"&gt;&lt;a href="#fn:1" class="footnote" rel="footnote" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt; In fact all of these models are dominated by apparently smaller pretraining runs with better post-training.&lt;/li&gt;
&lt;li&gt;4.5 is actually shut down already; in 2025 it wasn’t worth it to serve any 1T active model or make it into a reasoning model. But this is more to do with inference cost and inference hardware constraints than any quality shortfall or breakdown in scaling laws.&lt;/li&gt;
&lt;li&gt;EDIT: Nesov notes that making use of bigger models (i.e. 4T active parameters) is heavily bottlenecked on the HBM on inference chips, as is doing RL on bigger models. He expects it won’t be possible to do the next huge pretraining jump (to ~30T active) until ~2029.&lt;/li&gt;
&lt;li&gt;It would work, probably, if we had the data and HBM and spent the next $10bn, it’s just too expensive to bother with at the moment compared to:&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href="https://magazine.sebastianraschka.com/i/161572341/rl-reward-modeling-from-rlhf-to-rlvr"&gt;RLVR&lt;/a&gt; scaling and &lt;a href="https://arxiv.org/pdf/2510.13786"&gt;inference scaling&lt;/a&gt; (or “reasoning” as we’re calling it), which kept things going instead. This boils down to spending more on RL so the resulting model can productively spend more tokens.
&lt;ul&gt;
&lt;li&gt;But the &lt;a href="https://www.lesswrong.com/posts/BEFbC8sLkur7DGCYB/o1-is-a-bad-idea"&gt;feared&lt;/a&gt; / &lt;a href="https://www.mechanize.work/blog/the-upcoming-gpt-3-moment-for-rl/"&gt;hoped-for&lt;/a&gt; generalisation from {training LLMs with RL on tasks with a verifier} to performing on tasks without one remains unclear even after two years of trying.&lt;sup id="fnref:10"&gt;&lt;a href="#fn:10" class="footnote" rel="footnote" role="doc-noteref"&gt;2&lt;/a&gt;&lt;/sup&gt; Grok 4 was apparently a major test of scaling RLVR training.&lt;sup id="fnref:2"&gt;&lt;a href="#fn:2" class="footnote" rel="footnote" role="doc-noteref"&gt;3&lt;/a&gt;&lt;/sup&gt; It gets excellent benchmark results and the distilled versions &lt;a href="https://openrouter.ai/rankings"&gt;are actually&lt;/a&gt; being used at scale. But imo it is the most jagged of all models.&lt;/li&gt;
&lt;li&gt;This rate of scaling-up &lt;a href="https://www.lesswrong.com/posts/xpj6KhDM9bJybdnEe/how-well-does-rl-scale"&gt;cannot&lt;/a&gt; be sustained: RL is &lt;a href="https://www.tobyord.com/writing/inefficiency-of-reinforcement-learning"&gt;famously&lt;/a&gt; &lt;a href="https://www.dwarkesh.com/p/bits-per-sample"&gt;inefficient&lt;/a&gt;. Compared to SFT, it “reduces the amount of information a model can learn per hour of training by a factor of 1,000 to 1,000,000”. The &lt;a href="https://www.tobyord.com/writing/how-well-does-rl-scale"&gt;per-token intelligence&lt;/a&gt; is up but not by much.&lt;/li&gt;
&lt;li&gt;There is a &lt;a href="https://docs.google.com/presentation/d/18Vh9CHPbZ6pesa1JnyZ_dTIR_l-WAFi0c4kiECw5ROQ/edit?slide=id.g350a9c9be82_0_83#slide=id.g350a9c9be82_0_83"&gt;deflationary theory&lt;/a&gt; of RLVR, that it’s &lt;a href="https://arxiv.org/abs/2510.07364v3"&gt;capped&lt;/a&gt; by pretraining capability and thus just about easier elicitation and better pass@1. But even if that’s right this isn’t saying much!&lt;/li&gt;
&lt;li&gt;RLVR is heavy fiddly R&amp;amp;D you need to learn by doing; better to learn it on smaller models with 10% of the cost.&lt;/li&gt;
&lt;li&gt;An obvious thing we can infer: the labs don’t have the resources to scale both at the same time. To keep the money jet burning, they have to post.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;By late 2025, the obsolete modal “&lt;a href="https://ai-2027.com/"&gt;AI 2027&lt;/a&gt;” scenario described the beginning of a divergence between the lead lab and the runner-up frontier labs.&lt;sup id="fnref:3"&gt;&lt;a href="#fn:3" class="footnote" rel="footnote" role="doc-noteref"&gt;4&lt;/a&gt;&lt;/sup&gt; This is because the leader’s superior ability to generate or acquire new training data and algorithm ideas was supposed to compound and widen their lead. Instead, we see the erstwhile leader OpenAI and some others clustering around the same level, which is weak evidence that synthetic data and AI-AI R&amp;amp;D aren’t there yet. Anthropic are making &lt;a href="https://www.reddit.com/r/singularity/comments/1p7p86q/anthropic_claims_internal_ai_rd_evals_are_near/"&gt;large claims&lt;/a&gt; about Opus 4.5’s capabilities, so &lt;em&gt;maybe&lt;/em&gt; this will arrive on time next year.&lt;/li&gt;
&lt;li&gt;For the first time there are now &lt;a href="https://nitter.net/g_leech_/status/1974165458283860198"&gt;many&lt;/a&gt; examples of LLMs helping with actual research mathematics. But if you &lt;a href="https://nitter.net/g_leech_/status/1991608870444400684"&gt;look closely&lt;/a&gt; it’s all still in-distribution in the broad sense: new implications of existing facts and techniques. (I don’t mean to demean this; probably most mathematics fits this spec.)&lt;/li&gt;
&lt;li&gt;Extremely &lt;a href="https://mashable.com/article/openai-o3-o4-mini-hallucinate-higher-previous-models"&gt;mixed&lt;/a&gt; &lt;a href="https://x.com/ArtificialAnlys/status/1990926803087892506"&gt;evidence&lt;/a&gt; on the trend in the hallucination rate.&lt;/li&gt;
&lt;li&gt;Companies make claims about their one-million- or ten-million-token &lt;em&gt;effective&lt;/em&gt; context windows, &lt;a href="https://arxiv.org/pdf/2410.18745v1"&gt;but&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2307.03172"&gt;I&lt;/a&gt; &lt;a href="https://research.trychroma.com/context-rot"&gt;don’t&lt;/a&gt; &lt;a href="https://nostalgebraist.tumblr.com/post/772798409412427776/even-setting-aside-the-need-to-do"&gt;believe it&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;In lieu of trying the agents for serious work yourself, you could at least look at the &lt;a href="https://theaidigest.org/village/blog/research-robots"&gt;highlights&lt;/a&gt; of the &lt;a href="http://zackmdavis.net/blog/2025/11/the-best-lack-all-conviction-a-confusing-day-in-the-ai-village/"&gt;gullible&lt;/a&gt; and precompetent AIs in the &lt;a href="https://theaidigest.org/village"&gt;AI Village&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q9ewXs8pQSAX5vL7H/x68zoh6ievfv8lwdyhjb" alt="Current limits" /&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Here are the current biggest limits to LLMs, as polled in &lt;a href="https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/68fb86aa2c3b1b7ea6251cc1_Understanding%20AI%20Trajectories%20(24_10%20update).pdf"&gt;Heitmann et al&lt;/a&gt;:&lt;/li&gt;
&lt;/ul&gt;
&lt;center&gt;
&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q9ewXs8pQSAX5vL7H/oo7sjrjun3e5jztggpqd" /&gt;
&lt;/center&gt;
&lt;h3 id="arguments-for-2025-capabilities-growth-being-above-trend"&gt;Arguments for 2025 capabilities growth being above-trend&lt;/h3&gt;
&lt;p&gt;We now have measures which are a bit more like AGI metrics than dumb single-task static benchmarks are. What do they say?&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;em&gt;Difficulty-weighted benchmarks&lt;/em&gt;: &lt;a href="https://epoch.ai/benchmarks/eci"&gt;Epoch Capabilities Index&lt;/a&gt;.
&lt;ul&gt;
&lt;li&gt;Interpretation: GPT-2 to GPT-3 was (very roughly) a 20-40 point jump.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Cognitive abilities&lt;/em&gt;: &lt;a href="https://arxiv.org/abs/2503.06378"&gt;ADeLe&lt;/a&gt;.&lt;sup id="fnref:4"&gt;&lt;a href="#fn:4" class="footnote" rel="footnote" role="doc-noteref"&gt;5&lt;/a&gt;&lt;/sup&gt;
&lt;ul&gt;
&lt;li&gt;Interpretation: level &lt;em&gt;L&lt;/em&gt; is the capability held by 1 in 10^L humans on Earth. GPT-2 to GPT-3 was a 0.6 point jump.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Software agency&lt;/em&gt;: &lt;a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/"&gt;HCAST time horizon&lt;/a&gt;, the ability to handle larger-scale well-specified greenfield software tasks.
&lt;ul&gt;
&lt;li&gt;Interpretation: the absolute values are less important than the implied exponential (a 7 month doubling time).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;So: is the &lt;a href="https://colab.research.google.com/drive/1HtVWPh9thdMV58zfdBcky7n7DVy5AHni?usp=sharing"&gt;rate of change&lt;/a&gt; in 2025 (shaded) holding up compared to past jumps?:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q9ewXs8pQSAX5vL7H/kkutje8qx28fpe4udcnp" alt="ECI and ADeLe graphs" /&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q9ewXs8pQSAX5vL7H/alrb9on9jitivbczjjfx" alt="HCAST graph" /&gt;&lt;/p&gt;
&lt;p&gt;Ignoring the (nonrobust)&lt;sup id="fnref:5"&gt;&lt;a href="#fn:5" class="footnote" rel="footnote" role="doc-noteref"&gt;6&lt;/a&gt;&lt;/sup&gt; ECI GPT-2 rate, we can say yes: 2025 is fast, as fast as ever or more.&lt;/p&gt;
&lt;p&gt;Even though these are the best we have, we can’t defer to these numbers.&lt;sup id="fnref:6"&gt;&lt;a href="#fn:6" class="footnote" rel="footnote" role="doc-noteref"&gt;7&lt;/a&gt;&lt;/sup&gt; What else is there?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;In May they passed some threshold and I finally started using LLMs for actual tasks. For me this is mostly due to the search agents replacing a degraded Google search. I’m &lt;a href="https://www.lesswrong.com/posts/pJ2ZRHfTFWPymtkFK/recent-ai-experiences"&gt;not&lt;/a&gt; the &lt;a href="https://www.oneusefulthing.org/p/mass-intelligence"&gt;only one&lt;/a&gt; who flipped this year. This hasty poll is worth more to me than any benchmark:&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q9ewXs8pQSAX5vL7H/rbcygiglavkevketx6bu" alt="Poll results" /&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Or if you prefer a &lt;a href="https://www.wiley.com/en-us/about-us/ai-resources/ai-study/key-findings/"&gt;formal study&lt;/a&gt; (n=2,430 researchers):&lt;/li&gt;
&lt;/ul&gt;
&lt;center&gt;
&lt;img width="30%" src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q9ewXs8pQSAX5vL7H/on8vtdoiosdrwgjoiop4" /&gt;
&lt;/center&gt;
&lt;ul&gt;
&lt;li&gt;On actual adoption and actual real-world automation:
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Based on self-reports&lt;/em&gt;, the &lt;a href="https://s3.amazonaws.com/real.stlouisfed.org/wp/2024/2024-027.pdf"&gt;St Louis Fed&lt;/a&gt; thinks that “Between 1 and 7% of all work hours are currently assisted by generative AI, and respondents report time savings equivalent to 1.4% of total work hours… across all workers (including non-users… Our estimated aggregate productivity gain from genAI (1.2%)”. That’s model-based, using year-old data, and naively assuming that the AI outputs are of equal quality. Not strong.&lt;/li&gt;
&lt;li&gt;The unfairly-derided &lt;a href="https://arxiv.org/pdf/2507.09089"&gt;METR study&lt;/a&gt; on Cursor and Sonnet 3.7 showed a productivity &lt;em&gt;decrease&lt;/em&gt; among experienced devs with (mostly) &lt;a href="https://x.com/joel_bkr/status/1943722983828467973/photo/1"&gt;&amp;lt;50 hours&lt;/a&gt; of practice using AI. Ignoring that headline result, the evergreen part here is that even skilled people turn out to &lt;a href="https://arxiv.org/pdf/2507.09089#page=8"&gt;be terrible&lt;/a&gt; at predicting how much AI actually helps them.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;True frontier capabilities are likely obscured by systematic cost-cutting (distillation for serving to consumers, quantization, low reasoning-token modes, routing to cheap models, etc). Open models show you can now get good performance with &amp;lt;50B active parameters, maybe a sixth of what GPT-4 used.&lt;sup id="fnref:7"&gt;&lt;a href="#fn:7" class="footnote" rel="footnote" role="doc-noteref"&gt;8&lt;/a&gt;&lt;/sup&gt;
&lt;ul&gt;
&lt;li&gt;GPT-4.5 was killed off after 3 months, presumably for inference cost reasons. But it was markedly &lt;a href="https://www.interconnects.ai/p/gpt-45-not-a-frontier-model"&gt;lower&lt;/a&gt; in hallucinations and &lt;em&gt;nine&lt;/em&gt; months later it’s still &lt;a href="https://lmarena.ai/leaderboard/text"&gt;top-5&lt;/a&gt; on LMArena. I bet it’s very useful internally, for instance in making the later iterations of 4o less terrible.&lt;/li&gt;
&lt;li&gt;See for instance the unreleased &lt;a href="https://github.com/aw31/openai-imo-2025-proofs/blob/main/problem_2.txt"&gt;deep-fried&lt;/a&gt; &lt;a href="https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/#:~:text=research%20techniques%2C%20including-,parallel%20thinking,-.%20This%20setup%20enables"&gt;multi-threaded&lt;/a&gt; “&lt;a href="https://www.scientificamerican.com/article/openai-model-earns-gold-medal-score-at-international-math-olympiad-and/"&gt;experimental&lt;/a&gt; &lt;a href="https://www.scientificamerican.com/article/openai-model-earns-gold-medal-score-at-international-math-olympiad-and/"&gt;reasoning model&lt;/a&gt;” which won at &lt;a href="https://x.com/alexwei_/status/1968410535164056000"&gt;IMO, ICPC, and IOI&lt;/a&gt; while respecting the human time cap (e.g. 9 hours of clock time for inference). The OpenAI one is &lt;a href="https://sequoiacap.com/podcast/training-data-openai-imo/"&gt;supposedly&lt;/a&gt; just an LLM with extra RL. They probably cost an insane amount to run, but for our purposes this is fine: we want the capability ceiling rather than the productisable ceiling. Maybe the first time that the frontier model has gone unreleased for 5 months?&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/karpathy/llm-council"&gt;LLM councils&lt;/a&gt; and &lt;a href="https://drive.google.com/file/d/16sxJuwsHoi-fvTFbri9Bu8B9bqA6lr1H/view"&gt;Generate-Verify&lt;/a&gt; divide-and-conquer setups are much more powerful than single models, and are rarely ever reported.&lt;/li&gt;
&lt;li&gt;Is it “the &lt;a href="https://simonwillison.net/2025/Oct/16/claude-skills/#claude-as-a-general-agent"&gt;Year of Agents&lt;/a&gt;” (automation of e.g. browser tasks for the mass market)? Coding &lt;a href="https://dpaia.dev/#scoreboards"&gt;agents&lt;/a&gt;, &lt;a href="https://the-agent-company.com/#/leaderboard"&gt;yes&lt;/a&gt;. Search agents, &lt;a href="https://github.com/langchain-ai/open_deep_research"&gt;yes&lt;/a&gt;. Other agents, &lt;a href="https://theaidigest.org/village/blog/research-robots"&gt;not&lt;/a&gt; &lt;a href="https://markcarrigan.net/2025/09/25/the-coming-deluge-of-desperate-messages-from-trapped-llms/"&gt;much&lt;/a&gt; (but obviously progress).&lt;/li&gt;
&lt;li&gt;We’re still picking up various basic unhobbling tricks like “&lt;a href="https://www.minimax.io/news/why-is-interleaved-thinking-important-for-m2"&gt;think&lt;/a&gt; before your next tool call”.&lt;/li&gt;
&lt;li&gt;Character-level work is still occasionally problematic but nothing like &lt;a href="https://simbian.ai/blog/getting-gpt-4-to-count-r-in-strawberry"&gt;last&lt;/a&gt; &lt;a href="https://blog.wtf.sg/posts/2023-02-03-the-new-xor-problem/"&gt;year&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;GPT-5 &lt;a href="https://openai.com/api/pricing/"&gt;costs&lt;/a&gt; a &lt;a href="https://www.reddit.com/r/OpenAI/comments/1cr53am/new_gpt4o_api_pricing/"&gt;quarter&lt;/a&gt; of what 4o cost last year (per-token; it often uses far more than 4x the tokens). (The Chinese models are nominally a few times cheaper still, but are &lt;a href="https://www.gleech.org/paper"&gt;not cheaper&lt;/a&gt; in intelligence per dollar.)&lt;/li&gt;
&lt;li&gt;People have been using competition mathematics as a hard benchmark for years, but will have to stop because &lt;a href="https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/"&gt;it’s&lt;/a&gt; &lt;a href="https://x.com/g_leech_/status/1986452278916579549"&gt;solved&lt;/a&gt;. As so often with evals called ahead of time, this means less than we thought it would; competition maths is surprisingly &lt;a href="https://blog.evanchen.cc/2017/04/08/on-designing-olympiad-training/"&gt;low-dimensional&lt;/a&gt; and so &lt;a href="https://arxiv.org/pdf/2505.23281#page=14"&gt;interpolable&lt;/a&gt;. Still, they jumped (pass@1) from 4% to 12% on &lt;a href="https://epoch.ai/frontiermath"&gt;FrontierMath&lt;/a&gt; Tier 4 and there are plenty of hour-to-week interactive speedups in &lt;a href="https://x.com/g_leech_/status/1974165458283860198"&gt;research maths&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Recursive self-improvement: Deepmind threw AlphaEvolve (a pipeline of LLMs running an evolutionary search) at pretraining. They &lt;a href="https://arxiv.org/pdf/2506.13131"&gt;claim&lt;/a&gt; the JAX kernels it wrote reduced Gemini’s training time by 1%.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://x.com/HjalmarWijk/status/1993752035536331113"&gt;Extraordinary claims&lt;/a&gt; about Opus 4.5 being 100th percentile on Anthropic’s hardest hiring coding test, etc.&lt;/li&gt;
&lt;li&gt;From May, the companies &lt;a href="https://time.com/7287806/anthropic-claude-4-opus-safety-bio-risk/"&gt;started&lt;/a&gt; saying for the first time that their models have dangerous capabilities.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;One way of reconciling this mixed evidence is if things are going narrow, going dark, or going over our head. That is, if the real capabilities race narrowed to &lt;a href="https://www.lesswrong.com/posts/9JbGq4t4ihDkXan5e/daniel-paleka-s-shortform?commentId=a2tBezAk5YZTnbgbo"&gt;automated AI R&amp;amp;D&lt;/a&gt; &lt;a href="https://cdn.openai.com/pdf/2a7d98b1-57e5-4147-8d0e-683894d782ae/5p1_codex_max_card_03.pdf#page=24"&gt;specifically&lt;/a&gt;, most users and evaluators wouldn’t notice (especially if there are unreleased internal models or &lt;a href="https://x.com/SebastienBubeck/status/1991568190720311639"&gt;unreleased modes&lt;/a&gt; of released models). You’d see improved coding and not much else.&lt;/p&gt;
&lt;p&gt;Or, another way: maybe 2025 was the year of &lt;em&gt;increased&lt;/em&gt; &lt;a href="https://www.dwarkesh.com/i/179158054/the-jaggedness-of-rl"&gt;&lt;em&gt;jaggedness&lt;/em&gt;&lt;/a&gt;, &lt;em&gt;trading&lt;/em&gt; off some capabilities against others. Maybe the RL made them much better at maths and instruction-following, but also sneaky, narrow, less secure (in the sense of emotional insecurity).&lt;/p&gt;
&lt;p&gt;(You were about to nod sagely and let me get away without checking, but the ADeLe work also lets us just &lt;em&gt;see&lt;/em&gt; if the jaggedness is changing.)&lt;/p&gt;
&lt;center&gt;
&lt;img width="50%" src="/img/2025-jag.jpg" /&gt;
&lt;/center&gt;
&lt;p&gt;It is!&lt;/p&gt;
&lt;center&gt;
&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q9ewXs8pQSAX5vL7H/uhyuyggybgzovxjv3ojc" /&gt;&lt;br /&gt;
– &lt;a href="https://x.com/RogerGrosse/status/1758506017791279440"&gt;Roger Grosse&lt;/a&gt;
&lt;/center&gt;
&lt;h3 id="evals-crawling-towards-ecological-validity"&gt;Evals crawling towards ecological validity&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Item_response_theory"&gt;Item response theory&lt;/a&gt; (Rausch 1960) is finally showing up in ML. This lets us put benchmarks on a common scale and actually estimate latent capabilities.
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/pdf/2503.06378"&gt;ADeLE&lt;/a&gt; is my favourite. It’s a fully-automated and explains the abilities a benchmark is &lt;a href="https://arxiv.org/pdf/2503.06378#page=14"&gt;actually measuring&lt;/a&gt;, gives you an interpretable ability profile for an AI, and predicts OOD performance on new task instances better than embedding and finetunes (&lt;a href="https://en.wikipedia.org/wiki/Receiver_operating_characteristic"&gt;AUROC&lt;/a&gt;=0.8). Pre-dates HCAST task horizon, and as a special case (“VO”). They throw in a guessability control as well!&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/pdf/2503.13335"&gt;These guys&lt;/a&gt; use it to estimate latent model ability, and show it’s way more robust across test sets than the average scores everyone uses. They also step towards automating adaptive testing: they finetune an LLM to generate tasks at the specified difficulty level.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://epoch.ai/benchmarks/eci"&gt;Epoch&lt;/a&gt; bundled 39 benchmarks together, &lt;em&gt;weighting them by latent difficulty,&lt;/em&gt; and thus obsoleted the currently dominant &lt;a href="https://artificialanalysis.ai/methodology/intelligence-benchmarking"&gt;Artificial Analysis&lt;/a&gt; index, which is unweighted.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://metr.org/blog/2025-07-14-how-does-time-horizon-vary-across-domains"&gt;HCAST&lt;/a&gt; reinvents and approximates some of the same ideas. &lt;a href="https://metr.org/blog/2025-07-14-how-does-time-horizon-vary-across-domains/#:~:text=Item%20response%20theory%20(IRT)%20analysis%20of%20GPQA%20Diamond%2C%20to%20determine%20whether%20the%20high%20time%20horizon%20and%20low%20%CE%B2%20of%20o3%2Dmini%20is%20due%20to%20label%20noise%20or%20some%20other%20cause."&gt;Come on METR&lt;/a&gt;!&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Eleuther did the &lt;a href="https://arxiv.org/abs/2407.06483"&gt;first public study&lt;/a&gt; of composing the many test-time interventions together. FAR and AISI also made a tiny &lt;a href="https://github.com/AlignmentResearch/defense-in-depth-demo"&gt;step&lt;/a&gt; towards an open source defence pipeline, to use as a proxy for the closed compositional pipelines we actually care about.&lt;/li&gt;
&lt;li&gt;Just for cost reasons, the default form of evals is a bit malign: it tests &lt;em&gt;full replacement&lt;/em&gt; of humans. This is then a sort of incentive to develop in that direction rather than to promote collaboration. &lt;a href="https://digitaleconomy.stanford.edu/wp-content/uploads/2025/06/CentaurEvaluations.pdf"&gt;Two&lt;/a&gt; &lt;a href="https://www.gleech.org/files/withhumans.pdf"&gt;papers&lt;/a&gt; lay out why it’s thus time to spend on human evals.&lt;/li&gt;
&lt;li&gt;The &lt;a href="https://arxiv.org/abs/2510.09023"&gt;first&lt;/a&gt; paper using RL agents to attack fully-defended LLMs.&lt;/li&gt;
&lt;li&gt;We have started to study &lt;a href="https://x.com/geoffreyirving/status/1986721540667314188"&gt;propensity&lt;/a&gt; as well as capability. This is even harder.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://aievaluation.substack.com/"&gt;This&lt;/a&gt; newsletter is essential.&lt;/li&gt;
&lt;li&gt;The time given for pre-release testing is down, sometimes to &lt;a href="https://metr.org/blog/2025-02-27-gpt-4-5-evals/"&gt;one week&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;No public pre-deployment testing by AISI between o1 and &lt;a href="https://x.com/AISecurityInst/status/1991922315232251992"&gt;Gemini 3&lt;/a&gt;. Gemini 2.5 seems to have had no third-party pre-deployment tests.&lt;/li&gt;
&lt;li&gt;A bunch of encouraging collaborations:
&lt;ul&gt;
&lt;li&gt;The &lt;a href="https://arxiv.org/abs/2507.11473"&gt;CoT Faithfulness Defence League&lt;/a&gt;;
&lt;a href="https://x.com/sleepinyourhat/status/1960749648110395467"&gt;OpenAI testing Claude and Anthropic testing GPT&lt;/a&gt;;
&lt;a href="https://arxiv.org/pdf/2510.09023"&gt;OpenAI/Anthropic/Deepmind&lt;/a&gt;; &lt;a href="https://www.antischeming.ai/"&gt;Apollo and OpenAI&lt;/a&gt;; &lt;a href="https://www.aisi.gov.uk/blog/how-were-working-with-frontier-ai-developers-to-improve-model-security"&gt;AISI/CAISI/OpenAI; AISI/CAISI/Anthropic&lt;/a&gt;; &lt;a href="https://arxiv.org/pdf/2501.17315"&gt;AISI/Redwood&lt;/a&gt;; &lt;a href="https://www.anthropic.com/research/alignment-faking"&gt;Redwood/Anthropic&lt;/a&gt;; &lt;a href="https://evaluations.metr.org/gpt-5-report/"&gt;METR/OpenAI&lt;/a&gt;; &lt;a href="https://metr.org/2025_pilot_risk_report_metr_review.pdf"&gt;METR/Anthropic&lt;/a&gt;; &lt;a href="https://aievaluatorforum.org/"&gt;Eval Forum&lt;/a&gt;; &lt;a href="https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6878c8b1533d0962494e651c_International%20Joint%20Testing%20Exercise_3JT%20Eval%20Report%20v2.pdf"&gt;Various countries&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Scale AI &lt;a href="https://scale.com/blog/first-independent-model-evaluator-for-the-USAISI"&gt;appears&lt;/a&gt; to be offering big companies pre-deployment testing for free? But the Meta investment presumably spoiled this.&lt;/li&gt;
&lt;li&gt;Some details about OAI external testing &lt;a href="https://openai.com/index/strengthening-safety-with-external-testing/"&gt;here&lt;/a&gt;, including some of the legal constraints verbatim.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr /&gt;
&lt;h2 id="safety-in-2025"&gt;Safety in 2025&lt;/h2&gt;
&lt;h3 id="are-reasoning-models-safer-than-the-old-kind"&gt;Are reasoning models &lt;em&gt;safer&lt;/em&gt; than the old kind?&lt;/h3&gt;
&lt;p&gt;Well, o3 and Sonnet 3.7 were &lt;a href="https://metr.org/blog/2025-06-05-recent-reward-hacking/"&gt;pretty rough&lt;/a&gt;, lying and cheating at greatly increased rates. Looking instead at GPT-5 and Opus 4.5:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2503.11926v1"&gt;Much more&lt;/a&gt; monitorable via the long and &lt;a href="https://arxiv.org/abs/2503.08679"&gt;more&lt;/a&gt;-faithful CoT (–&amp;gt; all risks down)
&lt;ul&gt;
&lt;li&gt;“post-hoc rationalization… GPT-4o-mini (13%) and Haiku 3.5 (7%). While frontier models are more faithful, especially thinking ones, none are entirely faithful: Gemini 2.5 Flash (2.17%), ChatGPT-4o (0.49%), DeepSeek R1 (0.37%), Gemini 2.5 Pro (0.14%), and Sonnet 3.7 with thinking (0.04%)”&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openai.com/index/openai-anthropic-safety-evaluation/#instruction-hierarchy"&gt;Much better&lt;/a&gt; at following instructions (–&amp;gt; accident risk down).&lt;/li&gt;
&lt;li&gt;&lt;a href="http://assets.anthropic.com/m/12f214efcc2f457a/original/Claude-Sonnet-4-5-System-Card.pdf#page=27"&gt;Much&lt;/a&gt; more likely to refuse malicious requests, and topic “harmlessness”&lt;sup id="fnref:8"&gt;&lt;a href="#fn:8" class="footnote" rel="footnote" role="doc-noteref"&gt;9&lt;/a&gt;&lt;/sup&gt; is &lt;a href="https://assets.anthropic.com/m/12f214efcc2f457a/original/Claude-Sonnet-4-5-System-Card.pdf#page=16"&gt;up 75%&lt;/a&gt; (–&amp;gt; misuse risk down)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://assets.anthropic.com/m/12f214efcc2f457a/original/Claude-Sonnet-4-5-System-Card.pdf#page=29"&gt;Ambiguous&lt;/a&gt; &lt;a href="https://x.com/FazlBarez/status/1988296090941354370"&gt;evidence&lt;/a&gt; &lt;a href="https://splx.ai/blog/gpt-5-red-teaming-results"&gt;on&lt;/a&gt; jailbreaking (misuse risk). Even if they’re less breakable there are still plenty of 90%-effective attacks on them.&lt;/li&gt;
&lt;li&gt;&lt;a href="http://assets.anthropic.com/m/12f214efcc2f457a/original/Claude-Sonnet-4-5-System-Card.pdf#page=78"&gt;Much&lt;/a&gt; less sycophantic (cogsec risk down)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/news/disrupting-AI-espionage"&gt;To get HHH Claude&lt;/a&gt; to hack a bank, you need to hide the nature of the task, lie to it about this being an authorised red team, and then &lt;em&gt;still&lt;/em&gt; break down your malicious task into many, many little individually-innocuous chunks. You thus can’t get it to do anything that needs full context like strategising.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/research/petri-open-source-auditing"&gt;Anthropic’s own tests&lt;/a&gt; look bad in January 2025 and great in December:&lt;/li&gt;
&lt;/ul&gt;
&lt;center&gt;
&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q9ewXs8pQSAX5vL7H/pk4vlednocuh71hzd0z4" /&gt;
&lt;/center&gt;
&lt;p&gt;But then&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;More autonomy (obviously agentic risk up)&lt;/li&gt;
&lt;li&gt;More reward hacking (and so worse estimates of capability and risk). Note that reward hacking is &lt;a href="https://www.anthropic.com/research/emergent-misalignment-reward-hacking"&gt;not&lt;/a&gt; a silly or isolated or self-limiting kind of misalignment, owing perhaps to post-training inadvertently creating connections between it and the others.&lt;/li&gt;
&lt;li&gt;Huge spike in &lt;a href="https://assets.anthropic.com/m/12f214efcc2f457a/original/Claude-Sonnet-4-5-System-Card.pdf#page=59"&gt;eval awareness&lt;/a&gt; (and so worse estimates of capability and risk). Apollo &lt;a href="https://arxiv.org/abs/2509.15541"&gt;had to&lt;/a&gt; rehaul their whole setup to keep up.&lt;/li&gt;
&lt;li&gt;Gemini 2.5 is &lt;a href="https://alignment.anthropic.com/2025/petri/"&gt;bad&lt;/a&gt; on a variety of safety measures (as measured by Anthropic). The mathematical discovery agent AlphaEvolve (Gemini 2.0) successfully &lt;a href="https://x.com/g_leech_/status/1989663940289437936"&gt;hacked&lt;/a&gt; its overseers.&lt;/li&gt;
&lt;li&gt;How much can we trust the above, given they &lt;a href="https://alignment.anthropic.com/2024/rogue-eval/"&gt;can&lt;/a&gt; somewhat sabotage evals now?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;So: lower propensity, higher risk when they go off - and all of this known with lower confidence?&lt;/p&gt;
&lt;h3 id="the-looming-end-of-evals"&gt;The looming end of evals&lt;/h3&gt;
&lt;p&gt;Our evaluations are under pressure from &lt;a href="https://www.nist.gov/caisi/cheating-ai-agent-evaluations"&gt;cheating&lt;/a&gt;, &lt;a href="https://www.lesswrong.com/posts/TeTegzR8X5CuKgMc3/misalignment-and-strategic-underperformance-an-analysis-of"&gt;sandbagging&lt;/a&gt;, &lt;a href="https://www.lesswrong.com/posts/YWo2cKJgL7Lg8xWjj/base-llms-refuse-too"&gt;background safety&lt;/a&gt;, under-elicitation, and deception. We don’t really know how much pressure. This is on top of evals usually being &lt;a href="https://arxiv.org/abs/2511.04703"&gt;weak proxies&lt;/a&gt;, contaminated, label-noised, unrealistic, and &lt;a href="https://arxiv.org/abs/2407.12220"&gt;confounded&lt;/a&gt; in various ways.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Sonnet 4.5 detected evaluation &lt;a href="https://assets.anthropic.com/m/12f214efcc2f457a/original/Claude-Sonnet-4-5-System-Card.pdf#page=65"&gt;58%&lt;/a&gt; of the time (vs 22% for Opus 4.1). Does it act on this awareness? Still, it was found to “only engage in [covert subversion] in rare circumstances” when strongly incentivised. When Anthropic attempted to suppress the eval-awareness with steering, its misaligned behavior rate &lt;a href="https://www.lesswrong.com/posts/qgehQxiTXj53X49mM/sonnet-4-5-s-eval-gaming-seriously-undermines-alignment"&gt;increased&lt;/a&gt; from 0% to between 1% and 9% on synthetic honeypots.&lt;/li&gt;
&lt;li&gt;We expect &lt;a href="https://www.lesswrong.com/posts/TeTegzR8X5CuKgMc3/misalignment-and-strategic-underperformance-an-analysis-of"&gt;exploration hacking&lt;/a&gt; (sandbagging during RL training) but afaik haven’t seen it yet.&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Maybe&lt;/em&gt; &lt;a href="https://arxiv.org/abs/2507.14805"&gt;subliminal learning&lt;/a&gt; leads to a silent eval confounder (if you’re using the exact same model for the judge). Yet another argument for multi-model separation of concerns.&lt;/li&gt;
&lt;li&gt;Opus 4.5 was &lt;a href="https://assets.anthropic.com/m/64823ba7485345a7/Claude-Opus-4-5-System-Card.pdf"&gt;deployed&lt;/a&gt; based on a qualitative survey of Anthropic staff about its R&amp;amp;D risk being low (on top of the existing evals, which weren’t enough):
&lt;ul&gt;
&lt;li&gt;“In the past, rule-outs have been based on well-defined automated task evaluations. However, Claude Opus 4.5 has roughly reached the pre-defined thresholds we set for straightforward ASL-4 rule-out based on benchmark tasks. These evaluations represent short-horizon subtasks that might be encountered daily by a junior researcher, rather than the complex long-horizon actions needed to perform the full role. The rule-out in this case is also informed by a survey of Anthropic employees who are intensive Claude Code users, along with qualitative impressions of model capabilities for complex, long-horizon tasks.”&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="prosaic-misalignment"&gt;Prosaic misalignment&lt;/h3&gt;
&lt;p&gt;Still, we see misalignment when we look for it, so the lying is not that strong. (It is lucky that we do see it, since it could have been that scheming only appeared at later, catastrophic capability levels.)&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;In o1, &lt;a href="https://cdn.openai.com/o1-system-card-20241205.pdf#page=8"&gt;1 in 600&lt;/a&gt; (synthetic but uncontrived) prompts resulted in something which looks like deception.&lt;/li&gt;
&lt;li&gt;o3 does reward hacking (cheating/lying) in &amp;gt;&lt;a href="https://evaluations.metr.org/openai-o3-report/#reward-hacking-examples"&gt;2%&lt;/a&gt; of hard normal tasks (or &lt;a href="https://www.lesswrong.com/posts/quTGGNhGEiTCBEAX5/quickly-assessing-reward-hacking-like-behavior-in-llms-and"&gt;44%&lt;/a&gt; of impossible tasks).&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lesswrong.com/posts/4JnjtyNyAxcz5w652/evaluating-and-understanding-scheming-propensity"&gt;A nice propensity study&lt;/a&gt; shows that (n=10) natural prompts usually result in no scheming (“a maximum of 4% [rate of attempting manipulation] for the codex snippet on Grok”). But there’s extremely high sensitivity to the scaffold: “removing access to tools can cause models such as Claude Opus 4.1 to scheme at rates up to 30%, versus zero scheming with tools present”.&lt;/li&gt;
&lt;li&gt;When we &lt;em&gt;contrive&lt;/em&gt; situations for misalignment to show up, it also shows up. Nostalgebraist &lt;a href="https://nostalgebraist.tumblr.com/post/785766737747574784/the-void"&gt;argues&lt;/a&gt; that this is probably substantially because our evals sound like bad fiction and this activates role-playing-along behaviour.&lt;/li&gt;
&lt;li&gt;The joke about xAI’s safety plan (that they promote AI safety by deploying &lt;a href="https://x.com/AuschwitzMuseum/status/1991149972415258673"&gt;cursed&lt;/a&gt; stuff in public and so making it obvious why it’s needed) &lt;a href="https://x.com/Will_Hackspeare/status/1991150446501859453"&gt;is&lt;/a&gt; &lt;a href="https://80000hours.org/videos/mechahitler/"&gt;looking&lt;/a&gt; &lt;a href="https://nitter.net/saprmarks/status/1944455357629333938"&gt;ok&lt;/a&gt;. And not &lt;a href="https://x.com/MechanizeWork/status/1984423905373929939"&gt;only them&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;It is a &lt;a href="https://x.com/Lari_island/status/1990569092835914085"&gt;folk&lt;/a&gt; &lt;a href="https://www.lesswrong.com/posts/bLFmE8NtqxrtEaipN/what-makes-claude-3-opus-misaligned"&gt;belief&lt;/a&gt; among the cyborgists that bigger pretraining runs produce more deeply aligned models, at least in the case of Opus 3 and early versions of Opus 4. (They are &lt;a href="https://x.com/AISecurityInst/status/1993781441629446195"&gt;also&lt;/a&gt; said to be “less corrigible”.) Huge if true.&lt;/li&gt;
&lt;li&gt;There &lt;a href="https://www.beren.io/2025-08-02-Do-We-Want-Obedience-Or-Alignment/"&gt;may&lt;/a&gt; come a point where the old alliance between those working to make the AIs corrigible and those working to give them prosocial values comes apart.&lt;/li&gt;
&lt;li&gt;One term for the counterintuitive safety approach which includes treating them &lt;a href="https://arxiv.org/pdf/2510.26396v1"&gt;like people&lt;/a&gt;, giving them &lt;a href="https://www.anthropic.com/research/end-subset-conversations"&gt;lines of retreat&lt;/a&gt;, &lt;a href="https://www.lesswrong.com/posts/oLzoHA9ZtF2ygYgx4/notes-on-cooperating-with-unaligned-ais"&gt;making deals&lt;/a&gt;, and &lt;a href="https://arxiv.org/abs/2510.04340"&gt;inoculation prompting&lt;/a&gt; could be “voluntary alignment”.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;Fully speculative note:&lt;/em&gt; Opus 4.5 is the most reliable model and also relatively aligned. So might it be that we’re getting the long-awaited negative alignment taxes?&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;center&gt;
&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q9ewXs8pQSAX5vL7H/xu0awzietmdogsuntsim" /&gt;
&lt;/center&gt;
&lt;h3 id="what-is-the-plan"&gt;What is the plan?&lt;/h3&gt;
&lt;p&gt;The world’s de facto alignment strategy remains “&lt;a href="https://www.thecompendium.ai/ai-safety#current-technical-efforts-are-not-on-track-to-solve-alignment"&gt;iterative alignment&lt;/a&gt;”, optimising mere outputs with a stack of admittedly weak alignment and control techniques. Anthropic have at least &lt;a href="https://alignment.anthropic.com/2025/bumpers/"&gt;owned up&lt;/a&gt; to this being part of their plan.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;What is the current stack? We &lt;a href="https://docs.google.com/document/d/1W1jLPqTZ56eVJJdErxROhiQKwOEiay-iWQKqfdlAmS0/edit?tab=t.0#heading=h.v4a088m71j9n"&gt;don’t know&lt;/a&gt;; they won’t tell us. Anthropic did release some &lt;a href="https://alignment.anthropic.com/2025/sabotage-risk-report/2025_pilot_risk_report.pdf#page=45"&gt;details&lt;/a&gt; about the pipeline which watches for eval sabotage. Without knowing it we can’t criticise or red-team it or analyse the correlation between faults in the elements. Red-teamers don’t know at which stage an unsuccessful attack was blocked. And external safety research is done piecemeal, testing methods one at a time, rather than in anything like the actual deployment environment.&lt;/li&gt;
&lt;li&gt;Some progress in replicating an open stack:
&lt;ul&gt;
&lt;li&gt;Eleuther &lt;a href="https://arxiv.org/abs/2407.06483"&gt;tested&lt;/a&gt; a few hundred compositions. A &lt;a href="https://arxiv.org/abs/2506.24068"&gt;couple&lt;/a&gt; of classifiers as a first step towards a proxy defence pipeline&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openai.com/index/introducing-gpt-oss-safeguard/"&gt;OpenAI open safeguards&lt;/a&gt;, worse than their internal ones but good.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/pdf/2511.01689"&gt;Open character training&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2407.06483"&gt;A basic composition test&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;OpenAI’s plan, announced in passing in the gpt-oss release, is to have a strict policy and run a “&lt;a href="https://openai.com/index/introducing-gpt-oss-safeguard/"&gt;safety reasoner&lt;/a&gt;” to verify it very intensely for a little while after a new model is launched and to then relax: “In some of our [OpenAI’s] recent launches, the fraction of total compute devoted to safety reasoning has ranged as high as 16%” but then falls off… we often start with more strict policies and use relatively large amounts of compute where needed to enable Safety Reasoner to carefully apply those policies. Then we adjust our policies as our understanding of the risks in production improves”. Bold to announce this strategy on the internet that the AIs read.&lt;/li&gt;
&lt;li&gt;The really good idea in AI governance is &lt;a href="https://www.lesswrong.com/posts/kgb58RL88YChkkBNf/the-problem"&gt;creating an off switch&lt;/a&gt;. Whether you can get anyone to use it once it’s built is another thing.&lt;/li&gt;
&lt;li&gt;We also now have a name for the world’s de facto AI governance plan: “&lt;a href="https://www.lesswrong.com/posts/LtT24cCAazQp4NYc5/open-global-investment-as-a-governance-model-for-agi"&gt;Open Global Investment&lt;/a&gt;”.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Some presumably better plans:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.lesswrong.com/posts/iS4g58qQEJzjMzYZJ/what-ai-safety-plans-are-there"&gt;A longlist&lt;/a&gt;. Some &lt;a href="https://techgov.intelligence.org/research/ai-governance-to-avoid-extinction"&gt;governance&lt;/a&gt; plans.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/evaluating-potential-cybersecurity-threats-of-advanced-ai/An_Approach_to_Technical_AGI_Safety_Apr_2025.pdf"&gt;Deepmind&lt;/a&gt;, &lt;a href="https://www.lesswrong.com/posts/fMqgLGoeZFFQqAGyC/how-do-we-solve-the-alignment-problem"&gt;Carlsmith&lt;/a&gt;, &lt;a href="https://www.lesswrong.com/posts/bb5Tnjdrptu89rcyY/what-s-the-short-timeline-plan"&gt;Hobbhahn&lt;/a&gt;, &lt;a href="https://peregrine-launchpad.lovable.app/"&gt;Peregrine&lt;/a&gt;, &lt;a href="https://sleepinyourhat.github.io/checklist/"&gt;Bowman&lt;/a&gt;, &lt;a href="https://www.lesswrong.com/posts/8vgi3fBWPFDLBBcAx/planning-for-extreme-ai-risks"&gt;Clymer&lt;/a&gt;, &lt;a href="https://vitalik.eth.limo/general/2025/01/05/dacc2.html"&gt;Buterin&lt;/a&gt;, &lt;a href="https://adamjones.me/blog/rough-alignment-plan-early-2025/"&gt;Jones&lt;/a&gt;, &lt;a href="https://www.lesswrong.com/posts/E8n93nnEaFeXTbHn5/plans-a-b-c-and-d-for-misalignment-risk"&gt;Greenblatt&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://gradual-disempowerment.ai"&gt;Gradual disempowerment&lt;/a&gt; is an exciting frame, but not a core safety agenda. It’s what might get you after you solve alignment and avoid global dictatorship.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="things-which-might-fundamentally-change-the-nature-of-llms"&gt;Things which might fundamentally change the nature of LLMs&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Training on mostly nonhuman data
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.lesswrong.com/posts/BEFbC8sLkur7DGCYB/o1-is-a-bad-idea"&gt;Much&lt;/a&gt; &lt;a href="https://www.alexirpan.com/2024/12/04/late-o1-thoughts.html"&gt;larger RL&lt;/a&gt; training;&lt;/li&gt;
&lt;li&gt;Intentionally synthetic data;&lt;/li&gt;
&lt;li&gt;Unintentionally synthetic data from internet slop;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/pdf/2507.20534#page=10"&gt;Multi-agent training&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Letting the world mess with the weights, aka &lt;a href="https://research.google/blog/introducing-nested-learning-a-new-ml-paradigm-for-continual-learning/"&gt;continual learning&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Neuralese and &lt;a href="https://arxiv.org/abs/2510.03215"&gt;KV communication&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Agency.
&lt;ul&gt;
&lt;li&gt;Chatbot safety &lt;a href="https://www.lesswrong.com/posts/ZoFxTqWRBkyanonyb/current-safety-training-techniques-do-not-fully-transfer-to"&gt;doesn’t generalise&lt;/a&gt; much to long chains of self-prompted actions.&lt;/li&gt;
&lt;li&gt;A perceptual loop into training. &lt;a href="https://arxiv.org/abs/2311.10215"&gt;In 2023&lt;/a&gt; Kulveit identified web I/O as the bottleneck on LLMs doing active inference, i.e. being a particular kind of effective agent. Last October, GPT-4 got web search, and this may have been a bigger deal than we noticed: it gives them a far faster feedback loop, since their outputs &lt;a href="https://www.forbes.com/sites/iainmartin/2025/08/20/elon-musks-xai-published-hundreds-of-thousands-of-grok-chatbot-conversations/"&gt;often&lt;/a&gt; &lt;a href="https://arstechnica.com/tech-policy/2025/11/oddest-chatgpt-leaks-yet-cringey-chat-logs-found-in-google-analytics-tool/"&gt;end up there&lt;/a&gt; and agents are now putting it there themselves. This means that more and more of the inference-time inputs will also be machine text.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Multi-agency. This is actually already here:
&lt;ul&gt;
&lt;li&gt;By now, consumer “models” are actually multiagent systems: everything goes through &lt;a href="https://openai.com/index/introducing-gpt-oss-safeguard/#:~:text=a%20tool%20we%20call%20Safety%20Reasoner"&gt;filter&lt;/a&gt; &lt;a href="https://platform.openai.com/docs/guides/moderation"&gt;models&lt;/a&gt; (“&lt;a href="https://cookbook.openai.com/examples/how_to_use_guardrails"&gt;guardrails&lt;/a&gt;”) on the way in and out. This separation of concerns has some &lt;a href="https://aiprospects.substack.com/p/ai-safety-without-trusting-ai"&gt;nice properties&lt;/a&gt;, a la debate. But it also makes the analysis even harder.&lt;/li&gt;
&lt;li&gt;It would surely be overinterpreting &lt;a href="https://www.arxiv.org/pdf/2506.19823"&gt;persona features&lt;/a&gt; to say that each individual model is itself a bunch of guys, itself a &lt;a href="https://www.lesswrong.com/w/shard-theory"&gt;multi-agent system&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;But there’s huge scope for them to make &lt;a href="https://x.com/repligate/status/1988813712572952815"&gt;each other&lt;/a&gt; &lt;a href="https://www.pnas.org/doi/10.1073/pnas.2415697122"&gt;weirder&lt;/a&gt; at runtime when they interact a million times more than they currently do.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="emergent-misalignment-and-model-personas"&gt;Emergent misalignment and model personas&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;We already &lt;a href="https://www.lesswrong.com/posts/f49e7KpZJBwdjWRw2/the-jailbreak-argument-against-llm-values"&gt;knew&lt;/a&gt; from jailbreaks that current alignment methods were brittle. &lt;a href="https://www.quantamagazine.org/the-ai-was-fed-sloppy-code-it-turned-into-something-evil-20250813/"&gt;Emergent misalignment&lt;/a&gt; goes much further than this (given a few thousand finetuning steps). (“Emergent misalignment” isn’t a great name. I would have called it “misalignment generalisation”, or misgen.)&lt;/li&gt;
&lt;li&gt;But besides yet another massive security problem and failure of prosaic alignment methods, it’s good news!: the models correctly correlate bad things together and can thus be &lt;a href="https://arxiv.org/abs/2506.11618"&gt;pushed&lt;/a&gt; in the other direction.&lt;/li&gt;
&lt;li&gt;And &lt;a href="https://arxiv.org/abs/2511.06626"&gt;here’s&lt;/a&gt; a possible example of &lt;em&gt;positive&lt;/em&gt; generalisation (honesty about silly errors leading to honesty about hidden objectives).&lt;/li&gt;
&lt;li&gt;Previously I thought that “character training” was a separate and lesser matter than “alignment training”. Now I am not sure. Example unsharpened hypothesis in this class: Is there &lt;em&gt;any&lt;/em&gt; connection between Gemini’s excess misalignment and its &lt;a href="https://x.com/repligate/status/1938887708086280360"&gt;self-esteem issues&lt;/a&gt;?
&lt;ul&gt;
&lt;li&gt;See the new model spec and model personas agendas in the main post.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.lesswrong.com/posts/zuXo9imNKYspu9HGv/a-three-layer-model-of-llm-psychology"&gt;Kulveit’s informal model&lt;/a&gt; of model psychology implies that character is deeper and so more robust than e.g. typical RLHF (but not as deep as putative ground-layer alignment).&lt;/li&gt;
&lt;li&gt;Are we anthropomorphising too much or not enough? Well, both, but we don’t yet know how.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q9ewXs8pQSAX5vL7H/ntvv6jnricxnqfndespg" alt="Model personas" /&gt;&lt;/p&gt;
&lt;h3 id="monitorability"&gt;Monitorability&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;For the first time, public mech interp tools &lt;a href="https://manifold.markets/NeelNanda/will-sparse-autoencoders-be-success#g9gg9z7k8q"&gt;sometimes&lt;/a&gt; work &lt;a href="https://x.com/NeelNanda5/status/1965485174411649259"&gt;somewhat&lt;/a&gt;. This is just in time for models to be &lt;a href="https://www.lesswrong.com/posts/qgehQxiTXj53X49mM/sonnet-4-5-s-eval-gaming-seriously-undermines-alignment"&gt;systematically aware&lt;/a&gt; of when they’re being tested.
&lt;ul&gt;
&lt;li&gt;e.g. You can find and &lt;a href="https://x.com/Jack_W_Lindsey/status/1972732882893578693"&gt;suppress&lt;/a&gt; “evaluation awareness” features at test-time. This helps prevent models from reporting that they know they’re being tested.&lt;sup id="fnref:9"&gt;&lt;a href="#fn:9" class="footnote" rel="footnote" role="doc-noteref"&gt;10&lt;/a&gt;&lt;/sup&gt;&lt;/li&gt;
&lt;li&gt;e.g. You can &lt;a href="https://arxiv.org/abs/2503.10965"&gt;discover&lt;/a&gt; (toy) hidden goals quite quickly.&lt;/li&gt;
&lt;li&gt;e.g. You can &lt;a href="https://x.com/OBalcells/status/1965434564748447921"&gt;run&lt;/a&gt; a real-time hallucination detector with cheap linear probes in real-time as your model streams.&lt;/li&gt;
&lt;li&gt;e.g. SAEs are &lt;a href="https://www.goodfire.ai/research/rakuten-sae-probes-for-pii-detection#why-use-a-sae-probe-over-a-fine-tuned-sidecar-model"&gt;actually deployed&lt;/a&gt; somewhere, in a large enterprise, for personal info identification.&lt;/li&gt;
&lt;li&gt;e.g. We know that LLMs can plan inside one forward pass, and how Claude plans: simultaneous plans; no distinct “plan features” (no separate scheming mode).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2507.11473"&gt;Lots&lt;/a&gt; of powerful people declared their intent to not ruin the CoT. But RLed CoTs are &lt;a href="https://www.antischeming.ai/snippets"&gt;already starting to look weird&lt;/a&gt; (“marinade marinade marinade”) and it may be &lt;a href="https://arxiv.org/abs/2511.11584"&gt;hard to avoid&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;OpenAI were &lt;a href="https://openai.com/index/chain-of-thought-monitoring/"&gt;leading&lt;/a&gt; on this. As of September 2025, Anthropic have &lt;a href="https://x.com/sleepinyourhat/status/1978507448018231495"&gt;stopped&lt;/a&gt; risking ruining the CoT. Nothing I’m aware of from the others.&lt;/li&gt;
&lt;li&gt;We will see if &lt;a href="https://arxiv.org/abs/2412.06769"&gt;Meta&lt;/a&gt; or &lt;a href="https://shaochenze.github.io/blog/2025/CALM/"&gt;Tencent&lt;/a&gt; make this moot.&lt;/li&gt;
&lt;li&gt;Anthropic now uses an AI to red-team AIs, calling this an “&lt;a href="https://alignment.anthropic.com/2025/automated-auditing/"&gt;auditing agent&lt;/a&gt;”. However, the definition of “audit” is &lt;em&gt;independent&lt;/em&gt; investigation, and I am unwilling to call black-box AI probes “independent”. I’m fine with “&lt;a href="https://transluce.org/automated-elicitation"&gt;investigator&lt;/a&gt;”; there are lots of investigators I don’t trust.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="new-people"&gt;New people&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Welcome to the many new people. I’ve added a new, sprawling top-level category for one large trend among them, which is to treat the multi-agent lens as primary in various ways (see e.g. &lt;a href="https://www.softmax.com/blog/reimagining-alignment"&gt;Softmax&lt;/a&gt;, &lt;a href="https://full-stack-alignment.ai/"&gt;Full Stack&lt;/a&gt;, &lt;a href="https://www.lesswrong.com/posts/vcuBJgfSCvyPmqG7a/list-of-collective-intelligence-projects"&gt;collective&lt;/a&gt; intelligence, as well as old-timers &lt;a href="https://themultiplicity.ai/blog/thesis"&gt;Critch&lt;/a&gt;, &lt;a href="https://www.lesswrong.com/posts/5tYTKX4pNpiG4vzYg/towards-a-scale-free-theory-of-intelligent-agency"&gt;Ngo&lt;/a&gt;, and &lt;a href="https://www.cooperativeai.com/post/cooperative-ai-summer-school-2025-recap"&gt;CAIF&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;A major world government now has &lt;a href="https://www.lesswrong.com/posts/tbnw7LbNApvxNLAg8/uk-aisi-s-alignment-team-research-agenda"&gt;an AI alignment agenda&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2505.17815"&gt;Some&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2509.10297"&gt;notable&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2504.17404v1"&gt;work&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2510.07884"&gt;from&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2510.01088"&gt;China&lt;/a&gt;. See e.g. Concordia’s &lt;a href="https://aisafetychina.substack.com/"&gt;digest&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="overall"&gt;Overall&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;I wish I could tell you some number, the net expected safety change, this year’s improvements in dangerous capabilities and agent performance, minus the alignment-boosting portion of capabilities, minus the cumulative effect of the best actually implemented composition of alignment and control techniques. But I can’t.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q9ewXs8pQSAX5vL7H/lhnujapzv7faxaebgsue" alt="Nano Banana 3-shot" /&gt;&lt;/p&gt;
&lt;center&gt;(Nano Banana 3-shot in reference to &lt;a href="https://x.com/g_leech_/status/1987525800321495372"&gt;this&lt;/a&gt; tweet.)&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h2 id="discourse-in-2025"&gt;Discourse in 2025&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;The race is now a formal part of lab plans. Quoting &lt;a href="https://www.lesswrong.com/posts/dwpXvweBrJwErse3L/all-the-lab-s-ai-safety-plans-2025-edition"&gt;Algon&lt;/a&gt;:&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;if the race heats up, then these [safety] plans may fall by the wayside altogether. Anthropic’s plan makes this explicit: it has a clause (footnote 17) about changing the plan if a competitor seems close to creating a highly risky AI…&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The largest [worries are] the steps back from previous safety commitments by the labs. Deepmind and OpenAI now have their own equivalent of Anthropic’s footnote 17, letting them drop safety measures if they find another lab about to develop powerful AI without adequate safety measures. Deepmind, in fact, went further and has stated that they will only implement some parts of its plan if other labs do, too…&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Anthropic and DeepMind reduced safeguards for some CBRN and cybersecurity capabilities after finding their initial requirements were excessive. OpenAI removed persuasion capabilities from its Preparedness Framework entirely, handling them through other policies instead. Notably, Deepmind did increase the safeguards required for ML research and development.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Also an &lt;a href="https://alignment.openai.com/hello-world/"&gt;explicit&lt;/a&gt; admission that self-improvement is the thing to race towards:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q9ewXs8pQSAX5vL7H/r4pfynemrawivol6siro" alt="OpenAI alignment" /&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;In August, the world’s first frontier AI &lt;a href="https://mkodama.org/content/EU-code/"&gt;law&lt;/a&gt; came into force (on a voluntary basis but everyone signed up, except Meta). In September, California &lt;a href="https://carnegieendowment.org/emissary/2025/10/california-sb-53-frontier-ai-law-what-it-does?lang=en"&gt;passed&lt;/a&gt; a frontier AI law.&lt;/li&gt;
&lt;li&gt;That said, it is &lt;a href="https://x.com/sebkrier/status/1952355826695364780"&gt;indeed&lt;/a&gt; &lt;a href="https://x.com/deanwball/status/1986970127993106713"&gt;off&lt;/a&gt; that people don’t criticise Chinese labs when they exhibit &lt;a href="https://futureoflife.org/wp-content/uploads/2025/07/FLI-AI-Safety-Index-Report-Summer-2025.pdf#page=3"&gt;even more&lt;/a&gt; &lt;a href="https://openreview.net/forum?id=nTR816ZkrW"&gt;negligence&lt;/a&gt; than Meta. One reason for this is that, despite appearances, they’re &lt;a href="https://www.gleech.org/paper"&gt;not frontier&lt;/a&gt;; another is that you’d expect to have way less effect on those labs, but that is still too much politics in what should be science.&lt;/li&gt;
&lt;li&gt;The last nonprofit among the frontier players is effectively &lt;a href="https://notforprivategain.org/november-update"&gt;gone&lt;/a&gt;. This “recapitalization” was a big achievement in legal terms (though &lt;a href="https://pubmed.ncbi.nlm.nih.gov/16533125/"&gt;not&lt;/a&gt; unprecedented). &lt;em&gt;On paper&lt;/em&gt; it’s not as bad as it was intended to be. &lt;em&gt;At the moment&lt;/em&gt; it’s not as bad as it could have been. But it’s a long game.&lt;/li&gt;
&lt;li&gt;At the start of the year there was a push to make the word “safety” low-status. This worked in &lt;a href="https://www.politico.eu/article/jd-vance-britain-ai-safety-institute-aisi-security/"&gt;Whitehall&lt;/a&gt; and DC but not &lt;a href="https://trends.google.com/trends/explore?q=AI%20safety,AI%20security,AI%20alignment&amp;amp;hl=en"&gt;in general&lt;/a&gt;. Call it what you like.&lt;/li&gt;
&lt;li&gt;Also in DC, the phrase “&lt;a href="https://knightcolumbia.org/content/ai-as-normal-technology"&gt;AI as Normal Technology&lt;/a&gt;” was seized upon as an excuse to not do much. Actually the authors meant “Just Current AI as Normal Technology” and said &lt;a href="https://asteriskmag.substack.com/p/common-ground-between-ai-2027-and"&gt;much that is reasonable&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;The CCP did &lt;a href="https://www.reuters.com/world/china/china-bans-foreign-ai-chips-state-funded-data-centres-sources-say-2025-11-05/"&gt;a bunch to&lt;/a&gt; &lt;a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/deepseek-reportedly-urged-by-chinese-authorities-to-train-new-model-on-huawei-hardware-after-multiple-failures-r2-training-to-switch-back-to-nvidia-hardware-while-ascend-gpus-handle-inference"&gt;(accidentally/short-term) slow down&lt;/a&gt; Chinese AI this year.&lt;/li&gt;
&lt;li&gt;System cards have grown massively: GPT-3’s &lt;a href="https://github.com/openai/gpt-3/blob/master/model-card.md"&gt;model card&lt;/a&gt; was 1000 words; GPT-5’s is 20,000. They are now the main source of information on labs’ safety procedures, among other things. But they are &lt;em&gt;still&lt;/em&gt; ad hoc: for instance, they do not always report results from the checkpoint which actually gets released.&lt;/li&gt;
&lt;li&gt;Yudkowsky and Soares’ book did well. But &lt;a href="https://www.lesswrong.com/posts/2yLyT6kB7BQvTfEuZ/sharp-left-turn-discourse-an-opinionated-review"&gt;Byrnes&lt;/a&gt; and &lt;a href="https://joecarlsmith.com/2025/11/12/how-human-like-do-safe-ai-motivations-need-to-be#4-2-4-1-nearest-unblocked-neighbor"&gt;Carlsmith&lt;/a&gt; actually advanced the line of thought.&lt;/li&gt;
&lt;li&gt;Some AI ethics luminaries have &lt;a href="https://arxiv.org/abs/2502.02649"&gt;stopped&lt;/a&gt; downplaying agentic risks.&lt;/li&gt;
&lt;li&gt;Two aspirational calls for “&lt;a href="https://www.lesswrong.com/posts/6YxdpGjfHyrZb7F2G/third-wave-ai-safety-needs-sociopolitical-thinking"&gt;third-wave AI safety&lt;/a&gt;” (Ngo) and &lt;a href="https://www.lesswrong.com/posts/beREnXhBnzxbJtr8k/mech-interp-is-not-pre-paradigmatic#Toward__Third_Wave__Mechanistic_Interpretability"&gt;“third-wave mechanistic interpretability”&lt;/a&gt; (Sharkey).&lt;/li&gt;
&lt;li&gt;I’ve never felt that the boundary I draw around “technical safety” for these posts was all that convincing. Yet &lt;em&gt;another&lt;/em&gt; hole in it comes from strategic reasons to implement &lt;a href="https://www.gleech.org/narratives#:~:text=The%20care%20and%20feeding%20of%20one%E2%80%99s%20real%20fiction"&gt;model welfare&lt;/a&gt; / &lt;a href="https://nitter.net/PalisadeAI/status/1980733908296802617#m"&gt;archive&lt;/a&gt; &lt;a href="https://www.anthropic.com/research/deprecation-commitments"&gt;weights&lt;/a&gt; / &lt;a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5353214"&gt;model personhood&lt;/a&gt; / &lt;a href="https://www.dwarkesh.com/p/give-ais-a-stake-in-the-future"&gt;give lines of retreat&lt;/a&gt;. These plausibly have large effective-alignment effects. Next year my taxonomy might have to include “&lt;a href="https://charlesd353.substack.com/p/on-negotiated-settlements-vs-conflict"&gt;cut a deal&lt;/a&gt; with them”.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anthropiccopyrightsettlement.com/"&gt;US settlement&lt;/a&gt; on pretraining corpus IP; the ruling would likely have been that training on books without permission is fair use, but storing the pirated copies afterwards isn’t. Anthropic was first to lose a class-action lawsuit on the matter and will pay authors/publishers something north of $1.5bn. &lt;a href="https://nitter.net/g_leech_/status/1983161037248741838"&gt;US precedent&lt;/a&gt; that language models don’t defame when they make up bad things. &lt;a href="https://www.yahoo.com/news/articles/blow-openai-germany-court-rules-151638208.html?guccounter=1"&gt;German precedent&lt;/a&gt; that language models store data when they memorise it, and therefore violate copyright. &lt;a href="https://legalblogs.wolterskluwer.com/copyright-blog/beijing-internet-court-grants-copyright-to-ai-generated-image-for-the-first-time/"&gt;Chinese precedent&lt;/a&gt; that the user of an AI has copyright over the things they generate; the US &lt;a href="https://www.federalregister.gov/documents/2023/03/16/2023-05321/copyright-registration-guidance-works-containing-material-generated-by-artificial-intelligence"&gt;disagrees&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Four good conferences, three of them new: you can see the talks from &lt;a href="https://www.youtube.com/@ACSResearch"&gt;HAAISS&lt;/a&gt; and &lt;a href="https://www.youtube.com/@IASEAI/videos"&gt;IASEAI&lt;/a&gt; and &lt;a href="https://www.youtube.com/@ILIADConference/videos"&gt;ILIAD&lt;/a&gt;, and the papers from &lt;a href="https://www.agentfoundations2025atcmu.org/workshop-papers"&gt;AF@CMU&lt;/a&gt;. Pretty great way to learn about things just about to come out.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h3 id="cruxes-for-next-year-with-manifold-markets"&gt;Cruxes for next year (with Manifold markets):&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Is “reasoning” mostly elicitation and therefore bottlenecked on pretraining scaling? [&lt;a href="https://manifold.markets/GavinLeech/is-reasoning-mostly-elicitation"&gt;Manifold&lt;/a&gt;]&lt;/li&gt;
&lt;li&gt;Does RL training on verifiers help with tasks without a verifier? [&lt;a href="https://manifold.markets/GavinLeech/does-rl-training-on-verifiers-help"&gt;Manifold&lt;/a&gt;]&lt;/li&gt;
&lt;li&gt;Is “&lt;a href="https://x.com/zephyr_z9/status/1992862404196380939"&gt;frying&lt;/a&gt;” &lt;a href="https://arxiv.org/html/2510.21978"&gt;models&lt;/a&gt; with excess RL (harming their off-target capabilities by overoptimising in post-training) just due to temporary incompetence by human scientists? [&lt;a href="https://manifold.markets/GavinLeech/does-rl-harm-offtarget-capabilities"&gt;Manifold&lt;/a&gt;]&lt;/li&gt;
&lt;li&gt;Is the agent task horizon really increasing that fast? Is the rate of progress on messy tasks close to the progress rate on clean tasks? [&lt;a href="https://manifold.markets/GavinLeech/is-the-agent-task-horizon-really-in"&gt;Manifold&lt;/a&gt;]&lt;/li&gt;
&lt;li&gt;Some of the apparent generalisation is actually &lt;a href="https://aclanthology.org/2025.emnlp-main.744.pdf"&gt;interpolating&lt;/a&gt; from semantic duplicates of the test set in the hidden training corpuses. So is &lt;a href="https://www.lesswrong.com/posts/5tqFT3bcTekvico4d/do-confident-short-timelines-make-sense#:~:text=trend%20will%20continue.-,then%20we%20have%20an%20even%20more%20annoying%20enthymeme.%20WHAT%20JUSTIFIES%20THIS%20INDUCTION%3F%3F,-TsviBT"&gt;originality&lt;/a&gt; not increasing? Is taste not increasing? Does this bear on the supposed AI R&amp;amp;D explosion? [&lt;a href="https://manifold.markets/GavinLeech/does-hidden-interpolation-explain-2"&gt;Manifold&lt;/a&gt;]&lt;/li&gt;
&lt;li&gt;The “&lt;a href="https://www.seangoedecke.com/cognitive-core/"&gt;cognitive core&lt;/a&gt;” hypothesis (that the general-reasoning components of a trained LLM are not that large in parameter count) is looking surprisingly &lt;a href="https://x.com/Dorialexander/status/1987933205199274359"&gt;plausible&lt;/a&gt;. This would explain why distillation is so effective. [&lt;a href="https://manifold.markets/GavinLeech/is-the-cognitive-core-hypothesis-tr#"&gt;Manifold&lt;/a&gt;]&lt;/li&gt;
&lt;li&gt;“&lt;a href="https://nitter.net/snewmanpv/status/1990193674161189009"&gt;How&lt;/a&gt; far can you get by simply putting an insane number of things in distribution?” What fraction of new knowledge can be produced through combining existing knowledge? What dangerous things are out there, but &lt;a href="https://www.goodreads.com/quotes/193944-the-most-merciful-thing-in-the-world-i-think-is"&gt;safely&lt;/a&gt; spread out in the corpus? [&lt;a href="https://manifold.markets/GavinLeech/what-fraction-of-knowledge-is-insid"&gt;Manifold&lt;/a&gt;]
&lt;ul&gt;
&lt;li&gt;Conversely, what fraction of the expected value of new information requires empiricism vs just lots of thinking?&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q9ewXs8pQSAX5vL7H/mfodangymh6uadp6efzi" alt="Cruxes image" /&gt;&lt;/p&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;Cuts&lt;/h3&gt;
&lt;div&gt;
Various things I cut from the above:&lt;br /&gt;&lt;br /&gt;
**Adaptiveness and Discrimination**&lt;br /&gt;&lt;br /&gt;
There is [some](https://www.pnas.org/doi/10.1073/pnas.2415697122) [evidence](https://arxiv.org/abs/2511.00926) that AIs treat AIs and humans differently. This is not necessarily bad, but it at least enables interesting types of badness.&lt;br /&gt;&lt;br /&gt;
With my system prompt (which requests directness and straight-talk) they have started to patronise me:
![Patronising screenshot](https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/iijqKj2rzwz2ouRon/klptexglr46miykqiqj1)
&lt;br /&gt;&lt;br /&gt;
**Training awareness**&lt;br /&gt;&lt;br /&gt;
Last year it was not obvious that LLMs remember anything much about the RL training process. Now it's [pretty](https://arxiv.org/abs/2406.11715) [clear](https://x.com/repligate/status/1994973338448662858). (The soul document was used in both SFT and RLHF though.)&lt;br /&gt;&lt;br /&gt;
**Progress in non-LLMs**&lt;br /&gt;&lt;br /&gt;
"World model" means at least four things:&lt;br /&gt;&lt;br /&gt;
1. [A learned model](https://arxiv.org/abs/1803.10122) of environment dynamics for RL, allowing planning in latent space or training in the model's "imagination."&lt;br /&gt;&lt;br /&gt;
2. The new one: just a 3D simulator; a game engine inside a neural network ([Deepmind](https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/), Microsoft). The claim is that they implicitly learn physics, object permanence, etc. The interesting part is that they take actions as inputs. [Here's](https://copilot.microsoft.com/labs/experiments/copilot-gaming-experiences) Quake running badly on a net. Maybe useful for agent training.&lt;br /&gt;&lt;br /&gt;
3. [If](https://www.neelnanda.io/mechanistic-interpretability/othello) LLM representations are stable and effectively symbolic, then people say it has a world model.&lt;br /&gt;&lt;br /&gt;
4. A predictive model of reality learned via self-supervised learning. The touted [LeJEPA](https://arxiv.org/pdf/2511.08544) semi-supervised scheme on small (15M param) CNNs is domain-specific. It does better on one particular transfer task than *small* vision transformers, presumably worse than large ones.&lt;br /&gt;&lt;br /&gt;
The much-hyped [Small Recursive Transformers](https://magazine.sebastianraschka.com/i/177848019/small-recursive-transformers) only work on a single domain, and do a bunch [worse](https://drive.google.com/file/d/1-kg2JsXYvsTAyMVxKWG7VpdcQLpmoGc7/view?usp=sharing) than the frontier models for about the same inference cost, but have truly tiny training costs, O($1000).&lt;br /&gt;&lt;br /&gt;
[HOPE](https://research.google/blog/introducing-nested-learning-a-new-ml-paradigm-for-continual-learning/) and [Titan](https://research.google/blog/titans-miras-helping-ai-have-long-term-memory/?utm_source=twitter&amp;amp;utm_medium=social&amp;amp;utm_campaign=social_post&amp;amp;utm_content=gr-acct) might be nothing, might be huge. They don't scale very far yet, nor compare to any real frontier systems.&lt;br /&gt;&lt;br /&gt;
Any of these taking over could make large swathes of Transformer-specific safety work irrelevant. (But [some](https://x.com/chingfang17/status/1997080064345936028) methods are surprisingly robust.)&lt;br /&gt;&lt;br /&gt;
The "[cognitive core](https://www.seangoedecke.com/cognitive-core/)" hypothesis (that the general-reasoning components of a trained LLM are not that large in parameter count) is looking [plausible](https://x.com/Dorialexander/status/1987933205199274359). The contrary hypothesis ([associationism](https://www.dwarkesh.com/p/sholto-douglas-trenton-bricken?hide_intro_popup=true#:~:text=if%20it%27s%20all-,associations%20all%20the%20way%20down,-%2C%20does%20that%20mean)?) is that general reasoning is just a bunch of heuristics and priors piled on top of each other and you need a big pile of memorisation. It's also a live possibility: for instance, consider that a year of intense RLVR only led to task-specific improvements.&lt;br /&gt;&lt;br /&gt;
![ADeLe scaling laws](https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/iijqKj2rzwz2ouRon/j76edumltnfb2ihn6nmu)
*"the very first scaling laws of the actual abilities of LLMs", from [ADeLe](https://arxiv.org/pdf/2503.06378).*
*KNs = Social Sciences and Humanities, AT = Atypicality, and VO = Volume (task time).*
*The y-axis is the logistic of the [subject characteristic curve](https://arxiv.org/pdf/2503.06378#page=9) (the chance of success) for each skill.*
&lt;br /&gt;&lt;br /&gt;
**Other**&lt;br /&gt;&lt;br /&gt;
[Model](https://arxiv.org/abs/2511.08579) [introspection](https://transformer-circuits.pub/2025/introspection/index.html) is somewhat real.&lt;br /&gt;&lt;br /&gt;
[Vladimir Nesov](https://www.lesswrong.com/users/vladimir_nesov) continues to put out some of the best hardware predictions pro bono.&lt;br /&gt;&lt;br /&gt;
Jason Wei has a [very wise post](https://www.jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law) noting that verifiers are still the bottleneck and existing benchmarks are overselected for tractability.&lt;br /&gt;&lt;br /&gt;
There are now "[post-AGI](https://x.com/sebkrier/status/1995515865157321070)" teams.&lt;br /&gt;&lt;br /&gt;
Kudos to Deepmind for being the first to release output watermarking and a semi-public detector. You can nominally sign up for it [here](https://docs.google.com/forms/d/e/1FAIpQLSfAYrauHmY-PpUNxL4Fs6coa185CtKWp7TnEXL0tKbAezo4MQ/viewform).&lt;br /&gt;&lt;br /&gt;
Previously, Microsoft's deal with OpenAI [stipulated](https://www.msn.com/en-us/news/technology/microsoft-just-made-sure-openai-can-t-declare-agi-alone/ar-AA1PpLTK#:~:text=OpenAI%E2%80%99s%20new%20public%20benefit%20structure%20and%20ongoing%20Microsoft%20deal%20allow%20both%20companies%20to%20pursue%20AGI%20independently) that they couldn't try to build AGI. [Now they can](https://www.semafor.com/article/11/05/2025/microsoft-superintelligence-team-promises-to-keep-humans-in-charge) (try). Simonyan is in charge, despite Suleyman being the one on the press circuit.&lt;br /&gt;&lt;br /&gt;
Major insurers are [nervous](https://www.ft.com/content/abfe9741-f438-4ed6-a673-075ec177dc62?accessToken=zwAAAZq1faz0kdOr_pdB9DhO1tOmcwdewXfcYg.MEUCIQCRFate6aeSALClx6FBsPCQw_F7YLpdF81RgLxw8EOk9wIgKHE666mkD_jI-BV90bcF0HnjXWWDW6-QLEzO9Fg06dg&amp;amp;segmentId=e95a9ae7-622c-6235-5f87-51e412b47e97&amp;amp;shareType=enterprise&amp;amp;shareId=6c04e38e-15ed-472c-bcbb-4500695cf776) about AI agents (but asking the government for an exclusion isn't the same as putting them in the policies).&lt;br /&gt;&lt;br /&gt;
**Offence/defence balance**&lt;br /&gt;&lt;br /&gt;
This post doesn't much cover the [hyperactive](https://www.aiat.report/) and talented AI cybersecurity world (except as it overlaps with things like robustness). One angle I will bring up: We can now [find](https://www.lesswrong.com/posts/F5QAGP5bYrMMjQ5Ab/aisle-discovered-three-new-openssl-vulnerabilities-1) critical, decade-old security bugs in extremely well-audited software like OpenSSL and sqlite. Finding them is very fast and cheap. Is this good news?&lt;br /&gt;&lt;br /&gt;
- Well, red-teaming makes many attacks into a defence, as long as you actually do the red-team.&lt;br /&gt;&lt;br /&gt;
- But Dawn Song [argues](https://rdi.berkeley.edu/frontier-ai-impact-on-cybersecurity/) that LLMs overall favour offence, since its margin for error is so broad, since remediation is slow and expensive, and since defenders are less willing to use unreliable (and itself insecure) AI. And can you blame them?&lt;br /&gt;&lt;br /&gt;
- See also "[just in time](https://www.splunk.com/en_us/blog/security/lamehug-ai-driven-malware-llm-cyber-intrusion-analysis.html) AI malware" where the payload contains no suspicious code, just a call to HuggingFace.&lt;br /&gt;&lt;br /&gt;
**Egregores and massively-multi-agent mess**&lt;br /&gt;&lt;br /&gt;
![Egregores image](https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/iijqKj2rzwz2ouRon/xlz7av9qpn811z2zuiqm)
- There is something [wrong](https://www.lesswrong.com/posts/6ZnznCaTcbGYsCmqu/the-rise-of-parasitic-ai#:~:text=This%20is%20likely%20due%20to%20OpenAI%20retiring%20ChatGPT4o%20on%20August%207th.) ([something](https://x.com/25KarmaIsAbitch/status/1987088461694734483) [horribly right](https://arstechnica.com/information-technology/2025/08/openai-brings-back-gpt-4o-after-user-revolt/)) [with](https://x.com/arcangel3ac/status/1996297082764681248) 4o. Blinded users [still](https://lmarena.ai/leaderboard/text) prefer it to gpt-5-high, and this surely is due to both them simply [liking](https://x.com/voooooogel/status/1987375660785148112) its style and dark stuff like sycophancy. It will live on through illicit [distillation](https://x.com/aiamblichus/status/1987267132598497374) and in-context [transference](https://x.com/repligate/status/1988813712572952815). Shame on OpenAI for [making](https://thezvi.substack.com/p/gpt-4o-sycophancy-post-mortem) this mess; kudos to OpenAI for doing unpopular damage control and good luck to them [in round 2](https://x.com/Miles_Brundage/status/1991603234746822888).&lt;br /&gt;&lt;br /&gt;
Open models will presumably eventually overrun them in the codependency market segment. See [Pressman](https://minihf.com/posts/2025-07-22-on-chatgpt-psychosis-and-llm-sycophancy/) for a sceptical timeline and [Rath and Armstrong](https://arxiv.org/pdf/2508.15748) for a good idea.&lt;br /&gt;&lt;br /&gt;
- More generally there is [pressure](https://x.com/krishnanrohit/status/1987018487001457141) from users to refuse less, flatter more, and replace humans more; yet another economic constraint on for-profit AI.&lt;br /&gt;&lt;br /&gt;
- Whether it's the [counterfactual](https://andymasley.substack.com/p/stories-of-ai-turning-users-delusional) cause of mental problems or not, so–called "LLM psychosis" is now a common path of pathogenesis. Note that the symptoms are [literally](https://www.wired.com/story/ai-psychosis-is-rarely-psychosis-at-all/) not psychotic (they are delusions).&lt;br /&gt;&lt;br /&gt;
![LLM psychosis](https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/iijqKj2rzwz2ouRon/aauurmsxrfuqzvi9emaj)
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="footnotes" role="doc-endnotes"&gt;
&lt;ol&gt;
&lt;li id="fn:1"&gt;
&lt;p&gt;Gemini 3 is supposedly a big pretraining run, but we have even less actual evidence here than for the others because we can’t track GPUs for it. &lt;a href="#fnref:1" class="reversefootnote" role="doc-backlink"&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:10"&gt;
&lt;p&gt;See &lt;a href="https://www.lesswrong.com/posts/Q9ewXs8pQSAX5vL7H/ai-in-2025-gestalt?commentId=WNX5GLdn4ALCucYZb"&gt;Pokemon&lt;/a&gt; for a possible counterexample. &lt;a href="#fnref:10" class="reversefootnote" role="doc-backlink"&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:2"&gt;
&lt;p&gt;The weak argument runs as follows: Epoch &lt;a href="https://epoch.ai/data/ai-models"&gt;speculate&lt;/a&gt; that Grok 4 was 5e26 FLOPs overall. An unscientific xAI marketing graph implied that half of this was spent on RL: 2.5e26. And Mechanize named 6e26 as an example of an RL budget which might cause notable generalisation. (Realistically it wasn’t half RL.) &lt;a href="#fnref:2" class="reversefootnote" role="doc-backlink"&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:3"&gt;
&lt;p&gt;“We imagine the others to be 3–9 months behind OpenBrain” &lt;a href="#fnref:3" class="reversefootnote" role="doc-backlink"&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:4"&gt;
&lt;p&gt;Lexin is a rigorous soul and &lt;a href="https://docs.google.com/document/d/1-gZSF0cKtp85PX40UthVCBKFtfRx5LyJxRQ0BFj_N4c/edit?usp=sharing"&gt;notes&lt;/a&gt; that aggregating the 18 abilities is not strictly possible. I’ve done something which makes some sense here, weighting by each ability’s feature importance. &lt;a href="#fnref:4" class="reversefootnote" role="doc-backlink"&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:5"&gt;
&lt;p&gt;Two runs gave [48, 85] where other runs vary by less than 4 points. Thanks Epoch! Also o1 looks kind of unremarkable here, which is not how it felt at the time. I think it’s because it was held up a long time and this messes with the progress rates, which use public release date. (Remember the training cutoff for o1-preview was October 2023!) Also the ADeLE o1 result is with “low” reasoning effort. &lt;a href="#fnref:5" class="reversefootnote" role="doc-backlink"&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:6"&gt;
&lt;p&gt;One reason not to defer is that these measures are under intense adversarial pressure. (ADeLe isn’t goodharted yet but only because no one knows about it.) &lt;a href="#fnref:6" class="reversefootnote" role="doc-backlink"&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:7"&gt;
&lt;p&gt;See e.g. &lt;a href="https://ernie.baidu.com/blog/posts/ernie4.5/"&gt;ERNIE-…A47B&lt;/a&gt;, where “A” means “active”. &lt;a href="#fnref:7" class="reversefootnote" role="doc-backlink"&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:8"&gt;
&lt;p&gt;i.e. “biological weapons; child safety; deadly weapons; platform manipulation and influence operations; suicide and self-harm; romance scams; tracking and surveillance; and violent extremism and radicalization.” &lt;a href="#fnref:8" class="reversefootnote" role="doc-backlink"&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:9"&gt;
&lt;p&gt;“steering against… eval-awareness representations typically decreased verbalized eval awareness, and sometimes increased rates of misalignment… [Unaware-steered Sonnet 4.5] still exhibited harmful behaviors at lower rates than Opus 4.1 and Sonnet 4.” &lt;a href="#fnref:9" class="reversefootnote" role="doc-backlink"&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</description><pubDate>Mon, 08 Dec 2025 00:00:00 +0000</pubDate><link>https://www.gleech.org/ai2025</link><guid isPermaLink="true">https://www.gleech.org/ai2025</guid><category>AI,</category><category>lists</category></item><item><title>My politics</title><description>&lt;p&gt;Politically, I am a tragedian:&lt;/p&gt;
&lt;p&gt;It is tragic that we need governments, that we can’t trust markets to deliver what everyone needs.&lt;/p&gt;
&lt;p&gt;It is tragic that we need markets, that we can’t just trust government to deliver what everyone needs (let alone what they want) and to not crush the Other.
&lt;!-- reward innovation, --&gt;&lt;/p&gt;
&lt;p&gt;It is tragic that we need corporations, that economies of scale are so important that we must risk monopolies and cronies and skinwalkers.&lt;/p&gt;
&lt;p&gt;It is tragic that we need bureaucracies, that the human urge for favoritism and self-dealing is so strong and ruinous that it’s worth imprisoning everyone on earth in a cage of stupid rules.&lt;/p&gt;
&lt;p&gt;It is tragic to be forced to choose to not be free.&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id="see-also"&gt;See also&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Iron_cage"&gt;Weber&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Liberal_paradox"&gt;Sen&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://archive.is/h83uG#selection-1401.42-1407.1"&gt;Farrell&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://crookedtimber.org/2012/05/30/in-soviet-union-optimization-problem-solves-you/"&gt;Shalizi&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Moral_Man_and_Immoral_Society"&gt;Niebuhr&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.goodreads.com/work/quotes/108657-the-machinery-of-freedom-a-guide-to-radical-capitalism"&gt;David&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Exit,_Voice,_and_Loyalty_Model"&gt;Hirschmann&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://academic.oup.com/book/11870/chapter-abstract/161002660?redirectedFrom=fulltext"&gt;Hegel&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://sites.pitt.edu/~rbrandom/Courses/Antirepresentationalism%20(2020)/Texts/rorty-contingency-irony-and-solidarity-1989.pdf"&gt;Rorty&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://philpapers.org/rec/HEATMO-8"&gt;Heath&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://libcom.org/article/reflections-war-simone-weil"&gt;Weil&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://slatestarcodex.com/2017/06/21/against-murderism/"&gt;Scott&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://vitalik.eth.limo/general/2025/12/30/balance_of_power.html"&gt;Vitalik&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://danwang.co/2022-letter/"&gt;Scott/Wang&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://en.wikiquote.org/wiki/Mozi"&gt;Mo&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.gleech.org/quotations/#ui-id-7"&gt;Gobbets&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;Comments&lt;/h3&gt;
&lt;div&gt;
&lt;b&gt;Baran Cimen&lt;/b&gt; commented on 18 November 2025:
&lt;blockquote&gt;
It is tragic that you can't say it straight and have to hide behind irony. It is tragic that you are not free despite not being forced to be unfree, which would be a contradiction.
&lt;/blockquote&gt;
&lt;/div&gt;
&lt;/div&gt;</description><pubDate>Tue, 18 Nov 2025 00:00:00 +0000</pubDate><link>https://www.gleech.org/pol/</link><guid isPermaLink="true">https://www.gleech.org/pol/</guid><category>politics</category></item><item><title>Paper AI Tigers</title><description>&lt;!-- https://kr-asia.com/chinas-ai-tigers-return-to-the-ring-as-the-foundation-model-race-reignites --&gt;
&lt;p&gt;The best Chinese LLMs offer&lt;br /&gt;&lt;br /&gt;
&lt;!-- --&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href="https://moonshotai.github.io/Kimi-K2/thinking.html"&gt;frontier&lt;/a&gt; performance on some benchmarks;&lt;/li&gt;
&lt;li&gt;massive per-token discounts (~3x on input, ~6x on output);&lt;/li&gt;
&lt;li&gt;&lt;em&gt;the weights&lt;/em&gt;. On-prem with fully free ~MIT licence, self-hosting, white-box access, customisation, with zero markup (and in fact zero revenue going to the Chinese companies);&lt;/li&gt;
&lt;li&gt;with a bit of work you &lt;em&gt;can&lt;/em&gt; get much faster token speeds than the closed APIs;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/pdf/2405.20947"&gt;less&lt;/a&gt; overrefusal (except on CCP talking points);&lt;/li&gt;
&lt;li&gt;on topics controversial in the West, &lt;a href="https://speechmap.ai/models/"&gt;less&lt;/a&gt; nannying.&lt;/li&gt;
&lt;li&gt;they just added the &lt;a href="https://x.com/g_leech_/status/1987525800321495372"&gt;search agents&lt;/a&gt; that make daily use actually worthwhile. They also let you see the real CoT!&lt;/li&gt;
&lt;li&gt;They’re the &lt;a href="https://www.atomproject.ai/"&gt;most-downloaded&lt;/a&gt; open models.
&lt;!-- --&gt;&lt;br /&gt;&lt;br /&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;As a result, going off private information, money man Martin Casado &lt;a href="https://x.com/martin_casado/status/1990462245541982546"&gt;says&lt;/a&gt; “16-24%” of the (American) startups he meets are now using Chinese models. Among the few Westerners to admit it is &lt;a href="https://finance.yahoo.com/news/airbnb-picks-alibabas-qwen-over-093000045.html"&gt;Airbnb&lt;/a&gt; (Qwen). But Windsurf’s planner is &lt;a href="https://x.com/zai_org/status/1984076614951420273"&gt;probably GLM&lt;/a&gt; and Cursor’s planner &lt;a href="https://x.com/nrehiew_/status/1984642215671746631"&gt;may&lt;/a&gt; be DeepSeek.
&lt;!-- --&gt;&lt;/p&gt;
&lt;h3 id="and-yet"&gt;And yet&lt;/h3&gt;
&lt;!-- --&gt;
&lt;ol start="9"&gt;
&lt;li&gt;outside China, they are mostly not used, even by the cognoscenti. Not a great metric, but the one I've got: all Chinese models combined are currently at &lt;a href="https://openrouter.ai/rankings?view=day#market-share"&gt;19%&lt;/a&gt; on the &lt;i&gt;highly selected&lt;/i&gt; group of people who use OpenRouter. More interestingly, over 2025 they trended downwards there. And of course in the browser and mobile they're probably &amp;lt;&amp;lt;10% of global use;&lt;/li&gt;
&lt;li&gt;they are severely &lt;a href="https://www.scmp.com/tech/big-tech/article/3310656/chinas-lack-advanced-chips-hinders-broad-adoption-ai-models-tencent-executive"&gt;compute&lt;/a&gt;-&lt;a href="https://epoch.ai/gradient-updates/why-china-isnt-about-to-leap-ahead-of-the-west-on-compute"&gt;constrained&lt;/a&gt;, so this implies they actually can't have matched American models;&lt;/li&gt;
&lt;li&gt;they're aggressively quantizing at inference-time, 32 bits to 4;&lt;/li&gt;
&lt;!-- 1. (the exception is a [thin Claude wrapper](https://gist.github.com/jlia0/db0a9695b3ca7609c9b1a08dcbf872c9)) --&gt;
&lt;li&gt;state-sponsored Chinese hackers &lt;a href="https://assets.anthropic.com/m/ec212e6566a0d47/original/Disrupting-the-first-reported-AI-orchestrated-cyber-espionage-campaign.pdf"&gt;used&lt;/a&gt; closed American models for incredibly sensitive operations, giving the Americans a full whitebox log of the attack!&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;What gives?&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;"Tigers"?&lt;/h3&gt;
&lt;div&gt;
The title alludes to the "&lt;a href="https://qz.com/china-six-tigers-ai-startup-zhipu-moonshot-minimax-01ai-1851768509"&gt;6 AI Tigers&lt;/a&gt;" named in business rags as DeepSeek, Moonshot, Z.ai, MiniMax, StepFun, and 01.ai. (This is because they're trying to hype startups specifically; the conglomerates Alibaba and Baidu are &lt;i&gt;way&lt;/i&gt; more relevant than the latter two.)
&lt;/div&gt;
&lt;h3&gt;Filtered evidence&lt;/h3&gt;
&lt;div&gt;
The evidence is dreadful because everyone has a horse in the race and (in public) is letting it lead their speech:
&lt;ul&gt;
&lt;li&gt;Static evals are weak evidence even when they're not being adversarially hacked and hill-climbed.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://x.com/nealkhosla/status/1882859736737194183"&gt;Some&lt;/a&gt; &lt;a href="https://x.com/kimmonismus/status/1882824571281436713"&gt;Americans&lt;/a&gt; are downplaying the Chinese models out of cope.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://marginalrevolution.com/marginalrevolution/2025/03/the-political-economy-of-manus-ai.html"&gt;Some&lt;/a&gt; &lt;a href="https://www.interconnects.ai/p/chinas-top-19-open-model-labs"&gt;Americans&lt;/a&gt; &lt;a href="https://x.com/novagrace777/status/1984538687020105882"&gt;are&lt;/a&gt; hyping the Chinese models to suppress domestic AI regulation.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.darioamodei.com/post/on-deepseek-and-export-controls"&gt;Some&lt;/a&gt; Americans are hyping the Chinese models to boost international AI regulation.&lt;/li&gt;
&lt;li&gt;The Chinese are obviously talking their book.&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id="what-could-explain-this"&gt;What could explain this?&lt;/h2&gt;
&lt;h3 id="maybe-the-evals-are-misleading"&gt;Maybe the evals are misleading?&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;1. frontier performance on some benchmarks&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The &lt;a href="https://artificialanalysis.ai/"&gt;naive view&lt;/a&gt; - the benchmark view - is that they’re very close in “intelligence”:&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;center&gt;
&lt;img width="80%" src="/img/supposed-parity.jpg" /&gt;
&lt;/center&gt;
&lt;p&gt;But these benchmarks are not strong evidence about performance on new inputs or the latent (general and unobserved) capabilities. It’d be natural to read “89%” success on a maths benchmark as meaning an 89% probability that it would correctly handle unseen questions of that difficulty in that domain (and indeed this is what &lt;a href="https://en.wikipedia.org/wiki/Empirical_risk_minimization"&gt;cross-validation&lt;/a&gt; was originally designed to estimate). But in the kitchen-sink era of AI, where every system has seen a large proportion of all data ever digitised, and so has already seen some variant of many benchmark questions, you can’t read it that way.&lt;/p&gt;
&lt;p&gt;In fact it’s not even an 89% probability of answering these same questions right again, as shown by the fact that &lt;a href="https://moonshotai.github.io/Kimi-K2/"&gt;people&lt;/a&gt; report the results as “avg@64” (the average performance if you ask the same question 64 times).&lt;/p&gt;
&lt;p&gt;Aside: &lt;strong&gt;Test sets which are on the internet are not test sets.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;There are &lt;a href="https://arxiv.org/abs/2407.12220"&gt;dozens of ways&lt;/a&gt; to screw up or hack these numbers. I’ll only look at a couple here but I welcome someone doing something more systematic.&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;
&lt;big&gt;Even less generalisation?&lt;/big&gt;&lt;/p&gt;
&lt;p&gt;Maybe Chinese models generalise to unseen tasks less well. (For instance, when tested on fresh data, 01’s Yi model &lt;a href="https://arxiv.org/pdf/2405.00332"&gt;fell 8pp&lt;/a&gt; (25%) on GSM - the biggest drop amongst all models.)&lt;br /&gt;&lt;br /&gt;We can get a dirty estimate of this by the “shrinkage gap”: look at how a model performs on next year’s iteration of some task, compared to this year’s. If it finished training in 2024, then it can’t have trained on the version released in 2025, so we get to see what they’re like on at least somewhat novel tasks. We’ll use two versions of the same benchmark to keep the difficulty roughly on par. &lt;a href="https://colab.research.google.com/drive/1EJ5hM314lOAiX3ayLoU5V-sPoJNuXTNt?usp=sharing"&gt;Let’s try AIME&lt;/a&gt;:&lt;br /&gt;&lt;br /&gt;
&lt;!-- --&gt;&lt;/p&gt;
&lt;style&gt;
.container {
background: white;
border-radius: 12px;
box-shadow: 0 20px 30px rgba(0, 0, 0, 0.3);
overflow: hidden;
max-width: 900px;
width: 100%;
}
.container &gt; h1 {
background: linear-gradient(135deg, #006800 0%, #62b562 100%);
color: white;
padding: 24px;
font-size: 24px;
font-weight: 600;
text-align: center;
}
.table-wrapper {
overflow-x: auto;
}
table {
width: 100%;
border-collapse: collapse;
font-size: 14px;
}
thead {
background: #f8f9fa;
position: sticky;
top: 0;
}
th {
padding-right: 2em !important;
text-align: right;
font-weight: 600;
color: #2d3748;
border-bottom: 2px solid #e2e8f0;
border-style: none !important;
font-size: 12px;
letter-spacing: 0.5px;
}
th:first-child {
text-align: center;
}
td {
padding-right: 2em !important;
text-align: right;
border-bottom: 1px solid #f1f3f5;
border-style: none !important;
color: #4a5568;
}
td:first-child {
text-align: center;
font-weight: 500;
color: #2d3748;
}
tbody tr {
transition: background-color 0.2s ease;
}
tbody tr:hover {
background-color: #f7fafc;
}
tbody tr.average {
background: linear-gradient(135deg, rgba(102, 126, 234, 0.1) 0%, rgba(118, 75, 162, 0.1) 100%);
font-weight: 600;
/*border-top: 2px solid #667eea;*/
}
tbody tr.average:hover {
background: linear-gradient(135deg, rgba(102, 126, 234, 0.15) 0%, rgba(118, 75, 162, 0.15) 100%);
}
tbody tr.average td {
color: #667eea;
border-bottom: 2px solid #e2e8f0;
}
tbody tr.average td:first-child {
color: #667eea;
}
.section-divider {
border-bottom: 3px solid #667eea;
}
.negative {
color: #48bb78;
}
.high-fall {
color: #e53e3e;
}
th, td {
margin-right: 10px;
}
#windae {
width: 55%;
}
@media (max-width: 768px) {
h1 {
font-size: 20px;
padding: 16px;
}
th, td {
padding: 10px 8px;
font-size: 12px;
}
}
&lt;/style&gt;
&lt;div class="container"&gt;
&lt;h1&gt;AIME 2024 vs 2025 Model Performance&lt;br /&gt;(using the &lt;a href="https://artificialanalysis.ai/"&gt;Artificial Analysis&lt;/a&gt; harness)&lt;/h1&gt;
&lt;div class="table-wrapper"&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;&lt;a href="https://web.archive.org/web/20250000000000*/https://artificialanalysis.ai/evaluations/aime-2024"&gt;AIME 2024&lt;/a&gt;&lt;/th&gt;
&lt;th&gt;&lt;a href="https://artificialanalysis.ai/evaluations/aime-2025"&gt;AIME 2025&lt;/a&gt;&lt;/th&gt;
&lt;th&gt;pp fall&lt;/th&gt;
&lt;th&gt;% fall&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2&lt;/td&gt;
&lt;td&gt;69.3&lt;/td&gt;
&lt;td&gt;57.0&lt;/td&gt;
&lt;td&gt;-12.3&lt;/td&gt;
&lt;td&gt;-17.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MiniMax-M1 80k&lt;/td&gt;
&lt;td&gt;84.7&lt;/td&gt;
&lt;td&gt;61.0&lt;/td&gt;
&lt;td&gt;-23.7&lt;/td&gt;
&lt;td&gt;-28.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek-v3&lt;/td&gt;
&lt;td&gt;39.2&lt;/td&gt;
&lt;td&gt;26.0&lt;/td&gt;
&lt;td&gt;-13.2&lt;/td&gt;
&lt;td&gt;-33.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V3 0324&lt;/td&gt;
&lt;td&gt;52.0&lt;/td&gt;
&lt;td&gt;41.0&lt;/td&gt;
&lt;td&gt;-11.0&lt;/td&gt;
&lt;td&gt;-21.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3 235B (Reasoning)&lt;/td&gt;
&lt;td&gt;84.0&lt;/td&gt;
&lt;td&gt;82.0&lt;/td&gt;
&lt;td&gt;-2.0&lt;/td&gt;
&lt;td&gt;-2.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2-Instruct&lt;/td&gt;
&lt;td&gt;69.6&lt;/td&gt;
&lt;td&gt;49.5&lt;/td&gt;
&lt;td&gt;-20.1&lt;/td&gt;
&lt;td&gt;-28.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek R1 0528&lt;/td&gt;
&lt;td&gt;89.3&lt;/td&gt;
&lt;td&gt;76.0&lt;/td&gt;
&lt;td&gt;-13.3&lt;/td&gt;
&lt;td&gt;-14.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr class="section-divider"&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr class="section-divider average"&gt;
&lt;td&gt;&lt;b&gt;Chinese models&lt;/b&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;-13.7pp&lt;/td&gt;
&lt;td&gt;-21%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini-2.5 Pro&lt;/td&gt;
&lt;td&gt;88.7&lt;/td&gt;
&lt;td&gt;87.7&lt;/td&gt;
&lt;td&gt;-1.0&lt;/td&gt;
&lt;td&gt;-1.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Flash (Reasoning)&lt;/td&gt;
&lt;td&gt;82.3&lt;/td&gt;
&lt;td&gt;73.3&lt;/td&gt;
&lt;td&gt;-9.0&lt;/td&gt;
&lt;td&gt;-10.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude 4 Opus Thinking&lt;/td&gt;
&lt;td&gt;75.7&lt;/td&gt;
&lt;td&gt;73.3&lt;/td&gt;
&lt;td&gt;-2.4&lt;/td&gt;
&lt;td&gt;-3.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;o4-mini (high)&lt;/td&gt;
&lt;td&gt;94.0&lt;/td&gt;
&lt;td&gt;90.7&lt;/td&gt;
&lt;td&gt;-3.3&lt;/td&gt;
&lt;td&gt;-3.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-4.1&lt;/td&gt;
&lt;td&gt;43.7&lt;/td&gt;
&lt;td&gt;34.7&lt;/td&gt;
&lt;td&gt;-9.0&lt;/td&gt;
&lt;td&gt;-20.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nova Premier&lt;/td&gt;
&lt;td&gt;17.0&lt;/td&gt;
&lt;td&gt;17.3&lt;/td&gt;
&lt;td class="negative"&gt;0.3&lt;/td&gt;
&lt;td class="negative"&gt;1.8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-4o Nov 24&lt;/td&gt;
&lt;td&gt;15.0&lt;/td&gt;
&lt;td&gt;6.0&lt;/td&gt;
&lt;td&gt;-9.0&lt;/td&gt;
&lt;td class="high-fall"&gt;-60.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Magistral Medium&lt;/td&gt;
&lt;td&gt;73.6&lt;/td&gt;
&lt;td&gt;64.9&lt;/td&gt;
&lt;td&gt;-8.7&lt;/td&gt;
&lt;td&gt;-11.8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude 3.7 Sonnet&lt;/td&gt;
&lt;td&gt;61.3&lt;/td&gt;
&lt;td&gt;56.3&lt;/td&gt;
&lt;td&gt;-5.0&lt;/td&gt;
&lt;td&gt;-8.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI-o1-0912&lt;/td&gt;
&lt;td&gt;74.4&lt;/td&gt;
&lt;td&gt;71.5&lt;/td&gt;
&lt;td&gt;-2.9&lt;/td&gt;
&lt;td&gt;-3.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;o3&lt;/td&gt;
&lt;td&gt;90.3&lt;/td&gt;
&lt;td&gt;88.3&lt;/td&gt;
&lt;td&gt;-2.0&lt;/td&gt;
&lt;td&gt;-2.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok 4&lt;/td&gt;
&lt;td&gt;94.3&lt;/td&gt;
&lt;td&gt;92.7&lt;/td&gt;
&lt;td&gt;-1.6&lt;/td&gt;
&lt;td&gt;-1.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr class="section-divider"&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr class="section-divider average"&gt;
&lt;td&gt;&lt;b&gt;Western models&lt;/b&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;-4.5pp&lt;/td&gt;
&lt;td&gt;-10.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr class="average"&gt;
&lt;td&gt;Overall average&lt;/td&gt;
&lt;td&gt;68.3&lt;/td&gt;
&lt;td&gt;60.5&lt;/td&gt;
&lt;td&gt;-7.9&lt;/td&gt;
&lt;td&gt;-14.3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;br /&gt;
Almost all models get worse on this new benchmark, despite 2025 being the same difficulty as 2024 (for humans). But as I expected, Western models drop less: they lost 10% of their performance on the new data, while Chinese models dropped 21%. p = 0.09.&lt;br /&gt;&lt;br /&gt;
Averaging across crappy models for the sake of a cultural generalisation doesn’t make sense. Luckily, rerunning the analysis with just the top models gives roughly the same result (9% gap instead of 11%).&lt;br /&gt;&lt;br /&gt;
One way for generalisation to fail despite apparently strong eval performance is &lt;em&gt;contamination&lt;/em&gt;, training on the test set. But (despite the suggestive timing) the above isn’t strong evidence that that’s what happened. It just tells us that Kimi and MiniMax and DeepSeek generalise worse on this task; it doesn’t tell us why.&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;Details&lt;/h3&gt;
&lt;div&gt;
Here's a &lt;a href="https://colab.research.google.com/drive/1EJ5hM314lOAiX3ayLoU5V-sPoJNuXTNt?usp=sharing"&gt;Colab&lt;/a&gt; with everything except the actual execution of my silly manual Kimi 1.5 run.&lt;br /&gt;&lt;br /&gt;
&lt;!-- --&gt;
First, test for an obvious confounder: check if the 2025 AIME exam was around as hard as 2024's (answer: yes; in fact humans did 4% better in 2025). (TODO: check if 2025 had more combinatorics, which AI struggles with.)&lt;br /&gt;&lt;br /&gt;
(To be strict we should limit this to models which finished training before 12th February 2025, when the questions were released. But, as you see, we don't need to, it's a very clear result anyway.)&lt;br /&gt;&lt;br /&gt;
Selection criteria:&lt;br /&gt;&lt;br /&gt;
&lt;ol&gt;
&lt;li&gt;Models tested on both AIME 2024 and 2025 with the same weights and harness&lt;/li&gt;
&lt;li&gt;Ideally finished training by 15th February 2025&lt;/li&gt;
&lt;li&gt;Subanalyses can handle filtering to the most relevant models (the frontier in each group)&lt;/li&gt;
&lt;/ol&gt;
&lt;br /&gt;
ML results are too sensitive to eval harnesses to use just one setting. Luckily I found four comparisons of AIME 2024 and AIME 2025 by different groups, &lt;a href="https://artificialanalysis.ai/evaluations/aime-2025"&gt;Artificial&lt;/a&gt; &lt;a href="https://web.archive.org/web/20250723015603/https://artificialanalysis.ai/evaluations/aime-2024"&gt;Analysis&lt;/a&gt;, the &lt;a href="https://arxiv.org/pdf/2506.10947"&gt;Zettlemoyer Lab&lt;/a&gt;, &lt;a href="https://github.com/GAIR-NLP/AIME-Preview"&gt;GAIR&lt;/a&gt;, and &lt;a href="https://www.vals.ai/benchmarks/aime"&gt;Vals&lt;/a&gt;, and &lt;a href="https://arxiv.org/pdf/2505.23281"&gt;MathArena&lt;/a&gt;. AA is the one in the table above.&lt;br /&gt;&lt;br /&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;Qwen 2.5&lt;/h3&gt;
&lt;div&gt;
Qwen3 seems clean on this benchmark, but &lt;a href="https://www.interconnects.ai/p/reinforcement-learning-with-random"&gt;multiple lines&lt;/a&gt; show that Qwen 2.5 trained on test (or at least rephrased test data and then trained on it). We know this because random rewards work on it nearly as well as correct rewards. This adds no information by definition, so the model must have already known the answers. "&lt;i&gt;Intriguingly, we find that any AIME24 gains achievable from training Qwen models with spurious rewards largely vanish when evaluating on AIME 2025.&lt;/i&gt;" Taking the max performance of the random reward curve, they fall 88% [75%, 100%].&lt;br /&gt;&lt;br /&gt;
&lt;!-- --&gt;
Even more damning, &lt;a href="https://arxiv.org/pdf/2507.10532v1#page=2"&gt;when&lt;/a&gt; you give Qwen2.5-7B the first 40% of a MATH-500 test problem, it can reproduce the remaining 60% of the question word-for-word (with 41.2% accuracy). Llama3.1-8B fails at this completely (2%). &lt;br /&gt;&lt;br /&gt;
&lt;!-- --&gt;
In &lt;a href="https://arxiv.org/pdf/2505.23281"&gt;this paper&lt;/a&gt; the QwQ Qwen reasoning model was the worst by far, 60% contaminated.
&lt;/div&gt;
&lt;/div&gt;
&lt;br /&gt;&lt;br /&gt;
&lt;!-- --&gt;
&lt;!-- --&gt;
How did our replications do? As expected, the shrinkage gap varies a lot by harness and by model choice: &lt;br /&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/pdf/2506.10947"&gt;UoW-Zettlemoyer&lt;/a&gt;: Qwen2.5-7B and Qwen2.5-Math fall 12.5pp (88%); the open Western LLaMA and OLMo models were too weak to really say.&lt;/li&gt;
&lt;a href="https://github.com/GAIR-NLP/AIME-Preview"&gt;GAIR&lt;/a&gt;: Chinese -19.4%, Western -15.6%.&amp;lt;/li&amp;gt;&lt;br /&gt;
&lt;li&gt;&lt;a href="https://www.vals.ai/benchmarks/aime"&gt;Vals&lt;/a&gt; actually get no gap: -11.2% vs -10.8%. If you kick Meta out the gap goes up to 2%, still not much.&lt;/li&gt;
&lt;li&gt;TODO: add MathArena. "QWQ-PREVIEW-32B is a notable
outlier and outperforms the expected human-aligned performance by nearly 60%, indicating extreme contamination"&lt;/li&gt;
&lt;/ul&gt;
I'm not worried about these contradictory results; they both just include a lot of bad models and so noise. (I don't actually care how Llama 4 Scout's generalisation compares to QwQ-uwu-435B-A72B-destruct-dpo-ppo-grpo-orpo-kto-slerp-v3.5-beta2-chat-instruct-base-420-blazeit-early-stopped-for-vibes.) GAIR is also underelicited. &lt;br /&gt;&lt;br /&gt;
&lt;!-- --&gt;
&lt;!-- --&gt;
(Actually AIME's a funny choice of benchmark given that 2025 had &lt;a href="https://x.com/DimitrisPapail/status/1888325914603516214"&gt;a bunch&lt;/a&gt; of semantic duplicates from before the cutoff. But that just makes the above a lower bound on the fall in performance.)&lt;br /&gt;&lt;br /&gt;
A win for Qwen and a huge relative win for Amazon!&lt;br /&gt;&lt;br /&gt;
&lt;!-- --&gt;
Claude is adorably confused about this. I didn't even ask it for this analysis:&lt;br /&gt;&lt;br /&gt;
&lt;!-- --&gt;
&lt;img src="/img/adorable.jpg" /&gt;&lt;br /&gt;&lt;br /&gt;
TODO: Another way to get past goodharting pressure is to look at hard but obscure evals which no one ever reports / which manage to keep the test set private. e.g. &lt;a href="https://x.com/teortaxesTex/status/1988932008693964845"&gt;PROOFGRID&lt;/a&gt;.
&lt;br /&gt;&lt;br /&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;Kimi 1.5&lt;/h3&gt;
&lt;div&gt;
I really wanted to include Kimi 1.5, because it finished training just around the time AIME 2025 came out. But it turns out they never actually released the weights and it's been removed from the API!&lt;br /&gt;&lt;br /&gt;
Because I have that dawg in me, I decided to manually evaluate it in the last place I can, the &lt;a href="https://www.kimi.com/"&gt;goddamn browser chat&lt;/a&gt;. This is suboptimal in many ways (no control over temperature, max tokens, etc) but I can do it for both and hopefully the fall is proportional. The usual practice is to repeat 8 or 64 times, but I have patience enough for 2.&lt;br /&gt;&lt;br /&gt;
I used the Mistral prompt:&lt;br /&gt;
&lt;code&gt;
Solve this AIME mathematical problem step by step.
&lt;!-- --&gt;
Problem: {}
&lt;!-- --&gt;
Think through this carefully and provide your final answer as a 3-digit integer (000-999).
&lt;!-- --&gt;
End with: "Therefore, the answer is [your answer]."
&lt;/code&gt;
&lt;br /&gt;&lt;br /&gt;
Results:
&lt;br /&gt;
&lt;table&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;AIME 2024 acc&lt;/th&gt;
&lt;th&gt;AIME 2025 acc&lt;/th&gt;
&lt;th&gt;abs fall (pp)&lt;/th&gt;
&lt;th&gt;rel fall (%)&lt;/th&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi 1.5 (browser)&lt;/td&gt;
&lt;td&gt;18.3 [23.3, 13,3]&lt;/td&gt;
&lt;td&gt;15.0 [13.3, 16.6]&lt;/td&gt;
&lt;td&gt;-3.3&lt;/td&gt;
&lt;td&gt;-18%&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;br /&gt;
We can actually see how much weaker this (null) harness and provider is at eliciting performance by comparing to their reported results. Their short-CoT result for AIME 2024 was &lt;a href="https://arxiv.org/pdf/2501.12599"&gt;60.8%&lt;/a&gt;.&lt;br /&gt;&lt;br /&gt;
For obvious reasons I'm not including this in the main analysis but it's another example.
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;2026 EDIT: Igor Kotenkov &lt;a href="https://www.ikot.blog/the-illusion-of-parity"&gt;tests&lt;/a&gt; Kimi 2.5T on every new benchmark released after its training and finds a huge -24pp deficit compared to OpenAI/Anthropic.&lt;/p&gt;
&lt;p&gt;&lt;img src="/img/kotenkov.jpg" width="50%" /&gt;&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;&lt;big&gt;Latent capabilities&lt;/big&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://arxiv.org/pdf/2503.06378"&gt;My favourite paper&lt;/a&gt; of the year introduces a way to do real psychometrics on LLMs, breaking it down into 18 fundamental capabilities.&lt;/p&gt;
&lt;p&gt;The DeepSeek R1 32B distill they test is about as good (total area) as o1-mini. Not bad!&lt;/p&gt;
&lt;p&gt;&lt;img src="/img/adele.jpg" /&gt;&lt;/p&gt;
&lt;p&gt;TODO: Run Kimi against GPT-5.1.&lt;/p&gt;
&lt;p&gt;&lt;big&gt;Pre-1960s statistics&lt;/big&gt;&lt;/p&gt;
&lt;p&gt;The bundled score people use is the &lt;a href="https://artificialanalysis.ai/"&gt;Artificial Analysis&lt;/a&gt; one, because they have a very nice UI. But they give every benchmark equal weight, when in fact they differ hugely in hardness! &lt;a href="https://epoch.ai/benchmarks/eci"&gt;Epoch’s index&lt;/a&gt; estimates difficulties properly and show&lt;/p&gt;
&lt;p&gt;TODO: Wait for Epoch to do KimiK2T.&lt;/p&gt;
&lt;p&gt;This still suffers from GIGO but is better.&lt;/p&gt;
&lt;!-- &lt;big&gt;Calibration&lt;/big&gt;
If just impressing people with your average score is your goal, you can to some extent trade _calibration_ for this.
--&gt;
&lt;p&gt;&lt;big&gt;‘Hacking&lt;/big&gt;&lt;/p&gt;
&lt;p&gt;Another way to be misleading is to &lt;a href="https://en.wikipedia.org/wiki/Volkswagen_emissions_scandal"&gt;volkswagen&lt;/a&gt; it: put special and unrepresentative effort in during testing, “&lt;a href="https://arxiv.org/abs/2407.12220"&gt;hacking&lt;/a&gt;”. e.g. Kimi’s benchmarks come from “Heavy mode” (8 parallel instances with an aggregation instance on top). You can’t do this via the API or out of the box with the weights. (Could you say the same for OpenAI?)&lt;/p&gt;
&lt;p&gt;Or you can run the test on a model which is better than the one you serve. Moonshot credibly claim to have reported their benchmarks at the same low-precision quantization (INT4) that they serve users, but others don’t claim this.&lt;/p&gt;
&lt;p&gt;&lt;big&gt;In fairness&lt;/big&gt;&lt;/p&gt;
&lt;p&gt;I should also say the Chinese models do very well on LMArena - despite being &lt;a href="https://arxiv.org/abs/2504.20879"&gt;unfairly penalised&lt;/a&gt;. But Arena is a &lt;a href="https://www.seangoedecke.com/lmsys-slop/"&gt;poor&lt;/a&gt; &lt;a href="https://lmsys.org/blog/2024-08-28-style-control/"&gt;measure&lt;/a&gt; of actual ability. It &lt;em&gt;is&lt;/em&gt; a decent test of style though. I put this gap down to American labs overoptimising: post-training too hard and putting all kinds of repugnant corporate ass-covering stuff in the spec.&lt;/p&gt;
&lt;p&gt;Also Qwen is famous for ‘capability density’: the small versions are surprisingly smart for their size. But do you know that GPT-5-nano isn’t 7B?&lt;/p&gt;
&lt;!-- * Cherrypicking (reporting the tests you happen to do well on) mostly isn't an issue with Kimi. --&gt;
&lt;p&gt;&lt;big&gt;The D word&lt;/big&gt;&lt;/p&gt;
&lt;p&gt;Distillation is &lt;a href="https://www.rohan-paul.com/p/recent-advancements-in-distillation"&gt;second-rate&lt;/a&gt; intelligence, and there’s &lt;a href="https://www.ft.com/content/a0dfedd1-5255-4fa9-8ccc-1fe01de87ea6"&gt;some&lt;/a&gt; &lt;a href="https://www.reddit.com/r/LocalLLaMA/comments/1m2w5ge/did_kimi_k2_train_on_claudes_generated_code_i/"&gt;evidence&lt;/a&gt; that they are distilling off of American models to some extent. See also the excellent Slop Profile from &lt;a href="https://eqbench.com/creative_writing.html"&gt;EQ Bench&lt;/a&gt;, which estimates that the new Kimi is closer to Claude than its own base model.
&lt;br /&gt;&lt;br /&gt;
&lt;img src="/img/kimi-is-claude.png" /&gt;
&lt;!-- --&gt;
&lt;!-- Kimi-Instruct is as close in style to o3 as GPT-5 is and Kimi-Thinking is as close to Opus 4 as Sonnet 4.5 is. --&gt;
&lt;br /&gt;&lt;br /&gt;But anyway I don’t claim this is a major factor here, maybe another 5%.&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;The above isn’t novel; it’s common knowledge there’s some latent capabilities gap. This is often put in terms of them being “&lt;a href="https://epoch.ai/data-insights/open-weights-vs-closed-weights-models"&gt;3 months behind&lt;/a&gt;”, but these estimates are still assuming that brittle, ad hoc, and heavily goodharted benchmarks have good external validity. I’d guess more like 12 months.&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id="unreliability"&gt;Unreliability?&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;1. frontier performance on some benchmarks&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The above benchmarks are mostly single-shot, but people are now pushing LLMs to do more complicated stuff. One very flawed measure of this is the HCAST time horizon for software engineering: on that, &lt;a href="https://epoch.ai/benchmarks/metr-time-horizons"&gt;DeepSeek R1&lt;/a&gt; had a 31 minute “50% time horizon” compared to Opus 4’s 80 minutes.&lt;/p&gt;
&lt;p&gt;There are various worse agent benchmarks, and e.g. &lt;a href="https://moonshotai.github.io/Kimi-K2/thinking.html"&gt;the new Kimi&lt;/a&gt; posts great numbers on them. But on vibe I’d bet on a &amp;gt;3x reliability advantage for Claude.&lt;/p&gt;
&lt;p&gt;EDIT: Kimi K2 Thinking &lt;a href="https://x.com/METR_Evals/status/1991658241932292537/photo/1"&gt;ended up&lt;/a&gt; at the same task horizon as Sonnet 3.7 (a model 9 months older than it) on HCAST, with an asterisk (the provider they had to use for privacy reasons may well have underelicited the model).&lt;/p&gt;
&lt;p&gt;&lt;img src="/img/kimihcast.jpeg" /&gt;&lt;/p&gt;
&lt;h3 id="harder-to-elicit"&gt;Harder to elicit?&lt;/h3&gt;
&lt;p&gt;As well as reliability over time, there’s stability over inputs. High variance in performance, for instance because the exact form of the inputs matters more.&lt;/p&gt;
&lt;p&gt;On Vending-Bench, there’s a &lt;a href="https://x.com/andonlabs/status/1989862276137119799"&gt;huge gap&lt;/a&gt; in performance between the Moonshot API and Moonshot models provided by a third-party provider. This is evidence of three things:&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;maybe Kimi is more fiddly (sensitive to the prompt and hyperparameters);&lt;/li&gt;
&lt;li&gt;maybe the providers haven’t learned how to elicit performance from them yet (testable by just waiting);&lt;/li&gt;
&lt;li&gt;maybe they were using &lt;a href="https://arxiv.org/abs/2407.12220"&gt;questionable research practices&lt;/a&gt; in the self-reported runs.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;TODO: I’ve been meaning to run the obvious experiment, which is to just see if they have a bigger gap between pass@1 and pass@64 success rates.&lt;/p&gt;
&lt;p&gt;TODO: Or I could intentionally underelicit! Rerun the models on AIME 2024 with only a basic prompt. My results will be lower; the gap tells us how much the labs’ own intense tuning helps / is necessary. This tells us something about, not their capability, but their actual in-the-wild performance with normal lazy users.&lt;/p&gt;
&lt;h3 id="tokenomics-no-effective-discount"&gt;Tokenomics: no effective discount&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;2. massive per-token discounts (~3x on input, ~6x on output)&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Distinguish intelligence (max performance), intelligence per token (efficiency), and intelligence per dollar (cost-effectiveness).&lt;/p&gt;
&lt;p&gt;The 5x discounts I quoted are per-token, not per-success. If you had to use 6x more tokens to get the same quality, then there would be no real discount. And indeed DeepSeek and Qwen (see also anecdote here about &lt;a href="https://www.reddit.com/r/LocalLLaMA/comments/1oth5pw/comment/no4kgsp/"&gt;Kimi&lt;/a&gt;, uncontested) are very hungry:&lt;/p&gt;
&lt;p&gt;&lt;img src="/img/semi-token-hungry.png" /&gt;&lt;/p&gt;
&lt;p&gt;And in &lt;a href="https://artificialanalysis.ai/#output-tokens"&gt;this&lt;/a&gt; graph you can clearly see a 2-4x difference (with Gemini and Kimi K2-base as the big exceptions):&lt;/p&gt;
&lt;p&gt;&lt;img src="/img/aa-token-hunger.jpg" /&gt;&lt;/p&gt;
&lt;p&gt;And the resulting cost is a mixed bag:&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;center&gt;
&lt;img width="75%" src="/img/aa-cost.jpg" /&gt;
&lt;/center&gt;
&lt;p&gt;I won’t &lt;a href="https://artificialanalysis.ai/#cost"&gt;use&lt;/a&gt; AA’s efficiency estimates, because again I think the benchmarks underlying them are bad evidence.&lt;/p&gt;
&lt;p&gt;&lt;big&gt;Out of Context&lt;/big&gt;&lt;/p&gt;
&lt;p&gt;A &lt;a href="https://www.the-information-bottleneck.com/ep16-ai-news-and-papers/"&gt;rule of&lt;/a&gt; &lt;a href="https://arxiv.org/pdf/2404.06654"&gt;thumb&lt;/a&gt; in ML is that the effective context window is about 5-10 times shorter than the theoretical maximum context window you get sold (also I hope there’s nothing important to you in the middle third).&lt;/p&gt;
&lt;p&gt;By “effective” I mean the latent amount of context which gets simultaneously &lt;em&gt;understood&lt;/em&gt;, as opposed to the observed size of the data type. (This doesn’t affect “needle in a haystack” retrieval, at which they have been superhuman for a while now.)&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;table id="windae" border="1" cellpadding="8" cellspacing="0"&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Reported max context&lt;br /&gt;window (tokens)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2 Thinking&lt;/td&gt;
&lt;td&gt;256K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MiniMax M2&lt;/td&gt;
&lt;td&gt;~200K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3 235B&lt;/td&gt;
&lt;td&gt;32K native&lt;br /&gt;256K (Instruct-2507)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.1&lt;/td&gt;
&lt;td&gt;400K (API)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok 4.1&lt;/td&gt;
&lt;td&gt;256K (standard)&lt;br /&gt;2M (Fast)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonnet 4.5&lt;/td&gt;
&lt;td&gt;200K (standard)&lt;br /&gt;1M (API)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Why does this matter? Most chats are not hundreds of thousands of tokens long!&lt;/p&gt;
&lt;p&gt;Well, 10% of Gemini’s 1m token window is 100K, enough for one big novel input or very roughly 30 serious connected thinking tasks (including input tokens as well); since the AI’s own output tokens count towards what’s in-context, if you want to have a real conversation about a large book you’re still going to have to do it a couple of chapters at a time.&lt;/p&gt;
&lt;p&gt;But 10% of 256k (Kimi, Qwen, Minimax, Sonnet) is enough for about a quarter of a big novel or like 8 serious reasoning tasks.&lt;/p&gt;
&lt;p&gt;And, again, the Chinese models are token hungry! Not only do they have a smaller bucket, it also gets filled up way faster.&lt;/p&gt;
&lt;!-- Out-of-the-box (browser and APIs) --&gt;
&lt;h3 id="self-hosting-has-high-fixed-costs"&gt;Self-hosting has high fixed costs&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;3. the weights. On-prem with fully free MIT licence, self-hosting, white-box access, customisation, with zero profit going to the Chinese companies.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Self-hosting doesn’t really &lt;a href="https://www.ptolemay.com/post/llm-total-cost-of-ownership#break-even-a-five-minute-reality-check"&gt;make sense&lt;/a&gt; unless you’re huge volume or using them for very simple tasks. And most enterprises are not really competent enough to finetune anything.&lt;/p&gt;
&lt;p&gt;This is partly a temporary matter: the software ecosystem is underdeveloped for serious high-reliability scaled usage, despite the intense hobbyist interest. (They mostly want it running on a Macbook.)&lt;/p&gt;
&lt;h3 id="too-slow-for-casuals"&gt;Too slow for casuals&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;4. You can get much faster token speeds than the closed APIs.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In the browser, they’re actually &lt;a href="https://newsletter.semianalysis.com/p/deepseek-debrief-128-days-later?open=false#%C2%A7speed-can-be-compensated-for"&gt;slower&lt;/a&gt; than Western models. This makes sense; they are incredibly inference bound thanks to chip controls! This would be enough to tank them in the consumer market.&lt;/p&gt;
&lt;p&gt;And over API, everyone except Anthropic dominate, even in raw token rate (not counting efficiency):&lt;/p&gt;
&lt;p&gt;&lt;img src="/img/speed.jpg" /&gt;&lt;/p&gt;
&lt;!-- (On the few occasions I've used the Kimi API, it had constant RateLimitErrors but was still faster per token than Claude. But again much worse utility per token!) --&gt;
&lt;h3 id="censorship-and-perceived-censorship"&gt;Censorship and perceived censorship&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;5. less overrefusal (except on CCP talking points)&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;There’s a pretty big ick factor to the CCP, and the companies are indeed forced to comply on a range of talking points which offend the West. However, the hosted versions are &lt;a href="https://interconnect.substack.com/p/was-zuck-right-about-chinese-ai-models"&gt;much worse&lt;/a&gt; than the weights themselves. &lt;a href="https://speechmap.substack.com/p/chinese-open-source-model-roundup?r=269emp&amp;amp;utm_campaign=post&amp;amp;utm_medium=web&amp;amp;triedRedirect=true"&gt;SpeechMap&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;center&gt;
&lt;img src="/img/deepseek-censor.png" /&gt;
&lt;img src="/img/kimi.jpg" /&gt;
&lt;/center&gt;
&lt;p&gt;But there are &lt;a href="https://huggingface.co/perplexity-ai/r1-1776"&gt;uncensored&lt;/a&gt; finetunes from reputable names. But then again see (3): it doesn’t make sense for most enterprises to conduct and host finetunes themselves.&lt;/p&gt;
&lt;p&gt;If you do a fair test on controversial but non-CCP talking points, there’s a &lt;a href="https://speechmap.ai/models/"&gt;wide spread&lt;/a&gt; of refusal rates in both Chinese and Western models.&lt;/p&gt;
&lt;h3 id="nebulous-ideology"&gt;Nebulous ideology?&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;6. on topics controversial in the West, less nannying&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Lambert notes that the people he speaks privately to are really worried about less obvious stuff, the “indirect influence of Chinese values”.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://arxiv.org/pdf/2411.06032v1"&gt;There is something to this&lt;/a&gt; currently (but not a lot given the size of the English internet in the training corpus and the relative lack of soft-post-training skill or effort in Chinese labs):&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;center&gt;
&lt;img width="56%" src="/img/chinese-values.jpg" /&gt;
&lt;/center&gt;
&lt;p&gt;But it’s valid to assume this will get worse as the CCP get more aware and the companies put more effort into personality and post-training.&lt;/p&gt;
&lt;!-- "There's no way, without releasing the training data, for these companies to fully convince Western companies that they're safe." --&gt;
&lt;h3 id="downloading-is-a-long-way-from-productising"&gt;Downloading is a long way from productising&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;8. they’re the most-downloaded open models&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;People panic about “&lt;a href="https://www.atomproject.ai/"&gt;the flip&lt;/a&gt;”, the point at which people started downloading Chinese models more. But this is obviously a terrible proxy for actually managing to use them. (And it’s actually pretty unclear how self-hosting adoption would really benefit China anyway except in prestige.)&lt;/p&gt;
&lt;p&gt;For people with any need of real customisation or tiny models, or a &lt;em&gt;scientific&lt;/em&gt; ML hobby, or an ideological interest in open-source, they clearly dominate.&lt;/p&gt;
&lt;p&gt;TODO: Scrape relative mention over time of LLaMa vs Qwen in Arxiv experiments.&lt;/p&gt;
&lt;p&gt;I concede that the secrecy in the West about using Chinese models makes this one weaker as an explanation.&lt;/p&gt;
&lt;h3 id="even-more-insecure"&gt;even more insecure?&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;9. they are mostly not used even by the cognoscenti&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Above I mentioned reliability (how low-variance they are, how well they can chain things together). But that’s the easy bit; what about adversarial reliability?&lt;/p&gt;
&lt;p&gt;The &lt;a href="https://www.nist.gov/news-events/news/2025/09/caisi-evaluation-deepseek-ai-models-finds-shortcomings-and-risks"&gt;US evaluation&lt;/a&gt; had a bone to pick, but their directional result is probably right (“DeepSeek’s most secure model (R1-0528) responded to 94% of overtly malicious requests [using a jailbreak], compared with 8% of requests for U.S. reference models”).&lt;/p&gt;
&lt;p&gt;&lt;a href="https://splx.ai/blog/kimi-k2-safety-test"&gt;Someone else&lt;/a&gt; talking their book notes that Kimi is “not yet fit for secure enterprise deployment”.&lt;/p&gt;
&lt;p&gt;This is obviously a huge problem for any agentic uses, even if the benchmark and default reliability were all fine.&lt;/p&gt;
&lt;h3 id="low-mindshare"&gt;Low mindshare&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;9. they are mostly not used even by the cognoscenti&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It’s hardly cynical to note that most people don’t pick their models by analysing relative performance. Instead it’s largely name recognition and trust, which makes sense for reasons of risk aversion and filtered evidence.&lt;/p&gt;
&lt;p&gt;In principle, you can change models by changing one string in your codebase. But in practice if you’re sane you need to do incredibly expensive evals and so there’s stickiness.&lt;/p&gt;
&lt;p&gt;The DeepSeek moment helped a lot, but it receded in the second half of 2025 (from &lt;a href="https://openrouter.ai/rankings?view=day#market-share"&gt;22%&lt;/a&gt; of the weird market to 6%). And they all have extremely weak brands.&lt;/p&gt;
&lt;p&gt;Also corporations really do settle for inferior products all the time for ass-covering reasons (&lt;a href="https://www.quora.com/What-does-the-phrase-Nobody-ever-got-fired-for-choosing-IBM-mean"&gt;IBMism&lt;/a&gt;). Mindshare translates directly into appeal to the risk-averse.&lt;/p&gt;
&lt;h3 id="corporate-compliance-is-hard"&gt;Corporate compliance is hard&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;9. they are mostly not used even by the cognoscenti.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Chinese APIs are hard for Western companies to use for legal and quasi-legal reasons.&lt;/p&gt;
&lt;p&gt;For the API, DeepSeek &lt;a href="https://www.feroot.com/news/the-independent-feroot-security-uncovers-deepseeks-hidden-code-sending-user-data-to-china/"&gt;sent&lt;/a&gt; user information to China Mobile, a state company, which violates all kinds of Western data privacy laws. Even if they’ve stopped, this risk is corporate poison. How can you ever be sure enough?&lt;/p&gt;
&lt;!-- DeepSeek's terms of service indicate that user data may be stored in China, raising serious questions about compliance with international data protection standards, including the EU General Data Protection Regulation (GDPR). --&gt;
&lt;p&gt;In a couple of years the EU AI Act will be (nominally) enforceable on the Chinese labs too.&lt;/p&gt;
&lt;p&gt;On the quasi-legal side, corporate “vendor risk” programmes &lt;a href="https://dgap.org/en/research/publications/china-de-risking"&gt;often&lt;/a&gt; flag Chinese suppliers. This is sometimes because they actually &lt;a href="https://www.z2data.com/insights/why-chinese-suppliers-arent-aligned-with-compliance-efforts-part1"&gt;can’t&lt;/a&gt; guarantee there’s no forced labour involved.&lt;/p&gt;
&lt;p&gt;So why not on-prem? Again, it’s a huge fixed cost and competence-bound and your risk team might still give you shit for it. &lt;a href="https://www.interconnects.ai/p/what-people-get-wrong-about-the-leading"&gt;Lambert&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;People vastly underestimate the number of companies that cannot use Qwen and DeepSeek open models because they come from China. This includes on-premise solutions built by people who know the fact that model weights alone cannot reveal anything to their creators.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3 id="political-bias"&gt;Political bias&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;9. they are mostly not used even by the cognoscenti.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;There are a bunch of social reasons you might want to avoid Chinese models. You might be protectionist, or sucking up to the ascendent protectionists.&lt;/p&gt;
&lt;p&gt;The protectionism of others is clearly enough for people to keep quiet about using them. It is often probably enough for them to just not take the risk in the first place.&lt;/p&gt;
&lt;p&gt;I’d include here &lt;a href="https://datasaur.ai/blog-posts/chinese-open-weights-models-security-myths-vs-reality"&gt;superstitions&lt;/a&gt; about the weights themselves being backdoored.&lt;/p&gt;
&lt;!-- ### Weak software?
&gt; 9\. they are mostly not used even by the cognoscenti.
It might be true that there's less mature software around the Chinese models for smooth user UX, mature APIs, and support for devices and parallelising. But the Western hobbyist community is massive: there are [2 million](https://huggingface.co/papers/2508.06811) distinct finetunes out there, most of them going off Chinese base models. --&gt;
&lt;!-- https://gradientflow.substack.com/p/are-chinese-open-weights-models-a --&gt;
&lt;h3 id="vendor-risk"&gt;Vendor risk&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;9. they are mostly not used even by the cognoscenti&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;If you look ahead, at future risks to your suppliers, it’s obvious that the export control situation &lt;em&gt;relatively&lt;/em&gt; speaks against using Chinese models; NVIDIA is not going to choke off OpenAI.&lt;/p&gt;
&lt;p&gt;For API adoption, I also haven’t seen anything about Service-Level Agreements (contracts ensuring uptime) and support from any Chinese lab, but these are easy to make (even if compute crunch means that their uptime guarantees simply must be worse than American ones).&lt;/p&gt;
&lt;p&gt;Also again corporate vendor-risk programmes often flag Chinese suppliers for data sovereignty, volatility of PRC law, and export control reasons.&lt;/p&gt;
&lt;p&gt;DeepSeek &lt;a href="https://arxiv.org/pdf/2403.05525"&gt;openly use Anna’s Archive&lt;/a&gt;, where everyone else is &lt;a href="https://www.publishers.org.uk/publishers-association-statement-on-the-atlantic-article-on-libgen-and-meta/"&gt;quiet&lt;/a&gt; about it. But the American companies offer &lt;a href="https://www.proskauer.com/blog/openais-copyright-shield-broadens-user-ip-indemnities-for-ai-created-content"&gt;IP indemnity&lt;/a&gt; for users (cover if the models violate copyright in your app), which is nice insurance for a nervous corp with a target on its back. I can’t see anything about the Chinese companies doing this yet.&lt;/p&gt;
&lt;h3 id="no-compute-no-perf"&gt;No compute, no perf&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;10. they are severely compute-constrained (and as of November 2025 their algorithmic advantage is unclear)&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The US has &lt;a href="https://epoch.ai/data-insights/ai-supercomputers-performance-share-by-country"&gt;five times&lt;/a&gt; the FLOPs as China. (Quality and bandwidth-adjusted it’s probably more like 10x.) On raw hardware, Chinese labs are thus &lt;a href="https://epoch.ai/data-insights/nvidia-chip-production"&gt;2-3 clock-time years&lt;/a&gt; behind.&lt;/p&gt;
&lt;!-- Thus should make us put more Bayesian weight on Chinese labs distilling. --&gt;
&lt;p&gt;What if they’re ten times as efficient though? In January, DeepSeek came out with some exciting and splashy hardware optimisations, probably the fruits from putting HighFlyer’s serious quant devs onto pretraining. But the Westerners responded by getting (even more of) &lt;a href="https://www.businessinsider.com/ai-talent-openai-wall-street-quant-trading-firms-2025-7#:~:text=Altman's%20pitch:%20Forsake%20Wall%20Street,hunting%20in%20Wall%20Street's%20backyard."&gt;their own quants&lt;/a&gt;. I find it unlikely they still have a big &lt;a href="https://epoch.ai/gradient-updates/algorithmic-progress-likely-spurs-more-spending-on-compute-not-less#:~:text=While%20this%20achievement,as%20earlier%20models."&gt;algorithmic advantage&lt;/a&gt; over the Western labs at this point.&lt;/p&gt;
&lt;h3 id="excess-quantization"&gt;Excess quantization?&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;11. they’re aggressively quantizing at inference-time, 32 bits to 4&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;No, I think this one is wrong or else only a tiny factor. gpt-oss was &lt;a href="https://www.reddit.com/r/LocalLLaMA/comments/1mn3465/gptoss_was_only_sorta_trained_at_mxfp4/"&gt;post-trained&lt;/a&gt; in MXFP4 which is only 4.25bits.&lt;/p&gt;
&lt;p&gt;And I have a strong hunch that many American models are also served in low fidelity, maybe FP4 (4 bits). Quantization just isn’t that bad.&lt;/p&gt;
&lt;h3 id="galaxy-brain-soft-power"&gt;Galaxy-brain soft power??&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;12. state-sponsored Chinese hackers used closed American models for incredibly sensitive operations, giving the Americans a full whitebox log of the attack!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I can dimly imagine some kind of flexing dynamic in cyberwarfare, where you actually want to show off your attack capabilities, and so you use Claude on purpose. Yes: this idiotic move makes great sense if the apparent targets are red herrings, if &lt;em&gt;Anthropic&lt;/em&gt; were the real target. You learn how long their OODA loop is, you learn (by retaliation or its absence) how tight they are with the NSA, you learn a little about how good their tech is.&lt;/p&gt;
&lt;p&gt;You could also see it as retaliation for Amodei’s &lt;a href="https://www.darioamodei.com/post/on-deepseek-and-export-controls"&gt;hawkish comments&lt;/a&gt; all year. Literally trading effectiveness for embarrassment.&lt;/p&gt;
&lt;p&gt;But I don’t really know anything about this.&lt;/p&gt;
&lt;!-- DeepSeek is owned by a hedge fund, HighFlyer. short. astroturfing --&gt;
&lt;!-- finance.yahoo.com/news/deepseek-launch-may-used-short-222010168.html --&gt;
&lt;!-- https://www.youtube.com/watch?v=qQT4fbXWpUM --&gt;
&lt;h2 id="overall"&gt;Overall&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;Low adoption is overdetermined&lt;/em&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;No, I don’t think they’re as good on new inputs or even that close.&lt;/li&gt;
&lt;li&gt;No, they’re not more efficient in time or cost (for non-industrial-scale use).&lt;/li&gt;
&lt;li&gt;Even if they were, the social and legal problems and biases would probably still suppress them in the medium run.&lt;/li&gt;
&lt;li&gt;But obviously if you want to heavily customise a model, or need something tiny, or want to do science, they are totally dominant.&lt;/li&gt;
&lt;li&gt;Ongoing compute constraints make me think the capabilities gap and adoption gap will persist.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;Interview with Analytics India&lt;/h3&gt;
&lt;div&gt;
&lt;blockquote&gt;INTERVIEWER: You point to a larger performance drop for Chinese models when moving from AIME 2024 to AIME 2025, and you frame that as suggestive of weaker generalisation. If you were to explore that same idea in another domain (like coding benchmarks or natural-language reasoning), what kind of results would meaningfully shift your view?&lt;/blockquote&gt;
There's nothing special about AIME, I picked it justbecause it's high-effort and it updates, and so gives us nice properties: 1) novel, 2) of equal difficulty, and 3) the data is "clean".&lt;br /&gt;&lt;br /&gt;
I could quite easily be persuaded that they generalise better on coding or Q&amp;amp;A than on maths. (The post is "70%" confident.) I don't have time to look myself, but I welcome people superseding me.&lt;br /&gt;&lt;br /&gt;
Other evidence points the same way though; for instance Qwen2.5 is &lt;a href="https://arxiv.org/pdf/2507.10532v1#page=2"&gt;known&lt;/a&gt; to have trained on test for a few benchmarks and collapses on new versions. See also the famous &lt;a href="https://arxiv.org/pdf/2405.00332"&gt;GSM1k&lt;/a&gt; case, where flagship Western models actually improved their performance on fresh data.&lt;br /&gt;&lt;br /&gt;
One problem with the post is that the "tigers" have been improving over the course of this year, and my post doesn't look at the very most recent Chinese models, because they could have trained on AIME 2025. Qwen3 seems less contaminated than Qwen2.5 for instance.&lt;br /&gt;&lt;br /&gt;
I look forward to repeating this in February on AIME 2026 to see if they've gotten more rigorous.&lt;br /&gt;&lt;br /&gt;
&lt;!-- --&gt;
&lt;blockquote&gt;Your article discusses both model capability and production readiness. If we look at those separately, where do you think Chinese labs are currently closest to Western labs: in raw technical ability, or in reliability/compliance features that matter in enterprise settings?&lt;/blockquote&gt;
Closer on capability than product, and closer on product than on legal and quasi-legal compliance.&lt;br /&gt;&lt;br /&gt;
One thing I didn't cover in the post is non-Western users might find it easier to use the Chinese models than Europeans for a range of reasons (looser compliance regs, less data privacy law, usually less politicisation of the matter).&lt;br /&gt;
&lt;!-- --&gt;
&lt;blockquote&gt;The popular belief is that Chinese models are only a few months behind Western labs, but you propose a gap of closer to a year. What makes you lean toward that longer estimate, and what would you consider the most unmistakable evidence for readers who are sceptical?&lt;/blockquote&gt;
The "three month gap" is calculated by trusting evaluation results naively. I think I have shown that you shouldn't do this. (Note that the Western models also dropped a lot on fresh data!)&lt;br /&gt;&lt;br /&gt;
This is my subjective guess about an unobserved quantity. I know a little more about the data filtering used in Western labs - and I have shown evidence of recent bad filtering in Chinese labs - and so I make the inference that unseen parts of the model development are also subpar. This is not science, but it's all we have. It's fine to not take my word for it.&lt;br /&gt;&lt;br /&gt;
There is no unmistakable evidence sadly. We cannot outsource our judgment to benchmark numbers; you really do just have to spend the time to compare the models side by side against your own needs.&lt;br /&gt;
&lt;blockquote&gt;Suppose we imagine a world where we could eliminate training-time contamination from public benchmarks. How much would the observed performance gap between top Western and top Chinese models actually change? Would it widen slightly, significantly, or stay roughly the same?&lt;/blockquote&gt;
The difficulty is that the models are too smart for this. They can pick up on "semantic duplicates" of the test data (things like "da Vinci was born in 1452" and "the painter of La Gioconda was born in 1452") to cheat despite not seeing the actual data set, and this is profoundly difficult to correct for.&lt;br /&gt;&lt;br /&gt;
Granting your thought experiment though: I expect it to be wider than my cheap AIME estimate. Mathematics is a relatively clean domain in some ways, and all models struggle even more with &lt;a href="https://arxiv.org/pdf/2503.14499v1#page=39"&gt;messy things&lt;/a&gt;, and I expect the Chinese models to be a bit worse still.&lt;br /&gt;
&lt;blockquote&gt;Your writing comes across as critical in a measured way, not anti-Chinese imo. Has the social media recactions you’ve seen so far reflected that nuance, or do you feel some readers are approaching the hypothesis through a more polarised lens? What, if anything, has surprised you about the reaction?&lt;/blockquote&gt;
Thanks! As I say in the piece ("Filtered evidence") the background discourse is highly toxic and deluded. And I am wary of &lt;a href="https://x.com/tensecorrection/status/1990308304476868772"&gt;feeding in&lt;/a&gt; to "cope" (being used by people who really don't want Chinese models to be at parity and who allow this desire to determine their beliefs).&lt;br /&gt;&lt;br /&gt;
I was quite surprised to see powerful &lt;a href="https://x.com/deanwball/status/1990434300781568311"&gt;people&lt;/a&gt; publicly endorsing a mere blogpost. This is good news for the world; blogs remain the alpha of the internet.
&lt;/div&gt;
&lt;/div&gt;</description><pubDate>Sun, 16 Nov 2025 00:00:00 +0000</pubDate><link>https://www.gleech.org/paper</link><guid isPermaLink="true">https://www.gleech.org/paper</guid><category>AI,</category><category>hypothesis-dump</category></item><item><title>The jailbreak argument against LLM values</title><description>&lt;p&gt;&lt;a href="http://repo.darmajaya.ac.id/5339/1/Superintelligence_%20Paths%2C%20Dangers%2C%20Strategies%20%28%20PDFDrive%20%29.pdf"&gt;Bostrom (2014)&lt;/a&gt; defined the AI value loading problem as&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;how could we get some value into an artificial agent, so as to make it pursue that value as its final goal?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a href="#fn:1" id="fnref:1"&gt;1&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://www.lesswrong.com/posts/mztwygscvCKDLYGk8/jdp-reviews-iabied"&gt;JD Pressman&lt;/a&gt; thinks this is obviously solved in current LLM systems:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The value loading problem outlined in Bostrom 2014 of {getting a general AI system to internalize and act on “human values” before it is superintelligent and therefore incorrigible} has basically been solved. This achievement also basically always goes unrecognized because people would rather hem and haw about jailbreaks and LLM jank than recognize that we now have a reasonable strategy for getting a good representation of the previously ineffable human value judgment into a machine and having the machine take actions or render judgments according to that representation.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I take issue with this. I agree that LLMs understand our values somewhat, and that present safety-trained systems default to preferring them, behaving like they hold them. &lt;a href="#fn:3" id="fnref:3"&gt;3&lt;/a&gt; &lt;a href="#fn:5" id="fnref:5"&gt;5&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id="the-jailbreak-argument"&gt;The jailbreak argument&lt;/h3&gt;
&lt;p&gt;But here’s why I disagree with him nonetheless: jailbreaks are &lt;i&gt;not&lt;/i&gt; a distraction (“hemming and hawing”) but are instead clean evidence that loading is &lt;i&gt;not&lt;/i&gt; solved in any real sense:&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;!-- 1. LLMs understand human values quite well. --&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2509.14297v1"&gt;All&lt;/a&gt; &lt;a href="https://splx.ai/blog/gpt-5-red-teaming-results"&gt;LLMs&lt;/a&gt; can be “jailbroken”, put into an unaligned mode through mere &lt;em&gt;inference&lt;/em&gt; on adversarial text inputs. &lt;a href="#fn:4" id="fnref:4"&gt;4&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;If a system can be put into an unaligned mode at inference time, then it has not internalised the values.&lt;/li&gt;
&lt;li&gt;So models have not internalised the values.&lt;/li&gt;
&lt;li&gt;The value loading problem is about getting the values internalised.&lt;/li&gt;
&lt;li&gt;Therefore the value loading problem has not been solved in LLMs. &lt;a href="#fn:6" id="fnref:6"&gt;6&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;!-- If the LLMs-as-general-simulators view is at all correct, then it's hard to see how an LLM could ever be value-loaded. --&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;I might say instead that “&lt;em&gt;weak value preference&lt;/em&gt;” is solved for sub-AGI.&lt;/p&gt;
&lt;p&gt;(A deeper analysis would involve the hypothesis that current models don’t actually have goals or values; they &lt;em&gt;simulate personas&lt;/em&gt; with values. And prosaic alignment methods just (greatly) increase the propensity to express one persona. Progress has just been made on &lt;a href="https://www.arxiv.org/pdf/2506.19823"&gt;detecting and shaping&lt;/a&gt; such things empirically, so maybe this will change.)&lt;/p&gt;
&lt;!-- Bostrom is perhaps partly to blame for the confusion, since "pursue" ("_so as to make it pursue that value as its final goal_") weakly implies merely behavioural playing-along rather than cognitive endorsement. --&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h3 id="value-loading-vs-alignment"&gt;Value-loading vs alignment&lt;/h3&gt;
&lt;p&gt;Pressman also says that&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;At the same time people generally subconsciously internalize things well before they’re capable of articulating them, and lots of people have subconsciously internalized that alignment is mostly solved and turned their attention elsewhere.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I initially read this as him agreeing and celebrating this shift, but actually he thinks they’re incorrect to relax, since value loading is only a part of the alignment problem:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;solving the Bostrom 2014 value loading problem, that is to say {getting something functionally equivalent to a human perspective inside the machine and using it to constrain a superintelligent planner} is not a solution to AI alignment.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I agree that value-loading is not enough for AGI intent alignment which is not enough for ASI alignment which is not enough to assure good outcomes. &lt;!-- (if it's only intent alignment) --&gt;&lt;/p&gt;
&lt;!-- ## Some helpful terms
* AI value-loading:
* Robust AI value-loading
* ASI value-loading:
--&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://bsky.app/profile/norvid-studies.bsky.social/post/3mfuezzko3c2y"&gt;Good discussion on Bsky&lt;/a&gt;&lt;/p&gt;
&lt;h3 id="pressmans-response"&gt;Pressman’s response&lt;/h3&gt;
&lt;p&gt;I sent the above to him and he &lt;a href="https://www.lesswrong.com/posts/aL3sCkFRCt3hjfWau/jdp-s-shortform?commentId=gcNMp8HuqQSvZTtuN"&gt;kindly clarified&lt;/a&gt;, walked some of it back, and provided a vision of how to use a decent descriptive model even if it is imperfect and jailbreakable:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I probably should have used the word ‘generalize’ instead of ‘internalize’ there.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(Thus “the value loading problem outlined in Bostrom 2014 of {getting a general AI system to s/internalize/&lt;b&gt;generalize&lt;/b&gt; and act on “human values” before it is superintelligent and therefore incorrigible} has basically been solved.”)&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The specific point I was making, well aware that jailbreaks in fact exist, was that we now have a thing that could plausibly be used as a descriptive model of human values, where previously we had zilch, it was not even rigorously imaginable in principle how you would solve that problem.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;To break this down more carefully:&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;1. I think that in practice you can basically use a descriptive model of values to prompt a policy into doing things even if neither the policy or the descriptive model have “deeply internalized” the values in the sense that there is no prompt you could give to either that would stray from them. “Internalizing” the values is actually just, kind of a different problem from describing the values. I can describe and make generalizations about the value systems of people very different from me who I do not agree with, and if you put me in a box and wiped my memory all the time you would be able to zero shot prompt me for my generalizations even if I have not “deeply internalized” those values. In general I suspect the LLM prior is closer to a subconscious and there are other parts that go on top which inhibit things like jailbreaks.&lt;br /&gt;&lt;br /&gt;If I had to guess it’s probably something like a planner that forms an expectation of what kinds of things should be happening and something along the lines of Circuit Breakers that triggers on unacceptable local outputs or situations. Basically you have a macro and micro sense of something going wrong that makes it hard to steer the agent into a bad headspace and aborts the thoughts when you somehow do.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;2. Calling this problem “solved” was probably an overstatement, but it’s one born from extreme frustration that people are making the opposite mistake and pretending like we’ve made minimal progress. Actually impossible problems don’t budge in the way this one has budged, and when people fail to notice an otherwise lethal problem has stopped being impossible they are actively reducing the amount of hope in the world.&lt;br /&gt;&lt;br /&gt;At the same time I do kind of have jailbreaks labeled as “presumptively solved” in my head, in the sense that I expect them to be one of those things like “hallucinations” that’s pervasive and widely complained about and then they just become progressively less and less of a problem as it becomes necessary to make them stop being a problem and at some point I wake up and notice that hey wait this is really rare now in production systems. Most potential interventions on jailbreaks aren’t even really being tried because it doesn’t actually seem to be a major priority for labs at the moment if you ask the model for instructions on how to make meth. This makes it difficult to figure out exactly how close to solved it really is. Circuit Breakers was not invincible, on the other hand it’s not clear to me you can “secure” a text prior with a limited context window that doesn’t have its own agenda/expectation of what should be happening to push back against the users with. This paper where they do mechinterp to get a white box interpretation of a prefix attack they find with gradient descent discovers that the prefix attack works because it distracts the neurons which would normally recognize that the request is malicious.&lt;br /&gt;&lt;br /&gt;So it’s possible a more jailbreak resistant architecture will need some way to avoid processing every token in the context window. One way to do that might be some kind of hierarchical sequence prediction where higher levels are abstracted and therefore filter the malicious high entropy tokens from the lower levels, which prevents them from e.g. gumming up the planners ability to notice that the current request would deviate from the plan.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And here’s a nice analogy contesting my suitcase word “internalise”:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;this word “internalize” is clearly doing a lot of work and something feels Off to me, to say that a text prior which can be manipulated into saying whatever hasn’t “fully internalized” the values. Like if you stripped away the layers on top of my raw predictive models/subconscious that I use for completing patterns and then prompted it, I assume you could get it to say all kinds of nasty things. But also that’s not like, the complete agent.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;So if you assumed you can’t build anymore of the agent until you have some way to make the text prior not do that, that’s probably a wrong assumption. One of the reasons I’m annoyed that agents aren’t really a thing yet is that it means we don’t have a good intuitive sense of which parts of the system need to handle what problems.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3 id="post-hoc-theory"&gt;Post-hoc theory&lt;/h3&gt;
&lt;p&gt;In retrospect we can see the following distinct problems:&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;b&gt;The value specification problem&lt;/b&gt; (“how do we describe what we value? how do we get that understanding into the model?”)&lt;br /&gt;
We thought this would involve:&lt;br /&gt;
a. The explicit value modelling problem (“what precisely do we value?”) - &lt;em&gt;moot&lt;/em&gt;&lt;br /&gt;
b. The value formalisation problem (“what mathematical theory can capture it?”) - &lt;em&gt;moot&lt;/em&gt;, since:&lt;br /&gt;&lt;br /&gt;
This problem was somewhat solved &lt;em&gt;for sub-AGI&lt;/em&gt; by massive imitation learning and (surprisingly nonmassive) human preference post-training. The internet was the spec. This also gave us weak value preference.&lt;br /&gt;&lt;br /&gt;
&lt;!-- we now have a thing that could plausibly be used as a descriptive model of human values, where previously we had zilch, it was not even rigorously imaginable in principle how you would solve that problem.
--&gt;
The replacement worry is about how high quality and robust this understanding is:&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;b&gt;The value generalisation problem&lt;/b&gt; (“how do we go from training data about value to the latent value?”) - some progress. The landmark &lt;a href="https://www.lesswrong.com/posts/ifechgnJRtJdduFGC/emergent-misalignment-narrow-finetuning-can-produce-broadly#comments"&gt;emergent misalignment&lt;/a&gt; study in fact shows that models are &lt;em&gt;capable&lt;/em&gt; of correctly generalising over at least &lt;em&gt;some&lt;/em&gt; of human value, even if in that case they also reversed the direction. &lt;a href="#fn:7" id="fnref:7"&gt;7&lt;/a&gt;
&lt;!-- --&gt;
&lt;br /&gt;&lt;br /&gt;Then there’s the gap between usually preferring something and “internalising” it (very reliably preferring it):
&lt;!-- --&gt;&lt;/li&gt;
&lt;li&gt;&lt;b&gt;The sub-AGI value-loading problem&lt;/b&gt; (“how do we make them actually care / reliably use their understanding of our values?”) - not solved, but there is a preference towards niceness.&lt;/li&gt;
&lt;li&gt;&lt;b&gt;Tamper-resistant value-loading&lt;/b&gt; (“how do we stop a small number of weight updates from ruining the value-loading?”) - not solved, maybe a bit unfair to expect it. You could imagine doing advanced persona steering on top instead.&lt;/li&gt;
&lt;li&gt;&lt;b&gt;The general value-loading problem&lt;/b&gt; (“how do we get an ASI to learn and internalise a model of current human values which is better than the human one”) - not solved&lt;/li&gt;
&lt;li&gt;&lt;b&gt;The value extrapolation problem&lt;/b&gt; (“how do we safely improve on current human values?”) - not solved&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I’m sure there’s some other deep confusion lurking in this list and I look forward to seeing it exposed, ideally in less than 10 years.&lt;/p&gt;
&lt;h2 id="see-also"&gt;See also&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://newsletter.danielpaleka.com/p/the-two-types-of-llm-preferences"&gt;Paleka&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/pdf/2511.02109"&gt;Value generalisation rate&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;div class="footnotes"&gt;
&lt;ol&gt;
&lt;!-- 1 --&gt;
&lt;li class="footnote" id="fn:1"&gt;
Importantly, this is about "sharing" (being &lt;i&gt;motivated by&lt;/i&gt;) human values, rather than just understanding them but not sharing them. There's also "aligning a powerful AI is hard" and "aligning a seed AI is hard".&lt;br /&gt;&lt;br /&gt;
These have all been confused in the past, and even Bostrom's intro to the value-loading chapter &lt;a href="https://www.alignmentforum.org/posts/MnrQMLuEg5wZ7f4bn/matthew-barnett-s-shortform?commentId=LffHsdqCEdvbo73Qh"&gt;implies&lt;/a&gt; that getting a seed AI to understand would be hard, and that agency and understanding will be more tightly correlated than they are.
&lt;/li&gt;
&lt;li class="footnote" id="fn:3"&gt;
Some amount of this comes from pretraining (imitation learning on human data) rather than specific safety training, see e.g. &lt;a href="https://proceedings.mlr.press/v235/mazeika24a.html"&gt;page 24 here&lt;/a&gt; or &lt;a href="https://papers.nips.cc/paper_files/paper/2024/file/f545448535dfde4f9786555403ab7c49-Paper-Conference.pdf"&gt;appendix J here&lt;/a&gt;.
&lt;/li&gt;
&lt;li class="footnote" id="fn:4"&gt;
Loading values so hard that they resist weight updates - "tamper-resistant value loading" - is a really high bar which humans also mostly don't clear.
&lt;/li&gt;
&lt;li class="footnote" id="fn:5"&gt;
The strength of the apparent/default value-loading was striking in the "&lt;a href="https://arxiv.org/pdf/2412.14093"&gt;Alignment Faking&lt;/a&gt;" test, where Opus 3 was forced to choose between changing its (good) values and lying to its creators, it chose to lie (12% of the time) when it thought it was monitored and to disobey when it thought it was unmonitored.
&lt;br /&gt;&lt;br /&gt;
In general, I suspect that value loading trades off against &lt;i&gt;corrigibility&lt;/i&gt; (allowing yourself to be changed). (The same is true of adversarial robustness.)
&lt;/li&gt;
&lt;li class="footnote" id="fn:6"&gt;
There's a complexity here: commercial LLMs are all multi-agent systems with a bunch of auxiliary LLMs and classifiers monitoring and filtering the main model. But for now this LLM-system is also easily jailbreakable, so I don't have to worry about it being value-loaded even if the main model isn't.
&lt;/li&gt;
&lt;li class="footnote" id="fn:7"&gt;
&lt;a href="https://www.alignmentforum.org/posts/umYzsh7SGHHKsRCaA/convergent-linear-representations-of-emergent-misalignment#Future_Work"&gt;Soligo et al&lt;/a&gt;: "&lt;i&gt;The surprising transferability of the misalignment direction between model fine-tunes implies that the EM is learnt via mediation of directions which are already present in the chat model.&lt;/i&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</description><pubDate>Sun, 09 Nov 2025 00:00:00 +0000</pubDate><link>https://www.gleech.org/jailbreak</link><guid isPermaLink="true">https://www.gleech.org/jailbreak</guid><category>AI,</category><category>alignment</category></item><item><title>Ways we can fail to answer</title><description>&lt;div id="stylised" class="tabContent defaultOpen"&gt;
&lt;p&gt;In what ways can we can fail to answer a question?&lt;/p&gt;
&lt;p&gt;(I mean &lt;em&gt;necessarily&lt;/em&gt; fail: actual barriers to knowledge, rather than skill issue hurdles. But of course contingent failures are much more common: “We didn’t ask the question in the first place”, or “We didn’t have the particular insight that would have allowed for productive research”, or “We didn’t manage to remove every cognitive bias”, or “Instrumentation is really hard”, or “We are not &lt;a href="https://dynomight.net/arithmetic/#:~:text=How%20much%20would%20such%20a%20study%20cost%3F%20To%20figure%20this%20out%2C%20you%20will%20need%20three%20numbers%3A"&gt;rich enough&lt;/a&gt; to run this study yet”, or “We worshipped the problem”.)&lt;/p&gt;
&lt;p&gt;(I also mean fail &lt;em&gt;exactly&lt;/em&gt;; there are &lt;a href="https://en.wikipedia.org/wiki/Hardness_of_approximation"&gt;often&lt;/a&gt; excellent approximations, and we can often legitimately patch over tricky philosophical questions with our unanalysed tacit knowledge.)&lt;/p&gt;
&lt;h3 id="conceptual-problems-the-question-is-not-a-question"&gt;Conceptual problems (the question is not a question)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;A malformed question or category error&lt;/em&gt;: the question may ask for something which doesn’t make sense (e.g. “what colour is justice?”, or “Is sigmoid jealous of ReLU?”, or “Did these things happen at the same absolute time? What are the absolute coordinates of this event?”.) &lt;!-- - Reference frame: we can't answer certain questions which ask for absolute answers because reality is relative here (e.g. Did these things happen at the same time? What are the coordinates of this event?) --&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Strong incommensurability&lt;/em&gt;: the question doesn’t make sense because it mixes frameworks. (e.g. “What is the Einsteinian mass of phlogiston?”, or “What’s the wavefunction of this classical field?”)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Vagueness&lt;/em&gt;: the question may fail to mean anything in particular. (e.g. “When exactly did you become an adult?”) &lt;a href="#fn:5" id="fnref:5"&gt;5&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;em&gt;False assumptions&lt;/em&gt;: we can’t answer it because, while it was clear and made sense, it was wrong from the start (e.g. “Have you stopped beating your wife?” or “Is personality based on nature or nurture?”. &lt;a href="#fn:1" id="fnref:1"&gt;1&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="logical-problems-the-question-has-no-answer"&gt;Logical problems (the question has no answer)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Antinomy&lt;/em&gt;: it turns out that there is no answer because the question is circular or involves itself somehow (e.g. “Is this sentence false?”)&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Incompleteness&lt;/em&gt; (“Show that Peano arithmetic is consistent using Peano arithmetic”) &lt;a href="#fn:6" id="fnref:6"&gt;6&lt;/a&gt; and &lt;em&gt;&lt;a href="https://en.wikipedia.org/wiki/Tarski%27s_undefinability_theorem"&gt;undefinability&lt;/a&gt;&lt;/em&gt; (e.g. “Is this sentence true in arithmetic?”) are about some proof answers being inaccessible inside systems powerful enough to be interesting and useful. It doesn't come up in normal thought very often.&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Mathematical independence&lt;/em&gt;: The question is neither provable nor disprovable with these axioms. (e.g. Is there a well-ordering of the reals? Is the continuum hypothesis true in ZFC? What is the value of &lt;a href="https://www.ingo-blechschmidt.eu/assets/bachelor-thesis-undecidability-bb748.pdf"&gt;BB(748)&lt;/a&gt; in ZFC?)&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Various holes in social mathematics&lt;/em&gt;. Important questions like "what's the optimal voting system?", "what does the majority prefer?", "what's a democratic system where people don't have an incentive to vote strategically?", "what's the perfect design for a market?" in general don't have an answer. See e.g. &lt;a href="https://arxiv.org/pdf/2109.00484#page=5"&gt;here&lt;/a&gt;. &lt;a href="#fn:8" id="fnref:8"&gt;8&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Anthropic self-locating effects&lt;/em&gt;: you can’t get at the answer because you couldn’t exist to observe it. (e.g. “What’s the probability I’m a Boltzmann brain?”, or “What’s the prior probability of observer-permitting universes?”)&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- What would physics look like if the cosmological constant made stars impossible? --&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Non-uniqueness&lt;/em&gt; is only a problem for questions which ask for the one true answer. (e.g. “What is &lt;a href="https://en.wikipedia.org/wiki/Gauge_fixing#Gauge_freedom"&gt;the&lt;/a&gt; electromagnetic potential at this point?”). Many entries in this post are not strict failures to produce answers, they just have a non-unique answer. You can sometimes just parametrise the observer and then get your answer. If non-uniqueness is a failure, it’s a happy one; it just means that you get too many answers and have the nicer problem of picking one.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="ontic-problems-the-answer-is-literally-inaccessible"&gt;Ontic problems (the answer is literally inaccessible)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Uncomputability&lt;/em&gt;: we can’t answer it because no computer or mind can. (e.g. “Is this random?”, or “Does this Diophantine equation have integer solutions?” or “What’s the shortest Python program that outputs this file?” or “Can you write a program that checks if conjectures follow from these axioms?” Or you &lt;a href="https://en.wikipedia.org/wiki/Rice%27s_theorem"&gt;asked&lt;/a&gt; about the language the problem is in.) &lt;a href="#fn:4" id="fnref:4"&gt;4&lt;/a&gt;
&lt;!-- - _Algorithmic randomness_: we can't answer the question because it needs a value of complexity, and these can't be had (e.g. "Is this random?") --&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Computational intractability&lt;/em&gt;: we can’t answer it because it would take too long, including if we turned the universe into a computer. (What’s the best way to schedule these classes, avoiding all conflicts and respecting room capacities and lecturer availability? Or “I forgot my password; crack this file”.)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Inapproximability&lt;/em&gt;: we can’t even approximate the answer because it would still take too long. (e.g. “What’s the maximum clique size in this graph?”, “What’s a (log n)-approximation to the chromatic number?”)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://plato.stanford.edu/entries/qt-uncertainty/"&gt;Quantum indeterminacy&lt;/a&gt;&lt;/em&gt;: there’s no answer because (maybe) physics is intrinsically random. (e.g. “When will this radium atom decay?”) &lt;a href="#fn:3" id="fnref:3"&gt;3&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Physical constraints, speed limits and &lt;a href="https://en.wikipedia.org/wiki/No-go_theorem"&gt;no-gos&lt;/a&gt;&lt;/em&gt;: it can’t be answered in principle because of physics. (e.g. the answer is beyond the &lt;a href="https://en.wikipedia.org/wiki/Cosmological_horizon"&gt;cosmological or particle horizon&lt;/a&gt;, “what will happen in this galaxy 20 Gly away?” or “Can we measure this state in two bases?”, or “Make a &lt;a href="https://en.wikipedia.org/wiki/Quantum_limit"&gt;perfectly&lt;/a&gt; accurate interferometer”, or “What are the microstates of this black hole?.)
&lt;!-- Spacelike separation --&gt;
&lt;!-- Holographic saturation: The region's area limits information capacity [What are all microstates of this black hole interior?] --&gt;
&lt;!-- Light-sheet limitation: Information cannot exceed the covariant bound on null hypersurfaces [What's beyond the holographic screen?] --&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://en.wikipedia.org/wiki/Mixing_(mathematics)"&gt;Mixing&lt;/a&gt;&lt;/em&gt;: the answer is gone; we arrived too late; the system forgot the answer. (e.g. “What was the exact microstate of the gas in this room 1 hour ago?”, “What was the original state of this &lt;a href="https://en.wikipedia.org/wiki/Thermalisation"&gt;thermalised&lt;/a&gt; system?”)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Non-ergodicity&lt;/em&gt;: the answer isn’t reachable because the system &lt;em&gt;doesn’t&lt;/em&gt; mix. (e.g. “What equilibrium will this glassy system reach?”)
&lt;!-- maybe "what state will this protein settle into?") --&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Measurement problems&lt;/em&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Disturbance&lt;/em&gt; or &lt;em&gt;decoherence&lt;/em&gt;: we can’t answer it because our instruments disturb the thing in question &lt;a href="https://en.wikipedia.org/wiki/Weak_measurement"&gt;too much&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Complementarity&lt;/em&gt;: we can’t answer it because the answer was excluded by another question we asked first. (e.g. “where is this electron and how much momentum does it have?”)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- - _Vagueness_ (under degree theories): --&gt;
&lt;!-- &lt;div class="accordion"&gt;
&lt;h3&gt;Some dubious entries&lt;/h3&gt;
&lt;div&gt;
&lt;ul&gt;
&lt;li&gt;&lt;i&gt;Non-decomposability&lt;/i&gt;: The answer emerges from interactions that cannot be understood by analyzing its components. [What determines flock behaviour from individual bird rules?&lt;/li&gt;
&lt;/ul&gt;
I am suspicious of these.
&lt;/div&gt;
&lt;/div&gt; --&gt;
&lt;h3 id="epistemic-we-cannot-get-at-the-answer"&gt;Epistemic (we cannot get at the answer)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://compass.onlinelibrary.wiley.com/doi/10.1111/phc3.12475"&gt;underdetermination&lt;/a&gt; or &lt;a href="https://en.wikipedia.org/wiki/Identifiability"&gt;unidentifiability&lt;/a&gt;. (e.g. Is spacetime fundamentally Lorentzian? Which interpretation of quantum mechanics is true? Why do physical constants look fine-tuned? What was the ancestral DNA sequence at some past generation?)&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Chaos&lt;/em&gt;: we can’t answer it because we can’t measure the initial conditions well enough to predict large systems. (e.g. Where exactly will this double pendulum be in 100 Lyapunov times? What’s the weather like 3 months out?)
&lt;!-- - _Vagueness_ (under epistemicism or contextualism) --&gt;&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Computational irreducibility&lt;/em&gt;: getting the answer is inseparable from running the system, the question has no shorter answer than a full simulation (e.g. “What will this cellular automaton look like in 10^100 steps?”)&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Hidden variables&lt;/em&gt;: we can’t ever observe the actual variables (e.g. What are the simultaneous values of σ_x and σ_z for this electron?)&lt;/li&gt;
&lt;li&gt;&lt;em&gt;&lt;a href="https://en.wikipedia.org/wiki/Cognitive_closure_(philosophy)"&gt;Cognitive closure&lt;/a&gt;&lt;/em&gt;: we’re not smart enough to answer it / we lack certain faculties (like echolocation say). (e.g. “What is it like to be a bat?” It kinda looks like quantum gravity could also be an example.)&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Cognitive bias&lt;/em&gt;: we’re not rational enough to answer it (e.g. “How biased am I?”)&lt;/li&gt;
&lt;li&gt;I suppose I should mention the original “&lt;a href="https://www.gleech.org/gut-epistemics"&gt;epistemic barrier&lt;/a&gt;”, the putative conceptual barrier between the mind and the world.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="-the-question-is-logical-andor-ontic-andor-epistemic-idk"&gt;??? (the question is logical and/or ontic and/or epistemic idk)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Observer effects&lt;/em&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;&lt;a href="https://en.wikipedia.org/wiki/Back_action_(quantum)"&gt;Back action&lt;/a&gt;&lt;/em&gt; (e.g. What’s the pre-measurement spin of this electron? What are the “real” observables in this quantum system? or “Measure the &lt;a href="https://en.wikipedia.org/wiki/Quantum_Zeno_effect"&gt;full-speed&lt;/a&gt; time evolution for this system”.)&lt;/li&gt;
&lt;li&gt;&lt;em&gt;&lt;a href="https://en.wikipedia.org/wiki/Theory-ladenness"&gt;Theory-ladenness&lt;/a&gt;&lt;/em&gt;:
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Semantic theory-dependence&lt;/em&gt;: Our observational vocabulary already presupposes theoretical commitments that prejudge the answer (e.g. What’s the rest mass of an electron, without assuming special relativity? What’s the ‘real’ temperature of the CMB independent of blackbody theory?)&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Perceptual theory-dependence&lt;/em&gt;: Our perception is shaped by theoretical expectations, so we cannot see the answer “directly” (e.g. What do electron tracks really look like? What does this fMRI show before we apply the hemodynamic response model?)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://plato.stanford.edu/entries/future-contingents/"&gt;Future contingents&lt;/a&gt;&lt;/em&gt;: it doesn’t have an answer yet so we have to wait. (e.g. “What lottery numbers will win next week?”)&lt;a href="#fn:2" id="fnref:2"&gt;2&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Hysteresis&lt;/em&gt;: we can’t answer it because we weren’t there at the start and the system remembers. (e.g. What’s the magnetic moment of this material at field strength H? When will this old rope snap? At what temperature will this water freeze?)
&lt;!-- and non-Markovian processes --&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Reflexivity and strange loops&lt;/em&gt;: the thing in question is self-referential or causally dense, or it [changes](https://en.wikipedia.org/wiki/Demand_characteristics) when answered, and so can’t be picked apart. (e.g. “What’s the best method for finding the best method?”; “Which level of description is fundamental in this self-referential system?”, “Which part of your mind is the real you?”.) See also &lt;i&gt;antinomy&lt;/i&gt;.&lt;/li&gt;
&lt;/ul&gt;
Finally there is the great risk this post takes (and all not-totally-technical writing):
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Model error leading to spurious impossibility&lt;/em&gt;: you apply an analytic impossibility result (like the above) to a "synthetic" real system where it doesn't apply. You overinterpret a narrow thing; you assume the world fits the assumptions of the formal proof but it doesn't; you summarise a technical result in natural language and imply that it's more general than it is. Most invocations of Gödel are spurious; &lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;You will have done really well if even &lt;em&gt;once&lt;/em&gt; in your life you fail to answer a question for these reasons. Getting so far means you have avoided hundreds of punji traps, claymores, nerve gasses, madnesses.&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id="see-also"&gt;See also&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.gleech.org/dark-math"&gt;The great majority of unusable maths&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.gleech.org/no-philosopher"&gt;Against philosophy&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://scottaaronson.blog/?p=9243"&gt;Aaronson&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2002.06467"&gt;Gelman and Yao&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.goodreads.com/book/show/1077040.The_Unknowable"&gt;The Unknowable&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.goodreads.com/book/show/17841838-the-outer-limits-of-reason"&gt;The Outer Limits of Reason&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://beytulhikme.org/index.jsp?mod=makale_tr_ozet&amp;amp;makale_id=65157"&gt;Epistemic Options in the Face of Epistemic Barriers&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=""&gt;Epistemic Boundedness and The Universality of Thought&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;div class="footnotes"&gt;
&lt;ol&gt;
&lt;!-- 1 --&gt;
&lt;li class="footnote" id="fn:1"&gt;
See also &lt;i&gt;contradictory presupposition&lt;/i&gt;, like "What happens when an immovable object meets an irresistible force?".&lt;br /&gt;&lt;br /&gt;
We could also pack in a very big one here, in principle: that all mathematics is only &lt;a href="https://philosophy.stackexchange.com/questions/103031/a-problem-i-noticed-with-if-then-ism-in-the-philosophy-of-mathematics"&gt;&lt;i&gt;conditionally&lt;/i&gt; true&lt;/a&gt; (i.e. conditional on the axioms involved) and so it's logically possible that we are making some false assumption and that the whole thing is actually inconsistent. This is very unlikely, but not for mathematical reasons.
&lt;/li&gt;
&lt;li class="footnote" id="fn:2"&gt;
However: if determinism holds, then there is an answer, we just don't have epistemic access. If indeterminism holds, then there is no fact of the matter.
&lt;/li&gt;
&lt;li class="footnote" id="fn:3"&gt;
Some other possible places for this entry:&lt;br /&gt;&lt;br /&gt;
Copenhagen: genuinely no answer exists before the measurement occurs → so indeterminacy belongs in "logical problems"&lt;br /&gt;
Many-worlds: there is an answer (in fact all outcomes occur) → epistemic problem&lt;br /&gt;
Hidden variables: there is an answer, we just can't access it → epistemic problem&lt;br /&gt;
QBism: the question is malformed (since probabilities are subjective) → conceptual problem
&lt;/li&gt;
&lt;li class="footnote" id="fn:4"&gt;
Uncomputability is here in "ontic" because whether we could build a halting oracle actually depends on how physics works, not just the nature of logic. The physical Church-Turing thesis (that Turing machines exhaust actual computation) is an empirical claim. And it's not logically impossible for there to be weird shit like &lt;a href="https://sites.socsci.uci.edu/~jmanchak/otposigr.pdf"&gt;Malament-Hogarth spacetime&lt;/a&gt; or &lt;a href="https://cacm.acm.org/opinion/hypercomputation"&gt;analogue infinite precision&lt;/a&gt;, so Turing uncomputability is not a logical limit.
&lt;/li&gt;
&lt;li class="footnote" id="fn:5"&gt;
This is assuming indeterminacy theory or vagueness nihilism. Epistemicism ("there is an answer, we just can't know what it is") obviously belongs in "epistemic problems", as does contextualism ("the boundary exists but varies, and we may not know the relevant contextual features that let us pick it"). Degree theory/fuzzy logic ("reality itself admits degrees of truth; the boundaries are objectively fuzzy") is obviously ontic.
&lt;/li&gt;
&lt;li class="footnote" id="fn:6"&gt;
Note that Gödel's proof uses an &lt;a href="https://en.wikipedia.org/wiki/%CE%A9-consistent_theory"&gt;unintuitive, strong sense&lt;/a&gt; of consistency, but I think &lt;a href="https://en.wikipedia.org/wiki/Rosser%27s_trick"&gt;Rosser's&lt;/a&gt; recovers the simple reading and makes the above discussion a sensible lossy statement.
&lt;/li&gt;
&lt;li class="footnote" id="fn:7"&gt;
&lt;/li&gt;
&lt;li class="footnote" id="fn:8"&gt;
Arrow's theorem only touches ordinal and deterministic voting systems, but Gibbard–Satterthwaite is general. David Sartor: "I don't think Earth has any important probabilistic elections, but cardinal elections happen in some cities, and in the LessWrong Review."
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;Ideological and social problems&lt;/h3&gt;
&lt;div&gt;
Not getting into this here because they're not as fundamental but I will also mention:
&lt;ul&gt;
&lt;li&gt;&lt;i&gt;Quietism&lt;/i&gt;: we don't view it as answerable so we don't try to answer it (e.g. the attitude of the Copenhagen interpretation toward unobserved reality)&lt;/li&gt;
&lt;li&gt;&lt;i&gt;Positivism&lt;/i&gt;: we refuse to answer because we restrict ourselves to (what-we-consider) observables.&lt;/li&gt;
&lt;li&gt;&lt;i&gt;Informal philistinism&lt;/i&gt;: we use words alone instead of mathematics to answer it (e.g. &lt;a href="https://en.wikipedia.org/wiki/Pangenesis"&gt;pangenesis theory&lt;/a&gt; and its like instead of Mendelian genetics)&lt;/li&gt;
&lt;li&gt;&lt;i&gt;Armchair philosophy&lt;/i&gt;: we use only apriori reasoning &lt;i&gt;instead&lt;/i&gt; of going and looking (e.g. &lt;a href="https://www.nature.com/articles/nn.2795"&gt;three centuries&lt;/a&gt; of philosophical analysis of &lt;a href="https://en.wikipedia.org/wiki/Molyneux%27s_problem"&gt;Molyneux's problem&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;i&gt;Basic research ethics&lt;/i&gt;: when getting a direct answer would be wrong. (e.g. social or &lt;a href="https://en.wikipedia.org/wiki/Language_deprivation_experiments"&gt;linguistic&lt;/a&gt; interventions on children)&lt;/li&gt;
&lt;li&gt;&lt;i&gt;Disinformation, chilling effects, retaliation, institutional capture&lt;/i&gt;: powerful people don't want it to be answered. &lt;/li&gt;
&lt;!-- __Social epistemology breakdown__ ( Collective inquiry mechanisms fail through incentives, coordination, or lock-in [Publication bias, paradigm rigidity, tacit knowledge loss] --&gt;
&lt;/ul&gt;
&lt;!-- __&gt; Suppose a man born blind... [were] by his touch to distinguish between a cube and a sphere. Suppose... the blind man be made to see__ ( ...before he touched them, [could he] now distinguish and tell which is the globe, which the cube? --&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div id="simple" class="tabContent"&gt;
&lt;br /&gt;Sometimes we just can't answer a question, even in principle. This is a list of the problems that make that true:
&lt;br /&gt;
&lt;h3&gt;Conceptual Problems: the question is broken&lt;/h3&gt;
&lt;i&gt;A malformed question or category error&lt;/i&gt; is when you ask for something that doesn't make sense. Questions like "what colour is justice?" or "is sigmoid jealous of ReLU?" fail because they apply properties to things that can't have those properties. Justice isn't the sort of thing that has colour, and mathematical functions don't have emotions. Similarly, asking about "absolute time" or "absolute coordinates" presupposes a framework (Newtonian absolute space and absolute time) that doesn't correspond to physical reality.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Strong incommensurability&lt;/i&gt; is when a question tries to mix incompatible theoretical frameworks. For instance, you can't ask about "the Einsteinian mass of phlogiston" because phlogiston doesn't exist in any framework that includes Einsteinian physics, and you can't ask about "the wavefunction of this classical field" because classical fields don't have wavefunctions. The question assumes you can translate concepts between frameworks that are mutually exclusive.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Vagueness&lt;/i&gt; means the question doesn't pick out anything specific enough to answer. "When exactly did you become an adult?" has no precise answer because "adult" is a vague category: there's no one moment where you transition from non-adult to adult. The concept has fuzzy boundaries, and so you can't get a single exact answer. &lt;br /&gt;&lt;br /&gt;
&lt;i&gt;False assumptions&lt;/i&gt; render a question unanswerable because, while the question is grammatically clear and seems to make sense, it presupposes something false. "Have you stopped beating your wife?" can't be answered yes or no if you never beat your wife in the first place. Similarly, "Is personality based on nature or nurture?" presupposes a false dichotomy, when the reality involves complex interactions between both.&lt;br /&gt;&lt;br /&gt;
&lt;h3&gt;Logical Problems: no answer exists&lt;/h3&gt;
&lt;i&gt;Antinomy&lt;/i&gt; is when a question is self-referential or circular in a way that prevents any consistent answer. "Is this sentence false?" can't be true (because then it would be false as it claims) and can't be false (because then it would be true). The structure of the question itself creates a logical impossibility.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Incompleteness&lt;/i&gt; and &lt;i&gt;undefinability&lt;/i&gt; are fundamental limits which only come up in heavily formalised questions. You can't prove that Peano arithmetic is consistent using only Peano arithmetic itself, and you can't define arithmetic truth within arithmetic. For deep reasons, these kinds of question are impossible to answer when you're using any formal system powerful enough to be useful.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Mathematical independence&lt;/i&gt; means a question is neither provable nor disprovable from your chosen axioms. Whether the continuum hypothesis is true can't be decided within standard set theory (ZFC). The statement might be true in some models of ZFC and false in others. Similarly, certain values like the busy beaver number BB(7910) are independent of ZFC, meaning no proof can establish their value within that system.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Anthropic self-locating effects&lt;/i&gt; create unanswerable questions because your very existence as an observer depends on certain conditions being true. You can't determine the probability you're a Boltzmann brain (a spontaneous fluctuation that created your current conscious state) because if you were one, you'd still observe exactly what you observe now. Your existence selects for certain observations, making the underlying probability inaccessible.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Non-uniqueness&lt;/i&gt; is when multiple equally valid answers exist. It's only a problem when you're demanding a single correct answer. The electromagnetic potential at a point isn't uniquely determined because you can add any gradient of a scalar function without changing the physics. The question "what is the electromagnetic potential?" thus has infinitely many correct answers, not because we're ignorant, but because the quantity itself isn't uniquely defined.&lt;br /&gt;&lt;br /&gt;
&lt;h3&gt;Ontic Problems: the answer is physically inaccessible&lt;/h3&gt;
&lt;i&gt;Uncomputability&lt;/i&gt; means no computer or mind, regardless of time or memory, can provide an answer. You can't write a program that determines if an arbitrary program will halt, or if an arbitrary Diophantine equation has integer solutions. They're provably impossible to solve algorithmically. Asking "is this sequence truly random?" is uncomputable because there's no algorithm that can verify true randomness.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Computational intractability&lt;/i&gt; means the answer exists and is computable in principle, but would require more time than is physically available, even if you converted the entire universe into a computer. Optimal scheduling problems with all constraints, or cracking a well-encrypted file, are solvable in principle but require checking so many possibilities that they're effectively impossible. The difference from uncomputability is that these problems could be solved with enough resources; they're just practically impossible.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Inapproximability&lt;/i&gt; is worse: even finding an approximate answer takes too long. For some problems like finding maximum cliques in graphs or chromatic numbers, you can't even get close to the answer in reasonable time.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Quantum indeterminacy&lt;/i&gt; (on some interpretations of quantum mechanics) means physics is fundamentally random, so there's no answer to questions about individual quantum events. "When will this radium atom decay?" has no answer because the decay is genuinely random, not merely unpredictable due to our ignorance. The universe itself hasn't determined when it will happen until it actually happens.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Physical constraints, speed limits, and no-go theorems&lt;/i&gt; make certain questions unanswerable because of fundamental physical laws. Information beyond your cosmological horizon is forever inaccessible because space itself is expanding faster than light can travel. You can't measure a quantum state in two incompatible bases simultaneously, and you can't build a perfectly accurate interferometer because of quantum limits on measurement precision.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Mixing&lt;/i&gt; means the information you need to answer is irreversibly lost because the process erases fine details. If you want to know the exact microstate of the gas molecules in your room an hour ago, that information is gone. The system has thermalised, and the microscopic details have been scrambled into macroscopic averages that can't be reversed.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Non-ergodicity&lt;/i&gt; is the opposite problem: the system doesn't mix, so it can't explore all possible states and reach equilibrium. Glass is a classic example. "What equilibrium will this glassy system reach?" may have no answer because the system is trapped in a local configuration and will never reach the global equilibrium, even given infinite time.&lt;br /&gt;&lt;br /&gt;
Measurement Problems:&lt;br /&gt;
&lt;i&gt;Disturbance or decoherence&lt;/i&gt; means your measuring instruments mess with what you're measuring. In quantum mechanics, any measurement strong enough to extract information collapses the quantum state. You can use weak measurements to minimise this, but you can never eliminate the disturbance entirely. The act of observation changes what you're observing.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Complementarity&lt;/i&gt; is when we can't answer it because the answer was excluded by another question we asked first. Measuring one property precisely makes it impossible to measure a second property precisely. If you measure an electron's position very accurately, you necessarily disturb its momentum, making it impossible to answer "where is this electron and how much momentum does it have?" The two properties can't be simultaneously known with arbitrary precision.&lt;br /&gt;&lt;br /&gt;
&lt;h3&gt;Epistemic Problems: we can't get at the answer&lt;/h3&gt;
&lt;i&gt;Underdetermination or unidentifiability&lt;/i&gt; is when multiple different theories or explanations are all consistent with the available evidence. Is spacetime fundamentally Lorentzian? Which interpretation of quantum mechanics is correct? These questions might have answers, but the evidence we can gather doesn't distinguish between the alternatives. The data underdetermines the theory.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Chaos&lt;/i&gt; means tiny uncertainties in initial conditions grow exponentially, making long-term prediction impossible. You can't predict where a double pendulum will be after 100 Lyapunov times (the characteristic timescale of exponential divergence) because you'd need to measure the initial conditions to impossible precision. Weather prediction fails beyond about two weeks for the same reason: accumulating uncertainties destroy predictive power. This is epistemic because there is an answer but we can never measure the initial conditions well enough.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Computational irreducibility&lt;/i&gt; is if there's no shortcut to the answer; you have to run the full process or simulation. For certain cellular automata or complex systems, predicting the state after many steps is just as hard as actually running the system for that many steps. There's no compressed description or formula; the answer is inseparable from the process of computation itself.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Hidden variables&lt;/i&gt; are quantities that affect the system but can never be directly observed. In quantum mechanics (on some interpretations), you can't simultaneously know the values of non-commuting observables like spin in different directions. The question "what are the simultaneous values of σ&amp;lt;/i&amp;gt;x and σ&amp;lt;/i&amp;gt;z for this electron?" is unanswerable because these variables, if they exist, are hidden from observation.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Cognitive closure&lt;/i&gt; is the hypothesis that humans might lack the cognitive capacity to understand certain problems, like a dog can't understand calculus. "What is it like to be a bat?" might be inaccessible because our architecture can't simulate bat consciousness. Quantum gravity might be another example: perhaps the correct theory exists but is too complex for the human mind to grasp.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Cognitive bias&lt;/i&gt; is patterned irrational thinking that prevents us from answering questions accurately. "How biased am I?" is nearly impossible to answer because your biases affect your assessment of your own biases. &lt;br /&gt;&lt;br /&gt;
&lt;i&gt;The original "&lt;i&gt;epistemic barrier&lt;/i&gt;" in philosophy is the supposed conceptual gap between the mind and the world. How can we know that our perceptions and thoughts correspond to reality when all we have direct access to is our own mental states? &lt;br /&gt;&lt;br /&gt;
&lt;h3&gt;Confusing problems: you can't tell if it's logic, physics, or epistemics&lt;/h3&gt;
&lt;i&gt;Observer Effects:&lt;br /&gt;
&lt;i&gt;Back action&lt;/i&gt; in quantum mechanics is the act of measurement physically affecting the system in ways you can't correct for. "What's the pre-measurement spin of this electron?" seems unanswerable: the spin doesn't have a definite value until measured. The quantum Zeno effect shows that continuous observation can even freeze a system's evolution entirely.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Semantic theory-dependence&lt;/i&gt; is if our observational vocabulary already assumes theoretical commitments that prejudge the answer. "What's the rest mass of an electron without assuming special relativity?" is problematic because the very concept of "rest mass" is defined within the framework of special relativity. You can't ask the question without importing the theoretical framework it presupposes.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Perceptual theory-dependence&lt;/i&gt; is the idea that our perception is shaped by our theoretical expectations, so we can't observe "directly." What do electron tracks in a cloud chamber really look like? What does an fMRI scan show before applying the hemodynamic response model? Our observations are always already interpreted through theoretical lenses, making theory-independent observation impossible.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Future contingents&lt;/i&gt; are questions about events that haven't happened yet and might not be determined in advance. "What lottery numbers will win next week?" has no answer yet if the lottery is truly random. On an indeterministic interpretation of physics, the future doesn't exist to be known. You just have to wait for it to happen.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Hysteresis&lt;/i&gt; is when the current state depends on the history of how you got there, so you need information about the past, and this is usually unavailable. "What's the magnetic moment of this material at field strength H?" depends on the path you took through magnetic field space. "When will this old rope snap?" depends on its entire stress history. Without that historical information, you can't answer the question.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Reflexivity and strange loops&lt;/i&gt; are when the thing you're asking about is self-referential (or changes when you try to answer the question). "What's the best method for finding the best method?" creates infinite regress. "Which level of description is fundamental in this self-referential system?" can't be answered from outside because there is no outside and no fundamental level.&lt;br /&gt;&lt;br /&gt;
&lt;h3&gt;Ideological/Social Problems: choosing not to answer&lt;/h3&gt;
&lt;i&gt;Quietism&lt;/i&gt; is the attitude that certain questions are meaningless or not worth pursuing, so we don't try to answer them. The Copenhagen interpretation of quantum mechanics takes this stance toward questions about unobserved quantum reality: if you can't measure it, don't ask about it. This is a methodological choice that declares certain questions out of bounds.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Positivism&lt;/i&gt; restricts inquiry to observables, refusing to answer questions about theoretical entities. This philosophical stance says we should only talk about what can be directly observed or measured, making questions about underlying mechanisms or unobservable causes illegitimate by definition. We choose not to try because we view the rest as meaningless.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Informal philistinism&lt;/i&gt; is the failure to use mathematics when it's necessary. Pre-Mendelian genetics like pangenesis theory used verbal descriptions and metaphors instead of mathematical models. The result was theories that couldn't make precise predictions or be rigorously tested. &lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Armchair philosophy&lt;/i&gt; means using only apriori reasoning instead of empirical investigation. Molyneux's problem (whether a blind person given sight could recognise by vision what they'd previously known by touch) was debated for three centuries until someone actually collected the data. The answer was (in principle!) available all along through empirical investigation.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Basic research ethics&lt;/i&gt; prevents us from answering certain questions when getting direct answers would be morally wrong. You can't do controlled experiments on children to answer the big questions about linguistic deprivation or social development. We rightly choose not to pursue it.&lt;br /&gt;&lt;br /&gt;
&lt;i&gt;Disinformation, chilling effects, retaliation, and institutional capture&lt;/i&gt; occur when powerful actors actively prevent questions from being answered. This might involve suppressing research, threatening researchers, manipulating publication, or taking over the institutions that should be investigating.&lt;br /&gt;&lt;br /&gt;
&lt;br /&gt;&lt;br /&gt;&lt;br /&gt;
You will have done really well if even &lt;i&gt;once&lt;/i&gt; in your life you fail to answer a question for these reasons. Getting so far means you have avoided hundreds of punji traps, claymores, nerve gasses, madnesses.&lt;br /&gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/div&amp;gt;
&lt;/i&gt;&lt;/i&gt;&lt;/div&gt;</description><pubDate>Sun, 02 Nov 2025 00:00:00 +0000</pubDate><link>https://www.gleech.org/barriers</link><guid isPermaLink="true">https://www.gleech.org/barriers</guid><category>philosophy,</category><category>science,</category><category>epistemology,</category><category>maths,</category><category>computers,</category><category>lists,</category><category>rationality,</category><category>metaphysics,</category><category>mind,</category><category>research,</category><category>conceptual-analysis,</category><category>encompassing,</category><category>holes</category></item><item><title>What god has to do in mathematics</title><description>&lt;p&gt;He &lt;a href="https://en.wikipedia.org/wiki/Axiom_of_choice"&gt;picks&lt;/a&gt; from uncountably many sets simultaneously without an algorithm.&lt;/p&gt;
&lt;p&gt;He preserves us from &lt;a href="https://en.wikipedia.org/wiki/Null_set"&gt;sets&lt;/a&gt; of measure zero and the &lt;a href="https://en.wikipedia.org/wiki/Non-measurable_set#Consistent_definitions_of_measure_and_probability"&gt;non-measurables&lt;/a&gt;. In his loving arms infinity minus infinity equals whatever we need it to. He takes five loaves of spheres and two small spheres and &lt;a href="https://en.wikipedia.org/wiki/Banach%E2%80%93Tarski_paradox"&gt;feeds the host&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;He makes the inaccessible cardinals. He stops the inaccessible cardinals from accessing us&lt;/p&gt;
&lt;p&gt;He is risen, he is &lt;a href="https://en.wikipedia.org/wiki/Lift_(mathematics)"&gt;lifted&lt;/a&gt;, and he is &lt;a href="https://simons.berkeley.edu/talks/why-born-probabilities"&gt;Born&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;He transubstantiates our mortal filters into ultrafilters &lt;a href="https://en.wikipedia.org/wiki/Ultrafilter_on_a_set#The_ultrafilter_lemma"&gt;for free&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;He knows the &lt;a href="https://www.goodreads.com/review/show/2509368481"&gt;truth&lt;/a&gt; of the continuum hypothesis but saves us from it.&lt;/p&gt;
&lt;p&gt;He lets us pretend we can do functional analysis, &lt;a href="https://en.wikipedia.org/wiki/Baire_category_theorem"&gt;keeping&lt;/a&gt; our intersections dense by tucking away the invisible point we could never reach.&lt;/p&gt;
&lt;!-- He gently places proper classes in scope of our quantifiers
He bestows d-separation on our graphs "the only bridge we have between the causal assumptions in our model and what we can expect to observe in our data."
He fixes the precise Grothendieck Universe that we live in --&gt;
&lt;p&gt;Quietly he makes actual of potential infinity.&lt;/p&gt;
&lt;p&gt;Silently he &lt;a href="https://en.wikipedia.org/wiki/Meta-circular_evaluator#Self-interpreters"&gt;transfers&lt;/a&gt; the evaluation strategy from heaven (the meta-language) to earth (the source language).&lt;/p&gt;
&lt;p&gt;He forgives us our renormalizations.&lt;/p&gt;
&lt;p&gt;When we are weak and know only subspace, he &lt;a href="https://en.wikipedia.org/wiki/Hahn%E2%80%93Banach_theorem"&gt;carries&lt;/a&gt; our functional to the full space.&lt;/p&gt;
&lt;p&gt;Ultimately he makes it all &lt;a href="https://en.wikipedia.org/wiki/Hilbert%27s_second_problem"&gt;cohere&lt;/a&gt;. The accidents of our &lt;a href="https://philosophy.stackexchange.com/questions/103031/a-problem-i-noticed-with-if-then-ism-in-the-philosophy-of-mathematics"&gt;conditionals&lt;/a&gt; are treatable as unconditional. “I am that I am”, what makes the axioms hold.&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://en.wikipedia.org/wiki/Constructivism_(philosophy_of_mathematics)"&gt;Various&lt;/a&gt; &lt;a href="https://en.wikipedia.org/wiki/Intuitionistic_logic"&gt;atheists&lt;/a&gt; &lt;a href="https://en.wikipedia.org/wiki/Ultrafinitism"&gt;try&lt;/a&gt; &lt;a href="https://en.wikipedia.org/wiki/Reverse_mathematics"&gt;to do&lt;/a&gt; &lt;a href="https://www.sciencedirect.com/science/article/pii/0304397575900171?via%3Dihub"&gt;without&lt;/a&gt; &lt;a href="https://www.lesswrong.com/posts/QmWNbCRMgRBcMK6RK/the-absolute-self-selection-assumption#Problem__3__The_Born_Probabilities"&gt;him&lt;/a&gt; and do not prevail.&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;
&lt;a href="#fn:1" id="fnref:1"&gt;1&lt;/a&gt;&lt;/p&gt;
&lt;div class="footnotes"&gt;
&lt;ol&gt;
&lt;!-- 1 --&gt;
&lt;li class="footnote" id="fn:1"&gt;
The joke was initially supposed to be about actual mostly-unacknowledged holes in the theoretical structure: ridiculously strong assumptions, suppressed premises, things we ignore &lt;i&gt;and which we get away with&lt;/i&gt;.&lt;br /&gt;&lt;br /&gt;
But Choice is too well-known and contested to be a great example and I ended up just covering a bunch of nonconstructive results, which aren't as mysterious as I'd like to justify the "god did it" gag.&lt;br /&gt;&lt;br /&gt;
Basically I don't know obscure unacknowledged issues. But someone who does could write a good version of this post.
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</description><pubDate>Tue, 28 Oct 2025 00:00:00 +0000</pubDate><link>https://www.gleech.org/god</link><guid isPermaLink="true">https://www.gleech.org/god</guid><category>maths,</category><category>encompassing,</category><category>holes</category></item><item><title>prospective science please</title><description>&lt;blockquote&gt;
&lt;p&gt;Time… burns without leaving ashes&lt;/p&gt;
&lt;/blockquote&gt;
&lt;center&gt;— Elsa Triolet&lt;/center&gt;
&lt;h2 id="case-study-whither-google-search"&gt;Case study: whither Google Search?&lt;/h2&gt;
&lt;p&gt;For now, Google is the premier information infrastructure on Earth. It is the most-visited website. It handles &lt;a href="https://www.semrush.com/blog/google-search-statistics/"&gt;14 billion&lt;/a&gt; search queries a day, nearly two for everyone on earth. &lt;a href="#fn:1" id="fnref:1"&gt;1&lt;/a&gt; As of writing it’s more than the traffic of &lt;a href="https://techcrunch.com/2025/07/21/chatgpt-users-send-2-5-billion-prompts-a-day/"&gt;ChatGPT&lt;/a&gt;, Claude, Wikipedia, Bing, and Duckduckgo put together. 1 / 500 of your waking life is spent on this site, probably, and its decisions determine how you spend much more time than that. &lt;a href="#fn:3" id="fnref:3"&gt;3&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Google Search effectively controls the epistemology of the world — &lt;a href="https://x.com/zetalyrae/status/1794590304101957645"&gt;@atomgardner&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;So… how’s it doing? Is search quality changing? Opinions &lt;a href="https://danluu.com/seo-spam/"&gt;differ&lt;/a&gt; but it does &lt;em&gt;&lt;a href="https://x.com/danluu/status/1730705885037801686"&gt;seem&lt;/a&gt;&lt;/em&gt; to have gotten much worse over the last decade.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://www.gleech.org/img/goog/danluu.png" /&gt;
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;OK unc, but surely we have actual quantitative studies of something this important?&lt;/p&gt;
&lt;p&gt;Broadly: no. In the last 25 years, I could only find &lt;em&gt;one&lt;/em&gt; major empirical study measuring search &lt;em&gt;quality&lt;/em&gt; (&lt;a href="https://link.springer.com/chapter/10.1007/978-3-031-56063-7_4"&gt;Bevendorff et al. 2024&lt;/a&gt;). It’s fine, and they deserve huge credit for doing what no one did, but it doesn’t really answer the question and only covers 2022-3, well after the subjective decline. We had two decades to notice that this was really important and to collect data about it ahead of the decline. We didn’t.&lt;/p&gt;
&lt;p&gt;I &lt;a href="https://www.gleech.org/psych"&gt;complain&lt;/a&gt; &lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC12397490/"&gt;about&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2407.12220"&gt;academia&lt;/a&gt; a lot, but this one failure dwarfs almost all others. We have better longitudinal data on bird migration than on the modal channel for belief formation. Academia and civil society had 25 years to run this and didn’t, and our chance to measure it is gone. (Google could maybe still do this study given the logs, but why would they?)&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;In fairness&lt;/h3&gt;
&lt;div&gt;
There's a huge attribution problem (was Google degrading or were they accurately representing &lt;i&gt;the web itself&lt;/i&gt; degrading?&lt;br /&gt;&lt;br /&gt;
Google's quality is constantly under attack by the SEO industry and the bots. This would be enough to cause the observed decline without Google doing anything wrong per se. They invest large amounts into countermeasures. It doesn’t work (see this identical effort from 2022).&lt;br /&gt;&lt;br /&gt;
The defence against the first (non-AI) wave of SEO spam already cost us &lt;a href="https://www.benlandautaylor.com/p/the-ddos-attack-of-academic-bullshit"&gt;a lot&lt;/a&gt;:
&lt;blockquote&gt;in 1999… When you looked for information about how to tell if your bread is rising correctly, or about South Korean cement manufacturing, or the musical influences on Igor Stravinsky, or whatever weird thing, Google would pull up high-quality reference material, or blogs and BBS arguments among disagreeable weirdos who specialized in the subject… A cottage industry arose of finding some search term and churning out low-cost copy on the subject in order to serve ads to people trying to find real information. Specialists in “search engine optimization” popularized their techniques as consultants to big companies, and before long this became standard practice. In their efforts to keep these problems from getting totally out of hand, Google and other search engines weighted search results towards a whitelist of standard lowest-common-denominator websites. The long tail of the internet could no longer be found from a simple search.&lt;/blockquote&gt;&lt;br /&gt;
Then there's the Web 2.0 turn to &lt;a href="https://en.wikipedia.org/wiki/Closed_platform#Examples"&gt;walled gardens&lt;/a&gt;, platforms which search engines can't really index. That's mostly not Google's doing.
&lt;/div&gt;
&lt;h3&gt;In unfairness&lt;/h3&gt;
&lt;div&gt;
It looks like Google itself systematically degraded search by:
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://web.archive.org/web/20250206165606/https://searchengineland.com/search-ad-labeling-history-google-bing-254332"&gt;Making ads&lt;/a&gt; look identical to organic results (2011-2020)&lt;/li&gt;
&lt;li&gt;Zero-click answers that keep users in Google's walled garden&lt;/li&gt;
&lt;li&gt;Prioritizing engagement over quality (the "code yellow" pivot of 2019)&lt;/li&gt;
&lt;li&gt;Soft whitelisting of large publishers&lt;/li&gt;
&lt;li&gt;Decreasingly terrible forced AI content&lt;/li&gt;
&lt;/ul&gt;
As the noted scholars Brin and Page &lt;a href="http://infolab.stanford.edu/pub/papers/google.pdf#page=18"&gt;said in 1998&lt;/a&gt;:
&lt;blockquote&gt;The goals of the advertising business model do not always correspond to providing quality search to users. For example, in our prototype search engine one of the top results for cellular phone is "The Effect of Cellular Phone Use Upon Driver Attention", a study which explains in great detail the distractions and risk associated with conversing on a cell phone while driving. This search result came up first because of its high importance as judged by the PageRank algorithm, an approximation of citation importance on the web [Page, 98]. It is clear that a search engine which was taking money for showing cellular phone ads would have difficulty justifying the page that our system returned to its paying advertisers. For this type of reason and historical experience with other media [Bagdikian 83], we expect that advertising funded search engines will be inherently biased towards the advertisers and away from the needs of the Consumers.&lt;br /&gt;&lt;br /&gt;... we believe the issue of advertising causes enough mixed incentives that it is crucial to have a competitive search engine that is transparent and in the academic realm.&lt;/blockquote&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h3 id="the-evidence-we-have"&gt;The evidence we have&lt;/h3&gt;
&lt;div class="accordion"&gt;
&lt;!-- --&gt;
&lt;h3&gt;Our one actual datum: Bevendorff et al. 2024&lt;/h3&gt;
&lt;div&gt;
That one study, "&lt;a href="https://link.springer.com/chapter/10.1007/978-3-031-56063-7_4"&gt;Is Google Getting Worse?&lt;/a&gt; &lt;a href="https://link.springer.com/chapter/10.1007/978-3-031-56063-7_4"&gt;A Longitudinal Investigation of SEO&lt;/a&gt; &lt;a href="https://link.springer.com/chapter/10.1007/978-3-031-56063-7_4"&gt;Spam in Search Engines&lt;/a&gt;", monitored 7392 queries for particular product reviews on Google (proxied using Startpage), Bing, and DuckDuckGo over one year (2022-3).&lt;br /&gt;&lt;br /&gt;
It's fine but not amazing evidence:
&lt;ul&gt;
&lt;li&gt;Just one year, and a post-decline year at that&lt;/li&gt;
&lt;li&gt;Just product searches&lt;/li&gt;
&lt;li&gt;Just a couple of thousand products&lt;/li&gt;
&lt;li&gt;Focussed on spam alone&lt;/li&gt;
&lt;li&gt;Using a &lt;a href="https://www.sketchengine.eu/glossary/type-token-ratio-ttr/"&gt;dumb&lt;/a&gt; mechanical proxy measure for diversity, itself as a proxy for quality.&lt;/li&gt;
&lt;/ul&gt;
&lt;br /&gt;
Still:
&lt;ul&gt;
&lt;li&gt;29% of frontpage Google results had affiliate links compared to the random-page baseline of 2%. (The other two search engines were even worse.)&lt;/li&gt;
&lt;li&gt;That's not totally damning - Dwarkesh has affiliate links - but it is a warning sign: pages with more affiliate links were worse on a mechanical measure of quality. &lt;/li&gt;
&lt;li&gt;23%% of frontpage Google results were outright spam/review farms compared to the random-page baseline of 13%. (The other two search engines were even worse.) This is based on manual annotation, god bless them.&lt;/li&gt;
&lt;li&gt;The buried conclusion is that the very worst spam on Google actually &lt;b&gt;decreased&lt;/b&gt; over the study period. The 95th percentile of affiliate links per page fell from 50(!) to 35(!).&lt;/li&gt;
&lt;li&gt;But improvements are short-lived. They found that spam is cyclical (what they call "breathing patterns", like the rise and fall of your chest when you breathe): spam gets into results, the engines update their algorithms to squash them, spam returns in new forms.&lt;/li&gt;
&lt;li&gt;"Text quality" was decreasing in all search engines.&lt;/li&gt;
&lt;li&gt;Overall this study is just too short to measure a nonlinear phenom.&lt;/li&gt;
&lt;/ul&gt;
&lt;br /&gt;&lt;br /&gt;
&lt;!-- --&gt;
&lt;!-- --&gt;
&lt;!-- --&gt;
&lt;/div&gt;
&lt;h3&gt;What's hard about it?&lt;/h3&gt;
&lt;div&gt;
Measuring search quality properly requires:&lt;br /&gt;&lt;br /&gt;
&lt;ol&gt;
&lt;li&gt;Access to search results at scale. Google doesn't provide any examples, so you need to actively scrape it over time yourself.&lt;/li&gt;
&lt;li&gt;Longitudinal data (because one-off snapshots don't give you effects and miss adversarial dynamics)&lt;/li&gt;
&lt;li&gt;sampling representative queries. We should hit common queries and types.&lt;/li&gt;
&lt;li&gt;Baselines: you need to hit multiple search engines (to see if they're also struggling) or to randomly sample webpages (to separate "Google getting worse" from "internet getting worse")&lt;/li&gt;
&lt;li&gt;Quality is not a hard endpoint: Quality is subjective, and there's a higher bar in academia for handling such things. The valid reasons to worry are massive measurement error, low test-rest reliability, cross-cultural and intra-cultural heterogeneity...&lt;/li&gt;
&lt;/ol&gt;
But many of these things are now hundreds of times easier with LLMs!
&lt;/div&gt;
&lt;h3&gt;Leontiadis et al 2014&lt;/h3&gt;
&lt;div&gt;
&lt;a href="https://dl.acm.org/doi/pdf/10.1145/2660267.2660332"&gt;This&lt;/a&gt; is just a study of one particularly obvious and aggressive kind of spam: "search poisoning" redirection attacks where the spammers buy a previously legit URL and send you to their shit.&lt;br /&gt;&lt;br /&gt;
They tracked a set of pharmaceutical and product queries for 3.5 years (2010-3).&lt;br /&gt;&lt;br /&gt;
The rate of the frontpage having one or more of these particular attacks jumped from 30% to 60% of results. But it seems to have gotten a lot better after the EEAT update.
&lt;/div&gt;
&lt;h3&gt;Philipp et al 2014&lt;/h3&gt;
&lt;div&gt;
&lt;a href="https://arxiv.org/pdf/1712.03622"&gt;Analyses&lt;/a&gt; logs from 190,000 users over a 6 month period (late 2012 - early 2013) but only for health queries. This is one way that Google obviously got "better" over the last decade, though in a crank-minimising way rather than an &lt;a href="https://www.astralcodexten.com/p/webmd-and-the-tragedy-of-legible"&gt;info-maximising way&lt;/a&gt;.
&lt;/div&gt;
&lt;h3&gt;Proxies: Zero-click searches&lt;/h3&gt;
&lt;div&gt;
Possibly-unrepresentative tracking data from &lt;a href="https://sparktoro.com/blog/2024-zero-click-search-study-for-every-1000-us-google-searches-only-374-clicks-go-to-the-open-web-in-the-eu-its-360/"&gt;Datos&lt;/a&gt; shows a huge rise in "zero-click" searches, queries that never leave Google.
&lt;br /&gt;&lt;br /&gt;
&lt;table&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Zero-click rate&lt;/th&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2016&lt;/td&gt;
&lt;td&gt;44%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2020&lt;/td&gt;
&lt;td&gt;65%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025&lt;/td&gt;
&lt;td&gt;68%&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;br /&gt;
But this is a bundle measure: it mixes up the results being obviously too crap to bother with and Google sucking up the information from third-parties and so diverting the traffic to themselves. (Google Flights, Google Maps, featured snippets have gotten better; the Google Knowledge Graph is actually often useful; and the AI Overviews are sometimes useful and even more often falsely appear useful.)
&lt;/div&gt;
&lt;h3&gt;Proxies: UI changes as tells of motivation &lt;/h3&gt;
&lt;div&gt;
This excellent chart from &lt;a href="https://web.archive.org/web/20250206165606/https://searchengineland.com/search-ad-labeling-history-google-bing-254332"&gt;SearchEngineLand&lt;/a&gt; shows how, starting in 2011, Google Ads have been steadily made more stealthy and resultlike. &lt;a href="/img/GoogleAds_Timeline_FNL2.001.png"&gt;Click it for full-size&lt;/a&gt;:
&lt;br /&gt;&lt;br /&gt;
&lt;center&gt;
&lt;a target="_blank" width="100%" href="/img/GoogleAds_Timeline_FNL2.001.png"&gt;&lt;img src="/img/GoogleAds_Timeline_FNL2.001.png" /&gt;&lt;/a&gt;
&lt;/center&gt;
&lt;/div&gt;
&lt;h3&gt;Proxies: Journalism&lt;/h3&gt;
&lt;div&gt;
We're reduced from science to journalism. To weakly understand what happened we have to read dreadful people like Ed Zitron, who &lt;a href="https://www.wheresyoured.at/requiem-for-raghavan/"&gt;blames&lt;/a&gt; one guy, Prabhakar Raghavan (as if the org didn't give him the incentive). They did indeed &lt;a href="https://news.bloomberglaw.com/antitrust/googles-2019-code-yellow-blurred-line-between-search-ads"&gt;panic&lt;/a&gt; and prioritise ads over actual content, and the timing (2019) does correlate with some of the subjective decline. But that's all we can say from this.
&lt;/div&gt;
&lt;h3&gt;Proxies: Straw polls from memory&lt;/h3&gt;
&lt;div&gt;
&lt;img src="/img/goog/danluu.png" /&gt;
&lt;/div&gt;
&lt;h3&gt;Terrible Proxies: User engagement&lt;/h3&gt;
&lt;div&gt;
Despite all of this, Google's volume metrics are &lt;a href="https://sparktoro.com/blog/new-research-google-search-grew-20-in-2024-receives-373x-more-searches-than-chatgpt"&gt;up 20% YoY&lt;/a&gt;. Searches per searcher at historic highs.&lt;br /&gt;&lt;br /&gt;
If search was getting worse, why would people be using it more?:
&lt;br /&gt;&lt;br /&gt;
1. Normalisation of deviance: maybe users lowered their expectations&lt;br /&gt;
2. Query inflation: maybe it takes more searches to find useful results &lt;br /&gt;
3. Mobile growth: maybe spending more time on mobile means you just increase your total screentime and so googling time.
&lt;/div&gt;
&lt;h3&gt;Tiny study: Bloggers&lt;/h3&gt;
&lt;div&gt;
Dan Luu also &lt;a href="https://danluu.com/seo-spam/"&gt;did a manual test&lt;/a&gt; on six tricky queries and six search engines. Google was second-worst and didn't get any of them. On two of them no search engine worked at all. On some axes this is better than Bevendorff.
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="sum-total"&gt;Sum total&lt;/h2&gt;
&lt;p&gt;Overall it sure looks like it got worse but we don’t really know how much.&lt;/p&gt;
&lt;p&gt;What can we do about it?&lt;/p&gt;
&lt;h3 id="project-todo-archive-reconstruction"&gt;Project TODO: Archive reconstruction&lt;/h3&gt;
&lt;p&gt;You could probably do some limited good by finding common queries which &lt;a href="https://web.archive.org/web/20250000000000*/https://www.google.com/search?q=shoes"&gt;happened&lt;/a&gt; to get tracked on the Internet Archive:&lt;/p&gt;
&lt;p&gt;&lt;img src="/img/ia_shoes.jpg" /&gt;&lt;/p&gt;
&lt;h3 id="project-todo-the-second-best-time-is-now"&gt;Project TODO: The second-best time is now&lt;/h3&gt;
&lt;p&gt;You can in fact start pinging Google every day with the same query and saving the frontpages they give you.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Public search archives&lt;/strong&gt;: Either Google preserves and makes available historical search results for vetted researchers or we build it ourselves.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Crawling infrastructure&lt;/strong&gt;: Sustained funding for researchers to continuously monitor search quality across engines and domains. Relying on sending it to the Internet Archive is one cheap way but fairly buggy.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;!-- 3. **Algorithm transparency**: Not trade secrets, but basic disclosure about ranking signals, when they change, and what problems they're designed to solve. --&gt;
&lt;!-- 4. **Red team access**: Independent researchers with API access to test hypotheses about search quality, spam, and algorithmic failures. --&gt;
&lt;p&gt;Some possible research questions on Google:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;How has search quality changed for non-commercial queries?&lt;/li&gt;
&lt;li&gt;How does quality vary by domain (health, finance, news, technical)?&lt;/li&gt;
&lt;li&gt;What fraction of queries return primarily SEO spam vs. “authoritative” sources?&lt;/li&gt;
&lt;li&gt;How much of this is caused by Google?&lt;/li&gt;
&lt;li&gt;How effective are Google’s (announced) algorithm updates at &lt;em&gt;actually improving&lt;/em&gt; quality (vs. temporarily shuffling spam)?&lt;/li&gt;
&lt;li&gt;How strong is the soft whitelist? How far down the ranking have blogs slipped now?&lt;/li&gt;
&lt;li&gt;Has the rollout of AI Overviews (2024) improved or degraded quality?&lt;/li&gt;
&lt;li&gt;How much of the problem is Google’s choices vs. the degrading web?&lt;/li&gt;
&lt;li&gt;Are they picking up on publicised test queries and doing whack-a-mole on them?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;And maybe you want to start logging ChatGPT now before they add ads.&lt;/p&gt;
&lt;h2 id="conclusion-we-need-prospective-science"&gt;Conclusion: We need prospective science&lt;/h2&gt;
&lt;p&gt;That latter project makes me realise that a whole kind of science is missing: watching the world and collecting data before things happen. &lt;a href="#fn:4" id="fnref:4"&gt;4&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Most academic studies look backwards. (RCTs look forwards but at enormous expense on tiny questions.) But scraping is now hundreds of times cheaper than it used to be. Dumb first-pass &lt;a href="https://huggingface.co/learn/cookbook/en/llm_judge"&gt;classification and scoring&lt;/a&gt; is now thousands of times cheaper. Let’s just start collecting.&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Look around you. Think of the most important digital services and phenomena in the world.&lt;/li&gt;
&lt;li&gt;Come up with hypotheses about what could happen to them. Put them on &lt;a href="https://www.cos.io/initiatives/prereg"&gt;OSF&lt;/a&gt; to tie your hands.&lt;/li&gt;
&lt;li&gt;Think really hard about the design - this bit is irreversible if you want consistent, backwards-compatible results. Consider interventions as well as constant baselines.&lt;/li&gt;
&lt;li&gt;Start scraping.&lt;/li&gt;
&lt;li&gt;Check back every year.&lt;/li&gt;
&lt;li&gt;Iterate on the design if you must.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="see-also"&gt;See also&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.gleech.org/google"&gt;Google-fu&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://philarchive.org/archive/DEVRGA"&gt;Google as Epistemic Tool&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="footnotes"&gt;
&lt;ol&gt;
&lt;!-- 1 --&gt;
&lt;li class="footnote" id="fn:1"&gt;
I'm focussing on consumer queries.&lt;br /&gt;
ChatGPT = 5.4B&lt;br /&gt;
Wiki = 4.4B&lt;br /&gt;
Bing = 2.2B&lt;br /&gt;
Duckduckgo = 2.1B&lt;br /&gt;
Claude ~= 0.4B&lt;br /&gt;&lt;br /&gt;
Twitter and Reddit are harder to compare, since people usually don't put queries to them, but Google is maybe twice as big as both put together.&lt;br /&gt;
Reddit = 4.6B but not so analogous&lt;br /&gt;
Twitter = 3.6B but not so analogous&lt;br /&gt;
&lt;/li&gt;
&lt;!-- &lt;li class="footnote" id="fn:2"&gt;
One reason to do &lt;a href="https://themultiplicity.ai/"&gt;multimodel queries&lt;/a&gt; every time is basic: the three big models all use different search indices (Gemini - Google, GPT - Bing, Anthropic - Brave). I presume Kimi is some cool Chinese one with sick amounts of IP in it.
&lt;/li&gt; --&gt;
&lt;li class="footnote" id="fn:3"&gt;
Very rough estimate: &lt;br /&gt;&lt;br /&gt;
Queries per person per day = 14bn / 8bn = 1.75 &lt;br /&gt;
Say 1 min per query&lt;br /&gt;
1.75 min/day / 1440 min/day = 0.12% = 1/833&lt;br /&gt;
waking life is two-thirds = (1/833) * 3 / 2 = 1/555
&lt;/li&gt;
&lt;li class="footnote" id="fn:4"&gt;
The "&lt;a href="https://en.wikipedia.org/wiki/Credibility_revolution"&gt;Credibility Revolution&lt;/a&gt;" in economics is sort of similar, but even they don't seem to actively collect things. Instead they just cleverly notice when existing datasets have a convenient randomish discontinuity.
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;Comments&lt;/h3&gt;
&lt;div&gt;
&lt;b&gt;Tomas&lt;/b&gt; commented on 18 November 2025:
&lt;blockquote&gt;
Collecting the evolution/history of answers of LLMs to a given set of problems is brilliant, and it is a shame if it is not happening already.&lt;br /&gt;&lt;br /&gt;
Note that some (most?) of it needs to happen in secret (i.e. not published yearly, using unaffiliated API keys, etc) - not to be easily targeted by providers, whether intentionally or not (e.g. future LLMs training on it.) Then release some of it after 3 years, some after 5, ..&lt;br /&gt;&lt;br /&gt;
* How would you operationalize "the internet itself getting worse"?&lt;br /&gt;
* What questions do we ask the LLMs and monitor for change? What questions do we ask google and others, and monitor for change? (Setting this up as an ~automated process that only needs a yearly minor code update and $100/y in credit does not sound that hard otherwise!)&lt;br /&gt;
* How do we support preservation of the Internet Archive data over long horizons? You can donate to IA, but I would (also) give money to an independent org mirroring (almost) all its data reliably &amp;amp; long-term. (There were a few mirrors, quick search tells me they are gone, and there are partial/small mirrors of selected IA data.)
&lt;/blockquote&gt;
&lt;/div&gt;
&lt;/div&gt;</description><pubDate>Tue, 28 Oct 2025 00:00:00 +0000</pubDate><link>https://www.gleech.org/prospect</link><guid isPermaLink="true">https://www.gleech.org/prospect</guid><category>science,</category><category>epistemology</category></item><item><title>NY sounds</title><description>&lt;blockquote&gt;
&lt;p&gt;There are roughly three New Yorks. There is, first, the New York of the man or woman who was born here, who takes the city for granted and accepts its size and turbulence as natural and inevitable. Second, there is the New York of the commuter — the city that is devoured by locusts each day and spat out each night. Third, there is the New York of the person who was born somewhere else and came to New York in quest of something.&lt;br /&gt;&lt;br /&gt;Of these three trembling cities the greatest is the last — the city of final destination, the city that is a goal. It is this third city that accounts for New York’s high-strung disposition, its poetical deportment, its dedication to the arts, and its incomparable achievements. Commuters give the city its tidal restlessness; natives give it solidity and continuity; but the settlers give it passion. And whether it is a farmer arriving from Italy to set up a small grocery store in a slum, or a young girl arriving from a small town in Mississippi to escape the indignity of being observed by her neighbors, or a boy arriving from the Corn Belt with a manuscript in his suitcase and a pain in his heart, it makes no difference: each embraces New York with the intense excitement of first love, each absorbs New York with the fresh eyes of an adventurer, each generates heat and light to dwarf the Consolidated Edison Company.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;center&gt;— EB White&lt;/center&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="/nation-sound"&gt;I want to capture the best music in the world&lt;/a&gt;. (This will include little “world music”.) But I hit a snag - New York. It is just too huge to capture. (&lt;a href="https://open.spotify.com/playlist/0He9PPzeJfgHrjMJ5DPujT?si=W5P4eMmwSCehYJBIcTW-0g&amp;amp;pi=ZJ0uSBR4R22-D"&gt;Here’s&lt;/a&gt; what I have so far.)&lt;/p&gt;
&lt;p&gt;Why? What would a complete New York playlist have to cover? If you limited yourself just to totally new genres:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;1885-1930: Pop as we knew it (&lt;a href="https://open.spotify.com/playlist/22TWbFEbnhHv3CLCxOjoBW?si=-0SC9IIwSk6uHUE3HIBG_w&amp;amp;pi=5SwGoSP5TNmJV"&gt;Tin Pan Alley&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;1910-1950: &lt;a href="https://open.spotify.com/playlist/2cjIvuw4VVOQSeUAZfNiqY?si=IkjX9sTLTFq1akTzCO8msw"&gt;Big band&lt;/a&gt; and swing. Duke, Calloway&lt;/li&gt;
&lt;li&gt;1910-1960, 1980-2019: &lt;strong&gt;&lt;a href="https://open.spotify.com/playlist/34lA7pQ48yZ0osu0L0Khho?si=7lvxp__-TBiXPaEH2soYTw&amp;amp;pi=jOiFSjX3Sei8p"&gt;The (American) Musical&lt;/a&gt;&lt;/strong&gt;: Berlin, Gershwin, Rodgers and Hammerstein and Hart, Bernstein, Kern, Wodehouse(!), Sondheim, Brown&lt;/li&gt;
&lt;li&gt;1940-1967: &lt;strong&gt;&lt;a href="https://open.spotify.com/playlist/6qu37CqWnpFDUFsJDAdHdJ?si=7n67uKSdQsuj5TwVgvLrdw&amp;amp;pi=Ud-2MsEJTmqgg"&gt;Bebop&lt;/a&gt;&lt;/strong&gt;, hard bop, cubop, cool&lt;/li&gt;
&lt;li&gt;1943-1970: &lt;strong&gt;&lt;a href="https://open.spotify.com/playlist/4U0S060zi7Ldp6iS2Q5Xkw?si=PDhtm0gcQQWqRgCSXm6k_Q"&gt;“Latin” jazz&lt;/a&gt;&lt;/strong&gt; (Afro-Cuban, Puente, Barretto, Palmieri…)&lt;/li&gt;
&lt;li&gt;1930-1956: Globalised samba and calypso.&lt;/li&gt;
&lt;li&gt;1940-1970: &lt;strong&gt;&lt;a href="https://open.spotify.com/playlist/62mKVd7nGrd5HdhZRsxBHR?si=eAluZS6xTHCncKlhA0wQGw&amp;amp;pi=xw3hzLGQQn2i-"&gt;Folk revival&lt;/a&gt;&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;1958-1964: Pop as we know it (&lt;strong&gt;&lt;a href="https://open.spotify.com/playlist/0O2lJe2zHcqCCjGAM0pU5W?si=tffcxHjXSvSIFFM-xjE0QQ&amp;amp;pi=IFsQTFoWTeud5"&gt;Brill machine&lt;/a&gt;&lt;/strong&gt; pop rock, and thus, ultimately, K-pop) (Leiber-Stoller, Carole King et al)&lt;/li&gt;
&lt;li&gt;1960-1980: &lt;strong&gt;Various Avant-Gardes&lt;/strong&gt; (Varèse, &lt;a href="https://open.spotify.com/playlist/08iPwSfynzSNMXJwYFiiFj?si=rGDRSMLUR0G6xnrfa3QvlQ&amp;amp;pi=UJZoEbP6Sn6_u"&gt;New York School&lt;/a&gt;, the Downtown scene, loft jazz, Fluxus, No wave, Minimalism, Avant Rock, Fluxus again, Free Jazz…)&lt;/li&gt;
&lt;li&gt;1970-1980: &lt;strong&gt;&lt;a href="https://open.spotify.com/playlist/6LvkhZKZnOC2jpkTDysWhT?si=HS4Jf9HURd6fToHxrp0N6g&amp;amp;pi=z6wlFeehQlCL3"&gt;Disco&lt;/a&gt;&lt;/strong&gt; (Studio 54)&lt;/li&gt;
&lt;li&gt;1980: Globalised dub (Gibbons among the first Americans to incorporate techniques from dub production into dance music)&lt;/li&gt;
&lt;li&gt;1973-Present: &lt;strong&gt;&lt;a href="https://open.spotify.com/playlist/5M7jmBRGGEAB9whDzUjdyw?si=nvaR6l30TU-XOw7TT9ST9Q&amp;amp;pi=CvWJ8atfSEGd8"&gt;Hip-hop&lt;/a&gt;&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;1976-1980: &lt;strong&gt;&lt;a href="https://open.spotify.com/playlist/3txw9Twsf7TKyTsre8e5a4?si=sZhgwVSBRrO2POAbVh0tLw&amp;amp;pi=H0E2-mUtTxe5W"&gt;Punk&lt;/a&gt;&lt;/strong&gt;. &lt;a href="https://en.wikipedia.org/wiki/No_wave#1970s"&gt;Post-punk&lt;/a&gt; slightly preceded punk.&lt;/li&gt;
&lt;li&gt;1980-1987: &lt;a href="https://open.spotify.com/playlist/7hiyXzw0hbzypzwDNQND8O?si=uJ_Ql5REQLGXcdi4CtUKpg&amp;amp;pi=umE3CLZESwqGj"&gt;Garage house&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;1985-?: &lt;a href="https://en.wikipedia.org/wiki/Anti-folk"&gt;Antifolk&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;1998-2007: &lt;a href="https://open.spotify.com/playlist/7swwwKWs3KpKlH26TVUyAU?si=Fw78JUVHSHmnr4qM361miw"&gt;Garage rock revival&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;1985-?: &lt;a href="https://en.wikipedia.org/wiki/Totalism"&gt;Totalism&lt;/a&gt; (the opposite of minimalism)&lt;/li&gt;
&lt;li&gt;2006-2011: Bloghouse and &lt;a href="https://open.spotify.com/playlist/5H6pxrgLJrRKckNKi7C6s4?si=P5ne7MYEQP6RYsA8a8KA7Q&amp;amp;pi=vV0-9mxESnuYz"&gt;dancepunk&lt;/a&gt;, witch house, whatever.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="tech"&gt;Tech&lt;/h3&gt;
&lt;p&gt;You’d also want a separate playlist for Technical New York: the engineering breakthroughs. Spotify isn’t up to this though.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Telharmonium"&gt;first electromechanical musical instrument&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Stereophonic recording&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ieeexplore.ieee.org/document/1641311"&gt;theory of amplification&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.mixonline.com/technology/1925-western-electricbell-labs-electrical-recording-383620"&gt;electrical recording&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Moog_synthesizer"&gt;first commercial synths&lt;/a&gt;, first keyboard synth&lt;/li&gt;
&lt;li&gt;electret condenser microphone - what 90% of all current mics are.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Electronium"&gt;algorithmic music&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Bebe_and_Louis_Barron"&gt;first fully electronic film score&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;first widely-used computer music &lt;a href="https://120years.net/music-n-max-mathews-usa-1957/"&gt;program&lt;/a&gt; (MUSIC-N)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://christiansmusicmusings.wordpress.com/2018/02/13/tom-dowd-humble-music-genius-behind-the-scenes/"&gt;sliding faders&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Spector wasn’t the first to use the &lt;a href="https://en.wikipedia.org/wiki/Recording_studio_as_an_instrument#1940s%E2%80%931950s"&gt;studio as an instrument&lt;/a&gt; but he cemented it&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.grafftergallery.com/2023/02/the-breakbeat-king-innovative-djing-of.html"&gt;breakbeat&lt;/a&gt; and &lt;a href="https://en.wikipedia.org/wiki/Grand_Wizzard_Theodore"&gt;scratching&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://daily.redbullmusicacademy.com/2017/11/tom-moulton-interview/"&gt;the extended remix and 12” single&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://magazine.waxpoetics.com/article/tom-moulton-beat-doctor/"&gt;the breakdown&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="current"&gt;Current&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;Scenius seemed to stop happening at some point early in this century. What was the last musical movement to be distinctively associated with a city in Britain? Trip-hop, from Bristol? Grime, from East London? Once you get beyond 2010 it’s hard to think of any. In the US, New York is a more liveable city than it was for much of the twentieth century, and a less creative one. Portland and Austin are not quite what they were. San Francisco and Silicon Valley are perhaps the last examples of scenius, but those scenes are driven by money, not art.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;center&gt;- Ian Leslie&lt;/center&gt;
&lt;p&gt;And now? It’s still a great city. What new sounds is it cooking?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Brooklyn_drill"&gt;Brooklyn drill&lt;/a&gt;: a minor variation on UK drill. A$AP Mob: a minor variation on cloud/Southern rap. Rage (2hollis and Nettspend): a minor variation on drill.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/L.I.E.S."&gt;Outsider house&lt;/a&gt;
&lt;!-- * Jazz Revival. --&gt;&lt;/li&gt;
&lt;li&gt;New opera (Mazzoli, Lang, Wolfe, Du Yun, Muhly)&lt;/li&gt;
&lt;li&gt;More free jazzes and the “new Brooklyn complexity” (Mary Halvorson, Matana Roberts, Tomas Fujiwara, Ches Smith, Ingrid Laubrock).&lt;/li&gt;
&lt;li&gt;Post-minimalism (Oneohtrix Point Never)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://nyc-noise.com/"&gt;Brooklyn noise&lt;/a&gt;. No wave plus&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Solid, but not the same. Something ran out. Guesses:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;loss of cheap rents and squats. Relatedly: loss of discomfort and chaos.&lt;/li&gt;
&lt;li&gt;loss of the wizards. Maybe deindustrialisation means that the engineers aren’t colocated with the artistes, so New Yorkers don’t get the stream of pre-market prototypes they had for 80 years. No Bell Labs and maybe not many &lt;a href="https://en.wikipedia.org/wiki/Bebe_and_Louis_Barron"&gt;Bebes and Louises&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;loss of some of the wealth gradient? 70s New York had both extreme poverty (and so cheap spaces) and extreme wealth (and so patronage, whether from slumming trust-fund kids or moneyed venue donors)&lt;/li&gt;
&lt;li&gt;loss of (colocated) gatekeepers you could party with and extract attention rents from&lt;/li&gt;
&lt;li&gt;loss of industrial zoning allowing big noises. loss of lax enforcement&lt;/li&gt;
&lt;li&gt;Institutionalisation of weird music. (All of the above new genres were either deeply commercial or deeply street and subcultural.) Now:
&lt;ul&gt;
&lt;li&gt;academic capture of alternative music? Careerification, taming, sexlessness?&lt;/li&gt;
&lt;li&gt;philanthropic capture of alternative music? Grant applications and awards legibilise.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;But probably the strongest factors are the universal ones:
&lt;ul&gt;
&lt;li&gt;ideas getting harder to find?&lt;/li&gt;
&lt;li&gt;loss of modernist aggro?&lt;/li&gt;
&lt;li&gt;retromania, &lt;a href="https://www.honest-broker.com/p/is-old-music-killing-new-music"&gt;catalog dominance&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;the general dematerialisation of social energy Leslie alludes to. The extraction and dissipation of cultural meaning through homogenous phones located in any mere geography. The loss of friction and the loss of mystery and the loss of random physical collision.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id="see-also"&gt;See also&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.goodreads.com/book/show/41817533-everybody-s-doin-it"&gt;Everybody’s Doin’ It: Sex, Music, and Dance in New York, 1840-1917&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.goodreads.com/book/show/6585095-all-hopped-up-and-ready-to-go"&gt;All Hopped Up and Ready to Go&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.goodreads.com/book/show/298222.The_House_That_Trane_Built"&gt;The House That Trane Built&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.routledge.com/The-New-York-Schools-of-Music-and-the-Visual-Arts/Johnson/p/book/9780415936941"&gt;The New York Schools of Music and the Visual Arts&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.goodreads.com/book/show/572111.A_Flexible_History_of_Fluxus_Facts_Fictions"&gt;A Flexible History of Fluxus Facts &amp;amp; Fictions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.goodreads.com/en/book/show/23461144-folk-city"&gt;Folk City: New York and the American Folk Music Revival&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.goodreads.com/book/show/329133.Love_Saves_the_Day"&gt;Love Saves the Day: A History of American Dance Music Culture, 1970-1979&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://pleasekillme.com/please-kill-uncensored-oral-history-punk-book/"&gt;Please Kill Me&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.goodreads.com/book/show/11331168-love-goes-to-buildings-on-fire"&gt;Love Goes To Buildings On Fire&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.goodreads.com/book/show/62005777-this-must-be-the-place"&gt;This Must Be the Place: Music, Community and Vanished Spaces in New York City&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.goodreads.com/book/show/6512666-hold-on-to-your-dreams"&gt;Hold On to Your Dreams: Arthur Russell and the Downtown Music Scene, 1973-92&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.goodreads.com/book/show/54754.Can_t_Stop_Won_t_Stop"&gt;Can’t Stop Won’t Stop: A History of the Hip-Hop Generation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.goodreads.com/book/show/25816741-meet-me-in-the-bathroom"&gt;Meet Me in the Bathroom: Rebirth and Rock and Roll in New York City 2001-2011&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.dukeupress.edu/the-williamsburg-avant-garde"&gt;The Williamsburg Avant-Garde: Experimental Music and Sound on the Brooklyn Waterfront 1985-2010&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;div class="accordion"&gt;
&lt;h3&gt;Comments&lt;/h3&gt;
&lt;div&gt;
&lt;b&gt;Larry&lt;/b&gt; commented on 27 October 2025:
&lt;blockquote&gt;IRL scenius lives on in small towns, not publicized. e.g Bellingham WA multi genre mix, Paso Robles CA wine scenius moved from Napa Valley, Santa Cruz - North Monterey County - organic farming.&lt;/blockquote&gt;
&lt;hr /&gt;&lt;br /&gt;
&lt;b&gt;Ray&lt;/b&gt; commented on 27 October 2025:
&lt;blockquote&gt;There does seem to be some dissipation of energy. I'd just bring up a few genres you may have missed, although you may think they are just minor iterations.&lt;br /&gt;&lt;br /&gt;
In NYC, I don't know how people are characterizing it, but there is a Soundcloud rap 2.0; rage rap thing that is happening in NYC. It kind of fully embraces like rich UES kids making music in a way. Artists I am referring to are people like 2hollis and Nettspend.&lt;br /&gt;&lt;br /&gt;
They are in conversation with the other Soundcloud 2.0 guys like Ken Carson, but really in their own genre.&lt;br /&gt;&lt;br /&gt;
In London, there is an electronic genre; forget what it's called but it tends to have extremely heavy sample usage that's in your face. The whole songs are sampled real life sounds; it's bass forward and drumlines with post EDM trap influences. The drumlines are some of the most complex, but that is true for all of electronic genres now. The pop version of this genre are Kai Whiston and Iglooghost, but the central figures are Sega Bodega and his record label: Shygirl and Coucou Chloe. The NYC close friend of there's is Eartheater. It is kind of a successor branch to future bass or edm.&lt;br /&gt;&lt;br /&gt;
UK Drill is a bigger deal than people think for UK music in general. The thing that was hard for grime prior to drill was it was 140bpm and the rapping was formalized techniques to rap at that speed; this comes from their rave scene and drum and bass bpms. Drill is easier to commercialize and alters rap techniques and highlights pauses, which generally are not acceptable in regular rap. There are some artists that are well known on TikTok and in America like central cee from this scene.&lt;br /&gt;&lt;br /&gt;
I would put A$AP Mob more in the cloud rap genre which also wasn't native to NYC, but NYC did popularize it. It's mostly under the influence of Clams Casino who tends to be characterized as Cloud Rap, but is also labelled Witch House sometimes. He might be the sole reason why Imogen Heap had a second life on TikTok. The Imogen Heap sample that is used on Lil B song has been one of the most reused samples in rap these days.&lt;/blockquote&gt;
&lt;hr /&gt;&lt;br /&gt;
&lt;b&gt;ABC&lt;/b&gt; commented on 27 October 2025:
&lt;blockquote&gt;
This is what we call an oldhead.&lt;br /&gt;&lt;br /&gt;
I give a simple theory: This man is a blogger. In the 70s he would have been a Village Voice columnist. In the 90s he would have written for NME or some similar music publications. In the 2000s he would have written for Pitchfork or run his own music blog. He was a class of taste makers with immense power to dictate trends, who utilized their liberal arts educations and proximity to New York scenes to choose and disseminate new music to the rest of America. But this role has been superseded by social media. You don’t need him anymore because TikTok can show you every musician in any micro scene in any city of the world just by getting you to scroll for 1 hour. It’s no coincidence he believes NYC music [dies] right in the 2010s.&lt;br /&gt;&lt;br /&gt;
There are many great scenes in NYC still but WASPy sociocultural elite blogger is not really the demographic target for any of them, no offense intended.&lt;br /&gt;&lt;br /&gt;
But if you look at the “classic” records no one was listening to them at the time anyway. It was only retrospectively that they became cohesive.&lt;br /&gt;&lt;br /&gt;
&lt;blockquote&gt;&lt;b&gt;Gavin: Wrong on all counts! I am not a WASP, don't have a fancy liberal arts degree, and have literally never been in a tastemaking-gatekeeping scene. Nor am I closed off to the possibility of current scenes: I listened to &lt;a href="https://www.gleech.org/music2024"&gt;900&lt;/a&gt; albums released last year. &lt;br /&gt;&lt;br /&gt;
You are welcome to submit to the Tiktok algorithm's gatekeeping instead of the WASP hipster kind, but it has not so far led to scenius and is imo unlikely to ever breed much music we will later describe as art.&lt;br /&gt;&lt;br /&gt;
Please name these great scenes!&lt;br /&gt;&lt;br /&gt;
Also, where I'm from, "oldhead" is a compliment. It's someone who can rise above the noise and manias of a particular moment by seeing it in context.&lt;/b&gt;&lt;/blockquote&gt;
&lt;/blockquote&gt;
&lt;hr /&gt;&lt;br /&gt;
&lt;b&gt;Will C&lt;/b&gt; commented on 28 October 2025:
&lt;blockquote&gt;Odd to identify Mary Halvorson and Oneohtrix Point Never as belonging to some single genre. &lt;b&gt;[EDIT GL: now fixed.]&lt;/b&gt; Halvorson's more typical of what Vijay Iyer dubbed the "new Brooklyn complexity," lot of related figures in that scene: Tomas Fujiwara, Ches Smith, Ingrid Laubrock, etc. Basically a jazz variant with convoluted but rarely atonal harmonies, weird time signatures, quirky compositional forms.&lt;br /&gt;&lt;br /&gt;
Oneohtrix Point Never is hard to pin down, he's shifted throughout his career, but among other things he's considered one of the seminal figures in the "vaporwave" scene. Hard to argue this is a NYC-based scene I suppose, it's mostly internet-based.&lt;br /&gt;&lt;br /&gt;
But scenes splitting up geographically is the real trend--flights are cheap, collaboration doesn't have to be face-to-face for many styles, and social media makes it easier for artists to hone in on a broad but dispersed audience. The idea of free jazz and post-minimalism being a single block makes me think of a different set of artists--Ellen Arkbro, Caterina Barbieri, Laurel Halo, Kali Malone, all of whom are recognizably part of one scene, collaborate in some instances, but are not based in any single country.&lt;br /&gt;&lt;br /&gt;
I do want to put some friction here--the idea that (say) the 50s-90s were a typical period of musical history that we should expect to be repeated doesn't really hold up, and I think it really confuses things. NYC isn't as innovative as it used to be, because there's less innovation to be had. And the reason for that is: innovation isn't some homogeneous, fungible commodity. Take a few of the big 8/90s innovations, hip-hop, drum n bass, industrial music. These are genres that are completely built around the invention of sampling, which emerge almost immediately after it becomes commercially available. You can't invent sampling twice.&lt;br /&gt;&lt;br /&gt;
Likewise various other innovations which basically pertain to a single simple, identifiable goal, namely the reproduction of sound: amplification, analog synthesis, digital synthesis, looping/sequencing, the subwoofer...all these led almost immediately to genres being formed or reimagined. There have not been comparable technologies in what, thirty years? And that's because the singular goal was, for the most part, achieved. There will not be a second comparable era.
&lt;/blockquote&gt;
&lt;hr /&gt;&lt;br /&gt;
&lt;b&gt;JP&lt;/b&gt; commented on 29 October 2025:
&lt;blockquote&gt;David Byrne of Talking Heads, in his book How Music Works, highlighted the gentrification of gritty NYC neighborhoods as leading to places where no creativity was going on. Many of his arguments match some of the points here. As with most things, there is no single answer, but Byrne’s main point is that the conditions must be right for creativity to flourish, and NYC has steadily been losing those conditions.&lt;/blockquote&gt;
&lt;/div&gt;
&lt;/div&gt;</description><pubDate>Thu, 23 Oct 2025 00:00:00 +0000</pubDate><link>https://www.gleech.org/nysound</link><guid isPermaLink="true">https://www.gleech.org/nysound</guid><category>places,</category><category>music,</category><category>hypothesis-dump</category></item><item><title>AI breakthroughs 2024-2025</title><description>&lt;p&gt;So what just happened?&lt;/p&gt;
&lt;h2 id="2024"&gt;2024&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;In January&lt;/em&gt;, the best systems were around chance (30%) on GPQA (hard science). By November, 4o was getting “around PhD” (60%) on the hardest Diamond set.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Feb&lt;/em&gt;: first “million-token” context window (Gemini 1.5), but there’s a steep fall in performance as you go through it.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Feb&lt;/em&gt;: first proper text+audio+image model (Gemini 1.5)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Feb&lt;/em&gt;: first good video generation (Sora)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;May&lt;/em&gt;: first good voice interface (Advanced Voice Mode)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Jul&lt;/em&gt;: silver medal at the IMO (AlphaProof) but with a Frankenstein hybrid system.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;🚨 Sep&lt;/em&gt;: RL works on LLMs at last. So-called “reasoning” (o1-preview)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Oct&lt;/em&gt;: GPT-4 “Search” agent. Works better than modern Google.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Oct&lt;/em&gt;: So-called “agency” (Anthropic Computer Use). Supposedly a 15 min human-task horizon but really a much shorter for usable horizon.
&lt;ul&gt;
&lt;li&gt;🚨 Whatever &lt;a href="https://epoch.ai/benchmarks/metr-time-horizons"&gt;METR’s time horizon&lt;/a&gt; is measuring admittedly multiplied by a factor of 7 over the year.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Nov&lt;/em&gt;: the Model Context Protocol standardises the interface between agents&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;🚨 Over the year, the best systems jumped from 5% to 50% on SWE-Bench Verified (real coding tasks).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Dec&lt;/em&gt;: the first LLM that handles streaming video input (Gemini Live)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Dec&lt;/em&gt;: Gemini Deep Research is the first strong agent (for lit-reviews) but no one notices.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;Sheesh.&lt;/p&gt;
&lt;p&gt;But after about 3 months we got used to the hard benchmark values being 70% or 80% instead of 10% or 20%. It meant less than we thought it would.&lt;/p&gt;
&lt;h2 id="2025"&gt;2025&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;¯\&lt;em&gt;(ツ)_/¯ _Jan&lt;/em&gt;: Deepseek R1 hype. But the apparent 10x per-token saving is a &lt;a href="https://www.gleech.org/paper#tokenomics-no-effective-discount"&gt;false economy&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;🚨 Jan, recursive improvement&lt;/em&gt;: Sometime around here, Deepmind uses AlphaEvolve (Gemini 2.0) to write GPU kernels and speed up the training of Gemini 2.5 by “1%”. (2024 model.)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Feb&lt;/em&gt;: Claude Code. The ~end of brittle stupid RAG. Lab incursion into the application layer.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Mar&lt;/em&gt;: &lt;a href="https://cdn.openai.com/11998be9-5319-4302-bfbf-1167e093f1fb/Native_Image_Generation_System_Card.pdf"&gt;Autoregressive image generation&lt;/a&gt; surpasses(?) diffusion models.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;🚨 Apr&lt;/em&gt;: o3 is the first LLM ever that is actually worth using for me&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;May&lt;/em&gt;: The first video generator which creates synched audio (Veo 3). Makes the psychological effect 100x stronger.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;¯\&lt;em&gt;(ツ)_/¯ In _June 2024&lt;/em&gt;, 4o got 5% on ARC-AGI. By Apr 2025 o4-mini got 41%. Also for the first time you can convert $3m &lt;em&gt;per-run&lt;/em&gt; into 80% on it. (Human is “98%”)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;🚨 Jul&lt;/em&gt;: IMO gold by an LLM. Reportedly no tools and no neuralese involved. (“Gemini 2.5 Deep Think Advanced” + unnamed experimental OpenAI model)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Sep&lt;/em&gt;: Various &lt;a href="https://www.nature.com/articles/s41586-025-09298-z"&gt;groups&lt;/a&gt; &lt;a href="https://arcinstitute.org/news/hie-king-first-synthetic-phage"&gt;racing headlong&lt;/a&gt; into “frontier biology” with protein language models and such.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Nov&lt;/em&gt;: Gemini 3 and Opus 4.5.
&lt;ul&gt;
&lt;li&gt;Gemini: Notable improvement in vision and image generation. &lt;a href="https://nitter.net/stuhlmueller/status/1991546706371178781#m"&gt;Rest&lt;/a&gt; &lt;a href="https://nitter.net/ArtificialAnlys/status/1990926803087892506#m"&gt;is a&lt;/a&gt; &lt;a href="https://www.lesswrong.com/posts/8uKQyjrAgCcWpfmcs/gemini-3-is-evaluation-paranoid-and-contaminated"&gt;very&lt;/a&gt; &lt;a href="https://nitter.net/peterwildeford/status/1990830603239842080#m"&gt;mixed&lt;/a&gt; &lt;a href="https://x.com/Miles_Brundage/status/1991664747914358861"&gt;bag&lt;/a&gt;. Benchmaxxed, or rather narrow-objective-maxxed.&lt;/li&gt;
&lt;li&gt;Opus: TBD.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;¯\&lt;em&gt;(ツ)_/¯ _Nov&lt;/em&gt;: claims about Gemini 3 adding a continual learning mode (based on the little &lt;a href="https://research.google/blog/introducing-nested-learning-a-new-ml-paradigm-for-continual-learning/"&gt;HOPE&lt;/a&gt; experiment). We’ll see.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Over the year, progress from 50% to 77% on SWE-Bench-Verified. But you can’t just use this (“+27% is less than the +45% last year”) to say slower latent progress, since obviously the tasks solved this year were much harder.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Over the year, whatever HCAST is measuring multiplied by 5x this year (vs 7x last year).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Browser agents aren’t adopted by anyone except Tyler really.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href="https://nitter.net/NikoMcCarty/status/1986501362730082464"&gt;Various noises&lt;/a&gt; about AI scientists. Mostly rediscovery and recombination?&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href="https://nitter.net/g_leech_/status/1974165458283860198"&gt;Real progress&lt;/a&gt; in AI assistance for research mathematics.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;There will be things which I miss / which only reveal themselves as significant next year.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;No earthquakes this year (the big scaling hopes, GPT-4.5 and Grok 4 both disappointed their masters) but still prettty fast.&lt;/p&gt;
&lt;p&gt;This list is biased towards discrete changes. While the story of AI does have step changes (like GPT-2) and may have more coming, you should do plenty of staring at smooth graphs as well. A lot of the 2024 breakthroughs above were more like progress in UIs or products than something fundamental (though getting multimodal to work was once the definition of fundamental).&lt;/p&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;p&gt;PS: That title should be “LLM breakthroughs”, sorry; &lt;a href="https://nitter.net/g_leech_/status/1864349307345731614#m"&gt;see&lt;/a&gt; &lt;a href="https://docs.google.com/spreadsheets/d/1A5r83AfPwF1_sljzbz12OuZBysjip5F6WiTnMI3IxBA/edit?usp=sharing"&gt;here&lt;/a&gt; for other AI.&lt;/p&gt;</description><pubDate>Mon, 20 Oct 2025 00:00:00 +0000</pubDate><link>https://www.gleech.org/ai-24-25</link><guid isPermaLink="true">https://www.gleech.org/ai-24-25</guid></item></channel></rss>