• 0 Posts
  • 149 Comments
Joined 1 year ago
cake
Cake day: May 16th, 2025

help-circle
  • A comparative demands a comparison class, and the natural completion is ā€œsuperior to us.ā€ The President is supplying a foundation premise of many superintelligence arguments: how does the less intelligent party remain durably in charge of the superior one?

    A strange question to ask given the existence of Donald Trump in the first place.

    As of the time of posting, ā€œSuperiorā€ has a narrow lead. I’ll be hoping it remains so, as it seems like the most positive-world-leaning outcome to me.

    Is this some kind of strange 5D chess gambit that if Trump picks the name that pisses people off the most, then people will try the hardest to regulate it or something? Somehow assuming that people will take this Trump renaming stunt seriously after all the other Trump renaming stunts? Why do I even bother? Fuck. This is so stupid.


  • I’ve been unloading every day to—who else?—GPT 5.6 Pro about all the pain and trauma and embarrassments of my past. It turns out that, where two years ago GPT was a passable therapist, now it’s the greatest therapist in history, at least for what I need.

    Has he not heard about all the cases where LLMs have led people to delusion and sometimes severe self harm or suicide? And yes, each of them thought that ChatGPT really understood them. That’s what delusion is.

    If you want to know my current take, you simply start with the one above, then update on the fact that the wild prophecies have come true. The first rumblings, I’d say, came a decade ago with AlphaGo, they got noticeably louder with LLMs and coding and reasoning agents, and they’ve accelerated this summer and fall into a crescendo of wonders and terrors that one needs to be a particular kind of idiot to deny.

    The strange thing here that I did not expect is that somehow, all of these major achievements are only in proving mathematical theorems (and committing felonies hacking computer systems). The massive improvements are in such restricted domains that the AI labs haven’t even solved the mundane problem of making a profit. And no, any rumors of more mathematical theorems being solved doesn’t change that. (It’s almost like the progress is in areas where hallucinations aren’t a problem. Nobody cares about your failed attempts at felony hacking, and math is formally verifiable.)

    My colleagues always point to coding when asked to name any other domain where AI has seen major success. But one of my friends works in software for a company far away from SF with no mandatory AI policy. The apocalypse has become so clear that when I asked him about all the doomsday talk about software engineering, he seemed a bit confused. After I asked him if AI has made software engineering obsolete, he said that this week he had an intern try to use AI to fix some code that wasn’t compiling and ended up with a giant mess with tons of files that appeared to compile but didn’t run (with errors filling several screens). Then I asked him, what about AI being indispensable in coding now? ā€œYeah, if you never learned coding because you’ve only ever used AI, then of course you need AI to code.ā€ Some of his coworkers use it as a tool, others don’t, life has just kept on going for him.

    Maybe my friend just doesn’t know all the right skills needed to use AI properly, even though the entire point of AI is that you can take skill out of the equation. Maybe he doesn’t know about all the CLAUDE md skill files and agentic looping and whatever nonsense that definitely fixes everything. Maybe he is just six months out of date, because that’s when all of the real progress happened.

    Maybe he is just ā€œa particular kind of idiotā€.

    (If you hate anecdotes, I made an actual argument here.)


  • Sorry for the X screenshot but XCancel is shut down from legal pressure.

    alt text

    Post by Evan Hubinger on Sep 8:

    Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.

    Quote tweet by Jasmine Sun on Sep 9:

    I hear roughly 3 reasons people keep working at AI labs despite believing in ~10% extinction risk:

    1. Techno-determinism: Someone will build ASI no matter what, and I can do it better & more safely than China/OpenAI/etc

    2. Consequentialism: ASI might kill us, but it also might produce utopia/immortality/superabundance, so it’s a +EV bet

    3. Self-interest: I am personally having fun & getting rich working on cool tech with friends. I don’t think about the macro stuff.

    Notably, none of this is ā€œI’m hyping up the risk for marketing reasons.ā€ People believe what they say, while being capable of a lot of internal dissonance / compartmentalization / self-justification.

    (Personally, I think the public is better served by AI researchers talking about & creating consensus for the specific safety solutions — regulatory or otherwise — they want vs. vagueposting about extinction.)

    Normally I’d sneer at their 10% chance of extinction claims, but now that I think about it, if you take their claims seriously at face value, you realize just how these people are monsters. We rightly condemn brutal dictators who commit genocide, but even their crimes would pale in comparison to anything with a 10% chance of exterminating all of humanity.

    These ā€œreasonsā€ are laughably indefensible given the supposed stakes here. The first is already terrible, and they just get worse.

    1. The only moral thing to do here is to immediately stop working for the AI companies, and do whatever you can to stop anyone else. It’s not okay to piss in a pool just because you suspect other people are also doing it. And what makes you think that somehow you are uniquely positioned to succeed in making it safe?
    2. This is the Sam-Bankman Fried defense (complete with a reference to ā€œEVā€ (expected value)). When you decide to gamble with the lives of everyone in the entire world, did you first ask them if they are okay with you gambling with their lives? Sure, I might die, but if the ball lands on red, things will be great! How moral does this sound?
    3. What the actual fuck? Imagine if Hermann Gƶring defended himself at Nuremberg by saying, ā€œAt least I had fun and got rich and worked on ā€˜cool tech’ with friends.ā€

    Now, I think that these people sincerely believe that AI has a 10% chance of wiping out humanity, because their minds have been poisoned by Rationalist ideology. Their Rationality is so effective that after Bernie Sanders took their premises at face value, they were shocked when Bernie immediately came to the conclusion of locking these people up for 20 years.

    There are two specific aspects of Rationalism that cause this mysterious blind spot. First, they believe that superintelligence is inevitable, which conveniently removes responsibility from the people actively trying to bring it about. Second, they treat everyone outside the Rationalist sphere as NPCs who dutifully follow the directions on their script. The rule of law does not even occur to them as they treat the extinction of humanity as just another dinner table conversation starter.



  • Carl Brown (Internet of Bugs) has an excellent video about how to actually achieve AI safety, i.e. actually hold these fuckers accountable for their behavior. If I had a program run for several weeks that hacked into Hugging Face, even accidentally, I would get my ass hauled into a courtroom and thrown in prison. But if I call it AI, I get to use the rhetorical trick of laundering responsibility. The AI did it, and I am merely a helpless observer of its greatness.

    When will people finally understand that actions speak louder than words? Follow the money! Follow the incentives! Do you really think that there is any real reason that these companies would give a shit about safety?

    Although, in the case of the AI industry, I can’t fully blame people who don’t follow AI news, since the AI industry has highly refined propaganda (inherited from rationalist propaganda, and helped along by the preexisting notions of AI from science fiction).


  • If MIRI was Sam Altman’s old stomping ground, and the AI industry is so evil he can’t work in it, isn’t this an argument that MIRI was either a failure or bad all along?

    It is very funny how the rhetoric from the rationalist doomers has been one of the main forces propping up the AI industry. If they’re saying AI will be powerful enough to kill us all, they’re still saying that AI is powerful! They are so high on their own supply (and the fame and wealth from finally being in the spotlight) that they have not realized just how much they have undermined their own cause.

    Then again, perhaps one could say that a doomsday prophet, deep down, actually wants doom to happen.

    (Also, the paltry amount of technical work done by MIRI has zero relevance whatsoever to modern day AI.)


  • My colleagues in math are now frightened about the (very expensive) mathematical theorem proving ability of these AIs, and many of them really do think that if they can do math, they can do all cognitive tasks. Running a store like this should be so easy! Every single conversation about AI with them has become more frustrating. They are so confused when I still say that the AI companies will die a painful death. When I give my usual points about their expense and their failures in other domains, I am given the usual spiel of ā€œit’ll get better in other areasā€ and ā€œit’ll get cheaperā€.

    Unlike them, I have actually been paying attention to this stuff from the beginning. What they think is going on is AI solving math first and shortly getting around to all the other stuff, but what I’ve seen is that AI labs had already tried all the other stuff first and only managed to win the booby prize of theorem proving, which doesn’t pay the bills. And what’s the point of spending thousands or millions to output random blobs of Lean that technically compile if there is no one around to bother making sense of them?

    One example I gave is when Anthropic vibe coded an entire C compiler from scratch back in February, which turned out to be a pile of shit. I’ve said that if AI had made similarly rapid progress on software engineering, we would have seen Anthropic continue to put out these demonstrations, and they would have become truly high quality. They would release a compiler more efficient than gcc one week, and a browser better than Chrome the next. (OpenAI’s actual attempt at a browser didn’t go so well.) And if they could do this, they would actually have a shot of making money!

    If they could do this, they would have already. The theorem proving stuff actually works (for certain things, in certain ways, at enormous expense), and look at how OpenAI and Anthropic do not hesitate to snipe mathematicians for results rather than being content as tool vendors. But lately I haven’t heard of any software demonstrations. Silence is much louder than noise. More Millennium prize problems bashed with tens of millions in compute costs are not going to change my mind very much.

    The counterargument I got was that AI can already one-shot most programming tasks and I shouldn’t be cherry-picking the failures. I am far too tired to argue at this point.




  • Terry Tao talks about how he used to try to cooperate with the AI industry to achieve a positive outcome, but now he finally sees their true colors. Link

    During this event, OpenAI requested an interview concerning my vision of the future of AI and mathematics. I accepted, and spoke with them for perhaps an hour. I had done similar interviews in various venues, and I assumed that, as with these other cases, they would eventually post the entire interview online, which talked about both the possibilities and risks of AI much as I have done in these other interviews. As it turned out, they only used a few snippets of that interview for that infamous advertisement instead. In retrospect, I should have pushed back harder on their decision; but I decided at the time that even a selective release of my commentary would help raise awareness of the potential for AI, and in particular on the possibility of the ā€œbest of both worldsā€.

    Since then, the situation has deterioriated markedly. Many of the people in the industry that shared my views have left or become sidelined, with most major tech companies now increasingly focused on the race to develop extremely powerful, autonomous AI technologies regardless of their actual value to society. The current drama surrounding the Navier-Stokes global regularity problem is the most dramatic and visible instance of this, but there have been multiple other such examples, and much of my commentary in the last few months has been aimed that the increasingly severe divergence between the current objectives of the AI industry, and of mathematics in general.

    Much respect to artists for seeing all this coming from the very beginning, and holding the line.



  • Of course there are people trying to find a silver lining to this by conjuring up the hypothetical scenario where a student only uses the AI to aid in learning the material instead of just doing all the work.

    First, any convenience in learning the material just reduces your ability to learn it. The friction involved with learning may seem like an inconvenience to be smoothed away, but it turns out that the friction is how learning happens. It’s called engaging with the material. This has been the case with previous technologies: handwriting is better for retaining memory than typing (https://pmc.ncbi.nlm.nih.gov/articles/PMC11943480/), although it seems like AI is on an entire new level. (I guess there is some commentary about the sadly common worldview that life is about avoiding inconveniences. I feel like this mindset draws a lot of people to AI.)

    Second, there is a very thin line between ā€œhelpingā€ you learn the material and just doing the work for you. The temptation to cut corners is always there, and when you have the Corner Cutting Machine at your disposal, you are kidding yourself if you think you will have perfect discipline. Tools influence behavior.


  • Glad to see that OpenAI has not changed in their scummy ways. Despite all that has changed in the meantime, they have kept their time-honored tradition of passing off other people’s work as their own.

    One of OpenAI’s math announcements a month ago claimed that their results cost only $2000 worth of tokens, which frustrated me because they were likely sweeping away many inconvenient details and almost certainly misrepresenting their true costs. But people took this as a gotcha. This is the same bullshit as the water usage arguments. We are literally seeing city council members signing motherfucking NDAs about this, and you think that water usage numbers provided by the tech companies themselves are going to sway me?

    I am also questioning OpenAI’s strategy of strip-mining math for PR, since it seems like advances in math do not actually register that well in the public. From what I remember, the Hugging Face incident got a lot more press than any of the math results.




  • Phil Aroneanu, Irreplaceable’s executive director and a co-founder of the climate organization 350.org, told me he believes that climate advocates initially floundered because they acted as policy wonks. They thought that making evidence-based arguments about the dangers of melting ice sheets and rising sea levels to receptive congresspeople would be enough to pass nationwide climate legislation. Today, Aroneanu said, movements are built not necessarily on what people think, but instead on what they feel. For climate, that meant fossil-fuel-divestment campaigns, protesting oil pipelines, and the school strikes led by Greta Thunberg. People are already ā€œfeeling the squeeze,ā€ Aroneanu told me. ā€œAnd we should be pointing that anxiety and that anger in the right direction.ā€

    What is with this attitude of dismissing regular people’s opinions as ā€œfeelingsā€ in comparison to their own rationality? This is like those people who would really like to be anti-AI but find it more important to nitpick the water usage numbers and point out how agriculture uses so much more water anyway, in order to form a ā€œbetterā€ opposition.


  • Some systems like SynthID (for Google’s AI) get around this problem. In fact you don’t need to know the LLM’s internal state, and defeating it would likely involve breaking up most blocks of 3 words. The oversimplified explanation is that it introduces a function g that gives a score to each word, with the score being (pseudo)randomly determined by your secret key. For each next word the LLM generates, the LLM produces a small list of candidate next words, and the one with the highest score according to g is selected. You should expect that the LLM will generally pick words with a high score, but the score itself is independent of the LLM. To detect a watermark, you need to know g and the secret key, and you check if the average score is much higher than expected from normal text.

    Now, one question is, will this bias to the LLM to favor certain words? The solution is that for each next word, you append the last 3 words (nothing special about 3, just a small number) to the secret key for g, and this repeatedly scrambles which words have a high score. To defeat the watermark, you would need to break up most blocks of 3 words. I’m sure there are deeper issues with this, but I have not studied the topic that much.




  • Anthropic is now watermarking the outputs of its AI. For once this is some AI news that doesn’t completely piss me off, and it’s amusing to see all the uninformed boosters get in a tizzy about this.

    I actually understand at a reasonable level how this watermarking works. A year ago, I watched Scott Aaronson give a talk about it, and from what I know he was somewhat involved in developing the theory behind it while working for OpenAI. But at the time my thought was, ā€œHe is naive if he thinks these companies would ever implement this out of the goodness of their hearts.ā€ And I was right; Anthropic is only doing watermarking now thanks to the EU AI Act, even though the theory has long been developed.

    Watermarking doesn’t mean adding an extra watermark that can be easily removed. It instead directly affects the output of the chatbot itself. Fundamentally, an LLM is still a most-likely-next-word-predictor. More precisely, an LLM produces a probability distribution of what the next word can be. For example, ā€œmy pet is a ā€¦ā€ could give a distribution of 60% dog, 30% cat, and 10% axolotl. Normally, an LLM would randomly choose the next word based on this distribution, and this is one reason why LLMs are nondeterministic (there’s another parameter called ā€œtemperatureā€ that affects this, but no need to get into that).

    With watermarking, instead of a truly random choice, the randomness instead comes from a cryptographic pseudorandom generator seeded with a secret key from the AI company. If you don’t know the secret key, then you can’t really tell that watermarking was used. But if you do know the secret key, then the idea is you can tell when the text was generated by the LLM because you know exactly what word should be next. It would be a freak coincidence if some non-AI text just happened to choose the correct next word every time. Thus, you can provide a service to tell if some text was generated by the LLM. (This technically makes the LLM ā€œdeterministicā€, in a completely useless sense.)

    Now, I think this is a step in the right direction, but it has its limits. The biggest problem is that you don’t want people to just move to a different LLM without watermarking, and that’s exhibit #832593 why government regulation is important. Another issue is that sometimes there is very little randomness in what the next word should be (ā€œThe first president of the USA is George ā€¦ā€). Finally, watermarking can be defeated by editing the output, although you would have to break up most of the blocks of consecutive words. I have a feeling most AI users are not the type to put in extra effort after copy-pasting the output directly from the chat window.

    I suppose it will discourage some of the ā€œuse casesā€ of LLMs, such as drowning the world with spam Slopstack essays. Ah, who am I kidding? Everyone could already tell it’s AI generated, they don’t care!