• 1 Post
  • 91 Comments
Joined 6 months ago
cake
Cake day: February 9th, 2026

help-circle

  • This has always been a huge red herring. Llms are built on top of the transformer architecture which does text autocomplete (and yes we can combine text embeddings with other inputs like images). They have some interesting properties where they seem to be able to do text autocomplete in a bunch of different scenarios that they weren’t explicitly trained for, but they were never designed for precise dna analysis. It is their architecture that prevents them from other long horizon tasks like playing chess and the way that they represent text is why they can never count the letters in strawberry (most have this specific question hard-coded in their training data now).

    Anyone who believes that LLMs are going to solve cancer either has no idea how they work or has been one-shotted from talking to Claudia







  • I studied transformer architecture models and have played around with them (unfortunately) enough to understand how they work. Under the surface the model produces what look like XML tags <thinking> </thinking> to designate which tokens are thinking tokens and which are ā€œnormalā€ output. That is literally the only hard difference between the two output modes. The reinforcement learning might tune the thinking to be more like ā€œwhat a human would expect to see in a thinking blockā€ but it’s still the same RNG madlib process generating everything underneath and any attempt to ascribe intelligence to this process should be met with lethal force incredulous cynicism. Just like any claim that ā€œwe don’t know how they workā€ - actually yes we know exactly how they work. What we can’t comprehend is the exact numbers and weights inside the massive pile of probabilistic algebra being processed to generate your slop. If I flip 5 coins in a row and the observer’s belief is anything other than ā€œyou just got very luckyā€ most people would call them crazy rather than join the cult and worship the coin god…