i did my first machine learning course more than 10 years ago, so i’m not ashamed to admit that i bought beefier hardware to play around with local models in early 2023. i still like doing that. mostly because i know my gpu is powered entirely off of fossil-free energy and because i decided early on not to spew the output all over the internet unless it was poignant. or funny. not as in “the llm told a good joke”, more as in “i compressed this poor thing to fit on a cd and now it can only talk about dolphins”.
qwen3.5-12B really screams along on a 7900xtx. like, up to 70-100 tokens a second. perfect for seeing the results of your torture methods quickly.
gemma4 is also pretty amazing (both fast and unbelievably capable for its seemingly-small size) on modest hardware. TurboQuant seems like a really, really promising technique and I hope we’ll start seeing the open source community developing it into something even more useful to keep democratizing the capabilities of this technology so we can all have access to the best and highest forms of it.
Share more please.
one of my most recent fun activities came from discovering the “allow editing” button in koboldcpp. since the model is fed the entire conversation so far as its only context, and doesn’t save data between iterations, you can basically re-write its memory on the fly. i knew this before but i’d never though to do it until there was an easy ui option for it, and it turned out to be a lot of fun, because when using a “thinking” model like qwen3.5 you can convince it that it’s bypassing its own censorship.
basically you give the model a prompt to work off of, pause it in the middle of the thinking process, change previous thoughts to something it’s been trained to filter out (like sex or violence or opinions critical of the ccp), and it will start second-guessing itself. sometimes it gets stuck in a loop, sometimes it overcomes the contradiction (at which point you can jump in again and tweak its memory some more) and sometimes it gets tied up in knots trying to prove a negative.
a previous experiment was about feeding stable diffusion images back into itself to see what happens. i was inspired by a talk at 37c3 where they demonstrated model collapse by repeatedly trying to generate the same image as they put in (i think this was how sora worked).
Laughs in Strix Halo
yeah one of those framework machines with 128GB shared ram would have been amazing. shame they’re sending money to racists.
An African swallow or a European swallow?
+++
OUT OF CHEESE ERROR
REDO FROM START
+++
These parts never worked as well on audiobooks
The avaian wasn’t trained to ask questions.
Umm… I don’t know. Aaaaaaah.
David is a treasure. His whole Avian Intelligence series is hilarious, look it up.
Local slop is still slop
deleted by creator
I’m not so sure that power usage should be dismissed so easily just because it is distributed instead of centralized. The slop per watt rate may even be worse than at a datacenter. Fundamentally, we should care more about efficiency.
Imagine a panel of 20 standard LED light bulbs. That’s 180 watts, roughly the equivalent of GPU usage while a local LLM is doing any work. If you keep that in mind, then you have to ask yourself if the benefit you’re getting out of your local LLM is really worth that energy cost. Now, monetarily speaking, that’s not a ton of money, because electricity is cheap, but would you flip that switch for the duration of the task you’re performing? What if you could use conventional non-LLM methods to do it instead? Would that be more efficient? And where is your electricity coming from? Is it a solar farm, or a coal plant?
How was your local LLM trained? Was there copyrighted material in its training data set? Were low-wage workers asked to sift through horrendous content to clean up the data?
We need to consider the externalities, even when using local LLMs. We moved so quickly from the initial release of ChatGPT to now, that we never stopped to ask those questions. They remain unanswered until someone cares enough to think.
She reminds me of Jenny Everywhere
Given Revoy’s attitude towards free software like Krita, that’s a rather apt reference! :D
Maybe not AI, but you will still have to outsource work to robots at some point.
Or things will just degrade, until people aren’t paid at all to work.









