I’ve run Qwen 3.5 4b and Gemma 4 e2b on CPU only, this should be faster than those I think (fewer active parameters). If you have AVX512 or AVX10 then it should help a bit. Still slow compared to a GPU lol.
- 5 Posts
- 17 Comments
anyone try this? this might be good for my crappy laptop lol
is it good enough to use with Zoo Code? is it better than Qwen 3.5 4b?
EDIT: woa

https://artificialanalysis.ai/models/ling-3-0-tiny
But not yet supported in llama.cpp https://github.com/ggml-org/llama.cpp/pull/26608
BeefAndPoultry@lemmus.orgto
LocalLLaMA@sh.itjust.works•Thinking injection to modify model behaviour (making Gemma 4 less lazy)English
2·3 days agoActually funny he’s not asking it to work harder (that would be system prompt or user message), he’s forcing it to think that it will work harder
BeefAndPoultry@lemmus.orgto
LocalLLaMA@sh.itjust.works•Thinking injection to modify model behaviour (making Gemma 4 less lazy)English
3·3 days agoThat’s a really cool idea. It’s like inception for an LLM, you make it think it was the one that thought of this lol
BeefAndPoultry@lemmus.orgto
LocalLLaMA@sh.itjust.works•Thinking injection to modify model behaviour (making Gemma 4 less lazy)English
1·3 days agoHave you tried preserve thinking? https://lemmus.org/post/24365786
BeefAndPoultry@lemmus.orgOPto
LocalLLaMA@sh.itjust.works•llama.cpp in progress pull request for smart caching of MoE experts, 16% to 35% TPS boost for my RTX 2080English
1·7 days ago(Oops I got my Gemma and Qwen speeds mixed up, edited the post to fix it.)
But now with the new commits they added, with the same number of hot experts, Qwen is up to about 34. If I increase hot experts to 48 then I get around 37.
Gemma is still around 23 with just 10 hot experts. With 16 hot experts I get about 26 TPS. If I overprovision my VRAM (thanks to
GGML_CUDA_ENABLE_UNIFIED_MEMORY=1) then 24 hot experts can give me 29 TPS, and 32 hot experts 34 TPS.
BeefAndPoultry@lemmus.orgOPto
LocalLLaMA@sh.itjust.works•llama.cpp in progress pull request for smart caching of MoE experts, 16% to 35% TPS boost for my RTX 2080English
3·7 days agoIn a few minutes a significant performance improvement incoming
👀
this has been a crazy few weeks! lol
BeefAndPoultry@lemmus.orgOPto
LocalLLaMA@sh.itjust.works•Qwen 3.8 Max (2.4T-a95b) and 27B open weights being released next weekEnglish
3·8 days agotrue, it’s not perfectly clear
also I just saw this

BeefAndPoultry@lemmus.orgOPto
LocalLLaMA@sh.itjust.works•Qwen 3.8 Max (2.4T-a95b) and 27B open weights being released next weekEnglish
2·8 days agohave you tried Qwen 3.6 35b a3b? check my guide, it’s still relevant to you just with different numbers because you have 12GB
BeefAndPoultry@lemmus.orgOPto
LocalLLaMA@sh.itjust.works•Qwen 3.8 Max (2.4T-a95b) and 27B open weights being released next weekEnglish
2·9 days agoGemma is probably good for that, as long as it’s consistently succeeding at the tool calls.
BeefAndPoultry@lemmus.orgOPto
LocalLLaMA@sh.itjust.works•unsloth/DeepSeek-V4-Flash-0731-GGUF · Hugging FaceEnglish
2·10 days agomake sure that holds up with large context, you might need to step down to Q3 (which I’ve heard is still good for this model, many people are even using IQ2)
BeefAndPoultry@lemmus.orgto
Piracy: ꜱᴀɪʟ ᴛʜᴇ ʜɪɢʜ ꜱᴇᴀꜱ@lemmy.dbzer0.com•High constant bitrate content, like 128mbps for video and ~10mbps for audio is possible for free?English
5·12 days agoYou’re looking for “2160p remux” torrents
BeefAndPoultry@lemmus.orgOPto
LocalLLaMA@sh.itjust.works•unsloth/DeepSeek-V4-Flash-0731-GGUF · Hugging FaceEnglish
2·12 days agothis appears to be the untouched model in GGUF format, 180 GB, MXFP4
https://huggingface.co/bartowski/DeepSeek-V4-Flash-0731-GGUF
BeefAndPoultry@lemmus.orgOPto
LocalLLaMA@sh.itjust.works•unsloth/DeepSeek-V4-Flash-0731-GGUF · Hugging FaceEnglish
2·12 days agothe original model has a lot of parts that were natively trained in 4 bit, so those layers can’t go higher
BeefAndPoultry@lemmus.orgOPto
LocalLLaMA@sh.itjust.works•How to run Qwen 35b-a3b on 4GB to 8GB of VRAM (with 24+GB system RAM)English
2·13 days ago--n-cpu-moe 36 --spec-type draft-mtp --spec-draft-n-max 3does seem to speed up token generation for meCan’t use llama-bench for MTP. In a basic tests it seems to improve from about 26 to 30 tokens per second output. But it seems to hurt my input speed from about 1300 pp down to 800.

Old music doesn’t get re-recorded for digital releases, they pull it from the original master again, which is better and more authentic than vinyl.
I mean if you went to a concert would you want them to add hiss and pop to the performance so it sounds more like vinyl? Lol