This guy has the Qwen 3.8 177b MoE running on 12gb vram with 64gb of ram, getting 24.4 tokens/sec. A bit more resource hungry than the Qwen3.6 35b, but not by much and still well within the realm of home usage for what is essentially a Frontier model.

  • hotspur [he/him]@hexbear.net
    link
    fedilink
    English
    arrow-up
    7
    ·
    1 day ago

    This is an impressive implementation—the key system ram thresholds he identifies are 20ish for good answering, and 40ish for prompt reading. So having 40+ ram might be a stretch for most peoples home systems, but the performance seems pretty good overall regardless.

    Gonna have to try and set this up, but expect downloading it to take forever… plus he’s got some weird custom cache stuff going on.

      • hotspur [he/him]@hexbear.net
        link
        fedilink
        English
        arrow-up
        5
        ·
        1 day ago

        Ah wicked yeah sounded like he would have some instructions. Looking forward to trying it—I’ve tried a bunch of different local models with webui and lately Hermes agent, but have generally beefed up sorta disappointed—either the model is dumb and runs fast, or is somewhat smart but runs so slowly that it’s impossible to work with meaningfully. 24+ but capable sounds like a pretty sweet spot to actually do something with it.

        • JoeByeThen [he/him, they/them]@hexbear.netOP
          link
          fedilink
          English
          arrow-up
          3
          ·
          1 day ago

          Word. it may be worth it for you to check out his Qwen 3.6 35b on a 1060 vid as well to get something almost as capable but probably much more faster on your rig. I was getting 23 tokens/sec on that one with my 1080 and 24gb of ram. And it’s a pretty capable model as well. It’d probably fly on what you’re running.

          • hotspur [he/him]@hexbear.net
            link
            fedilink
            English
            arrow-up
            3
            ·
            1 day ago

            Yeah true—I liked his breakdown of slower thinker, then faster do-er model setup. I have 3.6 35b on LM studio, but need to look at his optimizations because even with 4090/64gb ram it was pretty slow. Will be a good project ;)