Permanently Deleted

          • SmokeyDope@lemmy.worldM
            link
            fedilink
            English
            arrow-up
            2
            ·
            1 year ago

            You’re welcome. Also, whats your gpu and are you using cublas (nvidia) or vulcan(universal amd+nvidia) or something else for gpu postprocessing?

              • SmokeyDope@lemmy.worldM
                link
                fedilink
                English
                arrow-up
                2
                ·
                1 year ago

                If you were running amd GPU theres some versions of llama.cpp engine you can compile with rocm compat. If your ever tempted to run a huge model with partial offloaded CPU/ram inferencing you can set the program to run with highest program niceness priority which believe it or not pushes up the token speed slightly

    • ffhein@lemmy.world
      link
      fedilink
      English
      arrow-up
      4
      ·
      1 year ago

      Exllamav3 is still in development so it’s not fully optimized and could have bugs, but I get 16k context with 4bpw (which has very similar perplexity as Q4_K_M, according to developer’s own measurements) using only 22GB VRAM, since I also run my desktop env on the same computer.

  • gencha@lemm.ee
    link
    fedilink
    English
    arrow-up
    2
    arrow-down
    8
    ·
    1 year ago

    People implement this on their calculator during class. This is the kind of thing you would write to learn programming, the definition of entry-level. You’re using a device that can execute billions of trigonometric calculations per millisecond to produce code that calculates X and Y coordinates for few dozens of points on a radial trajectory.

    What the fuck…