Anyone tested this yet?

Enable with --moe-cache-mib N. In experiments it was found 10% of the total expert size is good.

Baseline is master with the default --fit. With --moe-cache-mib, --fit counts the cache and keeps fewer whole expert layers on the GPU.

I’m still using

commit 38976372452722cf5e6846429aa33ee892c9d89a
Author: Miltos22

To run that commit, do this:

git remote add https://github.com/miltos22/llama.cpp-wackMall-merge-request.git miltos22
git fetch miltos22
git fetch miltos22 38976372452722cf5e6846429aa33ee892c9d89a:refs/tags/miltos22-expert-caching-1
git checkout miltos22-expert-caching-1

When I have to use a low expert-hot-s value because I don’t have enough VRAM for the model, I use like:

expert-hot-s = 13
expert-hyst = 1.35
expert-heat-decay = 0.9988
expert-dwell = 2

Higher hysteresis, lower decay, and a little bit of expert-dwell reduces the PCI-e back-and-forth bandwidth usage.

There were a bunch of competing PRs for expert caching.

  1. https://github.com/ggml-org/llama.cpp/pull/17044
  2. https://github.com/ggml-org/llama.cpp/pull/21609
  3. https://github.com/ggml-org/llama.cpp/pull/21614
  4. https://github.com/ggml-org/llama.cpp/pull/23170
  5. https://github.com/ggml-org/llama.cpp/pull/24524
  6. https://github.com/ggml-org/llama.cpp/discussions/24528
  7. https://github.com/ggml-org/llama.cpp/pull/25294
  8. https://github.com/ggml-org/llama.cpp/discussions/25469
  9. https://github.com/ggml-org/llama.cpp/pull/25932
  10. https://github.com/ggml-org/llama.cpp/pull/26414
  11. https://github.com/ggml-org/llama.cpp/pull/26563
  12. https://github.com/ggml-org/llama.cpp/pull/26824
  13. https://github.com/ggml-org/llama.cpp/pull/27861
  14. https://github.com/ggml-org/llama.cpp/discussions/28248
  • hummingbird@lemmy.world
    link
    fedilink
    English
    arrow-up
    5
    ·
    1 day ago

    Yup worth a try. I got from 13ts to 21ts for Qwen 3.8 Flash Next on a RX 9070 XT 16GB. As always, the additional memory demand requires some compromise. For me reducing context from full 262k to 200k and reducing kv cache down to q8_0 was the best option. moe-cache-mib at 6544. Cache hit rate at ~70%.

  • BeefAndPoultry@lemmus.orgOP
    link
    fedilink
    English
    arrow-up
    1
    ·
    edit-2
    1 day ago

    I’m pretty disappointed this one was merged, it doesn’t give good performance.

    I haven’t yet understood completely the details, so can’t provide much info or suggestions for improvement yet.

    Should be good to merge and iterate on it from master.

    WHAT? Why did they choose to merge this one out of many if they don’t even understand it?

    I was getting way better performance with the old PR from miltos22 https://github.com/ggml-org/llama.cpp/pull/26563

    This is quite a comment someone left in one of the PRs lol

    @am17an if 1000 line PR is too much of a scope for a project, the project is fubar. is llama project so broken it can’t handle 1000 line PRs? state so publicly, so that the community can stop wasting time contributing to a project like that and move on.

    The one they accepted was nearly 1000 lines anyways.