Anyone tested this yet?

Enable with --moe-cache-mib N. In experiments it was found 10% of the total expert size is good.

Baseline is master with the default --fit. With --moe-cache-mib, --fit counts the cache and keeps fewer whole expert layers on the GPU.

I’m still using

commit 38976372452722cf5e6846429aa33ee892c9d89a
Author: Miltos22

To run that commit, do this:

git remote add https://github.com/miltos22/llama.cpp-wackMall-merge-request.git miltos22
git fetch miltos22
git fetch miltos22 38976372452722cf5e6846429aa33ee892c9d89a:refs/tags/miltos22-expert-caching-1
git checkout miltos22-expert-caching-1

When I have to use a low expert-hot-s value because I don’t have enough VRAM for the model, I use like:

expert-hot-s = 13
expert-hyst = 1.35
expert-heat-decay = 0.9988
expert-dwell = 2

Higher hysteresis, lower decay, and a little bit of expert-dwell reduces the PCI-e back-and-forth bandwidth usage.

There were a bunch of competing PRs for expert caching.

  1. https://github.com/ggml-org/llama.cpp/pull/17044
  2. https://github.com/ggml-org/llama.cpp/pull/21609
  3. https://github.com/ggml-org/llama.cpp/pull/21614
  4. https://github.com/ggml-org/llama.cpp/pull/23170
  5. https://github.com/ggml-org/llama.cpp/pull/24524
  6. https://github.com/ggml-org/llama.cpp/discussions/24528
  7. https://github.com/ggml-org/llama.cpp/pull/25294
  8. https://github.com/ggml-org/llama.cpp/discussions/25469
  9. https://github.com/ggml-org/llama.cpp/pull/25932
  10. https://github.com/ggml-org/llama.cpp/pull/26414
  11. https://github.com/ggml-org/llama.cpp/pull/26563
  12. https://github.com/ggml-org/llama.cpp/pull/26824
  13. https://github.com/ggml-org/llama.cpp/pull/27861
  14. https://github.com/ggml-org/llama.cpp/discussions/28248