Anyone tested this yet?
Enable with --moe-cache-mib N. In experiments it was found 10% of the total expert size is good.
Baseline is master with the default
--fit. With--moe-cache-mib,--fitcounts the cache and keeps fewer whole expert layers on the GPU.
I’m still using
commit 38976372452722cf5e6846429aa33ee892c9d89a
Author: Miltos22
To run that commit, do this:
git remote add https://github.com/miltos22/llama.cpp-wackMall-merge-request.git miltos22
git fetch miltos22
git fetch miltos22 38976372452722cf5e6846429aa33ee892c9d89a:refs/tags/miltos22-expert-caching-1
git checkout miltos22-expert-caching-1
When I have to use a low expert-hot-s value because I don’t have enough VRAM for the model, I use like:
expert-hot-s = 13
expert-hyst = 1.35
expert-heat-decay = 0.9988
expert-dwell = 2
Higher hysteresis, lower decay, and a little bit of expert-dwell reduces the PCI-e back-and-forth bandwidth usage.
There were a bunch of competing PRs for expert caching.
- https://github.com/ggml-org/llama.cpp/pull/17044
- https://github.com/ggml-org/llama.cpp/pull/21609
- https://github.com/ggml-org/llama.cpp/pull/21614
- https://github.com/ggml-org/llama.cpp/pull/23170
- https://github.com/ggml-org/llama.cpp/pull/24524
- https://github.com/ggml-org/llama.cpp/discussions/24528
- https://github.com/ggml-org/llama.cpp/pull/25294
- https://github.com/ggml-org/llama.cpp/discussions/25469
- https://github.com/ggml-org/llama.cpp/pull/25932
- https://github.com/ggml-org/llama.cpp/pull/26414
- https://github.com/ggml-org/llama.cpp/pull/26563
- https://github.com/ggml-org/llama.cpp/pull/26824
- https://github.com/ggml-org/llama.cpp/pull/27861
- https://github.com/ggml-org/llama.cpp/discussions/28248


Would try it.