Anyone tested this yet?
Enable with --moe-cache-mib N. In experiments it was found 10% of the total expert size is good.
Baseline is master with the default
--fit. With--moe-cache-mib,--fitcounts the cache and keeps fewer whole expert layers on the GPU.
I’m still using
commit 38976372452722cf5e6846429aa33ee892c9d89a
Author: Miltos22
To run that commit, do this:
git remote add https://github.com/miltos22/llama.cpp-wackMall-merge-request.git miltos22
git fetch miltos22
git fetch miltos22 38976372452722cf5e6846429aa33ee892c9d89a:refs/tags/miltos22-expert-caching-1
git checkout miltos22-expert-caching-1
When I have to use a low expert-hot-s value because I don’t have enough VRAM for the model, I use like:
expert-hot-s = 13
expert-hyst = 1.35
expert-heat-decay = 0.9988
expert-dwell = 2
Higher hysteresis, lower decay, and a little bit of expert-dwell reduces the PCI-e back-and-forth bandwidth usage.
There were a bunch of competing PRs for expert caching.
- https://github.com/ggml-org/llama.cpp/pull/17044
- https://github.com/ggml-org/llama.cpp/pull/21609
- https://github.com/ggml-org/llama.cpp/pull/21614
- https://github.com/ggml-org/llama.cpp/pull/23170
- https://github.com/ggml-org/llama.cpp/pull/24524
- https://github.com/ggml-org/llama.cpp/discussions/24528
- https://github.com/ggml-org/llama.cpp/pull/25294
- https://github.com/ggml-org/llama.cpp/discussions/25469
- https://github.com/ggml-org/llama.cpp/pull/25932
- https://github.com/ggml-org/llama.cpp/pull/26414
- https://github.com/ggml-org/llama.cpp/pull/26563
- https://github.com/ggml-org/llama.cpp/pull/26824
- https://github.com/ggml-org/llama.cpp/pull/27861
- https://github.com/ggml-org/llama.cpp/discussions/28248
Yup worth a try. I got from 13ts to 21ts for Qwen 3.8 Flash Next on a RX 9070 XT 16GB. As always, the additional memory demand requires some compromise. For me reducing context from full 262k to 200k and reducing kv cache down to q8_0 was the best option. moe-cache-mib at 6544. Cache hit rate at ~70%.
Would try it.
I’m pretty disappointed this one was merged, it doesn’t give good performance.
I haven’t yet understood completely the details, so can’t provide much info or suggestions for improvement yet.
Should be good to merge and iterate on it from master.
WHAT? Why did they choose to merge this one out of many if they don’t even understand it?
I was getting way better performance with the old PR from miltos22 https://github.com/ggml-org/llama.cpp/pull/26563
This is quite a comment someone left in one of the PRs lol
@am17an if 1000 line PR is too much of a scope for a project, the project is fubar. is llama project so broken it can’t handle 1000 line PRs? state so publicly, so that the community can stop wasting time contributing to a project like that and move on.
The one they accepted was nearly 1000 lines anyways.


