Basically the title.
For me, a dense model around 20b, with an additional engram table of 10-20b, or a sparse MoE around 80b would be amazing.
27B dense was a good size for 3.8, but I’m guessing they’re going to add a few B parameters in order to make it noticeably better.
If they can make the KV cache size scale more efficiently with many tokens, like DeepSeek 4.1 Flash fitting 1 million tokens in 890MB, it might give them enough room to add a couple billion parameters of weights and still fit well in 32GB GPUs
Not necessarly, I’d say. We’ve seen many releases with multiple sizes and the 27B class has been around since at least the Gemma3 times.
Well, it’s just a guess. But also Gemma4 increased to 31B, so I think there’s a possibility that Qwen does the same.
Another 35B-A3B MoE or similar would be nice…
FYI, Yandex released an 80B MoE base model (i.e. not instruction tuned!) recently: https://huggingface.co/yandex/AliceAI-Foundation-80B-A3B-Base – haven’t experimented with it personally though.
I hope an instruction tuned will be released as well.
32GB RAM and 8GB VRAM here, so 35b a3b is about my maximum, maybe with added n-grams


