Basically the title.
For me, a dense model around 20b, with an additional engram table of 10-20b, or a sparse MoE around 80b would be amazing.
Basically the title.
For me, a dense model around 20b, with an additional engram table of 10-20b, or a sparse MoE around 80b would be amazing.
27B dense was a good size for 3.8, but I’m guessing they’re going to add a few B parameters in order to make it noticeably better.
If they can make the KV cache size scale more efficiently with many tokens, like DeepSeek 4.1 Flash fitting 1 million tokens in 890MB, it might give them enough room to add a couple billion parameters of weights and still fit well in 32GB GPUs
Not necessarly, I’d say. We’ve seen many releases with multiple sizes and the 27B class has been around since at least the Gemma3 times.
Well, it’s just a guess. But also Gemma4 increased to 31B, so I think there’s a possibility that Qwen does the same.