Basically the title.

For me, a dense model around 20b, with an additional engram table of 10-20b, or a sparse MoE around 80b would be amazing.

  • ffhein@lemmy.world
    link
    fedilink
    English
    arrow-up
    2
    ·
    14 hours ago

    27B dense was a good size for 3.8, but I’m guessing they’re going to add a few B parameters in order to make it noticeably better.

    • BeefAndPoultry@lemmus.org
      link
      fedilink
      English
      arrow-up
      1
      ·
      10 hours ago

      If they can make the KV cache size scale more efficiently with many tokens, like DeepSeek 4.1 Flash fitting 1 million tokens in 890MB, it might give them enough room to add a couple billion parameters of weights and still fit well in 32GB GPUs

    • robber@lemmy.mlOP
      link
      fedilink
      English
      arrow-up
      1
      ·
      13 hours ago

      Not necessarly, I’d say. We’ve seen many releases with multiple sizes and the 27B class has been around since at least the Gemma3 times.

      • ffhein@lemmy.world
        link
        fedilink
        English
        arrow-up
        2
        ·
        12 hours ago

        Well, it’s just a guess. But also Gemma4 increased to 31B, so I think there’s a possibility that Qwen does the same.