• brucethemoose@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      5 days ago

      Technically, LLMs (and most ML models) are deterministic with the same input and same seed.

      I get what you mean though.

    • dwalin@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      6 days ago

      There is a parameter in llms called temperature. If you reduce it down to zero it will become deterministic. And probably even worse.

      • frongt@lemmy.zip
        link
        fedilink
        English
        arrow-up
        0
        arrow-down
        1
        ·
        5 days ago

        Yeah but the output would be crap. Just use the same prng seed and you’ll get reproducible output.

    • sunbeam60@feddit.uk
      link
      fedilink
      English
      arrow-up
      0
      ·
      6 days ago

      Exactly this. If you made an LLM that had deterministic code output then I’m all up for saying “AI is a compiler for human language”. But until then AI is most definitely not a compiler.

        • sunbeam60@feddit.uk
          link
          fedilink
          English
          arrow-up
          0
          ·
          4 days ago

          No. Floating point arithmetic and ordering of operations won’t make 0.0 deterministic.

          • theunknownmuncher@lemmy.world
            link
            fedilink
            English
            arrow-up
            0
            ·
            edit-2
            4 days ago

            Um… yes. Matrix multiplication is deterministic, there’s nothing non-deterministic about an LLM, unless you artificially add randomness to them, like randomly selecting 1 of the top N ranked tokens.

            • sunbeam60@feddit.uk
              link
              fedilink
              English
              arrow-up
              1
              ·
              4 days ago

              The function is deterministic, agreed. The implementation is almost always not. Hardware floating point addition is not associative. So a GPU kernel that splits a reduction differently (ie interleaving it with anything else, like running your graphics, or sharing your work with other users on the same hardware) will produce different results over different runs even at temperature 0.0.

              Determinism is almost always impossible when dealing with floating point on a multi-process/multi-user system.

              There are attempts to create batch invariant language models (https://github.com/thinking-machines-lab/batch_invariant_ops) but all the major ones are not.