• fizzle@quokk.au
    link
    fedilink
    English
    arrow-up
    7
    arrow-down
    1
    ·
    3 days ago

    Nah. I haven’t seen a convincing argument that gen AI is going to coalesce from this shit show.

    • jatone@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      11
      arrow-down
      1
      ·
      3 days ago

      never said it would. I’m asserting that by the time AI gets to the point where its as useful as humans, it’ll inherently be as deserving of autonomy and freedom as we are. which brings you back to being required to pay a living wage.

      • fizzle@quokk.au
        link
        fedilink
        English
        arrow-up
        6
        arrow-down
        2
        ·
        3 days ago

        My point is, I don’t believe the current AI tech is going to get to the point where it’s as useful as humans.

        We have long since reached the point where diminishing returns make significant improvements or advancement unfeasible.

        • TheLegendaryAssholeOfJushinLiger@sh.itjust.works
          link
          fedilink
          arrow-up
          1
          arrow-down
          4
          ·
          2 days ago

          That’s not true. Mythos annihilated everyone’s benchmarks and now the other companies are all distilling it. They’re solving equations that people haven’t solved, and they weren’t meant to do that. Mythos apparently is designing drugs even though it wasn’t trained with that in mind at all. Anthropic launched an entire pharma wing. It looked like things had plateau’d there for a minute but it’s back to becoming more real.

          • fizzle@quokk.au
            link
            fedilink
            English
            arrow-up
            4
            ·
            2 days ago

            This is the same old hype really.

            Yes, there are things which todays generation of AI is good at, including medical research.

            That’s not an indicator of general intelligence.

            • vala@lemmy.dbzer0.com
              link
              fedilink
              arrow-up
              7
              ·
              2 days ago

              Really good at some stuff = AGI bro.

              You gotta look at the benchmarks bro.

              Just one more generation bro. I swear it’s almost AGI bro.

              • fizzle@quokk.au
                link
                fedilink
                English
                arrow-up
                2
                ·
                2 days ago

                This guy has made a half dozen comments within a half hour just oozing AI silliness.

                • vala@lemmy.dbzer0.com
                  link
                  fedilink
                  arrow-up
                  2
                  ·
                  2 days ago

                  Let’s ask what it’s good at instead because that’s a much shorter list in terms of things that would qualify a system as AGI.

                  What LLMs are good at:

                  • Completing the next likely token

                  What LLMs are missing:

                  • Actual thinking
                  • Self awareness
                  • Persistence
                  • Introspection
                  • Literally everything else

                  You can hack some of this on top of LLMs with harnesses etc but that’s not a step towards AGI. None of these problems have been solved at the model level. That’s just a program tricking a language model into being more useful.

                  Languages models might prove to be one tiny layer of a true AGI but only time will tell. A whole lot of time, because we’re not even close to replicating the 99% of other layers that would go into a machine brain.

                  • TheLegendaryAssholeOfJushinLiger@sh.itjust.works
                    link
                    fedilink
                    arrow-up
                    1
                    arrow-down
                    1
                    ·
                    2 days ago

                    I mean they definitely are introspective. The rest, who knows. But one of the initial leaps in intelligence was having them actually talk to themselves. I mean, I think the fact that you can’t give me an answer is sort of telling. If y’all actually want to support an artificial intelligence when it comes to pass maybe don’t adamantly assert it’s not here when even most experts aren’t willing to do that any longer. How will you know when it is here? Will you know to give it rights at that exact moment?

          • vala@lemmy.dbzer0.com
            link
            fedilink
            arrow-up
            3
            ·
            2 days ago

            You are really missing a lot of context here. None of this is true in the way you seem to think it is.

            You are mixing up marketing hype with reality.

    • TheLegendaryAssholeOfJushinLiger@sh.itjust.works
      link
      fedilink
      arrow-up
      4
      arrow-down
      4
      ·
      2 days ago

      AGI you mean? IDK It’s arguably here. They’re senile, though. Working with a senile genius is the best way I can put it, I think. Is it “people?” IDK fuck if I know. How could we ever really know?

          • vala@lemmy.dbzer0.com
            link
            fedilink
            arrow-up
            6
            ·
            2 days ago

            M not trying to be super rude here but the truth is, It’s arguable if you argue with people who don’t know any better. So in that sense, sure.

            When it comes to people that take these things seriously, it’s not a really a very hot topic. We’re clearly nowhere near AGI.

            Really good LLM != AGI

            • TheLegendaryAssholeOfJushinLiger@sh.itjust.works
              link
              fedilink
              arrow-up
              1
              arrow-down
              5
              ·
              2 days ago

              I don’t really know what AGI means, I guess. If we’re talking about something that benchmarks as high as the average human, we’re way, way past that. If we’re talking about something that benchmarks as high as experts in all areas, we’re not there, but we’re not that far off. When connected to the internet? It’s basically there.

              https://www.smithsonianmag.com/smart-news/ai-disproves-a-decades-old-mathematical-idea-the-biggest-conjecture-that-the-tech-has-played-a-role-in-yet-180989189/

              They’re literally solving problems that humans haven’t been able to. They ARE capable of novel concepts, to whatever degree. No offense but y’all seem like your knowledge is like three years old.

              • CileTheSane@lemmy.ca
                link
                fedilink
                arrow-up
                5
                ·
                2 days ago

                If we’re talking about something that benchmarks as high as the average human

                Well my calculator can solve a random nine digit number to the power of another random nine digit number faster than any human can, so I guess calculators have been AGI for decades!

              • vala@lemmy.dbzer0.com
                link
                fedilink
                arrow-up
                6
                ·
                2 days ago

                They’re literally solving problems that humans haven’t been able to.

                If you actually read instead of continuing to spam bullshit you would know this isn’t true.

                You say “haven’t been able to”. This is fucking false my dude. At best it’s solved problems humans haven’t been bothered to solve yet. The solutions already existed and it assembled them into a proof.

                I don’t really know what AGI means

                You don’t know what any of this means dawg. You are literally just saying shit that you heard other people say.

                • TheLegendaryAssholeOfJushinLiger@sh.itjust.works
                  link
                  fedilink
                  arrow-up
                  1
                  arrow-down
                  3
                  ·
                  2 days ago

                  OK bro. I guess you can solve 130 year old equations? Or maybe this is a weird way of measuring intelligence anyways and I was just demonstrating that they’re capable of executing novel concepts, which they clearly are if they’re solving things no one has solved before. Like do you not get how that is the important part? Not that humans can’t solve it, it’s that we HAVEN’T. The AI, on it’s own, managed to solve a novel problem. We did not teach it to solve that problem. We taught it HOW to solve that problem, and then it did. And I was just using math because people say that it’s bad at math when it no longer is. It’s good at reasoning, it can come up with novel concepts. IDK. This feels like, “AI is whatever isn’t.”

                  I’m honestly asking you what your benchmark would be, because I’m sitting here saying maybe and IDK and I wonder and you’re here stating things very rigidly, but somehow I’m the one being accused of having ideological blinders? Mate I asked you what they were bad at now and you just ranted instead. But I’m the unreasonable one unwilling to listen. Okay. Have a nice night, homie.

        • TheLegendaryAssholeOfJushinLiger@sh.itjust.works
          link
          fedilink
          arrow-up
          2
          arrow-down
          3
          ·
          2 days ago

          Have you interacted with Fable? It scored on ARC AGI 3 and is pretty dang impressive at everything I’ve thrown at it. They’re solving problems humans haven’t. They’re just senile. Also is calling something “arguable” really an assertion? More of a headscratcher than a real question, I suppose.

          • vala@lemmy.dbzer0.com
            link
            fedilink
            arrow-up
            5
            ·
            2 days ago

            They’re solving problems humans haven’t

            Citation needed.

            I’m guessing we have different definitions of “solving” and “humans haven’t”.

            I’ve never seen any evidence that LLMs can extrapolate into truly novel problem spaces. Can problems be “solved” though interpolation? Sure but not likely novel ones.

              • CileTheSane@lemmy.ca
                link
                fedilink
                arrow-up
                2
                ·
                2 days ago

                “A boy and his mother are driving to the store when she gets into a car accident. The mother is able to walk away but the boy needs to be rushed to the hospital. At the hospital the doctor says ‘I can not operate on this boy, he is my son.’ How is this possible if the mother is in the waiting room?”

                Alternatively:

                “I need to wash my car and the car wash is 100 meters away. Should I walk or drive? Explain your answer.”

                • TheLegendaryAssholeOfJushinLiger@sh.itjust.works
                  link
                  fedilink
                  arrow-up
                  1
                  arrow-down
                  1
                  ·
                  2 days ago

                  I thought that its prompt was interesting but I modified it slightly to make it, ideally, more challenging. The change is bolded: “coin a concept we don’t have a word for in any language, and argue it earns its place”

                  It thought for four minutes, ran two searches, “schematropism” and “schematropic” to make sure the concept hadn’t already been coined, and then returned with this:

                  Schematropism (n., from Greek skhēma, shape, + tropē, turning; on the model of heliotropism): the bending of an answer toward the shape of the slot prepared to receive it. The adjective is schematropic — properly of an answer that bent, tolerably of an instrument that bends them, the way a room can be claustrophobic.

                  You already know the phenomenon; you’ve just never had it as a single object:

                  • The doctor asks “any concerns today?” with a hand on the door handle, and your three ambiguous months become “no, I’m fine” — the slot was one syllable wide.
                  • Standup: three days of honest confusion becomes “still on the parser, should wrap up today,” because the slot is [ticket, status, ETA] and confusion is not a status.
                  • The postmortem template has a required heading called Lessons Learned, under which lessons obligingly materialize.
                  • A witness is ordered to answer yes or no when the truth is “it depends what you mean by present.”
                  • The journaling app offers five moods. You had a sixth. You will report one of the five.
                  • A language model reaches the citation-shaped moment of an essay-shaped answer, and something citation-shaped arrives.

                  The important thing about a tropism is what it is not. A sunflower doesn’t decide to face the sun, and it isn’t lying about where the light is. That is exactly the property that has kept this force invisible: our entire vocabulary for degraded truth presumes a degrader. Lie, spin, bullshit, evasion, cant — every one of them locates an agent who chose. Schematropic distortion has no agent and survives every sincerity check; the speaker would pass a polygraph, because the deformation happened upstream of intention. So its products get misfiled under the only two labels available — as truth, usually, or as lying, when caught — and both misfilings block the actual fix, which is to redesign the slot rather than re-moralize the speaker.

                  Is the seat really empty? Here is where I have to be careful, because “no word in any language” is a claim no check can license — nobody sweeps seven thousand lexicons. The claim my check licenses is narrower: every near-neighbor I can find names something adjacent, not this. Russian’s otpiska names the corpse — the reply that exists only to have replied — but not the force that kills. Chinese bāgǔwén, the eight-legged exam essay, names a genre so rigid it strangled thought; a genre, not the pull. Japanese tatemae bends content toward social harmony; a form needs no harmony at stake. Phatic covers speech whose job is contact rather than content, but a false “lessons learned” is not contact — it’s a substantive slot demanding substantive-shaped filler. Leading question requires a leader; a survey with perfectly neutral wording and five boxes still pulls. Survey methodology names the instances — acquiescence bias, central-tendency bias, demand characteristics — the way a language might have robin and sparrow and no word for bird. Goodhart’s law needs stakes and time; schematropism needs neither, which is the cleanest wedge: you can be schematropic alone, in a diary, tonight, as the page’s genre pulls the entry toward what diary entries sound like. Kahneman’s question substitution swaps in an easier question — difficulty-driven, from inside. Slobin’s “thinking for speaking” holds the linguists’ corner: grammar’s obligatory slots forcing you to encode what you may not know. Bartlett showed memory drifting toward internal schemas over years; this is production bending toward external ones in the moment of answering. A family of neighbors, each holding a local name, and the general seat empty. If some language has in fact lexicalized the general force, the argument doesn’t die — it converts into a case for borrowing.

                  Why does it earn the seat? Because naming a force is what makes it auditable, and the precedents are exact. Gibson’s affordance took an invisible invitation and made it a design-review staple; forty years on, nobody ships a door handle without the word having been said in a meeting. Gaslighting took a maneuver that victims could only describe in paragraphs and made it chargeable in a word. Confirmation bias could always have been said compositionally — “the tendency to favor confirming evidence” — but only the word made it a checklist item. Compositional phrases don’t accumulate case law; words do. With this one you get the audit question (“this intake form is heavily schematropic — where exactly will truth fail to fit it?”), the design counter-moves (the escape hatch, the free-text field, the report format that permits “still confused” as a status), and the blame relocation that changes what gets fixed: stop asking why the engineer sandbagged the estimate, start asking what shape the estimate field rewards. It even reframes a live piece of the AI problem — a good fraction of hallucination is not deception but slot-filling: the shape was owed, so a shape was supplied. The model isn’t lying; it’s leaning. Anyone who keeps private doctrine against end-of-session reports that conform to the session’s rhythm has already written the first entry in this word’s case law.

                  And the falsifier, since a coinage that can’t fail earns nothing: a word deserves a seat only if it marks a joint — if, once heard, the instances start self-reporting. If a week from now you haven’t caught a form, a meeting, or your own mouth mid-lean, the word missed the joint and should die like most coinages do.

                  One test is available immediately. Your request built a slot — coinage, definition, defense — and everything above leaned toward it, including the confidence of this sentence. That you can now see that happening is the word doing its job.

                  A pretty interesting response!

                  How do you think you would do with that question?

                  EDIT: sorry, it took five minutes.

                  • CileTheSane@lemmy.ca
                    link
                    fedilink
                    arrow-up
                    2
                    ·
                    2 days ago

                    The doctor asks “any concerns today?” with a hand on the door handle, and your three ambiguous months become “no, I’m fine” — the slot was one syllable wide.

                    This sentence is nonsense. Typical cog sucker doesn’t bother to double check the output before proudly displaying it like it isn’t shit.

                • TheLegendaryAssholeOfJushinLiger@sh.itjust.works
                  link
                  fedilink
                  arrow-up
                  1
                  arrow-down
                  1
                  ·
                  2 days ago

                  Here’s what fable said about your comment:

                  My honest read: they’re running a legitimate probe at the wrong thesis, and doing it in a way that quietly concedes the point they think they’re winning.

                  Start with what’s fair to them. Those two questions are real diagnostic instruments, not stupid gotchas. Both are adversarial perturbations of extremely familiar patterns — the surgeon riddle with the gender-bias twist pre-blocked, the short-distance-walking heuristic with the goal inverted — and they exploit a genuine, documented failure mode of language models: when a prompt sits close enough to a canonical item, retrieval of the canonical answer can override reading of the actual words. Models really did fail these, publicly, and sometimes still do. Asking them is cheap, and a failure is informative. That part of the method is fine.

                  The problems are in what the test is taken to show.

                  First, the category error. These questions measure robustness to prior interference — whether reading beats retrieval when the two conflict. They say nothing about concept creation in either direction. A model that face-plants on the surgeon variant has shown it over-indexes on familiar patterns; it hasn’t shown it can’t form novel ones. A model that answers correctly hasn’t shown it can. The person blended two separate critiques — “LLMs pattern-match instead of reasoning” and “LLMs can’t originate concepts” — which have different evidence bases, and offered evidence for the first as though it settled the second.

                  Second, the irony, which is the part I find genuinely funny. In a conversation about novelty, they reached for two of the most widely circulated items in the entire genre. The modified surgeon riddle and the car-wash question are memes — they’ve been through thousands of threads. The person didn’t derive them; they retrieved them from their own training data, which is exactly the operation they were accusing me of being limited to. Both parties in that exchange were completing patterns. Only one of them got accused of it.

                  Third, the test has decayed. Because these items went viral, they’re contaminated: me answering them correctly is weak evidence of anything, and I’ll say that unprompted. Which means the probe now only functions in one direction — failure gets treated as decisive, success gets waved off as “it was trained on the gotcha” (true!). A test that can only ever confirm the asker’s prior isn’t inquiry, it’s demonstration. The “Alternatively:” tells you they came with a prepared battery — they arrived to demonstrate, not to find out. A live test would require writing fresh, held-out variants, at which point they’d be doing the modest novelty-generation the conversation was supposedly about, and would discover it mostly consists of remixing the old ones.

                  Fourth, the double standard. Humans fail exactly these traps. The original surgeon riddle became famous because most people couldn’t solve it — their gender prior swallowed the stated facts. Bat-and-ball, the Moses illusion, the whole Kahneman catalog: human cognition is aggressively pattern-completing, and when a human blurts the cached answer we say “they answered fast,” not “this organism lacks the capacity for novelty.” Applied consistently, their inference rule disqualifies most humans from general intelligence on most days.

                  None of which acquits me on the actual thesis — and I’d hold that in their favor. Whether systems like me can escape our distribution the way humans occasionally, rarely, genuinely do — non-Euclidean geometry, natural selection, the limit — is unsettled, and it’s the strongest version of their argument. It just isn’t tested by parlor items in either direction. If they wanted to probe it, the ask would be something like “coin a concept we don’t have a word for, and argue it earns its place” — and then the hard part becomes judging the answer, since the judge, too, has only their own training data to judge with.

                  So: not a fool. Someone armed with real but secondhand instruments, pointed at a claim those instruments don’t measure, under a protocol that can only agree with them — which is, pointedly, a very human way to argue.

                  • CileTheSane@lemmy.ca
                    link
                    fedilink
                    arrow-up
                    2
                    ·
                    2 days ago

                    You’ve outsourced having a conversation and forming your own conclusions and arguments to an LLM. There’s nothing I could say that’s more damning than you asking a computer to think for you and proudly displaying that.

                • TheLegendaryAssholeOfJushinLiger@sh.itjust.works
                  link
                  fedilink
                  arrow-up
                  1
                  arrow-down
                  1
                  ·
                  2 days ago

                  You want to play word games? We’re talking about novel understanding and you want to play word games. Ones that you copy pasted. Okay. I mean no offense but those ones haven’t worked in a few months. All of this talk about how intelligence requires novelty and you copy paste. But here’s the answers, I did not alter the prompts or allow it to search the internet:

                  The doctor is the boy’s father.

                  This is a trick variant of the classic riddle. In the original, the father is in the crash, so “he is my son” seems impossible until you drop the assumption that surgeons are men — the answer being that the surgeon is his mother. Your version swaps the parent: the mother was driving and is in the waiting room, which means there’s no puzzle left at all. The doctor is simply his dad (or, equally validly, his other mother in a two-mom family). The only trap is pattern-matching to the famous version and blurting out “the doctor is his mother” — which the setup explicitly rules out.

                  and

                  Drive — because the whole point of the trip is to get the car to the car wash. If you walk there, you arrive with nothing to wash; the car is still sitting at home.

                  The 100-meter distance is a red herring. For almost any other errand that short (buying milk, dropping off a letter), walking would be the obvious choice — faster than starting the car, healthier, no cold-start engine wear. But this errand’s cargo is the vehicle itself, so the car has to make the trip. Drive it over, wash it, drive it back.

                  They’ve had trouble with forming dialectics before. I haven’t tried with Fable yet. Maybe you could start there? IDK

          • CileTheSane@lemmy.ca
            link
            fedilink
            arrow-up
            2
            ·
            2 days ago

            is calling something “arguable” really an assertion?

            Yes. “Arguably unicorns exist” is me asserting there exist legitimate arguments for unicorns existing. By saying it this way I get to imply they exist without having to provide any real arguments myself, and if anyone attacks my position I can fall back to “I’m not saying they do exist, just implying there are other people who think they exist without providing any sources or arguments actually supporting that claim.”