• FiniteBanjo@feddit.online
    link
    fedilink
    English
    arrow-up
    15
    arrow-down
    3
    ·
    edit-2
    21 hours ago

    Lol what? Of course it won’t, If the AI slop ends with recursive edits it’s just going to cause degradation and collapse. I swear techbros have reality confused with their favorite fantasy fiction books.


    EDIT: To demonstrate, 90% accuracy of 90% is 81%. Even the best most specific models on earth are not capable of self improvement because they will never reach much less exceed their training data’s capability even if the largest most perfect dataset existed. They might think that by simply adding more layers of machines running in parallel and killing off models which underperform creating a system similar to evolutionary adaptation that it might eventually reach that 91%, but our current approach and level of technology have never demonstrated that capability not even theoretically.

    • MangoCats@feddit.it
      link
      fedilink
      English
      arrow-up
      3
      arrow-down
      9
      ·
      19 hours ago

      What happened in the computer programming space (with testable outputs) is that the first pass 80% accuracy nailed down an 80% success rate - wrote code that successfully met requirements 4/5 trials. Then, the agents were able to repeat the 1/5 failing trials with “sufficient heat” to both find their problems and create workable solutions, again 4/5 trials - so 80% success rate becomes 96% success rate, and so on… Back in early 2025, programming LLM agents would get themselves caught in iterative loops - trying, failing, trying again, failing again, then trying the first approach again - failing indefinitely. By mid 2026, I don’t see that behavior anymore - if the first “light pass - quick attempt” solution doesn’t succeed, they dig in deeper - do more research specifically focused on the problem areas identified in the first failure and try again, generally successful by the 2nd try, almost always by the 3rd - I haven’t had to break a “trying the first unworkable solution again because I can’t think of anything else to do” loop in over 6 months.

      Not all problem spaces are as clear-cut as software creation, but many have similar rules that just take a bit more training to learn.

      • Zos_Kia@jlai.lu
        link
        fedilink
        English
        arrow-up
        1
        ·
        3 hours ago

        You’re entirely right. This won’t replicate to other fields like writing and creative arts in general, but software engineering is just not that hard and can basically be brute forced with a good harness.

        It’s a done deal and there is no world where people will write professional code by hand. I like it cause it really separates coding (the job) from coding (the art form). People will code by hand for aesthetic reasons just like people learn the violin instead of using a synth and we’ll have a generation of lovingly crafted stuff. But boring software will be entirely automated, if not generated on the fly based on immediate needs.

        • MangoCats@feddit.it
          link
          fedilink
          English
          arrow-up
          1
          ·
          3 hours ago

          software engineering is just not that hard

          I’ll disagree on semantics here, it’s precisely because software engineering is hard (not difficult, but rigid - objective) that makes it a good fit for LLM agent execution. Soft, squishy, ill-defined fields are going to be a worse fit for LLMs because the practitioners themselves can’t create clear cut (hard) definitions of what it is they expect out of their practitioners, they just “know it when they see it.” As for relative difficulty, the “soft” fields have a very sliding scale for that with a lot of allowance given to newbies that isn’t accepted “at the highest levels” whereas, software engineering just is what it is, it doesn’t get more difficult as you progress in the field. Your job as a software architect / engineer is actually to find the easiest workable solution(s).

          People will code by hand for aesthetic reasons just like people learn the violin instead of using a synth

          I think it’s more like: people will code C or Rust or Python by hand just like people still code assembly by hand - exceptionally rare stubbornness with an exceptionally small audience who could even understand what they have done to begin to care about it. Violin vs synth - most of the world can listen and appreciate and have an opinion even if a vanishingly small fraction could ever hope to have the patience, let alone skill, to compose or perform at the highest levels of either form. “Synth” is a very broad target these days, varying from direct composition to performance digital transformation, through interfaces of every description and complexity: simple contact closure keyboards through multi-dimensional velocity, attack angle, strike momentum, and many dimensions of aftertouch bends which allow more expressivity than even bow and fingers on strings do, if the performer cares to train in that popularly scorned field. Having done a little amateur composition to performance vs performance capture synth work, I’ll say: once you have trained to work with the complex input devices, capture of live performance is hundreds of times more efficient than specifying all the nuance of a real performance as notation in a composition. The main reason people hate synth performances is that most synth performances are hack level, because hack level is easier (read: possible) on synth than a minimally passable live performance on violin with strings and bow.

          Similarly, most people are hating on AI slop because it’s so easy to produce and so many untalented hacks are using it to produce sub-par whatever it is they are making: code, prose, art, music… used as a tool, with a high bar of standards required before publication and release, LLMs are a powerful tool that can accelerate many creative processes, not just produce a lot of slop quickly.

          • Zos_Kia@jlai.lu
            link
            fedilink
            English
            arrow-up
            1
            ·
            39 minutes ago

            it’s precisely because software engineering is hard (not difficult, but rigid - objective) that makes it a good fit for LLM agent execution

            Yes i think we’re actually in agreement here. I said “hard” (not difficult) as a reference to “hard problems”, a term that comes from complexity theory but is now commonly used to describe problems which can’t be reduced to an algorithm or evaluated objectively, and thus can’t readily be “solved”.

            Squishy subjects like music and sociology are full of hard problems, while solid subjects like math and coding are full of easy problems. Now the change introduced by LLMs is that as long as a problem is “easy”, it can no longer be so laborious as to be impossible. Every software problem is solvable, modulo the effort/computing power you can spend on it.

            Your job as a software architect / engineer is actually to find the easiest workable solution(s).

            You also get bonus point if your solution is average (standard, unsurprising etc…), which makes it particularly soluble in LLMs which, by definition, can only produce output that is within the distribution of their training set.

            As for relative difficulty, the “soft” fields have a very sliding scale for that with a lot of allowance given to newbies that isn’t accepted “at the highest levels”

            That’s not where i would put the difference. If you take a field like music, the problem is that it can’t “just work”. A nostalgic song may move the masses today but you can’t say “okay we’ve solved nostalgia let’s get to serenity next”. Soon enough you’ll need a new nostalgic song and by definition it will be out of distribution. You can’t find it in a high dimensional representation of past music, and, well, you can’t train on future data, so there is no way an LLM finds it and recognizes it for what it is.

            with a high bar of standards required before publication and release, LLMs are a powerful tool that can accelerate many creative processes

            I still believe they’ll never amount to much regarding artistic processes, and not just for the reasons i already mentioned. To make something good you need to sit with it and walk with it and spend some time in it doing all the tedious little tasks until it really feels like home and you can express yourself in it. You can’t achieve that if a machine speedruns all the little tasks for you.

      • FiniteBanjo@feddit.online
        link
        fedilink
        English
        arrow-up
        6
        arrow-down
        4
        ·
        18 hours ago

        They Don’t pass 4/5.

        They pass 0/5 because they are 80% (that number is way too optimistic btw) accurate to human output on every one of the five attempts.

        They also can’t be forced to learn and retake the trial because they don’t have any contextual awareness, they just guess the next word in a sequence.

        Even if a machine made 4 self edits sucessfully, it would be permanently disfigured by the one failure and no longer be capable of making good edits.

        • Zos_Kia@jlai.lu
          link
          fedilink
          English
          arrow-up
          2
          arrow-down
          1
          ·
          3 hours ago

          That’s cool but you’re describing the models from 2 years ago and also not considering harnesses, which account for most of the progress of the last year or so.