ChatGPT is full of sensitive private information and spits out verbatim text from CNN, Goodreads, WordPress blogs, fandom wikis, Terms of Service agreements, Stack Overflow source code, Wikipedia pages, news blogs, random internet comments, and much more.

  • MxM111@kbin.social
    link
    fedilink
    arrow-up
    3
    arrow-down
    12
    ·
    11 months ago

    That’s not true. ChatGPT does not have database - it does not have any memory at all. All it “remembers” is what you type on the screen.

      • tabarnaski@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        7
        arrow-down
        4
        ·
        11 months ago

        You remember some dialogue from your favorite movie. Does this mean your neurons store copyrighted work?

        • Excrubulent@slrpnk.net
          link
          fedilink
          English
          arrow-up
          8
          arrow-down
          2
          ·
          edit-2
          11 months ago

          Yes.

          Just because they’re in a neural network and not ASCII or unicode doesn’t mean they’re not stored. It’s even more apt a concept since apparently those works can be retrieved fairly easily, even if the references to them are hard to isolate. It seems ChatGPT is storing eidetic copies of data, which would imply what other people have said in this thread, that it is overfitting itself to the data and not learning truly generalisable language.

          • MxM111@kbin.social
            link
            fedilink
            arrow-up
            4
            arrow-down
            3
            ·
            11 months ago

            The claim is that it contains entire copies of the book. It does not. AI memory is like our memory, we do not remember books word to word.

            • Excrubulent@slrpnk.net
              link
              fedilink
              English
              arrow-up
              7
              ·
              edit-2
              11 months ago

              They are spitting out, as in the quote above, “verbatim text”, as in, word for word. That is copyrightable.

              And that’s not what you said. You said it has no memory. That’s clearly wrong.

              • anlumo@lemmy.world
                link
                fedilink
                English
                arrow-up
                1
                arrow-down
                1
                ·
                11 months ago

                It’s only under copyright if it’s a significant portion of the work. Single sentences are not enough, unless it’s a short poem.

    • NaibofTabr@infosec.pub
      link
      fedilink
      English
      arrow-up
      3
      arrow-down
      2
      ·
      11 months ago

      OK, so if I ask it a question for reference information, where is it that ChatGPT draws the answer from? Information is not stored in the model itself.

      • MxM111@kbin.social
        link
        fedilink
        arrow-up
        2
        arrow-down
        1
        ·
        edit-2
        11 months ago

        There is a memory, a storage, that would not be called a database, which encodes interaction “weights” of neurons. Those parameters where modified during training process and in some sense the information is somehow encoded there. But it is not possible to decode the whole book word to word. It is very similar to our memory in this sense. Do you remember any book word to word? The whole book?

        • NaibofTabr@infosec.pub
          link
          fedilink
          English
          arrow-up
          3
          arrow-down
          2
          ·
          11 months ago

          You understand that the neural network is not the entire picture, right? Like, yes you’re correct in general about how these models are trained, but ChatGPT does not operate in a vacuum. For instance, when it was connected to the internet that was just for information searching - the neural network in use was frozen, it wasn’t actively training on internet content.

          It’s a language system, it can operate as a search tool, it has to have access to a source of information in order to generate responses to queries. That source of information isn’t contained in the model itself, but it is connected to it and it’s part of the whole ChatGPT system.

          • MxM111@kbin.social
            link
            fedilink
            arrow-up
            1
            ·
            11 months ago

            ChatGPT4 now indeed can connect to internet and read the sites and summarize the data. But this has nothing to do with storing it he whole books in their memory. It read the internet sites exactly the same way as you and me do. I do not understand what is your argument here. Internet is external to ChatGPT.