this post was submitted on 04 Dec 2023
888 points (97.9% liked)

Technology

59466 readers
3638 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related content.
  3. Be excellent to each another!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, to ask if your bot can be added please contact us.
  9. Check for duplicates before posting, duplicates may be removed

Approved Bots


founded 1 year ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] Jordan117@lemmy.world 12 points 11 months ago (2 children)

IIRC based on the source paper the "verbatim" text is common stuff like legal boilerplate, shared code snippets, book jacket blurbs, alphabetical lists of countries, and other text repeated countless times across the web. It's the text equivalent of DALL-E "memorizing" a meme template or a stock image -- it doesn't mean all or even most of the training data is stored within the model, just that certain pieces of highly duplicated data have ascended to the level of concept and can be reproduced under unusual circumstances.

[–] gears@sh.itjust.works 13 points 11 months ago* (last edited 11 months ago)

Did you read the article? The verbatim text is, in one example, including email addresses and names (and legal boilerplate) directly from asbestoslaw.com.

Edit: I meant the DeepMind article linked in this article. Here's the link to the original transcript I'm talking about: https://chat.openai.com/share/456d092b-fb4e-4979-bea1-76d8d904031f

[–] lemmyvore@feddit.nl 10 points 11 months ago (1 children)

Problem is, they claimed none of it gets stored.

[–] TWeaK@lemm.ee 5 points 11 months ago

They claim it's not stored in the LLM, they admit to storing it in the training database but argue fair use under the research exemption.

This almost makes it seems like the LLM can tap into the training database when it reaches some kind of limit. In which case the training database absolutely should not have a fair use exemption - it's not just research, but a part of the finished commercial product.