Discover · megaphone.fm

How to train your data

The Vergemegaphone.fm

This is what was read here when it was published. The Verge no longer carries it in what it publishes now.

Training data is the raw material of the AI industry. Claude, ChatGPT, Gemini, and the rest are built on top of oceans of stuff. What is that stuff? Books. Blog posts. YouTube videos. Reddit comments. All of it and more, in virtually incomprehensible quantities. Alex Reisner, a staff writer at The Atlantic who has been investigating training data, explains how AI companies get all this data, why they'd really prefer you not know what's in it, and whether training data could ever be a fair trade. Further reading: Apple raises prices on Macs, iPads, and more by hundreds of dollars | The Verge⁠…

More from megaphone.fm

Also in Discover

  1. Sony execs behind iconic game-sharing meme recreate it, but fittingly, without a physical gameKaan Serineurogamer.net
  2. What will happen to your most cherished possessions when you die? You don’t want to knowRichard Gloversmh.com.au
  3. Once pop’s wiliest genius, Beck sounds in need of a proper sea changeBarry Divola, Ben Maddensmh.com.au
  4. MacklemoreWikipedia
  5. That NFL match actually cost double previous estimates. But the jubilation made it worth itStephen Brooksmh.com.au
  6. Trump says he is banning CNN and other news outlets from the White HouseMichael Koziolsmh.com.au