Row 18558

Row ID: 18558 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 18558 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

It will be useful in a lot of ways, but probably difficult to improve on past where it’s at.

Existing LLMs were created by hoovering up as much human creative output as they could download off the internet. It needs absurd amounts of examples to reliably train. They did a good job of that, so there isn’t a lot *more* data to train from.

What *did* happen was tons of people indiscriminately posting LLM output online. Often deliberately trying to obfuscate that it is AI generated. That problem is only going to get worse. Attempting to train nearly any AI system on its own output tends to fuck up the results pretty badly.

I believe on one level we’re going to see a “peak” to LLMs just in how little of human intelligence they mimic, but we’re also likely to see a ceiling caused by lack of new reliable training data.

I guess we’ll see if they can work around these issues, but honestly the bombast and rampant cashing in on display in AI circles doesn’t inspire a lot of confidence that they’ll actually pull it off.

FieldValue
text It will be useful in a lot of ways, but probably difficult to improve on past where it’s at. Existing LLMs were created by hoovering up as much human creative output as they could download off the internet. It needs absurd amounts of examples to reliably train. They did a good job of that, so there isn’t a lot *more* data to train from. What *did* happen was tons of people indiscriminately posting LLM output online. Often deliberately trying to obfuscate that it is AI generated. That problem i…
label r/technology
dataType comment
communityName r/technology
datetime 2024-05-21
username_encoded Z0FBQUFBQm5LakwtWHItcUVTRFVDZUlsZUFQTmxkN09DaGhGdzNmWHNHdVhHWFRyY25waUpzbC1ZTmFkY0ZjX2gyYTVZaDhURVkxWlhLUGRkWXJ5c09tSTY4SXFibXdRSkE9PQ==
url_encoded Z0FBQUFBQm5Lak9ORGFFMDRQcWRNbXhRc1dTam9BamlEY0F5YmYyUHVTVXFnN29rTGFvaXFVdE95YkFTdGIxQU1CNDNXRnpFQ2Nzd3ZYb0gtdXZ5Z3pKQ3NZY2lXb3NPNTN5N1dLaWVtVWhGZm5zUGFGS0szME1xNnIyZTVNbDlybWQtUHFDX1kxY0ZNV1E2T2ZURmZTNkFLV0JBVm9oTnhRdkczLTQxQnZ5MXdyT3JHZU1RZG83R25TX3FLenc5ekZZSHQ2cUs3ZjQxd3pxNmFFeTNoS2ZHdUo3TGZSX3d0UT09

Raw Record

{
  "text": "It will be useful in a lot of ways, but probably difficult to improve on past where it’s at.\n\nExisting LLMs were created by hoovering up as much human creative output as they could download off the internet. It needs absurd amounts of examples to reliably train. They did a good job of that, so there isn’t a lot *more* data to train from.\n\nWhat *did* happen was tons of people indiscriminately posting LLM output online. Often deliberately trying to obfuscate that it is AI generated. That problem is only going to get worse. Attempting to train nearly any AI system on its own output tends to fuck up the results pretty badly.\n\nI believe on one level we’re going to see a “peak” to LLMs just in how little of human intelligence they mimic, but we’re also likely to see a ceiling caused by lack of new reliable training data.\n\nI guess we’ll see if they can work around these issues, but honestly the bombast and rampant cashing in on display in AI circles doesn’t inspire a lot of confidence that they’ll actually pull it off.",
  "label": "r/technology",
  "dataType": "comment",
  "communityName": "r/technology",
  "datetime": "2024-05-21",
  "username_encoded": "Z0FBQUFBQm5LakwtWHItcUVTRFVDZUlsZUFQTmxkN09DaGhGdzNmWHNHdVhHWFRyY25waUpzbC1ZTmFkY0ZjX2gyYTVZaDhURVkxWlhLUGRkWXJ5c09tSTY4SXFibXdRSkE9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9ORGFFMDRQcWRNbXhRc1dTam9BamlEY0F5YmYyUHVTVXFnN29rTGFvaXFVdE95YkFTdGIxQU1CNDNXRnpFQ2Nzd3ZYb0gtdXZ5Z3pKQ3NZY2lXb3NPNTN5N1dLaWVtVWhGZm5zUGFGS0szME1xNnIyZTVNbDlybWQtUHFDX1kxY0ZNV1E2T2ZURmZTNkFLV0JBVm9oTnhRdkczLTQxQnZ5MXdyT3JHZU1RZG83R25TX3FLenc5ekZZSHQ2cUs3ZjQxd3pxNmFFeTNoS2ZHdUo3TGZSX3d0UT09"
}

Entry Information