Row 87438

Row ID: 87438 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 87438 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

Been doing a research project on an article about AI and from what the devs are saying, current technologies are going to plateau HARD very soon. The issue is there are really two main ways to get data. 1) Feed it the Internet. This is great volume wise, but there is so much useless crap out there that the data needs to be sanitized so much that incremental improvements cost boat loads. There really isn’t that much upside to parsing the next 5% of the Internet‘s text for $$$$$$$$$ with extremely little upside. 2) Feed hand-picked, quality websites, from a variety of sources. The problem with this is, there really isn’t enough good data out there in the first place and it’s a PITA and not worth it to comb through the Internet to tell ChatGPT what to learn. That’s really why GPT hasn’t improved much since the 3.5 model dropped.

If development continues on better technologies, we should be a bit scared. But currently, save for better APIs that work with LLM-based models (think better implementations of Adobe Firefly, for example), there really isn’t that much room to grow with the current models (at least they don’t think there is). Expect companies like Google and OpenAI to milk current technologies so they become more and more profitable a la Google Search until a vast improvement comes down the line.

FieldValue
text Been doing a research project on an article about AI and from what the devs are saying, current technologies are going to plateau HARD very soon. The issue is there are really two main ways to get data. 1) Feed it the Internet. This is great volume wise, but there is so much useless crap out there that the data needs to be sanitized so much that incremental improvements cost boat loads. There really isn’t that much upside to parsing the next 5% of the Internet‘s text for $$$$$$$$$ with extremely…
label r/technology
dataType comment
communityName r/technology
datetime 2024-05-24
username_encoded Z0FBQUFBQm5Lak1wZV92V3lZLU1rWUt3UlhnR2QyQXdVT09LWmtVTkpRaHVmT0JiVkF3c0hZVUFoZEhoQmx4NXhNSDRaSjg2MC1Pbm9vVVJ4N3kyUG9FbXVrOWhWUnd4MGc9PQ==
url_encoded Z0FBQUFBQm5Lak83WGNvV2JBLW5SYm5vdjIxMG84ZXpRWkpVOHBMX0s1U1ZFZEw2cTc2eWQwV0hneFkySXlrOWpJTW9OWlJPazROWXB6NF8yOVBaRGQwSnp2aEZ5X1pBU1E1WGZFQUJxMmRoVlJ3b2h3THZndjhNemFoTlcySnRtaTR2SDdoX0U1T05pUWdfZG9FaS13dEpNbEU3UjFXNDduYldfbnhqa0x1ZlJvUkFZZHZKb252NzJSTl9mVWdkazlYUXlfTFVkWVhkRlc1ZmR5bnFpTG1JMEZ0X01FbDJ0UT09

Raw Record

{
  "text": "Been doing a research project on an article about AI and from what the devs are saying, current technologies are going to plateau HARD very soon. The issue is there are really two main ways to get data. 1) Feed it the Internet. This is great volume wise, but there is so much useless crap out there that the data needs to be sanitized so much that incremental improvements cost boat loads. There really isn’t that much upside to parsing the next 5% of the Internet‘s text for $$$$$$$$$ with extremely little upside. 2) Feed hand-picked, quality websites, from a variety of sources. The problem with this is, there really isn’t enough good data out there in the first place and it’s a PITA and not worth it to comb through the Internet to tell ChatGPT what to learn. That’s really why GPT hasn’t improved much since the 3.5 model dropped.\n\nIf development continues on better technologies, we should be a bit scared. But currently, save for better APIs that work with LLM-based models (think better implementations of Adobe Firefly, for example), there really isn’t that much room to grow with the current models (at least they don’t think there is). Expect companies like Google and OpenAI to milk current technologies so they become more and more profitable a la Google Search until a vast improvement comes down the line.",
  "label": "r/technology",
  "dataType": "comment",
  "communityName": "r/technology",
  "datetime": "2024-05-24",
  "username_encoded": "Z0FBQUFBQm5Lak1wZV92V3lZLU1rWUt3UlhnR2QyQXdVT09LWmtVTkpRaHVmT0JiVkF3c0hZVUFoZEhoQmx4NXhNSDRaSjg2MC1Pbm9vVVJ4N3kyUG9FbXVrOWhWUnd4MGc9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak83WGNvV2JBLW5SYm5vdjIxMG84ZXpRWkpVOHBMX0s1U1ZFZEw2cTc2eWQwV0hneFkySXlrOWpJTW9OWlJPazROWXB6NF8yOVBaRGQwSnp2aEZ5X1pBU1E1WGZFQUJxMmRoVlJ3b2h3THZndjhNemFoTlcySnRtaTR2SDdoX0U1T05pUWdfZG9FaS13dEpNbEU3UjFXNDduYldfbnhqa0x1ZlJvUkFZZHZKb252NzJSTl9mVWdkazlYUXlfTFVkWVhkRlc1ZmR5bnFpTG1JMEZ0X01FbDJ0UT09"
}

Entry Information