Row 78155
Content Data
This page contains data entry 78155 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
There is a big difference between scraping web sites and getting access to a cataloged and tagged archive. This is a much cleaner and more extensive dataset to use for training, that alone will skew the distribution.
In addition, they referenced bringing realtime content feeds from the “masthead” directly to users, not just including it in training data.
Quit with your “provide evidence” nonsense, we don’t even have good tools for linking model responses to which parts of the training data influenced that response, so what you’re asking for is not possible.
| Field | Value |
|---|---|
| text | There is a big difference between scraping web sites and getting access to a cataloged and tagged archive. This is a much cleaner and more extensive dataset to use for training, that alone will skew the distribution. In addition, they referenced bringing realtime content feeds from the “masthead” directly to users, not just including it in training data. Quit with your “provide evidence” nonsense, we don’t even have good tools for linking model responses to which parts of the training data i… |
| label | r/openai |
| dataType | comment |
| communityName | r/OpenAI |
| datetime | 2024-05-24 |
| username_encoded | Z0FBQUFBQm5Lak1qckpkbjJPM1VpbkxnX1k2RWZvalBCR3RQZ1RXT1dWclgzZkI1U28welhkTjdZTW00dklFd1c0b00zYjZsend6SEVlblZIT29NYURHZ3RjbXh2d1hrcERqUmJYcUxlR3lMQjlGaFpPNEhpZE09 |
| url_encoded | Z0FBQUFBQm5Lak8wUmtWXzNodWVZYkZNaDVpVmV5VUR2YWVOUGRrdU1yNW5rOWdWUFptckZrRjJBS1kwMXpGR21CRlMwQno1VkZObUNtZmw4U2pkVUxNcGlvU2d3Yjd2a1EwZ1JHb1EwNVQ1YXY4RktweEdEZW1CdjRrMGpiWnpjQWV1N3BYYVZXTHhBMkdxbnRtTXVnVjluQW9lcWF4d0V1Z1FMX3B1ZTIzdm4tMGZzd1FwV3AyMkxoQWEwTmVjUFZOaHMzc2hjLVNvY3NtSmVLTnE2UEMtQW5qczR3amlzUT09 |
Raw Record
{
"text": "There is a big difference between scraping web sites and getting access to a cataloged and tagged archive. This is a much cleaner and more extensive dataset to use for training, that alone will skew the distribution. \n\nIn addition, they referenced bringing realtime content feeds from the “masthead” directly to users, not just including it in training data. \n\nQuit with your “provide evidence” nonsense, we don’t even have good tools for linking model responses to which parts of the training data influenced that response, so what you’re asking for is not possible.",
"label": "r/openai",
"dataType": "comment",
"communityName": "r/OpenAI",
"datetime": "2024-05-24",
"username_encoded": "Z0FBQUFBQm5Lak1qckpkbjJPM1VpbkxnX1k2RWZvalBCR3RQZ1RXT1dWclgzZkI1U28welhkTjdZTW00dklFd1c0b00zYjZsend6SEVlblZIT29NYURHZ3RjbXh2d1hrcERqUmJYcUxlR3lMQjlGaFpPNEhpZE09",
"url_encoded": "Z0FBQUFBQm5Lak8wUmtWXzNodWVZYkZNaDVpVmV5VUR2YWVOUGRrdU1yNW5rOWdWUFptckZrRjJBS1kwMXpGR21CRlMwQno1VkZObUNtZmw4U2pkVUxNcGlvU2d3Yjd2a1EwZ1JHb1EwNVQ1YXY4RktweEdEZW1CdjRrMGpiWnpjQWV1N3BYYVZXTHhBMkdxbnRtTXVnVjluQW9lcWF4d0V1Z1FMX3B1ZTIzdm4tMGZzd1FwV3AyMkxoQWEwTmVjUFZOaHMzc2hjLVNvY3NtSmVLTnE2UEMtQW5qczR3amlzUT09"
}
Entry Information
- Entry ID: 78155
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000