Row 43947
Content Data
This page contains data entry 43947 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
For this, I'm just looking at hobbyist amounts of data. Some scraped websites, maybe some images. Few hundred gigs max.
I'm definitely more into the learning aspect of all this. Last time I dealt with "big data" was in the blockchain days, when we pumped terabytes through Kafka, mostly into ElasticSearch and Postgres.
May I ask how you're handling storage / hosting? Fast disk space is still pretty pricey on cloud, but I can tell from experience that hosting your own databases is also not very fun.
| Field | Value |
|---|---|
| text | For this, I'm just looking at hobbyist amounts of data. Some scraped websites, maybe some images. Few hundred gigs max. I'm definitely more into the learning aspect of all this. Last time I dealt with "big data" was in the blockchain days, when we pumped terabytes through Kafka, mostly into ElasticSearch and Postgres. May I ask how you're handling storage / hosting? Fast disk space is still pretty pricey on cloud, but I can tell from experience that hosting your own databases is also not very … |
| label | r/datascience |
| dataType | comment |
| communityName | r/datascience |
| datetime | 2024-05-22 |
| username_encoded | Z0FBQUFBQm5Lak1PMTlnVDJZQjJoZ2lXR1FaczAwZHpwZFYxUEE3RDFnZm5Rd0ZldHlnV3RmZHNPM2JhNm5wOUp0d2YyWnhPU1llcnJBWDRBUWx0TXdMSUFsT0pMWnNDblE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9lTnhBenRPcnBGbnMtX2RUR0ZMYldRN1dxcWlvd1ptSHI4NHVwb0dMQXc4SG5EY3FrcDB1amNERUZaS1kxLVBmcTRhRFFQTVJJQXpFMlJsS1BLNHFwTVB6ZGxydFE0bTVReHZsT3AySEhsUlVxU2dpYkRzZ242TWN3QWhhTzl2Zk5xTzlFRWwteXF0TW14aVdTR0xvZlRJWnI5QjVZZWQzNmkwekcyNEFMdWNNYU5wek1Oa2tSMWJHU3dnejBNTjBQQ0tZNGNVeVZiYzAxcnBkYUMtYkgzZz09 |
Raw Record
{
"text": "For this, I'm just looking at hobbyist amounts of data. Some scraped websites, maybe some images. Few hundred gigs max.\n\nI'm definitely more into the learning aspect of all this. Last time I dealt with \"big data\" was in the blockchain days, when we pumped terabytes through Kafka, mostly into ElasticSearch and Postgres.\n\nMay I ask how you're handling storage / hosting? Fast disk space is still pretty pricey on cloud, but I can tell from experience that hosting your own databases is also not very fun.",
"label": "r/datascience",
"dataType": "comment",
"communityName": "r/datascience",
"datetime": "2024-05-22",
"username_encoded": "Z0FBQUFBQm5Lak1PMTlnVDJZQjJoZ2lXR1FaczAwZHpwZFYxUEE3RDFnZm5Rd0ZldHlnV3RmZHNPM2JhNm5wOUp0d2YyWnhPU1llcnJBWDRBUWx0TXdMSUFsT0pMWnNDblE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9lTnhBenRPcnBGbnMtX2RUR0ZMYldRN1dxcWlvd1ptSHI4NHVwb0dMQXc4SG5EY3FrcDB1amNERUZaS1kxLVBmcTRhRFFQTVJJQXpFMlJsS1BLNHFwTVB6ZGxydFE0bTVReHZsT3AySEhsUlVxU2dpYkRzZ242TWN3QWhhTzl2Zk5xTzlFRWwteXF0TW14aVdTR0xvZlRJWnI5QjVZZWQzNmkwekcyNEFMdWNNYU5wek1Oa2tSMWJHU3dnejBNTjBQQ0tZNGNVeVZiYzAxcnBkYUMtYkgzZz09"
}
Entry Information
- Entry ID: 43947
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000