Row 84127
Content Data
This page contains data entry 84127 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
I think a lot of analysts already works with petabytes of data in Spark/Airflow, so as a DA my data-engineering skillset was already pretty overdeveloped for optimizing pipelines and writing performant SQL.
> Basic stats but you really have to know it well.
This is the crux of it though. I definitely want to know testing/sampling methodologies inside-out so that I can make a recommendation explaining each step of the analysis and why specific choices were made.
| Field | Value |
|---|---|
| text | I think a lot of analysts already works with petabytes of data in Spark/Airflow, so as a DA my data-engineering skillset was already pretty overdeveloped for optimizing pipelines and writing performant SQL. > Basic stats but you really have to know it well. This is the crux of it though. I definitely want to know testing/sampling methodologies inside-out so that I can make a recommendation explaining each step of the analysis and why specific choices were made. |
| label | r/datascience |
| dataType | comment |
| communityName | r/datascience |
| datetime | 2024-05-24 |
| username_encoded | Z0FBQUFBQm5Lak1uQVctMHFuQ2Q2VTZja1NnT05POGZnVV8zRlBReUFkbHd5Tnd1NU1zZ1ZKbXZkN3B1aGtiZG9zY0w3X1Y4TVZHMnJiLUc2eTE5ZFdmR3gtam5mZmh5MEE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak80UE1OOGVtSFAwcjVjZTNYb3A3MlFMbWQtaFNHVDR1NXJhcFY1VFhFVmlwZUowYVM1bkZmS0RkWDBpaUpNUzZWODFadnozYjkwRkpTTVVOSnpCWkRFTHdXaU9sYTBiWC1ZZkpicGZkRWFWZWpUWXdIb2g4ZzZIYU5CUkItcUZpbWVmZEpLQVNQX1RSX2NSdF8xb2RFbDUtQkl4WURDRmw2b0xqenZzY2w1RlJENVQ3NlVuVWQtVUdqSnBmZVVmcHVCNllVd1diYkczVV9WVDhRVG1jc1pOcDFCNDkxUW9vSmJPQjJXMll6V1N6ND0= |
Raw Record
{
"text": "I think a lot of analysts already works with petabytes of data in Spark/Airflow, so as a DA my data-engineering skillset was already pretty overdeveloped for optimizing pipelines and writing performant SQL. \n\n> Basic stats but you really have to know it well.\n\nThis is the crux of it though. I definitely want to know testing/sampling methodologies inside-out so that I can make a recommendation explaining each step of the analysis and why specific choices were made.",
"label": "r/datascience",
"dataType": "comment",
"communityName": "r/datascience",
"datetime": "2024-05-24",
"username_encoded": "Z0FBQUFBQm5Lak1uQVctMHFuQ2Q2VTZja1NnT05POGZnVV8zRlBReUFkbHd5Tnd1NU1zZ1ZKbXZkN3B1aGtiZG9zY0w3X1Y4TVZHMnJiLUc2eTE5ZFdmR3gtam5mZmh5MEE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak80UE1OOGVtSFAwcjVjZTNYb3A3MlFMbWQtaFNHVDR1NXJhcFY1VFhFVmlwZUowYVM1bkZmS0RkWDBpaUpNUzZWODFadnozYjkwRkpTTVVOSnpCWkRFTHdXaU9sYTBiWC1ZZkpicGZkRWFWZWpUWXdIb2g4ZzZIYU5CUkItcUZpbWVmZEpLQVNQX1RSX2NSdF8xb2RFbDUtQkl4WURDRmw2b0xqenZzY2w1RlJENVQ3NlVuVWQtVUdqSnBmZVVmcHVCNllVd1diYkczVV9WVDhRVG1jc1pOcDFCNDkxUW9vSmJPQjJXMll6V1N6ND0="
}
Entry Information
- Entry ID: 84127
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000