Row 86047
Content Data
This page contains data entry 86047 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Because there is no ground truth to compare the results you obtain with such method.
For low dimensional, convex, distant to each other clusters of data K-Means with elbow method itself is precise: after all, that's what he's good at. But suppose that you have such method, that for 1000000 rows of 1000 features data you can find the best K no matter what. How? How can you make sense out of an information you don't have about a domain that doesn't have a ground truth? Data can come in various forms: you can have a cluttered instance space where the difference in the underlying class is given only by small values of a single feature in a high dimensional space, or data can have missing values, can have lots of outliers, noise, and on top of that you don't even have the Target feature. Good luck with that.
| Field | Value |
|---|---|
| text | Because there is no ground truth to compare the results you obtain with such method. For low dimensional, convex, distant to each other clusters of data K-Means with elbow method itself is precise: after all, that's what he's good at. But suppose that you have such method, that for 1000000 rows of 1000 features data you can find the best K no matter what. How? How can you make sense out of an information you don't have about a domain that doesn't have a ground truth? Data can come in various fo… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-24 |
| username_encoded | Z0FBQUFBQm5Lak1vQl9mZTJYd0NlMUMtSGp1dXA4QWtpcklDT0VPczlNOEwxRnY2Yi1ubXEwbVBCU29IYUJ4dk9UbDZ3VTJodHRHRUNpLUVzVV9zUExJck92YWd1Y3J3RFoxNWFYOU5TdU1DOWpEVXV0MVcwRVk9 |
| url_encoded | Z0FBQUFBQm5Lak82aTJIU1NUSWk0YWpVeUpISWpWYUR1N3VwZDdFMHEwQzIzbDVlYmg0emRMUlhRS1FiZ0hJaHhpOU0zOVdnRGxXMngyRUxxRFhlNHNRbThSVXhGZndVNUtZYXhkRVA0azQxeXhBNzJSME5JS2xsWm1Lb1Z0Q2RWekFhT3lYSzlmT05ReGtDbEtUVi1Rd3Q5WVdVODI2Zmhsc2NUUGg4ZXZDakUwVHhDVmNYa0lKSl9kcFFfM3NhSFR6bXM5TGRETGZiN3lGdWVoX3NkWkhxekxwcFlEYVJkUT09 |
Raw Record
{
"text": "Because there is no ground truth to compare the results you obtain with such method.\n\nFor low dimensional, convex, distant to each other clusters of data K-Means with elbow method itself is precise: after all, that's what he's good at. But suppose that you have such method, that for 1000000 rows of 1000 features data you can find the best K no matter what. How? How can you make sense out of an information you don't have about a domain that doesn't have a ground truth? Data can come in various forms: you can have a cluttered instance space where the difference in the underlying class is given only by small values of a single feature in a high dimensional space, or data can have missing values, can have lots of outliers, noise, and on top of that you don't even have the Target feature. Good luck with that.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-24",
"username_encoded": "Z0FBQUFBQm5Lak1vQl9mZTJYd0NlMUMtSGp1dXA4QWtpcklDT0VPczlNOEwxRnY2Yi1ubXEwbVBCU29IYUJ4dk9UbDZ3VTJodHRHRUNpLUVzVV9zUExJck92YWd1Y3J3RFoxNWFYOU5TdU1DOWpEVXV0MVcwRVk9",
"url_encoded": "Z0FBQUFBQm5Lak82aTJIU1NUSWk0YWpVeUpISWpWYUR1N3VwZDdFMHEwQzIzbDVlYmg0emRMUlhRS1FiZ0hJaHhpOU0zOVdnRGxXMngyRUxxRFhlNHNRbThSVXhGZndVNUtZYXhkRVA0azQxeXhBNzJSME5JS2xsWm1Lb1Z0Q2RWekFhT3lYSzlmT05ReGtDbEtUVi1Rd3Q5WVdVODI2Zmhsc2NUUGg4ZXZDakUwVHhDVmNYa0lKSl9kcFFfM3NhSFR6bXM5TGRETGZiN3lGdWVoX3NkWkhxekxwcFlEYVJkUT09"
}
Entry Information
- Entry ID: 86047
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000