Row 86047

Row ID: 86047 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 86047 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

Because there is no ground truth to compare the results you obtain with such method.

For low dimensional, convex, distant to each other clusters of data K-Means with elbow method itself is precise: after all, that's what he's good at. But suppose that you have such method, that for 1000000 rows of 1000 features data you can find the best K no matter what. How? How can you make sense out of an information you don't have about a domain that doesn't have a ground truth? Data can come in various forms: you can have a cluttered instance space where the difference in the underlying class is given only by small values of a single feature in a high dimensional space, or data can have missing values, can have lots of outliers, noise, and on top of that you don't even have the Target feature. Good luck with that.

FieldValue
text Because there is no ground truth to compare the results you obtain with such method. For low dimensional, convex, distant to each other clusters of data K-Means with elbow method itself is precise: after all, that's what he's good at. But suppose that you have such method, that for 1000000 rows of 1000 features data you can find the best K no matter what. How? How can you make sense out of an information you don't have about a domain that doesn't have a ground truth? Data can come in various fo…
label r/machinelearning
dataType comment
communityName r/MachineLearning
datetime 2024-05-24
username_encoded Z0FBQUFBQm5Lak1vQl9mZTJYd0NlMUMtSGp1dXA4QWtpcklDT0VPczlNOEwxRnY2Yi1ubXEwbVBCU29IYUJ4dk9UbDZ3VTJodHRHRUNpLUVzVV9zUExJck92YWd1Y3J3RFoxNWFYOU5TdU1DOWpEVXV0MVcwRVk9
url_encoded Z0FBQUFBQm5Lak82aTJIU1NUSWk0YWpVeUpISWpWYUR1N3VwZDdFMHEwQzIzbDVlYmg0emRMUlhRS1FiZ0hJaHhpOU0zOVdnRGxXMngyRUxxRFhlNHNRbThSVXhGZndVNUtZYXhkRVA0azQxeXhBNzJSME5JS2xsWm1Lb1Z0Q2RWekFhT3lYSzlmT05ReGtDbEtUVi1Rd3Q5WVdVODI2Zmhsc2NUUGg4ZXZDakUwVHhDVmNYa0lKSl9kcFFfM3NhSFR6bXM5TGRETGZiN3lGdWVoX3NkWkhxekxwcFlEYVJkUT09

Raw Record

{
  "text": "Because there is no ground truth to compare the results you obtain with such method.\n\nFor low dimensional, convex, distant to each other clusters of data K-Means with elbow method itself is precise: after all, that's what he's good at. But suppose that you have such method, that for 1000000 rows of 1000 features data you can find the best K no matter what. How? How can you make sense out of an information you don't have about a domain that doesn't have a ground truth? Data can come in various forms: you can have a cluttered instance space where the difference in the underlying class is given only by small values of a single feature in a high dimensional space, or data can have missing values, can have lots of outliers, noise, and on top of that you don't even have the Target feature. Good luck with that.",
  "label": "r/machinelearning",
  "dataType": "comment",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-24",
  "username_encoded": "Z0FBQUFBQm5Lak1vQl9mZTJYd0NlMUMtSGp1dXA4QWtpcklDT0VPczlNOEwxRnY2Yi1ubXEwbVBCU29IYUJ4dk9UbDZ3VTJodHRHRUNpLUVzVV9zUExJck92YWd1Y3J3RFoxNWFYOU5TdU1DOWpEVXV0MVcwRVk9",
  "url_encoded": "Z0FBQUFBQm5Lak82aTJIU1NUSWk0YWpVeUpISWpWYUR1N3VwZDdFMHEwQzIzbDVlYmg0emRMUlhRS1FiZ0hJaHhpOU0zOVdnRGxXMngyRUxxRFhlNHNRbThSVXhGZndVNUtZYXhkRVA0azQxeXhBNzJSME5JS2xsWm1Lb1Z0Q2RWekFhT3lYSzlmT05ReGtDbEtUVi1Rd3Q5WVdVODI2Zmhsc2NUUGg4ZXZDakUwVHhDVmNYa0lKSl9kcFFfM3NhSFR6bXM5TGRETGZiN3lGdWVoX3NkWkhxekxwcFlEYVJkUT09"
}

Entry Information