Row 88203
Content Data
This page contains data entry 88203 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
The reason why clustering is hard is there is no ground truth. Unsupervised learning is much harder than supervised because there's no non-subjective evaluation criteria.
What's the business(real) problem you are trying to solve? Is it marketing to these clusters? Are the smallest clusters reliable?
Have you looked at data elements near the center of each cluster? Is it "easy" to determine why those central points are different?
Can you label some data with the clusterID then predict the clusterID using any ML classification method? Can you predict cluster==X vs not X with ML?
With clustering it's usually better to decide how to decide first, or at a minimum getting input on the maximum number allowable. Get mgmt to agree on the criteria and then do that effort and stop. Of course you may wind up tweaking the rules a bit once you see the data, but then that conversation is a bit easier as you need to decide if that lift is worth breaking the established rules for picking the clusters.
| Field | Value |
|---|---|
| text | The reason why clustering is hard is there is no ground truth. Unsupervised learning is much harder than supervised because there's no non-subjective evaluation criteria. What's the business(real) problem you are trying to solve? Is it marketing to these clusters? Are the smallest clusters reliable? Have you looked at data elements near the center of each cluster? Is it "easy" to determine why those central points are different? Can you label some data with the clusterID then predict the … |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-24 |
| username_encoded | Z0FBQUFBQm5Lak1xRnNoX3c5UElBRE1wbU1ER2taWFdNd2lyQldETEhYOHVsdGlhTW9KR1N3TXZpcGxYaTdXMlpzM2haWExDY3o1MjZpV3FRRzF2ODI5dTJ5VzdlT1ExVkE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak83b0Zia0lwTF9SVFl1am55c3F3YjhURXNFX1RzcGhPYV9OMDhUeEc1RTBLWFEzaTc0WGU4QnQwSG1kUG4yOVhhQmJGcXdfOV9ZbWZIa2c0cTBMTVRYWDM2OGNUd2JBUnV0ZGxGMW15M1dLNjNPM2F5M0NrNVU1YkgzbUtWRTZfS3lLMlV5TTloTUUwS0g5bFJOQzJaSnkxY1RlRmRUa0IyR2prSEZWN0luNTZSako1VUFaOVp6UUpTeHc1OHJDR0d0WXlYUEcyU1ZYSGRPWm90UlBpczd2dz09 |
Raw Record
{
"text": "The reason why clustering is hard is there is no ground truth. Unsupervised learning is much harder than supervised because there's no non-subjective evaluation criteria.\n\nWhat's the business(real) problem you are trying to solve? Is it marketing to these clusters? Are the smallest clusters reliable?\n\nHave you looked at data elements near the center of each cluster? Is it \"easy\" to determine why those central points are different?\n\nCan you label some data with the clusterID then predict the clusterID using any ML classification method? Can you predict cluster==X vs not X with ML?\n\nWith clustering it's usually better to decide how to decide first, or at a minimum getting input on the maximum number allowable. Get mgmt to agree on the criteria and then do that effort and stop. Of course you may wind up tweaking the rules a bit once you see the data, but then that conversation is a bit easier as you need to decide if that lift is worth breaking the established rules for picking the clusters.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-24",
"username_encoded": "Z0FBQUFBQm5Lak1xRnNoX3c5UElBRE1wbU1ER2taWFdNd2lyQldETEhYOHVsdGlhTW9KR1N3TXZpcGxYaTdXMlpzM2haWExDY3o1MjZpV3FRRzF2ODI5dTJ5VzdlT1ExVkE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak83b0Zia0lwTF9SVFl1am55c3F3YjhURXNFX1RzcGhPYV9OMDhUeEc1RTBLWFEzaTc0WGU4QnQwSG1kUG4yOVhhQmJGcXdfOV9ZbWZIa2c0cTBMTVRYWDM2OGNUd2JBUnV0ZGxGMW15M1dLNjNPM2F5M0NrNVU1YkgzbUtWRTZfS3lLMlV5TTloTUUwS0g5bFJOQzJaSnkxY1RlRmRUa0IyR2prSEZWN0luNTZSako1VUFaOVp6UUpTeHc1OHJDR0d0WXlYUEcyU1ZYSGRPWm90UlBpczd2dz09"
}
Entry Information
- Entry ID: 88203
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000