Row 79617
Content Data
This page contains data entry 79617 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
it's called the elbow method because you look at the elbow of the curve and focus on that interval to find your best one.
Now you have this interval, suppose is [x:y], both integers and x<=y. In my experience, I usually don't stop there. Next steps could include: 1) use Silhouette technique to narrow the range. Silhouette converges to the optimal range of K 2) decide qualitatively which K fits your purposes, higher K leads to better WSS but it also cause fragmentation and worse generalization (I usually choose a value k<= (x+y)//2 ) 3) test it with different parameters now that your have a reduced range of possibile Ks, in order to find better results
I'll tell you that tho, 19 features... I usually prefer DBSCAN since it works better with non convex clusters, and searching for the right eps and min_samples is usually easier. It works better with outliers also!
| Field | Value |
|---|---|
| text | it's called the elbow method because you look at the elbow of the curve and focus on that interval to find your best one. Now you have this interval, suppose is [x:y], both integers and x<=y. In my experience, I usually don't stop there. Next steps could include: 1) use Silhouette technique to narrow the range. Silhouette converges to the optimal range of K 2) decide qualitatively which K fits your purposes, higher K leads to better WSS but it also cause fragmentation and worse generalization … |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-24 |
| username_encoded | Z0FBQUFBQm5Lak1rZVBCY0htOUx0LUp6LVZaaThHMWkwR2ZOYk83ZVJJbVdLaGNmVWVEYnc0MF92RUxrWmVEeHlWUkg2QTlaUjV5ak1CZlgzN21PdnNXSUhZY2tPeFVDVjNiRGhIcTdqV0dEMjRwazZtT1J0Nkk9 |
| url_encoded | Z0FBQUFBQm5Lak8xU1hFc1Y0T3FtSUJOSEZoTHdDbjBTMUV2ZDZqNDVvbXZCaTd6eHhMR1hlM3NFYURBOEZPczZoaHpkcjhTcXVwanJGMU4zMno0SUVBb0lrVTJKdUZ2ZWpZYUoybXEyUXVkcWpLUHhXVHFWSmFxU2M0NkVMdVZ2X1Z3Tk1iNV9UNzA5cHljNkRrY01SbHNudkY1MVZ3eVk5ZGNpY3ZxYzFiTm1haVNRTGtXdThoTFRGRUpMQlFZT2ZtOS1LRTdQSUhlcFFhOU8zSF85UjlMR0VER0kzaHRaQT09 |
Raw Record
{
"text": "it's called the elbow method because you look at the elbow of the curve and focus on that interval to find your best one.\n\nNow you have this interval, suppose is [x:y], both integers and x<=y. \nIn my experience, I usually don't stop there. Next steps could include:\n1) use Silhouette technique to narrow the range. Silhouette converges to the optimal range of K\n2) decide qualitatively which K fits your purposes, higher K leads to better WSS but it also cause fragmentation and worse generalization (I usually choose a value k<= (x+y)//2 )\n3) test it with different parameters now that your have a reduced range of possibile Ks, in order to find better results\n\nI'll tell you that tho, 19 features... I usually prefer DBSCAN since it works better with non convex clusters, and searching for the right eps and min_samples is usually easier. It works better with outliers also!",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-24",
"username_encoded": "Z0FBQUFBQm5Lak1rZVBCY0htOUx0LUp6LVZaaThHMWkwR2ZOYk83ZVJJbVdLaGNmVWVEYnc0MF92RUxrWmVEeHlWUkg2QTlaUjV5ak1CZlgzN21PdnNXSUhZY2tPeFVDVjNiRGhIcTdqV0dEMjRwazZtT1J0Nkk9",
"url_encoded": "Z0FBQUFBQm5Lak8xU1hFc1Y0T3FtSUJOSEZoTHdDbjBTMUV2ZDZqNDVvbXZCaTd6eHhMR1hlM3NFYURBOEZPczZoaHpkcjhTcXVwanJGMU4zMno0SUVBb0lrVTJKdUZ2ZWpZYUoybXEyUXVkcWpLUHhXVHFWSmFxU2M0NkVMdVZ2X1Z3Tk1iNV9UNzA5cHljNkRrY01SbHNudkY1MVZ3eVk5ZGNpY3ZxYzFiTm1haVNRTGtXdThoTFRGRUpMQlFZT2ZtOS1LRTdQSUhlcFFhOU8zSF85UjlMR0VER0kzaHRaQT09"
}
Entry Information
- Entry ID: 79617
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000