Row 79617

Row ID: 79617 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 79617 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

it's called the elbow method because you look at the elbow of the curve and focus on that interval to find your best one.

Now you have this interval, suppose is [x:y], both integers and x<=y. In my experience, I usually don't stop there. Next steps could include: 1) use Silhouette technique to narrow the range. Silhouette converges to the optimal range of K 2) decide qualitatively which K fits your purposes, higher K leads to better WSS but it also cause fragmentation and worse generalization (I usually choose a value k<= (x+y)//2 ) 3) test it with different parameters now that your have a reduced range of possibile Ks, in order to find better results

I'll tell you that tho, 19 features... I usually prefer DBSCAN since it works better with non convex clusters, and searching for the right eps and min_samples is usually easier. It works better with outliers also!

FieldValue
text it's called the elbow method because you look at the elbow of the curve and focus on that interval to find your best one. Now you have this interval, suppose is [x:y], both integers and x<=y. In my experience, I usually don't stop there. Next steps could include: 1) use Silhouette technique to narrow the range. Silhouette converges to the optimal range of K 2) decide qualitatively which K fits your purposes, higher K leads to better WSS but it also cause fragmentation and worse generalization …
label r/machinelearning
dataType comment
communityName r/MachineLearning
datetime 2024-05-24
username_encoded Z0FBQUFBQm5Lak1rZVBCY0htOUx0LUp6LVZaaThHMWkwR2ZOYk83ZVJJbVdLaGNmVWVEYnc0MF92RUxrWmVEeHlWUkg2QTlaUjV5ak1CZlgzN21PdnNXSUhZY2tPeFVDVjNiRGhIcTdqV0dEMjRwazZtT1J0Nkk9
url_encoded Z0FBQUFBQm5Lak8xU1hFc1Y0T3FtSUJOSEZoTHdDbjBTMUV2ZDZqNDVvbXZCaTd6eHhMR1hlM3NFYURBOEZPczZoaHpkcjhTcXVwanJGMU4zMno0SUVBb0lrVTJKdUZ2ZWpZYUoybXEyUXVkcWpLUHhXVHFWSmFxU2M0NkVMdVZ2X1Z3Tk1iNV9UNzA5cHljNkRrY01SbHNudkY1MVZ3eVk5ZGNpY3ZxYzFiTm1haVNRTGtXdThoTFRGRUpMQlFZT2ZtOS1LRTdQSUhlcFFhOU8zSF85UjlMR0VER0kzaHRaQT09

Raw Record

{
  "text": "it's called the elbow method because you look at the elbow of the curve and focus on that interval to find your best one.\n\nNow you have this interval, suppose is [x:y], both integers and x<=y. \nIn my experience, I usually don't stop there. Next steps could include:\n1) use Silhouette technique to narrow the range. Silhouette converges to the optimal range of K\n2) decide qualitatively which K fits your purposes, higher K leads to better WSS but it also cause fragmentation and worse generalization (I usually choose a value k<= (x+y)//2 )\n3) test it with different parameters now that your have a reduced range of possibile Ks, in order to find better results\n\nI'll tell you that tho, 19 features... I usually prefer DBSCAN since it works better with non convex clusters, and searching for the right eps and min_samples is usually easier. It works better with outliers also!",
  "label": "r/machinelearning",
  "dataType": "comment",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-24",
  "username_encoded": "Z0FBQUFBQm5Lak1rZVBCY0htOUx0LUp6LVZaaThHMWkwR2ZOYk83ZVJJbVdLaGNmVWVEYnc0MF92RUxrWmVEeHlWUkg2QTlaUjV5ak1CZlgzN21PdnNXSUhZY2tPeFVDVjNiRGhIcTdqV0dEMjRwazZtT1J0Nkk9",
  "url_encoded": "Z0FBQUFBQm5Lak8xU1hFc1Y0T3FtSUJOSEZoTHdDbjBTMUV2ZDZqNDVvbXZCaTd6eHhMR1hlM3NFYURBOEZPczZoaHpkcjhTcXVwanJGMU4zMno0SUVBb0lrVTJKdUZ2ZWpZYUoybXEyUXVkcWpLUHhXVHFWSmFxU2M0NkVMdVZ2X1Z3Tk1iNV9UNzA5cHljNkRrY01SbHNudkY1MVZ3eVk5ZGNpY3ZxYzFiTm1haVNRTGtXdThoTFRGRUpMQlFZT2ZtOS1LRTdQSUhlcFFhOU8zSF85UjlMR0VER0kzaHRaQT09"
}

Entry Information