Row 85774

Row ID: 85774 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 85774 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

As I said, silhouette converges to the optimal range of K, and the results obtained make sense: there was still a steady descent w.r.t. the WSSE in the elbow graph. So yess, one could say around 90 to 100 is your best K. But The fact no matter what you do you can never know the exact best K for the task, and that's the caveat with unsupervised techniques: they cannot classify data, rather let the human *explore* the data.

What's next, you might ask:

Next part is, in my humble opinion, to stop looking for answers in the model and start analyzing the data itself: What's the domain of the data? What are the scenarios we want to cover with the K clusters? E.g.: I have this sensor that magically retrieves features about flowers like size, petal length/width, weight, but not pictures: without them I cannot deploy a human to label this data. But I know that this sensor is in multiple flower shops of New York and the sensors are acquiring data around Valentine's Day, I can take an educated guess and speculate that we're expecting a large cluster (the red roses) and 10-15 more clusters about other flowers (can be close to the large one, and those are the white or black roses, or can be really distant like the generic flowers for a wedding). Outliers must be treated: I bet that some niche flower shop has these crazy exotic flowers that no other shop has, and those are going to have a cluster on their own or worse they are too distant to the rest of the element of the clusters they have been labeled to.

The model itself did its job and can only get you so far: you have your k and you have your nicely cohesive clusters. But the best K? No one knows it, that's the point of *unsupervised* learning

FieldValue
text As I said, silhouette converges to the optimal range of K, and the results obtained make sense: there was still a steady descent w.r.t. the WSSE in the elbow graph. So yess, one could say around 90 to 100 is your best K. But The fact no matter what you do you can never know the exact best K for the task, and that's the caveat with unsupervised techniques: they cannot classify data, rather let the human *explore* the data. What's next, you might ask: Next part is, in my humble opinion, to st…
label r/machinelearning
dataType comment
communityName r/MachineLearning
datetime 2024-05-24
username_encoded Z0FBQUFBQm5Lak1vMXR5T0dvcWpPMUl5UFZvVmN6SWtjUFd0N1RUb2Y1aE5tTFdxZjY1YmhuWmdzWHJSYjQtRlg3UFhJU2Vya2E5TjAxdDNkeDhuYXR1YmVuMzZvaWoyVFdLNm1ESG5HcU5pR3gzMV9DWkNpbkU9
url_encoded Z0FBQUFBQm5Lak81c3VaRnByeE9KcFZtazk3Uk01d0FJMlM5OFJBVUhyV1AyblJ0Z0VSRVBCNWlGRHMtRGc3WE1HY01lYTNUY3hMQlB1SmdyT0dxcWU1ajdyNnBLM2pXZzE3VE9XakVwSS1CaHlUUVlTU0dZRjVqcUJSaHpCNFgwM0ZFVWMtSGNJMGo2c0JIVHh1OWhtQV92NGhqVmN4R2drOTZWUWpMY0pzZXR0TFVUV0NIU09jTzlFa1k5WTQwbUJYdTZ2OGNJTGZUcWhBcldrYlY4SEtjLW1hVFNaRUFfUT09

Raw Record

{
  "text": "As I said, silhouette converges to the optimal range of K, and the results obtained make sense: there was still a steady descent w.r.t. the WSSE in the elbow graph. So yess, one could say around 90 to 100 is your best K. But The fact no matter what you do you can never know the exact best K for the task, and that's the caveat with unsupervised techniques: they cannot classify data, rather let the human *explore* the data.\n\n  \nWhat's next, you might ask:\n\nNext part is, in my humble opinion, to stop looking for answers in the model and start analyzing the data itself: What's the domain of the data? What are the scenarios we want to cover with the K clusters? E.g.: I have this sensor that magically retrieves features about flowers like size, petal length/width, weight, but not pictures: without them I cannot deploy a human to label this data. But I know that this sensor is in multiple flower shops of New York and the sensors are acquiring data around Valentine's Day, I can take an educated guess and speculate that we're expecting a large cluster (the red roses) and 10-15 more clusters about other flowers (can be close to the large one, and those are the white or black roses, or can be really distant like the generic flowers for a wedding). Outliers must be treated: I bet that some niche flower shop has these crazy exotic flowers that no other shop has, and those are going to have a cluster on their own or worse they are too distant to the rest of the element of the clusters they have been labeled to.\n\n  \nThe model itself did its job and can only get you so far: you have your k and you have your nicely cohesive clusters. But the best K? No one knows it, that's the point of *unsupervised* learning",
  "label": "r/machinelearning",
  "dataType": "comment",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-24",
  "username_encoded": "Z0FBQUFBQm5Lak1vMXR5T0dvcWpPMUl5UFZvVmN6SWtjUFd0N1RUb2Y1aE5tTFdxZjY1YmhuWmdzWHJSYjQtRlg3UFhJU2Vya2E5TjAxdDNkeDhuYXR1YmVuMzZvaWoyVFdLNm1ESG5HcU5pR3gzMV9DWkNpbkU9",
  "url_encoded": "Z0FBQUFBQm5Lak81c3VaRnByeE9KcFZtazk3Uk01d0FJMlM5OFJBVUhyV1AyblJ0Z0VSRVBCNWlGRHMtRGc3WE1HY01lYTNUY3hMQlB1SmdyT0dxcWU1ajdyNnBLM2pXZzE3VE9XakVwSS1CaHlUUVlTU0dZRjVqcUJSaHpCNFgwM0ZFVWMtSGNJMGo2c0JIVHh1OWhtQV92NGhqVmN4R2drOTZWUWpMY0pzZXR0TFVUV0NIU09jTzlFa1k5WTQwbUJYdTZ2OGNJTGZUcWhBcldrYlY4SEtjLW1hVFNaRUFfUT09"
}

Entry Information