Row 92736
Content Data
This page contains data entry 92736 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
If you take the cumulative graph you'll see that after 7/9 the "area" below the part you selected is very small compared to the total area below the graph. This means that 7/9 clusters would only explain a small portion of the total variability of the data. More simply : you lose a lot of information by flattening the space from 36000 to 9.
Depending on what you want to do with the clusters, losing a lot of information might be a non-issue, or it can be a dealbreaker.
| Field | Value |
|---|---|
| text | If you take the cumulative graph you'll see that after 7/9 the "area" below the part you selected is very small compared to the total area below the graph. This means that 7/9 clusters would only explain a small portion of the total variability of the data. More simply : you lose a lot of information by flattening the space from 36000 to 9. Depending on what you want to do with the clusters, losing a lot of information might be a non-issue, or it can be a dealbreaker. |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-25 |
| username_encoded | Z0FBQUFBQm5Lak1zTWFyalRoU3BIblV1Y1FHclpvMFJDZWhMUXpjSjF0amgxWjROSWQtU2t5UGppSWhHcFlwdU5EcHBHeUg4a3RJMldHUXIwcVJ2Z1Zaa0pEVWRzblRhREE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak8tNXNzdURKZk8zYkluRFl0VUE2VWZZYjhua0NqMjQxc0J6bnRhYUF6SmRINHI2S1pLQ2k5LVlnamhqemw2S2dFMDlWRzF0MkxnWkt2M2xuZ3dzQWthNEp4N1o1T3JibTdFZGxfOFZTaWk4cVU0cEN1MzdWTTFlSmttVk1JZUEwSngwQmJoV09aZ05qQlJCczcwcjhrUW1WQmJzZVdPR2JyVVFtZVVWVlN5TkQwYk1od3NnZW5HelVIVWtMbnZOMVhXTlFHU3VsdS1YNkhMekltZ3M4b0dzQT09 |
Raw Record
{
"text": "If you take the cumulative graph you'll see that after 7/9 the \"area\" below the part you selected is very small compared to the total area below the graph. This means that 7/9 clusters would only explain a small portion of the total variability of the data. More simply : you lose a lot of information by flattening the space from 36000 to 9. \n\nDepending on what you want to do with the clusters, losing a lot of information might be a non-issue, or it can be a dealbreaker.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-25",
"username_encoded": "Z0FBQUFBQm5Lak1zTWFyalRoU3BIblV1Y1FHclpvMFJDZWhMUXpjSjF0amgxWjROSWQtU2t5UGppSWhHcFlwdU5EcHBHeUg4a3RJMldHUXIwcVJ2Z1Zaa0pEVWRzblRhREE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak8tNXNzdURKZk8zYkluRFl0VUE2VWZZYjhua0NqMjQxc0J6bnRhYUF6SmRINHI2S1pLQ2k5LVlnamhqemw2S2dFMDlWRzF0MkxnWkt2M2xuZ3dzQWthNEp4N1o1T3JibTdFZGxfOFZTaWk4cVU0cEN1MzdWTTFlSmttVk1JZUEwSngwQmJoV09aZ05qQlJCczcwcjhrUW1WQmJzZVdPR2JyVVFtZVVWVlN5TkQwYk1od3NnZW5HelVIVWtMbnZOMVhXTlFHU3VsdS1YNkhMekltZ3M4b0dzQT09"
}
Entry Information
- Entry ID: 92736
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000