Row 90047
Content Data
This page contains data entry 90047 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
I recently reviewed a thesis where the author applied classic machine learning models (random forests, support vector machines, etc) to a set of equipment maintenance data. After training the system, selecting the best model and applying it to the test data, they then looked at the cases where the chosen model gave a different predicted result than the test data. In 90% of those cases, it turns out, after more careful evaluation, that the human generated labels in the data were wrong and that the ML method was actually correct. So in a sense, while it wasn't the purpose of the tool, it identified data quality problems in the source data.
| Field | Value |
|---|---|
| text | I recently reviewed a thesis where the author applied classic machine learning models (random forests, support vector machines, etc) to a set of equipment maintenance data. After training the system, selecting the best model and applying it to the test data, they then looked at the cases where the chosen model gave a different predicted result than the test data. In 90% of those cases, it turns out, after more careful evaluation, that the human generated labels in the data were wrong and that … |
| label | r/datascience |
| dataType | comment |
| communityName | r/datascience |
| datetime | 2024-05-25 |
| username_encoded | Z0FBQUFBQm5Lak1yR0U3NW5RdG1uXy1nTTdIbzFkbldNM2VFUk1pV2Q2X0tUODhMQ0tMZ3RYb1Fld3VTVzlXLWxYdjNEakNZbENrUFZHNVZsbWtXVlZFRV85SFRDcFdxUnc9PQ== |
| url_encoded | Z0FBQUFBQm5Lak84Y0l3ZW10RC1UVUJha0l0c2tfVTUzcEZ3bVk4VmVnbXdfelNXNHpEcW1SS1hncEFDcndvei1nNmE3V1Jua290ZktPc3JkbWxiRzcxdEU4c0tlU3U3U1VqdmVyZFFWQXk1SG5vbmlVbGN4cmdONmJFaWd2VU5nUGUwZ28yNDJOVEk5WkxjM2ZXRFZxUjJlVHRxYnFxQzN3bjlkM2FBNElBa0lhOEdzVlpiakZuMmxPNmE4NWhFb1lncDdTWTlfTHVfRHhGOUpkZjcxZTFBM1Q0NEFpWWpoUT09 |
Raw Record
{
"text": "I recently reviewed a thesis where the author applied classic machine learning models (random forests, support vector machines, etc) to a set of equipment maintenance data. After training the system, selecting the best model and applying it to the test data, they then looked at the cases where the chosen model gave a different predicted result than the test data. In 90% of those cases, it turns out, after more careful evaluation, that the human generated labels in the data were wrong and that the ML method was actually correct. So in a sense, while it wasn't the purpose of the tool, it identified data quality problems in the source data.",
"label": "r/datascience",
"dataType": "comment",
"communityName": "r/datascience",
"datetime": "2024-05-25",
"username_encoded": "Z0FBQUFBQm5Lak1yR0U3NW5RdG1uXy1nTTdIbzFkbldNM2VFUk1pV2Q2X0tUODhMQ0tMZ3RYb1Fld3VTVzlXLWxYdjNEakNZbENrUFZHNVZsbWtXVlZFRV85SFRDcFdxUnc9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak84Y0l3ZW10RC1UVUJha0l0c2tfVTUzcEZ3bVk4VmVnbXdfelNXNHpEcW1SS1hncEFDcndvei1nNmE3V1Jua290ZktPc3JkbWxiRzcxdEU4c0tlU3U3U1VqdmVyZFFWQXk1SG5vbmlVbGN4cmdONmJFaWd2VU5nUGUwZ28yNDJOVEk5WkxjM2ZXRFZxUjJlVHRxYnFxQzN3bjlkM2FBNElBa0lhOEdzVlpiakZuMmxPNmE4NWhFb1lncDdTWTlfTHVfRHhGOUpkZjcxZTFBM1Q0NEFpWWpoUT09"
}
Entry Information
- Entry ID: 90047
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000