Row 90047

Row ID: 90047 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 90047 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

I recently reviewed a thesis where the author applied classic machine learning models (random forests, support vector machines, etc) to a set of equipment maintenance data. After training the system, selecting the best model and applying it to the test data, they then looked at the cases where the chosen model gave a different predicted result than the test data. In 90% of those cases, it turns out, after more careful evaluation, that the human generated labels in the data were wrong and that the ML method was actually correct. So in a sense, while it wasn't the purpose of the tool, it identified data quality problems in the source data.

FieldValue
text I recently reviewed a thesis where the author applied classic machine learning models (random forests, support vector machines, etc) to a set of equipment maintenance data. After training the system, selecting the best model and applying it to the test data, they then looked at the cases where the chosen model gave a different predicted result than the test data. In 90% of those cases, it turns out, after more careful evaluation, that the human generated labels in the data were wrong and that …
label r/datascience
dataType comment
communityName r/datascience
datetime 2024-05-25
username_encoded Z0FBQUFBQm5Lak1yR0U3NW5RdG1uXy1nTTdIbzFkbldNM2VFUk1pV2Q2X0tUODhMQ0tMZ3RYb1Fld3VTVzlXLWxYdjNEakNZbENrUFZHNVZsbWtXVlZFRV85SFRDcFdxUnc9PQ==
url_encoded Z0FBQUFBQm5Lak84Y0l3ZW10RC1UVUJha0l0c2tfVTUzcEZ3bVk4VmVnbXdfelNXNHpEcW1SS1hncEFDcndvei1nNmE3V1Jua290ZktPc3JkbWxiRzcxdEU4c0tlU3U3U1VqdmVyZFFWQXk1SG5vbmlVbGN4cmdONmJFaWd2VU5nUGUwZ28yNDJOVEk5WkxjM2ZXRFZxUjJlVHRxYnFxQzN3bjlkM2FBNElBa0lhOEdzVlpiakZuMmxPNmE4NWhFb1lncDdTWTlfTHVfRHhGOUpkZjcxZTFBM1Q0NEFpWWpoUT09

Raw Record

{
  "text": "I recently reviewed a thesis where the author applied classic machine learning models (random forests, support vector machines, etc) to a set of equipment maintenance data.  After training the system, selecting the best model and applying it to the test data, they then looked at the cases where the chosen model gave a different predicted result than the test data.  In 90% of those cases, it turns out, after more careful evaluation, that the human generated labels in the data were wrong and that the ML method was actually correct.  So in a sense, while it wasn't the purpose of the tool, it identified data quality problems in the source data.",
  "label": "r/datascience",
  "dataType": "comment",
  "communityName": "r/datascience",
  "datetime": "2024-05-25",
  "username_encoded": "Z0FBQUFBQm5Lak1yR0U3NW5RdG1uXy1nTTdIbzFkbldNM2VFUk1pV2Q2X0tUODhMQ0tMZ3RYb1Fld3VTVzlXLWxYdjNEakNZbENrUFZHNVZsbWtXVlZFRV85SFRDcFdxUnc9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak84Y0l3ZW10RC1UVUJha0l0c2tfVTUzcEZ3bVk4VmVnbXdfelNXNHpEcW1SS1hncEFDcndvei1nNmE3V1Jua290ZktPc3JkbWxiRzcxdEU4c0tlU3U3U1VqdmVyZFFWQXk1SG5vbmlVbGN4cmdONmJFaWd2VU5nUGUwZ28yNDJOVEk5WkxjM2ZXRFZxUjJlVHRxYnFxQzN3bjlkM2FBNElBa0lhOEdzVlpiakZuMmxPNmE4NWhFb1lncDdTWTlfTHVfRHhGOUpkZjcxZTFBM1Q0NEFpWWpoUT09"
}

Entry Information