Row 17044
Content Data
This page contains data entry 17044 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
It depends on what you consider ML. Statistical techniques have had a huge impact on the efficacy of medical treatment, and the line between "statistics" and "ML" is pretty blurry. The first randomized controlled trial was published in 1948, and less sophisticated statistical analysis has been used in medicine for a lot longer than that.
However, I must take issue with your statement that ML is "unbiased". Compared to what? All models have have assumptions baked in, if only because of imperfect training data.
For a more specific example, I read a paper about using convolutional neural networks to detect lung cancer in CT scans. They described the ML models as being "hypothesis free", which is nonsense.
They trained the models with CT scans from hospital EHR data. Think about this. Someone at the hospital thought it was a good idea for these patients to get CT scans, and that's a hypothesis. Its not like they had 50 % random healthy controls who got scans just for the hell of it, and 50% patients with known malignant tumors. I'm not saying it made the data worthless, but its hardly "hypothesis free".
I could go on: who are the people who get CT scans? In America, they are people with access to good medical care. They tend to be rich and disproportionately white. That's going to be baked into the training data too.
As they say, all models are wrong, but some of them are useful.
| Field | Value |
|---|---|
| text | It depends on what you consider ML. Statistical techniques have had a huge impact on the efficacy of medical treatment, and the line between "statistics" and "ML" is pretty blurry. The first randomized controlled trial was published in 1948, and less sophisticated statistical analysis has been used in medicine for a lot longer than that. However, I must take issue with your statement that ML is "unbiased". Compared to what? All models have have assumptions baked in, if only because of imperfect… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-20 |
| username_encoded | Z0FBQUFBQm5Lakw5enJTN25RRmpfQldXaE9nU0VwMHE5Y0E5VnJ6WGJhMmU4VUJuVVJpZXM5QzVLc1FoY01NOFBLcUwyckdSUlNsRzRsaHhpcGhGNGZlVDBwclJqbEdyWGc9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9NTGRqSWNUN1JOS1lOaHlZY1Z6WFFQT0tZMXlucy1xRk02bFpZVzRhc1FzRll2M2FzQTE1ZUVjUnYySl9KSkNXdk5wZl9LNVlmRjkwUEJYQmczWW9jVXRBY2J5aWVGU2NoQmN6NWF1Z2NTR3RXU2x0QVZIOUZrdkViOEI0dWxRUU5oN1Y1eXlEMERHaFFGc3FCU1dIVV9JVjJuVWdYdzJaM2c1WU1TZ3hSNWFlQnoxMTlLWXYwaTV1SmVOWl9YbzctelNKWDQ3OHhUWkd5YmhDWUxxbm1SVkh0UW5LRFNvcU9tc2tqeHpqdVFYZz0= |
Raw Record
{
"text": "It depends on what you consider ML. Statistical techniques have had a huge impact on the efficacy of medical treatment, and the line between \"statistics\" and \"ML\" is pretty blurry. The first randomized controlled trial was published in 1948, and less sophisticated statistical analysis has been used in medicine for a lot longer than that.\n\nHowever, I must take issue with your statement that ML is \"unbiased\". Compared to what? All models have have assumptions baked in, if only because of imperfect training data.\n\nFor a more specific example, I read a paper about using convolutional neural networks to detect lung cancer in CT scans. They described the ML models as being \"hypothesis free\", which is nonsense.\n\nThey trained the models with CT scans from hospital EHR data. Think about this. Someone at the hospital thought it was a good idea for these patients to get CT scans, and that's a hypothesis. Its not like they had 50 % random healthy controls who got scans just for the hell of it, and 50% patients with known malignant tumors. I'm not saying it made the data worthless, but its hardly \"hypothesis free\".\n\nI could go on: who are the people who get CT scans? In America, they are people with access to good medical care. They tend to be rich and disproportionately white. That's going to be baked into the training data too.\n\nAs they say, all models are wrong, but some of them are useful.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-20",
"username_encoded": "Z0FBQUFBQm5Lakw5enJTN25RRmpfQldXaE9nU0VwMHE5Y0E5VnJ6WGJhMmU4VUJuVVJpZXM5QzVLc1FoY01NOFBLcUwyckdSUlNsRzRsaHhpcGhGNGZlVDBwclJqbEdyWGc9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9NTGRqSWNUN1JOS1lOaHlZY1Z6WFFQT0tZMXlucy1xRk02bFpZVzRhc1FzRll2M2FzQTE1ZUVjUnYySl9KSkNXdk5wZl9LNVlmRjkwUEJYQmczWW9jVXRBY2J5aWVGU2NoQmN6NWF1Z2NTR3RXU2x0QVZIOUZrdkViOEI0dWxRUU5oN1Y1eXlEMERHaFFGc3FCU1dIVV9JVjJuVWdYdzJaM2c1WU1TZ3hSNWFlQnoxMTlLWXYwaTV1SmVOWl9YbzctelNKWDQ3OHhUWkd5YmhDWUxxbm1SVkh0UW5LRFNvcU9tc2tqeHpqdVFYZz0="
}
Entry Information
- Entry ID: 17044
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000