Row 57549
Content Data
This page contains data entry 57549 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
It really depends on the project. Usually I explain to non-ML people that there is a reward curve that looks something like log(x). We can do linear or logistic regression and get an 80% solution most of the time, at the cost of a data pipeline. We can do more work and use an appropriate off the shelf medium sized solution, like xgboost. That might get us to 87%. If that’s not good enough, that’s where things get costly. We can train a vanilla deep learning model, but we need a lot more data and compute. Maybe we get to 91%. Or we can develop something bespoke via research, which is years of time and likely a lot of data and compute and risk. That might be 94% if we hit it out of the park.
In most settings linear models or xgboost is good enough for the execs.
| Field | Value |
|---|---|
| text | It really depends on the project. Usually I explain to non-ML people that there is a reward curve that looks something like log(x). We can do linear or logistic regression and get an 80% solution most of the time, at the cost of a data pipeline. We can do more work and use an appropriate off the shelf medium sized solution, like xgboost. That might get us to 87%. If that’s not good enough, that’s where things get costly. We can train a vanilla deep learning model, but we need a lot more data and… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-23 |
| username_encoded | Z0FBQUFBQm5Lak1YRVQ3QWthNXdSc2xBdHE5QUFxZFhhek04QkhHSjBVVkNkQlFPaW5Ca1pLMzVwMmdIdTdHTFlhVUdsRERpV25hZWZLNW9hdGJ0bEY2dHZ2c3VxTnVKTk12dnVpVkdjOXRhelBnZFlrbDN6Uzg9 |
| url_encoded | Z0FBQUFBQm5Lak9tUHZMVG5uU01jdlItMVBKYm5BV1NmcGgzUjFCMmRMenJqb0NkT1dBbk90cW9kNnVnUzUxNUlBZ0VFRnBWdUUtZUNoZ2hBeVp3clF3XzZYRmViRzNLMXZXTWl5a21KNTh6N2NqT3FaVkRybUM1MHZPR09WeVNZcm9HNUlUWnl4R29vYnVNeWRMLThOazFjdjBSUXhkVUNBUVFSMFJuNG9qVEZYNXJKdXZIYTdtRkxHdXVSNDkyYmdaMDV6Qk43SHdDX3hMWHEweHkwSmVQWUFSTVRQMXhRUT09 |
Raw Record
{
"text": "It really depends on the project. Usually I explain to non-ML people that there is a reward curve that looks something like log(x). We can do linear or logistic regression and get an 80% solution most of the time, at the cost of a data pipeline. We can do more work and use an appropriate off the shelf medium sized solution, like xgboost. That might get us to 87%. If that’s not good enough, that’s where things get costly. We can train a vanilla deep learning model, but we need a lot more data and compute. Maybe we get to 91%. Or we can develop something bespoke via research, which is years of time and likely a lot of data and compute and risk. That might be 94% if we hit it out of the park.\n\nIn most settings linear models or xgboost is good enough for the execs.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-23",
"username_encoded": "Z0FBQUFBQm5Lak1YRVQ3QWthNXdSc2xBdHE5QUFxZFhhek04QkhHSjBVVkNkQlFPaW5Ca1pLMzVwMmdIdTdHTFlhVUdsRERpV25hZWZLNW9hdGJ0bEY2dHZ2c3VxTnVKTk12dnVpVkdjOXRhelBnZFlrbDN6Uzg9",
"url_encoded": "Z0FBQUFBQm5Lak9tUHZMVG5uU01jdlItMVBKYm5BV1NmcGgzUjFCMmRMenJqb0NkT1dBbk90cW9kNnVnUzUxNUlBZ0VFRnBWdUUtZUNoZ2hBeVp3clF3XzZYRmViRzNLMXZXTWl5a21KNTh6N2NqT3FaVkRybUM1MHZPR09WeVNZcm9HNUlUWnl4R29vYnVNeWRMLThOazFjdjBSUXhkVUNBUVFSMFJuNG9qVEZYNXJKdXZIYTdtRkxHdXVSNDkyYmdaMDV6Qk43SHdDX3hMWHEweHkwSmVQWUFSTVRQMXhRUT09"
}
Entry Information
- Entry ID: 57549
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000