Row 95819
Content Data
This page contains data entry 95819 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
> HuggingFace, FastAI and similar frameworks are designed to lower the barrier to ML, such that any person with programming skills can harness the power of SoTA ML progress.
I strongly disagree. As someone that worked in ML for the past 10 years it feels to me that a lot of the tooling developed recently is adding more fragmentation and overhead than anything else.
- HF: why not just store models in s3/datasets and use the canonical torch/tf way of loading models/datasets? we could have invested the energy that went into HF to make those "canonical ways" that already existed 10 years ago better documented and avoid adding yet another layer.
- W&B: don't even get me started. 99% of what people use it is again, s3 and Tensorboard
- FastAI: so much great stuff in there but the reason why everyone uses Python now is because it was a nicer developer experience than other frameworks. When DS people arrived from Fortran/C (me included) Python was like whoa, I can actually understand other people's code and I am not going to spend a week figuring out what is going on
and the list goes on. To get going now you need to login into 3 different services. and understand 3 different opinionated ways of doing things. It wasn't this way, and in my experience it it harder now than it was in early Keras days to actually re-use a trained model.
bonus: if you work on niche stuff where the models, loaders etc you use always need some internal tweaking (and where business happens, not another blog post on RAG), these abstractions make life hell.
| Field | Value |
|---|---|
| text | > HuggingFace, FastAI and similar frameworks are designed to lower the barrier to ML, such that any person with programming skills can harness the power of SoTA ML progress. I strongly disagree. As someone that worked in ML for the past 10 years it feels to me that a lot of the tooling developed recently is adding more fragmentation and overhead than anything else. - HF: why not just store models in s3/datasets and use the canonical torch/tf way of loading models/datasets? we could have invest… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-25 |
| username_encoded | Z0FBQUFBQm5Lak11XzNVVFdNZlc5QlM4d09XMXlEc2VIUFlkWVhFaEVIUFV5UVZINXBqWFFwOGU4WW1mUlN1SlNnNy1BRndidWF6NmFBOXUyZ2ZDRnNsdi04VFM2d01Edmc9PQ== |
| url_encoded | Z0FBQUFBQm5LalBBbDJRNVJhSUtTTURReUhabTl6a1pheWRCRlpnZ1NuM21PUFFudXZ1US1jYmVGTzkyUEhXb0c2UExOMnd2TkZsU2dicEpSMWptN0RWeU9fenpOdFJYUnV3Y3htdm9tM1hveVJmY0VSdDBNVVhsWnB3ZVlhYlNkamdkQllWRnBRdzhUZE5FWFczS1BKQ2JiamdaMEZzNGxPMzdwWjhwcEJPYmRIbXROY2lpMjdoSVE5YzRmLXZra0ZPb1oweE0xUk5ydWtxTnF1YnRRRWhBLW5yaGxhWjdjRkw4aGpiTWNPR1ZXeHhVWllFUnQyOD0= |
Raw Record
{
"text": "> HuggingFace, FastAI and similar frameworks are designed to lower the barrier to ML, such that any person with programming skills can harness the power of SoTA ML progress.\n\nI strongly disagree. As someone that worked in ML for the past 10 years it feels to me that a lot of the tooling developed recently is adding more fragmentation and overhead than anything else.\n\n- HF: why not just store models in s3/datasets and use the canonical torch/tf way of loading models/datasets? we could have invested the energy that went into HF to make those \"canonical ways\" that already existed 10 years ago better documented and avoid adding yet another layer.\n\n \n- W&B: don't even get me started. 99% of what people use it is again, s3 and Tensorboard\n\n- FastAI: so much great stuff in there but the reason why everyone uses Python now is because it was a nicer developer experience than other frameworks. When DS people arrived from Fortran/C (me included) Python was like whoa, I can actually understand other people's code and I am not going to spend a week figuring out what is going on\n\nand the list goes on. To get going now you need to login into 3 different services. and understand 3 different opinionated ways of doing things. It wasn't this way, and in my experience it it harder now than it was in early Keras days to actually re-use a trained model. \n\n \nbonus: if you work on niche stuff where the models, loaders etc you use always need some internal tweaking (and where business happens, not another blog post on RAG), these abstractions make life hell.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-25",
"username_encoded": "Z0FBQUFBQm5Lak11XzNVVFdNZlc5QlM4d09XMXlEc2VIUFlkWVhFaEVIUFV5UVZINXBqWFFwOGU4WW1mUlN1SlNnNy1BRndidWF6NmFBOXUyZ2ZDRnNsdi04VFM2d01Edmc9PQ==",
"url_encoded": "Z0FBQUFBQm5LalBBbDJRNVJhSUtTTURReUhabTl6a1pheWRCRlpnZ1NuM21PUFFudXZ1US1jYmVGTzkyUEhXb0c2UExOMnd2TkZsU2dicEpSMWptN0RWeU9fenpOdFJYUnV3Y3htdm9tM1hveVJmY0VSdDBNVVhsWnB3ZVlhYlNkamdkQllWRnBRdzhUZE5FWFczS1BKQ2JiamdaMEZzNGxPMzdwWjhwcEJPYmRIbXROY2lpMjdoSVE5YzRmLXZra0ZPb1oweE0xUk5ydWtxTnF1YnRRRWhBLW5yaGxhWjdjRkw4aGpiTWNPR1ZXeHhVWllFUnQyOD0="
}
Entry Information
- Entry ID: 95819
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000