Row 29977

Row ID: 29977 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 29977 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

It definitely could be done, but you'd need to collect a large amount of data first.

The problem with model distillation is that you ideally want a similarly large dataset (if not larger) than the original model was trained on.

So, imagine a company trains their FSD model with a million hours of data on their custom sensor configuration.

Now you want to distill their model, but you probably won't be able to do that with a small dataset of a thousand hours of driving data. You might have to collect hundreds of thousands of hours of data at least...

But think about it, once you have collected a million hours of your own driving sensor data, why do you even need to distill the original model anymore? You could just train your own model on your labeled dataset that you were forced to collect!

If this were a simple image processing model, it would be more feasible because you can easily scrape lots of "unlabeled" images from the internet cheaply. But for car driving data, you basically have to collect your own labeled driving data anyways so distilling another model makes less sense imo.

TL;DR - The benefit of distilling someone elses models is that you can save costs on labeling and simply use the other model as your labeler. But for driving data, you would need a human driver to drive around and collect all that data anyways so you will be collecting your own labeled data which makes the "labeler" model less valuable/useful.

FieldValue
text It definitely could be done, but you'd need to collect a large amount of data first. The problem with model distillation is that you ideally want a similarly large dataset (if not larger) than the original model was trained on. So, imagine a company trains their FSD model with a million hours of data on their custom sensor configuration. Now you want to distill their model, but you probably won't be able to do that with a small dataset of a thousand hours of driving data. You might have to co…
label r/machinelearning
dataType comment
communityName r/MachineLearning
datetime 2024-05-21
username_encoded Z0FBQUFBQm5Lak1GR2lWWk5Zb2JGQXc5TDVlaU00RHVsRXNZUDM4MDYzVkNVTFRQTGJWMGRxRTVPclNsMDNQZDRLSGg2bTNrNnNsOWdlU3VkWExvZHFmeldmSlMtQmFTZ3c9PQ==
url_encoded Z0FBQUFBQm5Lak9VdjlUMlZpbFp0SUVPY3dhXzRXb05paW04clVXUjFBOUhFcUlzd3E1VFFoNTY5a1c3RTBXeXY5WDRGRFRTSHU5NGdxX2Njc3dGdDhLblJsR1NYN2NQMnplVFZ4TGFEd1k2Z3YtUlVKdnN4eU4yaXFNMGJrV1FyMjZWbUNaTHY5MnpkTzhPRHdSdUtzTnR6enlEX2lTZUZZeFdqYmoyak5MNFhCN3NJLXp0SF9Oa2lJaTc2QUZQSjFaMFFFcjBETXdFNDg4cHZ1eU5WS1UwUjZyc3JBMDQyZ2hZa21zV2RDTGlkQlZDX3g1TGdrcz0=

Raw Record

{
  "text": "It definitely could be done, but you'd need to collect a large amount of data first.\n\nThe problem with model distillation is that you ideally want a similarly large dataset (if not larger) than the original model was trained on.\n\nSo, imagine a company trains their FSD model with a million hours of data on their custom sensor configuration.\n\nNow you want to distill their model, but you probably won't be able to do that with a small dataset of a thousand hours of driving data. You might have to collect hundreds of thousands of hours of data at least...\n\nBut think about it, once you have collected a million hours of your own driving sensor data, why do you even need to distill the original model anymore? You could just train your own model on your labeled dataset that you were forced to collect!\n\nIf this were a simple image processing model, it would be more feasible because you can easily scrape lots of \"unlabeled\" images from the internet cheaply. But for car driving data, you basically have to collect your own labeled driving data anyways so distilling another model makes less sense imo.\n\nTL;DR - The benefit of distilling someone elses models is that you can save costs on labeling and simply use the other model as your labeler. But for driving data, you would need a human driver to drive around and collect all that data anyways so you will be collecting your own labeled data which makes the \"labeler\" model less valuable/useful.",
  "label": "r/machinelearning",
  "dataType": "comment",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-21",
  "username_encoded": "Z0FBQUFBQm5Lak1GR2lWWk5Zb2JGQXc5TDVlaU00RHVsRXNZUDM4MDYzVkNVTFRQTGJWMGRxRTVPclNsMDNQZDRLSGg2bTNrNnNsOWdlU3VkWExvZHFmeldmSlMtQmFTZ3c9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9VdjlUMlZpbFp0SUVPY3dhXzRXb05paW04clVXUjFBOUhFcUlzd3E1VFFoNTY5a1c3RTBXeXY5WDRGRFRTSHU5NGdxX2Njc3dGdDhLblJsR1NYN2NQMnplVFZ4TGFEd1k2Z3YtUlVKdnN4eU4yaXFNMGJrV1FyMjZWbUNaTHY5MnpkTzhPRHdSdUtzTnR6enlEX2lTZUZZeFdqYmoyak5MNFhCN3NJLXp0SF9Oa2lJaTc2QUZQSjFaMFFFcjBETXdFNDg4cHZ1eU5WS1UwUjZyc3JBMDQyZ2hZa21zV2RDTGlkQlZDX3g1TGdrcz0="
}

Entry Information