Row 25273
Content Data
This page contains data entry 25273 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
I think they all generally suck and are overrated. Where their value is however that they have useful embeddings (don't cite me all anecdotal evidence).What this allows is an easy combination of time series and tabular data as well as training xgboost models, which are quite good for tabular use cases with a decent amount of samples.
I would actually love to see even smaller models with less embedding dimensions (and possibly even worse accuracy), so that I could pair them up with models that excel at truly low sample settings, like the GP. Sadly, these often scale very poorly with increasing dimensionality so the number of currently used embedding dimensions is often way too high for this combo.
In any case, I think the space of time series problems does not have a clean and small manifold as the language problems so I don't think it is possible to build truly well performing models with the current architecture/compute resources.
| Field | Value |
|---|---|
| text | I think they all generally suck and are overrated. Where their value is however that they have useful embeddings (don't cite me all anecdotal evidence).What this allows is an easy combination of time series and tabular data as well as training xgboost models, which are quite good for tabular use cases with a decent amount of samples. I would actually love to see even smaller models with less embedding dimensions (and possibly even worse accuracy), so that I could pair them up with models that … |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-21 |
| username_encoded | Z0FBQUFBQm5Lak1DSGlKVGpCajNycXBpRjVoOVR5MG5kaU55emVJYUljWHd5bjMxbVpLRU5NUmt4Nk54NXprdjJnQWxfd0xPU1dXbHNOZXk1MFRSRkJra0VPMjZnQzlmekE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9TSlg0R21rVGwxTFZWVlhKYTNyREIxV0lLM2JHeF9HaFUzU1dLRnk3OGlTZzlOS2pXYmlYano0N2VsUWZzakpXUkE4NWhWUWE5cG5EbW1hdUtieXhnLWFia3VPam81Sy1xTlRrbUlYWk1CQk5VWDFDZEg5T3l0U0tIV0J0Y003ZVlLMFUxNGdDTlJKdHpJbmJvN0JqMjZpbFgzMUpGM0QzeF9PQlJJZHE5Z2UxTlZOazZ5bHFuMF8ydDd3TUVlTHpCTVlDemQxYUg4VlJvOVc4WENNVEYwZz09 |
Raw Record
{
"text": "I think they all generally suck and are overrated. Where their value is however that they have useful embeddings (don't cite me all anecdotal evidence).What this allows is an easy combination of time series and tabular data as well as training xgboost models, which are quite good for tabular use cases with a decent amount of samples. \n\nI would actually love to see even smaller models with less embedding dimensions (and possibly even worse accuracy), so that I could pair them up with models that excel at truly low sample settings, like the GP. Sadly, these often scale very poorly with increasing dimensionality so the number of currently used embedding dimensions is often way too high for this combo.\n\nIn any case, I think the space of time series problems does not have a clean and small manifold as the language problems so I don't think it is possible to build truly well performing models with the current architecture/compute resources.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-21",
"username_encoded": "Z0FBQUFBQm5Lak1DSGlKVGpCajNycXBpRjVoOVR5MG5kaU55emVJYUljWHd5bjMxbVpLRU5NUmt4Nk54NXprdjJnQWxfd0xPU1dXbHNOZXk1MFRSRkJra0VPMjZnQzlmekE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9TSlg0R21rVGwxTFZWVlhKYTNyREIxV0lLM2JHeF9HaFUzU1dLRnk3OGlTZzlOS2pXYmlYano0N2VsUWZzakpXUkE4NWhWUWE5cG5EbW1hdUtieXhnLWFia3VPam81Sy1xTlRrbUlYWk1CQk5VWDFDZEg5T3l0U0tIV0J0Y003ZVlLMFUxNGdDTlJKdHpJbmJvN0JqMjZpbFgzMUpGM0QzeF9PQlJJZHE5Z2UxTlZOazZ5bHFuMF8ydDd3TUVlTHpCTVlDemQxYUg4VlJvOVc4WENNVEYwZz09"
}
Entry Information
- Entry ID: 25273
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000