Row 25273

Row ID: 25273 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 25273 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

I think they all generally suck and are overrated. Where their value is however that they have useful embeddings (don't cite me all anecdotal evidence).What this allows is an easy combination of time series and tabular data as well as training xgboost models, which are quite good for tabular use cases with a decent amount of samples.

I would actually love to see even smaller models with less embedding dimensions (and possibly even worse accuracy), so that I could pair them up with models that excel at truly low sample settings, like the GP. Sadly, these often scale very poorly with increasing dimensionality so the number of currently used embedding dimensions is often way too high for this combo.

In any case, I think the space of time series problems does not have a clean and small manifold as the language problems so I don't think it is possible to build truly well performing models with the current architecture/compute resources.

FieldValue
text I think they all generally suck and are overrated. Where their value is however that they have useful embeddings (don't cite me all anecdotal evidence).What this allows is an easy combination of time series and tabular data as well as training xgboost models, which are quite good for tabular use cases with a decent amount of samples. I would actually love to see even smaller models with less embedding dimensions (and possibly even worse accuracy), so that I could pair them up with models that …
label r/machinelearning
dataType comment
communityName r/MachineLearning
datetime 2024-05-21
username_encoded Z0FBQUFBQm5Lak1DSGlKVGpCajNycXBpRjVoOVR5MG5kaU55emVJYUljWHd5bjMxbVpLRU5NUmt4Nk54NXprdjJnQWxfd0xPU1dXbHNOZXk1MFRSRkJra0VPMjZnQzlmekE9PQ==
url_encoded Z0FBQUFBQm5Lak9TSlg0R21rVGwxTFZWVlhKYTNyREIxV0lLM2JHeF9HaFUzU1dLRnk3OGlTZzlOS2pXYmlYano0N2VsUWZzakpXUkE4NWhWUWE5cG5EbW1hdUtieXhnLWFia3VPam81Sy1xTlRrbUlYWk1CQk5VWDFDZEg5T3l0U0tIV0J0Y003ZVlLMFUxNGdDTlJKdHpJbmJvN0JqMjZpbFgzMUpGM0QzeF9PQlJJZHE5Z2UxTlZOazZ5bHFuMF8ydDd3TUVlTHpCTVlDemQxYUg4VlJvOVc4WENNVEYwZz09

Raw Record

{
  "text": "I think they all generally suck and are overrated. Where their value is however that they have useful embeddings (don't cite me all anecdotal evidence).What this allows is an easy combination of time series and tabular data as well as training xgboost models, which are quite good for tabular use cases with a decent amount of samples. \n\nI would actually love to see even smaller models with less embedding dimensions (and possibly even worse accuracy), so that I could pair them up with models that excel at truly low sample settings, like the GP. Sadly, these often scale very poorly with increasing dimensionality so the number of currently used embedding dimensions is often way too high for this combo.\n\nIn any case, I think the space of time series problems does not have a clean and small manifold as the language problems so I don't think it is possible to build truly well performing models with the current architecture/compute resources.",
  "label": "r/machinelearning",
  "dataType": "comment",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-21",
  "username_encoded": "Z0FBQUFBQm5Lak1DSGlKVGpCajNycXBpRjVoOVR5MG5kaU55emVJYUljWHd5bjMxbVpLRU5NUmt4Nk54NXprdjJnQWxfd0xPU1dXbHNOZXk1MFRSRkJra0VPMjZnQzlmekE9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9TSlg0R21rVGwxTFZWVlhKYTNyREIxV0lLM2JHeF9HaFUzU1dLRnk3OGlTZzlOS2pXYmlYano0N2VsUWZzakpXUkE4NWhWUWE5cG5EbW1hdUtieXhnLWFia3VPam81Sy1xTlRrbUlYWk1CQk5VWDFDZEg5T3l0U0tIV0J0Y003ZVlLMFUxNGdDTlJKdHpJbmJvN0JqMjZpbFgzMUpGM0QzeF9PQlJJZHE5Z2UxTlZOazZ5bHFuMF8ydDd3TUVlTHpCTVlDemQxYUg4VlJvOVc4WENNVEYwZz09"
}

Entry Information