Row 32340
Content Data
This page contains data entry 32340 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
I am not an author, but I knew them.
I think your thought about GBDT's finetuning is not a "standard" setting. That means, sometimes larger search space is helpful, sometimes not. It is dataset-specific. But in papers, they have to set a kinda of "standard" consistent setting for all the datasets. They cannot adaptively select hyper-parameter tunning spaces for different datasets.
By the way, in those papers, in seems they did not directly use thousands of trees. It is upper limits. And it seems the early-stops are used in most of these approaches, FT-T, SAINT, Excelformer...
I think these papers only provide some standard approaches, exploring good inductive bias, etc. In application, you should use ensemble, feature engineering, all those tricks yourself if you're an expert. If not an expert, I personally think Excelformer maybe the best choice currently, especially you don't have computational source for hyperpara tunning. It is why I like this paper.
| Field | Value |
|---|---|
| text | I am not an author, but I knew them. I think your thought about GBDT's finetuning is not a "standard" setting. That means, sometimes larger search space is helpful, sometimes not. It is dataset-specific. But in papers, they have to set a kinda of "standard" consistent setting for all the datasets. They cannot adaptively select hyper-parameter tunning spaces for different datasets. By the way, in those papers, in seems they did not directly use thousands of trees. It is upper limits. And it see… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-21 |
| username_encoded | Z0FBQUFBQm5Lak1IUE5kM1ZSYmF6VkFWNDk0dWhDd1FpY3pjekNZUlZmVV80N2t5NnQ4cVV0RTlDaFdKYU90eUVJV1VMTjgxWElwYy1SV0wtc01QLUUyOWRxTVVGSXRMR1E9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9XdXlwVTN6My1TVEdCZFpjWFR3emxRUE8ybm5VbnJPMG5VYlFFOUdzdG9xZ1VHN1RQeDQ4bU11VW9kMm9hNlpOZWxQTHRqbXZnRUFablJaY0JJN3ZCSkxtTUdJQTFhNWlaZmQ1Ql95X2Itb20yUkE5NlRtV0hkZnhKVEJqZ29fU3dxd2EtX0xaN1dtT1RnNDJJZGZaaWJxLWJBU1dGTzFMdmQ2bVp0Q1ptOG8tNUJwSmxoOThxamNud1RlVzI0MlIzSG5WVVpaTUZ0LVJsay0ydkthTjRQVFEwN1NlYmZVMTJJQlFpOVB4YjYtRT0= |
Raw Record
{
"text": "I am not an author, but I knew them.\n\nI think your thought about GBDT's finetuning is not a \"standard\" setting. That means, sometimes larger search space is helpful, sometimes not. It is dataset-specific. But in papers, they have to set a kinda of \"standard\" consistent setting for all the datasets. They cannot adaptively select hyper-parameter tunning spaces for different datasets.\n\nBy the way, in those papers, in seems they did not directly use thousands of trees. It is upper limits. And it seems the early-stops are used in most of these approaches, FT-T, SAINT, Excelformer...\n\nI think these papers only provide some standard approaches, exploring good inductive bias, etc. In application, you should use ensemble, feature engineering, all those tricks yourself if you're an expert. If not an expert, I personally think Excelformer maybe the best choice currently, especially you don't have computational source for hyperpara tunning. It is why I like this paper.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-21",
"username_encoded": "Z0FBQUFBQm5Lak1IUE5kM1ZSYmF6VkFWNDk0dWhDd1FpY3pjekNZUlZmVV80N2t5NnQ4cVV0RTlDaFdKYU90eUVJV1VMTjgxWElwYy1SV0wtc01QLUUyOWRxTVVGSXRMR1E9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9XdXlwVTN6My1TVEdCZFpjWFR3emxRUE8ybm5VbnJPMG5VYlFFOUdzdG9xZ1VHN1RQeDQ4bU11VW9kMm9hNlpOZWxQTHRqbXZnRUFablJaY0JJN3ZCSkxtTUdJQTFhNWlaZmQ1Ql95X2Itb20yUkE5NlRtV0hkZnhKVEJqZ29fU3dxd2EtX0xaN1dtT1RnNDJJZGZaaWJxLWJBU1dGTzFMdmQ2bVp0Q1ptOG8tNUJwSmxoOThxamNud1RlVzI0MlIzSG5WVVpaTUZ0LVJsay0ydkthTjRQVFEwN1NlYmZVMTJJQlFpOVB4YjYtRT0="
}
Entry Information
- Entry ID: 32340
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000