Row 5814
Content Data
This page contains data entry 5814 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
When I read the KAN paper, I see physicists casually making fun of the uncertainties in MLPs:
- "The philosophy here is close to the mindset of physicists, who often care more about typical cases rather than worst cases" lol this went hard on NNs
- "Finite grid size can approximate the function well with a residue rate independent of the dimension, hence beating curse of dimensionality!" haha.
- "Neural scaling laws are the phenomenon where test loss decreases with more model parameters"
- "Our approach, which assumes the existence of smooth Kolmogorov Arnold representations, decomposes the high-dimensional function into several 1D functions"
Key Differences With MLPs: - Activation Functions: Unlike MLPs that use fixed activation functions at the nodes, KANs utilize learnable activation functions located on the edges between nodes. - Weight Parameters: In KANs, traditional linear weight matrices are absent. Instead, each weight parameter is replaced by a learnable univariate function, specifically a spline. - Summation Nodes: Nodes in KANs perform simple summation of incoming signals without applying non-linear transformations.
Advantages Over MLPs: - Accuracy: achieve higher accuracy with smaller network sizes compared to larger MLPs in tasks like data fitting and solving partial differential equations (PDEs). - Interpretability: Due to their unique structure, KANs are more interpretable than MLPs.
| Field | Value |
|---|---|
| text | When I read the KAN paper, I see physicists casually making fun of the uncertainties in MLPs: - "The philosophy here is close to the mindset of physicists, who often care more about typical cases rather than worst cases" lol this went hard on NNs - "Finite grid size can approximate the function well with a residue rate independent of the dimension, hence beating curse of dimensionality!" haha. - "Neural scaling laws are the phenomenon where test loss decreases with more model parameters" - "… |
| label | r/deeplearning |
| dataType | post |
| communityName | r/deeplearning |
| datetime | 2024-05-04 |
| username_encoded | Z0FBQUFBQm5LakwyT1dkR0NzTzg2WFdyYWhsZVhTeHJONkhwV2JJeWVMRG5YTmpzb1FmcTVNdFFMeVBzcDc1X1k4aGxadjBYSVNON0pYNTFRcHlQMHBjLTB4cE96U1Yzb3pKREhqZ2NaZ0JwR3ZjVEljb2pMc1k9 |
| url_encoded | Z0FBQUFBQm5Lak9HbDl4dnBDNkJ0bVdYQjNVZTVuNEt3UlpBM2tBcWxMWlB4b1RFZ1lYS2F0Qm90aXRiQmlmWVdMOFhfMUdRSndfQy1GNzA5UjJhV1F4U1lNTGNraEEtR0REZU9EOVg4LUw4YjhsbHBJTC1Da1ZURlk2VlpoLUxmNHMzUTZhaFcwY2h6NUw2bDcyaGxQQ3RLcHJrakZQaEVFWFZ4TGhoeC1zNXlmcWl2bWsxMU53ZzltYy1NZjMyMjFXOWh2SnlaZE1z |
Raw Record
{
"text": "When I read the KAN paper, I see physicists casually making fun of the uncertainties in MLPs:\n\n- \"The philosophy here is close to the mindset of physicists, who often care more about typical cases rather than worst cases\" lol this went hard on NNs\n\n- \"Finite grid size can approximate the function well with a residue rate independent of the dimension, hence beating curse of dimensionality!\" haha.\n\n- \"Neural scaling laws are the phenomenon where test loss decreases with more model parameters\"\n\n- \"Our approach, which assumes the existence of smooth Kolmogorov Arnold representations, decomposes the high-dimensional function into several 1D functions\"\n\nKey Differences With MLPs:\n- Activation Functions: Unlike MLPs that use fixed activation functions at the nodes, KANs utilize learnable activation functions located on the edges between nodes.\n- Weight Parameters: In KANs, traditional linear weight matrices are absent. Instead, each weight parameter is replaced by a learnable univariate function, specifically a spline. \n- Summation Nodes: Nodes in KANs perform simple summation of incoming signals without applying non-linear transformations.\n\nAdvantages Over MLPs:\n- Accuracy: achieve higher accuracy with smaller network sizes compared to larger MLPs in tasks like data fitting and solving partial differential equations (PDEs).\n- Interpretability: Due to their unique structure, KANs are more interpretable than MLPs.",
"label": "r/deeplearning",
"dataType": "post",
"communityName": "r/deeplearning",
"datetime": "2024-05-04",
"username_encoded": "Z0FBQUFBQm5LakwyT1dkR0NzTzg2WFdyYWhsZVhTeHJONkhwV2JJeWVMRG5YTmpzb1FmcTVNdFFMeVBzcDc1X1k4aGxadjBYSVNON0pYNTFRcHlQMHBjLTB4cE96U1Yzb3pKREhqZ2NaZ0JwR3ZjVEljb2pMc1k9",
"url_encoded": "Z0FBQUFBQm5Lak9HbDl4dnBDNkJ0bVdYQjNVZTVuNEt3UlpBM2tBcWxMWlB4b1RFZ1lYS2F0Qm90aXRiQmlmWVdMOFhfMUdRSndfQy1GNzA5UjJhV1F4U1lNTGNraEEtR0REZU9EOVg4LUw4YjhsbHBJTC1Da1ZURlk2VlpoLUxmNHMzUTZhaFcwY2h6NUw2bDcyaGxQQ3RLcHJrakZQaEVFWFZ4TGhoeC1zNXlmcWl2bWsxMU53ZzltYy1NZjMyMjFXOWh2SnlaZE1z"
}
Entry Information
- Entry ID: 5814
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000