Row 3983
Content Data
This page contains data entry 3983 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
I was trying to replicate results from [Grokking paper](https://arxiv.org/abs/2201.02177). As per the paper, if an over-parameterised neural net is trained beyond over-fitting, it starts generalising. I used [nanoGPT](https://github.com/karpathy/ng-video-lecture) from Andrej Karpathy for this experiment. In experiment 1 \[Grok-0\], the model started over-fitting after \~70 steps. You can see val loss \[in grey\] increasing while train loss going down to zero. However the val loss never deceased.
For experiment 2 \[Grok-1\], I increased model size \[embed dim and number of blocks\]. Surprisingly, after 70 steps **both train and val loss started increasing.**
Does anyone have a possible justification for this?
https://preview.redd.it/ndlq5vgsvorc1.png?width=2768&format=png&auto=webp&s=d5cbb6cbefaf39bc686d72334eab568639bf0201
| Field | Value |
|---|---|
| text | I was trying to replicate results from [Grokking paper](https://arxiv.org/abs/2201.02177). As per the paper, if an over-parameterised neural net is trained beyond over-fitting, it starts generalising. I used [nanoGPT](https://github.com/karpathy/ng-video-lecture) from Andrej Karpathy for this experiment. In experiment 1 \[Grok-0\], the model started over-fitting after \~70 steps. You can see val loss \[in grey\] increasing while train loss going down to zero. However the val loss never deceased.… |
| label | r/pytorch |
| dataType | post |
| communityName | r/pytorch |
| datetime | 2024-03-31 |
| username_encoded | Z0FBQUFBQm5LakwxZll3czQwSXFEdDRIbWhuVTlZeXlxMGdxX0hmdzRtLU1zME9pd0xUUkg4RDdnMGNYaHhOUnhhUC1WN0dBdzJuU0NHNElYN2lmeDhPSWVhT1Vud1ZhLWc9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9GenRQbTZQcXZFNmhfOEdEQm9WUi0xbnZmaDl5SEh5SGJtNG56MzVWb0RNdmZGTWdYQi1SanNyOW91N21EVXhobzJwNWY2R0IwTHBkSExXaUcyd3NjV28xUW9yOW9wN3lzMTVTYmZGOURZZXFoZmMxZXNlZnNhQVJkNUpyVVNZQTZhanBqZUU3Z2wycDdCanZWZXQ5WTRabDJ1VGYzX05sakE5UVNHYWVPRW9RPQ== |
Raw Record
{
"text": "I was trying to replicate results from [Grokking paper](https://arxiv.org/abs/2201.02177). As per the paper, if an over-parameterised neural net is trained beyond over-fitting, it starts generalising. I used [nanoGPT](https://github.com/karpathy/ng-video-lecture) from Andrej Karpathy for this experiment. In experiment 1 \\[Grok-0\\], the model started over-fitting after \\~70 steps. You can see val loss \\[in grey\\] increasing while train loss going down to zero. However the val loss never deceased.\n\nFor experiment 2 \\[Grok-1\\], I increased model size \\[embed dim and number of blocks\\]. Surprisingly, after 70 steps **both train and val loss started increasing.**\n\nDoes anyone have a possible justification for this?\n\nhttps://preview.redd.it/ndlq5vgsvorc1.png?width=2768&format=png&auto=webp&s=d5cbb6cbefaf39bc686d72334eab568639bf0201\n\n",
"label": "r/pytorch",
"dataType": "post",
"communityName": "r/pytorch",
"datetime": "2024-03-31",
"username_encoded": "Z0FBQUFBQm5LakwxZll3czQwSXFEdDRIbWhuVTlZeXlxMGdxX0hmdzRtLU1zME9pd0xUUkg4RDdnMGNYaHhOUnhhUC1WN0dBdzJuU0NHNElYN2lmeDhPSWVhT1Vud1ZhLWc9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9GenRQbTZQcXZFNmhfOEdEQm9WUi0xbnZmaDl5SEh5SGJtNG56MzVWb0RNdmZGTWdYQi1SanNyOW91N21EVXhobzJwNWY2R0IwTHBkSExXaUcyd3NjV28xUW9yOW9wN3lzMTVTYmZGOURZZXFoZmMxZXNlZnNhQVJkNUpyVVNZQTZhanBqZUU3Z2wycDdCanZWZXQ5WTRabDJ1VGYzX05sakE5UVNHYWVPRW9RPQ=="
}
Entry Information
- Entry ID: 3983
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000