Row 3983

Row ID: 3983 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 3983 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

I was trying to replicate results from [Grokking paper](https://arxiv.org/abs/2201.02177). As per the paper, if an over-parameterised neural net is trained beyond over-fitting, it starts generalising. I used [nanoGPT](https://github.com/karpathy/ng-video-lecture) from Andrej Karpathy for this experiment. In experiment 1 \[Grok-0\], the model started over-fitting after \~70 steps. You can see val loss \[in grey\] increasing while train loss going down to zero. However the val loss never deceased.

For experiment 2 \[Grok-1\], I increased model size \[embed dim and number of blocks\]. Surprisingly, after 70 steps **both train and val loss started increasing.**

Does anyone have a possible justification for this?

https://preview.redd.it/ndlq5vgsvorc1.png?width=2768&format=png&auto=webp&s=d5cbb6cbefaf39bc686d72334eab568639bf0201

FieldValue
text I was trying to replicate results from [Grokking paper](https://arxiv.org/abs/2201.02177). As per the paper, if an over-parameterised neural net is trained beyond over-fitting, it starts generalising. I used [nanoGPT](https://github.com/karpathy/ng-video-lecture) from Andrej Karpathy for this experiment. In experiment 1 \[Grok-0\], the model started over-fitting after \~70 steps. You can see val loss \[in grey\] increasing while train loss going down to zero. However the val loss never deceased.…
label r/pytorch
dataType post
communityName r/pytorch
datetime 2024-03-31
username_encoded Z0FBQUFBQm5LakwxZll3czQwSXFEdDRIbWhuVTlZeXlxMGdxX0hmdzRtLU1zME9pd0xUUkg4RDdnMGNYaHhOUnhhUC1WN0dBdzJuU0NHNElYN2lmeDhPSWVhT1Vud1ZhLWc9PQ==
url_encoded Z0FBQUFBQm5Lak9GenRQbTZQcXZFNmhfOEdEQm9WUi0xbnZmaDl5SEh5SGJtNG56MzVWb0RNdmZGTWdYQi1SanNyOW91N21EVXhobzJwNWY2R0IwTHBkSExXaUcyd3NjV28xUW9yOW9wN3lzMTVTYmZGOURZZXFoZmMxZXNlZnNhQVJkNUpyVVNZQTZhanBqZUU3Z2wycDdCanZWZXQ5WTRabDJ1VGYzX05sakE5UVNHYWVPRW9RPQ==

Raw Record

{
  "text": "I was trying to replicate results from [Grokking paper](https://arxiv.org/abs/2201.02177). As per the paper, if an over-parameterised neural net is trained beyond over-fitting, it starts generalising. I used [nanoGPT](https://github.com/karpathy/ng-video-lecture) from Andrej Karpathy for this experiment. In experiment 1 \\[Grok-0\\], the model started over-fitting after \\~70 steps. You can see val loss \\[in grey\\] increasing while train loss going down to zero. However the val loss never deceased.\n\nFor experiment 2 \\[Grok-1\\], I increased model size \\[embed dim and number of blocks\\]. Surprisingly, after 70 steps **both train and val loss started increasing.**\n\nDoes anyone have a possible justification  for this?\n\nhttps://preview.redd.it/ndlq5vgsvorc1.png?width=2768&format=png&auto=webp&s=d5cbb6cbefaf39bc686d72334eab568639bf0201\n\n",
  "label": "r/pytorch",
  "dataType": "post",
  "communityName": "r/pytorch",
  "datetime": "2024-03-31",
  "username_encoded": "Z0FBQUFBQm5LakwxZll3czQwSXFEdDRIbWhuVTlZeXlxMGdxX0hmdzRtLU1zME9pd0xUUkg4RDdnMGNYaHhOUnhhUC1WN0dBdzJuU0NHNElYN2lmeDhPSWVhT1Vud1ZhLWc9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9GenRQbTZQcXZFNmhfOEdEQm9WUi0xbnZmaDl5SEh5SGJtNG56MzVWb0RNdmZGTWdYQi1SanNyOW91N21EVXhobzJwNWY2R0IwTHBkSExXaUcyd3NjV28xUW9yOW9wN3lzMTVTYmZGOURZZXFoZmMxZXNlZnNhQVJkNUpyVVNZQTZhanBqZUU3Z2wycDdCanZWZXQ5WTRabDJ1VGYzX05sakE5UVNHYWVPRW9RPQ=="
}

Entry Information