Row 7900
Content Data
This page contains data entry 7900 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Recently I have been continual pre-training OpenELM-1.1B on a custom corpus w/ 2B tokens, and followed by finetuning it with custom instruction dataset w/ 4M samples.
I found that the loss (both training and eval) of PT is consistently higher than in the SFT stage when trained on similar amount of tokens, but the grad norm for SFT is higher than PT.
I assumed that the pre-trained model has a lower ppl on this custom domain, so it has a lower loss in the SFT stage. My question is, what contributes to the gradient norm being higher in the SFT stage?
| Field | Value |
|---|---|
| text | Recently I have been continual pre-training OpenELM-1.1B on a custom corpus w/ 2B tokens, and followed by finetuning it with custom instruction dataset w/ 4M samples. I found that the loss (both training and eval) of PT is consistently higher than in the SFT stage when trained on similar amount of tokens, but the grad norm for SFT is higher than PT. I assumed that the pre-trained model has a lower ppl on this custom domain, so it has a lower loss in the SFT stage. My question is, what contribu… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-18 |
| username_encoded | Z0FBQUFBQm5LakwzUkVsTWlYY1ZyQU5VQnQtMHBGamdaejlRa1dFVjRYRmpUOWtnUVVUeXNIclZsNzBrenpqd2pDT3FURnRZR2VhbGk4emxOdmJsNnF6OUl5Y01rd0Y0cUE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9ITEFfbElsbHRVaElhYUZ5TkJ0MkJkek9hczU0bXNnN21mdWhhdTBPa3lRTHBySmNWZS1MNHF5dUV4djg4YV9pU1NZUHFQNG5fbU9JZTMwR1FUMktGYVJiR1pOUGhQVnJFMjluempxXzJBeWxrQzNmZ25fdkxTZDZjZUxzcGtkUW9LRUhWT0JScmlOTE0zODRfWHFSN3BJRjlZTk9QUk0xcVpnVnZFVktfRWdUZXhqa2tLbmFrZ2w2emkwM2QzZkdrVmVmeW8tbWNCdGN1WlhWRHlUaFVVQT09 |
Raw Record
{
"text": "Recently I have been continual pre-training OpenELM-1.1B on a custom corpus w/ 2B tokens, and followed by finetuning it with custom instruction dataset w/ 4M samples.\n\nI found that the loss (both training and eval) of PT is consistently higher than in the SFT stage when trained on similar amount of tokens, but the grad norm for SFT is higher than PT.\n\nI assumed that the pre-trained model has a lower ppl on this custom domain, so it has a lower loss in the SFT stage. My question is, what contributes to the gradient norm being higher in the SFT stage?",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-18",
"username_encoded": "Z0FBQUFBQm5LakwzUkVsTWlYY1ZyQU5VQnQtMHBGamdaejlRa1dFVjRYRmpUOWtnUVVUeXNIclZsNzBrenpqd2pDT3FURnRZR2VhbGk4emxOdmJsNnF6OUl5Y01rd0Y0cUE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9ITEFfbElsbHRVaElhYUZ5TkJ0MkJkek9hczU0bXNnN21mdWhhdTBPa3lRTHBySmNWZS1MNHF5dUV4djg4YV9pU1NZUHFQNG5fbU9JZTMwR1FUMktGYVJiR1pOUGhQVnJFMjluempxXzJBeWxrQzNmZ25fdkxTZDZjZUxzcGtkUW9LRUhWT0JScmlOTE0zODRfWHFSN3BJRjlZTk9QUk0xcVpnVnZFVktfRWdUZXhqa2tLbmFrZ2w2emkwM2QzZkdrVmVmeW8tbWNCdGN1WlhWRHlUaFVVQT09"
}
Entry Information
- Entry ID: 7900
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000