Row 7900

Row ID: 7900 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 7900 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

Recently I have been continual pre-training OpenELM-1.1B on a custom corpus w/ 2B tokens, and followed by finetuning it with custom instruction dataset w/ 4M samples.

I found that the loss (both training and eval) of PT is consistently higher than in the SFT stage when trained on similar amount of tokens, but the grad norm for SFT is higher than PT.

I assumed that the pre-trained model has a lower ppl on this custom domain, so it has a lower loss in the SFT stage. My question is, what contributes to the gradient norm being higher in the SFT stage?

FieldValue
text Recently I have been continual pre-training OpenELM-1.1B on a custom corpus w/ 2B tokens, and followed by finetuning it with custom instruction dataset w/ 4M samples. I found that the loss (both training and eval) of PT is consistently higher than in the SFT stage when trained on similar amount of tokens, but the grad norm for SFT is higher than PT. I assumed that the pre-trained model has a lower ppl on this custom domain, so it has a lower loss in the SFT stage. My question is, what contribu…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-05-18
username_encoded Z0FBQUFBQm5LakwzUkVsTWlYY1ZyQU5VQnQtMHBGamdaejlRa1dFVjRYRmpUOWtnUVVUeXNIclZsNzBrenpqd2pDT3FURnRZR2VhbGk4emxOdmJsNnF6OUl5Y01rd0Y0cUE9PQ==
url_encoded Z0FBQUFBQm5Lak9ITEFfbElsbHRVaElhYUZ5TkJ0MkJkek9hczU0bXNnN21mdWhhdTBPa3lRTHBySmNWZS1MNHF5dUV4djg4YV9pU1NZUHFQNG5fbU9JZTMwR1FUMktGYVJiR1pOUGhQVnJFMjluempxXzJBeWxrQzNmZ25fdkxTZDZjZUxzcGtkUW9LRUhWT0JScmlOTE0zODRfWHFSN3BJRjlZTk9QUk0xcVpnVnZFVktfRWdUZXhqa2tLbmFrZ2w2emkwM2QzZkdrVmVmeW8tbWNCdGN1WlhWRHlUaFVVQT09

Raw Record

{
  "text": "Recently I have been continual pre-training OpenELM-1.1B on a custom corpus w/ 2B tokens, and followed by finetuning it with custom instruction dataset w/ 4M samples.\n\nI found that the loss (both training and eval) of PT is consistently higher than in the SFT stage when trained on similar amount of tokens, but the grad norm for SFT is higher than PT.\n\nI assumed that the pre-trained model has a lower ppl on this custom domain, so it has a lower loss in the SFT stage. My question is, what contributes to the gradient norm being higher in the SFT stage?",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-18",
  "username_encoded": "Z0FBQUFBQm5LakwzUkVsTWlYY1ZyQU5VQnQtMHBGamdaejlRa1dFVjRYRmpUOWtnUVVUeXNIclZsNzBrenpqd2pDT3FURnRZR2VhbGk4emxOdmJsNnF6OUl5Y01rd0Y0cUE9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9ITEFfbElsbHRVaElhYUZ5TkJ0MkJkek9hczU0bXNnN21mdWhhdTBPa3lRTHBySmNWZS1MNHF5dUV4djg4YV9pU1NZUHFQNG5fbU9JZTMwR1FUMktGYVJiR1pOUGhQVnJFMjluempxXzJBeWxrQzNmZ25fdkxTZDZjZUxzcGtkUW9LRUhWT0JScmlOTE0zODRfWHFSN3BJRjlZTk9QUk0xcVpnVnZFVktfRWdUZXhqa2tLbmFrZ2w2emkwM2QzZkdrVmVmeW8tbWNCdGN1WlhWRHlUaFVVQT09"
}

Entry Information