Row 66316
Content Data
This page contains data entry 66316 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Hello, I’ve been learning about ml and dl for about 6 months now and have found a “family tree” of machine learning models to aid in my learning. I have a pretty solid understanding up of what (and how and why) things were done in the pre-transformer era, however…
With transformers, typically they increase in performance as data size grows and model size scales. I am curious as to what and how the proportions of transformer blocks (in language modeling (encoder, decoder), vit) such as embed depth, n_layers, and n_heads effect the model.
1. what ratios are ideal / common in terms of model scaling? I may want a 100k param model 2. beyond GPT-2 what architectural changes have (phi, gemma, llama) made to the actual model? Should I use a GPT2 like? PHI-like? llama-like?
| Field | Value |
|---|---|
| text | Hello, I’ve been learning about ml and dl for about 6 months now and have found a “family tree” of machine learning models to aid in my learning. I have a pretty solid understanding up of what (and how and why) things were done in the pre-transformer era, however… With transformers, typically they increase in performance as data size grows and model size scales. I am curious as to what and how the proportions of transformer blocks (in language modeling (encoder, decoder), vit) such as embed dep… |
| label | r/deeplearning |
| dataType | post |
| communityName | r/deeplearning |
| datetime | 2024-05-23 |
| username_encoded | Z0FBQUFBQm5Lak1jUFJ0eU9vVkpwQ2V1bkYwNC1zQXpubHpGaEpRcDF4ZXhndWlIV3dMUXYtMTg4ZE45S3NXYU5LdXVsWl9iTkRBQ0dDOFdMZ1gtdEhrS3RFU2g5WHpMaFE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak90cW1kZVAybDFoTXgwVElzU3JpeVZtREZib21qNGl2S0R3akxHMC02QzBWUWRSSUljYlpWS1JZNXdBVVM3MlZrUWdkZWVQZDhaelg4MjdtTFh4RzQ3N0tfaFo1alBmZXRiNzBKQ0ZwNTdDeUJfUTZSbmdQS0RZNnVOdkttUmMtTWFvQklROHlzOVlVNXVWcXIyYW1YTDJJVmpqelQwVWdUd1lYbFZOd2wtUFl2M2VFWFAxY25jWlpCb1BFYUlJN21CS050b25reU1QVnlmc1kybG1hOFdGdz09 |
Raw Record
{
"text": "Hello, I’ve been learning about ml and dl for about 6 months now and have found a “family tree” of machine learning models to aid in my learning. I have a pretty solid understanding up of what (and how and why) things were done in the pre-transformer era, however…\n\nWith transformers, typically they increase in performance as data size grows and model size scales. I am curious as to what and how the proportions of transformer blocks (in language modeling (encoder, decoder), vit) such as embed depth, n_layers, and n_heads effect the model. \n\n1. what ratios are ideal / common in terms of model scaling?\n I may want a 100k param model\n2. beyond GPT-2 what architectural changes have (phi, gemma, llama) made to the actual model?\n Should I use a GPT2 like? PHI-like? llama-like? ",
"label": "r/deeplearning",
"dataType": "post",
"communityName": "r/deeplearning",
"datetime": "2024-05-23",
"username_encoded": "Z0FBQUFBQm5Lak1jUFJ0eU9vVkpwQ2V1bkYwNC1zQXpubHpGaEpRcDF4ZXhndWlIV3dMUXYtMTg4ZE45S3NXYU5LdXVsWl9iTkRBQ0dDOFdMZ1gtdEhrS3RFU2g5WHpMaFE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak90cW1kZVAybDFoTXgwVElzU3JpeVZtREZib21qNGl2S0R3akxHMC02QzBWUWRSSUljYlpWS1JZNXdBVVM3MlZrUWdkZWVQZDhaelg4MjdtTFh4RzQ3N0tfaFo1alBmZXRiNzBKQ0ZwNTdDeUJfUTZSbmdQS0RZNnVOdkttUmMtTWFvQklROHlzOVlVNXVWcXIyYW1YTDJJVmpqelQwVWdUd1lYbFZOd2wtUFl2M2VFWFAxY25jWlpCb1BFYUlJN21CS050b25reU1QVnlmc1kybG1hOFdGdz09"
}
Entry Information
- Entry ID: 66316
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000