Row 66316

Row ID: 66316 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 66316 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

Hello, I’ve been learning about ml and dl for about 6 months now and have found a “family tree” of machine learning models to aid in my learning. I have a pretty solid understanding up of what (and how and why) things were done in the pre-transformer era, however…

With transformers, typically they increase in performance as data size grows and model size scales. I am curious as to what and how the proportions of transformer blocks (in language modeling (encoder, decoder), vit) such as embed depth, n_layers, and n_heads effect the model.

1. what ratios are ideal / common in terms of model scaling? I may want a 100k param model 2. beyond GPT-2 what architectural changes have (phi, gemma, llama) made to the actual model? Should I use a GPT2 like? PHI-like? llama-like?

FieldValue
text Hello, I’ve been learning about ml and dl for about 6 months now and have found a “family tree” of machine learning models to aid in my learning. I have a pretty solid understanding up of what (and how and why) things were done in the pre-transformer era, however… With transformers, typically they increase in performance as data size grows and model size scales. I am curious as to what and how the proportions of transformer blocks (in language modeling (encoder, decoder), vit) such as embed dep…
label r/deeplearning
dataType post
communityName r/deeplearning
datetime 2024-05-23
username_encoded Z0FBQUFBQm5Lak1jUFJ0eU9vVkpwQ2V1bkYwNC1zQXpubHpGaEpRcDF4ZXhndWlIV3dMUXYtMTg4ZE45S3NXYU5LdXVsWl9iTkRBQ0dDOFdMZ1gtdEhrS3RFU2g5WHpMaFE9PQ==
url_encoded Z0FBQUFBQm5Lak90cW1kZVAybDFoTXgwVElzU3JpeVZtREZib21qNGl2S0R3akxHMC02QzBWUWRSSUljYlpWS1JZNXdBVVM3MlZrUWdkZWVQZDhaelg4MjdtTFh4RzQ3N0tfaFo1alBmZXRiNzBKQ0ZwNTdDeUJfUTZSbmdQS0RZNnVOdkttUmMtTWFvQklROHlzOVlVNXVWcXIyYW1YTDJJVmpqelQwVWdUd1lYbFZOd2wtUFl2M2VFWFAxY25jWlpCb1BFYUlJN21CS050b25reU1QVnlmc1kybG1hOFdGdz09

Raw Record

{
  "text": "Hello, I’ve been learning about ml and dl for about 6 months now and have found a “family tree” of machine learning models to aid in my learning. I have a pretty solid understanding up of what (and how and why) things were done in the pre-transformer era, however…\n\nWith transformers, typically they increase in performance as data size grows and model size scales. I am curious as to what and how the proportions of transformer blocks (in language modeling (encoder, decoder), vit) such as embed depth, n_layers, and n_heads effect the model. \n\n1. what ratios are ideal / common in terms of model scaling?\n            I may want a 100k param model\n2. beyond GPT-2 what architectural changes have (phi, gemma, llama) made to the actual model?\n            Should I use a GPT2 like? PHI-like? llama-like? ",
  "label": "r/deeplearning",
  "dataType": "post",
  "communityName": "r/deeplearning",
  "datetime": "2024-05-23",
  "username_encoded": "Z0FBQUFBQm5Lak1jUFJ0eU9vVkpwQ2V1bkYwNC1zQXpubHpGaEpRcDF4ZXhndWlIV3dMUXYtMTg4ZE45S3NXYU5LdXVsWl9iTkRBQ0dDOFdMZ1gtdEhrS3RFU2g5WHpMaFE9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak90cW1kZVAybDFoTXgwVElzU3JpeVZtREZib21qNGl2S0R3akxHMC02QzBWUWRSSUljYlpWS1JZNXdBVVM3MlZrUWdkZWVQZDhaelg4MjdtTFh4RzQ3N0tfaFo1alBmZXRiNzBKQ0ZwNTdDeUJfUTZSbmdQS0RZNnVOdkttUmMtTWFvQklROHlzOVlVNXVWcXIyYW1YTDJJVmpqelQwVWdUd1lYbFZOd2wtUFl2M2VFWFAxY25jWlpCb1BFYUlJN21CS050b25reU1QVnlmc1kybG1hOFdGdz09"
}

Entry Information