Row 6093

Row ID: 6093 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 6093 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

I’ve read that inference speed for models like Llama-2 70B is ~10 t/s at best. So that left me wondering how the extremely large models like GPT-4 (1T params?) do their fast 20 t/s inference. With 10x the params, they gotta have at least 3x the layers(?) So that should make its inference much slower. Am I missing anything? What kind of further improvements might these companies be doing to power their fast APIs?

Edit: I must mention that you cannot parallelize across GPUs to help with latency of a single example when the data has to pass through model layers sequentially.

And with the large model sizes, model parallelism, with its inter-GPU communication should make it even slower…

FieldValue
text I’ve read that inference speed for models like Llama-2 70B is ~10 t/s at best. So that left me wondering how the extremely large models like GPT-4 (1T params?) do their fast 20 t/s inference. With 10x the params, they gotta have at least 3x the layers(?) So that should make its inference much slower. Am I missing anything? What kind of further improvements might these companies be doing to power their fast APIs? Edit: I must mention that you cannot parallelize across GPUs to help with latency o…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-05-07
username_encoded Z0FBQUFBQm5LakwyRmk5Z29IRmk2Z2VIWnlvSFljd0FDN290OU11Tk9YYVZkdmtURkgzdEdseVE3WERybTJkYk1pdXVWYXlCVlA5OVlGbnlwYUppVTdDMUdZNlVWaEhOS0E9PQ==
url_encoded Z0FBQUFBQm5Lak9HQ2ZlVEJWbk5rUXFZenNSX0o2LUFiSEtLaVdyVzFwa1pCY3hmWWo1UEI4S215MVVtQmZydWE1QTR5OGZON3FpdFhKanlHSEgzOVFMbjQ1dGhYWVo1em1FaXlMVGZVM2pJNnBYWllsLTZxRm85bkRYazVwOHFnTHpIdEE3ZGJjUl80aEdMX3U3WGs5NGM0RVZFRGgtVldJUGtSaEFOY3pkMUxIeGpQNWNiRG9sRy1NN2JWZkRlNkxKVC15ZF9GOEhBLWVoTUlsRTctZVpwdzE0Q3VJVk9wQT09

Raw Record

{
  "text": "I’ve read that inference speed for models like Llama-2 70B is ~10 t/s at best. So that left me wondering how the extremely large models like GPT-4 (1T params?) do their fast 20 t/s inference. With 10x the params, they gotta have at least 3x the layers(?) So that should make its inference much slower. Am I missing anything? What kind of further improvements might these companies be doing to power their fast APIs?\n\nEdit: I must mention that you cannot parallelize across GPUs to help with latency of a single example when the data has to pass through model layers sequentially.\n\nAnd with the large model sizes, model parallelism, with its inter-GPU communication should make it even slower…",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-07",
  "username_encoded": "Z0FBQUFBQm5LakwyRmk5Z29IRmk2Z2VIWnlvSFljd0FDN290OU11Tk9YYVZkdmtURkgzdEdseVE3WERybTJkYk1pdXVWYXlCVlA5OVlGbnlwYUppVTdDMUdZNlVWaEhOS0E9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9HQ2ZlVEJWbk5rUXFZenNSX0o2LUFiSEtLaVdyVzFwa1pCY3hmWWo1UEI4S215MVVtQmZydWE1QTR5OGZON3FpdFhKanlHSEgzOVFMbjQ1dGhYWVo1em1FaXlMVGZVM2pJNnBYWllsLTZxRm85bkRYazVwOHFnTHpIdEE3ZGJjUl80aEdMX3U3WGs5NGM0RVZFRGgtVldJUGtSaEFOY3pkMUxIeGpQNWNiRG9sRy1NN2JWZkRlNkxKVC15ZF9GOEhBLWVoTUlsRTctZVpwdzE0Q3VJVk9wQT09"
}

Entry Information