Row 6093
Content Data
This page contains data entry 6093 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
I’ve read that inference speed for models like Llama-2 70B is ~10 t/s at best. So that left me wondering how the extremely large models like GPT-4 (1T params?) do their fast 20 t/s inference. With 10x the params, they gotta have at least 3x the layers(?) So that should make its inference much slower. Am I missing anything? What kind of further improvements might these companies be doing to power their fast APIs?
Edit: I must mention that you cannot parallelize across GPUs to help with latency of a single example when the data has to pass through model layers sequentially.
And with the large model sizes, model parallelism, with its inter-GPU communication should make it even slower…
| Field | Value |
|---|---|
| text | I’ve read that inference speed for models like Llama-2 70B is ~10 t/s at best. So that left me wondering how the extremely large models like GPT-4 (1T params?) do their fast 20 t/s inference. With 10x the params, they gotta have at least 3x the layers(?) So that should make its inference much slower. Am I missing anything? What kind of further improvements might these companies be doing to power their fast APIs? Edit: I must mention that you cannot parallelize across GPUs to help with latency o… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-07 |
| username_encoded | Z0FBQUFBQm5LakwyRmk5Z29IRmk2Z2VIWnlvSFljd0FDN290OU11Tk9YYVZkdmtURkgzdEdseVE3WERybTJkYk1pdXVWYXlCVlA5OVlGbnlwYUppVTdDMUdZNlVWaEhOS0E9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9HQ2ZlVEJWbk5rUXFZenNSX0o2LUFiSEtLaVdyVzFwa1pCY3hmWWo1UEI4S215MVVtQmZydWE1QTR5OGZON3FpdFhKanlHSEgzOVFMbjQ1dGhYWVo1em1FaXlMVGZVM2pJNnBYWllsLTZxRm85bkRYazVwOHFnTHpIdEE3ZGJjUl80aEdMX3U3WGs5NGM0RVZFRGgtVldJUGtSaEFOY3pkMUxIeGpQNWNiRG9sRy1NN2JWZkRlNkxKVC15ZF9GOEhBLWVoTUlsRTctZVpwdzE0Q3VJVk9wQT09 |
Raw Record
{
"text": "I’ve read that inference speed for models like Llama-2 70B is ~10 t/s at best. So that left me wondering how the extremely large models like GPT-4 (1T params?) do their fast 20 t/s inference. With 10x the params, they gotta have at least 3x the layers(?) So that should make its inference much slower. Am I missing anything? What kind of further improvements might these companies be doing to power their fast APIs?\n\nEdit: I must mention that you cannot parallelize across GPUs to help with latency of a single example when the data has to pass through model layers sequentially.\n\nAnd with the large model sizes, model parallelism, with its inter-GPU communication should make it even slower…",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-07",
"username_encoded": "Z0FBQUFBQm5LakwyRmk5Z29IRmk2Z2VIWnlvSFljd0FDN290OU11Tk9YYVZkdmtURkgzdEdseVE3WERybTJkYk1pdXVWYXlCVlA5OVlGbnlwYUppVTdDMUdZNlVWaEhOS0E9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9HQ2ZlVEJWbk5rUXFZenNSX0o2LUFiSEtLaVdyVzFwa1pCY3hmWWo1UEI4S215MVVtQmZydWE1QTR5OGZON3FpdFhKanlHSEgzOVFMbjQ1dGhYWVo1em1FaXlMVGZVM2pJNnBYWllsLTZxRm85bkRYazVwOHFnTHpIdEE3ZGJjUl80aEdMX3U3WGs5NGM0RVZFRGgtVldJUGtSaEFOY3pkMUxIeGpQNWNiRG9sRy1NN2JWZkRlNkxKVC15ZF9GOEhBLWVoTUlsRTctZVpwdzE0Q3VJVk9wQT09"
}
Entry Information
- Entry ID: 6093
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000