Row 11076
Content Data
This page contains data entry 11076 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Pretty cool! The fact that it's n\^3 is not super-great buuut might not be so bad if interactions are limited to within a certain range of the current token. A 4k context is reasonably common for n\^2 transformers; assuming the constant factors are similar, that would make a comparable context for trittention cbrt(4096\^2) = 256.
I think Google has a hybrid model that uses linearized attention for long context info and quadratic attention for windows. But you could imagine adding trittention for a shorter window into that scheme.
EDIT: I'm not sure I 100% understand the trittention cube method but if that table compares models of similar parameter counts it looks like the clear winner. Seems really promising!
| Field | Value |
|---|---|
| text | Pretty cool! The fact that it's n\^3 is not super-great buuut might not be so bad if interactions are limited to within a certain range of the current token. A 4k context is reasonably common for n\^2 transformers; assuming the constant factors are similar, that would make a comparable context for trittention cbrt(4096\^2) = 256. I think Google has a hybrid model that uses linearized attention for long context info and quadratic attention for windows. But you could imagine adding trittention f… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-20 |
| username_encoded | Z0FBQUFBQm5Lakw1cTB2Y3MtOGxLbUZmTWRhLWdmZ19VckNZSFNHNFQ4WUFnelQ4VDFNYlc0WXBidVdVdEo2OXBXLTMtRFNBb2pzUUVtcHZsejVTcC15UGl3bk9XVzVvZnc9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9KUFBhNTN2SXFaeDdydkpXY0R3OTFyek1fTGkzOWN1eXBnX21Jb2lUb3VKOUhfd3FWRjNaanhPRnhxMFpYcG5neEpYNmhnblpHcllfUGpab0hLQ1RVVmdJSzhEVjlsNTAtMkZJSG1PcUx5XzNueVR0QVNuLTZYSjE2amR0aUdLNEI5Qlo0czUwZFJKS0VodG13QlFBQTc5cVZjRDVWZXZpSEpmeEpiQ1VaQmdzcDhEMDNNLVlWZUtPbE1iM3lwWkJm |
Raw Record
{
"text": "Pretty cool! The fact that it's n\\^3 is not super-great buuut might not be so bad if interactions are limited to within a certain range of the current token. A 4k context is reasonably common for n\\^2 transformers; assuming the constant factors are similar, that would make a comparable context for trittention cbrt(4096\\^2) = 256.\n\nI think Google has a hybrid model that uses linearized attention for long context info and quadratic attention for windows. But you could imagine adding trittention for a shorter window into that scheme.\n\nEDIT: \nI'm not sure I 100% understand the trittention cube method but if that table compares models of similar parameter counts it looks like the clear winner. Seems really promising!",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-20",
"username_encoded": "Z0FBQUFBQm5Lakw1cTB2Y3MtOGxLbUZmTWRhLWdmZ19VckNZSFNHNFQ4WUFnelQ4VDFNYlc0WXBidVdVdEo2OXBXLTMtRFNBb2pzUUVtcHZsejVTcC15UGl3bk9XVzVvZnc9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9KUFBhNTN2SXFaeDdydkpXY0R3OTFyek1fTGkzOWN1eXBnX21Jb2lUb3VKOUhfd3FWRjNaanhPRnhxMFpYcG5neEpYNmhnblpHcllfUGpab0hLQ1RVVmdJSzhEVjlsNTAtMkZJSG1PcUx5XzNueVR0QVNuLTZYSjE2amR0aUdLNEI5Qlo0czUwZFJKS0VodG13QlFBQTc5cVZjRDVWZXZpSEpmeEpiQ1VaQmdzcDhEMDNNLVlWZUtPbE1iM3lwWkJm"
}
Entry Information
- Entry ID: 11076
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000