Row 11076

Row ID: 11076 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 11076 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

Pretty cool! The fact that it's n\^3 is not super-great buuut might not be so bad if interactions are limited to within a certain range of the current token. A 4k context is reasonably common for n\^2 transformers; assuming the constant factors are similar, that would make a comparable context for trittention cbrt(4096\^2) = 256.

I think Google has a hybrid model that uses linearized attention for long context info and quadratic attention for windows. But you could imagine adding trittention for a shorter window into that scheme.

EDIT: I'm not sure I 100% understand the trittention cube method but if that table compares models of similar parameter counts it looks like the clear winner. Seems really promising!

FieldValue
text Pretty cool! The fact that it's n\^3 is not super-great buuut might not be so bad if interactions are limited to within a certain range of the current token. A 4k context is reasonably common for n\^2 transformers; assuming the constant factors are similar, that would make a comparable context for trittention cbrt(4096\^2) = 256. I think Google has a hybrid model that uses linearized attention for long context info and quadratic attention for windows. But you could imagine adding trittention f…
label r/machinelearning
dataType comment
communityName r/MachineLearning
datetime 2024-05-20
username_encoded Z0FBQUFBQm5Lakw1cTB2Y3MtOGxLbUZmTWRhLWdmZ19VckNZSFNHNFQ4WUFnelQ4VDFNYlc0WXBidVdVdEo2OXBXLTMtRFNBb2pzUUVtcHZsejVTcC15UGl3bk9XVzVvZnc9PQ==
url_encoded Z0FBQUFBQm5Lak9KUFBhNTN2SXFaeDdydkpXY0R3OTFyek1fTGkzOWN1eXBnX21Jb2lUb3VKOUhfd3FWRjNaanhPRnhxMFpYcG5neEpYNmhnblpHcllfUGpab0hLQ1RVVmdJSzhEVjlsNTAtMkZJSG1PcUx5XzNueVR0QVNuLTZYSjE2amR0aUdLNEI5Qlo0czUwZFJKS0VodG13QlFBQTc5cVZjRDVWZXZpSEpmeEpiQ1VaQmdzcDhEMDNNLVlWZUtPbE1iM3lwWkJm

Raw Record

{
  "text": "Pretty cool! The fact that it's n\\^3 is not super-great buuut might not be so bad if interactions are limited to within a certain range of the current token.  A 4k context is reasonably common for n\\^2 transformers; assuming the constant factors are similar, that would make a comparable context for trittention cbrt(4096\\^2) = 256.\n\nI think Google has a hybrid model that uses linearized attention for long context info and quadratic attention for windows. But you could imagine adding trittention for a shorter window into that scheme.\n\nEDIT:  \nI'm not sure I 100% understand the trittention cube method but if that table compares models of similar parameter counts it looks like the clear winner. Seems really promising!",
  "label": "r/machinelearning",
  "dataType": "comment",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-20",
  "username_encoded": "Z0FBQUFBQm5Lakw1cTB2Y3MtOGxLbUZmTWRhLWdmZ19VckNZSFNHNFQ4WUFnelQ4VDFNYlc0WXBidVdVdEo2OXBXLTMtRFNBb2pzUUVtcHZsejVTcC15UGl3bk9XVzVvZnc9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9KUFBhNTN2SXFaeDdydkpXY0R3OTFyek1fTGkzOWN1eXBnX21Jb2lUb3VKOUhfd3FWRjNaanhPRnhxMFpYcG5neEpYNmhnblpHcllfUGpab0hLQ1RVVmdJSzhEVjlsNTAtMkZJSG1PcUx5XzNueVR0QVNuLTZYSjE2amR0aUdLNEI5Qlo0czUwZFJKS0VodG13QlFBQTc5cVZjRDVWZXZpSEpmeEpiQ1VaQmdzcDhEMDNNLVlWZUtPbE1iM3lwWkJm"
}

Entry Information