Row 91996

Row ID: 91996 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 91996 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

Two examples of what I believe is the SOTA multimodal pretraining technique is in the Llava paper and the Qwen Audio paper. Essentially, they freeze the LLM during pre-training, and create an encoder that encodes the stuff other than text into the frozen LLMs input space. Then the LLM is finetuned on multimodal instructions. This way the LLM can "understand" multimodal data without forgetting its text understanding.

FieldValue
text Two examples of what I believe is the SOTA multimodal pretraining technique is in the Llava paper and the Qwen Audio paper. Essentially, they freeze the LLM during pre-training, and create an encoder that encodes the stuff other than text into the frozen LLMs input space. Then the LLM is finetuned on multimodal instructions. This way the LLM can "understand" multimodal data without forgetting its text understanding.
label r/machinelearning
dataType comment
communityName r/MachineLearning
datetime 2024-05-25
username_encoded Z0FBQUFBQm5Lak1zWUlkTVZsMXEtZFdxUUFqT016MFRqdGFFc2M1RHkwQUltc08tQ3VmQlFxeXNFRnFwd3E5SDBvNDFkWmJfUE92Rnp1ME5QeGtuaU5RSEFFMW9xMXdoUlhRVW8xTVlGTnpoMmEwOVBDWUJCOGs9
url_encoded Z0FBQUFBQm5Lak8tZzVhT1c0TXRlbHFia2RIM01pMUh0WWl3b24yNE1tVnJFNUEyQ3kwMlNMRThHdDhpZElseEZ3OTZPN2FLSHN2VG9DdjVKSjlLVlZjNnhDcUNZQ3p0REFMRnlkYWFLdkpPTEZQZG1jN1RsR3p2SGRaTHNaMFh5ZlJ5OGRCOXVzN2hVczMza0lGLWNHMVJrVm1SaGtGal9MNVB6b19QWnJGMjdwSUJFSzI2ZXdCWmNyd1BldDJxVE1ROEwzaHRSN015QUZXeVNFYWRfZU1iYzdUNWhzbTVyQT09

Raw Record

{
  "text": "Two examples of what I believe is the SOTA multimodal pretraining technique is in the Llava paper and the Qwen Audio paper. Essentially, they freeze the LLM during pre-training, and create an encoder that encodes the stuff other than text into the frozen LLMs input space. Then the LLM is finetuned on multimodal instructions. This way the LLM can \"understand\" multimodal data without forgetting its text understanding.",
  "label": "r/machinelearning",
  "dataType": "comment",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-25",
  "username_encoded": "Z0FBQUFBQm5Lak1zWUlkTVZsMXEtZFdxUUFqT016MFRqdGFFc2M1RHkwQUltc08tQ3VmQlFxeXNFRnFwd3E5SDBvNDFkWmJfUE92Rnp1ME5QeGtuaU5RSEFFMW9xMXdoUlhRVW8xTVlGTnpoMmEwOVBDWUJCOGs9",
  "url_encoded": "Z0FBQUFBQm5Lak8tZzVhT1c0TXRlbHFia2RIM01pMUh0WWl3b24yNE1tVnJFNUEyQ3kwMlNMRThHdDhpZElseEZ3OTZPN2FLSHN2VG9DdjVKSjlLVlZjNnhDcUNZQ3p0REFMRnlkYWFLdkpPTEZQZG1jN1RsR3p2SGRaTHNaMFh5ZlJ5OGRCOXVzN2hVczMza0lGLWNHMVJrVm1SaGtGal9MNVB6b19QWnJGMjdwSUJFSzI2ZXdCWmNyd1BldDJxVE1ROEwzaHRSN015QUZXeVNFYWRfZU1iYzdUNWhzbTVyQT09"
}

Entry Information