Row 91996
Content Data
This page contains data entry 91996 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Two examples of what I believe is the SOTA multimodal pretraining technique is in the Llava paper and the Qwen Audio paper. Essentially, they freeze the LLM during pre-training, and create an encoder that encodes the stuff other than text into the frozen LLMs input space. Then the LLM is finetuned on multimodal instructions. This way the LLM can "understand" multimodal data without forgetting its text understanding.
| Field | Value |
|---|---|
| text | Two examples of what I believe is the SOTA multimodal pretraining technique is in the Llava paper and the Qwen Audio paper. Essentially, they freeze the LLM during pre-training, and create an encoder that encodes the stuff other than text into the frozen LLMs input space. Then the LLM is finetuned on multimodal instructions. This way the LLM can "understand" multimodal data without forgetting its text understanding. |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-25 |
| username_encoded | Z0FBQUFBQm5Lak1zWUlkTVZsMXEtZFdxUUFqT016MFRqdGFFc2M1RHkwQUltc08tQ3VmQlFxeXNFRnFwd3E5SDBvNDFkWmJfUE92Rnp1ME5QeGtuaU5RSEFFMW9xMXdoUlhRVW8xTVlGTnpoMmEwOVBDWUJCOGs9 |
| url_encoded | Z0FBQUFBQm5Lak8tZzVhT1c0TXRlbHFia2RIM01pMUh0WWl3b24yNE1tVnJFNUEyQ3kwMlNMRThHdDhpZElseEZ3OTZPN2FLSHN2VG9DdjVKSjlLVlZjNnhDcUNZQ3p0REFMRnlkYWFLdkpPTEZQZG1jN1RsR3p2SGRaTHNaMFh5ZlJ5OGRCOXVzN2hVczMza0lGLWNHMVJrVm1SaGtGal9MNVB6b19QWnJGMjdwSUJFSzI2ZXdCWmNyd1BldDJxVE1ROEwzaHRSN015QUZXeVNFYWRfZU1iYzdUNWhzbTVyQT09 |
Raw Record
{
"text": "Two examples of what I believe is the SOTA multimodal pretraining technique is in the Llava paper and the Qwen Audio paper. Essentially, they freeze the LLM during pre-training, and create an encoder that encodes the stuff other than text into the frozen LLMs input space. Then the LLM is finetuned on multimodal instructions. This way the LLM can \"understand\" multimodal data without forgetting its text understanding.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-25",
"username_encoded": "Z0FBQUFBQm5Lak1zWUlkTVZsMXEtZFdxUUFqT016MFRqdGFFc2M1RHkwQUltc08tQ3VmQlFxeXNFRnFwd3E5SDBvNDFkWmJfUE92Rnp1ME5QeGtuaU5RSEFFMW9xMXdoUlhRVW8xTVlGTnpoMmEwOVBDWUJCOGs9",
"url_encoded": "Z0FBQUFBQm5Lak8tZzVhT1c0TXRlbHFia2RIM01pMUh0WWl3b24yNE1tVnJFNUEyQ3kwMlNMRThHdDhpZElseEZ3OTZPN2FLSHN2VG9DdjVKSjlLVlZjNnhDcUNZQ3p0REFMRnlkYWFLdkpPTEZQZG1jN1RsR3p2SGRaTHNaMFh5ZlJ5OGRCOXVzN2hVczMza0lGLWNHMVJrVm1SaGtGal9MNVB6b19QWnJGMjdwSUJFSzI2ZXdCWmNyd1BldDJxVE1ROEwzaHRSN015QUZXeVNFYWRfZU1iYzdUNWhzbTVyQT09"
}
Entry Information
- Entry ID: 91996
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000