Row 9953
Content Data
This page contains data entry 9953 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
My understanding of multimodal LLMs, such as LLaVA, use vision & text encoders to relate the two. Vision is taken a step further by introducing a foundational model to extract features in an image and organize the classes of detected objects into some sort of textual logic.
Now, I assume this idea is how desired text is trained to be 'discovered' in an image. After the text is 'discovered', however, is the LLM using a more standard OCR recognizer under the hood (such as in [this paper](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4636121)) to interpret the text? Or is there something else being done?
Thanks in advance!
| Field | Value |
|---|---|
| text | My understanding of multimodal LLMs, such as LLaVA, use vision & text encoders to relate the two. Vision is taken a step further by introducing a foundational model to extract features in an image and organize the classes of detected objects into some sort of textual logic. Now, I assume this idea is how desired text is trained to be 'discovered' in an image. After the text is 'discovered', however, is the LLM using a more standard OCR recognizer under the hood (such as in [this paper](https://… |
| label | r/datascience |
| dataType | post |
| communityName | r/datascience |
| datetime | 2024-05-20 |
| username_encoded | Z0FBQUFBQm5Lakw1OXJnUlRVenRPT09PaWZpQXBsZlZuZUx0cEZ5WXRiN1N0dTJuVU1pY2l2Tmc2MnBlRUNWSXczMmpNYVJjMGxaYzV0NGZEOW5hWlVpRzU4bHNIUHZZN2c9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9JNVVReVBudFc5RHRpVVl1aWF1UERmQWU1WllPWEJhdkhzZFJ4ODd0NjJHUlRZM2h6VzZmamFOMFBIZy02OWtIN2pRQnhyWnFsY0oyUG13LTZqYXBtOWEza1JkRVAwcmVTOExpTnVDR2lXUjlJaVFZMDYwUHhWQkVSUUE4WWV4UURveGlocHZUeWlRWi00VE9tNGcwUGhSREVQTFlVQnc5WUxMVk5aTHZpUFRFMllFQWkweTBLdWQya295cG9sSnVtQ0E3b1J2QzlhQkpLWnB3MFFweGlRQT09 |
Raw Record
{
"text": "My understanding of multimodal LLMs, such as LLaVA, use vision & text encoders to relate the two. Vision is taken a step further by introducing a foundational model to extract features in an image and organize the classes of detected objects into some sort of textual logic.\n\nNow, I assume this idea is how desired text is trained to be 'discovered' in an image. After the text is 'discovered', however, is the LLM using a more standard OCR recognizer under the hood (such as in [this paper](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4636121)) to interpret the text? Or is there something else being done?\n\nThanks in advance!",
"label": "r/datascience",
"dataType": "post",
"communityName": "r/datascience",
"datetime": "2024-05-20",
"username_encoded": "Z0FBQUFBQm5Lakw1OXJnUlRVenRPT09PaWZpQXBsZlZuZUx0cEZ5WXRiN1N0dTJuVU1pY2l2Tmc2MnBlRUNWSXczMmpNYVJjMGxaYzV0NGZEOW5hWlVpRzU4bHNIUHZZN2c9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9JNVVReVBudFc5RHRpVVl1aWF1UERmQWU1WllPWEJhdkhzZFJ4ODd0NjJHUlRZM2h6VzZmamFOMFBIZy02OWtIN2pRQnhyWnFsY0oyUG13LTZqYXBtOWEza1JkRVAwcmVTOExpTnVDR2lXUjlJaVFZMDYwUHhWQkVSUUE4WWV4UURveGlocHZUeWlRWi00VE9tNGcwUGhSREVQTFlVQnc5WUxMVk5aTHZpUFRFMllFQWkweTBLdWQya295cG9sSnVtQ0E3b1J2QzlhQkpLWnB3MFFweGlRQT09"
}
Entry Information
- Entry ID: 9953
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000