Row 9953

Row ID: 9953 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 9953 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

My understanding of multimodal LLMs, such as LLaVA, use vision & text encoders to relate the two. Vision is taken a step further by introducing a foundational model to extract features in an image and organize the classes of detected objects into some sort of textual logic.

Now, I assume this idea is how desired text is trained to be 'discovered' in an image. After the text is 'discovered', however, is the LLM using a more standard OCR recognizer under the hood (such as in [this paper](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4636121)) to interpret the text? Or is there something else being done?

Thanks in advance!

FieldValue
text My understanding of multimodal LLMs, such as LLaVA, use vision & text encoders to relate the two. Vision is taken a step further by introducing a foundational model to extract features in an image and organize the classes of detected objects into some sort of textual logic. Now, I assume this idea is how desired text is trained to be 'discovered' in an image. After the text is 'discovered', however, is the LLM using a more standard OCR recognizer under the hood (such as in [this paper](https://…
label r/datascience
dataType post
communityName r/datascience
datetime 2024-05-20
username_encoded Z0FBQUFBQm5Lakw1OXJnUlRVenRPT09PaWZpQXBsZlZuZUx0cEZ5WXRiN1N0dTJuVU1pY2l2Tmc2MnBlRUNWSXczMmpNYVJjMGxaYzV0NGZEOW5hWlVpRzU4bHNIUHZZN2c9PQ==
url_encoded Z0FBQUFBQm5Lak9JNVVReVBudFc5RHRpVVl1aWF1UERmQWU1WllPWEJhdkhzZFJ4ODd0NjJHUlRZM2h6VzZmamFOMFBIZy02OWtIN2pRQnhyWnFsY0oyUG13LTZqYXBtOWEza1JkRVAwcmVTOExpTnVDR2lXUjlJaVFZMDYwUHhWQkVSUUE4WWV4UURveGlocHZUeWlRWi00VE9tNGcwUGhSREVQTFlVQnc5WUxMVk5aTHZpUFRFMllFQWkweTBLdWQya295cG9sSnVtQ0E3b1J2QzlhQkpLWnB3MFFweGlRQT09

Raw Record

{
  "text": "My understanding of multimodal LLMs, such as LLaVA, use vision & text encoders to relate the two. Vision is taken a step further by introducing a foundational model to extract features in an image and organize the classes of detected objects into some sort of textual logic.\n\nNow, I assume this idea is how desired text is trained to be 'discovered' in an image. After the text is 'discovered', however, is the LLM using a more standard OCR recognizer under the hood (such as in [this paper](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4636121)) to interpret the text? Or is there something else being done?\n\nThanks in advance!",
  "label": "r/datascience",
  "dataType": "post",
  "communityName": "r/datascience",
  "datetime": "2024-05-20",
  "username_encoded": "Z0FBQUFBQm5Lakw1OXJnUlRVenRPT09PaWZpQXBsZlZuZUx0cEZ5WXRiN1N0dTJuVU1pY2l2Tmc2MnBlRUNWSXczMmpNYVJjMGxaYzV0NGZEOW5hWlVpRzU4bHNIUHZZN2c9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9JNVVReVBudFc5RHRpVVl1aWF1UERmQWU1WllPWEJhdkhzZFJ4ODd0NjJHUlRZM2h6VzZmamFOMFBIZy02OWtIN2pRQnhyWnFsY0oyUG13LTZqYXBtOWEza1JkRVAwcmVTOExpTnVDR2lXUjlJaVFZMDYwUHhWQkVSUUE4WWV4UURveGlocHZUeWlRWi00VE9tNGcwUGhSREVQTFlVQnc5WUxMVk5aTHZpUFRFMllFQWkweTBLdWQya295cG9sSnVtQ0E3b1J2QzlhQkpLWnB3MFFweGlRQT09"
}

Entry Information