Row 16365
Content Data
This page contains data entry 16365 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
This is an interesting question that shouldn't be too tough to test out! Create a dataset from your multimodal vectors with the labels being the modality, and then train a MLP to predict modality from vectors.
I would guess it depends on your specific model/task (and like the other poster mentioned, whether you are feeding extra information in to represent modality, like sentinel tokens) but I would guess that in most cases there should be artifacts in the vector that come from each encoder in a way that a model could use to tell them apart.
| Field | Value |
|---|---|
| text | This is an interesting question that shouldn't be too tough to test out! Create a dataset from your multimodal vectors with the labels being the modality, and then train a MLP to predict modality from vectors. I would guess it depends on your specific model/task (and like the other poster mentioned, whether you are feeding extra information in to represent modality, like sentinel tokens) but I would guess that in most cases there should be artifacts in the vector that come from each encoder in… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-20 |
| username_encoded | Z0FBQUFBQm5Lakw5UnpQQ2wtVktKS0tnbi0yNjdIZDdaYXFQQXBrcElxOGVGdktyVkZqdDIyUXRSb0RYUzRKTkNfSnR0dGkzQktxMzZwVVhadWNvTG0taVliV3VjSGpXM2c9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9NZGVpeVhZdmNITFl2aVVZSXEzcWlENDJCeGV2NGx1a205UjB4ZmtSRkdILTAwVmJzMXpUeXV4bF93bTAzVXZQWHJBeU9GSjU4ZkRoMGN3LUM0QnpaSTZFWTZ3WFRLTDNXUmNLczFPam5vNGdiUkZJMWtxMUhLUzNRdVdxdndoSDFDTmtfblJHZmZ6c0xzR0Npa3AtRkFiZDhIV21rcGM3T0UzbGZqWnFlWHNNN3hLWF8xUkJTakxDWG1HVV9vMDJySlhUSGw0UUhMMERvbEl5TjNPWHJaLXFpV0RCYjJSVE9rZ0M0OUNsRnZfWT0= |
Raw Record
{
"text": "This is an interesting question that shouldn't be too tough to test out! Create a dataset from your multimodal vectors with the labels being the modality, and then train a MLP to predict modality from vectors. \n\nI would guess it depends on your specific model/task (and like the other poster mentioned, whether you are feeding extra information in to represent modality, like sentinel tokens) but I would guess that in most cases there should be artifacts in the vector that come from each encoder in a way that a model could use to tell them apart.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-20",
"username_encoded": "Z0FBQUFBQm5Lakw5UnpQQ2wtVktKS0tnbi0yNjdIZDdaYXFQQXBrcElxOGVGdktyVkZqdDIyUXRSb0RYUzRKTkNfSnR0dGkzQktxMzZwVVhadWNvTG0taVliV3VjSGpXM2c9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9NZGVpeVhZdmNITFl2aVVZSXEzcWlENDJCeGV2NGx1a205UjB4ZmtSRkdILTAwVmJzMXpUeXV4bF93bTAzVXZQWHJBeU9GSjU4ZkRoMGN3LUM0QnpaSTZFWTZ3WFRLTDNXUmNLczFPam5vNGdiUkZJMWtxMUhLUzNRdVdxdndoSDFDTmtfblJHZmZ6c0xzR0Npa3AtRkFiZDhIV21rcGM3T0UzbGZqWnFlWHNNN3hLWF8xUkJTakxDWG1HVV9vMDJySlhUSGw0UUhMMERvbEl5TjNPWHJaLXFpV0RCYjJSVE9rZ0M0OUNsRnZfWT0="
}
Entry Information
- Entry ID: 16365
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000