Row 60977

Row ID: 60977 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 60977 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

But surely you realize that this is another, difficult task. first of all you need to learn to make any sense of the auditory and visual signal. Then you need to be able to use the correlation of both to be able to do source separation, then you need to realize that the source holding you close is probably communicating with you, while the bird outside is not. Then, for the example with youtube, you have to realize that the other signal further away might also be language (or more likely, ignore it because it is not correlated with any "parent" entity or any other entity that has a direct visual presence in the room).

You are right these are auxiliarly tasks, but all of these tasks are pre-solved for LLMs that get well curated english texts as input. Learning an LLM from raw audio recorded somewhere is much harder.

FieldValue
text But surely you realize that this is another, difficult task. first of all you need to learn to make any sense of the auditory and visual signal. Then you need to be able to use the correlation of both to be able to do source separation, then you need to realize that the source holding you close is probably communicating with you, while the bird outside is not. Then, for the example with youtube, you have to realize that the other signal further away might also be language (or more likely, ignore…
label r/machinelearning
dataType comment
communityName r/MachineLearning
datetime 2024-05-23
username_encoded Z0FBQUFBQm5Lak1aZW5zanlrWkxKRi1QS1NDTDc4YVdXbWt2Mm96S0xwVGNweE5PUjVCeTJHMEp1dEJrY0xDQ0VIMXJiMEE5Z25RaHFYbDhXVmxPNElDbHRTQTNoM0VUNGc9PQ==
url_encoded Z0FBQUFBQm5Lak9wQktxT1AzbmIxMUJmUXNQci1LY29wM20yWkhhZWM2OE9XTVJrS3pJSEhoN3k0QjNLN2F3QzJmd0tKTUh3Y21WMXFwQUZ5LUZESDFIMFpHcVY1aFNtdEx3aUMzVlAzZXlTNGZQaTRlQi16VTgtazFZTlNZckg5TENiOXdnOWRHdGl6cVhSd3NsNF9xd1NvQ2FjelB3dkFDLTVzeVpCNFpTcVpiY01LelN5dExhQU5NN2l0WmpZeWFtYTVJTll0MzJXY1RWWWJMUEFuUXJteGVxUmhlbEs2Zz09

Raw Record

{
  "text": "But surely you realize that this is another, difficult task. first of all you need to learn to make any sense of the auditory and visual signal. Then you need to be able to use the correlation of both to be able to do source separation, then you need to realize that the source holding you close is probably communicating with you, while the bird outside is not. Then, for the example with youtube, you have to realize that the other signal further away might also be language (or more likely, ignore it because it is not correlated with any \"parent\" entity or any other entity that has a direct visual presence in the room).\n\nYou are right these are auxiliarly tasks, but all of these tasks are pre-solved for LLMs that get well curated english texts as input. Learning an LLM from raw audio recorded somewhere is much harder.",
  "label": "r/machinelearning",
  "dataType": "comment",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-23",
  "username_encoded": "Z0FBQUFBQm5Lak1aZW5zanlrWkxKRi1QS1NDTDc4YVdXbWt2Mm96S0xwVGNweE5PUjVCeTJHMEp1dEJrY0xDQ0VIMXJiMEE5Z25RaHFYbDhXVmxPNElDbHRTQTNoM0VUNGc9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9wQktxT1AzbmIxMUJmUXNQci1LY29wM20yWkhhZWM2OE9XTVJrS3pJSEhoN3k0QjNLN2F3QzJmd0tKTUh3Y21WMXFwQUZ5LUZESDFIMFpHcVY1aFNtdEx3aUMzVlAzZXlTNGZQaTRlQi16VTgtazFZTlNZckg5TENiOXdnOWRHdGl6cVhSd3NsNF9xd1NvQ2FjelB3dkFDLTVzeVpCNFpTcVpiY01LelN5dExhQU5NN2l0WmpZeWFtYTVJTll0MzJXY1RWWWJMUEFuUXJteGVxUmhlbEs2Zz09"
}

Entry Information