Row 61277

Row ID: 61277 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 61277 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

There is substantial scholarship that language is [not learned through passive exposure](https://www.kqed.org/mindshift/60988/can-babies-learn-from-ms-rachel-and-other-baby-tv-shows). So all those youtube videos and background conversations are completely meaningless to the child. It's like training on data that has a random error function, a background hum that does not amount to any salient neural weights.

The relevant training data for speech is direct interaction, actually playing with the child, responding to its babling with meaningful answers, words uttered in relation to a physical or visual activity etc. Depending on the child, the level of caregiver involvement and the age when such interactions become possible (probably no sooner than 4-5 moths), we are talking about no more than a few hundred hours of very low density speech that must be parsed along with the corresponding multimodal visual and tactile input, all of which are alien to the child.

If you think that is low efficiency, then by all means I challenge you to create a model that, handed a few hundred hours of mp3 data (which roughly corresponds to the cochlear neural inputs) and an associated video stream, can produce the mp3 spectrogram of the word "mama" when an unknown video of that person is fed in. Of course, all of this would be fully unstructured learning, the only allowed feedback would be summing up the output spectrum to the input spectrum (listening itself speak), as well as video of a very happy mama when the first "ma" is uttered.

If you can really prove this is a simple problem than in all honesty you have some papers to write instead of wasting time on Reddit.

FieldValue
text There is substantial scholarship that language is [not learned through passive exposure](https://www.kqed.org/mindshift/60988/can-babies-learn-from-ms-rachel-and-other-baby-tv-shows). So all those youtube videos and background conversations are completely meaningless to the child. It's like training on data that has a random error function, a background hum that does not amount to any salient neural weights. The relevant training data for speech is direct interaction, actually playing with the …
label r/machinelearning
dataType comment
communityName r/MachineLearning
datetime 2024-05-23
username_encoded Z0FBQUFBQm5Lak1aY3BZWHA5XzZaR2l5M3dKLUVKNFJGdzB0TU5qem5yNUJQUEdTbEFBRjQ3YnJwdXgtWFh3SjdyVDBKVnhMWmtYXzNGaTduTFN6bEpYcVJucDVESEdmZXc9PQ==
url_encoded Z0FBQUFBQm5Lak9wR0pHMjRyNms1RjF5Z20xSDRweWpmUFAwZW0yOXp1eW5tSV8wbW5ld3R1X3EwLW1xY2I4aUEyWHItODJ5RWtPTG4yazVuTkpVTHpxMkJ0WGdYOW10ZjctQWpaa3E4ZjFFMWZmU3U2S3MtTHVZYzV5eFdOLVl5U29DVVdKOEd6WXNCbGZ6bHhNdXh0RjNjMWhLUEFQcExWSGJCQTFhSVhOcGdIN2FIeGo3QTdNc1JnSG10c1hqRGRWQjdpV1VVelRqbExaZGFGdTd6N01IQTBnNVpteGwzdz09

Raw Record

{
  "text": "There is substantial scholarship that language is [not learned through passive exposure](https://www.kqed.org/mindshift/60988/can-babies-learn-from-ms-rachel-and-other-baby-tv-shows). So all those youtube videos and background conversations are completely meaningless to the child. It's like training on data that has a random error function, a background hum that does not amount to any salient neural weights.\n\nThe relevant training data for speech is direct interaction, actually playing with the child, responding to its babling with meaningful answers, words uttered in relation to a physical or visual activity etc. Depending on the child, the level of caregiver involvement and the age when such interactions become possible (probably no sooner than 4-5 moths), we are talking about no more than a few hundred hours of very low density speech that must be parsed along with the corresponding multimodal visual and tactile input, all of which are alien to the child.\n\nIf you think that is low efficiency, then by all means I challenge you to create a model that, handed a few hundred hours of mp3 data (which roughly corresponds to the cochlear neural inputs) and an associated video stream, can produce the mp3 spectrogram of the word \"mama\" when an unknown video of that person is fed in. Of course, all of this would be fully unstructured learning, the only allowed feedback would be summing up the output spectrum to the input spectrum (listening itself speak), as well as video of a very happy mama when the first \"ma\" is uttered.\n\nIf you can really prove this is a simple problem than in all honesty you have some papers to write instead of wasting time on Reddit.",
  "label": "r/machinelearning",
  "dataType": "comment",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-23",
  "username_encoded": "Z0FBQUFBQm5Lak1aY3BZWHA5XzZaR2l5M3dKLUVKNFJGdzB0TU5qem5yNUJQUEdTbEFBRjQ3YnJwdXgtWFh3SjdyVDBKVnhMWmtYXzNGaTduTFN6bEpYcVJucDVESEdmZXc9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9wR0pHMjRyNms1RjF5Z20xSDRweWpmUFAwZW0yOXp1eW5tSV8wbW5ld3R1X3EwLW1xY2I4aUEyWHItODJ5RWtPTG4yazVuTkpVTHpxMkJ0WGdYOW10ZjctQWpaa3E4ZjFFMWZmU3U2S3MtTHVZYzV5eFdOLVl5U29DVVdKOEd6WXNCbGZ6bHhNdXh0RjNjMWhLUEFQcExWSGJCQTFhSVhOcGdIN2FIeGo3QTdNc1JnSG10c1hqRGRWQjdpV1VVelRqbExaZGFGdTd6N01IQTBnNVpteGwzdz09"
}

Entry Information