Row 57712
Content Data
This page contains data entry 57712 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
I’ve always been curious of this notion. I have a one-year-old who is yet to speak. But if I would give a rough estimate on the number of hours she has been exposed to languaged music, audiobooks, languaged videos on YouTube, and conversations around her, it must amount to an enormous corpus. And she has yet to say a word. If we assume a WPM of 150 for an average speaker and assume 5 hours of exposure a day for 365 days, that’s about 15 million words in her corpus. Since she is surrounded most often by conversation, I would assume her corpus is both larger and more context-rich. The brain seems wildly inefficient if we are talking about learning language? Her data input is gigantic, continuous and enriched by all other modes of input to correlate tokens to meaning. All that to soon say “mama.”
| Field | Value |
|---|---|
| text | I’ve always been curious of this notion. I have a one-year-old who is yet to speak. But if I would give a rough estimate on the number of hours she has been exposed to languaged music, audiobooks, languaged videos on YouTube, and conversations around her, it must amount to an enormous corpus. And she has yet to say a word. If we assume a WPM of 150 for an average speaker and assume 5 hours of exposure a day for 365 days, that’s about 15 million words in her corpus. Since she is surrounded most o… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-23 |
| username_encoded | Z0FBQUFBQm5Lak1YOE1FVmt1cGxjdE90QkExcjZuT3hGMzR0Vm1PNWcxVTZvMFBOcXVaMk1VbExweEhHTmx5d18tUzBCaGdydFpjUEdCNDZCTG5ZLUJXSWJlMjlrU1RiT29uNHJ2S0hlVmZGT3NfSk1fUktfWEE9 |
| url_encoded | Z0FBQUFBQm5Lak9ubmVVaEJ2TmNGMXg3TDR1N2o2UGFVTVNRcWN1OGt0Q1F4NG1UejRlSEs2V3NjZTVkZ2xxc0JmODJJcGNObXo0NE9ILThtWTAxOXJlcE1WZjc5S0VSRl9kQzBJalRqLUVuNmk1MVpQZ1lVOWp0WDYwWWFrcFNVemM3WkRTMk9oYzVkVkxZMG4xellYVkdCd2NCb2s5N2QwSXlDVm5sdzNOa0FteldGT1k4anRNdkZZdFhHNHZZZTdmdE9JbGlYVVRibUVIcFlLdW5Wc2d4MUxZLWprN050UT09 |
Raw Record
{
"text": "I’ve always been curious of this notion. I have a one-year-old who is yet to speak. But if I would give a rough estimate on the number of hours she has been exposed to languaged music, audiobooks, languaged videos on YouTube, and conversations around her, it must amount to an enormous corpus. And she has yet to say a word. If we assume a WPM of 150 for an average speaker and assume 5 hours of exposure a day for 365 days, that’s about 15 million words in her corpus. Since she is surrounded most often by conversation, I would assume her corpus is both larger and more context-rich. The brain seems wildly inefficient if we are talking about learning language? Her data input is gigantic, continuous and enriched by all other modes of input to correlate tokens to meaning. All that to soon say “mama.”",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-23",
"username_encoded": "Z0FBQUFBQm5Lak1YOE1FVmt1cGxjdE90QkExcjZuT3hGMzR0Vm1PNWcxVTZvMFBOcXVaMk1VbExweEhHTmx5d18tUzBCaGdydFpjUEdCNDZCTG5ZLUJXSWJlMjlrU1RiT29uNHJ2S0hlVmZGT3NfSk1fUktfWEE9",
"url_encoded": "Z0FBQUFBQm5Lak9ubmVVaEJ2TmNGMXg3TDR1N2o2UGFVTVNRcWN1OGt0Q1F4NG1UejRlSEs2V3NjZTVkZ2xxc0JmODJJcGNObXo0NE9ILThtWTAxOXJlcE1WZjc5S0VSRl9kQzBJalRqLUVuNmk1MVpQZ1lVOWp0WDYwWWFrcFNVemM3WkRTMk9oYzVkVkxZMG4xellYVkdCd2NCb2s5N2QwSXlDVm5sdzNOa0FteldGT1k4anRNdkZZdFhHNHZZZTdmdE9JbGlYVVRibUVIcFlLdW5Wc2d4MUxZLWprN050UT09"
}
Entry Information
- Entry ID: 57712
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000