Row 47382
Content Data
This page contains data entry 47382 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
**From the article:**
Soon after OpenAI released GPT-4o on Monday, May 13, some Chinese speakers started to notice that something seemed off about this newest version of the chatbot: the tokens it uses to parse text were full of spam and porn phrases.
On May 14, Tianle Cai, a PhD student at Princeton University studying inference efficiency in large language models like those that power such chatbots, accessed GPT-4o’s public token library and pulled a list of the 100 longest Chinese tokens the model uses to parse and compress Chinese prompts.
Humans read in words, but LLMs read in tokens, which are distinct units in a sentence that have consistent and significant meanings. Besides dictionary words, they also include suffixes, common expressions, names, and more. The more tokens a model encodes, the faster the model can “read” a sentence and the less computing power it consumes, thus making the response cheaper.
Of the 100 results, only three of them are common enough to be used in everyday conversations; everything else consisted of words and expressions used specifically in the contexts of either gambling or pornography. The longest token, lasting 10.5 Chinese characters, literally means “\_free Japanese porn video to watch.” Oops.
| Field | Value |
|---|---|
| text | **From the article:** Soon after OpenAI released GPT-4o on Monday, May 13, some Chinese speakers started to notice that something seemed off about this newest version of the chatbot: the tokens it uses to parse text were full of spam and porn phrases. On May 14, Tianle Cai, a PhD student at Princeton University studying inference efficiency in large language models like those that power such chatbots, accessed GPT-4o’s public token library and pulled a list of the 100 longest Chinese tokens th… |
| label | r/artificial |
| dataType | comment |
| communityName | r/artificial |
| datetime | 2024-05-22 |
| username_encoded | Z0FBQUFBQm5Lak1RVGwtcFN3d3kyUWJHUlItLWxTZHVuNlV6RE5GSnZTek5UQUNiQXZjcnNSNnVINU92RVFUUFFOTExUVEVOT2pPMEZvd0wyTTdxQUs3TlRtVGZTRVZRUXc9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9naFFpd2JDRFpFNV9aTTNIdVR2eFN2VE16Y1VlN1lhelFkTEhTUnVYZ1kzZEtGR3hpVng2YTgtNi1TcHU5Tnh4QzE3alJLaFlPcVpmcC15cXQ4dlItZHkwQVQ5bXd1dm96eWloSjBNSzRHU2U1MTEwUEczNTNHM19leXFUMl9KTlZTTlpNZmY5OUF6a2FCajhuMEROU2ZBTGNoNFpDYlQyN3U3MmR3S1hqZlA3bWRRUFJ6OHVYYVNXNnBqc1doZEctNWE1UFFKNy1KZmtlYVZyVFlpa3FPZz09 |
Raw Record
{
"text": "**From the article:**\n\nSoon after OpenAI released GPT-4o on Monday, May 13, some Chinese speakers started to notice that something seemed off about this newest version of the chatbot: the tokens it uses to parse text were full of spam and porn phrases.\n\nOn May 14, Tianle Cai, a PhD student at Princeton University studying inference efficiency in large language models like those that power such chatbots, accessed GPT-4o’s public token library and pulled a list of the 100 longest Chinese tokens the model uses to parse and compress Chinese prompts. \n\nHumans read in words, but LLMs read in tokens, which are distinct units in a sentence that have consistent and significant meanings. Besides dictionary words, they also include suffixes, common expressions, names, and more. The more tokens a model encodes, the faster the model can “read” a sentence and the less computing power it consumes, thus making the response cheaper.\n\nOf the 100 results, only three of them are common enough to be used in everyday conversations; everything else consisted of words and expressions used specifically in the contexts of either gambling or pornography. The longest token, lasting 10.5 Chinese characters, literally means “\\_free Japanese porn video to watch.” Oops.",
"label": "r/artificial",
"dataType": "comment",
"communityName": "r/artificial",
"datetime": "2024-05-22",
"username_encoded": "Z0FBQUFBQm5Lak1RVGwtcFN3d3kyUWJHUlItLWxTZHVuNlV6RE5GSnZTek5UQUNiQXZjcnNSNnVINU92RVFUUFFOTExUVEVOT2pPMEZvd0wyTTdxQUs3TlRtVGZTRVZRUXc9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9naFFpd2JDRFpFNV9aTTNIdVR2eFN2VE16Y1VlN1lhelFkTEhTUnVYZ1kzZEtGR3hpVng2YTgtNi1TcHU5Tnh4QzE3alJLaFlPcVpmcC15cXQ4dlItZHkwQVQ5bXd1dm96eWloSjBNSzRHU2U1MTEwUEczNTNHM19leXFUMl9KTlZTTlpNZmY5OUF6a2FCajhuMEROU2ZBTGNoNFpDYlQyN3U3MmR3S1hqZlA3bWRRUFJ6OHVYYVNXNnBqc1doZEctNWE1UFFKNy1KZmtlYVZyVFlpa3FPZz09"
}
Entry Information
- Entry ID: 47382
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000