Row 47382

Row ID: 47382 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 47382 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

**From the article:**

Soon after OpenAI released GPT-4o on Monday, May 13, some Chinese speakers started to notice that something seemed off about this newest version of the chatbot: the tokens it uses to parse text were full of spam and porn phrases.

On May 14, Tianle Cai, a PhD student at Princeton University studying inference efficiency in large language models like those that power such chatbots, accessed GPT-4o’s public token library and pulled a list of the 100 longest Chinese tokens the model uses to parse and compress Chinese prompts.

Humans read in words, but LLMs read in tokens, which are distinct units in a sentence that have consistent and significant meanings. Besides dictionary words, they also include suffixes, common expressions, names, and more. The more tokens a model encodes, the faster the model can “read” a sentence and the less computing power it consumes, thus making the response cheaper.

Of the 100 results, only three of them are common enough to be used in everyday conversations; everything else consisted of words and expressions used specifically in the contexts of either gambling or pornography. The longest token, lasting 10.5 Chinese characters, literally means “\_free Japanese porn video to watch.” Oops.

FieldValue
text **From the article:** Soon after OpenAI released GPT-4o on Monday, May 13, some Chinese speakers started to notice that something seemed off about this newest version of the chatbot: the tokens it uses to parse text were full of spam and porn phrases. On May 14, Tianle Cai, a PhD student at Princeton University studying inference efficiency in large language models like those that power such chatbots, accessed GPT-4o’s public token library and pulled a list of the 100 longest Chinese tokens th…
label r/artificial
dataType comment
communityName r/artificial
datetime 2024-05-22
username_encoded Z0FBQUFBQm5Lak1RVGwtcFN3d3kyUWJHUlItLWxTZHVuNlV6RE5GSnZTek5UQUNiQXZjcnNSNnVINU92RVFUUFFOTExUVEVOT2pPMEZvd0wyTTdxQUs3TlRtVGZTRVZRUXc9PQ==
url_encoded Z0FBQUFBQm5Lak9naFFpd2JDRFpFNV9aTTNIdVR2eFN2VE16Y1VlN1lhelFkTEhTUnVYZ1kzZEtGR3hpVng2YTgtNi1TcHU5Tnh4QzE3alJLaFlPcVpmcC15cXQ4dlItZHkwQVQ5bXd1dm96eWloSjBNSzRHU2U1MTEwUEczNTNHM19leXFUMl9KTlZTTlpNZmY5OUF6a2FCajhuMEROU2ZBTGNoNFpDYlQyN3U3MmR3S1hqZlA3bWRRUFJ6OHVYYVNXNnBqc1doZEctNWE1UFFKNy1KZmtlYVZyVFlpa3FPZz09

Raw Record

{
  "text": "**From the article:**\n\nSoon after OpenAI released GPT-4o on Monday, May 13, some Chinese speakers started to notice that something seemed off about this newest version of the chatbot: the tokens it uses to parse text were full of spam and porn phrases.\n\nOn May 14, Tianle Cai, a PhD student at Princeton University studying inference efficiency in large language models like those that power such chatbots, accessed GPT-4o’s public token library and pulled a list of the 100 longest Chinese tokens the model uses to parse and compress Chinese prompts. \n\nHumans read in words, but LLMs read in tokens, which are distinct units in a sentence that have consistent and significant meanings. Besides dictionary words, they also include suffixes, common expressions, names, and more. The more tokens a model encodes, the faster the model can “read” a sentence and the less computing power it consumes, thus making the response cheaper.\n\nOf the 100 results, only three of them are common enough to be used in everyday conversations; everything else consisted of words and expressions used specifically in the contexts of either gambling or pornography. The longest token, lasting 10.5 Chinese characters, literally means “\\_free Japanese porn video to watch.” Oops.",
  "label": "r/artificial",
  "dataType": "comment",
  "communityName": "r/artificial",
  "datetime": "2024-05-22",
  "username_encoded": "Z0FBQUFBQm5Lak1RVGwtcFN3d3kyUWJHUlItLWxTZHVuNlV6RE5GSnZTek5UQUNiQXZjcnNSNnVINU92RVFUUFFOTExUVEVOT2pPMEZvd0wyTTdxQUs3TlRtVGZTRVZRUXc9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9naFFpd2JDRFpFNV9aTTNIdVR2eFN2VE16Y1VlN1lhelFkTEhTUnVYZ1kzZEtGR3hpVng2YTgtNi1TcHU5Tnh4QzE3alJLaFlPcVpmcC15cXQ4dlItZHkwQVQ5bXd1dm96eWloSjBNSzRHU2U1MTEwUEczNTNHM19leXFUMl9KTlZTTlpNZmY5OUF6a2FCajhuMEROU2ZBTGNoNFpDYlQyN3U3MmR3S1hqZlA3bWRRUFJ6OHVYYVNXNnBqc1doZEctNWE1UFFKNy1KZmtlYVZyVFlpa3FPZz09"
}

Entry Information