Row 7213

Row ID: 7213 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 7213 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

OpenAI is wrong. Their claim of supporting over 90 languages with their Whisper module is inaccurate. Here is the proof 👇

Last year, I developed [ToText](https://totext.ai/), a free online transcription service using the Whisper module, which is an AI-based open-source speech-to-text module developed by OpenAI.

My aim was/is to provide non-technical users with an easier and smoother transcription service without the need for coding. However, shortly after its launch, I began receiving negative feedback from users regarding the transcription accuracy of various languages. Some languages were performing poorly, and others weren't functioning at all.

Testing each language integrated into the ToText platform became imperative. To achieve this, I proposed a survey study to the capstone students in my department. Fortunately, it was selected by a capstone team (shown in the picture), and I started supervising those students as they conducted a survey of transcription accuracy for 98 languages included in ToText.

These students did an exceptional job and obtained significant results. One of them was the disproval of OpenAI's claim of supporting over 90 languages. In reality, the critical question to ask is, "What level of transcription accuracy does the whisper module provide for each language?" If nearly half of these languages are transcribed poorly, is it accurate to claim support for them?

Yes, this is what happened to ToText. I had to remove 48 languages out of 99 languages from ToText and only 51 languages were retained for user access.

Whisper comes in various sizes such as tiny, base, small, medium, and large. ToText currently uses the base size (trained with 74 million parameters). While OpenAI could argue that their claim refers to larger sizes like the large size (trained with 1.5 billion parameters), there has been no clear statement from OpenAI regarding this.

# Survey Results

https://preview.redd.it/zq6kre0zgl0d1.jpg?width=1140&format=pjpg&auto=webp&s=1a325620a1f270bc9c6bee2523bd2f6a3d0d4427

Here is the summary of these results:

* 2 languages had an average score of 5, which is excellent (perfect transcription). * 10 languages had an average of 4 which is very good (very correct transcription). * 15 languages received an average between 3 and 4 which is good (correct transcription). * 24 languages obtained an average score between 2 and 3 which is average (medium transcription). * 33 languages received an average score between 1-2 meaning the transcriptions were minimally correct (poor transcription). * The rest of languages had an average score below 1, meaning the transcriptions made no sense at all (terrible transcription). * 1 language (Hindi) would not transcribe but translate instead.

# Final Thoughts

Whisper (base size) is a good tool for homogeneous languages, especially for romance languages known as the Latin or Neo-Latin languages. Many times for languages that are not based in Latin or don’t have a similar alphabet to it, the model will just return a phonetic transcription which is much less useful. It is possible that some tweaking needs to be done so the model can have a better definition of what a transcription actually is. Whisper is fine for personal use for most people who reside in a Western country but for larger-scale projects, it would need a lot of work, as it is not perfect even for the romance languages.

These results could be beneficial for OpenAI for improving their whisper module to have a better transcription service, especially for those low-performing languages.

If you're interested in learning more about this survey, you can visit [this blog article](https://blog.merjek.com/2024/05/14/openai-is-wrong-they-do-not-support-over-90-languages-with-their-whisper-module/).

Let me know about your opinions about the whisper module.

FieldValue
text OpenAI is wrong. Their claim of supporting over 90 languages with their Whisper module is inaccurate. Here is the proof 👇 Last year, I developed [ToText](https://totext.ai/), a free online transcription service using the Whisper module, which is an AI-based open-source speech-to-text module developed by OpenAI. My aim was/is to provide non-technical users with an easier and smoother transcription service without the need for coding. However, shortly after its launch, I began receiving negativ…
label r/gpt3
dataType post
communityName r/GPT3
datetime 2024-05-15
username_encoded Z0FBQUFBQm5LakwzeHhsOEJod0dSTklmZHZkX1lqSTdjdi0yeHFhMkJuZTQ4MHRFaFVTd3pnLWh5ck03eVcwUmxYcWs1VDhWcTI4SmQ5TlR4MGV1S2dWemdOakcwcEFqZlE9PQ==
url_encoded Z0FBQUFBQm5Lak9IZGViekIxc3Y4SVN5V0pkeUxrOS1iWjhkbzVqeFlJRjI5ai1UNFprQm1oUjJNeUtyZktwT01WT1AzcGZ6NFlWcnZ5MTJ6UHltd0NIY2dEMnE2eUdTbWNzQlpLQ0hjNzRfSWZtd3BmY3p2cWFoU19XN1pmMzNkSVRLaC15VTg0VUo3cTZyODc3ck13WWZraEFLSlFZLW9CWlplenBOOTlMZk05cTZLZVl6UnpDemRnVmlVV2VpWEQzN21rVGdRVi1D

Raw Record

{
  "text": "OpenAI is wrong. Their claim of supporting over 90 languages with their Whisper module is inaccurate. Here is the proof 👇\n\nLast year, I developed [ToText](https://totext.ai/), a free online transcription service using the Whisper module, which is an AI-based open-source speech-to-text module developed by OpenAI.\n\nMy aim was/is to provide non-technical users with an easier and smoother transcription service without the need for coding. However, shortly after its launch, I began receiving negative feedback from users regarding the transcription accuracy of various languages. Some languages were performing poorly, and others weren't functioning at all.\n\nTesting each language integrated into the ToText platform became imperative. To achieve this, I proposed a survey study to the capstone students in my department. Fortunately, it was selected by a capstone team (shown in the picture), and I started supervising those students as they conducted a survey of transcription accuracy for 98 languages included in ToText.\n\nThese students did an exceptional job and obtained significant results. One of them was the disproval of OpenAI's claim of supporting over 90 languages. In reality, the critical question to ask is, \"What level of transcription accuracy does the whisper module provide for each language?\" If nearly half of these languages are transcribed poorly, is it accurate to claim support for them?\n\nYes, this is what happened to ToText. I had to remove 48 languages out of 99 languages from ToText and only 51 languages were retained for user access.\n\nWhisper comes in various sizes such as tiny, base, small, medium, and large. ToText currently uses the base size (trained with 74 million parameters). While OpenAI could argue that their claim refers to larger sizes like the large size (trained with 1.5 billion parameters), there has been no clear statement from OpenAI regarding this.\n\n# Survey Results\n\nhttps://preview.redd.it/zq6kre0zgl0d1.jpg?width=1140&format=pjpg&auto=webp&s=1a325620a1f270bc9c6bee2523bd2f6a3d0d4427\n\nHere is the summary of these results:\n\n* 2 languages had an average score of 5, which is excellent (perfect transcription).\n* 10 languages had an average of 4 which is very good (very correct transcription).\n* 15 languages received an average between 3 and 4 which is good (correct transcription).\n* 24 languages obtained an average score between 2 and 3 which is average (medium transcription).\n* 33 languages received an average score between 1-2 meaning the transcriptions were minimally correct (poor transcription).\n* The rest of languages had an average score below 1, meaning the transcriptions made no sense at all (terrible transcription).\n* 1 language (Hindi) would not transcribe but translate instead.\n\n# Final Thoughts\n\nWhisper (base size) is a good tool for homogeneous languages, especially for romance languages known as the Latin or Neo-Latin languages. Many times for languages that are not based in Latin or don’t have a similar alphabet to it, the model will just return a phonetic transcription which is much less useful. It is possible that some tweaking needs to be done so the model can have a better definition of what a transcription actually is. Whisper is fine for personal use for most people who reside in a Western country but for larger-scale projects, it would need a lot of work, as it is not perfect even for the romance languages.\n\nThese results could be beneficial for OpenAI for improving their whisper module to have a better transcription service, especially for those low-performing languages.\n\nIf you're interested in learning more about this survey, you can visit [this blog article](https://blog.merjek.com/2024/05/14/openai-is-wrong-they-do-not-support-over-90-languages-with-their-whisper-module/).\n\nLet me know about your opinions about the whisper module.",
  "label": "r/gpt3",
  "dataType": "post",
  "communityName": "r/GPT3",
  "datetime": "2024-05-15",
  "username_encoded": "Z0FBQUFBQm5LakwzeHhsOEJod0dSTklmZHZkX1lqSTdjdi0yeHFhMkJuZTQ4MHRFaFVTd3pnLWh5ck03eVcwUmxYcWs1VDhWcTI4SmQ5TlR4MGV1S2dWemdOakcwcEFqZlE9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9IZGViekIxc3Y4SVN5V0pkeUxrOS1iWjhkbzVqeFlJRjI5ai1UNFprQm1oUjJNeUtyZktwT01WT1AzcGZ6NFlWcnZ5MTJ6UHltd0NIY2dEMnE2eUdTbWNzQlpLQ0hjNzRfSWZtd3BmY3p2cWFoU19XN1pmMzNkSVRLaC15VTg0VUo3cTZyODc3ck13WWZraEFLSlFZLW9CWlplenBOOTlMZk05cTZLZVl6UnpDemRnVmlVV2VpWEQzN21rVGdRVi1D"
}

Entry Information