Row 44674

Row ID: 44674 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 44674 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

While we are still figuring out ways to improve the agent's emotional response to OpenAI GPT-4o, we have already made significant progress in aligning OpenAI's TTS performance. To begin this experiment, we collected 10 hours of OpenAI TTS data to perform supervised fine-tuning (SFT) on both the LLM (medium) and VITS models, which took approximately 30 minutes. After that, we used 15 seconds of audio as a prompt during inference.

Demos Available: [here](https://firefly-ai.notion.site/OpenAI-Examples-34975ae263a9496c84e89fb7b1ea25a4?pvs=4).

As you can see, the model's emotion, rhythm, accent, and timbre match the OpenAI speakers, though there is some degradation in audio quality, which we are working on. To avoid any legal issues, we are unable to release the fine-tuned model, but I believe everyone can tune [fish-speech](https://github.com/fishaudio/fish-speech) to this level within hours and for around $20.

Our experiment shows that with only 25 seconds of prompts (few-shot learning), without any fine-tuning, the model can mimic most behaviors except details like timbre and how it reads numbers. To the best of our knowledge, you can clone how someone speaks in English, Chinese, and Japanese with 30 minutes of data using this framework.

Repo: [https://github.com/fishaudio/fish-speech](https://github.com/fishaudio/fish-speech)

FieldValue
text While we are still figuring out ways to improve the agent's emotional response to OpenAI GPT-4o, we have already made significant progress in aligning OpenAI's TTS performance. To begin this experiment, we collected 10 hours of OpenAI TTS data to perform supervised fine-tuning (SFT) on both the LLM (medium) and VITS models, which took approximately 30 minutes. After that, we used 15 seconds of audio as a prompt during inference. Demos Available: [here](https://firefly-ai.notion.site/OpenAI-Exam…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-05-22
username_encoded Z0FBQUFBQm5Lak1PRjBSaldpOENBVndTbzcxZ2FOZzRndGEtMEdmbnFNQjdlWUhEcnpjVXl4OXRhclBmbnZuLWdtdGxhZHN5VHVNc3puS3NUVl9CSmdqSDZmOUdHdDlGd3c9PQ==
url_encoded Z0FBQUFBQm5Lak9lUjNGODJkNjRnMy1BLWd0UTQwdjlXTjZBV3ZvUmF5eTVDdk43YzFvdm9KRnlobGhFbHdPdVctMURMemtpb0p2SXIwaGFZU3l0Q1ZlZUNfNEh5ODVDa1lkTVcyQnFIUzJVbE1ra016RjdNT3JfZ3BpZE1tYzJZV0wxRzRrQXY2ZzZUSWlPSWZBRTlKUUo4YlVRQ3lYTmNVNlRQN0dvYU9hNF9mWHhncExNRFBpdElwcUJKS24zbjVzY3dydi0wVHA5WFpMcE15dl96SWhJLTZVS205MHlxZz09

Raw Record

{
  "text": "While we are still figuring out ways to improve the agent's emotional response to OpenAI GPT-4o, we have already made significant progress in aligning OpenAI's TTS performance. To begin this experiment, we collected 10 hours of OpenAI TTS data to perform supervised fine-tuning (SFT) on both the LLM (medium) and VITS models, which took approximately 30 minutes. After that, we used 15 seconds of audio as a prompt during inference.\n\nDemos Available: [here](https://firefly-ai.notion.site/OpenAI-Examples-34975ae263a9496c84e89fb7b1ea25a4?pvs=4).\n\nAs you can see, the model's emotion, rhythm, accent, and timbre match the OpenAI speakers, though there is some degradation in audio quality, which we are working on. To avoid any legal issues, we are unable to release the fine-tuned model, but I believe everyone can tune [fish-speech](https://github.com/fishaudio/fish-speech) to this level within hours and for around $20.\n\nOur experiment shows that with only 25 seconds of prompts (few-shot learning), without any fine-tuning, the model can mimic most behaviors except details like timbre and how it reads numbers. To the best of our knowledge, you can clone how someone speaks in English, Chinese, and Japanese with 30 minutes of data using this framework.\n\n  \nRepo: [https://github.com/fishaudio/fish-speech](https://github.com/fishaudio/fish-speech)",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-22",
  "username_encoded": "Z0FBQUFBQm5Lak1PRjBSaldpOENBVndTbzcxZ2FOZzRndGEtMEdmbnFNQjdlWUhEcnpjVXl4OXRhclBmbnZuLWdtdGxhZHN5VHVNc3puS3NUVl9CSmdqSDZmOUdHdDlGd3c9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9lUjNGODJkNjRnMy1BLWd0UTQwdjlXTjZBV3ZvUmF5eTVDdk43YzFvdm9KRnlobGhFbHdPdVctMURMemtpb0p2SXIwaGFZU3l0Q1ZlZUNfNEh5ODVDa1lkTVcyQnFIUzJVbE1ra016RjdNT3JfZ3BpZE1tYzJZV0wxRzRrQXY2ZzZUSWlPSWZBRTlKUUo4YlVRQ3lYTmNVNlRQN0dvYU9hNF9mWHhncExNRFBpdElwcUJKS24zbjVzY3dydi0wVHA5WFpMcE15dl96SWhJLTZVS205MHlxZz09"
}

Entry Information