Row 5549
Content Data
This page contains data entry 5549 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
I'm working on a Speaker Verification project wherein I'm exploring different techniques to verify the speaker via voice. The traditional approach is to extract the [MFCC](https://medium.com/@derutycsl/intuitive-understanding-of-mfccs-836d36a1f779), Filterbanks, and prosodic features. Now this method seems to be outdated as most of the research is focused on making use of pre-trained models like [Nvidia's TitaNet](https://huggingface.co/nvidia/speakerverification_en_titanet_large), [Microsoft's WavLM](https://huggingface.co/docs/transformers/en/model_doc/wavlm), SpeechBrain also a model for this. Now these pre-trained models give **Embeddings** as output which represent the speaker's voice regardless of what he said in the recording.
Now my doubt is what do these **Embeddings** represent? One of the architecture's makes use of MFCC's and later passes them to NN like LSTM to capture the pattern.
| Field | Value |
|---|---|
| text | I'm working on a Speaker Verification project wherein I'm exploring different techniques to verify the speaker via voice. The traditional approach is to extract the [MFCC](https://medium.com/@derutycsl/intuitive-understanding-of-mfccs-836d36a1f779), Filterbanks, and prosodic features. Now this method seems to be outdated as most of the research is focused on making use of pre-trained models like [Nvidia's TitaNet](https://huggingface.co/nvidia/speakerverification_en_titanet_large), [Microsoft's … |
| label | r/deeplearning |
| dataType | post |
| communityName | r/deeplearning |
| datetime | 2024-05-01 |
| username_encoded | Z0FBQUFBQm5Lakwybk8zZEw0UmF2QzBKSXVubTE2QzJsWVhocFBKRzVVS0JRWWkzRFVYWnZVcVVVUVZmT3psOTkwcmlJVVpzZ1YwTnpMSHhmLXRPeXA2VjA4NjFUUTdWazk1WkNMOEl3M1BaMVpyek4tbDlZV1U9 |
| url_encoded | Z0FBQUFBQm5Lak9HeW9Icnlueko2X3FMV0VIdVZMbEE0VEU5WjZMTm1YTDBKT2hTVVFJcG5nUk9NUlYyd0dvZlZLYWNrc1ZQNExuZUVuLVFOcDg0NFJhZlRzUmtWeHVaLVl3TGxPekFwc1N4XzBua244UjNkN1ZOMjctMTNQbk43SGNJbFc3NmFFaklmeWdKclRpdGg5S0p3Q29jRnY2ZHZ5Vi1pODBfMFFUVTZWallPejhvdnQ3MVQ1ZEtsTDN1Z1Mwbk9zNjUwaGNTenBvTUVfN0xJR2VCaUtaa2VIV1lIUT09 |
Raw Record
{
"text": "I'm working on a Speaker Verification project wherein I'm exploring different techniques to verify the speaker via voice. The traditional approach is to extract the [MFCC](https://medium.com/@derutycsl/intuitive-understanding-of-mfccs-836d36a1f779), Filterbanks, and prosodic features. Now this method seems to be outdated as most of the research is focused on making use of pre-trained models like [Nvidia's TitaNet](https://huggingface.co/nvidia/speakerverification_en_titanet_large), [Microsoft's WavLM](https://huggingface.co/docs/transformers/en/model_doc/wavlm), SpeechBrain also a model for this. Now these pre-trained models give **Embeddings** as output which represent the speaker's voice regardless of what he said in the recording.\n\nNow my doubt is what do these **Embeddings** represent? One of the architecture's makes use of MFCC's and later passes them to NN like LSTM to capture the pattern.",
"label": "r/deeplearning",
"dataType": "post",
"communityName": "r/deeplearning",
"datetime": "2024-05-01",
"username_encoded": "Z0FBQUFBQm5Lakwybk8zZEw0UmF2QzBKSXVubTE2QzJsWVhocFBKRzVVS0JRWWkzRFVYWnZVcVVVUVZmT3psOTkwcmlJVVpzZ1YwTnpMSHhmLXRPeXA2VjA4NjFUUTdWazk1WkNMOEl3M1BaMVpyek4tbDlZV1U9",
"url_encoded": "Z0FBQUFBQm5Lak9HeW9Icnlueko2X3FMV0VIdVZMbEE0VEU5WjZMTm1YTDBKT2hTVVFJcG5nUk9NUlYyd0dvZlZLYWNrc1ZQNExuZUVuLVFOcDg0NFJhZlRzUmtWeHVaLVl3TGxPekFwc1N4XzBua244UjNkN1ZOMjctMTNQbk43SGNJbFc3NmFFaklmeWdKclRpdGg5S0p3Q29jRnY2ZHZ5Vi1pODBfMFFUVTZWallPejhvdnQ3MVQ1ZEtsTDN1Z1Mwbk9zNjUwaGNTenBvTUVfN0xJR2VCaUtaa2VIV1lIUT09"
}
Entry Information
- Entry ID: 5549
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000