Error-Driven Fixed-Budget Asr Personalization For Accented Speakers

Abhijeet Awasthi, Aman Kansal, Sunita Sarawagi, Preethi Jyothi

DOI

SPS

Members: Free
IEEE Members: $11.00
Non-members: $15.00

Length: 00:12:21

11 Jun 2021

We consider the task of personalizing ASR models while being constrained by a fixed budget on recording speaker specific utterances. Given a speaker and an ASR model, we propose a method of identifying sentences for which a speaker's utterances are likely to be harder for the given ASR model to recognize. We assume a tiny amount of speaker-specific data to learn phoneme-level error models which help us select such sentences. We show that speaker's utterances on the sentences selected using our error model indeed have larger error rates when compared to speaker's utterances on randomly selected sentences. We find that fine-tuning the ASR model on the sentence utterances selected with the help of error models yield higher WER improvements in comparison to fine-tuning on an equal number of randomly selected sentence utterances. Thus our method provides an efficient way of collecting speaker utterances under budget constraints for personalizing ASR models.

Chairs:

Shinji Watanabe

Tags:

signal processing society

IEEE icassp 2021

virtual conference

2021

sps

virtual conference icassp 2021

june 6-11 2021

icassp 2021

Error-Driven Fixed-Budget Asr Personalization For Accented Speakers

Abhijeet Awasthi, Aman Kansal, Sunita Sarawagi, Preethi Jyothi

Value-Added Bundle(s) Including this Product

ICASSP 2021 Virtual Conference - Presentation Videos Product Bundle

More Like This

To What Extent Can Plug-And-Play Methods Outperform Neural Networks Alone In Low-Dose Ct Reconstruction

Mc-Pdnet: Deep Unrolled Neural Network For Multi-Contrast Mr Image Reconstruction From Undersampled K-Space Data

Dstunet: Unet With Efficient Dense Swin Transformer Pathway For Medical Image Segmentation

Join an IEEE Society