Robustness Analysis of TANDEM-Based Speech Recognition System
Analýza robustnosti moderních rozpoznávačů řeči na bázi TANDEM architektury
Authors
Supervisors
Reviewers
Editors
Other contributors
Journal Title
Journal ISSN
Volume Title
Publisher
České vysoké učení technické v Praze
Czech Technical University in Prague
Czech Technical University in Prague
Date of defense
Abstract
Tato práce se zabývá analýzou robustnosti rozpoznávače řeči na bázi TANDEM architektury. Cílem je zjistit, jaký vliv na úspěšnost rozpoznávání mají různé varianty příznakových vektorů, s užším zaměřením na příznaky odhadovanými vícevrstvými sítěmi. K implementaci je použit široce používaný balíček nástrojů Kaldi. Pro splnění cíle práce byl vytvořen tzv. recept, který využívá zavedených konvencí Kaldi nástrojů k modulárnímu sestavení experimentů. Základním zdrojem řečových signálů je databáze SPEECON, která obsahuje signály nahrávané v různých prostředích čtyřmi mikrofony. Pro každé prostředí jsou tedy dostupná data ze čtyř různě kvalitních kanálů. Robustnost je testována na všech dostupných prostředích databáze SPEECON. Pro většinu prostředí bylo dosaženo uspokojivých výsledků, kde se TANDEM systém ukázal jako robustnější a úspěšnější než standardní řešení a to v průměru o přibližně 5 %.
This paper deals with the analysis of robustness of a speech recognizer based on the TANDEM architecture. The main goal is to find out which types of the TANDEM architecture feature vectors improve the classification accuracy. The influence of the multi-layer artificial neural network feature vectors is observed in more detail. The implementation is based on the free, widely spread tool called Kaldi and based on its conventions, the Kaldi recipe was created. The main source of data is the SPEECON database which contains the signals recorded in the different environments with the four microphone channels of a different quality of recording. The robustness is tested on all the available environments of the SPEECON database. The satisfying results were achieved for the most of the environments where the TANDEM architecture outperformed the standard approach about 5% WER on average.
This paper deals with the analysis of robustness of a speech recognizer based on the TANDEM architecture. The main goal is to find out which types of the TANDEM architecture feature vectors improve the classification accuracy. The influence of the multi-layer artificial neural network feature vectors is observed in more detail. The implementation is based on the free, widely spread tool called Kaldi and based on its conventions, the Kaldi recipe was created. The main source of data is the SPEECON database which contains the signals recorded in the different environments with the four microphone channels of a different quality of recording. The robustness is tested on all the available environments of the SPEECON database. The satisfying results were achieved for the most of the environments where the TANDEM architecture outperformed the standard approach about 5% WER on average.
Description
Citation
Underlying research data set URL
Permanent link
Rights/License
A university thesis is a work protected by the Copyright Act of the Czech Republic. Extracts, copies and transcripts of the thesis are allowed for personal use only and at one`s own expense. The use of thesis should be in compliance with the Copyright Act.
Vysokoškolská závěrečná práce je dílo chráněné autorským zákonem. Je možné pořizovat z něj na své náklady a pro svoji osobní potřebu výpisy, opisy a rozmnoženiny. Jeho využití musí být v souladu s autorským zákonem v platném znění.
Vysokoškolská závěrečná práce je dílo chráněné autorským zákonem. Je možné pořizovat z něj na své náklady a pro svoji osobní potřebu výpisy, opisy a rozmnoženiny. Jeho využití musí být v souladu s autorským zákonem v platném znění.