Projects:2018s1-103 Improving Usability and User Interaction with KALDI Open-Source Speech Recogniser

Project Team

Students

Shi Yik Chin
Yasasa Saman Tennakoon

Supervisors

Dr. Said Al-Sarawi
Dr. Ahmad Hashemi-Sakhtsari (DST Group)

Abstract

This project aims to refine and improve the capabilities of KALDI (an Open Source Speech Recogniser). This will require:

Improving the current GUI's flexibility
Introducing new elements or replacing older elements in the GUI for ease of use
Including a methodology that users (of any skill level) can use to improve or introduce Language or Acoustic models into the software
Refining current Language and Acoustic models in the software to reduce the Word Error Rate (WER)
Introducing a neural network in the software to reduce the Word Error Rate (WER)
Introducing a feedback loop into the software to reduce the Word Error Rate (WER)
Introducing Binarized Neural Networks into the training methods to reduce training times and increase efficiency

This project will involve the use of Deep Learning algorithms (Automatic Speech Recognition related), software development (C++) and performance evaluation through the Word Error Rate formula. Very little hardware will be involved through its entirety.

Projects:2018s1-103 Improving Usability and User Interaction with KALDI Open-Source Speech Recogniser

Contents

Project Team

Students

Supervisors

Abstract

Introduction

Background

Research and Development

Results

Conclusion

Navigation menu

Personal tools

Namespaces

Variants

Views

More

Search

Navigation

Tools