September 13, 2019
We propose a fully convolutional sequence-to-sequence encoder architecture with a simple and efficient decoder. Our model improves WER on LibriSpeech while being an order of magnitude more efficient than a strong RNN baseline. Key to our approach is a time-depth separable convolution block which dramatically reduces the number of parameters in the model while keeping the receptive field large. We also give a stable and efficient beam search inference procedure which allows us to effectively integrate a language model. Coupled with a convolutional language model, our time-depth separable convolution architecture improves by more than 22% relative WER over the best previously reported sequence-to-sequence results on the noisy LibriSpeech test set.
Publisher
Interspeech
October 02, 2026
Aykut Arslan
October 02, 2026
October 02, 2026
Andres Barei Bueno
October 02, 2026
October 02, 2026
Anindya Dey, Gabriel Herczeg, An Huang, Nicolas Jaramillo Torres, Jacob H. Swenberg
October 02, 2026
October 02, 2026
Joseph Phillip Brennan, Milana Golich
October 02, 2026
