RESEARCH

SPEECH & AUDIO

Sequence-to-Sequence Speech Recognition with Time-Depth Separable Convolutions

September 13, 2019

Abstract

We propose a fully convolutional sequence-to-sequence encoder architecture with a simple and efficient decoder. Our model improves WER on LibriSpeech while being an order of magnitude more efficient than a strong RNN baseline. Key to our approach is a time-depth separable convolution block which dramatically reduces the number of parameters in the model while keeping the receptive field large. We also give a stable and efficient beam search inference procedure which allows us to effectively integrate a language model. Coupled with a convolutional language model, our time-depth separable convolution architecture improves by more than 22% relative WER over the best previously reported sequence-to-sequence results on the noisy LibriSpeech test set.

Download the Paper

AUTHORS

Written by

Awni Hannun

Ann Lee

Qiantong Xu

Ronan Collobert

Publisher

Interspeech

Related Publications

October 02, 2026

RESEARCH

Tightness of the Cycle-Based Relaxation for Completed Length-Three Alpha-Cycles

Aykut Arslan

October 02, 2026

October 02, 2026

RESEARCH

On Solvable Evolution Algebras and a Conjecture by García-Martínez and Pérez-Rodríguez

Andres Barei Bueno

October 02, 2026

October 02, 2026

RESEARCH

String Two-Point Function = Height Function on a Curve

Anindya Dey, Gabriel Herczeg, An Huang, Nicolas Jaramillo Torres, Jacob H. Swenberg

October 02, 2026

October 02, 2026

RESEARCH

Semiabelian Groups Need Not Be Monomial

Joseph Phillip Brennan, Milana Golich

October 02, 2026

Help Us Pioneer The Future of AI

We share our open source frameworks, tools, libraries, and models for everything from research exploration to large-scale production deployment.