Why predicting durations as well as tokens allows transducer models to skip frames and achieve up to 2.82X faster inference.
Page 2 of 6