# Research & Innovation | Speechmatics

Source: https://speechmatics-website-git-preview-speechmatics.vercel.app/research

Powering the world's most accurate speech technology by embedding state-of-the-art deep learning technology into Speechmatics' API.

## AI innovation in speech technology and beyond

At Speechmatics, our relentless pursuit of innovation ensures unparalleled accuracy and broad language inclusivity, redefining industry standards.

[SSL-animation layers](https://assets.ctfassets.net/yze1aysi0225/2rl2Wt8JXsFeh4FkQqRZvF/3e6e88ad255ebe139e49948a163831d2/SSL-animation_layers.lottie)

## Self-Supervised Learning — Expanding accuracy through innovation

Enabling autonomous learning to enhance our understanding

At Speechmatics, self-supervised learning (SSL) serves as a transformative approach in training speech models, harnessing unlabeled data to enhance our speech recognition systems. 

This technique allows us to autonomously identify patterns in vast amounts of data, significantly expanding the diversity of speech variations our models can learn from and improving accuracy across multiple languages.

- [Find out more](https://www.speechmatics.com/company/articles-and-news/boosting-sample-efficiency-through-self-supervised-learning)

- [FastCompany](https://www.fastcompany.com/90846670/most-innovative-companies-artificial-intelligence-2023)
- [National Innovation Awards](https://nationalinnovationawards.org.uk/)
- [Go:Tech Awards](https://www.gotechawards.co.uk/2023-winners/)

## Tirelessly pushing speech technology forward...

Throughout the years, Speechmatics has remained at the forefront of speech recognition research and innovation.
We are consistently pushing the boundaries of what's possible in speech-to-text technology. 

- years:
  - 2015:
    - Published Research: Scaling Language Models:
      - description: Our groundbreaking research, as detailed in the paper Scaling Laws for Neural Language Models, explores how the performance of neural language models improves predictably with scale. By analyzing models ranging from small to very large, we established key scaling laws that guide the efficient design and training of advanced language models, leading to significant enhancements in their accuracy and capabilities.
      - link: [Read the paper in full](https://arxiv.org/pdf/1502.00512)
  - 2019:
    - Feature: Neural Punctuation:
      - description: The first heavyweight, scaled-up punctuation model on the market for use in transcription. This advancement uses sophisticated neural networks to understand context and linguistic nuances, dramatically improving the readability and coherence of automated transcripts.
      - link: [Learn more about Neural Punctuation](https://www.speechmatics.com/company/articles-and-news/speechmatics-launches-advanced-punctuation-to-transform-the-usability-of-transcripts)
    - Published Research: Texture bias of CNNs limits few-shot classification performance:
      - description: Delving into the inherent texture bias in Convolutional Neural Networks (CNNs) and how it impacts their ability to generalize in few-shot learning scenarios. Our research highlights that while CNNs excel at recognizing textures, this bias can limit their performance when classifying new, previously unseen classes, pointing to the need for more robust approaches in few-shot learning.
      - link: [Read the paper on CNN texture bias](https://arxiv.org/pdf/1910.08519)
  - 2020:
    - Parntership: what3words:
      - description: Our collaboration with what3words marks a significant leap in voice technology - enabling users to enter precise what3words addresses through voice commands. This integration combines our advanced speech recognition capabilities with what3words' innovative geolocation system, offering seamless and accurate navigation solutions.
      - link: [Our parntership with what3words](https://www.speechmatics.com/company/articles-and-news/what3words-speechmatics-partnered-people-enter-what3words-address-by-voice)
  - 2021:
    - SSL Model: Hydra:
      - description: Our pioneering self-supervised model, Hydra, designed to harness vast amounts of unlabeled data. By leveraging advanced self-supervised learning techniques, Hydra autonomously learns from diverse datasets without the need for manual annotations.
      - link: [Find out more](https://bold-awards.com/project/hydra-project/)
    - Feature: Advanced Speaker Diarization:
      - description: This innovative approach accurately distinguishes and segments multiple speakers in audio recordings, even in complex scenarios with overlapping speech. By leveraging advanced algorithms and large-scale data processing, we have dramatically enhanced the clarity and accuracy of transcribed conversations, making it a game-changer for industries relying on precise audio analytics.
      - link: [See more about our Speaker Diarization](https://docs.speechmatics.com/features/diarization#speaker-diarization)
    - Feature: Numerical Entity Formatting:
      - description: Improvements in transcript readability, for example adding conversions from “four million pounds” to “£4m” and three sixteenths to 3/16. Whether dealing with dates, times, currencies, or other numerical data, this feature intelligently formats and standardizes numerical entities, enhancing the clarity and accuracy of your transcripts.
      - link: [Find out more](https://docs.speechmatics.com/features/entities)
  - 2022:
    - Feature: Language Identification:
      - description: Our state-of-the-art Language Identification technology sets a new standard in accurately recognizing and distinguishing between multiple languages in audio streams. Using advanced neural networks and deep learning techniques, this feature swiftly and precisely identifies the spoken language, even in challenging environments with mixed languages or dialects.
      - link: [Detect the prominent language spoken](https://docs.speechmatics.com/features-other/lang-id)
    - Feature: English Finance Language Pack:
      - description: Specifically engineered for the finance sector, this advanced language pack leverages cutting-edge speech recognition technology to accurately capture and transcribe industry-specific vocabulary.
      - link: [Unlocking accurate undestanding of financial terms](https://www.speechmatics.com/company/articles-and-news/speechmatics-unlocks-accurate-understanding-of-financial-terms)
    - Feature: Language Model Adaptation:
      - description: Offering a pioneering approach to customizing and refining language models for specific use cases. By dynamically adapting to new data and context, this feature enhances the model's ability to understand and accurately transcribe domain-specific terminology and nuanced language.
      - link: [See our full list of features](https://www.speechmatics.com/product/features-and-deployments)
  - 2023:
    - New ASR Engine: Ursa:
      - description: Our groundbreaking speech-to-text engine that sets a new benchmark in transcription accuracy. Our next-generation deep learning system trained on millions of hours of audio, is designed to capture spoken words with unparalleled precision, even in noisy or challenging environments.
      - link: [The world's most accurate speech-to-text system](https://www.speechmatics.com/company/articles-and-news/introducing-ursa-the-worlds-most-accurate-speech-to-text)
    - Feature: Bilingual English - Spanish:
      - description: A global approach to language transcription that simplifies the complexity of engaging with multilingual audiences. This feature effortlessly transcribes speech that switches between English and Spanish, ensuring seamless and accurate text output.
      - link: [Unlocking bilingual transcription](https://www.speechmatics.com/company/articles-and-news/mastering-bilingual-communication-transcribe-english-and-spanish)
    - Feature: Universal GPU transcription:
      - description: Our revolutionary approach to deploying all languages on GPU dramatically enhances the speed and efficiency of our speech recognition systems. Whether for real-time transcription, translation, or analytics, running all languages on GPU empowers businesses with faster, more reliable, and scalable language solutions.
      - link: [Discover more](https://www.speechmatics.com/company/articles-and-news/autoscaling-with-gpu-transcription-models)
    - Feature: Translation:
      - description: The most accurate real-time speech translation API, delivering highly accurate translations across a broad spectrum of languages. Whether converting speech-to-text in real-time or pre-recorded, our system ensures that the nuances and meanings are preserved, facilitating seamless and effective cross-language communication.
      - link: [Breaking down language barriers](https://www.speechmatics.com/company/articles-and-news/our-new-unified-speech-translation-api)
    - Feature: Summaries:
      - description: Harnessing the power of LLMs, Speechmatics can extract even more value from your audio data by providing insights and actionable takeaways. By automatically generating clear and accurate summaries, this technology enables users to quickly grasp the essence of lengthy texts or meetings without missing critical details.
      - link: [Leverage AI-powered Summaries](https://www.speechmatics.com/product/summarization)
    - Feature: Sentiment:
      - description: Understand the emotional tone behind the spoken word. Ground-breaking qualitative insights, all built on accurate transcription.
      - link: [Uncover emotions at scale](https://www.speechmatics.com/product/sentiment)
  - 2024:
    - Feature: Audio Events:
      - description: Improving accessibility by identifying and labelling non-speech sounds in media - helping understand more than speech. This innovation enhances the context and usability of audio data, allowing businesses to gain deeper insights and automate responses to specific audio triggers.
      - link: [Discover more about Audio Events](https://www.speechmatics.com/product/audio-events)
    - New Conversational AI: Flow:
      - description: Flow gives you best-in-class speech technology to ensure every speech interaction is finally as natural as human conversation.
      - link: [The ultimate API for voice interactions](https://www.speechmatics.com/flow)
    - New ASR Engine: Ursa 2:
      - description: 18% reduction in word error rate (WER) across 50 languages compared to our previous Ursa model, with significant gains in Vietnamese (63%), Arabic (48%), and Polish (39%).
      - link: [Accurate, real-time ASR across 50+ languages](https://www.speechmatics.com/speech-to-text)
  - 2025:
    - Feature: Speaker Diarization breakthrough:
      - description: 48% reduction in speaker identification errors at 1-second latency, and real-time speaker tracking in milliseconds. 31% more accurate speaker labels than our closest competitor.
      - link: [Discover more about Speaker Diarization](https://www.speechmatics.com/company/articles-and-news/what-is-speaker-diarization-and-why-does-it-matter-in-voice-ai)
    - Partnership: Ambarella:
      - description: A collaboration with Ambarella that brings AI-powered natural language interactions to edge applications. By combining Ambarella's edge AI SoCs with Speechmatics' speech technology, users can now experience seamless, natural device interactions; even in environments without the internet.
      - link: [Our partnership with Ambarella](https://www.speechmatics.com/company/articles-and-news/speechmatics-collaborates-with-ambarella)

## Our published research

### Hierarchical Quantized Autoencoders

##### “Hierarchical Quantized Autoencoders”, Will Williams, Sam Ringer, Tom Ash, John Hughes, David MacLeod, Jamie Dougherty. February 19, 2020.

Speechmatics’ paper was submitted and accepted to the most prestigious ML conference – NeurIPs. The paper is about a type of lossy image compression algorithm based on discrete representation learning, leading to a system that can reconstruct images of high-perceptual quality and retain semantically meaningful features despite very high compression rates.

- [Read Our Research Paper](https://arxiv.org/abs/2002.08111)

### Texture Bias Of CNNs Limits Few-Shot Classification Performance

##### “Texture Bias Of CNNs Limits Few-Shot Classification Performance”, Sam Ringer, Will Williams, Tom Ash, Remi Francis, David MacLeod. October 18, 2019.

Speechmatics published the paper at NeurIPS 2019 presenting in the meta-learning workshop.

- [Read Our Research Paper](https://arxiv.org/abs/1910.08519)

### Discriminative training of RNNLMs with the average word error criterion

##### “Discriminative training of RNNLMs with the average word error criterion”, Remi Francis, Tom Ash, Will Williams. November 8, 2020.

In this paper, Speechmatics demonstrates how you could improve recurrent neural network language models by optimizing for downstream speech recognition accuracy directly, rather than the usual generative approach which tries to model the probability of the next word in a sequence.

- [Read Our Research Paper](https://arxiv.org/pdf/1811.02528.pdf)

### The Speechmatics Parallel Corpus Filtering System for WMT18

##### “The Speechmatics Parallel Corpus Filtering System for WMT18”, Tom Ash, Remi Francis, Will Williams. Machine Translation (WMT) October 31 – November 1, 2018.

Speechmatics published the paper at Workshop on Statistical Machine Translation (WMT) 2018 and presented a translation proof of concept.

- [Read Our Research Paper](https://www.statmt.org/wmt18/pdf/WMT100.pdf)

### A Framework for Speech Recognition Benchmarking

##### “A Framework for Speech Recognition Benchmarking”, Franck Dernoncourt, Trung Bui, Walter Chang. Adobe Research. Interspeech 2018.

At Interspeech 2018 in Hyderabad Speechmatics referred to as one of the most accurate providers of ASR after some evaluations, such as one done by Adobe Research. We demonstrated that our continued focus on innovation and to drive new R&D maintains our position in a growing and increasingly challenging field.

- [Read Our Research Paper](https://www.isca-speech.org/archive/pdfs/interspeech_2018/interspeech_2018.pdf)

### Scaling Recurrent Neural Network Language Models

##### “Scaling Recurrent Neural Network Language Models”, W. Williams, N. Prasad, D. Mrva, T. Ash, A.J. Robinson. ICASSP 2015.

This is the first paper that shows that recurrent net language models scale to give very significant gains in speech recognition and it describes the most powerful models to date and some of the special methods needed to train them.

- [Read Our Research Paper](https://www.semanticscholar.org/paper/Scaling-recurrent-neural-network-language-models-Williams-Prasad/ac973bbfd62a902d073a85ca621fd297e8660a82)

4

- items:
  - Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought:
    - heading: Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought
    - subheading: James Chua, Edward Rees, Hunar Batra, Samuel R. Bowman, Julian Michael, Ethan Perez, Miles Turpin. March 8, 2024.
    - description: While chain-of-thought prompting (CoT) has the potential to improve the explainability of language model reasoning, it can systematically misrepresent the factors influencing models' behavior--for example, rationalizing answers in line with a user's opinion without mentioning this bias. To mitigate this biased reasoning problem, we introduce bias-augmented consistency training (BCT), an unsupervised fine-tuning scheme that trains models to give consistent reasoning across prompts with and without biasing features.
    - link: [Read the paper in full](https://arxiv.org/abs/2403.05518)
  - Debating with More Persuasive LLMs Leads to More Truthful Answers:
    - heading: Debating with More Persuasive LLMs Leads to More Truthful Answers
    - subheading: Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, Ethan Perez. February 9, 2024.
    - description: Common methods for aligning large language models (LLMs) with desired behaviour heavily rely on human-labelled data. However, as models grow increasingly sophisticated, they will surpass human expertise, and the role of human evaluation will evolve into non-experts overseeing experts. In anticipation of this, we ask: can weaker models assess the correctness of stronger models?
    - link: [Read the paper in full](https://arxiv.org/abs/2402.06782)
  - Hierarchical Quantized Autoencoders:
    - heading: Hierarchical Quantized Autoencoders
    - subheading: Will Williams, Sam Ringer, Tom Ash, John Hughes, David MacLeod, Jamie Dougherty. February 19, 2020.
    - description: Speechmatics’ paper was submitted and accepted to the most prestigious ML conference – NeurIPs. The paper is about a type of lossy image compression algorithm based on discrete representation learning, leading to a system that can reconstruct images of high-perceptual quality and retain semantically meaningful features despite very high compression rates.
    - link: [Read the paper in full](https://arxiv.org/abs/2002.08111)
  - Texture Bias Of CNNs Limits Few-Shot Classification Performance:
    - heading: Texture Bias Of CNNs Limits Few-Shot Classification Performance
    - subheading: Sam Ringer, Will Williams, Tom Ash, Remi Francis, David MacLeod. October 18, 2019.
    - description: Speechmatics published the paper at NeurIPS 2019 presenting in the meta-learning workshop.
    - link: [Read the paper in full](https://arxiv.org/abs/1910.08519)
  - Discriminative training of RNNLMs with the average word error criterion:
    - heading: Discriminative training of RNNLMs with the average word error criterion
    - subheading: Remi Francis, Tom Ash, Will Williams. November 8, 2020.
    - description: In this paper, Speechmatics demonstrates how you could improve recurrent neural network language models by optimizing for downstream speech recognition accuracy directly, rather than the usual generative approach which tries to model the probability of the next word in a sequence.
    - link: [Read the paper in full](https://arxiv.org/pdf/1811.02528.pdf)
  - The Speechmatics Parallel Corpus Filtering System for WMT18:
    - heading: The Speechmatics Parallel Corpus Filtering System for WMT18
    - subheading: Tom Ash, Remi Francis, Will Williams. Machine Translation (WMT) October 31 – November 1, 2018.
    - description: Speechmatics published the paper at Workshop on Statistical Machine Translation (WMT) 2018 and presented a translation proof of concept.
    - link: [Read the paper in full](https://www.statmt.org/wmt18/pdf/WMT100.pdf)
  - A Framework for Speech Recognition Benchmarking:
    - heading: A Framework for Speech Recognition Benchmarking
    - subheading: Franck Dernoncourt, Trung Bui, Walter Chang. Adobe Research. Interspeech 2018
    - description: At Interspeech 2018 in Hyderabad Speechmatics referred to as one of the most accurate providers of ASR after some evaluations, such as one done by Adobe Research. We demonstrated that our continued focus on innovation and to drive new R&D maintains our position in a growing and increasingly challenging field.
    - link: [Read the paper in full](https://www.isca-archive.org/interspeech_2018/dernoncourt18_interspeech.html)
  - Scaling Recurrent Neural Network Language Models:
    - heading: Scaling Recurrent Neural Network Language Models
    - subheading: W. Williams, N. Prasad, D. Mrva, T. Ash, A.J. Robinson. ICASSP 2015. February 2, 2015.
    - description: This is the first paper that shows that recurrent net language models scale to give very significant gains in speech recognition and it describes the most powerful models to date and some of the special methods needed to train them.
    - link: [Read the paper in full](https://www.semanticscholar.org/paper/Scaling-recurrent-neural-network-language-models-Williams-Prasad/ac973bbfd62a902d073a85ca621fd297e8660a82)
  - One billion word benchmark for measuring progress in statistical language modeling:
    - heading: One billion word benchmark for measuring progress in statistical language modeling
    - subheading: C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, A.J. Robinson. Interspeech 2014. December 10, 2013.
    - description: This paper with Google presents a standard large benchmark so that progress in language modeling may be measured. Prior to this paper there was no open, freely available corpus that was large enough to be representative for modern language modeling tasks.
    - link: [Read the paper in full](https://www.semanticscholar.org/paper/One-billion-word-benchmark-for-measuring-progress-Chelba-Mikolov/5d833331b0e22ff359db05c62a8bca18c4f04b68)
  - Connectionist Speech Recognition of Broadcast News:
    - heading: Connectionist Speech Recognition of Broadcast News
    - subheading: A. J. Robinson, G. D. Cook, D. P. W. Ellis, E. Fosler-Lussier, S. J. Renals, and D. A. G. Williams. Speech Communication, 37(1), 2002.
    - description: This paper provides an overview of the 2002 state-of-the-art methods to perform speech recognition using neural networks.
    - link: [Read the paper in full](https://www.semanticscholar.org/paper/Connectionist-speech-recognition-of-Broadcast-News-Robinson-Cook/4eb4c82eb0dfc4f1fb5ebbf76cfad0549dde3d11)
  - Recognition, indexing and retrieval of British broadcast news with the THISL system:
    - heading: Recognition, indexing and retrieval of British broadcast news with the THISL system
    - subheading: A.J. Robinson, D. Abberley, D. Kirby, and S. Renals. Proceedings of the European Conference on Speech Technology. volume 3, pages 1267–1270, September 1999.
    - description: Here we show that speech recognition can be used to find information in audio in much the same way that web pages can be found with a search engine.
    - link: [Read the paper in full](https://www.semanticscholar.org/paper/Recognition%2C-indexing-and-retrieval-of-british-news-Robinson-Abberley/2ae7a16cbcc7636af096e2dc0f2505d5d7a24d3a)
  - Time-First Search for Large Vocabulary Speech Recognition:
    - heading: Time-First Search for Large Vocabulary Speech Recognition
    - subheading: A.J. Robinson and J. Christie. ICASSP, pages 829–832, 1998.
    - description: Here we fundamentally change the main mechanism in speech recognition to make it both faster and more memory efficient (also US patent 5983180).
    - link: [Read the paper in full](https://www.semanticscholar.org/paper/Time-first-search-for-large-vocabulary-speech-Robinson-Christie/76382db33cb014ceb79ee73526c398c49afeefea)
  - Forward-Backward Retraining of Recurrent Neural Networks:
    - heading: Forward-Backward Retraining of Recurrent Neural Networks
    - subheading: A. Senior and A.J. Robinson. Advances in Neural Information Processing Systems 8, 1996.
    - description: This presents the first “end-to-end” training paper for tasks such as speech recognition.
    - link: [Read the paper in full](https://www.semanticscholar.org/paper/Forward-backward-retraining-of-recurrent-neural-Senior-Robinson/c25e9ebd8fe9d761f4738f7936ef114f7f6afe5d)
  - The Use of Recurrent Networks in Continuous Speech Recognition:
    - heading: The Use of Recurrent Networks in Continuous Speech Recognition
    - subheading: A.J. Robinson. Automatic Speech and Speaker Recognition: Advanced Topics, chapter 10.
    - description: Recurrent nets applied to large vocabulary speech recognition for the first time.
    - link: [Read the paper in full](https://www.semanticscholar.org/paper/THE-USE-OF-RECURRENT-NEURAL-NETWORKS-IN-CONTINUOUS-Robinson-Hochberg/03bc854feaee144b54924b440eff02ed9082cc6b)
  - The Application of Recurrent Nets to Phone Probability Estimation:
    - heading: The Application of Recurrent Nets to Phone Probability Estimation
    - subheading: IEEE Transactions on Neural Networks, 5(2), March 1994. A.J. Robinson.
    - description: Recurrent nets are demonstrated to give the best performing system on a well-established phoneme recognition task.
    - link: [Read the paper in full](https://www.semanticscholar.org/paper/An-application-of-recurrent-nets-to-phone-Robinson/c6629770cb6a00ad585918e71fe6dbad829ad0d1)
  - A Recurrent Error Propagation Network Speech Recognition System:
    - heading: A Recurrent Error Propagation Network Speech Recognition System
    - subheading: A.J. Robinson and F. Fallside. Computer Speech and Language, 5(3):259–274, July 1991.
    - description: The first application of recurrent nets to speech recognition.
    - link: [Read the paper in full](https://www.academia.edu/30352226/A_recurrent_error_propagation_network_speech_recognition_system)
  - Dynamic Error Propagation Networks:
    - heading: Dynamic Error Propagation Networks
    - subheading: A. J. Robinson. PhD thesis, Cambridge University Engineering Department, February 1989.
    - description: This PhD thesis introduces several key concepts of recurrent networks, several different novel architectures, the algorithms needed to train them and applications to speech recognition, coding, and reinforcement learning/game playing.
    - link: [Read the paper in full](https://www.academia.edu/30351419/Dynamic_Error_Propagation_Networks)

## Technical spotlight

### Technical — Sparse All-Reduce in PyTorch

- [Sparse All-Reduce in PyTorch](https://blog.speechmatics.com/Sparse-All-Reduce-Part-1)

The All-Reduce collective is ubiquitous in distributed training, but is currently not supported for sparse CUDA tensors in PyTorch. 

In the first part of this blog we contrast the existing alternatives available in the **Gloo**/**NCCL** backends.

- David MacLeod — Machine Learning Architect

### Speech Intelligence — The.Shed: Speechmatics Capabilities in real-time

- [The.Shed: Using Speechmatics Capabilities in real-time](https://blog.speechmatics.com/rt-speech-capabilities-demo)

Imagine being able to understand and interpret spoken language not only retrospectively, but as it happens. This isn't just a pipe dream — it's a reality we're crafting at Speechmatics.

Our mission is to deliver Speech Intelligence for the AI era, leveraging foundational speech technology and cutting-edge AI.

- Aaron Ng — Machine Learning Engineer

### Technical — An Almost Pointless Exercise in GPU Optimization

- [An Almost Pointless Exercise in GPU Optimization](https://blog.speechmatics.com/pointless-gpu-optimization-exercise)

Not everyone is able to write funky fused operators to make ML models run faster on GPUs using clever quantization tricks. However lots of developers work with algorithms that **feel** like they should be able to leverage the thousands of cores in a GPU to run faster than using the dozens of cores on a server CPU. 

To see what is possible and what is involved, I revisited the first problem I ever considered trying to accelerate with a GPU.

- Andrew Innes — Chief Architect

## **Be part of our mission**

We are actively seeking talented individuals to join our collective team of ambitious, problem solvers and throught-leaders, paving the way for inclusion in speech recognition technology.

- [See available roles](https://speechmatics-website-git-preview-speechmatics.vercel.app/company/careers/roles)
- [About Us](https://speechmatics-website-git-preview-speechmatics.vercel.app/company/about-speechmatics)
