# AI Speech Technology | Speech APIs powering Voice AI

Source: https://speechmatics-website-git-preview-speechmatics.vercel.app/

A speech-to-text API for real conversations, not clean demos. Build Voice AI that understands accents, multiple speakers, and multilingual speech.

## Homepage hero

### Speech APIs powering Voice AI

#### Low-latency speech-to-text for multilingual, multi-speaker conversations

## Innovative Companies

Powering the world's best companies

### Logo AI Media

- [How AI-Media are advancing real-time captioning](https://www.speechmatics.com/product/case-studies/transforming-live-captioning-how-ai-media-are-advancing-real-time)

AI Media

Delivering 120X more with voice AI

Powering live content through AI-powered transcription, built on industry-leading voice AI

- Pipecat

- mediQuo logo

### LiveKit logo

- [Livekit link](https://www.speechmatics.com/company/articles-and-news/build-ai-agents-that-understand-who-said-what-livekit)

LiveKit

Enabling 100,000+ developers with leading speech recognition

Pairing LiveKit’s flexible agent framework with Speechmatics to build world-class agents

- Adobe logo

- NVidia Inception Program

- Vapi

- Humetrix

### NCI

- [NCI case study](https://www.speechmatics.com/product/case-studies/nci)

Redefining real-time captioning

How NCI delivered a 99% increase in usage of automated captioning

- Veritone Logo

### Media Track

- [Media Track case study](https://www.speechmatics.com/product/case-studies/media-track-enhances-global-media-monitoring-with-speechmatics)

Delivering a 20% leap in accuracy improvements

Improved transcription performance across more than 20 languages for their global clients

- Content Guru logo

- Docbuddy

- Cekura logo

- Jambonz

- Zylinc logo

### Prosodica

- [Prosodica case study](https://www.speechmatics.com/product/case-studies/vail-systems-prosodica)

Driving better conversations at scale

Leveraging speech recognition to track customer interactions, highlight key insights, and raise contact center performance

- Ubisoft

- Edvak logo

- ACA Group

- Stenoly

- Nabla logo

- Speech Intelligence - 3playmedia

- Vodex logo

## Accurate. Secure. Global.

Speech technology **built for companies** with global reach and uncompromising standards for quality.

### For use cases that can't wait

Live transcription

Real-time speech-to-text is here

High accuracy and low latency. STT in less than 1 second, without compromising accuracy and understanding.

- [Real-Time](https://speechmatics-website-git-preview-speechmatics.vercel.app/product/real-time)

### STT you can trust

Secure

Deploy anywhere, no data logging

Run Speechmatics on device, on prem and in the cloud depending on your privacy needs. We don’t log your data as standard.

- [Features and deployments](https://speechmatics-website-git-preview-speechmatics.vercel.app/product/features-and-deployments)

### Find new markets

Languages

56+ languages

We cover over half the world's population with our language coverage, helping businesses expand globally.

- [Languages](https://speechmatics-website-git-preview-speechmatics.vercel.app/languages)

## Every voice, every industry

### Voice AI that works where it matters most

Use Cases

From healthcare to live media, Speechmatics delivers real-world Speech APIs with low latency, multilingual capabilities, and built for scale.

### Medical & healthcare

MedTech

Support ambient scribe and dictation with our Medical Model, cutting errors on key terms by up to 50%.

- [Read more](https://www.speechmatics.com/use-cases/medical-transcription)

### Voice agent builders

AI voice agents

Sub-second, speaker-aware STT and TTS across 56+ languages. Plug in fast with a flexible API and native integrations to power AI voice agents.

- [AI Voice Agents](https://www.speechmatics.com/use-cases/ai-voice-agents)

### Live captioning

Media & Broadcast

Deliver accurate captions for live events, sports, and news — real time, at scale, and with accuracy that holds up in the spotlight.

- [Media Distribution & Captioning](https://speechmatics-website-git-preview-speechmatics.vercel.app/use-cases/media-distribution-and-captioning)

### Contact center analytics

CCaaS

Reduce wait times, increase agent productivity and improve customer experience in contact centers with Speechmatics' voice AI.

- [Contact Center Solutions](https://speechmatics-website-git-preview-speechmatics.vercel.app/use-cases/contact-center-solutions)

### Legal transcription

Courtroom

Speech recognition built for court reporters, legal professionals, and law firms who need unmatched accuracy across every accent and speaker — in real time.

- [Legal Transcription](https://speechmatics-website-git-preview-speechmatics.vercel.app/use-cases/legal-transcription)

### Meeting platforms

Note-taking

Build a meeting platform that makes a real difference to your end users with automated note taking, a comprehensive feature set that covers 56+ languages.

- [Note-Taking & Meeting Assistants](https://speechmatics-website-git-preview-speechmatics.vercel.app/use-cases/meeting-platforms)

## Why developers choose our speech-to-text API

### Accurate where it matters most

Accuracy

Deliver reliable transcription across accented speech, background noise, and multi-speaker conversations, not just clean recordings.

- [speech-to-text page](https://speechmatics-website-git-preview-speechmatics.vercel.app/speech-to-text)

### Multilingual by design

Languages

Transcribe 56+ languages with a single model, even when speakers switch between languages naturally throughout a conversation.

- [Explore languages](https://speechmatics-website-git-preview-speechmatics.vercel.app/languages)

### Built for real conversations

Conversations

Speaker diarization and custom vocabulary are included, so transcripts identify who's speaking and capture your terminology accurately.

- [View Docs](https://docs.speechmatics.com/)

### Match the model to the use case

Models

Select the model that best fits your use case, whether you need maximum accuracy, multilingual performance, or lower latency and cost.

- [Models - Docs](https://docs.speechmatics.com/speech-to-text/models)

## Uncompromised, enterprise-level security

Enterprise security tools and controls, built for privacy-critical workflows.

### ISO 27001

Privacy and compliance built in with our ISO/IEC 27001:2022 accreditation.

- [View Trust Center](https://speechmatics.safebase.us/)

### GDPR

Compliant with privacy and compliance directives such as GDPR.

- [View Trust Center](https://speechmatics.safebase.us/)

### HIPAA

Fully compliant with the Health Insurance Portability and Accountability Act (HIPAA).

- [View Trust Center](https://speechmatics.safebase.us/)

### SOC 2 Type II

Trusted where privacy matters most - SOC 2 Type II-certified.

- [View Trust Center](https://speechmatics.safebase.us/)

## Speech-to-text API built for every language and accent

### Speech-to-text API built for every language and accent

Reach more users with a speech recognition API that handles the way people actually talk: regional accents, multiple speakers, and 56+ languages, including conversations where speakers switch mid-sentence.

One voice AI API powers live captions, voice agents, meeting notes, and contact center analytics across international markets. For the full technical breakdown, see our [speech-to-text API](https://speechmatics-website-git-preview-speechmatics.vercel.app/speech-to-text).

- [Explore languages](https://speechmatics-website-git-preview-speechmatics.vercel.app/languages)

## Speech-to-text API pricing that scales

Start with $100 in credit, no card required. With volume options and enterprise support when your product is ready for more.

### Get started with our speech-to-text API

Free

Test transcription quality, language support and integration options with $100 in credit to get started — no credit card required.

- [Start building](https://portal.speechmatics.com/signup)

### Demanding products and growing needs

Pro

Usage-based pricing for batch and real-time transcription — start with $100 in credit, add a card when you're ready, plus 20% savings available.

- [Start building](https://portal.speechmatics.com/signup)

### Large-scale, regulated or sensitive deployments

Enterprise

Custom pricing, unlimited scale, with flexible deployment options and enterprise support.

- [View pricing](https://www.speechmatics.com/pricing)

## Integrate our speech-to-text API in minutes

### Integrate our speech-to-text API in minutes

Start building with clear documentation, flexible APIs and the tools developers need to move quickly. Create your key, connect your workflow and bring accurate speech-to-text API performance into your product.

- [Check out Docs](https://docs.speechmatics.com/integrations-and-sdks/)

Check our Docs

### Integrations column

#### Integrations

- docsUrl: https://docs.speechmatics.com/integrations-and-sdks
- docsLabel: Full documentation
- cards:
  - Vapi integration:
    - href: https://docs.speechmatics.com/integrations-and-sdks/vapi
    - logoAlt: Vapi
    - logoWidth: 62
    - logoHeight: 40
  - Pipecat integration:
    - href: https://docs.speechmatics.com/integrations-and-sdks/pipecat/
    - logoAlt: Pipecat
    - logoWidth: 70
    - logoHeight: 40
  - LiveKit integration:
    - href: https://docs.speechmatics.com/integrations-and-sdks/livekit
    - logoAlt: LiveKit
    - logoWidth: 40
    - logoHeight: 40
  - Zapier integration:
    - href: https://docs.speechmatics.com/integrations-and-sdks/zapier
    - logoAlt: Zapier
    - logoWidth: 81
    - logoHeight: 40
- snippets:
  - Python:
    - language: python
    - code: # Install the speechmatics package using the command "pip install speechmatics-rt"

#!/usr/bin/env python3
"""Real-time transcription with microphone."""

import asyncio
import os
from dotenv import load_dotenv
from speechmatics.rt import (
    AsyncClient,
    ServerMessageType,
    TranscriptionConfig,
    TranscriptResult,
    OperatingPoint,
    AudioFormat,
    AudioEncoding,
    Microphone,
    AuthenticationError,
)

load_dotenv()


async def main():
    api_key = os.getenv("SPEECHMATICS_API_KEY")

    transcript_parts = []

    audio_format = AudioFormat(
        encoding=AudioEncoding.PCM_S16LE,
        chunk_size=4096,
        sample_rate=16000,
    )

    transcription_config = TranscriptionConfig(
        language="en",
        enable_partials=True,
        operating_point=OperatingPoint.ENHANCED,
    )

    mic = Microphone(
        sample_rate=audio_format.sample_rate,
        chunk_size=audio_format.chunk_size,
    )

    if not mic.start():
        print("PyAudio not installed. Install: pip install pyaudio")
        return

    try:
        async with AsyncClient(api_key=api_key) as client:
            @client.on(ServerMessageType.ADD_TRANSCRIPT)
            def handle_final_transcript(message):
                result = TranscriptResult.from_message(message)
                transcript = result.metadata.transcript
                if transcript:
                    print(f"[final]: {transcript}")
                    transcript_parts.append(transcript)

            @client.on(ServerMessageType.ADD_PARTIAL_TRANSCRIPT)
            def handle_partial_transcript(message):
                result = TranscriptResult.from_message(message)
                transcript = result.metadata.transcript
                if transcript:
                    print(f"[partial]: {transcript}")

            try:
                print("Connected! Start speaking (Ctrl+C to stop)...\n")

                await client.start_session(
                    transcription_config=transcription_config,
                    audio_format=audio_format,
                )

                while True:
                    frame = await mic.read(audio_format.chunk_size)
                    await client.send_audio(frame)

            except KeyboardInterrupt:
                pass
            finally:
                mic.stop()
                print(f"\n\nFull transcript: {' '.join(transcript_parts)}")

    except (AuthenticationError, ValueError) as e:
        print(f"\nAuthentication Error: {e}")


if __name__ == "__main__":
    asyncio.run(main())
  - JavaScript:
    - language: javascript
    - code: // npm install @speechmatics/real-time-client @speechmatics/auth

// Paste your API key into YOUR_API_KEY in the code.

import https from "node:https";
import { createSpeechmaticsJWT } from "@speechmatics/auth";
import { RealtimeClient } from "@speechmatics/real-time-client";

const apiKey = "YOUR_API_KEY";
const client = new RealtimeClient();
const streamURL = "https://media-ice.musicradio.com/LBCUKMP3";

async function transcribe() {
  // Print transcript as we receive it
  client.addEventListener("receiveMessage", ({ data }) => {
    if (data.message === "AddTranscript") {
      for (const result of data.results) {
        if (result.type === "word") {
          process.stdout.write(" ");
        }
        process.stdout.write(`${result.alternatives?.[0].content}`);
        if (result.is_eos) {
          process.stdout.write("\n");
        }
      }
    } else if (data.message === "EndOfTranscript") {
      process.stdout.write("\n");
      process.exit(0);
    } else if (data.message === "Error") {
      process.stdout.write(`\n${JSON.stringify(data)}\n`);
      process.exit(1);
    }
  });

  const jwt = await createSpeechmaticsJWT({
    type: "rt",
    apiKey,
    ttl: 60, // 1 minute
  });

  await client.start(jwt, {
    transcription_config: {
      language: "en",
      operating_point: "enhanced",
      max_delay: 1.0,
      transcript_filtering_config: {
        remove_disfluencies: true,
      },
    },
  });

  const stream = https.get(streamURL, (response) => {
    // Handle the response stream
    response.on("data", (chunk) => {
      client.sendAudio(chunk);
    });

    response.on("end", () => {
      console.log("Stream ended");
      client.stopRecognition({ noTimeout: true });
    });

    response.on("error", (error) => {
      console.error("Stream error:", error);
      client.stopRecognition();
    });
  });

  stream.on("error", (error) => {
    console.error("Request error:", error);
    client.stopRecognition();
  });
}

transcribe();
  - .NET:
    - language: csharp
    - code: // Install-Package Speechmatics

using System;
using System.Collections.Generic;
using System.Diagnostics;
using System.IO;
using System.Text;
using Speechmatics.Realtime.Client;
using Newtonsoft.Json;
using Speechmatics.Realtime.Client.Config;

namespace DemoApp
{
    public class Program
    {
        private const string SampleAudio = "2013-8-british-soccer-football-commentary-alex-warner.mp3";

        private static string ToJson(object obj)
        {
            return JsonConvert.SerializeObject(obj);
        }

        private static string RtUrl
        {
            get
            {
                var host = Environment.GetEnvironmentVariable("TEST_HOST") ?? "wss://api.rt.speechmatics.io";
                // Port 9000 for Speechmatics docker containers
                return host.StartsWith("wss://") ? host : $"wss://{host}:9000/";
            }
        }

        public static void Main(string[] args)
        {
            var start = DateTime.Now;
            Debug.WriteLine("Starting at {0}", start);
            var builder = new StringBuilder();
            var language = Environment.GetEnvironmentVariable("LANG") ?? "en";
            Console.WriteLine(language);

            using (var stream = File.Open(SampleAudio, FileMode.Open, FileAccess.Read))
            {
                try
                {

                    var config = new SmRtApiConfig(language)
                    {
                        AuthToken= Environment.GetEnvironmentVariable("AUTH_TOKEN"),
                        // GenerateTempToken = True <- set this to True for accounts from portal.speechmatics.com
                        OutputLocale = "en-GB",
                        AddTranscriptCallback = s => builder.Append(s),
                        AddTranscriptMessageCallback = s => Console.WriteLine(ToJson(s)),
                        AddTranslationMessageCallback = s => Console.WriteLine(ToJson(s)),
                        AddPartialTranscriptMessageCallback = s => Console.WriteLine(ToJson(s)),
                        ErrorMessageCallback = s => Console.WriteLine(ToJson(s)),
                        WarningMessageCallback = s => Console.WriteLine(ToJson(s)),
                        CustomDictionaryPlainWords = new[] {"speechmagic"},
                        CustomDictionarySoundsLikes = new Dictionary<string, IEnumerable<string>>(),
                        Insecure = true,
                        EnablePartials=true,
                        TranslationConfig = new TranslationConfig() {
                            TargetLanguages = new [] {"de"},
                            EnablePartials = true
                        }
                    };

                    // We can do this here, or earlier. It's not used until .Run() is called on the API object.
                    config.CustomDictionarySoundsLikes["gnocchi"] = new[] {"nokey", "noki"};

                    var api = new SmRtApi(RtUrl,
                        stream,
                        config
                    );
                    // Run() will block until the transcription is complete.
                    Console.WriteLine($"Connecting to {RtUrl}");
                    api.Run();
                    Console.WriteLine(builder.ToString());
                }
                catch (AggregateException e)
                {
                    Console.WriteLine(e);
                }
            }

            var finish = DateTime.Now;
            Debug.WriteLine("Starting at {0} -- {1}", finish, finish-start);
            Console.ReadLine();
        }
    }
}

## **Power your products with enterprise-grade Voice AI**

We handle the speech, you deliver conversations that matter.

- [Get Started](https://portal.speechmatics.com/signup/)

- [Contact sales](https://speechmatics-website-git-preview-speechmatics.vercel.app/speak-to-sales)

## Speechmatics speech-to-text API: frequently asked questions

### What is a speech-to-text API and how does it work?

A speech-to-text API converts spoken audio into written text inside an app, product or workflow. Speechmatics can process recorded files or live audio, then return transcripts for captions, search, analytics, records and voice AI.

### Can I try Speechmatics for free?

Yes. New accounts get $100 in credit with no card required, so you can test transcription quality and language support before moving to a paid plan.

### How accurate is Speechmatics on real-world audio?

In independent testing by Pipecat (as of August 2026), Speechmatics returned a 1.07% pooled word error rate, the lowest of the 12 services benchmarked, including Deepgram, AWS, and Azure. Read more in [Speed you can trust: the STT metrics that matter for voice agents](https://speechmatics-website-git-preview-speechmatics.vercel.app/company/articles-and-news/speed-you-can-trust-the-stt-metrics-that-matter-for-voice-agents).

### How many languages does Speechmatics support?

56+, with coverage across accents, dialects, and code-switching, where speakers move between languages mid-sentence.

### Can Speechmatics tell speakers apart in a conversation?

Yes. Speaker diarization identifies and separates individual speakers in meetings, calls, legal proceedings, and healthcare consultations.

### Is Speechmatics GDPR-compliant and suitable for regulated industries?

Yes. Speechmatics supports cloud, on-premise, and on-device deployment, and is ISO 27001, GDPR, HIPAA, and SOC 2 Type II compliant.
