Gladia
Gladia delivers speech-to-text and audio intelligence APIs tailored for developers who need accurate, low-latency transcription inside their products. It combines asynchronous batch transcription with real-time streaming, multilingual support, and a growing layer of AI add-ons such as summarization, sentiment analysis, and entity extraction.
Make a decision about Gladia
- Already use it? Add Gladia to Stack Autopsy — check cost, overlap and safe cancellations.
- Thinking about buying it? Ask AI Advisor — compare fit, alternatives and trade-offs.
- Want to leave it? Find a replacement — see feature losses and migration risks.
Pricing
- Free Tier: Up to 10 hours of transcription per month at no cost for experimentation and low-volume projects.
- Self-Serve (Pay-as-you-Go): Asynchronous transcription from $0.61 per hour of audio and real-time from $0.75 per hour, with core features like diarization included.
- Scaling Plan: Volume-focused pricing, with asynchronous transcription from $0.50 per hour and real-time from $0.55 per hour plus flexible concurrency and discounts.
- Enterprise: Custom pricing for large deployments, including SLAs, premium support, and tailored hosting or data-retention guarantees.
Features
- Solaria universal STT model: Proprietary Solaria-1 powers transcription across more than 100 languages, handling accents and domain-specific jargon while maintaining high accuracy for names, numbers, and other key entities.
- Real-time streaming with ultra-low latency: Real-time APIs target latency under 300 milliseconds, with a Partials feature that returns partial transcripts in under 100 milliseconds for live assistants and voice agents.
- Audio intelligence add-ons: Beyond raw transcripts, Gladia offers diarization, automatic language detection and code-switching, sentiment analysis, named entity recognition, word-level timestamps, summarization, and any-to-any translation.
- Telephony and communication stack support: The API is tuned for SIP and common telephony protocols at 8 kHz, and works with WebRTC and popular communications platforms, fitting easily into existing voice pipelines.
- Developer-first experience: REST and WebSocket endpoints, language-agnostic integration, lightweight SDKs, playground, status page, and Discord-based community support keep integration quick and transparent.
- Enterprise-grade security and compliance: Gladia is GDPR, HIPAA, SOC 2 Type 2, and ISO 27001 compliant, with options for custom, on-premise, or air-gapped hosting and strict controls around model training and data retention.
Use cases
- Virtual meeting and collaboration platforms: Turning meeting audio into searchable transcripts, notes, and summaries for internal knowledge and productivity.
- Contact centers and CCaaS vendors: Powering real-time agent assist, QA analytics, and compliance monitoring over telephony-grade audio.
- Sales enablement and CRM enrichment tools: Capturing names, emails, intent, and objections on calls to feed downstream AI coaching and automation.
- Media, podcasting, and streaming platforms: Producing subtitles, captions, and searchable archives from large audio and video libraries.
- Specialized sectors (finance, legal, healthcare): Handling sensitive conversations where transcription fidelity and compliance are non-negotiable.
- Uncommon Use Cases: Used in voice UX research labs to analyze user interviews at scale; adopted by education platforms to transcribe multilingual lectures and tutorials.