Skip to main content

Overview

MiraiTTSService converts text into speech using Mirai’s streaming text-to-speech API, built for Indian languages: Hindi, Hinglish (Devanagari with English words in Latin script) and Gujarati. Audio streams over HTTP with about 100 ms to first audio, and is resampled from Mirai’s 48 kHz to your pipeline’s output rate, including 8 kHz for phone calls. The package also includes apply_output_lead(). It stops audio breaking up on phone calls over Pipecat’s websocket transports when the bot’s server is busy, and works with any TTS service.

Source Repository

Source code, examples, benchmark and issues for the Mirai integration

Mirai

Learn more about Mirai’s text-to-speech API

Installation

This is a community-maintained package distributed separately from pipecat-ai:

Prerequisites

Mirai Account Setup

Before using the Mirai text-to-speech service, you need:
  1. Mirai Account: Sign up in the Mirai console
  2. API Key: Create an API key (sk_live_...) in the console
  3. Voice Selection: Choose ashu, neha, shruti or sameer

Required Environment Variables

  • MIRAI_API_KEY: Your Mirai API key for authentication

Configuration

str
default:"None"
Mirai API key. Defaults to the MIRAI_API_KEY environment variable.
MiraiTTSService.Settings
default:"None"
Runtime-configurable settings: voice (default "neha") and model (default "mira-tts").
int
default:"None"
Output sample rate. Defaults to the pipeline’s audio_out_sample_rate.
str
default:"https://sandbox.voice.miraiminds.co/v1"
API base URL.

Usage

Phone calls

Pipecat’s websocket transports send audio at exactly real time, so a stalled event loop starves the phone provider and the caller hears the voice break. Let the transport run slightly ahead:
On 8 kHz Twilio-protocol calls with event-loop stalls, measured speech stretched by breaks dropped from 6.6% to 0.08% at 3 concurrent calls and from 29.1% to 0.01% at 10.

Compatibility

Tested with Pipecat v1.8.1 and v1.12.0.