Skip to main content

Overview

Giggy makes expressive Voice AI accessible to more developers. This integration brings Giggy’s paid streaming speech API into Pipecat, so developers can give their assistants a distinctive voice while keeping their existing transcription, language model, and transport providers. GiggyHttpTTSService connects Pipecat to Giggy’s HTTP speech endpoint, using Giggy voice UUIDs and progressive mono PCM16 at 24 kHz. The integration is maintained by Giggy, the speech API provider. It supports runtime voice/speed updates, first-audio and usage metrics, and cancellation on interruption.

Installation

Requires Python 3.11 or later. Last tested with Python 3.12 and Pipecat 1.12.0. The package currently pins that Pipecat version.

Prerequisites

For a step-by-step account, API key, voice and Streaming setup walkthrough, see the Giggy Pipecat setup guide.
  • A Giggy account and API key.
  • A Giggy voice UUID from GET /v1/voices or GET /v1/my-voices.
  • Streaming credits. This endpoint uses paid Streaming admission; see Giggy pricing.
  • Set GIGGY_API_KEY and GIGGY_VOICE_ID in your application environment.

Configuration

Additional keyword arguments pass to Pipecat’s TTSService. Runtime updates use GiggyHttpTTSService.Settings(voice="voice-uuid", speed=1.1) inside a TTSUpdateSettingsFrame(delta=...). The model is fixed to giggyspeech.

Usage

The application owns the session and must keep it open until the pipeline stops. Your transport/turn-detection processors generate InterruptionFrame and your output transport clears queued playback. Cancellation closes the active HTTP response; it does not guarantee immediate GPU preemption. Failed requests are not retried. This TTS service provides no microphone transcription or word timestamps. See the source repository for the single-file runnable WAV example, timeout behavior, maintenance and changelog.

Demo

Watch the Angela Fowler demo in Giggy’s repository. It shows real speech synthesis, a simulated user interruption and resumed speech. The displayed text is synthesis text, not microphone transcription.