Overview
BitHumanVideoService takes your bot’s TTS audio and turns it into a talking avatar.
It pushes OutputImageRawFrame (RGB) and the matching TTSAudioRawFrame, so the mouth
and the voice leave together. The avatar renders in your own process through the
bithuman Python SDK. Expression 2 animates any character from one portrait;
Essence 2 renders a photoreal person from one portrait.
Interruptions are handled: on InterruptionFrame the avatar drops the reply in flight
and goes back to idle.
Installation
Prerequisites
bitHuman account setup
- Create an account at www.bithuman.ai and create an API secret.
- Download an avatar model (
.imx). See docs.bithuman.ai.
Required environment variables
BITHUMAN_API_SECRET: your bitHuman API secret.BITHUMAN_MODEL_PATH: path to the.imxmodel (or passmodel_path).
Configuration
Constructor parameters forBitHumanVideoService:
str
default:"None"
Path to the avatar’s
.imx model. Defaults to BITHUMAN_MODEL_PATH.str
default:"None"
bitHuman API secret. Defaults to
BITHUMAN_API_SECRET. Never logged.bool
default:"True"
Sets
sync_with_audio on each image so the transport shows it after its audio.bool
default:"True"
If the avatar fails, forward TTS audio unchanged so the bot keeps talking.
float
default:"2.0"
Release a held
TTSStoppedFrame after this much quiet with no end-of-speech.float
default:"30.0"
On
EndFrame, the longest wait for queued speech to finish.Settings yet.
Session time is metered while the avatar is open (talking or idle). It closes on
EndFrame, CancelFrame or cleanup.
Usage
video_out_enabled=True on the transport. The service logs the frame size on the
first frame; set video_out_width and video_out_height to match.
A complete example is in the
repository.