Skip to main content

Transports

Transports exchange audio and video streams between the user and bot.

Serializers

Serializers convert between frames and media streams, enabling real-time communication over a websocket.

Speech-to-Text

Speech-to-Text services receive and audio input and output transcriptions.

Large Language Models

LLMs receive text or audio based input and output a streaming text response.

Text-to-Speech

Text-to-Speech services receive text input and output audio streams or chunks.

Speech-to-Speech

Speech-to-Speech services are multi-modal LLM services that take in audio, video, or text and output audio or text.

Image Generation

Image generation services receive text inputs and output images.

Video

Video services enable you to build an avatar where audio and video are synchronized.

Memory

Memory services can be used to store and retrieve conversations.

Vision

Vision services receive a streaming video input and output text describing the video input.

Analytics & Monitoring

Analytics services help you better understand how your service operates.