Evom Labs

Realtime voice infrastructure

Build products that
listen and speak

Turn speech into text with Loli 2.0 and generate natural speech with Loly 3.5. Add multilingual voice features to your product over API and WebSocket.

One voice pipeline, from input to response

Combine Loli 2.0 and Loly 3.5 to build voice assistants, conversational applications and real-time interactive experiences.

One voice, end to end

  1. 01

    Voice

    Audio input

  2. 02

    Loli 2.0

    Speech to text

  3. 03API

    Your application

    Process and compose a response

  4. 04

    Loly 3.5

    Text to speech

  5. 05

    Audio response

    Played back to the user

One platform, many ways to build

For developers

Add speech recognition and speech generation to your product over API and WebSocket.

  • Speech-to-Text

    Recognise and transcribe speech from audio files or streams.

  • Text-to-Speech

    Turn text into audio and receive the voice data as a stream.

  • Conversational apps

    Combine Loli 2.0 and Loly 3.5 to build interactive voice experiences.

  • WebSocket integration

    Process continuous audio for applications that need real-time responses.

  • Voice Library

    Create, manage and use custom voices inside your product.

Core capabilities of Evom Labs

Multilingual

Handle voice content in many languages on a single platform.

Real time

Streaming support for products that process and respond continuously.

API and WebSocket

Flexible integration for on-demand tasks and continuous audio streams.

Custom voices

Create and reuse voices from an audio sample you provide.

Two models for one complete voice experience

From speech recognition to spoken response, Evom Labs provides the pieces you need to build conversational products and voice content.

Loli 2.0Speech-to-Text

Recognise and transcribe multilingual speech in real time. Works with many audio file types, streaming and WebSocket connections.

Explore Speech-to-Text
Loly 3.5Text-to-Speech

Generate multilingual speech from text, with streaming, several output formats and support for custom voices.

Explore Text-to-Speech

Bring voice into your product

Evom Labs offers APIs for Speech-to-Text, Text-to-Speech and the Voice Library, plus WebSocket connections for experiences that process audio in real time.

  • Speech-to-Text API
  • Text-to-Speech API
  • Voice Library API
  • WebSocket realtime
View the API docs

A voice of your own for your content

Voice LibraryCreate and manage custom voices

Create a custom voice from one short audio sample, with no transcript required. Once created, a voice can be managed and reused across everything you make next.

  • Audio sample of 5 to 10 seconds
  • No transcript required
  • Voices managed in one place
  • Reusable across your content
Explore the Voice Library

Only create voices you have the owner's permission to use.

Integrate the way
that fits your product

Use the API for on-demand processing, or WebSocket for experiences that stream audio continuously.

Connect Speech-to-Text, Text-to-Speech and Voice Library tasks to your application.

Stream and process audio continuously for applications that need real-time responses.

Start building voice experiences with Evom Labs

Talk through your use case

Share how you plan to use voice, the scale of your product and your integration approach with the Evom Labs team.

Start building

Explore the voice products and start integrating Evom Labs into your application.

Build the future of voice experiences

Learn about Evom Labs