Our technology

Building production-ready voice assistants requires far more than connecting speech to an LLM. Real-world capabilities like actively participating in live conversations, identifying multiple participants, maintaining context across interactions, and recovering seamlessly from dropped calls require infrastructure that most conversational AI frameworks don't provide. That's why we built our own custom voice assistant harness—designed around real-world workflows and production reliability.

Built for Real-World Use Cases

Our technology is designed to solve the challenges that arise when AI leaves the demo environment and starts interacting with real people over the phone.

Betula voice assistant technology stack — built for real-world phone calls
  • minimizing conversation and action latency
  • handling carrier-grade call routing and spam blocking
  • recovering seamlessly from dropped calls
  • orchestrating outbound calls, texts, and reminders
  • enabling real-time owner collaboration during calls
  • maintaining persistent memory across interactions
  • interacting with legacy phone systems via DTMF
  • ensuring accurate timezone-aware scheduling
  • producing natural speech for numbers and phone calls
  • safeguarding caller privacy and data

Custom Harness for Voice Assistants

We built a purpose-built harness (on top of LiveKit) tailored for voice assistants—not a generic conversational framework repurposed for speech. Our stack provides:

  • ensuring reliable orchestration of inbound and outbound calls
  • maintaining persistent, searchable call history, transcripts, action items, reminders, and journal entries
  • ensuring consistent memory and context across calls and text interactions
  • decoupling speech and language models from application logic
  • continuously testing and monitoring for production reliability

Things other platforms abstract away, we built from scratch. Every layer—from carrier interconnects to LLM orchestration—is ours to inspect, tune, and control.

Powered by Best-in-Class AI

We use state-of-the-art models for language understanding, transcription, and speech synthesis. Our voice infrastructure is powered by LiveKit, Cartesia, Deepgram, and ElevenLabs, selected after extensive evaluation for reliability, latency, and conversational quality.

For healthcare applications, our platform also supports avatar agents and visual perception capabilities.