Aura Health

Voice First AI Companion

Voice First AI Companion

Aura Health partnered with us to build a real-time conversational voice experience, embed emotionally-aware AI dialogue, and scale personalized voice-based support to a global audience.

Aura Health partnered with us to build a real-time conversational voice experience, embed emotionally-aware AI dialogue, and scale personalized voice-based support to a global audience.

Company Overview

Aura Health's premium offering is a native voice-first AI companion, designed to feel like talking to a present, attentive guide rather than tapping through an app. Delivered through a dedicated native iOS experience, it lets users speak naturally and receive real-time spoken responses, shaped by conversational memory, mood tracking, and reflective journaling. It extends Aura Health's decade of wellness expertise into a fully voice-driven modality.

The Challenge

  • Latency vs. presence - a therapeutic voice conversation only feels real if responses arrive with near-zero lag; any stutter breaks trust in a deeply personal moment.

  • Emotional continuity - a single conversation isn't enough; the companion needed to remember prior context, mood patterns, and journal entries to respond with real relational depth over time.

  • Voice that doesn't feel synthetic - text-to-speech needed to carry warmth and nuance, not robotic cadence, to support genuinely sensitive conversations.

  • Safety at conversational speed - every new AI behavior needed rigorous evaluation before reaching real users, without slowing the pace of feature iteration.

  • Bridging voice and structure - free-flowing conversation had to connect back to structured wellness tools such as mood check-ins, affirmations, guided journaling, and onboarding, not exist as an isolated chat feature.

Scope

We partnered with Aura Health to architect their real-time voice AI core, from low-latency audio infrastructure through to a safe, evaluation-gated system for shipping new conversational capabilities, built to support natural, ongoing voice relationships at scale.

What We Delivered

  • Real-time conversational voice engine - low-latency, two-way voice conversations powered by OpenAI's Realtime API over WebSocket, so responses feel spoken, not typed-then-read.

  • Natural voice synthesis and low-latency audio transport - ElevenLabs text-to-speech paired with LiveKit's WebRTC infrastructure for responsive, natural-sounding audio in real time.

  • Persistent memory and context system - dedicated conversation, dialogue, and memory layers so the app carries context across sessions instead of starting fresh each time.

  • Integrated emotional tracking - mood and emotion signals captured directly within voice conversations, feeding personalized responses and longer-term insight.

  • Reflective tools woven into voice - AI-assisted journaling and affirmations that connect what's said out loud to written reflection.

  • Guided voice onboarding - a conversational first-run experience that introduces the app's voice interaction model naturally, rather than through static screens.

Impact

  • Voice is now the majority conversation mode, accounting for 55% of completed sessions.

  • 43% of users return for a repeat voice conversation within 7 days.

  • 50% increase in average session length from voice-based conversations vs. text-only interactions.

  • 18% of voice conversations extend into connected journaling or affirmation activity within the hour (10% journaling, 8% affirmation).

Conclusion

Aura's voice AI transformation extended its wellness mission into a new modality: real-time, emotionally continuous conversation that feels present rather than transactional. By pairing low-latency voice infrastructure with a rigorously evaluated rollout system, the app delivers natural, trustworthy voice support at scale, backed by the same real-time emotional intelligence that has made their core platform a leader for millions of users.

Company Overview

Aura Health's premium offering is a native voice-first AI companion, designed to feel like talking to a present, attentive guide rather than tapping through an app. Delivered through a dedicated native iOS experience, it lets users speak naturally and receive real-time spoken responses, shaped by conversational memory, mood tracking, and reflective journaling. It extends Aura Health's decade of wellness expertise into a fully voice-driven modality.

The Challenge

  • Latency vs. presence - a therapeutic voice conversation only feels real if responses arrive with near-zero lag; any stutter breaks trust in a deeply personal moment.

  • Emotional continuity - a single conversation isn't enough; the companion needed to remember prior context, mood patterns, and journal entries to respond with real relational depth over time.

  • Voice that doesn't feel synthetic - text-to-speech needed to carry warmth and nuance, not robotic cadence, to support genuinely sensitive conversations.

  • Safety at conversational speed - every new AI behavior needed rigorous evaluation before reaching real users, without slowing the pace of feature iteration.

  • Bridging voice and structure - free-flowing conversation had to connect back to structured wellness tools such as mood check-ins, affirmations, guided journaling, and onboarding, not exist as an isolated chat feature.

Scope

We partnered with Aura Health to architect their real-time voice AI core, from low-latency audio infrastructure through to a safe, evaluation-gated system for shipping new conversational capabilities, built to support natural, ongoing voice relationships at scale.

What We Delivered

  • Real-time conversational voice engine - low-latency, two-way voice conversations powered by OpenAI's Realtime API over WebSocket, so responses feel spoken, not typed-then-read.

  • Natural voice synthesis and low-latency audio transport - ElevenLabs text-to-speech paired with LiveKit's WebRTC infrastructure for responsive, natural-sounding audio in real time.

  • Persistent memory and context system - dedicated conversation, dialogue, and memory layers so the app carries context across sessions instead of starting fresh each time.

  • Integrated emotional tracking - mood and emotion signals captured directly within voice conversations, feeding personalized responses and longer-term insight.

  • Reflective tools woven into voice - AI-assisted journaling and affirmations that connect what's said out loud to written reflection.

  • Guided voice onboarding - a conversational first-run experience that introduces the app's voice interaction model naturally, rather than through static screens.

Impact

  • Voice is now the majority conversation mode, accounting for 55% of completed sessions.

  • 43% of users return for a repeat voice conversation within 7 days.

  • 50% increase in average session length from voice-based conversations vs. text-only interactions.

  • 18% of voice conversations extend into connected journaling or affirmation activity within the hour (10% journaling, 8% affirmation).

Conclusion

Aura's voice AI transformation extended its wellness mission into a new modality: real-time, emotionally continuous conversation that feels present rather than transactional. By pairing low-latency voice infrastructure with a rigorously evaluated rollout system, the app delivers natural, trustworthy voice support at scale, backed by the same real-time emotional intelligence that has made their core platform a leader for millions of users.

Tech Stack

OpenAI

ElevenLabs

Gemini

NodeJs

ExpressJS

PostgreSQL

Google

Firebase

HumeAI

Swift

LiveKit

See other case studies

See other case studies