Inworld AI

Inworld AI · Launch Video Breakdown: Hook, Pacing & Motion Design

The Realtime AI Company.

AI AgentsProductionSeptember 2, 2026@inworld_ai
0:00 · The Hook · Vision for Voice AI
0:00 / 0:00

Scene-by-scene timeline & spoken transcript

  1. The Hook

    Vision for Voice AI

    “Voice AI is the new default user interface, but what matters to people isn't how it sounds, it's how it makes them feel. We built real-time AI for consumer-facing applications because we want everyone to experience the magic of voice AI.”

    On screen
    Voice AI is the new default user interface, but what matters to people isn't how it sounds, it's how it makes them feel. Kylan Gibbs, CEO We built real-time AI for consumer-facing applications because we want everyone to experience the magic of voice AI.
    Camera
    Tracking shot following CEO through office lobby into elevator and to desk.
    Motion
    Cinematic shallow depth-of-field live action with natural camera movement.
  2. Product Reveal

    Introducing Realtime TTS-2

    “Hey, I'm Igor, Chief Science Officer at Inworld, and today I want to introduce you to our best TTS model family yet, the Inworld TTS-2. We've rebuilt the entire serving stack, and the model's architecture in order to support ultra-low latency in real-time production applications. TTS-2 is the most versatile speech synthesis model out there. Similarly to large-language models, it's now promptable, and understands the conversational context. It hears what the users say.”

    On screen
    Igor Poletaev, Chief Science Officer Realtime TTS-2
    Camera
    Medium office portrait cutting to medium close-up, static tripod.
    Motion
    Clean typographic lower-third card and title card overlay.
  3. Feature Teaser

    Voice Steering & Cross-Lingual Power

    “A good voice actor knows how to adapt to the moment. With voice steering, you can instruct the model in natural language, as you would an actor. Cross-lingual support allows dynamic switching between over 500 dialects, with the same speaker identity. And it all happens under 150 milliseconds to feel organic and real-time.”

    On screen
    A good voice actor knows how to adapt to the moment. With voice steering, you can instruct the model in natural language, as you would an actor. Cross-lingual support allows dynamic switching between over 500 dialects, with the same speaker identity. And it all happens under 150 milliseconds to feel organic and real-time.
    Camera
    Direct-to-camera medium shot of CEO seated at desk.
    Motion
    Synchronized burned-in subtitles with precise audio alignment.
  4. Feature Teaser

    Live Multilingual Office Demonstrations

    “Hey Prashanth, can you come and try this out? Are you excited for the launch? [in an excited tone] Of course! I'm very excited to launch TTS-2 today! Impressive. Anastasiya, you want to check this out? Okay, let's see. Are there many bugs left in you? [in a sarcastic tone] After your checks? Not a single one survived! That's good. Hey Andreas, can you come and try it? Yo. I'm having a rough day, do you have any advice? [in a calm and reassuring tone] Calm down, breathe, everything is going to be okay. That was crazy! We should make ads in all languages now!”

    On screen
    Prashanth Seralathan, Engineer Anastasiya Korol, QA Lead Andreas Assad, BizOps [in an excited tone] Of course! I'm very excited to launch TTS-2 today! [in a sarcastic tone] After your checks? Not a single one survived! [in a calm and reassuring tone] Calm down, breathe, everything is going to be okay. [excitedly] That was crazy! We should make ads in all languages now!
    Camera
    Handheld over-the-shoulder panning between speakers and laptop screen.
    Motion
    Screen recording capture of 3D generative audio particle orb visualizer on laptop display.
  5. Problem Agitation

    Customer Social Proof

    “Hi, my name's David. I'm cofounder and CTO of Livekit, and Livekit is a platform that helps developers to build voice agents. I was really impressed with how steerable the model is. The ability to prompt the voice to sound different given the same text, but having a full range of different types of emotions it could display is really impressive. It really felt like talking to a person. If you're using another TTS model and have not checked out Inworld yet, I would highly recommend giving them a look. They went from launching their first TTS model, to topping artificial analysis charts in just about a year.”

    On screen
    David Zhao, Cofounder, CTO, Livekit Customer
    Camera
    Static corporate documentary head-and-shoulders interview framing.
    Motion
    Subtle depth-of-field interview lighting with branded lower-third text.
  6. Call to Action

    Launch Day CTA & Partner Grid

    “TTS-2 is live today, and you've probably already spoken with it along with hundreds of millions of users. Try it today at realtime.ai.”

    On screen
    LiveKit fal deepinfra fonio.ai latitude slingshot AI state Strella Talkpal AI Inworld Realtime TTS-2, ranked #1 on Artificial Analysis Learn more at inworld.ai/realtime-tts-2
    Camera
    Static front shot transitioning to clean graphic end-card.
    Motion
    Floating glassmorphism logo pills appearing around subject, leading into animated brand logo reveal.

Related AI Agents Startup Launches

Explore all 1064 AI Agents launches →
OpenAI
Hook 92.0137.9M
OpenAIAI Agents

To ensure that artificial general intelligence benefits all of humanity

@OpenAI
Anthropic
Hook 92.057.6M
AnthropicAI Agents

Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use. Its capabilities exceed those of any model we’ve ever made generally available. https://t.co/2AvmEjHIX8

@claudeai
xAI
Hook 9.256.9M
xAIAI Agents

To understand the true nature of the universe.

@bot