Inworld AI · Launch Video Breakdown: Hook, Pacing & Motion Design
The Realtime AI Company.
Scene-by-scene timeline & spoken transcript
The Hook
Vision for Voice AI
“Voice AI is the new default user interface, but what matters to people isn't how it sounds, it's how it makes them feel. We built real-time AI for consumer-facing applications because we want everyone to experience the magic of voice AI.”
- On screen
- Voice AI is the new default user interface, but what matters to people isn't how it sounds, it's how it makes them feel. Kylan Gibbs, CEO We built real-time AI for consumer-facing applications because we want everyone to experience the magic of voice AI.
- Camera
- Tracking shot following CEO through office lobby into elevator and to desk.
- Motion
- Cinematic shallow depth-of-field live action with natural camera movement.
Product Reveal
Introducing Realtime TTS-2
“Hey, I'm Igor, Chief Science Officer at Inworld, and today I want to introduce you to our best TTS model family yet, the Inworld TTS-2. We've rebuilt the entire serving stack, and the model's architecture in order to support ultra-low latency in real-time production applications. TTS-2 is the most versatile speech synthesis model out there. Similarly to large-language models, it's now promptable, and understands the conversational context. It hears what the users say.”
- On screen
- Igor Poletaev, Chief Science Officer Realtime TTS-2
- Camera
- Medium office portrait cutting to medium close-up, static tripod.
- Motion
- Clean typographic lower-third card and title card overlay.
Feature Teaser
Voice Steering & Cross-Lingual Power
“A good voice actor knows how to adapt to the moment. With voice steering, you can instruct the model in natural language, as you would an actor. Cross-lingual support allows dynamic switching between over 500 dialects, with the same speaker identity. And it all happens under 150 milliseconds to feel organic and real-time.”
- On screen
- A good voice actor knows how to adapt to the moment. With voice steering, you can instruct the model in natural language, as you would an actor. Cross-lingual support allows dynamic switching between over 500 dialects, with the same speaker identity. And it all happens under 150 milliseconds to feel organic and real-time.
- Camera
- Direct-to-camera medium shot of CEO seated at desk.
- Motion
- Synchronized burned-in subtitles with precise audio alignment.
Feature Teaser
Live Multilingual Office Demonstrations
“Hey Prashanth, can you come and try this out? Are you excited for the launch? [in an excited tone] Of course! I'm very excited to launch TTS-2 today! Impressive. Anastasiya, you want to check this out? Okay, let's see. Are there many bugs left in you? [in a sarcastic tone] After your checks? Not a single one survived! That's good. Hey Andreas, can you come and try it? Yo. I'm having a rough day, do you have any advice? [in a calm and reassuring tone] Calm down, breathe, everything is going to be okay. That was crazy! We should make ads in all languages now!”
- On screen
- Prashanth Seralathan, Engineer Anastasiya Korol, QA Lead Andreas Assad, BizOps [in an excited tone] Of course! I'm very excited to launch TTS-2 today! [in a sarcastic tone] After your checks? Not a single one survived! [in a calm and reassuring tone] Calm down, breathe, everything is going to be okay. [excitedly] That was crazy! We should make ads in all languages now!
- Camera
- Handheld over-the-shoulder panning between speakers and laptop screen.
- Motion
- Screen recording capture of 3D generative audio particle orb visualizer on laptop display.
Problem Agitation
Customer Social Proof
“Hi, my name's David. I'm cofounder and CTO of Livekit, and Livekit is a platform that helps developers to build voice agents. I was really impressed with how steerable the model is. The ability to prompt the voice to sound different given the same text, but having a full range of different types of emotions it could display is really impressive. It really felt like talking to a person. If you're using another TTS model and have not checked out Inworld yet, I would highly recommend giving them a look. They went from launching their first TTS model, to topping artificial analysis charts in just about a year.”
- On screen
- David Zhao, Cofounder, CTO, Livekit Customer
- Camera
- Static corporate documentary head-and-shoulders interview framing.
- Motion
- Subtle depth-of-field interview lighting with branded lower-third text.
Call to Action
Launch Day CTA & Partner Grid
“TTS-2 is live today, and you've probably already spoken with it along with hundreds of millions of users. Try it today at realtime.ai.”
- On screen
- LiveKit fal deepinfra fonio.ai latitude slingshot AI state Strella Talkpal AI Inworld Realtime TTS-2, ranked #1 on Artificial Analysis Learn more at inworld.ai/realtime-tts-2
- Camera
- Static front shot transitioning to clean graphic end-card.
- Motion
- Floating glassmorphism logo pills appearing around subject, leading into animated brand logo reveal.
Motion design primitives
Synthesizable in Remotion · key components of the executive-announcement-live-demo template.
AudioReactiveOrb
A circular particle sphere undulating dynamically in response to speech frequency and amplitude.
+ Code construction recipe
Generate a Canvas element in Remotion; compute simplex noise offsets modulated by audio amplitude extracted via useAudioData.
FloatingPartnerPills
Semi-transparent rounded badges floating on screen with spring-based entrance and subtle sinusoidal vertical drift.
+ Code construction recipe
Use spring() from remotion for entrance stagger, map frame count with Math.sin() for subtle floating position.
MinimalEndcardReveal
Centered monochromatic logo mark with clean serif/sans typography fading in with slight upward translation.
+ Code construction recipe
Interpolate opacity from 0 to 1 and translateY from 12px to 0px using Easing.out(Easing.cubic) over 25 frames.
Style DNA & aesthetic system
Visual aesthetic
Typography
Motion
Special effects
Pacing & rhythm
Conceptual story rhythm
- 1
Visionary Opening
Purpose. Establish voice AI as the primary future interface.
Execution. Walking executive tracked into office environment.
- 2
Technical Introduction
Purpose. Unveil TTS-2 architecture and sub-150ms performance.
Execution. Chief Science Officer direct address with title card.
- 3
Multilingual Proof
Purpose. Demonstrate emotion steering and cross-lingual fluency across Hindi, Russian, Spanish, and Mandarin.
Execution. Casual desk interactions with colleagues reacting to live UI.
- 4
Enterprise Validation
Purpose. Provide trusted third-party benchmark proof via LiveKit CTO.
Execution. Professional boardroom documentary interview.
- 5
Ecosystem Grid & Call to Action
Purpose. Showcase ecosystem traction and direct users to live trial.
Execution. Floating partner logo badges followed by minimalist end card.
Remix in Animatiq Studio
Master remix prompt · agent-ready
Sample remix
Objective
Launch of a real-time multilingual AI customer support engine.How the recipe adapts
Replace internal office workers with support agents testing instant voice switching across Portuguese, Japanese, and English, concluding with enterprise integration badges.Remix blueprint · executive-announcement-live-demo
Scene durations (6 scenes, 148s)
Key Remotion components
Related AI Agents Startup Launches
Explore all 1064 AI Agents launches →To ensure that artificial general intelligence benefits all of humanity

Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use. Its capabilities exceed those of any model we’ve ever made generally available. https://t.co/2AvmEjHIX8
To understand the true nature of the universe.