Consumer AI runs on Inworld.

Mountain View, CA
Pinned Tweet
We’re excited to announce that @ultravox_dot_ai is now part of Inworld. Ultravox is the platform developers use to build real-time voice agents. We've worked with the team for a while through our TTS partnership, and today members of the team that built it are joining Inworld to keep developing it.
83
100
285
520,264
Join us in Berlin on Friday, October 23 for Backchannel Berlin, an afternoon of conversation about the research on speech-to-speech models. Our team will share how we're building and serving full-duplex speech-to-speech in realtime, including the things that haven't worked. Then the floor opens: bring your own work or a problem you're stuck on, or just come and listen. We'll be getting into turn-taking and interruptions, cascaded versus end-to-end, keeping an LLM's reasoning when you teach it to speak, cross-lingual transfer, and what we should actually be measuring for conversations. If you work on speech-to-speech, TTS, ASR, audio codecs, conversational modelling, LLM serving or evaluation, we'd love to have you. PhD students welcome. Food and drinks throughout. Register at luma.com/backchannel-berlin
2
2
13
1,224
GPT-6.1 Sol is live on Inworld Realtime Router. - OpenAI's upgraded Sol tier - Near-Astra intelligence at a fifth of the price, per OpenAI - 1M token context, 128K output - $2.00 in / $10.00 out per 1M tokens. Provider rates with no markup. Swap the model string and you're good to go!
2
11
843
We’re excited to announce that @ultravox_dot_ai is now part of Inworld. Ultravox is the platform developers use to build real-time voice agents. We've worked with the team for a while through our TTS partnership, and today members of the team that built it are joining Inworld to keep developing it.
83
100
285
520,264
Together, we can improve the full conversation: understanding what a user says, helping them get something done, and responding in a voice that fits the moment. The first improvement is live today. Every built-in Inworld voice on Ultravox now runs on Realtime TTS-2, with existing voice IDs unchanged, no code changes required and no additional cost.
1
11
596
This is a big step in our work on speech-to-speech experiences that bring speech understanding, reasoning, and expressive voice generation closer together, and we're only getting started. Full post on why we came together, what you get today, and what's next: tinyurl.com/inworldultravox
10
407
Last week, Entrepreneur covered the launch of Inworld Realtime TTS-2. The article looks at how developers can direct voice delivery in natural language, from tone and pacing to pauses and nonverbal sounds like a laugh or a sigh, so the same line can be performed differently depending on the scene. It also covers voice design from a text description, cloning from 5 to 15 seconds of authorized audio, and carrying a voice across languages. On the model side, it walks through when to use TTS-2 for the highest voice quality versus Realtime TTS-2 Flash for high-volume, latency-sensitive applications, with a 25ms time to first byte independently verified by Coval. It also includes results from Talkpal's four-week A/B test and perspectives from LiveKit and ARX Media on what responsiveness and believability mean in live interactions. Read the full article: entrepreneur.com/business-ne…
8
3
17
2,075
TypeSafe's Jev is live on the Inworld Realtime Router. Jev is TypeSafe's new decision model. It doesn't generate text, you hand it a state and ask typed questions. Is this message spam? Which plan should this user be on? Does this code change need a human to look at it? It returns a yes, a pick, or a score, each with a probability attached. TypeSafe puts it at up to 200x faster and 400x cheaper than asking an LLM the same thing. Available now for no markup using the same API key you already use with Inworld.
5
2
11
961
Grok 4.7 is live on the Inworld Realtime Router. - xAI's new flagship, a larger base model with longer RL training - 500K token context - $2.00 in / $6.00 out per 1M tokens - Provider rates, no markup - Same API key as every model you're already routing
1
1
9
1,032
Inworld AI retweeted
Inworld Realtime TTS-2 debuts at #2 of 17 on the Voice Arena US English 🇺🇸 TTS Leaderboard at 1068 Elo, above Google DeepMind's Gemini 3.1 Flash TTS and inside the statistical band of Cartesia's Sonic-3.6 at #1. Realtime TTS-2 is Inworld's latest real-time TTS model, built for live consumer applications. The version on the board is the research preview, which Inworld has since taken to general availability with notable improvements on quality, stability, speed and consistency. Every ranking on Voice Arena comes from blind, head-to-head votes by vetted native speakers, scored with Bradley-Terry Elo. Realtime TTS-2 enters a 17-model field at #2 placing it above every model on the board from ElevenLabs, Microsoft, OpenAI and xAI, with clear statistical separation from all four. Results backed by 600+ blind, head-to-head listener votes from native US English speakers. Congratulations to @inworld on the release! See below for the exact audio raters judged and unedited listener verdicts 🧵
2
5
16
1,268
Inworld AI retweeted
inworld-tts-2 from @inworld is live on Speko. Of the 7 best-sounding models on our board it is both the fastest and the most robust: 116ms to first audio, and zero currency or number errors. Try it: platform.speko.ai
1
4
11
1,105
Inworld AI retweeted
A month ago, AGI House builders got early‑access hands‑on with Inworld’s latest TTS model. It’s now officially GA! Hear all the details in our new podcast episode ↓ @KylanGibbs is the CEO and co-founder of @inworld , and previously worked as a product manager on early LLM and text-to-speech efforts at Google and DeepMind. After watching that technology land almost exclusively in enterprise use cases, he set out to build the infrastructure that would let consumer-facing AI actually reach everyone. In this conversation with @jelares (CTO of AGI House), Kylan breaks down why voice is quickly becoming the default interface for AI, what it actually takes to serve that experience to millions of users, and why Inworld chose to build an end-to-end stack rather than stitching together providers. Timestamps 00:00 Intro 00:34 Kylan's Background & Founding Inworld 02:10 Inworld's Focus 03:41 Why Voice Is the Most Natural Interface 04:53 Overview of Inworld's New TTS Model 06:15 Natural Language Steering for Voice 08:04 Voice Cloning & Voice Design 08:38 Scaling Voice to Millions of Users 09:35 Personalizing Voice by Use Case 10:53 Cost & Latency at Scale 12:32 Why Latency Feels Like Meaning to Users 12:47 End-to-End vs. Piecing Together Providers 15:52 Performance Benefits of a Unified Stack 17:35 Inworld's Research & Inference Teams 20:20 Build vs. Buy: The Case for Using Inworld 23:33 Why Voice Adoption Is Reaching an Inflection Point 30:11 What Consumer AI Apps Look Like in 2029 36:04 Closing Thoughts
4
2
14
2,881
Inworld AI retweeted
Igor Poletaev, Chief Science Officer at @inworld, joins the Summer Signal '26 stage. Inworld just launched Realtime TTS-2 and TTS-2 Flash. He'll be talking through what it takes to make voice AI feel like it's in the conversation and customize it to varied scenarios. Register to attend 👇
1
3
12
1,091
Inworld's newly released Realtime TTS-2 is the new #1 on the Artificial Analysis Controlled Voice Arena, just ahead of Cartesia Sonic 3.6, and ranks #2 on our Provider Voice Arena behind Sonic 3.6 Realtime TTS-2 is a new Text to Speech model from @inworld that supports over 100 languages, including English, Hindi, Spanish, French, German, Chinese, and Japanese. Language can be set explicitly or detected from the text, including multiple languages in a single request (e.g., "I'll grab a coffee. ¿Quieres uno? お疲れさま。"). It also supports delivery instructions written in plain text alongside the input (e.g., [speak tired but warm, like she just got home from a long day]). Key takeaways: ➤ Controlled Voice: Realtime TTS-2 takes #1 on the Controlled Voice Arena with an Elo of 1,123 (+16/-16) across 1,292 appearances, 4 points ahead of Cartesia's Sonic 3.6 at 1,119. Sonic 3.5 follows at 1,096, then Inworld’s own Realtime TTS-2 Flash - Research Preview at 1,075, and ElevenLabs Eleven v3 at 1,062 ➤ Provider Voice: Realtime TTS-2 takes #2 on the Provider Voice Arena with an Elo score of 1,252 (+18/-18) based on 1,094 arena appearances, placing it ahead of Alibaba Qwen-Audio-3.0-TTS-Plus at 1,241 and Speechify Simba 3.2 at 1,240, but behind Sonic 3.6 at 1,282 ➤ Throughput: The model processes 106 characters per second of generation time, compared to 124 for Sonic 3.6, 97 for Simba 3.2, and 40 for Eleven v3 ➤ Pricing: Realtime TTS-2 is priced at $20.83 per 1M characters, lower priced than Sonic 3.6 ($49), Eleven v3 ($100), and Qwen-Audio-3.0-TTS-Plus ($27.59), but more expensive than Simba 3.2 ($10) See more details and listen to samples in the thread below ⬇️
6
18
189
24,472
Inworld AI retweeted
"Expressiveness" is the current buzzword in voice AI. And I have a simple question. WHAT. DOES. IT. MEAN. Thanks to @inworld for sharing stuff that I could not have just Googled and making my content better. Check out their new TTS-2 model! 00:00 Intro 02:32 The Blizzard Challenge for voice synthesis 05:13 Naturalness vs expressiveness 06:40 Evaluation via humans 08:38 Speech arenas 10:32 AI as a judge 14:58 Deterministic proxies for expressiveness 16:17 How do you train expressive models? 17:11 Text markers for expressiveness 18:48 Alignment post-training 21:35 Goodhart’s law 24:22 Expressiveness is not always the goal
7
10
100
7,123
Realtime TTS-2 is GA today, now the #1 model on Artificial Analysis and fastest TTS in the world of its class. The research preview got far more usage than expected. That experience went back into research and today's GA model represents the biggest leap in Inworld's history.
175
218
876
1,214,343
With Realtime TTS-2, you feel like a director working with a supernatural voice actor - allowing new levels of control and responsiveness. @livekit has been a close partner on TTS-2, with CTO David Zhao calling it “a real step forward in emotionally expressive voice synthesis.”
1
19
2,024
Realtime TTS-2 is live at inworld.ai Available through AudioStack, Cloudflare, DeepInfra, GMI Cloud, Livekit, Mastra, Pipecat, Runware, Stream, Telnyx, Tencent RTC, VoiceRun, and Voximplant Retweet the first tweet + comment “TTS” for 100 hours of free audio credits
5
384