One really cool use case I’ve been exploring with Jev: a semantic VAD for real-time voice AI.
I plugged it into @AgoraIO ConvoAI and used the live transcript + recent conversation as context to decide whether the user has actually finished speaking.
That makes Jev surprisingly good at handling things like hesitation, unfinished thoughts, and those moments where you stop talking for a second because you’re still figuring out what you want to say.
It’s a nice example of where Jev fits really well in a real-time AI loop.
Here’s what it looks like in action 👇
Some cool things developers have been building with Agora lately 👀
☕️ Spidey is a friendly AI companion that helps you manage tasks by voice. By @HaimantikaM .
⚔️ Voice Arena brings two voice agents with different turn-taking and interruption settings into a live battle. By @HarishKotra .
🎙️ Jev VAD explores when a voice agent should listen or keep talking. By @SoCalJayF .
🎧 An interactive podcast where you can jump in and talk to two AI hosts. By @zicojzc.
Built @typesafeai Jev with @AgoraIO to control my Mac by voice - play a video, open folders, navigate the web, and close tabs.
Here’s a quick demo 👇
Thank you @BhosalePratim for the inspiration and @hermes_f for the video editing cool tricks from @OpenAIDevs
I've always wanted a way to talk to my Mac and do things on the go.
Until now, it was super slow.
So I built Macbrow.
Voice-controlled browser and computer-use agent.
Built using @GradiumAI@browser_use and @typesafeai
Still early and a work in progress.
github.com/timpratim/macbrow
One really cool use case I’ve been exploring with Jev: a semantic VAD for real-time voice AI.
I plugged it into @AgoraIO ConvoAI and used the live transcript + recent conversation as context to decide whether the user has actually finished speaking.
That makes Jev surprisingly good at handling things like hesitation, unfinished thoughts, and those moments where you stop talking for a second because you’re still figuring out what you want to say.
It’s a nice example of where Jev fits really well in a real-time AI loop.
Here’s what it looks like in action 👇
Tone and delivery shape how spoken communication is understood, and Google’s latest Gemini 3.8 Flash TTS and Flash-Lite TTS models address a persistent limitation in voice AI by giving developers finer control over how generated speech sounds and adapts throughout a conversation.
We’re excited about the new models’ ability to direct emotion, pace, character, and accent – plus a library of 20,000+ voices!
Our thoughts on what this opens up for real-time voice builders: agora.io/en/blog/a-new-gener…
Create and deploy custom audio with our new text-to-speech models:
🔵 Gemini 3.8 Flash TTS: Design unique voices with distinct accents and characteristics.
🔵 Gemini 3.8 Flash-Lite TTS: Built for efficiency and scale, choose from your created styles or our expansive production-ready library.
Vercel will be officially sponsoring tailwindcss.com. That's a given. We as a community and industry owe @adamwathan and team a lot. Tailwind is foundational web infrastructure at this point (it fixed CSS 😉). I've also reached out to Adam to explore how we can make this a longer-term commitment.
Argentina wins the World Cup!!!!!!
Messi is the greatest of all time!!
Legendary player. Legendary game. The greatest sporting event I've ever seen in my life.
I love you all!!!!!!!!!!
Wondering why your Thanksgiving groceries cost more this year? It’s because greedy corporations are charging Americans extra just to keep their stock prices high. This is outrageous. businessinsider.com/big-comp…