I build Custom Voice AI & Agents for businesses DM for Work / Paid Technical Collabs โ˜• Shipping AI Products in Public @debugginglife25

Hyderabad / Bengaluru / SF
Sarvam AI opened Voice Agents today, so I built a Hyderabadi sabzi aunty ๐Ÿ˜ญ Called her and bargained in Hindi + Telugu + English, kept interrupting, switched languages mid-sentence, and she still remembered my order, gave me a final total, and even confirmed my (fake) UPI payment. Took me ~15 mins to build. It was fun to recreate the lost art of vegetable bargaining, but this time with an AI agent.
Today, weโ€™re making Voice Agents available to everyone. With Voice Agents, you can build and scale intelligent, human-like agents that understand context, remember past conversations, and get better at achieving your business outcomes with every interaction. Our Voice Agents have already powered more than 350 million conversations across enterprise deployments. Operating at this scale has helped us refine the platform, deepen its capabilities, and build out the complete stack needed to handle the full complexity of voice interactions. The platform brings together the tools and systems to personalise, orchestrate, govern, and launch Voice Agents at scale. Start building your Voice Agent today: links.sarvam.io/Voiceagents
Paid partnership (ad)
60
117
1,534
176,292
Vishal Singh ๐Ÿฅ‘ retweeted
Last month OpenAI launched GPT-Live-1 for realtime voice. I wanted the same interruptible conversation but running fully locally, so I built an open-source version of it. Meet OpenGPT-Voice๐Ÿ˜Œ The loop stays on my machine: Silero VAD โ†’ Faster-Whisper โ†’ GLiNER/Laya โ†’ local LLM โ†’ Supertonic TTS โ†’ WebSocket The harder (fun) part was the messy edges of real conversation: - Barge-in cancels the active LLM stream and queued audio immediately. - Playback-aware history: interrupt after 2 words of a 100-word reply and only those 2 words enter memory. - Timers and external tasks can finish and speak without derailing the current turn. - The router waits through pauses and self-corrections instead of calling the LLM on every silence. Built for people who want the voice loop inspectable and replaceable: local-first experiments, offline use, and edge setups where you donโ€™t want every utterance sent to a hosted realtime API. Still a POC. It breaks in funny ways. Iโ€™m still fixing them (constantly) ๐Ÿ’ช Demo + GitHub repo in the comments
GPT-Live-1 is now available in the API. Bring ChatGPTโ€™s natural back-and-forth to your app, with voice agents that listen while they speak and work with the models and harness you choose.
8
2
12
942
Last month OpenAI launched GPT-Live-1 for realtime voice. I wanted the same interruptible conversation but running fully locally, so I built an open-source version of it. Meet OpenGPT-Voice๐Ÿ˜Œ The loop stays on my machine: Silero VAD โ†’ Faster-Whisper โ†’ GLiNER/Laya โ†’ local LLM โ†’ Supertonic TTS โ†’ WebSocket The harder (fun) part was the messy edges of real conversation: - Barge-in cancels the active LLM stream and queued audio immediately. - Playback-aware history: interrupt after 2 words of a 100-word reply and only those 2 words enter memory. - Timers and external tasks can finish and speak without derailing the current turn. - The router waits through pauses and self-corrections instead of calling the LLM on every silence. Built for people who want the voice loop inspectable and replaceable: local-first experiments, offline use, and edge setups where you donโ€™t want every utterance sent to a hosted realtime API. Still a POC. It breaks in funny ways. Iโ€™m still fixing them (constantly) ๐Ÿ’ช Demo + GitHub repo in the comments
GPT-Live-1 is now available in the API. Bring ChatGPTโ€™s natural back-and-forth to your app, with voice agents that listen while they speak and work with the models and harness you choose.
8
2
12
942
$5k/month looks big until you do the math 500 customers at $10 this app has 500M+ active users you don't need to win the internet just be useful to 500 humans thatโ€™s it.
1
14
787
I kept wondering: Why is @Cerebras so damn fast? So I went down the rabbit hole ๐Ÿ˜… Inference, memory bandwidth, SRAM, wafer-scale compute, TTFT, agent latency I finally started to understand how the pieces fit together, so wrote down what I learned ๐Ÿ‘‡
Article

I Finally Understood What Makes Cerebras So Fast

I've been seeing Cerebras everywhere lately. The crazy token/sec numbers caught my attention first. But I realized I didn't actually understand why it was so fast. I had this very simple mental model:

2
2
8
2,299
wanted to test Gemini 3.8 Flash + I hate making Google Forms manually so built baatsheet ๐Ÿ˜ญ ~ talk to it, get a production-ready form + AI response insights in seconds fun little voice ai project Quick demo + GitHub repo in the comments ๐Ÿ‘‡
4
1
16
4,192
What it solves: - โณ Hours work โ†’ seconds โ€” a form that took 20 minutes of clicking now takes just one sentence via Voice. - ๐ŸŒ Language barrier โ€” speak in Hindi, get a complete form made entirely in English (questions and options). - ๐Ÿง  Writer's block โ€” don't know what to ask? Gemini designs context-aware questions and realistic options for you. - โŒจ๏ธ No typing required โ€” speak your prompt hands-free in Hindi or English (added 4 Indian languages support for now but more can be added). Here's the GitHub repo ๐Ÿฅ‚: github.com/vishalsingh2972/bโ€ฆ
1
363
Me and Claude Opus 5.5 tonight
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
2
2
42
6,212
ok this is insane. at effectively the same frontier intelligence tier, grok 4.7 is - 8.3ร— cheaper output than GPT-6 Astra - 8.3ร— cheaper output than Fable 5.1 - 5ร— cheaper input than GPT-6 Astra - 5ร— cheaper input than Fable 5.1 to put this into perspective, $100 of grok 4.7 output would cost $833 on GPT-6 Astra or $833 on Fable 5.1
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
5
1
11
9,070
ok i was wrong this model is good on paper but the outputs are horrible, hopefully things get better with 4.8 ๐Ÿคž
1
114
fuck this is pretty cool ๐Ÿ˜‚ It finds the things sitting in your inbox that you said youโ€™d get back to, figures out what needs to happen, and actually moves them forward, research, docs, slides, meetings, even pulling context together from different emails. this is the AGI Iโ€™ve always wanted, my inbox desperately needed this lol ๐Ÿ˜ญ
The obvious is missing. So we built Sol - hellosol.app Sol finds the work itself, does it, and comes back for your approval. Every day in our emails we say "Iโ€™ll shareโ€, "I'll reviewโ€, "I'll get back" - then repeat the exact same thing to an AI. Why? Sol finds everything you said youโ€™d do & gets them started for you. It does the research, creates the doc, builds the slides, finds the time, connects the dots across multiple emails, doing everything it takes to get the job done - but doesnโ€™t send, schedule, or share anything until you approve. Sol runs on its own computer, uses a browser, and has a library of skills that automatically get assigned to the work that needs to get done. No setup. It just starts working. We've raised $4M from General Catalyst, Nexus Venture Partners, DeVC, PeerCheque, Kunal Shah, and a few others. Extending early access now. @generalcatalyst @nexusvp @DeVC_Global @peercheque @neerajarora @b_jishnu @kunalb11 @miten @RTinkslinger @Rahul_J_Mathur @AkarshS27 @SiddhantD06 @RajatAgarwal167
Paid partnership (ad)
1
7
1,727
came across Halo and this caught my attention: take the Hugging Face model you already have and scale training without rewriting it for another framework distributed training without giving up the HF workflow SFT, distillation and RL on the same stack too pretty neat ngl
Introducing Halo, the best framework for post-training of open-source models. Halo delivers up to 2.8x the throughput of stock TRL with less peak memory, while models stay in their native HuggingFace format. Star us on GitHub: github.com/whitecircle/halo
Paid partnership (ad)
14
1,184
โ€œWhat if @elonmusk was born in Bihar?โ€ Made this AI audio episode in 2 min ๐Ÿ˜‚ Sherpa turned it into a full series with Trump as the villain, Sam Altman + Dario Amodei as the jealous cousins, and Elon's journey from Bihar to Mars ๐Ÿš€ possibilities especially for ๐Ÿ‡ฎ๐Ÿ‡ณ short storytelling is insane with this Quick 2 min demo ๐Ÿ‘‡
Introducing Sherpa: the most advanced fiction writing AI We accelerated from $250M in ARR to $500M because Sherpa helped increase content production by 1200% in 1 year Sherpa was trained on 5.5B hours of playtime with minute by minute dynamic retention data. 550K+ creators have produced 2.6M hours of content annualised using it Pocket FM is like Netflix for audio-only dramas, with our own pool of one-person studios. 10% of eligible writers on Pocket FM make >$200K One blockbuster produced >$100M in revenue 3 writers have become millionaires in <2 yrs We built Sherpa to enable anyone to make >$1M by writing world-class fiction stories: 1. The Idea: Drop a 1-2 sentence concept. Sherpa interrogates it like a veteran editor on tension, stakes, and psychology 2. World & Characters: It builds out the complete lore, tone, and character psychologies 3. Sub-Plot planning: Breaks the premise into arcs, arcs into episodes, and episodes into scenes 4. Scene-by-Scene Generation: Outlines and drafts entire episodes, with you able to steer, rewrite, or override anytime 5. Editorial Review: Stress-tests every draft for pacing, engagement drop-offs, prose, and coherence before it locks 6. One-Tap Production: Pick a voice, convert to audio drama, and publish directly to Pocket FMโ€™s millions of listeners 7. Global Scale & Monetization: Revenue-share on performance, with automatic localization so you earn across international markets Test Sherpa for free here: pocketfm.com/sherpa _____________________________________________ Generic LLMs fail at serialized fiction because they lack a long-horizon narrative reward function. Sherpa solves this through three core technical leaps: 1. Narrative World Model (State Tracking & Retrieval): Context windows degrade over long runs. Sherpa constructs an evolving semantic knowledge graph tracking character states, secrets, and plot dependencies. High-speed retrieval surfaces exact context on demand, maintaining zero continuity decay across hundreds of episodes 2. Hierarchical Story Planner: When writing a 500-episode story like Naruto, you need to plan 100s of sub plots. Rather than generating linearly, Sherpa decomposes narrative across discrete levels: season -> arc -> sequence -> episode -> scene. Rather than generating everything upfront, like a generic LLM, Sherpa uses progressive planning and dynamic replanning. As the story evolves, it identifies what changed, traces the downstream impact, and replans only the affected parts. 3. Prose Engine (Trained on series' retention data): LLMs write robotically, but serial fiction needs emotion, tension, pacing, and dialogue that sounds like real people. Sherpa's Prose Engine was designed specifically for storytelling. It was built on 1B+ tokens of Pocket's own stories, trained by learning from what listeners engage with, where they drop off, and what keeps them hooked. Feedback is taken from specialized evaluator models that measure every scene against a 40-item checklist. (Evaluator models were benchmarked against human reviewers and matched them 80โ€“90% of the time.) _______________________________________________ Owning distribution and creation puts us in a very unique spot. More shows -> More data -> Sherpa becomes better -> more creator success -> more creators -> more shows Pocket FM has already seen one $100M IP. I believe Sherpa will soon lead to dozens of single-person studios creating billion-dollar shows. Most people are scared of AI but I think it'll unlock more human creativity, help creators earn more, and bring the next great IPs to life. This will create millions of jobs and new income streams.
Paid partnership (ad)
3
8
1,549
jev is hottest topic in tech rn ๐Ÿ”ฅ but how would you describe jev to someone non-technical? lemme explain it in the simplest way possible ๐ŸŒน jev is really good at answering multiple choice questions (and telling you how confident it is) and a surprising number of real-world problems can be broken down into lots of small multiple choice questions, solved one after another so software can just act on the answers. thatโ€™s basically the main idea, awesome work @typesafeai ๐Ÿ‘
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? Iโ€™ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev โ€ข 20-200x faster โ€ข 40-400x cheaper (w/ output tokens free) โ€ข Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
Paid partnership (ad)
4
14
1,997
Tested Gemini 3.8 Live on 3 of my favorite books and the results were insane ๐Ÿ”ฅ I pointed the camera at: ๐Ÿ“– Bhagavad Gita โ€” asked it to explain a Sanskrit sloka + make it relevant to a 20s Gen Z developer ๐Ÿง  Atomic Habits โ€” showed it a diagram and asked for a quick explanation ๐Ÿ˜ด Chicken Soup for the Soul โ€” gave it a page and asked for a warm bedtime-story version No PDFs. No uploads. Just the phone camera + my ugly voice ๐Ÿ˜‚ I genuinely didnโ€™t expect it to understand the context this well. Awesome job @GoogleAI Team โค๏ธ Demo ๐Ÿ‘‡
Say โ€œhiโ€ to our most advanced audio models from @GoogleDeepMind yet built for natural, production-ready voice applications. ๐Ÿ”ท Gemini 3.8 Live ๐Ÿ”ท Gemini 3.8 Live Extended Thinking With these models, you can speak naturally, collaborate easily, and tackle complex tasks using just your voice.
Paid partnership (ad)
3
2
42
9,993
What if an AI agent could actually use a website autonomously? Not just look at it or click around. But actually, understand what it can do, and do it. I still donโ€™t fully know where this goes, so I built a little experiment project, Agent-Atelier. Itโ€™s a luxury fashion storefront powered by native WebMCP, exposing 34 tools through document.modelContext. The agent can: โ†’ search & filter products โ†’ compare items โ†’ check size & stock โ†’ manage the cart โ†’ navigate the site โ†’ go through multi-step shopping flows I also hooked the same tools up to chat + voice, which made the whole thing way more interesting. I have so many more things to figure out here(trying to do rn): โ†’ what should the agent be allowed to see? โ†’ how should capabilities change with state? โ†’ when should a human have to confirm something? โ†’ what happens when the agent needs to chain multiple actions? I got these in mind but no idea how to implement them still figuring out tbh. But watching an agent discover the website and actually do things on it for the first time was kinda insane ngl. Quick demo + GitHub repo in the comments ๐Ÿ‘‡
3
2
8
1,919
Built a Voice AI agent in 5 mins, this is illegally good. Love it ๐Ÿ˜ป Opened CallKaro just to see how fast I could ship something real. Made a used-car buyer qualification agent for @myspinny ๐Ÿš—: prompt โ†’ conversation flow โ†’ Hinglish โ†’ budget/intent extraction โ†’ guardrails โ†’ post-call summary Then I tried to break it like an actual buyer: โ€œsix fiftyโ€ โ€œautomatic better haiโ€ โ€œ7 lakh ke andar milega?โ€ โ€œdiscount kitna doge?โ€ โ€œloan approve ho jayega na?โ€ - It caught the vague answers. - Asked the right follow-ups. - Didnโ€™t hallucinate discounts, availability, or loan approval. โœ… Thatโ€™s where it gets super interesting. Making an agent talk is easy. But getting conversation logic + guardrails + structured extraction + natural voice to hold together in one call is the actual job. Pronunciation still needs some work, 5 mins was not enough for me to teach it everything lol ๐Ÿ˜ญ Zero โ†’ a working conversational voice agent ready to deploy in ~5 minutes. ๐Ÿคฏ Honestly, crazy how fast this stuff is getting.
Paid partnership (ad)
5
3
14
1,804
Quick weekend Voice AI project ๐Ÿ”ฅ Built Jaldi Bol + tested it with a real visa processing/travel-tech company @atlys ๐Ÿ“Š My ๐—Ÿ๐—ฎ๐˜๐—ฒ๐—ป๐—ฐ๐˜† ๐—•๐—ฟ๐—ฒ๐—ฎ๐—ธ๐—ฑ๐—ผ๐˜„๐—ป: ๐Ÿ—ฃ๏ธ Deepgram Nova-3 โ†’ ~125ms ๐Ÿง  Gemini Flash Lite โ†’ ~180ms TTFT ๐Ÿ”Š Cartesia Sonic 3.6 โ†’ ~105ms TTFA ๐ŸŽ™๏ธ Silero VAD โ†’ ~180ms โšก LiveKit WebRTC โ†’ ~50ms Put it all together โ†’ I got ~520ms p50 end-to-end latency (which I felt was pretty great ngl ๐Ÿง˜) I know how annoying it is when you donโ€™t get a reply back, especially when youโ€™re waiting for an answer and the clock is ticking. ๐Ÿ˜ญ So, I made Jaldi Bol to quickly handle: โ†’ Visa status โ†’ Document issues โ†’ Urgent cases โ†’ Human escalation, if needed The flow here was pretty simple: Visa delayed โ†’ Jaldi Bol checks โ†’ answers โ†’ escalates if needed. Next, went a little deep on the voice loop part: โœ… Full-Duplex Barge-in: Instant interruption cutoff (<300ms) with zero echo feedback. โœ… Multi-Agent Orchestration: TriageAgent โ†’ VisaAgent โ†’ EscalationAgent โœ… Zero Hallucination: 100% grounded against visa database schemas. โœ… Short 1โ€“2 sentence responses for natural conversations Biggest learnings (after aggressive testing): โ†’ Streaming everything matters more than just a faster LLM. โ†’ Barge-in is huge. Interrupt it โ†’ it shuts up and listens. โ†’ Shorter responses sounded way more human. Started this as a latency test. Wanted to see firsthand how fast I could push a voice agent before it actually started feeling conversational. Demo + GitHub repo in the comments ๐Ÿ‘‡
Paid partnership (ad)
9
5
21
1,529