Byung-Gon Chun (Gon) retweeted
Do you always need a frontier proprietary model for coding agents? 89.6% at $0.02 per task vs 87.0% at $1.00. That's GLM-5.3-Flash next to GPT-6 Astra on SWE-bench Verified. The cheaper model solved more tasks. We ran 12 configs of 9 models on accuracy, cost, and time. Where open models win and where they don't, check it out 👉 friendli.ai/blog/open-vs-clo…
1
7
119
Byung-Gon Chun (Gon) retweeted
AI agents are taking on increasingly ambitious tasks, even Millennium Prize Problems 🫢. As tasks get longer and more complex, cost per token tells less of the story. Cost per completed task starts to matter much more. Cheaper tokens ≠ cheaper tasks. GLM-5.3 and GLM-5.3-Flash both beat GLM-5.2 on accuracy at a lower cost per task. Flash gets there with cheap tokens. GLM-5.3 gets there with 12–45% fewer agent turns, at the exact same price as 5.2. Read the full blog 🔗 friendli.ai/blog/glm-5.3-vs-…
1
6
280
Byung-Gon Chun (Gon) retweeted
Switching your coding agent to a new inference provider shouldn't mean digging through JSON, TOML, and YAML files. Introducing FriendliLink 🔗: An open-source CLI (frlink) that points the coding agent you already use — Claude Code, Codex, Cursor, OpenCode, Pi, Hermes Agent, DeepSeek Harness — at open models on FriendliAI Model APIs. One command to connect, one command to put it back. What it offers: → No config hunting — frlink claude on writes the key and base URL into the agent's own settings → Fully reversible — off restores your original config byte-for-byte → Reasoning that works — thinking controls pulled from Friendli's live catalog, translated to what each agent expects → No proxy, no daemon — after on, your agent talks to Friendli directly 👉 Get your API key on friendli.ai/suite 👉 Full blog: friendli.ai/blog/friendlilin…
1
8
251
Byung-Gon Chun (Gon) retweeted
FriendliAI is now a BYOK provider on @vercel AI Gateway. 🎉 Already on the Gateway? Route to FriendliAI with no SDK swap — just bring your Friendli API key. Pricing and credits carry through. ✨ GLM-5.2 and Gemma 4 listed today. Get a Friendli API key: friendli.ai/docs/guides/suit…
1
2
7
747
Byung-Gon Chun (Gon) retweeted
🚀 Day-0 support for GLM-5.3 is live on Friendli Model APIs. @Zai_org ’s latest open-weight model delivers major gains in coding and long-horizon agentic tasks, including open-weight SOTA on Terminal Bench 3.0 and Agents’ Last Exam. 🔗 Try GLM-5.3 now friendli.ai/models/zai-org/G…
2
6
145
Our team put two of the latest frontier open-weight models, GLM-5.3 and Kimi K3, to the test.
GLM-5.3 vs. Kimi K3: Which coding agent actually delivers? 🤔 We evaluated both models across standard coding agent benchmarks and an open-ended game development task. Read the full breakdown 👉 friendli.ai/blog/kimi-k3-glm… Get on the GLM-5.3 and Kimi K3 drop list 👉 qg2m6.share.hsforms.com/2qG7…
3
75
Byung-Gon Chun (Gon) retweeted
GLM-5.3 vs. Kimi K3: Which coding agent actually delivers? 🤔 We evaluated both models across standard coding agent benchmarks and an open-ended game development task. Read the full breakdown 👉 friendli.ai/blog/kimi-k3-glm… Get on the GLM-5.3 and Kimi K3 drop list 👉 qg2m6.share.hsforms.com/2qG7…
2
8
399
Byung-Gon Chun (Gon) retweeted
Interviewing our wonderful panel from last night! Thank you to all who showed up and listened in. We enjoyed hosting @nvidia, @OpenHandsDev, and @ollama at our SF office. Stay tuned for the next one! 👀
1
3
10
585
Byung-Gon Chun (Gon) retweeted
Your response_format validates. The output is still useless. 🤷 Constrained decoding only masks invalid tokens — it doesn't pick the right one. Structure ≠ intent. Bonus: up to 21% faster on FriendliAI via speculative decoding. Full breakdown 👇
Article

Structured Output Requires More Than Guided Generation

More deep dives on inference internals like this one: friendli.ai/blog → Structured output is a response from an LLM that follows a predefined schema, such as a JSON schema. This allows

2
3
169
Byung-Gon Chun (Gon) retweeted
🎂 Stack your AI coding cake! Inspired by @JensenHuang's “AI is a five-layer cake” framework, this game lets you build the full AI coding stack with four essential layers: 🧠 AI Agents 🤖 Open-weight Models 🔵 @friendliai — Inference 🟢 @nvidia — GPUs Inference is where it all comes together—and what FriendliAI does best. Play the game, screenshot your result, and share your AI coding cake in the comments. 👇 stack-your-ai-coding-cake.fr…
1
4
91
Looking forward to the event!
This Thursday, August 20th, we’re bringing together four leaders at the forefront of open models, coding agents, and AI infrastructure: Kyle Kranen, Senior Manager, Dynamo at @nvidia Yunmo Koo, Founding Engineer at @friendliai Dong Chen, Software Engineer at @ollama Saurya Velagapudi, Principal Engineer at @OpenHandsDev Together, they’ll explore what it takes to move AI coding agents to production and how open models are changing the economics of AI-native engineering. Join us for the conversation, followed by audience Q&A and networking. Spots limited! 📍 San Francisco 📅 Aug. 20 RSVP: luma.com/o9tv62y5
1
34
Byung-Gon Chun (Gon) retweeted
This Thursday, August 20th, we’re bringing together four leaders at the forefront of open models, coding agents, and AI infrastructure: Kyle Kranen, Senior Manager, Dynamo at @nvidia Yunmo Koo, Founding Engineer at @friendliai Dong Chen, Software Engineer at @ollama Saurya Velagapudi, Principal Engineer at @OpenHandsDev Together, they’ll explore what it takes to move AI coding agents to production and how open models are changing the economics of AI-native engineering. Join us for the conversation, followed by audience Q&A and networking. Spots limited! 📍 San Francisco 📅 Aug. 20 RSVP: luma.com/o9tv62y5
5
3
15
20,585
GLM-5.3 is out. Top-tier coding and agentic performance!
Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting a new standard among open models Tech Blog: z.ai/blog/glm-5.3
4
152
@OpenHandsDev, @ollama, and @friendliai take on whether open models — Kimi, GLM, MiniMax — have closed the gap enough to run production coding agents. One moderated panel, 30 minutes of open Q&A, and networking over food and drinks. If you own the budget, the headcount, or the vendor strategy, this conversation is for you. Grab your spot while there's still room! Aug 20, 6 PM → luma.com/o9tv62y5
27
Byung-Gon Chun (Gon) retweeted
Open Models, Production Agents: The New Economics of AI-Native Engineering August 20th, @friendliai is hosting @OpenHandsDev and @ollama for an event full of bringing together leading voices at the intersection of open models, coding agent infrastructure, and agentic engineering. Featuring Saurya Velagapudi, Principle Engineer at @OpenHandsDev and Jeffrey Morgan, CEO at @ollama Spots are limited, grab yours while they’re available! luma.com/o9tv62y5
1
2
6
497
Byung-Gon Chun (Gon) retweeted
MiniMax H3: Omni-Reference, Commercial-Grade Generation, Unbeatable Cost Efficiency, Open Weights Your creative destiny, on your terms. Now Live at HailuoAI.video & MiniMax API.
It's coming. @Hailuo_AI #MiniMaxH3
71
88
3,419
5,041,305
Byung-Gon Chun (Gon) retweeted
@friendliai has signed the open letter “Open Weights and American AI Leadership,” alongside @Microsoft, @nvidia, @Meta, @OpenAI, @ollama, and 50+ other organizations. Open weight models are the foundation of an accessible, competitive AI ecosystem. They let any team — startup, enterprise, university, or public institution inspect, adapt, and run advanced AI on their own terms, and they keep the benefits of AI broadly shared rather than concentrated in a few hands. We believe openness is one of the most important paths to innovation, competition, and even safety in AI. Glad to stand with this community. Read the letter: microsoft.com/en-us/corporat… #OpenWeights #OpenSourceAI
1
4
9
358
Byung-Gon Chun (Gon) retweeted
Welcome to FriendliAI, Brian Cho! 🎉 Joining as Director of Engineering with 6 years at @Meta and 4 at @AptosLabs under his belt — right as our engineering team scales to match the pace of the agent era.🚀
2
9
169
Byung-Gon Chun (Gon) retweeted
NVIDIA Nemotron 3 Embed is now on FriendliAI. Open embedding models built for agentic retrieval: → 8B: frontier accuracy, #1 on the RTEB leaderboard → 1B: designed to retain 95%+ of the 8B's accuracy, built for scale → 16K context, text + code, fully open weights, datasets, recipes, and fine-tuning guidance Deploy via Friendli Dedicated Endpoints, OpenAI-compatible API. 👉friendli.ai/models/nvidia/Ne… 👉friendli.ai/models/nvidia/Ne… 👉friendli.ai/models/nvidia/Ne… Read full blog 🔗friendli.ai/blog/nvidia-nemo…
2
2
14
1,007
Byung-Gon Chun (Gon) retweeted
Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost 🔹 Built for long-horizon agentic coding and self-evolving workflows Kimi K3 is now live on on Kimi.com, Kimi Work, Kimi Code, and the Kimi API. Open Weights by July 27, 2026. 🔗 API: platform.kimi.ai 🔗 Tech blog: kimi.com/blog/kimi-k3
1,737
7,584
56,995
25,542,289