I have 57 hours left to burn through $600 worth of Kimi K3 that I paid for with $48.
Kimi K3 is now 50% off on Zro. Same model. Half the price. Run it in the coding agents you already use. New to Zro? Bring your next project, we’ll handle the inference. Available through September 24. zro.moonmath.ai
2
7
644
Better than OpenCode Go
DeepSeek 4.1 at current Zro pricing: 3× plan → 6× usage 5× plan → 10× usage zro.moonmath.ai
2
1
10
1,067
Replying to @grok
actually, if we divide the weekly limit by 7 and multiply by 365, the number would be a bit higher... and you reach 376x if you also squeeze in the value of X Premium Plus.
57
This is the cheapest Kimi K3 in the market right now!
Kimi K3 is now 50% off on Zro. Same model. Half the price. Run it in the coding agents you already use. New to Zro? Bring your next project, we’ll handle the inference. Available through September 24. zro.moonmath.ai
1
6
588
This can be very useful for verifying ZK proofs on Bitcoin
1/ Today we're releasing Lattice Jolt: a post-quantum version of the Jolt zkVM, built on lattices instead of elliptic curves. It's _faster_ than curve-based Jolt and has the shortest proofs of any post-quantum zkVM — under 100 KB. Post: a16zcrypto.com/posts/article…
1
1
37
2,335
He is not wrong actually.
The following number divides RSA-270: 233108530344407544527637656910680524145619812480305449042948611968495918245135782867888369318577116418213919268572658314913060672626911354027609793166341626693946596196427744273886601876896313468704059066746903123910748277606548649151920812699309766587514735456594993207
1
10
2,857
Oh, this is @ChatGPT 6 Astra. It can ride Space Mountain in @Disneyland. Videos from @LMGVids
5
649
Zro is providing DeepSeek v4 Flash and GLM-5.3 Flash for free for the next hour. API keys below (!) Try them! Specifically, let GLM-5.3 Flash do a security code review for you. Can GLM-5.3 break your project within an hour? You are about to find out.
we are live, and for the next few hours we are free! npm install -g @moonmath-ai/zro zro login --api-key sk-gjnjRysmpJbmdzcSKWpHXw zro launch opencode --model deepseek-v4-flash-0731 npm install -g @moonmath-ai/zro zro login --api-key sk-gjnjRysmpJbmdzcSKWpHXw zro launch hermes --model glm-5.3-flash
5
624
Due to unforeseen circumstances, this performance of coding with SuperGrok, Claude, Cursor, ChatGPT cannot continue. Grok Bot is also gone. My surviving coding plans are @zroai_ and the China-only subscriptions of Qwen and WorkBuddy.
1
8
1,440
is this some new GitHub virus to open-source projects?
3
15
4,020
I start to wonder if a large number of inference providers, not limited to OR, including OpenAI and Grok and Cursor, have the similar issue, related to how to bill when the client stream stops and cannot finish the current turn. The best defense is probably to detect unusual amounts of disconnections.
1/ Here’s a demo showing how I generated tokens for free on @OpenRouter using the @FireworksAI_HQ endpoint for GLM-5.3
1
1
2
1,005
The solution is easy: use Kimi K3 in Cursor for Sol (when you really need it) and Grok 4.6 for everything else. Kimi K3 doesn't have the Sol Ultracode Subagent Mania issue, which is a good thing...
We’re ending our partnership with Cursor following its acquisition by SpaceX. Under our proposal, Cursor’s direct access to our models would end on November 12. We know that the people most affected by this decision are the developers who rely on OpenAI models in Cursor. We care about their experience in this transition and we’re ready to go above and beyond to support them. openai.com/index/our-decisio…
1
7
966
Weikeng Chen retweeted
GLM‑5.3 Flash is now live on Zro. We’re giving it special treatment from day one, with a focus on fast, reliable serving. More performance improvements are coming over the next few days.
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5… API: docs.z.ai/guides/llm/glm-5.3… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai
2
5
15
1,197
So DeepSeek did the SFT (to some extent, but overfit on the harness) and didn’t have enough time doing RL?
Stanford这帮人可能找到了今年最狠的Claude、ChatGPT平替路线,我直接把手里所有闭源模型订阅全停了,这可能是今年开源AI最狠的一次逆袭。。 Stanford的人用DeepSeek V4 Flash, 同一个任务生成5个答案, 再让同一个模型自己当裁判打分, 挑最高分的那个交卷。 Terminal-Bench 2.1, 79%直接干到88%, 超过了Claude最新的Fable 5, 成本只有对方的十一分之一。 全网都在喊开源吊打闭源, 但90%的人没看懂真正的信号。 生成端早就卷到头了。 真正拉开差距的是验证。 以前大家拼模型参数, 拼训练数据, 拼基准测试高0.5。 现在发现, 同一个便宜模型多跑几次, 再自己挑出最好的, 就能干翻贵11倍的旗舰。 这就是test-time scaling的真正玩法。 模型不用做大, 推理时多花一点算力, 用生成加验证换效果。 验证本身还能继续缩放, 更细的评分标准, 多验证几轮, 把评价维度拆得更碎。 框架已经开源, 任何人都能拿DeepSeek V4 Flash直接复现。 模型是引擎, 验证才是刹车和方向盘。 引擎够猛就行, 没有刹车和方向盘谁也不敢开快。 我觉得接下来所有agent框架都会默认加上自验证这一层吧 hhh
619
Today is a good day to run GLM-5.3 on all your repos for a security audit with OpenCode Go.
GLM-5.3 now available in Go text · 1M context · same pricing as 5.2
2
9
896
I have been working on this for a few months. It currently supports DPF-PIR, HarmonyPIR, OnionPIRv2, and now also ORAM with TEE client proxy. Lightning integration for paying for PIR and Nostr server discovery on the way.
.@weikeng is building Private Information Retrieval (PIR) for bitcoin: query the UTXO set without revealing what addresses you are interested in. It’s a rare way to gaze into the bitcoin abyss without it gazing back.
8
13
49
5,138
Before: GPT 5.6 Sol Ultracode + /goal (ran out of usage limit in 2 days) Now: GPT 5.6 Sol High + “use Terra for subagents not Sol”, “let subagents do coding”, “my usage limit is tight please have mercy”, “I will notify you when CI is done, please do not loop wait” (still at 99%)
5
494
I was one of the users who sent that idea to gdb back then when he was collecting ideas :) I wonder if OpenAI can gift a banked reset for predicting what eventually goes into products. @thsottiaux
OpenAI appears to be working on a feature that would let you purchase a rate limit reset for Codex Source: public checkout pricing config & ChatGPT web app assets
1
434
My current use of Zro: (1) Kimi K3 as Sol, (2) GLM-5.2 as Terra, (3) DeepSeek v4 Flash as Luna. The inference is fast, especially for Kimi K3. When you need independent code review, GLM-5.2 can help.
Today we added @AMD MI325X as a new serving backend for @zroai_. Kimi K3 is the first model running on it, powered by @sgl_project. We believe this is the first production support for the official Kimi K3 weights on CDNA3. With 256 GB of HBM per GPU, the full model fits on a single 8×MI325X node. This is likely the lowest-cost hardware configuration capable of serving Kimi K3. The benchmark below shows solid serving performance, even before adding any of MoonMath’s custom kernels. Huge thanks to @digitalocean for their partnership. ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Here is everything we changed to make Kimi K3 work on CDNA3 with SGLang: 1. Rerouted MXFP4 experts to AITER’s Triton GEMMs: AITER’s default 4-bit MoE path selects a FlyDSL kernel on gfx942, then fails during compilation. Its A4W4 kernels are gfx950-only. AITER states this in the source, and AMD’s MI35x tests note that “MXFP4 does not register on gfx942.” Without this change, K3’s MoE has no working execution path on CDNA3. We enable the Triton route automatically on gfx942. 2. Fixed an AITER GEMM crash on sliced activations: K3 passes the tuned GEMM path a 1,536-wide slice of a 2,112-wide tensor. The launcher rejects its strides and throws; under graph capture, that becomes fatal instead of falling back. The tricky part is that PyTorch ignores the stride of size-1 dimensions, so the view reports as contiguous and .contiguous() does nothing. The fix explicitly checks for the canonical strides required by the launcher. 3. Fixed a recent regression in K3’s SiTU activation: An August 1 upstream commit added an unguarded CUDA-only include, breaking the kernel build on every AMD GPU. Because the failure appears during graph capture, it looks like a graph issue rather than a missing kernel. We route ROCm to SGLang’s equivalent Triton implementation. Upstream, a two-line ifndef USE_ROCM guard would fix MI355X as well. 4. Selected the required graph-capture mode up front on ROCm: Upstream replaced graph recapture with a validator that raises an error. Under speculative decoding, capturing a mode that is too weak now kills the run instead of triggering a recapture. We force the correct hidden mode before capture begins. 5. Added support for 12 MLA heads per rank: With TP8, K3’s 96 heads become 12 heads per rank. AITER has no MLA kernel for that shape: it rejects the configuration at startup and can later call abort() from C++ without a Python stack trace. We bypass the startup assertion and zero-pad the query heads from 12 to 16 so a supported kernel can run. 6. Enabled the configuration the stack actually requires: The working setup needs: SGLANG_USE_AITER=1 SGLANG_AITER_K3_OPT=1 <head-padding flag> --trust-remote-code SGLANG_USE_AITER=1 enables the AITER MoE route. Without SGLANG_AITER_K3_OPT=1, expert weights are padded to 256 and TP8 runs out of memory during loading. Missing --trust-remote-code fails only after loading roughly 1.42 TiB of weights.--kv-cache-dtype fp8_e4m3 is not inherently required by K3 or ROCm; our current attention path requires it because it does not yet include a BF16-KV kernel. All credit goes to the team - I am here simply to report their work 💐
1
2
5
639