im not a guy. i love radiohead. prev:physics, redis kernel, cloud infra, robotics learning. now:model, inference, frontier ai research. hot nerd ig:roninkamuy

sf, palo alto
can’t believe it’s October already. 2025 is ending soon.
14
1
105
6,680
🦭went a beach party
18
223
18,716
💀🦭zelda and overwatch changed my dark life.
realizing how narrow my experiences were compared to others here (in addition to boomer ofc)
14
88
14,461
i said i'd study Miles and so here is my first little pr, using H100s on @sfcompute's givemeanode, now Autoresearch. thank you my friend @evanjconrad <3🦭 Autoresearch is the easiest GPU platform i've used! i used to write and maintain a lot of scripts to use rented compute efficiently. with Autoresearch, my coding agent could create nodes, run the Miles container, inspect failures, retrieve results, and shut down the H100 node when I no longer needed it - all in the same workflow as editing code. It saved me a lot of manual node setup, file management (and money!) the Diffusers version Miles pins could override the deterministic setting and wrapped FA3 in a custom op without a registered backward. my patch calls FA3's public autograd interface and preserves the setting through dispatch. on H100s, the attention tests matched numerical references and produced bitwise-identical outputs and gradients across 20 fixed-input runs per configuration, including 2- and 4-GPU Ulysses. i also added a tiny Wan FSDP2 training test comparing SP=1 and SP=2. GPU charges for this validation came to about $8, including setup and retries(very efficient!). draft PR: github.com/radixark/miles_di…
11
10
130
6,177
awesome work and awesome tech report! ever since minimax h3, there’s been so much acceleration work around omni models lately - everyone wants to make them real-time. it’s really interesting to watch these approaches converge. sparse and low-bit attention feel especially complementary: sparse attention reduces how many interactions you compute, while low-bit attention reduces the cost of each interaction. real-time omni models may end up being one of the strongest forcing functions for attention and inference efficiency.
Meet PixVerse R2, our new real-time world model. Explore living worlds. Control and edit them with prompts. Shape the story. Meet characters that remember and respond.
1
23
3,023
Asuka Zheng🦭 retweeted
Introducing Limite 1B - Violetto. A model for high-frequency mathematical intelligence.
64
228
1,824
286,860
Just released Parakeet Redux! A ternary speech-to-text model, built by compressing NVIDIA's Parakeet model from 1.2GB to 178MB. Runs at 113x realtime on CPU, and beats the base model on the 25-language FLEURS benchmark while staying within 0.3 WER on English.
2
4
116
6,440
2030: Your AI lawyer thinks you’re guilty, so the prosecutor subpoenas its activations.
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
2
45
3,656
>Their entire hacking cost less than $3,000 in tokens. >OpenAI awarded them $6,500.
OK, this is a big deal: 3 researchers used Claude Opus 5 to turn an image upload bug into an OpenAI employee account takeover, then had the compromised employee’s Codex open a PR in OpenAI’s internal monorepo. Their entire hacking cost less than $3000 in tokens. Opus 4.8 struggled with the exploit. Then Opus 5 dropped and cracked it within hours. AI-powered cyberattacks are becoming common and cheap. The best defense is to put the best AI in the hands of defenders too.
7
1
411
25,765
Following MiMo-V2.6’s live RL run. Useful to see learning curves alongside sampling stats, queues, costs, and restarts! The2B tokens per step is a substantial workload. What I’d like to understand is the compute allocation: at a fixed total budget, what produces the most improvement on unseen tasks? A few questions this run brings up: - More prompts, more rollouts, or better feedback? The dashboard shows ~1,568 prompts × 16 rollouts per batch. Covering more tasks and exploring more trajectories per task serve different purposes. Grader compute adds another choice: when is it worth generating fewer trajectories and spending more on evaluating them? The best allocation may shift as the policy improves. - What does agentic in-group credit assignment actually do? Test cases and rubrics provide different kinds of feedback. I’m curious how the grader compares attempts within a group, and whether it assigns credit to whole trajectories, stages, or individual actions. The ablation I’d like to see is whether additional grading cost pays for itself through more efficient learning. - Critic/RLOO choices also affect scheduling. RLOO uses sibling rollout returns as a baseline, a learned value function predicts return from the current state. If advantage computation requires the full group, slow trajectories can delay it even in an asynchronous pipeline. A value-based estimator that permits independent trajectory processing could reduce that dependency, but adds critic inference, training, and estimation costs. Retaining an RLOO component may retain the group dependency too. This is a design question, not a claim about MiMo’s implementation. The dashboard’s critic/* metric names alone don’t establish that it uses a learned critic. - Throughput needs to be read alongside policy lag. More trajectories per hour don’t necessarily mean faster learning. By the time a long-running task finishes, its behavior policy may be several versions behind. How the system handles that gap - and avoids systematically favoring short tasks - matters alongside raw throughput. The comparison I’d most like to see: time and total cost to reach the same held-out performance, accounting for rollouts, grading, environments, and any critic. Appreciate the MiMo team making the run visible. There’s a lot to learn from the intermediate behavior of a training system. mimo.xiaomi.com/rl/ Le Critique: arxiv.org/abs/2608.16739
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: mimo.xiaomi.com/rl/
4
1
76
10,053
my bet: future agent harnesses will be orchestrations of models, with trained smaller models handling routing, context isolation and memory updates. models deciding what other models get to see, remember and spend compute on.
Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier The AI industry has spent a decade optimizing along a single axis: build bigger, more expensive models. But intelligence has never been a monolith. It is a collective, distributed system. Humanity itself is a collective intelligence. We built Sakana Fugu on this conviction: the most powerful AI systems will not be isolated giants, but collaborative ecosystems that learn to coordinate. Evolution innovates under constraints, and the future belongs to systems that know not just how to solve a problem, but which machinery to deploy for the lowest possible cost. Today we are releasing Fugu Max and Fugu Ultra v2. Fugu Max expands the Pareto frontier outward, orchestrating our largest pool of open and specialized models to deliver frontier-grade results at a fraction of the token spend. Fugu Ultra v2 pushes that frontier upward, achieving peak performance on complex multi-step tasks without depending on the very frontier models it competes against. Together, they show that orchestration does not force a choice between cost and performance. It can push both at the same time, advancing the Pareto frontier. Relying on a single company’s model for critical infrastructure is a massive risk. As recent export controls have shown, access can disappear overnight. Collective intelligence is the practical hedge against this concentration of power. An orchestration system simply routes around vendor restrictions by relying on an entirely swappable agent pool. I am incredibly proud of our team for shipping this. By orchestrating the world’s models, we are building the resilient infrastructure required for AI sovereignty.
12
12
209
12,791
Asuka Zheng🦭 retweeted
la vida antes de este tweet:
today we launched ChatGPT. try talking with it here: chat.openai.com
186
41,706
441,170
7,506,491
I heard it's the best post-training framework right now, with really good engineering taste. Gotta study and try it out carefully this week! 👀
The full technical report for Miles is now available. Miles is built for production-level post-training, with stability, efficiency, and flexibility at its core. The report covers its system design and how it works in practice.
1
9
117
11,242
Asuka Zheng🦭 retweeted
So Astra is able to identify sounds from mel spectrograms zero-shot. I don't think we've scratched the surface of what this model can do (and this is light reasoning btw)
207
441
6,396
1,157,678
so again, most tasks can be considered coding tasks, and the results are natively programmable.
GPT-6 Astra built this Manhattan world in Unreal Engine over the course of a week. It was literally able to go street by street to make each one perfect.
7
2
73
6,274
rsi 😂
GPT-6 ASTRA IS GODLY STUPID IN CREATING 3D GAMES. this guy figured out how to create insane graphics using Astra and the trick was just image gen. steps to recreate: > connect Codex to blender mcp. > paste in the concept for the game. > tell Astra to use the Codex image gen skill to generate concept images of the target art style, then iterate until in-game screenshots look as close as possible to those, at 60fps set the reasoning to high. this was also one shotted in 45 minutes btw.
1
28
5,217
Asuka Zheng🦭 retweeted
Wow, this brought back a project Woojeong Kim, @srush_nlp, and I worked on in 2024. A strong model read a task description, turned it into "ability vectors", and passed them to a weak model. Sasha even came up with a great name for it: Strong-to-Weak Ability Transfer (SWAT) 😆 We never published it, but the idea eventually developed into ProgramAsWeights, which compiles task descriptions into small reusable neural programs that run locally. Very cool to see @aimalysheva and the Mostik team make a related idea work so well at this scale. programasweights.com
meet @mostik_ai! what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, which I can't say much about while the competition is still running. and this, which I can. everyone's arguing about whether open models will catch up to frontier models. we think it's the wrong question. here's the one we pose: why does a frontier model have to generate your answer at all, when the only thing you need from it is the reasoning? we do this by enabling models to communicate in latent space. through our protocol, hidden states pass straight from a frontier model into a small one running on your infrastructure -- no text between them, and neither model is fine-tuned. two models from different families, sharing reasoning, both left untouched. how do we know it works? we tested it on a setup where a 753B model reads the problem, and a 4B edge-class model writes the answer. with this approach, we get results 80% as accurate as the frontier model, but at 20x faster performance. we're committed to preventing frontier model lock-in and are already partnering with inference providers to accelerate open-weight adoption. we've done this between 15 of us, in four months, 12 PhDs and a Fields medalist, backed by @generalcatalyst WIRED has the first external account of the company and the work: wired.com/story/russian-star… full writeup, the setup, and all the numbers: mostik.ai/read-more
2
12
86
10,102
this is actually huge. sota models could directly provide hidden states to smaller models, allowing a 4B parameter model capable of running on edge devices to perform decoding. it's a new way to do fast and cheap inference. shut out to people who first made it. will agent talk in latent space in the future? probably.
meet @mostik_ai! what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, which I can't say much about while the competition is still running. and this, which I can. everyone's arguing about whether open models will catch up to frontier models. we think it's the wrong question. here's the one we pose: why does a frontier model have to generate your answer at all, when the only thing you need from it is the reasoning? we do this by enabling models to communicate in latent space. through our protocol, hidden states pass straight from a frontier model into a small one running on your infrastructure -- no text between them, and neither model is fine-tuned. two models from different families, sharing reasoning, both left untouched. how do we know it works? we tested it on a setup where a 753B model reads the problem, and a 4B edge-class model writes the answer. with this approach, we get results 80% as accurate as the frontier model, but at 20x faster performance. we're committed to preventing frontier model lock-in and are already partnering with inference providers to accelerate open-weight adoption. we've done this between 15 of us, in four months, 12 PhDs and a Fields medalist, backed by @generalcatalyst WIRED has the first external account of the company and the work: wired.com/story/russian-star… full writeup, the setup, and all the numbers: mostik.ai/read-more
8
5
115
15,824
I'm really curious about how Jürgen sees the world. The guy is standing so high up and looking so far ahead. The more you read his work, the more you realize it. A kid, an artist, and a master.
This "recurrent depth" is essentially what's in Sec. 5.3 of the 2015 paper: On Learning to Think: Algorithmic Information Theory for Novel Combinations of Reinforcement Learning Controllers and Recurrent Neural World Models arxiv.org/abs/1511.09249. This paper went beyond the inefficient millisecond by millisecond planning of my 1990 neural world models, addressing planning and reasoning in abstract concept spaces. The 2015 control network C is a prompt engineer that learns to create a chain of thought: to speed up decision making, C learns to query its separate neural world model for abstract reasoning. The prompts and the answers are internal self-generated sequences of vectors that don't have to represent natural language.
3
2
153
22,365