PhD student @MIT in ML + robot learning | UC Berkeley EECS w/ Honors | Prev @nasa, @amazon, @berkeley_ai | Astronaut+Goldwater Scholar, NSF GRFP

There’s lots of buzz around agentic harnesses for robots, particularly for long horizon tasks that require complex reasoning and memory. But what will it really take to turn a reasoning agent into a reactive, reliable and low-latency robot policy? In our new paper, workspace models, we design a new memory architecture that acts as a latent harness for stronger reasoning models. Here’s why we think this might be the way forward: (1/9)
7
54
200
28,913
We emphasize that the training framework proposed in this paper can be extended to many capabilities outside memory as well! In particular, any useful computation identified by a reasoner can be amortized at train-time into a “latent harness” to support better control. Here's an example on a real Franka robot! (8/9)
1
1
13
798
Just took our G1 to the father-daughter dance. Maybe danceability is the new humanoid benchmark?
3
3
55
4,285
Crazy awesome humanoid skills 🔥
Taking humanoid soccer out of the lab: no instrumentation, no controlled environment. Real fields, around people. Trained in simulation. Deployed in reality. No post-training. Just pre-training. Dribbling. Passing. Turning. Scoring goals. Work led by Avi, Khai, @nolan_fey along with @martinpeticco @JohnMarangola, and Venki.
9
641
Fascinating and inspiring idea as we begin thinking about very-long-context continual learning!
In-context continual learning requires models to accumulate experience and reuse it later in the same sequence. But an RNN compresses an ever-growing history into a fixed-size state, where each token gets a single write into memory. We study dynamic compression: letting the model revisit the past and reorganize its state as it discovers what needs to be reused.
4
693
Cool application of SDFT!
Continual learning running on a phone. Not inference. Not a fine-tune you kick off with a dataset. A small language model (SLM) that keeps learning from its own experience while it's being used, with every gradient step on the handset. How it works (lin826.github.io/SLM-Online-…):
7
467
Great results from my even greater friend!
Excited to share SPD: simulation pre-training for dexterity. We pre-trained a policy in simulation and fine-tuned with less than 2 hours of real data (with @sarthakkamat)
8
2,023
Turns out using VLMs to guide exploration in prompt space can rapidly improve VLA performance!
You already know prompting can change what an LLM does. Turns out it can change what a robot does too — we made a robot learn a task it kept failing just by rewriting the prompt. No retraining. No new data. Better prompt in, better robot out. 🧵 1/8
10
1,353
Off policy BC + on policy DAgger for training memory! Awesome work
We never really knew how to train nonlinear RNNs well… BPTT struggled with vanishing grads (no long-range memory) and sequential rollout (hard to parallelizable). What if instead an oracle told us the optimal memory state m_t at each step? Then the RNN could do one-step supervised learning on (m_t, x_{t+1}) → m_{t+1} labels. We call this Supervised Memory Training (SMT): a replacement for BPTT that trains RNNs without unrolling them. SMT is time-parallelizable and solves vanishing gradients. Website: akarshkumar.com/smt/ arXiv: arxiv.org/abs/2606.06479
4
478
This is awesome! The less humans intervene on how to reason, the better. Let visual reasoning emerge via compression.
What if the best visual reasoning steps are ones humans can’t specify? 🤔 Existing VLM reasoning is often constrained by language, pixels, and human-designed intermediates. We introduce Latent Implicit Visual Reasoning, where we show that VLMs can discover the best visual reasoning steps by themselves — no bboxes, no intermediate images, no extra supervision. Presenting this week at CVPR! (1/n)🧵
6
446
Cool message on how we should use RL to train for diverse outcomes that cover many rewards. During test time, we can exploit what we know. Awesome work @RyanBoldi !!
Your RL post-training may be sabotaging your LLM’s test-time scaling! Conventional RL pretends that you can collapse all reward signals *upfront* into a single *scalar reward*. We introduce Vector Policy Optimization (VPO), which natively maximizes *vector-valued* rewards, boosting test time search performance, even on the original scalar.
3
338
Exciting to see the culmination of so much hard work from the @EkaRobotics team! Congrats @pulkitology @haarnoja @srinathm1359 and many more! These demos are taking rapid and dexterous manipulation to the next level 🚀
Eka means unity -- “one,” in Sanskrit and “first” in Finnish. We’re building intelligence for the physical world in its native language: forces. Until now, robotics faced a tradeoff — generality or speed. The real world requires both. Robotics also faced a data problem. Our Vision–Force–Action (VFA) model — the first of its kind — breaks the generality-speed tradeoff and the data barrier. It's a new foundation uniting performance, generality, and safety for putting capable robots in everyone's hands. Today, I am excited to share our journey of pushing robots beyond human limits. Today, dexterity becomes scalable. Today, I welcome you to the Era of Eka. Co-founded with @haarnoja, and so thrilled and grateful to be working with a dream team at @EkaRobotics. Learn more: ekarobotics.com
4
415
Nitish Dashora retweeted
From sorting chicken nuggets to screwing in light bulbs, Eka’s robots are eerily lifelike. But do they have real physical smarts? wired.com/story/when-robots-…
2
4
17
13,416