musing in ai + robotics. documenting the journey @tofufolabs

dreams are important 🌄
Pinned Tweet
mood of life right now.
4
2,042
just getting started
3
14
1,032
Replying to @tata_neu
@tata_neu is such a scammy app. you try and book flights and they charge you 30% higher than what you could find on MMT. They showing 10k for 16th Oct, none of the flights match the price
3
1
129
pls suggest better credit cards
15
assembling the so-101 is the hardware version of installing CUDA before chatgpt. It makes your soul scream.
1
110
Some exciting personal news: I'm starting a newsletter in the Understanding AI family: Understanding Robots! (@binarybits will be an occasional AV correspondent) Up first, I explain GPT-Astra 6's impressive robot capabilities and what it might mean for the future of robotics
10
21
196
30,758
ChatGPT has gone so good. Pura apple vibes
59
booth #1102
Happy #IROS2026! Stop by the RAI Institute booth #1102 to grab some stickers and check out some of our custom-built robots including the UMV, Roadrunner, and Koala grippers. You can also catch Roadrunner in action in the robot demo area. Check the Exhibit Hall demo schedule for more details!
3
452
IROS 2026 is ongoing. So here's the top 10 papers/posters you can try exploring.. 1) Scaling Cross-Embodiment World Models for Dexterous Manipulation Human hands + very different robot hands are represented as 3D particles, giving one world model a shared space for planning across embodiments. arxiv.org/abs/2511.01177 2) AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models A latent world model scores candidate robot actions during offline VLA post-training instead of requiring endless physical rollouts. arxiv.org/abs/2603.08519 3) OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Models Instead of forcing a VLA to learn every camera viewpoint, reconstruct the scene in 3D and render canonical views first. og-vla.github.io/ 4) RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation A nice middle ground between cheap 2D video models and expensive full-3D world models: use a geometry-aware 2.5D representation. arxiv.org/abs/2510.09036 5) TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation Touch tokens are activated when contact happens, instead of feeding tactile information into the VLA all the time. arxiv.org/abs/2603.12665 6) AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint-Robust VLAs Move the robot camera and the policy can break. AnyCamVLA re-renders the new observation toward the viewpoint the VLA was trained on. arxiv.org/abs/2603.05868 7) LangGap: Diagnosing and Closing the Language Gap in VLAs Some VLAs can score extremely well while barely using the instruction. LangGap keeps the scene fixed and changes the language to expose that shortcut. arxiv.org/abs/2603.00592 8) VG3S: Visual Geometry Grounded Gaussian Splatting for Semantic Occupancy Prediction Frozen vision foundation models provide geometric priors that help build better Gaussian-based 3D occupancy representations. arxiv.org/abs/2603.06210 9) EquiBim: Learning Symmetry-Equivariant Policy for Bimanual Manipulation If you mirror the scene and swap the robot’s arms, the action should mirror too. Simple physical symmetry turned into a learning constraint. arxiv.org/abs/2603.08541 10) Learning Tactile-Aware Quadrupedal Loco-Manipulation Policies The robot does not just predict where its hand should move. It also learns what the desired tactile interaction should feel like over time. arxiv.org/abs/2604.27224
3
9
77
5,878
bro wtf is going on .. why wont it stop
1
75
This blog post explains why hand-eye calibration is mostly a problem of getting your coordinate frames right. > the goal is simple: connect what the camera sees to the robot’s coordinate system. > Samarth keeps the frame notation explicit, like ^A T_B, so you can actually see what every transform means. > the core calibration problem becomes: AX = XB robot motion on one side, observed camera/target motion on the other, unknown extrinsic transform in between. > one subtle point: move the camera from fixed-in-the-world to mounted-on-the-arm. and the robot poses used to construct A need to be inverted. > invert different inputs and you can solve for different physical transforms using essentially the same algebra. > and the AX = XB solution does not have to be the final answer. the post suggests using its least-squares result as initialization for nonlinear refinement. samarth-robo.github.io/blog/…
1
5
533
3,000 retail robots are turning physical stores into datasets. > Simbe says it now has more than 3,000 autonomous Tally units under contract. > these robots don’t pick or manipulate anything. they patrol store aisles and continuously scan what is happening on the shelves. > inventory availability prices promotions product locations merchandising conditions > Simbe says Tally has already captured more than 45B shelf images and completed 5M hours of autonomous operation. > the interesting idea is that the robot is basically a moving sensor platform. instead of asking store workers to manually audit shelves... the physical store becomes a continuously updated dataset. > Simbe describes the 3,000+ contracted units as the largest commercially committed shelf-intelligence fleet it has identified. simberobotics.com/about/news…
5
367
A parcel robot has already picked 1.75 million real packages. > BEUMER just launched robotpick, a system built to turn chaotic piles of parcels into a clean stream for sorting. > it uses computer vision + adaptive vacuum gripping to find, pick and separate boxes, bags, polybags and letters. > the important part: this is not just a lab demo. BEUMER says robotpick has already processed more than 1.75M parcels in a live customer environment. > under the right parcel mix and operating conditions, the company says it can automate 99.9% of bulk parcel singulation. > this is exactly the kind of robotics deployment that is easy to overlook: one repetitive physical bottleneck automated at massive scale. beumergroup.com/press/beumer…
1
1
1
243
PolyUMI: record sight + touch + sound with one handheld robot gripper > PolyUMI is an open-source, wireless device for collecting robot manipulation demonstrations. > while a human performs a task, it records 4 synchronized streams: wrist RGB video optical tactile images contact audio motion + proprioception > the interesting part is that touch and sound capture contact events that vision can easily miss. slipping, tapping, scraping, pressure changes. > the same tactile sensing finger can then be moved from the handheld gripper onto the robot. > so the sensing setup during human demonstrations stays much closer to what the robot sees during execution. > this feels like an important direction for manipulation data: not just showing robots what happened... but recording what the interaction actually felt and sounded like. polyumi-vista.github.io/
3
3
32
2,243
O-ID wants humanoid joints to be hot-swappable like server parts. > the Tokyo startup just raised $1.2M to build a fully modular industrial humanoid. > joints, limbs, and compute modules are designed to be replaceable. > so if one joint fails on a factory floor: don’t ship the entire humanoid back for repair. swap the broken module in minutes. > the idea is basically: humanoid hardware → field-serviceable infrastructure > O-ID is now moving its prototype toward production for Japanese factories. > it has also signed an LOI with Sumitomo Electric. > humanoid reliability may end up being less about never breaking... and more about making failures cheap and fast to fix. o-id.net/
2
6
632
Feather Robotics is building a $30K “Android of robotics” instead of another closed humanoid. > Feather sells a modular wheeled humanoid developers can actually build on top of. > starting price: $29,990 > instead of forcing its own foundation model, Feather can run models from: NVIDIA Skild Physical Intelligence and others > developers can swap compute, cameras, hands, and other hardware depending on the application. > TechCrunch reports Feather has already crossed $1M in revenue. > its robots are already being used for things like: cooking in restaurants cleaning science labs > the bet is simple: don’t build the entire robotics stack yourself. build the hardware platform that thousands of Physical AI companies can deploy their intelligence onto. techcrunch.com/2026/09/24/me…
1
8
640
10 NeurIPS 2026 papers in robotics / embodied AI worth watching: 1) Embodied-R1.5 one model for physical reasoning, grounding, planning + robot control arxiv.org/abs/2606.11324 2) LaST-R1 reinforcing the latent physical reasoning inside a VLA arxiv.org/abs/2604.28192 3) LessMimic humanoid interaction using local geometry instead of continuously copying human motion arxiv.org/abs/2602.21723 4) Point4D long-range 4D motion reconstruction through occlusions arxiv.org/abs/2609.09145 5) Grasp-Then-Plan with Failure Attribution figure out whether the grasp failed or the motion policy failed arxiv.org/abs/2606.03385 6) OxyGen shared KV caches for much faster multi-task VLA inference arxiv.org/abs/2603.14371 7) PACE help robots understand how far through a long-horizon task they already are zhuochenn.github.io/ 8) TDBench tests whether VLMs really understand top-down / bird’s-eye scenes arxiv.org/abs/2504.03748 9) Where to Attend / PaPE positional encoding designed around visual geometry instead of 1D language arxiv.org/abs/2602.01418 10) H-GenPO hierarchical generative policy optimization for Physical AI sunghoonim.github.io/ really interesting batch across robot reasoning, humanoids, VLA post-training, 3D perception, long-horizon manipulation and efficient inference.
4
38
246
16,903
6LPA for experienced in this economy ?
1
87
This blog post explains why camera calibration can fail even when the math and optimizer look perfectly fine. > one overlooked source of error is the calibration target itself. a large printed checkerboard can be off by several millimetres. mounting it can introduce warp too. > your optimizer assumes those 3D control points are exactly where you told it they are. if the physical target is wrong, the calibration is solving the wrong problem. > calibration is also a joint nonlinear optimization over: camera parameters + one target pose for every image so initialization matters a lot. x` > Zhang-style initialization works well when distortion is modest. wide-angle lenses can violate those assumptions badly enough that you need a better initialization strategy. > and real datasets contain bad corners. robust losses + explicit outlier rejection belong inside the calibration pipeline, not as an afterthought. good calibration isn’t just “detect corners and call OpenCV.” the physical target, initialization, data geometry, and solver all matter. skiprobotics.com/articles/ca…
2
83
This blog post explains how noisy pairwise camera poses can be turned into one globally consistent set of poses. > suppose you estimate lots of relative transforms: Xᵢⱼ ≈ Xᵢ⁻¹Xⱼ now you need to recover every global pose Xᵢ. > simply chaining transforms along a spanning tree is cheap, but every noisy edge adds error. longer chains → more drift. > group synchronization instead uses all the relative measurements together. > there’s also a gauge freedom: apply the same global transform to every pose and all the relative measurements stay unchanged. so one pose has to be fixed as the reference. > FAST-Sync turns the hard nonlinear initialization problem into a structured linear approximation, solves it efficiently, then projects the result back onto the Lie group. > important: it doesn’t replace nonlinear optimization. it gives Gauss-Newton / Levenberg-Marquardt a much better place to start. nice bridge from relative pose estimation to global SfM and pose-graph optimization. gtsam.org/2026/08/12/fast-sy…
1
4
139
This tutorial explains what modern RANSAC actually looks like beyond the simple version usually taught in class. > the full pipeline is closer to: minimal solver → sample hypothesis → score it → refine promising models → pass the result downstream > the minimal solver matters because it determines what model you can generate from the smallest possible sample. 5 points → essential matrix 4 points → homography 3/4 points → many PnP variants > modern RANSAC then gets much smarter about: which samples to draw how to score hypotheses when to refine them when to stop > the tutorial also covers differentiable + learning-based alternatives and where robust estimation still sits inside larger vision pipelines. > useful follow-up once you already know F, E, homography, and PnP. danini.github.io/ransac-2025…
2
71