When SceniX joined World Labs, we said spatial intelligence was never only about perceiving and generating virtual and physical worlds, but also interacting with them. Today, we’re sharing early results from that vision: building worlds that train robots. 🌎🤖↓

Jul 28, 2026 · 4:13 PM UTC

40
132
910
335,314
With the help of generative world models, a real-to-sim-to-real (R2S2R) simulation engine turns one physical task into many controllable, reusable worlds, helping robotics teams train policy models and test changes faster, uncover failures earlier, and reduce costly experimentation on hardware.
7
24
142
26,507
The Real-to-Sim part of R2S2R transforms physical robots, sensors, environments, objects, and interactions into simulations that preserve task-relevant observations and dynamics - not only how the world looks, but how it acts when the robot interacts with it. These results set a new bar for sim-real alignment in robot manipulation.
3
19
144
26,225
The Sim-to-Real part of R2S2R uses these aligned worlds for training. The policies were trained entirely in simulation with zero real-world data, transferred directly to diverse robot platforms, and operated autonomously for hours without failure or human intervention. Our engine is policy- and embodiment-agnostic, allowing us to serve customers with different robots, sensors, policy stacks, and deployment needs.
2
4
44
5,184
Sim-to-real aligned simulation also makes evaluation scalable. Here, the same failure behavior and outcome appear in both virtual and physical worlds. The simulation captures more than the task setup or final success label; it reproduces the conditions that push a policy toward success or failure. The result is faster iteration, broader coverage, and substantially lower cost.
4
6
47
11,369
In our taxonomy of world models, we called the simulator the linchpin: the place where agents can act, learn, and be evaluated. Our R2S2R engine can move robot development beyond slow, expensive, hardware-bound iteration toward more scalable and cheaper training and evaluation. worldlabs.ai/blog/real-to-si…
6
6
75
14,298
Sort replies: Relevant Recent Liked
Replying to @drfeifei
RL is all about envs. Real2sim2real is one of the best ways to scale envs for physical RL. Congrats!
1
5
71
9,111
Replying to @drfeifei
飞飞教授,如果编程智能在自我迭代上够快,会不会取代空间智能的数据呢?
1
430
Replying to @drfeifei
🚨 Analysis: “Can a Language Model Learn Facts Continually in Its Weights?”, Baseten’s Latest Paper They sequentially force-fed 247 invented facts into Qwen3 weights, watched most of them become unreachable after a few dozen later writes, and still framed the whole exercise as “continual learning of facts.” What they actually demonstrated is that fine-tuning is cosmetic on a probability landscape whose mass is still dominated by pre-training. [ What the study says ] Bare-statement writes give near-perfect recitation that collapses under any real use. Study-style data does better but still loses most of its reach after twenty sequential updates. “Forgotten” facts keep the majority of their original log-probability lift. Put the same statement back in context and accuracy jumps up to 80%. Two written facts almost never work together. Self-retrieval of anything the model was just trained on sits at 34%. Context barely degrades while the weights tear themselves apart. Full paper:
arxiv.org/abs/2607.11020 [ What the transformer actually does ] It samples from a conditional distribution whose overwhelming mass was fixed in pre-training. A fine-tune is a superficial local edit. It can temporarily raise the probability of a target pattern under a narrow set of queries, but it does not remove the underlying density. The pre-training prior eventually reasserts itself. Every later write is another tear: local overfitting that redirects query geometry, produces interference, and leaves the bulk of the original distribution intact. That is why the “forgotten” fact is still sitting there in log-prob space and why context restores it instantly. Context never tore the landscape; the weight updates did. [ Exhibits ] • Weight writes create question-keyed patches, not addressable facts (a keyhole)
• Later writes do not erase the earlier probability mass; they simply redirect the queries that used to reach it
• 70% of failures under bare-statement training just spit out the newest patch instead
• Context recovers a “lost” study fact up to 80% with zero weight change
• Joint use of two written facts collapses; the model cannot list its own recent training
• Capability damage tracks KL divergence from the original model, every edit costs global coherence All of it is the pre-training distribution reasserting dominance over superficial local bias. Weights are not a reliable knowledge substrate. [ Conclusion ] This is not a partial success of continual learning. This is a controlled demonstration that fine-tuning is mostly BS for durable knowledge. You are not writing facts into the model. You are temporarily scarring the surface of a distribution whose mass still belongs to pre-training. The scars fade or get overwritten; the original landscape wins. Continual learning is the latest industry chimera, built on a category error that projects human traits like memory, knowledge and cognition onto a system whose every output is merely an isolated, local, query-constrained rollout. [ The real problem ] We keep treating weight updates as if they were knowledge installation. They are not. They are high-risk cosmetic interventions that introduce tears, interference, and side-effects while the pre-training prior remains the default. ICL and context work better precisely because they leave that prior untouched. [ What actually works ] Stop pretending fine-tunes install durable facts. Put the information in context when it must remain reachable. Treat every weight edit as a temporary, lossy patch with measurable global damage. Verify the sampled output against the actual distribution, not against the fantasy that you rewrote the model’s knowledge. The knowledge store was never there. The fine-tune did not create one. It just tore the landscape a little more. Apply at your own risk. ☕️ How many more papers are we going to publish that treat superficial weight edits as if they were permanent knowledge updates?
342
Replying to @drfeifei
> Me : Think again. More patiently. > Claude: Flagged . . .
158
Replying to @drfeifei
Hey everyone — I just shared a post on LinkedIn about a simple but effective idea for memory management in real-time AI systems. If you work with AI models, streaming, or anything that involves long-running sessions, this topic is definitely worth checking out. It covers a common issue: AI models slowing down, freezing, or losing consistency after extended conversations.The post explains a lightweight “memory rotation” approach that keeps the system stable without needing bigger servers or additional resources. It’s written in both Arabic and American English so anyone can jump in.Feel free to check it out and join the discussion. linkedin.com/posts/abdulaziz…
96
Replying to @drfeifei
wow
1,277
Replying to @drfeifei
this is genuinely very interesting! congrats on the release!
622
Replying to @drfeifei
i've seen sim gains vanish at contact physics, which metric catches that early
206
Replying to @drfeifei
@theworldlabs always keeping us well fed.
5
537
Replying to @drfeifei
building synthetic worlds to train robots is only as good as how well those worlds lie about the real one.
2
437
Replying to @drfeifei
Simulation can predict whether a policy will succeed. Governed execution determines whether and when it may act in the real world. The deeper opportunity is a loop of simulation, execution outcomes and learning that builds evidence for policy, model trust and bounded autonomy.
1
1
456
Replying to @drfeifei
Literally, that picture looks so perfect. Absolute perfection. Speechless.
137
Replying to @drfeifei
How i join world lab @drfeifei
1
265
Replying to @drfeifei
really curious how far the interaction layer can go once robots start learning inside these generated worlds
60
Replying to @drfeifei
When should we expect them to develop an accurate, internal spatial representation of a chess board? chessbench.ai/timeline
735
Replying to @drfeifei
This is the future we believe in, where spatial intelligence acts in the real world. Excited to see SceniX and World Labs building the foundation for robots!
189
Replying to @drfeifei
not just perceiving worlds or generating them, learning to interact with them, that's the breakthrough 💀
493
Replying to @drfeifei
Exciting progress bringing AI world models closer to robotics
62
Replying to @drfeifei
👍👍💪🏻💪🏻
316
Replying to @drfeifei
建数字世界训练机器人,可现实里连买个充电器都要实名,这平衡点在哪?
31
Replying to @drfeifei
我更在意的不是感知、生成与互动。我关注的是它对物理世界规则的理解与运用。
54
Replying to @drfeifei
Worldbuilding, literally.
17
Replying to @drfeifei
The real-to-sim fidelity here is next-level — not just visual, but dynamics and failure modes too. Policy- and embodiment-agnostic simulation that actually closes the gap this well could become the new default for scaling robot learning. Excited to see how far this goes.
308
Replying to @drfeifei
The block principle is the basis of the brain's cognition of everything in the world. It is as important as living things are made of cells. Block principle: The brain's cognition of things consists of neurons that record sound, words and physical images respectively. When any part of this block is triggered, all neurons in the whole block will be activated. A language composed of sound and words is used to bind physical images. When you hear the pronunciation, you will activate the neurons of both the text and the physical image; When you see words, you will activate neurons of both pronunciation and physical images; When you see a physical image, you will also activate both sound and text neurons. This is the basis of the brain's cognitive composition of all things in nature. Take the object "Sun" as an example. When you hear the pronunciation of the word "Sun", it is bound to the object and words, and the brain perceives the corresponding words and physical images. When you see the words "the sun", although it is not read, its pronunciation and physical image will appear in your mind; Similarly, when you see the sun, you know its pronunciation and words in your mind. This is because when the word "sun" is converted into EEG signal by vision, the brain's EEG passes through the neurons of its pronunciation and still image at the same time, and the pronunciation of the text and its object image appear in consciousness. So no matter which part is triggered, the text, sound and physical image of the whole block will appear in the perception system at the same time. 3D animation is based on the concept of planarization, which helps to better understand the concept of block principle. In fact, the projection of any complete object by the brain requires multi-dimensional and a large number of brain regions to participate, and neurons that record words, sounds and physical images are not piled up in the same area. Just like when we look down at the earth from a panoramic perspective in space, it is a blue plane sphere, but in fact it is multidimensional and colorful. When studying difficult things, we must simplify and conceptualize the whole universe, the most complex organ, in order to deduce the movement mode of EEG between neurons from a macro perspective through thought experiments without any instruments and data. Simplicity is the symbol of genius-Leonardo da Vinci The goal of science is simplification, not complexity-Richard Feynman. The simplest theory is often the most powerful-Stephen Hawking.
157