I’m excited to share that I am co-founding World Mechanics, a frontier lab working on interpretable foundation models for the physical world, together with
@soniajoseph_.
A big reason I’m excited about physical AI is that it gives us a rare opportunity to rethink interpretability from the ground up. For LLMs, much of interpretability is necessarily post hoc, where we take a model that has already been trained and try to reverse-engineer its internal representations, often without clear ground truth for what it should have learned.
For physical AI, we’re still early enough to make interpretability a first-order objective of training itself. Physical systems also give us unusually useful data-generating processes for interpretability. We can intervene on systems in simulation or on real machines, observe how they respond, and make use of known physical structure. This gives us a path toward representations that faithfully abstract the underlying system and remain controllable under intervention.
At the same time, the interpretability problem itself changes. Video and sensor data are not directly legible in the way language is, and models that reason about the physical world learn very different representations from those we see in LLMs. Many interpretability methods developed for language models already do not transfer directly, leaving a lot of fundamental work to do.
This matters because safe and generalizable behavior of physical AI models depends on them learning faithful abstractions of the systems they act on, including the right semantics and causal structure. Current black-box evaluations give us only indirect evidence that these abstractions are correct, since they cannot cover every scenario or explain why a model fails in a particular one. Interpretability-based white-box evaluations, on the other hand, would allow us to understand the model’s representations directly, explain failure modes, and use those insights to improve the reliability and safety of physical AI models.
Because the field is still early, I also care a lot about how the research community around it develops. We want to do as much of our work in the open as possible and collaborate closely with researchers in academia and elsewhere.
If you’re working on related problems, or this sounds like something you’d want to help build, reach out!