PVL Lab @ Princeton | Memory and Perception | Anthropic Fellow | thebestworstcase.substack.co…

New York, USA
Researchers are actively improving memory for LLMs/VLMs. But are we measuring it right? A system can score 100% by rereading the input for each question. If a human did that, we’d call it terrible memory! We test beyond accuracy with ECC: Efficiency, Compression, Calibration. 🧵
4
7
22
982
Researchers are actively improving memory for LLMs/VLMs. But are we measuring it right? A system can score 100% by rereading the input for each question. If a human did that, we’d call it terrible memory! We test beyond accuracy with ECC: Efficiency, Compression, Calibration. 🧵
4
7
22
982
We also compared the memory properties of different model backbones—Transformers, Mamba, recurrent models, TTT, and others—on the same task. Several alternative backbones, including Mamba and TTT, showed better compression–calibration tradeoffs than RoPE Transformers. 8/9
1
4
63
We hope ECCBench provides a general framework for evaluating memory in settings where retrieval accuracy alone is insufficient—from language and video models to robotics and autonomous agents. Check out the paper for the full results! arxiv.org/abs/2609.00103 9/9
4
46
Shmuel Berman retweeted
Dynamic camera intrinsics often appear in robotics and 3D vision—yet training and evaluation data for predicting them from RGB is surprisingly rare. Introducing InFlux++: our new real-world benchmark and synthetic training dataset for dynamic intrinsics prediction. 🧵1/6
1
17
127
10,750
Shmuel Berman retweeted
We have released an initial Infinigen2-Flying-Indoors dataset! It provides 2000 stereo videos, 24 frames each, with ground truth for depth, flow, normals, segmentation, albedo and more. Each group of 4 trajectories provides random synchronized views of the same procedural dynamic 3D scene. Code to replicate it will be released soon as infinigen2.0.0a2. See full youtube preview & huggingface link below ↓
2
8
39
15,738
I am excited to share my work as an @AnthropicAI Fellow! Can language models control robots? ⬇️
8
6
99
8,890
I had a lot of fun working on this project! Thank you to Michael Ilie, Daniel Freeman, my advisor Jia Deng, and everyone at Anthropic who offered advice and support.
1
9
502
Shmuel Berman retweeted
our summer project on AI Alignment princeton-superalignment.git… hope to announce two exciting results by end of ICML :-)
1
4
102
11,298
LLMs make human behavior more predictable. Prediction markets incentivize people to change outcomes they have control over in order to make a profit. I'm worried that these two technologies might have dangerous synergies that could be disastrous for a stable society. 1/3
1
5
294
Right now there is a small set of actions both beneficial and controllable by any individual. Such actions are bad for society (like insurance fraud), but because they are obvious, we can ban them. 2/3
1
1
97
What does a society look like where people are incentivized to change anything in their power? Chaos. Read more here: thebestworstcase.substack.co…
62
Shmuel Berman retweeted
Introducing Goedel-Architect: an open-source framework for formal theorem proving in Lean 4. Using the open-weight DeepSeek-V4-Flash (284B-A13B), it reaches state-of-the-art results, rivaling proprietary systems at a fraction of the cost. It solves 4/6 on IMO 2025, 11/12 on Putnam 2025, and 3/6 on USAMO 2026. On PutnamBench it solves 88.8% (597/672) at just ~$1.65 per problem. Paper: arxiv.org/abs/2606.06468 Project page: goedelarchitect.github.io/
4
20
58
6,961
Shmuel Berman retweeted
1/ Now that we're running out of data, how do you optimally scale multi-epoch pretraining to hundreds of epochs? Our first paper from Q! q0 trains a population of models, instead of single model that saturates fast, reaching a dramatically lower loss at *every* epoch budget. w/ @bishmdl76 @akshayvegesna @ShmuelBerman
16
54
265
31,258