Come hang out at our #COLM2026 workshop on Oct 9! LSEI asks how LMs can learn by interacting w/ the world and other agents! We’ve got an exciting speaker lineup and plenty to discuss. Bring your questions & hot takes! 🔥 📍 Hilton SF Union Sq, 15&16 🔗 learning-situated-interactio…
4
24
2,535
Check out our paper at COLM 2026. We make (world-model based) controller language-instructable. It needs no manual labels with rollouts + our pos-hoc annotation method. The result is a flexible controller interface that any VLM/human can plug into. zinengtang.github.io/instruc…
5
47
7,879
Zineng Tang retweeted
Excited to announce the first workshop on Learning from Situated and Embodied Interaction @ #COLM2026! 👥🤖🎉 What can interaction with environments, humans, and other agents teach language models that passive text cannot? 🔗 learning-situated-interactio… 📅 Submit by June 30th!
2
7
30
7,868
Zineng Tang retweeted
Join our COLM workshop on learning from situated and embodied interactions! Now accepting papers 📢 Interaction with environments, humans, and other agents can serve as an important learning signal for language models. We welcome paper submissions on anything related to learning from interactions (including embodied, multi-turn, multi-agent contexts), cooperative AI, social learning, pragmatics, and more!
Excited to announce LSEI @ COLM 2026: Workshop on Learning from Situated and Embodied Interaction, happening in San Francisco on Oct 9! We study how interaction can serve as a learning signal for language models. Deadline: June 23, 2026 CFP: learning-situated-interactio… (1/2)
1
1
18
2,632
Zineng Tang retweeted
Excited to announce LSEI @ COLM 2026: Workshop on Learning from Situated and Embodied Interaction, happening in San Francisco on Oct 9! We study how interaction can serve as a learning signal for language models. Deadline: June 23, 2026 CFP: learning-situated-interactio… (1/2)
1
3
16
5,447
Excited to announce LSEI @ COLM 2026: Workshop on Learning from Situated and Embodied Interaction, happening in San Francisco on Oct 9! We study how interaction can serve as a learning signal for language models. Deadline: June 23, 2026 CFP: learning-situated-interactio… (1/2)
1
3
16
5,447
🚨Announcing Zebra-CoT, a large-scale dataset of high quality interleaved image-text reasoning traces 📜. Humans often draw visual aids like diagrams when solving problems, but existing VLMs reason mostly in pure text. 1/n
1
27
129
18,128
Zineng Tang retweeted
CoT transformed text reasoning. What about multimodal? 🤔 Check out our new dataset of interleaved text and image reasoning traces. We also show interesting visual CoT examples generated inherently by the model finetuned on our dataset!
🚨Announcing Zebra-CoT, a large-scale dataset of high quality interleaved image-text reasoning traces 📜. Humans often draw visual aids like diagrams when solving problems, but existing VLMs reason mostly in pure text. 1/n
2
11
1,828
Zineng Tang retweeted
Excited to share our new work! DOVE 🕊️: a dynamic vision encoder that adapts token count to image complexity. Fewer tokens, same fidelity—outperforming fixed-length AEs tokenizer on classification & VLM tasks! Arxiv: arxiv.org/abs/2506.03643 Web: dove-encoder.github.io/dove-… #AI #CV
1
23
109
11,361
Excited to share our new work! DOVE 🕊️: a dynamic vision encoder that adapts token count to image complexity. Fewer tokens, same fidelity—outperforming fixed-length AEs tokenizer on classification & VLM tasks! Arxiv: arxiv.org/abs/2506.03643 Web: dove-encoder.github.io/dove-… #AI #CV
1
23
109
11,361
🔥 DOVE uses 68 % fewer tokens but has better FID than VQGAN/TiTok and +10–12 pts on VQA/ImageNet/CIFAR. It achieves significantly stronger performance on classification, probing, and VLM tasks. DOVE also brings emerging properties—PCA heatmaps reveal sharper segmentation.
1
3
464
Big thanks to my undergrad intern Lingjun for delivering such impressive work, and to Rudy for the thoughtful co-advising!
3
293
We are thrilled to announce TULIP! 🌷 tulip-berkeley.github.io/ A state of the vision language encoders coupled with generative model for stronger representation learning.
7
66
295
31,315
TULIP achieves state-of-the-art performance across multiple vision and vision-language benchmarks. It significantly improves zero-shot classification on ImageNet-1K, enhances fine-grained object recognition, and boosts multimodal reasoning scores. Compared to existing methods, TULIP shows up to a 3× improvement on MMVP and a 2× boost in fine-tuned vision tasks.
1
1
5
985
Also thanks to, @LongTonyLian, Seun Eisape (seuneisape.github.io), @XDWang101, @roeiherzig, @Yalatweets, @alsuhr, and @trevordarrell for their great efforts!
1
5
682
Zineng Tang retweeted
Automating AI research is exciting! But can LLMs actually produce novel, expert-level research ideas? After a year-long study, we obtained the first statistically significant conclusion: LLM-generated ideas are more novel than ideas written by expert human researchers.
94
744
3,577
1,124,614
Zineng Tang retweeted
We announced Phi 3.5 series today! 1️⃣ Multilingual Mini 3.8B: huggingface.co/microsoft/Phi… 2️⃣ MoE 16x3.8B (active 6.6B): huggingface.co/microsoft/Phi… 3️⃣ multi-frame vision LLM: huggingface.co/microsoft/Phi…
We released phi 3.5: mini+MoE+vision A better mini model with multilingual support: huggingface.co/microsoft/Phi… A new MoE model:huggingface.co/microsoft/Phi… A new vision model supporting multiple images: huggingface.co/microsoft/Phi…
There's a new version of this post
5
41
4,182
CoDi-2 is selected as #CVPR2024 Highlight. Come joint us in today’s poster session Arch 4A-E #314 in 5pm to 6:30 pm! codi-2.github.io @yzy_ai @nlpyang @ChenguangZhu2 @mohitban47
🔥Excited to introduce CoDi-2! It follows complex multimodal-interleaved in-context instructions to generate any modalities (text, vision, audio) in zero/few-shot interactive way! codi-2.github.io huggingface.co/papers/2311.1… @yzy_ai @nlpyang @ChenguangZhu2 @mohitban47 🧵👇
10
18
5,698