Machine learning lab at Columbia University. Probabilistic modeling and approximate inference, embeddings, Bayesian deep learning, and recommendation systems.

New York, NY
Blei Lab retweeted
Proud to be part of this adventure with such a fantastic team—@kchonyc, @keunwoochoi, @HenriDwyer, and @elmanmansimov. So excited for what’s ahead. Come join us!
We’re excited to announce Ortet, a frontier AI lab for health. A fundamental gap in care today is that the industry and its scientific tools focus on discovering what works at the population level, while it is individual patients we must focus on to deliver the best care. Individualized care requires a broad view of the patient, the care they receive, and the system around them. Yet the information needed for such a view remains fragmented. At Ortet, our goal is to create continuously improving intelligence for every patient: the Ortet Health Model, bringing together biological, clinical, and operational dimensions to better predict what happens next and which actions are most likely to lead to better outcomes. To deliver it, we are creating frontier AI from the ground up—across compute, data infrastructure, and models—and building and operating at every layer of AI. We are scientists and engineers behind foundational AI advances, including the attention mechanism, GRU, and neural text-to-image generation. Across the team, our work spans frontier AI research, health, life sciences, and large-scale data and compute infrastructure—from building AI for lab-in-the-loop molecular design and co-founding Prescient Design to establishing Genentech’s frontier research team and building the largest GPU cluster in life sciences. We imagine a future where everyone has access to affordable, precise, and better care. If you want to help us build that future, join us: ortet.ai
7
4
50
4,968
We have a new paper on topic models! We introduce Mechanistic Topic Models (MTMs). MTMs model SAE feature counts rather than words or plain embeddings, combining the benefits of LLM-based and probabilistic topic models. Paper (published in TACL) : arxiv.org/abs/2507.23220
6
14
125
7,027
We show that MTMs are better than traditional and neural baselines at finding salient and abstract themes, while costing less than LLM-based methods. Moreover, working with SAE features allows you to construct topic steering vectors to steer LLM generation toward MTM topics.
1
4
513
1/ Two new preprints on OOD generalization. Shared lesson: the goal should not always be to find what is invariant. For good prediction in new environments, it can be better to model environment variation explicitly, then marginalize it out. arXiv:2604.26128 arXiv:2606.05365
1
6
18
1,057
Announcing our next talk with NYU Professor @_romain_lopez_: "Learning from Millions of Cells with Deep Generative Models" Registration: eventbrite.com/e/ml-nyc-spea…
1
3
4
614
Hypothesis testing 🤝 mechanistic interpretability
The circuit hypothesis proposes that LLM capabilities emerge from small subnetworks within the model. But how can we actually test this? 🤔
2
13
2,344
I am on the academic job market this year! My research advances probabilistic machine learning and its application to biology. I'm looking for faculty positions in stats/CS/applied math and in bio departments. Academic website: eweinstein.github.io/
2
15
79
19,836
Check out our new paper using hierarchical causal models and representation learning to estimate the effects of immune proteins
I'm thrilled to present a new paper with @blei_lab and @lizbwood about estimating the causal effects of T cell receptors in humans. arxiv.org/abs/2410.14127
2
18
2,587
New paper from @blei_lab member @EliWeinstein6 on manufacturing samples from generative protein models. The idea is to approximate a complex distribution with simpler ones that are easy to sample from in the real world, using stochastic chemical reactions.
I’m thrilled to announce: at @jura_bio we’ve constructed a new kind of generative protein model, which allows us to manufacture its generated designs in the real world at petascale. Paper: static1.squarespace.com/stat… Blog: jura.bio/blog/variationalsyn…
3
11
2,864
Excited to announce that ML-NYC is back this semester. Our first speaker is Bin Yu from UC Berkeley on Monday Sept 23rd 4pm. Register here: eventbrite.com/e/ml-nyc-spea…
5
6
1,655
We're excited to welcome Daniel Lee as our next speaker on Wednesday September 20th at 4pm! The event will be followed by a happy hour hosted by the Flatiron Institute Register here: eventbrite.com/e/ml-nyc-spea…
1
1
1,827