Computational Biology | Machine learning research scientist at @GoogleDeepMind

Pinned Tweet
Prioritizing and interpreting disease-associated genetic variants remains one of the greatest challenges in human genetics. Today, we’re thrilled to introduce AlphaGenome Atlas 🧬, a genome-wide platform providing precomputed predictions for the regulatory effects of all ~9 billion possible single-letter changes and >100M observed indels in the human genome. Here is what Atlas delivers: 1. Variant Prioritization via AVI To prioritize variants, we developed the AlphaGenome Variant Impact (AVI) score. AVI predicts a unified score per variant, where higher values indicate greater disruption. It achieves state-of-the-art performance across diverse benchmarks. As a proof-of-concept with our collaborators at @broadinstitute, AVI prioritized a deep-intronic variant in DNM1, helping solve a previously unexplained rare epileptic encephalopathy case by revealing a brain-specific cryptic splice site.   2. Multi-Layer Molecular Interpretation Variant prioritization is only half the battle; researchers also need to understand why a variant matters. Atlas decomposes variant effects across multiple interpretable layers: •Feature Attributions which decompose each variant’s score into specific biological modalities driving the impact.  •Cell-Type Specificity: Precomputed predictions with AlphaGenome across hundreds of biosamples reveal the exact cellular context in which a variant acts.  •Regulatory Grammar: Over 2,600 de novo DNA motifs (and >250B genome-wide instances) show when variants directly disrupt critical regulatory binding "words".   Explore the resource: 🌐 Interactive browser & precomputed data: alphagenome.google/atlas - 🎥 Video: piped.video/U0aToL5C-bQ - 📖 Blog: goo.gle/4heuCvn - 📄 Preprint: storage.googleapis.com/deepm…
8
29
103
12,034
Jun Cheng retweeted
Very happy to announce that our team @GoogleDeepmind has pushed the boundaries of generative biology, achieving the successful synthesis of AI-designed proteins that are both functional and watermarked. This proof-of-concept watermarking of the building blocks of life is enabled by SynthID Bio, our new protein watermarking method. It is designed to safeguard the new era of AI-powered generative biology and strengthen global biosecurity. You can read my thoughts here on why watermarking AI-designed proteins is an important research breakthrough: x.lingyaoai.com/pushmeet/status/210531…
89
292
2,287
660,329
Prioritizing and interpreting disease-associated genetic variants remains one of the greatest challenges in human genetics. Today, we’re thrilled to introduce AlphaGenome Atlas 🧬, a genome-wide platform providing precomputed predictions for the regulatory effects of all ~9 billion possible single-letter changes and >100M observed indels in the human genome. Here is what Atlas delivers: 1. Variant Prioritization via AVI To prioritize variants, we developed the AlphaGenome Variant Impact (AVI) score. AVI predicts a unified score per variant, where higher values indicate greater disruption. It achieves state-of-the-art performance across diverse benchmarks. As a proof-of-concept with our collaborators at @broadinstitute, AVI prioritized a deep-intronic variant in DNM1, helping solve a previously unexplained rare epileptic encephalopathy case by revealing a brain-specific cryptic splice site.   2. Multi-Layer Molecular Interpretation Variant prioritization is only half the battle; researchers also need to understand why a variant matters. Atlas decomposes variant effects across multiple interpretable layers: •Feature Attributions which decompose each variant’s score into specific biological modalities driving the impact.  •Cell-Type Specificity: Precomputed predictions with AlphaGenome across hundreds of biosamples reveal the exact cellular context in which a variant acts.  •Regulatory Grammar: Over 2,600 de novo DNA motifs (and >250B genome-wide instances) show when variants directly disrupt critical regulatory binding "words".   Explore the resource: 🌐 Interactive browser & precomputed data: alphagenome.google/atlas - 🎥 Video: piped.video/U0aToL5C-bQ - 📖 Blog: goo.gle/4heuCvn - 📄 Preprint: storage.googleapis.com/deepm…
8
29
103
12,034
To make AlphaGenome Atlas even easier to use in computational biology workflows, we also built an AlphaGenome Agent Skill 🤖 Instead of writing custom query scripts or parsing raw tables, you can interact with AlphaGenome Atlas in natural language directly with Google Antigravity! 🤖👇 piped.video/b2qw3rDNX0Q
1
1
6
507
A huge congratulations to the entire team and all our incredible collaborators! Sincere thanks as well to the open-source genomics community whose foundational tools helped make this work possible. We can’t wait to see the discoveries the community makes with AlphaGenome Atlas 🧬
6
205
Jun Cheng retweeted
🚀Please help share widely! We're excited to launch the new #AI4BIO #Fellows Program @SCSatCMU to recruit exceptional early-career scientists pursuing bold, independent research at the intersection of AI and biology. Fellows will be supported by The Center for AI-Driven Biomedical Research (#AI4BIO) @SCSatCMU and co-mentored by two CMU School of Computer Science faculty members, with opportunities spanning AI models, computational biology, and autonomous science, including engagement with the CMU AI Science Foundry (ai-science-foundry.cmu.edu/). We are looking for truly exceptional candidates who want to help define new directions for AI-driven biomedical discovery. 📅 Apply by November 15, 2026. 🧬 Program information: cmu.edu/ai4bio/apply/index.h… 🤖 Application via Interfolio: apply.interfolio.com/190197
6
66
227
18,530
Jun Cheng retweeted
We’re hiring in my group at Calico. We build Borzoi and its successors—deep learning models that predict how every nucleotide shapes cell-type-specific gene regulation—and apply them to interpret human genetic variation.
1
32
163
22,228
Great to see this, congrats!
Scaling laws are powering AI. It’s time to scale biology. Today we’re launching the Virtual Biology Initiative to generate the data to unlock scaling laws in biology and build accurate predictive models of the cell. Digital representations of proteins are already expanding our understanding of life at the molecular level, and accelerating the design of molecules and medicines. Accurate digital representations of the cell could reveal the mechanisms that are responsible for disease, and show how to reverse them. The protein data bank, and worldwide repositories of protein sequence biodiversity were created through decades of work by the scientific community. The advances in artificial intelligence for proteins would not have been possible without them. The cell is orders of magnitude more complex, and we will need to create the data in just a few years rather than decades. This will require a coordinated global effort. We're partnering with Broad, Wellcome Sanger, Arc, Allen, Human Cell Atlas, Human Protein Atlas, NVIDIA, and Renaissance Philanthropy. Biohub is contributing to this effort as both a funder and a builder. We are developing microscopy to observe millions of cells in living organisms, and cryo-ET to resolve the cell in atomic detail. We're building instruments that expand the range of modalities and parameters that can be simultaneously measured. We’re developing molecular, cellular, and tissue engineering to create models of disease and design interventions. The data we generate will be available to the worldwide scientific community. We’re also committing $100M over the next five years to support work beyond Biohub. We invite other scientific teams and funders to join. Link: biohub.org/news/virtual-biol…
1
661
Join us at #AIxBIO 2026🧬8-10 June, Wellcome Genome Campus UK! We have amazing lineup of speakers and scientific committee. From gene regulation to protein design to drug discovery, abstract deadline 27 April.
5
430
Jun Cheng retweeted
Excited to co-host #AIxBio 2026 (8–11 June) at @wellcomegenome / @sangerinstitute in Cambridge, UK! Join an amazing speaker lineup alongside co-hosts , @ferruz_noelia, Jun Cheng, Jussi Taipale @BenLehner, @deboramarks, The conference will explore five themes: 🔬 Solving the gene regulatory code 🧪 Solving proteins — data, structure and design 💊 Solving chemistry and therapeutics 🫀 Solving cells, tissues and organs ⚙️ AI methods development Abstract submissions are open — deadline 27 April 2026. Selected abstracts will be featured in the main programme alongside invited talks. Register and submit your abstract here: shorturl.at/ZlvzP #AIxBio #AI #Biology #MachineLearning #Genomics
2
16
44
14,209
Jun Cheng retweeted
@Nature's @ewencallaway writes about how hackathons using AlphaGenome are solving rare disease mysteries. "The event helped diagnose someone with a rare condition called Rothmund–Thomson syndrome... [which] affects eyes, skin and other organ systems." Discoveries like this are why we @GoogleDeepMind believe advancing science is one of the most meaningful motivations for developing AI. Read more: nature.com/articles/d41586-0…
14
97
3,903
AlphaGenome paper and models are out today! nature.com/articles/s41586-0…. We have seen great community engagement since we released the API. We hope to open more use cases with the model weights! github.com/google-deepmind/a… We updated the supplementary note with some very interesting case studies. I encourage everyone using AlphaGenome for non-coding variant interpretation to read it. It was a huge team effort, I’m very proud of all of our co-authors @googledeepmind!
4
68
303
20,063
This is a great idea, congratulations!
Modern GWAS can identify 1000s of significant hits but it can be hard to turn this into biological insight. I'm excited to share our new work combining genetic associations and Perturb-seq to build interpretable causal graphs, out today in @Nature:
7
1,608
Excited to share #AlphaGenome, a start of our AlphaGenome named journey to decipher the regulatory genome! The model matches or exceeds top-performing external models on 24 out of 26 variant evaluations, across a wide range of biological modalities.1/6
14
207
903
87,562
One of the most exciting parts of our #AlphaGenome work is the ability to directly predict splice junctions from sequence and also use it for variant effect prediction. This is enabled by modeling the competition between splice sites and junction supporting reads.4/6
2
3
12
3,388
AlphaGenome model architecture. follows a U-Net-like structure with an Encoder, a central Transformer Tower, and a Decoder. The architecture has inspiration from Borzoi and AlphaFold.
3
674
Jun Cheng retweeted
Excited to launch our AlphaGenome API goo.gle/3ZPUeFX along with the preprint goo.gle/45AkUyc describing and evaluating our latest DNA sequence model powering the API. Looking forward to seeing how scientists use it! @GoogleDeepMind
5
16
76
4,318
Jun Cheng retweeted
Happy to introduce AlphaGenome, @GoogleDeepMind's new AI model for genomics. AlphaGenome offers a comprehensive view of the human non-coding genome by predicting the impact of DNA variations. It will deepen our understanding of disease biology and open new avenues of research.
19
208
1,228
138,284
Having worked on RNA splicing models for years, this is a significant milestone for me. It is a real privilege to work with an extremely talented team. With the new API, I’m excited to see what researchers can find with #AlphaGenome.6/6
1
1
13
2,347