accelerating medicine at @EdisonSci | PhD with @BintuLab and @WJGreenleaf at Stanford | views/opinions my own and do not reflect those of my employer

San Francisco, CA
Michaela Hinks retweeted
Excited to share Minerva, our approach using genome language models for biological discovery! Using Minerva, we find that UG27 reverse transcriptase systems encode variable arrays of diverse ncRNAs with a shared structure, each templating a short DNA hairpin. With @garykbrixi.
20
114
473
118,790
Today, Pioneer Labs is announcing our first step towards terraforming Mars. 🚀🌼 With equipment that fits in just a single rocket launch, we can convert Martian dirt, water, and air into enough building materials to construct a small city on Mars. To do it, we made the first microbe for Mars. We found the best microbe on Earth and used evolution to teach it how to source all of its nutrients directly from Martian materials. The first astronauts will be greeted with safe shelter already filled with water, oxygen, and rocket fuel for the return journey. This is the first step toward using biology to make Mars a friendly place for life. It lets us live off the land and helps us build the next great frontier. It's the first of five organisms we need to green Mars ⬇️
554
1,437
11,373
2,481,302
Michaela Hinks retweeted
The community response to the Millennium Problems for Biology has been great. A few updates: There is clearly a deep hunger for a more comprehensive list of big unsolved problems in biology and medicine that can serve as a beacon for the field, and we have seen some very good suggestions. Over the next few weeks, FutureHouse will run community-led process to assemble a longer list of grand challenges in biology and will publish this expanded set as additional to the original set, with attribution. In addition, the Millennium Problem list and specifications also needs to be updated in response to a few issues that people have pointed out. That will happen in the coming days. Finally, FutureHouse will convene a panel to adjudicate success, and will also announce cash prizes associated with the problems. We will have more to share soon. Thank you to everyone who has participated in the conversation! x.lingyaoai.com/SGRodriques/status/210…
In light of the progress in mathematics, we at Edison Scientific and FutureHouse have assembled a set of Millennium Problems for Biology. They are chosen to be very hard to solve but very easy to validate in a simple laboratory environment. Any of these, if solved, would mark a major advance in biotechnology, and most of them would contribute materially towards curing disease. These are, in some sense, the “last reasonable eval” for AI in biology. This was work primarily by @MichaelaThinks and myself, with contributions from many others. Short descriptions below. The full descriptions of the problems with acceptance criteria are at the Bio Millennium Problems website, linked in the next post. Share more if you have ideas. If they meet our criteria, we’ll add them to our list (with attribution and permission).
8
14
153
15,734
When I read that an internal OpenAI model solved the Navier-Stokes problem, I realized it would be very hard for us to know if that model possesses similar "millennium intelligence" in biology (even if it does!). "Curing disease" is a noble goal, but we cannot validate proposed treatments rapidly enough to use it as a frontline eval for new models. Here, @SGRodriques and I propose Millennium Problems for Biology - meaningful, difficult challenges in biology whose solutions should be straightforward and quick to validate in a wetlab. If a model were to solve any of these, I would know that machine intelligence was capable of fundamentally reshaping humanity's ability to make, measure, and model biology. Please see details of the problems in the thread below, and see our website for the permanent home of the list. I encourage the community to red-team these challenges to make sure they are adequately specified and difficult to cheat, and propose additional challenges to add. We will subsequently update the list in response to feedback.
In light of the progress in mathematics, we at Edison Scientific and FutureHouse have assembled a set of Millennium Problems for Biology. They are chosen to be very hard to solve but very easy to validate in a simple laboratory environment. Any of these, if solved, would mark a major advance in biotechnology, and most of them would contribute materially towards curing disease. These are, in some sense, the “last reasonable eval” for AI in biology. This was work primarily by @MichaelaThinks and myself, with contributions from many others. Short descriptions below. The full descriptions of the problems with acceptance criteria are at the Bio Millennium Problems website, linked in the next post. Share more if you have ideas. If they meet our criteria, we’ll add them to our list (with attribution and permission).
13
37
364
31,505
Michaela Hinks retweeted
I agree with @DavidRBellamy that people are totally miscalibrated on the risk of AI designing dangerous viruses. We should not be talking about AI-engineered bioweapons like it is the literal end of the world. The upside from medicine is going to be so much larger than the downside from engineered bioweapons. In addition David's points, there is actually an even more fundamental point here in our favor, which is evolution. As soon as you release an engineered virus into the wild, the virus is no longer under your control: it will evolve however it wants. And, what viruses want is to maximize their ability to replicate. The thing that maximizes their ability to replicate is to infect as many people as possible, which means being extremely contagious and not killing their hosts (dead hosts don't spread virus). The viruses that are most evolved for human biology are the common cold viruses: extremely contagious and not at all lethal. Viruses that kill humans do so by accident, usually because they are new to human biology (e.g. COVID, when it first jumped, or flu when it jumps from birds). If you stick a bunch of machinery into the virus to kill humans, I guarantee you that machinery will disappear from the virus very quickly. In response to this, I hear people say things like "the AI could engineer a kill switch so that the virus doesn't kill the humans initially but then once it has spread through the entire population the AI will hit the switch and kill all the humans." No. What would actually happen in practice is that the kill switch would accumulate deleterious mutations because there would be no evolutionary pressure to preserve it, and would quickly become non-functional. Evolution is fundamental. There is no way around it. For AI readers, trying to engineer a virus by setting its initial genetic code and then releasing it into the wild is like trying to train a model by setting the initial weights and then hill-climbing on a hidden training mixture you have no control over. And you're not even allowed to run any experiments in advance! You may be able to influence the behavior of the model in the first few iterations, but you will quickly lose control. There are big dangers. AI will be great at making one-off, non-replicating biological weapons, which are also scary (but much less scary than replicating bioweapons). AI will be great at helping people to weaponize existing pathogens, which is also a major danger. Also, a misanthropic model could get creative: it could release viruses repeatedly to counteract the effects of evolution, for example. Many people could die this way. Sensible surveillance is important, as is having proportionate controls on wet lab equipment. But we don't live in the dark ages anymore, we're not going to have a smallpox or black death-style pandemic where 50% of people die. COVID killed ~0.1% of humanity. A pandemic 100x worse than COVID would be horrific, but it would also not be civilization-ending. I get the sense that many AI researchers (with the notable exception of Dario personally) actually expect that AI in biology will do more harm than it will do good. This could not be further from the truth. As someone who likewise falls into the very small group of people who have actually physically made viruses with their own hands and has also worked on frontier AI, the upside here is huge, and the downside is not anything like what is being portrayed.
12
25
156
71,115
New Online! A practical guide to studying genome function using single-molecule genomics dlvr.it/TVT9gy
3
34
134
7,339
Michaela Hinks retweeted
In the Life Sciences and Curing Diseases... well, I'm biased on this one. We think AI can transform science and health. We're going directly after the highest burden diseases plaguing humanity, as well as supporting the complement to AI progress - high quality scientific data:
1
11
71
32,166
much better to get in the driver's seat than complain on twitter
Arguing over disease cure timelines is exhausting IMO. It doesn’t matter if you think it’s pure grandstanding or somewhat legitimate for certain disease clusters. The fact is that multiple of the largest companies on this planet, who as a group raise an order of magnitude more capital per year than the entirety of biotech, are placing healthcare atop their list of priorities. This quantum of capital flowing, even inefficiently, towards healthcare is a generationally important boon. Don’t root for it to fail.
1
1
18
2,788
Michaela Hinks retweeted
Replying to @wc_ratcliff
I think there's a bigger difference. Math and especially physics are taught in a way that doesn't really describe how a framework is devised; rather, just "shut up and compute" (quite literally). At the frontier of biology, you have to actually develop the conceptual framework yourself. That is not easy, and we generally don't get much training for it.
2
4
27
3,019
Michaela Hinks retweeted
I can not emphasize enough how much love and care has gone into this resource. Each of the 3865 models has been individually optimized and thoroughly vetted. It’s a truly foundational resource, and holds many many stories that are begging to be told.
Very excited to announce ENCODE GRAMMAR (Genomic Regulatory Atlas of sequence Models, Motifs, Annotations & Rules): 3,865 experiment-specific deep learning model sets and sequence annotations for decoding human regulatory DNA. 1/
1
24
131
13,817
Michaela Hinks retweeted
Very excited to announce ENCODE GRAMMAR (Genomic Regulatory Atlas of sequence Models, Motifs, Annotations & Rules): 3,865 experiment-specific deep learning model sets and sequence annotations for decoding human regulatory DNA. 1/
12
201
735
60,927
Michaela Hinks retweeted
Today, we're announcing Edison Advances, a home for our research announcements, engineering blogs, open-weights models and benchmarks. At Edison, we’re building Kosmos, an AI Scientist, to solve foundational research problems and make truly novel discoveries across scientific domains. Alongside that work, we’re developing better ways to measure and evaluate AI scientists, push their ability to work on long-horizon tasks, and identify gaps in their reasoning. Edison Advances brings our existing work together in one place and we will use it to publish new work as our research progresses. Follow along as the work develops, and get in touch if you’re interested in contributing.
14
39
328
14,919
Michaela Hinks retweeted
We've released a new open-weights vLLM for reading chemical structure patents. This is a hard problem. These "Markush" structures are like graph regex: patterns that match many molecules. We used autoresearch and I expect many new models like this to emerge soon in science 1/4
9
41
209
18,407
Michaela Hinks retweeted
Here are a few of the beliefs that guide our culture at Edison: 1. AI will unlock a new era of scientific progress. 2. Developing medicines is a serious business that requires deep humility and skepticism. It is not enough to simply assert that AI will accelerate medicine; we have to prove it. 3. Trust is critical, and we can’t build trust with our customers if we compete with them. 4. Generic concerns about safety do not justify broad restrictions on science. When there are specific, credible threats, we should mitigate those threats transparently and proportionately. Otherwise, we should trust our customers to use their judgment. 5. Open-weights models are good for science, since they improve reproducibility and transparency. 6. There are real human lives on the line here. Every day counts.
5
16
142
10,427
Michaela Hinks retweeted
I have spent my entire life working on this and thinking about this for the past 4 years. I don't know what will happen in 20 years, but I can promise you that on the 5-10 year timescale, scientists are not out of their jobs. AI is going to massively accelerate the pace of science, increase productivity, let individual scientists make way more discoveries way faster, and is going to make science overall more fun. But the model is going to be collaboration between humans and AI, not replacement. The key difference here between science and e.g. software engineering is that science is not verifiable in any rapid/convenient way (unlike software), unlike programming. We still need humans for their scientific taste.
Today we all lost our jobs..... Three Nature papers showing that scientists in the conventional sense are obsolete At least read the first one.... the AI replaced all things that the scientist does .... nature.com/articles/s41586-0…
40
170
1,049
163,895
Michaela Hinks retweeted
I’m starting to see a lot of posts where people are giving LLMs their genomes. Directly—yeah, absolutely! Right now—worrisome! The clinical interpretation of genomes is always done in context with your phenotype. Which LLMs don’t really have. You need to understand how all of these factors contribute. Doing it in isolation misses most of the picture. Privacy? Hmm. Not sure I’d feel comfortable with this yet. LLMs are removed from the data generation. They’re not calibrated to the protocol used to sequence (or genotype) your DNA. Arrays have non-trivial false positives, which is what the vast majority of people have who have their ancestry data. This risks over and under-treatment. I’m also worried about people feeling they’ve got a protective phenotype and adopting some more cavalier risk-posture. I’m unsure how well LLMs currently fulfill the role of a genetic counselor. Are they going to recontact you when the variant gets its status altered in ClinVar. No, not yet. Can it explain penetrance and prevalence and absolute/relative risk? Okay—so yeah consumer-facing LLMs will be the main surface area between humans and their health in the future. Yes. Definitely. They’re really good right now at taking your genetic data and helping you craft questions and understand what’s going on. But making health decisions based on them + your DNA in isolation. I would be very cautious at the moment.
9
5
55
10,721