Code & Cure
Decoding health in the age of AI
Hosted by an AI researcher and a medical doctor, this podcast unpacks how artificial intelligence and emerging technologies are transforming how we understand, measure, and care for our bodies and minds.
Each episode unpacks a real-world topic to ask not just what’s new, but what’s true—and what’s at stake as healthcare becomes increasingly data-driven.
If you're curious about how health tech really works—and what it means for your body, your choices, and your future—this podcast is for you.
We’re here to explore ideas—not to diagnose or treat. This podcast doesn’t provide medical advice.
Code & Cure
#57 - If We Can Invent New Viruses Should We
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
AI can now “autocomplete” DNA, and researchers are using that ability to design real, working viruses in the lab. We dig into a Science paper that trains genome language models to generate brand-new bacteriophage genomes, then validates them by building the phages and testing whether they actually infect E. coli. If you’ve been curious about AI in genomics, synthetic biology, or what comes after today’s large language models, this is a concrete example of design, not just prediction.
We walk through the key ideas without assuming you’re a biologist: what a bacteriophage is, why ΦX174 is a useful starting point, and how models like Evo 1 and Evo 2 learn the “grammar” of genomes from massive pretraining. Then we get specific about the engineering: supervised fine-tuning to a narrow phage family, prompting with a short nucleotide prefix, and the practical filters that keep generated sequences from turning into biological nonsense. We also talk about host targeting, novelty constraints, and why diversity matters when you’re trying to outmaneuver bacterial defenses.
The clinical angle is impossible to ignore. Antibiotic resistance keeps rising, and phage therapy could become a more precise way to kill dangerous bacteria, especially when standard drugs fail. But we end where everyone’s mind goes sooner or later: if we can generate novel viruses quickly, what prevents misuse, accidents, or designs we don’t fully understand yet? Subscribe for more clear-eyed conversations about AI and medicine, and if this raised your blood pressure or your hope, share the episode and leave a review with your take on where the guardrails should be.
References:
Generative design of bacteriophages with genome language models
King et al.
Science (2026)
Credits:
Theme music: Nowhere Land, Kevin MacLeod (incompetech.com)
Licensed under Creative Commons: By Attribution 4.0
https://creativecommons.org/licenses/by/4.0/
DNA Letters And The Big Idea
SPEAKER_02G A G T T T T A T C A T C G T A C C Hello and welcome back to Coding Cure, the podcast where we discuss decoding health in the age of AI.
SPEAKER_00My name is Vasant Sarati. I'm an AI researcher and a cognitive scientist, and I'm here today with Laura Hagopian, an emergency medicine physician.
SPEAKER_02And uh we were just using some abbreviations for nucleotide bases in DNA. I don't know if those actually make anything when you put them in a row.
SPEAKER_00I'm sure they're not. I just made a blind.
SPEAKER_02Did we encode did we encode life? It is unclear.
SPEAKER_00We would have to do that several million more times.
SPEAKER_02Yeah, maybe we need like an AI system to help us generate what an actual genome would look like.
SPEAKER_00Wait, what an idea that is! Coincidentally, that is the paper we're gonna be talking about today. It's a great paper, it's super interesting. It's in science, and we'll we'll we'll link it in the show notes. But this is I'm I'm pretty excited about this paper for many reasons, and I can't wait to talk about it.
SPEAKER_02Yeah, and um neither of us are biologists, however.
SPEAKER_00Yes, it's a big that's a big note caveat asterisk, whatever you want to call it. Yes.
SPEAKER_02But I think there's
What Bacteriophages Do To Bacteria
SPEAKER_02a lot that we can learn here, um, both from a medical standpoint and from an AI standpoint, about where where can we go with this and how does this work and why did they even do it? So um, yeah, I think let's let's dive right in.
SPEAKER_00Right. It's it's just interesting, right? So it's all about this notion of what's called a bacteriophage, which uh is is an interesting word, and phages are things that um occcupy another cell, right? They're things that take over another cell. So a virus is an example of this. So a bacteriophage is a type of thing, type of virus, type of virus that infects that infects a b bacteria.
SPEAKER_02Yeah, exactly. And um and and so they decided to see, hey, like, can we use AI to make new bacteriophages, new viruses that infect bacteria?
SPEAKER_00Why would they want to do that?
unknownRight.
SPEAKER_02I mean, I feel like there are tons of applications of this. There could be applications in biology itself, but also in therapeutics. Um one of the things I think about a lot is uh antimicrobial resistance. So I don't know if you've ever had an experience where you get treated for something with an antibiotic and like it doesn't work. Right. And then you have to go up to the bigger gun. And there are some cases where there's no antibiotics that work, but there's a lot of cases where there are bacteria that are resistant to antibiotics. And so one concept, at least from the medical angle, is hey, these new phages could help overcome any sort of bacterial resistance and provide an opportunity for new treatments.
SPEAKER_00Right, right. And I think also just in general, I think the idea of being able to generate phages like this means that you're generating DNA sequences, which in turn has um some effects that are go beyond potential benefits that go beyond just you know, this example you gave. But there could be other benefits to gene therapies in general, right? Um being able to look at the whole um sequence and being able to generate entire sequences rather than just look at individual pieces or tiny parts of it.
SPEAKER_02Yeah, exactly. And we're talking about like thousands and thousands of these, you know, nucleotide bases in the DNA sequence. So even if you think about a single gene, that's huge in terms of the amount of DNA that's encoded there. Um, and then you're like, okay, well, it's even larger to make a whole genome. It's gonna have genes in it, it may have like other sequences inside of it that are gonna regulate or special sequences that are for recognition. And then there's the question of like, how do how do these all interact with each other? And what do they all encode? And scientists don't even necessarily know all that they encode and how they interact. And so when you're trying to like do something new with, say, a bacteriophage or or a virus, right? You're like, I don't know, is changing this gonna make it better? Is it gonna make it worse? We're not even like, do we know what it's gonna do? Sometimes yes, sometimes no. And so, well, then the idea becomes could a large language model ingest all of these genomes from all of these viruses and kind of figure it out, right? Like maybe we can't see inside that black box, but maybe there are patterns that could be noticed in these huge swaths of data that the large language model could then create these sort of fresh genomes to
Why Synthetic Phages Could Matter
SPEAKER_02do something.
SPEAKER_00Yeah. And that's kind of so so step taking one step back. I mean, I think this is where we can dive right into the paper itself. Um, and we said that phages are viruses that infect bacteria. Um, and what it is is is sort of a protein shell that is sometimes comes with a tail apparatus that is wrapped around a genome. And one example of this is something called phi X174, uh, phi, the Greek letter phi X174. And uh this is one of the phages that have been used for uh that that is that is the that attacks um E. coli um cells. And uh this is what this paper focuses on in uh in more detail as well. And it the reason they focus on this one is that it's because it's a tiny, it's pretty tiny, it's single-stranded, it's um you know, it's only got 11 genes, and it's got um uh what they call 5.4 kilobases. I think that's 5,400 bases.
SPEAKER_01Yeah, it's tiny. It's just 5,400 little little nucleotide bases. That's it.
SPEAKER_00And and so the idea is that uh they took that as an example starting point. But uh, but to your point earlier, their base starting point was actually like a language model. Just like ChatGPT, just like um a quad or whatever else we use for text, uh, we have a language model, the same underlying neural network architecture um for generating sequences, because that's what these transformer models do, right? The whole point of a transformer model is you give it a bunch of tokens of text and it produces the sequence that's the next set of tokens that matches a pattern that it's learned. In this particular instance, uh, that pattern that the language model has learned is DNA sequences. And there are two big models called Evo 1 and Evo 2, both are called genome language models. Um, and they work just like a text language model. They predict the next token. Um, except, of course, these individual tokens are individual DNA nucleotides rather than words. Um and so So that's why I was saying all those things at the beginning.
SPEAKER_02C, T, G. Like there are there's four of them, right? It's like which letter from the alphabet do you want? C is for cytosine, A is for adenine, T is for thymine, G is for guanine. So it's just predicting, hey, which one would I, which one of these four, right, would I expect to come next?
SPEAKER_00Yeah. And and just to be clear that there are differences between uh a genome type transformer model and a language, a normal text language transformer model, these line, these sequences are long, right? These day DNA sequences are super long. So they have to like change the architecture a little bit behind the scenes to allow for really long inputs. Uh and so they had to do that. But but besides that, that's it's basically the same idea. And so what they do is they do the step of pre-training. Much like um in our LLMs, we pre-train on all of the web data. Uh, they pre-train on a massive corpus of genomes across all domains of life. Um not just like virus, not just like similar viruses. That includes 2.7 million bacteriophage genomes. Oh my gosh. Um, and this is kind of where it absorbs the general kind of grammar of how these genomes um are built, right? And so it's got a massive, it's absorbed all the stuff, right? It knows how DNA, I shouldn't say it knows how, but it it has a pattern for how DNA sequences work. Um and then what they did was uh they did a step, which is often done even in text-based models, which is called supervised fine-tuning, which is say you have a large language model, but you really only want to use it for one task and you care about a set of data, and you have, you know, you have sort of right answers for a whole bunch of example problems for your set of data. So you want to just for it to do that job really well. So you do what's called fine-tuning, which is you train the model that's already been pre-trained, you train that model again on this smaller set of data, and that'll change the knobs inside of the neural net just a little bit. It'll fine-tune it, literally, if you want to think about it that way, um, to do better at this one task. So they did that here. They fine-tuned it to only focus on uh 15,000 um genomes that belong to the family of the uh Phi X174 phages, um, just so that they can specialize it on that specific phage.
SPEAKER_02Um, I mean, that makes sense. We don't really want something that's gonna like infect us, right? We want we want something that's gonna that's gonna work on E. coli.
SPEAKER_00That that's fair. That's fair. But again, I mean, bear in mind that the AI system isn't making phages. The AI system is just returning letters. C A, G, T, whatever.
SPEAKER_02Yeah, but the humans did make phages out of some of these. Oh, yeah. So we'll talk about that. We'll talk about that. We'll talk about that. I'm getting ahead. I'm getting ahead. Sorry.
SPEAKER_00No, no, it's that's that's one of the most exciting parts of this paper is that it's just not a theoretical idea, but they actually went ahead and made some of these, right? So um, but once you have the language model that's specialized and trained, how do you use it? Well, you prompt it, right? Yeah. That's right. So they can they also just prompted it. Of course, the prompting for a genomic language model
Genome Language Models And Fine-Tuning
SPEAKER_00is different. What they do is they give it the first few nucleotide sequences of um um of the the phage they want, the X, you know, phi X174, um, to get it going, kind of get in, you know, bootstrapped up with some initial set of um DNA letters, right? So I think the specific letters were G A G T T T A T C G. So um, so they, you know, and of course, you know, they also did some things clever. They um when they trained it, they also arranged all the um special when they did the fine-tuning, they also arranged all the letters in a way that was, you know, kind of predictive of the next sequence. Kind of they organized it before they fine-tuned it. Okay. Uh, which is kind of interesting too, um, so that it will be able to recognize that intro, right? That intro few um letters. So that's it. And that's the uh input, and out pops out the rest of the sequence. And of course, now you're wondering, okay, now now what do you do with this? Well, you do the biology of it and you build it and so on.
SPEAKER_02But before which ones, yeah, exactly.
SPEAKER_00So in theory, you can run this repeatedly over and over again, infinitely, right? You can keep running this, and it's going to be because it's a neural network, it's going to be um, and if you set the, you know, um make it like allow it to give you different outputs each time, it's gonna be slightly different each time. And and that's a good thing. And in fact, in this case, they produced a whole bunch of them. But that's that I think this is where it gets very interesting, also, which is it's not enough to just have a collection of a ginormous collection of these sequences. You can't make them all. Um, but also some of them might be garbage. Of course, yeah. So, like, how do you know what's a good one, what's not? So they employed a series of filters, constraints, to ensure that first of all, the genome that's produced is something that is plausible, right? Um, of course, it'll produce the correct letters, but it shouldn't be.
SPEAKER_02I was gonna say, mean you mean you don't want like a P in there?
SPEAKER_00Yeah, exactly. So, you know, they made sure that there were constraints set up so that it didn't produce garbage. And one kind of garbage is simply absurd um polymer, homopolymer runs like 30 identical bases or something, something weird like that. Right. Right.
SPEAKER_02T T T T T T T T T T T T T T. Yeah, okay.
SPEAKER_00It must actually encode recognizable proteins, right? So that's a check that they can do after the fact, after it's generated a whole bunch. Um and they had to build apparently a custom gene annotation tool um because there's some limitations with the standard ones. But in any event, that's one, they had to do some quality control on the on these on these phages, right? So on these um sequences that were produced. So then the next thing they had to do was also um what they call tropism constraints, um, in which they had to ensure that the um uh produced phage will actually grab on to the specific host bacterium surface receptors.
SPEAKER_02Oh, interesting.
SPEAKER_00And so so that means it must keep within greater than 60% identity um of the um identity to the phi one x174 spike. So that means that the it they were imposing a constraint to for this specific type of phage, but also on the specific host type.
SPEAKER_02Um okay, so like make to make sure they thought it could actually like get into the E. coli bacteria.
SPEAKER_00Yeah, not only that, they wanted to specifically target E. coli C, right? And so they were able to verify that the generated phages infected E. coli C only, but not six other E. coli strains.
SPEAKER_02I mean, that's good. We like don't want some random virus like running rampant infecting other things. So I think that's actually like a good safety check they did on this.
SPEAKER_00Right. They also wanted to enforce novelty and diversification constraints. And what that means is um they wanted to generate not just near copies of the Phi X174, but interesting varieties of them, right? Um, and so they wanted to make sure that there was some variety in the outputs, and that would make it stronger set a set of uh potential candidates.
SPEAKER_02This is so interesting because if you think about what happens in like in reality, not not in this AI reality, but in reality, it's like if you think about um evolution and natural selection, it's usually like one small thing changing at a time. And here it's like, well, you could produce something that's, I mean, similar, but has a bunch of changes. And that's that that to me is like a very interesting finding because it's not replicating how things happen in the natural world, and you can see a lot more changes and because of something like this all at once, rather than having like change a little change here, another little change there, another change there. Like you can have it all happen all at the same time. Exactly. Exactly.
SPEAKER_00So all of that is to say
Filters That Make Designs Testable
SPEAKER_00it produces a list of letters.
SPEAKER_02It does, lots of letters that then humans have to make them. Yes, right? And that's what they did. I think they made almost 300 of these viruses, these bacteria phage viruses. And then they said, okay, well, let's test them out, let's see if they work. Do they actually infect the E. coli and like kind of like burst, burst the bacteria uh and kill it? Um, and and for 16 of them, the answer was yeah.
SPEAKER_00Wow.
SPEAKER_02Yeah, they did. And they were completely novel. These were not viruses that existed in nature before this study was done. And so they created novel viruses that could infect E. coli, and they did some more testing on them and they found that like some of them worked better than others, right? Some of them were like faster at infecting the bacteria than others, some of them were slower. Like, and then one of the interesting findings from a from a clinical angle is that some of them worked on resistant E. coli. So there were there's like E. coli out there that's resistant to, you know, whatever. Like it could be resistant to antibiotics. We worry about those things because, say you have something that's resistant to antibiotics, you get an infection, you have no good way to treat it. So the question was like, hey, we don't know if this will work, but is there a possibility that these can infect resistant E. coli? Yeah. And the answer was yes.
SPEAKER_00The answer was yes, right. They were, they did a bunch of resistance experiments and they found that in fact the the designed cocktail of phages um, you know, was able to overcome the resistant E. coli. Um, and those especially those E. coli that the natural sort of collection of phages couldn't.
SPEAKER_02Exactly. And so that is something that, you know, you you snap your fingers almost not uh they didn't really snap their fingers, but like pretty quickly created something that m would take maybe uh a longer time from an evolutionary innovation standpoint. Exactly. And instead, we have this very genetically diverse, you know, set of phages that were created that in theory, at least, could make it so that you could treat someone with a resistant E. coli bug.
SPEAKER_00Yeah. And in fact, that's exactly I wanted to mention this because that's one of my also one of my favorite parts of this paper was that um you can do this sort of directed evolution and and sort of basic engineering and make something that's otherwise slow and incremental to something that's a very fast process. And so because the generative model, the EVO one and EVO two and super, I mean fine-tuned, supervised fine-tuned with the fade stuff, um, can propose thousands of viable-ish candidates, right? And uh and this is across a huge swath of sequence um space, and it can do this overnight, it can do this very, very fast. Um, and the best part is it can find architectures, uh, phage architectures that evolution hasn't happened to try yet.
SPEAKER_02Yeah, that's crazy, right? It's like sane.
SPEAKER_00So, like one of the phages they used a capsin capsid protein borrowed from an evolutionarily distant relative, and it's still folded and assembled. And I I really think that this kind of long-range recombination um is hard to reach just by tinkering. Like you need, I mean, the something like this can generate these kinds of results really quickly.
SPEAKER_02Yeah, and I there's so much data to sift through. Like, this is a great application for it. Um, but at the same time, I was elephant in the room. I know. I I know we have to talk about it.
SPEAKER_00Why are we making viruses, right? I mean, at the core of
Resistance Breakthroughs And Biosecurity Fears
SPEAKER_00it.
SPEAKER_02Or like what else could you what else could this technology be used to do? And like how safe is it really? Like this, these viruses were just viruses that infected E. coli. They can't infect humans, they can't even infect other strains of E. coli. Like, great.
SPEAKER_00However, but those are all human imposed constraints, like I talked about before.
SPEAKER_02Right.
SPEAKER_00That was the responsible researchers putting in those constraints.
SPEAKER_02Right. But like, what if you had what if you had someone who wasn't responsible researcher? Exactly, and wasn't thinking through those things. Or what if you have someone that does want to create like the next, I don't know, super influenza or something that could really hurt people? Like you can see that this could be used for good, but then at the same time, I'm like, you flip the coin and you see, like, oh, geez, we could create something that's more virulent, that spreads more easily, that does bad things. And so I think this is where the big question mark comes in my head is like, should we be doing this kind of stuff? And how do you govern it? And what sort of policies do you put around it? And how do you stop people from creating something that's bad, especially when you don't even know what bad is? Like it's right, we were talking about how it creates things that like really haven't been seen before. These are novel viruses, or they're pulling something from like, you know, a very different virus and adding it to this one or whatever it is. It's like, how do you even, how do you know? Like, how do you know what it's gonna do? And how do you stop it from from doing bad things?
SPEAKER_00Yeah, no, that's exactly right. And um, there's also the issue of even if you're not produce even if you produce some carefully in the lab, they can still um be accidentally, you know, let out and things like that. So like there's the these can go crazy in the wild, right? I mean, this is not you are producing biological specimens at the end of the day.
SPEAKER_02Exactly. And you create one and it could sort of like combine with others out there and be out in the world, right? It just they can exchange DNA, etc. And so I I think this is not a question that we have an answer to, right? I think this is a part of an article like this that scares me. I think this article is really cool, and at the same time, I think it like raises a lot of red flags.
SPEAKER_00Yeah, you know, it's it's it's both exciting and terrifying at the same time. I mean, this is true of any type of technological and you know innovation in that sense, but you know, there's always a bad side to it, right? I mean, there's There could be. Yeah, yeah.
SPEAKER_02And then there's a question of like, how do you stop that bad side from showing? How do you make sure that there's you know safety mechanisms in place, that there are policies in place um to prevent this from being used for something bad.
SPEAKER_00Yeah, exactly.
SPEAKER_02A somber ending to this episode. Very cool, very cool study. Very cool that they were able to create these novel viruses with pieces that we'd not not seen before in nature. But at the same time, I think it's very sobering to think about, oh gosh, they actually did create things that have never been seen before in nature. And if you think about other applications, not being sure what these things will do, it is very sobering.
SPEAKER_00Thank you for joining us.
SPEAKER_02I'll see you next time on Code and Cure.