In August 2026, researchers at Stanford and the Arc Institute published, in the journal Science, the first bacteriophages -- viruses that infect and kill bacteria -- ever designed from scratch by an AI genome-language model rather than found in nature.[1] The AI side of the work was fast and cheap by biology's normal standards. The part that actually mattered took a lab, real time, and a lot of failure.
The team fine-tuned genome-language models called Evo 1 and Evo 2 and used them to generate roughly 700,000 candidate genome designs for phiX174, a well-studied bacteriophage that infects E. coli.[1] Generating those 700,000 candidates was the fast part -- a computational process, not a biological one. Researchers narrowed that pool down to about 300 designs worth actually synthesizing and testing in a real lab. Of those 300 synthesized genomes, 16 produced a genuinely viable bacteriophage: one that could infect an E. coli cell, replicate inside it, and burst it open, confirmed under electron microscopy.[1] Some of the 16 outperformed the natural, evolved version of the same virus; a cocktail of several of them overcame E. coli's resistance to any single one.
The AI model didn't get it right and then get verified. It generated a huge field of plausible candidates, and reality did the actual sorting. Roughly 5.3% of the synthesized genomes turned out to be viable -- and that 5.3% only exists because researchers were willing to synthesize and test 300 real candidates rather than trust the model's own confidence about which ones would work. A genome that looks complete and plausible to a language model is not the same claim as a genome that can survive contact with a real bacterial cell, and the gap between those two claims is exactly the 284 candidates that didn't work.
Why does this matter? This is a genuine scientific achievement, not a cautionary tale -- 16 new, functional, AI-designed life forms is a real result, published in one of science's most selective journals. But the actual shape of the achievement is worth being precise about: the model's contribution was a large field of candidates generated almost for free; the discovery was which ones were real, and that step still required a lab willing to synthesize and physically test hundreds of genomes that mostly didn't work. Generation got radically cheaper. Verification against physical reality did not.
Companion piece on this outlet: "Researchers Found 34 Million Distinct Concepts Wired Into One AI Model..."