Daniel McKinnon's son Owen died of a rare lung disease. The genetic cause, a deletion of 91,000 DNA letters, was missed by the lab that sequenced his genome. A human specialist found it later. Afterwards, while his family was looking for answers during another pregnancy, McKinnon built an AI pipeline of his own and fed it the old data. It found the deletion again.
That experience led McKinnon to found Gamow Labs, a company that uses AI agents to interpret genomes and revisit cases that clinical labs left unresolved. He told the story on The Cognitive Revolution, in an episode hosted by Nathan Labenz and Prakash Narayanan and published on 1 October 2026.
Three negative genomes
McKinnon described "basically five years of heartbreak" before the birth of his son Warren, now a healthy 13-month-old. Between losing Owen and having Warren, the family also lost a second pregnancy very late, for genetic reasons. Because of that history, their doctors watch closely. When something looked slightly questionable on a 16-week pregnancy scan, their maternal-fetal medicine doctor told them she normally wouldn't flag it, but in their case they needed to look carefully.
The family had the fetus's whole genome sequenced. It came back negative. McKinnon said this was the third negative whole genome he had seen since Owen's death, and he did not trust the result. "I'm heartbroken, but I'm mad, and I'm going to do something," he recalled thinking.
He called the labs and asked for his family's raw data, which he said patients can get through their HIPAA rights. Then he "vibe coded" his own interpretation pipeline. Vibe coding means building software by describing what you want to an AI model and iterating on the code it writes. McKinnon said this was harder then than now, before tools like Claude Code, and took a lot of work on his part. He built it hoping for some comfort about the pregnancy. To his surprise, it also performed well on the family's earlier cases, including Owen's.
Why the deletion was missed
Owen's deletion did not break a gene directly. It removed an enhancer, a stretch of DNA that helps switch a gene on, and that enhancer sits about a million DNA letters away from the gene it controls.
McKinnon explained that the lab, like many others, used a filter for structural variants, the larger deletions and rearrangements that are messy to detect. Today's sequencing breaks a genome into roughly a billion fragments of about 150 letters each, which then have to be pieced back together. The filter ignored anything more than 1,000 letters upstream or downstream of a gene. McKinnon called that "a totally reasonable trade-off if you have humans looking at all of this stuff." Gamow Labs' published case study describes the same failure: an earlier clinical filter excluded a deletion in an enhancer of the FOXF1 gene because it lay more than 1,000 letters from coding regions.
His idea was to put OpenAI's o3 model, which he called the first real breakthrough among agentic models, into a loop and have it keep looking. The first pass checked the obvious things, the parts of genes that code for proteins, using bioinformatics tools that he says were crude by today's standards. When that turned up nothing, it went round again, and again. A clinical lab cannot work that way, he said. Staff can spend only so much time on each case, and a case without a conclusion is reported as non-diagnostic. In McKinnon's account, most whole genomes come back non-diagnostic, even for infants with suspected genetic disorders.
He sees this as a time problem that machines can take on, since they can work nonstop and in parallel. He was careful to say that people have long built software to speed up genome interpretation. What he thinks he spotted early was that the task suits agentic AI: it takes many steps over a long time, answers can be checked, and current systems are far from solving it. He forecasts it could become a task like maths, where people generate harder and harder problems and models keep improving until it is "basically solved."
Sequencing is cheap; interpretation is not
McKinnon recalled that when the human genome was sequenced 25 years ago, Bill Clinton and Tony Blair said it would revolutionise medicine. Since then, he said, the field has left a "graveyard of genomics companies." Sequencing itself has become cheap: below $1,000 a genome, then $500, and some high-throughput labs now advertise less than $100. In his view, what was missing was the interpretation layer. That means working out what a patient's variants mean, ranking them, explaining to doctors why they matter and how the patient might be treated.
He pointed to RareBench, a benchmark his company built in which a system must rank the genetic variant behind a case. He said the best traditional tool for ranking variants, LIRICAL, scores "something like 10%" on it. LIRICAL is a program that weighs how likely each observed symptom is under different candidate disorders, optionally combined with genome data. McKinnon said Opus 5.5 running in Claude Code, a vanilla setup, scores "something like 50%."
Those figures differ from the company's published RareBench 0.1 report from August. It covers 122 cases, and its best full-benchmark scores were 35% for Grok 4.6 and 34% for Claude Opus 5. McKinnon also mentioned "other benchmarks we have," so his numbers may come from a different or later run. Labenz said the 50% is a benchmark score, not the share of patients diagnosed. The report itself presents the benchmark as a target for improving models and tools, not a measure of clinical diagnoses.
Diagnosis, not screening
McKinnon separates screening healthy people for possible future problems from investigating an illness that is already there. Gamow Labs does the second. Sequencing already has a market, he said: every baby in top neonatal intensive care units is sequenced, and babies in other units would be if hospitals had the resources. The question is better defined, too: "baby has pulmonary hypertension, figure out why."
The company works in two ways. In what he calls "online" work, a case comes in after traditional labs gave doctors no useful answer, the patient agrees to reanalysis, and Gamow Labs reviews the data. Sometimes, he said, "but not always," it finds something that helps decisions about a pregnancy or a newborn. The second is bulk reanalysis for rare-disease centres that hold thousands of genomes but have only a handful of genetic counsellors and bioinformaticians. Because a machine does most of the work, he said, a batch of 80, 100 or 200 genomes can come back in a week, work that would take a centre six months or a year. The company's website describes interpretation as ongoing: old cases can be reread when new research links a gene to a disease, without sequencing again. McKinnon said the company is four months old, has no revenue yet and is not focused on a business model: "we're just trying to diagnose more sick kids."
He said the company has diagnosed many previously undiagnosed children. The company's own blinded study, run with a specialist at Baylor College of Medicine, covered a narrow set: 46 people who were either infants lost to interstitial lung disease or their healthy family members. These were hard cases involving complex structural or intronic variants that clinical labs had initially missed. Its system recovered all 19 diagnoses the specialist's lab had already reached and contributed two newly resolved ones. It returned no diagnosis, or only a variant of uncertain significance, for all 20 healthy relatives, which the study counted as correct, and left five cases unresolved. The results come from one disease group and do not show how the system performs across rare diseases generally.
Stuck at 'uncertain'
McKinnon said most of the hard cases his team sees end at a variant of uncertain significance, or VUS. Variants are scored using a standard points system from the American College of Medical Genetics and Genomics (ACMG), and a variant needs six points to count as "likely pathogenic," meaning likely to cause disease. That label can open up treatment and insurance options. Being in between, he said, is "very bad."
A family has two ways out. One is to wait until other patients turn up with similar symptoms and the same variant, which adds points. McKinnon said this explains news stories about a child finally diagnosed after ten years of epilepsy, but it is slow and does not scale. The other is to persuade a university lab to run a functional study, testing the variant in cells that resemble the affected tissue. For alveolar capillary dysplasia, a rare lung disorder, the usual choice is IMR-90, a fetal lung cell line. That matters, he explained, because the relevant gene is active only between weeks 16 and 20 of development, so a cancer cell line will not do. Researchers can then edit the genome or insert DNA to see whether the variant breaks the gene. By McKinnon's figure, such a study can add two to four points. Expert ClinGen guidance says the weight depends on the experiment's controls, replication and relevance to the disease.
His team has bought what was left of a company called Arpeggio and hired four people to run these experiments. When they spoke, the team was three weeks in and running its first experiments that day. McKinnon said such testing is not commercially available. He hopes to offer it to clinical labs trying to move a variant up to likely pathogenic, and also to use the results to improve the models. "Right now, the machines aren't learning," he said. A VUS is a dead end, but an experiment can show that a stretch of the genome matters for a disease and check the models' predictions.
Asked by Narayanan how to make experiments cheap enough to run at scale, McKinnon pointed to a mix of AI and robotics rather than one invention. He described a robot arm running 384 experiments at a time, cartridges holding thousands of plates, and AI helping to design experiments and primers (short DNA pieces used in testing). He said using AI for this requires special permission. AI is also needed to sort through the very large datasets that result.
'I want to help'
McKinnon said everyone he has met at frontier AI companies has told him "I want to help." Whether that means a "trillion dollar company" can pivot to his problem is another matter, he said, having worked at such companies himself. He would still be surprised if the top labs did not spend time on it. He also urged them to tell this story, if only for selfish reasons. He contrasted Dario Amodei's recent tweet about curing cancer in five years with "We're diagnosing kids today." In his view, people unhappy with data-centre construction or worried about AI taking jobs should hear about a concrete way AI is helping now. He also thinks mid-career tech workers who have spent years on ads ranking or data cleanup want work that helps children.
The cost of slowing down
After the conversation, Labenz said honesty requires admitting that slower frontier progress would trade away something here. He still thinks pacing is "probably a good idea and probably worth it," and does not believe the frontier companies would be unable to grow revenue or be overtaken if they paced their efforts. But AI is still climbing the hill on problems like this one. "That 50% number that Claude gets to today," he said, "obviously leaves half of the questions unanswered." For individual families those questions matter enormously, and he called pacing "a costly compromise."