October 1, 2026
Heard in AI

Rare-disease AI founder says his edge is harnesses, not models

Gamow Labs founder Daniel McKinnon says startups like his can't compete on building models. His team combines frontier models, builds tools and tests the whole system, and he says that layer is thinning as models improve.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on The Cognitive Revolution, episode published October 1, 2026

In one rare-disease case, McKinnon says, an AI model named a gene as the cause and then swapped it for another. Daniel McKinnon, founder of Gamow Labs, which uses AI to interpret patients' genomes, recalled how Grok 4.6, then the top model on his company's benchmark, said a gene such as NRF2 was responsible and then "just changed the name of the gene to something totally different." The case was counted as a miss. "I'd never seen that before," McKinnon said. "And I was just like, this is dumb. And our harness and our tools prevent the agents from doing dumb things."

McKinnon told the story on The Cognitive Revolution, in an episode published on 1 October 2026 and hosted by Nathan Labenz and Prakash Narayanan. His point was that a company like his makes its contribution in the software and testing around the AI models, not in the models themselves. His account shows what a specialised AI startup does on top of models built by the big labs.

A harness, tools and evals

"We're basically a harness, a tools and an evals company," McKinnon said. A harness is the software that runs a model through a task: it decides what the model sees, which tools it can use and how its steps fit together. Tools are the programs the model can call, such as a literature search. Evals are structured tests that score how well the whole setup performs.

McKinnon expects this pattern across "vertical AI", the startups that apply AI to one industry or problem. He has worked inside frontier labs. He said he started on the first language-model project at Meta, OPT-175, and worked on evals at both Meta and Google. His conclusion is that unless you are "somebody who is very famous with very deep pockets", you are not going to compete at the model layer, because intelligence is progressing so fast. He added that he did not want to bet against the new labs trying to.

Mixing models

The first thing Gamow Labs does is combine models, which McKinnon called ensembling. He said it helps with cost as well as performance. Gamow's published benchmark, RareBench, shows that the cases cluster: Claude is good at some kinds, Astra at others, and even Gemini, which he said is "not on the frontier", does well on certain cases. There are "very dumb ways" and smart ways to combine them, McKinnon said. He described the routing, meaning which model gets which case, as an important edge that Gamow and many others are working on.

The RareBench 0.1 report, published in August 2026, covers 122 cases. For each case the model gets the patient's symptoms and genome information and tries to identify the causal variant. The scores are top-k recall: a case counts as solved when the correct causal variant appears among the system's top-ranked hypotheses. The report also tracks separately cases where the system found the correct variant but did not rank it highly enough to make the top-k cutoff, and cases it missed entirely. On this measure the report has Grok 4.6 solving 35% of cases, just ahead of Claude Opus 5 at 34%. The benchmark is meant as a target for improving models, harnesses and tools, not as a measure of diagnoses delivered to patients.

Reading the traces

The second thing is evaluating the whole system. "It's not just the model at this point, it's the whole system," McKinnon said. Without structured measurement, "you don't know what to hill climb", meaning which step-by-step improvements are actually working.

That includes what he called pretty robust trace analysis. A trace is the record of every step an AI agent took, every tool it called and every piece of text it read. Gamow compares how models perform with and without its harness. "This sounds stupid, but like not that many people actually look that closely at data," McKinnon said. The renamed gene was one such finding.

Gamow's July 2026 case study on a blinded reanalysis of lung-disease genomes lists similar failures by ChatGPT. In one, the model did not check a deletion that was present at 40% frequency. In another, it mixed up coordinates from two versions of the human reference genome, hg19 and hg38, and recognised the mistake only after being prompted with the correct hypothesis.

Testing tools without the model

Narayanan asked the harder question. If Gamow keeps upgrading to newer model families, how does it measure the value of its own layer separately?

"I honestly don't have like a great answer to that question," McKinnon said, because his team co-designs the tools with the models. He gave two examples involving Astra. In the first, a literature-search tool had tagged papers for the model to read. Astra said it had read them, but the answer was in a paper and it missed it. Looking through the trace and the context window, the text the model actually had in front of it, the team thought: "I don't think you actually read these papers." McKinnon calls problems like this "little paper cuts": a case missed for one reason here, another there. The tools exist to put the model "on these better rails."

In the second, the first time Gamow ran its eval on Astra, the model "just stopped". Some of the tools made it want to think about things too much. McKinnon compared it to people online who ask a model to do research and come back 30 minutes later to find it looking at flights from Dubai to Calcutta. The eval caught the problem early, the team changed how Astra called some of the tools, and it worked again.

McKinnon was careful not to downplay the models. "I don't want to communicate that these models aren't amazing. They absolutely are," he said. But he argued that without a specific eval it is hard to tell whether an impressive-looking result was actually good.

Why not just use more computing power?

McKinnon said that if he worked at OpenAI or Anthropic with infinite compute and could run millions of rollouts, meaning separate attempts at the same case, the agents might converge on an answer and make much of this work unnecessary, especially since they can now write their own tools.

The obstacle is cost. The goal, he said, is "consistently and repeatedly and, you know, honestly, affordably getting to the right answer." At $10 or even $100 a case, "who cares." But if it took a million rollouts, at $1,000 or $10,000 a case, it would not make sense. His team is researching how those cost curves work.

The RareBench report shows how much costs can differ. Grok 4.6's 35% top-k recall cost about $1.98 per case, and Opus 5's 34% cost $9.35 per case, nearly five times as much for a similar score. DeepSeek V4-Flash scored about 20% at $8.55 for all 122 cases.

Bullish and bearish on his own layer

Prakash asked what McKinnon most wants the frontier labs to improve. "I want them to hill climb my task," McKinnon said. If models get really good at clinical genetics, "everyone wins", and Gamow probably has less to do at the harness layer.

That is why he called himself "both bullish and bearish" on vertical AI. He sees lots of value in owning the customer relationship and said OpenEvidence has shown this. But the layer he builds "is getting thinner." When he started, vanilla OpenAI o3 would not do the task: he said it scored 0% on his benchmark when he re-tested it "just for fun", which is why he needed so much scaffolding. Today he puts the gap at "20, 30 percentage points or something" and said it is "definitely like shrinking." He did not give an exact figure.

He is not attached to the harness for its own sake. His goal, he said, is to build "this great AI native clinical diagnostics company", and he wants the best person to do the work.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02

From the conversation

Podcast episodes

The Cognitive Revolution

AI:AM: Was Trump-Xi Anything? What Counts as Utopia? + AWS GPUs Cost 3X & AI Diagnoses Rare Diseases

Episode published This article draws on 2:59–3:42 and 1:12:20–1:19:06 (approximate times)

Article history

Updates to this article

Tags