27 September 2026
Heard in AI

Moonshots panel weighs 12 testable targets for AI-driven biology

On Moonshots, recorded September 25, 2026, the hosts welcomed FutureHouse and Edison Scientific's twelve "Millennium Problems for Biology": hard lab goals with pass-or-fail tests. The panel discussed scaling the idea to 100 targets, each worked on by 1,000 AI agents, with everything the agents do made open source. Host Peter Diamandis disputed the list's exclusion of age reversal. He drew on his experience designing the XPRIZE Healthspan competition and argued that age reversal could be measured in days. That competition's published rules, however, rely on controlled trials with a one-year treatment period, and participants are expected to take part for at least 14 months. Anthropic's agent-assisted enzyme search showed the current state of such tools: agents ranked candidates, humans ran the experiments, and the enzyme system's function is still unknown. Alexander Wissner-Gross's plan to search simulated "digital twin" cells for cures is his forecast, not a working capability.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Moonshots with Peter Diamandis, episode published 27 September 2026 (recorded 25 September 2026)

The idea on the table was to stop asking AI to "solve biology" and give it a list of tests instead. Near the end of a Moonshots episode recorded on September 25, 2026, host Peter Diamandis brought up the "Millennium Problems for Biology". His co-hosts Alexander Wissner-Gross, Dave Blundin and Salim Ismail and guest Emad Mostaque treated the list as a practical answer to a question the AI boom keeps raising: if machines are going to make discoveries, how does anyone know when a discovery has been made?

Hard to solve, easy to check

On September 23, Sam Rodriques and Michaela Hinks of FutureHouse introduced twelve biological challenges, put together with Edison Scientific. The name echoes mathematics' famous Millennium Prize Problems. The authors picked goals that are very difficult but can be checked plainly in a laboratory. The list includes the origin of life, cryopreservation (freezing living things so they can be revived), limb regrowth, a better version of Rubisco (the enzyme plants use to capture carbon), new nitrogen-fixing enzymes and bacteria that produce gene therapies. The authors call it a living project that can be revised and expanded.

Each problem comes with a written pass mark. To solve the cryopreservation problem, intact adult mice would have to stay frozen or vitrified for at least 24 hours and then recover with more than 99% viability and no permanent damage. To solve limb regeneration, adult mice would have to regrow limbs whose movement and sensation match those of control animals so closely that blinded observers cannot tell which limb regrew. These are proposed criteria. No one has reported meeting them.

Diamandis summed up the design: the problems "should be hard to solve and very easy to check," and, in his reading, verifiable in a standard wet lab (a lab for hands-on biological experiments) in a couple of days. He tied it to a formula from a book he wrote with Wissner-Gross: "Pick your targets, build your harnesses, set up your benchmarks, and go." Mathematics, Diamandis said, was already "incinerated" by AI, and Wissner-Gross added "thoroughly". Biology was the next target.

When Diamandis put the list to Mostaque, the reply welcomed it but pushed for far more: around 100 such problems, each attacked by 1,000 AI agents, with everything they do made open source. Agents here are AI systems that plan and carry out multi-step tasks on their own.

Blundin said his daughter, who works in biotech at Moderna, called it the right list. He argued that a framework for measuring progress matters because, as he put it, the Nobel Prize is becoming "completely irrelevant" as a yardstick.

The fight over aging

Diamandis's favorites included the origin of life, cryopreservation and regrowing limbs. But he said the list "missed age reversal." Blundin replied that the easy-to-check rule was "why age reversal probably isn't on the list." The FutureHouse authors say much the same: they left out broad goals such as solving aging because success would be hard to judge.

Diamandis rejected that reasoning with an account from his own prize design. When he was setting up the $101 million XPRIZE Healthspan competition, he said, Peter Thiel and Aubrey de Grey first proposed a longevity prize. He did not see how to run one without a 20- or 30-year time horizon. Geneticist George Church, he said, told him to "forget about longevity, measure age reversal". If a treatment makes someone functionally 20 years younger, Diamandis argued, "that's measurable in days."

The competition's published rules reflect the switch from lifespan to function, but they show that judging it takes much longer than a few days. The seven-year competition asks teams to restore muscle, cognitive and immune function together in people aged 50 to 90. Treatments are tested in controlled phase-II trials (mid-stage clinical studies) with a before-and-after, single-crossover design. Each participant is tested three times over a three-month baseline period, treated for a one-year intervention window and then tested again, so the trial measures change within each person against their own baseline. Teams must also include time controls, standard of care or another suitable control, and participants are expected to take part for at least 14 months. Results are measured against normal age-related decline. Main awards of $61 million, $71 million or $81 million go to teams that reach 10, 15 or 20 years of restored function. A single functional test may be quick, but under these rules proving that age reversal worked still depends on a controlled clinical trial, not on a two-day lab check.

A real discovery with an open question

The panel had just discussed a concrete example of AI-assisted discovery. Diamandis described how Anthropic gave its Claude model a single instruction: search a huge DNA database for interesting reverse transcriptases, enzymes that copy genetic material from RNA into DNA. He said 950 Claude agents searched for 21 hours and found a repeating DNA pattern that looks like CRISPR, the bacterial system that became the basis of gene editing.

Anthropic's own account, published September 23, confirms roughly 950 agents working over 21 hours. It also shows a division of labor that "working autonomously" does not capture. The agents surveyed more than 200,000 reverse transcriptases, flagged about 3,500 candidate systems and chose twenty for detailed reports. Along the way they reviewed the literature, reproduced known results and criticized candidates, and many proposals were dropped before any testing. People did the lab work. The finding, which Anthropic calls an array-associated reverse transcriptase, pairs a previously known reverse transcriptase with a neighboring helper gene and an overlooked repeating DNA array, found mainly in viruses that infect bacteria. Human experiments showed that the array produces distinct short RNAs. What the system does in nature is still unresolved, and more experiments are under way.

Diamandis said the researchers expect it to resemble programmable enzymes used in gene editing. He then asked whether Nobel-level work would soon arrive weekly. Wissner-Gross said it would. He argued that when AI can produce "a thousand Nobel Prizes in one year," the prizes will shift to larger bodies of work.

The Anthropic case shows what the Millennium list is trying to change. The agents found something new before anyone understood it, and the question of its function is still open. Under the FutureHouse approach, a result counts only when it meets a test agreed in advance.

Not everyone treated the pace as simply good news. Ismail noted that once CRISPR let scientists edit DNA "as easily as you can edit a Word document," combining it with AI sends both the possible benefits and the possible downsides "through the roof." Wissner-Gross called it hypocritical for AI companies that warn about AI's dangers to open wet labs. Still, he said he backs Anthropic and other leading labs doing so, because without superintelligence he does not expect the top 5,000 diseases to be solved.

Wissner-Gross's next step: searching digital cells

Wissner-Gross described where he thinks this goes over the next few years. Language models learned from much of humanity's text and images online. In the same way, he expects "digital twins" of the cells in our bodies, meaning simulations trained on huge datasets that include experiments where cells were deliberately altered. Researchers would then search those simulated cells for ways to cure disease, starting from a diseased cell and looking for a route to a healthy one.

His model for that search is AlphaGo, Google DeepMind's Go-playing system. As DeepMind explains, AlphaGo combined neural networks with search. One network suggested promising moves and another estimated who would win, so the system could spend its effort on the most promising lines of play. In its 2016 match against Lee Sedol, its unusual Move 37 in game two helped set up a win. Wissner-Gross called the goal of cell search "a move 37": an unexpected intervention that moves a cell from sick to healthy.

He argued that pairing such simulations with easy-to-verify goals turns each hard biology problem into something like a mathematical conjecture, a statement that is either proven or not, with a clear target. In his view, biology becomes "a search problem" and so "tractable for the first time in history." He framed this as a path still ahead, beginning "when we have digital twins for cells." In the one real example the panel discussed, the computer search ended with humans at the bench, running experiments on an enzyme system whose job no one yet knows.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

Connected ideas and articles

From the conversation

Podcast episodes

Moonshots with Peter Diamandis

Why Jensen and Zuck think the doomers are wrong (plus AI get’s a rebrand) | #294 MOONSHOTS Live

Episode published (recorded )This article draws on 43:54–50:51 (approximate times)

Article history

Updates to this article

Tags

Useful AI agents needn't top benchmarks, Blundin says of Meta's Muse

Investor Dave Blundin says Meta's Muse personal agent shows a split in AI. On one side are consumer agents that handle everyday tasks. On the other are frontier models that compete for the top benchmark scores. Blundin chairs EverQuote, whose shares he said fell 15% on back-to-back days. Investors feared Muse would take over insurance and mortgage shopping, which he thinks won't happen. Emad Mostaque argued that models are now smart enough and the priority is avoiding mistakes. Blundin warned that people who start with today's usually correct models may trust answers that are still sometimes wrong.

5 min read

Aaron Levie sees open-weight tokens and lab revenue growing together

On Training Data, Box CEO Aaron Levie describes how his customers actually pick models: a default for asking questions of their files, and hard-nosed accuracy evaluations for the high-volume extraction work where most tokens are spent. He endorses Decagon founder Jesse Zhang's argument that mature workflows migrate to open-weight models, and explains why the big labs' revenue and open-weight token volume can climb at the same time.

7 min read

Box CEO Aaron Levie's two rules: beat generic agents, then let them in

On Sequoia's Training Data podcast, Box CEO Aaron Levie said any company sitting on customers' data now has two obligations: build an agent measurably better than an off-the-shelf one at its own workflows, and expose the same capabilities to outside assistants like Claude and ChatGPT. He described the tuned search-and-retrieval harness behind Box's agent, the evaluations that track model progress, and his bet that within five years roughly 90% of enterprise tokens will be spent on work nobody asked for directly.

8 min read

After Navier–Stokes, a panel asks where to point 100,000 agents

OpenAI's claimed Millennium Prize result used roughly 10,000 agents on a problem that was, as one entrepreneur on Moonshots put it, unusually easy to specify. The panel's argument: as the price of that kind of compute falls, the scarce skill becomes writing the target — and today's models, asked for ten ideas to cure cancer, produce a bad list.

6 min read