October 9, 2026
Heard in AI

Periodic Labs bets failed experiments can teach AI to do science

Periodic Labs' Liam Fedus says its records of failed experiments, and of the whole process behind each result, give its AI data that published research lacks. He argues even stronger future models will still need physical experiments.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Latent Space, episode published October 8, 2026

A materials experiment can fail in more than one way. Liam Fedus, who built Periodic Labs with Ekin Doğuş Çubuk, described the plainest case on the Latent Space podcast: "we intended to produce some structure. And then all of the evidence points to us not producing that structure." Sometimes, he said, the miss is more interesting. The team set out to make one thing, and "something else emerged from the data," a new structure that hadn't been identified before.

Periodic is an AI research company that is building automated laboratories where AI systems help decide what to make, make it and work out what came out of the furnace. In the episode, published October 8, 2026, Fedus joined hosts swyx and Brandon to discuss why these misses might be some of the most valuable things Periodic produces. They also discussed whether a much smarter future AI model could skip the lab entirely.

The data the literature leaves out

Not every "failure" is an accident. The conversation touched on negative controls, experiments run on purpose so that a particular result should not appear, as a check on everything else. It also covered unwanted byproducts. An impurity phase, meaning a stray crystal form mixed into the sample, can "kill your property" or even be toxic, and Fedus agreed.

The discussion then turned to a basic problem in machine learning. A classification algorithm learns to sort examples into categories, and it cannot learn to say "no" if it has only ever seen "yes." Materials science is a particularly bad case. Researchers usually publish the crystals they managed to make and rarely publish the ones they couldn't. A failure in a paper also doesn't prove that a crystal can never be made. It might only reflect the synthesis method or technology that was used.

When Periodic runs its own experiments, it records negative results together with the context of what was tried. That makes it possible to train a classifier on them. Periodic's website makes a related point, saying the company generates experimental data that includes the unsuccessful outcomes that rarely reach the published literature.

The conversation also noted an irony. In a lab like Periodic's, most individual experiments probably fail, so a cold start could face too many negatives rather than too few. The answer was to label differently. Labeled one experiment at a time, the data is mostly negative. Labeled by campaign, meaning a whole series of attempts aimed at one goal, it could be more balanced, because some campaigns do succeed. The AI's "reasoning traces," its written step-by-step working, could then span entire campaigns.

How a string of failures becomes know-how

Fedus said the path to a success is valuable too. When a run of negative results ends in a positive one after repeated adjustments to the process, that gives what he called "process engineering type data": a record of how the right result was actually reached. That matters, he said, because so much of materials science is ambiguous about how something was made. Even a known material "can be highly non trivial to replicate." Some recipes are in a high school textbook, while others are at the frontier. Periodic is building a system that, given a string of failures, works out how to reach the success.

For Fedus, this is the main point. "This type of data basically doesn't exist anywhere else," he said. The company tries to capture what he called "the full lineage of the scientific process": the conversations, the intuitions, what was run in the lab, which computations were done and what code was written. "Rather than training on the final output of science," he said, "you're training on the process of doing science."

Why more computing power can't fix noisy data

The discussion also covered the quality of the data that does get published. It was said that most experimental results in the literature have such a high noise floor that they are usually less accurate than density functional theory (DFT), the standard way of calculating a material's properties on a computer. When different labs run experiments in different parts of the world at different times, results fluctuate widely. One hope for Periodic's standardized workflows is that they will lower that noise.

The conversation raised the idea that the trade-off between spending on compute (computing power) and spending on data may be shifting toward data. Fedus tied this to the noise problem: "If you throw a huge amount of compute against a bunch of noise, you're not going to have a good thing emerge."

Asked how Periodic uses AI models, Fedus said the company mixes open-source and closed models. In many areas, he said, the big commercial models have too much latency or cost too much. But it isn't only about cost. With data that no one else has, Fedus said, Periodic can in some cases go beyond what frontier models can do even at their highest reasoning settings, and it can be more compute efficient than they are. He called that "really instrumental" to the program.

Can a smarter model skip the lab?

The conversation compared Periodic's approach to recursive self-improvement, the idea of AI systems that help build better AI. Here, a model is trained to run better experiments instead. The loop would close at a higher level if Periodic's physics models led to better chips, which in turn led to better AI.

That raised a harder question. Could a future frontier model, tuned on reasoning traces like Periodic's, simply do the work without a lab? Fedus gave two reasons to doubt it. First, he said, interfacing with the physical world brings different challenges from the noise that already exists in running machine-learning loops. Second, putting knowledge into a model's weights, the internal settings adjusted during training, still matters. A model can also improve by reasoning longer when it answers. But "if inference time reasoning was sufficient," he said, "all of the frontier labs would have stopped training at like GPT-4." Because the physical sciences differ from machine learning, he expects that training this knowledge into Periodic's own weights will lead to different types of systems and capabilities.

The discussion added a more basic objection. Machine learning is good at what it has been trained on, but scientific discovery is, almost by definition, what a model has not been trained on. So even a much better model would still have to run experiments. Periodic is building its labs so that open or closed models alike can "tinker with the universe," because a big discovery is unlikely without trying things. Fedus put it bluntly: no one, he said, is going to "zero shot" a room-temperature superconductor, meaning get it right on the first try without any experiments. A room-temperature superconductor is a material that would carry electricity with no resistance under everyday conditions, and it remains a long-sought goal.

The discussion also raised skepticism about zero-shotting scientific results. Fedus said it is important to distinguish the results seen in math and theoretical physics from the physical world. They are, as the exchange concluded, two very different things.

Share this article

Go to the original

Sources & further reading

  1. 01

Connected ideas and articles

From the conversation

Podcast episodes

Article history

Updates to this article

Tags