October 9, 2026
Heard in AI

How Periodic Labs says it trains AI against a noisy physical lab

Liam Fedus of Periodic Labs says rewarding AI for finding a superconductor is too slow and noisy. The startup instead rewards it for identifying what X-ray data show a lab actually made, though ambiguity remains.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Latent Space, episode published October 8, 2026

Rewarding an AI model for discovering a room-temperature superconductor sounds like the obvious goal for a company that wants to find new materials. Liam Fedus of Periodic Labs called that approach "completely infeasible." On an episode of the Latent Space podcast published October 8, 2026, Fedus and his colleague Ekin Doğuş Çubuk explained what the startup does instead. Periodic is building AI systems, simulations and automated labs to discover new materials. Its AI learns from smaller, checkable steps inside the experimental process, using data that comes out of its own labs.

The method behind this is reinforcement learning (RL). A model attempts a task, receives a reward when it does well, and the strategies that earned rewards are strengthened. In math and coding, where much recent AI progress has come, an answer can be checked precisely, and a developer can run as many attempts, called rollouts, as computing power allows. Fedus and Çubuk explained why a materials lab offers neither.

A lab is not a math problem

Fedus said Periodic's RL environments and data come from its physical labs, which he called "our ultimate truth." Optimizing against answers already known from papers or textbooks is not enough, he said, because the company is trying to go beyond them. Math offers high precision. A lab offers variance and aberrant measurements. A sample does not come out of the furnace labeled, and even labeling it can be noisy. Different machines can disagree. Telemetry, the logged record of what an instrument did, can be incomplete: a run thought to be at one temperature may have been somewhat off.

Much of the standard toolkit still applies, Fedus said: midtraining, reinforcement learning, tool-using agents and variance reduction. But the system also has to work accurately under uncertainty and make efficient use of limited data. Unlike in a digital environment, Periodic cannot simply add more environments or rollouts.

Çubuk said every scientific experiment requires some reduction in dimensions. In math or code, all the needed context can sit in front of a person or a model. Physics starts with more atoms than any computer could store, so scientists compress them into a few descriptors. He pointed to thermodynamics, where five or six variables turned out to explain steam engines well, and called it "the beauty of physics."

He listed the practical problems. A furnace degrades because material from the samples evaporates and coats the heating element, so it gets worse with every use. Its temperature is not perfectly uniform, so a sample's position changes the result. Optical instruments are sensitive to vibration. During his PhD, he recalled, one lab traced erratic results to someone walking upstairs at a certain time of night, which shook a laser setup. Çubuk argued that most real-world tasks that require intelligence look more like science than math, with noise and missing context. Fedus added that the best reasoning strategies for math or theoretical computer science may not be the best ones for science.

Why the end goal cannot be the reward

Fedus described Periodic's discovery loop in three steps. First, decide what to make: atoms that will hold together and are expected to have the property of interest. Second, work out the processing conditions that will actually synthesize it. Third, characterize the product to find out what was made.

Rewarding only the final outcome would mean starting an experiment, waiting a couple of days, and updating the model based on whether it had found a superconductor. Fedus said that fails on several counts. There are not enough agents to reduce the variance, rollouts stall while samples move between instruments, and furnace time has physical limits. "It's just too slow, too noisy," he said. Instead, Periodic builds agents around the data its labs have already produced.

Periodic's September 15, 2026 announcement makes the same point. It says running more simultaneous experiments requires more equipment, power and engineering, and that single experiments can take days. The company trains on accumulated experimental data rather than keeping graphics processors (GPUs) idle while each experiment finishes. A companion infrastructure post says scientific agent rollouts can last hours, so the company runs training and inference on separate GPU allocations that do not wait for each other.

Reading the X-ray fingerprint

Characterization, the third step, gives the cleanest reward. Periodic shoots X-rays at a sample. Their wavelength is comparable to the spacing between atoms, so the rays scatter into a diffraction pattern that Fedus described as a fingerprint of the crystal structure. The model is rewarded for correctly identifying the phases present, meaning the distinct crystal arrangements in the sample. It is penalized for naming spurious phases or ones with no real chemical plausibility. Fedus called this a small, self-contained task, and said similar environments can be built for each piece of the loop and stitched together into the full discovery process.

Identifying phases is difficult, and Fedus said doing it automatically also removes a bottleneck. A lab running many experiments quickly gets stuck on understanding what it has produced.

The conversation then turned to what an experiment actually yields. To make a material, the lab mixes precursors (starting crystals) and heats them until the atoms react and form a new crystal, which changes the X-ray diffraction (XRD) pattern. In a real campaign, the first attempt usually fails, and the result is not a clean yes or no. The sample is typically a mixture: leftover precursors, some amorphous (non-crystalline) material, and phases nobody intended to make. Fedus added that some phases have never been recorded in any paper or database. The AI then has to use its tools to work out which arrangement of atoms explains the pattern. That matters for steering the work. If the target phase makes up 1% of a sample, the team needs a reliable measurement to know whether changes are raising its purity.

Periodic has published its own measure of this skill. In a September 15, 2026 research report, the company describes FrontierXRD as 134 difficult diffraction measurements from its labs, with an average of five phases per accepted solution. Periodic reports that its model Neon succeeded on 55.3% of them, compared with 2.7% for the model it started from. A panel of AI judges did the scoring, and the company says it calibrated them against PhD experts. Periodic reports that human experts agreed with each other 77.2% of the time on calibration patterns. These are the company's own figures.

Ambiguity does not go away

Asked whether choosing phases as the target makes the signal unambiguous, Fedus said no. "A very dumb thing is like replicates," he said: the lab repeats experiments. "Two different phases can actually be consistent with the same pattern," he said, so the system needs "chemical intuition" to rule out explanations that are unlikely given the synthesis conditions and prior knowledge.

The discussion named thermodynamics as the biggest prior, followed by physics, and then other instruments. Phases that look alike in XRD can differ in electrical or magnetic properties or in their shape under an electron microscope. Asking a human to analyze ten kinds of measurements consistently at once is hard, the discussion went, but not for an AI. Thermo Fisher's explanation of scanning electron microscopy shows why this helps: different detector signals reveal surface shape, composition-related contrast and elements present. Fedus said the system combines signal across instruments, across experiments over time and across replicates, "rather than just over indexing on one measurement from one instrument."

The hosts asked why biologists can reconstruct a protein from its crystal data while materials scientists cannot do the same with XRD. The answer was that the pattern shows average behavior. Atoms in a crystal may shift slightly to one side or the other while sitting in the middle on average, and the pattern hides that. Fedus said moving toward single-crystal XRD can help. A 2008 review of powder diffraction by David and Shankland supports that distinction. Powder diffraction compresses three-dimensional information into a one-dimensional pattern, so signals overlap. Single-crystal measurements keep more of them separate.

Freezing the world at a date

Fedus described a second kind of environment. As Periodic gathers more experimental data, it can timestamp the state of the world: here is all the evidence the lab had up to a given date. The model is then asked what the scientist chose to do next, or what the next experiment produced.

This guards against a known weakness. If a pretrained model has already memorized a published result, an RL task built on that result lets it "fake work," Fedus said. It reaches the right answer without doing the physics reasoning or the simulations, and training then reinforces those shortcuts. "Those reasoning strategies will not generalize to novel systems," he said. Because Periodic's lab records are new and keep the full lineage of each experiment, Fedus said the company can build environments that are not possible elsewhere.

One basket of data, many environments

Asked how the lab could scale, Fedus said Periodic does not simply run an experiment and wait for an agent to roll out all the way to the final result. Its campaigns produce data across many instruments. From that data the company can build RL environments by the date of an experimental or computational campaign, or from a subset of instruments: one instrument alone, or several combined to reach a reward. He acknowledged that this does not speed up the experiments themselves. From a machine-learning perspective, though, he said a "finite basket of data" viewed in these different ways can be expanded for training.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

Connected ideas and articles

From the conversation

Podcast episodes

Latent Space

Synthesis Superintelligence: from Semiconductors to Superconductors — Periodic Labs’ Liam Fedus and Ekin Dogus Cubuk

Episode published This article draws on 2:49–13:27, 19:44–25:12, 37:37–38:16, 38:47–39:32, 40:38–41:28 and 1:05:54–1:07:55 (approximate times)

Article history

Updates to this article

Tags