Some assembly required, with machine learning

WashU scientists are training AI systems to “run” experiments to check viability of predictions

Leah Shaffer 
WashU scientists take machine learning to the next step in expanding its utility for designing new materials and chemical compounds. Training an AI tool to model out synthesis of compounds could vastly improve speed of discovery. (Image: Shutterstock)
WashU scientists take machine learning to the next step in expanding its utility for designing new materials and chemical compounds. Training an AI tool to model out synthesis of compounds could vastly improve speed of discovery. (Image: Shutterstock)

Machine learning doesn’t replace human intelligence, but it can outlast human endurance, which makes it a very helpful tool for chemistry and material discovery. Scientists know the machine learning models can make predictions based on the vast reams of data it’s trained on, but can it take it a step further and massively scale up testing out those predictions?

“We already see that AI is powerful in terms of predicting new structures,” said Zhiling Zheng, assistant professor of chemistry in the Arts & Sciences at Washington University in St. Louis.  “But for most of the bench chemists or material scientists, we actually are more interested in making the materials themselves.”

In a recent prize-winning essay for the journal Science, Zheng proposes how AI systems tackle that next step: “At the heart of this platform is the AI’s ability to read chemistry like a chemist,” wrote Zheng in the essay.

Christopher Cooper, assistant professor of energy, environmental & chemical engineering in the McKelvey School of Engineering, is on the same page. He recently published a paper in the journal Matter documenting how to curate troves of data for polymer synthesis.

Data curation is the first major step of this work. The data needs to be collected and converted to a form that the machine learning models can easily digest.

The machines have been fed a full diet of the known rules of chemistry. That’s how they make their predictions. What’s missing is the application of those rules, following through to run simulations on making those molecules. For that part, it needs the “recipes” of chemical synthesis, the instructions buried in journals, textbooks and footnotes over the century, and that is what must be collected, translated and “fed” to the machine.

“You want a model to be able to understand those instructions, mash them together and say, ‘this is higher likelihood of being successful,’” Cooper said.

How to create a self-driving lab

Zheng comes from a background in studying metal organic frameworks (MOFs), which are built by metal ions as nodes and organic blocks as linkers. It is shaped like a cube with metallic elements on each corner, lending itself to endless tinkering and potential uses, a LEGO set for chemical synthesis. The problem is it is too open ended.

“There are just so many different possibilities,” Zheng said. The different ways to design metal organic frameworks could number in the millions of variations — no human could take the time to run those experiments to see if they work.

A self-driving lab, or autonomous lab, is an idea has been in academia for decades, Zheng noted. But now he is proposing a new methodology that was otherwise impossible before the advent of large language models, or LLMs. Researchers can now train AI much like they would train a graduate student: Give them a pile of instructions and let them practice and learn from mistakes as they go. To test out his idea, Zheng and his team trained LLMs on a literature-based dataset of some 4,000 linker transformations of MOFs. Think of it as 4,000 different LEGO sets to build. To that, they train an AI agent to further filter that 4,000 down based on chemical constraints. Those design agents, using computational simulations, found 10 new viable materials that demonstrate stronger water harvesting performances than state-of-the-art aluminum-based adsorbents. These discoveries were achieved not through the usual “brute-force” linker screening, a labor-intensive process, but through “targeted design guided by model suggestions, learned from literature-based community knowledge,” he wrote.

A similar project is ongoing with Cooper and polymer design, but the first step is to collect that community knowledge.

How to feed a self-driving lab

The actual work of translating notes and journals into something the AI can effectively “read” is a technical challenge unto itself, but Cooper has a plan for how to approach it.

In research published this year in the journal Matter, Cooper and co-author Kathryn Miller at the National Institute of Standards and Technology, share how they created an “automated approach to curating a materials science library.” Gather all the data on constraints and practicalities, what polymers were used, what dynamic bonds were used, what properties and applications were discussed, tag that for the machine to read, then ask the machine to keep filtering it down from there. Run the test, find the flaws in the execution, fix the flaws, run the test again and repeat. All to find the most viable candidates for dynamic polymers that might otherwise been lost in a sea of options.

They focused their approach on this emerging field of dynamic polymers. This class of materials has huge potential for application thanks to its self-healing properties, its ability to respond to stimuli, its 3D printability, underwater adhesion features, degradability and recyclability.

The result was the Dynamic Polymer Annotated Library (DPAL) which they successfully ran through a couple case studies.  

Moreover, with funding from the NSF, Cooper is also working on new ways to represent these polymers in a computationally efficient way for machine learning applications.

Polymers are very disperse, he said, in the same way there are so many variations with the MOFs, there could be millions of different molecules present in average polymer sample. To describe that pile, they use sets of probability distributions, essentially a way for the machine to quickly read and understand the potential likelihood of success for a predicted design. 

“Breaking into those probabilities gives us ways to physically represent a system and can accelerate model prediction and accuracy,” Cooper said.

At the end of the day all these machine learning tools vastly increase the efficiency of those human scientists like Cooper and Zheng. Instead of toiling through trial and error, let the machine run through laborious processes, leaving the human to do the final checks and connect the dots.

“You can be a manager rather than working in the lab for 10 hours,” Zheng said.


Zheng Z. Reprogramming synthesis. Science 393,253-253(2026). DOI: https://doi.org/10.1126/science.aeh4807

Miller K, Cooper C., The Dynamic Polymer Annotated Library: An automated approach to curating materials science literature. Matter. Vol. 9, Issue 3, 2026, 102624, DOI: https://doi.org/10.1016/j.matt.2025.102624.

Click on the topics below for more stories in those areas

Back to News