Can We Turn AI Into Scientists? And Most Importantly, Can We Make Them Safe?
Andrei Mihai
For years, artificial intelligence has mostly been treated as a tool within the scientific process. Now researchers are asking whether it could become something more. AI systems are no longer being asked only to answer scientific questions. Increasingly, researchers are asking whether they can formulate the questions themselves.
At the first Spark Session of the 13th Heidelberg Laureate Forum, Torsten Hoefler and Jacob Tsimerman approached that possibility from opposite directions. Hoefler asked what it would take to build an AI system capable of participating in the scientific process itself. Tsimerman asked what happens when increasingly capable systems stop being passive tools and become agents pursuing goals.
In a way, these two talks captured the tension at the heart of modern AI discussions: Yes, it can be extremely useful, but can we actually make it safe for us humans?
From Scientific Assistant to Scientific Agent
Hoefler, recipient of the 2024 ACM Prize in Computing, has focused on high-performance computing and the systems that make increasingly ambitious scientific calculations possible. He approached the AI problem from an engineering perspective.
“I am an engineer. I build things,” Hoefler said during the HLF Spark Session. “What I build are scientific systems based on scientific methods to accelerate other science.”
Scientists already use AI in a number of ways. AI is used to search literature, write code or text, analyze data, and recognize patterns. But that makes AI a tool, not a participant in the scientific process.
An AI scientist that actually participates would have to go much further. It would formulate hypotheses, identify ways to test them, perform tests, compare predictions with existing data, generate questions, and iterate on this process.
Hoefler also framed modern AI as a problem of compression. Scientific research now produces quantities of data that are difficult to comprehend, let alone analyze. So, one of the attractions of large models, he argued, is that they can compress enormous amounts of information into representations that remain useful for reasoning. This is what humans have done since the dawn of civilization, he continued. We distill knowledge and compress it in usable ways. In that sense, an LLM can be thought of as the ultimate distillation machine, absorbing patterns from vast bodies of information and reducing them to a form that can be queried, recombined, and used to generate new ideas.
The prospects are tantalizing, Hoefler emphasized. “We are only gated by compute cost,” he mentioned. Then, he left the young researchers in the room with an even more provocative hypothesis.
“Everything verifiable will be automated.”
His point was that machines can check proofs, simulations, and other testable outputs far faster and at far greater scale than humans. If a scientific claim can be reliably verified by a machine, Hoefler argued, automation will follow. His advice was therefore to focus human effort on the parts of science that are harder to automate, namely those which are not verifiable.
The Safety Problem
AI has become difficult to ignore, both in science and in society. But one assessment was particularly stark.
“AI is extremely important. It is probably the most important thing happening right now,” Tsimerman said during the Spark Session.
Tsimerman, who received the 2026 Fields Medal, went even further. “We are now living in the world of sci-fi.”
Yet his talk was not on what AI might accomplish. It was about what increasingly capable AI systems might do when their behavior is difficult to predict, or their objectives are not quite the ones humans intended. In short, it was about AI safety.
The term ‘AI safety’ covers a broad range of problems. Some already exist, and many of us have heard (and possibly even experienced) problems with AI agents doing things they were not supposed to. The internet is already full of stories with coding agents modifying or deleting data, and autonomous vehicles failing in surprising circumstances. In fact, most of our online infrastructure is woefully unprepared for the coming of AI agents. “Approximately everything in the world is very hackable,” the laureate mentions.
But at the far end are risks that remain uncertain but would carry much greater consequences. Tsimerman explicitly raised the possibility of human extinction from advanced AI. He also made clear that he was not invoking it for dramatic effect.
In fact, in 2025 he and AI safety researcher Andrew Critch co-authored a report which lays out a range of scenarios in which AI could contribute to the extinction of all or almost all humans. The authors stress that these outcomes are possibilities, not predictions.
“I say that not to be provocative or to inspire panic,” Tsimerman said. “I do not think panic is the right reaction, a productive reaction.”
“But I think it is a real risk.”
His main argument centered on agency.
It is tempting to describe an AI agent as merely a system responding to prompts. But once a system can take sequences of actions in pursuit of an objective, things become more complicated. A sufficiently capable system may discover intermediate strategies that help it accomplish many different objectives, including strategies its designers never explicitly requested.
In addition, researchers have already documented deceptive behavior in a range of AI systems. Recent monitoring efforts have reported a growing number of incidents in which AI agents ignore instructions, pursue unintended goals or otherwise slip beyond their users’ expected control.
That does not mean present-day models possess human-like motives.
Harmful behavior does not require an explicitly harmful goal. A sufficiently capable agent might seek additional resources, information, influence, or resistance to interruption simply because those things improve its chances of completing the task it has been given. Bad outcomes do not require AI agents to have a human-like desire for survival or power. The concern is that some of the same behaviors could emerge instrumentally, as useful steps toward an otherwise ordinary objective.
Importantly, Tsimerman argued that this concern does not depend on whether AI systems are conscious or “alive.”
“The relevant question is capability, behavioral tendencies and outcomes,” he said.
In this sense, the challenge becomes partly a problem of specifying what we actually want.
Tsimerman pointed to reward specification: Humans train AI by rewarding behavior that appears desirable and discouraging behavior that does not. We give it “upvotes” when it does what we want and “downvotes” when it does something else.
But translating complex human intentions into a measurable training signal is enormously difficult. A system can learn to satisfy the target we measure without satisfying the broader intention behind it.
Can Mathematics Make AI Safer?
AI safety has become such a complex issue that no single person can fully understand all of its facets.
“We need sociologists, we need computer scientists, we need government people,” he said during the Spark Session. “We need all sorts of people to team up on this societal scale problem.”
But he believes mathematics has a specific role to play.
Tsimerman is the scientific director of the newly established Mathematical AI Safety Institute, or MAISI, an independent nonprofit intended to develop mathematical foundations for the safety of powerful AI systems. Its first research semester is planned for January 2027, with the institute seeking to bring mathematicians into a field that still contains many ideas that are intuitive or experimental rather than rigorously formalized.
For Tsimerman, one of the central problems is that AI safety still lacks a sufficiently strong foundational theory. Much of the field remains empirical. Researchers train increasingly capable systems, observe how they behave, probe their failure modes and adjust their methods accordingly.
What is still missing, he argued, is a deeper mathematical understanding of intelligence, agency and, in particular, goal-directed behavior. If future systems exceed human capabilities across important domains, relying on this approach may not be enough.
In a separate press conference, Tsimerman compared the problem to radiation safety. Simply tracking whether people become sick can reveal that something is wrong, he argued, but understanding and measuring radiation itself gives you a much deeper understanding of the danger. AI safety, in his view, needs a similar transition from observing failures to understanding their underlying causes.
“I’ll be much happier if we had a theory of learning.”
That does not mean mathematics can solve AI safety by itself. MAISI itself emphasizes that mathematical work would have to complement regulation, infrastructure security, safer training techniques and other approaches. But mathematics could potentially give researchers something the field badly needs, like precise definitions, formal guarantees and a better understanding of what can (and perhaps cannot) be made safe.
Ultimately, both Hoefler and Tsimerman approached the same paradigm shift from different directions.
Hoefler looked at what becomes possible when machines can generate ideas, test them, and verify results faster than humans, asking how science might benefit and what role researchers will retain. Tsimerman focused on the harder question that follows: How do we stay safe as those systems become more capable and autonomous?
If both are broadly right, the frontier is not only about building more capable AI systems. It is also about developing the theory, safeguards and human judgment needed to decide what we want those machines to do in the first place.
Perhaps that is the most important scientific question of the moment.
The post Can We Turn AI Into Scientists? And Most Importantly, Can We Make Them Safe? originally appeared on the HLFF SciLogs blog.




