Chinese Academician Wang Jian: Let Scientific Data Become the “First-Class Citizens” of Foundation Models

Avatar 0

On July 17th, Wang Jian, an academician of the Chinese Academy of Engineering and director of Zhejiang Lab, proposed during his keynote speech at the 2026 World Artificial Intelligence Conference (WAIC) that artificial intelligence is facing a new paradigm shift, and this time, the driving force is no longer just language data, but scientific data. He stressed that what truly matters in the future isn’t training a bigger language model, but making scientific data the “first-class citizens” of foundation models.

He pointed out that the widely discussed large language models (LLMs) are essentially foundation models built on massive amounts of text data. Wang Jian noted that humanity’s processing of text is nothing new. “Just like ancient Chinese texts and Old Latin originally had no punctuation or spaces, the process of adding punctuation and spaces is fundamentally similar to the core tokenization logic in today’s large language models. LLMs have taken our ability to understand and use text data to an unprecedented level,” he said.

However, a great deal of critical information in scientific research doesn’t exist in text form. If AI’s vision is always limited to analyzing papers and code written by humans, it can never truly touch the essence of science. Wang Jian used geoscience as an example, noting that over 70% of key geoscience information comes from raw physical data like spectra and seismic waves, not from the text of academic papers. “A picture is worth a thousand words, but a spectrum might be worth a million pictures,” he said.

Wang Jian stated that future foundation models should be built directly on scientific data, achieving unified modeling of text, code, and scientific data.

He illustrated this “paradigm shift” with a case from astronomy. In November 2024, a single-author study published in an American astronomy journal discovered nearly 1.5 million previously uncatalogued space objects. The sole author was an 18-year-old high school student, and the data came from a retired satellite that had been monitoring asteroids. Wang Jian pointed out that our understanding of data is nowhere near where it needs to be. Quoting from “The Structure of Scientific Revolutions,” he emphasized that “discovering new problems in old data” is the basic logic of scientific change, just as Galileo used data accumulated by Tycho Brahe rather than collecting his own new observations.

“Scientific breakthroughs don’t just come from new experiments; they can also come from re-examining existing data,” Wang Jian said.

Wang Jian then traced the origin of the “Foundation Model” concept. He noted that professors at Stanford University first explicitly coined the term in 2021, and reiterated that LLMs are essentially text-based foundation models. He used the example of ancient Chinese text lacking punctuation and Old Latin missing spaces, showing that text processing existed long before AI—adding spaces and punctuation is conceptually similar to today’s tokenization. LLMs have brought humanity’s ability to process and use text to an incredible new level.

Returning to science, Wang Jian recalled a conversation from 12 years ago with geoscientists about using AI to help their field. They had three simple requests: get all the scientific data, get all the relevant research papers, and have the infrastructure to do it. He noted that today, AI still falls far short of the first requirement. Most current “AI for Science” work only has AI read the text of scientific papers, and it barely touches the actual scientific data. In geoscience, over 70% of the information isn’t in text; it’s in signals like spectra, seismic waves, and sound waves. “A picture is worth a thousand words, but a spectrum is worth a million pictures,” Wang Jian repeated with emphasis.

Wang Jian introduced that Zhejiang Lab defined the “Scientific Foundation Model” a few years ago. It’s built on scientific data, not on scientific paper text, with the goal of making scientific data a first-class citizen of large models. The project is codenamed “021” (Zero to One). He stressed that this scientific foundation model must place text, code, and scientific data into the same space and the same model to handle all tasks. It’s not a vertical-domain model; it’s an independent technical model.

He gave the example of GeoGPT, a project done in collaboration with the team of Academician Shen Shuzhong from Nanjing University. Geoscientists study rocks, aiming to “make rocks talk” and understand the Earth’s evolution. Earth was born about 3.8 billion years ago, and time scales are measured in millions of years. Scientists want to improve time precision from millions of years down to tens of thousands of years—a leap of several orders of magnitude. Fossils have the “Signor-Lipps Effect”—the last disappearing species doesn’t become a fossil, and fossilization is random. Finding a fossil doesn’t mean that species disappeared at that exact moment, making precise dating extremely difficult.

On July 1st of this year, at the 5th Stratigraphy Congress, the team led by Academician Shen Shuzhong, referencing 100,000 biological species and nearly 20,000 fault data points, used GeoGPT to place data on a unified timeline, constructing the world’s most complete and accurate geological timeline to date. Last week, GeoGPT was included in the “UN Decade of Sustainable Development Recognition List.” Wang Jian pointed out that governance structure is crucial. The model set up a global governance committee over four years ago, which has played a key role in the work’s global impact.

Wang Jian further stated that the scientific foundation model marks a very different turning point for AI. In the future, AI will become as foundational as mathematics, no longer being a standalone field. He concluded by saying that the role mathematics has played in the history of human science will certainly be the role AI will play in scientific research over the next 50 years.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

Log In / Sign Up

Enter your email to receive a secure code. No password needed.