OpenAI has discovered the latest Fields Medal winner, who is ByteDance’s newly announced scientist planning to target?
This year’s Fields Medal is considered the “last pure blood” award – perhaps future breakthrough research will be achieved through the combination of AI.
The process of this transformation may be faster than we imagine: at the press conference of the Fields Medal, someone asked one of the winners, Canadian mathematician Jacob Tsimerman, what his next job would be, and he said he had started to turn to AI security research and would soon be working in OpenAI’s security department.
On the other hand, OpenAI’s Chief Research Officer Mark Chen immediately welcomed it.
Scholars at the forefront have regarded top AI companies as platforms to showcase their aspirations. Recently, more and more top scientists, including mathematicians, physicists, biologists, etc., are flowing towards top model companies. In the first half of this year alone, Nobel laureate in chemistry John Jumper, economist Chad Jones, philosopher Harvey Lederman, and others have joined Anthropic.
Behind this, top AI companies have also turned their attention to the same direction: they have invited scientists to the site of AI research and development. The deep integration of AI and cutting-edge science is transforming from optional to mandatory.
Domestic AI giants are also following suit. On Thursday of this week, Byte’s Seed Edge team released the Seed STEM Scientist Program, opening up 100 collaborative seats for scientists and doctoral students in cutting-edge scientific fields such as mathematics, physics, chemistry, and biology, inviting them to explore unknown problems in basic science with Seed. From publicly available information, this is the first large-scale plan in China to invest in AI to accelerate cutting-edge scientific discoveries.
In the second half of the big model, the competition is no longer about running points
In the current fast-paced competition in the field of AI, focusing on long-term basic AI research may not be as intuitive as polishing code capabilities and improving evaluation scores. The latter’s ability enhancement can quickly translate into product experience and commercial revenue, while the feedback cycle for basic research and AI accelerated scientific discovery is much longer and the exploration path is full of uncertainty.
But looking at it the other way around, once a real breakthrough is made on this path, the paradigm level value it brings will far exceed ordinary product iterations.
Yao Shunyu previously mentioned a viewpoint to the team: there is no magic in training large models. The real challenge lies in doing the most fundamental and certain things that can be done right. This judgment also echoes Anthropic CEO Dario’s long-standing view that the core elements driving AI progress have always been computing power, data quality and scale, training time, and scalable objective functions. “All smart methods and techniques are actually not that important.
Looking back at true SOTA models like Claude and Seedance, there are no shortcuts behind them. The core has been anchored early on, and a lot of foundational work has been done on data quality and underlying architecture.
Anthropic focuses its resources highly on improving the code capabilities of its models. Through reinforcement learning training, the models can autonomously understand and generate high-quality code through repeated experimentation, and form a self accelerating evolutionary flywheel based on user feedback.
In the field of video generation, which many large companies find difficult to sustain, Seedance has transformed AI video generation from a “card pulling” experiment into a more controllable and creative industrial production tool through innovative model architecture and precise data processing.
It can be seen that what truly determines the upper limit of a model’s ability is never the accumulation of skills, but the consolidation of data, the stability of infrastructure, and the long-term goals, while maintaining strategic determination and sufficient patience.
This is not an easy task in today’s environment, but we can see that more and more companies are willing to settle down and take such a difficult but correct path on the domestic big model track.
This year, DeepSeek’s mHC (manifold constrained hyperconnectivity) architecture is considered a breakthrough in the field of large model infrastructure. It is also one of the important architecture upgrades of DeepSeek V4, replacing traditional residual connections and running through the entire model design. It is worth mentioning that the mHC research refers to the Hyper Connections (HC) proposed by the ByteDance Seed team.
Prior to this, residual connections were widely regarded as the default configuration in the industry as the basic skeleton of Transformers. HC was the first to break this single stream paradigm, significantly improving the model’s feature representation ability without increasing computational complexity. From the pioneering proposal of HC to the engineering implementation of mHC, it has become a microcosm of the technological innovation of Chinese AI teams in the field of underlying architecture, which is not commonly seen in the development of large-scale model technology in the past.
At the beginning of last year, ByteSeed established a team called Edge, hoping to conduct research on topics with lower certainty and uncertain results in the short term. In order to provide researchers with a stable environment, they even adjusted the performance evaluation cycle – there was no evaluation mechanism in between, and unified evaluations were conducted once results were available.
Since its establishment, this team has been relatively low-key. Recently, I saw that they released the ultra long distance evaluation set EdgeBench on X, which for the first time in the industry systematically defined the learning rules of agents in real environments and found a new Scaling law for environmental learning dimensions. This has also sparked discussions among many industry experts.
It is said that OpenAI is also integrating this evaluation set to test the long-range capability of its own model.
In the field of AI, people often focus on the big models themselves. Building benchmarks actually requires a long cycle and a large amount of manpower. Designing and open sourcing benchmark tests for advanced AI capabilities is not only to measure their own progress, but also to guide the entire field towards a more transparent and reliable direction by establishing industry standards and exposing model weaknesses. Top AI laboratories invest heavily in building benchmarks.
OpenAI has open-source the Evals framework, becoming a standardized tool for evaluating large language models and LLM based systems in the industry. OpenAI has also designed a series of dedicated benchmarks for specific cutting-edge capabilities, including SimpleQA, BrowseComp, MLE bench, etc.
Anthropic has done a lot of work around SWE bench, and they have identified Claude’s code proficiency gaps; They also released Terminal Punch to measure the ability of agents to complete long link tasks in real command-line environments.
The relationship between these evaluation works and the strengths of various models is not accidental.
It can be seen that the judgment of cutting-edge AI teams on the development path of AI has undergone a transformation. The evolution of the ability of large models has moved out of the pre training dividend period that relies solely on parameters and data stacking. The core growth momentum of the next stage will inevitably come from the continuous interaction, feedback, and autonomous evolution of models in real open environments.
The forefront of basic science is the highest form of open unknown scenarios. The Seed STEM Scientists program just launched by ByteDance is aimed at identifying and breaking through research level problems with scientists, exploring the possibility of AI accelerating scientific discoveries. Scientists provide not only questions, but also problem awareness, judgment criteria, and feedback methods.
Why Top AI Companies
Are they all competing for field scientists?
Why collaborate with scientists? The answer is that the remaining ‘questions’ are no longer difficult enough for a large model.
Nowadays, the mainstream model evaluation system mostly revolves around closed tasks with standard answers. The problem has a definite solution, the evaluation has fixed standards, and the path to improving model ability is clear. But real scientific research follows a completely different logic: it lacks a clear path to solving problems and requires researchers to independently propose hypotheses, design validation plans, handle abnormal and noisy data, and derive underlying mechanisms. The entire process is full of trial and error, repetition, and uncertainty.
The tasks in the real world are the ultimate standard for testing the capabilities of AI models. After pre training scaling and testing scaling, the industry is generally looking for the next direction of scaling, and “continuous interactive learning in real open environments” is one of the most promising paths.
Therefore, when developing cutting-edge models, the world’s top AI teams no longer rely solely on AI teams to polish models behind closed doors, but instead introduce top researchers from various disciplines, making real domain problems the driving force for the evolution of model capabilities.
In North America, many university researchers have recently joined large model companies. In early July, Jelani Nelson, the head of the EECS department at Berkeley, announced an academic leave and joined Anthropic. Since the beginning of this year, at least 22 professors and researchers have temporarily left or reduced their work at universities such as Stanford, Berkeley, and Harvard, and switched to OpenAI, Anthropic, DeepMind, and Meta. These four companies currently have at least 80 current or former university professors gathered.
Behind the gradual addition of big cows, there are also systematic projects.
At OpenAI, the Residency program has become a regular channel for recruiting high-end talents every year. It is specifically aimed at talents who are not currently primarily focused on AI research, allowing participants to join OpenAI’s cutting-edge application and research teams from day one and engage in practical cooperation. This project is equivalent to an “accelerated version of doctoral training”, which can help cross disciplinary experts quickly complete disciplinary integration. In several milestone projects such as ImageGPT and Codex, resident researchers have been deeply involved. The breakthrough in scientific research ability of GPT series models in mathematical reasoning and biological mechanism deduction also relies on the continuous input of interdisciplinary talents.
At Anthropic, the STEM Fellows Program, specifically designed for experts in the fields of science, technology, engineering, and mathematics, will undergo systematic and large-scale expansion in 2026. The project lasts for several months and specifically invites top experts in non AI fields such as physicists, biologists, mathematicians, economists, etc. to join Anthropic internally. They can bring their own unsolved problems or complex workflows in their field and directly use the yet to be publicly released Claude model and internal evaluation tools.
In such projects, material scientists will build a dedicated evaluation process for Claude’s phase stability inference, while climate scientists will integrate professional atmospheric modeling tools with the model in depth. During the residency period, over 80% of the researchers produced high-quality academic achievements, and the feedback from real disciplines directly became training signals for model iteration, driving Claude’s evolution in scientific reasoning and complex tool invocation.
The value of such projects has long exceeded the scientific research itself, and top AI companies have placed equal importance on competing for field scientists and computing power.
Of course, for the scientific community, collaboration is equally valuable. Big models are becoming a new research infrastructure: from formula derivation and numerical simulation to the analysis of massive literature and experimental data, AI is reshaping the paradigm of scientific research, making it possible to advance topics that were previously limited by manpower and computing power.
Being invited to join the Big Model team for on-site research is as significant as astronomers obtaining exclusive access to the next generation of super space telescopes. This type of project can bring many advantages to scientists, not only obtaining a large amount of AI computing power and cutting-edge model experience, but also utilizing AI as an “external brain” to shift energy to defining the problem itself.
For these scientists flocking to Anthropic and OpenAI, such deep collaboration is both an amplifier of research efficiency and a lever to leverage new discoveries. At this moment, when we turn our focus back to China, we find that the gap in the intersection of AI and basic scientific research is being filled by ByteSeed.
AI endgame
We are fighting for long termism
When Fields Medal winners decide to join OpenAI, and as more and more STEM scientists begin to come together with model teams, the next stage of competition for AI is shifting from short-term product iterations, ranking runs, and single point task breakthroughs to longer-term, more complex, and cutting-edge fundamental propositions.
Scientific research is crucial because it stands at the forefront of human knowledge. If we can establish a foothold in such high complexity open environments, it means that the next generation of AI truly has the ability to participate in real complex tasks.
This path is difficult to deliver impressive results quickly, and it is also difficult to form a clear business cycle in the short term. It is slow, heavy, and full of uncertainty. But it is precisely because of this that it can best test whether a team is willing to truly immerse themselves in the foundation and cutting-edge.
The ultimate answer that big models need to submit is whether they can work together with humans to advance those unsolved problems.
The progress of science requires long termism, and the implementation of AGI is no exception.