Qiushi Engine tops AI research benchmark, beating Claude Code

Share:
Audio Loading voice…
Qiushi Engine tops AI research benchmark, beating Claude Code

Synopsis

A Zhejiang University-led AI agent called Qiushi Engine has overtaken Anthropic's Claude Code on the ResearchClawBench leaderboard — a benchmark testing full-cycle autonomous scientific research — marking a striking milestone for Chinese academic AI in a domain dominated by Western labs.

Key Takeaways

Qiushi Engine , developed by a Zhejiang University -led team, ranked first overall on the ResearchClawBench leaderboard as of Tuesday, 21 July 2026 .
Open Science Desktop finished second and Anthropic 's Claude Code finished third on the same benchmark.
ResearchClawBench was created by a team led by the Shanghai Artificial Intelligence Laboratory to test whether AI agents can independently replicate or surpass human-authored research conclusions.
Qiushi Engine was officially launched last week and is described by its developers as capable of 'end-to-end autonomous scientific discovery' in real physical environments.
The benchmark result intensifies competition between Chinese academic institutions and Western AI labs — including OpenAI and Anthropic — in the autonomous research agent space.

Zhejiang University's Qiushi Engine has claimed the top position on the ResearchClawBench leaderboard for autonomous scientific research as of Tuesday, 21 July 2026, surpassing Anthropic's Claude Code and a field of leading AI agents in a benchmark designed to test whether systems can genuinely conduct independent scientific inquiry.

What Qiushi Engine Does

Qiushi Engine, officially launched last week by a Zhejiang University-led team, is a large language model-based agent built to perform scientific research in real physical environments. Its developers said the system is capable of 'end-to-end autonomous scientific discovery' — a distinction they drew against existing systems that are reportedly limited to narrower, task-specific operations.

The agent secured the first overall position on ResearchClawBench, with Open Science Desktop finishing second and Claude Code in third.

Why It Matters

ResearchClawBench was created by a team led by the Shanghai Artificial Intelligence Laboratory to evaluate whether AI agents can independently replicate or exceed the conclusions of human-authored reference papers. Unlike conventional capability benchmarks, it targets the full research pipeline — hypothesis, experimentation, and conclusion — making top placement a meaningful signal of autonomous reasoning depth.

The result positions a Chinese academic institution at the frontier of agentic AI, a domain where Anthropic, OpenAI, and other Western labs have been investing heavily. OpenAI's GPT-5.5 and other prominent agents were also part of the competitive field evaluated on the benchmark.

The Competitive Backdrop

The leaderboard outcome arrives amid an intensifying global race to build AI systems capable of replacing or augmenting human researchers. Academic and state-backed Chinese institutions — including the Shanghai Artificial Intelligence Laboratory — have accelerated investment in foundation models and agentic infrastructure, increasingly challenging US-headquartered incumbents on third-party evaluations.

Autonomous research agents are seen as a critical next frontier: if validated at scale, they could compress scientific timelines across drug discovery, materials science, and engineering — sectors where both China and the West have declared strategic priorities.

What's Next

The durability of Qiushi Engine's lead will depend on how rapidly competing labs update their agents; leaderboard rankings in this space have historically been volatile. Observers will also watch whether the benchmark itself — and the Shanghai Artificial Intelligence Laboratory's methodology — gains broader acceptance as a credible evaluation standard outside China.

For Anthropic and other Western developers, the result underscores the urgency of advancing agentic capabilities beyond coding and reasoning into full-cycle scientific discovery.

Point of View

Particularly in domains — autonomous research, reasoning, coding — where Western incumbents have staked reputational claims. What mainstream coverage underweights is that ResearchClawBench itself was built by the Shanghai Artificial Intelligence Laboratory, raising legitimate questions about benchmark design neutrality that independent replication will need to resolve. The deeper story is the commoditisation of agentic AI: as top-of-leaderboard positions rotate rapidly, the moat is shifting from benchmark performance to deployment infrastructure, data access, and regulatory clearance — areas where the competitive dynamics look very different from a public leaderboard.
NationPress
21 Jul 2026

Frequently Asked Questions

What is Qiushi Engine and who built it?
Qiushi Engine is a large language model-based AI agent designed for autonomous scientific research in real physical environments, officially launched last week by a team led by Zhejiang University . Its developers describe it as capable of 'end-to-end autonomous scientific discovery', distinguishing it from systems limited to specific tasks.
What is ResearchClawBench?
ResearchClawBench is an AI benchmark created by a team led by the Shanghai Artificial Intelligence Laboratory that tests whether AI agents can independently carry out scientific research and match or exceed the conclusions of human-authored reference papers. It is designed to assess the full research pipeline rather than narrow, task-specific capabilities.
How did Claude Code perform on ResearchClawBench?
Anthropic 's Claude Code ranked third on the ResearchClawBench leaderboard as of 21 July 2026 , behind Qiushi Engine in first place and Open Science Desktop in second.
Why does this benchmark result matter for the AI industry?
The result signals that Chinese academic institutions are now competitive at the frontier of autonomous agentic AI — a domain attracting heavy investment from Anthropic , OpenAI , and other Western labs. If autonomous research agents prove reliable at scale, they could accelerate scientific timelines in fields like drug discovery and materials science.
What should we watch for next in autonomous AI research agents?
Leaderboard positions in agentic AI have historically shifted quickly as labs update their systems, so Qiushi Engine 's lead may be contested soon. Independent validation of ResearchClawBench 's methodology by institutions outside China will also be a key signal of how credible the ranking is as a global standard.
Nation Press
The Trail

Connected Dots

Tracing the thread behind this story — newest first.

8 Dots
  1. Latest Yesterday
  2. 3 days ago
  3. 4 days ago
  4. 1 week ago
  5. 4 weeks ago
  6. 1 month ago
  7. 2 months ago
  8. 2 months ago
Google Prefer NP
On Google