Beijing surgeon solves 20-year-old Crouzeix conjecture using GPT-5.6-Sol
Synopsis
Key Takeaways
Jin Shanmu, a postdoctoral researcher and resident neurosurgeon at Peking Union Medical College Hospital in Beijing, has reportedly proved Crouzeix's conjecture — a two-decade-old open problem in numerical linear algebra — using a 16-hour autonomous run on OpenAI's GPT-5.6-Sol model via the ChatGPT Work platform. The breakthrough, announced in August 2026, stunned the global mathematics community and reignited debate over AI's role in frontier research.
The conjecture, explained
First proposed by French mathematician Michel Crouzeix in 2004, Crouzeix's conjecture holds that the norm of applying any function to a matrix is no larger than twice the function's maximum value on that matrix's numerical range. Despite its abstract framing, the problem sits at the intersection of matrix analysis and functional analysis, and had resisted proof by specialists worldwide for over 20 years.
Jin, a self-taught mathematics enthusiast, stumbled into the field while conducting research on transcranial ultrasounds — a tool used in brain surgery. The conjecture's connection to signal processing in ultrasound analysis drew his attention before the scope of the problem became clear.
Why it matters
The proof, if independently verified, would mark one of the most significant AI-assisted breakthroughs in pure mathematics to date. It places GPT-5.6-Sol alongside a short list of AI systems — including Google DeepMind's AlphaProof — credited with advancing formal mathematical reasoning beyond the reach of conventional tools.
The result also underscores a widening pattern: domain experts outside mathematics, armed with frontier AI models, are increasingly capable of engaging with problems that once required years of specialised training. Jin's medical background, rather than being a disadvantage, appears to have provided the applied framing that motivated the research direction.
The competitive backdrop
OpenAI's GPT-5.6-Sol, deployed via the ChatGPT Work platform, was the instrument of record here — a notable data point as Anthropic's Claude and other frontier models compete aggressively on reasoning benchmarks. The autonomous, extended-session format — 16 hours of continuous operation — points to agentic AI workflows as a meaningful new surface for scientific discovery, distinct from single-prompt query interactions.
Researchers at institutions including Cornell University, where mathematician Alex Townsend has worked on related problems in numerical analysis, are expected to scrutinise the proof closely. The mathematics community typically requires independent verification before a conjecture is formally considered solved.
What's next
The immediate question is peer verification: whether the proof holds under formal mathematical scrutiny will determine whether this becomes a landmark moment in both AI capability and pure mathematics. If confirmed, it will intensify pressure on research institutions to integrate long-horizon agentic AI tools into their workflows — and raise fresh questions about attribution, authorship, and the future of human-led mathematical discovery. Observers will also watch whether this accelerates attempts to apply similar AI-assisted approaches to harder open problems, including the Riemann hypothesis.