AI research
Can AI actually do maths now? Explained
AI solved problems that stumped people for decades. Here's what that really means, simply.
The answer
Yes — in 2026 AI solved real, decades-old maths problems, checked by special software.
You may have seen headlines that AI 'solved famous maths problems' in 2026. Is it real? Mostly yes — and here's what it means, in plain English, with the hype dialled out.
Why these problems make a good test
The maths problems AI tackled are called Erdős problems, named after the mathematician Paul Erdős who posed hundreds of them over his career. What makes them a good test for AI? Three things. They're easy to read — you can understand the question without a maths degree. They're hard to solve — clever people couldn't crack them for decades. And there's no answer to look up, because nobody had solved them yet. That last point is key: it rules out the possibility the AI just retrieved a memorised solution. If it produced an answer, it had to work it out for real.
In May 2026, two big AI labs — OpenAI and Google DeepMind — both reported cracking Erdős problems, within days of each other. OpenAI's model disproved an idea about geometry that had been open since 1946. Google DeepMind's system, AlphaProof Nexus, solved nine open Erdős problems — two of which had been unsolved for 56 years — plus 44 further mathematical conjectures. Two separate companies, two separate approaches, one week apart. That's why it felt like a moment.
How we can trust the answers
This is the clever part — and it's the bit the headlines usually skip. AI can sound completely confident and still be wrong. (If you've ever seen a chatbot confidently make up a fact, you know the feeling.) So how do you know these maths answers are actually right?
Google's approach pairs the AI with a proof checker called Lean — software that flat-out refuses to accept a proof unless every single logical step is airtight. Think of it like an examiner who can't be fooled, can't get tired, and won't give partial credit for nearly-right. The AI suggests the proof; Lean acts as the unbribable gatekeeper. That's the difference between 'an AI said so' and 'the maths genuinely checks out'.
OpenAI's result was checked a different way — by real mathematicians, including a very well-respected one named Timothy Gowers. That's genuine evidence, but it's a different level of confidence from machine-checking:
| OpenAI | Google DeepMind | |
|---|---|---|
| Checked by | Human mathematicians | Lean proof checker (software) |
| Can be fooled? | In theory, yes — experts sometimes miss things | No — the machine either accepts or rejects |
| Status | Formal review still ongoing | Each proof is formally certified |
Neither is fake. But when you see the word 'solved' in a headline, it's worth asking which kind of guarantee you're looking at.
AlphaProof Nexus addresses AI hallucination by pairing an AI model's generative capacity with formal proof-checking through the Lean proof assistant. The AI proposes a proof, and a separate verification system checks every logical step.
Does this mean AI is smarter than people?
No — and the experts are the first to say so. The AI solved specific, hard problems with human guidance and strict checking software. The people who built it describe it as a powerful assistant for particular kinds of maths, not a replacement for mathematicians. Google DeepMind's own boss said the system is 'still not AGI' — that's the phrase for 'generally human-level intelligence'. Think of it as an extremely capable specialist, not an all-rounder.
Hassabis moved quickly to temper expectations, saying the system is 'still not AGI' even as it points toward a more practical role for AI in verified mathematical research.
Why it might matter to you down the line
Here's the quietly exciting part. The same 'AI suggests, software checks' trick that worked in maths could, in time, make AI much more reliable in other areas where answers can be verified — think software that's automatically tested for bugs, or scientific models with checkable outputs. Fewer confident mistakes, more results you can actually rely on. That's a long way from being achieved across the board, but the plumbing for it now demonstrably exists. For now, the honest summary is this: AI just became a genuinely useful helper for some of the hardest maths around. The answers are carefully checked. The people who built it are the first to say it's a tool, not a miracle. Impressive, and appropriately modest — a rare combination.
Frequently asked questions
Has AI genuinely contributed new maths?
What does 'proof checker' mean in simple terms?
Why do these maths problems matter?
Should I worry AI is getting too smart?
Could this approach be used outside maths?
Sources
- An OpenAI model has disproved a central conjecture in discrete geometry — OpenAI, 20 May 2026
- Solving open problems with AlphaProof Nexus (preprint, arXiv:2605.22763) — Google DeepMind / arXiv, 21 May 2026
- OpenAI's milestone math breakthrough played to AI's strengths — Understanding AI, 22 May 2026
- Google DeepMind's AlphaProof Nexus solves 9 Erdős problems and 44 conjectures — Crypto Briefing, 26 May 2026