When AI Doesn't Prove Mathematics, It Changes It
Key Takeaways
- A paper posted to arXiv on July 29, 2026 shows the 150-year-old Maxwell Conjecture is false: five point charges can produce at least 24 non-degenerate critical points, against a conjectured maximum of 16.
- The counterexample’s construction was suggested by GPT-5.6 Sol. The verification, formal analysis, and write-up were done entirely by the three human authors.
- This is a different role for AI than “solved the problem.” It’s closer to: proposed a direction worth trying.
- Mathematicians have long described flashes of intuition preceding proof: Poincaré and Ramanujan both wrote about it. What’s new is that the intuition can now come from outside the mathematician.
- If AI systems are going to contribute usefully to discovery, we may need to evaluate them on the quality of the directions they propose, not only on whether they can independently verify those directions to be true.
We usually ask whether AI can prove mathematics. A result from late July suggests we’ve been asking the more boring question.
Here’s what happened: a 150-year-old result in electrostatics, the Maxwell Conjecture, just fell, and the construction behind the counterexample was suggested by GPT-5.6 Sol. The proof is human. The initial idea wasn’t.
The more interesting question isn’t whether AI can prove mathematics. It’s whether AI can change what mathematicians choose to investigate in the first place.
What Is the Maxwell Conjecture?
The Maxwell conjecture is a 150-year-old guess about electricity: if you scatter n electric charges in space, the conjecture says there are at most (n−1)² spots where the resulting force field cancels out to nothing. Those cancellation spots are the “critical points”: places where a test particle would feel no push at all.
In 1873, James Clerk Maxwell (better known for unifying electricity and magnetism) made this as a side observation about electrostatics: the potential field generated by n point charges seemed to have at most (n−1)² non-degenerate critical points, the equilibrium points where the field’s gradient vanishes. For five charges, that bound predicts at most 16. It went unresolved for over 150 years, sitting in that category of problem too simple to attract sustained attack and too stubborn to fall by accident.
Was the Maxwell Conjecture Disproven?
Yes. A paper posted to arXiv on July 29, 2026 proves the Maxwell Conjecture false: a five-charge configuration produces at least 24 non-degenerate critical points, well past the 16 the conjecture allowed.
Philip Arathoon, Gavin Ball, and Matthew D. Kvalheim posted “The Maxwell Conjecture is False” to arXiv. The counterexample is almost embarrassingly small: place three unit charges at the vertices of an equilateral triangle in a plane (a configuration with four equilibria, one at the center and three along the edges), then add two smaller charges just off that plane, symmetric above and below it. That perturbation causes a degenerate critical point to bifurcate, and the resulting five-charge configuration has at least 24 non-degenerate critical points. Sixteen was the ceiling. Twenty-four is what showed up.
Three charges. Four equilibria: one central, three along the edges.
The proof (the Taylor expansions of harmonic polynomials, the Hessian classification confirming each point is genuinely non-degenerate, the transversality argument tying it all together) is conventional, rigorous, human mathematics. What’s unusual is where the construction came from: the authors state directly that the key geometric idea was suggested by GPT-5.6 Sol.
What Did the AI Actually Contribute?
Precision matters here: “AI solved a 150-year-old conjecture” is not what happened, and the real story is more specific and more interesting.
GPT-5.6 Sol didn’t produce a proof. It produced a suggestion: take a symmetric base configuration, break the symmetry with small off-axis charges, and watch what happens to a degenerate equilibrium under that perturbation. That’s a direction to search in (a hypothesis about where a counterexample might live), not a verified mathematical fact. Everything that made the hypothesis into mathematics (checking that the bifurcation actually produces 24 distinct non-degenerate points, not some artifact of approximation) was done by Arathoon, Ball, and Kvalheim using standard analytic tools.
That distinction between an exact result and a numerical approximation of one is doing quiet work in this story, and it is worth pulling out on its own. A quantity can be pinned down completely and still have no finite decimal expansion, which is why √2 is an exact number even though its decimals never end, and why a system that stores 1.4142135623730951 has already thrown away something a system reasoning from x² = 2 still has.
So the division of labor looks like this: the model proposed a construction; the humans determined whether the construction was true, and did the work of proving that it was. Neither step is dispensable. A suggestion nobody verifies is just a guess: the same discipline of holding a claim in doubt until it’s earned belief that applies to any unverified claim, human or machine-generated. A verification process with nothing to verify has nothing to do.
Is This a New Stage in How Mathematics Gets Made?
The traditional picture of mathematical progress runs roughly: observation, conjecture, proof, theorem. Somewhere in that chain, an insight has to arrive from somewhere: usually from a mathematician’s accumulated intuition, sometimes from a flash that seems to come from nowhere.
Tap a stage above to see what happens there.
Henri Poincaré wrote about exactly this in his 1908 essay on mathematical creation: he described working on a problem, setting it aside, and then having the solution arrive fully formed while stepping onto a bus, the product (he argued) of unconscious work that had continued after he’d stopped consciously attending to it. Ramanujan’s biographers describe something similar but stranger: results he said arrived through intuition he couldn’t fully explain, which he’d then hand to others, or later prove himself, sometimes years after first writing them down.
What both stories share is that the source of the idea and the justification of the idea were separate acts, often separated by real time and sometimes by different people. The Maxwell Conjecture counterexample follows that same shape (construction from one source, verification from another), except this time the source generating the initial direction wasn’t a mathematician’s unconscious. It was a language model, prompted by mathematicians who were themselves searching for a way in.
That’s not “AI replacing intuition.” It’s intuition’s supply chain gaining a second source.
What Exactly Did the Model Contribute: Knowledge, or Something Else?
It’s tempting to reach for a single word (knowledge, creativity, search) and none of them quite fit alone.
It wasn’t knowledge in the sense of retrieving a known fact; as far as the authors’ own account goes, no one had previously written down this construction. It has some of the shape of creativity, in that a genuinely new geometric idea got produced. But it also has the shape of search: a model trained on enormous mathematical content proposing a plausible perturbation to try, out of an enormous space of perturbations it could have proposed instead.
Maybe the honest answer is that these categories were never as separable as we treat them in ordinary conversation. A mathematician’s “creative insight” is also a search: through a lifetime of learned pattern, intuition about which symmetries tend to break interestingly, memory of similar problems. What’s new isn’t that a machine searched. It’s that the search space it searched was mathematical construction space, and what it returned wasn’t noise.
What If Most of the Suggestions Are Wrong?
Here’s the harder question underneath the celebration: suppose a system like this proposes ten thousand constructions, and one of them turns out to matter. Did it discover mathematics? Or did it just expand the number of directions available for humans to check?
I don’t think that’s a rhetorical question with an obvious answer, and I don’t think it needs one to be useful. Even “it expanded the search space available to humans” is a real, non-trivial contribution: mathematicians have finite time and attention, and most conjectures don’t get solved because nobody happened to try the right construction, not because the right construction was unreachable in principle. A system that reliably nominates promising, previously-unconsidered directions changes the economics of that search, even if it’s right only rarely and even if a human still has to do all the confirming.
At these settings, the model's suggestions surface about 2 promising directions, less than what one mathematician gets through by hand in the same stretch. Most of the 200 proposed still go nowhere. That's the point: the cost of a wrong suggestion is low, so the value only has to show up occasionally.
That framing also explains why “AI solved another conjecture” is the wrong headline. The theorem still belongs to the three mathematicians who proved it. What changed is upstream of the theorem: in what gets tried.
How Should This Change the Way We Evaluate These Systems?
This is where the story connects to a question I’ve written about before, from the other direction: how do you evaluate an AI system whose value isn’t fully captured by “did it get the right answer”?
Most benchmarking of frontier models still centers on correctness: did the model produce the right proof, the right code, the right final number. That’s the right frame when a system is meant to deliver a verified result on its own. It’s the wrong frame for a system whose actual contribution is a well-chosen hypothesis that something else will verify. A model that proposes one brilliant, previously-unexplored construction and nine unusable ones might be far more valuable to a research program than one that proposes ten thousand safe, unremarkable ideas nobody would have missed anyway, and a pure accuracy score won’t distinguish them.
If hypothesis generation is going to be a real category of AI contribution to science, it needs its own evaluation axis: not just “was this true,” but “was this worth someone’s time to check,” measured against what a domain expert would or wouldn’t have tried on their own. That’s a harder thing to score than a proof, but it’s closer to what actually happened on July 29th than any pass/fail metric would be.
Closing Thought
For most of the history of mathematics, the central question has been whether a given statement is true or false. The Maxwell Conjecture just answered that question about itself.
But the more durable question this result raises isn’t about truth. It’s about origin: where do the ideas worth checking come from, and how many of them are we currently failing to try?
The next shift in mathematical discovery probably won’t be machines replacing mathematicians at the business of proof: that part of the story, so far, hasn’t changed at all. It may instead be machines changing the population of ideas mathematicians consider worth their time to pursue.
That boundary, where a machine genuinely contributes to discovery and where the work still belongs to a human, is the thread running through these essays on AI in science and mathematics.
That’s a smaller claim than “AI does mathematics now.” It might also be the more consequential one.
It’s also not a claim about whether anything understood the mathematics along the way. That’s a separate question, and Searle’s Chinese Room argument settled it in advance, regardless of how the proof turned out.
Recognizing which ideas are worth the pursuit is itself the difference between reacting and reflecting: the slower, effortful mode that decides what deserves the search in the first place.
Further Reading
- Arathoon, P., Ball, G., Kvalheim, M. D. (2026). “The Maxwell Conjecture is False.” arXiv:2607.27197. The source paper; states directly that the counterexample’s construction was suggested by GPT-5.6 Sol.
- Poincaré, H. (1908). Mathematical Creation, from Science and Method. Poincaré’s own account of unconscious incubation preceding a flash of mathematical insight.
- Hadamard, J. (1945). The Psychology of Invention in the Mathematical Field. Princeton University Press. Surveys similar accounts across mathematicians, including Poincaré’s.
- “The Maxwell Conjecture Is False”, Hacker News discussion thread.
More essays at Call to Think · About this project
Frequently Asked Questions
Why did the Maxwell Conjecture stand unresolved for 150 years?
It sat in the awkward middle: simple enough to state that it never attracted a sustained research program, but stubborn enough that it was never going to fall by accident. No one had systematically searched the space of charge configurations for a violation, so the counterexample that existed all along went unfound.
What did GPT-5.6 Sol actually contribute to disproving it?
Not a proof. It suggested a geometric seed: start from a symmetric three-charge configuration, then perturb it with two smaller off-axis charges and watch a degenerate critical point bifurcate. The paper's three human authors then did the hard analysis (Taylor expansions of harmonic polynomials, Hessian classifications, and a transversality argument) to confirm and formally write up the resulting configuration of five charges with at least 24 non-degenerate critical points.
Does this mean AI can now prove theorems on its own?
No. This case is explicitly the opposite of that story: the construction was a hypothesis, not a proof. It still had to be verified, formalized, and defended through conventional mathematical argument by human mathematicians before it counted as mathematics. What changed is where the initial idea came from.
What does this change about how we should evaluate AI systems?
It suggests outcome correctness isn't the only thing worth measuring. A system that proposes a promising, previously unconsidered direction (even one it can't itself verify) may be doing something valuable that a narrow accuracy score won't capture. Evaluating hypothesis generation as its own capability, distinct from proof or execution, is a genuinely different measurement problem.