Keywords: LLMs, cultural technologies, skateboarding, creative communities, AI industry.
Update, September 13th, 2026. Terry has kindly reposted this essay on his blog.
A note on the timing of this essay. I started writing this in the middle of the ICM, and it has taken me this long to finish writing, see also my shorter post from August. In any case the bulk of the essay was done in August and shared with a few close friends for feedback. Earlier this week, Tristan Buckmaster announced a breakthrough finding, done in collaboration with Levent Alpöge, of a finite-time blow up for 3D incompressible Euler with forcing. He also made serious allegations of misconduct by OpenAI (see announcement). The next day, OpenAI announced a finite-time blow up result for 3D Navier-Stokes with forcing, obtained from their internal LLM. There has been lots of press coverage, I recommend Kenneth Chang’s reporting. I have decided not to change the opening of the essay and instead make a note of the news here. Tristan gave considerable credit to the ideas of Córdoba and Martínez-Zoroa in the blow up construction; I found this important detail to be quite validating of the perspective on LLMs I advocate in this essay, and found that to be further reason to leave the content of the essay largely unchanged.
What would be the significance if one day we woke up to the news that output from a Large Language Model contains a definitive answer to the question of blow up for the incompressible Navier-Stokes equations, and to the question of where the zeros of the Riemann zeta function lie? (Or, to mention problems dear to my heart: a proof of a nonlocal version of the ABP maximum principle or a resolution of the Mahler or Komlós conjectures?)
Many of my fellow mathematicians consider this a kind of nightmare situation, and I think this is a very valid position. Here I would like to make a different case, however. Personally, I have come to believe that the news of a Large Model output containing an interesting, insightful, novel mathematical idea can and should be received as a source of communal pride for mathematics and for mathematicians. I also believe there can be a world where LLMs are not harmful but rather supportive of mathematicians as we follow our prime directive: understanding things.
My intention in this essay is to elaborate on what I mean by this. I believe this perspective can assuage fears many of us feel amid the daily hype and doom narratives around AI and the alleged impending obsolescence of mathematicians. I will not be talking about the less existential / philosophical and more practical / material concerns here, such as challenges to the funding of the mathematics occupation, how journals will work, changes to the training mathematicians, and so. Those issues are also important, but I think clarifying the perspective I push here will make the discussion of the more practical matters easier.
Bill Thurston touched on the same idea in his famous 1994 essay, “On proof and progress in mathematics”, which is more relevant than ever. Thurston observes that mathematicians (as a group) prove theorems and solve problems, but that is not all they do: they also absorb these solutions, and find new uses for them once they are proven. This “digestion”1 in turn allows for the next wave of breakthroughs, and thus advances our understanding of mathematics.
This is not secondary to what mathematicians do, but our very purpose. The value of mathematics to humanity lies not in determining whether a statement or conjecture is true or false, but rather in the understanding gained as we make those determinations (even if the facts themselves are also of great value). To put it briefly, mathematicians are people who try to understand the causes of things.

Felix, qui potuit rerum cognoscere causas
LLMs, the technology $\neq$ LLMs, the products from large companies
The conversation about LLMs and generative AI within the mathematical community at large must deal with an obstacle before we can get to the substance of the matter. That obstacle is the current and unfortunate identification of generative AI, the technology, with generative AI, the specific product and vision from the large AI companies. It is imperative for us to make that distinction and separate the two. The private AI lab vision is what most people have in mind when discussing AI, and this is no accident, for this vision has been aggressively pushed into every part of our daily lives.
It might seem impossible at the moment, but we can imagine an alternative world with many parallel approaches to the use of LLMs developed by the grassroots, reflecting the variety of circumstances across different communities, industries, and institutions. A seed for this world is the emergence of open-weight models. This is a matter for a different essay.
So in all that follows, I will crucially rely on the above distinction, so that when I say LLMs or “Large Model” I am talking about the technology itself, and very much not any of the major AI companies or their products.
With all this throat clearing behind us, let me get to the matter at hand. Why should the output of an LLM containing new and exciting mathematical ideas be a source of pride for mathematicians, and not of anxiety? My answer owes a lot to this very interesting article by Henry Farrell, Alison Gopnik, Cosma Shalizi, and James Evans. I quote here one of their central assertions:
…Large models should not be viewed primarily as intelligent agents but as a new kind of cultural and social technology, allowing humans to take advantage of information other humans have accumulated.
I will do my best to apply the ideas in their article to mathematics and the work of mathematicians. If, rather than continue reading this, you stop and go read their article and Thurston’s article instead, I would consider this essay a success.
Cultural and social technologies
Humans have always lived in a complex world and their survival has hinged on their ability to process information in amounts beyond what any individual can handle – and it is likely this has been so nearly as long as humans have existed. Herbert Simon’s concept of “bounded rationality” captures this well: in real life, unlike in the simplest control theory models, we have limitations on the time available to decide and on the accuracy of our information. This means that an agent making decisions must give up optimality and settle for “good enough.”
Simon saw that one implication of bounded rationality is that humans have to (and do) develop systems to process information, as a group. We see that humans have built states, stories, bureaucracies, markets, libraries, and more. These are cultural and social technologies. They are means by which humans (not individually but as a group) navigate amounts of information so vast that they cannot be handled by any individual on their own. An important feature of these technologies is that they are to some extent forgetful: you ignore or downplay smaller details in order to make a bigger picture easier to comprehend.
Markets, governments, and bureaucracies mean each human can focus on a few things and still get the aggregate benefits. Thanks to markets, governments, and bureaucracies, I can rely on being able to buy fresh fruits and vegetables around the block from my Manhattan apartment. Thanks to markets, governments, and the field of modern medicine, we now have high confidence that a newborn child will live to be an adult – in stark contrast to what was the norm for most of human history. Cultural technologies are powerful and awe-inspiring things. It is understandable that we tend to anthropomorphize them so often.
This week I started teaching my graduate-level optimal transport course. I wrapped up the first class by showing one neat application of optimal transport: a “one-page proof” of the isoperimetric inequality in $\mathbb{R}^n$. Of course, the proof being just one page depends on what we are assuming as given.
This is standard, not only when teaching advanced courses but also very much in day-to-day research, and it is made possible by the power of cultural technologies. In the proof I don’t have to explain from scratch the notion of Lebesgue measure or sets of finite perimeter. When in the middle of a proof I say “I will do the details only when the domain has a smooth boundary, and I will assume the Lebesgue measure is $1$ without loss of generality thanks to scaling,” I am relying on a shared trick that I can trust a first-year PhD student to be familiar with.
All of this is to my benefit and the students’. We can focus on “the one page” proof with the interesting new idea: how Brenier’s theorem provides a map through which the isoperimetric inequality is essentially reduced to the arithmetic-geometric mean inequality.
So, here is basically what I showed on the blackboard: assuming $E \subset \mathbb{R}^n$ has Lebesgue measure $1$, we want to prove that
\[|\partial B| \leq |\partial E|\]where $B$ is the $n$-dimensional Euclidean ball with the same volume as $E$, so $|B|=1$. Then, I tell my class how Brenier’s theorem says there is always a map $T$ from $E$ to $B$, that this map is volume-preserving, and that $T = \nabla u$ for some convex function $u$. Then, since $\det(D^2u(x))=1$ and all the eigenvalues of $D^2u(x)$ are non-negative, the arithmetic-geometric mean inequality says that $1 \leq \frac{1}{n}\Delta u$ everywhere in $E$, and so
\[1 = |E| = \int_{E}1\;dx \leq \frac{1}{n}\int_{E}\Delta u(x)\;dx.\]Then, the divergence theorem and the fact that $||\nabla u(x)||\leq r := (1/\omega_n)^{1/n}$ (where $\omega_n$ = volume of the unit ball) guarantee that
\[1 \leq \frac{1}{n}\int_{\partial E}\nabla u(x)\cdot \nu(x)\;d\sigma(x) \leq \frac{1}{n}(1/\omega_n)^{1/n}|\partial E|.\]For a ball of volume $1$, scaling shows that $|\partial B| = n\omega_n^{1/n}$, so we conclude that $|\partial E| \geq |\partial B|$, as we wanted.
In interacting with this proof, we are engaging in a conversation through space and time with many mathematicians, but especially with Knothe (1957) and Gromov (1986), followed later by separate works of Brenier, McCann, and Trudinger – and even with my younger self in a set of lecture notes with McCann, where we go over the above proof (see concretely Section 1.6). The conversation is mediated by a common mathematical culture, which encompasses the common definitions and methods I can count on the graduate students to know. I don’t have to prove the divergence theorem or the change of variables formula to justify these computations, and I don’t have to worry about the smoothness of $T(x)$, since I can work by approximation.
Without having to go into all that detail, I presented on one blackboard the idea that the Brenier map takes you from the arithmetic-geometric mean inequality to the isoperimetric inequality, and I did this in the last 15 minutes of the class. These are cultural technologies in action! Which cultural technology in particular? In this case it is the technology of a scientific community with standards in what is taught, the mathematical literature, and so on.
Large Language Models as cultural technologies
Farrell et al.’s main point is that LLMs are cultural technologies, just like markets, bureaucracies, and scientific disciplines. Accordingly, the framework of cultural technologies is a very useful one as we seek to understand and navigate the impact of Large Models on our disciplines.
LLMs indeed are a lot like the example from my class, and in a way they are “wrappers” around an older cultural technology: the scientific literature. When you ask an LLM a question about the isoperimetric inequality, the ABP maximum principle, or a PDE, you are effectively interacting with every person who has thought about and worked on those things – provided the data the LLM was trained on contained those contributions in one way or another. Surely the lecture notes I reference above are in the training data of most LLMs in use today, just as a vast portion of the digitized mathematical literature (across all disciplines) was part of the training data. LLMs are lossy, so while a model might not be able to reproduce a given paper exactly as in the training data, it can reproduce large parts of it.
What sets LLMs apart from older cultural and social technologies is the fact that we have digitized all this data, and can now explore it, reproduce it, and recombine it with ease by sending text commands (prompts) from a computer terminal, or even a phone. LLMs are, in essence, top-tier information-processing tools. They are also an excellent tool for generating conversational level-text, and for coherently combining patterns of information even if they come from very different sources.
By a product-design choice at a few private labs, the output of most LLMs is conversational in tone and made easy to anthropomorphize. These properties have amplified an ELIZA-like effect that reaches even people with advanced training in mathematics and computer science. In a way, anthropomorphizing this technology – which, as some have argued, did not have to be the case – does everyone a disservice: the conversational setup of the LLM gives the appearance of a single “thinking” entity, and buries the work by vast groups of people and which are being recombined. This appearance is a feature not of LLMs, but of the LLMs as the AI companies want them. We are not dealing with a very smart, genius-level intelligence; we are interacting with something greater (and yet somehow more human): the aggregate creations of millennia of human ingenuity and creativity, digitized and efficiently compressed – all mediated by state-of-the-art statistical and generative algorithms.
In particular, any mathematical proof found in an LLM output is the collective outcome of lifetimes of work by mathematicians. It must be welcomed as the fruit of our shared mathematical heritage, belonging to all of humanity, carried through millennia by people who at various times went by names other than mathematician (natural philosophers, astronomers, geometers). This applies very much to LLM output containing what we might consider novel proofs or ideas, for in the end, it is all an outgrowth of that gigantic body of knowledge. Now there lies a scaling law that has been going for centuries. The recently arrived software is just a kind of harness or wrapper for it.
Now, let’s reconsider the “nightmare scenario” I opened this essay with. One day we wake up to the news that the output of a Large Model contains a resolution to the Riemann Hypothesis. We are now equipped with the ideas of Farrell et al., and especially the idea of cultural technologies. Reading the news under this framework, we are learning a new mathematical fact, but we are also learning that this fact is the fruit of the collective efforts of researchers in the field up to this moment. The problem was cracked by the ideas of one or two (or maybe many) researchers to the point that a resolution could be reached by the sampling and recombining of ideas an LLM is good at.
This leaves us with a delicate question: what ideas are reachable from the totality of ideas in the mathematical literature at a given time? I am going to borrow from geometry and loosely use the metaphor of “convex combination” of ideas to discuss this.
Convex hulls of ideas

A convex body containing another, from Geometry and the Imagination by Hilbert and Cohn-Vossen
The capabilities of LLMs have progressed quite a bit in the last couple of years, and dramatically so in mathematics this year alone. The rapid improvement makes this discussion difficult. The question “what can LLMs do?” is a rather moving target, at least for the time being. This being said, let’s try to think of a related question that might not be as LLM-dependent, and has a chance of being useful, even if it is still vague.
Ideas sometimes are very clearly compositions of other, more fundamental ones. Sometimes a new idea really stands apart from what came before. What would the finding of a solution the Riemann Hypothesis, or of the 3D Navier-Stokes smoothness problem, within the output of an LLM, tell us? One toy model (which might not reflect empirical reality) is that the result might belong to some sort of “closure” or “convex hull” of the available literature. In broader terms, one could pose the following question: given a collection of ideas, what ideas could be said to stand squarely beyond them, and which are lying closer to their “convex hull”?
I will not give a more concrete definition for this, because I am not able to. However, I can discuss by way of example what this definition ought to cover. I also do not use “convex hull” to mean that the point is derivative or that the finding of a point of something inside a convex hull to be easy – high dimensional convex sets are deeply complex objects!
At the very least, we can agree that the proof of the infinitude of primes $p \equiv b \mod a$ with $a,b$ coprime is certainly not in the “convex hull” of number theory that preceded Euler and Dirichlet: using an infinite series to quantify the “amount of primes” is precisely the sort of idea that stood squarely outside of what was considered number theory until it was proposed. However, the question of the infinitude of primes of the form $an+b$ ($a,b$ coprime) could be posed within the number theory that existed before Euler. In fact, this was known already for special choices of $a$ and $b$, say for primes of the form $4n+1$. Likewise, number fields, ideals, and their unique factorization are all ideas lying well beyond all the number theory that preceded them. However, once we have these ideas together, it is slightly less outside the convex hull (or maybe it is even inside the convex hull) to develop a proof of the fact that every ideal class contains infinitely many prime ideals – a combination of Euler’s and Dirichlet’s ideas with the idea of, say, Gaussian integers.
From what I have seen (so far) in research areas that I am closest to, no LLM has produced anything remotely like Euler’s introduction of the methods of analysis in the study of primes. One possibility is that LLMs are not even being used so far for such things. In fact, one could put it like this: Euler was not simply answering a well-known question, but rather saying “here is actually a whole new question that we did not know we could ask, and is actually a very important one that connects to older problems!” The framework of LLMs as cultural technologies, to the extent that it reflects existing LLMs, suggests that this is a type of contribution we cannot expect to find in their output.
According to Farrell et al., cultural and social technologies are preceded by individual capabilities, which are then aggregated and transmitted by the technology. In their essay, they note that “without innovation, there would be no point to imitation.” Suppose Euler had access to an LLM trained on all of the mathematics known up to the moment he started working in number theory. Without an individual first introducing the analytical perspective on prime numbers, it would be very hard for our hypothetical LLM to innovate analytic number theory into existence, and then use this innovation to solve the “conjecture”2 on the infinitude of primes that are congruent to $b$ modulo $a$, for any $a,b$ coprime. All the same, the framework is compatible with the happy possibility that a person working with an LLM might have a eureka moment where they, the person, arrive at this innovation.
Here is another example, of a slightly different character. One of the most important theorems in PDE is the Krylov-Safonov theorem for non-divergence form elliptic and parabolic equations. For those not in PDE, these are versions of the Laplace equation and the heat equation, except the “coefficients” in front of the partial derivatives are not constant but rather change from point to point.
Such equations, linear and nonlinear, appear in the study of stochastic processes as well as in optimal control and geometry. The Krylov-Safonov theorem produces a Hölder regularity estimate for solutions of the PDE that only depends on the uniform ellipticity – what is notable is that this estimate works even if the coefficients are potentially very discontinuous. Consider, for concreteness, a linear equation such as
\[\sum_{i,j=1}^n a_{ij}(x)\partial_{ij}u(x) = 0\]When the coefficients $a_{ij}$ are smooth, the regularity of solutions to this equation can be understood from the theory for the Laplace equation. Heuristically, the coefficients being smooth means that as you zoom in near a point, the coefficients are closer and closer to being constant, which puts you just an affine change of variables away from Laplace’s equation. Therefore the good regularization properties of Laplace’s equation carry upward to the larger scales and so one proves $u$ is smooth. This estimate, however, depends very much on the modulus of continuity of the functions $a_{ij}(x)$.
Now, I cannot go on at length about this here, but obtaining estimates for $u$ that do not depend on the continuity of the coefficients $a_{ij}(x)$ is extremely important. In fact, the importance of this question cannot be overstated – such estimates have profound implications in the theory of stochastic processes, in geometric analysis and differential geometry (say, for the many PDE related to curvature), in stochastic homogenization, and more. For this reason the Krylov-Safonov theorem is one of the most important PDE results we have.
The Krylov-Safonov theorem, however, depends on a mathematical fact that comes from a very different area: convex analysis. There is an estimate by Aleksandrov that says the following: for $B\subset\mathbb{R}^n$ denoting the unit ball, given any function $h:B \to\mathbb{R}$ convex in $B$ and with $h=0$ on $\partial B$, we have
\[\|h\|_\infty \leq C_n |\partial h(B)|^{1/n}\]where $\partial h$ denotes the subdifferential of the convex function. This estimate holds for general convex bodies, not just $B$, and in this generality it is equivalent to the reverse Blaschke-Santaló inequality from convex geometry. The estimate is the key component in what is known as the Aleksandrov-Bakelman-Pucci (ABP) estimate, which in turn is essential for the Krylov-Safonov theorem.
To state it in simple terms: without this fact from convex geometry the Krylov-Safonov theorem cannot be proved, not in any way we know today. This is a significant statement, as many important theorems in analysis have multiple, independent lines of attack that do not necessarily use the same tools. This would be like asking how to prove results in geometric analysis and probability without knowledge of the Poincaré inequality. Without the advances in convex geometry done many decades earlier, the theory of fully nonlinear elliptic PDE, which seems at first sight so remote from convex geometry, could not have advanced the way it did onward from Krylov and Safonov.
Imagine a parallel universe where we came to LLM technology before we knew about the Krylov-Safonov theorem, so that the question of regularity for non-divergence elliptic equations without dependence on coefficients is considered an important open problem. The cultural technologies framework suggests that in this universe an LLM (at least at the current level of power) would not be able to settle this open problem if the mathematical literature in this parallel universe does not have a well-developed field of convex geometry/analysis. By the same token, if in this parallel universe there is a well-developed convex geometry literature, including the Aleksandrov estimate, then it is plausible an LLM could bridge the two things and arrive at the ABP estimate and Krylov-Safonov’s theorem.
Thinking in terms of “convex hull of ideas,” one sees that whether an LLM can solve a given mathematical problem is in large part a function of the status of the mathematical literature at any given time, and is not just a function of how large the LLM is. With this perspective, the event of an LLM output containing the solution of a hard problem ought to be welcome as a happy occurrence: an LLM helped us dig out a gem buried somewhere in our mathematical heritage, or in the convex hull, if you will.
In fact, this has already happened: less than a year ago, we heard news of “Erdős problems” solved by an LLM, which then turned out to be instances of the LLM finding solutions in old (and in some cases, forgotten!) papers. At the time, this realization was seen as a defeat for LLMs. Now, that situation illustrates the perspective I am advocating here, in miniature.
Creative communities during turbulent times
The framework of LLMs as cultural technologies, together with the “convex hull” of ideas it inspired, has addressed my own philosophical concerns with LLMs. It has reinforced my optimism that the mathematics I value will endure. However, this does not change the fact that the mathematical community is facing serious and urgent challenges. The AI companies’ aggressive and narrow-minded push of their vision is hurting the field, exacerbating old fault lines and dysfunctions of our profession, and creating new ones.
This is not the first time a creative community has faced a turbulent period or a big change brought by new technologies or economic realities. There are lessons we can learn from past instances, in mathematics and in other disciplines. One can think of painters facing the advent of photography, musicians facing the arrival of the synthethizer, or mathematicians themselves – think for instance of the creation of the arXiv in the 1990s. However unusual it might seem, chapters from the history of professional skateboarding present uncanny parallels to the circumstances in mathematics, and this is a history with many teachings and warnings for mathematicians.

Rodney Mullen performing “the Impossible” trick in an episode of the Physics Girl (Dianna Cowern)
By the 1980s, skateboarding had been around for several decades. It had professional contests, sponsorship deals, and a specialized press. Like mathematics, skateboarding had an ecosystem that allowed practitioners to develop their craft and advance it, and even make some money along the way. It also had subfields, like freestyle, vert, and street skating. In 1984, Kevin Harris, who came from the subfield of freestyling, observed street skaters with only “six months of experience” were already earning more than Rodney Mullen, who at the time had tens of thousands of hours of experience and a nearly perfect record in professional competitions. Harris has recalled his reaction to this observation as: “Okay, this is the death of freestyle and it’s going to hurt vert skating hugely.”
The recession around 1990 dealt the finishing blow to the professional skateboarding ecosystem as it existed at the time. Within a few years, the discipline lost many of the venues and big public events it was built around, even largely losing its version of “funding” – sponsorships. This crisis drove many pros into early retirement. One could argue that with championships and sponsorships gone, the metric for measuring success was gone, and with it all the incentives for participation in competitive skateboarding. However, this was not the only metric, nor even the main one, for skateboarding continued and evolved in the coming years.
Without public contests with large audiences and related sponsorships, skateboarders had to find some other means to communicate and show their work. The economic and social problems in this case were solved with a technology that already existed. Skating videos became the new medium by which skateboarders recorded and communicated their work. Before, if a skater spent countless days perfecting a delicate trick, they could showcase it to at most a few hundred, maybe a thousand people during a championship. Now, a videotape showcasing the work of several skaters could be seen by tens of thousands of people. Before, a skater who did not win several championships was likely to go unappreciated. Now, a skater needed only one good “part” –a single skater’s segment within a video– to be recognized and contribute new tricks to the skating community. The most successful “parts” would introduce a new trick or a new twist on an old one that would get replicated and adopted by others.
I find it notable that the resolution to the crisis in skateboarding made it a bit less a competitive sport and a bit more a collaborative, creative field, with videotapes playing the role of journals3, and a “part” playing the role of the articles in the journals. This turbulent transition turned out to be for the best for skateboarding as a discipline, while it also turned out for the worse for many individual skateboarders. This ecosystem remains largely unchanged today (with internet videos replacing videotapes). The motivation for a skater to participate in the community, aside from their love for skating itself, was that their contributions could be showcased in video, and their ideas could be shared with others who could try to repeat them or recombine them in new ways.
Mullen said that an invention only becomes an innovation once the community adopts it. This provides a measure of the value of a work. In mathematics we often talk about “the impact” of a paper or idea, and this is completely analogous to a trick being impactful: has the community taken this idea, and used it, and modified it further? Has this work made it easier for newcomers to the field to understand the state of the art?
Reclaiming our mathematical heritage, and the future of mathematics
One lesson I take from cultural technologies is that the most precious component in a Large Language Model, the part most responsible for whatever value lies in its output, is the human-generated work, the cultural and technical heritage encoded in it. Throughout the summer the largest AI companies have been treating the solution of high-profile mathematical problems as a form of PR for their products. We as mathematicians need to lean into our curiosity and love of math and welcome this new knowledge, without letting its provenance ruin our enjoyment of it. We must counter the damaging misconception that “AI has beaten mathematicians at solving problem X,” which is being promoted by the corporate AI labs. Instead, we can celebrate how our understanding of mathematics, together with our creativity, lies behind the gems of mathematics LLMs are helping us uncover.
Earlier in this essay I touched on the importance of separating LLMs (the technology) from LLMs (the product and vision of the AI companies). Well, we need to also make this a reality for our mathematical infrastructure. It needs to be a top priority of the community to promote and work towards a world with open LLMs – think of the LLM equivalent of free software. It would be quite bad if our profession became dependent on the whims of the AI industry, whose leaders inhabit a moral universe so distinct from ours in science and mathematics that nothing fruitful or sincere can emerge from associating with them4. We must remember that the combustion engine was never the heritage or exclusive creation of car makers in Detroit; and railroads and locomotives were not the exclusive heritage of the robber barons of the Gilded Age; and so it is also true that LLMs do not belong in any fundamental way to large industry labs in the Bay Area. The convex hull perspective in this essay suggests it is worthwhile to look into lighter, mathematics-specific LLMs with well curated data, and asses the scale of the technical (and financial) obstacles for such a thing. It would demostrate that the big labs’ “moat” is not as deep as they’d like it to be.
We also ought to approach the use of LLMs with an open attitude, experimenting and exploring until we find those LLM practices that best serve our values: advancing human understanding of mathematics, as Thurston would put it. This also means working as a community to minimize the disruption to mathematical careers, so that as few people as possible share the fate of those skateboarders who had no choice but to retire early. This must also involve policy advocacy: to dramatically increase public funding for science, and policy advocacy for serious public oversight of the large AI labs. We must also experiment with the creation of new social structures and professional arrangements within mathematics, so that we can allow as many individuals as possible to remain part of the mathematical community. If we are successful in this, the mathematical community will come out stronger from this turbulent time, serving as a model for other communities and industries working to transition to a world of ubiquitous LLMs.
To return to where we started, an LLM’s output “solving” the nonlocal ABP or the Mahler conjecture would be, circumstances aside and in a strict sense, a triumph for the discipline. That does not mean that we cannot have mixed feelings about it. It is not pleasant to be scooped by a living, thinking person, and it is no more pleasant to be scooped by the collective knowledge of the field, processed and recombined via an LLM. Rather than being shocked, we ought to be in awe, not of the technology, but of the immensity and power of the mathematical knowledge we have built to this day. Large Models have made this enormity apparent and impossible to miss. Mathematicians must claim such developments as the rightful fruit of our discipline, and accordingly take pride in them, understand them, adopt them, and continue advancing mathematics. Fortunate are we, able to know the causes of things.
-
“Digestion” is a term Terry Tao proposed in his ICM lecture on mathematics in the age of AI, part of a metaphor of LLM generated proofs as bringing us to an era of proof abundance, as in mass produced food. ↩
-
I say “conjecture” only in the context of this hypothetical scenario of Euler and his 18th century LLM. ↩
-
In mathematics, meanwhile, we are actually approaching a situation where we will have to switch towards whatever comes after journals, or at least journals as they are presently understood. ↩
-
Here I am very purposefully borrowing a phrase I have long been fond of. It was written by Bertrand Russell in a response to Oswald Mosley. ↩