Technology

Terence Tao: Understanding 6 essential mathematical concepts

2026-08-21 04:00
575 views
Terence Tao: Understanding 6 essential mathematical concepts

Big Think Home For Business Membership Open search Open main menu For Business Learning & Development Develop leaders in your organization with Big Think+ classes Custom Content Create premium thought...

This content is locked. Please login or become a member.

Terence Tao: Understanding 6 essential mathematical concepts

Terence Tao has spent decades solving problems most people can't even state, and he says he's never had a single eureka moment. In this conversation, the UCLA mathematician and Fields Medalist breaks math down to six essential ideas: numbers, algebra, geometry, probability, analysis, and dynamics.

My name is Terence Tao. I'm a professor of mathematics at the University of California, Los Angeles. And I have a forthcoming book, Six Math Essentials. Today on Big Think, I'll be talking about six essential pillars of mathematics: how math interacts with science historically and how it has anticipated many of the great developments in the sciences, and how new developments in AI will impact math and science going forward.

Chapter 1: The six essential elements of mathematics

I decided to organize my book around six really fundamental concepts that have origins from thousands of years ago or centuries ago, which are very familiar in the early stages to most people, but mathematicians have developed over time to become extremely sophisticated. Numbers is the first concept, and then algebra, geometry, probability, analysis, and dynamics. These are all very basic concepts, but they've evolved into very sophisticated mathematics. But if you strip away all the technical complexities, they really are just extremely intuitive concepts, and the mathematics that we developed is just a precise language to describe it really, really carefully, and in a way that allows you to think really clearly about these concepts.

Numbers

Numbers are one of the oldest mathematical inventions and still the most useful, really. We have records of carvings on bones that predate the alphabet or other writing, and it was invented multiple times by multiple civilizations. If we didn't have numbers, we would have to always speak poetically when trying to describe any situation, and it would always be a little bit imprecise, and the next person who's communicating what you're saying would get it slightly different. Numbers allow for precision. They are placeholders for concepts like quantity and size and magnitude, into very portable things that you can communicate to other people who may not have directly interacted with the objects you're describing. And once you try to describe a very complicated scenario with many, many moving parts, you need numbers and everything else that's born on top of that.

Humans are not really wired to think in numbers. If you don't have the ability to think quantitatively, to measure both the benefits and the costs of an action and which one is bigger, you can make some life choices that you'll regret later — that you spend a lot of resources for very little gain. The first step in being able to think more quantitatively about these things is to understand numbers, and then more advanced mathematical topics like probability and algebra on top of that. Basic things like agriculture or trade could not have happened without numbers to measure large quantities of grain, and actually the ability to tax — you can't have a civilization without taxation, unfortunately — and that requires mathematics and numbers.

"Of course not everything is quantitative. If you want to go on a date, you shouldn't be measuring the costs and benefits of your prospective partner. Some things should still be very subjective and personal."

But there are increasingly, in this modern world, lots of decisions — for example in finance or in medicine — where some quantitative thinking is very helpful.

The thing about numbers is that they take on a life of their own, because once you have the concept of number, you can study numbers abstractly, divorced from their actual application, and you find patterns, and you find that it's very natural to extend the number system that you have to create new numbers which you wouldn't have thought would be applicable to your original context, but they fit very well into the number system.

If you're counting sheep, sometimes you want to add sheep and you want to subtract sheep. And so very soon you develop the notions of addition and subtraction. But you realize if you are only counting numbers — one, two, three, four — you find that you can always add these numbers together, but you can't always subtract. If you subtract four from three, it doesn't make any sense — you can't take away four sheep from three sheep. But the patterns in the numbers themselves are so regular that if you just blindly apply the rules of arithmetic, it feels like you should be able to take away four from three and have a new number. And it took a while, but eventually people realized that you can invent these negative numbers, and you can add negative one, negative two, negative three to your number system, and zero. Zero took a long time to actually realize was a good addition to the number system. And you still get all the nice laws of arithmetic — for example, if you take a number a and you subtract b and then you add back b again, you get back a. That's one of the laws of arithmetic, and it still works even when you have negative numbers.

Similarly, we learned to divide numbers by another, and we created fractions, and fractions fit very well into this number system. And then there was a shock. We found that there were numbers somehow between all these rational numbers, all these fractions — numbers like the square root of two, which could not be expressed as any ratio. And this was a big shock — these numbers are literally called irrational numbers, which is Latin for "insane," not "unreasonable." But they do exist, and they're very useful. You can never write down all their digits on a finite sheet of paper, but it is very, very useful to have all these extra numbers lying around.

And then eventually we tried to take square roots of negative numbers, which we couldn't do in a regular number system, and we invented complex numbers, and that turned out to be extremely useful for electromagnetics and quantum mechanics. It's remarkable that these number systems, which were often invented just so that we could be better at solving equations and solving practical problems, end up actually being the most natural language to describe very, very complicated phenomena in the real world, like quantum mechanics.

Algebra

Algebra is the second layer of abstraction over numbers. So with numbers, we took concrete things like a bunch of sheep or a quantity of water, and we replaced these quantities with numbers that you can then apply operations to — addition, subtraction, division, and so forth. Algebra goes one step further and tries to not look at specific numbers like seven or 17. First of all, it replaces numbers by even more generic placeholders, giving them names like x and y, but it also studies the operations themselves — plus and times and all these other operations — and asks what properties the operations have, not just the numbers.

And so people discovered that these operations that are so useful in arithmetic themselves have many fascinating properties. Addition has a property called commutativity — if you add a to b, that's the same as adding b to a. And that turns out to be an extremely useful property for helping you solve problems involving addition. Similarly, multiplication and all the other basic arithmetic operations obey these very simple laws.

Later on we discovered that these laws also hold for other operations. For example, if I want to take an object and rotate it — I can rotate it by 30 degrees and then rotate it by another 60 degrees. But if I rotate it in the other order, rotate it by 60 degrees first and 30 degrees next, I end up with the same position I started with. These two rotations are commutative. So even though that operation has nothing to do with addition or multiplication in a traditional sense of numbers, it has the same algebraic structure. On the other hand, some things do not obey the commutative law. If I put on my socks and then I put on my shoes, I get a different outcome than if I put on my shoes first and then my socks. Those two operations do not commute.

Certain operations obey nice laws like the commutative law, and certain operations don't. Sometimes you see that the laws that are present are very similar to laws that we already understand for, say, numbers. And because of that, we can take intuition and ideas and proofs from a theory of numbers and transfer them to a different setting. For example, matrices are a much more complicated concept than a number — not just one number, but a whole square array of numbers. But it turns out that matrices obey very similar laws of algebra to numbers. And if you are very good at manipulating numbers, you can start to manipulate matrices the same way. Many of our modern technologies, for example large language models, are based on being able to manipulate matrices very, very efficiently.

Once you have abstracted to numbers and then to algebra, you are working with equations that involve variables like x and y, and x may have some physical meaning and some specific value. But often it can clarify your thinking to not focus on the specific values of these numbers or what they represent, and just manipulate these equations by pure algebra, just by moving symbols around. When you first learn this, it feels very disconnected from your actual experience, but it is a very powerful technique.

One early use of algebra was the story of Johannes Kepler, the astronomer and scientist who was one day walking down the streets of his hometown and saw the wine market. Wine sellers were selling wine by the barrel — some large barrels and some small barrels — but they were able to figure out how much wine was in each barrel so that they could pay the wine sellers for these barrels. If you have a barrel and you want to compute how much wine there is, you could pour it out into cups and so forth, but it was very tedious. What fascinated Kepler was that the person in charge of the market had a very efficient way to measure the volume of the barrel. He just had a stick with various markings on it, and there was a bunghole in the middle of the barrel, and he just poked the stick down the barrel into the corner and saw how far the stick went, and based on that marking he could say, "Oh, this is 30 gallons of wine," or whatever, and then they could price it.

This astounded Kepler — how you could just take this one measurement, of just this one sort of diagonal, and work out the shape of the volume, because some barrels could be very tall and skinny or short and wide, and somehow this one measurement was able to compute the volume. He did the math. This was a puzzle to him. So he went home and wrote out some equations — assume the radius of the barrel is r and the height is h. Nowadays, with modern algebra, this is a question you can assign to a high school student. It's almost precisely one of these word problems that we love to give our students, and you can compute the volume and this length. He found that the length did not completely determine the volume, but a wine seller would want to sell as much wine as possible, and you would want to maximize how much volume you could get for a given length. Reasoning that all these merchants were trying to maximize their profit, even though he couldn't quite solve the equations right away, if he added this profit incentive, he did some rudimentary version of what we would now call calculus. And that turned out to almost exactly match the shape of the barrels that were actually sold in the marketplace. His formulas did actually match what people used, with very high precision. He had explained these rules, which I think the marketplace had come up with over time just by empirical measurement. But he had found a very satisfactory explanation, and I think this was part of the inspiration for the calculus developed by Newton and Leibniz a few centuries later.

Geometry

Geometry is literally Greek for "measurement of the earth." Since antiquity it was important to know how many miles it was to travel from one place to another, and how to navigate in the ocean or out in the wilderness by the stars. We needed to understand how to use things that we could observe, like angles and distances, for very practical problems like transportation.

Just like numbers have got various patterns — like a plus b equals b plus a — once you start measuring distances between different points and angles, there are lots and lots of relations between all those measurements as well. So geometry obeys laws just like numbers obey laws. One law, for example, is similarity. Once you know that two shapes are similar — they have the same angles — then all their sides are proportionate. Once you know that the side of the big triangle is, say, two times as big as the side of the small triangle, then you know that all the other sides of the big triangle are also two times as big. The proportions are equal.

The reason why this is so useful is that it allows you to predict the measurement of distances or scales that you couldn't directly reach — you couldn't directly measure. You can look at a distant mountain and estimate how far away it is because you know something about how tall it is and the angle of elevation — you can use a similar triangle and make a scale model. So now you can navigate and see how to get to a distant location even if you can't see it directly from your sailboat. And even in ancient Greek times, they were able to measure distances to the moon and the sun, which they could not possibly measure directly, but just through the laws of geometry — similar triangles and things like that, which they had already worked out at that time — they could actually get reasonably good measurements without any satellites or advanced technology. Once you know geometry, you can really extend your senses well beyond what you can just touch and measure directly.

Probability

The fourth math essential is probability, which is the standard way that mathematicians try to encapsulate one of the basic features of the real world: uncertainty. In your primary or high school classes, when you teach mathematics, we often present very sanitized, predictable word problems — Annie has 30 apples, she gives half of them to James, how much does James have — where everything is precise and you know all the information. But in the real world, there is uncertainty and unpredictability.

When you flip a coin, maybe it is heads, maybe it is tails. Now, if you knew exactly how much force you applied to a coin and all these measurements, you could maybe do enough of a simulation to actually predict exactly which way the coin will land. But this is extremely difficult, and often you don't have that data. What was eventually realized is that rather than try to compute exact answers all the time for every single outcome, sometimes you just have to accept that there's a range of outcomes to a given measurement. But what's important is which ones are more frequent and which ones are less frequent.

The first people to realize this was important were gamblers, because gamblers would bet on certain events — like betting that certain dice rolls would sum up to a certain number. If they could calculate the odds correctly, they could make money, and if they didn't calculate odds correctly, they would lose money in the long run. The mathematics of probability was created by several letters that gamblers wrote to their mathematician friends, asking for help trying to optimize their gambling strategies.

"Probability has blossomed well beyond its gambling roots. Anytime we have a system too complicated to model all the way from first principles, there's going to be some stochasticity, and we want to have a probabilistic model."

So whether it's the stock market or whether a drug is going to be successful, we turn to probability. Now, quite often, we don't know what the odds are. Probability works best when it involves an event that happens over and over again — thousands and thousands of trials, and we can start getting a good measure of the odds. If it's an event that only happens once in a century, it may not quite be the right mathematics, and we're still trying to work out mathematics that is a better replacement for probability for extremely rare events.

There are some miracles that make it effective. You may think that every time you do a different experiment — instead of doing a medical trial, you're trying to understand the outcome of a die roll, or what genetic traits are going to emerge from evolution — you would think that in every different circumstance you get all these different distributions, some heavy-tailed and some narrow. But there are these funny laws in probability called universality laws, where even very general types of random systems produce various common shapes. The most famous of which is the Gaussian, or what's called the bell curve shape — many distributions, like the heights of men or women, form almost a perfect bell curve. We actually have good explanations now — we have understood probability well enough that we can explain quite a few of these universal laws from mathematics, although some are still mysterious.

Analysis

Analysis is how mathematics deals with two things. One is inaccuracy in our measurements — that sometimes it's not so much randomness but just imprecision. Analysis is sort of the mathematics of error bars, where we've realized that sometimes we need to understand not only what the numerical value of various quantities are, but how much uncertainty we have — what are the plus and minus error bars around them. These are concepts that you can't even talk about with only a qualitative language of "big" and "small" — that only works up to a point. We may measure the length of an object and it's roughly two meters, but maybe two meters plus or minus 10 centimeters. There's an approximation and there's an error, and ideally you want the errors to be zero, but in the real world we can't always make the errors entirely zero. But sometimes if we make the errors smaller and smaller, and if we keep making more and more precise measurements, we can make the errors shrink to zero. But it may take an infinite amount of precision, or an infinite amount of time, to get the error all the way down to zero.

Analysis is also about how we take limits and how we deal with infinities. In algebra, the laws of algebra work very well when you're just taking a finite number of operations. If you're just taking five things and adding them together, there's no problem — you can rearrange them in any order. But once you start trying to work with an infinite number of operations, there are some funny paradoxes that show up — sometimes you can rearrange an infinite number of objects and end up with a different sum than before you rearranged, which doesn't happen with finite sums.

For example, if you're always betting on, say, roulette — red or black — you win 50% of the time and you lose 50% of the time. There is a theorem that there is no strategy that will allow you to constantly win, that will guarantee you a win in the long run, as long as you only have a finite amount of money. No betting strategy can beat the house, basically. There's a strategy of always doubling down when you lose and betting bigger and bigger numbers, and the moment — as long as you win at least once — you will get your dollar back. So there was some way to constantly beat the house. But the problem is that it assumes you have an infinite amount of money. What the strategy is doing is compressing all the risk of losing money into this very, very small event where you're always losing. At some point you're betting millions of dollars, and at some point you become bankrupt.

So analysis helps you understand exactly what these tail risks are, and how to reason with infinities in a way that avoids all these paradoxes. There was a lot of very inaccurate mathematics that predated analysis, where people were just saying, "Oh, if I do this infinitely often, I can get rid of all these problems." And it took a while to realize that infinity is a very dangerous beast if you're not trained to deal with it properly. But we need to deal with infinities all the time in the real world — well, maybe not infinities, but we need to deal with very large numbers, and infinity is a very good approximation for understanding how we deal with large numbers. But it has to be used with care.

One famous demonstration of the unintuitive nature of infinity is what's called the infinite monkey theorem. The common way to phrase this is that if you have an infinite number of monkeys in a room — or maybe just one monkey typing forever on a typewriter, hitting keys at random — then usually the monkey will just type nonsense and gibberish, but every so often the monkey will type a word, or a sentence.

"The infinite monkey theorem is that if you wait long enough, or you have enough monkeys, you're almost certainly guaranteed eventually that the monkey will type whatever you like — the complete works of Shakespeare, or Hamlet, or Wikipedia, or anything."

You can prove this mathematically — as long as the probability of the monkey doing it at least once is positive, it doesn't matter how small; if you wait long enough, eventually the probability that this particular pattern gets hit will eventually go to one. It's like if you play Russian roulette and you only have one bullet in the revolver and you keep firing — it doesn't matter how many chambers you have, eventually it will fire. You'll get your hit for any given word or sentence or paragraph, but the time taken grows exponentially with the size of the text. If it's just a four-letter word, it might take an hour or so of typing before the monkey gets it. But if it's already a seven-letter word — the word "Shakespeare," for instance — that may already take years. A sentence might take millennia. And to get even a fraction of Hamlet — think even just a page — would take way more than the age of the universe before you'd actually see it.

So infinity is actually just a placeholder for a number that could potentially be far larger than any fixed number that you could come up with. Even though in the real world you don't have infinitely many monkeys, you don't have infinite budget — often we reason with infinity first, as an idealized situation, to see what is possible. And then from there we can turn to the more quantitative questions of exactly what we can do with finite resources. But first you understand what to do with infinite resources.

When I was a kid, I used to play a lot of computer games. For many computer games, there were certain cheats you could put in — you could give yourself infinite health or infinite ammunition. And it was sometimes helpful to play the game first with those cheats, where you didn't have to worry about managing your health potions or your ammo, and just see how to solve the game, and then you could play it on a harder mode and see how to do things more efficiently.

So many problems in math are actually solved this way. One thing that distinguishes math from other disciplines is that we have the freedom to fail, because failure is very cheap in mathematics. If you're running a business and you make a bad business decision and your company goes bankrupt, that's a terrible mistake. If you're a surgeon and you cut the wrong thing, that's a terrible mistake. But if you're trying to solve a math problem and you make an incorrect assumption and it doesn't work, it's not really that much of a bad mistake — you just try again. It's actually a good move, when you're trying to solve a math problem, to first make an idealized assumption — assume that you have an infinite amount of energy, or zero friction, or some other unrealistic assumption. Solve the problem then, and then try to get from the infinite world back to the finite world. And this is where analysis comes in, to carefully see what features of infinite mathematics still work in the finite world and which ones break down.

Dynamics

Dynamics is the mathematics of change over time — the study of the rules of incremental change, how a state evolves from one time to the next, and all kinds of emergent and interesting behavior which you may not expect from the initial rules. We have found that even very simple rules can generate extremely complicated emergent behavior if you iterate them long enough.

Evolution is a good example in biology. You have a bunch of organisms and they reproduce, and the fitter ones survive more often than the weaker ones, and the organisms have certain traits that they can pass down to descendants. These are very simple rules, and it turns out that you can create a massive diversity of species and predator-prey relationships and incredibly complicated dynamics.

If you're on the freeway and you have all these cars, and each car is just trying to move as fast as it can given the car in front of it — if there are too many cars in front, you'll slow down; if there aren't many cars, you speed up — and each individual car is not doing anything very complicated, it's just trying to optimize its flow. But when you put all the cars together and see what it does to the whole network, you get these amazing emergent phenomena, like traffic waves. You get waves that can compress and grow, kind of like a slinky — if you mess around with a slinky, you can have a compression wave and an expansion wave, and it leads to phenomena you wouldn't initially expect. If there's a traffic shock, like an accident, and the cars pile up, then even when the accident is removed and there's no obstruction to travel, if you iterate the dynamics you can actually do the math and see that it takes a while for the compression wave to dissipate. Living in Los Angeles, this is a phenomenon I've encountered quite often — sometimes I encounter a slowdown in traffic and there's no accident, no immediate cause, because many hours ago there was something that caused a slowdown, but the dynamics take a certain amount of time for the wave to dissipate.

Once you understand the dynamics really well, you can do modeling and simulation, and then you can make predictions — like, if I add another lane to this freeway, will the traffic get better? In fact, sometimes it doesn't. There are paradoxes where closing off certain lanes of traffic can actually make the global traffic flow faster.

Now, some dynamics are predictable. Sometimes we have equilibria, which are states that just stay the same for all time, and sometimes these equilibria are stable — if you move a little bit away from that state, you come back to it. If you have a pendulum going straight down, that's a stable equilibrium — if you perturb it, it will move a little bit and then get back toward the stable equilibrium. But if you make a pendulum upside down, balancing on the tip, it could technically stay in that position forever. But any slight perturbation will cause it to move away from the equilibrium over time. So it's important to know which equilibria are stable and which ones are not.

We are now facing a world of climate change, where we have lived for 10,000 years in a climate pretty close to equilibrium — it would get hotter or colder some years, but it would bounce back to equilibrium. And we're now actually in danger of leaving that equilibrium for much less stable dynamics, which is scary, but it needs to be modeled, and we may need to figure out how to adapt and change our agriculture and all our other practices.

Understanding dynamics, and which systems are stable, which ones are not, which ones are chaotic, which ones are predictable — it's actually extremely important. There are very mundane things, like predicting the weather. We take for granted that we have accurate weather predictions for the next seven days — this was actually an amazing achievement of atmospheric scientists. They collected lots and lots of data, but they also solved a lot of dynamical systems problems that got the error rate down to a point where we can actually reliably predict weather forecast a week in advance. It's still not completely 100% accurate, but it's much more accurate than just guessing. Systems that involve a lot of humans are still very unpredictable — the dynamics of the stock market or politics are well beyond the ability of current dynamical systems theory to model, but natural systems and some human systems, like traffic, we can actually model.

So it's a fairly advanced area of mathematics. We often need a lot of computer simulations, and we need to solve very advanced differential equations, but it can give some very valuable insights. One of the discoveries of dynamical systems is that most systems exhibit what's called chaos. And this came as a surprise. When Newton introduced his law of gravitation, one of the great successes of the theory was that it explained the motion of the moon around the earth and the earth around the sun — he could explain retroactively all these funny laws of Kepler, like why planets move in ellipses. He solved what we now call the two-body problem — if you have two massive objects, like the sun and the earth, and you move them around, governed by a single law of motion, Newton's inverse-square law of universal gravitation, he could solve the equations using his newly derived theory of calculus, and he had perfect formulas for the orbits — they were perfect ellipses, exactly verifying Kepler's theory. It was an amazing achievement.

"Once he solved the two-body problem, it was very natural for many of Newton's successors — and I think also Newton himself — to try to solve the three-body problem. I think Newton once said that this was the only problem that ever gave him a headache, because no matter what he tried, he could not get an exact solution."

Leading societies of the time offered major prizes for anyone who could write the solution. This was considered one of the major open problems in mathematics. We still do not have an exact solution for these equations, and the belief now is that there isn't really one that you can write down as a nice, neat formula. But when you actually look at the numerics, you see that it is not some nice periodic pattern — it often stays periodic for a long period of time, but then suddenly it will change to something a little bit different, and then it will change yet again.

We suspect now that our solar system, which currently has eight planets, in the past had other planets that were in the system and mostly moved in sort of elliptical orbits, according to Kepler. Every so often, the little interactions between the gravitational force of Jupiter or Mars would jiggle these planets a little bit out of their usual orbit, and occasionally they would just veer off completely, and sometimes two planets would collide, or one would escape the solar system — for example, there's an asteroid belt which we believe is a remnant of a collision from millions of years ago. Even the most stable of systems, like the solar system, which looks like it hasn't changed for millennia, has long-term instabilities. Once you move beyond the simplest of systems, there are lots of little, tiny, unpredictable, or very hard to predict deviations that occasionally can pile up — just like occasionally a bunch of monkeys can sort of write the works of Shakespeare, occasionally gravitational perturbations can set an entire planet off course.

So often, actually, in the most advanced forms of dynamics today, even if you start off with a completely deterministic system with no unpredictability whatsoever, we find that the best way to model it is to approximate it using probability, and just assume that there's going to be some random fluctuation back and forth, and eventually your predictions will just get blurrier and blurrier — which just seems to be a fundamental feature of chaos, which many systems have.

So these six essentials don't describe all of mathematics, but they do describe six of the great themes that mathematics tries to encapsulate, and there's much more precision to it — this is just a taste of what goes on these days.

Chapter 2: How math solves the problems of science

I view STEM as a whole ecosystem. At the bottom there's basic research — mathematics and some other fundamental sciences — where we pursue things mostly driven by curiosity. We see a phenomenon that is crying out for an explanation or further study. It may not be a phenomenon that we urgently need to solve right now for an immediate problem, but it's something that looks like it should have an interesting answer. Mathematics is almost entirely curiosity-driven like that — there's some pattern in numbers, some pattern in shapes, that people just observed while trying to do something else, and we want to understand it better. At some point, other scientists are able to connect that pattern to something they're studying, and some mathematical numerical pattern might show up in the behavior of insects in a swarm, or in a stock market, or whatever. And then, sometimes, once you understand it, you can actually convert it into some useful technology, or you can have a company that makes money out of some service related to exploiting that phenomenon.

We often don't see that — that's much further down the pipeline. What you do need is that the people doing the basic sciences have to talk to the people doing applied sciences, and they have to talk to people who are doing engineering, and they have to talk to people in industry. If you didn't have one of these communities, you wouldn't have this pipeline of getting from curiosity-driven questions to actual commercial results — like the ability to communicate across the planet with almost zero cost.

The unreasonable effectiveness of mathematics

This is part of what Eugene Wigner calls the unreasonable effectiveness of mathematics in the physical sciences. He observed that mathematicians often discover concepts — such as complex numbers or curved space — just because they seem to be a natural extension of the mathematical objects they're already studying. And then, 10, 20, 50 years later, some scientists discover that these concepts, introduced for fun or for play, were in fact exactly what was needed to understand some new type of science. So that's a really amazing phenomenon, and we still don't have a good explanation for why that actually works.

The parallel postulate and non-Euclidean geometry

One historical example of how curiosity-driven mathematics led to a really deep scientific advance was the story of the parallel postulate. Euclid, in the third century BC, introduced the notion of proof — being able to explain complicated results in geometry from simpler axioms. For example, that the sum of angles of a triangle always adds up to 180 degrees — he was able to explain that in terms of simpler axioms. He laid out five axioms of geometry, which he thought reduced all the other facts he knew about points and angles and lines to these five statements, and four of them were very straightforward — like, if I give you two points, there's always a line you can draw between them. These are very straightforward, non-controversial axioms. But there was this one axiom, called the parallel postulate, which gave him a lot of grief. His original version was very complicated — it got simplified, but even the simplified version was always controversial.

The simplified version is that if you have a line and a point, and the point is not on the line, then there's exactly one line you can draw through that point which is parallel to the first line — meaning it never crosses the first line. So that was this axiom, that you can always draw a parallel line through any other point, and there's only one — you cannot draw two parallel lines. Once you have that, you can derive the angles of a triangle adding up to 180 degrees, and all the other classic results of Euclidean geometry. But it was a very ugly axiom compared to the other four, which were really elegant.

Eventually they realized that there actually were multiple geometries beyond Euclidean geometry. There's something called spherical geometry, where there are no parallel lines at all — instead of lines you have great circles, like the equator or a longitude, and these great circles on the sphere always intersect. You can never make two parallel great circles. So there are no parallel lines. And then there's this weirder geometry, harder to visualize, called hyperbolic geometry, where lines actually diverge from each other — lines that start off looking parallel but move further and further apart, and in fact there are actually multiple parallel lines you can draw from one point to a given line. Those two geometries are entirely self-consistent, and eventually it was just accepted that there were more geometries out there than just Euclidean geometry. These were the first two non-Euclidean geometries to be discovered: spherical geometry and hyperbolic geometry.

But once we were freed of our notion that there's only one geometry, this opened all the floodgates, and people studied all kinds of other geometries — curved spaces, spaces shaped like donuts or with twists in them. There are geometries where, if you're right-handed, you can go off and explore the universe and come back left-handed — where you can change your orientation just by travel, which is very unintuitive — or you can come back smaller or larger than what you started with. People developed all these geometries and a very nice language for describing all of them, called Riemannian geometry, after Bernard Riemann.

But it was a curiosity — these abstract curved spaces — the universe we lived in seemed completely flat. But then Einstein, when he was trying to understand gravity, eventually came to the conclusion that what gravity was doing was bending space and time in a certain way, and he needed a language to describe how space and time could bend in such a way that light rays would become not straight, and sometimes they would hit each other or diverge. He asked his mathematician friend if there was any existing mathematics that would describe this, and he said, "Oh yeah, there's this bright chap, Bernard Riemann, who developed this theory." And it turned out to be almost exactly the right language to describe the Einstein equations. There's a notion of curvature that some space can have — positive curvature, negative curvature — in Riemannian geometry, and the Einstein equations turned out to be extremely simple to state in this language: basically, mass and energy create curvature, and the curvature of space and time is proportional to how much mass and energy you have in your system. That's basically the Einstein equations. Now, solving them is a different matter — they're extremely hard to model. Even the question of how to model two colliding black holes, we can barely do it with modern supercomputers. But stating the equations is actually extremely natural once we had this language.

Sphere packing: from cannonballs to wireless networks

Another example of how mathematical curiosity led to really practical developments centuries later is the story of sphere packing. There was a British sailor who was curious about the question of a certain number of cannonballs they had to stack in the hold of their ship. These cannonballs are round, not square, so when you stack them there's a certain amount of wasted space, and he was curious what the most efficient way to pack cannonballs was, so you could get the most cannonballs into a certain amount of space. He asked a physician friend, who happened to know Johannes Kepler.

Kepler eventually proposed that the most efficient packing should be the same packing you see nowadays in supermarkets when you pack oranges — what's called hexagonal close packing. You pack layer by layer — each layer is sort of a triangular grid of cannonballs or oranges — and then you stack another triangular grid on top, shifted by a little bit, and then stack another back. There's a regular pattern, which is the most natural pattern, and it's about 76% efficient. Kepler thought this was the best you could do — that there was no clever way to squeeze in any more space — but he couldn't actually prove it. This became known as the Kepler conjecture, and it was one of the most famous unsolved problems in geometry for centuries.

In two dimensions — packing discs in a plane — that's a simpler problem, solved by about 1900. There's a similar triangular lattice, and that was relatively easy to prove was optimal. Three dimensions is just too many possibilities — there was no way there. In fact, we still do not have a nice, simple proof of the Kepler conjecture that humans can completely understand by themselves. The conjecture was eventually solved and published, I think, in 1998, but it required computers — it was one of the first computer-assisted proofs. There was a team of referees who said they could not verify all of the computations, but they at least believed the strategy was correct, though there were still lingering doubts. It was only much more recently, in 2014, that the proof was converted into what's called a proof-assistant language — a computer language specifically designed to check proofs with 100% certainty. So the Kepler conjecture is now formally verified — we are now 100% certain it is true.

But mathematicians weren't content with just the three-dimensional problem, so they also asked what happens in four dimensions, or five, or six. Here, of course, there is no practical application — there are no four-dimensional oranges or cannonballs you'd like to pack — but people still ask this question. People also asked what happens if, instead of a continuous space, you have a discrete space. In particular, once computer science became developed, we realized that in addition to the geometry of regular space, where x, y, z coordinates are given by real numbers, we're interested in studying the geometry of strings of bits. This is now very divorced-sounding from the original sphere-packing problem, both because now you have many, many dimensions — thousands and thousands — and space is now discrete rather than continuous. But still it is geometry, and many of the techniques to understand sphere packing still work in this setting.

And then it turned out that this problem — packing spheres as efficiently as possible onto this big, huge cube of bit strings — turns out to be extremely practical. When cell phones became digital, every signal you send, like an image or a text, is encoded as some bitstring, which is then sent over some wireless network. But other people are also sending signals, and you don't want your signal to be corrupted by interference and mistaken for someone else's signal. So you want to keep each different signal as separated from each other in this space of bitstrings as possible. And it turns out mathematically that this problem of separating all these signals, so there's no way one can be confused for another, is almost exactly the sphere-packing problem, except in high dimensions and discrete.

A lot of the mathematics used to understand sphere packing could be used to design really efficient codes — and not just to design codes, but also to tell engineers what the theoretical limit of communication was: what is the maximum number of bits per second you could possibly hope to send in a certain wireless spectrum. That gave really good benchmarks to measure how efficient your protocol was — you could price how many billions of dollars you should pay for a certain wireless spectrum, because now you know exactly how much data you can push through it. The entire wireless telecommunication industry is based on being able to pack oranges in really high dimensions.

Compressed sensing: from MRI to broadband

One of the applications of mathematics that I was involved in, that I'm most proud of, is the story of compressed sensing. I was once at an interdisciplinary program at a math institute here in Los Angeles, and I met with a friend of mine who is a statistician, and he was working with an electrical engineer trying to improve medical imaging — specifically MRI scans. At the time, MRI scans were quite slow — you had to sit in the scanning machine for like three minutes so that you could collect enough data from all different angles that the scan could reconstruct a good image of your body and be able to pick up tumors or cysts or anything else that's medically important to resolve.

But if you only sat in the machine for a short period of time, like half a minute, you would not get enough data. If you tried the standard reconstruction algorithm — using what's called least squares approximation, the standard technique at the time — you would get an image so blurry and so low-resolution that you could not tell anything useful for diagnosis purposes. So you had to sit in these machines for minutes and minutes, and if you were a kid, sometimes you had to be sedated, because after the second minute or so the kid would wriggle around and not follow instructions.

So they were trying a new technique — not least squares, but something called total variation minimization. They had a hunch that this other method might perform a little bit better, so they tried it on some test data. They were expecting a slightly sharper image than the least squares approximation, but they got perfect resolution — they got back almost exactly the correct image, even though they only took a few measurements.

"It would be like giving someone a crossword puzzle where you had only filled in 10% of the letters, and suddenly they could fill in all the other letters without having to look up the clues."

They couldn't explain this, and they showed it to me, and my first instinct was: you made a mistake. You could not possibly have done what you said you did. And in fact, I was going to prove to them that there was not enough information in the data they took to make this measurement possible. So I went home that night and tried to write down a proof that there was no way to correctly guess the right image from the small amount of data they were measuring. And while writing it down, I found one of my steps didn't work — and in fact, it showed the opposite: that if a certain measurement matrix had a certain property, then what they were doing was actually going to work. And then I checked that their measurement matrix actually did seem to obey this property. So I actually understood how the method worked, and I went back to them the next day and explained this, and they got very excited, and we wrote a couple of papers, and that got everyone else excited.

This method that they had stumbled upon was not completely new. Seismologists had discovered a similar method — they had a different problem, trying to understand the fault lines to locate the fault lines of a crust based on a small amount of seismic data. And astronomers had a similar problem, trying to measure the location of stars using a very small amount of observed data. There were a couple of other disciplines where a similar problem — trying to extract a high-quality image from a very small amount of signal — had occurred. In each case they had found some ad hoc fix that could squeeze more resolution out of the data they had, but they could not explain mathematically why it worked. The seismologists thought, "Here's a trick, but it only works for seismographs." The astronomers had a trick, but it only worked for astronomy. But once we found the mathematical explanation, we found that this was a general technique that we now call compressed sensing. It's useful for MRI, but also for wireless broadband, and for certain types of sensor networks. Once we figured out the underlying mathematics, we could see all the other applications it's useful for. And so now compressed sensing is taught in textbooks right next to least squares — sometimes least squares is the right thing to do, and sometimes compressed sensing is the right thing to do, and sometimes neither. It's now a very well-developed theory. I was quite pleased to be involved at the very beginning of that.

Why the same short explanations work for math and physics

It's a fascinating interplay between mathematics and science. We have this unreasonable effectiveness of mathematics, where mathematical discoveries often end up being the best way to explain physical phenomena. Philosophers and historians have debated why this is the case. One of my theories is that whenever we learn anything — whether it's math or science or any other subject — the first theories or explanations we make are often not the best. We don't understand what the cleanest way to express something is. Before you understand something completely, you might have an overly elaborate explanation for why something is true. But often the true explanation is more elegant and shorter and simpler than our initial attempts to describe the phenomenon. But finding the short explanation takes time, because you have to unlearn certain assumptions you might have that turn out to be incorrect.

For example, with Einstein's theory of relativity, one of the key assumptions people had before Einstein was that time was universal — everyone had the same notion of time, that an hour for me is the same as an hour for you. That mindset really blocks you from finding the right way to explain gravity in particular. But once you accept that everyone has their own relative notion of time, you can find the right language to explain things properly.

Mathematicians also try to take phenomena they first understand using very inefficient language and condense them — try to find the most concise, elegant explanation of a mathematical phenomenon. And I think, just because there's only so many ways you can say things concisely, it just happens that often the concise way to describe some mathematical phenomenon is also a concise way to describe a physical phenomenon. So that is my theory. Unfortunately, we only have one timeline of science — we have maybe a hundred turning points in science, and that's a little bit of data, and you can make some theories. I would love, in the far future, if we meet other civilizations, to see their history of science and see whether they also had their version of Kepler and Einstein and Newton, and whether they followed a similar track or a completely different one. I don't know.

The nature of mathematical discovery

The funny thing about doing mathematics is that there's this stereotype that we're all geniuses, and we're stuck at a problem, and then you get this eureka insight — a light bulb goes on and you get this genius idea out of nowhere. I would love for that to happen to me, actually. This does not happen so often to me.

What does happen when I work on a problem is that I try something and it doesn't work. I try something else, and it kind of works, but it gets stuck at a certain point. But now at least I know there's at least one obstacle, and I need to find some tool that will help me deal with that obstacle. So maybe I will identify a sub-problem that has the same type of difficulty but is simpler, and I try to solve that one first. And if I can do that, I can try to scale back up to the original problem, and I go back and forth. Often a lot of what you're doing is exploring the negative space of the problem — all the techniques that don't work — and eventually, if you have enough of the negative space, the path forward becomes clear almost by elimination, that there's only sort of one thing to do that could work. Sometimes there's nothing that can work, and then you give up on the problem.

But there's this repeated — what you might call failure, but it really is just really understanding the limitations of what different approaches can do. And then, after weeks or months of this, the answer becomes clear. But by that point, it's so internalized what the difficulties are that it doesn't feel amazing to you anymore — it feels natural. Like, of course you had to do this, because there's this difficulty, you must go around this pothole, you must do this step first because in five lines you're going to need this hypothesis to solve this problem. You just become so attuned to the problem that everything becomes natural. Or sometimes you never solve the problem, because you never attune, and you give up.

"The feeling I get is never so much eureka, but it's always, 'How come I missed this earlier? I was so stupid.' But often it's the constant experimentation and failure that really primes you to accept the right solution."

There's this dramatic contrast between the standards we assign to outcomes and the standards we assign to process. For outcomes, mathematics very famously has a very high standard of correctness — for a given problem, there's a correct answer and lots of incorrect answers. When we grade our students' homework, if you get your sign wrong and get the wrong answer, you get all these negative marks — you get criticized if you make math mistakes in your final answer. One consequence of that is that many students who go through, say, a high school level of mathematics become very averse to making any mistakes whatsoever when they try to approach a math problem.

Paradoxically, the process of arriving at the answer is almost the complete opposite — you almost have to make mistakes over and over again, and you have to try the stupid things first to appreciate why the clever things work. There's a quote by Niels Bohr, the physicist, who said that an expert is someone who has made all the mistakes that can be made in a very narrow field. You don't publish these mistakes — you execute these mistakes in your process in order to locate the correct answer, but only after exploring a lot of incorrect answers first. It's very important, I think, to normalize failure in the process, to disclose that behind every successful solution to a problem there are dozens of incorrect attempts, and it's not because the people trying these problems were stupid — it's often just part of the learning process.

Chapter 3: How AI is changing math and science forever

Science and mathematics has changed a lot over the centuries. Traditionally in science, the two major paradigms were theory and experiment. You would create a theory — Kepler might create a theory of how planets move, or Newton might create a theory of gravity — and then there's experimental data — you would run an experiment and see what happens, and then you try to see if the theory and the experiment fit.

Math was a little different, in that it was almost entirely theory. There are very, very few experiments that you would do purely in mathematics — there were a few, for example, Gauss famously computed the first 100,000 prime numbers, and that was a data set he used to make predictions — he predicted what we now call the prime number theorem. But science and experiment were the two major forms of science.

Then later on, simulation came along — you didn't have to run a big, expensive experiment; sometimes you could just simulate, let's say, a hurricane in a supercomputer instead of in real life. And then later on, big data came along — rather than just do a small number of experiments to try to confirm or deny a specific theory, you could take megabytes or petabytes of data and try to discern patterns, try to extract laws from just massive data sets. That's a more emerging type of science.

But now all these modes of science are being transformed, because we now also have AI to help us. In the past, every one of these ways of doing science had to be done by human scientists — you had to have someone perform the experiments, or someone to do the theoretical calculations, or run the simulations, or go through the data. You could use computers for some of that, but you had to program the data analysis tool, or whatever, and you still needed a lot of expertise.

We have automated labs that can perform experiments automatically. You can get a coding agent to run a simulation for you, and you can try to run automated data analysis. And increasingly you can also do automated theory — you can take some mathematical problem and ask what the consequences of these hypotheses and axioms are, what conclusion you should get. These tools can now be done at scale, much faster — an AI could potentially run many more theoretical analyses than any one human scientist could.

On the other hand, this is not the only thing we want. There's value in doing things the slow way — a scientist spending hours and hours working things out on pen and paper, doing experiments in the field with their bare hands, actually debugging the simulations that show up — they often learn a lot of extra insight beyond just getting the answer they're trying to seek. They can discover new phenomena, see connections, see similarities to some previous thing that's been studied elsewhere in the literature, and they can communicate what they're finding to other people.

So there is this paradox: on the one hand, AIs are becoming more powerful and more capable, making fewer mistakes, and they are ostensibly achieving a lot of the goals we think scientists are trying to do — they're running experiments, analyzing data, writing papers. But it may be that it comes at the cost of the AI picking up some skill, but no human scientist getting any better at doing the science. No human can communicate exactly what just happened, and why this scientific discovery is interesting, why this proof is new, what features it has, and how it connects. We may have to redesign our conception of what science is and what we actually want out of it. What exactly is science for, and what are we trying to do — and is there a danger that we are optimizing the wrong thing when we point our AI tools at science?

The hiking analogy

One analogy I've given in the past is that science is a little bit like going on a hike. You've heard there's some interesting waterfall, some beautiful waterfall out there, so you decide to hike with some friends to find it, but you need to make a map. You get a little lost, but maybe while getting lost you discover something else interesting, and you make a note of it. On the way to this waterfall, you find an even more spectacular waterfall in the distance — you can't get there yet, but maybe some future hiker will figure out a way to get there too. So there's a whole process to get to your goal, which is also very valuable.

"These AI tools, they can be like helicopters that will just fly you directly to this waterfall, and you can see it, and then you fly back — but you learn nothing about how to get there."

You may not see any other interesting phenomena than the specific thing you asked for. And so even though technically you achieve your goal much more efficiently, there may be something that is lost.

How large language models actually work

Modern AIs are powered by a type of algorithm known as machine learning, which is trying to predict patterns in data. A very simple example of machine learning is regression — if you have some inputs and some outputs, like, let's say you observe that if you feed some animal more food, they get bigger. You can plot how much food you give various animals and plot their weight, and you get some dots on a graph, and if you're lucky they will fit some line, and that line becomes your prediction. So if you give this dog this much food, they would gain this much weight.

Now, in the real world, you don't always get these nice linear relationships. Often there are many inputs and many outputs, and the relationship can be really complicated. But sometimes the data has a shape, and we now have all kinds of clever ways to detect this shape and try to fit curves to these input-output pairs.

What large language models — which power chatbots and things like that — are doing is playing the game of naming the next word in a sentence. Roughly speaking, if I say "roses are red, violets are blank," what is the next word to fill the sentence? You can probably guess the answer is "blue." That's an output. You can imagine this giant plot where the inputs are all these incomplete sentences and the outputs are the words you want to complete, and you've got all these dots in this high-dimensional space, and you want to fit some curve that will try to explain what is the most likely word to come out. Sometimes there's more than one answer — "Hello, my name is" — there could be many names you could put after that. So you don't always get a single answer, but you could try to get the most plausible answer.

People have tried this — the autocomplete feature on your phone does this, you text something and it will suggest the next word, and sometimes it's kind of right, sometimes it's silly. Once you have any kind of operation like this, it creates some dynamics, and many people have just played on their phone, pressing autocomplete over and over again, but you get these gibberish sentences — you get monkeys typing on typewriters.

"The magic of large language models is that if you train them on enough data — trillions and trillions of data points — and you really try to fit as good a curve as possible, and this takes millions and millions of dollars of computing power and months and months of time, then suddenly, even when you iterate, it stays coherent. It begins to sound not like monkeys, but like a human speaking."

And somehow we don't fully understand why that's the case. But what seems to be true is that language — like English or other natural languages — contains a lot of hidden patterns that we're not consciously aware of. We know some of the laws of English — laws of grammar and things — but there are sort of unspoken, unwritten rules of language that humans pick up. A human child, even though they're not taught what a noun is or what a verb is, can pick up what order English words go in just by continual exposure to the language.

It seems like you can teach these models to also pick up patterns in language, to the point where you can give them math questions — "the answer to 2 plus 3 is," and they will say five. They have been trained to get the correct answer to at least simple math questions. Once you have a little bit of ability to speak English, you can kind of go in loops and check your work and make fewer mistakes, and you can prompt these models to proceed step by step and not say something unless it's been double-checked, and so forth. And so they become a little bit smarter, quote-unquote, to the point where they can solve many, many complicated tasks — but they're still just guessing the next word to say. It's not really grounded in any deep understanding of the real world. It's just that they've seen the patterns in the English language, or other languages, so well that they can mimic people who are speaking in an intelligent fashion, and they can present as being intelligent long enough that they can fool us — but long enough that they can actually do useful things.

So we can now solve certain math problems by asking the LLM to provide a proof, and sometimes the proof is complete rubbish, but if you loop it enough and have enough checks, you can actually start having a positive success rate.

"It's a very strange way of solving problems — completely orthogonal to the way we normally think of intelligence as being very grounded, methodical thinking from first principles. It's like having someone who knows a lot and is slightly drunk, and is sort of throwing out ideas, but with enough guidance you can actually extract useful output."

It's not the most advanced mathematics out there, actually, but you give it a lot of data and a lot of time and a lot of other band-aids and things, and it actually works pretty well.

AI vs. human mathematicians: breadth vs. depth

I find that some of the debate on AI's role in science and other disciplines defaults to a one-dimensional view of thinking — there's easy tasks and hard tasks and very hard tasks, and humans can do tasks up to a certain level, and AIs can do tasks up to a certain level, so which one is better. That's kind of a one-dimensional way of thinking. But what I found, when working with AIs and comparing their way of solving problems to humans' way of solving problems, is that they are really quite complementary.

Human experts will focus on depth — a human mathematician who will solve thousands and thousands of problems over their career, they will pick one or two problems that they think are difficult but not so difficult that they're impossible, but difficult enough that the exercise of trying to make a bit of progress toward them will reveal all kinds of insights that they can share, and maybe their students or other collaborators can build upon.

When we point AIs at really difficult problems, where none of the standard techniques apply, they are still very bad — they're just randomly guessing. But they excel at breadth. If you point them at a thousand problems of various difficulties, some may be just too hard, but there will be some that are actually within reach of existing methods, and there's some method out there in the literature that will solve your problem, or maybe you have to combine two separate methods. There just aren't enough human experts to look at all these problems, and the human experts who do look at them may not realize there was this obscure paper from a journal in 1970 that actually has the key idea that will solve this problem. They don't have the patience or the time to go through all the different combinations of which technique might work on which problem.

But the AIs will somewhat randomly take educated guesses as to what techniques might work for a problem, and some of them will be stupid, but some of them might work, and through all these combinations we're finding that sometimes they can catch a solution that the rest of the humans have missed. Occasionally the conventional wisdom of the experts is wrong — we all think a problem has a positive answer, but actually it has a negative answer, and we just didn't look at the negative case too much because everyone thought the answer was true. But an AI may not have that preconception. So sometimes the AI just serves as an independent pair of eyes, and some problems that we thought were very difficult had a surprisingly simple solution, which in retrospect maybe we should have gotten ourselves.

They're beginning to become successful when you point them at a very broad range of problems and they solve some percentage of them — maybe you point them at a thousand problems and they solve 5% of those problems. That's still 50 problems solved — you can already have tools that in some sense outperform human mathematicians by raw number of problems solved. Now, the 50 problems that get solved may not be the 50 problems you most want solved — they could be 50 random problems — but still it is very impressive. What I think we will have to do as a profession is find ways to incorporate this new capability to solve some problems at broad scale, and somehow figure out how to make that mesh with our existing capability to solve a few deep problems very slowly.

Kepler's lesson for the age of AI

Kepler's story of how he found his famous laws of motion is a fascinating one — it shows how important the process is. Kepler learned of Copernicus's theory of the motion of the planets. Copernicus had roughly worked out how far the Earth was from the Sun, how far Mars was, and so forth, and Kepler noticed that the ratios of these orbits looked a little bit like certain ratios that showed up in geometry. Eventually he proposed that if you take spheres — one sphere for every planet, and he had six planets known at the time — he could inscribe five Platonic solids, like a dodecahedron and a cube and a tetrahedron, between these six spheres, and get a perfect fit. This explained the shape of the solar system in terms of the five Platonic solids. This was his beautiful geometric idea.

It was only after he managed to get his hands on some really high-quality observational data from Tycho Brahe — which he had to fight for, and possibly even steal — and tried to fit his theory to this data, that he found it didn't actually quite fit, with the precision that Tycho's data offered — he could not quite get these spheres to fit. In fact, he discovered from that process that the orbits of Mars and Earth could not be circles at all — that there had to be some other shape. He spent many years figuring out what to do. I don't know how long he held on to this theory of the Platonic solids, but you can see in his writings he tried many other things — he tried to make the circles off-center, and at some point he landed on the ellipse, and then suddenly everything fit.

It does show that there is an interplay between theory and experiment — you can pose a theory, and if it doesn't fit the data, it may not be a good theory. But it's more complicated than that, too. Before Kepler, one of the criticisms of Copernicus's theory was that Copernicus himself acknowledged his measurements were worse than the best predictions available at the time. The best models were the geocentric models, developed by the Greeks and then the Arabs and Indians — there were many adjustments and fine-tuning, and they had a very precise model that could predict in a very complicated way where all the planets would be. Copernicus's model was worse. Just knowing agreement with data is not necessarily the only metric. It was only after Kepler found his revised model, where the orbits were not circles but ellipses, that the heliocentric model became more accurate than the geocentric model.

"If Kepler and Copernicus had AIs, and they asked them to predict a model for the universe, it could be that the AI that generated the correct heliocentric model would be discarded, because initially its predictions were not as good as the geocentric ones."

It takes time to really digest all these theories and see how they fit with everything else we know about planets and motion and gravity and everything. One concern, actually, is that AIs are too fast — there's a danger that these AIs will do what's called overfitting, and create a very complicated model that has nothing to do with what's actually going on, but just fits your data extremely well without extrapolating beyond that data set.

How we incorporate AI into the scientific discovery process will be a challenge. It can certainly accelerate individual steps of the process — you can make experimentation faster, write code faster, write your papers faster. But science as a whole may not necessarily accelerate just because every single component gets faster. There's a danger that we will optimize the wrong thing when we point AI at science, and we will, on paper, get all these amazing successes, but find out that science is not actually advancing in the way it used to. But we will find out it's still better to have these tools than not have them. We're still learning how to use them most efficiently.

Proof indigestion

Part of what we do is solve problems and try to find solutions to problems and prove things. Proofs go through a certain life cycle. First you have to generate a proof or solution — this used to be quite hard — but some of the proofs you generate are incorrect, so then you have to verify them, check which ones are correct, and that also used to be quite tedious. But both of these tasks are becoming more and more automated. So we're beginning to see more and more proposed solutions to various problems, and many of them are actually correct, but proofs are also getting longer, and when they're written by AIs they are often not very pleasant to read. An AI-generated proof might spend a lot of time talking about something very trivial and very little time on the most interesting portion of the paper — I think because the AI can't distinguish what is hard, what is difficult, because by brute force, everything takes the same amount of time for them. A human who has naturally had to struggle at the most difficult step of a paper would naturally spend a lot of time on that step.

You need to write out the paper, and a proof, in a way that reads well and can be explained to other people, and then other people have to get excited by it — they have to accept it as, "Oh, this is really interesting, this will help me solve my own problems," or, "This really clarifies why this phenomenon was true." It has to be accepted, and this is where we traditionally have the peer review process — we send papers to referees, and if the referees are excited by the result, the paper gets accepted. But there could be papers that are technically correct and readable and fine, but answering a question that no one cares about. And then finally it needs to be completely polished and put into textbooks and taught to students. Often the first version of a proof is not suitable for writing into textbooks — it's often done in a very inefficient way, and the ordering of steps is not quite logical. There's a certain digestion process, where someone has to spend a lot of time thinking very hard about what is completely the right way to organize and edit the paper, so it flows in the same way — a little bit like how you would edit a documentary or a movie.

So what we're finding is that AI tools are accelerating the early stages of this process, but not the late stages. We are now generating many proofs, verifying a bunch of them, but the pace of understanding them and putting them into final textbook form is still done by humans. In fact, we're now experiencing what you might call proof indigestion, where suddenly there are lots and lots of pending solutions to problems that should be understood and should go into textbooks, but we're just flooded now with too many of them, and we have to pick and we have to triage. This is something that has never had to happen before. It used to be that solutions came out so rarely that if a solution to a major problem got solved, all the experts would sort of drop everything and read it and try to digest it as quickly as possible, because it was so rare and so valuable that it was worth doing. And now we're just getting flooded with all these possible solutions. I myself have had to stop trying to stay current with all the latest developments in my field — sometimes there's just so much going on, I can't promise now to read every single development that shows up. This was already beginning to be a problem before AI, but AI has really accelerated the sheer volume of content being generated. So we're going to need much better curation and filtering. It's a good problem to have — it's better to have too much food, more food than you can eat, than not enough food to eat. But it is still a problem.

AI's rapid progress in mathematics

AIs have become increasingly capable in mathematics. For people like me who have been following the developments for the last three years, there's been a kind of steady progression — four years ago they could solve middle school math problems, then high school math problems, then high school Olympiad-level problems, then some graduate-student-level qualifying exam problems, and then they started solving a few of the minor unsolved problems that maybe someone like Paul Erdős would have proposed but no one really looked at — a lot of low-hanging fruit. And then, just recently, there have been one or two occasions where they managed to solve some problems that people actually really did try very hard to solve. Somehow, collectively, all the humans were taking the wrong turn, and the AIs, which had a different set of biases, managed to cobble together a solution that was quite clever and has already had some impact — there have been some nearby problems to the unit distance problem, for instance, which have also been solved by humans who adapted the AI's technique. I found that quite exciting.

For some of my colleagues it was very concerning, especially if they hadn't been following the previous developments and didn't realize where things were at. If a colleague had only seen what ChatGPT could do in 2023, and if you asked it a difficult math question, it would give you complete rubbish — they are quite different now. It's still fundamentally the same technology, but they have found ways to reduce the error rate and become genuinely useful.

Now, it's still unclear how replicable this is. Many of these achievements are done by private companies, who are not disclosing how much resources they spent — was it $100,000 of compute? A million dollars? We don't really know, and we don't know their success rate — was this problem that they solved the only problem they looked at, or did they look at 10 problems, 100 problems? While the results are impressive, we don't have enough data to really gauge whether this will become a completely regular occurrence going forward, or whether it's only if you spend $100,000 over several months with a team of ten people that you can get results like this. Maybe it's only 1% of all the problems we care about that are amenable to this method — we don't know. But there are efforts to more properly benchmark this in a scientific way. The most recent of these challenges is called the FrontierMath challenge — they tested the latest models against a test set of 10 research-level questions, and the best models could do like five or six out of ten of these medium-level-difficulty math problems, which already had a solution, but the solution was kept secret.

There's a lot of routine tasks that we do every day in our research, and some percentage of those can now be done by AIs. It could be expensive — many of these tools require a couple hundred dollars to run before they can get a solution, and sometimes they fail — they spend all this compute and end up with nothing useful. We are seeing now in programming that many expert programmers are reporting that their ability to write code has increased by a factor of five, or 10, or 100 with these tools. But they are also feeling themselves losing the ability to code by hand, and sometimes they cannot review the code that comes out of these agents. There's a trade-off — speed is not everything.

Collaboration, human and artificial

I very much like the collaborative aspect of mathematics. I didn't realize how important it was until relatively late in my career — a lot of the mathematics I've learned, I learned after grad school, by working with mathematicians and scientists in different fields. I teach them what I know, they teach me what they know, and I become much broader as a result.

When you're working with a collaborator you've been working with a long time, there's a point where you become almost mentally attuned — perhaps like if it's a really close friend or family member you talk to for a long time, sometimes you can complete each other's sentences, you know what the other person's going to say. And you can sometimes get that when you collaborate — you can throw out an idea, and before you even finish the sentence, the other person gets it and can run with it.

People have tried to use AIs like this, but AIs — you can't quite converse with them the same way. Of course they make mistakes, sometimes they're sycophantic and only tell you what you want to hear. Also, the interactions right now with these tools are kind of personal — I've tried to collaborate with people in person and also have an AI present, but it breaks up the rhythm. These tools aren't really conversational, not as fluid a conversation as with human collaborators, yet — maybe they will get there. Until recently they don't learn from your conversations — with a collaborator, you can resume the next day and pick up very quickly, and sometimes even for a colleague you haven't met for years, you can pick up some very old threads. AIs have a certain amount of context, and they can remember some things, and they can make notes and kind of simulate this memory, but you can't attune to an AI the same way you can to a really close collaboration — not yet.

I actually don't use these tools so much for the actual problem-solving process. I've found to date that these tools are much better at secondary tasks — like doing literature searches, or checking a proof, or writing some code, or proofreading something I wrote to see if there's any opportunity to make things a little bit tighter. The rhythm of working with an AI is not quite the rhythm I prefer working with a human collaborator — but that could just be the current state of current technology. Maybe future AIs will be much more conversational and much more human to interact with.

The risk to the next generation of scientists

We're at a somewhat risky point in the structure of funding the scientific enterprise in general, because on the one hand, these tools are allowing us to create the outputs of science — or what seems to be the outputs of science — at a much accelerated rate. But it could come at the cost of nurturing our seed corn for the next generation of scientists. For example, there is a real concern that the training problems we give our graduate students to work on, as their first projects, to get a little bit of recognition and career training and experience — these are the types of problems that AIs can now replicate many of the papers for. But if you replace the grad students with these AIs, the AIs will generate these grad-student-level papers, but then we won't get the next generation of students.

But if we don't continue this process of digesting all this AI output to build the next base of knowledge for the next generation of humans and AIs to build upon, we may end up stagnating as a scientific society — we'll be able to optimize everything we can do with our current technology, but we may not actually develop really original new ideas anymore.

We will need to really have a much more open discussion about what basic science is and what it's useful for, and why it's important to still have curiosity-driven research — why we still need a community of humans to explore things, sometimes slowly, sometimes in ways that are not as efficient as the latest AI models, and also the insights that we gain. I think we should share them more, and we should do more outreach to the general public. I think the general public today can see the visible outputs of science — they have a cell phone, they have the internet, they have GPS — but many people don't see the whole process, and how a basic understanding of math and science actually makes the world around them a lot less scary, and just a lot clearer. I think a lot of people now are just living in a state of anxiety — the world is so complicated, and we haven't emphasized these softer values of science as much as the hard, technological outputs. But science does add a certain amount of clarity to one's thinking, and these are valuable things, and they need to be supported.

Source: Steph McClain · bigthink.com