
The Quanta magazine podcast has an excellent episode this week titled, Is AI Reasoning Right for the Wrong Reasons? I don’t want to get too far into the details of it, but one statement stuck out to me. This isn’t an exact quote, but it was along these lines:
The podcast cites a paper published by Apple Research where the investigators replaced intermediate explanations with wrong assertions and the model still arrived at the correct answer in the end.
I found this fascinating, though not for the reasons given in the podcast. Rather, it called to mind my experiences both as a high school student and as a teacher of mathematics.
First, as a student. Sadly, I am somewhat unimaginative when it comes to getting work done: you give me the problem and, unless it is simply impossible to complete in a reasonable time with the tools I have at hand, I will work hard at it until I come to the answer, trying to justify each step with logical, convincing reasoning. As I sometimes say about my PhD in Mathematics, “I don’t have a PhD because I’m smart. I have a PhD because I was too dumb to quit.”
Compare this to a couple of my friends. They’d be the first to say that they didn’t work hard in school. (That’s what they used to say, anyway.) One of them liked to say that he did well in his classes not because he knew the material, but “because I’m a good bullshitter.” (Yes, that’s a verbatim quote.)
Next, as a teacher of mathematics, both in high school and at the university, I saw a similar phenomenon: students would often arrive at a correct answer, but even the good ones often failed to explain why it was correct. Press them on it and they'll mumble some jumble of nonsense. In other words, they have to be taught how to explain their reasoning, even if it’s correct.
The casual reader might wonder why it matters that a student explain how he arrived at his answer. After all, he has the correct answer; shouldn’t that suffice? Well, yes and no. In the strictest sense, if you’re designing a bridge and you intuit the correct result to a crucial calculation on which lives depend, then sure, it may not matter if you can explain it. In fact, it won’t matter at all, because no one will want to build it unless you can explain why your answer is correct. (It might frighten you how many construction engineering majors really didn’t care to learn how to do that.) When it comes to putting money on the line, never mind their lives, people can be a bit distrustful.
Even in something less critical like business, good luck bidding on a contract if you can’t explain how you arrived at the quoted cost. You might get away with it if you’re a relative or close friend of the contracting official, but in a free country with a solid respect for the rule of law and the free market you’re likely to face lawsuits as a consequence, so you’d better have factored that into your quote.
If the point isn’t clear, it’s that humans themselves tend to do precisely what the LLM’s are doing here: they mumble nonsense, hoping to stumble on something that will halfway make sense and that they can pass off as an answer. We don’t by nature go looking for The Truth; we go looking for whatever will get the people we care about to approve and/or esteem our answers. We have to be taught to reason; the rules of Aristotelean logic do not in fact come naturally to us.
And just like human beings, they’re often willing to change their minds and come around to your point of view if you, the authority they wish to please, challenges them well enough. I have more than once argued a prestigious LLM into changing its mind on an ethical matter.
In sum, are the LLM’s getting the right answers for the wrong reasons? I’d challenge the premise itself: how can you say the answer is correct when the reasons given are wrong? The reasoning is part of the answer. You may prefer to be lucky than to be right, but you can’t be lucky every time, and it’s more likely than not that you will be unlucky. Whereas if you know how to reason properly, then you will be correct every time you do so — even if the best you can arrive at is, “I don’t know.”
The only reason this isn’t understood is that our educational system has become so terribly degraded by the cult of self-esteem and left-wing indoctrination (or, at a select few universities, right-wing indoctrination) that we struggle even to teach math majors how to reason properly, because quite frankly, in my experience, mathematicians are the only people left using deductive reasoning as the method of discerning the truth. (Sadly, a lot of mathematicians also want to abandon that.)
I will concede something to the gathering mob: you can’t always know The Truth. In fact, that’s the case most of the time. The best you can have is various gradations of certainty: I am very certain that this is true, or only slightly certain that this is true. What might be the difference between those two?