Back to blog listing

The Fundamental Questions of the AI Revolution


Originally published on The Study Substack, November 1, 2024.

Not An Expert

I preface all this with the fact that I am not an expert. I’m not a research scientist working on pushing the boundaries of AI. Although, I wouldn’t hate doing that. It’s one of the few things I think I would actually enjoy doing as my day job. At least for a time.

Point being, I refuse to wait until I have the world’s deepest set of qualifications to talk about these ideas. Don’t put that on me, it isn’t fair! The things I want to talk about here are some of the most complicated things I’ve ever even heard of. If the metric for discussing any of this was to first achieve expertise in all the relevant topics, then I daresay no one on Earth has earned the right.

If my ā€œqualificationsā€ do matter to you at all: I have a degree in Computer Science, I’ve worked for over 11 years as a Software Engineer, and 3.5 years of that were at ā€œBig Techā€ (Meta, in this case). I don’t think it matters, but I was asked to add this by someone who was reading an early draft.

Since this is a relatively long read compared to most of my past writings, it may be worth mentioning my target audience. Really I’m writing this for myself as I push my understanding of these topics. This essay will likely be more interesting to those who are not already deeply familiar with the fundamental theory of AI. I expect real AI experts/researchers to pop a few of my assumptions like a balloon. As my friend recently told me when I mentioned some insecurities about posting my writing: ā€œPut yourself out there first, worry about it later.ā€

The (Potential) Limitations of AI

AI is obviously the ultimate hype in the modern era. I can’t imagine most relatively online people can go a single day right now without hearing or thinking about something AI-related. So much so that I don’t even see a purpose in breaking down what AI stands for as I write this (just in case, AI = Artificial Intelligence). And yet I feel, in the sphere of common conversation, there’s a massive gap. One that misses touching on the potential deep and fundamental limitations of AI.

My goal here is simply to lay down some thoughts and ideas that I think are useful for understanding what the hell is going on. There are a lot of people yelling very loudly about what is and isn’t possible with ā€œAIā€. The quotations there will hopefully make a lot of sense as I go on here.

Before going any further, I do want to make one caveat quite clear. What we commonly refer to as Artificial Intelligence (AI) is really just applied Machine Learning that uses Neural Networks. It can sometimes be difficult to properly track what definitions people are using for these terms in various conversations. So let’s break them down super clearly and simply for the purposes of this conversation.

  • Artificial Intelligence (AI) = The attempt to make computers do human-like stuff.
  • Machine Learning (ML) = One potential approach to achieving AI.
  • Neural Networks (NN) = One single ML approach among many.

All of what we commonly refer to as ā€œAIā€ in the modern era is really just Neural Networks. One single approach to one single subset of the entirety of AI. There could potentially be infinitely many different approaches to ML, and ML could potentially be just one of infinitely many different approaches to AI. So, as a baseline, let’s just pinpoint the fact that what we are all calling ā€œAIā€ is just one tiny (albeit extremely effective) little drop in a potentially massive ocean.

The Big Divide

We could spin our wheels all day on the common questions.

ā€œIs AI even useful?ā€ Well, I think it very obviously is, and anyone who doesn’t must be trying very hard to not be informed about what AI is capable of already. But my opinion here isn’t super relevant, and neither is yours. For each individual, they might be able to find something useful in the current offerings of the AI landscape, or they might not. Not a particularly useful question for what I’m trying to get at today.

What I really want to talk about is ā€œArtificial General Intelligenceā€, or AGI. AGI would be a very general type of AI. As opposed to a very specific type of AI. Not just something that can generate images, or something that just recommends movies to watch. An AI that can basically do anything that any other specialized AI could ever do. An AI that can learn to do any of the things doable by AI. The AI hype has inadvertently generated a big wall, whereupon advocates and detractors tend to fall into one of two hard camps. AGI advocates typically say outright ā€œAGI is possible and is coming soon!ā€. AGI haters typically say ā€œAGI is outright not possible and the AI hype is BS!ā€

If you are already bored and have to stop now, please walk away with just this one idea: the real, honest truth of AGI is that both camps I’ve described above are equally full of shit. Lots of people have opinions on whether or not AGI is possible, or when we will or won’t get to it. But the total and complete reality of the world we currently inhabit is this: no one actually has the real answer to this question!

To be extremely clear. AGI may or may not be possible. No one knows. NO ONE KNOWS. Not a single human being that has ever lived has definitive, hardcore, scientific evidence that AGI is or is not possible beyond a reasonable doubt. Literally no one. It is an unsolved research problem. Worse, it’s a research problem that may not even have an answer. The question of whether or not AGI is possible may itself be impossible, on a fundamental level, to actually answer.

If anyone ever says anything remotely like ā€œAGI is possibleā€ or ā€œAGI is not possibleā€, the only real possibility is that they are either uninformed, or intentionally being misleading. Of course, it may be ironic that I type these words as someone, somewhere, is preparing to drop the most intense mathematical proof showing that AGI is definitively possible or not. I look forward to being exposed as a fool, for a fool I am.

I am not an AI/ML expert. Not even close, not even remotely. The above truth (that no one actually knows) is the last truth I will try to push with any confidence. And of that particular truth, I am absolutely and unabashedly sure. From here, we will talk about what I find to be some neat ideas, mostly in the form of questions. Because we do not have to be experts to reason about a lot of these things. We can just ask some fundamental questions, ideally yes or no questions, and then think a bit about what either answer (yes or no) might mean.

Unfortunately, I do need to rattle off a few things first, and it might take a minute. Hopefully you find it all interesting enough that it gets you excited to read more about some of it. These aren’t my ideas, which is why I’m not structuring them as questions quite yet. For now I just need to lay down some widely accepted truths in the world of modern math and science.

Limited Results, Everywhere

You may or may not be aware that there are massive holes in both Mathematics and Computer Science. Without getting too bogged down in the details, both systems are broken. And unlike how the possibility of AGI is definitely not proven, the broken nature of Math and Computation is definitely proven.

As I discuss these limitations I won’t be giving concrete or thorough examples, mostly for simplicity and brevity. I’ll try and include a big list of references at the end so that an interested reader can push further in trying to understand these things. For now, let’s keep things as high level as possible.

Math Is Broken

Gƶdel’s Incompleteness Theorem is a Mathematical finding from the early 1930’s that says (oversimplifying): ā€œAny formal Math system will have statements which are true but can never be proven to be true.ā€

We often think of Math as being this extremely precise thing where statements either are TRUE or FALSE. And really, it is. Every Mathematical statement does have what we call a truth value. x = 5 is a statement. And we can ask, in some situation, is this true? Does x actually equal 5? For a beautifully large portion of mathematics, we can figure out the answer to that question. It either is or it isn’t, and we can prove it. How do we prove it? Using math, of course!

The problem is, some of those statements can NEVER BE PROVEN to be either true or false. That doesn’t mean we can’t feel really strongly that the statement is true. It could absolutely be true, with us stranded on a big island of unprovability, cursed with intuition but unable to ever possibly get a definitive answer.

This isn’t some weird edge case of mathematics, it is a fundamental property of the entire thing, inside and out. Math is fundamentally broken, and it always will be.

Computer Science Is Broken

Computer Science fares no better. Shortly after Gƶdel cursed math as broken, Alan Turing essentially formalized computer science AND broke it, back to back.

There’s an idea in computer science called Computability. What it means is that, just like our unprovable math statements, there are things we might want to compute, but literally just cannot. This is completely distinct from the idea that certain things are easy or hard to compute. Easiness vs Hardness is what we call Complexity Theory. But complexity only applies to things we actually can compute, which is only a subset of all the things we might want to compute.

Computer Science is broken, and it always will be.

But We Still Have Technology…

So math is broken, and computing is broken, but we still have all this amazing tech, right? Yes. YouTube and all the stuff behind it falls into the category of things that are totally computable. That’s hype, I love YouTube, and the fact that it IS possible makes me happy. Same with videogames, warehousing and inventory software, accounting software, banking software, operating systems, and even AI. So it’s not all bad news. Even with these fundamentally broken and limited systems, we can still do a lot of things.

AI might be just like that. Imagine some big list of all the human-like things we might want AI to be capable of. Some of those things might be literally impossible. We, as a species, don’t really know the answer to that question.

Quantum Fuzziness

There’s another idea, this time in Physics, called Quantum Mechanics. What Quantum Mechanics states is that reality, the universe we inhabit, at the tiniest levels of matter, is really strange. When you have big stuff like planets and stars, or people and cars, they behave in a structured way (well, maybe not people). However, according to quantum mechanics, when you have very small stuff, like an electron, it doesn’t.

It’s not that we don’t know how it behaves, or that we can’t prove or observe how it behaves. It just doesn’t behave in a very structured way. There’s a baked in randomness, or fuzziness. The Heisenberg Uncertainty Principle says that it is impossible to know a certain amount of that state of small things at any point in time. Impossible, not as in beyond us or our capabilities as a species right now, but as in completely and totally fundamentally not a thing. All of reality, at the smallest levels, is just ablaze, boiling, roiling, with this randomness. All the time, everywhere.

The simple example, although it won’t sound so simple if you aren’t super into physics, is the position and momentum of an electron. Electrons have positions. Electrons have momentums (they move). They cannot, according to Heisenberg, have a DEFINITE position AND a DEFINITE momentum. If the position is 100% decided, the momentum could be any value, and vice versa. There’s a fixed amount of ā€œdefinitenessā€ in the system. The more certain the position, the less certain the momentum. Maybe you kinda know both, 50/50. Or you are pretty sure of the position, but have no idea about the momentum. And so on.

To quote the guy in the MIT OpenCourseWare lectures on Quantum Mechanics: ā€œThe universe really behaves this way. It’s not strange that electrons have this fuzziness. That’s just how the universe really is. What’s incredible is that if you take 10^27 electrons together they behave like cheese!ā€

Quantum Mechanics is the result of a big revolution in physics in the early part of the 20th century (see: The Thirty Years that Shook Physics). A lot of important realizations happened back to back to back. Physicists are still trying desperately to make sense of what it all actually means. One of the problems we face even talking about it here is that there is no consensus about what it all means. There are even multiple ā€œinterpretationsā€ of quantum mechanics, ranging from ā€œthere’s randomness at the smallest levels of realityā€ to ā€œthe multiverse is realā€.

Why am I talking about Quantum Mechanics in what you thought was an essay about AI? How can these two things possibly be related?

Something that is computational is something that follows a rigid process. This process might be extremely simple or extremely complicated, or anything in between. No matter where on the spectrum the complexity of the process lies, it lies somewhere. This rigid process is called an algorithm, which is really just a set of mechanical steps. ā€œDo A, then do B, then add the results of A and B to get C, etc.ā€ You could write an algorithm to build a house. Or bake cookies. Or find the shortest path of flights from San Francisco to Atlanta. If a process can be structured in such a rigid way, then that process is computational. Being able to be structured in such a rigid way is what computational means.

If reality is a big ball of fuzzy, uncertain, random stuff happening, then it couldn’t possibly be articulated as one of these rigid processes. Imagine trying to write a little computer simulation of some of this quantum weirdness, like what’s happening in an electron. You couldn’t actually do it, because you would need some of this baked in randomness in the simulation to even get close.

You might be saying, wait, can’t computers generate random numbers? Couldn’t we use that to do the simulation of the quantum weirdness?

No! They really can’t, actually. Well, they kind of can. But how are those random numbers being generated? By a computer! Computers are, by definition, only capable of computing things that are computable. If a computer is doing something, literally anything that it does, that thing is a ā€œcomputational thingā€. The process of generating seemingly random numbers with a computer is still one of these ā€œrigid processesā€ we just described.

And this, my friends, is where we start to get the questions I promised so long ago.

Is Reality Computational?

If quantum mechanics is right, then reality has this baked in randomness and is therefore not computational. Rather than assuming we know, let’s pose it as a question and just reason about the ramifications of potential answers. Remember, even the physicists, a group of people that have been studying quantum mechanics (and making quite good progress) for the last 100 years, don’t agree.

If quantum mechanics IS computational, then we really don’t have much to worry about. The physicists got it wrong! We don’t need the randomness stuff, we just need the physicists to figure out what’s really going on. Then we can write our little simulation and everything is copacetic again. No stress, no tears, no randomness.

If quantum mechanics IS NOT computational, then stuff gets spooky again. Now we have this division into two groups. Computational things (like the stuff that happens in a computer) and NON-Computational things (like whatever the hell is going on in quantum systems). If that’s the case, then reality itself is not a computational process.

Quick sidebar: I first got a lot of the ideas that follow from Stephen Wolfram’s short book What is ChatGPT doing?. More on that in the outgoing thoughts at the end of this essay.

WTF is ChatGPT doing?

ChatGPT is an excellent example of some AI stuff that seems to be going quite well. You can jump on and ask it questions, or have a conversation, and it generally seems to produce good output. We all know about some of the issues that generative AI has. Extra toes in that image of Gandalf riding a unicorn turkey dragon. An AI-generated song that has parakeets squawking in the background instead of a hi-hat. Or ChatGPT, making up new words.

Despite all these issues with GenAI, it still does a pretty damn good job at some things. A convincing job I might say! How does ChatGPT do what it does? It’s trying to simulate human language. After all, AI is all about trying to make machines do human-like things. In the case of ChatGPT, the human-like thing is writing stuff. So we ask the question: Is writing stuff (language) a computational process?

What if Language is Computational?

Let’s say it is a computational process. Whatever is going on in our brains when we speak or read or write is just a biological computation, albeit an extremely, mind-explodingly complex one that I don’t understand. It’s a rigid process. And then the reason ChatGPT can do it is because ChatGPT itself is a computation process, running on computational machines.

It makes sense that it can do language-y stuff if language is computational. In this case, we have the hope that we can probably figure out the real computation laws behind language and come up with a much more energy efficient way to do it. Rather than doing billions and billions of linear algebra computations per word ChatGPT is generating, maybe one day we can code up a simple little function that efficiently produces essentially the same results (or better) as what ChatGPT does using the GDP of a small nation worth of energy costs daily.

After all, you could build a super crazy complicated Neural Net that gives you the results of the function x + 5. It probably wouldn’t ever be perfect for all possible values of x, but it might get pretty close. Although, why would you do that when you can just write a simple function that returns the exact value in all possible cases? It helps to know how to write the function! We definitely don’t know how to write the simple computational version of ChatGPT, at least not yet.

It could also be that language is computational, and ChatGPT does it well for this reason, AND Neural Nets are really the only way to do it well. In other words, maybe we really do need a Dyson Sphere for GPT-6.

What if it isn’t?

But what if it isn’t? Let’s say we take quantum fuzziness at its word. Randomness exists, non-computational stuff exists. Let’s say, then, that language is one of these non-computational fuzzy things. Then how is ChatGPT able to do what it’s doing?

Well, maybe ChatGPT is just approximating. Maybe it isn’t really doing the whole language thing. It just does a decent enough job that we kind of perceive it as doing the real thing. It does these billions and billions of computations on a GPU and we get a word back, and it turns out that the word is good enough that we kind of buy the whole process, even if it isn’t real language stuff.

What if it’s actually super spooky?

Okay, but. For the sake of exposing where the real spooky shit would be. Let’s say all of the following are true. Quantum Fuzziness is real. Non computational stuff is real. Language is one of these fuzzy non-computational things. AND ChatGPT is actually really doing the real, non-computational, fuzzy thing.

If that doesn’t give you goosebumps, think back to one of the earlier things I laid out about computation. ChatGPT runs on computers (specifically GPUs, mostly, but computers nonetheless). These machines can only do computational things. There’s no randomness, no fuzziness. Somehow, someway, these neural nets are pushing past that limitation, and doing this crazy non-computational thing with nothing but computation.

If that’s what’s happening, we are in for a whole world of crazy stuff on the horizon.

Is Consciousness Computational?

Here we are at last. Maybe the real big question. When we talk about AI or AGI, what do we really mean by ā€œintelligenceā€? It’s poorly defined, no matter where you look. Does it just mean ā€œseeming to do human-like thingsā€? Or does it mean, specifically, ā€œReally doing human-like thingsā€?

In the case of AGI, I think people typically mean ā€œCan really do ALL the human-like things.ā€ Pick a definition you like and stick with it as you think about these things. If you’re reading someone else’s thoughts, maybe start by being really sure what you think they mean. You may find that they don’t know themselves.

Furthermore, what’s the relationship of intelligence and ā€œconsciousnessā€? Is consciousness a superset of intelligence, whereby intelligence is a prerequisite for consciousness? Or is there no connection necessary? Can something be intelligent but not conscious? Can something be conscious but not intelligent? Are all humans both? Does IQ matter? Is a human intelligent, in the way we mean intelligence when we say AI or AGI, regardless of how ā€œsmartā€ they are?

I won’t beat this to death, but even a cursory glance shows us that the language here is fighting against us. It seems difficult to classify a baseline definition for these key terms. My personal definition of Intelligence in AI is mostly what I also think ChatGPT is really doing, approximating a more complicated human behavior.

Roger Penrose famously predicted that the Quantum Fuzziness we talked about earlier is the actual thing that brings about consciousness. That the fuzzy, random stuff is what separates consciousness (a therefore non-computational process) from the other totally computational stuff. If ā€œAIā€ is just Neural Nets, and Neural Nets are just computational stuff, and consciousness is ā€œnon-computationalā€ stuff, then AI will at best only ever approximate consciousness.

But again, we don’t know. We really don’t know. We truly do not know. Not me, not you, not anyone.

You might be tempted to use this all as fuel for an argument that we don’t actually need to be concerned about AI with respect to the whole doomsday scenario where AI takes over. I actually think the two are entirely separate. I think a sufficiently smart AGI system could probably do great harm, or great good.

What do I mean by ā€œsufficiently smartā€? Well, let’s look at the whole ā€œcomputabilityā€ stuff from earlier. Not every mathy thing is truly mathable. Not every computy thing is truly computable. Maybe not every AI-y thing is truly AI-able. So let’s say that sufficiently smart AGI is one that can learn, quickly, to approximate any of the human-like stuff that is computationally approximate-able. We don’t need to jump over any spooky, fuzzy boundaries to get there. If an AGI could do all the human-like stuff possible as well as ChatGPT does language, girlfriend we screwed.

Parting Thoughts

Monkey man finish long essay. Monkey man tired. Monkey man try to excite you with neat, shiny ideas. Monkey man hope he succeed.

I’ve taken a lot of these ideas from other places. I’d be lying if I said I had a perfectly formatted list of references to pull from right this second.

What I will explicitly not do is apologize endlessly for some poor characterization or analogy I’ve made, or for saying things before being an expert. As I laid out as my first premise, normal humans have the right to talk about these things. Expertise notwithstanding. That said, I would love to be corrected if I’ve made a mistake, or have someone point me towards some new and relevant information. Please leave a comment if you spot something!

Resources and References

General AI and AGI

  • Artificial Intelligence - Wikipedia
  • Artificial General Intelligence - Wikipedia
  • Stuart Russell’s Artificial Intelligence: A Modern Approach - a standard textbook on AI, covering foundational topics and modern approaches
  • Lex Fridman Podcast with Nick Bostrom on AGI and the Future - in-depth discussion on AGI possibilities and challenges
  • Stephen Wolfram - What is ChatGPT Doing? - a detailed but accessible explanation of ChatGPT’s inner workings

Machine Learning and Neural Networks

  • Machine Learning - Wikipedia
  • Neural Networks - Wikipedia
  • Deep Learning Specialization on Coursera by Andrew Ng - covers neural networks, deep learning, and their applications
  • 3Blue1Brown: Neural Networks Series (YouTube) - visual introduction to neural networks and backpropagation
  • The Universal Approximation Theorem - Wikipedia - the theorem showing the theoretical power of neural networks

Gƶdel’s Incompleteness Theorem

  • Gƶdel’s Incompleteness Theorems - Wikipedia
  • Stanford Encyclopedia of Philosophy - ā€œGƶdel’s Incompleteness Theoremsā€
  • Gƶdel, Escher, Bach: An Eternal Golden Braid by Douglas Hofstadter - classic exploration of Gƶdel’s ideas and their intersections with art, music, and cognition
  • Numberphile: Gƶdel’s Incompleteness Theorem Explained (YouTube)

Computability and Turing’s Work

  • Computability Theory - Wikipedia
  • The Halting Problem - Wikipedia
  • ā€œOn Computable Numbers, with an Application to the Entscheidungsproblemā€ by Alan Turing - Turing’s original paper defining computability and the limits of computation
  • Computerphile: The Halting Problem (YouTube)
  • Introduction to the Theory of Computation by Michael Sipser - a popular textbook on computability, complexity, and automata theory

Quantum Mechanics and Quantum Fuzziness

  • Quantum Mechanics - Wikipedia
  • The Heisenberg Uncertainty Principle - Wikipedia
  • The Thirty Years that Shook Physics by George Gamow - a classic introduction to the revolutionary concepts in quantum mechanics
  • MIT OpenCourseWare: Quantum Physics I (Lectures)
  • PBS Space Time: ā€œIs Quantum Randomness Real?ā€ (YouTube)

Philosophy of AI and Consciousness

  • Consciousness - Wikipedia
  • The Turing Test - Wikipedia
  • Stanford Encyclopedia of Philosophy - ā€œPhilosophy of Artificial Intelligenceā€
  • The Emperor’s New Mind by Roger Penrose - Penrose’s arguments against AI consciousness based on quantum mechanics
  • Lex Fridman Podcast with Roger Penrose on Quantum Consciousness
  • Mind and Cosmos by Thomas Nagel - a philosophical exploration questioning whether consciousness can be explained by science alone