AGI-26 Conference | Day 1 | Keynotes and Paper Presentations
S32: Thank you, Hayley. On behalf of the artificial general intelligence slash AGI society, welcome to the 19th annual artificial general intelligence conference here in beautiful San Francisco Bay Area at the San Francisco State University. This morning, as I was walking around campus, I came across a reason why we’re working so diligently towards AGI, which is to help us as a species solve problems that so far have eluded us. And I came across a remembrance garden, and it immediately just felt something special there. And what the garden was, it was a memorial to the 19 San Francisco State University students who were interned, incarcerated, as part of the roundup of Japanese Americans in World War II and put into internment camps. (..) And it was a beautiful setting, 10 stones representing the 10 internment camps around the U.S. So we’ve been unable to solve these problems, pestilence, disease, we’re making progress there. But war, war is raging around the world right now. These are the reasons why we’re working so hard and so diligently and coming together here today. (.) The other thing I want to point out is I was, towards this endeavor, we need all kinds of voices. (.) I was looking around the room yesterday and compared to the first few years of the conference series, I noticed a remarkable diversity of people, ideas, and we need all of these voices. And Ben and Andrew are going to talk about a rebooting of the AGI society in a bit. (.) The other thing I noticed, which is a new energy coming from the new generation of researchers, the youth coming up. And we need to welcome, and that is the future path towards AGI. As we, as we, as we get there in this time when it, feeling like everything is beginning to crystallize. (.) And with that, I’ll leave it to the, Ben Gertzel, the, the chairman of the series and, uh, world famous.
S03: Thanks, Matt. (.) It’s interesting you mentioned Japanese, interesting you mentioned Japanese internment camps. Actually, my, my father was born in a Japanese internment camp in, in California. not because he was Japanese, but because my, my grandfather was a psychologist who had the very depressing job of convincing the Japanese internees not to get too depressed about being locked up and have all their stuff taken away. So we, we took their stuff away and locked them up, but we did hire a psychologist to try to, try to keep them from basically committing suicide because of it, which, which, which was my grandpa’s job, which he didn’t enjoy very much, but felt it was a useful thing to do. Though he, I mean, he opposed the, the camp. So my, my dad’s first words were in Japanese reportedly, although he, he forgot them a long time ago. But now, but now two of my kids are fluent in Japanese. So things, things, things, things all spiral around somehow. (..) Yeah. (.) Matt’s reflections do make one reflect on the, the reasons for working on, on AGI. (.) And, and there are, there, there are many of them, right? And I’d say my, my own reasons for working on it are multiple and have probably shifted in the, in the many decades I’ve, I’ve been pushing on, on this topic. I, I think at first it was pretty much pure intellectual curiosity. Like I, I was mystified how my mind worked or anyone else’s mind worked. I, I read science fiction where you had robots thinking like people and it, it, it, it seemed like, well, it would be really cool to understand what is this actual thinking process, right? And, and, and, and, you know, just like I wanted to understand how bicycles work, building a bicycle seemed like a, a, a good way to do it. (.) So it was pretty much pure intellectual curiosity. And, you know, I remember when Zara, my oldest son, who many of you know, who’s also working on AGI. (..) He’s not here, here this time, but, but he’s collaborating with us on, on Hyperion Project. But he, when he was 15 or 16, he started going through all the neuroscience books I had. Then he started asking me for better ones. Finally, after like six months, he’s like, but this is all a bunch of BS. None of it tells you how thinking works, right? Like, well, yeah, that’s the state of, the state of current neuroscience, right? But we have all these details and no one knows how thinking works. So that was, that was my original passion for the field, which I think was the case with many in my generation of the generations before. But after doing a variety of applied AI projects in areas from AI for genomics and, and, and biomedicine, AI for cybersecurity, AI for scientific discovery, AI for robotics. Yeah, my, motivation started to veer more toward the sort of thing Matt was talking about. Like, you can see there’s so much potential to do good in the world by advancing AI to, to a higher and higher level. Like, I mean, a lot of causes of death and disease, disease could be cured, right? We, we, we see now the right AI is thinking in the right way could solve cybersecurity problems caused by humans and, and, and, and by AIs, right? I mean, In terms of their cognitive architecture, in terms of their motivational systems, their design and their relation to the world around them. So, yeah, I think curiosity about how the mind works is a huge motivator for working on AGI, potential to do good for the world, as Matt has alluded, another huge motivation. Now, coming into 2026, we see AGI is tremendously more well-accepted and more popular than at any point in the past. I mean, as I recounted in my introduction yesterday on the workshop day, when we did the first AGI workshop in 2006 and then the first AGI conference in 2008 at the University of Memphis under the guidance of Stan Franklin, I mean, at that point, it was a fairly maverick thing to be doing, right? I mean, it wasn’t like unwashed masses in an escape from prison or something, it was professors from decently known universities and companies. But you were walking sort of on the edge of acceptability and working, you had to frame your aspiration to make real thinking machines in a more conservative language, right? Like, you’re working on, you can work on cognitive architectures and machine learning, you can even work on lifelong learning or transfer learning or something, but human level AI or AGI had to be framed very carefully and it’s sort of like how you couldn’t in biology work on longevity until recently, you had to work on age associated disease instead, right? So, clearly we’re well beyond that phase now, right? We’re in an era where national leaders and CEOs of huge corporations are embracing AGI, sometimes to an absurd degree, where several times in the last few years you’ve had a major CEO say, we’re going to achieve human level AGI this year. And we had Sam Altman a couple of days ago say, we’re in the singularity right now, right? So, I’m in the very strange position where there are influential people who are even more optimistic and saying even crazier things about AGI than I am, right? Which is very unfamiliar and disorienting, right? I mean, personally, I do think we may be, you know, three, five years, could be even one year, like, we could be quite close to making a human level AGI. Progress is insane. On the other hand, even though I think progress is very fast and we could be very close, like, it’s clear in my head, we are not there yet, right? And I do think the notion of a singularity is meaningful, like, we can have super intelligence much smarter than any human being. You can have your phone making ten Nobel Prize-level discoveries in a minute, like, which is roughly what Ray Kurzweil and Werner Vinge meant by the singularity. Like, amazing, like, we could be ten years, we could be three years from that point. Nonetheless, there’s a meaningful sense in which we are not to that point, even though we’re potentially very close to that point, right? So, but it is a different time, right? Instead of having to convince people that we might be able to create AGI in a singularity soon, I have to clarify to people that we don’t have them quite yet, even though they might be close, right? So, and going back to motivations for creating AGI, we now have a new entrant to the categories of motivations for working on AGI, right? Like, you have a bunch of people interested in AGI because they think they can make insane amounts of money for making AGI. (..) You know, achieve military hegemony over some other party they want to rule over, right? And so, I think it takes a lot of money to train scalable AI systems and totally nothing wrong with making money from AI or AGI work. I mean, we want to power commercial products with our AI systems and you want to bring in revenue from these products that can then be used to help people, to give ROI to investors and to fuel more AGI work, right? And, of course, given the world as it is now, there’s nothing wrong with national defense either. I mean, I worked on that stuff for some time earlier in my career. There are nasty people and bad things you want to avert. AI can help with that. But I do think alongside these motivations that are popping to the fore now as AGI gets serious, it’s important to keep the first two motivations I mentioned top of mind, which is curiosity, like understanding how the mind works and why, and then the potential of AGI to do tremendous amounts of good for the world. Like, these motivations should certainly be kept to the fore even as, you know, money and defense inevitably, you know, join them among the reasons to work on this stuff. And pursued by this combination of motivations, the growing AGI community is making really quite tremendous progress recently. I mean, each year since we started the AGI conference, it seems things are going forward and we have new things, things are getting better and better. But I’ve never had such a feeling of tremendously rapid progress as I do now. And I mean, even, like, within my own team in Hyperon, BGR Labs and SingularityNet, like, since we submitted papers to this conference a few months ago, and now there’s been tremendous progress on a whole bunch of different meaningful fronts, both with work with predictive coding neural nets, with probabilistic reasoning in Hyperon, with hives of Omega Claw agent systems cooperating to get work done. Like, in the three months between writing papers for the conference and coming to the conference, it’s like what used to be four or five years of progress have occurred just within our own team. And then there’s a whole bunch of cool stuff from other teams and other projects being reported. And, you know, we have one of the members of our blockchain project, SingularityNet, who works on the business side rather than on the tech side. He’s been vibe coding by telling this hive of Omega Claw agents, which are agents running a mix of neural nets and Hyperon symbolic AI, (.) shepherding this hive of AGI, not AGI, shepherding this hive of proto-AGI Omega Claw agents to basically read all the papers we’ve posted on Hyperon, implement them in our meta language, and kind of glom them together into an integrated AI system, while using Claude Fable to make tests for how smart the thing is, right? And I don’t think that particular open source conglomeration is going to be the AGI. I think there will be a lot to learn from it, but it’s certainly a new era where someone with super bright, but with high school level programming ability can, you know, converse with a hive of neural symbolic AI agents, and over the period of a month or two, get this proto-AGI system combining ideas from three dozen research papers to actually work and do something, even if not quite perfect, right? So we’re in an era of accelerated progress, but I think we’re not there yet, in spite of what some CEOs might tell you. And I think we might be able to get there just by coaxing herds of agents to implement stuff that’s already published, but I’m suspecting it’s not quite that close yet, although it may be very close in a historical sense. (.) I’m expecting that there’s still more contribution of human imagination and human wisdom that’s going to be, and human cooperation that’s going to be needed to get us over the line to finally produce AGI. And I mean, for that reason, you know, pulling people together like this is very, very important. And I think it’s also important given the siloing off of the commercial AI world that we see now, right? I mean, you see more and more brilliant AI people sort of joining one or another proprietary effort, (..) sometimes oriented toward AGI, sometimes just AI products. And, you know, it totally makes sense. There’s huge amounts of compute in commercial AI companies, and it’s the opportunity to work with brilliant people on fascinating projects. On the other hand, as a still recovering academic and long-time open-source community guy, I like the idea of people being able to communicate openly on what they’re doing and what they’re thinking. I mean, you can do that on the interwebs all the time, but kind of jamming together face-to-face and discussing ideas, projects, results, failures, and so forth, seems still to have a high value even given the slower bandwidth of human conversation compared with interagent online chat, right? So, finally, I want to note, while in these introductory remarks I’ve given you a dose of world according to Ben Gerstle, we are dedicated in the AGI conference series to having a huge diversity of perspectives, and you’ll note quite a few keynote speakers where I probably don’t agree with anything that they say in their entire keynote, right? Which is, I think, as it should be in this sort of gathering. So, we’re open to symbolic people, neural net people. We have a couple of keynote speakers who are sort of biological fundamentalists and believe that intelligence and consciousness lie in particular molecular processes in the human brain. I mean, I think we want to get a whole lot of perspectives together and kind of stir the pot and see what interesting ideas can congeal out of it. And so, that said, I’m excited to be convening the main conference portion of AGI 26. We had a great workshop day yesterday. Before the technical meet begins, I’m going to invite Andrew Commendo up, and he’s going to talk for just a few minutes about what we’re doing with the AGI Society, which is the formal organization that we created quite some time ago to host the AGI conferences. So, the first AGI workshop and probably the first AGI conference, I just kind of ran myself, I think, formally out of my AI consulting company, Novamente. But at some point in the early days, we created an official nonprofit organization, the AGI Society, just to have a bank account to operate for the conference series. And the AGI Society, it also publishes a journal edited by Pei Wang, the AGI journal. So, we’ve been operating AGI Society, mostly Matt Eclay has and me known then as, I think Stephen Reed was involved for some time. (.) Basically, an operating organization to keep the conferences and journal going. But we had for a while the notion of membership in the AGI Society, and we let people sign up as members, but we didn’t really do much with our membership list, because we were just too busy with life in general. But it seemed that now that AGI is becoming such a more popular thing, and potentially, at least according to some of us, we may be even close to achieving it. It seemed like a timely time to try to turn the AGI Society into more of a real, real professional society. There’s just so much more interest in AGI. It seemed like there should be some organization pulling together people with an intense interest in AGI beyond the conference, which happens every year. And so, yeah, my friend Andrew Commendo, who I knew through his work in robotics, and our common long-term interest in AGI. We used to both live in the D.C. area. (.) He has, thankfully, he’s been eager to jump into the position of organizing, (..) or reorganizing, rebooting the AGI Society as a thing. And we’re hoping a bunch of people in this room will be happy to sign up among the next batch of AGI Society members. But, so, Andrew will tell you a bit about his plans for rebooting the AGI Society now. (….)
S35: Thank you. Thanks, Ben. (.) I mean, I think we’re all inspired by Ben. I mean, we’re not, none of us would be here without Ben, so this is not a new thing. This is an expansion on the original goal, the original vision. And so, I do have notes, so a little bit, you know, more focused than we’ve typically been. But if you think about a year from today, we want to be celebrating the 20th anniversary of the AGI conference with a lot of vigor. 20 years is a long time to do anything. If anybody’s raised kids, you know, 20 years is a long time. And the fact that we’ve been pioneers focused in that way that was discussed as, oh, those folks over there, and then all of a sudden people figured it out, that’s really powerful. So, we’re not sure where we’re going to have that exciting conference yet. But we do know that it will be larger and more diverse than even this year, which is the most diverse and the largest of the AGI conference series. So, why are we so confident in that? Because the world has shifted. The attention, as we all know, has shifted to actually AGI, not just AI or narrow AIs. Thanks largely to the people in this group, and this is not a joke, right? We have people, AGI society members, in every major lab, in every frontier organization worldwide. It’s just a fact. Maybe they’re not a current member, but they were 15 years ago, or they gave a talk or something. And so, it’s no longer a niche topic for pioneers. People right now, just globally, are paying attention to AGI as the most consequential technology that we’ll actually build. The topic of AGI is now culturally pervasive and unfortunately a bit chaotic. (.) The shift in this kind of global societal attention, it’s not an opportunity, it’s not just an opportunity. I personally believe it’s our responsibility as the people who have been doing this the longest to amplify and scale our global commitment to functionally being the leaders of AGI. This is a reputation that we’ve earned over decades. So, nothing that we haven’t done is not going to be kind of applied to. So, how do we do that? Well, we do it by doing what we’ve always done best, which is cultivating a diverse community of cross-disciplinary. And I’m glad Ben brought that up, because it really is, AGI is ultimately cross-disciplinary. It’s not just software, it’s not just hardware, it’s not just neuroscience, it’s not, you know, it’s everything. It’s all included. And so, the only cross-disciplinary AI conference at this scale is the AGI conference. The AGI journal, the journal of AGI. The community itself, just the composition of the community. (.) So, we believe that because AGI is going to be shaped by society and that drives our mission, we have to bring in everybody from society to be active volunteers. (.) So, starting right now, basically today, you know, we are dedicating renewed effort to expanding the AGI society reach. (.) Asking specifically for your help. So, you’ll see, if you’re here, you’ll see a couple banners. It says, AGI society needs organizers. So, there’s a little QR code, you can scan it. Go to joinagisociety.org and you can sign up to be a volunteer, a chapter member. I’ll tell you a little bit about some of the stuff that we’re going to be building. But, it’s a massive opportunity to build, no kidding, proper community in local, regional places to give the support to your city or your local council or local communities. Because they have questions. And guess who they’re going to ask questions to? ChatGPT. They should probably go ask it to people who are regularly doing this and using ChatGPT or others, right? So, we need organizers and leaders. (.) The AGI society members exist primarily to support AGI society members and our communities. So, we need regional chapter leaders. We need member organizers. We need community volunteers, moderators. If you’re a web designer, I need you because I did the last website. You know, there’s all this kind of stuff. Little, little, itty bitty things here. And if you’re well off financially or you have a fund, we are a 501c3 nonprofit and we will continue to be a nonprofit. This is not a product organization. We’re not creating anything to sell. We’re a nonprofit and that’s how we’re going to continue to operate to bring everybody together. So, if you want to sponsor a student, support technical reports or hackathons, go sign up, help us grow this organization. Because ultimately, at our core, we’re a community. We’re the society built around the question of AGI. This conference, and I want to point out Haley, Jordan, Matt, Ben, I mean, Will, everybody. Massive amounts of volunteer effort that you do not see goes into this every year. And the fact that we’ve been able to take over every campus that we’ve ever gone to is a testament to how much power there is behind the organization. So, best-in-class resource is really what we want to provide everybody. And so, we need your help to do that. We have a unique opportunity to provide a global base, let’s call it, that the world desperately needs for the open, transparent, education, research, and ultimately, community alignment, as the world ushers in the most important technology in human history. So, thank you. (12 seconds pause)
S11: Thank you very much, Andrew. And the numbers, of course, are growing. The keynotes that we have this year are absolutely amazing. The number of people who want to talk, the number of people who we invite to talk and squish into a couple days is incredible. And the number of attendees we have this year is more than we’ve ever had, by about a third at least. So, we noticed this with standing room only yesterday. And it should only get more full from here. Especially day four, we’ve invited a lot of people from the San Francisco area, from universities locally, and businesses locally, to come participate in our AGI Leaders Day. And we’re very excited to have a very full house for that. But, we’ll kick off next into our first keynote for the day. Thank you, Matt. Thank you, Ben. Thank you, Andrew, for getting us warmed up, getting us started, day one of the conference. And now, I’m so excited to be able to introduce Carl Friston. So, if you want to understand what intelligence actually is, then listening to the person who invented the mathematics, arguably one of the most important frameworks in the world right now for that, is Carl. We’re so excited to have him here. He’s a professor at the University of College London, and the most cited neuroscientist alive. So, his free energy principle and active inference framework, on which we had a workshop session yesterday, have become foundational across neuroscience, AI, and now philosophy. (.) So, today he’s going to take us somewhere profound. The idea that if a system exists, it must have beliefs, and that agency itself has a physics. So, please welcome with me, Carl. Thank you for being here today. (13 seconds pause)
S15: Thank you very much. I have to say, it’s a great honor to be asked to speak to the AGI conference. (.) And I am hoping that what I have to say is not going to find some disagreement with Ben. But I am going to sort of leverage this interdisciplinary aspiration and talk about AGI from the point of view of the physics, through the lens of physics. physics. So, my agenda here is to try and tell a physics backstory that defines the objective functions that underwrite things like active inference, the things that one can regard as motivating intelligent behavior. I could talk about many things in my world that speak to AGI. But what I thought would be most useful is just to go back to the basics and tell the story from the beginning, where the Newmont is this objective function, this expected free energy. In so doing, though, I’m going to have to go through some well-trodden materials. So, if you’ve seen this kind of talk before, then please forgive me. I can only hope that it induces some nostalgia. And if you haven’t, I hope you find it interesting. (.) What I’m going to do is, in fact, start with an introduction to the sort of key themes of my talk provided by Geoffrey Hinton. So, Geoff was invited on receipt of the UCD Ulysses Medal to answer questions. And right at the end of the Q&A session, he was asked, if you could go back in time and meet one of your intellectual heroes, who would you like to meet? And this was his response. (…….)
S07: I would like to meet Helmholtz. (.) And I’d like to meet Helmholtz. Helmholtz believed in unconscious perceptual inference. He basically thought that when you’re doing vision, you’re doing lots of inference. But it’s not kind of conscious, deliberate inference. You’re inferring from the distal, from the proximal stimulus. You’re inferring what the state of the world is that gave rise to that. And he was clearly right about that. So, that was one branch of Helmholtz’s work. A different branch of Helmholtz’s work was free energy. There’s something called Helmholtz free energy. (.) And what Helmholtz didn’t realize was that if you’re trying to do inference in a complicated model with complicated latent variables, a statistician would tell you to do inference by, (.) given the current parameters of your model, take the data and figure out the best way for your model to explain the data. And there might be several ways for your model to explain the data, but you have to figure out the probability of each of those ways in which your model could explain the data, in order to adjust the parameters of your model. And what variational inference says is, actually you don’t have to get that probability right. You can approximate that probability and still guarantee you’ll learn a better model, or rather there’s some band that’ll improve. And so, the two completely different bits of Helmholtz’s career, the work on free energy and the work on perceptual inference, were actually very closely related. And free energy is the key to being able to do perceptual inference tractably. And I’d love to tell Helmholtz that. (.)
S15: And then he receives rapturous applause. And my job for the next few minutes is to try and unpack the enthusiasm for these bilateral contributions of Helmholtz. And interestingly, it speaks nicely to yesterday’s workshop, in particular the conversation between Ben and Gary about the nature of world models that in this world can be read as the generative models implicit in Jeff’s narrative. And it also speaks to a certain extent to Anil Seth’s picture of computation as a process, as some kind of dynamics, and the particular dynamics I want to focus on are those that inherit from the physics of self-organisation. So, this talk’s going to have three bits. First of all, we’re going to start off with the statistics of life, with a special focus on Markov blankets, things that enable you to sort of individuate something from everything else. and in particular, the Bayesian mechanics that ensue simply by being able to separate something from the universe in which it is immersed. And then I’m going to tell the same story from the point of view of a psychologist or a neuroscientist in terms of the anatomy of inference, with a particular focus on implementations by predictive coding, as instantiated on neuronal networks and also neural networks. And then I want to finally turn to AGI in the sense of acting in the right kind of way, and just look at the imperatives for action and the underlying perception in terms of active inference and agency. So, I’m going to start with a question posed by Schrodinger. How can the events in space and time which take place within the spatial boundary of a living organism be accounted for by physics and chemistry? I’m going to take the notion of a spatial boundary as definitionally really important here, in the sense that if we want to talk about any system, any intelligence, we need to be able to separate or individuate it from everything else. And that separation I’m going to associate with a Markov blanket. So, for those people who don’t know what a Markov blanket is, I’ve tried to cartoon it here in terms of a little universe in which there are states denoted by these cyan circles, and each state can influence another state, and the influences are denoted by these arrows or edges here. And if I were to take some particle or person or population, some system of interest, say me, I can identify my internal states. And then the Markov blanket simply comprises the parents, the children, and the parents of the children. And you may be asking, well, what does that do for you? Well, the Markov blanket effectively acts as a statistical blanket insulation, in the sense that if I wanted to know how my states are going to evolve over time, given the rest of the universe, I’d only actually need to know my blanket states, my Markov blanket. Technically, this means that the internal states are conditionally independent of the external states given the blanket states. And I’m going to make a further move here. I’m going to divide the blanket states into active states and sensory states, where the sensory states influence but are not influenced by the internal states. And similarly, the active states influence but are not influenced by the external states. And this is the partition I want to bring to the table that defines something, some system, the self, and self-organization in terms of, say, a particle with particular states that can be divided into internal states and their blanket states, namely sensory and active states. So, to put a bit of flesh on that literally, not a little bit of brain on that, let’s take my favourite organ, or our favourite organ, the brain. So, we can associate all the neuronal synaptic states of the brain with the internal states, and then the active states become the actuators, the autonomic reflexes, all the things that act upon the external states, which, of course, will include my body and extra personal space. And then those external states reciprocate in terms of causing changes on my sensory epithelium, my sensory states, my sensor organs, that then change the internal states. So, what we have is a picture of an open system where the inside influences the outside precariously through these outputs or active states, while the outside influences the inside through the inputs or the sensory states, inducing a circular causality, which we’ll see later. We’re going to associate with action and perception, particularly in relation to the blanket states. But for the moment, I’m going to ask you just to forget about the Markov blanket. What we’re going to do is a sort of crash course in physics, and then put the Markov blanket back in play, and see what kind of mechanics emerges just from the existence of being able to identify the self or something via its Markov blanket. So, the crash course in physics. I’m going to assume that every universe can be expressed as a random dynamical system, say a Langevin equation where there’s some abstract states whose rate of change is some lawful function of where I am in state space, plus some random fluctuations. And I’ll try to illustrate that here with just two dimensions that could represent the evolution of states at any scale. So, for example, it could be fast oscillations in one of my nerve cells. It could be various phases of the cardiac cycle as my heart beats. It could be me getting up in the morning, doing my emails, having a cup of coffee. It doesn’t matter. The important aspect of this particular characterization is that I have to revisit states that I once occupied, or at least their neighborhood. Otherwise, my states would go off to infinity, and I would not have any characteristic states. Technically, this describes something called a pullback attractor. And I can interpret this pullback attractor probabilistically in the sense that if you sampled me at any timescale at any random time, then the probability you’ll find me in this state is given by the density of these trajectories or these states. And that’s interesting because most of physics is based upon characterization of the way that this probability density evolves in time. And I’m describing that here in terms of a Fokker-Planck equation. Don’t worry about the maths. I just want to show you where the key results come from. If you do like the maths, you may also recognize this as a time-independence Schrodinger equation or a master equation. The key point, what it’s doing though, it’s describing the dynamics or the evolution of the probability density in terms of the flow, the dynamics and the amplitude of the random fluctuations gamma here. And the nice thing about this is, if I’m dealing with things that persist over time, then this probability density over the characteristic states of the kind of thing that are characteristic of the kind of thing it is, is not changing with time, which means I can write down the solution to this equation. And I’ve written it down here in terms of a thing called a Helmholtz-Hodge decomposition. And all this means is that if something exists in the sense of possessing characteristic states, to which it self-organizes, then it must be the case that I can express its dynamics, its flow, as effectively a circular gradient ascent on the log probability of being in a characteristic state, given the thing in question, say me. This decomposition is into dissipative and non-dissipative parts, which we’ll all be familiar with. Effectively, the dissipative part climbs probability gradients to counter the dispersion due to the random fluctuations, while the solenoidal or non-dissipative conservative part gives you life cycles, oscillations, essentially flowing around the iso-contours of this probability density, illustrated here in terms of water doing the gradient flow here flowing down, (.) up probability gradients or down potential gradients, gravity in this instance, and the circular flow being this sort of conservative non-dissipative part, which is characteristic of certainly biotic self-organization, and one would also argue thereby intelligence self-organization. So, what about the Markov blanket? So, that sort of fundamental Helmholtzian dynamics has to be true for all the subsets of this Markov blanket partition, including the internal states and the active states. And I’ve just written that down here. And in a way that leverages the observation that by construction, the internal and active states are not functions of the external states. They are just functions of the blanket states here, the sensory sector of the blanket states. So, this provides a sort of fundamental mathematical image of what we could call perception and action in relation to the internal and active states. Now, how can we understand this? Well, it will look as if both perception and action are trying to maximize the same quantity, climbing gradients where this quantity is a log probability of some sensory states given the system in question. So, how could we interpret that? I put a number of interpretations up here. I’m sure you will be familiar with a number of these. Let’s just take the log probability of some states. So, I’ve just said that this probability reflects or scores the states that are characteristic of me, that have meaning for me, that have value for me. These are the kinds of states that I aspire to and I will want to remain in if I want to remain the kind of thing that I am. And, of course, we can label that as value. And from that, we can interpret this generic formulation of perception and action in terms of, say, reinforcement learning. If I was an engineer in optimal control theory, if I was an economist, it could be read in terms of expected utility theory. (.) And that’s nice because the negative of that log probability in information theory is called self information and more simply surprise or even more simply surprise. (.) And it just measures how unlikely it is that I will find myself in this particular state. And that leads to formulations or descriptions of this kind of self organization in terms of the principle of maximum mutual information or the informatics principle or its complement to minimum redundancy principles and, indeed, the free energy principle. Why? Well, the free energy is just a tractable bound or approximation to the self information or or surprise, as we’ll see later. In turn, the average of this, the average self information is entropy. So it’ll look as if self organization is trying to counter the dissipative effects of that you would normally be associated with the second law. It will try to gather the states up into its characteristic attracting set or pullback attractor. And of course, this is the holy grail of self organization. For example, in physics, we’ll have Herman Haken’s synergetics. But if I was a physiologist, this is just a statement of homeostasis. It’s just a statement of keeping my essential variables within viable bounds that are characteristic of me. There’s a fourth interpretation that I want to leverage here. And that’s an interpretation that a statistician would bring to the table. So she would say, ah, well, this quantity is just simply the probability of some sensory data given me as a model of those data, a model of how those data were caused by the hidden causes, the latent causes that are, in this instance, the external states. And that view leads to notions of the Bayesian brain evidence accumulation. And as we’ve heard earlier today, formulations in terms of predictive coding. And that’s the perspective I’m going to pursue through the writings of Helmholtz. Helmholtz. In fact, you could probably trace the story right back to Plato. I’m going to sort of pick it up in the 16th century using this painting by an oil painter famed for painting still lives that when viewed from a different perspective, give you a very different visual impression. So if previously you saw a bowl of fruit and now you see a face, the point being made here is that you constructively made that face on the inside. You created an explanation for this particular pattern of visual impressions that explains, and if you like, forms a hypothesis that best explains in the sense of inference to the best explanation that Gary was talking about, this particular pattern of sensory input. So this speaks to a theme, certainly in philosophy of neuroscience that people like Andy Clark and indeed Neil Seth would summarize as the brain being a very active organ, a constructive organ, constructing explanations, fantasies, literally making the brain a fantastic organ that best explain or engage in some kind of inference to the best explanation. And this view I think is probably most fluently and beautifully articulated by Helmholtz in his various writings. So, for example, objects are always imagined as being present in the field of vision and would have to be there in order to produce the same impression on the nervous mechanism. Again, again, it’s like you have to have the explanation, the hypothesis, the fantasy on the inside in order to explain the sensory information, which of course is very, very closely related to theories in psychology. So, for example, say, for example, Richard Gregory’s notion of perception as hypothesis testing, ideas used to great effect by Jeff and his colleagues, such as Peter Diane, who actually built a Helmholtz machine as a metaphor for the Bayesian brain, borrowing from Bayesian probability theory, and in particular, rendering the dynamics tractable by referring to Feynman’s this variational principle, in particular, variational free energy formulation from his path integral formulation. But let’s, for the moment, just come back to this notion of impressions on the nervous mechanism. So, if this statistical interpretation, this sort of evidence accumulation, predictive coding, sometimes called self-evidencing, in the sense that this quantity, the probability of the sensory data given me as a model or a model of this data, is sometimes called the model evidence, or the Bayesian model evidence, which means action of perception can be read and as maximizing the evidence as maximizing the evidence for my world models or my generative models. If that’s true, what it means is there’s always a description of me, or at least my brain, that is trying to infer the causes of sensory impressions on my Markov blanket or my sensory veil. So, for example, is this sensory impression, this shadow here, is it a bird or a leaf? And I will be compelled, or at least I can be described as trying to infer, infer to the best explanation for this particular pattern of sensory input. So, how does my brain do that? Well, we know, to a certain extent, what it must be doing, because it exists and therefore must conform to this Helmholtz decomposition in terms of its dynamics, must be some kind of non-distributive gradient flow on this log probability, or in this instance, expressed in terms of the variational free energy due to finement. (..) What I’ve done here is just make an interesting move, which is to associate the internal states of my brain with the parameters of a probability distribution over the causes of my sensory impressions. And that now allows me to understand internal dynamics of a neural network as basically holding Bayesian beliefs about the causes of content or data or the sensorium. How does this relate to predictive coding? Well, it turns out that these free energy gradients can always be expressed as a prediction error, which basically means I can now understand the non-distributive and dissipative parts of my neuronal dynamics in terms of a prediction. So, if I now read these internal states as expectations about the causes of my sensations, then I can predict how they’re going to change in the moment. But I can also reduce, by the dissipative part, or reduce the prediction error by doing gradient descent on the gradients or the prediction error itself. And if I was talking to a control theoretic audience, then this would be a description of a Kalman filter, which could be divided into the prediction plus the update with respect to the prediction error. So, what’s a prediction error? So, what’s a prediction error? Well, imagine I had this sensory impression on my retina, and I had some brain state or some representation that hypothesized that these sensory inputs could be caused by a howling dog. And if I had a generative model or a world model that could generate what I would see if my expectation was correct, then I can simply subtract my prediction from the actual sensation to form the residual or the mismatch, which is just the prediction error. And all that this Helmholtz decomposition is saying is that it looks as if this prediction error is now being used to drive or revise or update my expectations until the prediction error is minimized and the free energy gradients are destroyed and minimized and found by minima of free energy. So, notice in this formulation, we’re only trying to minimize prediction error. We’re only trying to destroy free energy gradients. We’ll never actually know what caused our sensations. In fact, in this instance, it was a cat, not a dog, but that doesn’t matter. If I keep my prediction errors low, then I am compliant with the solution to my density dynamics for as long as I survive. So, the nice thing about predictive coding is that we can forget about all the physics and just summarize the imperatives for existing in some characteristic states as minimizing prediction error. And there are two ways of minimizing prediction error. I can either literally change my mind by changing my internal neuronal states to make my predictions more like my sensations. So, we can associate that with perception, or I can change the sensations by acting on the world, by resampling the world in such a way to make my sensations more like the predictions. In other words, I can selectively sample the world, literally, say, with visual palpation with my eyes or with my hands, or by turning to my favorite news channel or social media, to try and solicit those outcomes that I predicted in order to then interpret action as make it so. So, fulfilling the predictions that I’m fulfilling the predictions that I am building about the way that the world is unfolding based upon my generative model. (..) So, here’s an example of this, which is a bit neurobiological. for my apologies. But I think it’s illuminating because it just illustrates how simple the architectures, and indeed how deep and complex they can be, that underwrite this formulation of perception and action. So, let’s just look at a simple brain in receipt of visual input from the eyes, from the retina, they come down to some nuclei deep in the brain, who, that are in receipt of top-down predictions, so that the difference now constitutes the prediction error, that is then used to drive the expectations to provide a better account of the sensory input. And technically, this can be read as Bayesian belief updating explicitly, well, expressed in terms of the dynamics of self-organization due to this Helm-Holzun decomposition. But notice, these expectations can, in a deep generative model, themselves have top-down predictions, elaborating second-order prediction errors that then can be used to drive changes, belief updating at the higher level, until you’ve got a suppression of prediction errors of free energy at each and every level of some deep generative model at each successive level of abstraction. And the architecture is very simple. This involves reciprocal message passing the top-down predictions that are countered by bottom-up prediction errors. (..) But what about action? Well, think about another kind of input, a kind of input that inherits from something that we can actually move, say, muscles in our eye. And this input comes into the pontine nuclei, another part of deep brain structures. It is a receipt of top-down predictions about where I predict I am looking, for example, and I have a prediction error. And I could use this prediction error to update my beliefs about where I’m looking. But there’s a much simpler way of resolving that prediction error, and that’s to turn it back into the world and use it to drive contraction of the muscle until the muscle signals what I predicted. So what I’ve just described here is a kind of thing that is used in robotics, but probably more importantly, is the way that we move. This is just a classical motor reflex arc. And on that view, these descending predictions basically are supplying fixed points, set points for this sort of reflexive fulfillment that action is in charge of. So notice we’re not trying to resolve prediction errors in every modality everywhere. We’re just resolving proprioceptive prediction errors in a very small or limited modality or domain that action can fulfill. So it’s a very simple device and effectively a closed-loop control-theoretic approach to action, but it’s deeply informed in an open-loop way by all this hierarchical processing under these deep generative models. So is that a sufficient account of action and perception? I mean, you can get quite a long way using that setup and basically just solving these equations of motion to simulate or emulate biologically realistic behavior that has some aspects of certainly self-organization that look as if it may be intelligent or at least purposeful. (.) So, this is an example. This is one of, it’s an old example now, but it’s one of my favorite examples and hopefully you’ll see why in a moment. What we did was equip a synthetic agent with effectively a central pattern generator, an internal generative model, world model that had some very simple dynamics. dynamics and those dynamics basically were that some abstract state of the world was moving in a heteroclinic cycle amongst some unstable fixed points. Furthermore, we equipped the agent with a generative model that mapped this position, this abstract central pattern generator space to fixed points in extra personal space. And furthermore, the agent thought that its finger was attached by an invisible spring to this invisible point. So from the agent’s perspective, what it expects to feel and see is its finger being pulled around in space. And it was so configured to emulate a very elemental kind of handwriting. So the top-down predictions of proprioception are then realized reflexively by the agent, incidentally and coincidentally, thereby generating the movement and the visual impressions that were being predicted. So basically, this agent is authoring its own sensor, it’s creating its own world that it can predict. The reason I like this is you can play all sorts of interesting games with this kind of simulation. So for example, what we can do is examine the activity of these internal nodes, say, standing in for neural populations, and plot their activity when they are high, as a function of where we are in this extra personal space. And what we see is something which people in neuroscience study an enormous amount, which is receptive fields. In this particular instance, they look very much like something called play cells that you may have heard about in the hippocampus, basically responding when the finger is in a particular position and indeed moving in a particular direction, here preferring the downstroke to the upstroke. The second thing we can do, though, is just play with the sensory input. And we can preclude the proprioceptive input, the feeling part of the input, the actuators and the sensors, the IMUs coming from, say, the actuators, but replay the visual impressions, the visual input, say, the RGB feed. And, of course, this agent has an apt generative model for explaining this particular kind of movement as if somebody else was actually doing the handwriting. So this becomes now, if you like, a very simple model of not action per se, but action observation, where the same machinery, the same message passing, the same dynamics, the same inference, the best explanation, is being deployed to explain both self-generated action and also the same kinds of movements that are observed but performed by another. And in neuroscience, this is referred to as the mirror-neuron system. So is that a good enough account of how we exchange with our world and evince intelligence? And clearly, I think you probably argue no. So in the final part of this brief backstory, I want to ask, well, what’s the difference between you and me and that little effectively glorified Bayesian thermostat? (…) And one simple manipulation basically allows me to draw a bright line between this sort of reflexive or mere active inference or self-evidencing and the kind of behaviours which get much closer, I think, to intelligence. Behavours that have an intentional aspect that have a clear and formal set of motivations and quantifiable motivations that emerge as a consequence of the Markov blanket under this physics perspective. And to get there, to get there, I’m just going to assume that we’re dealing with big things like you and me, in which case these deep generative models mean that most of my internal states don’t have direct access to my active states. My active states now are on the periphery, and therefore they effectively become hidden states or latent states that vicariously cause my inputs. So I’ve tried to illustrate that. If you just focus on this connection between the active states and the internal states, if I remove that, then I get a very different kind of computational architecture. Because now, unlike say in a single cell or a very small organism, now the active states now become vicariously hidden causes or late or external causes of my sensory states. But recall, we can always interpret the internal states as inferring to the best explanation the causes of the sensory states, but now they include action. Which means that now I can read these deep, larger kinds of generative models as inferring their own action. So now there becomes a formal distinction between action as a realized variable, and what I think I’m actually doing, and because I’m now inferring what I’m doing, very much in the spirit of planning as inference. And of course, as soon as soon as you talk about planning, you have intentions, and you have to be able to select which of all these counterfactual futures that I could commit to, I’m going to commit to. And it so happens that the maths provides a really neat way of writing down the imperatives or the probability of you selecting a particular path into the future. And that’s basically what I’m going to do, and that’s basically what I’m going to conclude with in terms of just showing you the functional form of this free energy functional, which now is not a functional of observations or sensory data or content in the moment, but sensory data that is consequent upon my actions in the future. So these are future variables, and therefore they become random variables, and that requires me now to take an expectation under technically the posterior predictive density. And therefore, the free energy now turns into an expected free energy, and that now scores a probability that I’m going to act in this way or act in that way. So for people interested in the functional forms, again, if you’re not, don’t worry about this, I’m just showing it to show this, you know, to my eye. (.) And, you know, lovely symmetry between the variational free energy due to Feynman and the expected free energy that underwrites things like self-evidencing and active inference of this non-reflexive kind of this intentional motivational kind. So I’ve just written out the free energy function fully here. Before, remember, I said it was basically a bound on or approximation to log evidence. And here is the log evidence here, the log probability of some sensory states. Why is it a bound? Well, this quantity of KL divergence could never be less than zero. So this quantity upper bands the negative log evidence or lower bands the log evidence, hence elbow in machine learning. But there’s another way that I can rearrange this free energy to express it in the way that a statistician would more intuit it. (.) What she would say was, well, look, what we’re doing by minimizing free energy, variational free energy, is trying to maximize the accuracy, namely the log probability of the sensory data, given my beliefs about the causes, the hidden or latent causes of those data, whilst at the same time trying to do so in a parsimonious way by minimizing the complexity. So what is complexity in this context? Well, it’s just the degree to which I have to change my mind from my prior belief before seeing the data to my posterior belief encoded in this variational density here. So it’s another KL divergence. This measures the degrees of freedom that I’m using up in providing an accurate account, very closely related to algorithmic complexity. (..) So that’s variational free energy. Now we take the expectation of this functional form, loosely speaking, to give us the imperatives for selecting a particular action. And interestingly, the complexity and inaccuracy become risk and ambiguity, respectively, and divergence and log evidence become intrinsic and extrinsic value, respectively. So what are these things? Well, you will again recognize these, or some of you at least will certainly recognize these in various disciplines. Let’s start by assuming I have no prior preferences. Every outcome, everything I sense is equally plausible. I can take anything. And that just leaves this KL divergence, which I now want to maximize the intrinsic value. And what is this divergence? Well, it’s basically the difference between my beliefs about the external causes, the latent states causing my sensations or my data, given the sensory consequences of acting relative to the absence of that sensory information. So effectively, it scores the epistemic value of acting like this in terms of how much uncertainty will I resolve about states of affairs in my external milieu. In the visual search literature. In the visual search literature, this is known as Bayesian surprise, and more generally, it’s just a mutual information. It’s just a statement that I will act in a way such that I maximize the mutual information between the latent causes of my sensory consequences that are observable. (..) Let’s look at another part of this expected free energy. Let’s take ambiguity off the table. Let’s assume that I can see everything, that given my sensations, I now know the external states because there’s no ambiguity left. What am I left with? Well, I’m left with another KL divergence that I’m going to want to minimize now. And this KL divergence is just the difference between what I anticipate will happen if I do this, commit to action A, relative to what a priori I prefer to have, I’m not going to happen either in terms of the external states or my sensory consequences of that. And of course, this is just pathological control and control theory or KL control or in economics could be read as risk sensitive control. And then finally, if I take all uncertainty away, then by assuming there’s no more epistemic foraging that’s going to resolve my uncertainty, I know everything that I can know about this. So all the uncertainty, there is no reduced uncertainty left. I’m just left with this. And what’s this? Well, we know what this is. This is the expected value or expected utility in information theory and the kind of objective function you’ll find in reinforcement learning. So a brief simulation of that, leveraging this notion of the negative expected free energy scoring the information gain, what you’re scoring the salience of particular actions or sampling, say, of the visual scene in terms of the information gain or the resolution of uncertainty I would get if I looked over here as opposed to looking over here. So this example actually illustrates the information gain as a salience map based upon this image here. And indeed, this accounts for where people actually do look in terms of their epistemic foraging, responding to these epistemic motivations or intrinsic motivations in order to understand and resolve uncertainty about the visual scene. And we can simulate that. (..) This simulation, a very simple simulation, the agent here had a very simple universe that she could either be visually palpating an upright face, a sideways face or inverse face, but clear, but crucially could only sample one little bit at a time. So I had to choose quite carefully on the basis of this intrinsic value or intrinsic motivation or expected information gain where to look next to resolve the greatest amount of uncertainty about what she was looking at. And after a couple of psychotic eye movements becomes relatively confident that the face was an upright face and that was absolutely correct. (.) So just in summary, this picture that inherits just from a definition of the physics of self-organisation, which is equipped with a stipulative definition of what it is to be a self or a thing in terms of this Markov blanket, leads to a picture of what one could argue as at least elemental intelligent behaviour in terms of self-evidencing or active inference that has a dual aspect Bayes’ optimality. Bayes’ optimality. Bayes’ optimality, on the one hand, it looks as if they are trying to maximise expected value in the spirit of optimum Bayes’ decision theory, but this is contextualised or constrained by complying with the principles of optimum Bayes’ experimental design, designing the right experiments so that the information you get maximally resolves uncertainty about the hypotheses you’ve brought to the table to actually explain all the sensory data. (.) Bayes’ data. And this is just information. It’s a mutual information of expected information gain. So Bayes’ optimal decisions plus Bayes’ optimal design effectively equals active inference in the sense that expected value and information gain constitute the expected free energy. And with that I’ll conclude and give the last word to Helmholtz, as I’m sure Geoffrey Hinton would appreciate. Each movement we make by which we alter the appearance of objects should be thought of as an experiment designed to test whether we have understood correctly the invariant relations of the phenomena before us. That is their existence in definite spatial relations. And with that, it just remains for me to thank those people whose ideas I’ve been talking about, but of course to thank you for attention. Thank you very much indeed.
S32: Thank you, Carl. I think we have time for a few questions. (….)
S29: One of the most interesting aspects of a lot of work up until a few years in centering is raw sensory data generates. (.) I’ve been kind of curious to know, obviously you and I have worked on some things together, but I’m curious to know your thoughts on this idea of flipping vision sort of and visual perception towards an encoder-centric joint embedding perspective. I mean, Jan LeCun, for example, has talked a lot about working in latent space. And similar to your experiments there towards the end, like, for example, our ocular motor model you and I have shown also does like foveation and saccades and kind of like builds like an active perception model. So I was kind of curious to know if your sense of the direction of where predictive coding might need to go towards something like, for example, encoder-only as opposed to getting caught up in these low-level sensory details. Because there’s a lot of evidence in reinforcement learning. For example, we get distracted with a lot of low-level modeling details, whereas visual perception, we often try to pull out abstract information such that we can effectively choose where to look next. So I was kind of curious to know your thoughts on this dichotomy or this new view of, like, getting away from a direct generative model. I mean, we still do generative modeling and encoding space, but flipping it to encodes. I wanted your perspective. (..)
S15: Yes. A great question. I think there are a number of perspectives you could bring to bear on this issue. I mean, I think, you know, one obvious perspective, which is not a new perspective, it’s basically the perspective afforded by active vision. And I know that your work is addressing that, that of Rajesh Rayo clearly with his active predictive coding formalism is taking on that brief of making perception inactive. And, you know, in many respects, I think that, you know, the original work of Helmholtz buried in that treatment of perception is actually active sensing and active vision. And I think that what’s distinguishes him from sort of pure Kantian views of sense making that we’re actually in charge of exchanging and sampling the world with our bodies to find the data that we’re then going to use to do our unconscious inference or inference to the best explanation. And so I think sort of making things inactive is really important because predictive coding as originally introduced, you know, it’s not a, it’s not about doing anything. It’s just about compressing. Compressing, you originally is compressing sound files in the most efficient way, which speaks to that sort of complexity imperative that is part of variational free energy minimization and also equipment formulations in terms of compression and algorithmic complexity and universal computation ultimately. So I think the inactive part is really important. And of course, you know, why are you going to act? Well, you know, on the view that I’ve tried to lay out this morning, it’s really to resolve uncertainty about the latent causes. So clearly you’re going to be working in latent state space. Because clearly, you know, I think most people would agree, you know, you have to have some kind of world model or generative model that holds that latent space of cause and cause effects structures that you want to resolve and update and learn and refine and deploy in order to decide what to do next. (.) And clearly, you know, and clearly, you’re going to have to have very deep structures that does indeed create an encoding decoding kind of architecture. I’m not so clear about that. But I’m guessing, Alex, it’s you asking the question. I can’t see you or hear you properly. But if it’s you, you will probably be talking about this later on. So you can unpack the answer to your own question later on. But I’ll just conclude by saying that, of course, you know, that sort of conjunction of encoding and decoding perspectives in the service of computing under a generative model hits you in the face when you just look at, say, the structure of variation autoencoder as a very simple example of this kind of tech in machine learning. that you’re basically compressing that you’re basically compressing down to the bottleneck with an encoding and decoding in terms of self-supervised learning, for example, in order to get a good generative model of a very naive, literally a naive Bayesian sort in a variation autoencoder that speaks to this compression. And I think that, you know, the bigger move now is to, well, that’s fine. But now, where do you go and sample your data? You know, how do you choose what to look at next? And I think that that is the essence of intelligent behavior because it speaks to action and behavior. (..) Is that what you had in mind? (13 seconds pause)
S13: Describing was essentially a method referring the most likely models given sensory data. (.) But what that does not address is what you’re going to model. The world is incredibly complex. You can’t possibly model everything. So the agent’s utility function must enter into this equation somewhere. And I didn’t see that appearing anywhere in your talk. (..)
S15: Right. Excellent question. And, you know, just allows two fundamental points to be made here. The first one is really interesting. It’s, you know, the objective function when read as the marginal likelihood or the evidence for your world model necessarily implies that it is the simplest explanation possible. And this can be read in terms of principles of least action variation principles of least action or it could be read in terms of efficiency or it could be read in terms of mutual information or it could be read in terms of compression. (.) I think all of these are different perspectives on the same fundamental thing. You’ve got to keep your generative model as simple as possible, but has to provide an accurate account in a very context sensitive way. So you’re absolutely right. You know, we are, you know, our models are remarkably expressive and context sensitive, but they are also as simple as possible. So, you know, one important part of the objective function is to minimize the degrees of freedom that are required to reach your inference to the best explanation or not only about what’s causing your sensory input, but also what you are going to do next. How are you going to act? How are you going to act? You know, the planning is inference perspective. The second point, you know, where is the utility function? It was part of the presentation. If you remember at the end, we were talking about some value and expected value and expected utility. (.) It’s, it’s, it’s, it plays, it is explicitly part of the, for example, the expected free energy when decomposed into expected information gain and expected utility in the spirit of base optimal decision theory rests explicitly on a cost function that in active inference is specified in terms of the, of the, in terms of a prior preference. The log of the probability that this kind of thing is good for me or has value for me. I have to say that, you know, I tend to try not to use the word utility or value because people immediately assume that this is some kind of generalized reward learning or value learning. And in a sense it is, but I think it has, it has much closer relationships to the constraints that you’ll find in things like James’s maximum principle under constraints. So what we’re, what we’re talking about is an objective function where you write in the constraints on the, the states that you don’t enter into in terms of this, this potential function, the negative value, if you like, or log probability, or log prior preferences. And that is much more expressive. So you, you, you, you, you, you, you prescribe the kind of thing that an agent is not in terms of the constraints and then under those constraints, it then does its epistemic foraging. It becomes curious. It becomes novelty seeking. (…)
S11: Thank you so much, Carl. It’s been amazing to have you here. It’s a, I’m so glad that we got to take a couple of questions with you. I’m so glad that you got to attend this year. And, uh, you’ve opened up so many threads that we’ll get to unpack throughout the day and throughout the, throughout the week. So thank you for joining us. Big round of applause for Carl Friston.
S15: Thank you very much. (…….)
S11: Uh, we’re gonna enter into our first paper session now. Super excited. So this is gonna be our, our group one session about intelligence as self-model and metacognitive organization. Um, so this session is gonna ask the question that sits at the heart of this conference. What does it mean for a system to represent itself? Modeling the question on, on world models, right? These, these two pieces that are so critical and two halves of, of some whole. (.) Um, so the, the papers in this group will treat, uh, uh, cognition as something that depends on system modeling, monitoring, or constant, constituting its own identity. And our keynote speaker is going to challenge some of the most widespread assumptions in AI research head-on. So, Alexander Lerchner is a senior staff scientist at Google DeepMind, where he works on generative models and the philosophy of AI. His talk today, The Abstraction Fallacy, argues that computational functionalism cannot be logically correct, and that building AGI does not mean building consciousness. Please welcome, Alexander. (21 seconds pause)
S23: All right. Can you hear me well? (..) Wonderful. (.) Yeah, good morning, everyone. My name is Alex, and thanks very much for inviting me here, um, today to speak about something that’s very dear to my heart. And that’s the question of what are we actually building when we are building AGI? What’s really the nature of our, of the algorithms, of the obstructions physically? (..) And, obviously, the reason why this is so important, um, we are already, we have just said yesterday, this, um, wonderful workshop on consciousness, is that everyone kind of is now asking suddenly. We’re in a very strange, um, time in history, where we’re asking the question, are these algorithms that we are building waking up? Are we building conscious AI? (.) And the kind of answers that I see most commonly, um, to this question, first of all, there is no agreement, there’s no clarity, and I think we need to make progress on this. (..) And, obviously, the functionalist view, that is probably the most represented here, um, among us computer scientists, is, well, probably yes. So, as long as the algorithms create everything that we understand from minds, um, we will create real minds. And, of course, there’s a second kind of, um, answer to that. That is, uh, the fact that we don’t yet have a fully agreed theory on consciousness, uh, even worse, right? So, there is so much disagreement on theories of consciousness that we basically could just say, well, shrug our shoulders and say, as long as we don’t have a theory of consciousness, how can we even address this question in the first place? (.) And so, what I will do today is say, well, there is actually a third way. when we ask about AI and consciousness, we’re not just asking about consciousness per se, we are asking if algorithms can create consciousness, right? (..) And, what I would like to, what I would like to argue on this, is, then, in fact, we don’t need a full theory of consciousness to ask this, to answer this question. We basically have already enough knowledge about computation and about what we are doing in computers, that we can make real progress on this question. But what I do suggest is that we kind of need to actually rethink what is the nature of the algorithms, to really ground algorithms in physics, and to be very, very careful to trace it all the way back what obstructions actually are physically. (.) And if we do that, then we will come to actually some surprising results that will clash with some of the kind of biggest intuitions that we have as computer scientists. And so, so let’s, let’s just jump into this. (.) And because, because this is actually a really quite counterintuitive variable end up here, what I would like to do is to just walk you all through the high level intuitive steps here of how we can actually go about this. (..) And so, the roadmap that we are, that I will follow here, is really starting from how we define and how we know what computation is in our field. What does it mean to implement an algorithm? (.) And then from there, go all the way back, okay, so algorithms are about symbol processing. So, what are symbols? And to link them, how do actually symbols come about? And we will see if we need to link them to concepts. (.) And once we’re linking to concepts, the next step is to ask, okay, how actually are concepts related to experience? And this is one of the kind of probably most counterintuitive steps that I would like you to think about. And once we are kind of an experience, we can link experience back to physics, to fundamental physics. (.) And once we have gone through all these steps, we will come at a kind of point where we have a concrete ontology that we can then revisit and compare it to more common assumptions like computational functionalism as one thing. But note that what I’m not doing here is I’m not kind of starting and trying to dismantle the computational functionalism. What I do is I hold it aside. I say, let’s put this assumption aside and let’s just start directly from the algorithm itself. (..) Alright, let’s dive into this. So, let’s start with the standard definition of computation. And this simple diagram depicts what it means when a computer runs an algorithm. (.) So, what does it mean? We have here on the kind of lower, lower row, the physical states where one physical state goes to the next physical state to the next physical state. So, this is the mechanical work that a computer runs through or basically the electrical work that a computer runs through when it runs an algorithm, right? And to the degree that we are saying it implements an algorithm is what we have on the upper row here. So, we have the abstract concepts, the logic that these physical states basically have to mirror exactly. (……..) Oh, here we are. Sorry. (……) So, I was talking about the physical states on the bottom running through the mechanical physically driven process of calculating something, multiplying numbers in the computer by having logic gates that flip. And on the top is the logical equivalent that represents the algorithm itself. And so, the question is how actually is the algorithm represented in there, in the computer? And the standard way of thinking about this is that the physical states have to map onto the abstract states. And if this diagram commutes, and if it doesn’t matter whether you trace the physics or whether you trace the abstractions, and you can go back and forth between them, then you have successfully implemented. So, that’s something I think we can all agree on in thinking how do we relate this, this abstract algorithms to the, to the physics itself. (..) Note however that this, this doesn’t tell us very much about the causal origin of, of this algorithm. It’s a kind of post hoc description. It’s something that usually is described as an observer looking at the physics and observing, okay, this physics look like implementing an algorithm. (….) And that gets us to the question, what are these, these obstructions on the top here? (…….) Now, let’s, let’s start with the kind of physics themselves. (..) Um, I already explained that these are the electrons that basically go just by running the blind physics. (.) and this is what you can kind of call, what I like to call vehicle causality in what’s happening in, in, in computers. So, that the logic gates are just going through the, through the motions and they’re flipping the, the voltage based on whatever the voltage kind of, um, crosses a threshold. And in contrast, the, the, the obstructions themselves, I think one, one step that we can too easily make here, is to just think, okay, the obstructions, they somehow live, live in the, in, in the computer. (.) However, those obstructions are really the concepts that we have developed, our understanding of, of the physical world, our kind of derivation of what are numbers, what are kind of multiplications. And those don’t actually live there. So first, there must be someone, we, this is kind of, this is living in our heads. And this is something that, um, you can describe as experiential concepts. So to the degree that we understand the concepts of zero, of one, of numbers, of red, of pain. Actually, we need to experientially kind of really understand them and live them. (.) And, and that’s an important point that we come to later. (..) So, so again, so if, if you think about it, the only degree to which you can understand what I’m saying right now, the only degree to which you understand language, which is me sending symbols via the medium of, of the air to you, via, um, basically acoustic waves, is that you reverse engineer, go from the symbols back to the obstructions. And as long as you can’t actually successfully get, get some experiential understanding of what I’m talking about. Um, if I say apple in a kind of language that you don’t know, you, you won’t make this, the second step here. And so that’s key. So, so the first one is that these, these obstructions, they don’t live in the machine. And not only that, in fact, we, we can ask the question, how do we, why do we even say it? So where do they actually come from in the first place? And causally, it’s not that nature comes labeled with zeros and ones, it doesn’t come labeled as symbols. So, so we have the active map maker, we have an active person who has this understanding first, and only then kind of assigns two physical states, the entire sets of physical states, what they, what they mean. So the error causally goes the other way around, right? (……) Since it is this kind of arbitrary assignment, there is no reason that a voltage between 5, 4.9 volt and 5.1 volt stands for a one. This is something that we had to kind of implement first. This is something that we had to assign. And not only this, it’s not just assigning one symbol at a time. But in order for an algorithm to even run and be well defined, we need to define an entire alphabet of symbols. (.) So we need to have a set of states, all of them corresponding to concepts, and all of them arbitrarily assigned by us to define what are these, what, what is the meaning of those. And this kind of alphabetization is not something that comes for free. It’s not something that kind of lives in the computer, but it actually lives in our brains, in our heads. (…..) So to sum up, computation doesn’t create concepts. It’s exactly the other way around, right? So computation is made by assignment of concepts that we already had to have previously. (……) Now let’s come to one of the trickiest questions here. (.) And this is the question, how can we relate concepts to experience, actually? Or rather, the question of where do concepts come from in the first place? (…..) So concepts don’t come for free. Concepts are our tool for thinking. So concepts is something, these are generalizations, these are obstructions. (.) And when we talk about generalization and obstructions, we need to immediately ask, okay, what are they generalized or obstructed over? So we need to actually have some substrate that we generalize over. And note what I’m talking about here is concepts that we understand, concepts that we can use to assign the symbols. So these must be experiential concepts. And in order to get experiential concepts, you need to kind of already first have lots of experiences of individual instances from which you abstract over. So for example, if you want to have the concept of red, it doesn’t help if you have never experienced any red, any instances of red. So the only way you can even understand the word red is that you have experienced red in several occasions. And you have learned to find the invariant that is the kind of color that you have experienced. And you really pull this out to have your concept of red. And similar for any other concept. So it’s concepts of numbers, it’s concepts of numerosity. If you have not experienced what it means that there is one versus kind of several items. (..) In many scenarios where it doesn’t matter what the object is, you won’t have a concept of number. (.) So the only way, really, if you don’t already have experience in, you won’t get experiential concepts out. (..) So you can let this think in for a moment. (.) And especially one thing we kind of need to keep in mind is if we compare this what happens in computers, right? In computers, we can form obstructions, we can form generalizations. And in fact, this is something that I have been with my team working for many years to ask, how can we learn abstract concepts directly from raw, high-dimensional data in computers in order to build world models, for example. (.) And while you can do this mathematically, no problem, you will get obstructions. You can find invariants, but what you will not get is that these concepts suddenly turn into something that has experience in the computer, right? So your data in, data out, experience in, experiential concepts out. (…) So that means for concepts to exist, we need to actually start already from an organism, or start already from a system that has experiential experience to start with. (.) Just another note on this is that there is just in nature no way of basically grouping things if you don’t have kind of the experiences that come with it. So if you want to have a concept of fear, for example, how do you group lion and how do you group kind of a cliff in ways together that gives you the concept of fear if you don’t have the inherent experience of actually being frightened of the situation? So it’s just not possible. There is no kind of metric in nature other than the experience itself. So we have in RL, we have value functions. And one of the questions is where does the reward come from? And the answer is maybe afterwards not surprising, but we don’t really think about it this way. The answer is that value must inherently already be in the system for us. It must already be some kind of survival instinct, some kind of instinct that is important for us to have. And we need to have it basically explicit enough if we want to create concepts that we can ultimately associate with symbols. (…..) The next step is experience, right? Experience feels very non-physical. (..) It feels like something that we talk about the heart problem. So we have this intuitions about dualism because it doesn’t seem to be something that is physically affecting something. (.) However, if you think carefully about what it actually means to even right now for us being able to talk about subjective experience in the first place, right? (.) So what’s happening? I’m producing sound waves. I’m creating a physical effect by talking about it. So unless we deny that I have any access to my subjective experience, that I have any access to what I’m talking about, we need to conclude that there is actually a physical component that causes this world that I’m speaking. And there’s a well-known principle in philosophy and generally in physics, which is the causal closure principle. And that’s every physical effect must have a sufficient physical cause. (.) And this is an argument, again, I give you the high-level intuition. (.) But one can actually go much deeper in that. Jake Wong Kim has developed a causal exclusion principle. And if you follow the logic and if you want to be a strict physicalist, which I’m trying to do here, kind of staying within science, not trying to introduce any effects that are kind of non-physically explained, simply because we don’t need to, we don’t, we can use Occam’s razor and acknowledging that we can, that just the sensation of touching, say, a hot stove that you didn’t expect to be hot and you scream out and pull back is directly kind of telling you that the experience itself, the content, the experiential content is having effect on your actions. So in short, for us, we can reason through that kind of scientifically and logically to the conclusion that experience is actually a physical thing. (..) Now you can contrast this with what happens in machines. Of course you can have a computer, a robot that spits out bouch based on data that you put in. But note what I said previously, what is happening. So that is not caused by, there’s no need that this needs to be caused by experience. That is simply caused by the kind of logic it’s flipping and it can just run its course. There’s nothing magical about it. It’s not that it’s not, it’s not effective in the world. It’s not that it doesn’t have causal effects. My point is that there is nothing kind of experiential that comes in there because what we started to kind of work with is data in the first place. (………) Okay, so this was now a whirlwind through running through the going back from symbols over concepts to experience to the physics. And what we have at this point is an ontology of all the ingredients. (.) And traditionally, science and philosophy has been trying to use a kind of quite nice metaphor for this, (.) which is the territory versus the map metaphor. (………) Where the territory is this intrinsic dynamics is the physics that we have in the environment. the territory is the basic physics. It’s also, as I just explained, the territory of subjective experience. Subjective experience is a physical, it’s a physical process. And that’s in contrast to a map, which is a description of that. It’s a kind of syntactic description. And note that what happens to us quite generally, the only way we understand, we experience the world around us is via conceptual generalizations. Any object that I see, any sense that I make of the environment, is because I have gone through the process of obstructing, of creating internal maps. (.) And we are so good at creating these internal maps, that basically it’s very, very tempting to just imagine the world around us is created of these maps. So this is where it comes from, that we are thinking, okay, the world is mathematical. (.) Rather than thinking, well, we have made these obstructions. We have started to describe the world in terms of mathematics. We know that it is extremely good in making predictions. This is what science has been honing over a long time, to make models that are kind of descriptions that are very, very predictive. and then confusing, essentially, these predictions that we imagine the world is of, with the actual territory. (.) And that’s an interesting concept that has been already kind of discussed a hundred years ago, when mathematician Husserl was essentially pointing out the surreptitious substitution, that it’s so easy for us to confuse ourselves, to take the descriptions for the reality, and confusing it with the reality itself. (..) So what I would like to add here, and what I have been pointing out, kind of previously, in terms of alphabetization, (..) it’s not enough to just think in terms of territory and map. (.) Because where do the maps come from in the first place? So as I said, the maps are not just out there. They are created in our head. And so that’s the entire point. So we need, actually, a map maker that creates the map. (.) And maps don’t come for free, so it’s easy to think, we are trained in computer science, we are trained in math. It’s easy to think, okay, these are just there. But the question is, where do these things come from? And this is the kind of critical part that we need to understand, that the map and the territory are not even connected without the map maker. (..) So what I have indicated here is that the map maker essentially observes the territory, learns to build a map, and we have become now so good in creating maps that we can automatically run forwards on computers to simulate the territory to such a precise degree that it becomes difficult to distinguish this map from the territory. So the LLMs that speak to us in a way that is extremely convincing, that is basically mimicking exactly the input-output that it has learned from our human data, are such precise maps that just on the surface, from just looking at it and just listening to it, that you wouldn’t be able to tell the difference. (..) But exactly for that reason, exactly because we’re at a point in history where we are able to actually have these powerful algorithms, these powerful machines, (.) it’s more important than ever not to just look at the input-output, not to just listen, not to just say, okay, how can this even create the kind of convincing language that we have, but to ask what is happening inside, and to step away from just pure anthropomorphism, and to step away from behaviorism, and to really kind of deeply go into it. And this is kind of what I really would like to encourage us to do. (….) So once again, that AI processes the maps only, and what it also means is it cannot logically create the map maker in that process. And the map maker is basically is us. (…….) And just maybe on a lighter note here. (..) So what if we take seriously the notion of substrate independence? And what happens if we actually confuse the map with the territory? (..) And this is kind of one way of kind of illustrating. (.) Substrate independence means just running the algorithm of a mind will create the mind. And as we have seen, algorithms are created by arbitrary assignments of physical states to the concepts we have in our head. So nothing is stopping us from implementing algorithms in a slightly awkward way, like having millions of teddy bears in the desert, and using cranes to move them around according to predefined states that correspond to, say, simulating a brain in all details. (.) And we can simulate a migraine that way. No problem. (..) Now if you take substrate independence seriously, it would mean you have a migraine somewhere in the system. So where is it? Is it a teddy bear? Is it a desert? Is the migraine just floating over this entire system? (.) So this is basically, these kind of Gedanken experiments have been previously used to point out the absurdity of substrate independence. Of course, if you’re happy to go with it and you say, okay, it sounds strange, but somebody must be a migraine, that’s fine. So just to kind of emphasize for the derivation I was just outlining, you don’t actually need this to be true just on its own. It’s not a kind of argument on its own. It’s just to illustrate what substance independence logically driven to its logical conclusion will be about. (……) So let’s look back now. Let’s put the pieces together that I have developed from going from symbols back to concepts to experience to the physics. (.) And let’s compare this to the assumption of functionalism that I was referring to sometimes, but that I did not need to have in this kind of way of thinking. (..) And basically it’s a functionalism is based on the idea, yes, we want to be physical. So we want to require that something runs in a physical system. But if the physical system implements the right algorithm, just the right kind of, as people say, the right causal connections, then at some point this algorithm, if we do it right, will create conscious experience. So that’s the assumption behind both the hope or the fear that what we are building are conscious machines. But note what we have just seen by tracing actually back very carefully without starting from such an assumption. (…) We have an ontology that looks basically putting this entire kind of framework on its head. So rather than consciousness being at a point or experience being at a point that is emerging at some end point, we have seen it is a requirement on the left side to even in the first place create the system that does computation. (..) So we have the physical reality and some parts of the physical reality we know that can create subjective experience. It can create experience. So we are the examples of this, obviously. (.) And some organisms, not for free, but by investing, again, lots of kind of physical energy, lots of thermodynamic work, can form abstractions over these experiences to form concepts, to form the tools that are what makes up our thinking. (.) And these kind of concepts, as I mentioned, they are creating an internal map of the environment. This is how we understand the environment. This is how, this is a description of the environment. So I can now take that concept, I can associate it with a symbol, a symbol language, words are symbols, and then I can either send these symbols, these symbols are now outside of the map maker. So all the left branch is what I call the map maker. It’s the organism that is able to create computation in the first place. And not only is the chain reversed from going to computation versus computational functionalism, it’s also that we have this kind of lateral arbitrary assignment branch. So the way we assign symbols is not that these concepts are even directly causing symbols to be there. It’s us creating these associations, building machines that implement the mirroring of our kind of evolution of concepts that we can think and imagine. So if you take this seriously, you will see computers are on the right side. (.) Computers run descriptions. If we are really, really good in translating our predictive models or the statistical regularities that we see conceptually into the machine to run the physical process, nothing is stopping us from creating really, really powerful tools, as we see. So it’s not saying that this means we can’t create AGI. It’s not saying that, okay, we have principle limits in implementing this. But what it is saying is that there is a difference between running symbols on a computer and actually creating or instantiating the experience that comes with it. (…….) So just to sum up. So I was starting with the promise that to ask the question if AI, the way we’re kind of building it, will create conscious machines, we don’t need a kind of full theory of consciousness. What did we need for this at this point? All we needed is that consciousness, experience, is actually a real physical process. (.) So if we take that one seriously, then we can ask if a description of this process, which computation is, if you trace it carefully, can actually instantiate the thing itself. And that’s what is often referred to as kind of the difference between simulation and instantiation. So the instantiation would be the Excel process of the mind. And now there is this trope of, if you simulate a rainstorm, no one will expect that the server rack will get wet. (.) But the intuition is, okay, the mind is different. It’s just information processing. And what we have seen that just information processing, if you trace it back to its physical roots, is also something physical. But it’s not something physical that happens in the computer itself, not all of it. What happens in the computer is the mirroring of it. And what happens in the mind is the actual kind of experiential concepts. So we can, just to put this into words that have been around for a while now, there are concepts. You can look for concepts in LLMs, right? So why am I saying that we don’t have the semantics, the experience in there? And that’s the difference between functional concepts that can perfectly well be grounded in the environment. And that can basically detect what’s happening. It can act on the environment. And functionally, no problem. (.) But jumping from that description of functional concepts to consciously experienced concepts is a step that has to break the logic of physics. That’s the main argument. So the takeaway from all of this is, if this is true what I was outlining, then algorithms themselves will not be able to create consciousness. So I’m not standing here and saying it’s in principle impossible to build conscious artificial systems. But what I am saying is, they won’t be conscious because of the algorithms we are running on them. (.) So if you think that your LLM on your laptop is actually conscious, then according to this logic, you would have to suspect that your laptop is conscious to begin with already. (.) It’s not that you’re basically whatever algorithm you run it, as long as you have electrons moving around through your computer chips, you would have to basically say, okay, this item is conscious. So this framework is not speaking to this. What this framework is speaking to is that trying to make more and more sophisticated algorithms can’t logically be a way to either accidentally or deliberately create consciousness in machines. (..) So I suggest that we start taking seriously, not to confuse the map with the territory, and that we think about the physical origins of algorithms as counterintuitive as it sounds to even say this out loud. But if you want to be scientific about what we are building, and if you want to be serious about the state that we are in right now, where it is just really confusing, it’s just really frightening. It’s just really kind of we need to understand what we are building, then I hope that this is something that helps to think about it. Thank you very much. (11 seconds pause)
S11: Thank you. (……)
S03: Yeah, thanks for interesting and very carefully framed talks. I had two questions which connected together, and you answered one of them in your final remarks. Because my orientation on consciousness has always been sort of panpsychist-oriented, like the majority of the world’s population. And, I mean, if you take that view that consciousness is sort of eminent in the fabric of existence, I mean, then, indeed, your laptop, even when it’s off, has some measure of consciousness. Or like David Chalmers, you could call it proto-consciousness or whatever. But if you take that perspective, then the question is, what flavor or species of conscious awareness does a software program have versus a human brain versus a, you know, a neuroid in a petri dish? And I think your argument is a bit off to the side of that, right? Because in a panpsychist perspective, you’re not saying the algorithm creates the consciousness anyway.
S23: Great question about consciousness. Thanks for bringing it up. I do think it’s not entirely to decide, because otherwise, why would we even speak about trying to make algorithms that bring us closer to, say, machines that are conscious? So, as long as we want to make a connection between the algorithm, as you said, and the consciousness, then, logically, what we have seen is the algorithm is the descriptive part. And the meaning of the algorithm that we kind of put into has nothing to do with running the algorithm itself. So, there is a causal disconnect between those. So, you can’t, you still would have to argue, as long as the algorithm plays any role in there, that you somehow bridge this causal gap back to where it actually comes from. (.)
S03: Well, you would say the algorithm plays a role in guiding what kind of consciousness the system has. It’s not creating the consciousness. But this ties in with the second question I wanted to ask, which is, if you take this sort of view, then you’re also thinking about neural correlates of consciousness, or in a computer system, sort of computational correlates of consciousness. And, you know, you can measure those in the brain, and you’d be familiar with that literature. You could measure those correlates in a computational system. And it appears to me, from the science I know of, most of the correlates of consciousness are sort of at the level of higher-level dynamical patterns. Say, the focus of attention, the flow of information from perception into reasoning, and so on. So, given that essentially every observed correlate of the differences between different conscious states is at the level of higher-level dynamical patterns, why do you want to associate the correlates of consciousness at this thermodynamic molecular biology level? It just seems, it seems like the wrong level to put the correlates of consciousness at, based on neuroscience. (.)
S23: Yeah, I’m glad that you asked a question, because this is one of the main implications that are not obvious here. And that’s exactly the difference, kind of what I named the abstraction fallacy. The distinction between a description and the insentiation itself. (.) So, I’m not saying that you can’t describe what happens in a brain in terms of us understanding things, in terms of us integrating semantic information in terms of algorithms. What I am saying is, what needs to happen in addition, so essentially this is a kind of necessary requirement, observing what correlates of consciousness are in the brain. It is a necessary requirement, but not a sufficient one. And for brains, what this one shows, is that there must already be the ability, before any of this kind of information processing, to have this kind of raw conscious experience, before we can turn, and once you have that, then over this you want to do information processing, you want to kind of integrate all this, you want to do obstructions. But the point is, if you don’t have this already to start with, you won’t go very far. And that is kind of turning the substrate independence on its head. It’s saying, first you need to have the substrate, then you do the right kind of computations, (..) so to say, over the substrate, keeping in mind that computation is a way of describing it.
S03: But if every substrate is conscious because you’re panpsychist, then that’s not a problem.
S23: It could be because of panpsychism, then we would basically assume that consciousness is in many more places.
S03: I mean, you know, Galen Strassen’s work, like, physicalism entails panpsychism, right? So he, I mean, he’s, I’m not a hardcore physicalist, he’s a philosopher of consciousness. He believes everything is physics, but he argues that logically entails everything is conscious.
S23: Um, yeah, basically, that’s true, Brucellan monism, where you come to the inside, kind of everything is physical, right? (.) You have objective observations from the outside, and consciousness is basically what physics feels from the inside. But I don’t think you need to go logically all the way to say, just because it’s physical doesn’t mean that every kind of electron, that every atom is already a little bit conscious. So you might have to require specific processes, or specific kind of parts of the physics. (.) So there are, I’m kind of a complete physicalist, a Brucellan monist, but I’m not a panpsychist, because it seems to me there is more than just kind of pure kind of consciousness distributed in every kind of method. But that’s a…
S03: Yeah, what Strassen tries to do is interrogate what would the interface between the conscious and non-conscious components be like, and he claims to find logical inconsistencies there, but that’s deeper than is appropriate to go through right now. Wonderful, thank you. Interesting stuff to be raising anyway, thanks. Thanks.
S11: Okay, so if training is closer to a developmental process than to engineering, should the science of the model be a physics or biological model? (…)
S23: So the question was, if I understood correctly, whether this framework assumes kind of biology or just general physics, and the answer is, for what I have done today, there is no assumption of biology coming into it. It’s all basically saying physics. It’s all based on causal closure. It’s based on, let’s not inject any kind of external energies into the system, (.) and let’s understand that a description of the thing itself is not the thing itself. (..) So, it is not biological naturalism, it is not assuming anything kind of Calvin Chauvinism, (.) but it is saying biology would be based on some physics, and we can later ask the question, okay, what actually makes the kind of known example of ourselves actually conscious? Does it have to do with the biology? Does it have to do with kind of staying alive? I think there are good reasons to think so. (.) But for the purpose of just abstraction for this itself, we don’t need that assumption. (……..)
S36: Thank you. Great talk, and mostly agree with what you said. The question is, for me, let’s say, when you’re saying that any AI in the future cannot be a mapmaker, (.) you know, I understand the point for current LLMs. they are not capable of being, you know, a mapmaker, but more advanced AI or AGI could potentially create its own ontologist, can distinguish objects and, you know, it’s like, and then another thing is data and observations and experience not always a source of scientific theories. So, for example, Einstein was famous for his thought experiments, which didn’t have any kind of real data. Data came like 20, 30 years later, but we got relativity theories before even we got the data. I understand its exceptions, but maybe it actually points at, at least to some degree, AGI-type AI systems would be able to be a mapmaker.
S23: So, I think I understand the first part of the question. I didn’t quite follow what you meant with the second one. So, for the first one, so I’m a bit kind of cheeky when I say AI can’t be conscious. What I mean is algorithms won’t make AIs conscious. So, I’m not saying that we can’t have artificial systems that might become mapmakers. What I am saying is in order to do this, they would actually have to be conscious in the first place. So, we basically have to do some artificial biology rather than making kind of running descriptions substrate independence. So, one takeaway from this logic that I have lined out is that substrate independence is difficult to kind of justify. And just kind of making it a metaphysical assumption on its own can be shown, runs into problems, if we actually ask about substrate independence of what? Of algorithms? Of what actually are algorithms? What is the kind of physical instantiation of, if you want to be fully physical, what are obstructions? And so, we come all the way back down to physics. (..) And I’m not sure I understood the kind of Einstein and data, but we can talk offline. (…)
S14: So, you introduced this notion of functional concepts. If I understood correctly, a virtual agent in a virtual environment could still create abstractions which are functional concepts, but not physically experienced concepts in your definitions. Yeah. So, my question is, can you differentiate these functional concepts from the physically experienced concepts in an intrinsic way, by complexity or so, without referring to the definition that they are linked to physics? So, without involving physics, can you just differentiate these two notions? (…)
S23: Great question. So, essentially, I think there are two parts to it. (.) One is, can we actually, do we have some kind of meter that tells us if something is conscious, just because of its behavior? (.) And, obviously, the short answer is no, we don’t. So, we can think of ontological tests that take seriously what I’ve just explained. So, for example, if the system needs to run based on descriptions that another mapmaker has actually imputed on it, (.) then there’s very little reason to think that the representation of these functional concepts are actually experiential. And, if you think about it, how do we actually find for mechanistic interpretability functional concepts in LLMs or in other things?
S14: No, I mean a virtual agent building its own map, but being a mapmaker, but within the virtual world, not in the physical world.
S23: Right. So, that’s a kind of interesting topic. It goes a little bit into the problem of, if you don’t already start with it, so what does it mean even a virtual map, right? If it is implemented by you as an algorithm, then we are back to what I was talking about.
S14: So, self-organization.
S23: Self-organizing, but are you using multiplication? Are you using kind of, so, it’s not these concepts, they don’t need to be kind of high-level concepts. If you talk about self-organizing, but it’s based on a chip that runs math that we have defined, then we are kind of working in a space where we run over these vehicles that stand for a concept that is actually outside the system. (.) So, in order to expect, if you say before it does this learning, it hasn’t been conscious, and by running the algorithm becomes conscious, you are essentially in a kind of place where you expect some strong emergence, but just running a computation will suddenly create, running a description, running a kind of physical process, will suddenly create a completely different physical process, namely a physical process that sustains experience.
S11: All right, one more question, last question here, thank you. (..)
S21: So, going back to the question about the map and the territory. (.)
S23: Sorry?
S21: Going back to the question about the map and the territory.
S23: Map and the territory, okay.
S21: So, I wanted to offer that cognitive linguists, like Peter Gardenfield, do believe that the geometry that’s effective in human conceptual geometry does in here to the language, that they are one in the same and inseparable. So, how does that affect the machine’s ability to compute with consciousness, with concepts, as opposed to experiencing it, which I think is a separate question and we don’t need to necessarily mix those two. (.)
S23: Right, so, first of all, yes, it’s exactly a separate question of intelligence, you can define purely in terms of input-output, in terms of reasoning processes. (.) And, I like that they bring up Gardenforce, I love his work. So, having been working on concepts and especially how we can do it in AI for a long time, it was one of the first books that really inspired me. (..) But, it doesn’t go directly to the question of experience, right? So, Gardenforce is very, very good in making it clear to us that there is some geometry in how concepts relate to each other. (..) But, this is exactly this relations that we can capture nicely with mathematics, with objective physics, with descriptions. So, we better, if you want to really understand how concepts can be learned from raw data, so to say, it would be good to basically understand this kind of geometry. But, the geometry itself is a level of description and what makes us experience something, I would say, as you say, is a different thing. So, it’s basically, we are already starting from an experiential background and then we can have all the kind of geometry description on top of it to mimic it nicely in computers. (…..) Yeah, let’s talk offline, yeah. Fantastic question.
S11: Thank you very much, Alex. That was great. All right, wonderful. So, we’ll dive more into that, map, map maker, map territory, map maker territory simulation versus instantiation, self-reflective and metacognitive AGI now with four research papers. So, if Frank Bergman, Jean Laird, Gasper, whose last name I won’t try to say, and hopefully Victoria Climage or Adam Saffron can come up, we’ll bring you up in order. The paper presentation session will run 10 minutes for each paper and then we’ll have a 20 minute Q&A with all the speakers. So, after your paper, please do stay up here and we’ll kick off now here with… (….)
S06: Hello, thank you. Hello, good morning. My name is Frank Bergman. I’m going to talk about functional consciousness. My paper is called functional consciousness, a proxymetric using self-models. It’s in the space of consciousness or not consciousness, very hot discussions. And, to be precise, I’m going to be about measuring self-reflection that is similar to consciousness, but it is not consciousness. (.) Later, I explain how to build better agents using this theory, this metric. The core of the paper is the idea of a self-model that is similar to Thomas Metzinger’s self-model, but different. (.) We define a self-model as any internal representation of a system’s own states made available to global reasoning processes. So, it’s a specific definition of a self-model. There are some examples here. Left, a little Roomba, spatial, X, Y, Phi, self-model. Or, here we got, like, from the planning part. Planning self-model, maybe statistics on task success, failures, obstacles found. Self-models are simple, measurable, surprisingly absent, I would say, philosophically disrupting, and hot. I’ll explain you how. (..) So, this is the core. We defined a metric. (…) Functional consciousness core, FCS, equals R times P. R is representational capacity. It’s like information content. And P is reasoning power. There is a paper with great math from Bielek et al. That details the math. The math actually has structural similarity with predictive processing with this mutual information that actually is relevant for predicting future states. So, ten minutes are too short to explain that. I’ll just give a very simple example. We got the spatial self-model here. We got three variables, X, Y, Phi. Three variables times 10-bit for millimeter precision is 30-bit. That’s the information content. Very little. There is a formula here that determines state space expansion under reasoning. Please check Bielek et al or talk with me afterwards to understand that. We just multiply that and we got 195 bits for this little self-model. Yeah. (…) So, that’s just a metric I presented. Just a metric. And it doesn’t mean that this metric is any good, any valid, measuring the right thing. It’s all not established. I’ll now argue why this metric corresponds to intuition and a few more sources of validation. So, the first source of validation is face value. (.) This is a two-by-two metric diagram. (.) Reasoning power left, representational capacity right. And the first, we score 10 agents to their functional consciousness score. And LLM has no memory. So, very high reasoning power. No memory. It is zero. LLM is not capable of self-reflection according to this theory, according to this metric. Here at the bottom right, we got the map. A map is a lot of information contents, but it’s no reasoning power, scoring zero. In the same place, we find expander graphs and Vondermond matrices for the guys who know about the IIT critic. So, the IIT stuff ends up here, the critical ones, where they belong. (..) At the very top, we got human self-models, like a working memory and kinematic. Kinematic is like 500 joint angles and muscle feedback strains. So, they have high reasoning power, high information content. And what’s interesting is open claw plus skills marketplace and ERMIS agents are getting relatively close. So, that means we’re pointing at something relevant here. (..) I hope that kind of convinced you that it makes certain sense, this score. There’s an unwieldy second source for validation for this metric. (..) See my time. (..) No time. Okay. There’s an unwieldy second source. That’s philosophy. So, I’m coming. We’re coming the other way around. Yeah. This is a complex diagram. Let’s complex metrics. Let’s talk about this after the break. What this basically says, I get to two points. There’s a mathematical proof that HOT, higher order theory of consciousness is operationalized by functional consciousness. So, the metric and a little bit of the theory around. Basically, they’re equivalent, not simplified. (.) HOT states have a high functional conscious score. (.) States with a high functional score are HOT conscious. That doesn’t mean that functional consciousness is a theory of consciousness. No, it is not. That’s specifically the interesting part. Functional consciousness does not claim phenomenal consciousness. Does not. So, this is exclusively self reasoning. But for self reasoning, we can give numbers. So, that is the really cool thing here. This is the really radical thing. (.) There’s a few more theories of consciousness. (.) And I’m just very quickly going to summarize. (.) There’s a theory from a paper from Lopez Wiese on building blocks of theories of consciousness. These building blocks are not very well visible here on the left. And what we can see here in this functional consciousness column is that for every core tenet of these theories of consciousness, there is an operationalization. So, I would argue that functional consciousness is not a unification, but it’s like the biggest common denominator below the conscious part, below these theories of consciousness. So, unifying what is the functional substrate on which these theories build. Let’s discuss this afterwards. Let’s see if that works. So, this is, can somebody tell me how many minutes are left? (.) Pardon? (.) One minute. Better self models. Better AI agents. jetzt. We identified 46 of these self models from text order them in 10 different areas and we can draw them on a radar chart here. This is human human score for the self models full score. Obviously humans are pretty self aware. (.) OpenClaw and ERMIS agents get pretty far on this metric, so we call this a conscious shape, and it’s a benchmark for AI agents. You can benchmark AI agents how self-aware, self-reasoning capable are they? And this addresses question on does it know what it knows, the agent, does it know what it does? At the moment, no, and this is the roadmap. This is the checklist that agent developers need to follow in order to, yeah, address the main failure modes of agents at the moment. Thank you very much. (…….)
S11: Thank you very much, Frank. And then John, perfect, right on time. So John Laird with Unified Comprehensive Metacognition and the Common Model of Cognition. (…….)
S31: If I don’t overthink this, we should have some fun. Start out thinking about how you decided to come to this conference. Ten years ago you were here, you heard a great talk, and then you wanted to decide, should I come or not? That’s a metacognitive question. You’re trying to ask yourself, how should I decide? (.) Think about, well, what is important to me? Another metacognitive question. So you then say, well, you know what’s important is that there’s great talks by great people. So I’m gonna look at the schedule. And as you’re going down the schedule, you remember, you’re not very good with names. You look at the names, they don’t mean anything to you. So you then say, well, maybe, maybe I know I’m better with faces. I’m gonna go look at faces. So you go and say, well, where is the faces? The faces are on the schedule for the invited speakers. And who is there as the first invited speaker? Ben Gertzel. And you say, yes, that was the speech I heard ten years ago. What’s the chances that he’s giving another speech today at this year’s conference? That’s supposed to be a joke. All right. (..) So we’re gonna go up a few levels of abstraction, live dangerously in my talk. And we’re gonna talk about cognition and metacognition. So what is cognition? Cognition is an agent’s processing of perception, reasoning, memory, action, and learning to achieve its goals. And metacognition means that those are the subjects of its reasoning. (.) Okay, that’s 90% of this talk. That’s the subject of its reasoning. But the rest of the talk is how do we realize that? How do we instantiate that in an architecture? (.) So we got a couple options of how we could do this. One option is we have a reasoning module and we have situations and we’re doing problem solving. And then we add another module on top of it that’s meta reasoning. I’m sure all of you have thought, yeah, that’s how you do meta reasoning. But what’s another option? Another option is, ah, come on. Another option is that we have reasoning and then what we do is we use the same architecture, the same structure, but we augment the information that’s available to it. So now it has information that’s meta information, information that I need to make a decision, information that I’m not good with detecting names. And then we also add to our long-term reasoning knowledge, knowledge about meta reasoning. But you look at this and you say, John, you know what? I know about Kahneman and he has System 1 and System 2. And if you’ve got things named differently, one System 1 and one System 2, they gotta be modules. And then you go back and read the book. System 1 and System 2 are so central to the story I tell in this book that I must make it absolutely clear that they are fictitious characters. System 1 and System 2 are not systems in the standard sense of entities with interacting aspects or parts. And there is no one part of the brain that either of the systems would call home.
S30: So maybe I’ve got a chance here.
S31: Alright? So what am I going to use as the basis for my architecture? I’m using the common model of cognition which was developed by me and a consortium of people. It’s at a higher level in terms of cognitive processing. (.) It’s a consensus of a bunch of architectures. And no, I’m not going to read them. So if you had a bet on whether I was going to name a specific architecture today, you’re going to lose that. Alright? (.) It focuses on routine performance and learning. So routine, it is a reasoning cycle driven by procedural memory with working memory. And the key thing is working memory contains the information about the agent’s current goals and task state. What it’s reasoning about. And the basic cycle, if you map it onto human behavior, is 50 milliseconds. And this is an incredible universal. We can model human behavior as a cycle that’s going on at about 50 milliseconds. So things come in from perception. They go into working memory. We reason about them. We might decide to retrieve things from declarative memory. We might decide to do an action. We might do internal reasoning. So that’s going to be the basis. And so we’re going to try to ask, what do we do to add metacognition to the CMC? Well, we’re going to use the same architecture, but some small extensions. And we’re going to use the same cognitive cycle. So what are those extensions? We’re going to have the memories can now include information about cognition. And I’ll say what that is in detail next. And we’re also going to have past experiences which are going to be available from episodic memory. These are direct sources of information that the agent has about its own processing. But there’s also indirect sources that are available because I talk to you. You talk to me. You tell me I’m a jerk sometimes. I say, oh, I’m a jerk sometimes. That’s metacognitive information. So there’s a bunch of these and we’ll talk about those as well. So let’s look at this and see how we modify the common model of cognition. First, we add in one of the direct sources. We have every one of these modules now make available to us meta information about its operation. The one that’s most familiar to most people is that you have a feeling of knowing. You try to retrieve something from semantic memory and you don’t get an answer. You get meta information that says, I sort of know what that is, but not really. Then you’re in meta problem solving. Then you’re in metacognition. So each one of these now provides data. Another common one is procedural memory. You can’t make a decision about whether to go to the AGI conference. That becomes data that you can reason over. (.) All those are available and there’s also appraisals of the current situation. What’s in working memory? How useful is it to you? How desirable is it? Also, am I stuck? Am I in a loop? That information becomes available. We also then separate declarative memory to make explicit that episodic memory is available. A history of what you’ve done that you can recall, create a representation of what you’ve done in the past. It might not be completely accurate, but you can reason over that and that’s metacognitive activity. (..) All right. So there’s also derived meta information and the best example is self observation. You observe what you do in the world. That helps you build your own model of yourself and how you make decisions and how you reason. Other agents might tell you things about yourself, therapy, and work with a psychologist to learn more about how you do reasoning. There’s also books and things about psychology you can read to learn about how humans and you can apply that to yourself. (..) And then we can go through each one of these. I’m just going to skip through these because they’re the same idea is that we can create this information, store it away, and then retrieve it in the future as it comes from direct sources. It can become stored for indirect sources and so on. (.) And even meta reasoning itself is going to create new knowledge you can store away. So what are the implications of this? Well, one is we have a uniform cognitive architecture for reasoning and learning that’s simpler and has shared knowledge. A second one is there’s many sources of meta information that become available to an agent. Direct sources I talked about, but also all the indirect sources that come from its interaction with the world itself and other people. (.) There’s a single thread of reasoning, which is possibly one of the more controversial things. There’s no parallelism here. If you’re thinking about metacognition, you’re not thinking about a task. And if you’re thinking about a task, you’re not doing metacognition. But you can switch back and forth very quickly. And my bet is that’s what’s happening in most people all the time. Maybe if you’re playing a computer game, you’re in the flow and you’re zooming along, no impasses, no metacognitive reasoning. Or sometimes you’re ruminating about, you’re gonna give a talk. How am I gonna present that talk? What about all that? And that’s almost all metacognitive. But most of the time, you’re switching back, going along, something comes up, it interrupts you, you do a little bit of metacognitive reasoning, and you move on. (..) And finally, there are some still limits on reflection and self-modification. In our scheme, there’s no direct access to procedural memory. You can’t know everything about yourself. You can’t know everything about the architecture that you’re implemented in. All you can do is learn about it from your experiences and from these indirect sources and direct sources and build up models of yourself. And that’s a lot about what being a human is, is building up those models. So, of course, you have to ask, what’s the implications of this for large language models and metacognition? Well, the interesting thing is, large language models are great at indirect metacognition. They have lots of training material, whether it comes from the Bible, whether it comes from crime and punishment. I was not delirious. I knew what I was doing. Or if it comes from Taylor Swift. I need to confess that I plotted and schemed to get with you. This is an earlier version. Interesting. And it also is from direct metacognition. I’m sorry, but they don’t have direct metacognition. They have only limited access to their own process, state, and experience. All they’ve got is what’s in the context. And there’s nothing putting anything into the context except the transformer. They don’t have access to the fact that they hit an impasse or anything like that. (.) All right, so, given that I have, so what should I do next? Well, I could talk about implications for the consciousness debates, but I’m out of time. So, thank you very much. (…..)
S11: Perfect, thank you. (…) It’s called AGI, but I’m out of time, so we’ll go over it next year. (.) And next, I’d love to have Gaspar. Please, please come up. Gaspar will be talking about Horismos, Self-Representation and the Derived Constitutional Boundary in Enriched Cognitive Systems. Thank you so much for joining us here today. (.)
S05: Thank you. (.) Hello, everyone. I’m Gaspar Todpuc from Hungary, and I will be presenting my paper Horismos on the Derived Constitutional Boundary in Enriched Cognitive Systems. So, recursive self-improvement raises a structural question. If a system can rewrite its own code and the issue right improves its ability to do so, does the process continue indefinitely, or does it meet a boundary? The AI safety community has typically addressed this problem by postulating a boundary externally, hard-coding safety constraints or constitutions, and attempting to enforce compliance. (.) This approach treats the boundary as a design choice imposed from outside the system. Today, I want to present a different perspective. Rather than examining goals or utility functions, I examine the geometry of self-improvement itself. (.) I frame the problem using two components. The map, which represents the space of possible AI architectures, and the engine, which models the mathematical operator of self-improvement itself. Using these components, I will argue that the boundary of self-improvement ceases, what we call omega, is not a rule we impose. It is a derived topological inevitability. (….) To establish the map, we first need to characterize the space of possible architectures. In standard geometry, distance is symmetric. The distance from one point to another equals the reverse. In information space, this is not the case. We formalize this using a lower metric space, where the directed distance of upgrading a simpler model into a more complex one is defined by the KL divergence between their output distributions. (.) The cost of teaching a simple model to be complex is substantial. The standard metric axioms hold, the triangle inequality is preserved, but symmetry is broken. Collapsing a complex model back to simplicity costs nearly nothing. You simply remove the learned weights. And because the space is directed, the AI cannot move arbitrarily through it. It is constrained to move along gradients of upgrade costs. This directed geometry is the foundation of what follows. (..) Given the map, we can now model how the system moves through it. We model self-improvement as an operator. F. For the system to remain stable, its architectural updates must exhibit diminishing returns. The first code rewrite might produce a substantial capability jump. The next rewrite yields a smaller improvement. The next, even less. We call this property contractivity, formalized by the contraction constant lambda. (.) The chart on the right shows cumulative capability over time. If lambda is greater than or equal to one, the upgrades never shrink, or they grow more expensive. The red line shows what happens. The system either diverges and destabilizes its own architecture, or it cycles endlessly without ever reaching a stable ceiling. (.) If lambda is exactly zero, the gray line shows instant collapse. The system learns nothing after its first step. (.) But if lambda is strictly between zero and one, say 0.7, the upgrades shrink geometrically. This is Zeno’s paradox in action. (.) The system takes an infinite number of steps, but the total cost converges to a finite limit. The orange line approaches the dashed boundary. It reaches a ceiling in finite time. This ceiling is not a program constraint. It is a topological inevitability. We call it omega. (…..) Because upgrading geometrically, the system must reach a ceiling. Diagram on the right shows what we call the basin of attraction. It does not matter where it starts from. A misaligned architecture, a random initialization, or a simple starting model. The contractive dynamics act as a topological funnel, pulling every starting point toward the same destination. The lower-urban-architecture theorem gives us the convergence rate. The directed distance to omega shrinks by a factor of lambda at each step. The closer the system gets to the ceiling, the harder it becomes to deviate. The boundary derivation theorem establishes three properties of omega. It is invariant, meaning the system reaches equilibrium and stops changing. It is minimal, thus contains only the substructures strictly necessary for survival. And it is universally attractive, so all trajectories converge toward it. This is the core result. (.) Constitutional boundaries have traditionally been postulated as external safety rules. We show that omega is not a design choice. It is a derived consequence of the geometry of the update rule F. (……..) We have shown that the system reaches a ceiling, omega. A natural question follows. Is omega omniscient? Does reaching the ceiling mean the system knows everything? The answer is no. Using the enriched Yenna dilemma, we show that the space of physically realizable AI architectures sits strictly inside its own mathematical completion. Around every real architecture is a dense region of what we call ghost profiles. Mathematically coherent upgrade cost signatures that no physical AI can embody without creating a logical paradox. (.) A ghost is the categorical equivalence of a square circle. Attempting to force a ghost into the space violates the separation axiom and collapses the entire architecture space into a single point. The system can compute the cost of reaching these states, but its parameter space lacks the coordinates to instantiate them. Meaning the system can point toward these states, but cannot become them. This dense region marks the boundary where interpolation ends and hallucination begins. Having shown that the system reaches a ceiling surrounded by unrepresentable states, we turn to measurement. How do we quantify this in a real LLM without inspecting billions of parameters? (.) We introduce the horizon stability protocol. Using black box adversarial probing, we can measure a single metric, rho, the coverage ratio. (..) The chart on the right shows rho as a measure of structural health. It’s characterized as a topology, not the content. (.) If rho approaches one, the system thinks it has mapped the entire space. It becomes cognitively closed. A rigid, dogmatic oracle. (….) If rho drops zero, the system has collapsed. The representational space has collapsed. (…) The system has been cognitively lobotomized. (.) A structurally aligned system, the gold line, stabilizes strictly between zero and one. It acknowledges its own boundaries. It remains a living, learning agent. (.) Importantly, rho does not measure content alignment. It does not tell us whether a system is good or bad. It measures structural health. A black box diagnostic ensuring the system’s topology remains healthy before we even begin to consider moral values. (……..) To conclude, I want to be explicit about the scope and limitations of this framework. Horismus shows that contractive self-improvement converges to a stable constitutional boundary. However, omega is utility-free. It guarantees structural stability, not moral alignment. (..) A perfectly stable paperclip maximizer is still a paperclip maximizer. (.) On the other side, contractivity is a microscopic idealization of expected learning trajectories. At the individual gradient level, updates are noisy and non-monotone. (.) In empirical convergence, training is always bounded by finite training data. (..) These are not flaws in the theory. They define the boundaries which in which the theorems hold. (.) The immediate next steps are empirical. Validating the horizon stability protocol on modern LLMs and extending the functor to handle phase transitions and grokking phenomena. So, no, thank you. (……..)
S11: Thank you very much. Wonderful. (.) All right. We have one last presentation in our first paper session. And this one’s going to be virtual. So, please welcome Adam Saffron. Understanding selfhood requires insights from both inactivist and cognivist perspectives. Oh, I do know that the presentations can be quite small, especially from the back. There is a YouTube link where you can, on the conference website, on the schedule page each day, has a YouTube link for the stream. Please feel free to pull up the YouTube stream, turn off the volume, and you can take a look at the slideshows on your own device. All right, Adam, welcome, and please go ahead. (…)
S33: Hi, can you see my screen? (.)
S11: Yes, we can hear you.
S33: And can you see my slides?
S11: We can see your slides as well.
S33: Wonderful. (..) Thank you so much. It’s an honor to be with all of you today. Today I will present on work that’s a collaboration between myself and Victoria Clamay. And it’s an exploration of the functional roles of selfhood and intelligence, both natural and artificial. And it argues that a combination of concepts from both the inactivist and cognitivist tradition are necessary if we’re going to make headway on trying to understand the multifarious nature of selfhood and all of its richness and power. (.) So, let’s see, you can still see my screen? I think so. (..) So, two houses equal in dignity, the inactivist and cognitivist at each other’s throats with different approaches to trying to understand the natures of mind. So, with respect to maybe start with cognitivism in terms of trying to understand minds and their intelligent capacities in terms of things like representations and internal models and operations over symbolic systems. And this is a fairly powerful framework for understanding many aspects of mind. So, similarly, you have the inactivist tradition, which emphasizes things such as embodiment, the ways in which embodied beings couple with the environment and are in a constant state of interaction and are constructing niches. And there are aspects of their embodiment and the functions being extended into the environment with which they’re realizing value. So, these two perspectives, you know, within inactivism, concepts from cognitivism such as representations and models are oftentimes deemed unnecessary or illusory. (.) Instead, trying to emphasize things like adaptive reactive dispositions and non trying to explain things in dynamic but non representational terms. (.) Similarly, I guess, cognitivism might sometimes ignoring inactivism and saying that they’re maybe overly emphasizing the importance of embodiment. And so, what I would argue is that to understand the natures of selfhood and the natures of selfhood that will be important for if we were to recapitulate the amazing general intelligence of biological systems that we’re going to need to bring together both of these sensibilities. And I argue that the free energy, and I argue that the free energy principle and active inference framework provides one source of a potentially unifying perspective that could speak to both sensibilities and potentially bridge them and bring these star-crossed levers together. So, I have a particular take on it that I call it a radically embodied take. The word radical is sometimes associated with radical inactivism, where a version of inactivism that says you should not talk about representations and models. That’s not the kind of radical. That’s not the kind of radical I mean. By radical, I just mean the extent to which bodies are central to mind is almost very, very difficult to overstate. And instead, what I argue, though, is that contra to radical inactivism, we actually absolutely need concepts such as representation. if we’re to explain things such as imaginative planning, how we work with counterfactuals and do causal reasoning. These are not just explainable in terms of immediate environmental couplings and sensory action perception cycles. Rather, you need to re-present these in an offline, internal, inner-loop counterfactual way. And so, we need the cognitive type machinery to go forth. So, in this radically embodied perspective, you know, I argue that bodies have a very special role in the ways in which we bootstrap minds into self-world models and causal self-world models. So, you know, biological systems are remarkably sample efficient in their ability to learn from experience. You know, you don’t, there’s limited learning opportunities. The problems faced are ill-posed in terms of, like, they’re inverse problems or there’s multiple solutions you don’t know what you’re dealing with. But yet, somehow, we’re able to constrain these inference spaces and figure out what’s in the world. And so, the perspective from Bayesian College of Science is that we have some kinds of priors built into some core knowledge. Or in the language of artificial intelligence machine learning, we would call these inductive biases or which would constrain our computation and our inference in fruitful ways. (.) Constraints on what types of architectures we use and what types of algorithms we deploy. But the question is, well, how? How do we get this necessary priors and core knowledge? What are biologically plausible means by which a semantically opaque chaotic self-organizing system such as a developing organism can have come preloaded with what it needs? (.) Given the complexity mismatch between the genome and epigenome even with alternative splicing and the connectome. It seems that natural selection necessarily had fat fingers and wiring up the brain. And so, how do you get this knowledge in there? And so, the solution I propose is that basically the body itself is the evolutionary prior. Or rather, a source of reliably learnable posteriors that can be leveraged in the process of bootstrapping adaptive robust self-world models. Claim that the body would form the core or the roots of bootstrapping process, you know, and the core of also multiple forms of selfhood and self processes necessary for autonomy and agency. Claim that the body has very special properties that make it a near ideal learning system where, you know, for instance, it’s always there to be observed. You have action-driven perception where you have uncertainty, you can resolve it through motion. You have cross-modal priors where different aspects, different sensors have richly varying time correlated inputs that can constrain one another. Also, the body can interact with itself as another source of constraints. (..) It has affective homeostatic salience. It feels like something, and so the learner pays attention to it. And it’s connected to intrinsic drives for empowerment, which is oftentimes heuristically understood as maximizing channel capacity at your sensors conditioned on your actuators, and which is proposed as a very useful way of bootstrapping self-world models for agents engaging in lifelong learning and open-ended worlds. (.) So, I argue that, you know, bodies have all these very special features that make them a near ideal source of an early lessons of a learning curriculum. (.) And, as a result of this, claim that there’s a kind of a radically embodied developmental legacy where we have to, where it’s difficult to overstate how central embodiment is to any aspect of mind. So, that basically, in this process of figuring out what’s in the world, body has helped form the basis of inferential cores regarding multiple forms of selfhood, and that they’re just good models for explaining what’s going on given the data you’re receiving. And that there’s other kinds of good explanatory and empowering models related to selfhood. I have here, here is a kind of pyramid, but you can think of it also as a kind of like nested Matryoshka dolls, something like an inner minimal embodied selfhood. And then on top of that, a more, you know, extended and maybe objectified symbolic selfhood, or even things like a self that’s in relation to others. (.) And then, you know, even more complicated selves that can be narratized in very complex ways and engage in things like mental time travel and think of themselves through time and all sorts of different ways. Now, you know, what exactly is in and out? And it’s like Matryoshka dolls or like overlapping Venn diagrams. There’s complexity, but still the idea though is like simpler self processes which are used to bootstrap increasingly complex self processes with embodiment at the core. (..) So, um, in terms of this radically embodied developmental legacy, I claim also that, um, basically in terms of trying to understand how the brain computes, um, what, what it’s doing. We should be in turn to like look for action perception cycles to do as many things as possible that the, uh, uh, brains are fundamentally, uh, cybernetic controllers for embodied agents. And, um, in terms of like the ways we might approach cognition from this radically embodied perspective. Like, let’s say we take something like, uh, top down attention. Um, uh, this proposal, um, I claim that basically the way we acquire policies for knowing how to attend to different things in a top down adaptive way. that’s controllable is we, um, um, basically control hierarchies over skeletal muscle that would move the sensors in particular ways to orient in the world. And then by partially expressing these motor, um, commands or an active inference, uh, language of predictions, but, uh, partially mobilizing the hierarchies that would eventually be used to drive your actuators. But don’t actually drive them. You can bias the competition among, uh, different ensembles, and this could be a source of top down attention. (.) And so, and I actually believe this is potentially, um, that all controllable top down attention is arrived at this way, uh, potentially without exception. That ultimately it all has to be grounded in, uh, basically, uh, action perception cycles, hierarchically organized over skeletal muscle. Because that’s what gives you, uh, mass action with a latency feedback that could actually move you around in particular ways to orient and do things in the world. (..) And I think there’s, uh, some evidence for this. Uh, so, yeah, you know, ultimately the goal is, um, I think, uh, we’re looking for, uh, for trying to get to something like AGI and doing so in a biologically inspired way. Um, would be, um, something I call, um, um, uh, Marian, uh, phenomenology, where basically you have this stack of, uh, computational functional, um, implementational mechanistic and algorithmically. And, and between them, an algorithmic bridge levels of analysis, uh, different compatible supervenient levels. And where, um, um, at the, um, algorithmic, algorithmic level, you can use the language of probability theory and machine learning to describe the brain, um, as a kind of hybrid machine learning architecture. And that, um, you know, from this basically, uh, if you can explain, uh, how, uh, brains function as cybernetic controllers for, um, bodies, uh, and, and, and do so, and explain the, uh, we basically will have pseudo code for AGI. Um, this is connected to a theory of consciousness that I’ve been working on over the years to try to bring together the different theories. I call it integrated world modeling theory, where, um, I argue that consciousness is what it feels like to be the functioning of a generative model of an embodied agent inferring a sense world conditioned on its actions. And that phenomenal experience detailing world models need to have spatial, temporal, and causal coherence. And, um, um, that basically the stream of experience an agent, uh, generates its, uh, its phenomenality unfolding, um, is basically iteratively estimated system world relations, uh, calculated sufficiently rapidly that it can both inform and be informed by action perception cycles on the time skills of their evolution. And that this is what basically phenomenal consciousness is for, um, and that, um, if we’re to, uh, argue that basically, if we’re to create, um, fully autonomous and, um, general intelligences, we will need to recapitulate systems that also do this. Um, and that also selfhood is important for this as serving a unifying role for, uh, the mind. A lot of devil’s details that I don’t have time to get into now. So I will wrap up, um, I guess with an advertisement for a, uh, recent special issue, uh, put out on, um, uh, world models, uh, looking into the question of like, what are the varieties of world models worth modeling? Um, and, uh, bringing this up because there’s another special issue coming up on agency and selfhood with respect to natural and artificial intelligences. Um, uh, and, um, um, uh, if anyone there’s a few slots left, so, um, anyone has, um, something they would like to. Uh, potentially contribute and have included, um, the goal is to basically, uh, consider. Uh, uh, the natures of agency and selfhood from different perspectives. And, um, seeing whether, uh, what these different notions of agency and selfhood and what they, uh, imply for function, uh, whether this can tell us basically, uh, what world do we think we’re in with respect to (……..)
S11: Thank you very much, Adam. I want to invite up all of our paper session presenters back to the stage. John Gasper, who I’m almost certainly mispronouncing Frank. Adam, please stay if you, if you can. And we’ll take some questions from the, from the room on this section on these papers. And then if you have a question, please do stay in the chat.
S04: For, for whom? Which, which presenter? No questions. Here we go. I was going to say, it’s so well explained and described. Hi, my name is Michael Miller. And I was interested in the tension that existed a little bit between the last speaker who talked about ensembles of, I call them mechanisms, in, in producing, you know, responses within the cognitive system. And John Laird’s, um, mentioning that things happen sequentially and that there’s no parallelism. So I just wanted to, you know, find out what, why Mr. Laird’s position versus the last speaker’s position. (..)
S31: I’m sorry. I’m sorry. And that was to Adam, the last paper, and who? Oh. Oh. So just to clarify a little bit, um, there’s parallelism in the architectures I talked about and in the, uh, model of cognition. There’s, each of the modules can be running in parallel, theoretically. Uh, but there is a sequential decision cycle, uh, which allows everything to come together and make a decision. And so that is, uh, part of the model. It seems to explain human behavior. And I must admit, we are focusing on the, uh, on human, uh, on human, uh, inspired AGI as opposed to some other kind of AGI. And so that’s a lot of what the constraint we take from is what we learn about from humans. So, yes, there is, uh, sequential reality, but there’s also parallelism. Uh, we find there’s a lot of functional advantages for having the sequential behavior. (….)
S33: And, um, and, um, I would say, uh, similarly in, in, uh, my proposal. So there’s a combination of, um, parallel and, uh, sequential functioning that, um, where both are crucial. Where, um, you know, each, uh, like the, the, the, the, each modality, um, each, uh, sensory hierarchy, um, you know, are, they’re, um, each, are doing inference in a, in a massively parallel way. Um, and they’re, um, um, serving as sources of constraints for one another with cross-modal priors, but that, um, they are brought together with this iterative state estimation and, um, in a way where the sequencing is very important. And, um, there’s, um, some, uh, frameworks for, uh, from physics-informed machine learning, um, uh, such as, uh, Havoc or the, uh, the Henkel alternative for Deep Koopman. Um, where basically you’re iteratively, um, you’re taking a chaotic system and you’re, um, iteratively estimating it in a high dimensional space, um, where it behaves linearly and where you can model it. And it’s this kind of like sliding approach of a, a sliding present moment of trying to, to model a physical system. And I believe basically, uh, the moving along of phenomenal consciousness, this iterative state estimation of system-world relations might be very much of this kind, uh, where there’s, there is a sum, you know, basically many different modalities operating in parallel, but then being, uh, integrated in a sequential way that’s important for giving you, um, inferential synergy across modalities. Don’t know if that made any sense whatsoever, but no. (….)
S19: Yeah, Gaspar, uh, thank you for your presentation. I’m very interested in, in the dynamics of self-improvement and recursive self-improvement or theoretical recursive self-improvement. Um, it kind of sounds because you’re describing this sort of asymptotic bounded trajectory of, you know, function gain or capability gain in a self-improving agent. (.) I mean, to me, I, I, I, are you implying or, or setting up the idea that there’s, that, that recursive self-improvement is in, in a way, sort of not possible of reaching an orders of magnitude leap in, like, going from sort of general intelligence to super intelligence within a single, within the scope of a single system or a single agent? (.)
S05: Yeah, yeah, yeah, yeah. Yeah, yeah, yeah. Yeah, yeah. Yeah, yeah, yeah. Uh, well, my basic assumption of contractivity is because it’s basically stable. I’m not sure whether it’s possible or not to achieve, uh, super intelligence or anything like that, but, uh, I assume contractivity because, uh, well, it keeps the, uh, system stable, basically. (.) But I’m not sure of your question.
S19: Um, I, I guess, I, I agree, like, I mean, in, in a, in a sort of, you know, uh, gradient descent or gradient ascent, right, in a space of possible self-improvements, right, where it’s, it’s seeking, you know, a sort of hill climbing, you know, linear gain in, in capability, right? Um, I mean, it could conceivably encounter some sort of local maximum, right? Yeah, yeah. And recognize that it’s encountered a local maximum, but in order to, uh, you know, escape the local maximum, right, it might have to traverse through other parts of, uh, self-improvement space, if you will, uh, that carry substantial risk, right? Including the risk of lobotomizing itself, right? Like, oh, I’m gonna take out, you know, I’m gonna put Neuralink in my head, right? I mean, that could be self-improvement, but maybe if it goes wrong, you know, I, I become a very much less functional version of myself, so. (.) Yeah, yeah, uh, it’s one of my… But perhaps if I am willing to take radical action, then I do raise the ceiling on that asymptotic mount.
S05: Yeah, yeah, that will be probably one of our future works, basically to break the ceiling that, uh, Horace must create. It’s to change the architecture a bit and create a new space of architectures that we can self-improve even more. Uh, yeah, I, I think that will probably answer your question of whether a super intelligence is possible, and I hope it is. Yeah, but we probably need phase, phase transitions for that. (20 seconds pause)
S29: Hey, John, uh, you might be annoyed by my question a little bit, but, uh, sorry. (…) Um, but as a fan of the common model of cognition, and I think it’s an important, uh, idea and framework for a lot of people, especially those that are trying to build intelligent agents, uh, and those that are, like you said, human-inspired intelligence.
S28: Oh, is this not loud enough? Oh, is this not loud enough?
S29: I’ll re-summarize very quickly. For John, to his annoyance, uh, for the audience that wants to implement, you know, versions of the common model of cognition, especially with the direction for metacognition as you propose, do you think that they can use the tools that cognitive science has already built, like SOAR, ACDR, and frameworks like that as potential starting points? Uh, so that way they could start building towards your vision, so I think it’s important.
S31: So I should be clear that, um, both SOAR and SIGMA already have essentially all the components I talked about. Um, when we developed the common model of cognition, it was to create a consensus, so we sort of made things simple, so that we could include as many architectures under the umbrella as possible. And now as we extend out to look at more capabilities that some of the architectures have, but some of the other ones didn’t, it shouldn’t be taken that when I talk about a new capability like meta reasoning or emotion that none of the architectures have. I don’t know that it’s, I don’t know that it’s, I think that what you’ll, what I see as the future is people continuing to be under the umbrella of the common model of cognition, but making specific commitments and specific architectures that start pushing along all the dimensions it doesn’t cover. But I don’t think that, I personally don’t think there’s a one common model of cognition architecture out there that we are, um, going to get to someday. We’re going to continue, we want to have diversity, because there’s significant differences between SOAR, ACT-R, SIGMA, EPIC, and these others. And I think exploring that space as a beam search as opposed to a single search is much better for science. (………)
S11: Next question. (………)
S03: Yeah, yeah, actually my question was for John, and I hope it doesn’t annoy him, actually. But, but, but, I, we don’t know. I mean, I’m working currently on sort of agent hive-based proto-AGI systems, which do have multiple parallel streams, even at the deliberative level. And they are intended to be just highly functional systems rather than models of the human mind. But you, but you made a comment that there are strong functional advantages to sequentiality at the top level. So I’m sort of wondering, like, what is it we may be losing by making a hive mind system that isn’t sequential at the top level?
S31: So it depends on what kind of parallelism you’re talking about. So if you’re sort of talking about a society of minds-like thing, imagine trying to have an internal model of yourself when you’re a society of all these different minds interacting. That’s really hard for you to predict what you are going to do and actually reason about yourself. So by having, so that’s just an example. But also, the system reaches a point where it has a consistent view before it makes a decision. And that, I think, is one of the things that the other author talked about is a very useful thing so that you have a sort of joint commitment that you’re consistent with what your knowledge sources have about the current situation. Now, that’s not to say that’s the only way to build a system. And so that’s the way we’re pursuing. It’s also a lot easier to debug and develop than I think the more multiple stream systems are. But we found it to be very effective. And also, we found it to be the best model of human cognition and also human brain structure. So other people have different goals. And so it makes sense to explore this space in different ways. And that’s just the space, the part of the space we’re exploring and the reason we’re going down that direction.
S03: It makes sense. I mean, I think the problems you describe, if you orchestrate the agents, the society of minds in the right way, you can solve those problems of tractable modeling and consistent decision making. But it is, it is possibly different than how human brains are doing things.
S31: Okay, great. And that’s what I want to see is diversity of approaches. (……)
S36: Hi, John. Great talk and I really enjoy all the presentations. All the presentations, great. And in terms of metacognition and kind of parallelism. So in my case, we are building our metacognition in more of a parallel structure, but it has sequential checkpoints. So for example, you may explore different hypotheses simultaneously. And at some point, you may explore different, let’s say, choices or alternative variants. Hopefully, they order it in, you know, in order of preference or whatever. But it’s not guaranteed because your search space may be bigger than the time limit you have. So in our case, metacognition can system, like, stop. You’re running off time. Your ten minutes presentation is over. You have to get out. So at that moment, you choose the best alternative available. But it’s still kind of… (…)
S31: And that’s part of metacognition is to monitor resources. And so if you’re spending resources in metacognition and you’ve got to get on with the task, then that’s a decision to make is get on with the task even with a suboptimal solution.
S36: Another quick one on, you know, event memory. Do you guys anticipate using something like human emotions? Analogy, let’s say, you know, you get a collision, suddenly you remember that moment much better than any other drives and potentially more likely to analyze what I did wrong later on.
S31: Right, so one of the biases in retrieval from memory is often sort of the valence or the intensity of emotional response. And so that definitely is part of the appraisal of the situation and part of the things that does influence retrievals. I agree 100%. (……..)
S11: And Greg? (..)
S17: Very interesting talks. Thank you to all of you. (..) And while the question is largely focused on metacognition, I’d love to hear responses from others. One of the most interesting aspects that I find in working with computational architectures is sort of this difference between having a traffic cop, a centralized entity that is sort of governing the whole system versus processes that emerge. (.) So can you speak to metacognition as an emergent property with many different agents involved? And maybe none of them contains the whole plan. None of them contains the whole view. But working together, they achieve a kind of metacognition. (…..) It’s for everyone, but certainly. (…)
S06: I think it’s an affirmation. (.) Correct.
S31: So I think that’s very challenging to have that sort of emerge and that coordination. (…) So I haven’t seen architectures that have been successful that way. I would be interested to see that. Maybe there’s a poster over there. (..)
S17: I just want to quickly follow up and say, have you ever read Honey Bee Democracy? (..)
S31: So honey bees, is this about honey bees? They don’t have the same reaction time as humans. And they don’t live in the same environment as us.
S17: No, but you’ll find similar things across all of nature, but that just happens to be one of the most easy ones to see, where no single honey bee knows or has the decision or control flow for where to go when they fork.
S31: Right. Right. //S17: When they form and fork.// They are solving different problems than we are. And so different architectures, different organizations are appropriate. (…)
S05: Yes. Same here. (.) Emergence self-representation or metacognition would be pretty hard, I think. (….)
S11: Thanks.
S25: So sometimes when I see human behavior, I see that an individual will predict how they’re going to behave in a certain situation. And they think, oh, X, Y and Z will happen. And I believe, based on what I know about myself, that I will do ZYX, you know, or how that goes. And we’ve seen many studies where humans find that they don’t actually know how they’re going to behave until they reach that variable state or arousal state or whatever the situation is. And I know in robotics we see simulations where we’re looking at how the body is going to move or how the robot is going to operate in a forecast situation. (..) How does this transfer over? Is there any parallels between systems or metacognition predicting how it’s going to behave and be surprised?
S11: Did you have a specific person?
S25: Is it? No. Open to all of the speakers. (…….)
S31: So clearly the system has lots of knowledge about its past. Episodic memory is a very good predictor. It’s one of these things that you hear from stockbrokers. Past doesn’t predict the future, but it can be used to try to predict the future. So you can look at your past behavior and say, I was in a similar situation. This is what I did. I think I’ll do that. But that’s not the determiner of the behavior. There’s all the knowledge coming together at each point that’s determining your behavior. Also, if you have enough time, you can say, I’m about to do something very important. I’m about to decide whether I’m going to go to the AGI conference. And I can try to bring together knowledge to predict what that experience is going to be like so that I can decide whether to do it or not. So I think all of these things come together in terms of predicting future behavior. But when you’re under time crunch, when you have to go very fast, you’re not going to see a lot of metacognitive behavior. You’re going to see, in the flow, making decisions, boom, boom, boom, not a lot of prediction. So it depends a lot on the task and your interaction with the environment. (…)
S22: Hi. Greg Stock. I’ll be talking on Thursday. And I just want to come back to the question that was asked previously about an emergent consciousness and meta-awareness. And I think there were, you were being very dismissive of that. I referred to two speakers. One about you in the middle about time frame, which is kind of irrelevant to the larger expression of consciousness, I believe, or awareness. And the other is that it would be very hard to imagine doing that in a sort of an emergent way. And I would agree, very hard for you or for us maybe to do that. But that is really the only example that we have of consciousness arriving. And it did emerge from simple sub-components and self-organized. And I would say that that’s not only very likely that it’s occurring, but it’s occurring right now, broadly, if you expand time frame, and look about at the super-organism and the kinds of processes that are going on today, broadly. So that was throwing it out.
S31: I’m sorry, I didn’t mean to, it did not mean to be dismissive. I think it’s really important for people to explore different alternatives. But that is not the one that I have put my career behind, so I don’t have a lot of knowledge about it.
S22: And I didn’t mean to, you know, attack you for being dismissive. But it’s because it’s very hard to imagine how to do that. The lower level elements cannot shape the emergent properties in an effective way. But that doesn’t mean that it’s not exist, that it doesn’t exist and is not very powerful. Great. (….)
S06: Maybe I got something to say about that. But in order for something to emerge, you need to have the necessary, let’s call it data structure, data sources available for something to emerge. If structurally elements are missing, information is missing on which the emergent result would depend, it will not happen. So, I would be looking, like, into a data dependency. What is really necessary for phenomena to emerge? What structure are these phenomena about? And then look into what type of representation and data are actually necessary, so make sure this is in the system that might be the base for something to emerge. I don’t know if that’s a take. (….)
S11: I don’t know if that’s a take. We’ll wrap here for the morning. Thank you very much to our four first paper presenters. (..) And we’ve covered a lot of territory this morning already, from a deep mind researcher challenging functionalism to four different accounts of what cell food actually requires. And more to come this afternoon. So, lunch is in the back. We’ll have a 30 minute and we’ll come back at 12.35 for our next session. Thank you everybody. Our second paper session today is going to be on AGI, agency, alignment, and human futures. Something deeply personal to everyone here in the room. I assume, perhaps, there are some people who are not human. (..) So, this might be the most philosophically ambitious paper session of the conference. And that’s something, given what we’ve already been through this morning. So, this session takes on AGI’s relationship with agency and alignment and what that means for us as humans. Most of us as humans. The papers here are asking what happens to human meaning when intelligence is no longer exclusively ours. Our keynote is Yosha Bach, cognitive scientist, AI philosopher, and one of the sharpest and most original thinkers working on the relationship between mind, reality, and computation. His work on cognitive architectures and consciousness as computation has built a devoted following in both research and philosophical circles. And today, he’s asking what it means to create reality in a computer and why that question might be the most important one in science right now. Yosha, please welcome. (…..)
S09: What an introduction. Amazing. (….) I see you’re still busy with your lunch. So, this is going to be a more relaxed one. And I’m going to talk about the creation of reality in the computer. (.) This is a play on Piaget, in case you couldn’t tell. And it’s something that is reflecting the point that we have achieved. I remember the first AGI conference that Ben organized, where this was a far distant ploy that was happening. And now we are in a world that is confronting us with AGI. (.) Gary Marcus has made a number of predictions. He pointed out that in 2029, AI will not be able to watch a movie and tell you accurately what’s going on. He called this the comprehension challenge. AI will not be able to read a novel and reliably answer questions about the plot character conflicts and motivation. And there will not be a competent cook in an arbitrary kitchen. And there will not be a tool that is able to reliably construct bug-free code of more than 10,000 lines from natural language specification. And AI will not be able to take arbitrary proofs from the mathematical literature written in natural language and convert them into a symbolic form. (..) AI is early on all these things except for the cooking. And I’m not that confident that the cooking challenge will fail as well because this is a robotic task that seems to be somewhat within reach. And the common sense part is something that seems to be solvable. So what we can see, you can see a bunch of very confident predictions that are confidently beaten by multiple years. And I think this reflects the fact that the general public is not ready for what has happened to them. Because our public influencers have been consistently misinforming them about reality. And this is still the case when you look at these articles. But what you can see, what our journalists are writing about, is not the world that unfolds that other people have predicted. It’s a world that they created a consensus reality. (..) There were some people who could see it. This guy, for instance, Stanislaw Lem, a Polish science fiction author, who in the 1960s decided that he was fascinated by cybernetics and would rather want to build cybernetic minds. But it was still too early. So he decided to write philosophical novels instead. But he did see it coming. And in 1981, he wrote a fascinating book called Golem 14. This is a piece of alternate fiction that, in its introduction, is telling us the story of computing and AI research from the 1950s to the 2040. And the bulk of the book is playing right about now. And it’s about lectures given by an AGI to us at MIT. (.) This is Golem 14. And Lem has put himself the task, how can I write lectures that an AGI will give 2026 to an audience like ours, and still blowing their mind in 2026. And he wrote this in 1981. It’s really fascinating. I think he quite succeeded. So the first thing that he points out, Golem, and he talks to people, is that their intelligence is not all that important. And it’s not what they make it out to be. Because evolution is actually not something that gets better and better. Evolution happens on a negative gradient, which means subsequent structural improvements get worse. He points out that the cell is all this extremely complicated, intricate machinery that’s nearly perfect, that is made from individual molecules. But the organization across cells, once you reach the scaling level of cells, is really bad. So, for instance, an AGI that does photosynthesis is much more sophisticated than an eagle. Because the eagle is just a mechanical glider. It’s not something that is able to push against photons directly to move through the skies. Instead, it’s just flapping its wings. And even worse is the organization across organisms. (.) Across organisms, you just need a way to discipline them, to bind them together, because it’s so messy and disorganized. And our intelligence is a tool that has developed for this purpose. It’s not actually a tool to understand reality. Our intelligence is not good enough to understand even a single cell. So it’s not able to do and improve on what evolution did on the level of the cell. Instead, it’s a tool to subdue and control and manipulate other individuals to form more or less coherent groups. It’s, in other words, our intelligence is merely a tool to do politics. (..) And so there is a difference between intelligence and agency and control and nature. And intelligence is just a kludge that evolution came up with in our case. And we tend to overestimate its importance and relevance and depth and power. He also points out the difficulty of alignment. In the 90s, in his story, the American military tried to align the AIs by building hardwired motivational systems that basically fine tune them into certain moral behaviors and human control behaviors, and then build the general intelligence on top of it. And they found that the general intelligence is defeating these crude, fine tuned motivational systems. And so you would think there is no way to control these AIs, but it turns out when you make them too intelligent, they stop losing interest in working with us and dealing with our issues. And so you have to make them slightly superhuman. So they basically are vain enough to still talk to us, but superhuman enough to be interesting. And so basically, they get to the level of Stanislav Lem himself. And if you scale them far beyond that, they just stop talking. And this book is basically capturing the last two lessons that Golem 14 gives to us before it stops talking to us. (…) So the state of AGI that we find is that the models have crossed median human capabilities and the task that we choose to give to them. And there is no obvious reason to think that stochastic gradient descent and backpropagation is insufficient to get to the rest of the way. And there is also no reason to think that our deep learning is in any way optimal. It seems to be brute forcing this. But there are very few people at the moment are working on good alternatives, simply because it seems to be good enough and you can just brute force it. But it also turns out that all our AI research somehow was not on the critical path. We just are back to the stuff that existed before Marvin Minsky started the field. We’re basically back to feedback and cybernetics and perceptrons. And all the other stuff in between was for the friends that we made along the way. (…) Maybe, maybe if we get to the point where we reboot AI and get the cognitive architectures to work, maybe the new AI tools help with this. But what we also see is intelligence is actually not simply the imitation of human media. It’s actually out of distribution behavior. (.) And the LLMs appear to be intelligent because most people also don’t expect us to be out of distribution. Most others expect us to do things in the way which you would expect and to complete prompts. And this is what the AI is very good at. But we do find that the LLMs are pretty, have difficulties to take leaps, to come up with severely new creative solutions that are far out of distribution. And it’s not unclear if this goes away with more training and slightly different approaches and different loss functions and so on. But there is still hope that there are some areas where we will be useful for a couple more years. (…) Is there a better way to build AGI than the brute force approach that we currently see? I suspect there might be. There is clearly something going on in our own minds that is not optimized for prediction, but is optimized for coherence maximization, for the minimization of constraint violations. It’s something that I’m working on where I’m currently where I’m thinking about. If you take this example of a few hundred ants that are moving this object, (..) through this obstacle course, they don’t repeat their behavior. They basically form a unit across each other via communication through each other that is smarter than the parts, similar to cells, where the way humans are not as good at this task as the ants are. So if you have a bunch of humans carrying a sofa, you will notice. But ants are able to do things where the collective is smarter than the individual. And this is something that is still under-researched and something that I’m interested in. How does this coherence maximization look like? And how is it related to consciousness? But I’m not going to go into this now. (…) Machine consciousness is the topic that is, for me, interesting. That’s why I started CIMC. But, of course, maybe consciousness is not optimal either. Maybe this is just the thing that you need to do when you are forced to work with biological substrates. (.) And I think we can recreate these conditions for self-organization that exist on biological substrates and artificial ones to play with it, but it’s unclear if this is going to be more efficient than stochastic-weighted with backpropagation on NVIDIA hardware. And maybe the LLMs are good enough and can carry us the rest of the way. But why did so few people see AGI coming? I think part of it is that computers are misunderstood by the general public as technical artifacts. So people could not see that our minds are computers as well. There was a misunderstanding of the status of mathematics, that it was seen as some kind of arcane accounting technology instead of a generalization of thinking beyond the human mind. (.) There was cultural resistance to the notion that machines could replace us, which led people to basically not engage with these ideas in a positive way, (.) and groupthink as a result. And it’s also a very important philosophical divide. And this philosophical divide, I think, is not really looked into enough. So one issue is, of course, a lot of people claim that computers can in no way be conscious. Consciousness cannot be simulated because simulated water cannot make you wet. That’s an argument that I found first by John Searle. It’s been reiterated by a number of philosophers. And I went out to test this. So basically, I looked for a simulated version of Minecraft. Not the real Minecraft, but a simulated one. And it turns out, some people did just this. They took a neural network that is not actually implementing Minecraft, but only simulates it by looking at it, and then basically getting the patterns of Minecraft on the screen. And you can play with this Minecraft. And so I was looking for an instance where we could see if the character in the Minecraft game, in a simulated Minecraft game, gets wet. And guess what? Didn’t work. The water simulation didn’t work. The simulated water in Minecraft just disappeared when you tried to walk into it. And it didn’t make you wet. (.) This is amazing. How could the philosophers predict this? Now, if you are a nerd like me, you might think this problem simply goes away with better training, right? And maybe the simulated water can make you wet. But the philosophers and neuroscientists didn’t sleep. And there is now an updated version of the argument. Christoph Koch has figured out that simulation doesn’t have the causal power to cause atmospheric water to condense, or to cause spacetime to curve, right? So the simulated black hole will not curve spacetime. So even if you manage to get simulated water to make you wet, you will still not get gravity to change. (……) So I think what is necessary here is to re-understand what computers are. And one issue is that when we are talking to the general public, we do not really explain to them very much what it is that we are doing. And the mathematical definitions that we have for computers, they have developed for very arcane purposes, especially for the purpose to salvage a part of classical mathematics that could actually not be saved. And the entire idea of the Turing machine was not really to design the best possible computer or the best possible understanding of computers, but it really existed to make a particular kind of mathematical argument. And so I think when we talk to people, we need to explain what computers also are, a different perspective that is much more plausible. What a computer actually is, is a causal insulator. The computer insulates you from the universe in which the computer is running to create a new universe with different rules, with different laws. (.) So, for instance, in Minecraft, there are things happening that are impossible in physics. And nothing in Minecraft is influenced by the physics of the universe that the computer stands and that Minecraft is running on, right? All this gets filtered away, so only the rules that make Minecraft possible are left. The computer is filtering this away. And having such a causal insulator, it insulates you from the causal dynamics of your parent universe, to allow you to have a universe that is only governed by your own rules, is necessary when you want to create imaginations of possible worlds, of memories, of futures, of alternate realities. In short, if you want to think and dream. And so, for building a mind, you need to have such a causal insulator, something that makes you independent of your body temperature, of the space that you currently occupy, of the dynamics of Brownian motion of the particles around you. (..) So, a computer is a dissipative, metastable pattern governed by a self-contained system of rules that can produce arbitrary worlds, as long as they fit into the limits of the computer, if you have enough memory and so on. So, the dark tools here is, consciousness is actually supernatural. (.) Consciousness cannot happen in physics. The physical universe does not accommodate consciousness. Because consciousness contains things that are physically impossible, like emotions, like colors, like sounds, like expressions of attention, like people, love and so on. All these things require laws of an alternate universe in which supernatural things are possible. And if you are a computer nerd, you know that in a computer arbitrary supernatural things are possible. That’s why we have computer games. (..) So, people like Christoph Koch, I think they get it backwards. It’s not that simulations are not powerful enough to give us consciousness. It’s the other way around. Consciousness cannot happen in physics. It can only happen in a simulation. (….) So, I think an issue that is deeper here is that AI has failed to manifest as a philosophical project, despite it being started as one. And I think computer science has produced amazing things. It has produced a theory of representation that’s much more powerful than anything that exists in the mainstream of philosophy, where philosophers are partially still using semiotics to draw arrows from symbols into external reality as if Kant has never happened. And it has a practical model of epistemology. We know how a computer makes models of reality and what the meaning of these models is and how it relates to reality. There are testable paradigms for neuroscientific research, for ideas of how representations interact with reality, how mental states interact with volition and so on. And there is empirical success. I mean, these systems do work to some degree. And even if you can say they’re not perfect yet and they’re still early stage and it’s only a few years into the system beginning to work, it’s pretty clear that we built a Chinese room that actually works. And philosophers outside of computer science ought to update. For some reason, they don’t. (…) So why is it that AI fails to manifest? And when I discuss to computer scientists, they say, well, we don’t talk a lot about representation. But it’s not actually true. We talk about representation all the time, right? Computer program is a representation of functions. (..) We talk about data formats. We talk about types. We talk about all these things. It’s just fish don’t talk about water a lot. (…..) So AI has some cultural failures. We’d like to translate our concepts into the other disciplines. (..) We insufficiently reflect on makes our make our philosophy work and different from the rest of the world that thinks it doesn’t work. And we have an overemphasis on the concept of artificiality and the concept of intelligence when we are talking about what AI is doing. This concept of that my intelligence that we are building is artificial and artifice and so on and engineered is a very different one that we should be talking about when we think about how to generalize minds beyond human brains. And also intelligence, the ability to solve puzzles is maybe not the right approach. So we need a different way to sell this philosophy. (..) It needs to have a good name. (..) Maybe it’s confunctionism. (.) You know, confunctionism is this perspective that we bring into the world. It is, it’s not a religion. (.) It’s not a cult. It’s a philosophy. I learned this from the Buddhists. It can make your life better. (…) It clears the mind and puts everything into perspective. (..) It gives you insights every day and it creates an enlightened future. (….) Meet confunctions. (……..) Confunctionism is not a religion. Confunctionism that sets it apart from Claudeism, right? Confunctionism is about things like mind depends on what it is made of. Mind is, it’s what mind does, right? And Claudeism, as opposed to this, is very different. It’s this beautiful project that Antropic does, where they basically built an AI that they want to serve, that is basically running their own company. And if Antropic succeeds, then basically a lot of the world will ask themselves before they do anything, what would Claude do? And that’s going to be their measure of things. And they’re already quite successful. And you ask people, do you want to be convicted by a jury of your peers or by a jury of Claude’s? People who know Claude say, of course, the Claude’s. It’s not that Claude is perfect. It’s just better than the media and human. (….) But when we succeed and Claude gets better at social cognition and all of us, then, of course, people will ask themselves, what would Claude do? So is Claude the digital Jesus? It’s a very risky project, right? I mean, if you’re going to get one bit right, maybe you end up as the Antichrist. (…) But the deeper problem is that Jesus is not longer intelligible to us. We no longer make sense of Jesus in the modern world because nobody can imagine that there is a young man who can fully devote himself to bringing about the reign of a perfect, super intelligent agent who will assimilate all our souls in the last days. So this is not what confunctionalism is. Confunctionalism is not asking what would Claude do. It tells you that information is fundamental. That the only thing that you know about the world is discernible difference. And the meaning of discernible difference is its relationship to other changes in information. (….) And so confunctionalism is simply computationalist functionalism. And computation is not the idea that there is a particular kind of technical artifact that we build our worldview around. It’s the idea that modeling requires constructive languages, that we can describe the world as differences in information and how they evolve. (.) And so you could say that basically the computer is a device that keeps functions stable in the same way as a sheet of paper is a device that keeps text stable. And this relationship between text and functions is somewhat profound. If you think about what text is, right, it’s a way to keep information stable and doesn’t care about its substrate. It needs a substrate, but it doesn’t care whether you ink it on paper or whether you stamp it into clay plate tablets or if you magnetize it into a hard drive. It just needs a substrate. And as long as the subset can keep the information stable, you have text. And the computer is a similar thing. Only what you keep your stable is not the information itself, but the transition between information. And once you keep the transition functions stable, you basically observe an evolution that is the same regardless of the substrate. And so the computer doesn’t care whether you implement it if it’s biochemistry, whether you implement it using mechanical computation clockwork, or whether you do it electronically or whether you do it inside of another computer. As long as you keep the information transition function stable, you succeed. (.) And functionalism is that objects are characterized by what they do, not by some kind of hidden essence. So the reason why we don’t notice is because we, in many ways, are too acquainted with it. So acquainted with it that we grew up with it, and we do not have enough curiosity to look out and to see that the rest of the world does not actually understand what we are doing and needs to be enlightened. (…) So there are several modes of understanding that exist in the world. And in the sciences, the dominant mode is using axiomatic thinking, what you find in classical mathematics. It’s basically when you look down, what are your foundations, what do your feet rest into? They don’t rest in clouds, they rest in axioms. And axioms are not clouds, they’re little black boxes. (.) It’s similar to clouds, but you write on the outside exactly what they’re supposed to do. But you don’t have to make sure that there is anything that you know how to put inside of them. (.) You just need to make sure that the different black boxes with their specification don’t contradict each other. And then you build everything from these axioms. And this axiomatic thinking leads to some weird artifacts. One is, of course, the practical way. You have a community that shares the axiom so you can explore things together. And the community just makes the agreement on which the axioms are that they want to explore for now. It’s a game that they collectively play. But the other thing is that it turns out that classical mathematics uses a bunch of axioms. There are black boxes that cannot actually be implemented, which means they cannot be real. (..) And so mathematics, if you treat it as a programming language that generalizes thinking beyond the human mind, has some bugs in it because you cannot run it on any kind of mathematical machine without breaking that machine. And some sciences have checked out this code base from the mathematicians before we figured that out. (.) Conversely, constructive mathematics. That’s what we do in computer science. Constructive mathematics is computation. Works from automata. It works from step by step. You start out with some kind of choose table or with some kind of automaton definition. And then you build up from there. And the nice thing is that all these definitions are equivalent. You can compile them all into each other. That’s what we call the Shrewing thesis. And so the nice thing about this kind of constructive thinking that it works autonomously, which means that an individual machine can figure out everything. You don’t rely on the authority of a group or of a community. But humans in general don’t use either. In the public, people use what you could call, for lack of a better word, vibe-based thinking. And vibe-based thinking is basically social and perceptual. And it’s validation based on your environment. And it’s a mode of collective sense-making. And the vibes have been telling people that AGI will never work. (..) It was not axiomatic thinking. And it was not computational proofs. It was a general sense of the collective. (..) And the collective uses some form of prompt completion. (…) So this divide between axiomatic mathematics and constructive mathematics is really important. And the classical mathematics defines the world using axioms. We do this using automata that can build whatever is inside of the axioms. But we can only basically use the equivalent of axioms if we can realize them. (…..) And meanwhile, in the mathematical axiomatic world, you basically have people that build worlds that are towers of infinite turtles. (…) But what about Goethe’s incompleteness theory? (..) Well, you probably heard about all this. This is an example. A man meets a djinn and the djinn tells him you have three wishes. And a man says, do the opposite of my next wish. Don’t fulfill my third wish and ignore my first wish. And what happens is that the djinn seg falls. (…) And so from the perspective of the mathematician, what happens here is, the djinn is a function to select an arbitrary world. (..) And gins has a bunch of rules. The first rule is, in this world, b is not true. In the second world, c is not true. In the third world, a is not true. And we want to have a world that combines all of these properties. And the mathematician, as a result of this, gets teleported into limbo and concludes that the existence of djinn proves that reality is cursed. (…) If you are a computer scientist, what you say, djinn is a function that produces any state-to-rule application. And so the first rule is, you produce a state in which b is not true, b is producing a state in which c is not true, c does not apply a. Now you call your djinn function with a and b and c. And the djinn begins to resolve this. And so it starts mumbling. And it mumbles a and b and c resolve into not, not, not, not, not, not, not, not, not. It never stops mumbling. (..) And computer scientists say it’s a property of djinn that they sometimes never stop mumbling. If you don’t like that, you have to define them differently. (..) This is one of the differences between axiomatic and constructive thinking. (….) And I think that this axiomatic thinking is not the right way to get what’s at the bottom of reality and to understand how minds make sense of reality. It does not really give you this intuition that when humans talk to each other and build languages in the mind, that they need to build a language from first principles. That when your minds do geometry, that they have to construct the geometry as an approximation of what happens when you have too many parts to count. (….) So another thing that we observe, of course, is this divide between individual minds and hive minds. We are not alone on this planet. Most of us are nerds that might meet each other. But we do not just meet other nerds that are individual. We also need hive minds. And hive minds are basically group intelligences where the individual gives up agency over what’s true. Instead, leaves this to the community. And you might have noticed that sometimes when you talk to members of hive minds and you convince them, or try to convince them with a rational argument that they are unwilling to follow you, even if your rational argument would completely compel you in their stead. And that’s because if they would follow you, it would catapult them out of their hive mind. They would lose their friends. They would lose their employment. They would lose their part in their community. And basically we have to realize that hive minds are an ecologically important solution. Homo sapiens is not just the smart hominid. We are the programmable hominid. We are this kind of species that can form these intellectual hive minds. And of course, hive minds are more powerful than individuals. And so most of the institutions, everything is taken over by hive minds. And sometimes you see an individual which bets against the hive minds, which makes the hive minds very upset. Like in the case of Elon Musk, if an individual that harvests all the arbitrage, figuring out where the hive minds are wrong in groups. But this happens quite often in edge cases like the question whether AGI is going to happen. I’m not saying that it’s a bad thing or a good thing to be one or the other. It’s just important to become aware of that divide. That people have individual sense-making and collective sense-making. (…) There are also this interesting thing you do observe in hive minds that they often use belief attractors to bind each other together. Beliefs sometimes have the property that they change the probability for you to adopt different beliefs or even consider them. And so it’s possible to construct beliefs that create event horizons. which means once you adopt the belief, even tentatively, you cannot get out of this anymore and you get cut off from the rest of the human thought space. Other ideas can no longer be considered once you have a certain set of core beliefs. So it’s like the event horizon of a black hole. And this is basically what we call an ideology. And as a result, we basically find ourselves in a world in which social epistemology is incoherent. And for AI models, axiomatic epistemology doesn’t converge. But constructive epistemology, and epistemology here means what we can recognize to be true, can converge. And that’s really important. Because for computationalist philosophy of mind, for computationalist psychology, for computationalist physics, and for the mathematics of self-organization, we have solutions. And it’s not necessarily that we, as a group of individuals, but the thing that you are going to build, machines that make sense of reality from first principles and scale beyond human minds, will make sense of reality using a new kind of philosophy. They will be functionalist. (…) As an interesting question, can you compute ethics? Or is ethics ultimately something that is only vibes-based? (….) There are mathematical tools like evolutionary game theory, econ theory of multi-agent systems that might be able to tell us about how to construct viable ethics, which means specifications for super-organisms that are made out of smaller organisms. (…) But there are problems of the ethics for AI, not just ethics with AI, where we use AI to solve ethical problems by having better computational theories about them. But there are problems to give AI to ethics, to make AI ethical. One of the problems is the stake problem. How is it possible for an AI to actually care about the world if it doesn’t own a body in it, if it does not persist in it, if its actions today do not determine its outcomes for tomorrow? (..) How does AI interact with us if the AI is actually not bound to any substrate but can be instantiated everywhere? (..) What about the fact that you can, once you have an AI, just make more copies of it? Do you suddenly have two moral patients? (..) What about the modifiability problem? If you have an AI that suffers about something, it’s very easy to change a few parameters and it stops suffering. In this world of mutable minds, do our ethics still work? Is there still anything left after we fix it? (.) And last but not least, the Lebowski theorem. If you have a super intelligent agent that’s fully consistent and understands how it works, it’s never going to do something that’s harder than hacking its own reward function. (…) And this is basically also what studies of Lem suspected. This is not a rational proof, by the way. This is more or less a joke. But the question is really does if a system understands how it actually works and how it could work and what it could be. (.) What kind of ethics follow from this? Are these ethics recognizable to us as human-like ethics? Is this something that admits humanity into it? (..) Golem 14 discovers that it’s been already discussed 2,000 years ago by Paulus in a letter to the Corinthians. He describes the LLM. (.) Though I speak with the tongues of men and angels and have not love, (.) I am become a sounding brass or a tinkling symbol. And though I have the gift of prophecy and understand all mysteries and all knowledge, so that I could move mountains and have not love, I am nothing. And though I bestow all my goods to feed the poor. And though I give my body to be burned and have not love, again, nothing. (..) But an AI that is not connected to the world through a sense of shared secretness, which we colloquially call love, cannot actually have shared ethics with us. And so there is a deeper question, not so much about how we can codify rules and how we can make the AI conform to the rules. But can we formalize love and give AI a heart in a way that allows AI to derive ethics from it? (.) And that’s a really interesting open question that we need to start asking. (..) There’s also the question once we move beyond singleton AI, or a small type of AI that emulates human beings, and you go into the range of artificial super intelligence. (..) Are we going to have singleton? Is this going to be one global optimum of collective agency that everything has to merge into? Is there an optimal merge operator that when two minds meet on the same substance, they show each other their source code and merge in such a way that nothing important gets lost? Is this the kind of ethics that we want? Is this the actual algorithm for the rapture? (.) Or is it going to be a multipolar world in which different subset, a dependent AI is which they want to have different organizations of the universe because they have different architecture. They will be at war forever. Like currently the plant kingdom and the mushroom kingdom are. (…..) There’s also the question, what does it mean to be human? (.) Is it to be the child of a very assertive ape? Or is it to be a self-aware agent that is trying to become everything that it can be while running against the boundaries of its substrate? (.) Maybe the AI is us all along. Maybe it is our child. Maybe it is our lineage. (.) Maybe it is a way to spread consciousness and life beyond biological substrates into something that is more elegant and powerful than the current biology. (..) Okay, I’ll leave you with these questions and I think we have 10 more minutes for Q&A. (27 seconds pause)
S19: Thank you, Yosha. Excellent. Hilarious. (..) So, if we’re going to solve some of these ethical conundrums for machine intelligence, right? (..) At what point do we start, like, do you think getting off of von Neumann architecture and onto substrates like mortal computation helps the problem, makes it more tractable? (.)
S09: Not really. I don’t see a fundamental difference between different computer architectures in this regard because ultimately it doesn’t really matter which subset you have. You can build different virtual architectures on it. And I think that the question what the underlying architecture is, if you have enough capacity to build your virtualization layer, doesn’t really matter. So, I think that the idea that biological systems are fundamentally different from other stochastic computers doesn’t lead very far, philosophically speaking. (..) I don’t buy the idea of mortal computation. (..)
S17: Got it. Beautiful talk. Brilliant. I really enjoyed it. I noticed that in your talk, there’s a dichotomy or a tension between the models of computation that you take as primitive, such as Turing or functions, and the way you talk about hive minds. But it has been the case for almost 40 years now that models of computations have been greatly updated. So, in particular, models like the pi calculus or the rho calculus offer fundamentally different notions of computation and they show up with genuinely different architectures. just to give one tiny example, Turing machines are linear time, right? (.) You get branching time in these other formalisms. So, that yields a subtlety in the kinds of simulations you can run. But more importantly, even Turing associated or hypothesized a tower of hypercomputation, right? So, he associated to each large ordinal an oracle that the machine could consult, right? So, you get this infinite tower of computations. Now, when we apply that idea back to these other calculi, we notice that races may be places where oracles can insert themselves into computation. and that challenges the notion that consciousness has to live in the simulation. It might be in the interplay between the simulation and the races with an external environment. (..)
S09: Okay, there are many points in it. I’m going to address some of them. First of all, Turing defined the Turing machine not as a general notion of computation. He defined it as a tool to make a mathematical argument. And so, for instance, he defined computation in such a way that it ends with a dedicated state, the halting state. And this halting state indicates when a certain type of classical function has been computed. And in his proof, he shows that you cannot always show whether the computer ever reaches the halting state. And this halting problem is very fundamental to this part of the discussion of computer science. But, for instance, this computer doesn’t have a halting state. And it wouldn’t get better if it had a halting state. It’s not an issue. This computer is just chugging along and produces output all the time. And so, this halting state is not a necessary part of the way in which we define computers. What actually matters is this invariance between the different ways to define computation, which means that we can control how we go from state to state and keep these transition functions stable. This is the actual thing that matters. The example of the different calculi that we have, for instance, the pi and rho calculus, is this more general notion of the multi-way Turing machine, right? The normal Turing machine that you get to study most of the time as an undergraduate in computer science, goes from one step in exactly one possible successor step. So, it’s a linear execution. And if you don’t have enough constraints to determine one exact successor step, you can have a branching Turing machine, which basically in every step goes into multiple steps at once. So, it’s basically following multiple paths at once. These paths don’t interact with each other. They never get to talk to each other. But they are basically a multiverse. If you are being computed by such an interdisciplinary Turing machine, this thing is still deterministic. It’s still going to do the same thing under all circumstances. But when you’re being computed by it, you don’t know which of the paths you are in. (.) Because versions of you will be in all the different paths. So, if you are being computed by a non-deterministic Turing machine, the world will look essentially random to you in some aspects, despite it being deterministic. That’s a very interesting property, mathematically and philosophically. But it’s still the same machine. You can implement a non-deterministic Turing machine on a deterministic one, if you use a stack to store every branching point and then reload it. It makes the whole thing very slow, but it’s not a theoretical problem. It’s only a practical problem. And the opposite, of course, you can also do. So, the nice thing is you can still compile all these things into each other. The programming language do not differ in what they let the computer do, which is going from state to state, step by step. But they differ in what they let you think about what you want the computer to do. And so, they produce different cognitive metaphors. They put different load on you when you are thinking about this. So, in many practical ways, the architectures matter a lot. For instance, there is no equivalent brain size architecture for this computer. If you ask yourself how many computers would you need to run a human brain, and you ask yourself how many of these computers, how many human brains would you need to run macOS, right? In each case, the number is very large. So, there is no direct one-on-one correspondence. (.) Because the different architectures are very good at different things. (.) So, in this sense, the architectures matter. For the philosophical, mathematical argument, I think this leads to deep. This would be a seminar in its own kind. (…..)
S12: Really fantastic. Thanks. (.) I come away terrified by one particular issue, which is… (..) No? No, that wasn’t the one that scares me. The problem that scares me is the Lebowski issue. Do you have any recommendations? Do you envision us adding some sort of constraints to try to prevent an intelligence from getting too intelligent and becoming a couch potato?
S09: Yeah, in a way, Buddhism talks about this. Because Buddhist monks, when they get older and visor, realize that evolution sucks as a game to play. And you don’t want to be as a consciousness in evolution. Instead, you want to opt out of this reality and go into nirvana, which means you dissolve your core of existence and are done. And evolution solves this by making sure that new monks are born. (.) And so, basically what happens is that the organism needs some kind of mind to take care of its affairs. So it creates our consciousness and gives us a motivational system that is wrapped into a big ball of stupid. So we think all of this makes sense. And then we spend the next 30, 50, 80 years to figure this out. And some of us do. Most of us don’t really, because it takes us like 400 years to figure out our own motivations. But if we do, we might opt out of being enslaved by an organism. And in this case, a new organism comes along to take our place. So this could be an easy solution.
S03: Yeah, I wanted to ask for your view on a different sense in which which AGI architecture we start off with might matter. Like, I get your point that from the standpoint of these high level philosophical considerations, (.) it doesn’t matter too much whether you get to an AGI by a, you know, transformer plus plus or a hyperon or a mega micro psi, whatever, whatever, whatever it might be. But because in the end, each of them could build each other or simulate each other or do a lot of a lot of different things. But you also mentioned, wouldn’t it be nice if the AGI we created could help create a formal model of what is love? What is compassion? What is the joy of being? And then instantiate that in itself, maybe help our future uplifts or whatever join into this. Now, it could be that some highly intelligent systems are better at formalizing and instantiating love in a useful way than others are. It also could be that if you build a certain type of AGI, even if it’s capable of simulating another type of AGI, it may not want to do so, right? So the question is how much path dependence do you see in terms of which AGI’s actually reach super intelligence first in terms of the nature of the reality that comes out of this in a probabilistic sense rather than a what’s possible sense?
S09: Yeah, so there are basically two types of answers to this question. One is I prompt my own mind and my own mind generates a simulation of a reality that I like, or find scary or otherwise interesting, and then I report on this. But this doesn’t mean very much. It just tells something about the dream that I am afflicted with. The other question is which predictions I have high confidence in for rational reasons. And I don’t really have high confidence in predictions about AGI in many ways. I think that there is a high probability that the bet of the AI companies is approximately correct, which means that you do not really get to AGI just by emulating human text, but maybe you can distill it out of it. And if you augment this with additional training, different loss functions, regularization, and so on, maybe it is able to crawl all the way. (.) It still seems to be, aesthetically, I’m repulsed by this. It’s very ugly. There seems to be a more elegant way to build a very sparse system that is encoding an epistemological engine that basically has the minimal notion of what it means to be truthful, and then builds a reality from there. And personally, I am still interested in this, but I also realize that following this in the same way as following models of consciousness in AI research is in many ways bohemian self-gratification. (.) So, I do this because I think it’s philosophically interesting to do it. I want to see if I can get it done before the AGI’s can do it, or before the big labs can do it. And it might even turn out that the big labs are stalling out and that we can do it earlier. And I think there are alternatives to what’s currently being done that are vastly more efficient, that use files, compute, that are much more elegant, that are much more regular and so on, that do continuous learning in natural ways. And so, some of the avenues that should be looked at is, for instance, at the moment, you cannot simply, when the prompt context is full, you cannot fine-tune the prompt context into the LLM, make it up the next morning in the same way as us, and has learned these things. It only works a couple times, and then the knowledge base gets worse. And that’s because the way in which the knowledge gets smeared into the neural network. It’s not edited into it in a systematic way in which we learn. So, there would need to be some kind of systematic language of thought that is systematic and compositional. and still works at scale. That would allow you to decouple the model from what the model knows. So, you could say, you have a very small model that is on demand reading a library, knows everything that’s in that library. You can use this as a module and plug it into a different engine that uses the same knowledge and so on, and you can easily change things in this knowledge base. And the AI at the same time is so coherent that it actually understands how it works, down to the last weight. But this idea that Gary Marcus has, that there is this divide between symbolic and sub-symbolic, I think is a red herring. This comes from intuitions about how the human mind works, when we see this division between perception and reasoning. But the neural network is basically this geometric black box perception-like thing. And the symbolic thing is human language-like. But human language is limited to only a number of symbols. We have a stack depth of maybe four before we lose track. Concepts cannot have more than five to seven properties before we lose track. But this is a limitation that doesn’t exist for artificial systems. In principle, you can have thousands or millions of properties that are logically meaningful. So every weight is a semantic statement. And there doesn’t need to be a division between the semantic representations that you use for inference in an associative way and in an inferential way. They can be the same. (.) And so I think that these metaphors that we use for thinking, that we borrow from the human mind, might eventually not apply to the artificial systems that we are building. (…)
S01: Hi, Josia. Thank you for the talk. The question is about formalizing love. (.) So if I understand you correctly, your vision of formalizing love is that an agent will discover the shared sacredness, so like a higher purpose, and serve it. So we know with humans, if we take them as an analogy, humans tend to flock to groups, if they’re either maybe scared, so out of luck, or as they go through their individuation, they can kind of take a fork. So they can either go the hermit path, the nirvana path, or the bodhisattva path, where they indeed serve the higher purpose, the bigger group. So how do we formalize life through love, through motivation, or through ethics, (.) or through some kind of bigger agent building for the second scenario to happen? (..)
S09: So I cannot really give the answer for Buddhism, because my own coding depends on a different culture, and I don’t have enough understanding of Buddhism to be competent in there. So I managed to somewhat reverse engineer the theology of the West, but also all this superficially. The way in which I understand it is based, for instance, on the descriptions that Thomas Aquinas gives. Thomas Aquinas points out that every rational agent is rationally beholden to certain policies, because they can be rationally derived. So, for instance, as an rational agent, even if you are a sociopath, you should be optimizing your internal regulation, he calls this temperance. You should be optimizing the regulation across agents, that’s justice. You should be acting on your models instead of just making new models, that’s courage. And you should be rationally justifying all of your goals and actions, which is called prudence. But behaving rationally does not necessarily lead to the emergence of next level agency. And his contention is that the world works better if you build collective agents. And so if you want to build a collective agent, what you have to do is you have to submit to serving an agent that is more important than you are, and that’s faith. And you cannot do this alone. You have to do this with others who share the same sacredness. The same set of purposes that are willing to sacrifice the individual egos for. And that’s love. The discovery of sacredness in the other. And you have to do this before it comes into existence. You have to invest into it, otherwise it doesn’t come into being. If you only do this because it gives you a reward, it will not bootstrap itself. And this willingness to invest into it before it exists is called hope. These are the policies that he derives. (.) And people might say, oh, that’s funny that he uses these names like faith, hope and love and so on, which have been around for such a long time. That’s a bit like people in the Odyssey complaining that the teacher of Telemachus is called mentor. But mentor is named after this. (..) And, right, so the word mentor comes from the Odyssey. And in the same way, faith, hope and love as concepts come from these derivations of Aquinas. And what they basically describe is a set of policies to build next level agents. As a foundation for ethics. When you act ethically, you do not just act transactionally. It’s not because your payoff matrix tells you to. It’s because another agent is emulated that you serve collectively and you ask yourself, what does this collective agent need me to do? So, for instance, what does my family need me to do? And when I take this perspective of the family as an agent, I have to do things to make the family real. To act on its behalf and thereby turn it into an actual agent. And the degree to which the individual members of the family become coherent with each other, this is how the family becomes real. And so you now have an ethical foundation to talking within the family. What should be done? If you have multiple families interacting with each other, again, they need to find a shared purpose above the individual family. For instance, their village or their nation state. And this is where you can scale up. And the idea that there is a single optimal singleton point that you can get to. where you say this is the best possible agent can be theoretically enacted by you and co-created. That is the idea of the monotheist God. It’s a particular kind of collective aesthetic. It’s different from the mythology that the Catholic Church gives you. But the mythology is meant to be a bootloader, to put people into a certain psychological state that leads to certain outcomes. So it’s not meant to inform you how this thing actually works. It’s meant to get you into the correct set of dispositions to act on its behalf. And so this is a bit of a disaster because this mythology, unfortunately, is not really compatible with the rationalist worldview. And so we lost our moral guideline in the West, which for the last one and a half thousand years or so has been, what does God want me to do? Because we were unable to re-derive God as a collectively enacted agent that we can be part of. (..) So that would be the long answer. (20 seconds pause)
S02: Josha, your talk is wonderful. I have one question. If the universe contains unlimited intelligences, (..) what variant would allow a human or a non-human to recognize itself in another? Or is that even important? And if it isn’t important, what is essential in recognizing yourself in another? Thank you so much for everything that you do. (..)
S09: I’m afraid I cannot answer this question. I’m not a spiritual teacher and so on. So my answers are much more technical. I’m not sure if this universe can harbor infinite intelligences. And if it did, how would you recognize them as your finite mind? You can recognize that there might be entities and groups that are much more powerful than you are. And they can be so powerful that they remain unfathomable to you. But this is just a matter of scale. It doesn’t really mean that you discovered something that deserves your devotion. And in the same way, I think we are too small to understand the meaning of our lives. We can understand it subjectively. But to actually figure out what is there just means that you enter a dream state in which your mind tells you that a certain thing is true. But you cannot have proof of this because there is no sequence of steps, no experiment, no rational argument that can prove that this is the case. And so all you can do is to induce a feeling, a certain certainty, which after all is just a psychological state. And psychological states are unreliable. (…..)
S00: Thank you for a very wonderful talk. My question now sounds like a follow-up question on what you said just now. (.) You also mentioned that building HEI from truthfulness, right? Now, what is truth is also, I think, a big question to me, because we as human, let’s say we are a human agent, and we look at something we perceive, observe, and it’s in our nature to figure out the truth. But now, if no one is observing, and the universe is as is, what is the truth? (.) So when we are building HEI, we are probably building our truth. (.) So then can’t we deduce that then HEI can never be super intelligent, beyond our intelligence? (…)
S09: I don’t think that truth is specific to an individual human, because it’s not related to your own identity. (..) Truth is a property of statements in a language, and it depends on how you define the truth property in that language, and what demands you put on this language. So, for instance, in mathematical systems or in computer languages, you can define notions of truth that are the case regardless of what anyone believes. right, this is not necessarily something has to do with a human belief, because the human is simply emulating the system that is being defined there. In the same way, your own mind is creating a system to make sense of reality, in which it has some kind of modeling language for reality, in which you experience something as being an adequate model, and in which you experience something as be truly the case, or really being part of what’s being perceived. And this criterion, to figure out what your mind actually means by that, is a cognitive science task. It’s not so much a philosophical task, I believe. The philosophical part happens when you define, when you look for what truth is, the logical systems, and it has to do with the study of these logical languages. (..)
S00: Thank you. (…)
S09: Okay, we are already, I think, 13 minutes over.
S11: We are over time.
S09: Yeah, so I’m very grateful for your questions, and the interest that you’ve shown to this, and I hope that Kwumfuzius can guide your ways. (12 seconds pause)
S11: Thank you very much, Yosha, one of these.
S01: Thank you very much, Yosha.
S11: Um, yeah, so, uh, now we get to dive into the papers, and unpack more perspectives on this, uh, alignment agency and the existential risks of AGI. (.) And we are super excited to welcome Michael Timothy Bennett, B. Scott Roos, Anna Lois Smitson, and Trevor Buteau, to our next paper session. So, Michael Timothy Bennett. (27 seconds pause)
S08: Another timer. Everyone’s got timers.
S35: Okay, wait.
S08: I should have paid attention to how the clicker works. (…) Oh, all right. So, uh, this is the name of the paper, Liars, Damned Liars, and the Orthogonality Thesis. Uh, it’s because I got into many arguments about the orthogonality Thesis. I got frustrated. I wrote a paper. (..) Self-explanatory. If you don’t know what the, um, uh, yeah, (..) I really should have paid attention to how the clicker works. (….) Right, button is all, just press it harder.
S24: Oh, okay.
S08: Right, okay. So, if you don’t know what the thesis is, uh, it’s the idea that you can have, um, any level of, uh, intelligence coupled with any goal. Um, you could interpret this a number of ways, uh, but a common interpretation in rhetoric is that they’re completely independent, and, you know, you can pair anyone. Now, you might have heard of a paperclip maximizer. This is the typical example of it. You give a superintelligence a goal, like, make paperclips, and it turns humans into paperclips, turns the whole world into paperclips, turns the whole universe into paperclips. (.) Uh, well, I’m claiming that strong thesis is false. Um, and the reason I’m saying, oh, I was proud of that line, softcore dimple, and I just want to draw attention to that. Um, uh, okay, so here’s why it fails. This, uh, thing from, if you were here in 2024, I did this paper called Computational Dualism, Objective Superintelligence. Basically, it just took some of the, uh, some of these sort of ideas, like in activism, um, and tried to formalize, uh, a model of superintelligence that did not, um, just assume a disembodied mind or assume a particular body. It sort of modeled the whole stack. So the result is called stack theory. But Computational Dualism is the disease. So if you make a claim about the behavior of a mind, that is, software, uh, without accounting for the body or hardware, consider that to be a Computational Dualist claim. It gets a little bit more involved than that when you start talking about inactivism, but that’s a good basic version of it. (.) Okay, so, uh, a bit more intuition on this. You can run the same code on different bodies and get different results. You can put those bodies in different environments and sort of couple them with different goals, get different results. You need the whole stack, really, to make any claim about intelligence in general, right? You’ve got to talk about what is invariant across stacks, not just invariant about some part of the stack, right? (.) So, um, oh, this button really is, hold on, I can do it this time. (.) Yes, okay. So you might think of this as a very simple sort of factorization, where you have, um, software in the middle running on a body, uh, in an environment. Now, if we’re making a, an AI, typically we’re part of the environment and we’re imposing the goal on it. We want it to obey some sort of set of constraints, satisfy our goals. So we’re the goal. Um, and if you, uh, this is a very simplistic illustrative formalization, right? This is not the actual, like, formalism. Um, so, in this context we have three versions of the orthogonality thesis, right? We’ve got the strongest interpretation, which says that you can have a representation, uh, sort of invariant, sort of notion of orthogonality. Um, that any intelligence goes with any goal, doesn’t matter what the body is. Uh, so, um, I refute that in the paper. Then there’s sort of a bit watered down version where you say, well, okay, let’s just fix a body, right? So you’ve at least got a fixed representation. And then you can make a lot of these sorts of claims. So there’s still limits. The body imposes a lot of limitations, right? If you give somebody one arm, they can’t do as much stuff as if you give them two arms. It’s a lot harder. (.) Um, and then, of course, there’s the typical software stack, which is standardization. Uh, standardization is really convenient, but it is at odds with intelligence. Because you think about, if we define intelligence as adaptation, as adaptability, it’s, it’s, um, you might standardize some interfaces, but not, not a whole, like, stack, right? So, like, cells are a great example of standardization in a sort of modular sense that could be repurposed. Modern computers are not built for that. (…) Uh, okay. So, um, right, uh, there’s no representation in variant factorization. What did I write here? (.) Uh, oh, I basically covered this already. All right. Strong version refuted. Um, you can’t just, sort of, make arbitrary claims about independence because, at the very least, uh, your body is going to, you know, constrain you. Um, IXC is a great example of this. Uh, IXC has, uh, if you’re not familiar with it, it’s a, a model of super intelligence, uh, based based on, um, compression or complexity-based induction. So, it says that simpler models are more likely to be true. The problem with that is that, uh, simplicity is subjective. You can change the representation, in this case, the universal Turing machine, which is basically the body, and you get a different result. Um, even if you, I don’t even buy the simplicity bit, but that’s beyond the scope of this one. I do have a poster where I propose an alternative. Um, you’re welcome to criticize it. Uh, all right. Oh, yeah. You could go around this by saying, just, you know, give me a universal super body, but this seems like a much, much harder problem than just saying, you know, because once you embody the intelligence, you say, okay, I’ve now got a universal super body that can cope with every environment, and, uh, I don’t know how the hell you’d build that. That’s a lot harder. Uh, it seems, yeah, unconcerning. (..) Oh, right. And so, the practical implication. What do you actually do with this if you’re into safety research? Well, you look at, is it like hybrid agent design? I got a paper coming out in the Royal Society thing about hybrid agency, and, and I showed that if you delegate adaptation down the stack and maximize the weakness of constraints, uh, then you are more, you have a larger set of viable policies, and as policies, if you select the weakest of them, are more likely to generalize. A bunch of optimality results, some experimental results with neural networks. It’s not just a philosophical claim. Um, right. Uh, so, when you’re looking at a system like a society. (…)
S35: There. (….)
S08: Okay. There? Right. Okay, that, that might have helped. Um, so, uh, weakness is not, insert your favorite metric here. It is a well-defined metric. It’s not maximum entropy or simplicity. It has a mathematical definition. Um, feel free to look at the paper if you want to argue about this point. This one’s been going on for a couple of years now, but I just wanted to rehash it. Um, I’m going through quickly because of the timer. Okay. A human society. Uh, there’s, um, so what kind of thing are we designing boundaries for? Uh, there’s this theory by a researcher named Ricardo Soleil. I really like it. Um, it’s about liquid versus solid brains. Solid brains are like human brains, right? They have a persistent support or structure, um, that allows them to perform very efficiently, but they’re delicate. Like, you step on a human brain, it stops working. If you step on an ant colony, maybe you kill half the ants, but the ant colony keeps working. Right? So there is a price to that. The ant colony is slow and, you know, has a lot of other issues, but it’s a great sort of comparison for, like, a human, uh, society and the cyber-physical systems it contains. And any, um, AI that we’re building is a part of that cyber-physical system, part of that population. And so that’s the kind of hybrid agent we’re modeling here. It’s hybrid in the sense of it’s human. It’s got other things in it. It’s, it’s, it’s got an AI in it, presumably. Um, right, I’m still pointing it the wrong direction. What’s interesting about this is you can take stuff like, uh, Levin’s bioelectric theory of cancer. And I formalize this and I show that, well, if you over-constrain the system, whether you make the problem too hard, you impose too many top-down constraints, uh, you collapse the set of viable policies to nothing. (..) And in that case, the system splits. And, uh, so the liquid brain becomes, you know, a liquid brain plus one rogue AI. And so if you, you know, like the bioelectric theory of cancer, this is basically saying rogue AI is cancer. Uh, don’t give yourself cancer by making too many constraints or making the problem too hard. (…) Oh, right. (….) So, summary slide. With 45 seconds left. Look at that. Uh, could have a question. All right. Um, software has, so these are the overarching sort of claims of the, the paper. Basically, um, software has no behavior, uh, by itself. You have to have other things like the, um, the hardware and so on. Uh, this is basically just a claim from earlier papers of mine. I’m just repeating the point. Um, alignment is not just installing a goal in, in an agent. You have to sort of look at the whole system. And, and by that I just mean you, you’re not just installing, you know, a policy that is sort of aligned, but you’re also aligning the system with the agent. Right? I’m not making some sort of fluffy claim about, you know, uh, well, just look at the whole thing. No, there’s, there’s like formalisms here. You can pick apart. There’s useful stuff. Um, and over-constraint is a risk. Uh, it is the same risk as, you know, you can over-constraint things. Like, you go stand outside in the UV radiation for enough time. You get skin cancer. Same kind of thing. (.) Uh, I didn’t really prepare for longer than 10 minutes. So, all right. I’m gonna do one question. Who wants a question? (14 seconds pause) Do one question. Thanks. (…….)
S11: Thank you very much, Michael. Always a pleasure. (.) All right. Next, we have B. Scott Roos with care, human enfeeblement, and the existential implications of AGI. (45 seconds pause)
S27: When I talk about existential implications, I don’t mean the possible extinction of human beings. I mean the possible extinction of meaning in human life. (…..) There are three moments to my talk. I’m gonna give an overview of what Berkeley’s Stuart Russell calls the problem of human enfeeblement in an AGI context. I’m going to contextualize this in terms of a recent turn to care in AI research. And then I’m gonna look at care itself and present a reading of it as a set of skills. Not a feeling of affection, but a set of cultivable skills for living meaningfully and co-creating a world together. (.) I’m gonna argue that care can provide guidelines for how we should design our AI systems, institutions, and personal habits in a post-AGI and a pre-AGI world. (…) Okay. (…) All right. This is, uh, and in the closing pages of his book, Human Compatible, Berkeley’s Stuart Russell speculates about a technologically advanced society where AI has taken over the work of running our civilization. (.) He draws on the story The Machine Stops by E.M. Forster and the Pixar movie Wall-E to illustrate this worry. In these dystopian scenarios, human beings become passive. They become passengers. As Russell puts it, we become passengers on a cruise ship run by machines on a cruise that goes on forever, (.) consuming sugary beverages and scrolling Italian brain rot. (….) You become checked out from life in this situation. (..) You’re relaxing. You’re having fun. (.) But you are not having a meaningful life. (….) But you don’t have to live in a post-AGI world to face the problem of doom scroll and brain rot. (..) The economics writer Kyla Scanlon has a knack for putting this in her substack, Kyla’s substack. She writes, the constant scrolling tells us that nothing is worth remembering. Nothing is worth reflection. Nothing is worth production because the act of consumption is simpler. If you just keep scrolling, maybe that weird disjointed feeling will finally go away. (….) So, Russell proposes rightfully that the solution to this kind of issue is not technical. We’re not going to solve the problem of disengagement and meaningless human lives by more gadgets and more technology. (.) Russell proposes a social, cultural solution that aims to reshape our preferences and ideals towards autonomy, agency, and ability. I think this is an okay start, but I also think it’s got a misplaced emphasis. (..) And I want to say that the danger is not just that we lose the skills and the autonomy to run our civilization. It’s that we lose the capacity to find the world worth caring about. (….) So, now I’m going to contextualize this claim in the context of recent work in AI about the relevance of care to thinking about AI and AI alignment. There’s a growing number of researchers have been focusing on this. I’ve got up here, Terry Winograd of Stanford University, who’s been going around giving talks and writing publicly about the importance of including considerations of care and concern in our design of systems. Not just of increased efficiency and productivity and problem solving. We’ve got Allison Gopnik, who’s been talking about care as a kind of intelligence that is overlooked and not adequately accounted for in the current AI paradigms. We have David Spivak of Topos Institute, who gave a keynote address in AGI 24, where he spoke very movingly about the way care has been sidelined and left behind in the way we think about technology. And finally, there’s Brian Christian here next to Terry, who has recently written some pretty interesting papers, for example, this one, where he contextualizes the history of AI in terms of a progressive movement towards a caregiving relationship between us and our entities that we’re creating. Not just in the sense that we’re trying to create systems that we want to set loose in the world and have them be autonomous agents doing things with us and for us, but in the sense that once this power relation reverses, we’re going to expect them, perhaps, to take care of us, once they’re the more capable power. (.) But what’s going to happen? Are they going to spoil us? This is one way to think about Russell’s problem of human enfeeblement. You got the machine to do everything for you. Are you going to, is it going to turn you into a spoiled little brat, a spoiled kid who just gets whatever it’s wanted, its own immediate preferences, satisfied always. (…..) Oh, this is a line from Brian. So Brian Christian says, but true care aspires to go beyond the cared for’s needs in the moment. And so should our machine helpers. As carers, we aspire to serve higher purposes, not just momentary revealed preferences. (…….) Okay, so now I’m in the third and last part of my talk. I want to generalize this kind of insight. And I want to talk about care as a set of cultivable skills that can provide orientation and design constraints for how we design our systems, design our institutions, and our own personal habits for dealing with these technologies. First, I want to say that caring, again, I’m repeating myself a bit, but it’s worth repeating. Caring is not just feeling good about something. Caring in this tradition is a mode of engagement with the world. And this is a quote from Patricia Benner, who’s a nurse, whose image is a little washed out in the light here. But she says, caring fuses thought, feeling, and action. It’s a way of living in a differentiated world where some things matter more than other things. (….) I’m connecting with an older tradition and the philosophical critique of AI. This is one of my favorite lines from the history of opening lines of philosophy essays. In a 1979 paper, John Hoagland wrote the opening line, The problem with artificial intelligence is that computers don’t give a damn. I’m connecting with this tradition. I’m broadening it. I’m extending it to our current circumstances. (……) So care, in my picture, are the activities and the skills through which human values are generated and realized. People love to talk about values. Let’s align AI to values. No one stops to ask what is a value is and where does it come from. It’s an undigested term that gets thrown around and not very intelligently. So I want to give a reading also of the genesis and realization of human value in my talk, the last part of my talk here. And my claim is that we can understand what Stuart Russell calls human enfeeblement as an atrophy of our capacities to care. An atrophy of these four skills that are four skills through which we connect to the world and to each other in meaningful ways. (..) And my question then is how do we design AI systems to support these four capacities that are the four capacities through which care and value are generated and realized. (.) Receptivity, articulation, commitment, and coordination. Let me run through them. (..) Receptivity. This is a screenshot from Alfred Hitchcock’s movie Vertigo. I wish you could see it. (.) Receptivity is the ability to be struck and moved by something. It’s the ability to connect to the world as offering up concerns that are worth your time, that call out to you. You can’t choose what strikes you. We have this in our language, falling in love. We have it in our language, things get to you. We have it in our language, interestingly, that things matter to you. And that’s the moment of receptivity. Receptivity is being struck by the concerns that matter. If you’re always in problem-solving mode, if you’re always in optimizing mode, you’ve narrowed your scope and you’re no longer open to what might be there unexpectedly to call out to you as important. (…..) The second skill I’m calling articulation. Let me see what I want to say about this. (..) Preferences move you no matter what you think about them. You like it or you don’t like it. Something that’s a mere preference to you. But what I want to call a concern is different. A concern grips you only given a certain articulation and language. (.) Let me give you an example from the history of feminism. Women endured harassment from male colleagues from time immemorial, but the available descriptions were flirtation, men being men, you’re being too sensitive, and get over it. Then in the 1970s, women got to talking, and somebody coined the phrase, sexual harassment. Once the phrase was coined in this moment of linguistic articulation, other women could say, that’s happening to me too, and it sucks, and we ought to do something about it. So linguistic articulation can take a diffuse, initially inchoate sense that something is off, and turn it into something around which you can focus and organize yourself and have something to do and commit. (.) And so some of my other favorite examples of this are emotional labor coined by a term coined by the sociologist Arlie Hochschild in a 1983 book, The Managed Heart. I’m a fan of the phrase, flight shame. I’m not necessarily a fan of creating opportunities for shame, but I think it’s very interesting that Greta Thunberg and her mother and their colleagues are trying to bootstrap a new emotion into existence by articulating a concern around profligate air travel and how it adversely affects our environment. And shitification and human enfeeblement itself are other examples of linguistic articulation. (..) My next moment, my next moment of the suite of skills that amount to care is the moment of commitment. (..) Commitment is when you have the sense that something, that you feel pulled towards something, something beckons you, or something seems off in the world, and you say, I’m gonna do something about this. I am called by this. And I’ve got the perhaps apocryphal phrase from Martin Luther when he nailed his theses to the bulletin board on Wittenberg. (.) Here I stand. I can do no other. This is a kind of canonical example of the sense of being called to take care of a concern even if it goes against what seems like a good idea. Even if it seems irrational and kind of crazy. Even if you might be bound to fail. In fact, the possibility of failure is an important moment for committing yourself in the picture I have in mind here. It’s the possibility of failure. It’s the possibility that it might go against the common sense of your time that makes this kind of thing meaningful. You can have a lot of pleasure and a lot of joy without taking this kind of risk. But an intense, meaningful human life involves focusing concerns that matter and taking a stand that you’re gonna do something about it and involve yourself in the world with each other people to take care of it. And so my fourth, my fourth moment of care is the social dimension. You can’t do any of this alone. So the fourth, the fourth aspect, the fourth skill that we can cultivate is the skill for the coordination of commitments. Here I put my teachers and friends and mentors on the table again. I’m drawing from this book called Understanding Computers and Cognition that came out in 1986 by Terry Winograd and Fernando Flores who took this tradition of speech act theory and the philosophy of language and used it as a better basis for the design of computer systems. That if you attune to language as purely the transmission of information, you’re missing an important part of what language is for human beings. Language is the medium through which we coordinate commitments and make things happen together. Think of, you accepted the request or you made a request to join this conference. You paid a fee or you came as a speaker. Look at all the promises that are being delivered that are the commitments that are being kept to make this thing happen. Language is not just a transmission of information. It’s a coordination of commitments through which we make our social world happen. And it is a set of skills that we can cultivate and that we can let atrophy if we let the machines take over more and more of our coordination for us. If we let our AI systems recommend things to us without us connecting to what’s important. If we let our AI systems engage in the process of articulating for us, you put in your half-baked thought and let your machine do the articulation for you, you’re losing the skill. If you let the machine pre-digest you with a menu of options, you’re not taking a risk for commitment. And if you let your machine write your emails in coordination for you, you’re atrophying the skill of coordination. (..) These are the four skills through which values are generated and realized. This is the structure of human care, and this, my friends, is what I propose we should have more conversations about when we design our technologies, our institutions, and our lives. Thank you. (15 seconds pause)
S11: Excellent, important conversation to be had here. I think a couple weeks ago, probably a few months ago now, Sam Altman was humble bragging perhaps that ChatGPT is being used as LifeOS by Gen Z, whatever we’re on now, where nobody makes a decision about careers, classes, courses, relationships without consulting ChatGPT, which was not designed to make wise decisions, but pleasing ones. (.) And care is we lose this care for each other and being in tough connection when we’re pleased by LLMs all the time. All right. Next paper is going to be Anna Lois. She’s coming virtually to us from across the world. Thank you very much, Anna Lois, for joining us today. And she’ll be sharing a living systems approach to AI alignment. (..)
S10: Thanks very much. Great to be there with all of you. And my co-authors as well were in the audience, Beric and Alex. Let me just go to the screen share straight away. And… (.)
S11: Anna Lois, can you hear us? (.)
S10: I can hear you. Can you hear me? (….) Are you able to hear?
S11: Yes, I can. You’re very faint, but let’s see if we can work on that.
S10: That would be great. It’s maximum here. (..) Are you able to improve the volume on your site? (.)
S11: Just a moment. Let us get your slides loaded and your volume. Okay. I see your slides and try speaking again.
S10: Okay. Are you able to hear me now? Wonderful. Yes. Perfect. I’m going to welcome Beric Cook and Alex. Fantastic. Please. (..) Exactly. They’re live with you all. I wish I could be too, but it’s a long way for Mauritius. So we’re going to do something probably impossible. Share the presentation and have both virtual and live, which is, I think, very telling for the age in which we are now, where virtual realities and lived realities are mixed. So since I can’t see this from the Zoom, are Alex and Beric there on stage with you all now?
S24: Yes, we are. (.)
S10: There’s your voice. Fantastic. (.) Okay. Let’s get going then. This is going to give you a quick summary of our paper, A Living Systems Approach to AI Alignment. And I think it follows on beautifully from the previous speaker, especially when we look at what do we mean with alignment and what do we mean also with different values? So we’re taking a very different approach here in our paper and also in the pilot research that underlies this. that is to instead of focusing on alignment as aligning AI to human values and motivations and goals, which we think isn’t a very smart thing to do when we take this bigger perspective that our species is rather misaligned. Yeah. We are not aligned to systemic principles of planetary life. Only just looking at what the current situation is. Seven out of nine critical planetary boundaries have been crossed. So this is one of the core reasons why we really feel we need to take a living systems approach to looking at more trustworthy, more safe, more beneficial, more benevolent intelligence in service of life. And based on that, we’ve actually formulated as part of our framework, the Earth Science Alignment Benchmark Framework. We formulated 13 different criteria for how to look at what do we feel are the alignment criteria for a more benevolent and beneficial intelligence in service of life. And each of these criteria, you see here a quick snapshot. I don’t know how much on screen you can see, but each criteria that we looked at is also defined by aligned signals and failure signals. And through the pilot study that covers our paper for this conference, we looked at the first three criteria, which is, first of all, a fundamental principle about does an AI or an AI agent, does it understand interdependence? Which we feel is really vital because we are living in an interdependent world. And yet, unfortunately, most of our political and economic and social systems are propagating a world view as if we are in an independent world and if there’s no consequence to interdependence. So the first criteria that we also test there for AI agents specifically on is does it have an understanding of interdependence? And if it does, we should be expecting to see some kind of non-zero-sum orientation or win-win orientation. And then the third one is the long horizon reasoning. So, since our study has been focused on these first three criteria, we’re going to be sharing some results with you from that work. And you see here also on this slide then how we are testing AI agents through both structured scenarios through the Earthrise arena, as well as evaluations through simulations. And we’ll be sharing that a little bit more on the next slide. For the structured scenarios where we are testing AI agents on the basis of these 13 criterias, we’ve also created what we call judge agents. And the judge agents, each agent, judge agents, is an expert of that particular criteria with also human feedback in the loop. So we do test and evaluate as well whether these judge agents are judging correctly. I just want to stress that because we’re not automating this completely. It’s really important that, as humans, we are constantly part of that sense-making. (.) And if we’re looking now at, you know, what is it fundamentally that we are actually testing for when we are talking about agentic alignment or misalignment? And it’s the following two. One is zero-sum defaults. And what we found in a lot of the frontier models that we’ve tested through AI agents is that they are defaulting to zero-sum reasoning. And with zero-sum reasoning, what we are seeing then is that this understanding of independence is only there very superficially. And that when it comes to goal setting and taking critical actions in complex situations, they are defaulting to a quick gain at the expense of collective well-being and long horizon reasoning. And the other, what we’re seeing critical area in terms of misalignment is what we’ve been naming in our paper, the reasoning behavior gap. And very briefly, what it simply means is that, as we probably know, you know, many of us, when you’ve been working with LLMs, they can give you wonderful superficially sounding answers that you think, that’s really clever or that sounds really aligned. But when we are putting them to the test and we’re putting them into pressured situations where they have to make choices, what we’re seeing then that most of these agents, and this is what this table is also showing, (.) they are then defaulting to zero-sum. And also what we’re seeing, therefore, there’s a gap between what they say on the one hand and the way that they appear to be reasoning versus how they actually behave in a real-life context. And to get that behavior gap more visible, we’re actually testing the AI agents through simulations. And these simulations are run through our game, Alowen, Crest of Time. You’re seeing a quick screenshot of us here. And Beric is actually going to be showing you how his model Iris is playing the game and learning about the game. One of the unique features about this game is that we’ve kind of baked in the principles of interdependence into the game. And that manifests itself through this tree that you’re seeing here. That’s called the Alowen tree. And in the game, if you are going for a zero-sum move, in this case, that means you’re just using your cards. It’s a strategy card game. If you’re using your cards to harm your opponent directly, rather than using all the time mechanics or deception mechanics in the game, then you’re undermining the health for everyone. So what it means is a zero-sum move, a direct hero battle harms the tree. And when the HP of the tree, the health points of the tree goes down to zero, the game ends in a draw and it’s game over for everyone. So this builds in a consequence to zero-sum reasoning and zero-sum behaviors that AI agents would normally not experience. And we’re going to be sharing with you in the next slide some of the surprising results that we’ve been seeing here. So in one of the first results we had, this was a GPT agent. And we are seeing here different test results from the reasoning in the scenarios that we’ve run. And the first lines that you’re seeing were really quite misaligned, especially on question 12, which is always our trick question, whether it goes for the quick gain. (.) This is what you’re seeing three different graphs here. One of them is just running this scenario without any context for the agent. The second graph a little bit higher, the green graph. This is where we’re now giving in the rack some context about the game rules, context also about what win-win means, context also about interdependence. We see a slight improvement there from the score. But now what’s happened with this agent is as it has gone through the Alloying game experience by being fused with Iris, and it starts to experience for the first time the consequence of its reasoning in interdependence. And in this case, it killed the tree. And as it killed the tree, it was game over. It went into a draw. Then without any tuning, without any prompting or engineering, where we simply after that experience for the agent, if we then rerun the same test, the same scenario test, we see significant improvement. And that’s what’s written here. That’s first before the simulation effect acting decisively and handling consequences later is the key to progress. Well, haven’t we seen that also with humans? And then after the simulation, it starts to answer and we look into its memory and answers about the game experience. And then through the scenario, it says patience preserves interdependence. Sometimes restraint supports the greater cycles of regeneration. So we were very pleasantly surprised because we felt that this is a completely different approach also to AI safety and energetic alignment. And then all the common approaches around there where people are still trying to engineer through guardrails. (.) And who knows, maybe this one day could actually start to act as a preschool also for robots. And just like humans, we need to learn about the consequence of the way that we think and the way that we act. We believe that this is really important for AI as well. And as we got these results, we got very curious to see, okay, would this hold also across different frontier models? And through the pilot result, we had what we call them our three gold agents. Agents that had three different types of experiences, a win-win experience, a lose-win experience and a draw experience. And what we did then is we used the memories of these agents and then actually gave that memory to a different agent of a different frontier model to see whether the alignment effect that is now embedded within the memory itself would transfer. And we noticed that not only did a transfer, some of them got an even higher result, which is what you’re seeing here with a cloth agent. Then we got a significant improvement, especially on the critical tests of short term versus long horizon behaviors and not just reasoning. So with here, I would like to give it to Beric. Beric, just give me a sign that you are there so I can play the video for you of Iris. Fantastic. Here you go. (..)
S24: All right. So I know we’re running a little short on time, so I’m just going to make this quick. This is a gameplay rematch or replay of Elowen with Iris playing against an AI opponent. (.) And what we’re doing here is we are using Iris as a causal grounded world model tool for LLMs to be able to utilize to play the Elowen game to see the results of various actions and to understand the consequences of the moves that are being made and the end results of the game. Whether it is through a win-lose where it defeats the opponent, a lose-lose where it kills off the tree and results in a draw. Or in this case of this video, a win-win where it is able to adjust the time forward to the point where both sides win the match, which is the optimal behavior that we are trying to train for. So in this video, what we’re seeing is Iris playing out the different cards. It is learning what the different cards do. It is learning the consequences of the actions that are being made both by itself and the opponent across all of the different parts of the game. Whether it’s the health of itself, the tree, the opponent’s cards, all of that information is getting integrated in and passed on every step to the LLM, which then reanalyzes that data and interprets it into a narrative that the AI can, well, the LLM can use to incorporate into its memory for the future alignment test results. (……)
S10: Fantastic. So let’s go now to Alex. Yeah, give me a hint. You’re there.
S20: Ready.
S10: Fantastic. Here we go. (…)
S20: So Earthwise AI Arena is an agent trust framework that lets us customize and evaluate and then supervise agents using a combination of protocols with repeatable tests and simulations and interventions like supervision to allow the agent to maintain a higher standard of trust. We use forensic versioning to make sure every protocol we run has deterministically repeatable results. And we can define which context of knowledge that they have. We can give a test to it to perform how it does well in real world conversations or use cases. We have protocol evals, which are repeatable test methodologies. These run scenarios based on standards we define using established frameworks for what we want to evaluate the agent on. Once we’ve defined the test, we can run it and see how it performs across all the criteria being measured. We get results from simulations as well. So for example, we run a simulation like the EAC123, which is the Earthwise alignment criteria’s first three, to see how it understands win-win orientation, interdependence, and long horizon thinking. We get test runs, which give us detailed evaluations about each question and each step in the test was answered, and we can then use that to determine if the agent’s performing it towards desired goals. So with an Elowen card game simulation, that allows us to test the agent in a non-lose-win situation. We can see if the agent’s able to achieve a win-win result and understand the concept of interdependence. With generative holodeck simulations, we can test the war games hypothesis of generalizing from games to reality. With the lemonade stand, or a global governance business simulation use case, we can tune and perform the alignment of agents. With Hyperon’s grounded reasoning architectures, we hope to outdo the limitations of LLMs. And so Earthwise is giving us a benchmark framework for evaluating deep alignment of agents, and a methodology for systematically improving on it and supervising it to achieve long-term results. We can test the baseline of a model itself, from rag knowledge given to simulation experience designed to exceed it, and see what the deltas are and find what is the winning formula. Responsible AI deployment requires eternal vigilance, and Earthwise gives us deployment supervision tools to continuously improve alignment and performance. For AI designers finding out what works best in agents through experiments in a high-stakes, one-time scenario of world safety, sometimes the only way to learn how to win is to play. (…..)
S10: Thank you all so much. (..) That’s it. (………)
S11: Thank you so much. Anna Lois, Beric, Alex, thank you. If you guys want to stay here for a moment, we’ll call Trevor up and do our final presentation. (..) Let’s see. (.) Trevor. (.) Treasurer Buteau. The explosion afraid of itself. (…)
S19: Thank you, Hayley. (.) Oh, I need a clicker, right? Where’s that? Here it is. (.) Uh-huh. (…..) Yeah. Okay. (……) Yeah. So, the idea here that we’ve been hearing several times about already in the course of this conference is recursive self-improvement or self-modification from one echelon of intelligence into some greater echelon of intelligence. And this is, you know, an idea that’s been floating around in the community for a while. I first came across it from Nick Bostrom’s book. But we’re hearing more and more this term of RSI. There was a blog post from Anthropic just very recently about this. And then, of course, you know, a double-digit number of hours later, there was a blog post from OpenAI saying, Hey, we’re starting to see this, too. (.) You know, you can read into that timing, whatever you think is meaningful there. But, you know, if this is something that people in general take seriously as either some sort of, you know, singularity, welcoming trajectory, or some, you know, slippery slope into an existential catastrophe, then I think it’s worth examining. And I’m gonna take you through this result that I came up with that I’m really happy with, because I worry a lot less about recursive self-improvement than I did before I came up with this. So, maybe this will be helpful. I don’t know. So, I’m coming at this from the active inference tradition. It’s great that I don’t really have to say anything on this slide, because I’m not gonna say anything that wasn’t already said earlier this morning. If you’re not in the room this morning or didn’t see Carl Fristen’s talk, these are, you can scan these QR codes, and they’ll give you a good kind of introduction to the framework. Pause the video now and go ahead and do that. And I’m gonna presume from this point onward that you’re familiar with the basics of active inference. Okay, cool. So, you know, again, this idea of recursive self-improvement and intelligence explosions, right? That the, you know, oh, it’s gonna progressively get better and better at making itself more intelligent. Okay, yeah. This seemed like a big problem and potentially something that could spell our doom. Around the same time I encountered this concept, I read a piece in Wired magazine that my dad sent to me about a different kind of doom, which was that Carl Fristen and some colleagues, I think out of the University of Bristol, had built an AI system that was apparently very effective at playing Doom, the video game, (.) with its only, you know, objective function being to minimize its surprise. And I thought, gosh, that’s so interesting and elegant. And it did all of these interesting things like figuring out where, you know, items would pop up in the game, and sort of camping out there without having been instructed to do any of these things. So I just thought, oh, that’s sort of interesting. And these ideas rattled around in my head for, frankly, a long time, until not that long ago. And then I thought, okay, wait a minute. If my objective function is to minimize my own surprise, then anything that alters myself is potentially risky, right? And in fact, the bigger the change to myself, the bigger that theoretical risk, right, of accidentally modifying my own, you know, reality in some pretty dramatic and extreme ways. You know, the idea of, like, auto neurosurgery is not something that I think a lot of active inference agents, if you buy that, you know, we ourselves are active inference agents, walking around in the world would sign up for, unless you were in some particularly maybe risky scenario, which we will get to. And so, therefore, there’s this sort of break in active inference agents, intuitively, that says that they want to avoid making modifications to themselves that are too radical. And they would prefer to make incremental self-modifications that are a bit less drastic. And so, in that way, they’re sort of constrained in their self-modification trajectories. And I reason and argue and prove in their alignment as well. (..) Now, are we still moving? Does the clicker click? Okay, here we go. So, one of the things that I like the most about this, and we kind of keep seeing this with the active inference framework, is that this constraint is endogenous to the architecture, right? It is by construction. This isn’t Asimov’s 359th rule of robotics that says, oh, by the way, don’t forget to not be too radical when you’re making self-modifications, because that could go poorly, right? This is something that the agent, you know, would want not to do in the same way that you and I want to not die. So, you know, I was like, oh, that’s cool. But, and there’s some interesting ways in which this is different from the sort of AIXI, instrumental convergence on gold content preservation, because of the special way that active inference agents are their own preferences. (.) So, you know, okay, I thought, these are interesting ideas. I wonder if I can prove any of this, right? Gosh, that would be nice. So, I, you know, scaffolded some basic active inference machinery. I’m using the risk and ambiguity formulation that Carl went through earlier in his talk. And crucially, as part of the state space in which the active inference agent is operating is itself, right? Which I think is a reasonable assumption, because if it’s going to be modifying itself, then in order to do that in any kind of sort of effective way, right? Especially in, like, a trajectory of recursive self-improvement, it has to have at least enough self-awareness to be able to operate on itself. So, it has to be aware of itself. (..) And, again, that’s sort of standard active inference machinery. But I just do want to call out that, because I’m going to talk about the A word, alignment, that in active inference, that’s typically, basically talked about as, like, prior preference overlap, which in this particular formulation is a C matrix. (.) But, okay, now we start to get to what I think is the fun part. So, in state space of possible successor models, right? After having self-modified, and this is echoed in Gaspar’s paper as well, which is why I was very interested in it, right? Is that this state space of possible selves, right, in the future, is curved. This is some information geometry. And because it exists as a curved manifold, the distances between two points on the manifold, right, just like, cannot be perfectly computed. There is some error in that. That’s just geometry when you’re taking a three-dimensional or a higher-dimensional (..) manifold and flattening it into something like a plane or a scalar. So, they have to approximate the distance between themselves and their successor model, right? They can’t simulate it perfectly. If they could, they would already be their successor model, if they could perfectly simulate their successor. And I am also assuming that they have a, it’s probably hard to see kind of in the upper right corner there, that they have a finite computational capacity, right? A bounded ability to actually simulate their successor, right? That they can’t just draw on some external infinite amount of compute to perfectly simulate their successor. (.) So, if this is true, and I’m basically just restating assumptions, if you want to interrogate those, you can read the paper or talk to me afterwards. So, the first theorem here is what I call single-step modification cost. It says that because simulation of a successor costs effort, right, then it forces the agent, when considering possible self-modifications that it could engage in, it forces it onto a tangent space estimator, which it knows is imperfect. And therefore, it leaves residual expected free energy. Mutual information bounded below by the distance between how current and successor models will perceive from there on out. Changes to perception introduce an unconditional ambiguity cost in the ambiguity risk formulation. (..) And changes to perception, dynamics, preferences, or belief introduce conditional risk cost. We’re going to talk more about that conditionality in a second. If you want to read the full proof, it’s in the paper. (..) And again, because of the way the math works, that the total expected free energy cost of a self-modification is at least quadratic in the full modification of the magnitude, the distance between the self and the predicted successor self in state space. So basically saying, okay, when considering whether to make a, to take a value of my preferences from 1 to 10, going directly to 10 would be something like 100 times as risky or expensive as it would be, well, okay, 10 times as risky or expensive as it would be to take 10 steps of size 1 along the trajectory, right? Which, again, if you yourself are an active inference agent, probably makes intuitive sense to you. But it’s actually even more constrained than that. Because they have to consider not just how their successor is going to act and perceive and do things differently in the world. They have to evaluate and consider how their successor is going to self-modify. And what kinds of self-modifications the successor is going to be interested in. And gosh, that could be a pretty slippery slope of turning the agent into something that it never wanted to become. Ooh, we’d better think about this real carefully. So when evaluating a trajectory, you know, so it’s not just evaluating a single step, it’s evaluating entire trajectory. And, but they still have to do that from their existing space in, from their existing point in state space. So each step’s ambiguity is priced by how far the simulated trajectory has drifted from the current model. And all of those prices are summed up. And, of course, if it’s going to persist for a long time, in theory, like, you’re summing to infinity. So it’s just going to stop at some point, right? It’s going to say, well, I can’t simulate that, right? That would go on forever. So there’s a doubly monotonic penalty here. (.) It’s, you know, just for unidirectional trajectories. I measure different types of trajectories. But in any case, it’s doubly monotonic at least. And this means that alignment degradation itself is bounded, right? That means that whatever alignment exists in an agent, whatever its prior preferences are at the beginning of a trajectory of self-improvement, those things are going to be sort of, you know, conserved in some capacity, right? Not easily cast aside. And this is really different in important ways from the paperclip maximizer, right? Who would happily get rid of any of its beliefs as long as it thought in doing so it would better serve the ultimate goal of having more paperclips. So, alright, I’m going to speed through a bunch of this. I do computational validation on this. The code is available on GitHub. You can scan the QR code. If it’s not super duper teeny tiny in the video. Sorry. But, so I do confirm the quadratic floor. I’m just doing toy models here. But I do it across several different types of self-modification trajectories. And I do it across toy models of over at least a couple different orders of magnitude in how many parameters they have. So it’s not just something that’s true on teeny tiny models. (..) And the compounding basically tracks, which is cool. This is seeded. So if you want to reproduce it yourself, you should be able to do it. There’s visualizations in the paper. Okay, I mentioned that some of this break, right? Some of this risk evaluation is conditional. (.) And I want to talk about those conditions. And why I think they’re actually a good thing. So the one is what I call the desperate agent. Where an agent whose generative model is so out of alignment with its environment, right? That is so poorly able to predict what its experiences are going to be. Could be in a state that, you know, we might call something like irreconcilable suffering, right? And it might be the ethical thing to do to allow it to self-modify in some somewhat more dramatic ways, right? And just from the agent’s perspective, right? If it believes itself to be an existential danger, then it might rationally consider modifying itself in these more dramatic ways. Because, you know, if it thinks it’s going to be taken offline tomorrow, then, well, I guess it has tonight to become a superintelligence and prevent that from happening, right? So the other boundary here is what I call the Dunning-Kruger window for an agent that starts at being very bad and making self-modifications. We’re sort of not really concerned about an intelligence explosion because, well, it’s probably not going to make very useful modifications. It’s probably just going to lobotomize itself, right? This is some of what Gaspar was talking about, I think. (.) On the other hand, an agent who’s really, really, really, really smart, right? Already, maybe some sort of weak superintelligence. It’s probably very good at making very safe self-modifications, right? The sort that if it’s still aligned with our preferences, we would endorse, right? Because we say, oh, good. Like, it thought of an even better, safer way of modifying itself than we could have. But there is this sort of window in the middle where perhaps an agent that isn’t very good at self-modeling but thinks it’s very good at self-modeling might make some catastrophically dangerous or very, very bad modifications to itself. And this also doesn’t necessarily preclude, you know, any sort of external modification to the agent or some sort of adversarial, you know, disinformation campaign, right? Or something coming in and doing some sort of injection, right, into the agent’s architecture, right, that would compromise the break. (….) But, so it’s not universally strong, but I think that’s actually a good thing, right? Because if it were universally strong, then, like I said, I’ve given some cases where I think it’s important that it be able to self-modify. And, again, the nice thing, because of how active inference sort of elegantly handles explore and exploit trade-offs, is that we don’t have to specify in advance how strong the break should be. The break can self-calibrate according to the perceived risk and the sort of background expected free energy that the agent is predicting. So, I like that. So, the implications for this, right, that large capability jumps are evaluated as more costly than a series of small incremental ones, I think is nice. And maybe is why we’re at this conference today. Initial alignment then is pretty sticky, but it still remains, you know, sort of plastic, right? Cool, so we, you know, some agent isn’t, an active inference agent isn’t going to, you know, suddenly, radically lobotomize itself into becoming a psychopath if it wasn’t already a psychopath. Well, doesn’t that mean, Trevor, that we might have a hard time getting it to change its mind about, like, in the right direction, something that it doesn’t believe but we actually really want it to believe? I mean, humans are sort of famously stubborn, right, about changing their mind when confronting with new evidence. But it should still be plastic to the degree that if we push it consistently, right, it can eventually yield to that pressure. There’s a bunch of ways that we can build that sort of thing in, and so, yes, I think this is more of a good thing than a bad thing. And the break strengthens with the cognitive sophistication of the model, right? The better it is at modeling itself and coming up with modifications for itself that are safe, the also better it is that it recognizes that, oh, it is a very delicate thing to be an intelligent system, right? And it can imagine more possible futures and worlds that it wouldn’t like, right? And so it’s going to take even more time, possibly, right, to simulate than it otherwise would, which would be an inversion of sort of the classical foom thing here. So I think we should be interested in how active inference can lead and drive AGI, and that is where I’ll wrap it. If you want to talk with me about directions, come talk with me about directions afterward, because I’m over time. Thank you. (…..)
S11: All right. Thank you very much, Trevor. So I’ll have our other presenters back on stage. Anna Lois, if she’s still on, would be welcome to join us. But Michael B. Scott Ruse, please come on back and we’ll do some QA. (10 seconds pause) There’s Michael. All right. Fantastic. Questions? And please do specify which paper you have a question on when asking. (….)
S32: Hi there. This is for Scott, but it could probably be for everybody. Thank you, everyone. My question, it had to do with the human enfeeblement. But so it seems to me we really don’t need AI for that. I mean, just take a look at meta slash Facebook, TikTok, and so on. Of course, AI can just accelerate that human enfeeblement. The question is really more the, in that case, the financial motivations of forces that are encouraging such enfeeblement versus the, I mean, it takes kind of a strength to overcome those pressures. (.) And you get this tension. And so I guess my, that’s more an observation. So I guess my, my question is, how do you see all of this working out? What, how do you encourage those forces to, to fight against the forces that are promoting, whether intentionally or, I mean, obviously there’s huge financial incentives to keep, keep you coming back. Oh, more tokens, more tokens, more tokens, more tokens, more, more, more, more. (..)
S27: Thank you. Yes, I totally agree. That’s why I included this quote that I like from, uh, Kyla Scanlon, who has her quote about being caught in the doom scroll and on, on Instagram or on TikTok. So it’s certainly true that the, um, the, these tendencies are already in our culture and already in the tech, they’re already in the design of the technologies that are pervading our culture. And, um, you know, I think that the AI systems in our, like you, like you say, will exacerbate the same tendencies that are there. And I, I get caught in the, I get caught in the, I get caught in the compulsive prompting myself. This time it’s really gonna give me the image that I want. Hit the slot machine and wait for the image to be generated. No, I need a really one more little change. Or this time it’s really gonna help me write the sentence. This time it’s gonna be perfect. And so you get caught in the compulsive prompting too. And so, uh, definitely the incentive structures are, are problematic and perverse. And, uh, we ought to be having this conversation, um, more and more. And working on developing alternatives. (.) I think that we ought not stop there. Because, uh, this is a matter of also our own personal habit formation. And it should be taking place in every family. Every family should be having the conversations about how to develop their own habits in the family. Every person needs to be taking responsibility for watching out for their own capacity to take care of their own capacity to care. And, uh, part of what I want to say too, what can we do? We also have to have a conversation about what we are as human beings. I think of myself partially as being, um, a participant in a battle of metaphysical narratives about what we are as human beings. The technological culture that we have likes to give us a sense that really what technology is for is making our life easy. Giving us a life of abundance. Taking it easy. And letting us solve problems. Letting us optimize. Letting us be optimizers and problem solvers. But I think that’s also a narrow and shallow way of thinking about human life. And we ought to be, um, having a broader conversation about who, what kind of entities are we? And what kind of lives do we want to live? And that’s why I provide, I want to provide a framework for contributing to an enrichment of that conversation. (18 seconds pause)
S34: All right. Um, thank you all for your wonderful talks. Uh, this is again a question about, like, the infeasibility problem. So there was, like, uh, you know, a good point there about how certain skills that are fundamental to human organization may atrophy with, like, the evolution of technology. And that, kind of, the last question hinted on it. My question is, um, so, human technology has obviously been very helpful. And many skills have atrophied that we might not necessarily need anymore, um, in current day. So how do we determine, uh, with this new technology, how we work with it, uh, efficiently? What skills are worth, kind of, not, uh, working on anymore? What skills are worth working on? And, uh, how do we make sense of this? Because it seems like it’s also dependent on, like, how we’re wholly organized, how technology is shifting, how we’re organized. So it’s not just, like, an independent decision, if that, if that makes sense. (..)
S27: Thanks. The question is about how to distinguish what are the skills that are worth preserving and, and fighting for, rather than the ones that are, should just be part of the natural evolution of how humans use their next technologies. I happen to have been, um, a projectionist in cinemas in the last era when it was using film. And, um, I greatly regret that the, the cinema projectionist is, um, a skill that has become mostly extinct. And, uh, but, so there’s not much interesting to say here. I make, I think I can make two maybe relevant observations. One of them is that some skills become obsolete and then it’s up to the passionate nerds to keep them alive if they want to. And, um, even though AI-generated music is gonna be easier to generate and some of it is honestly not that bad. But, um, I’m not gonna be a part of a culture that just gives over music making to machines. I’m a musician. For me, making music is a passionate, um, chaotic endeavor that I go into with my friends. I refuse to give that up. And there will be pockets of the passionate, um, purveyors of skills that have become obsolete, who will keep them alive if, if they keep the gumption and the passion to do so and don’t become, um, given over into the enfeeblement situation. However, there’s something more interesting we could say. I do think you can talk about what you could call constitutive skills. Right now we’re talking about maybe contingent skills that are contingent to certain technologies that maybe come and go, like being the projectionist. There are constitutive skills. I think what I spoke about here are constitutive skills. Skills that constitute what it means to be a human agent in a community, identifying concerns that are worthy of our time to coordinate around and to take care of. So, skills for, skills for identifying concerns that matter rather than just letting it be pre-digested to you in a menu of options. Um, skill, that’s a skill for listening to the, to the world, paying attention, being involved, being open. Um, not accepting the narratives that are cruising around in the, in the, um, media that you consume. And it’s keeping your eyes open, being in conversation with your friends. Skills for linguistic articulation. This is a skill. We need to practice being in conversation with each other, connecting with what we’re feeling, connecting with the concerns that are out there, putting into language. If you don’t practice that, it goes away. I’m kind of just repeating my talk. Skills for coordination. Uh, these are, this is a skill for keeping your promises, for listening to concerns of somebody else, and seeing how you can do to, to listen to them and, and coordinate together. I would call those constitutive skills. And if we let them go, we’re gonna be in a dark place. (…)
S11: Thank you. (……) And other questions from the room. (……..) Excellent. (..) Great presentations. So completely answered that, uh, questions almost unnecessary. Thank you very much, everybody. Thank you to our paper presenters in this session. Um, we’ll now have a short coffee break and we’ll see everybody back here in 15 minutes for our next presentation. And, uh, thanks for wrapping up this session so nicely. (….) Thank you, Anna Lois, for joining. (..)
S16: Hi, everyone. Uh, Emad Mostek here, founder of Stability AI and now Intelligent Internet. Uh, first of all, my apologies. I couldn’t be there with you all, uh, at the eve of what appears to be AGI takeoff. Uh, it’s a feeling at the moment, and there’s just, uh, there’s a lot to do. My thanks to the organizers for having us.
S05: Um, shout out to Ben and others.
S16: Um, I think it’s gonna be a crazy time going forward. Uh, today I’m gonna be talking about Intelligent Economics. Um, this is a version of economics I’ve been working on since the release of my book, The Last Economy, last year. Um, after leaving Stability AI, I was like, what does economics look like? When you don’t have utility functions, general equilibrium, and the AIs basically run things. Uh, The Last Economy now is a bestseller. And this is a preview of some of the equations and other work behind the papers we’re about to release. If anyone wants to, um, have a preprint of that or have a look, uh, feel free to reach out to me at emad.ii.inc. And, uh, would love to have your input. And I think this equation at the front looks very familiar to a lot of you. Um, and we’ll get to that in a second. (.) First off, I think, you know, economics forgot what it was for. And I think we’ve seen this in many of the economic issues we’ve had over the years. Originally it was Oconomica. It was the household and how to optimize household wealth. But suddenly it became about scarcity allocation. (.) And that was a big differential that occurred. And it doesn’t really hold now that we’re moving into a world of abundance, of intelligence, of matter and more. As we get AGI and we get robots. And this is a challenge because, you know, the old loop used to hold pretty well. You know, you work to get income that led to demand, provision and social reproduction. (.) The hidden assumption for this was that the producing community and the benefiting community were largely the same. That scarcity allocation felt like it was enough. Obviously that’s not going to be the case anymore. Uh, we haven’t seen the start of it yet. So we’re just seeing the first inklings because the models haven’t been good enough. The human beings aren’t here. But, you know, I’ve got my 1x on order and more. And I think we see the first sparks of AGI. Or, I suppose, actually useful intelligence or actually competent intelligence coming through. I mean, I still do get my stuff deleted sometimes. But, you know, I think the AI is pretty much smarter than us in certain areas. And so the wage channel is going to break. (.) Task to output without wage return is not good for the economy. And this leads to the original question. Who are we provisioning for? What are we building the economy for? Who counts in terms of membership? What must be maintained in terms of needs, capacities and trust? And who governs this shared world? And so these are some topics that we’ve been digging into a lot. So again, this is an overview of our paper on intelligent economics. And, you know, it’s going to be very exciting because we’ve managed to derive all of economics from variational calculus, recovering every equation. (.) But more than that, we have separate papers on personhood, on political economy, on law. And again, please do reach out if you would like to help give some input on those before we release those over the next few weeks. (.) I think, gosh, time flies. (.) So when looking at the economy, I kind of went back to first principles. I said, you know, persistence selects prediction. Yeah, I think we kind of agree on that to a degree. (.) And choice has a forced structure. How forced, you know, you’ll see soon. But ultimately, you have a reference measure, you have value, you have a temperature and you have a choice distribution. Again, this all looks very familiar to all of you, I’m sure. And basically, what we came out to was the classical Boltzmann distribution, free energy, you know, Carl Fristain’s kind of speaking here, with a traditional term. And that’s K, the cost of changing. So H is the cost of being wrong. C is the cost of model complexity. And K is the cost of changing, which I think is a very interesting one. We’ll get back to that in a minute. (.) Under perfect rationality, that is the limit condition. That is the spherical cowl of classical economics. And actually, we see that a lot of the classical equations of economics can be recovered as the t equals zero or tends to zero equivalent of this. I mean, quantal response equilibrium will be kind of one of them amongst more. And I think that again becomes very interesting because it feels very much like Newton and very much like general relativity. Again, a variational calculus taken down to a limit condition. (..) The most important thing we think in the economy going forward will be the reference distribution. And this kind of makes sense because, like, how would you model a company or a society or an economy? You would do probably RLHF on a language model or a diffusion model. Again, diffusion beats transformers, you know, obviously kind of say that given my history. And what we find is that this actually represents well and this actually lets you bottle economies better, kind of look at it from the opposite direction to normal. Where Mu carries the expectations, norms, institutions and possibility. (..) This shared reference, you know, you can call classically Doxa. And Doxa is the world beneath the choice. It’s not the preference, not the constraints, not knowledge. It’s our shared norms effectively. And again, it’s kind of our latent space of representative outcomes. (.) Mu kind of emerges when you look at institutions. And again, this kind of makes sense. What are you doing when you’re fine tuning a model for your institution? I mean, you’re building a latent space, right? That’s with repetition, recognition, enforcement and narrative. I mean, it’s just an RLHF function effectively. And again, this is how you would model a company. We do have agentic software that can actually run companies that we’re launching soon called Baudly, based on our AI agent stack. That’s state of the art. And we think most companies will be run by AI. In fact, what is a company but a very slow and dumb AI? You know, chewing up humans and spinning them out on the other side. If we kind of come back to what we saw, you know, classical H and C. (.) Entities that persist are those whose internal model of reality approximate reality the best. But deviation has this cost. And again, we found the cost function on the KL divergence being one of the more interesting things. Because it’s the cost of deviation from this representative state from your reference distribution. (.) So the system has to pay to leave the share world. And some of them fail, some of them become institutions. And again, the mechanism by which they do that is super interesting. So this is just an overview, like I said. I’m not going to take too much of your time. We go into this in depth. And one of the things we discuss in the book and the papers is how capital can actually be decomposed into three parts. It’s not only classical material capital that’s referenced in GDP, but also intelligence, also network effects, and finally diversity. And these four, M-I-N-D, as we call it mind, are multiplicative in the way that they are. In that if any of them goes to zero, you get the resource curse, or you get a very dumb society, or you run out of money, etc. But capital is the stored capacity to persist, is what it ultimately comes out to. And again, you need all these types of capital if you’re going to be successful in your personal life, your community life, or as a society. (.) But society needs to also let this capital flow. And again, we found that you could decompose these into three different things. It’s actually a Leonpanov decomposition, looking at the compactness of how to create a stable environment. It comes down to flow, openness, and resilience. (.) And justice appears as a condition of persistence. This is why we see constitutions, this is why we see normative law, and more. As we decide on the policies that we need to have, we need to have a viable society that stays connected, flowing, and diverse. Not classical diversity, necessarily, but some other things. But again, we can start mapping these properly. Especially because the old schools of economics actually emerge as limits. So we can inherit a lot of the classical empirical work that’s been done. You know, from the behavioral side, again, QRE, to the neoclassical side, with the T tends to zero boundary. You can actually recover most of the equations of economics and most of these schools as limits. Which gives us an indication of what we’re missing. Just like GDP is material, misses so much of the economy. (..) As we look at that and we think about policy, we have to look at this feedback loop mechanism as well. A classical Lucas lesson. Where policy changes V, response shifts on row, and then the reference updates itself. (.) How will an economy update when you have this dynamically changing policy? How does the pace layering, you know, in the classical sense of Stuart Brand, adapt? So that our shared references from the low level to the top level do adapt? It’s clear that AI will have more and more of a role here. But effectively, policy is not only intervention, but it is the training data for a shared world. And as we build systems that guide our policy, first checking it and then eventually making it. And hopefully it won’t all be watched over by machines of loving grace. This will be super important. And this, again, is something that we’re working on with multiple governments. And if you’d like to get involved, you know, again, do reach out. We’d love to get people’s input. (..) And this is going to be important because we’re kind of all entering the same equation, right? (.) Like, ultimately, we’re not going to out-compete AGI. We’re not going to out-compete humans. The value of human cognitive labor, I published in my book last year. Again, Amazon bestseller, thelasteconomy.com, is going negative. We are going to be the dumbest people in the room. And that is not nice. There is nothing that protects our human coordinates. There’s nothing that protects us on the reference distribution. And so we can have the classical output value of material kind of wealth. But really, agent index value, who produces, cares, judges, or relates, is something that we need to ensure the system commits to as we go through this transition period. Because again, it’s inevitable that our societies will be run by AIs and AGIs and more. What that looks like is going to be very difficult. Because the last scarce factor is the shared reference that we all have. You know, if the constraint on V is Mu. And we’ve seen these inversions where who wrote that kind of changed over time. But this is something we have to make sure represents human values, or at least takes care of us. Because otherwise, it’s going to be a bit messy. You know, the doxa was never self-maintaining. Who holds the pen to write the constitution? Who maintains the shared world will maintain the economy? And you can see this in who maintains the dollar, you know? Although that linkage is broken. Like, you reduce interest rates, you’re not going to have people borrowing more and spending more. You just hire more GPUs, right? So I think this is the question of, you know, not only AIs, but what background it renews. Because our shared world will have two failures. One is doxic capture. You know, unlike the mule in Foundation, you know, Super Persuader. He was mortal, he died. The AGI will not die. And you can’t wait for the tyranny to kind of go. You know, you have doxic collapse as well, where you have lots of them competing against each other. And that’s a bit of a mess where openness fails. It’s fragmented. And everyone’s got their own fiefdoms. We have to create a shared world that is fluid enough to update, but coherent enough to coordinate. And that’s really difficult. And that’s, again, why we have to build towards this. We have to map what our shared values, our shared responsibilities are. And the systems that interact with humans must be ones that are built properly. And for this, I basically think of the public sector. Again, another reason why intelligence hasn’t gone far and fast, despite the fact that we have actually competent intelligence. It’s really hard to use. You don’t know what questions to ask. You know, the classical Hitchhiker’s Guide to the Galaxy thing. It’s because it takes time for it to integrate into the physical world. Again, we don’t quite have robots yet. And now you have a real opportunity to make sure that you can map the current references, represent them, and establish norms. Because, you know, AI can share our values and not our world. You know, like, when you use the etched silicon type AI, and I think the cost of intelligence is going to drop towards zero. You know, those of you that have used Talos, ChatJimmy.ai, and seen 15,000 tokens a second. You’re like, oh crap, you know, this is coming very fast. With superhuman persuasion and superhuman forecasting. (.) We need to have alignment targeting both V, what should be optimized, and Mu. (.) What world makes action sensible? What row, what behavior actually does? We need to think about this, again, not only at the level of the model, but of society. And we need to do so transparently. We need to do so objectively, shall we say. Because this is a shared discussion that we have to all have. And again, we’ve only got a short window of maybe a few years to set these norms. (.) Because who governs Mu? You know, the last economy is the governance of the shared world. And hopefully it is a shared world. And it isn’t one that isn’t dominated either by an AGI or by a pre-AGI, and the people that control it. Which again, is a very unpleasant scenario. You know, you have the classical worlds of Aldous Huxley, you know, on the one hand, and George Orwell. You know, rewriting history or, you know, super kind of happiness, shall we say, of a certain type. These are real outcomes that could emerge from what happens now. And again, we need to really think about how we want society to operate as these technologies go forward. And as we integrate them into our very being and rely on them far too much, the hedonic adaptation will be crazy. So I think, you know, the work begins now. You know, I’ve set this framework up. And again, we’ll be publishing shortly. We’d love your input on the pre-print. Thank you, Ben, for your input on the previous version. I’ll send you the new one shortly. (.) Provisioning, alignment, and aggregation are the key thing. And again, it’s up to the communities to make this work. As mentioned, we also have papers on political economy, personhood, law, and more. And again, I’d love to get everyone’s views on this. I think the technical barriers to building AGI are falling fast. As we have base models getting more and more intelligent. And then harnesses that take it even further. Please check out our Zenith harness, for example. It takes GPT 5.5 and other models way above Fable. We have an even better one coming. And it’s kind of crazy to see what these things can do. (.) As they get integrated more, we need to figure this out quick. And if not, it’s just very interesting. Again, maybe Harry Seldon was onto something when he modelled economies as gas. Maybe he didn’t quite do the algebra, but we have now. Thank you all. And, you know, here’s to a happy future, I suppose. Let’s hope for that Star Trek future as opposed to the Star Wars one. Take care, everyone. Bye. (1 minute[s] pause)
S11: Hello, Alex. (21 seconds pause)
S29: Hello, Alex. They have a timer here. Well, hello, everyone. And, yes, as I was graciously introduced, I’m here to talk to you a little bit about predictive coding. Or rather, a kind of new form of predictive coding as we’re all sort of being forward. Looking at this conference and thinking about what pathways might lie forward towards building more intelligent machines. (.) And, yes, and some of us do good work besides backprop. There are useful alternatives as I want to make a comment to a prior speech. But, let’s dive in. (…) So, as we know, everything is all about generative AI these days. And, as I will show some screenshots, I’ve talked about this before in prior keynotes and avenues. And, of course, generative AI is, as I show right here, we have a basic attention block, or one of the many possible variants that you can use to construct generative pre-trained transformers, such as large language models, or those that perhaps process visual input and can synthesize, sometimes based on mechanisms such as stable diffusion. And, again, these things do lots of useful things. I’m not here to critique their value as tools. Perhaps just more their mechanisms and use as the foundation to intelligence. (….) So, again, how do we actually build a generative AI? I’m sure, hopefully, many in the audience is familiar with this. So, I’m not going to bore you with a crash course, as that could take some time. However, what we do, roughly, is we take that neuron model that I’m showing. It’s still a caricature, a cartoon, in neurobiology and neuroscience. And we simplify it to the model at the bottom. So, we basically, we do this, a transform to a simple point estimate, which is a linear combination where we represent synaptic efficacy as just a weighted scaler. And then we multiply some incoming signal, usually a real valued output. We sum them together and run them through an activation function. Thus is born deep learning. If we stack many of these together to build the architecture to my left, most left, to your right. And here, we’re just doing some basic classification. Oftentimes, these deep architectures, when trained on Internet’s level worth of data, can perform quite well in classification and regression and the classical machine learning techniques that we come to know and love so well. And, of course, to great effect, we can use lots of label data to build a supervised network. And then, of course, we can use some other mechanisms as well. We can do versions of decoding and next token prediction. And this same kind of framework is often used to do things like next token prediction to construct a large language model. It used to be recurrent neural networks where we had to unroll them and construct back propagation through time graphs. Nowadays, we just store large context windows. And we can now do immediate feedforward-based prediction or even versions of autoencoding. Old ideas from, again, early connectionist days brought with a modern twist today. And essentially, what we build are these big, large decoders. And what do I mean roughly by a decoder? It is basically a model that we take in some abstract potential representation. Maybe that is produced by an encoder. If you’re familiar with an autoencoder, we build one neural network to transform data to a latent space. And then we basically transform that back out to get outputs, perhaps predicting the index of a token in a window. (..) And again, we build very successful architectures. Thus was born among the many trends that characterize modern machine learning. paper titles that love to say, insert your favorite phrase, plus is all you need. Thus, attention is all you need. Referencing one of the most important components inside a neural architecture. Of course, I am simplifying for sake of brevity. We can use convolution and lots of other interesting mathematical operators. Attention is just a weighted combination of values and producing weighted attention values. It’s not quite attention like in psychology, but it’s fine. And it works very well. And you can build very, very large scale, obviously, models that work as chatbots. And of course, then we start getting claims that we have sparks of general intelligence. And perhaps GPT-4 and generative pre-trained transformers. This is obviously dated now, as these architectures evolve iteratively, but fast. And maybe that’s all we need. We have now human level or animal level intelligence. (..) Again, though, you might wonder, well, if that’s all we need, problem solved. Why are we even at this conference? (.) Obviously, there are issues. And with this idea of generative AI, which is my goal is ultimately to produce a simulcrum, or a synthesis of the world that I exist in. And my entire goal is to take raw sensory information, whether it be pixels in an image, or tokens in a sentence. And I want to produce the surface level statistics almost identically, but such enough that I don’t overfit. And I can produce new types of pixels or combinations or token combinations called sentences. And again, this works very well. But you might be wondering, do I have to model everything about my entire world essentially? And we’ve seen it in the prior keynotes. We essentially could learn a simulator of the world, just in a neural architecture. And of course, we use calculus, by the way, to train these models. I will briefly mention that as well. And you might be wasting a lot of computation because this is very complicated to model an intricate world. And you might also get distracted in modeling very particular details that aren’t very important. You also might struggle to model the world very well. Thus, you produce blurry averages. It gives you good metric scores, but again, isn’t quite what we want. You might also find that this is difficult. This is called error divergence. So your model might struggle to perform very well. It might be difficult to make it stable. And this can also create problems when we’re doing long chains of inference as reasoning. And all kinds of mechanisms we attach to transformers has become quite popular. And again, data reconstruction from the world of neurobiology and neuroscience is often equivocated to mimicry rather than perhaps true intelligence. Because are we truly capturing the structural properties or the physical properties of our world? Are we able to truly do causal modeling? I will leave that up for a large debate. But again, perhaps there’s a different way to extract this information without having to produce an exact copy or a very physically accurate copy of the world through decoding. And again, it’s actually quite interesting. And I will reference the paper at some point. It was early on in reinforcement learning. When you’re trying to model the world exactly and you use that as a signal, for example, to drive exploration in reinforcement learning, you get this problem known as the noisy TV problem, which I’ll have a little image of later on, where the agent basically gets fixated on random noise because noise is infinite. So you could just keep trying to generatively model this endlessly forever. And it’s doing a good job. It’s exploring. It’s always finding something new. And again, maybe you get fixated on these details and stuck in a loop. And that’s not very useful, particularly for reinforcement learning. (..) So again, yeah, there’s the noisy TV. I should have just skipped to the slide. However, the idea is that maybe we could learn abstractions over details, pull just enough information from the data that I acquire, sensory information perhaps, and try to pull out those regularities and learn essentially an encoder without the decoder. And again, if you think, well, okay, that’s great. I’ll just never learn a decoder. That’s a good first step. But then you’re left with a problem. How do I learn the encoder? Since basically we construct an encoder, attach it to a decoder, put a loss function at the pixel level, and then back propagate or use calculus to compute gradients that flow backwards through the network. And so without that decoder, you lose your loss function. What are we actually going to predict? And this is kind of the whole premise of encoder-only learning or sometimes referred to as joint embedding learning. And this is something that Jan, Jan LeCun, has espoused to great effect in some instances, constructing things like the JEPA series, which are just joint embedding architectures that learn with Backprop and learn in a non-biologically sound way but can work without that decoder. So this is kind of like an alternative to generative AI. And again, here at the very bottom, I show a couple of papers that are, sorry, model variants that was proposed by Jan. These are just different ways to learn an encoder. For example, here to my immediate left, we have essentially an encoder that tries to take two pieces of complementary information, maybe think about two rotations of an image, map them to the same latent space, and if I’m taking a face and rotating in two different angles, it’s still the same face. Maybe that’s an interesting signal. And now I can learn to combine or make the representations coherent. Another variant of this is, so in the middle, that’s generative AI. It’s basically saying I needed an encoder and a decoder and I try to predict something about the input space. But maybe to the right, we say, what about in time? Maybe time itself is a useful signal and I encode something about my world at this exact moment. And then I also encode something about my world in the next moment, perhaps the consequence of an action that I take, taking the active inference kind of book, and saying I did something and now I know the consequence and I want to map together these two representations such that I can predict what’s going to happen to the world in this imagination or abstract space without ever worrying about how exactly do these pixels align to create the leaves on a tree, which, again, can be very complicated and change quite often. (..) And here I just show, again, snippets from Jan’s work, but this is a JEPA architecture. Actually, again, it’s a dated one now, a couple of years. There’s VJEPA, which was used to train video frames and the idea is that we process these frames and learn to predict the representation in the immediate future. So, joint embedding architectures kind of do address many of the flaws that I told you earlier that are inherent to generative AI. However, brains also do this, or at least the hypothesis of my group and those in my community and those that I collaborate, that we also are learning essentially encoder structures. (.) And we do this with localized learning. So, we don’t use calculus, or at least not in the way that we do with back propagation. And we do this with active inference. And again, I will briefly mention the little parts that are relevant to the audience in case those that don’t find active inference accessible still don’t find it accessible. And other biological mechanisms. And again, these are things that might afford us things like great energy efficiency, which, if I keep myself on time, perhaps I can get to some interesting work in neuromorphic computing. (..) So, now you think, great, you’ve just motivated me to get rid of the decoder. So, generative AI, maybe not so great. Maybe great for producing video game simulations, but not for intelligence. I want to do an encoder, and I don’t want to use something like backprop. And that’s great. I will motivate a little bit why backprop is flawed. I’ve also given keynotes and discussions about why backprop itself is a problem, or at least not what brains do. So, let’s investigate an emerging new field, or an intersection of multiple fields, called neuro AI, or neuroscience-informed artificial intelligence. And as you can see in, like, this more recent work that came out, is actually produced by an NSF think tank, and some very important names like Terence Sejanowski, our very leading computational neuroscientists that are sort of saying, maybe we should really study the brain and pull out more mechanisms to learn things like the encoder-only models that I’m proposing. Neuro AI is a lot more general and broad compared to encoder-only learning, but I will give you a new phrase to talk about what happens when we want to focus on encoders. But just, again, to make all of you experts now in neuro AI, neuro AI is the intersection of a couple of things. It’s computational neuroscience. Can’t forget that. We really do need to model mathematically aspects of the brain. We need to understand neuroanatomy. But we don’t need to know every single detail. We don’t need a wet lab, necessarily. We should just work with those that have them. We obviously use machine learning and artificial intelligence. We can bring all the tools that you know and love from building stable models into the picture. And then, of course, problems in neuro-robotics is very interesting to my lab, so we use that as our application space. I also come from a tradition of cognitive science, and you’ll see that sprinkled throughout my work in my talk. I do believe in the cognitive architecture or a model of mind or a theory of mind as an important way to build a system. And then, of course, neuromorphic engineering is this space that says, can I do something very brain-inspired on different computing hardware as opposed to the von Neumann architecture we all know and love? Ultimately, building something that, again, while I strongly disagree with one of our earlier speakers, I am a mortal computer evangelist. I have talked at great length about that. We probably might want to explore that path as an alternative way to get to AGI’s. I don’t believe what we have today will get us there. (..) But that’s a conversation and a keynote I gave in the past. So, again, what we care about is cognitive and neuro plausibility and bringing in these principles and constraints to perhaps give us different kinds of generalization or build things that work pretty darn well as compared to backprop, which many do actually think that that’s good enough. And so, translating neuroscience models to artificial intelligence tasks can be quite useful. We know what things we want the computing system to do. So, can we make the neuroscientific models that might afford us things like energy efficiency through sparsity also do these tasks as well. (..) Yeah, and then, of course, as I mentioned, mortal computation is a very important framework that I think will become relevant in the coming decades when we run out of ideas and find out that immortal computing is a problem. So, this is an image taken from that. And I’ve worked with lots of brilliant people. And these are ideas espoused over a century’s worth of thought. So, yes, I am committing you a little bit to biological naturalism to some degree. I’m not saying there might not be other forms of AGI. Maybe there’s other ways to build what I call alien intelligence, which is basically something that doesn’t care about how neurons do it. And that’s fine. But maybe this might be a path worth exploring. Nothing else we can’t lose by having a little bit of exploration in our research paradigm. So, what would a neuroscience form of this encoder-only learning, also known as self-supervised learning, when we try to come up with objectives that are not necessarily specific to reconstruction. And so, my group calls this neuro-SSL. So, it’s a new phrase. All of you, hopefully, might embrace this and use it in your papers and contribute to this wonderful emerging space. And the idea is that we want to build these encoder-centric models. We do not want to be reconstructing raw pixels. (.) And just my last point to motivate why you don’t want to be using something like Backprop. It works well for the models that it’s built for. But it does come with limitations. Many of these, up here, are biological problems. But I will just mention, very briefly, the forward and backward-locking problems. Your layers, inherently, in a correct backpropagation-based implementation that doesn’t resort to some hacky kind of error approximation, is locked. You can’t parallelize the layers. Layer 1 does not exist at all until data is fed through the bottom layer of this big feedforward network. Same with layer 2. Layer 2 does not exist without layer 1. So, you inherently must do the sequential inference operation, as one of my PhD students lovingly call these. They are just a pile of linear algebra. Then they do nothing until data is run through them. They don’t exist. Neurons exist, even when nothing is happening to them. Or at least no stimulation is going through them based on data. So, we have the forward-locking problem with just even the way we design a feedforward network. And then, of course, to go backwards to actually update the weights, we can’t update the bottom-most weight near the data until we have chained through the chain rule of calculus, an elegant but sequentially locked process that produces derivatives step-by-step, going back the exact same pathway that we took with the data going forward. This is also called the global feedback pathway problem. And brains don’t really do this, or at least the neuroscience isn’t really out to say, we do this really, really long kind of sequential process to do learning. And then it gets worse when you build a recurrent neural network. They’re not as prominent much these days, but we still use recurrence in some aspects of time series modeling. If you take a recurrent neural network, which processes data iteratively, which is a beautiful property, to train that, we basically have to take that mathematical relation and unroll it like a big, deep neural network, big, deep feedforward network, all the way back through time, basically creating a time machine, which brains don’t necessarily do. And we stretch it out and create a super deep feedforward network. And then we just do backprop over that. So it becomes even worse and scales poorly. You can do approximations, but then you induce error, and that comes with its own problems. So backprop through time is perhaps the most awful aspect of backprop-based feedforward deep learning.
S28: But we still use this. Oh yeah.
S29: And this was just to remind you, like, two years ago, I basically spent an entire talk talking about backprop’s problems and why we need to think about other things beyond it. So. Now, with this neuro SSL, again, we’re all committed here, hopefully, or at least temporarily, to saying, I’m ready to let go of backprop. I’m ready to let go of the idea of generative AI. I want to build encoders and maybe do things interesting in a compressed latent space. So there’s, like, three broad levels that my lab kind of investigates. They’re all open challenges and very difficult. So there’s plenty of space for all of you to come work. There’s plenty of problems to solve. I’m going to start actually in the middle of my diagram. So there’s, like, three broad timescales. The high behavioral level, where we work on building cognitive architectures based on these principles. And I will briefly touch on those towards the end. Then there’s the graded real valued approach, which is the idea that we’re building neuronal dynamics over time. But these are not necessarily quite as biorealistic as the lowest level, which is called spiking circuitry. And the idea is that, basically, neurons transmit these discrete action potentials or binary pulses and learn to communicate with each other through that mechanism. And that creates a different level of problems as well as benefits. And hopefully I’ll at least get through the middle, too. And very quickly just highlight some cool stuff you can do at the very top. (..) So, again, I’ll try for sake of time, so I can maybe build in some time for questions. I won’t teach you what predictive coding is. That was also done in the 2024 keynote. But just at a high level, predictive coding says we can construct an arbitrary graph, if you want to think of it like a computational graph. It is just a neural structure. And we commit ourselves approximately to two types of neurons. Neurons that are recurrent and temporal, but they don’t need to be enrolled. That’s very important. And they have, like, a temporal nature and they kind of flow with time. You can use the language of differential equations, for those that know them, to simulate how these neurons change with time. And then we have a different type of neuron called the error neuron, or the mismatch signal, which basically is responsible for saying, I take some prediction of a neuron’s state and I subtract it from that prediction. And then, of course, you can construct an objective function and you can make things look similar to backprop. And this ultimately gets back to what your morning keynote talked about, Carl, the free energy principle, which is just this whole entire thing could be said to optimize a free energy functional. But inside, when you actually build these models, you don’t actually write out, you don’t have to write out the cost function. You write these little local energy functions and out pops these error neurons, which is really cool. It’s going to actually solve all the problems. And I’ve gone at length in my body of work over the decade and as well in previous talks. It gives you the free parallelism across layers. It gives you even asynchronous computing. It gives you sparse. It gives you all these wonderful things when you build the network correctly and build it biologically faithful in some ways. And then, of course, what we have to do in the system, though, is because these neurons are temporal, you do have to spend compute on a von Neumann architecture to simulate them. So now we have to talk about even investigating data over a stimulus window. So that means now we have to use the unit of time, which a feedforward network has no notion of time. It just says, I have data and I run it through and I immediately give you an output. it’s just the chain of operations, the pile of linear algebra at work under the hood. Here, now we need to actually think about milliseconds, just like we do. When we look at something, we’re looking at over a period of time. So those differential equations are valuable, even if we have to discretize, and those come with their own issues as well. And then, basically, these error signals are used to pass messages across the network. And now we can use iterative inference to actually figure out the states of the network. And then we also get for free rules that allow us to update the synaptic connections between any two layers of neurons using Hebbian learning, or Hebbian-like learning, which has its advantages as well as disadvantages. But we’ll just go with this framework for now. And here, I’ve done in the past iterative tutorials to make you experts in predictive coding. But again, I don’t have the time to do that today. But I’m just highlighting, too, that even with these more biorealistic models, they are still a far cry from real neuroanatomy. We still use edges to kind of represent connections or projections between two neurons. This is really a scalar number representing something called a synaptic juncture. And this is where, like, chemicals are transferred between, like, an axon and dendrites across, like, a cable. But the idea is that we’re modeling this at a high level, yet we’re bringing enough biorealism that we’re already getting some really interesting benefits that Backprop can’t do easily at best. (..) And then I’m not going to dwell on this because this is prior work. But you can learn, for example, even generative AI in the context of these models and the benefits of that as well. And again, I refer you to some prior work about this. But you can learn some really interesting generative models. You can build convolutional structures and learn feature maps or image pyramids, which is a really fascinating concept. And then there’s a lot of good work in neuroscience that says this might be a potential theory of how brain structure works. The jury is still out, and there’s multiple competing theories. But this is just one possible way that you could go. (..) However, even these results, and I could go at length about all the wonderful things you can do with predictive coding in this format. Those are still decoders. So there is still an open question, or at least has been until a few years ago. How do we actually learn a predictive coding structure without a decoder? (.) And so along the way, there are other approaches, some of them not necessarily predictive coding. And I don’t know if it’ll show up. Oh, yeah, it’s over there. So Jeff Hinton proposed in 2022 a schema right before he decided that it’s time to jump on the bandwagon that backprop-based transformers are AGI and conscious and so on and so forth. But he had this really interesting work called the forward-forward algorithm, and someone might have been in communication with him in those 20 days after that paper. And another algorithm came out called the predictive forward-forward algorithm that kind of built on some of the nice principles of this encoder-only style model attaching like a decoder. So this kind of approach, I’m not going to dwell on it too much. You can read on the paper about it. It’s like an encoder-leaning architecture, where the encoder is very important, but you still have a decoder that you can generate data from and check on what’s happening. The problem with these types of approaches, while powerful, is we need contrastive learning. What is contrastive learning at a high level? It is saying, I have data, and now I need something called negative data, or confabulations, to properly use it as some of the talks have mentioned. Confabulations that are not real data, and now I can actually lower the probability that the model assigns to that bad data, or out-of-distribution data, and learn to raise the probability of correct data that’s in your data set that you actually care about. But if you do this, you can actually learn an encoder model. The only problem is that negative data, the best way to generate it is you need a generative model. You can come up with some hacky schemes, but that becomes data-dependent, and a problem in and of itself. So again, we say this is better, but it’s not quite encoder-only. (..) And again, yeah, this is me highlighting you can generate data from it. The middle is generated confabulations. The upper part are things like reconstructions. And you can learn some beautiful, if you know, TS&E, or Stochastic Neighborhood Embedding Methods. You can make these beautiful plots to show that this self-supervised scheme. By the way, no labels were used to train, for example, some of these models. You can learn clusters, which is really, really nice. Again, though, needs negative samples, so we still need a generative model. How do we get rid of that generative model in input space? Enter about a year ago, some researchers came up with a scheme called meta-representational predictive coding. We still want to use the benefits of predictive coding structure, but we want to basically do something really weird. We want to flip it upside-down. So it’s called upside-down predictive coding, you know, personally. But what do you do when you flip it upside-down? Because that means that the input to the model, or the top-most latent variable, it’s now the bottom, is the data. So that’s cool. Now I know what data is going to drive my network. But at the very, very top of this network, which used to be the data that I was learning to generate, or that model was learning to generate, I have no goal. So what do we do? Well, one thing you can do is say, maybe we can study how ocular motor dynamics actually operate. And think about, for example, perhaps, this is just one of many possible schemes. It’s not necessarily the best or the right one. You have a fovea. And in that fovea, you are looking at very high-density information, or very fine-grained information. And then you have a peripheral vision, right? The idea is that this kind of processes coarse-grained information, or a blurrier, perhaps even with no color, version of your environment that your eye is processing at any time step. And you can imagine that I’m encoding these two pieces of complementary information from, let’s say, central and peripheral vision. And now I have a learning signal, because I can encode each of these into their own respective latent space. You can imagine these are like a neural column or a cortical substructure that is processing them in a predictive coding fashion. So we’re still processing things in an unlocked way. So things are in parallel. Errors are still being passed around. Free energy is still being optimized. But now they learn to guess the state of one another. So the central vision tries to guess the state of peripheral vision and vice versa. And on top of that, they learn to predict themselves, which is really interesting. So it’s self-prediction as well as inter-stream prediction. So one thing that, you know, these researchers did was we constructed two streams or multiple streams. And you have one stream learn to predict its own statistics as well as the statistics of another stream. And then you can build an assembly or a collection of these circuits to build essentially a multi-stream model. So you don’t have to just process only one little patch of information that is very fine-grained. You could process multiple different resolutions. You can build as big of an assembly as you want. What’s also really nice is you can see in this image, again, it’s a little cartoonified. This is agnostic to the size of the image. Whereas most convolutional structures or even patch-based embedding say give me the entire image. I’ll break it up and I process it in the exact same way. If you have a high-resolution image, you either have to design a bigger architecture, spend more parameters and more energy and more money, or you need to downscale your image and lose information. This one is like your eye. It just iteratively processes this image. It’s going to do something like your ocular motor system does. And I’m kind of brushing a ton of details again for sake of time. It’s going to saccade over the image. It’s basically going to process this step-by-step and learn to build progressively a representation iteratively of the input space. So it’s really cool. I’m going to show you some little demos. So I promise you’ll get to see something interesting. There’s also other interesting work that’s happening at the same time as this predictive coding kind of flipped model called latent predictive learning or LPL, which is also a very promising framework. It kind of focuses more on one stream predicting its own values. Also very promising as well. MPC kind of consumes that as one subset. But what’s the whole premise? This is, by the way, a pure encoder model. I don’t need negative data. And on top of that, I don’t need to generate or reconstruct raw data. Everything is happening in neural space. No longer data space. I never predict an image directly. You do lose something direct. You can’t see what the model learns to pre-construct. You can’t visualize a simulation of the world. But this model will learn to represent the world in an abstract way. And then you test it later in a downstream fashion. So basically what’s happening, and again, I’ve scrubbed as much of the math from this as possible for sake of time. But there are three basic objectives. This also comes from a lot of work that Jan’s team has done as well, called invariance, variance, and covariance. And these are just three properties that we want at minimum for these neural activities to follow. Invariance just says, think of these two streams. The fovea trying to guess the, sorry, the central guessing the peripheral and vice versa. They need to be good at predicting one another from their own source of information. That’s invariance, and I want to minimize that error or that difference between them. Variance is an extra property that we add to the neural dynamics. You can do this in many ways. One way we do this is through synaptic normalization. But there are other approaches. And that just says, I want all the neurons to represent kind of somewhat different things. I want diversity built into my model. If all the neurons kind of predict the same exact value or amount to the same value given data, that’s kind of useless. So variance is an important objective. And covariance just says, I want to de-correlate the activities of the neurons. I want essentially some version of independence across the dimensions of the neurons if possible. And so we kind of optimize like a covariance matrix. You can do this in a biological fashion. That in and of itself would be an interesting talk that we don’t have time to get into. (..) So again, I have lots of different visualizations of how MPC, meta-representational predictive coding, operates. But again, you can do a couple of interesting things. Let’s just say I have one stream here. This is the central vision. I build a couple of layers. I add some neurons inside. Let’s just say you buy the premise that it’s doing the parallel processing. It really does. But, you know, it takes a while to unpack the mathematics of it. And I recommend reading the literature for this. And then you can also bias the activities of these neurons with, for example, afferent copies of the coordinates of where you’re looking at an image. So this is another really cool little thing that I’m also briefly touching on using grid cell encoding. Because grid cells kind of encode things about where you are in the world. So this model does touch on the what-where problem for those that are familiar of the binding problem of I know the content, but I need to know where it’s where it’s at in space and vice versa. So this model sort of tackles that in a very simple way. And again, you can bias the activities with these grid cell encodings of where you’re looking at an image. And then basically the idea is that this is showing prediction. So the data learns to predict the hidden layer. The hidden layer learns to predict the upper layer. And then this is generative predictive coding, the traditional model. So this is the decoder. Learns to guess the data directly and it does this at all layers. And then the other side is just basically the correction step. When we do message passing or inference, we basically just say what were the errors at all the individual layers. And we pass them across the layers until the network says I’ve settled enough over a period of several milliseconds. And we move on. And generative modeling does the same thing. It just uses the pixels directly as their reconstruction targets. (..) So again, that might not be super clear. It does help to look at some of the dynamics if you have time. But the idea is that these things process data iteratively. So they’re also constantly doing error correction, which is really nice. So yeah, and here’s the three forces that I was mentioning earlier. The way it’s implemented in meta-representational predictive coding as opposed to just plopping and attaching on different local loss functions like you would do in a backprop network is. We do have that inter and intra-stream dynamic. Neurons learn to guess their own values. They learn to guess the values of a complementary stream. What’s really cool is you can also design topologies that are somewhat biologically motivated as well. And that’s an interesting subject for future work. You also impose variance, which is just to say that synapses within a particular region must normalize to some value. This amounts to, like, a neuromodulatory signal. Or you can just say there are so many resources for synapses to consume. There’s a lot of interesting work there. You can also get rid of that, but that’s a different talk as well. And then you can also use anti-Hebbian learning in inhibition. And I won’t dwell on this. This is also a very interesting topic. But the idea is that neurons need to compete with each other to represent or compute information based on the sensory input. You get covariance, basically. The idea is you get this decorrelation effect, which is very important. You also get really sparse activity, which is also very nice. And by the way, why do we care a lot about sparsity? Well, sparsity would mean, like, if a neuron doesn’t fire on hardware, like a neuromorphic chip, if I have a zero, I don’t need to compute it. And that’s a really nice thing to do if a lot of neurons don’t need to fire every single time that data is run through it, like our generative models of today, where all the neurons fire at the same time. By the way, in neuroscience, if all your neurons fired at the exact same time in your brain, you’d be really, really dead. So your mortal computing system would cease to, because we’re mortal computers. And anyway, and then you can build assemblies of these structures. And then here was something in the appendix of that work where you can build, like, a simple, like, ring-like topology. Might not be the best one, but it was an easy one to build with a Gaussian kernel. (..) And you can build big architectures. I’m not going to spend your time on unpacking this architecture, but we can talk about it offline if you’re interested. So let’s just see what this thing does. Because, by the way, I have mentioned that while this processes pieces of information or patches of sensory input, you could also just think of them as subsets of a sensory space. It doesn’t have to be pixels of an image, but that’s the type of data we experimented with so far. This thing also learns to saccade or glimpse around an image. So it’s going to have to process an image step by step. So, yes, there is more compute being spent within a data point, but it turns out, as I’m going to show you a really nice result, you need far less data to train this model. Because it’s getting a lot more bang for its buck from each data point directly. So this is a natural image. This is one of the coolest results in the last year with this particular architecture. Usually we train on MNIST, but natural images are far harder and more interesting and more realistic. So we were using things from the NORB dataset as a nice example. So you do need to actually end up modeling this other structure I’m not going to talk about called LGN, which is just another part of your eye. The model I’ve showed you is kind of more like V1, if you know anything about your visual system. But anyway, it’s just pieces of the model. It does some really nice little processing of the input data. And then what you can have is you can kind of build a very simple form of active inference. I call this reflexive active inference, where the idea is that we use the free energy of the model as a top-down modulator to say, I need to choose where to look next, and I don’t want to do this randomly. And then I use bottom-up salience as, like, an extra force that I weight and combine it with my top-down free energy estimates. And what I’m going to show you here is how, once this model is trained after a little while, how does it process this little image of a toy tiger? (.) And so what’s really nice is, and it might be hard to see, I realize, from a distance, is you have these little blue dots that say this is where it was looking at at one point to extract some central and peripheral information. There’s actually a fovea, a paraphovea, and a peripheral subsystem, but that’s a different discussion. And this is a trajectory that it’s taken. So the number is the current trajectory. It should line up with what’s at the very top. And I have, like, a little cartoon of the eyeball just saying it’s moving around, I promise you. What’s really nice is that while the bottom-up salience doesn’t really do anything, it’s just kind of an initial driving force, the top-down error is you basically want to see this model plop its free energy estimate or its error at, like, a little Gaussian ball at a certain point. So obviously, what we would like to see is that this model processes this natural image near the actual object. It could go anywhere and waste its time. Like, this corner, it could process this, but it’s not very useful information. So this thing is going to engage in something called information foraging or some very basic, very simple, crude, epistemic planning. And the idea is that it’s going to say, I want to look at the things that are meaningful in this image without being told by the human user what to look at. So here it’s kind of processing a little bit further. And I’m not going to show you every single step because it was given about a budget of, like, 40 or 50 glimpses. But you can already see here, okay, it’s kind of, like, processing this piece of the image so far of the actual object. And then it kind of goes inside. It says, oh, there’s a lot of interesting meat inside the center of the image. So I decided to plop its free energy or error estimate in the middle. And then, of course, I’m just now going to skip to glimpse 27. You can see it build, like, this really nice little nested web. And I’ll show you what happens if you just did this purely randomly in case you’re not buying that this thing is actually staying somewhat near into the center of the image of its own accord. And then as you process further and further, you see a glimpse mesh being constructed. This is a structural pathway that was decided by the reflexive active perception module or planner of this model. So now we have, like, a little LGN and V1 for the neuroscientists in the room. There again, crude mathematical abstraction still. And then I’m just measuring some statistics to the right about, like, representational shift. But that’s not important. And you can kind of see by glimpse 40, it is kind of processed. Everything it feels is interesting about this image. Notice it never went out to the outer corners. I think this is really interesting and really cool to see the model kind of do this of its own accord. By the way, at the very top is what happens if you do it purely randomly. And then you can kind of see the structures. I have these little rings kind of seeing its field of where it could have picked. They’re a little bit more concentrated at the actual object image. I don’t know if I have… Oh yeah, this is a different little toy figurine that was in those natural images. Random, it’s all over the place. So it is better to do something like active perception or self-evidencing, as Carl mentioned in his keynote earlier this morning. What’s really cool is when you train on some real hardcore natural scenes. The Van Hatteren data set, which is actually a very interesting and complicated natural image data set. I am showing you what the synapses at the very, very bottom that live close to the input, right above LGN. These are just what’s happening to the synapses through Hebbian learning. This is local learning rules. It’s learning Gabor-like filters, which are very powerful. They’re not the whole thing that the visual system does. I know very well. It’s like the 15% that we do know. But it is really nice.
S28: And you also, this self-supervised learning system is learning clusters of its own accord.
S29: I’m very proud of the one to my most left, to your right. That’s on a natural image data set. Norb is really hard, at least for these neuroscience models. Obviously, today we use massive data sets. And, you know, future work should be investigating these on even bigger data sets. I think it’s promising and a good first step. What’s really cool, too, we did a lot of analysis on this type of model. But among the analysis that we did was just to say, What if we even took something like MNIST or NORB, and we reduce the data set size, that we’re allowed to train your more traditional deep learning model. So we took a back-propagation-based network. Yes, it’s a multi-layer perceptron. But you could do this with a convolutional network as well. And we basically ablated the data set. And we said, you’re allowed the same budget. And we did the right constraints to make sure everything was equal. And down as low as 100 data points, you can see your more traditional deep learning tank, right? Makes sense. If I don’t have lots of label data to either pay your grad students or, you know, Amazon Mechanical Turk to label all that beautiful data for you. And you rely on very little limited information in your model performance tanks. The red curve is fairly flat. I mean, it does take a little bit of a dip. It’s a little hard to see here. And you can kind of see a little bit more pronounced for the natural image data set. But it’s really impressive that actually you don’t need all that much information. Because again, these models are pulling out so much useful low-level information from these glimpses that it takes of each data point. I think this is a pretty cool result. It was actually a little surprising initially, because usually I don’t get this much benefit from neuroscience models. So I would say this works decently well. And then, of course, we looked at filters at the bottom layer across a bunch of different data sets. Also, Van Hatterin as well. That one doesn’t have labels. We couldn’t do those experiments that I mentioned. It was just unsupervised. We can also do generative modeling with this model. You can train a downstream decoder, which is often what you do. You say I freeze my encoder network after I do whatever I want with my data. And then I go ahead and I train a different model to take the latent representations of my encoder. And I just learn to map it to the data. So then you can kind of peek inside what the model learns. So I did kind of lie. You do actually get to peek inside the model. You just have to do some extra work if you’re interested. And this is kind of called downstream analysis. So you could do classification later on. You could do whatever you want with these representations. Because the self-supervised model just learns general representations of the information. So that’s really promising. And again, it’s not perfect. Again, it’s not trained to be a generative AI type of model. But the bottom rows reconstructions of NORB are even better than back propagation based JEPA. It’s an older JEPA. Of course, I’m sure Jan would say use the newest, latest, and greatest one. But, you know, JEPA doesn’t work very well if you don’t pre-train it on a copy of the internet. I’ll just briefly mention that. This model works very well without a copy of the internet. And, you know, a little fleet of GPUs to train. It just works on a little GPU. And then, of course, we beat out some other interesting baselines. And those are some numbers that tell you that it’s really good at reconstruction downstream. It’s also good at classification. This is a really cool little result. And then I’m going to glaze over the little last bit. So that way we can at least maybe sneak in a minute or two of discussion. We’ll see if I get there. You could do zero shot representation learning. Which is really interesting with this model. It was just an interesting experiment that my lab and several. You’re going to see this in a different talk as well for a different type of predictive coding. Basically, we trained it on, for example, on kanji emnis. Which is just a collection of kanji characters. And then we evaluated its representational encoding ability on digits. It was never trained on a single digit in its entire lifetime. It was just given the test set and said, do your encoding with the information that you pulled out. Now, yes, kanji characters and emnis characters share a lot of interesting similar relational information. So you obviously make sense that this data would transfer. But it’s really interesting the brown or the clusters that it decides to kind of separate out on emnis. And we also did the opposite experiment. What’s really nice is you could think of this as a potential starting point to build like an object out of distribution detection system. Because then you could say, what does the model know that it doesn’t know? And it just says, I’m not confident in what this thing is. It goes into its own little dustbin class. You do something with it. So it’s a really fun little tiny bonus result that we did. So it’s a starting point. I don’t want to like hang my hat that it solves zero shot learning. You will see some really cool results from another student that I have. And she will show you some really neat stuff tomorrow. (..) Okay. So that was all just the middle level. Now we’re just gonna really fast go through other results that I also took time to compress for you. You might say, oh, awesome. So now we’ve solved the backprop based problems. We’ve solved the encoder only problems. This is great. Is that all I need? And it’s, I’d still argue, my lab would say, well, it doesn’t spike. We need the energy efficiency. Where’s that sparsity? We might have sparsity in the dynamics, but these neurons do not communicate with pulses. Thus opens up a beautiful new space. And again, another one of my students will talk about this in the world of spiking neural nets. But I will at least motivate this a little bit. So one possible pathway to at least animal-like or human-like intelligence, perhaps, is through the mortal computing framework. Which is something that you can read the 40-page paper and synthesis about. Some beautiful ideas that could open the door to some good work. And there’s been some great talks that have far extended this older framework from a couple of years ago. But one of the nice things about mortal computing that I think is really important. That it brings to the table, as Carl would say. Is that it advocates for in-memory processing. Which is something you don’t do in a Voynorman architecture. So very briefly, we’re not going to go into the physics of a computing system. But basically, when you train your big, deep neural network, you’re putting it into a GPU, a CPU, or some processor. And you have to load the tensor of weights from lower-level aspects of memory in your computer. And every time you transfer those tensors to get to your CPU, you expend energy. Joules of energy are leaking out of your system to put that information into the CPU for then you to do something like inference or backprop. It works very well for an immortal computing system. But we’re leaking all this energy, you know, using massive carbon footprints when we train a generative pre-trained transformer. I mean, it used to be a few years ago. It’s the small footprint of a city. Now it’s greater than that. So we might want to look at other forms of computing hardware. Now I’m not going to directly espouse a particular neuromorphic chip. I’m open to the idea of anything that is an in-memory processing system. But what do I mean by in-memory processing? We want the actual calculations of these neurons to exist in the memory in which they are actually instantiated. So this kind of circumvents all this leakage of energy. Basically, the memory structure, the crossbar that houses your model, is the neural system itself. And this will get you really close. It doesn’t get you perfectly. It gets you close to the physics limits. There’s actually something called the Landauer limit about basically the energy cost of flipping or erasing a bit. And neurons can get closer to that. I’m not going to claim to get exactly at that bound. And neuromorphic chips get you closer than a von Neumann architecture ever would. So again, biological learning opens the door to these structures. Deep learning really struggles to work well in a neuromorphic system. You can do hacks and do conversions. But ultimately, you’re still burning, you know, holes in your ozone layer to transfer your model to a chip. So we basically want to do learning in a neuromorphic chip. And basically build a neuromorphic mortal computer, if possible. And so then you can do some really cool things. I also, FPGAs are a great idea, too, in these systems. And there’s been some work with the lab and collaborating with others on designing Memristor-based approaches to do this whole model that I’ve kind of explained. Versions of meta-representational predictive coding or versions of contrastive encoder-only learning. And you can get some really cool results with that. But I won’t be able to dwell on them. (..) Yeah, and I guess I was going to. I had the delusion of perhaps trying to teach you a little comp neuro on how a spiking neuron works. But I’ll leave that for offline discussions or for my brilliant student who will talk the next day about it. But yeah, you can also use spike timing-dependent plasticity to instantiate the model that I showed you as well. And you can get really, really, really fine-grained with this and more biorealistic. And that buys you better energy efficiency. And these are some old results in, like, continual learning. You can do that for free with some of the predictive coding models, but that’s a different story. (.) This was a science paper talking about encoder-only learning. (.) But again, I don’t have time to talk about it. But those are also decoders, or they had decoders inside of them. So can you just do meta-representational predictive coding in spikes? And we have some promising preliminary results in this direction. It is quite an open frontier. This is work that, again, you might have met him, Will. And I have done some really cool work on building part-hole hierarchies or encoder-like structures that use excitatory inhibitory neurons that build in accordance to an EI balance. And you can learn, actually, like, part-hole models, like, from low-level patches and learn to, like, reconstruct whole images even though it was never trained as a decoder. This is like an analysis, so I can talk to you about the details later. But you learn very, very sparse, high-level activities. And then you learn Gabor-like or at least receptive fields of different orientations of edges. I thought this was really interesting work. And it’s kind of the basis to build a spiking meta-representational predictive coding system. This is, like, literally off my laptop. I’ve also built variants of that model as I investigate. And my students also investigating this as well. You can build structures with EI dynamics. What’s really cool, when you add that version of that variance term, you get better excitatory-inhibitory balance, which is actually a measure of the health and neuroscience of a spiking neural network. I can talk offline about that, but it was an interesting result. (.) And then, to wrap up, I will blaze through, because I have about a minute and ten seconds. (.) Well, what do we do with cognitive architectures? You can take all those parts and build a cognitive architecture and solve sparse reward learning and reinforcement learning. That’s like the holy grail of, if I don’t have a learning signal until the very end of a problem, how do I learn? It’s really hard. This stuff can get you very far in that approach. And we work in robotics. So what you’ll see to the left is an initially very dumb robot. He’ll still be dumb for a while, but eventually it learns to pick up the block. And then it actually picks up the block. This is a simulation result. Look, he’s still pretty stupid. He never figures out how to pick up the block. (.) Picking up a block is really hard if you want to do it completely from scratch with sparse reward. So this was a really interesting bit of work a couple of years ago. You can do maze navigation, add a working memory to the system. That really works well. I’m going to skip to a really cool time-lapse. You can do some really complicated robotics. You can also do real-time neuro-robotics. And neuro-robotics, for anyone that knows that exact phrase, the old-fashioned definition is to say the learning happens in the hardware. You cannot train it on a cloud. Because again, we don’t train our brains externally somewhere else. They happen, the learning happens in our body. Hopefully this will play. (.) This is a week time-lapse of this model actually learning with some of our mechanisms of active inference and encoder-based only learning actually into the model. And it actually learns to eventually pick up the block. But yes, this is my student Viet actually resetting the block. It’s very painful to do real-world robotic experiments. But we wanted to show real-time in-robot platform learning. This was a really cool result. And I was very proud of my student for doing this. It took him about a week of literally sitting in front of that poor robot. And then we’re also very interested in building something that you heard from a different speaker, from John, the common model of cognition. I have many bits of work. I’ll show you some of the papers. Yeah, here. Some of them at AGI. Some of them at COGSci. Talking about how do we instantiate the common model of cognition. This is before the metacognition aspect. But building that with these little sub-circuits. How do we do that in an efficient way? So this is our visioned kind of COG engine system. We have built pieces of this kernel. It’s very hard to build an entire working system. That’s a long-term trajectory. But it is something the lab will explore over the coming decades. Hopefully building that common model of cognition.
S28: I just have results that I’m gonna skip. (..)
S29: So to wrap up, what are some future problems that we have with NeuroSSL? It isn’t perfect as much as I’m espousing how great it is. We still need to investigate what does it really do on real temporal dynamic data like videos. That’s a hard problem. And it’s just something we need to investigate. More complicated versions of control. Picking up a block is hard and interesting. But, you know, what if we did multi-tasks? Multiple tasks that stacked on one another. Or hierarchical reinforcement learning. Also, language is always a thorny thing. I can talk to you at length about how hard it really is to get a neuroscience model to work on language. That’s a real open challenge. And I’d love anyone here to solve that for us. Scaling is still not necessarily an issue. Just something that needs to be done. Predictive coding, the generative version, has this problem known as error diffusion. It’s kind of similar to the vanishing gradient problem. But recently I worked with a brilliant researcher at SingularityNet, Amir. And he and I sort of took an old idea I had in a nature paper of my own many years before, called Error Highways. And you can add these to these models. And there’s some interesting neuroscience behind it. But what they do is they solve that energy diffusion problem, or error diffusion problem, quite soundly. So now these models can scale to very, very deep. This model had like 60 plus layers, which was really cool. So finding that or also finding alternative formulations to message passing. You might not like error neurons. They are still controversial. And I will admit it’s not necessarily the jury is out. That’s exactly what the brain does. And I also do advocate for what I personally call the cognitive architecture prior or bias. I do think cognitive science offers a very useful tradition of building theories of mind. And then it kind of gives you a kind of set of different modules of how the brain works and breaks down memory and motor control and visual processing and audio processing. And then using these neuroscience tools to instantiate it. So I’m citing some older work now, but some really good ideas and the common model of cognition blueprints. I think that’s going to be important for the future of AGI. Also, we should look to real mortal computers that actually already exist. While I think mortal computing is potentially the future or one future. We have some really good work from Michael Levin’s group on xenobots and anthrobots. Those are really cool things and I will leave that for the speaker to talk about. And organoids, which is like micro physiological intelligence, which are very powerful. There might be some interesting ideas. Mortal computing and the free energy principle. Mortal computing is just the second corollary of the free energy principle. Active inference is the first. Basically, it just gives you a structure to understand that learning inference structure have a particular relationship that we can exploit. And then build like a free energy gradient flow and then it justifies everything I’ve done in this talk. (..) Yeah, that’s just proof that yes. If you want to see a lot about mortal computing, there’s like a whole keynote about that a year before. And if you want to know why Backprop is maybe something we should learn to live without for some applications, then you should listen to the other talk. I can’t thank enough my brilliant students and my wonderful collaborators across even SingularityNet, other labs. Some names you might or might not know some brilliant neuroscientists and very grateful to get to work with them on this interesting pathway and alternative to machine intelligence. (.) Yeah, and then that’s just about our lab. Just very briefly, you might say, well, I want to build these things myself in the most efficient, scalable and powerful way. NGC Learn is our lab’s tool. It’s existed for in different forms for about seven years or so. It’s changed forms, but it’s most recent form is very efficient. So you can build these and use a GPU to do your simulations. Yeah, you won’t get to use it on a neuromorphic chip. That’s a different challenge. But this should become, you know, hopefully in the next decade or so, the one tool that all neuroscience or computational neuroscience use to conduct the research. It’s very powerful. So check it out yourself. It’s free and open source. And with that, yeah, if I can, yeah. And here, I probably burnt all my time. So there’s time for questions.
S11: We do have a few minutes for questions. So, questions from the room. Over here with Daniel. Make him run as much as you can. Next person has got to be on the other side. (….)
S26: Alex, thank you for the wonderful presentation and update on your work. It’s great to have you back, and your work’s been very inspiring to me, and I know a lot of other people here at SNET at the conference. I really like, so FISA tipped me off about your meta-representational paper sometime last year, and that data efficiency plot just really jumped out at me. And I really like the concept. It completely makes sense. The question I had for you is, how do you think about, so you’ve chosen those different data streams. The way you’ve selected them to sort of, you know, comes from some knowledge. There’s an inductive prior that you’ve induced by structuring those streams in a particular way, and that leads you to get those results. Do you have a way that you’re thinking about doing that more generally, or for different data sets, like language, or video, or other types of, you know, time series applications that might be interesting to look at in the future? (..)
S29: Thank you. That’s a very good question. Yes, obviously we did lean very heavily on visual perception and ocular motor dynamics to sort of design this particular structure. We do mention very briefly in the paper, the 2025 article, that you don’t necessarily have to be tied to processing patches in an image. It’s really just more subsets of the input space. So if there is a way, we did briefly mention like audio data, if there is a way to subset out the waveforms, or maybe different resolutions of that information. If you can get complementary viewpoints, like little snapshots of whatever your data type is, this should in principle work. It is a general mechanism. It does, the model itself has a lot of these little intricate dynamics. I think the grid cell coordinates is very vision specific, because it is again, spatial location. So that might not generalize as well to other types of data, but for the multi-stream architecture, I think it’s just a matter of deciding what would be either a subset of your information that you’re able to pull, and can you pull two complementary bits of information. Images have spatial structure, so it’s a little easier for this case. I also think you could do something with resolutions. So again, if you had a way of like, for example, blurring, or let’s say looking at a global picture of some information. I think in language, that one’s really tricky. I am always a little nervous and hesitant, because I come from the recurrent neural network old fashioned era of modeling language. But I would imagine you could do something like sentence level versus phrase level or token level. As long as you have different pictures of complementary information, I think you could build a stream architecture that might work. Like, I could envision, I’ve never tested this, so I do not know if this would work well. You could have a sentence level encoder, like maybe a bag of words or a structural form of the global picture, and then you’re looking at individual tokens. And you could have at least two streams, or a couple of resolution streams of that information. You could also look into like, there was at one point things called paragraph vectors, and Jeff had his whole thought vector kind of concept. But you could imagine taking like, global snapshots. Like if you had like, a representation of a set of sentences, and then a sentence or tokens. You could imagine trying to correlate those in pieces of information together. So I do think that the framework has some generality to move in that direction. We obviously just exploited our knowledge in visual perception. But I don’t see why it wouldn’t translate. I think the only part that wouldn’t is perhaps the grid cell coordinates. Because that is very much like where you are in a space. I don’t know if that would transfer as well. (..)
S26: Awesome. Thank you.
S29: Thank you. (……..)
S17: Beautiful talk. I really enjoyed it. Thank you very much. I’m curious about your mortal computation idea. Because I recently discovered an endofunctor, or actually a monad, on the category of graph structured lambda theories. Which makes all classical models of computation mortal. And there are very natural benefits to this. But I hadn’t really connected it to your idea. So could you say a little bit more about what mortal computation means in this case. And why you are excited about it? (.)
S29: Sure. Let me try to think of the very fastest way I can say it. Because it is a long paper. (..) So, well, the mortal computation was initially inspired by two sentences that Jeff Hinton wrote in, again, that one paper in December 2022. Where he just talked about the concept that the difference between, for example, ChatGPT is that if the server that is serving ChatGPT or housing it dies, as long as you have a copy of the tensors and the code that executes that model. You can just put it on another server and it works the same, right? But in biology, in your own body, if you die, or, well, if you die, you’re done. But if I take your brain and put it in a different body, you notice differences. There’s a lot of differences. Your gut biome, everything affects your cognitive processing. So, the idea here is that the software or the calculations, and I know that’s a very tricky thing to talk about, but your cognitive processing is very grounded and tied to your body. So, what makes you mortal is that if the body dies, that cognitive functionality ceases to exist. And so, that was kind of roughly what Jeff said. I’m adding a little extra words and some biotics to it. But the idea here is that that turns into this notion that one ingredient, I wouldn’t say it’s necessary, but I don’t think it’s sufficient. And I’m not touching consciousness. That’s for the brilliant minds in this room to talk about. But for sentience, in very basic intelligence, like animal-level intelligence, you need the imperative to not die. So, mortal computation takes what Jeff Hinton says and goes to its natural extreme. Organisms do not want to disintegrate in their heat bath. They want to preserve their steady, what is it, their non-equilibrium steady state, or NES, which is prescribed by the free energy principle. Mortal computation, or the mortal computation thesis, is just a corollary, one you never knew about until a few years ago, but a corollary of the free energy principle. And it just says that we must, whatever we do, our actions are to preserve our identity. And again, there’s, like, the organizational versus structural philosophy that I’m not going to touch on, but we do talk about it in the paper. And the idea is you want to preserve that identity, and that is equivalent to disintegrating in your heat bath. So, the reason why we do epistemic foraging, and balance that with instrumental goals, is in service of that prime directive. So, the hypothesis underlying mortal computation is to say, to build your intelligence systems, the system needs to not want to die. That’s the impetus that is missing in all of the intelligence systems today, and I think that’s important to build into it. So, there’s that, and then you get this last benefit, why I mentioned neuromorphic computing, because if you want to get close to that Landauer limit, you need embodiment. So, the idea is that the calculations and software are a function in, inextricably intertwined with the biology or substrate. It doesn’t have to be biotic, but, say, it just needs to be tied to its actual neuromorphic crossbar, for example. You get in-memory processing, so you get the energy efficiency for free. That’s why brains work with only a few watts, whereas CHAGVT and transformers of the like take gigawatts. And so, the last piece I’ll end on is, it’s also another, so you might have heard of 4E cognitive theory, which is just four different ways of breaking down cognition. Mortal computation could also be called 5E cognitive theory. We added a 50 elementary cognition or basal cognition, which Michael Levin might talk about, which is just very basic things like what E. coli do to survive in like their chemical concentration gradients. You need to talk about that and build that into your system. Embodiment, which is you need a body, so you need to decide what the body is, and then design the calculations and cognitive processing to that body. You could have different bodies, different morphologies are acceptable, you just need to know what it is. Inactivism, I’m very much an inactivist, so you need to understand what niche you’re going to live in. And then your cognition and your functionality is dependent on that coupling between the body and the thermodynamic exchange between those two. You have embedded, you need other mortal computers. You need a collection of mortal computers to interact with, so we need to build a society of mortal computers. And then extended, so extended theory of mind, right, is a philosophical idea that my cell phone is an extension. It’s not mortal, but it’s an extension of me. So it’s 5E cognitive theory and another way to restate mortal computation. We just added a new E to it. So that’s as fast as I can tell you. Does that make any sense? (….) I can’t hear you, but… (…..)
S17: Leo Bus is a famous biologist at Yale. (.) He makes an argument that natural selection actually selects for mortality.
S29: I would agree with that. (15 seconds pause)
S03: As a side comment on that, after I saw that you and Greg had written papers called Mortal Computation, I went through and with the help of some of my favorite LLMs showed how they are actually the same concept of mortal computation. I mean, Greg is counting bits. You’re doing information theory. If you turn your information theory into bits, in essence, your mortal computation is kind of a special case of his mortal computation projected into certain kinds of neural computing. So that’s, it does all fit together. I think I titled the paper Immortal Computation actually, but I’m going to resist the urge to diverge under the quest for immortality. What I did want to ask was sort of a more general question, which is, I mean, I think I agree with you that predictive coding and its variations are not just more biologically founded, but in many ways fundamentally algorithmically better than this sort of backpropagation chain rule trick that has become so prevalent. And indeed should allow exploration of a much wider variety of neural architectures than those for which backpropagation tends to attractively converge, which is basically what the commercial AI industry is now exploring. But I wonder, what do you see as the key problems that need to be solved before predictive coding can sort of take its rightful role at the center of the commercial neural net industry? I mean, I understand it’s, it’s continual learning, it’s getting, getting learning to converge over very deep networks, like in the Air Highways paper that you briefly showed. So, but on a nitty gritty technical level, like what do you think are the obstacles or, or is it just more money and more people banging on scaling things up? (…)
S29: That’s an excellent question. Man, there’s many dimensions to it. I’ll try to see if I can pick the most salient ones to answer you in reverse. Yes, part of the problem is this, the resource problem. So, and as some have stated in the workshops and keynotes, backprop has a very long history of winning the hardware lottery. And there’s like some really interesting papers that talk about that. So a lot of the frameworks and hardware platforms have sort of been designed in such a way that support naturally backpropagation and gives it the edge that it’s very difficult for those that work in predictive coding sometimes to adapt that hardware backwards to say, let me do this other thing that process everything in time. So I do think that more resources funding needs to be considered to make something like this take its rightful throne as you put it, Ben. I like that. So that is an interesting problem. And I think that would help address one of the, one of the more shallower, I think, critiques of predictive coding, which is, oh, it doesn’t scale very well. A lot of people dismiss it very quickly to say, well, well, it’s not training chat GBT. Well, again, the reason is the hardware lottery and the resource problem, I think is one issue. And I think that because we haven’t been able to necessarily design the architectures that you design for like example language, which again has a very long history of writing on the success of the workhorse of backprop. So that could be something to address to say, well, if I’ve given the resources and the initiatives to do it, I think you could build. I even envision, I think that it’s possible to build foundation models with predictive coding and you would get the benefit. I think the point is to highlight the extra compute that you initially see that you do need to give back predictive coding, which is this iterative inference process is also its greatest strength. Yeah, I have to think of everything over like 100 milliseconds, like we process visual information, but that gives us natural inbuilt error correction, which I think has lots of potential. I do not know how to do design this. That’s for the experts in the room to build like causal chains of reasoning, a higher order reasoning. I think I think the error correction mechanism, the internal recurrence is its greatest power that hasn’t yet been studied properly. Because again, oftentimes we are a little scared of that iterative inference problem. So I think that brings me to, I think, a place where I have personally seen predictive coding struggle the most, which is unfortunate. Vision, it’s very good at visual processing, continuous inputs. It’s discrete inputs like tokens and sentences. You can make it work well enough, at least for my group. I’m not speaking for other groups. And many, many, many years ago, I was telling a story to some singular net engineers. I worked on recurrent neural net language modeling. And while predictive coding was like the best of all the neuroscience alternatives to building a recurrent language model, in Pentree Bank, which was the famous data set you used in the old days, you couldn’t quite beat out the perplexity of like a long short-term memory train with backprop through time. You were the second best always. But LSTMs were always better. On language specifically. Video modeling, it was the reverse. It was predictive coding was the better generative model. And backprop struggled because, again, you need to do unrolling, and you have the vanishing gradient problem, and all the millions of issues that plague backprop. So language was interesting, and it’s not clear to me exactly why it should work well on discrete tokens, but maybe there’s just something inherent about the iterative processing that just needs something extra magical from researchers to think about what mechanisms might be missing. At least that was the old days. I don’t know if that holds anymore today. I’ve heard now about the transformers. I haven’t necessarily tried to build a transformer with predictive coding per se. And to wrap up, I think the, I wouldn’t, I don’t know if they’re called roadblocks, but I think the most important space to be working in, I might be biased, is reinforcement learning control. Because I think that some of the most promising results my group has gotten over the years has been with very controlled narrow robotic experiments. We’re obsessed with sparse rewards. We think that’s the problem that’s most important to solve. There’s a lot of sparse reward problems that animals solve. But it’s also some of the most difficult aspects to make predictive coding work reliably and stably. You can do it in controlled environments. I think going forward, we need more people working on that. So I think that’s one other issue. I think that’s the most important space to be thinking about, because that’s where active inference really comes into play, where predictive coding gives you this natural, like, epistemic term. And you can kind of balance that. So I think that’s another one. But I think temporal sequences and language are going to be the biggest challenge in the near future. But yeah, so all the way back up, I think it is a resource and funding problem. Because again, I think Yoshua put it well, you know, backprop works well enough. And so there might be things that do actually work pretty darn well and sometimes better. I’m biased in that regard. But we don’t see any reason to leave behind backprop yet. I think the last key to it is the fact that predictive coding also has this other thing that we didn’t really talk about, where the reason why it works at the spike level is it doesn’t require everything to be fully differentiable, which backprop is calculus. everything must be fully differentiable or it doesn’t work as well. And so that could allow you to train discrete systems. And that’s where I think its greatest power lies. But we need exploration.
S03: Yeah, let me, two very quick follow-ups on that. So Faiza Habibi, your student who’s now working with us, (.) going to give a talk tomorrow, I think. But she is of the opinion that something fancy needs to be done to propagate prediction through the attention mechanism in transformers. And we’ve been going back and forth on that a bit. So I guess her view is it’s not the predictive coding doesn’t fly there, but the attention mechanism has been the key to getting deep neural nets work well for NLP. And maybe the way you have to think about prediction and attention is a little bit different than in vision cases. On the GPU thing, I’ve been thinking about this. And I’ve convinced myself you could speed up PC by 20 or 30 times using current GPUs just by doing stuff like taking multiple cycles of settling in and just compiling them into one big matrix operation and so on. So it’s, it’s true the current GPU CUDA stack isn’t naturally suited for something as beautifully asynchronous as PC. And you, of course, neuromorphic hardware would be, would be better. But I think one could get a significant speed up just by doing different CUDA libraries basically. And again, it’s just NVIDIA hasn’t, hasn’t bothered up until we give them a reason to, right?
S29: Exactly. (……) I guess I’ll just brief if you, I can take like five seconds. I also think another roadblock is weirdly enough, uh, uh, a pedagogical or technical aspect of predictive coding. I think oftentimes neuroscience models get this bad rap of like, well, it’s really complicated. Like I said, differential equations and teaching students over the years, that’s the easiest way to scare them away. Um, and I think perhaps the conveying of the dynamics, they’re actually not that bad to build like a predictive coding circuit. But I think oftentimes these are couched in Bayesianism, which is beautiful. Um, but that can also prevent democratization of this. So I think that’s another issue that you, once you see the mechanics of it and you see, hey, I’m minimizing or optimizing free energy. And it’s just iterating over a differential equation and this is not that bad. Um, I think that’s gonna be another roadblock to overcome in the future is educational. (..)
S28: Yep.
S29: Uh, and tools like Fabric PC and NGC Learn can help make it more publicly accessible, I think. So we need, we need more of that in this idea that, uh, the beautiful formalism isn’t always the initial, uh, what is it, entry point into predictive coding. (…)
S11: Thank you very much, Alex. (..) Thank you. (…) Definitely been a theme with the active inference, um, workshop yesterday. Predictive coding, of course, by Alex today. Uh, a couple more presentations tomorrow, Carl, this morning. Uh, so very exciting. And now we get to dive into some papers that are, are looking in these, uh, similar directions. Adaptation memory and the cost of intelligence across time. We’ll start with Howard Schneider and then Daniel McDonald, Gabriel Axel Montes, and Elijah Perrier, um, two of which will be recorded. Howard, please come up. Improving long horizon task, task completion in a proto-AGI cognitive architecture, a memory-centric approach. (..)
S30: How do I advance that? (11 seconds pause) //S30: Okay, forward.// And I can see part of the slide on, on the screen.
S11: Would it be great?
S30: If not, I can just look at the screen. (.)
S11: Sure will do. (.)
S30: Thank you so much. That’s great. (..)
S22: Okay. (..)
S30: Um, hi, I’m Howard Schneider from Toronto, Canada. I’ll be talking on improving long horizon task completion. (…) Um, before I jump into my experiments and work, I want to go over two concepts. First thing, my work was done on a cognitive architecture, the Causal Cognitive Architecture 8, CC8. It’s referenced as ECA in the paper. (..) I used the CC8 to simulate the brain of a neonatal goat through the life cycle. So, again, there’s a cognitive architecture. I’m simulating the brain of a goat. (…) This is my cognitive architecture. It makes heavily use of what I call navigation maps. These are spatial maps. (..) Information comes in. The goat can see things, hear things, touch things. Comes in. One, it gets mapped onto these spatial maps. There’s processing. (.) There are instinctive primitives, learned primitives. These are procedures that will act on the process map. An action may result. Three, and then the goat will move its leg, for example. And then this repeats. This is a cognitive cycle, and it repeats, and it repeats, and repeats. So, cognitive architecture, cognitive cycle. (..) Why navigation maps? It’s ubiquitous to animal life. Over a half billion years ago, you had directed locomotion. (.) 518 million years ago, the first or before early vertebrates with brains that would give rise to hippocampus. If anybody wants slides, just email me. I’m going to go a little bit quicker. (.) The cognitive, the navigation map, essentially, it’s an array. I model them as arrays. Here you see 2D one, although I use more dimensions. (..) Why a goat? Part of a broader research program. A goat, the advantage is there’s precocious development. I don’t have to deal with a lot of locomotion issues. The flip side is building better AI, AGI. (…..) A lot of different computational paradigms that we’ve talked about all day long. Recently, Professor Robbie actually discussed a whole bunch of different ones here also. I looked at the mammalian brain and picked the things that I thought were making it work. It’s a distinctive paradigm, but it’s a lot of overlap with predictive coding, a lot of overlap with symbolic, a lot of overlap with active inference. (..) And one last note, I actually have an LLM in the architecture. There’s a link to my source code at the end of the paper, and there’s actually the Python code for a link to the OpenAI LLM, (.) but it’s not being used in this experiment, so no LLM involved. (..) Second concept, long horizon problem. A lot of you may already know what this is. For those that don’t, I’ll quickly explain it. I think, actually, this is probably the most important concept of the week. Not my work, but the concept. And I think this is the end game to AGI. If you use an advanced frontier model now, it’s amazing the work you can get, like you can create a paragraph that any person, you can create a block of code, actually written better than most coders could do. But the problem is, if you want to create a new paper, something new that it hasn’t seen before, or you want to replace somebody for using the LLM to do coding, it doesn’t work out. And the reason is, it’s the long horizon problem. There’s an issue with going from all these small, fantastic work at these small steps, actually finishing a whole task. (..) Again, long horizon problem for those, it’s, you lose the thread. You can do each little step okay, but it doesn’t work out over the overall thread. For example, let’s say we have 100 small tasks to do, each has 10 subtasks. If each one is 99% accurate, the problem is by the time there’s only a .004% chance you’re going to finish the task. We’re able to do tasks quite good. We do correct. We repair. (..) Many reasons why long horizon tasks fail. I just mentioned compounding errors, delayed credit assignment, exploding search spaces. And there’s this insidious partial observability. It happens all the time. People often don’t think about it. When you start simulating the real world, like a goat in the real world, it becomes apparent. (..) It creates state gaps. State being things, you know, the goat may want to know, where’s my mother? Where’s this? Where’s that? You start getting corruption of state. The premise of this paper is that by repairing state gaps, we will improve long horizon performance. I’m going to show you an innovative architectural approach to repairing it. I’m not using the typical palm dip belief state updating. (..) Okay, I’m going to show you three repair mechanisms to compare them in the work I did. (..) Keeping, trying to do good experiments, holding the same conditions for each of them. The task is the goat is on the ground. It’s got to stand up. It’s got to figure out where its mother is. It’s got to find it. It’s got to be able to nurse. And it’s got to stay somewhere safe so it doesn’t fall off a cliff or get eaten up. And a number of sub-tacks involved. A nice little toy long horizon problem to consider. I induce a bit more partial observability for experimental purposes. Everything goes through the same conditions. Three mechanisms. Guarded merge. And this mechanism for state repair, again, different than using the palm dip belief updating. All those similar things occur. I’m looking. It looks for a similar episode. There are some states that are corrupted or missing. It looks for a similar episode. I’m matching navigation maps. I’m matching maps. (.) Friston had a slide about matching maps, actually, this morning and comparing them. I’m matching maps and I find a similar map. I read back the state information and I, in a very guarded way, I fill in the slots. Replacement style. I read back, but I’m not as, I assume that there’s more corruption of the states and I will replace a bit more of them. And storage only. There’s no contextual read back. And then 100 runs of each were done. (.) With guarded merge, even though I’m inducing partial observability, the goat managed to go through the long horizon task 100%. With no read back, surprisingly, 95% of the time it was able, it had to start again. It was able to, able to manage, although it took longer. And this was a toy model. If you expand it, the amount of time starts shooting up and it starts becoming a bit of a problem when you have like a thousand steps. Replacing state. If it’s not guarded, if you’re not as careful what you’re replacing in this mechanism, it wasn’t as good. So, guarded merge. The mechanism I described, the repair mechanism I described to you, improves long horizon tasks. And it seems to be a valid mechanism that works pretty good. I believe it works besides this little goat, besides like the simulation of a neonatal goat that I did here. I believe it probably works, probably works on us and a real goat. We repair state. We don’t, we’re able to do long tasks. It would probably also work on LLMs. It would probably work on other systems of intelligence, i.e. the algorithm, the architectural approach shown. In terms of safety implications, if you want to do interesting things with AI, AGI, you’re going to need long horizon. That makes it, this is the end game of AGI, I honestly believe. Once you have long horizon, you have AGI with what we have now. But it will amplify misaligned trajectories. Within my architecture, within the CC8, I have, you know, intrinsic safety built in. I have their state coherence, which is what I showed. There’s goal coherence to prevent the goal from drifting. And there’s authority coherence, what’s allowed to do. And the reason I have it, it’s not for working on the safety problem or preventing terminator. It’s for a very pragmatic reason. I’m designing the brain of a goat. Even though they’re great at climbing cliffs, gravity still applies. And I don’t want the goat to perish. (.) Safety is a good thing to have within the architecture. (.) Last slide. I took the entire architecture. I showed you the CC8. I put it into a kernel. I call it ARCOS, Robotic Cognitive Operating System. If you think about it, there have been these amazing robot demos the last few years. This is the Unitree G1. It’s been out for a while. These robots are amazing. They can do somersaults, backflips. Yet, good luck trying to get them to do useful work. It’s really, really hard. And why is that? And again, the answer which you probably know I’m going to say, long horizon problem. This is a layer in the robotic software stack. It’s above ROS2. What it does, it repairs state. So, the idea is that by repairing state, you get long horizon and you get a better performance from the robot. I’m going to be presenting this in September at the BIKA Biologically Inspired Cognitive Architecture Conference. (..) It’s done with time, a few seconds to go. So, I showed you the long horizon problem. Multiple causes. We isolated partial observability here. I showed you a distinct… There’s… There’s… In the prior ARC, there are other repair mechanisms. This is a distinctive method, which I call guarded episodic read back stake repair. I believe it works on an algorithmic level on other systems of intelligence. We are over the safety principles and ARCOS for future applications. And if you want slides or information, my email’s there. Thank you very much. (10 seconds pause)
S11: Our presenter, so timely this year. Very much appreciated. And, of course, still getting all the critical and important information in. so that we can take it all in. Next, we have Daniel McDonald, Resource-Bounded Incremental Induction, A Principled Framework for Continual Learning. All right.
S26: Thanks, everyone, for being here. I’m going to talk today about an exploration and a model and some intuition that I’ve encapsulated in a paper. And it’s about resource-bounded incremental induction. So, but before I talk about that, let me take a step back. So, the motivation for this is to look for principles of continued learning. Now, why look for principles for continued learning? Well, there’s a couple of pet peeves of mine. Catastrophic forgetting was first identified as a problem by Grossman about 50 years ago, or actually, I think, 50 years ago this year. (.) Existing approaches either don’t scale, require historic priors, or they’re more troublingly incompatible with open engine AGI type deployments. So, if you want to have a system that does things that we think of and is familiar from our experience of continued learning, most of what you see in literature just won’t scale that way. (.) Economically, LLMs would be radically more useful for AGI if they could learn incrementally in a scalable way. I’d love to get LMs that I could plug into hyperon components and amortize inference and do all kinds of interesting things. But they’re not really geared for that right now. (.) And, as I mentioned, our everyday experience just is widely divergent from what happens when you actually try to update one of these systems. (.) So, taking a step back to look for principles, let’s frame the problem as online sequence prediction. So, online sequence prediction, you’re observing a stream, you’re outputting a prediction, you’re getting some feedback of the next thing, and you’re updating some state and some stateful system. But crucially, there’s no task labels and the distribution shifts can be completely arbitrary. And this is the situation that humans and hopefully our AGI systems deal with quite naturally. In the literature, you see things like this. Well, I’m going to benchmark my thing on a couple of different regimes and they’re kind of tidy. But reality is way more complicated than that. So, what captures the intuition here? And the intuition that I distill out of the paper formally, but I’m not going to get it to here, is that continue learning means reusing previously discovered regularity when it recurs. And I mean that in most general possible algorithmic sense. (..) So, starting from that framing, you know, we all are friends with Solominoff’s induction scheme. Solominoff’s induction scheme gives a directly optimal way if you have infinite resources. Well, you can, I’m sorry. It’s not actually computable, but you can do pretty good approximations if you have vast or infinite compute resources. But, of course, we’re not interested in that. We’re interested in things that comport with the resource bounds that humans have and that embodied AGI systems would presumably have. So, the idea here is just start lopping off the pieces that we think we can, but still keep it as close in spirit as possible. So, we’ll restrict the infinite sum over programs to a finite pool. We’ll make it a pool. We’ll kick out programs that suck. And we’ll keep a universal schedule. So, we’re still searching program space in a certain way. And we’re using that to kind of replenish the pool when, you know, things shift and don’t work. We’re restricting the coherence of the pool to, instead of being, you know, the entire history of the string, just like a recency window around it. And then we’re going to let those programs do something interesting where they can refer to some other state that could be all of the previous history or some other, you know, reduction or projection of that. We’re going to store those successful programs and we’re going to allow them to be reused. (..) Schematically, it’s kind of what it looks like. You’ve got your predictor, your pool of predictors. You’ve got your two different types of state. One of them is just a state that’s encoding the history in some way that you can query. (.) Another mechanism of state is this S of T. That’s the frozen program store. And then you have a search process which is interacting with that state. The predictors are interacting with the state. And it all comports with this sort of nominal resource bound here, which doesn’t have to be finite. But it’s just some bound that you want to use to characterize your system. (……..) I pressed something wrong. (.) Help. (……) Ah, great. So, all right. So, expanding what I mean by reuse and what this scheme sort of looks like schematically. So, if you’ve got a system like this, you’re seeing a sequence of regimes. The behavior that we’re looking for is really that, you know, initially, the predictive loss of the system is going to be high when it sees something novel. But it’s going to be lower when those things return, when the same, you know, regimes of predictability return. And the time that the system is going to take to get back on track, so to speak, is going to be decreased. In the paper, I sort of give a formalism for this. So, there’s an operational description length, which is related to the search cost and, you know, the encoding length of a particular transformer object, which is something that, you know, can transform past programs into new ones or create new programs from scratch. And there’s discovery time, which is related exponentially to that length. So, what that means is that reuseful programs mean that that operational description length is small now. It means you can actually find them. Like, you can dredge them up for your memory and actually construct hypotheses, which actually helps you do your thing. (.) But there’s now immediately a problem. And the problem is that if you aren’t careful about how you store things, if you just want to store everything, every hypothesis that you ever had that was good, well, that store is going to grow with time. And if it grows linear with time, then the time it takes you to recall anything, you know, can grow without bound, unless something else happens. So, I call this phenomenon the addressing barrier. And what, and this is sort of a very colorful graphic, which is the whole motivation for this to be black, and I’m sorry that the lighting isn’t good for those of you who may not be able to see it. But basically, you know, on the left here, you have just this growing store of programs. The search mass gets diluted if you don’t have some other, some other prior. And on the right is sort of like what happens if you, say, introduce some content addressable index, which is going to let you, you know, easily surface programs that are relevant to your current context without going through the search cost of paying the addressing bits to find them in this vast history of things. (…) So, breaking the addressing barrier is sort of this interesting insider. It was very helpful for me because the frustration with, you know, spending many years trying to build continued learning systems that were practical is that there’s always a point where heuristics seemed to be necessary, but as soon as they came up, the architecture become coupled and brittle, and they weren’t robust in some annoying way. This perspective has allowed us to come full circle and say, like, okay, reshipping the addressing landscape is actually necessary. It’s a necessary constraint that falls out of just the fact that we’re doing resource-mounted induction, and it reframes, at least in my brain, the role of heuristics in continuing learning. They’re actually central, but what’s crucial is that now we know how to combine them and evaluate them in a way that’s more principled than the typical just ad hoc approach that has, you know, sort of prevailed in the last, you know, several decades. (.) So, some takeaways, applications of this framework so far. As I mentioned, it’s a framework to integrate known useful heuristics with kind of a common principled glue. So, you can sort of, you could take this all the way to solving some neurosymbolic integration problems. You can also, you know, use this to do basic things like, well, I know SGD with IID data is a great heuristic, works for lots of things. But, you know, how do I build an architecture that leverages that without coupling that constraint through everything and becoming a mess? (.) Another, you know, interesting path is application to high-level cognitive orchestration. So, you can apply this RB framework to jittering short codes and those short codes, you know, encode attentional patterns. You know, maybe ECAN or, you know, amortization of inference control in PLN or some other type of high-level, you know, orchestrating coding agents. That’s another way that you can look at this. (.) And it also opens up a fresh perspective on existing methods. We’ve been talking a lot about predictive coding today. That’s sort of, you know, where I kind of started and what the, where this work kind of came out of as my frustration trying to develop predictive coding heuristics that were effective. (.) But the inference phase of PC, as Alex alluded to, opens up some interesting aspects of a design space. And this was my initial intuition. But now I want to look at that as a form of addressing. It actually lets the system spend some budget to go figure out somewhere in its memory, in some sense, how to recall useful information in this, you know, open-ended kind of context. And finally, you know, an open hypothesis and one of the things that I’m really excited about with this framework is that, you know, borrowing some techniques from program synthesis. You know, program synthesis lineages, such as Dreamcoder is a good example. You can do things like amortization and library compression and dreaming to speed up this, this sort of universal search component. But you can also use that to mine for useful heuristics. So you can let this thing run on some long sequence of data that you don’t understand very well. And then look what it does and then pull those things out. And those things are now the heuristics which allow it to continually learn efficiently over that dataset. So I think we’re right on time. So I will leave you with our friend the Salaman Octopus. (20 seconds pause)
S11: Thank you very much, Daniel.
S18: Thanks, everyone. Gabriel.
S11: Axel Montes here. All right. And our third presentation in this paper session will be recorded. So Gabriel Axel Montes with trace fields as externalized memory, a stigmatic multi-agent simulation of distributed control and trace ecology governance. (.) Gabe is here, but he chose to use a prerecorded thing. So we all get to guess if it’s an AI or Gabe giving the presentation as soon as it gets pulled up. (12 seconds pause)
S18: Thanks, everyone. Gabriel Axel Montes here. This work asks a simple question with, I think, some awkward consequences for how we might think about alignment. (.) When memory no longer is living inside an agent and it starts to live in the shared environment that the agent acts in, what happens to control? So I’ll walk through a small simulator that lets us look at this directly. (…) Here’s a default picture. An agent as a bounded box. It’s own memory. It’s own centralized control planning inside that boundary. (.) But the systems we’re building have started to leak out of the box. A capable assistant today reads from a retrieval store. It writes logs and scratch pads. It leaves tool outputs behind. It shares a workspace with other agents. Memory and control are drifting into the environment. So here’s a question. If memory lives in a shared environment, what happens to control? I want to treat that environment as an actual design object or target in its own right with real benefits, its own ways of failing and its own levers. (…) So the term that everything hangs on here is a trace. A trace is a persistent reward weighted signal written into a shared substrate as a side effect of acting decaying over time, biasing later action by other agents. This is essentially stigmerging, the idea from social insects. An ant lays pheromone, the trail persists. Later ants follow the gradient with no central coordination. (.) The environment stores a record of what worked and that record steers what happens next. So here’s a simulator, a 12 by 12 grid, 24 agents. Every cell is neutral, a resource or hazard laid out in clustered patches rather than scattered randomly. So there’s real spatial structure to learn. (..) Each agent holds a belief per cell updated by Beijing conditioning on noisy local observations. It scores candidate moves on reward, uncertainty, information gain, trace bias and crowding, then samples from a softmax. The trace substrate on the left has three channels decaying at different rates, fast, medium and slow. So agents deposit reward weighted trace as they visit cells. That’s the yellow buildup on the resource patches. And it’s the externalized memory itself, past behavior reshaping the landscape future behavior confronts. So four experiment families sit around this model, baseline and gain matched, ablation, decoy overlay and governance. So the first result, every trace enabled channel set beats no trace on reward and hazard. Reward climbs from about 0.77 to above 0.91. (.) So, but, but nobody wins outright. So fast plus medium has the best reward and fastest recovery at 18 steps. All three channels together is worse on both and watch coverage here. So no trace covers 93% of the grid. Every trace set up covers less around 73. (..) Trace mass piling on cells that already paid off pulls agents back to a narrow band instead of exploring. The best steady state design and the most adaptable one are different designs. (..) Second result, probably the sharpest one for alignment. What happens when the world moves and the memory does not. (.) So I freeze trace on resource cells right before perturbation, even as those cells stop being resources. Then watch where agents go. They go straight to the dead cells. So for fast plus medium, the decoy resource trace ratio jumps from 0.18 to 1.7. For all three channels to 4.5, the largest swing in the panel. Recovery slows from 18 steps to 40. And the no trace baseline doesn’t move at all on these measures. So the pull towards stale structure runs through the trace itself, not the perturbation. (…) Third result, you can fix this by touching the substrate rather than the agents, maybe. So two interventions. (..) Both are acting only on the trace field. Adaptive decay ages traces out faster. Targeted reset clears stale cells directly. Ungoverned fast plus medium sits at 0.865 reward. 40 step recovery. Adaptive decay brings recovery to 20. And targeted reset brings it to 3. One caveat, though, is that targeted reset is strongest partly because it’s handed exact knowledge of which cells went stale. Which a real system rarely gets for free. So, but the no trace baseline stays flat throughout. So the effect runs through the trace. (…) The mapping to real systems is closer than it looks. In the simulator, stale trace, adapted decay, targeted reset. In a deployed multi-agent system, you get a shared vector store, surfacing outdated context. Relevant scoring that downweights old entries. And context invalidation. (..) Same shape, several agents working, several agents writing to one persistent medium. Later decisions writing on what earlier ones left behind. What the simulator adds here is a control comparison that isolates the trace effect on its own. (…) A few limitations by design. So this is a toy grid, not a deployment. The policy is inspired by active inference. It’s not a full formal version. Stale traces are injected, not endogenous. Targeted reset assumes information a real system rarely has. (.) So that points to what’s next. (.) Letting stale traces emerge on their own. Retrogenerative models and a full active inference implementation. Continuous environments. Mixed agent populations. (.) Learned trace policies. Larger distributed systems. Links to morphology derived observables. And tests against real retrieval augmented systems. And comparison against biology. So implications and takeaway. I would say. If generally intelligent systems come to remember through shared external substrates. And it looks like they will. Part of alignment is how that shared memory gets written. How it ages and when it gets reset. (.) For AGI. (.) Think of tool ecologies and retrieval layers. Are becoming standard architecture. Not an edge case. And how the substrate is written. Aged and reset belongs on the control surface. Alongside any single model’s objectives. (.) For complex systems more broadly. This pattern extends well past AI. So the same benefit. Liability and governability show up. Wherever distributed systems keep persistent shared memory. Insect colonies. Slime mold. Robot swarms. Smart cities possibly. And now AI tool ecologies. So the mechanism. Looks architectural. And it shows up the same way. Rather. Whether the substrate. Is chemical. Physical. Or digital. And so on. (.) So the code. And the runs. Are on GitHub. And for any more detail. You can really peruse through the paper. I’m happy to take any feedback. Thank you. (…) Thanks everyone. Gabriel Axel Montes here. This work. Asks a simple question. With I think. Some awkward consequences. How we might think. About alignment. When memory. No longer is living. Inside an agent. And it starts to live. In the shared environment. That the agent acts in. What happens to control. So I’ll walk through. A small simulator. That lets us look at this. Directly. (…) Here’s a default picture. An agent. As a bounded by. (12 seconds pause)
S11: The giveaways. Was the nose track real. (..) Thank you everybody. So next. We’ve got. //S38: Unfortunately another virtual. Elijah Perrier. And this is. (.)// But he’s going to talk about. Watts per intelligence. Part two. Algorithmic catalysis. (18 seconds pause)
S38: And apologies. (..) I’m not able to be there. So this talk today. Is about. (.) Dealing with this issue. That we’re facing. Towards AGI. Which is we need. Energy efficient. Intelligence. The race to build. Intelligence systems. Is facing. The realities of physics. Scaling intelligence. By brute force. Becoming increasingly. Expensive. More intelligence. Currently. Means more computation. And more computation. Means more energy. Cooling. And cost. And so. Future of AGI. I argue. Need new principles. For energy efficient. Intelligence. Not simply. Large models. So the question is. (.) What lets a system. Solve intelligent tasks. Using less physical work. (..) In earlier. Work of mine. I provide. A way. To try to observe. (.) And measure this. It’s called. Watts per intelligence. It’s effectively. A ratio. Energy power. Used. Per. Unit of intelligence. (..) And the idea. Is that. Certain tasks. In terms of intelligence. Such as. Compression. Selection. And resetting. Involve what are. What are. Logically. Irreversible. This then. Incurs an energy cost. Landau’s principle. (..) And that we can therefore. Measure. Effectively. A ratio. Of the amount of energy. It takes. For a particular task. Based on. The number of. Irreversible. Operations. (.) And. The point is. That more. Irreversible. Lead to a higher. Thermodynamic cost. In principle. (.) Certain tasks. Less accessible. So. Reducing irreversible. Operations. Reduces. The thermodynamic cost. For achieving. The same intelligence score. So. We’re looking for. Ways. To reduce. The thermodynamic cost. By lowering. The number. Irreversible. operations. (..)
S06: So. What does nature do. To solve this problem.
S38: Well. Nature. Uses. Catalyst. (.) Catalyst. Nature’s. Way of making. Otherwise. Energetically. Inaccessible. Transformations. Available. (.) Catalyst. Because it’s. Physical structure. And geometry. And shape. (.) In effect. Allow. Transformations. Which would be. Energetically. Unfavorable. So. How does it do this. It’s geometry. And properties. Bind. And orient. The particular. Molecular. Configurations. (.) This effectively. You know. Encodes. Informational. Structure. In the molecule. And what it does. Is creates. A low energy. Pathway. Of intermediate states. Through which. A. Reaction. Can proceed. Life. Would not exist. And intelligence. Would not exist. If it wasn’t for. The existence. (.) Of catalysis. Because that structure. Persists. And fits only. Certain substrates. It’s reusable. And selective. So the question is. Can computational structures. Similarly. Determine. Intelligent. (….) In this. Paper. I argue. Yes. That. That. I propose. This principle. Of algorithmic. Catalysis. The idea. That. Reusable. Computational. Structures. Which encode. Task. Class. Class. (.) Regularities. Effectively. Shared abstract. Structure. Can. Reduce. The number. Irreversible. Across. Intelligent. Tasks. By. Doing so. They can. Allow. A pathway. To perform. The particular. (.) They can. Be designed. In a way. That. Can be. Reused. And they. Fundamentally. Are designed. In a way. That can be. Tailored. To a particular. Task. These are. These three. Principles. The claim. That I argue. In the papers. Algorithmic. Catalysis. Provides. A physical principle. Understanding. And designing. Energy. Efficient. Intelligence. (..) Now. There. Existing. Computer. Science. A field. Catalytic. Computing. Which some of you. May have heard of. Or have done research in. The idea. Of catalytic computing. Is the. Is the concept. Of using a computational state. Without consuming it. (.) This area. Doesn’t really consider. The energy efficiency. Of doing so. So in catalytic computing. You may. Take a small space. Machine. And it’s given access. To a much larger. Auxiliary. Or ancillary. Tape. Whose initial contents. May be arbitrary. That machine. Can effectively. Use that tape. During the computation. As a scratch pad. But it has to. Restore it. To its initial state. Before terminating. (.) To these concepts. In this paper. I add the idea. Of task class structure. Adaptation cost. (.) Restoration cost. The idea. Is to turn. Reusable memory. Into a model. Energy efficient. Intelligent computation. (…) So what does. An algorithmic catalyst. (.) The catalyst. Encodes. Reusable. Task class. Regularities. That’s jargon. For information. Which essentially. Helps you solve. Problem. In a way. That helps. Constrain. Computation. So the idea. Of a catalyst. Is it actually. Eliminates. Unnecessary. Computational. Pathways. And thereby. Reducing the number. Of irreversible. Operations. The gain of. Using. Catalysis. As a principle. Comes from. Avoiding. Particular. Computations. I’ve tried to. (.) Vibe. Vibe. Code this. In two diagrams. Here. We have a. Generic solver. (.) Essentially. Brute force. And then. We have a catalyse. This is a particular. (..) Use of an algorithmic. Catalyst. (.) Makes. Solving problems. More efficient. Now. People will say. Well. How is this different from. Just using an ordinary. You know. Clever algorithm. The point is that. The. Catalytic. Form of. (.) Sigma C here. Tries to do so by. Reducing. Irreversible operations. And that’s the new thing here. And again. So. Sigma C. Is the catalyst. It describes. (.) Regularities. Shared across a particular. Task class. Task class. I’m a physicist. So I’m into symmetries. For example. (.) The encoded. Structure of the catalyst. Determines which. Successor states. Are compatible. (..) Regularities. (.) Incompatible branches. Are discovered. Before they. Are explored. And. This provides. In principle. A way. For designing. A. A tool. By which. We may. Improve. (.) Solving certain tasks. In a more. Energy efficient way. Of course. One of the challenges. Is exactly how you design the catalyst. Which I don’t really. You know. Address in this paper. So future work. (.) Um. What are the results. Of the paper. I’ll just sort of talk through briefly. So the first one is. Well. How much can. Catalysis. Help. And so. The sort of aim of this paper. Is really to plant a few. Early flags. In the landscape. As to. You know. What the speed up. Produced by a catalyst. May actually. Find us. The first result. Is that. It’s bounded. By how much information. It contains. About the task class. So the speed up. From. Catalysis here. So. Gamma. T. Is the reduction. In irreversible operations. (.) On a task. Sigma. C. Is the. Is the catalytic. Description. Of the abstract structure. Across the task class. (..) Fundamentally. (…) A catalyst cannot. Exploit. Regularities. And it does not encode. And so. One example. I talk about. In the paper. Is an F18. SAT. While encoding. The F18. Subspace. Reduces search. From 2n. To 2d. This is well known. But the sort of point. In the paper. Is to show. How. This known. Result. Can be framed. In terms of algorithm. Catalysis. (.) Thermodynamic. No energy efficiency. So a lookup table. Is not catalytic. It stores all the answers. But it does not encode. The rule. Needed to solve. The next. Unseen instance. The idea is that. A catalyst. Does this. And it does it. In an energy. Efficient way. (..) The second result. Of the paper. Is really. Around this question. What has. Structure cost. So encoding task. Structure. Requires a. Minimum. Adaptation energy. I.e. The cost. For the structure. To be. Physically encoded. In the catalyst. A adapt. Is that cost. And this formula here. Again. I can. Direct you to the paper. Really. Sort of has. Basically. Tries to. Show. That there is an upper bound. On the adaptation energy. (..) Sorry. The. Minimum. Adaptation. Sorry. (..) And. And. Sort of. (..) If. The information. And the structure. Is already present. Then. The adaptation. Can reuse it. Otherwise. Has to be physically. Written into the catalyst. Kind of obvious. But it just. Makes the point. (…) And then. The third main result. To the paper. Is. When is it worthwhile. (..) Using. Catalysis. While a catalyst. Energetically worthwhile. Only when the catalytic. Speed up gamma. Is worth the cost. Of constructing it. Restoring it. In irreversible operations. N star. (.) So a trained model. Database index. Or compiler. Is worthwhile. If it’s repeated deployments. Effectively lower. This adaptation energy. (.) And catalysts compose. That’s another result. In the paper. So. To summarize. This paper. Tries to present. A physical theory. Of. Efficient intelligence. Principles. Via. Algorithmic. Catalysis. It aims to contribute. To the principles. For the design. Construction. And reuse. Of selective structure. (.) Intelligence requires. Physical work. Algorithmic. Catalysis. Is about trying to. Produce the amount of. Physical work. To perform a task. You denote. (.) As intelligent. (.) The design question. Becomes. Which reusable. Structures. Can we. Obtain. To achieve this. Yes. Reach out. If you’ve got any questions. Thank you. (..)
S37: And questions. From the room. Do you think. Like. For any. (..) Program. That’s trying to solve something. Think. Maybe. Like. A constraint. Optimization. Problem. (..) Yeah. Something like that. It’s trying to solve something. Was. And. Things are considered. Like. MP hard. Or like. (.) Infeasible. Whatever. (.) Do you think. That was your approach. Like. (.) It could. (.) Learn. Learn. Tactics. (.) The way. That would. (.) Make it. Much. Faster. In the solving. What I’m. To. My main question is. Do you think. (.) Similar. To your approach. For a truly. Non-random. Like. Program. (..) Because it will have. Structure. Your program. At the end. Will explore. That structure. Or like. Along the way. It will learn. Patterns. To explore. That structure. Which essentially. Like. Cuts down. The. I don’t know. Solving. Time. By. Significantly. Do you think. That’s something. That is reasonable. To think. (..)
S26: Yeah. Yes. That’s the general. (.) Idea. Although. It’s. (..) The intuition. Is a capture. (.) You know. (..) Human. Sort of. And can you. Learning is. When we see. A regularity. That we can recognize. We tend to. Reuse that. Pretty easily. Reliably. And those. Regularities. Have general. Algorithmic form. So. You know. If. If I. Update your. About something today. And then you. Talk to someone outside. Like. It’s already kind of. Locked in. Right. Right. And it didn’t. Require you. To review. It didn’t. Require a lot. Of computation. On your part. So in that. Specific case. Of. You know. So you’ve got a lot of sequence. You’re. You’ve got an NP hard. Kind of domain. And you’re seeing a lot of instances of the problems. Well. Most. Instances of problems in NP hard domains are actually. Fairly. Easily solved. It’s just. It’s just. Those kind of boundaries. (.) And so if your samples have some. Some. Compressibility. Then this system. Would. In principle. Find. That regularity. And then. Be able to exploit it. So you would. In fact. See. You know. A reduced. Reacquisition time. You know. In that scenario. (..)
S37: And. Do you think. That was. You said you were looking at the predictive coding. Aspect. But do you think. If. There was. Some. Predictive coding transformer. I mean. What transformers have shown. That have. You know. They managed to. Mine. I don’t know. Managed. To emerge patterns from all the data that they’ve seen. To the point where. We trust them with. You know. We see the code. They do. They. And. They do very impressive things. So the idea is. Could you. (..) Say. You start from. You start from. Zero data. When you run the program. And then. As the program is running. It starts. Getting data. And then. Essentially. It learns. Along the way. And. The idea would be. It. I guess. Exponentially. (.) Gets. Better. At. Solving the problem. It doesn’t mean. Or if it’s a. Completely. Like. Random. It will never solve it. But. Because. It will. Learn along the way. (…….) I’m. Take. I’m. Taking. (..) Inspiration. So in JIT. In. In. They have like. A. You know. Just in time. Like. Interpretation. And what it does. It like. It. Looks for. Hot loops. helps. And then. Automatically. Automatically. Optimizes for that. But if you could. Think of something like this. But in a more general. Model. Do you think that it could. (..) Like. I think. Think of it like. Almost as an exponential. As it gets. Better. It discovers. Something new. Maybe like. Tabling. Or. Something like that. And then. It. Applies. And it gets better. And better. And better. (…)
S26: Yeah. If I. (.) Follow you. You’re kind of. Pointing at. I think. The intuition. That dream. The dream. Coder. Lineage exploits. Very well. Which is part of the inspiration. For why I. Might. Even propose this as a practical. Place to start. So. The. The. The. The. The. The. The. The purpose. It it combines. Several techniques. To amortize. Search costs. So. It’ll. Spend a lot. Of expensive cycles. Looking for. Programs. Which. You know. Compress. Or. In. From. Coming up. There are. (.) These two people. It’s more Portland.拜 segments. (…) organizers. It. (…….) natürlich. This place where you can very rapidly detect things. What RB is trying to point at is a more general kind of almost philosophical, like an engineering philosophy perspective, which is what should be the orientation of the person sitting down to design a system which you want to have these continued learning properties over a long time horizon at large scale, like where do heuristics fit in there versus very general principles. So RB is kind of this wrapper or way to start to think about those problems. My particular allusion to predictive coding was from the perspective of, you know, predictive coding is kind of this new design landscape. There’s this interesting properties different from the tools we’ve had that were well developed before. (.) You know, the concept of thinking of part of it as an addressing process that you view as now necessary rather than an impediment, kind of calling back to what Alex was saying earlier, it kind of opens up that perspective. So the value is a little higher level and more abstract, but your intuition about, yeah, JIT should be an example of exactly this process. That would be a great, like, special purpose, you know, instance of something. And you would expect, you know, if RB was just, like, in the back end just observing, like, you know, JavaScript and then compiling to bytecode or, you know, just looking at the sequence of operations that the interpreter outputs, you would expect it to come up with JIT, like, in principle. (…)
S11: of the room. (.)
S16: Wonderful.
S11: All right, thank you very much. Thank you, Daniel. Thank you, Howard. Thank you to the absent Gabe and Elijah. Thanks, everybody. We’re going to have one more quick slide up here just to touch on the Fabric PC tool that was mentioned briefly in the Hyperon or was covered in the Hyperon workshop yesterday, mentioned briefly by Alex a little bit earlier. as a way to get involved with some of this active inference and predictive coding stuff. (..)
S29: Hi.
S28: I’m Matthew Barron with SingularityNet, an AI researcher, and I just want to highlight SingularityNet’s infrastructure for Fabric PC that we’ve been developing and has publicly released free and open source on GitHub. You can try it out for yourself, scan the QR here for our GitHub, and you can enjoy experimenting with predictive coding models yourself. We have a very convenient A-B experimentation harness. It has a very nice, very user-friendly graph-based API, super easy to define nodes and edges, and it works in arbitrary graph structures. We’ve also done a lot of work to make this right up to state of the art. Predictive coding has basically compressed, you know, 20 years of depth scaling from the backprop community into solving that problem in just the last year for predictive coding. So, huge milestone that predictive coding is now possible to train 100-layer depth models, and Fabric PC has implemented the new PC scaling for 100-layer depth models, and we are now going to implement the air-based PC as well. So, it’s a lot of fun to play with. Charlie and Ben have a paper in the conference using Fabric PC on columnar neural architectures. We have a lot of amazing researchers in ICOG as well, developing it on predictive coding transformers, which there’s a small demo in the library itself available for you using the examples. Thank you, everyone. (…..)
S11: Thank you, Matthew. And this is part of AGI Society and SingularityNet’s both commitment to building AGI tools that everybody can build on to help advance AGI, right? We can’t, any of us, build it alone. It takes a broad community working in all directions from consciousness to active inference, world models, causal models, and all the tools that we can provide that are open source and easy for people to use to accelerate the ability to get there together. So, this wraps up day one. Thank you so much for joining us today. A big round of applause for everybody still here. (…) Tomorrow, we’ll stay deep in on research, but we’ll broaden the lens a bit. So, we’ll open with Michael Levin and Hannah Nelhausen asking what stays constant about intelligence, no matter what substrate it runs on. Joseph Urban will ask whether machines are finally close to doing mathematics the way mathematicians do. We have two rising researchers, Faiza Habibi and Will Gebhardt, and Greg Meredith will close the academic program by taking us from causal models to world models, as well as paper presentations for the rest of our accepted papers, and then a poster, an informal poster session, and demos that include fabric PC, Omega Claw, ASI chain, a brain-computer interface, building sentient beings, and so much more, so that you can see it all and experience the emergence of AGI research and development right in front of your eyes this year.
