AGI-26 Conference | Day 2 | Keynotes and Paper Presentations

S34: AGI 26, again, our 19th annual conference covering general intelligence for machines, synthetic intelligence, building adaptive problem solver that can reason beyond their training data and work in uncertain environments over insufficient resources and information to help us do all the tasks or live their own goals and autonomous lifestyles as they want. The range of goals for AGI is huge and vast. We’ve seen that over the past days, covering so many aspects of what we’re all building here. Today, I’m super excited to introduce Michael Levin and Hananel Hazen, who will be doing a virtual and live presentation. So this kicks off our next paper session. This session is looking at intelligence as grounded world engagement. And this is one of the deepest questions in the field. What is intelligence really when you strip away the assumptions that has it look like a human brain? So these papers understand cognition through action, perception, environmental interaction, intelligence as something that emerges from engagement with the world. Before we kick off into the session, though, I’d like to introduce or bring Andrew Commendo, AGI Society Executive Director, for another invitation to all of us to get more involved in questions like these around AGI and the dialogue and AGI in the world. And just tell us a little bit more about AGI Society’s forward-looking steps. (..)

S08: Thank you, Haley. Is this working? Is this working? Yes? Yes? No? Good. Okay. Thank you, Haley. So just a couple of points from yesterday. 2,223 live streams. That’s a record for the AGI conference. So congratulations to everybody who’s doing that. Awesome job for everybody online. Thank you, Haley. I appreciate your work. Even though I am. So thank you for your attention, earn your attention to the program, I would like to jump in, but we’re just a little bit better. Also have it, this is, again, the largest conference we’ve had so far, over 300 in-person attendees, and a lot of good fun. We had the hosted speaker’s dinner last night, which was fantastic, so thank you everybody who attended as well as sponsors for that. One major point is that the diversity we talked about we’re actually realizing. So we have lots of high school students here, and we have a very young researcher who’s seen and giving presentations. I think that’s a big deal that we should celebrate. So, ultimately, you know, ultimately, this organization is going to be run by that generation, eventually, and so we need to do our best to support them as they’re young and burgeoning and so they understand the power that’s in their hands. So, with that, the last point I’ll make is at the lunch break, I will be having the AGI Society collective meeting where we’re going to discuss the future. It’s about 30 minutes, about roughly. It’s going to be very informal, but we’re going to talk about possible roles, how do you start a chapter, what is this going to look like when you want to host an event, things like that. How does the society, chapters, these types of things. And, largely, it’s going to be informational and, you know, how do we make this start working over the next year. If you have seen, if you see a banner and you see the QR code, scan the QR code, go to joinagisociety.org and sign up for what you want to do. If you don’t see anything there, and I’ve seen most people who’ve signed up, and we’ve had over a dozen people now sign up as regional organizers, people are putting, make my own role. That’s perfectly great, because we need one of everybody. There’s nothing that we don’t need. So, there’s a place for everybody in the society. And with that, thank you, Haley.

S34: Thank you very much, Andrew. So, yes, at lunch, seek out Andrew and get involved building local chapters, local discussions, and bringing this discussion larger from the people who have been looking at AGI longer than anyone else in the world. So, this next presentation, as I said, dual, virtual, and live, which I love, because we’re looking at the idea of substrate independence. So, two researchers, two substrates, one conversation. Michael Levin is a professor at Tufts University whose work on bioelectricity, morphogenetic intelligence, and what he calls diverse cognition, has changed how everybody thinks about the boundaries of mind. I was going to say scientists, but everybody. He joins us virtually, and alongside him, in person, is Hanan El-Hazan, whose work on spiking neural networks and biologically inspired computing bridges Levin’s biological insights to artificial systems. So, together, they’re going to ask, what stays constant about intelligence, no matter what it’s running on? (..) Please welcome these speakers with me. (.)

S19: All right. Well, thank you very much for that introduction. Very excited to talk to you. So, I will do the first half of the talk, and then Hanan El is going to show you a specific example of some of the things that we’re talking about. So, in my group, we study a variety of unconventional substrates of intelligence, molecular networks, passive materials, cells, tissues, slime molds, ant colonies, all sorts of things. And, and you can see, and you can see more of that here at this website, and the white paper discussing some of the ideas that I’m going to talk about today is here at this, at this preprint. So, in my group, we sort of do this cycle between biology and life as it could be, and different embodiments of intelligence, which I think of as mind as it can be. And today, what I want to address are two questions. First of all, where does intelligence, which I will generalize to patterns of competent navigation of spaces, and these are all kinds of problem spaces. So, not just the three-dimensional world or the linguistic space, but the physiological state space, anatomical state space, patterns of gene expression, all kinds of patterns of competent navigation. Where does it come from? Where does it come from? Where do those patterns come from? And talk a little bit about what may or may not be unique about life, and think about, really, the notion of bio-inspired and inspiration in general. (.) So, we have three ways that the conventional paradigm understands putting in effort towards getting some sort of a good controller. So, if you have something that is competent, be it a living organism, which I’ll just start off with, and then I’ll talk about the technology. If we have something that is good at certain tasks, there’s pretty much just three ways that it could have gotten there, either bio-engineers designed it, evolution selected for it, or it was trained in some way. So, rational design, evolution, or training. Those are the three ways we know to put in effort. And so, what I want to talk about are the delta that’s left over when you account for all of these things. I think there’s something very important that tells us something about the origins and nature of intelligence that does not come from any of these three ways of getting competency. So, first, let’s look at some biological examples of surprising outputs, and they’re surprising given the conventional origin story. So, if you’re thinking of design, evolution, and learning, you’d be hard-pressed to explain some of these things. So, let’s think about that. So, first of all, you know, we all know that here we are on Earth to evolve, but if somebody had come to you originally, you know, let’s say you’re an alien running some sort of investment fund, and they say, we want some investment, here’s what we’re going to do. We’re going to select for survival on this planet called Earth, on the African plane, just survival, nothing else. And what we think is going to happen eventually is that we’ll get some quantum chemists, we’ll get quantum chemists, we’ll get a moon landing, we’ll get all this amazing stuff, but the only thing we’re going to select for is survival. What do you think? How does that sound as a likely thing? So, nowadays, we don’t pay much attention to it because, obviously, that’s more or less what happened, but if someone had suggested this, I wouldn’t have put any money into this. So, I’m not sure if any of you would, it seems unlikely, it seems that if you only select for survival, it’s unclear that you’re going to get the kinds of things that we have here in humanity. So, okay, so maybe selection for survival in some way shaped the human brain, and for some reason, it became good on seemingly very different tasks than what it was evolved for. There are limitations of head size during birth, which minimize redundancy in the brain. So, given what happened, what performance would you expect from a human brain with a drastically reduced volume? So, ask yourself that, if I said to you, I’m going to take a human brain, I’m going to drastically reduce the volume of computational medium in that brain, what performance do you expect? Well, most of the time, you get radically reduced performance, except sometimes there are these interesting clinical cases, and you can study about them here in this paper where we review this, where humans have radically reduced brain structure, but normal or above normal intelligence. Now, it’s not that you can’t shoehorn this into sort of hand-waving about redundancy and things like this in neuroscience, but none of our paradigms predict this. We have no theories in neuroscience that actually suggest that this is ever going to happen, that you could have massive reductions and yet still have normal intelligence. So, something is missing in this accounting. So, let’s look at some other examples. Suppose, here’s some model systems. Suppose I wanted a tadpole of a frog, and I wanted to form an eye on its tail. I want it to be able to see out of that eye, but I don’t want it connected to the brain at all. And by the way, when the tail dies during metamorphosis, I want the eye to know that it’s different and ignore all those cell death signals. What would I need to do to engineer such a thing? Now, we could calibrate. I don’t have time to tell you what, you know, grant reviewers said ahead of time when we suggested some of these things. And they represent the standard expectations. Something like this should be incredibly difficult. You should need lots of either evolution or neuroengineering or something. You should have to put in a lot of effort to make this happen. Turns out you don’t need any of that. If you take a tadpole, this is the tadpole of the frog Xenopus laevis. Here’s the mouth. Here are the nostrils. Here’s the brain, spinal cord. So, you’ve noticed we prevented the primary eyes from forming. We put an eye on its tail, and then we use this machine to verify that, in fact, they can see out of these eyes. This device trains them for visual learning. So, they can see out of these eyes. These eyes do not connect to the brain, okay? When the tail inevitably dies during metamorphosis, the eye ignores it completely, rides the degenerating tail up and lands on the butt of the frog. If nothing else today, you could say you’ve seen a frog with an eye on its butt. So, this is fairly remarkable. Why isn’t there any lengthy round of adaptation needed? Why didn’t you need mutation selection? All the kinds of things that are typically required to produce an animal with a radically different sensory motor architecture. It just works out of the box, you know, zero shot. (..) Here’s another example from neuroscience. Suppose I wanted to take a shy, slow reptile, and I wanted it to operate at the speed of a cat. What would I need to do? Now, just think about that. If I told you I had a turtle and I wanted it to work like a cat, what would you have to do? You might think, okay, millions of years of evolution or maybe some, you know, neural engineering that we have no idea how to do. It turns out you don’t really need to do much. So, somebody put a turtle on a skateboard, immediately unlocked this kind of operation speed. The turtle is not only able to keep up with the cat, but also seems to want to play. It’s trying to engage with the cat, you know, it has all of these capabilities that you would think would be incredibly difficult, all unlocked with a tiny little prosthetic, a passive prosthetic. This isn’t even some smart, you know, AI or anything like that. It’s just a, it’s a passive prosthetic to its change to its embodiment. So, so what was the latent space of possibilities for this being? You know, turtles have been around for a really long time prior to this. What, you know, what, what was this, was this competency already in there? Why, why was it so easy to unlock? And we can realize that these kind of, these kind of changes are a sort of periscope to find the space of possibilities for us. Suppose, suppose I wanted an autonomous motile construct, which would contain no neurons. I wanted to make copies of itself from cells in the environment, like von Neumann’s dream of a robot that finds materials and duplicates itself. I wanted to respond to sound, but I insist on using just wild type frog cells, no training, no selection, no synthetic biology. How would we do that? You know, if you think about this, if this was the mission statement for a project, you would say there’s no way, there’s no way this is going to work. Turns out, actually, it’s, it’s pretty easy. All you have to do is liberate some skin cells from the frog embryo, and you get something we call a xenobot. The xenobot self-assembles. If you provide it with loose epithelial cells, it will assemble them into these little balls. The little balls mature into the next generation of xenobot, and guess what they do? They run around and make copies of themselves. So multiple generations. So this is kinematic self-replication. If you put a speaker under the dish, you find out that they respond to sound. Normal frog embryos do not do that. (..) So how did we get this? This is a standard genome. We did not engineer it. We did not put in any synthetic biology circuits. The environment is exactly the environment that frog embryos develop in. It’s weak saltwater. That’s it. There’s nothing else in it. There’s no instructive information in the environment. It was kind of engineering by subtraction. We simply released the constraints of the other cells that keep these cells confined to a boring fate of being a kind of a two-dimensional surface covering on the frog embryo. And this is what you see. And I could show you many, many interesting behaviors that these guys have. Now, suppose I want this thing. And remember, there’s no neurons in here. But I wanted to have an information architecture resembling that of human brains. Okay. So information structure. What would I need to do to it? Turns out, again, nothing. If we do calcium recordings, here you see calcium signaling. We can analyze it using the same information theory tools that are used to analyze brains compared to null models. And you see that for xenobots, the difference between them and their null models is not much different than the difference between human fMRI data and their null models. And so what I’m not saying is that these things are like brains. I am instead saying that both brains and all kinds of other materials can have similar information architectures for a very deep reason that I’ll talk about momentarily. And, okay, and if these things are active and they have this kind of information architecture, well, then I’d like them to learn. I’d like them to form memories. What would we need to do? You might think that because there’s no neurons, these things have never existed before, so they weren’t selected for learning capacity. We didn’t engineer them for that. You might think we’d have to put in some synthetic biology circuits. You know, people have been designing synbio circuits for memory, for being able to learn. (.) Turns out you don’t actually need to do any of that. They already do, so we discovered that, actually, they can distinguish two different kinds of stimuli. They can hold on to that memory for 24 hours, which you can then read out in their behavior, in their physiological signature, meaning the calcium signaling, and in gene expression. Okay, so at least a 24-hour memory that’s already there for some reason. We didn’t have to do anything to put it in. We just had to find it, which is something interesting from the perspective of the observer has to be clever enough to identify these things. So, so what if you wanted to swim in much more complex patterns? Well, here’s one. (.) What you could do, apparently, is just stick some random neurons in there. This is what, this is what was done. My postdoc in my lab, Halay Forawat, who basically just stuck a bunch of neurons in these guys. So normal xenobots don’t have any neurons. These do. And you find out that they actually do have an interesting nervous system architecture here. And then they have these, these, these novel behaviors. No muscles here, by the way. This is all ciliary motion. So the little hairs on the skin that normally redistribute mucus are now being repurposed for being able to swim in the water. So, okay, well, suppose I want the exact same thing, but I want it to be made of adult cells, not embryonic. So I can’t rely on embryonic plasticity. And I want it to be able to heal neural wounds. No synthetic biology, no training, no selection, not going to do any of that. But I want it made out of adult cells, and I want it to heal neural wounds. How would I build this thing? You can imagine, I’m trying to, trying to write a project about what, what kind of steps you would take to do that. Turns out, you don’t need to do much. Again, you can just liberate some adult human cells, not embryonic adult, from tracheal epithelial samples in, in, in, when, when removed from the body, they organize themselves into this little thing, which we call anthrobots. You could never guess from watching this. You would never know that they had a human genome. By sequencing the genome, you would see just normal homo sapiens. You would actually never see that, you know, you would never learn that, that this was the, the, the form and function that they had. You would absolutely not know until you tried. That if you do grow a bunch of human neurons on a plate, you take a scalpel, put a big scratch through it, and then sprinkle some, some anthrobots in there that they would collect into a super bot cluster and then start knitting together across the gap. And if you lift them up four days later, this is what you see right, right underneath where they were. They’re taking the two sides of the, of the neural wound and pulling it together. You’d have no idea they’re capable of doing that. What did we do to cause this? Nothing. We didn’t, we didn’t train them. We didn’t evolve them, uh, and we didn’t, uh, engineer. So, so, so there’s a fundamental, uh, and very interesting thing about biology here. We know when we paid the computational costs to have a good frog or a good human, it was during the, the eons that, uh, this, this system was bashing against the environment and being selected for. That’s the story we tell about the origin of standard plants and animals and things like that. This is how, this is when we say, uh, how did we get these amazing things? Well, it’s the history of selection that got us there, but there’s never been any xenobots. There’s never been any anthrobots. There’s never been any history of selection for them. And the, and yet we have a developmental sequence. We have novel behaviors. We have, uh, novel transcriptomes. They express different genes. Uh, when did we pay the computational cost for designing all of this? And you can’t just say that at the same time that you selected for being a good frog, you also got xenobots because the whole point of evolution was that there should be specificity. There should be a correspondence between the features that you see in a, in a creature now and the history of environments that, that got you here, right? You can’t, you can’t, you know, you can’t just say, uh, it just, it just sort of happens. We have to understand where these things, uh, where these things come from, these normal competencies. (.) Um, and so, so you might say, okay, so, so let’s, let’s look at where we are. So, so I’ve shown you a bunch of examples where if you think ahead of time, see, the problem is people are very good at moving the goalposts. So once you see that example, you get used to it and you say, well, I guess that’s what biology does, or you say it’s emergence or something like that. But, but fundamentally, if you, if, if you ask ahead of time, would I expect this? Like the, for these things, the answer is no, because none of these things, uh, were predictable from the stories we tell about the usual source of competency of selection of design or of, uh, of training. So, okay, uh, looks like biology has some kind of magic to it. Maybe it’s evolution does this somehow, maybe it’s the great complexity. Maybe there’s some sort of quantum magic here. Uh, but the standard view of, uh, bio-inspired is that, okay, biology has these amazing features. And what we should find out is, is what the biology is doing. And then that can inspire our technology, right? That the biology has unique features and that is what’s going to inspire AI, robotics, uh, you know, neural networks, design, all of that. I want to argue that, that this is, uh, this is not the right way to look at this, or at least not the best way to look at this. Um, we have now, so we, we know there are sources of, sources of information that do not come from physics or biology. Here’s a very simple example. So what you’re looking at is a Halley plot of a function in complex numbers as simple as this. So it’s only a few bytes of, of information here that, uh, that serve as a kind of, uh, pointer to an incredibly rich data structure. It doesn’t hurt that this thing’s very sort of organic and biological looking, but, uh, when you, when you ask the question, so why specifically this pattern? If this was biological, you would want to tell a story about selection, about, about selection forces that made it be exactly like this versus some other, uh, some other pattern. If this was, um, if this was a piece of physics, you would want to talk about some, some physical forces that, um, impinged on some kind of material and, and got you, got you to this pattern versus some other kind of pattern. But in, but, but here we can’t use either of those explanations. This, this thing exists in this particular shape and no other because of properties of mathematics, because of certain mathematical, truths that do not rely on any kind of a history of selection, they don’t use any features of the physical world. There’s nothing you learn in, in, in physics that will, uh, tell you why this thing is the way it is. There’s nothing you could do in physics to change this, this pattern. You could tweak all the constants at the beginning of the big bang. This thing would not change. So, so it’s interesting. We already know that in some cases there are sources of pretty rich information that do not come from, uh, from physics or biology. They do not require design selection or training. There is another source of specific information. Now, in this case, it’s just a static pattern or so we think. And the question is whatever the source of information is, what else might that, uh, might that source have other than static patterns? And so now, um, just, just briefly, and I can only show you a couple of examples. Uh, there’s, there’s a lot more primary work coming from our lab on all of this. So I’ll show you a couple, uh, whereas in biology, we talked about, uh, having bioengineers, having evolution or having learning in computer science. You have the same three sources where competency comes from. There’s engineering, right? So you can make a good controller or robot because somebody knew how to write an algorithm. And so they did, uh, evolutionary search. So nobody knew, but we generated trillions of variants or so, and, and threw out everything that didn’t work. And so we found, we found a good one or training, meaning that your system had contact with the problem space before and thus, uh, you know, sort of took its lumps and learn how to do it. So these are, these are the kinds of, uh, uh, uh, uh, routes by which we put in effort to reach something that is highly competent. When we say we created an intelligence or we made an algorithm that does something, or, you know, we, we engineered a robot of some kind, this is what we mean. We put in effort along these three lines, but even in very simple computational systems, we have a phenomenon that, uh, that I’ve been calling free lunch for it’s a, it’s a free lunch. And it’s probably not entirely free. It’s probably a heavily discounted lunch, but, but the point is that we get more out than we put in, you get scenarios, just like in the biology where you did not put in effort in along these three, uh, areas, like you thought you might need. And yet something amazing, uh, happened. So let’s, uh, let me, let’s, uh, let’s show you a couple of examples. So, so first of all, suppose I want a short deterministic algorithm, completely deterministic, uh, no, no, no quantum, anything, and no stochastics, a short deterministic algorithm. I want it to be, uh, it, it needs to be designed for an assumption of reliable hardware. In other words, it needs to have no extra code for what happens if my actions failed. I tried to execute something and it didn’t work. So the hardware is assumed to be reliable, no checks of any kind of, uh, as to whether what you did worked or not, but I want it to also work when it’s commands occasionally fail. I, I want to design it for reliable hardware, so no checking on anything, but I’m going to deploy it on, uh, on unreliable medium, uh, and I want it to still work. Also, at the same time that the short algorithm is doing whatever it’s supposed to do, I want it to do something else, something interesting. For example, maybe I wanted to implement homophily. Homophily is a biological, uh, sort of drive to hang around other creatures that are like you. So for example, the, the urge to, um, uh, uh, not be surprised. So, so surprise minimization in the sense of Carl Fristen’s work, uh, might drive multicellularity because hanging out with copies of you is the, you know, having a copy of yourself next to you is the least surprising thing in the environment, right? So, so you can imagine multicellularity being driven by, by the goals of, uh, of surprise minimization. So that’s a, that’s a theory that, um, Chris Fields put out a little while ago and I, about, uh, about driving multicellularity. So that, that’s called homophily. So, so, so I want this algorithm to work on an unreliable medium. I want it to implement homophily. Oh, and I also want it to do delayed gratification. What’s delayed gratification? It’s basically the marshmallow test. There’s this idea that you’re smart enough to go away from your goal. So, right, so, so back up the, back up the gradient that points towards your goals in order to later recoup, uh, better gains. Yeah. So, so, so, so it’s a, it’s a, it’s a reasonably sophisticated cognitive property, delayed gratification. They test human children on it and not all of them can do it. Um, so I want all of this, but I’m not adding any extra code for that. I want all of this, but I’m not adding any extra code. Now that sounds crazy. When, when we design algorithms, when we engineer algorithms, we don’t work like this. You, you list the things you want it to do, and then you write code to make sure that that’s what it does. I’m not putting in any code and any code for any of that. How, how, what, you know, how is this, how is this going to work? No problem. Uh, simple bubble sort already does this. And I don’t have time to go through all the details, but there’s a, there’s a paper here and some, um, blog posts on my, in my blog explaining how this works. Uh, good old bubble sort, which, which all CS students study in their, in their first, uh, in their first year for probably the last 80 years or whatever, has these incredible properties that are nowhere in the code. You, you, you know, there’s only a few lines of code to these sorting algorithms and nowhere does it say anything about these things that they actually do. If you know how to look, they do delay gratification. They do homophily. It’s, it’s really quite, quite astounding. So, so, so, so, so next example, uh, suppose I want learning capacity in my process. I want, when I say learning capacity, I mean, something out of the standard behaviors textbook, like Pavlovian conditioning, maybe habituation sensitization, maybe anticipation, maybe associative conditioning, which is Pavlovian learning. Uh, I’m not springing for any memory though. I want learning, but I, but I’m not going to put in any memory. I’m not going to put any algorithms for this. There’s no, there’s no memory medium. Now that sounds crazy too, except that as it turns out, um, and again, uh, we have, we have, uh, published work on this. Uh, I don’t have time to go through details, but, uh, networks like gene regulatory networks, uh, systems of predator prey equations. So, so ecological simulations, sentences of logic. So mutually referential sentences and even random electrical circuits all show different kinds of examples of learning in random instances, meaning not selected, not designed, not trained, uh, all just random networks already do all of this. You get it completely for free from the substrate. And these are very, these are highly diverse, um, diverse substrates, but you don’t need to do anything for these kinds of, uh, these kinds of capacities to come, to come forward. So, so, so, so just in the last couple of minutes, uh, I want to, uh, talk about what I think all of this means. And then there’s, you know, there’s way more to be said, but, but just, uh, just really briefly, if, if, if different kinds of patterns, patterns of anatomy, patterns of physiology, patterns of behavior, right? So all, all of these different kinds of patterns, navigation of different spaces are really instances of the same thing. Yeah. So there’s this class of patterns that are kind of, uh, uh, behavioral propensities in different spaces. Then maybe we can think about inspiration, uh, as a, as a kind of a degree of effort. See, the whole story I’ve been telling you now is a mismatch between the degree of effort that you’ve put in versus the payoff that you got. And one of the things that, that people say a lot about AIs, which bears on AGI and so on, is that, look, these are just machines. They’re following an algorithm. I, I have inspiration. I can, you know, I can, I can really sort of have, have new ideas that, uh, that are not, uh, laboriously calculated step-by-step by an algorithm. I, I, I have, you know, we as humans have some kind of an access, some kind of access to, uh, to, to something that is not, uh, uh, the product of, of, of kind of step-by-step calculation. I can be, I can be, I can be inspired. I can have intuitions. And so what I’m going to claim is that, yeah, we, we, we do, but, uh, but, uh, don’t feel, um, uh, don’t feel too special about it because I think this is a phenomenon that reaches all the way down. I do not think this is a biological issue. I think this is, uh, this is quite generic and I, even, even so-called machines have access to this is what it is, is my, is my view. So, so let’s just talk about that for, for a minute. Um, inspiration, uh, formalize me what, what is inspiration are many different kinds of inspiration, but, but one thing that inspirations have in common is that you seem to get more out than you put in. So here are three examples along that spectrum. So, so if you’re sitting there laboriously calculating step-by-step between where you start and the, uh, answer you’re trying to get to, that’s conventional. That, that’s what we expect. You expect to put in effort and you sort of grind through step-by-step and eventually you get where you’re going. And then you’ve got, you know, uh, mathematical, um, uh, uh, uh, prodigies like von Neumann who looks at the problem and he doesn’t need to do all the steps. He just sort of says, oh, I see it’s this and then it’s this and then, and then he’s got it. He doesn’t need to observe that effort. Then, then you have people like this. So this is Ramanujan who his personal experience was that, um, you know, he heard, uh, he heard these theorems whispered into his ear and, uh, and, and that’s it. Right. So, so what we’re seeing here are different ways to navigate a kind of landscape of ideas, a landscape of patterns. Some of these patterns are solutions to problems. Some of these patterns are behavioral, uh, repertoires, uh, propensities, and so on. And you navigate that landscape either laboriously, you crawl along it step-by-step, or you sort of jump along it if you’re better, or you leap across it if you have some kind of enormous intuitive capacity. And, and, and here we tried, um, uh, Robert and I tried to, uh, kind of quantify, uh, the effort in traversing the space of ideas. And we now have, uh, there’s now some tools that, that we can do that. So, so, so, so this is what I’m proposing is, is really important is to look at any, any agent, any kind of material, any kind of, uh, uh, process with a beginner’s mind. We, we shouldn’t assume that, that, uh, what, what the cost of, of all these things are because, because we’re so used to putting an effort in these three ways, we should actually check. We should actually experimentally look to see what else, what else is happening and how much effort are they, are they actually putting in. And so we now have a research program on this, and there’s a bunch of papers coming, looking at, uh, how even simple, um, simple, very simple agents, you know, simple robots that really have almost nothing in them. In other words, they don’t, they’re not trained, they’re not selected, they’re not engineered to do anything. There’s no algorithm, but they are the beneficiaries of a kind of intuitive leap. They don’t even have sensors in some cases, you know, they’re navigating these mazes. Um, uh, uh, they, they, they are able to, to, to receive actionable information outside of the normal, uh, channels that we usually deal with. And this, and this, I think has, has massive implications. So I, I would say that, uh, and this is a very active research program now, I would say, yeah, we’re, we’re the beneficiaries of all kinds of patterns ingressing from a latent space that, that give us more than we put in. But if, but, but, but, but we’re not part of a special club, it doesn’t take biology, it doesn’t take complexity, it doesn’t take evolution to, uh, to get on that, that’s that spectrum. Now, now, now being really good at it might, might take those things, but we don’t know. And so, you know, this, this, this, this question of, uh, these, these things look like free lunches, as I said, because there’s discrepancy between the way our current paradigms expect to pay for them and what you actually get. So what do you get? Well, we know you get static patterns from xenobots and anthropots. Apparently you get pattern trajectories as well. We know now you get behavioral policies in some cases that I haven’t talked about. You also get goal states, right? So intrinsic motivations. What else do you get? We don’t know, maybe, maybe compute, maybe perspectives, maybe something else, but all of this is relative to our still fairly meager ability to predict and, and actually test these things when they, when they show up. I think our intuitions are currently really bad, uh, about, about this. So, so my, my model is basically this, that, uh, there is a, there’s a latent space of patterns. Um, these patterns are, for lack of a better word, ingressing into physical interfaces. Those interfaces are simple machines, biobots, uh, robots, uh, uh, uh, you know, cells, embryos, you, you, you name it. All of these things are kind of front end thin client. That hosts, uh, that hosts, uh, various, uh, various patterns. We are blind to most of these. We assume we make these things. I don’t believe that we make them. I think we facilitate them. And I think that there, all of this is most obvious in Platonist mathematics. And I think math really is the behavioral science of just one layer of that space. It’s the layer that contains things amenable to certain classes of, of, you know, precise of formal models, but biology, psychology, AI, and so on is the behavioral science of other kinds of patterns that we would all recognize as kinds of minds. They would be on the spectrum. They would be complex, uh, agentic forms of different kinds of behaviors. Once they get a body, you don’t see any of these things until you give them a body. And we’ve been doing that, as I said, by, by, uh, using robot, robotic, uh, um, embodiments, uh, giving the, giving those kinds of embodiments to, uh, these patterns as controllers. And that’s, that’s what I think is happening. So fundamentally, I don’t believe that biology is somehow on top of our technology. I think that both biology and our technology receive inspirations literally from a latent space that we need to understand. I think that, um, passive matter, active matter, computational matter, and even the agentic material of life are just, uh, front ends for all kinds of, uh, patterns of form and behavior that all engineering is to various degrees, reverse engineering and mostly behavioral science, I think. And so I’m just going to, um, conclude here with a couple of, uh, summary points. I think that patterns of form, behavior, physiology, computation are an invariant class across, uh, a number of related disciplines. I think that design, evolution, and learning do not tell the whole story. Emergence is just a way of sweeping surprise under the rug. It doesn’t even begin to cover, I think, what’s going on here. We have to characterize the latent space from which these patterns are drawn. We can’t just say it’s emergent and that’s sort of the end of it. The research program now is to quantify ways in which our accounting of effort put in versus what we get out are deficient. It’s, uh, you know, you make the same errors as if you were studying a thin client and you didn’t know there was a whole server out back, you’d make, you’d make the same mistakes as, as we have made. I think, uh, I think we do receive, uh, these, these patterns of thought, um, as inspirations. And I think, uh, all systems do it to various degrees. And in the end, I think we don’t make intelligence either biological or technological. We facilitate its entry into the physical world, which enables us to interact with it. And I think that has massive implications for understanding and developing AI, bioengineering, cyborgs and all of that, lots of humility warranted about many, many assumptions that, that permeate these discussions, um, today. And so, and so just to, you know, just to, just to end here, I would say that because, you know, this is not a, a discussion of, of, of humans versus AIs, every combination of evolved material, engineered material, and software is some kind of viable embodiment for these patterns, you know, cyborgs and hybrids and every kind of thing. Some of these already exist. Some of these we will be living with shortly. (.) And so we really need to understand how to reach a kind of ethical synthbiosis, not only with, uh, you know, language models and things like that, but with all kinds of beings for whom we are now making embodiments, uh, you know, from, from swarm robotics to social and financial structures that embody minds that we can’t even begin to, uh, to, uh, to, to predict. And this is, I think, kind of an existential step for us to go beyond the current paradigms and understanding what we actually get when we make some of these things. So I’ll, I’ll, I’ll stop here and just, uh, thank the people who, um, who did all this work and our, our many collaborators, um, and the disclosures of a couple of companies that have licensed some of these ideas. And I’ll hand over to Hananel, who’s going to show you some, uh, some really cool, uh, recent findings. (..) Okay.

S26: So thank you, Mike. I’m Hananel. Um, and we’re going to continue a little bit from, uh, on biology and continue more on the machine learning part. and what we’re going to talk about is noise. (.) I’m going to talk about noise in biology, noise in, uh, computational. (.) It’s a lot of noise and I’m going to create more and, oh, more echo or that noise. (..) So we’re going to talk about noise. We talk about noise in biology, noise in, uh, computation, noise in, uh, most of the subsets that we are, uh, uh, working with. And we, the question that we’re going to ask is, can we use this noise similar to what Mike asked in his, uh, uh, in show a lot of examples in biology, we are going to continue with those, uh, biology, uh, examples and to show some application using those noise. So before I’m continuing, there are two, there are those characters that we’re going to see all of the presentation. We’re going to talk about noise, undirected background noise that exists everywhere. And the system that we are using need to deal with those noise background. And we’re going to talk about the data, our pressure, our treasures, because we use that values in order to process the information that we need. And we’re going to talk about the energy, the effort that we, uh, uh, put in, in order to, uh, uh, use those, uh, bits in order to do some computation, something, and to the, to get to that desired goal. So in biology, there are many parts in biology that use background noise. For example, this kind of very nice, cute motto that exists in our body that shuttling chemicals from one place to another, those motors are using noise in order to shuttle those chemicals from one place to another. They using noise, they using ATP, the energy in order to resist the noise that they don’t want. And they carry around with the noise that they do not want. So they tweaking a little bit the noise to their direction. We are in California, so they can say they’re riding the waves, they’re resisting the waves that they don’t need, and they’re riding the waves that they do need. And eventually they’re reaching the goal. This process exists in our body all the time. That’s how what shuttling a lot of chemicals in our body, and they use thermal energy and use many kinds of energies and randomness and noise. So another thing is in mutosis cells in our body contain the same genetic process, the same genetic code. However, even in the same cells, the same new body cells, they are divided. They’re divided between each other. They think they are expressing the genome slightly different. The body invest a lot of effort in order to correct those mistakes in the cloning. However, those mistakes exist. Some of the mistakes are banal. None of them are noticing them. Some of those mistakes can be cancerous, can be diseases. And some of the mistakes are the result of evolution. And that’s the mutation that we get in order to survive. (…) So some of those noise affect the mitosis of the cells. And some of them, until we get the slide back, some of them affect us and some of them doesn’t affect us. (..) Back to computers. In computers, noise is the enemy. Noise is the one thing that we want to reduce. We don’t want that in our system. There is many, many forms of noise. There are thermal noise. There are computers. Your computers, your energy that we affect in order to do the processes in the computer create heat. And the heat creates noise. And this is something that we want to dissipate. And we invest a lot of energy in order to solve that problem. So in the extreme case, for example, is one of the forms of quantum computers, which require zero, almost zero, absolute zero temperature in order to work. Because that’s the only place that that process can happen with the minimal amount of noise. And we invest so much energy in order to reduce that noise, in order to get a good quality computing. However, can we use that noise for us? And that’s where I’m going to talk about today. And before of that, I’m probably doing something here. Oh, yeah. Thank you. (..) So before of that, just to be on the same page, this is neural network. This is what neural network doing in during the process. They propagate the information from one side to another, and they activate those weights. Those weights are those the parameters that we want to tweak. And the reason that GPU are so good is because we found a way to represent those weights in a in a matrix form. And those weights are those and the GPU is doing the computation on those form of the matrix computations. And that’s why GPU are so good. In a way, our parameters in neural network are defined by how many connections we have in the neurons in the neural network. And the way that we train that neural network is by using back propagation with these hands off. This is the back, the best algorithm to train a neural network. There is no other algorithm to do that. However, as much as I hate to admit, it’s the best algorithm, but it’s not biology. It’s not how biology work. And I will give you more example to why it’s not how biology work. However, this algorithm is the algorithm that see all of the network in a God mode view. And when you present in a data point and you want to train the neural network to do that, this algorithm go over all of the weights in the in the connection of that neural network and change that. That’s mean if we have GPT style, a 3 billion parameters, meaning that we have 3 billion dimension, 3 billion parameters that this algorithm tweak in every time that it wants to change something to adapt the network to the training set. And the bad part of this is that the bad part of this is when we train the neural network on one data set, and we want to change that data set, and then we have an update on that data set, we can’t really train only the update. We need to start all over again. We need to train everything from scratch. And that’s the huge drawback. That’s not how, definitely not how human brain work. So, and the point that I’m trying to say is if you look at the state of the progress in the last 10 years, we can see that our demand for more parameters, bigger neural network, is just doubling every two years by a factor of 400. However, the hardware is not keeping up with it. The hardware, the memory that we have in the hardware is not keeping up with that demand, and it’s only twice every two years. So, it’s not sustainable, and the parameters that we need to tweak in order to get our Claude or our ChatGPT-favorite agent work, it’s enormous. So, can we do something about that? So, apparently, if one, another, another clue to why the neural network in our brain is not doing the same thing, we have 86 billion neurons in our brain. The fact that we learned today that this building exists here, and we know how to get here, we didn’t rewrite the whole neural network in our brain. We just modify some parts of it, not all of it. And you know what? If we are on that subject already, there is a couple of problems in the neuroscience that they talk about, which is not really clear here. In that paper, it’s the dark matter, the dark neurons in the neural brain. In the brain, meaning that most of our activity, regardless if it’s, I’m thinking about what I’m going to eat in the next meal, or if I’m calculating a differential equation, our activation in a brain is almost the same. Regardless of the hardness of the hardness of the problem, which is strange, and not only that, most of the neurons in the brain is not activated during those times. So, what are those neurons doing? And if they really activate, most of the activities that we see in the brain is mostly noise. There are many, most of it not correlated to anything that we can see, but only very few of those parts are activated, which clearly say that backpropagation is definitely not one of the algorithms that we want to get if we want to get biologically inspired. But you know, that’s what we have today. So, one of the new updates in the last years is that parameter effective tuning, which saying, if you take a neural network, and you’re going to train a neural network on two classes, for example, you’re going to train the neural network on houses and trees, and you succeed to train that neural network perfectly to classify between trees and houses. So, if you’re going to train the neural network, you’re going to train the neural network. If you want to add cats to all of this, you can’t use the current algorithm. You can’t use the current weights. You spend so much energy in training that, and you can’t reuse it. However, what you could do, and this is low-rank adaptation, you can freeze that neural network. You’re not going to touch it. Nothing. And to add two vectors that only those two vectors you train in order to add the cat on top of what you already had. So, you add your two vectors on top of the frozen parts, and then you get a neural network that learn both of the concepts. Also, the previous one, houses and trees, and the new concept, also cats. So, we ask ourselves, wait a second. That means that the information from the cats exists solely on those weights, on those two A and B vectors. Meaning, because all of the other parts are frozen, so that means that we can remove that part? Let’s say that I only want to train only on the new information, only cats. Why do I need the whole thing? So, we try that. We try that on the simplest test that we think of is this maze algorithm. We took a maze algorithm, the state-of-the-art maze algorithm, and there is an agent here that see the environment and try to get from the beginning to end. And this is the neural network that govern this agent, and it’s only 10,000 parameters, very small neural network, not even getting close to what we know on other kind of neural network. And instead of the fully connected neural network, we just throw them, and we replace them with two vectors similar to what we saw. Instead of looking at, instead of in the back example we learn the cats, now we learn this kind of behavior. And what we discover is that we increase the rank from one, two, three, to four, and none of them succeed. The last one that we can see here is, the last one you can see here is the fourth rank, and it did succeed to do something and then get inside some loop. Apparently, it’s not enough for him to learn. So, I thought, okay, we can conclude, that’s it. This is fantastic. The LoRa idea is fantastic to do some modification using a frozen neural network that you already pre-trained, previously trained. And on top of it, you can use something. So, we talked about that, and I talked with Mike, and Mike told me, you know what, we see that in nature there is a lot of noise. Can we use noise here? (.) Maybe we can use noise here. So, we do exactly that, and this is what I want to show you. If we add noise, and for that thing we call Lota LoRa, we add noise instead of the pre-trained neural network. Meaning, we don’t train any neural network. We just dump a lot of noise, and we start with the frozen noise backbone. And that noise is a Gaussian noise, completely random, you can find it everywhere. And using two vectors, you will see something interesting. And the reason we call it Lota LoRa is because not only the Lota randomness, it’s also remnants of the Lota, the lottery ticket hypothesis. Meaning, that drive, that saying is that every random substrate, there is a subset of that random that it’s actually the answer that you’re looking for. You can use that subset. And those adapters, the rank, is the earnest of those randomness. Try to use that randomness to your inventor. So, back to the previous example, we switch the fully connected neural network with two parts. One frozen, and one is the two adapters. And we use the same thing in order to train the agent. And what we can’t, in order to see the same thing, rank four, solve the task, completely solve the task. So, what we can conclude with that? Noise can help. (…..) Now we can see that we have two levers here. One is the amount of noise, and the amount of rank. Can we do something about that? So, we explore the whole possibilities. From one part, we don’t have any noise, to another part that we do have full noise, full matrix noise. And you can already see that the part that we don’t have much noise, any number of rank will not get us much. We’re not getting too, we’re getting some activity with the network is getting some information. However, we just need one rank. Not a lot. One rank with full randomness in order to solve that task. (.) And that’s something that motively return in the continuous studies. And what’s important to see here, that once we get to that point, adding more rank on top of that didn’t improve it. And that’s one of our open questions. But for now, I can say that, although maybe you don’t see it that way, that, Carlos, rank one is 7% of the training variables. We just succeed to solve the same maze with 729 parameters. That’s it. That’s what our algorithm search. The neural network that solved it completely needed 10,000 parameters to solve it. However, we tried it, we succeed to do that in 7% of it. So, let’s advance. You came to tell me already, oh, that’s 10,000 neurons. Come on, that’s very, very small. So, let’s go bigger. We didn’t want to jump way too high. So, start with this. We start with C410, convolution neural network. We start with feed-forward neural network. We graph neural network with decision transformers. All of those from 92,000 parameters to 24 million parameters. Also, LSTM and recurrent neural network. And to our surprise, we get almost the state-of-the-art parameters, the state-of-the-art performance with fraction of the parameters. We just don’t need so many parameters in the brain or in the neural network. We just need noise. Noise to ride those randomness. The adapters ride the randomness. And you see here something very interesting. Each one of those tasks have a different percentage, which seems odd for us, but again, still working on that. But if we look at the trajectory of all of this, we see something very interesting. We see that as we add more parameters, we’re getting diminishing return. We’re not getting the full success on the network. (..) And most of them getting to the SOTA, meaning we not only some of them we pass, but not by much, but some of them pass. So that point raise a lot of other questions. And this question is, are we looking at compression? Because if you look at the Kalmagorov complexity point of view, everything is compression. Are we looking at compression? So I think we don’t, because if we take a crazy number, we add huge amount of parameters. We only see not only diminishing return, at some point we see even less success. So meaning adding more parameters, it’s actually hurting us, not giving us better results. So what is it? It’s optimization. It’s maybe architecture. Maybe it’s different architecture. We don’t know. And currently now we’re working on that. And our most annoying problem is why are we not passing the state of the heart? Why are we not passing the 100%? If we clearly solve the task with much less parameters, which a lot of noise, but with much less parameter, why adding more parameters? We don’t get more better performance. So I don’t know. But how can we not adding an ideal hypothesis? Because we start with hypothesis. So let’s start a new one. Maybe we can do something about the rank of the task. And we create or coined the terms that the minimum rank hypothesis. That’s shown that you can solve any task that with those neural network, you need to sweep the rank. And the rank tell you, the higher the rank, it tell you that the task that you’re trying to solve is much more complex than you think it is. So some tasks require much lower rank, rank one, meaning you don’t need a lot of parameters to solve it. It’s probably very, very easy. The other one, probably you need much higher parameters. And it’s very close to the PCA, meaning you have a principle components analysis or principle components parameters that can define your problem in the same way that SVD, single value decomposition, can compose the metrics of the neural network. We just need the rank of that decomposing. And that will tell us how hard this task is. So you’re saying, oh, that’s cool. And this is enough for us. But what could we do with it? There is more, another more application we can do with that. So one application that we think about is you can train a one neural network and give to all of your client. You know what? Similar. Tesla. I know maybe there are many, there are many examples of it, but Tesla is one of them. You’re buying a model of Tesla and the Tesla, the car come with all of the feature. But some of it, or it depends how much you pay. Most of it are, will be disabled. So you need to pay in order to enable this feature. You don’t, they don’t take your car and put the machine inside. You have already the machine. They just switch a software switch in order to turn that on. So we can do the same thing here. We can train one neural network with three different keys. And each keys, we can give you a three different features. And you train only one. So depend on your client, we can give a different client a different keys. And you only train one neural network. That’s it. You ship that neural network. Once you get that key, this information is not leaked to the other keys. And so this is one of them. And this is a form of polycomputing with one set of computing depend on the keys. You can get a different activation, different form of knowledge that locked in that neural network. Another application of it is, and I don’t know how you didn’t stop me. What about GPT? You show me 20 million models. (.) What about 300 million? 900 million. So we didn’t go higher than 900 because Tuft computers are too expensive. However, we saw that having the neural network, the same Lota LoRa on GPT didn’t hurt in our performance. Not only that, we can replace the noise with not giving a float 32 noise. We can use thernary, three bits, or binary as a substrate of the noise. And we don’t need to shuttle so much amount of information between the GPU to the CPU. And the best part of it, 900 million model is only way 100 megabyte. The reason is we don’t need to save much because noise we can generate. Noise is one seed. It’s one number, one integer. You can generate it on the fly. The only thing that you need to save is the ranks. And the information of the rank is not a lot. We just see it’s half the fraction of the real neural network. So, for example, 900 million parameters, we took about two and a half gigabyte. We can ship it in 100 megabyte. That’s it. That’s the footprint on the neural network. So what we learned? We learned that noise is not the enemy. We can use it as a resource similar to what biology already doing. We can see that Lota LoRa can solve 90% or 97% of the performance with fraction amount of variables. We also see that random metrics, the substrate that we use, those adapters, cost practically nothing. It’s 64-bit. It’s like just one integer. And the saturation, the rank of the situation that we saw, tell us about the complexity of the problem. So we can assess how hard the task is to the neural network. So thank you. And I am here to ask any question if you want. Thank you. Yeah.

S16: So thank you both. Yeah. The question is regarding the ternary approach that you mentioned. Can you elaborate more on that? And regarding the quantization, have you tried with quanticized networks? How it work, your approach?

S26: Yes. We tried that with quanticized neural network and the adapters, if you meant. Yes. It works the same way. We lose a little bit because, come on, we have less floating point on the adapters. But it’s not, it’s almost the same as you’re losing the same full dense neural network. (.) So we don’t lose much. And we do lose a little bit. (..) What? (….) Internal? (…..) Yes.

S23: Yes.

S26: The same thing. (..) Apparently, the noise is not the, the representation of the noise is not the issue. the dimension of the noise, how strong the noise is, is the issue. (…..)

S01: Hi, Michael. My question is, if you’re talking about these capabilities as being intrinsic, not the product of direct selection, can you differentiate what we’re seeing there with the idea of a spandrel or something that’s just a side effect?

S19: Sure. Yep. Great question. I actually had a couple of slides on this that I removed for reasons of space. (..) So, yes, certainly, there are many situations where what we get is not the direct target of selection. We call those spandrels. (.) Now, there’s a couple of features to this. First of all, we have to be careful not to make this a kind of catch all term that just sort of hides our ignorance, because in the end, we have to be able to predict and control these. We can’t just sort of get whatever we get and say, well, that was a spandrel, you know, especially in some of the really important systems that we now are creating. Even if it was just a spandrel, the real question is, okay, but what is the space from which these spandrels show up? And can we learn to predict them? But the other thing is, there’s a spectrum here. It’s one thing if the spectrum is some sort of minor sort of feature. feature. But, you know, for example, if I told you, like, let’s just agree that some magnitude of spandrel becomes something else. If I told you, yeah, well, I wrote a little graph traversal algorithm, or I evolved it, but as a spandrel, I also got Microsoft PowerPoint. You would say, okay, that’s not a spandrel. There’s something else going on here that doesn’t, isn’t really captured by this idea of, you know, just sort of the extra features. So we can call them spandrels if we want, but the set of cognitive tools that we use to deal with spandrels are sort of surprise and just kind of shoulder shrugging and say, well, you know, sometimes these things happen. We don’t need to worry about them too much. But the problem is that what I’m pointing out is that sometimes what you get are not simple little features you can ignore. They’re behavioral propensities or problem solving competencies that when we make systems, you really want to know what these are. You know, we make social structures, financial structures, robotics, internet, you know, kinds of things. And if we can’t predict the level of cognition they’re going to have and the goals that they’re going to have, because that’s the other thing that’s at stake here. When you ask, what are the goals of a biological system? You say, ah, whatever evolution, you know, evolution put in. And when we make things that never existed before, be they AI, cyborgs, or anything else, knowing what kind of goals these systems are going to have is pretty critical. And the toolkit of thinking about spandrels doesn’t help us with this at all. So what I’m proposing is that we pull together a better set of ways to think about this that actually say, okay, really, currently unexpected, really complex, really impactful, potentially high-level problem-solving competencies. And we need to be able to predict and manage them. In some cases, you want to minimize them. In other cases, you might want to enhance and facilitate them, you know, that they can be amazing. These freelancers can be amazing. They can also be really problematic. And, yeah, the concept of spandrel is fine, but it doesn’t even begin to give us the tools we need to go into the future with this stuff. (..)

S25: So about this concept of noise, I mean that there’s an echo there. (.) It has been well known in neural networks that noise is essential, for example, for initializing networks or in graph neural networks, you have this random node initialization. (.) So could you capture at which point your noise approach differs from the standard noise approaches? So what’s your new essential insight? (……)

S00: Yeah. (..)

S26: For what I understand, if there is a difference between the different type of noises in the neural network. Yeah.

S25: And the well-known role of noise in neural networks, what’s your new insight with respect to the well-known noise? (..)

S26: I’m sorry, the echo makes it hard. (..)

S25: Yeah. (.) Okay. It’s noise is in use for neural networks since decades. So noise is not a new thing. Yes. So can you characterize what your new noise, new use of noise is actually? What part of the noise it used, the network? The new thing, your innovation with respect to the standard use of noise. In which new way do you use noise?

S26: I’m trying to get your question. I’m trying to get your question. You’re saying there are two different, what kind of different randomness affect the already exist noise in the system. That’s what you’re asking.

S25: Yeah. (.) So, yeah, you have impressive results, but normal neural networks use noise for decades. (.)

S26: Neural network what?

S25: Standard neural networks use noise already for initializing waves. //S26: Correct. Yes.// For initializing noise nodes in graph neural networks. //S26: Yes.// So, at a more high level, can you say, what’s your innovation here? What’s your innovative use of noise that leads to these impressive results?

S26: So, in a regular neural network, let’s say that we take the regular example with 10,000 neural network, you start with initializing those weights with chiming randomness or some other your favorite randomness initialization. And the algorithm, backpropagation, will use all of it in order to tweak all of the weights and starting from that point, reaching to the target that the training data that you want, to trend on your data. In effect, backpropagation is a very wasty algorithm. It will use every noise, every weight that you will give it. You will give it a small neural network. You will give it a big neural network. It will use everything. This method, the Lota-Lora method, squeeze the neural network or squeeze the backpropagation to use only what it needs and to use the fixed random weights or substrate in order to supplement what it already exists. And if you will see, if you think about it, the Lota-Lota-Ticket hypothesis suggests that in that substrate of that noise already exists a structure that will help you to solve, just by chance, it will help you to solve your problem. The only thing that you need is to filter out the other noise that doesn’t serve your purpose and to amplify the noise that already exists there to use that noise in order to harness that noise to solve your target. And that’s what the adapters are doing. The adapters, their whole job is to substrate or stress that bad noise and use the good noise to your advantage. depend on the hardness of the task or depend on the size of the randomness. You require more or less rank. (.) And that’s our finding. (..)

S34: Last question. Thank you. (….)

S15: I have a question about the problem you are trying to solve. There are categories of problems where random algorithms produce feasible good quality solutions. And when I listen to what you are doing, you are injecting randomness into your solution, into your algorithm. And you actually run the danger of replacing your solution with a randomized generic algorithm. So did you look at the categories of problems and problems which are solved by random algorithms versus other problems and how your approach performs in those spaces?

S26: So let me reform that question. You are saying that each one of the types of problems that I showed in the presentation have a different type of information that it needs to solve. And how the randomness helps for that. So the question that this is a good question. We just don’t really know what which one of those leverage or switches we need to turn off and turn on. It’s very complicated because the architecture that the convolution neural networks use is very different than the recurrent neural network and very different from feedforward neural network. And each one of them have a different limitation or the reason that we use convolution to images is because that architecture works better than other architecture. So that it’s clear that the architecture is part of the solution. We don’t just don’t know how much. And on top of that, the optimization, not every optimization under Adam algorithm is built to one part of of optimization of the of the regular weights of optimization. However, we don’t know how good it is in this kind of squeezing it only to the to the rank. So we have also the optimization problem. And on top of that, we have the randomness and the randomness also like the other gentleman here said, depend on the randomness that you present. We get a different solutions. So we’re working on that. That’s the short version. (…)

S34: Wonderful. Thank you so much. Thank you. Thank you for joining us here today. Thank you for these important thoughts and inputs on intelligence substrate independence. (..)

S19: Thank you. Thanks, everyone. (….)

S34: And with this opening, we’ll kick off into the papers for this session. So we have four papers now on cognition as grounded embodied world engaged processes. And we’ll start with a live a or the air. I apologize. A Georgian with origins of grounded semantics.

S28: So emergent of comportement grounded semantics. I am Olivier Georgian from the Catholic University of Lyon. And this work was done with Pierre-Emmanuel Marel and Cook. So there are three main concepts that I want to present today. The concept of non-representational sensory signal. The concept of interactional motivation. (..) The concept of schema mechanisms 2.0. (.) And then I will play some demonstration. (……) So non-representational sensory signals. You probably all know about Markovian sensory signal. When the agent has a full access. So can observe the full state of board games. (.) In contrast, non-Markovian sensory signal. There are two kinds. There are partially representational sensory signals. (….) Markovian precision process, POMDPs. The agent receives an observation. That is an observation of a representation of a probe. So it could be deterministic in the form. Observation is a function of the state. Or stochastic in the form. Observation is given by a distribution of probability over the state. But in both cases, it’s partially representational. (.) Non-representational sensory signal. Is when the agent receives a sensory signal. That is not a direct function of the state of the environment. So a typical example is a feedback from the action. In this case, the observation of sensory signal is given by a distribution of probability over the decision of the agent and the state. And more generally, a non-representational sensory signal is an event of interaction. (…) So we propose this model called the inactive decision process to represent the agent that receives non-representational sensory signal. So the agent first makes a decision at time t, dt. The environment transitions to state ST plus one. And the agent receives an event of interaction, IT, that is linked to a feedback from the decision. (.) And there is no representational reward signal as opposed to POMDP models. (.) Which brings me to the second concept, interactional motivation. So I’m going to refer to the paper by Pierre-Yves Houdet and Frédéric Kaplan. What is intrinsic motivation? They make a distinction between external motivation. (..) For example, in board games, the goal state is defined as properties of the environment. (…) Or another example is a reward signal that is given by a function of the state of the environment. In contrast, there are internal motivation. The typical example is the free energy principle. (..) But we propose interactional motivation. Which consists of associating a valence, numerical valence, to events of interaction. (.) And it can result in intrinsic motivation. If the agent just enact a task for the sake of enacting a stream of interactions that have positive valence. Or it can result in extrinsic motivation if the agent enact a task for the sake of a final positive interaction. On the third concept is schema mechanism 2.0. So you probably know about Gary Dreschel’s schema mechanism that goes back to the 1990s. But it uses representational sensory signals. (.) So more recently, we proposed another approach to schema mechanism. (.) with Filippo Perotto and Chris Torrison, whom you met in Reykjavik last year. That it’s a schema mechanism that uses non-representational sensory signals. And the schemas are events of interaction. (..) Primitive events are decisions on feedback. So, for example, the agent moves forward and bumps into a wall. That’s an event of interaction. (.) And it learns composite schemas in a hierarchy. And the fact that the schemas are not representational allows for recursive sequence learning in a bottom-up way. (….) So, comportement grounded semantics. We draw inspiration from theory of inaction and constructivist epistemology that suggests that biological organisms learn from non-representational sensory signals. (.) And the first habits of interaction. This refers to Piaget’s first developmental stage called theory of sensory motor apparatus. (..) So you will refer to Pei Wang’s paper on experience grounded semantics. in which tokens, or events, emerge from their outcomes, from their role in the agent stream. (..) And I refer to Ronson’s paper on symbol grounding, (.) who argues that the meaning is grounded in comportments, (…..) which do not require a final goal. It’s just learning habits of interactions. (…..) Okay, so my demonstration is called the SNF experiment. It’s inspired by the behavior of newborn mammals that tend to ascend a gradient of odor to reach the milk of their mother. So it’s an instinctive behavior. (.) All mammals do that when they were born. But it still involves some kind of learning. (…..) So in this experiment, the agent has six possible decisions. It can smell to the left, smell in front, smell to the right, turn left, move forward, turn right. And if it tries to move forward in a wall, it will bump. But it has no idea of what this decision means. It ignores that it exists in a grid. (.) And after making those decisions, it receives a feedback, which can be a decrease of smell, stable smell, increase of smell, or bump. (.) And it has an interactional motivation. So the SNF events have a valence of minus 1. They are slightly costly. The bump event has a valence of minus 10, as if the agent hated bumping into walls. Moving forward with decrease of smell has a valence of minus 5. And moving forward with increase of smell is the only positive valence, which causes the agent to tend to ascend the gradient of smell. (….) So here is the demo. This is the agent. This is a target. And here you have a representation of behavior of the agent. (…) So, of course, it’s going to find the target because it has this tendency to ascend the gradient of smell. The gradient of smell is visualized by the gradient of green color. But it doesn’t know what it is doing. It picks a decision randomly because it doesn’t know their meaning. And it receives some feedback. (..) So it smells. It sniffs around. And, of course, we want it to… So what we observe is that it has learned to sniff in front before moving forward in order to avoid bumping into walls because it dislikes bumping into walls. So we can see that the behavior is getting slightly progressively organized. So it’s not only learned to avoid moving forward when it sniffs a wall, but more importantly, it learns to actively use the sniff possibilities of interaction as an active perception, if you wish, or to gain some… (.) So it’s an epistemic behavior. So after a while, it learns the correct behavior and it reaches the target. (…) So, again, what’s interesting is not that the fact that it managed to reach the target, but the fact that its behavior organizes as if it understood what sniffing means, what sniffing in front, sniffing to the left, sniffing to the right means, even though this semantics was not coded at the beginning. (……) From this schema mechanism, we trained a little transformer just to visualize the semantics. (.) And it shows here that… So we analyzed the self-attention head matrix, and it shows that the event move forward increase of smell is dependent on the sniff-left increase of smell and turn left. Of course, because it sniffs to the left, turn left, and then it can predict what’s going to happen when it moves forward. So it does not depend… (….) So here there is a second sniff-left increase here. But of course, after turning, sniffing to the left doesn’t tell you anything about what will happen when you move forward. So it shows, it visualizes the emergence of this semantics. (…) Okay, so in conclusion, it’s a simple demonstration of the emergence of compartment-grounded semantics. We see it as a first step toward higher-level cognitive development. We expect that the next step would be to learn spatial-sequential schemas instead of only sequential schemas. And then the agent would have to construct a body model, displacement models, spatial-sequential world model in terms of possibilities of interactions. (.) And only after that, we expect that the agent would have to learn logical rules and to create rules. (..) Okay, that’s it. (10 seconds pause)

S34: Thank you very much, Olivier. (.) So next we have Roger Harrelson. He’s unable to attend last minute, and so was also unable to pre-record. As a substitute, a poor substitute, I’ll just read the abstract, give everybody a flavor of this paper and how it contributes. The Springer edition of the conference proceedings is expected to be out on Friday. It will be electronic, and we will link it to the conference site and send a link of the proceedings to all attendees so that you can read the papers in full depth at that time. But till then, the paper is unified cognition from a single optimization principle, active inference with meta, in meta, with emergent reasoning, planning, and self-knowledge. This is Roger Harrelson, independent researcher. So the paper. We present a cognitive architecture for artificial general intelligence in which perception, action selection, causal reasoning, temporal planning, and self-knowledge all derive from a single optimization principle. Minimize expected free energy. (.) Minimize expected free energy. Implemented in approximately 12,000 lines of meta, the architecture unifies three forms of inference, induction, deduction, and abduction, as hypotheses generators under a uniform metabolic selection mechanism. Hypotheses that predict correctly earn energy. Those that do not are removed. This yields emergent behaviors including spontaneous investigation of hidden causes, the Sherlock Holmes effects, principled retreat under a viability threat, and honest uncertainty grounded in actual query failure. An adoptive beam search reduces planning complexity from mathematical equation, achieving an 80,000 times reduction of 50 actions. A magnitude banded covariance pruning scheme reduces a pairwise Hebbian scaling from OE of 2 to OE to B, achieving a 19.4 per pair reduction at 1,000 observables with zero significant false negatives. An integrated 14 observable reef monitoring evaluation demonstrates emergent causal discovery, metabolic pruning, and, via ablation, that abductive inference is necessary for investigative behavior. So, again, that’s just the abstract unified cognition from a single optimization principle, active inference and meta with emergent reasoning, planning, and self-knowledge by Roger Harrelson. Again, proceedings will be out on Friday. We’ll send the link out to everybody so that you can read the full papers of all these paper presentations being given here today. Okay. (..) So, next in the session is Sergei Rodinov with executable world models for ARC AGI 3 in the era of coding agents. (..)

S27: Okay. ARC 3 challenge. (.) Probably you already know about this. is it’s some set of environments with unknown rules and unknown goals. (..) And so, to get 100% score, agents should solve this environment with less actions than humans. (.) And they tell us that it’s actually simple for humans, but it’s not completely true. Like, I invite you to try it. It’s not actually simple for humans. Like, it’s quite complicated. Guys, could you remove this panel? Does anybody hear me from the organization? (..)

S34: I apologize, Sergei. Oh, yeah. Oh, it works.

S27: Yeah. So, we built this, our proof of concept agent for solves this ARC 3 challenge. And this agent is based on three, like, ideas. The first idea is executable world models. So, we ask agent to maintain world models in the code. (..) Second idea is simplicity. So, it’s like Occam razor. So, we ask agent to periodically improve, simplify world models to make it more simple and more general. And the third is verification. As we all know, LLMs, like, in general, are very good in situation when they can verify the solution against something. And so, in environment, we can ask agent to verify against previous observations. So, we ask agent to reproduce all previous observations exactly and verify against them. And this proof of concept agent worked surprisingly well. And we get, like, 60% of performance. And we solve, like, 15 games out of 25. (.) Okay. But, each of these ideas can actually help or hurt. (..) Executable world models, of course, like, code is more rigorous. And it’s theoretically can support offline planning. But it’s costly to maintain. And maybe the text is enough. Second, simplification. I mean, theoretically, it should help. But maybe in some situation, it could destabilize a partially correct model. And replay verification. And replay verification. Of course, like, it can expose contradiction. But it can also encourage other engineering over, like, some non-important details. (…..) Okay. And to test which idea is, like, works and which doesn’t, we can make this experiment with four agents which are different as little as possible. So, first object is a textual variant. So, we just ask products to maintain world models in the text. Second is executable variant. We just add, like, a requirement to maintain world models in the code without any specification of what this model should be. Then we add simplification on top. And the last variant is a verification variant where we, like, ask agent to reproduce all previous observations exactly. (.) Great. And we get, and we get following results. We tested it with, we tested this agent with GPT 5.4 and GPT 5.5 with high and X-high reasoning. (….) And, uh, and, uh, and we get following results, uh, or, like, GPT 4.4, GPT 4.5 with, uh, different reasoning. And, uh, results were quite, like, uh, unexpected to us because, uh, uh, difference between variants, uh, like, uh, variants, uh, like, where much smaller, when we initially expected. (.) And there is a clearest result that we have, like, improvements with, uh, with improvement of model, with improvement of, uh, like underlying KLM or increase, increasing efforts. And if you go a little bit deeper, you can see that simplification usually helps as it should be, but executable requirements, it’s actually usually hurts. So textual variant usually work better than simple executable variant, which was unexpected, but verification variant works better in all cases. But it’s the most expensive one to run and difference is not so high actually with other variants. (..) Great. And what about GPT 5.6? And with GPT 4.6, 5.6, we saturated public dataset. So we tested our verification variant and our textual variant with X-high and max reasoning efforts. And our verification variant in both cases solved all levels, got 99% score. And if you check the number of actions which agents spend to solve all games and all levels, it’s actually half than, less than half than human baseline. So almost all games, it’s solved with superhuman performance, like far better than humans. And even our textual variant works not so bad. (.) We still have this trend that we have improvement with reasoning efforts from 92 to 95. 95 and with maximum reasoning effort, like this agent solve all games and almost all of them with superhuman performance. Yeah. (..) Okay. (.) The real question is, the real question is, what does it tell us in the context of AGI? (..) First, of course, these results should be verified with private dataset, but it will not happen because our team refuses to collaborate with us like in any way. (.) But let’s assume it’s true and it’s not only contamination, then what are the results? Yes. First, we have very, very rapid improvement of performance from GPT 5.4 to GPT 5.6, like very fast improvements. (.) And so it seems coding agents solve ARC3 and it means, like, we can hypothesize that we already have AGI, but for a low dimensional and deterministic environments. So coding agents already seems to be able to solve, like, arbitrary problem in a low dimensional deterministic environment better than humans. It is a hypothesis, yes? It is a hypothesis, yes? But, like, of course, for real AGI, we still need to solve, like, undeterministic environments and we need to deal with high dimensional environments and to build obstructions. Okay, even though we don’t have access to, like, verification that Hata said, I still can, like, have some, I still have some arguments that coding agents actually already solve ARC3. First, ARC team, they report results for different LLMs, but they do it for chain of thought, like, agents, it’s not real agents, for chain of thought algorithms. And, like, they do it like a caveman in 2022. So, uh, and, uh, even for chain of thought agents, like, algorithms, Claude Opus 5 already reached 30% performance, yeah? Uh, even for chain of thoughts. Then, if we compare performance of, uh, the, the, the, the, the, the chain of thought, uh, algorithm with, uh, uh, compare with our actual variant, which is, like, almost bare codex agent, yeah? Uh, then with GPT 5.6 for public games, for the same games, they got, uh, uh, 13%, and our, our agents had, like, almost, like, 96%. So, agents works better than chain of thought. Who would have thought about this? Like, yeah. Uh, so, and, uh, so, what agents could do with Cloud Opus 5? That’s almost certainly 100%. Then our result is not only contamination, because, uh, GPT 5.4 is actually predates on, uh, almost all games, and our agents already got, like, 53%. (..) And the last argument that, uh, even with chain of thoughts, uh, on, on results, which is reported by ARC team, they have very steep, uh, improvements from, uh, GPT 5.4 to GPT 5.6 for their private, uh, data sets. So, for their chain of thought agents for GPT 5.4 got 0.2%, and for GPT 5.6, they got 8%. So, we can, like, it’s kind of expectable that GPT 4.6 agentic system, they will saturate ARC 3 fully. I, like, I, like, I cannot, I cannot proof it, but, uh, it needs to be verified, but they don’t want to verify it. Uh, and, uh, there is links to, uh, project repository and, uh, two papers. Yeah. So, thank you. (…)

S34: Thank you very much, Sergey. All right. Last paper in this session, we have Gabriel Axel Montes, uh, again, virtually pre-recorded. So, uh, I don’t know how many were here yesterday, but we’re playing the game of, since Gabriel’s here, but not presenting, is this an AI or Gabe presenting? So, we’ll, we’ll take a show of hands afterwards. So, Greg’s paper is, Structural Morphology as an Architectural Prior for Stable Cognitive Basin Formation in AGI Architectures. Please enjoy. (1 minute[s] pause)

S24: Uh, I’m, uh, I’m here, and I’m talking about, uh, and then he’s telling you, and, uh, we should, for example, (….) on the cliff. (..) We should, uh, speak about technology. (..) The views, and the views, and the views. (..) See that, by, uh, we should talk about, promoting these follow-up wishes, (.) and, uh, the views on government, and that is, especially, now, that by software, by, to, So, you can be able to learn a lot from the same stuff, and the brain, basically, is the system of a favorable result for the integration of a very good student community. So, if you represent an opposition of the community, it’s a projection, and you start to have a, and ask the letter of the community to come to the park, and your, uh, picture of the community. So, you’ve got a different thought that it’s one by one in connection. That’s a big part of the community. And, can you imagine the results of the people, uh, low outcomes made in that amount of time, and that’s 5,000 people. The season presents a very difficult time for parents, but only for the whole operation, and their costs are very small. (.) so, and to be honest, I’m still most part of a platform. (.) I’m just a part of the Resίο Team i mentioned in my experience of having a Homer and Mac ¡Dong! (.) Have a breakfast! I constituted him late or as I was website,ào my American number inания, etc. Next question is about, uh, what do you mean by the public and the public side of the public? Well, the media show advice from the public manifest in the public. That’s the problem from the public. (..) The public is, uh, the public. Well, you send me a number of possibilities in the public. This is so far, you know, the public’s public. So, you know, the public’s public is, uh, the public’s public. Well, the components are those shapes that you can set up a lot of questions, you have multiple lines of different characters, such as the server’s query, and the above, and the reading numbers, and the system’s article. (..) The other thing, when the product passes on the surface, you can set up a few, and the server component is in the app app app. (10 seconds pause)

S22: Third case, robotics, plus something that appeared after this paper was accepted. Robotics, so body plan and control topology, uh, change the action today itself, upstream in performance. The anchor results embeds kinematic constraints directly into the policy, so a skill transfers across robots with different bodies while staying feasible and safe. (.) It’s the one case where a bundle component of robustness is improved by a design decision rather than only diagnosed. (..) Now the note, and I’ll flag it here as preliminary, so a system called Omega Claw, which many of you may be familiar with, shipped recently, a neuro-symbolic agent from the Hyperon stack. It reads as a convergence this framework predicts, an open costile agent framework whose substrate is now Hyperon agent. Its own documentation names the same scenes. Separate channels, memory, and profile layers, a continuous loop and session of refresh, auditable proof trails, and explicit security policy service. So I’m not claiming a new, um, spectral result on this, just that the predictive shape. Predictive shape is showing up in a writing assistant. (…) So, put the three cases side by side, and the bundle’s purpose becomes clear. It exposes mismatches rather than crowning a symbol. (.) Open Claw has a high trace richness, but only medium scene resistance. Records are abundant, coordination still bottlenecks through narrow routing. And Hyperon has the opposite tension. (.) Broad, shared state design, but low scene integration because the pieces haven’t converged yet. Public service. Robotics is the case where robustness gets engineered directly into the policy. (.) Recurrence is undetermined everywhere. (.) Static structure alone, not probate, and that’s just an honest gap. (….) The safety point that I probably care about most is that morphology doesn’t predict failure directly. It flags architecture. It flags architecture. (…..) So, a case, and they’re requesting a mix of the being. Supposed an agent that passes every safety vector, right? The record uses no drift. (.) A morphological audit shows that one scene of public manifest to run timeout as the single highest 20s connector from a whole architecture. All capability extension passes through one narrow, external, durable surface. Anyone who can edit and work in manifest gets outsized leverage without touching a lot of weights or tripping the evals. So, the benchmarks are assigned to it by construction, and the morphology shows it directly. (.) Caveats. So, this is a prior, not a proof. There’s only seam integration that’s fully operation-wise at this stage. You can use it comparatively. The cases here are software-centered, so biological substrates are open. (..) So, implications and takeaway. The shape of an architecture works under, deserves the same scrutiny as what it produces. Two implications can fall from that. So, for AGI broadly, heterogeneous systems, LLM agents, shared substrate neuro-symbolic stacks, embodied controllers, these all become comparable in one vocabulary. Seams, traces, time scales, robustness, before anyone can really claim the realized dynamics. For evolving systems, the same analysis works as a developmental audit. (.) Re-run it over successive releases and watch whether scenes thicken and integration moves from the repo boundaries towards genuine internal coordination. Omega Claw is probably the first live chance to do just that. (.) You can find the analysis artifacts on GitHub. And for any more detail and depth, you can refer to the paper. I’m very happy to hear any feedback on this. Thank you. (…..) Thanks, everyone. If you’d relax a month as well. This paper argues that when we compare AGI architecture…

S34: That was great.

S22: We usually measure not necessarily…

S34: That’s going to be the presentations for Dave. We do some of our two presenters that we can do a quick Q&A with. So, Sergey, Olivier, if you could come back up. (…..) I have good news about the proceedings, as Gabe said. Looking forward to people reading them and giving feedback. Is that they are ready and will be published to the AGI website shortly. So, that’s good news. All right. Questions for Sergey? (..)

S00: There we go.

S34: Or Olivier? Hi. (.) And please specify who you have a question for.

S13: The first question is for Sergey. Can you hear me? (..)

S25: A little bit.

S13: Can you hear me? Okay. The first question was for Sergey. Both talks were excellent, by the way. (.) I was looking for an architecture guide with them. I wanted to see, actually, what the contribution of your team was versus the LLM. You know, what actually you put in to have… Because it seems like the LLM is driving the interaction between Arc AGI 3 test. And so, can you explain what your team actually did in order to make it work? (..)

S34: Sorry. (…)

S27: Difficult to… (…) I mean, it was very difficult to hear. (….) Like… Yeah. (……..)

S13: What did the architecture… What was the architecture of your system? There was no diagram.

S19: I…

S13: And so, what did your team do versus what the LLM? It looked like the LLM was doing all the work. //S27: So, what did you do?//

S27: There is… There is… (10 seconds pause) Okay. (..) First about architecture diagram. You can check it in the second article. (.) There is very, like, precise description of architecture. But… But… I will… I will describe it very shortly. So, it’s basically codex agent. So… (..) Like… The main architecture… We have a codex… (.) Codex… Which we instruct to play a game. So, codex interacts with the game. We hire clients. So, it… Run client and… Fun game. So, like… The answer to your question… Actually, everything is done by codex. (.) What we do… We only, like, hand some prompts. So, our architecture is very thin controller on top of codex. We only send the main prompts. Like, explaining what it needs to do. And then, we send some continuation prompts. Like, simplification prompts. Asking… Okay. You have both models. Please simplify it. And then… (.) Yeah. So, it’s… Everything is done by codex. (13 seconds pause) That is, like… (..) That is, like, fundamentally different. It’s… It’s, like… Like… Your… (.) Codex agent. Like… Or… When you speak to ChatGPT. Who don’t speak to LLMs. You speak to agents. Which has, like… They have access to Python. They have access to… (.) System. So, it’s, like… It’s, like… It’s fundamentally different system. It’s not, like, LLM which simply, like, generate… (..) Some text. It’s, like… It’s LLM which can run code. Which can run everything in the system. Which can read files. It can write files. It can write files. So, it’s… (.) It’s fundamental difference.

S23: Yeah. I mean… I mean… He’s asking the LLM to generate a procedural model of pieces of the game, which is in Python. And you can then run that model. And you can formally verify aspects of that model. So, I mean… It’s LLM, Python, interpreter, formal verifier. And you can use codecs to orchestrate all these things. But the… The subtler question is, with the latest LLMs and the fact that the results get better with the latest LLMs. Now, indeed, the latest LLMs are still bad at solving the problems on their own. But the fact that wrapping the latest LLMs in this sort of agentic loop with a code interpreter and a verifier, that this now solves the problems better. Is it because the latest LLMs are smarter? Or is it because the Arc AGI-3 problems were somehow baked into the latest LLMs in their training, but in some way that still doesn’t let them solve the problem directly very well, right? And we… We don’t know the answer to that. The… I guess at least when we last talked, we didn’t know the answer to that, right? I mean… I… The way to solve that would be true out-of-sample problems from that same problem class that you knew were not in the training data. of the latest LLMs, we could generate our own true out-of-sample problems, but then there’s the question, like, why should you trust us to generate true out-of-sample problems for our own algorithm? You would like the Arc AGI-3 organization to now spit out some true out-of-sample problems that you know aren’t baked into the latest LLMs and they haven’t done so. (…) Is that accurate enough, Sergey?

S27: Yes, yes, yes, yes. Yeah.

S13: And for Olivier, could you elaborate a little bit more about intrinsic motivation versus the, you know, extrinsic, external? (..)

S28: So, more about intrinsic motivation. Yeah, I think the term intrinsic motivation is a bit ambiguous because we never know if it refers, if it is intrinsic to the agent or intrinsic to the task. And that’s two different things. So, that’s why I like the first distinction between internal or external motivation. It means it’s internal to the agent or external to the agent. So, the way I would describe it is that any biological agent only has internal motivation. (.) Only artificial systems like chess players can have external motivation because the motivation is defined externally as a property of the state, such as check-matting your opponent. It’s an external property. It’s an external property. So, for biological system, it’s always internal. But for artificial system, it could be both. But we are working on internal motivation. And then internal motivation could be either intrinsic to the task or extrinsic to the task. Intrinsic to the task means you are playing music. You do it for the sake of playing music. It’s the task itself that gives you your motivation. Extrinsic to the task is you work for money. (.) You don’t like working, but you do it for money. But it’s still internal in the sense that you value money. (..) That’s it. (…)

S16: Thank you, Sergey. (.) Regarding, you have tried in the Kaggle competition, your model, the world model. Have you tried in the Kaggle competition with your world model?

S27: It’s an interesting question. It’s actually like on Kaggle competition. So, it’s actually what Arc team allow you to do is only participate in Kaggle competition. They’re kind of annoyed when you try to do something else. But, because, like, I… But in Kaggle competition, they allow very little resources. It’s like, it’s incredibly little. So, and if you remember my, like, plot with performance, and each performance difference you have with GPT 5.4 and GPT 5.5, so, and the resources which they allow to use, it’s covered over there, yeah? So, it’s like, it’s almost nothing. You have one GPU and, like, 10 hours to solve 50 games. It’s like, yeah, it’s interesting engineering to ask. But, it’s, for my opinion, it has nothing to do with AI research. So, I even don’t try. It’s, I think it’s pointless. But, yeah.

S16: Yeah, you tend to disagree. We, in fact, for example, with a new reaction model that we presented in the paper, we already achieved, like, 0.25 just in 15 of the games. You know, of course, in the Kaggle competition, you use 55 games instead of the 25 that you have offline. And we thought the use of GPUs. We’re just using just a new neural model. Maybe it’s another approach, you know, using LLMs, no? Good, good.

S27: It’s just, like, how much to achieve, like… How much what? (.) Which score? (..)

S16: 0.25. But it’s better than TPT 5.4, for example. It’s still very low. Very low. But, yeah. (28 seconds pause)

S24: But, it’s…

S27: I don’t know. He can discuss. But, I think it’s, like… Yeah. Maybe if you can do the same with so little resources, it’s cool. But, have you, like… I, uh… I, uh… Had this results on, uh… Kaggle competition was around 2%. But, it’s like, it’s nothing. He… He have 100. So, you claim… You have 25. (..)

S16: 0.25. Yeah. //S27: Huh?// 0.25. For example, GPT 5.4 scores less than 0.20. So, this is a bit better than using the LLM with the standard tests that they do. So, they spend some money by using the LLM. But, of course, in the competition, you can use LLM because it’s just GPUs and some compute. We don’t use GPUs. We use CPUs, by the way. Internally computing instead of binary. But, I think it’s just…

S27: Maybe it’s another approach to, to, to H.I. As I told, like, for my opinion, it seems coding agents, they can already solve, like, (.) deterministic, low-dimensional environment better than humans. So, it was my point, and I think Kaggle competition is another story. (…)

S34: Thank you. (…) Other questions from the room? (….) Anything for Olivier and Sergey? (….) All right. Wonderful. Thank you very much both. And, to our paper presenters who could be here as well. (…..) So, now we’re going to dive straight into our next paper presentation. There is coffee, of course. But, we’re going to go straight into the next session without an official break. (.) And, I’m really excited for this next session. I’m excited for every session, but not playing favorites. But, this is going to be about intelligence as formal structure and inference. So, from intelligence grounded in a world to intelligence grounded in mathematics, we’re moving now into how to learn intelligence in formal structures. Logical operations, category theory, non-axiomatic reasoning, the architecture of inference itself. And, our keynote is Joseph Urban, a professor at the Czech Technical University in Prague, one of the world’s leading researchers in automated theorem proving and machine learning for formal mathematics. His title today references the QED manifesto with an extraordinary question that we’re getting close on. (.)

S23: Yeah, and I think Haley has said most of what needed to be said. But, I couldn’t resist saying a word of introduction for Joseph also. So, I think it’s an honor to have Joseph at the AGI conference again. is that this guy has really been leading the field of AI for theorem proving for some years now, going back to the deep math paper, a collaboration with Google that sort of put neural nets and theorem proving on the map, and then way, way before that. So, we’re seeing now, I mean, latest model LLMs are, in some senses, able to do math at the level of a human mathematician, not in every sense. Connecting LLMs with formal verifiers like Lean4 and other alternatives is doing amazing things also. But, there’s still a lot of path forward at the intersection of AI, AGI, and theorem proving. And, there’s still more to be done in the domain of having AI prove math theorems. And, there’s also a lot to be done with AI for theorem proving in domains other than traditional math, like common sense reasoning as an aspect of general intelligence, or automation of different aspects of science pursued by automated theorem proving methods. And, we’re gonna hear about all that as soon as everything here is working. (..)

S09: Thanks a lot, Ben. I apologize for my lateness. I flew in from the Federated Logic Conference, where I had three talks. (..)

S34: We actually, we do need five minutes. So, feel free to coffee break, and we’ll be back in just a moment. I apologize. (3 minute[s] pause) Talk worthy of two introductions. Please welcome Joseph Urban. And, a talk showing how mathematics is benefiting as much as other domains or more from this explosion in AI power. Please, Joseph. Thank you.

S09: Thank you. Excellent. I have my slides. So, again, I apologize for my lateness. And, you will hear a bit more. I just flew from the main logic reasoning conference, which is happening every four years this year from Lisbon. And, the slides are a bit quickly written and co-written by all sorts of AI tools. But, I’ll explain what I mean by QED, Utopia, and Machine Understandable Science. (…..) So, eight years ago, at my first AGI, which was in Prague, I gave this funny titled talk, No One Shall Drive Us From the Semantic AI Paradise of Computer Understandable Math and Science. (.) And, at the time, it was kind of a bit of a provocation from me and a bit of a pun on a famous quote by a famous mathematician 100 years before that. But, my talk this year is, like, this has kind of really almost arrived or is, like, very seriously arriving. (….) So, I’ll show you that basically what happened in the last six, seven months is that we are really now turning human language mathematics, like, really large pieces of mathematics into symbolic formally proof-checked artifacts. (…) And, it’s done in some kind of very interesting loops, which are between the kind of statistical machine learners and neural machine learners. and the very strong logical tools and the libraries of knowledge. And, like, this loop is, like, a very interesting object that is now evolving. And, it’s really interesting how far these loops will get and what will become of them in the relatively near future, given kind of the gradient. (..) So, the QED in the QED utopia comes from quote era demonstrandum, which is like a long ago invented way how to close mathematical proofs. But, it’s also like a 32-year ago manifesto, written kind of after in the Bourbaki style by a group of people in Formal Math mostly, who, who, who. Oops, now you have seen all my slides. (……) Is there a way how to go to the beginning? (………) I’m trying. (….)

S24: Oh, now it’s stopped. (10 seconds pause)

S09: Sorry. (…) Oh, thank you. (…) All right, I’ll start kicking forward. Yeah. (….) So, this is the kind of QED manifesto slide where people like 32 years ago basically started to dream about representing all of mathematical knowledge in a giant verified, very strongly formally verified knowledge base or database, which would kind of serve as a as a basis for basically all kinds of human, math, science, development and it would kind of keep, like for example, software verification and hardware verification and all these fields which are kind of using all sorts of mathematics and reasoning would profit and contribute to this, like, securely, logically verified knowledge base. Yes. (…) And so, so we had a workshop in 2014 where Michael Beeson gave this talk where he, funnily enough, talked about the QED singularity and he, he gave a bunch of blockers at the time. Uh, the, the, the, the main of them would be that this is just too much work. Like, let’s say 10, uh, years ago, we, we kind of had the technology, but it was like really costly to, to basically transcribe, uh, mathematical textbooks and articles into, into this formally, uh, verified and proof check knowledge. So, so, so, so he’s, he’s, he’s this, uh, Mike’s, uh, slightly pessimistic, uh, slide, uh, from 2014. (..) And, uh, uh, uh, and at AGI 21, I, I kind of, uh, wrote, uh, this kind of more optimistic slide when, uh, we were, uh, really kind of progressing with this field of automated formalization. (.) basically thanks to neural language models and their combination in some kind of feedback loops with the strong, uh, mathematical proof checking, uh, that we, that we had. So, so in, in my talk five years ago here, I, I said that it’s, it’s almost, uh, possible, but not yet possible. Uh, so there was a talk by Sam Alexander, which I may, uh, show you, uh, about this particular, uh, AGI, uh, theory. This, which he did this, uh, Marcus Hutter, which was like reasonably mathematical talk. Uh, which had like, like a bunch of theorems, which, uh, they are really proving in, in, in the paper. And, and my point was like, like maybe five years from then, we will really, um, be, uh, proving like, like on demand in, in real time. We will be basically fact checking, like, like very strongly fact checking, like basically in real time transcribing what, uh, Sam was talking about, uh, into formally verified, uh, math. So, so today, uh, I, I did it. It took me half an hour with chat GPT. So, so, so, so, so the systems really in, let’s say the last, uh, half year or so. (.) became so good that it took me 30 minutes to, to instruct chat GPT to, to download a particular proof checker. I, I told it how to run the loop, how, how to kind of try to formalize and, and proof check it because I have been now doing it for, uh, more than a half a year. And how to ultimately basically transcribe the, the paper of Alexander and Hutter into like fully formally, uh, verified code. So, so what I’m showing you is, is the kind of automatically produced formally checked, uh, artifact. So, so, so, so basically if you want to remember one thing from my talk, it’s that like, this is now really possible. What I was talking maybe five, eight years ago ago, I can now do in half an hour. And it’s really amazing. It’s, it’s kind of changing the, the field of, of mathematics and probably also, uh, science. (..) So, so, so, so one, uh, uh, result of, of, of this is that, uh, this is now causing quite a bit of motion in all sorts of research, uh, communities. So, so, so, so last week there was this federated logic conference in Portugal. And apart from like a bunch of talks about combining, let’s say, learning and proving and conjecturing systems, there was suddenly also like a bunch of talks about really throwing, uh, mathematical textbooks into, into such workflows and getting them transformed into, into this formally, uh, formally checked artifacts, uh, with various cost and time. And, and, uh, like a very impressive, uh, uh, recent thing is now, this is now basically moving into industry. So, so, so Dominic Mulligan from the AWS automated reasoning group, uh, gave a talk about fully forming, uh, formally, uh, formally, uh, verifying the, the Graviton 5, uh, Nitro hypervisor, which, which they deployed without people basically knowing it. So, so now, like a very important component of the Amazon cloud is, is fully, uh, verified this hundreds of thousands of line of formal, uh, proof. Um, so, so, so, so that was kind of an intro and I will probably go through some of these examples, uh, now. (…) Um, so, so, so, so what I did myself in, let’s say December, uh, January, when. when, uh, basically is our girls all started to tell me that some of these agentic flows are becoming, uh, pretty competent. It was to take, uh, reasonable, reasonably well-known, uh, textbook on general topology like this monkeys. And I, uh, chose, uh, very simple, uh, formal proof assistant by Chad Brown called, called Megalodon. (.) And, uh, through, uh, kind of coded an environment and, and prompt, uh, in which I kind of could throw the whole latex of that book. to, to, to, to the coding agent equipped with the proof checker. And some, some more tools which I, which we, uh, coded. (.) And, but, but basically I kind of programmed like one of these long running, uh, feedback loops between the LLM, which, or the coding agent, which reads the latex, uh, is instructed to, to produce, to produce the, uh, formal, formal logic, uh, out of the latex. Like the definitions, theorems and proofs. And is also instructed to call the proof checker, which in this case works as a compiler, as a relatively fast compiler. (.) Very often to, to debug what the LLM, uh, did. And like, like one remark is that the LLM, at least in December, was pretty bad at this. So, so if you, if, if you hear that the LLMs are perfect and that they can one shot something. Well, definitely not in this case and not in December, uh, the LLM was doing mistakes basically all the times. But like the, the, the kind of magic bullet was that the LLM calls the, the, the checker very frequently. And the checker gives it the, the relatively good error messages. Like, like, hey, you made a mistake on, on, on line 530. Uh, I do not understand how, how you wanted to prove this statement on line 700. And the, the, the LLM has, uh, enough common sense reasoning to it. (.) That it, it, it can use this kind of, uh, feedback. And after enough, uh, trying, basically get the, get the proofs and, and, and the theorems and the definitions, uh, right. Like, like, obviously with some, with some caveats. So, so after, I’ll say two weeks of, of running this experiment. Then I basically fully automated it. And I was running with the heavy chat GPT pro subscription. It produced 130,000 lines of, um, um, verified, uh, formal theorems and their proofs. Including some, like, pretty serious, uh, theorems in, in topology. Like the Urizona lemma, the Urizona metrization, the TISA extension theorem, et cetera. So, so this was like something that we then, uh, started to, uh, replicate. So, so this is my talk from, uh, ITP last year. So this is the kind of growth of the formal, uh, proof lines there. (.) And, uh, like, like, like the very, uh, uh, interesting thing about this, how, how embarrassingly simple it is. So, so you really have just this, like, very simple, uh, loop, uh, that says, like, try to translate the LaTeX, um, around the proof checker. If it works, go further. If you get an error, uh, try, try to repair the, uh, errors. Yes, um, uh, yeah. So, so, so, so then we started to replicate it with different, uh, systems, which are, like, more popular in the theorem, more proof checking community. So with John Harrison, we extended it to his system, Holite, which is also used at AWS, uh, for basically industrial verification of cryptography. (.) And just like that, basically, John, uh, formalized the full book of probability that he, that he chose, uh, sometime in, in March, in, in Holite. So it’s just one commit in, in his GitHub repo. So, and, like, uh, like, uh, this is not what he does, but basically, like, just take a book and, and formalize it. Uh, then, uh, I, with, uh, three colleagues, uh, took, uh, well-known proof assistant called Isabel Hall. (..) And, uh, uh, basically repeated the experiment with general topology. And, uh, in, uh, something like 24 active days, we have basically done fully, like, without any gaps, all the theorems in the 39, uh, sections of, of this book. Um, of, of, of, of this general topology part of, of, of, of, of that book. (.) Uh, then, then we took, like, four agents and tried to parallelize this, this, this business to, to kind of speed it up using some kind of bounty-based system. So, so, so, so you can imagine that the agents, uh, get some initial money, some initial budget, and their goal is to compete, but also collaborate in, improving, uh, the theorems in the book. And they can issue sub-bounties for the, uh, for the hard, uh, proofs that they want to do so, so, so that other agent can help them, like, especially if the other agent specializes. uh, in, uh, uh, uh, something. Uh, so, so that way we kind of managed to, to speed it up quite, quite a bit. (.) And then the, the very last experiment, which I did myself recently, uh, was to take this relatively well-known book on algebra by Godeman, which has, like, I think about 600, uh, pages. And I ran it with the latest ChatGPT 5.6 on extra high reasoning. (.) And in less than nine days, it basically did the full book and resulting in something like 700,000 lines of formal proof. So this was, like, five days ago when it finished. But it kind of shows you how easy it is now. Like, I basically didn’t have to do almost anything. Like, I just really took codex with this extra high reasoning, gave it the LaTeX of this big book. And nine days later, it’s basically done. (.) So not to speak just about our group. So there has been, like, a very, like, significant formalization done recently, which was, like, a collaboration of humans and machines on a project that is about fields medal result by Marina Vyazovska. Yeah, so this was done, this was a bit controversial. So if you want, you can read this paper because the humans and the company which was doing the formalizations had some disagreements. And, like, like, I guess the upshot is that the company kind of finished the proof, which the humans started together with the company. But it’s a really, like, a massive fields level result. (..) And then there is this result from Facebook, I think, from May, where they took 26 open access books and tried also some kind of agentic setup. And I think they claim they basically created a huge library out of these 26 open access books. (..) So this is, like, from their paper. (..) So I took John Harrison’s slides because he’s really, like, the most senior person in formalization of mathematics and also formal verification of software and hardware. (.) So he wrote a bunch of slides from the point of view of Claude, talking about how Claude works with John on the formalization. And one of the funny remarks is that this, like, very simple, stupid feedback loop really seems to be working. So here is the next slide saying that basically what matters most rather than some kind of specialized coding of various tools is the general progress in the coding agents or the LLMs under them. (..) A bit more about this AWS hypervisor or isolation engine. (..) So it’s 330,000 lines of machine-checked math. (..) And they really verified quite a lot of important properties about it, which is obviously important for Amazon’s cloud business, like, for people not to be able to see what somebody else is doing on the virtual machine next door. So there is really, like, this has been a project that has been running at Amazon for, I think, like, four years with a very careful design, et cetera. Yeah. But, like, one thing which Dominic Mulligan, who was, I think, the leader of the project, at least on the formal verification side said, is that by, let’s say, January, the project had something like 250,000 lines of code. And just, let’s say, in this next half year of the project, they basically automatically added something like 100,000 lines of code using the coding agents and this kind of feedback loops between the coding agents and the proof checkers. So his comment was that now writing the proofs, which really used to be the bottleneck, which was kind of the main bottleneck preventing us from this QED utopia, is no longer the bottleneck. Like, writing the proofs became easy for the coding agents. Yes. (.) And the bottleneck is now the humans, so they are hiring like crazy, who need to say what it is that has been proved and how does it relate to the project, to your specifications, et cetera. (…) Yeah. (..) Yeah, I guess I said all of this. (..) So one more report from basically, I think, three days ago in Portugal. (.) So there was the annual competition of automated theorem previous, which is kind of getting a bit questionable because today you hear about all these LLM based or not just LLM based, but let’s say agentic setups solving hard math problems, which I may get into a bit. But like the symbolic ATPs as obviously a long history, it’s this kind of somebody calls it the good old fashioned AI. But like one update is that the competition this year was won by a system called Vampire, which is no news, like Vampire has been winning the competition pretty regularly. But the new version of Vampire has a very fast neural guidance in it, which is running on a CPU. So it runs basically under the same competition rules on the same CPUs as all the other symbolic solvers. (.) And this kind of trained neural guidance improved Vampire by, I don’t know, something like 20% or so. So you can see that the LLM based, let’s say guidance or learning based guessing or guidance for the proof search is obviously not the only alternative for automated theorem proving. So apart from, for example, so apart from, for example, this like very fast new neural guidance of Vampire where other improvements reported this year at the conference where, for example, people use, again, some kind of neural symbolic tools to create better abstractions, like better definitions, better lemmas, which they throw basically into the theorem or better conjectures or better conjectures, which they add into the theorem proving task and then they help the symbolic solvers in various ways. So, and obviously to make it even more interesting, a lot of this stuff can now be kind of developed or vibe coded with the help of the LLMs and the coding agents. (..) So, so, so that, that was like this bunch of examples. And then I guess I will get into some kind of theories and speculations now. (.) So, so, so, so one, one way how I’m thinking about the LLMs and I, I guess this is a big debate. So I kind of might not be super best informed person about all this huge amount of research done on LLMs. but like, for example, 30 years ago when I was kind of starting my research and combining automated reasoning and machine learning. Like one of the kind of high profile projects was Doug Lennart’s psych where basically Lennart, after doing a bunch of interesting things in mathematics, like AM for example, decided decided to build this huge system containing all the common sense rules and heuristics and facts kind of explicitly manually and symbolically. (.) And at the time I was thinking that this is like a very ambitious and maybe like fragile and that you, you could kind of try to replace the, that huge manual effort by basically machine learning by, by learning from let’s say theorem proving corpora the, the various, the various problem solving heuristics rather than programming them manually. (.) But, but it seems to me at least, but you can obviously discuss it, that what this huge recent LLMs achieved is basically something like Lennart was trying to create. that by basically reading the Wikipedia and all of the internet and whatnot. (.) They, they basically internally created some kind of compressed representation of, of, of these, let’s say, common sense knowledge basis and rules and heuristics. And, uh, I, I would say, uh, like, it’s, it’s kind of interesting how, uh, we can use these things, right? Like, like people at some point invented this kind of chain of thought, which, or tree of thought, which is, which is some kind of, uh, soft, uh, search on, on, on, on top of this huge, uh, database. And what, what we are now kind of doing, like, like when we use, for example, the LLM as a coding agent is, is to, is, is to just add some, some kind of proof checking, uh, component on top of that. And that way we, we, we kind of augment like, like, like this fuzzy, soft, incomplete, and possibly buggy, uh, database more and more with something that, that, that, that is like solidly, logically verified. And it, it, it, it, it may be thought of maybe as, as, as some kind of bootstrapping procedure that, that we started with some very buggy internet or Wikipedia. But by, by this kind of relatively simple process, we will now basically more and more develop, like, very non buggy or much, much less buggy, uh, knowledge basis, which will be represented in, in these, uh, knowledge base. (…) And, like, like, like one interesting example of, of, of, of, of the, like, very soft or fuzzy loops that people, uh, started to do is from last year, IMO, the, the International Math Olympiad, where two basically unknown guys, like, no, no big clap, took a commercially available model. I think initially they took, uh, Gemini. And they, they just programmed the feedback loop in, like, with one, uh, language model. So, so one was behaving as a solver and the other one as a judge or a critic. And, and, and, and they basically played this kind of ping pong between them, like, like, when the, the, the solver proposes some, some kind of part of the proof and, and, and the judge looks at it. And the judge is kind of prompted to, to be a, a very, like, meticulous mathematical expert. And, and, and, and he looks for bugs there. And if he, he, if he does, then, then it’s, he sends it back to, to the, to the solver. And, and this way, like, like this, the surprisingly simple, you could say soft loop, because there is no kind of hard, uh, logical proof checking involved. Really, really improved, uh, the, the commercial Gemini, uh, model to the level of, I, I think they got the golden, uh, medal, basically. (.) So, so, so, so this was discussed, uh, quite a bit. And I would say one caveat about that is that, for example, DeepMind or probably everybody, uh, trained their models on, on this, like, relatively limited high, high school mathematics, uh, quite a lot. (.) And it’s kind of unclear whether this kind of soft, uh, feedback loops could, uh, scale up to, like, research level math on which the models haven’t been, uh, trained, uh, so much. So you could try to argue that we, like, having kind of the hard, uh, logic-based, uh, proof checking is, is really, uh, necessary. Uh, to, I, I, I, I’m kind of guilty of one of these first kind of hard, um, uh, learning proving loop. So that’s this machine learner for automated reasoning or, uh, malaria from, which is today, uh, 20 years old. So, so, so, so the idea is really like this very, uh, simple feedback loop between the machine learner who, who guides some kind of proving procedure and, and the prover who, um, who tries to finish the proofs. And if, if it finishes the proof, then the machine learner basically learns from, from these additional proofs. and, uh, uh, again, like one impressive thing, which, uh, at least impressed me in 2025 is that basically like, like the kind of deep seek reasoning, uh, revolution to, to some extent, maybe, maybe not to such level of, uh, hardness, but quite a bit more than like this kind of soft loops, uh, entered this territory. And you, you, you, you basically, uh, have some kind of feedback loop, uh, like, like that. When you, you, for example, have like a curriculum of harder and harder mathematical or programming problems. Where you can like relatively easily check whether, uh, the problem, whether the solution, uh, with which the language model came up is the right solution or not. And that kind of gives you some signal. And even though you could kind of say that maybe this is not a perfect signal because the, the model could have guessed the, the solution in, in, in some kind of suspicious, uh, way. Uh, it’s, it’s, it’s a, it’s a pretty strong signal. And, and, and the kind of cool thing that you get out of that is that this, this deep seek or generally like with these big LLMs, you are basically no longer constrained by any logical calculus. So, so, so, so, so, so the game kind of evolved from, uh, the theorem provers, which are always constrained by, by a particular, uh, very well-defined logical calculus to this like totally unconstrained, uh, world where basically the, the LLM can basically design its own reasoning. Like, like, whatever, whatever will, um, get it to the solution will get reinforced, uh, during, like, these feedback loops. (.) Uh, so, yeah, again, this is just a slide from the deep seek paper. (..) Um, yeah, I, I just wanted to mention, like, one more experimental because two years ago I was talking here about our conjecturing group for inventing solutions for the online cyclopedia of integer sequences. So, so, so, so, so, so that feedback loop is, again, in some kind of a constrained scenario, but, but with some kind of constrained domain specific language for proposing the explanation for the integer sequences. Yeah. But in something like 1000 iterations, but in something like 1000 iterations, with a very, very tiny, uh, neuro, you can call it language model. But it’s like really like 10 million parameters, uh, it has now solved like almost 40 percent of the OEIS, um, corpus. So, it’s, like, 135,000 OER sequences. And, like, one thing that you can add to it is the theorem proving task. When, basically, when you get multiple explanations for the same integer sequence, you can ask the questions, can I prove these two programs to be equivalent, or can I disprove them to be equivalent? So, actually, in 2025, we have kind of augmented this conjecturing feedback loop to a loop, which has, like, an additional theorem proving aspect to it. (..) And I’ll just mention, like, in this kind of constraint slash unconstrained, so the previous thing was constraint, but, like, we have a constraint DSL. In the DeepSeq world, the thing is unconstrained. Like, you only have some kind of a signal, which, and the DeepSeq can do whatever it wants to do to reach the solution. And we are kind of hoping that, because it’s so hard to get to the solution, it will get there rationally, right? Like, whatever it will do will be some kind of a reasoning process. So, I’m kind of thinking, like, so this has been in the news, like, one or two weeks ago now. (.) So, somebody just unleashed the latest anthropic model, I think Fable or Opus 5, on a really old mathematical conjecture while watching the world championship in soccer. (.) And the model basically came up with a counterexample that has been later confirmed by other mathematicians. (.) But my point here, and this is no longer an isolated example, so in 2026, we are really getting solutions to open problems. And, obviously, people can quarrel, like, how close are these solutions to existing literature, or how hard are the problems, the quote-unquote open problems, et cetera. But, like, what I see here is something similar to this DeepSeq phenomenon from 2025, that basically, as the models are consuming the knowledge and training more and more on our data in better and better ways, because they are kind of approaching this level where they will basically be producing new results, which we will be able to verify, either the mathematicians will verify it or we will translate it to using these auto-formalization tool chains that are now working pretty well. And that will provide more and more data, like, like, like, high-quality data for training the language models or other machine learning systems even more. So, I see this as some kind of beginning of this feedback loop basically in the wild now and I think it’s kind of unstoppable. (..) Yeah, so this is the paper about this. (..) Yeah, so I wrote a couple of slides about why I think, like, this hard logic-based proof-checking still matters. And it’s mostly, again, based on this John Harrison slides where, even though he says that the LLM can do, like, very advanced things, it can also do, like, very elementary logical mistakes, it behaves a bit like some kind of brilliant intern, which benefits from being coupled with a grumpy accountant. (..) So, so, so there’s, like, a longer slide from John about that. (.) And here’s another funny, funny slide from John where you, you, you could, so, so, so the kind of way how humans have been formalizing mathematics have been developing, like, a fancier and fancier user interfaces and all sorts of tools that kind of help, help the humans, etc. etc. But what, what we have seen, like, recently with these feedback loops is that basically the simpler the setting is the better for the progress in this kind of auto-formalization feedback loops. Yes. (..) What is also interesting is that what we see in these, like, whole book auto-formalization experiments is that it’s, it’s not, like, magically perfect from, from the beginning. But the whole process is some kind of a huge debugging process where, let’s say you initially read the book and the book is kind of, or the LLM is slightly negligent or the book doesn’t have it in it and some theorem or definition misses some assumption. And you, you, you, you really only learn that when you are using that assumption or in, in some kind of proof. So, so, so, so, so what seems to be really important when we are doing this is not to do just the auto-formalization of the specifications of, of, of, of the top level definitions and statements. But to really, uh, basically, uh, basically, uh, basically, uh, run, run, run the proofs, which in some kind of symbolic way you could say back-propagate to, to the statements and definitions and, uh, like, like, basically do debug your specifications. (…) Um, yeah. (.) So, so here is, uh, like, some, like, one speculation. Uh, so, of course, like, like this invention of symbolic logic, uh, is relatively new. Like, it’s like 100 years or so. And people have obviously come up with all sorts of fancy stuff already. Like, I know, starting with first old logic, high old logic, type theory, set theory, et cetera. Uh, but I think, uh, like, we will now basically think, think to these agentic loops. (.) Uh, and again, looking a bit at, for example, this, uh, crazy deep seek style, totally unconstrained, uh, RL frameworks. We, we will see a development of, like, new logical, high-level, uh, languages, like, languages, which will still be precisely specified and, and kind of compilable to, to the more, uh, low-level languages. Uh, but they, they, they will, they will have all sorts of advantages and maybe also map, uh, better to, uh, to, to things like the human natural language or at least the human, uh, mathematical language. So, for example, there is this project Malinka in, uh, France where they have a textbook where basically their goal is to start understanding more and more and higher and higher level of, of the mathematical grammar or mathematical language. And I, I, I, I almost see, see, see this as some kind of vengeance or revenge of, of the linguists who you probably know, know these sayings by Fred Jelinek that each time he hires a linguist, et cetera. So, so, so, so I, I think now, like, like kind of paradoxically and assisted by the LLMs and, and the coding agents, the, the, the linguistic theories and, and the theories of, of grammar and, uh, how, how the language has, uh, kind of explainable structure and, and semantics will basically prosper, uh, again, but that’s kind of a speculation. Um, yeah. Um, yeah. And I probably already mentioned it. So, so, so I, I, I think there, there, there will now, now be like a proliferation of, um, languages and formal languages and not, not just, um, languages, but also proof checkers. So one colleague of mine now basically wipe coded a totally new, uh, theorem prover and proof checker, and there will be like, uh, all sorts of on demand translators and presentations. So, so, so, so for example, I have wipe coded very recently, uh, presentation of, of, of this huge, uh, algebra book that I have, uh, formalized. And it seems really easy to, to kind of create, uh, presentations on different levels. So, so that people with different, uh, level of expertise can, can benefit from, from that. Um, yeah. (..) Um, yeah. And I, I, I, I probably already said that. So, so, so I think the, the, the, the agents, like, like, for example, in this kind of deep seek unconstrained world, uh, are, or will be able to, to, to basically do this generalized unconstrained search. which includes the, these, these action spaces like write a dedicated solver for a subproblem or, um, propose some new definitions and representations, et cetera, et cetera. (..) Um, so, so, so, so I had a idea that this could, to some extent relate to Schmidt-Huber’s Gettle machines, which is kind of an AGI-ish topic. Uh, but I will probably not say much about it. So, so, so in, like, like the original Schmidt-Huber paper, which I may have read long, long time ago, a bit. You, you have, uh, a machine, which basically, uh, rewrites itself when it can prove that the modified system is better than the current system. Uh, so we already have had that in, uh, let’s say automatic. automated and interactive theorem proving for a while. So Maireen, Magnus Maireen and Jared Davis, that is Milava and Jitava, uh, I don’t know, maybe 10, 15 years ago. Where the machine can upgrade itself if it kind of proves that some kind of useful extension is, uh, sound. But, but it’s not exactly what the, uh, Gettle machine is doing because it’s, it’s not proving that it will be better. It’s just proving that the extension is, uh, sound. (..) Um, but, but like what, what, what I’m again seeing, like, like this kind of unconstrained, uh, space, uh, which, let’s say these feedback groups are creating today. is that, uh, we will have this kind of self-improvement when the, the soundness is, is being proved and the utility is probably not proved, uh, but this kind of measured, uh, heuristically. So, so, so that, that’s my kind of, uh, estimate in, in this kind of Gettle, uh, Gettle machine idea. And I already mentioned that we have basically done some, something like this. And, and, and this, uh, loop, which we have over the online cyclopedia of integer sequences. So, so, so for example, one thing which I was presenting two years ago is that in this one, uh, 1,000 iterations of, of, of this conjecturing, uh, loop, we have something that we call the human jumps or the human cost jumps and, and the automatic jumps. So the automatic jumps is when some symbolic breakthrough is discovered, like the prime numbers or whatnot. Whereas the human jumps is when we, the programmers improve the programming language by, for example, giving it arrays or something like that. But, but, but, but, but I think like the, these two things with this, like big machines, which are the LLMs will now basically merge and, and this will be kind of co, co developed. Uh, yeah. So, so, so again, speculation. Will we have specialized tools or big, uh, chat GPTs? I, I, I think we will, we will have like a ecosystem of containing, uh, all, all these things and basically everybody, uh, will prosper from, from that. (..) And like here, here are some kind of caveats. So, so for example, in this auto formalization business, there are all these, uh, things that we have seen. In, in, in, in the, in the, in the last half year, which are some, some kind of caveats that things are not perfect yet. Uh, and as I said, for example, the Amazon people are now hiring more and more, uh, formal proof people, not to write the proofs, but, but to audit the proofs, et cetera. (.) Uh, yeah. So, so, so here, here is a bit of a forecast. So here are things which already work. (.) And yeah, I, I don’t know, some, some things are pretty, uh, speculative and that’s probably it. So, so, so, so, uh, is the QED utopia upon us? Uh, so, so in some sense, in some sense, no, like we are not yet in, in this like super comfortable verified, uh, world, but, but I would say in some sense, it’s, it’s almost here. (..) Uh, and I, I, I think if, if, if we are not there to today, we will be there like by the end of 2026 or 2027. So that’s it. Thank you.

S34: Thank you so much, Joseph. So we’ll start with questions. (…)

S17: First of all, thank you very much for your talk. This being the 70 year anniversary of logic theorist, the dawn of AI, the anthropic code leak, telling us all these things are selling are not just pure DNNs. Your talk is quite timely. Of course, you know that. Um, I did think, and I was sharing this with my friends here at the table, that I would be the only person that was at Flock and this conference. I’m very happy to see that I was mistaken about that. I didn’t have to present three. I didn’t have to do three talks like you, fortunately. Uh, I just had to show up for a one day workshop. So I just want to mention that. Uh, my question’s coming. Uh, and I, we also have a theorem prover, uh, that we use called shadow prover, which I’d love to get your feedback on. My question, a little broader and all that. Um, do you believe that the singularity coming to pass is logically defensible? Because I’ve seen a lot of proofs that are out there that say it’s not. Uh, I haven’t seen any proofs that say it’ll singularity will indeed come to pass. So I’d love to get your thoughts on that. Thank you. (.)

S09: Oh, could you, before you hand over the microphone? Uh, you, you mean the general singularity or the QED?

S17: David Chalmers singularity.

S09: Uh, yeah, you need to update me a bit. I’m, I’m, I’m, I’m not so much of an AGI person too.

S17: Oh, Ben probably would summarize it better, but we reach human level, level intelligence, exceed it. Then that just keeps replicating itself, getting more and more intelligent. That’s my understanding. The singularity will just magically have machines that keep getting smarter and smarter and smarter. Uh huh. I think that’s pretty much the same.

S09: So, so, so I guess the, um, the alternative would be that we don’t even reach the human intelligence or it somehow plateaus and like it, it doesn’t really grow linearly or, or whatnot. So I, at, at this point, uh, the, the gradient is, is really quite interesting. So I may, maybe a couple of years ago, I, I would be saying other things. Uh, but, uh, to, to, to today, like I, I’ve been doing the experiments on these small mobile models. Like for example, in automated theorem proving or in, in this conjecturing from the online encyclopedia of integer sequences. I, I, I can see, uh, for example, quite a bit of linear growth, like even let’s say after 1000 iterations of this learning conjecturing loop on, on the online encyclopedia of integer sequences. I, I still see kind of linear improvement. So, so, so now if I kind of believe that we are basically now in such a self improving loop with kind of everything or with mathematics, like for example, this Jacobian, uh, problem being solved. And more and more problems now being solved by mathematicians. And the data being basically immediately fed to chat GPT and anthropic, like they, they are getting massive amounts of theorem proving data from my experiments. Like I, I give them terabytes of data for free. Uh, I, I, I, I think at least in the short term that there is a potential for quite a lot of self improvement. And I, so, so that’s estimate based on my kind of small scale experiments. Like I, I, I don’t see what happens in the long horizon. Yeah. (….)

S25: Yeah. Thanks for this presentation of really impressive results. My question is, um, if you consider this bottle net that you mentioned at AWS that you need humans to write informal specifications of say software requirements. And also if you have these informal specifications, you need to formalize them. And I think there you, your loop does not apply because you don’t have kind of, it’s, it’s not verification. It’s validation of informal against formal. So does this impressive work help with first, uh, validating formal against informal and perhaps even with writing the specifications or writing new math books? You know, where are we there?

S09: Yeah. So again, I will shamelessly cite my OEIS experiment, which, which I had, uh, talk here about, uh, two years ago, where we can see that like this very simple conjecturing plus, uh, testing loop. You, you basically test whether the conjectured program explains some sequence. You, you, you don’t even do the theorem proving, uh, has, has the potential to, to come up with new ideas or like, uh, so, so for example, after iteration 40, you, you suddenly come up with the program for primes, which you didn’t have before. So I, I, I have like, since I have been running that experiment, I have been always replying to, to people who say that the LLMs are just memorizers and, and they don’t have the ability to, to generate anything new. I, I, I, I tell them, like, look at, at this, like very simple experiment that I have been running like this, like very small, uh, language models and how, how far it got there by some kind of self, self play. Uh, with, uh, with like zero human, uh, with like zero human, human knowledge. So I’m, I’m, I’m, I’m not very pessimistic about the machines not being able to come up with new ideas, new, new specifications. (.) Especially if we have the feedback loop, which validates them, which, which kind of proves that they are good and, uh, at least test them, et cetera. The, the, the kind of first part of your question, like whether we will be able to replace the human validators who translate, let’s say the, some kind of informal specification into like the formal specifications. Uh, I, I would say it would be very dangerous for somebody like AWS now to, to do that. That’s why they are hiring people like crazy. Uh, but I, again, like I, I wouldn’t be so pessimistic. Okay. I, I would say you can develop a lot of tools, which will help automate this, this kind of stuff. (.)

S25: Thank you. (…..)

S33: Oh, thank you for a talk. So one of the questions I have is about the cost.

S24: So how much would it cost to prove a thousand lines of code? (…)

S09: Oh, so, so that’s very interesting. In, in the first paper I did, uh, in December, uh, like the headline was that I, that there’s 130,000 lines of general topology. for $100 because I did it with the chat GPT pro subscription, which costs 200 and I did it in two weeks. But then I measured the API cost by running some tool over the session file. And the API cost is, I don’t know, somewhere between $5,000 and $10,000. And now you, you get two numbers which differ by two orders of magnitude. And like, which of them is the right number. So, so one way how to think about it is that, for example, Amazon is running also deep seek. And that’s like an open weight model, which Amazon doesn’t need to pay anything for. So, so presumably you, you get only the, the kind of energy cost plus some margin, uh, for, for the Amazon people. And like my estimate is that, I don’t know, running deep seek is maybe like 10 times cheaper in the API costs than, than running some anthropic model, et cetera. So, so, so maybe the truth is somewhere in between. Maybe the truth is like 10 times more. The sub subscri, subscription cost and 10 times less than the API cost. But that, that’s kind of just a speculation for me.

S24: Thank you so much.

S02: This matches my observations as well. Okay. (……)

S27: I have an additional questions. (…) I have questions. Yes. Uh, about, uh, your experiment with the harmonization of groups. Do you actually, do you file station to language people’s correct or you don’t do it by hands? And the second question is more abstract. Can we look at that? (.) Can we formalize all? (…) Can we formalize all mathematics? (…)

S09: Yeah. So I, I do try, especially in the first experiment, but, but also in most of the other experiments, I do try to check the, the main definitions, the main theorems, et cetera. Like one, one advantage of formal math is that, let’s say, if you check, if you have some headline theorem, like Fermat, Fermat’s last theorem, it’s very easy to state. So you can just check the definitions that feed into the statement. And for, for Fermat, you, you don’t need to check basically anything else. If you trust the formal proof checking, uh, technology, which we kind of do because these systems are crazily, uh, developed today. So, so, so, so the, this, this is like a major factor which helps us with checking the, these hundreds of thousands of lines of generated proof because we know that it’s enough to, to only check some of the things and, and the rest is guaranteed by, by the formal proof setting. Uh, the, the, the, the second question is, if I understand it correctly, when will we formalize all of the math? Uh, so I’m really, so, so this last experiment, which I did, I am not Facebook. I am not open AI. I am not deep mind. I, I just paid for nine days of usage to, to open AI. AI. And I formalized the 600 page book on algebra. (.) So if people like me do do this, we, we, we, we might have like thousands of books, uh, formalized this, this, this year. And it’s, it’s, it’s, it’s, it’s getting better and better. Like what take, what took me a month, half a year ago, took me nine days or, or even less, uh, now. So I, I, I think of, of course, like there are some blockers. So for example, you can say algebra is relatively easy. (.) Algebraic topology is harder. You have many pictures. Geometry is, is even crazier. So it’s not like we have the, the, there are some mathematical proofs, which might not even count as proofs, right? Like there are gaps in them, et cetera. So, so there, there is lots of issues, but I, I would say I’m, I’m pretty optimistic. At this point. (………)

S23: Yeah. Thanks. Thanks for this. I want to follow up on the previous question because it’s where automated fear improving intersects with AGI the most, really, because I think auto formalizing is great. And one thing we’re seeing is it’s, it’s not an AGI hard problem, right? Like LLMs combined with formalizers and automated fear improvers can do it, which is, which is awesome, right? And some conjecturing, which you touched on recently, is really where automated theorem proving seems to intersect with AGI more so, like coming up with like, what are the interesting new theorems to look at? And of course that intersects with theorem proving if you’re trying to prove that the Reimann hypothesis or something really hard where you have to make a bunch of weird new conjectures for what the, what the lemmas might be on route to the, to the proof. So mostly I wanted to prompt you for some daydreaming or speculative thinking on like what, what new ingredients do you think need to be introduced to the sorts of frameworks you’re talking about to handle creative conjecturing, which is, is half the fun of doing math. Right, right, right, right, right, right, right. At least. (..)

S09: Yeah. So I, I wipe coded a conjecture, like, like five, five different neuro symbolic conjecturing architectures about a month ago. So, so I’m kind of actively experimenting with wipe coding, various conjecturing architectures. And I would say, like, one thing which kind of fits the AGI conference, I guess, is that I think the, of course, like, the fact that the LLM people realize that reasoning is good, like this kind of deep-seek revolution, like, they need to be praised for that, that they kind of abandoned this kind of nonsense that the LLM just learns all sorts of algorithms and da-da-da. (.) But I think this kind of combinations of, like, the, let’s say, symbolic abstraction ideas, and they can be implemented differently in different neural architectures, or they can be emulated by the LLMs, which can, like, through the chain of thought emulate practically anything, like, up to some epsilon, I would say. (.) So I think, so one totally trivial idea is that, like, you have these kind of prediction-making based on related theories, right? Like, what, what holds in, in ten different theories, which are differently close to, to, to your, uh, to your target, uh, theory. And people in automated theory are really doing, like, pretty amazing stuff using, like, this very, very simple idea, and, and for many years. And this somehow, at least, as far as I know, hasn’t been, uh, um, uh, merged yet, or implemented yet, uh, at least not explicitly with, for example, the, the various LLM co-conjecturing workflows. So, and, like, like, like, you, you can kind of push it further, so one thing we did with this agentic auto-formalization is basically some kind of idea of prediction market. So, so, so, so in, in, in the extreme, you could, like, like, people have proposed this long time ago, you could have both, um, models or AIs and, and humans, uh, betting on, on, on, on, on the conjectures on basically how, how, how the, uh, proof will most likely, uh, be achieved. Uh, so, so, so, so that some kind of infrastructure for, like, the economics of, of, uh, conjecturing. So, yeah, I, like, like, my, my speculation is that there are still, like, many ideas coming from, like, non-neural approaches, like, like, abstraction, uh, et cetera, which will allow us, especially in combination with the statistical machine learning, uh, make quite a bit of a bit of progress in conjecturing. (…..)

S23: Yeah, that, that makes sense. I guess when you’re saying finding common elements among proofs of similar things, I guess the question is, at what level of abstraction do you find that commonality? If it’s a very low level, then that is a common ATP thing already. If you’re trying to, if you’re trying to form higher level abstractions from the proofs, that’s where things get interesting, right? Because it’s like, it is the, is the knowledge representation you want for proof strategies, is it the same as a knowledge representation for, for small proof tactics, right? And that, that, that is, it’s not entirely clear to me. I don’t know if it’s clear to you. (…)

S09: Yeah, I, I, I, I guess I agree. Like, if you’re, uh, live coding or code coding some very low level, uh, let’s say, automated theorem prover working in some kind of resolution, superposition, calculus, uh, conjecturing, then it’s likely to be quite different than, than if you are conjecturing for, uh, the Riemann hypothesis or, uh, something like that. So, so, yeah, I, I agree that the representations will be different and, like, like, like, I, I, I probably didn’t speak much about this whole hierarchy of, of, of, of the problem solving in, in, in this domain. So, so let’s say the symbolic solvers currently are pretty good and can be even improved in, in doing, like, this very fast, but quite low level proof search. But we also do have some methods how, how, how to make them kind of more and more high level. But then, obviously, you, you can also really play, like, these conjecturing methods on a really high level. Yeah. (..) Well, like, basically in natural language, right? Like, so, so a lot of the kind of LLM based, you can call them theorem provers are actually doing that, right? Like conjecturing in, in some very high level space. And, and then if you want to do a formal proof, then, then you gradually, uh, specify that. (…)

S34: The last question right here. (.)

S32: Uh, hello. So, yeah, I can say hello. (.) I, I thought the talk was very interesting. Um, I was kind of thinking about, uh, next steps kinds of thoughts. For example, it’s very, I mean, it’s, it’s incredible that we can, we can formalize a book or we can, we can prove conjecture. But when, when you hear that the formal, like, proof is 600,000 lines of code, obviously for a human, that sounds huge. For a machine, less so. But, um, what’s the, like, logical detour, you know, factor compared to what you’re actually trying to get? And how much do you think we can improve this? Cause a lot of, uh, human, like even mathematics itself is this whole work of compressing, compressing, compressing until we can, you know, make use of it. So. (.)

S09: Yeah, I, I totally agree. Like I, I, so, so there are two things like, like I kind of quickly mentioned that I have already wipe coded some kind of high level presentation mode just for myself to, to, to be able. To, to, to, to see what’s going on in, in the formal proof. So, so that I exactly don’t need to go through 600,000 lines of code. But, uh, the, the other thing which I mentioned and which I think will be happening more and more is that like we, we really have been doing research in formal proof and symbolic logic for some 100 years. And compared to the evolution of natural language, it’s, it’s, it’s like nothing. So, so, so I, I think we will really now see a lot of the development of like higher level logically grounded languages, which at least some of them might kind of address this problem. So, so for example, already the, the Amazon people who verified this hypervisor, they, they, they are kind of evolving a particular automation tactic, which has already seen a lot of development in, in their project to fit more, more and more problems. And it’s to, to some extent manual, but you, you can kind of assume that this will be like more and more automated. So, so you will, you will really have tactics which will say something like very high level, like by analogy or something like that. And, and, and they will have a very precise meaning. You, you will be able to look at their code and ultimately they will be grounded in, in the kind of low level logical calculus. Yeah. (…)

S34: Thank you very much. Thank you very much, Joseph. That was wonderful. (…..) The emergent ability of LLMs trained on language to do math makes me wonder after, uh, how to talk if the internet isn’t just very high quality noise. (.) Um, so we’re going to have a break for lunch. We’ll be a 30 minute break. We’ll be back at quarter after one and we’ll continue the paper session there so that we get to look at intelligence as formal structures and inference. Something to tantalizingly look forward to as, as you eat. Thank you very much. (……..)

S06: Hey, welcome to my talk. I have the honor of, uh, having the talk right after lunch, uh, where you are still chewing your pizza. (.) And, uh, I’m, I’m, I’m sorry to distract you from this. (…..) Yeah, but my hope is to make it through before the sugar spike kicks in so that you can still listen. The other, uh, strategy to attract your attention, uh, I decided to choose this, uh, great title implication as algorithmic containment tantalizing, isn’t it? (..) And, uh, so I want to talk to you about reasoning. Um, we have major, two major approaches like neural nets and like some, some symbolic reasoning and, um, symbolic reasoning. Well, neural nets often are quite bad at it because there’s a lot of hallucination or confabulation in LLMs. And, uh, if you think about, uh, how people reason. If I, if I show you this bottle and ask you, well, if I open it and tip it and you know there’s water in it, you’re gonna assume the water will flow out. So how, how do you do that? It appears like you imagine this bottle, even if you don’t see it, and play this video in your mind. And then you see that water flows out and you do perception on your imagined picture. And in this perceptual process, you extract the property that water is then on the ground. (.) So this is what I mean by grounded reasoning. And I would like to suggest a third alternative. So apart from the way it’s done in neural nets and also not in the way it’s done in symbolic systems. (……) So in symbolic systems, what, what is a symbol actually, right? Like, what is a symbol dog? How does it relate to a real dog or the, or the image of a dog? (.) Well, I don’t really know how, apart from the letters dog being a cue in the human mind, that remind us what a dog is. So it’s basically cues a representation of the human mind. Essentially a symbol is, uh, the mean, so the meaning of the symbol is actually extrinsic to the AI. It’s outside of the computer. It’s inside the human mind. So for the AI, you could actually replace the symbol by any other symbol. And that, uh, notion is, uh, expressed by, uh, by Harnett in the famous symbol grounding problem. How can the semantic interpretation of a formal symbol system be made intrinsic to the system, rather than just parasitic on the meanings of, uh, in our heads? How can the meanings of a meaningless symbol, tokens manipulated solely on the basis of the arbitrary shapes, be grounded in anything but other meaningless symbols? (..) So I want to argue that symbols are not really representations, because they do not contain any information about what they represent. What do I mean by that? (..) Well, um, the alternative, a representation of an image of a dog, would be a bit string, where the image of a dog is also a bit string underneath. And, and when one bit string has information about another bit string, what do we mean by that? Well, it means that there is a relatively short program, like a Turing machine, that takes the representation and computes what it represents in an easier way than it would, than without the representation. So the, in, in the language of algorithmic information theory, the Kolmogorov complexity, oops, what did I do there? (.) The Kolmogorov complexity, which is the shortest program computing X from R, would be shorter than the Kolmogorov complexity of X, so the shortest program of computing X. So, in some sense, a representation has to be somehow helpful in computing that thing that it represents. (…) So, what can we do with this? (…….) Symbolic versus grounded reasoning here. The other problem that symbolic reasoning has is, is this combinatorial explosion. If you want to make, if you have a rule from A follows B, from B follows C, from C follows D, there are all sorts of other distracting rules that you may want to, you have to traverse in order to find D. So, you have a combinatorial explosion. What I rather want to suggest is, if you have a generative model A, so properties should be generative models or world models, then you can just sample from it an instance of your, of what, what you want to represent. In this example, an equilateral triangle. And then, you do perception on that instance. And you perceive the consequence and derive the implication. So, how could this might work? How can we be sure that this is always true or something like that? So, in the formalization, here, F would be the premise. P, the parameter. Everything else about this triangle, like its position, size, its orientation. (.) And X is just a particular instance of the triangle. (…) And we just search, for example, if you want to search for another properties, like in the statement, every equilateral triangle is an isosceles triangle. G would be the isosceles triangle that we would like, and we derive that property from this instance. (.) So, how might this work? So, first of all, I would like to show you how in my system that I’ve been developing over the last year called William. It’s a data compression system based on program synthesis. So, in this example, you give a triangle example of an equilateral triangle, like a list of coordinates of three lines. And then it discovers 0.123. And then it discovers a difference vector. And every step, it compresses the data. (.) And then it discovers, transfers the difference vector in polar coordinates. And then it discovers that the lengths are equal. Two of the lengths of the triangle are equal. And then, finally, it discovers that all angles are equal to 60 degrees. And the compression process basically discovers this property of equilateralness. (.) And the same thing you can do with the isosceles triangle, where basically the same properties are discovered here, except for that the angles are all equal. So, we have just two sides of equal lengths. That is what this discovers. And you can notice from the shape of the graph that it’s discovered, that the description of the isosceles triangle is already kind of very similar of the equilateral one. (.) And that gives you a hint that one is implied by the other. (…) Let’s move to the main theorem here. What do I mean by algorithmic containment, then? Algorithmic containment of the string G in the string F means that there’s a short program computing G from F. So, here in this question, in this example, you can take the F is 0, 1, 0, 0, 0, 1. And G is just a flipped version of it, 1, 0, 1, 1, 0. And there’s a short program computing one from the other. And that is what algorithmic containment means. (.) In our context, it means that we don’t really have to write down the consequence of our rule. If we have a premise, we can start discovering the consequences from the premise. Because we can compute them from the premise by using this process. (.) That’s what implication actually originally means in the colloquial sense. (.) Etymologically, implication derives from Latin implicare, which is involve, entangle or to fold in. So, it’s there without being explicitly stated. (…) My suggestion is that we come back to that original meaning. (.) And in terms of the theorem, if you have a feature, which is a property that compresses your data, X. (.) This is a Turing machine, takes some parameters, P. And if you have another feature, G, that is contained in X. That is true for all features, as I show in my theory of incremental compression. All features are contained in the instance, like, my body size is contained in the image of me, for example. Because you can compute my body size by taking an image, a picture of me. (.) And additionally, if the parameter does not contain some… If you don’t smuggle any information in the parameter about the consequent G that you want to conclude, then you can actually follow that your consequence is contained in the premise. (.) And how does it work? Like, schematically, if you have split the information inside your sample into F and P, happens automatically when you do incremental compression. (.) the information inside X is split into two pieces that do not overlap. (.) And if P and G do not overlap as well, then G is also contained in X. So the green information has to be contained in the black box, but it cannot overlap with P. Well, then it must be contained in F. (..) That’s a graphical representation of how the theorem works. And we can see that how it can work also in practice. (…….) What happened to my slides? (15 seconds pause) All right. What we have established now is that for some particular instance, we have this implication. But ideally, right, if you want to really do replace the implication rule, we would have to show it that that’s for all instances is true. Well, here we can just do a simple thing. We just sample. If you want to know that every equilateral triangle is also isosceles, you just keep sampling several random equilateral triangles, discover the same thing about them. And then you can show that the fraction of cases where this is true converges to one with this speed. So it converges to one exponentially if you pick a couple of really non-random samples. So if you don’t, if you just sample random parameters, random equilateral triangle, you will cover the majority of the space. It doesn’t guarantee you that there are no exceptions, but you cover large parts of the space. (…..) And here’s how it works in my system, William. Here’s the graph for the collateral triangle. Here’s the one for the isosceles. And you sample a parameter. And then, oh, there should be an animation. (..) Okay. (.) There should be values propagating from here to there. And this should be inverted here. So basically, the system discovers or verifies the isosceles property. If you do it vice versa, if you have here an isosceles triangle and here an equilateral, it will fail here at the last point where it discovers usually that all angles are equal. Well, in this case, they won’t. So the propagation fails. That’s how the system actually verifies the statement. (..) So, and another cool thing about it, the question is how do we discover what actually follows from our premise, right? Since we don’t have it given. Well, if you have an isosceles triangle here and drop a perpendicular on its base, you discover, oh, it cuts the base in two. (.) Well, that’s not a random thing. If you take a random split at the random position, it’s not going to be split in half. So, you have here a non-random property. In terms of algorithmic information, you will be able to compress the data a little bit. (.) And that amount of compression, here, the average number of bits that you can compress the data with, (.) means that the number of samples you need to test your implication is going to be lower. (.) The better you compress, the fewer samples you need to test your implication. So, you can just use compression criteria to come up with implication statements that are likely to be true. (…..) So, now the important question, well, what if they are rare examples? (.) Well, for example, three points form a triangle. Well, unless they are collinear. (.) Or two random lines intersect unless they are parallel. Or two circles intersect in zero or two points, well, unless they are tangent to each other. So, do you notice that this, the exception, has always somehow some regularity to it, right? (.) So, the distance equality or equal slope or collinearity. (.) Regularity means compressibility. (.) So, you actually, there’s the hypothesis, well, maybe in many cases, exceptions have a compressible structure. So, by looking for compressible properties, you can actually find exceptions to your conjecture. (..) And there’s also a theorem to it. I’m running out of time, I’m afraid I’m not going to go into it. But here is an example with system William again. Let’s test the statement that all angles in the triangles are smaller than 179 degrees. Well, if you random sample triangles randomly, most of the time you won’t discover an exception to that rule. So, random sampling requires 116 samples to discover a contradiction. But if you sample the triangle points, the coordinates, in a simplicity-biased way, that means instead of picking random numbers like here, (..) you have simple numbers here. Where, for example, all three x values are zero. That is a simple thing. You don’t need much description because all three x values are the same. Your description length is slower. Well, what you have discovered, well, if they’re all x values are zero, they’re all on the line. It’s a collinear, these are collinear points. So, it’s a degenerate triangle and exactly the case where the statement fails. So, even in this example, in the first sample, it disproves. (…) So, I summarize how should we do the grounded reasoning? How do I suggest to do it? And discover properties by compressing it. (.) Then connect those properties by sampling. Just take one property, make some samples, discover subsequent properties, and make several samples to test it. (.) If your consequent properties compresses data well, you need fewer samples to test it. It’s a big hint that the implication might actually be true. (.) And challenge, in the end, your implication by simplicity bias, search for exceptions. (.) And then, it’s, of course, you cannot never be sure that you have really 100% certainty that the implication is true. It’s not a panacea. Exceptions are not always simple and are not always even computable. But, let’s remind ourselves, why do we want 100% certainty? (.) It’s because we do chaining. We want, in formal logic, we want to chain one rule after the other. And if we have even a small uncertainty, the error will multiply during the chaining. (.) And you get that problem where? Exactly in the non-empirical disciplines like math and philosophy, where you’re not allowed to look at the world. (…) But, for working AI systems, we are allowed to look at the world. We can say, we don’t need chaining. We just directly sample our hypotheses and test them. So, we can afford to be wrong 1% of the time. (..) So, good compression is also not the same as semantic clarity. That’s also a drawback. It doesn’t mean that it’s completely, clearly interpretable what you get if you have a program that compresses things. But, in symbolic logic system, you have no connection, a priori, between the symbol and what it represents. So, just because you write down a symbol that has clear semantic clarity for you, you still haven’t connected it to the world. So, all the work is still ahead of you. (…..) So, what about complex concepts? (.) Not just Euclidean geometry or abstract concepts. (.) Well, these are difficult things. With grounded reasoning, it doesn’t give you the solution for that. But, what it does, it forces you to actually connect your reasoning system with the world. And sometimes, if the world is complex, you have to have complex models that represent it. But, that’s what reasoning requires if you really want to reason about the world. (.) So, okay. (..) Yeah, I think I’m basically done. I think where the AI is moving now towards world models, actually, is a good direction. (.) And if we really add compression criteria to them, that could help us make great advantages in reasoning. (.) Thank you for your attention. (……..)

S34: Thank you very much, Arthur. Next, we’ll have a pre-recorded talk. (.) Fernando Corbacho, Pablo de los Riscos, and Mikael Arvib. Proposal for an AGI formal comparative framework based on category theory. (..)

S33: So, I’m Mikael Arvib. A proposal for a formal comparative framework for AGI architecture based on category theory. (..) The central question is simple. Can we describe and compare very different agentic architectures within one common mathematical language? (..) Today, AGI research is extremely diverse. For instance, we have formal learning, active inference, causal agents, LLM agents, among others. Every community uses its own concepts and mathematical formalisms, and we have a lot of different benchmarks, cognitive architectures, and even the common model of cognition. But, we still lack a precise language for comparing architectures themselves. Our claim is that category theory can provide exactly that language. Not to replace existing approaches, but to study their properties and describe their structural relations using a common mathematical framework. (..) So, what is our answer to this problem? Instead of viewing an architecture as a particular algorithm or implementation, we view it as a structural theory. In other words, an architecture specifies how computation is organized, what knowledge exists inside the agents, and how both are related. (.) We capture these ideas through three complementary layers. The syntactic layer, the knowledge layer, and the constraint layer. (.) The syntactic layer describes the computational structure of the agent, its operational modules, interfaces, and how information flows between them. The knowledge layer describes the internal knowledge of the agent, and the operations over them, like for example, creation, transformation, and reuse of the knowledge. And finally, the constraint layer connects these two worlds by specifying the design principles that every implementation of the architecture must specify. (..) This separation allows us to compare architectures independently along each of these dimensions, rather than treating them as a monolithic system. (…) Now, let’s look at the mathematical foundation behind these layers. We model both the syntactic and the knowledge layer as three hypergraph categories. Why hypergraph categories? Because they provide a natural language for describing compositional systems, where independent components can be connected, reused, and combined through structured wiring diagrams. (.) In particular, the Frobenius structure naturally captures operations such as copying, merging, creating, and deleting information, which are fundamental in many agent architectures. (.) Although both layers share the same mathematical formalism, they describe different aspects of an architecture. (.) The syntax layer represents the compositional organization of the computational process, while the knowledge layer represents the internal knowledge structures and all the transformation that can act on them. (..) This common foundation will allow us to relate both layers in a principal way, through the constraint layer. (..) So far, we have described the syntax and the knowledge layers independently. However, an architecture is not just two isolated structures. The key question is how they interact. (.) This is precisely the role of the constraint layer. It consists of two complementary components. (.) The first one is the relational interface, which we formalize as a profound tool. (.) Rather than defining how knowledge is represented, it simply specifies where syntactic components are allowed to interact with knowledge. The second one is the constraint system, an indexed family of mathematical constraints that every realization of the architecture must satisfy. (..) These constraints capture the assumptions that distinguish one architecture from another, such as Bellman consistency in reformed learning, causal assumptions in causal agents, or Solomanov priors in universal AI. (.) Together, these two components connect the operational workflow with the knowledge, while remaining independent of any particular implementation. Once these three layers are defined, we can give a precise mathematical definition of an agent architecture. An architecture is simply the triple consisting of its syntax layer, its knowledge layer, and its constraint layer. (.) More importantly, we can also define morphings between architectures. These morphings are mappings that translate the layers of an architecture into another, specifying the transformation in both. (..) These definitions lead us naturally to the category of Art A. (..) The figure on the right illustrates what we call a topological space of architectures. Here, each node represents an architecture, and each edge corresponds to architectural morphings, such as introducing causal reasoning, factorization, or embodiment. (.) For example, reinforcement learning can be enriched with causal reasoning to obtain causal reinforcement learning. And further transformations lead us toward richer architectures, such as schema-based learning. (…) We have been talking about architectures at an abstract level. However, architectures are not agents themselves. They specify the structural blueprints from which many different agents can be built. (.) To obtain concrete agents, we interpret an architecture through a semantic function into an implementation category, such as set, stock, or other categories. (..) Each implementation gives rise to a concrete agent, and all implementations, compatible with the same architecture, naturally form an agent category. (.) Moreover, architecture morphings induce corresponding transformation between these categories of agents. In this way, our framework connects abstract architectural designs with concrete implementations, in a mathematically coherent manner. (..) To illustrate the framework, let me briefly compare two well-known architectures, reinforcement learning and universal AI. Reinforcement learning represents the simplest taste. It has a relatively flat operational structure, a single knowledge carrier, and a small set of constraints, such as Bellman consistency and Markov assumptions. (.) Universal AI, on the other hand, introduces a much richer internal knowledge organization. (.) It includes explicit memory, hypothesis spaces, bayesian updating, universal prediction, and epistemic constraints. (..) Despite these differences, both architectures are described using exactly the same three-layer framework. A syntax layer, a knowledge layer, and a constraint layer. (.) This shows that the framework is flexible enough to capture both relatively simple and highly expressive agent architectures within a mathematical layer. Once architectures are represented in this way, moving from one to another becomes a matter of adding or refining a structure, rather than creating an entirely new formalism. (..) Our third case study is causal reinforcement learning. Our framework describes it as a structural extension of reinforcement learning. The key idea is the introduction of an explicit causal knowledge carrier, separating causal reasoning from policy optimization. (.) This requires extending all three layers of the architecture. The syntactic layer incorporates new components for causal intervention and learning. The knowledge layer introduces an additional representation for the causal world model. And finally, the constraint layer adds the assumptions required to preserve causal semantics while keeping all the constraints already present in reinforcement learning. As a result, instead of introducing a completely new formalism for causal reinforcement learning, the framework shows precisely which new structure must be added to reinforcement learning. (..) Our final case study is schema-based learning, which I present here only as a brief introduction. Within our framework, SBL does not appear as an isolated architecture, but as the result of a sequence of structural transformations applied to causal reinforcement learning. (.) First, knowledge becomes factorized into multiple specialized models instead of remaining centralized. Second, learning is decentralized into multiple cognitive models, each with its own functional role and with access into the memory. Finally, the introduction of a body layer separates raw sensory processing from higher-level cognitive representations. (.) Together, these transformations produce a much richer and more modular architecture, showing that increasingly sophisticated architectures can be understood as excessive structural refinements. Let me conclude with the main message of this work. We have proposed a categorical framework for describing and comparing architectures through three complementary layers, syntax, knowledge, and constraints. And using these, different architectures can all be represented and related through structured transformations. (.) Beyond comparing existing architectures, we believe this framework also provides a foundation for study structural properties of architectures, or for systematically design new ones. (.) If you are interested in the full formal development of this work, the standard archive version is available through the QR code. And, if you have any questions, comments, or would like to discuss possible collaboration, please feel free to contact us. Thank you very much. (……)

S34: On to our third paper of the session, Quantum Logic Networks, by Ben Gertzel. (.)

S23: Yeah, so, in case you haven’t heard enough of me already, here we go again. And tomorrow morning, first thing, for those who show up for the opening, I’ll give a half-hour talk on Hyperon Path to AGI, which is the main thing I’m working on now, aiming to actually create real thinking machines and all that, connected with predictive coding neural nets and so forth. Now, this talk is something, I won’t say totally different, but a bit different. Given the wonders of our community of referees, the handful of papers I submitted to the conference this year, on the actual AGI work I’m doing now, all got rated a little lower, and you can see a few posters over there, right? So, we have posters on Quantel weakness as a wide-ranging approach to Occam’s Razor-esque heuristics for AGI, on TransWeave, which is a novel mathematical method for transfer learning from different dynamic programming-ish AI processes, and on fluid dynamics-based approaches to neural networks. But all that are posters. And the one paper I submitted that our referees thought was good enough for an oral presentation, probably because they couldn’t understand it, was on quantum logic networks. And this is something I think is kind of where AGI was when I started my career, in the sense that it’s obviously important, it’s obviously going to happen. It’s not quite feasible to do, given the hardware technology that we have available right now, but you can work out the math, you can sketch out algorithms, you can sort of figure out how you will build it when the hardware is there. And that’s, you know, that’s how I was thinking about AGI for a long time. Now, I think, in my personal view, for human-level AGI on commodity, GPU, CPU, hardware, even though the von Neumann architecture isn’t really right for AGI, it’s come along very well. Well, I think it’s good enough. I think we can put our AGI algorithms on the hardware, and things will get even better when we get next generation of hardware. For quantum computing, we do now have quantum computers that you can use and really do things. And, like, I’ve been playing with different quantum algorithms online on, like, the Dirac 3 optical quantum computer through their API. So it’s amazing we can do real quantum computing now, you know, just sitting on your laptop, connecting by the API to some room-temperature quantum machine. On the other hand, you can’t do on any current quantum computer a non-trivial version of what I’m going to go through in the next five minutes or so. You would need more logical qubits than are available on any existing quantum computer, but the math doesn’t care, right? And I do think even if you don’t need quantum computers to get to human-level AGI, the human-level AGI will rapidly set about advancing the field of quantum computing so it can port itself onto those quantum computers and massively upgrade its intelligence, right? So, what I’m going to race through here now, and I’d encourage you to look at the paper that was submitted, which goes through it a little more leisurely. I’m going to race through the notion of porting PLN, the probabilistic logic network framework that’s part of the Hyperon AGI platform and part of the Omega-Claw agent and so forth. So, how would you do PLN in a way that fundamentally leverage quantum computing? And that’s a step from PLN to QLN. I haven’t figured out what RLN will be, but we’re advancing, right? So, PLN, it’s an uncertain logic engine that developed the first version of in 2000, 2001 or something. But the basic idea with PLN is you can take probability distributions or approximations thereof associated with premises, observations, statements, and you can propagate them through mathematical proofs. And you can do it with first-order logic, higher-order logic, paraconsistent logic, blah, blah, blah. But you’re tagging your axioms with probability distributions, and you crunch the distributions through the proofs, and you get distributions at the end. Now, the problem you meet is your distributions get more and more entropy as you go through, because you’re accumulating more uncertainty as you do the proof steps, right? And porting this to quantum mechanics, I mean, you get deep into some details, but the basic math apparatus is not so hard. You have the notion of a qubit, which can live in a superposed state between off and on, not just either off or on. And we’ll have a Hilbert space, which in what I’m considering here will be a finite dimensional complex vector space of possible states, right? And then we’re acting on these density matrices rather than on individual truth value numbers. And when you reduce the quantum density matrix to one dimension, QNM will reduce to PLN. Now, the formalism doesn’t say what size the density matrix has to be. (.) If it’s small, you can just simulate the quantum mechanics on classical machines anyway. When the density matrix gets into the thousands of dimensions, simulation blows up and you actually need a quantum computer. And I’ll talk a bit at the end about when using a quantum computer might actually have some practical value. (..) The elements of PLN that I treat in the paper are inference rules that originally came from Pei Wang’s NARS system, which is a non-probabilistic reasoning approach, and we gave them a probabilistic analog within PLN. So you do deduction, which is basically transitive reasoning, A implies B, B implies C, therefore A implies C, induction, abduction, and revision. And we have simple mathematical forms for these rules in PLN, which I don’t have time to go over now. But what I did in this paper was take the PLN truth value formulas for each of these rules and then figure out, like, what’s the most natural analog in the quantum domain? Well, you’ve replaced the PLN truth values with quantum density matrices, right? So instead of a PLN simple truth value, which is, in essence, a probability and then a weight or a count value telling you how much evidence that probability value is based on, you have a density matrix and then a count value. And the dimension of the density matrix sort of depends on how big is the context in which your inference is being done. You could do it very small. Like I said, it’s when the dimension gets into the thousands that classical simulation doesn’t work well and you have more of a reason to really be doing things in the quantum domain. And this slide summarizes a bunch of math, which is given in the paper, and I cannot read it from over here. (….) So basically, what I did is go through each of these four inference rules and figure out which quantum math corresponds to the PLN inference rule. So deduction in PLN is a sort of probability chain rule in quantum mechanics that corresponds to a certain trace formula. And if you’ve done quantum mechanics that will look obvious, otherwise it will look like voodoo, right? And abduction and induction are sort of Bayes rule type things. Well, there’s something called the PETS recovery map that sort of tells you the best operator to use as the inverse of a given operator. And there’s a bunch of known math about that that just tells you how to reverse a quantum operator and you can put that in place of Bayes rule. And you can, for say, revision, instead of just a weighted average of numbers, you’re doing a convex combination of these matrices and then taking the trace, right? So it’s gone over in the paper and this sort of presentation doesn’t have enough time to go through the math. But the basic methodology was you can take each of these real number formulas and you find an analog in the algebra of quantum operators that has the same symmetries and the same mathematical properties as the PLN rule. And so this blown up a bit, this is the deduction rule. And you can see, like, if you look at the chain rule used for deduction in the upper left corner, which is a very simple probability rule formula for like A implies B, B implies C, therefore A implies C. So if you know any probability theory, you can see why that rule should be the thing. If you look at that rule and you then look here, like, it’s pretty much the same thing. You have a tensor product in there to make sure everything has the right dimension. You have the trace of the operator, which is the way to get rid of the middle term B and A implies B, B implies C, A implies C. It’s sort of the same form as the PLN rule, but you’re just putting the quantum operator algebra in there. And you can then prove a bunch of nice properties about this. Like, you can prove the correct quantum rules. They don’t add any new evidence. Like, you can’t hallucinate if you do the rules exactly. The inferential processing can lose information. It doesn’t create information. The entropy can only increase as you step through. So you can prove that it has simple properties. You can also come up with mathematical bounds on how much the result of the inference will vary when you reorder the premises. Like, quantum theory, not everything commutes, and the order of premises can mess up the conclusion or enhance the conclusion. You can bound how much that will happen in terms of the density matrices involved. Now, an interesting question here is when you actually need to use quantum inference. Like, when is it of any value, right? And then, unsurprisingly, the case where it’s of value is where there’s a complex pattern of quantum coherence among your premises, and there’s no way to sort of do a change of basis to get rid of it and make everything look diagonal, right? So if you have quantum interference between a whole bunch of different premises in your knowledge base, I mean, then doing this sort of thing with a sizable density matrix is going to give you different answers than you get from PLN. And, I mean, how important that is in everyday reasoning is unknown. I have an argument that it may be. Of course, if you’re reasoning about how to build a molecular nano assembler, it’s probably very, very useful, because then you’re actually reasoning about, you know, molecules that are operating in the quantum domain, and doing logical reasoning using regular probabilistic inference is going to mislead you. And this, it sort of cashes out, at least at the math level, I thought I had for a long time. Like, for everyday Newtonian stuff, we have a sort of naive physics of basketballs and baseballs and walking around, and we don’t have to derive everything about the everyday world from Newton’s laws, right? I mean, you can, but it’s not what we do every day. In the quantum domain, we derive everything from the math, because our intuition is bad, right? Now, on the other hand, an AGI could be given sensors at the quantum level, and could then make a sort of folk or naive physics of the quantum world, complementing the standard mathematical physics. I would say, QLN could be a route to that, right? Like, if you have quantum data coming into an AGI system, this is like a heuristic, imprecise, AGI-ish way of informally reasoning about the quantum world, which should be a quite interesting thing, right? And it would suggest a sense in which the AGI may understand quantum mechanics at a gut level that we don’t. Now, at risk of wasting even more time on something that we can’t build this year, I want to take one more minute to mention an allied observation I made that is not in the paper for this conference, but I wrote up more recently. I constructed an interesting argument that if you have a small system, in a certain sense, which is observing a large system, even if they’re both systems you would normally describe using classical physics, then that small system should model the larger system using the quantum algebra of observables. And the basic reasoning is that quantum logic is what you use to reason about something you cannot, in principle, know. And a system with algorithmic information, 10, cannot, in principle, know a lot of things about the states of a system with algorithmic information, 10,000. So, if this is the case, it would suggest that your deliberative mind should model your unconscious mind using quantum logic, which is an amusing thought that I haven’t fully digested. But it could be that something like QLN is actually then useful within the classical domain for a small subsystem of an AGI system to reason about a larger one. And if that seems to hold up as math, and I’ve written it down in Fable and GBT56 Pro, didn’t find the mistakes in my arguments yet, right? If that makes sense, it could imply that doing something like this with smaller dimensional density matrices, like 10 or something, could actually be of relevance in current digital AI systems. Because if the density matrices aren’t such high dimensional, you can simulate it all classically, right? So, anyway, that’s a few thoughts on quantum logic networks. I apologize for skipping all the interesting details, but there just wasn’t time. But they’re in the paper. Thank you all. (..) And be sure to show up early tomorrow morning. You can hear about HyperROM, which we’re actually now building. (.)

S34: Thank you very much, Ben. (…) Again, we’re getting the proceedings out so everybody can. This is the preview of coming attractions, so just think of it that way. (.) Our fourth and final paper here in this paper session is by Bo Wen Zhu, Bo Yang Zhu, and Pei Wang, Vision by Logic, a Naive Visual Recognition Model with Non-Axiomatic Logic, and it will be a virtual presentation. So please welcome Bo Wen with me.

S03: Okay, so I’m happy to share our work, Vision by Logic. (.) So there are basically two issues here. The first one is, can logic be applied to modeling vision? If so, how? And another question is, why do we need to use logic now that deep learning has gained huge success in computer vision? I will not discuss the second question, because that deserves another paper. But for the first question, I will show you our newest study on how to use logic to model the vision process, and we build a very simple visual system using that non-Axiomatic logic. (.) So traditionally, you know, in classic logic, like first-order predict logic, it’s very, it’s a symbolic and rigid kind of approach. So everything here is a symbol and you have some axioms or a knowledge base, you know what is true and what is false. But, you know, in visual procedure, it’s usually called sub-symbolic, and it’s full of ambiguity, uncertainty, flexibility, and so on. So every, you know, every belief in vision is non-binary, so it’s a matter of degree. (…) So usually in computer vision, so people use statistical approaches or probabilistic graph models and so on. And so over the decades, so people found that deep learning performs well on computer vision. (.) and usually logic, when logic is applied to vision, people use neural symbolic approaches. Like they can, so one approach is that they can extract features using a neural network, and then, you know, convert the vector into a symbol, and then you can reason on the, using a symbolic system. So we can call this paradigm perception before reasoning. (.) But, you know, if, if, if one consider a non-except logic, it’s a different paradigm. We call it a perception as reasoning. So because, you know, in nail, there’s no absolute truth, and every belief and desires are learned, and truth value is represented as, you know, frequency and confidence, it can represent uncertainty. So those properties exactly matches, you know, the, you know, what I need in vision. (..) And so there are some previous attempts, you know, they tried to use non-except logic to build a visual system. But I think the previous attempts, you know, there are some flaws. So, so, you know, there’s no, there’s no effective approach to vision based on nail before. (.) So in this work, we adopt this kind of knowledge representation. representation. (.) This representation was proposed in our last year’s AGI conference paper. The core idea is that the presence or absence of any part should be non-decisive, because there are many noise in, in the visual process. So this property, the non-decisive property can make sure that the system is robust to noise. And so in the previous paper, we represent the part-whole relation using inheritance relation in nail. for example, for example, a bicycle has a wheel or a wheel is a part of a bicycle. Then you can use the inheritance, inheritance statement in non-except logic. And this kind of representation also applied to low-level or sub-symbolic concepts. Like you have a certain feature, we call it feature one. It contains, contains a certain pixel with a certain color, like a color white. Then you can represent it as a, you know, feature one and arrow and white. And using this kind of part-whole relation, we can, we can represent an object as a combination of them with this form. (..) And we can call this, you know, in this form P can be called a object prototype. (..) And we can use the existing inference rules in non-except logic. Like we can use abduction for recognition. For example, a prototype has a certain component. and we observe something, it’s an instance, it also has that component. Then through the abduction rule, we derive that that instance is, (.) it’s a kind of, that prototype. (.) So this is, this is basically what we can call a recognition. (.) And for learning, we can use the induction rule. So, for example, for example, if we observe that assist, an instance has a certain component. And we were, we, you know, assign that instance to a certain prototype. (.) Then we can say that that prototype has that component using the induction rule. And we can use revision rule. And we can use revision rule to accumulate evidence. Like, you know, we have several parts in an object. And then through the abduction rule for each part, we derive that a certain instance belongs to that prototype. And each part can generate a certain statement. Then we can combine all the conclusions and use the revision rule to accumulate the evidence. And this knowledge reputation and inference rules are proposed in the, in the previous paper published last year. So in this year’s work, the core challenge is how to learn new concepts of object in a visual system. (..) To, to, to build a complete AGI system, visual system, there are, there are many factors. Like, the system should be subjective, active, and unified. And there should be a sensory motor loop and active perception. And the system should work in real time. But in this, in the initial study, we just designed a simplified visual system using NAIL. So the purpose is just to verify the feasibility of modeling vision using non-axiomatic logic. So we make many pragmatic simplifications. Like we, you know, assume a constant number of prototype concepts and the system just to passive observation without an active, you know, sensory motor loop. And the goal of the system is just to constructing, um, constructing part-hole relations using a convolutional like structure. (..) So this is the overall pipeline. (.) Um, the image, an image is encoded into concepts. And then, um, through abduction and revision, the system can do a prototype matching. And then, um, some, um, some prototypes are recognized. And then, um, using, with that, a recognized prototype concepts, we can just refine them a little bit using the induction and revision rule. And then, uh, so to the next image, then do, do this procedure again. Um, so for the first step, concept encoding, um, we can just split, uh, color into several levels. Like, uh, you can split, you know, use the grayscale, um, approach. Like you can have a black color, black color concept, white color concept, and some levels between them. And also for a, for any color, we can, um, separate them into three channels, red, green, blue. And then for each channel, we can use the same method to, uh, you know, encode them into concept and then use the combination of the three channels. And in this paper, we just, uh, adopt the most, the simplest, uh, one. Like we just use a black and white concepts. Um, and after encode, uh, encode an image into concepts. The next is to, um, you know, do reasoning and to recognize and to, to learn new prototypes. So, uh, we design a, you know, we call it convolutional concept network. So for each layer, um, there are some lower, lower layer stimulus. And, uh, it is represented using, uh, Narciss, which is the, you know, a statement in, um, in a non-eximatic logic. Uh, for example, in, in this figure, so there’s a certain anonymous concept, which belongs to a certain prototype. And then, and then this, um, this instance also occurs in a certain patch in a certain position. And then, uh, we can, you know, construct a new anonymous concept with that patch. And then that, that patch has a certain component according to, you know, the, you know, the lower, lower layer stimulus. And using the, um, abduction rule, we can, you know, do a step of match. (.) And, um, you know, for each layer, there are many prototypes. For example, there are N prototypes. Uh, each prototypes, each prototype is, uh, analogous to a certain, like, uh, convolutional kernel in deep learning. And, um, this matching process can, you know, um, can repeat in the next layer. (.) And after a certain, um, prototype is recognized, then, you know, the induction rule is applied according to, you know, using the, the actual input. And, uh, the, the, the corresponding, uh, the corresponding prototype is then revised using the induction rule. So, uh, this is the overall idea. Uh, due to the limited time, I can, I cannot, uh, you know, provide too much details. So, here’s the, uh, experimental results.

S34: Oh, and I apologize, but we’re running, uh, your, your session’s running a little bit.

S03: Okay, uh, just, just to, yeah, okay, I’ll be quick. So, uh, for the first layer, there are several prototypes learned in this layer. You can see, uh, the red color corresponds to white and blue color, sorry, red color corresponds to, uh, black. And the blue color corresponds to, uh, to white. And in the first layer, there are some prototypes learned here corresponding to some edges. these. And in layer two, we just reconstructing the representations just to visualize them. And you can see some, um, fragments are learned. And then we finally use, use a linear classifier and found that in the MNIST digits, it gets, uh, you know, 95 accuracy. Uh, but, uh, but here, the accuracy is not important. We just need to check whether the prototypes are learned. (.) And, um, yeah, the major conclusion. (.) So, vision can be effectively, effectively modeled by non-explanatory logic. And there are also some future work. Like, uh, we can improve the model and the experiments and do the complete AGI perception model. (.) So, uh, thanks for listening. And here’s the, you know, the paper link and the source code. If you, you are interested, you can know more details through these links. Thanks. (..)

S34: Thank you, Bowen. (….) I apologize that you can’t hear the applause that you just got. (..) So that wraps our papers for this session. Now I’d like to call up our two live presenters and Bowen if you can stay on, um, to answer any questions that we have. So Ben, Bowen, and Arthur. Now Ben’s here, and Arthur’s just a question from the room. (18 seconds pause)

S00: Alright, so this question is primarily for Ben. Ben, in your talk on quantum logic networks, you made the point that the von Neumann architecture is probably not the right architecture for AGI. That’s a statement I think I agree with. Quantum computers, on the other hand, look interesting. But quantum computers, while they exist, they’re not exactly what you would call ubiquitous or readily available. Especially not if you need a large number of qubits. Given those two things, I’m wondering, do you see any useful role for analog computers? (..)

S23: Um, in principle, but I think in practice, my best guess is we can get to human level AGI, the boring, old-fashioned way with networks of multi-GPU, multi-CPU machines. I think if we use an architecture like Hyperon or a predictive coding neural net, you don’t need a hyperscaler server farm. You can use a mix of smaller server farms and then a larger sort of periphery of random machines in a network. I mean, you’re fighting against the infrastructure not being MMD parallel and not being what you want in a bunch of other ways, right? But that, I mean, it doesn’t mean you can’t do it, right? So, I mean, obviously, given some simulation cost, one computer architecture can simulate another one, which can simulate another one. And you’re just eating some inefficiency, right? And quantum computers, obviously, that’s a huge field of possible architectures, right? It’s cool, like I said, we have some we can play with now. You’d probably need 5,000 or 10,000 logical qubits, not 100, to be able to implement bits of an AGI architecture on your PPU along with your GPU. (.) It seems like if things go fast but not super intelligence level fast, we’re, I don’t know, what, 7 to 15 years from having quantum computers like that. On the other hand, I founded SingularUnet eight years ago, right? So that doesn’t, and I’ve been working on AGI for many decades. So that, that’s not that insanely far off when you think about it, which gives some level of realism to this sort of math. (.) Analog computing, I mean, for sure, if we had a scalable, affordable, general-ish purpose analog computer, it’d be very interesting. You can run neural nets. I mean, even PLN is symbolic logic, but all the truth value formulas are floating point computations, right? I mean, you could even do various sorts of symbolic AI leveraging analog computers. But, I mean, that’s, you know, in the 90s, I was working with a connection machine with 64,000, 128,000 MMD parallel processors. It was very cool, but I wish the, I wish the hardware world had gone that way. Neuromorphic chips, I mean, we’ve been talking about this with predictive coding. (.) PC would be very natural for various kinds of neuromorphic chips, but then what kind? Are they gonna put spiking neurons? Will they put chaotic Izykiewicz neurons on the neuromorphic chip? I mean, it’s a, in any case, they’re not available at a reasonable price and scale at this moment. And I think we could make PC run, like, 20 or 30 times faster just by making different GPU libraries right now. So, I sort of think, we can think on both those timescales. Like, we can optimize the use of the current suboptimal compute infrastructure. We can also figure out how to make optimal algorithms for the computers we’ll have in five or ten years. And if we get to AGI, maybe that will bring about 10,000 logical qubit machines, like, a year after the human level AGI emerges. And whether classical analog computers will have a role then is interesting to think about. And you have, you have things like the D-Wave, which is marketed as an adiabatic quantum computer. But it’s not clear how much of what D-Wave does is basically analog computing using thermodynamics versus to what extent it’s leveraging quantum effects. I think it’s clear there are quantum effects there. But how much of their compute speed is due to the quantum versus merely cool analog stuff, I don’t think we even know, right? So, the boundary isn’t even that strict. (.) But, yeah, there’s a lot more to say about that, but it’s more than enough. (12 seconds pause) We’ve baffled everyone in this submission. (..)

S29: Hello. My name is Gerald, and I’m here with my daughter, LaVie. And LaVie committed. She’s 10 years old, and she said she wants to build up her AI company. She’s 10 years old. So, my question to both of you is what would be the first steps for her to encounter? What would be good resources for her to start? (..)

S23: To start an AI company?

S12: Yeah.

S23: Maybe I’ll let Haley. (..)

S34: Well, it’s interesting. We have a number of emerging young researchers. In fact, our next two presentations are rising researchers. We have a 15-year-old who presented a paper, has a poster or two in the back. We have an 11-year-old and now a 10-year-old competing in this space. So, yeah, building an AI company is a very worthy and approachable endeavor. LLMs are making it more accessible than ever. There’s a lot of cool technology going on. And to build the company, you know, in Silicon Valley, you need an idea and a good pitch and get out there in front of investors. I think that I’ve been inspired today by Michael Levin’s presentation where I think that one of the things he did very well was not be constrained by assumptions, which I think our young researchers, our 10-year-olds are going to be better at. You know, what looks possible and doesn’t look possible and imagining the unimaginable, building the unbuildable. And I think that there’s a leg up there versus us who have ideas.

S23: Let’s add on to that. So, I think there’s a lot of ways to make a product or a lot of ways to make a company. But one interesting approach to take is figure out something that you would like to have, that you would like to use, that you would like to see out there because you want to play with it or you want to use it to get something done. And then, you know, in the current era, sometimes you can actually build that. And often, AI will play some role in helping you build that thing, right? So, what’s led me to experiment with these omega hives I’ve been talking about is I wanted to automate some of the research I was doing. So, I’m like, how can we, how can I piece together something I can use as an automated research assistant? And I have, you know, as a music keyboardist, I have loads of ideas for how to tweak the keyboard interface so it can be more expressive. Once you create the AGI and I have nothing better to do, I’ll start building some of those, right? So, I mean, whether it’s a game you would like to play, like something you’d like to have to help with school work, something to chat with people differently, something much more amazing and novel that I’m not thinking of. Like, if it’s something you would like to use and play with or people you know would, you know, maybe now with vibe coding, you can actually build it. And an awful lot of software products have interesting roles for AI when you start to think about it. And you can feel free to brainstorm, you know, 25 possible ideas, sort of chat about them with your friends and family and see which one seems like might be something you could actually do given the technology that’s available. And then the other key point is don’t be afraid of failure, right? And that’s, having lived all around the world, that’s one of the things I’ve come to appreciate about American culture. Like, we allow each other to screw up and then just get up and keep doing something else again. And that’s not so true everywhere else. Like, brainstorm a bunch of ideas, pick some you can try out, try to make something, play with it. If you like it, go with it. Maybe you do, make a company around it. If you don’t like it, chalk it up to experience and pop the next idea off the queue, right? So that, off the cuff is my thought about it. (….)

S34: Great question. Thank you. Wonderful. Thank you very much, everybody. Oh, and join the AGI Society, of course. It’s a joinagisociety.org link right here. (..) All right. Mathematics is the skeleton of mind. Our paper session wrapped. Thank you very much for that. All of our presenters, thank you very much, Bowen. And we’re gonna move, keep moving right along. The next session is one of my favorite things about the conference. We’ve done this a couple years now. We saved some time to invite researchers that are early in their careers. Because, again, just like this, ideas coming out of the next generation of AGI researchers are not incremental. And they’re so important. And encouraging that, growing that, and hearing from these voices is so critical. So, first, I’d love to introduce a dear friend. Faiza Habibi is an AI researcher at Singularity, coming out of the computing lab. (…) I’ll switch here, just in case. Her work sits at the intersection of predictive coding and generalization. We’re all here, the hard problem we’re looking to solve. So, what does it mean for a system to extract principles from experience that hold true in futures it has never seen? She’s going to make the case that this is the actual path to AGI. So, joining us virtually, Faiza. Welcome. Welcome.

S04: Hello. (..) When intelligence generalizes. (…) We are going back to the Greek ancient legends. (..) Anyone here watched Odyssey movie by Christopher Nolan? (..) Please raise your hand. (…) Good. (…) 10 years of war, and now it’s over. (..) And on the beach at Troy, there are ships. (.) And on those ships, there are men. (..) The men who all want one thing. Just one thing. (.) And it’s home. It’s their home, Ethica. (…) And the man leading them is Odysseus, the king of Ethica. (..) Odysseus was the bravest man in all ancient legends of Greek. (.) Odysseus was so brave that he had done something no living man does. (..) Odysseus has gone down into the land of death. (.) He has stood in the dark and spoken with the gods. (…) The gods gave him the map to home. (.) But there was one warming. (..) There is an island ahead. (…) Tyranesia. It’s the name of the island. On that island, graze the cattle that belongs to Halas, the god of the sun. (..) Sail past them, and you will finally find your way home. (.) Or touch them. And you will lose your man, your sheep. And you will never ever again see home. God said. (..) The lesson we learned from Greeks is that gods don’t lie. (….) And they sail. (..) And in one evening, sea beaten and starving. (.) It’s a beautiful land on the horizon. (.) Tyranesia. (..) It is evening, and they are so tired. So they land there, and Odysseus told them the gods’ warning. (..) And Odysseus falls asleep. (…) Up on the hill, there were cattle. (.) Fat, slow, and shining cattle. (..) The home was over the horizon, sounding like a dream. But the stakes were right there, up the hill. (.) Much closer, much believable. (…) The temptation was too strong to resist that they killed the cattle for meat, sacrificing their future prosperity. (..) Odysseus wakes up, and he smells the smoke. (..) And he was so desperate, so hopeless about getting home. (..) He was as hungry as any of them. (.) But he trusted the truth. (..) He had faith in what he had seen. He had faith in what he had heard from gods. (….) And they sailed off. Zeus shattered the sheep with a thunderbolt. (..) And every man went into the water. (.) Every man died. (..) Exactly as Odysseus has warned them. (….) Everything they had done, everything they had survived, was for just one thing. (..) And they were so close. (.) And they were so close, but they never arrived. (…) Only one man arrived home, who alone survived. (.) And it was Odysseus. (..) But he arrived late. (..) He arrived alone, on a stranger’s ship. (…..) Today, we are all on the same ship. (..) Sailing on the same island. (.) For one thing. (..) There is one thing that we all want. (.) A goal that we have bet our career on. (..) And we all hope one day our ship takes us to AGI. (…) It is a very big goal. (..) And we are so close. (.) But we must be very careful about the warnings. (…) Welcome everyone. I am Faiza Habibi. (.) And I am AGI researcher at the SingularityNet. (.) I am coming from academic lab. From the NAC lab. (.) Where my research has been entirely focused on one thing. (.) On the predictive coding learning algorithm. (..) As a researcher, my job is to be bold. (..) To go to the land of unknowns. (..) And to explore every possible path that can take us to our AGI. (..) To show you. To show the people in the ship with me. (.) The path that promises us home. (…) And today, I stand here holding a warning. (..) I hold the warning that science has shown to me. (..) And we all know that the science never lies. (…) To warn you. The path today that we think it is taking us to AGI cannot carry us home. (…) There are days when I feel I’m Odysseus. (.) Not that I’m so brave or smart. (.) I feel I’m Odysseus living in a world that everyone here is thinking about steaks. (..) And today, I’m going to tell you why eating those steaks would cost us our wished home. The AGI. (..) I’m gonna tell you where we are coming from. From the beginning of the AI. (..) And I’m going to tell you that the glory is past. (.) The path that we are currently at is flawed. (..) But the ethical is still here. (..) There is a shining path that we can take to get us there. (..) Now, let’s begin. Let’s begin with the past. (…..) Fantastic team we have seen today are machine learning. (..) They are machines that learn data distribution. But they are all essentially black boxes. (…) The question is, can these black boxes take us to AGI? (……) Black boxes. There are systems. There are neural networks that somehow find this complex mapping. That machine that matches their input to their output. (..) Now, I want to ask you one thing. (.) That as you listen, I want you to ask yourself one thing throughout my talk. (..) If these networks work the way we believe they can. (..) I think the answer is more concrete and more exciting than you might even expect. (….) So, there are three scenarios. There are three main scenarios that machine can learn. So, what they are learning. Depending on these three scenarios, they encode different information. (..) The first one is learning is different machines. Multiple models learning multiple distributions in parallel. (..) The second one is one single machine that learned from a sequence of distribution. And the third one, a large model learning on multiple distributions. Shuffled together pretending that it’s only one single but complex task. (..) But what matters here is what they learn from these data. (…) The first scenario, they train in parallel. Each model fits on a single distribution. Narrow and specialized on a single distribution. The second model, a single model trained on a sequence of distributions. And it overwrites prior knowledge for the favor of the current task. (..) And the last one, a giant model learning complex distribution that shuffles all the information together. (..) Also narrow, but this time specialized on a broader domain. (……) Time that the training stops. The second model, the sequential model catastrophically forgot the data it has seen. (.) And only remembered the last distribution. (..) So, it cannot be reliable performing well on all the data. (.) So, it cannot be our solution. (…) Now, let’s explore the two others. (.) Can they bring us the AGI? (…) Let’s see. (……) Let’s take as many models as possible. And train them on as many tasks as exist in this world. (.) So, they become specialized, but narrow in only one task. So, we have multiple models. (..) Now, the question becomes, how do we bind their speciality? How do we connect all of these narrow domain AIs? (.) To make the AGI. (……) Keep this aside, so we can also see the other approach. And we can compare them together. (.) The other ones seem to be working without having the binding problem. (…) A single giant model that learns all the information together. (..) So, let’s say we take the whole data in this planet, in the internet. And we shuffle them all together. And we hope that it can take us to AGI. (…) So, we make the model, we start with the small, and we make the model large enough, so it can fit all the information within it. (.) And we give it infinite computational power. (.) Infinite data, infinitely large model, and infinite computational power. (..) So, what happens after? (.) Let’s say they became AGI. (…….) So, here we have AGI. (.) Let’s organize our mind. So, we have two models here. We have a model that collects many specialized AI. (..) And we somehow could figure out how to wire them together. (.) And the second one is an infinitely large model that’s supposed to be our AGI. (14 seconds pause) In the world, in the world that they have learned from, comes a new information that they have to handle somehow. (….) One of the biggest problems in AI today is they cannot transfer their knowledge into new environment. Essentially, they lack generalization in out of distribution. It seems that they remain narrow, only confidently specialized in their training domain, and unreliable otherwise. (…..) They learn data somehow, to have all the knowledge human needs. (.) Let’s begin with the first model. (…) There are two scenarios for the first model. (….) Either it overrides one of its modules, its prior knowledge, which means it forgets the previous understanding of the world. (..) Or, it must add a new module and keep the rest untouched. And it comes the problem with binding the information, with somehow connecting all of these existing modules with the new one. (.) So, it cannot be our solution. The solution that we are looking for, it cannot be adaptable easily. (.) Let’s see the second approach. (….) The next hope is the biggest model. (…) But the problem is, same as the first model, it also encountered some kind of forgetting. (.) It may not be as strong as the previous one, but still, it finds a very strong bias to the new information. And it cannot be our solution neither. (…….) Forgetting is very expensive. (.) It translates to a never-ending investment. (.) And AGI is off the table. Or at least, with the current methods that we know for AI today. (..) The problems. (.) The problems are, none of the current systems are reliable, or sustainable, or either a good business investment. They are not reliable because they cannot handle new environments. (..) They are not sustainable because you have to keep investing on the retraining. For those of you in the room who are investors, I’m going to tell you the cost of retraining and retraining the large models is very huge. And it never ends when the world changes. (.) And our world is always changing. (..) The current practice we see in big tech companies is to start from scratch. (.) To start from scratch with retraining. (.) They shuffle the new with the old and retrain the whole network again. (..) And it’s not a sustainable business model or an investment. So far, I showed you that machine learning is not going to take us to AGI. Or at least the current approaches with machine learning. (..) And that was step one. (.) Step two is how machine learns and what makes their intelligence always remain narrow. (15 seconds pause) Machine learnings, when they want to learn the data, they have one deterministic forward mapping from input to output. They couple all the parameters under a global single objective that entangles knowledge across the entire parameter space rather than creating separatable and reusable units. (…) This creates problems. (..) This creates two critical failures. (..) The first one is that it overwrites prior. New learning overwrites old configuration, making continual learning structurally impossible. (..) And the second one, it requires retraining forever. (.) Every novel feature combination requires retraining from scratch. (.) And that’s very expensive. (………) A different question. (.) We ask what if more computational power is not all we need. (.) Machines learn big data narrowly. But children and humans learn little data widely. (.) Humans learn from few examples and generalize across different tasks, situations, and environments. (…) We need something that doesn’t override. We need a system that retains what it learns. (..) A system that transfers its knowledge to unseen environments and adapts in real time. The capability that AI today is missing to get to Asia. (……) The results show that predictive coding is what we are looking for. Predictive coding is a computational model of the brain. It’s a theory on how might brains work. (..) Predictive coding, when it comes to learning a data, it separates inference from learning. (..) It predicts its internal representation locally. And it corrects the mismatch error locally. It recombines its internal structure of known pieces to explain the stimuli before learning or overwriting anything. (.) Predictive coding is a promising path. (.) A novel input is an unseen combination or unseen configuration. (..) And one deterministic forward mapping can never give us the chance to recombine the internal knowledge. Now, I want to show you why predictive coding is uniquely so powerful. I’m gonna show you some beautiful results that I’m so excited about. (….) I told you that the hope is to have a machine that learns from small data widely, similar to children. (…) In our research, we take 20 images of natural sin. Only 20 images. Nothing more. (..) And we train a three-layer predictive coding network on those samples. (..) The task here is learning the representation of the data. (……) Training is completed. We show samples from unseen distributions. (..) MNIST, Japanese characters, three-dimensional shapes, or zero-shot benchmarks. Using what it knows, the predictive coding network build a version of the data for itself to present the data representation. And it can adapt quickly when data changes. (….) So, let’s see one of the examples. One of the representation it learned for 3D, three-dimensional shapes. (…) Taking this representation, we can train a network on reconstruction tasks, and we can do the downstream tasks for reconstruction. (.) This representation can be used to classify, and it reaches 95% accuracy on the classification of NORB dataset. It detects novelty and clusters distributions. (..) And there are many other tasks it can do. (..) So, what matters here is that the predictive coding was able to build the representation of the data. So, they are semantic, and they are useful to generalize for any downstream task. (….) Jen, I want to highlight for you the efficiency of our framework that we are building on top of the PC, or predictive coding. (.) Our framework operates in a fully zero-shot, meaning it requires no retraining, no fine-tuning, and no additional cost when it encounters new domain. So, it can get adapted to out-of-distribution samples. (..) This model promises to produce more reliable output. (.) It dramatically reduces the need for data. So, it is very much data efficient. (..) And it’s sample efficient, meaning that it doesn’t require huge internet-sized datasets. It learns only from 20 images. And it reduces the computational cost, and it is sustainable. Therefore, it is a very, very good bet. It’s a very good investment. (…) At the beginning of this talk, I told you about Odysseus and his main journey. (..) And a warning they ignored that cost them home, Ithaca. (..) And I shared that predictive coding is providing to be a part of the desired future. And this is exactly what we were looking for. (..) And getting there, getting in the future of AGI, it requires us to take bold moves. (..) It takes us courage. (.) The courage to walk in the direction, while everyone else seems to be moving somewhere else. (..) AI today is the cattle on the hill. (..) It feeds us today. It smells incredible. But it will never get us to Aetica. (…) There is part of the story that I never liked. (…….) I never liked the part of the story that Odysseus gets home alone. (.) Odysseus had the secret of the gods. And he was so right. But he arrived with no man standing next to him. (..) I don’t want that ending. I want you to remind you that the Aetica is still here. (..) At the SingularityNet, we have a ship and we are sailing towards the AGI. (..) We firmly believe that PC, predictive coding learning algorithm, being the answer we were looking for. (…) There is a road forward. (.) We are putting our efforts and our resources to build the road. (.) We are spending our efforts to scale PC, to scale the right thing. We are building frameworks, NGC Learn and Fabric PC, for this road. (..) We are trying to solve neurosymbolic processing through PC circuitry. (..) And we want to solve AGI with solving continual learning. (….) The decision is yours. (.) Do you want a steak now or you want home? (..) That is your decision to make. (..) But before it’s too late, join us to be a part of that ship now. (.) Or you will reach there alone, late or on a stranger’s boat. (.) Thank you. (12 seconds pause)

S34: Absolutely wonderful. I wish you could have heard the applause here. But we very much appreciate you being able to present virtually at the very least. (.) Your descriptions, your story, your narrative are moving on this journey that we’re all on together. So, from generalization to the substrates of intelligence itself, let me welcome our second rising presenter. (.) Will Gevert is a fifth-year PhD student at the Rochester Institute of Technology as a part of the Neural Adaptive Computing Laboratory. (.) His work has focused on biologically plausible credit assessment methodology in spike use. (..) And his current work focuses on expanding this to self-supervised representation learning. So, this is about pushing that frontier. So, Will Gevert with spiking representation learning.

S25: Okay. (11 seconds pause)

S18: See if I can not mess this up too badly. (..) Just badly enough. Yeah, that’s the trick. (..) Alright. (…) There we go. (.) So, unfortunately, I don’t have a movie to preface my talk with. You’re just stuck with a gray background and, you know, me. So, I’m here today to talk to you guys about spiking representation learning. And some of the open questions that come with it. And perhaps a potentially novel approach to how we want to actually tackle these sorts of problems. But before I can truly talk about spiking neural networks, representation learning, or any of that. We first really have to focus on, well, encodings. So, for us, an encoding is the representation of our data once it has gone through some transformation. Where we are trying to reduce the amount of resources required to actually store the information. While maintaining the ability to extract the information from the image. In this case, back out of it. So, here, I’m gonna be representing all of our encodings and representations as colors. This is mostly so that way the people in the back don’t try and have, don’t have to read some tiny number. You just have to see a slightly washed out color palette. So, here we are, we have two different images here. Well, we have two different patches of an image. I’m gonna, important to stress here that in the work that we do, we work in patch hole hierarchy models. So, generally speaking, what this works out to is, exactly like in biology with a human, you don’t see the entire world in front of you at a time. If you hold your thumb out in front of you and focus on your fingernail, that is roughly the amount of space that your brain ever actually looks at as your eyes sweep around a room. And so, importantly with that, that means that we only ever need to actually produce representations on things roughly that size. You don’t have to produce a representation for your entire world view space that you can possibly see. You only need to do it for the little part that you focus on at a time. So, in this case, we have a cat. And in the top image, we are looking at the face of the cat. And we produce this lovely color embedding from it. Now, in the bottom, we’re looking at a different part of this cat. Its foot. Now, as I’m sure we’re all aware, feet are very different from heads in most animals. And so, therefore, it’s gonna produce a different embedding space. (.) And so, the important thing is that we would be able to go backwards out of these embedding spaces. But when you try and just take these and fit them directly into a model, you can run into a couple of challenges. (.) Specifically, you run into sample-to-sample collapse, which is it doesn’t matter what I am looking at. I will always produce the same embedding, which isn’t technically wrong. If we determine that we can fit all of our information into ten colors, that’s great. It’s even better. I can fit it into a singular color, and it’s just zero. (.) And so, with this, basically, as a collapse problem, is it’s not useful, though, right? If I just give you, if you ask me to give you an embedding, and I always just give you the same embedding back, I have succeeded in my task of giving you an embedding, but I have failed to give you an embedding that is useful in the context of doing anything with it. And so, this is one of the challenges that you’ll run into if you ever try and naively implement representation learning or just creation of embeddings in general. (.) The other problem you can run into is informational collapse, which is, okay, I’ve successfully given you different embeddings for whatever you’re looking at, but it’s just all about the same number. And that causes other problems because now even if I have a hundred features in my embedding space, I really only have effectively one feature in my embedding space, right? In this case, it’s orange. It’s a truly useful feature. And so, the thing about just having a singular value in your embedding space is that dramatically reduces the amount of information that you can store in that embedding space, right? So, if I’m dealing with integers, and let’s just say 8 bits, so we go to 256, instead of having 256 to the 10 possible values, here I have 256 total possible values. And so, we really want to also compare this informational collapse on a singular embedding scale. (.) Now, luckily, if we look at deep learning, people have actually, you know, solved this. (.) Representational learning is nothing new. I’m sure we’ve all been doing it. In fact, we just heard about predictive coding that produces embeddings and produces representations. I’m going to look at a mechanism that’s a little different than predictive coding. We’re going to look at a modern solution of self-supervised learning known as VicReg, or variance, invariance, covariance, regularization. This comes from a paper published by Adrian Bards, Jean Ponce, and Jan LeCun back in 2021. Since 2021, they have published a couple of more papers on this. (.) Sorry, the whole stage is shaking. (..) And so, like, they’ve published more work on this, but this all works back to the same general principle, which is you can fight both forms of collapse if you have a variance term, an invariance term, and a covariance term in your loss function that you’re trying to minimize. And we’ll talk through what each one of these terms actually does. (..) So, what is our variance term? So, for here, let’s just say that Z1, Z2, Z3, Z4, and Z5, these columns that we have here, are representations pulled from different patches in the same image. Now, this variance term says that for any individual element, so say the nth element of a representation, we want to maintain that the standard deviation across this batch of patches that we’ve pulled out of the image stays above some threshold. (…) There’s a phone call behind me. And so, what this allows us to do, if we maintain above this threshold, that means that on a sample-to-sample value, we’re not, we are fighting that the representational collapse of from sample-to-sample. Because we won’t have to worry about the first element being the same across all of our possible embeddings, and therefore being a useless dimension. (.) And so, this is great. It does help, but you still run into a problem. This doesn’t fix the internally in a sample, the value-to-value collapse. Because one of my embeddings could all be one. The next one could all be two. The next one could all be three. And they would have a very large standard deviation. And they would have, it would be great. But we still would run into this continuous problem of sample-to-sample collapse. So, okay, sorry, I apologize. These are out of order from what I thought they were. We’re gonna talk about invariants first. So, there’s one other part about this that I didn’t, that I elected not to mention at the beginning, which is, we embed things in our brain the same way. Like, when our brain thinks about something, we don’t care how we’re looking at it, right? If I tilt my head to the side, now everybody’s face is rotated a little bit. But I don’t care. I still know it’s your face. And I don’t have a different representation in my brain for a triangle rotated at every degree. I don’t have any of these problems. And so, when you’re doing representation learning, a large part of it comes from introducing transformations to your input. So, if you have our input image here that’s a cat, maybe I grayscale the cat. Maybe I rotate the cat. Maybe I blur the cat a little bit. And the point is, is that all of these might be slightly different and might produce slightly different embeddings, but we want to try and minimize the distance between them. And so, this is that invariance. We want to make sure that our model is invariant to any sort of transformation done to the input information that actually will, but it’s still, like, not just destroying the information, right? If I just remove 90% of my image, I don’t expect it to produce the same embedding. (.) So, that right there works out to be our invariance term. (.) Next up, we have this covariance term, which is minimizing the correlation between the variables on an embedding. And so, what this is trying to do is that if I look at the covariance matrix between all of the values across a singular embedding, we want it all to go to zero, because that is saying that no, that is basically pushing all of the elements of our embedding to be orthogonal to one another, to where they never actually are impacted by another one, because this will maintain as much information as possible in the smallest amount of space, because there’s going to be no overlap. (.) So, we follow this all up. So, we have now our variance, invariance, covariance terms, and we have all those, and that’s great. But my talk is not on representation learning or self-supervised learning. It’s on spiking representation and spiking self-supervised learning. So, we actually have to move towards spike trends. So, as a very brief primer on what a spiking neural network is versus just a deep neural network, or pretty much most networks that have been talked about at this conference, is that they’re stateful models that exist through time. So, in your traditional deep neural network, you have a singular input, it propagates through, and then you have a singular output. Or, if you’re in, like, a recurrent neural network, you have inputs over time, and they, so you have an input, and it processes it over time, and you get an output. That’s the closest equivalent to what a spiking neural network is. So, in a spiking neural network, it is temporal over time. I don’t look at an image once, I look at an image for an amount of time, right? The time scales that we work with all vary. Some of them go down to the scale of about 0.1 milliseconds per inference loop, some of them up to three or five seconds per loop. It doesn’t really, it matters, but, like, for the sake of an embedding, it doesn’t matter as much. The other big thing that, when it comes to embeddings, is, so we’ve been dealing with colors, which are just me emasking away of a floating point value. Spiking neural networks don’t output floating point values. They output either 0 or 1. They’re based entirely on neurons in our brain, which communicate with one another through discrete events. Either a signal is traveling or a signal is not traveling. And this also causes a lot of problems for anybody who likes backpropagation. You’ll notice our lab doesn’t like backpropagation very much. Is that these are not differentiable, practically speaking. There is a part of spiking neural networks that likes to kind of hack them together to get you to have a differentiable function. But we’ve left the realm of biology. And if you’re going to be working in spiking neural networks trying to move towards biologically plausible and biomimetic systems, why are we trying to train them in non-biological ways? (.) So, that’s kind of where we’re at. And so they produce spike trains, which are just a train of 0s and 1s. So, what doesn’t a temporal encoding actually look like? So, if I have, in this case, 10 values in my output layer, at every single time step, they’re either emitting 0 or 1s. And what this means, and they are color-coded, so you don’t have to actually figure out what’s a 0 and what’s a 1. And what this means, though, is that I’m going to go through. And on the left here, I have one sort of spike. I have one time step, the next time step, and the next one. And then when they’re colored, there’s a 1 in there. The purple one has two 1s, and it’s just in the same spots that the red and the blue one spiked. Now, on the right side of this screen here, we have a different sequence of spikes coming out. But it’s the same spikes, just shuffled. And so here’s where you run into a problem. Spiking neural networks are stochastic. There’s nothing that says that the same input will always produce the same output. Now, if you run the same input repetitively, you will, in fact, produce a sampling range. So it will produce similar outputs, but they’re not always the same. So one of the problems that we have to tackle is, well, are these the same encoding? And traditional methods will tell you yes, if this goes. (….) Right, traditional methods mean that I can work technology. (…..) It was working. (…) Right, I’ve broken technology. Try again. (..) There we go. Okay. (.) So, traditional methods will tell you that yes, they’re the same. Because they convert to what things are known as rate code or graded values. Which is basically, we add up all of the spikes out of every neuron over the same fixed window, and we compare them. And in this case, these are the same. Now, I hold the belief that if we are working in spikes because we’re trying to build biological systems, why are we going back to rate to, like, floating point values? Because usually we divide this by the number of time steps. And so if we’re going back to these floating point values and doing all of our logic and training and inference on these floating point values, what’s the point of spikes? I’ve just built a system that is more tedious to implement sometimes and then thrown out all of the temporal advantage that it gets me. And so I don’t think that this is the future for spiking neural networks. I think we need ways of actually correlating the spike trains together with one another. And there are ways of doing this. They’re coming about. But there’s some interesting things that pop out of spiking neural networks, specifically the topology that naturally comes as a result of building spiking neural networks. (.) So we get decorrelation for free. So inside of a spiking neural network, generally there’s a lot of positive pressure, right? So we have positive voltages going into our neurons and we’re not getting into how LIF neurons work in earnest. But basically we have a bunch of neurons and they’re getting positive signals. And then eventually if they get enough positive signals, they emit a spike. Now the problem is, is that if I have an input and it’s just mapped to all of my neurons, they all get roughly the same positive signal. Which means that they’ll all start firing at the same time. Now one of the nice things about this and to combat this and something we already do in spiking neural networks is we implement lateral inhibition. So lateral inhibition is a very biological aspect of our brain and pretty much trying to build a spiking neural network without it is not going to work. And so lateral inhibition just works as on the right here we can see. We have our input. We have our bunch of excitatory neurons. They are laterally connected to a bunch of inhibitory neurons and then they go back. And so these are connected in roughly a 4 to 1 or an 80-20 ratio. And this is what it is in the brain. Is you have a pool of neurons, 20% of them are inhibitory, 80% of them are excitatory. And they’re just sparsely connected in a loop. (.) And so one of the things that these do is when you get an excitatory neuron fires, or you get enough excitatory neurons that fire, it’ll cause inhibitory neurons to fire. Which will basically quiet down all of the other neurons. And so what these do is if I have too many neurons that fire in my excitatory layer, the rest of the excitatory layer just stops. And so this naturally forces the neurons to not all produce spikes at the same rate. Or even in the same amount. I mean, even at all. And so we get this decorrelation for free out of a biologically, uh, you know, mapped system. And you don’t have to do anything extra for this. Like, there’s no fancy learning rule that you need to get this to work. Like, you can just use pretty much any learning rule that exists for spiking neural networks, and you’ll get this decorrelation out of it. And so that kind of brings me, uh, to one of my more chords, the neuromorphic representation for all of this. Which is basically, we have a lot of people working on credit assignment. We have a lot of methods that talk about how to get these things to learn based on stimuli, reward structures, all of that. We have dramatically less people looking into what does our topology get us, right? We, we get a, we have this constant battle of, well, everybody’s been building these this way. How do we get this to actually, like, what could we change to make it better? And obviously, inhibition does something, right? And inhibition gets us one of these three terms. And so my question then, that I really want to answer here, and the thing that I would really love to pose to everyone else, is find the neuroscientist that studies the brain, and then ask them, well, what are other parts of our brain, right? People like to throw around the word astrocyte as just does stuff with neurons. And they do. But we don’t know exactly how they work. So they do tend to be the cop-out answer of, well, the astrocyte does something. And then we throw math at it. And so I want to see, can we get other topological structures for our spiking neural networks that will naturally produce a embedding space that fits all of the requirements that we know we need to produce good embeddings. (..) And really kind of drive that home as, I think we could get a model that doesn’t rely on specific learning algorithms. We could get models that model after our brain more accurately. Because your brain doesn’t have representation collapse for most things. You don’t think that every single thing in front of you is a chair. You know? People like to think that you can sit on everything. But at the end of the day, we all know that they’re not chairs. (..) But at some point we learn this because every child thinks the thing in its hand is food. We have choking hazard warnings on everything. Because what does a child do as soon as it holds onto something? It puts it in its mouth. And it’s doing that to get information, for sure. But I think that eventually children learn what is and isn’t food. And so it doesn’t, not everything collapses down to it as food. And so I think that we can look to the brain and just biology and topology to solve a lot of problems. And we don’t have to leverage pure math. Which is saying something, seeing as my entire background is in math. (.) And so I’m just going to leave everybody with this. I’m apparently getting the conference back on schedule. (.) The open question of what other biologically systems naturally produce functionality from Vicreg or just produce functionality that we try and get with math. And I’ll open that up to questions as well. (11 seconds pause)

S34: Thank you. (..) And if you are a goat, it is all food. A skill I’ve always been jealous of.

S18: What was that?

S34: If you are a goat, it is all food.

S18: True, true. Everything is food if you try hard enough.

S34: Question. (..)

S12: Hi, Riza Rasool from Quiet. Early on in my career, I worked on the cochlear implant and on the Biodify project. (..) So you were asking, hey, what are the parts of the brain? Well, these are inputs to the brain. This is work back in 1997. There’s lots of literature on what the signals look like to make, to stimulate the immune spiral organ inside the cochlear implant. And it’s not voltages. They’re spikes. Yeah. They’re spike trains. You’re absolutely right. And then when we took that linear stimulating array, made it two-dimensional, and then surgically implanted them into humans to restore sight. They were not pixel values that were used to stimulate. They weren’t static molecules. They were spike trains. Yeah. And then since then, so that was in 1999. Since then, the bionic eye work has gone to become a cortical implant back of the head. And so there’s lots of literature you can find, and there’s probably lots of ground truth data that you can find to look at those areas.

S18: Yeah, I’m sure, I’m sure that’s a great resource. I’m just getting into this project, so that is definitely someplace I’ll be looking. (..)

S34: Thanks, Risa. And over here? (………)

S20: Thanks for the talk. You talked about different structures. (.) It seems like you’re still doing this. (….) I wonder critical. (……)

S18: Can you hold the microphone a little bit closer?

S20: Have you looked at thematic cortical loops in the brain? It seems like a huge structure for learning sequences. (…)

S18: So as far as, I have not looked at thematic cortical loops specifically. I do know that we have looked at structures and formats, and I’ve definitely had done investigation into structures and formats that are not just as simple as a feed forward loop. It’s like a feed forward network. The concept of recurrence, skip connections, loop feedback going into other things are definitely parts that we need. And honestly, I mean, you can look at things where you can just take, like, a giant pool of neurons, right? And then they work incredibly well together, and I don’t actually need to process all of them in some feed forward fashion in order to get useful results out of them.

S11: Thank you. Thank you. (.) Thank you. Thank you so much for the brief presentation. Thank you. I wanted to ask one, also, one possible, you know, application, which I’m not sure if your lab has, you know, tried it or not, but I feel like there may be a connection to the map there, which is, I wonder if, you know, spiking neural networks and the predict coding? Yes. You know, non differentiable kind of update, can be applied to a, the, the task of neural structure learning or causal structure learning. So I have had a fortunate chance to work on that, which I believe could be, should be a very important kind of subsystem for any AGI system. But basically, the task over there is that you are, you are given, say, observational table, observational data in the form of a table, you know, describing, say, ID observation of a system with many, many variables. And that’s the input. And the task is to do structure learning, by which I mean you produce, usually, say, a directed acyclic graph or a depth that describes the, say, the independences or the, the, the, kind of, kind of, feeds the data. (.) So that’s, that, that’s the task.

S18: Like pulls out features of the table that you’re being described or? Yes.

S11: So you are, you are giving a list of features and for each, all of this feature you have, say, 10,000 or 100 or whatever, let’s say, ID observational, observational data. And the task is to try to ask, okay, what is the possible, you know, causal structure? You know, causal structure in the sense of structure causal model from, you know, parallel definition. And so the output is kind of this, like, directly acyclic graph, describing the plausible causal connection from one variable to the other.

S18: So, I have not looked into that. That sounds very interesting. What I can say is that generally the way that spike, well, not all of them, but one of the ways that I’ve generally focused at learning, at training spiking neural networks is entirely, like, correlational or, like, correlational style learning. So, I would imagine that building models where causal loops are a part of them would be something that would fit in very naturally with that.

S11: Yeah. Yeah. That sounds, I agree with that. And, actually, the current existing work is also doing correlational based learning. Okay. The gap I’m seeing the current method is that this program is fundamentally combinatorial search program. And because of the, you know, mainstream machine learning and neural style learning, one of the way that people are trying to do it, they’re actually trying to force this discrete structure, the discrete depth, into a continuous matrix. The matrix form of representing a weighted graph. Okay. And then people are finding ways to express, kind of, designing functions, describing the level of a cyclicity in the graph. And then designing the function in such a way that it’s differentiable and then, you know, doing, and then people can do, you know, bad propagation. But I can see there’s many problems in that, like, numerical instability, you know, local minima and such. For sure. So I was just thinking, you know, spiking neural network and predictive coding because you are doing local, also discrete updates. It feels to me that it’s, it may be a natural fit to kind of address this kind of problem just from a totally different perspective of bringing something that’s, you know, fundamentally new. So.

S18: Yeah. It sounds like it would line up. Find me after this and show, and hand me a paper because I would need to read more about what we’re trying to target. But yeah, from what you’re saying, it would absolutely feel like it would line up.

S11: Thank you so much. Absolutely. That’s a talk offline.

S18: Thank you. (..)

S34: Thank you very much, Will. That was a fantastic talk, a fantastic presentation. (…) So here we’ve had two rising presenters at the frontier of AGI research. And I’m very grateful for their contributions here. A theme this year appears to be the young presenters and the young innovators in AGI. So on that theme, on that topic, I’d like to introduce a very special announcement. (.) The launch of something, I think, that just speaks to how seriously the academic community and the world are now taking AGI as a field of study. So let me introduce Gabe Axel Montes, Ben Gertzel, and Matty Clay to talk about the CIHS program launching this year. (……)

S23: Yeah, this will be a super brief announcement. (…) But while I’ve been in industry in the open source world for a long time, I started my career as an academic in math, computer science, and cognitive science. And I taught AI among other topics. While I taught classes in cog-sci, I never taught a class in AGI specifically because just when I was an academic, it wasn’t recognized enough, right? And I’ve, over the last five, ten years, I’ve had conversations with a few universities about doing AGI courses or AGI degree programs. Now, it seems the time has finally come, and I’ve been collaborating with a small private university based in Southern California, the California Institute for Human Science, to design a master’s degree program in AGI. It’ll be offered initially online, somewhere down the road. There might be a face-to-face version, but it’s online for now. And we’ve had other universities interested in doing similar programs, but we’re starting this here for real, right? And it should be quite cool. I mean, the first main faculty for the first year are up here on stage. We’ve got the intro to AGI, AGI and consciousness, and algorithms and data structures for AGI. And best of all, CIHS has brought on Dr. Gabriel Axel Montes as program director for the master’s program. So I will now turn it over to him to tell you anything I’ve left out. (…..)

S21: Thanks, Ben. Yeah, so just very brief, really. This is a long time in the making, and it really just took a, you know, a nimble institution to make something like this happen. An institution without a lot of largesse, you know, that would go through all these crazy long processes. (.) And now we have the first credited master’s degree program in AGI. So the TLDR is that there’s an info session on August 13th that’ll explain everything in detail. Myself and Ben will be giving this info session. It’ll be all online. It’ll be recorded. There’s a website with this QR code. We’ll lead you to it with all the information, the general curriculum, the list of all the courses. But basically, it’s a year and a half. It’s fully online. (..) And it’s, you know, there’ll be tons of material available. We’ll be covering the HyperOn stack, non-HyperOn material as well. But there’ll be a lot of opportunities for students to engage, to learn about all these different facets of cognitive synergy that make up, you know, the would-be seed AGI. And so this, there’s many different ways, you know, this degree could be useful to people. But it will cover consciousness and AI and AGI. So CIHS has a special aspect of it, which is it’s been studying consciousness for decades in its own right. So now, coupling that with AGI, it’s gonna be pretty exciting. So it’s gonna be, there’s gonna be technical work covering the technical foundations of AGI, hypergraphs, probabilistic programming, pattern matching, and so on, cognitive and neural systems. There’ll be electives. It’s about 58 quarter credits over six quarters. So that’s, again, a year and a half. And the first cohort starts in the fall, in late September. So basically, spread the word for the good of the future of humanity, essentially, and the budding AI minds out there, that we’re starting in the fall. And there’ll be an info session August 13th, and you can find more information online. Would you like to say anything, Matt? (….)

S25: Just simply, I’m excited to be a part of this groundbreaking effort on the part of Ben and Gabe.

S23: Yeah, thanks. I mean, I think AGI education is increasingly important. It’s a bigger topic, and we should be undergraduate, should be PhD level, should be down to kindergarten, right? And we’ll be rolling out some open course material for folks who don’t want to enroll in a, or aren’t in a position to enroll in a formal degree program. But I do think aggregating a formal degree program in a cohort of students like this is one interesting and valuable thing to do. And part of my motivation for doing this is wanting to spread the good word about AGI. Part is to further legitimate AGI just in the established world. And part is, I’ve had an almost finished AGI curriculum sitting around a long time. And this gives me motivation to actually finish it and deliver it, right? So I’m, I’m, I’m quite, quite, quite excited about this.

S21: Thank you very much.

S34: Yeah, and again, just, just sharing a taste of how more, how much more accepted and popular and accessible AGI is. We’ll, we’re gonna head into our final paper presentation here in just a minute. I just wanted to take a pause to note a couple of important upcoming items. So, at the end of the day, after this last paper session, we’ll have an informal poster session and demo session. I’ve seen tons of people at our posters already. We have so many good posters. Unfortunately, not everything could be accepted as a paper. There simply isn’t time. We’ll also have demos on ASI Chain, on Omega Claw, the Hyperon-ish agent. We’ll have a BCI demo. We’ll have Building Sentient Beings. They’ll be at tables set up along the side here. And then tomorrow, of course, we’ve got our, our general audience day, our AGI leaders day. We’ll have David Eagleman talking on what intelligence even is and how we’d know. Alison Goffnick on what human childhood reveals about what AI still cannot do. Hod Lipson on machine self-awareness. Neil Gershonfield on physical foundations we’re going to need. And real demonstrations, we’ll have Cody the robot up here. We’ll have Longevity with Raju Bio. Again, Omega Claw talking about these agentic systems and what they’re powering now. So, tomorrow will be an absolutely impactful day. Moving to a less academic center, to a more real world impacts of AGI, decentralization, social and economic impacts. And everything that follows on what we’re doing here these last two and three days of building AGI. So, I just want to make sure that that’s highlighted on everybody’s agenda. That they attend early for Ben’s talk on Hyperon and how SingularityNet’s looking to get to AGI. And then to understand what AGI means in a deeper sense to the world, to humans and beyond. (.) So, to finalize the academic portion here today, we’re going to go into our last paper session, intelligence as a physical and reproducible artifact. So, this is a session that asks what it actually means to build intelligence, not just theorize about it. So, this group is about the material realization of intelligent system, construction, verification, portability, and concrete technical processes through which intelligence gets instantiated. (.) That’s who I’m about to introduce. (.) So, our keynote for the session is Greg Meredith. Much anticipated. A mathematician, computer scientist, and one of the architects of Rho Calculus, a Firefly Distributed Computing Platform. His work on process calculi and the formal structure of computation has deep roots in the question of how intelligence maps on the physical substrates. Today, he’s going to take us from causal models to world models, showing how the mathematics structures and computation itself generates a natural notion of what it means for an agent to make sense of the world. So, Greg, I welcome you up to the stage. And then we will finalize our final four, three or four papers for the day. (31 seconds pause)

S30: Hey folks, do you have the latest link to the slides? (…….) Sorry. Got to make things complicated. It’s no fun if you don’t. Right. Well, first of all, I want to thank Ben and the organizers of AGI for inviting me to give this talk. Um, and I want to thank all of you for coming to attend. Uh, you know, I, I confess that I, um, when I began preparing for this talk, this is clearly not the latest version of the slides, that’s okay. Um, yeah, I, I, I, I was gonna go off script anyway, so it might as, might as well begin early. Um, I, I, I, when I, when I began thinking about preparing this talk, I, I realized that there’s such a, a gigantic iceberg of context, uh, that I have, um, that I think a lot of folks in the AI, uh, sector do not have. And so I was sort of spinning my wheels about, you know, which tip of the iceberg do I actually show you? Um, so, so, um, I, I thought I might, I might get into it, uh, in a way that I’m, I’m hoping will be a bit more, uh, um, experiential, but, but perhaps not. So, um, how many people in the room were here yesterday, um, during the metacognition session? All right, most, and, and how many people in the room were here during Yosha’s talk? (..) Okay, excellent. So, so you, you, you may have recalled, I, I asked a couple of questions, right? (..) So, so the question I asked the metacognition folks was, could you imagine, and I’m, I’m asking the same of you, could you imagine, um, that metacognition would emerge? So you, you have a bunch of agents, and none of them know the whole picture. They don’t know how the end-to-end system actually works. So, like, like bees in a beehive, right? None, none of the bees, none of the scouts, none of the, uh, not the queen, none of them know, um, (.) where they’re going to fly to when the, when the, when the hive splits into a swarm that’s gonna take off. Right? So none of them have the whole picture. Um, but they, but they come to a consensus about how, how, how to, how to go next. So could you imagine, um, metacognition arising in that way? (.) Right? So I’ve just, I’ve just given you an, an example of how that sort of thing arises in the natural world. There are tons and tons of examples in the natural world of how that happens. (.) Um, aha! There’s, wow. (…) So, so, so, so this is the preamble I wanted, I wanted to, uh, I wanted, I wanted to give. Um, and, and so the, so, so that, that’s part of it. But the other part was I, I noticed that Yosha, so he, he was very attracted to these kinds of things that happen in the natural world. None of the, none of those ants actually know the whole picture, right? There, it’s the, it’s the coordination amongst the ants that gives rise to this larger protocol, right? But when you, when, but the formalisms that Yosha was reaching for are very far from being able to write down what’s happening here. So I’m gonna walk you through a model of computation which makes it possible to do this. And if you remember when I asked the metacognition folks about this, two out of three folks thought it would be very hard. And in my experience working with folks with these models of computation, they, they, uh, they find it very hard, right? But if you, if you use the model in the way a spider uses its web to offload a lot of processing onto the tool out in the environment, it allows you to jumpstart this way of thinking. So that’s kind of the preamble. So let me, let me walk you through some, some, some basic ideas. (..) All right. So, so what I want to do is to, is to introduce you to this idea of a programmatic representation of environment, right? (.) And, and, and this, this actually showed up in the history of computing. But what we want is to be able to represent the agent as a program and the environment as a program, right? (.) And that means that you have to have a split, right? between, you have to divide the system into the agent, agentic part and the environmental part. So that split, that divide turns out to be very important when we’re thinking about computation. (…) So in particular, suppose that we wanted to make this model compositional. (.) What does it mean for a model to be compositional? I’m not sure if you’re, if you’re kind of in touch with, there’s a whole movement throughout computer science and mathematics about the importance of compositionality. (..) So in this picture, the way compositionality unfolds is that we would, we would take that upper level structure and we would, we would recurse it. So that would mean that program inside it would divide into program and environment. And environment would decide, divide into program and environment, right? So if we think about this is actually, this actually happens physically, right? So, so I, I’ve used the, the astronaut and the earth for a particular reason, but inside the astronaut we find cells, right? And, and, and environment for say your skin cells is a bunch of other skin cells, at least, at least in part, right? And then if I dive inside the, the cell, um, that we find, for example, the environment of the DNA includes the RNA. (.) And the environment for the RNA includes the DNA. (.) And you can play several games like this over on the environment side, but we don’t have time for me to do that. (….) But you, you’ll notice in, in, in that picture, I didn’t have a bottom layer. It didn’t bottom out, right? (.) So where would it bottom out? And this model of computation proposes that the, the, the atoms of this way of organizing computation are, are recognizers. (.) So like, like, you know, dendrites, um, or, or, or neurons, uh, and, and, um, emitters. So like molecules. Now, you might ask yourself, well, if my, if my pattern recognizer, uh, is going to have, it’s going to sort of look for a pattern off of a channel, and then do something having received that, that signal, what is the thing that it’s going to do? Well, how about more program? That’s a part of making this model compositional. (…..) Right? (…..) So now I’ve, now I’ve, I’ve, I’ve, I’ve, I’ve, I’ve, I’ve circled the loop there. (…..) Now, before I tell you how to, how to, how to ground the loop out, I just want you to observe that if we were to make this vertical bar, this divide between program and environment, be both associative and commutative, then what you really get is a mixture of atoms. Right? The tree flattens out into a big soup, right? (.) And moreover, if you do that kind of flattening, I’m not saying that it’s always a good idea, but if you do that kind of flattening, what happens is that program and environment become the same kind of thing. (.) So it’s, it’s programs all the way down, right? Which is, again, very similar to the, the sort of proposition that Yosha was making, right? Which is that everything is computation. It’s all a simulation in a certain sense. (….) Okay. Now, but we can do, we can go a little bit farther with that. Um, so if I were to say program in, in composition with program and composition with program, blah, blah, blah. If I just write that as pi program, right? Where pi, you know, big pi is, is just like a, uh, in area, in area, uh, parallel composition. Um, then that soup of recognizers and emitters refactors into two chunks. (..) There’s a bunch of recognizers on one side and a bunch of emitters on the other. Now, if you’re following along the mental picture that I’ve, I’ve suggested, right? Then, then you notice that we did something very funny, right? So we, we moved, we, we, we moved all the recognizers inside cells and cells and everything to one side. And all the substances that might represent emitters onto the other side. So we’ve erased all those boundaries. (.) But turns out that you can put all those boundaries back in a nice way if you were to assume that channels have a certain kind of structure. Then the structure on channels would re, reassert those boundaries. (….) I should have, I should have mentioned before I left. By the way, the recognizers form a kind of procedural memory. (…) So at, at a, at each channel, again, if you’re having, if it feels very abstract, think about a channel as like a location in space. It’s like that, it’s not exactly that, but it’s like that. So, so I, I’ve got this big coordinate grid. The, the, the, the points on this coordinate grid are channels. And at each point in this coordinate grid, I either have a recognizer or I have an emitter, right? (.) Now, the reason it’s not a coordinate grid is because the grid’s rigid and this structure is fluid. (.) And in fact, the communication topology reconfigures itself as it, as it, as it, as it computes. But we’ll get, we’ll get into that. (.) The, the, the, the, the, the key idea, though, is that we have a procedural memory and we have a fact base. (.) So that’s, that, that just falls out of the structure of this model of computation. We don’t have to do anything. We didn’t break a sweat. (.) Now, again, if I’m going to make this model compositional, I have to say what a channel is, right? I’ve left channel as unexplained. (..) And here, I’m going to borrow, I’m going to steal from giants or stand on the shoulders of giants. So, so, so Kurt Gödel observed something that life also observed 4 billion years before him. That coding is a fundamental aspect of computation, right? (.) So, so what Kurt Gödel did was to say, hey, I could, I could code up logic in arithmetic. And then if I have logical propositions about arithmetic, I can have logical propositions about logic, right? So that, that, that loop that Gödel constructed was necessary in order to, to, to state and prove his, his famous incompleteness theorems. (.) But life figured this out a lot earlier, right? It figured out how to code in proteins, how to make proteins, right? (.) So, so we’re, we’re just going to say, that’s a neat trick. We want to do the same trick. So we’re going to have channels be the codes of programs, right? (….) And the other thing that I, I left unexplained was the patterns. So what we’re going to do with, how we’re going to explain the patterns in terms of this structure is to simply say that patterns are programs with holes in them. (..) So this is a, this is a, also a trick. So while, while I have Kurt Gödel up in one corner, down in the other corner is Connor McBride, who figured out sort of the mathematics of, of poking holes in programs and how to make that to be a crisp, clean mathematical operation. (….) All right, so, so summarizing, what did we just do? I just told you what programs are. That’s the top line. That gives you the structure of programs. I haven’t told you how they execute. So the way they execute is if it turns out that when you, you can match the data that’s being sent. So in this case, Q is being sent on the channel, right? In this, in this top rule here. (.) When Q is sent on the same channel that the receiver is receiving on, and if the Q matches the pattern that we’re looking for, then that creates a substitution. So the data that’s in, the pieces of data that are in Q that fill the holes are substituted for the holes. So we rep, we write down that substitution as sigma, right? (.) So we apply that substitution to the rest of the program, P, and that’s the workhorse of this model of computation. This is what Yosha was calling the transition function. (.) Okay, so, so that’s one of the rules. And it says that that vertical bar, remember, we divided, we divided computation into program and environment. So this, this interaction between program and environment, this vertical bar is about interaction. That’s what the first rule says. Yeah, these are not the most recent slides, but it’s okay. So the next rule says that interact, that the vertical bar is also about autonomy. It says that if a program could make some progress, then if you compose it with an environment, say Q, it can also make progress on its own without having to interact with Q. Now, now, this rule says, is sort of limited because it still says we would do one of these interactions at a time. You can add a second rule, which is in the other version of the slides, the more updated version of the slides, that allows to have as many, as many programs as can make progress, will make progress. Okay, and that’s important because in a moment, I’m going to talk about how that relates to kind of breaking out of the simulation in a certain sense. And then the third rule just says, hey, there’s a bunch of structure in the syntax that doesn’t really mean anything with respect to computation. So we’re just going to, we’re going to say that we’re only really doing computation over equivalence classes. So you don’t have to understand this rule that well, because I haven’t introduced how to erase this structure, but that’s basically what it’s saying. All right, so now, yeah, before I get to this slide, I want to point out something. This structure allows me to have races. So I can have two emitters and one recognizer. (..) I could have two recognizers and one emitter. (.) And this model does not say which of those, which of the, what are the winners of those races? (.) It doesn’t say. In order to run this model, you have to supply a source of entropy. (..) That’s important because what people are finding is that when you add entropy to computation, you get new expressive power. And what this model does is to say, here’s the API for entropy. (.) The API is this vertical bar. It’s races. (..) Now, this slide is all about how we can make it even more biological, right? So in the case of a neuron, you have all of the branches outside the dendrites, right? And so here, this is saying you could actually guard the program that’s waiting on input with multiple independent interactions with a bunch of emitters. This makes it, this makes the model even closer to the biological sense, you know, that we’re attempting to emphasize here. Okay. So, now, lest you think this is just an abstract model of computation, it is far from that. So, my company, Firefly, has built a high performance implementation of this architecture. (.) Now, we have started way at the top of complexity, right? So, instead of trying to drive this, you know, down against the metal initially, what we’re doing is to say, hey, you can just use this for coordination of software components. So, SingularityNet and the artificial super intelligence alliance have built their ASI chain on top of this infrastructure that my company, Firefly, has built. But it doesn’t stop there. Boring Financial has also built their infrastructure on top of this. The world’s largest systems integrator, Tata, has just signed, the press release went out this week. They just signed an agreement with us to build on top of Firefly. We also just signed an agreement with DataCrypto. So, this is a very serious thing at this gross level of coordination. But what we’re doing is we’re marching towards the metal so that you can run this at the same scale and speed that you currently get with transformers and other kinds of networks. So, think of the, think about this as a framework in which to do, in which to begin to, in which to begin to write protocols where metacognition and other kinds of phenomena are emergent. So, but this is not yet a physical model. It’s a causal model, right? (.) So, what it really describes is a network of events. What are the events? The events are when an emitter and a recognizer actually interact and the data is transmitted. And this is, there’s a transactional interpretation here that is very powerful and why it works for coordinating agents and software tools and things like that. But, what it gives rise to is a network of these communication events which you can see as a causal representation of events in the world. Now, if you, so what it gives you are kind of the building blocks where you might begin to understand the notion of a light cone, space-like and time-like separation. So, those notions that we find in physics are, are, you can find analogs of them in this model. Now, if there are card carrying physicists in the room, I’m sure we’ll have a lot to talk about. But, but let, let me, let me go a little bit farther. So, how do we get closer to physics-like behavior from this causal behavior? Well, one of the things that we’re going to do is we’re going to add energy. So, interaction will cost energy. Then we’ll add history. And, and then we’re going to make interaction stochastic. And then the process of making interaction stochastic, you’ll start to see some of the things that we normally find in neural networks. (…..) So, you know, I think we all agree that, like, when you’re looking at a mainstream programming language, say Python or Rust or whatever, those programming languages are completely ignorant of the fact that running the programs spews heat out into the world and costs energy. It’s not baked into the model of computation. It’s like, it’s the generalization of the old joke about list programmers, right? List programmers know the value of everything and, and the cost of nothing, right? So, that turns out to be truer of the entire realm of mainstream programming languages than people like to admit. (…) So, what we’re going to do is we’re just going to decorate all our programs with the cost to run a particular program. (.) And then we’re also going to overload the way we construct programs with also being able to supply located sources of energy. (.) Now, I can’t, you know, since you don’t have a lot of familiarity with this model, going through this slide is a little bit challenging. Happy to talk about it afterwards. But the key takeaway from here is that there’s, this is an endofunctor, right? So that it’s not just this operation of adding energy to the transition functions turns out to be something you can do to every classical model of computation, right? (..) And if you do it twice, it flattens down to the same thing. So it really is a monad on a particular category. Now, one of the interesting takeaways from this is that energy automatically makes computation mortal. (.) Like, one of the things that we know about computation that is clearly not physical is that programs can run forever. I’m sure more than more than one of us in this room have written non-terminating loops or written deadlock situations or all kinds of things where programs don’t terminate. And we’ve spent tons and tons of skull sweat and other kinds of things to reason about when does a program terminate. Well, this operation guarantees termination for all programs. You decorate your programs with an energy supply and do the cost accounting and every program either runs to completion or runs out of energy. (.) So mortal computation turns out to be an endofunctor on a particular category. (…) Dual to energy in very much the same way that we find that energy is dual to time evolution and physics is adding history. So one of the things that’s weird about physics is that it has a terrible accounting of cause. (.) Why is that? Well, because physics is totally sold on this idea that I can run time forward and backward. (..) But that makes it really, really hard to explain why there’s an arrow of time. This is a fundamental discussion across all of physics. Why do we find an arrow of time that corresponds to our own experience? (.) Well, it turns out that I just told you this approach gives you a causal arrow. It puts time inside the model, right? (..) And now, how could we recover reversible time? Child’s play. You add a tape recorder, right? (..) You just record the causal events. That’s what we’re doing with history. We just record the causal events. And guess what? When you spend energy to do your computation, you record what you spend over in the history. And that’s why history is dual to computation, right? (..) And those two monads do not swap. There’s no distribution law. So you have to spend energy to make a tape recording. (.) But you could take a history of that, right? So there’s this infinite tower of cost, history, cost, history, cost, history, right? That says something about the simulation hypothesis that nobody’s talking about. (…..) So both of these are monads on a particular category. I don’t have time to go into it. But it’s a general construction. It is not limited to the model I just showed you, right? (…..) Now, stochasticity is inherent in the interplay between micro and macro states, right? So that’s, like, we know that we experience stochasticity in the world, right? You put the snowball at just the right place and you’re not quite sure how it’s gonna roll, right? So that’s, so you want that to be explicit in the computational model, right? (.) It’s not explicit in Python. It’s not explicit in Haskell, not Idris, but there’s a way to make it explicit. (….) So why would you want to do that, right? So let’s, let’s think of, let’s place this in our current moment. There are over 7 billion smartphones in circulation. There are 1.4 billion tablets, 2 billion laptops in circulation, right? The average user uses about 9, 10 apps daily, and there are over 100,000 apps that are released across the major platforms each month. So if you then also factor in all the protocol stacks and the servers in the data centers, the glue that basically gives us our current experience of the internet, what you’re talking is roughly 3 billion apps running, sorry, 3 trillion apps, let me do my math right, 3 trillion apps running simultaneously to give you the experience of the internet that you know and love. Well, you know what? That puts us in the, within an order of magnitude of the number of cells in the human body, about 30 trillion cells. Now think about what this means, right? If we don’t want to keep having events like the CrowdStrike event that, that grounded planes and, and, and crash telephone services and everything that we saw in the summer of 2024, if we don’t want that to be a regular occurrence, we have to be able to lift up and reason about programs like we reason at the, about biology. So we need these kinds of statistical methods, right? (…) So, so what we’re going to do is we’re going to describe classes of programs, right? (.) So, so in the case of the model that we were just talking about, we identify formulae that pick out channels, formulae that pick out P’s and Q’s, right? And, and that gives you both, the formulae give you the, the macro states, the ensembles of programs, whereas the, the individual programs give you micro states. And I’ve just now related the, the micro and macro structurally. that turns out that that structural mapping is even closer than we thought. There’s a way to auto generate the macro descriptions from the micro descriptions. So that turns out to be a monad, right? So in the case of this, we, we want to, let, let me show you how it works out just, just briefly in terms of the model that we’ve just been talking about. So we, we, we want a formula for the channel, a formula for the, for the, the, the rest of the program, the continuation P, P, a formula that, that picks out the Q’s, right? And we’re going to give each triple a weight, right? So a weight map will be a collection of these triples mapped into a weight, right? And now what we want is that when you transition, you’ve actually done one of these, these, then you update the weight map. Sound familiar, right? So it turns out that this, this structure is enough to auto generate a simulator. (.) So Microsoft, by the way, I did the first generation of this computational model inside Microsoft. It was called BizTalk Process Orchestration. It made 75 million in its first year in terms of coordinating software components across the industry. By the, by the, by its fourth year in production, it was voted best software product of the year, right? So this is, this is not just math. This is real practical stuff. (.) So anyway, the, the, the, the, the point, the point here is that we can, inside Microsoft, a group took this same kind of approach and applied it to, to what’s called the stochastic pie calculus and use that to describe biological processes, which were subsequently used, for example, to, to discover processes in the, in MS, right? So the, the leukocytes end up rolling and tethering along the cell wall. And that’s one of the factors that gives rise to, to the inflammation. Now it turns out that in, in that setting, the differential equation models did not make those predictions, (.) but the, the, the stochastic pie calculus models did make those predictions, and they were subsequently verified in the lab. So again, this is serious stuff. This is not, this is not just a mathematical abstraction. But where did the formula come from? So they come from a family of algorithms called OSLF, Operational Semantics in Logical Form. So, basically what it does is, you feed in, as input, a model of computation. You get out, as output, a, a spatial behavioral logic that is adequate for the model of computation. By adequacy, what I mean is, there are, two programs in a given model are bi-similar, and we should talk about that, if and only if they satisfy the same set of formulae. In terms of, of, of, say, for example, Yosha’s talk, you can think of bi-simulation as the formalization of functionalism. (..) Bi-simulation is, in math, the functionalist paradigm. (….) All right, so I’ve, I’ve, I’ve, I’ve, I’ve just gone over the contents of this slide. This is, just gives you an example of the kind of logic that we automatically generate. (.) So we auto-generate this logic from the model of computation. (…..) So, now, the reason this is important, one of the reasons this is important for AI is that, first of all, what is AI’s value prop to the scientific community? That useful, salient, even important aspects of intelligence, and almost always AI folks mean human intelligence. There are lots of different kinds of intelligence in the world. (.) It can be faithfully rendered to computation. (..) There are interesting points about computation that aren’t made clear. For example, computation is ontologically isolated. (..) Think about that for a minute. Suppose I have a Java program, and I want to wire a sensor up to it. Well, if I don’t have a Java programmatic representation of the events from the sensor, that program cannot utilize the data. And this is true generally. Computation is in a box. (.) It can’t actually see the outside world. (.) But there is a world that it does see, and that’s the world of the bi-simulation equivalence classes. So the bi-simulation, bi-simulation is the finest relation, equivalence relation, respecting behavior on programs. So there’s a fixed objective ontology for every model of computation. (…) So if we wanted to build a science bot, if we wanted to know what it’s like to be a bot, let’s think about this. What does our science bot have to do? Our science bot needs to be able to formulate hypotheses about its environment. What we just said, programs are ontologically isolated. So the only environment a computation can have is other computations. (..) So to formalize a hypothesis about its environment, it needs one of these logical formulae. So these logics that we just generated is the language that is in perfect conformity with the computational phenomena. (.) So it expresses those in the, it expresses its hypotheses in these logical formulae. And now it needs to make an assay or a test to determine whether or not its environment is a witness for the formula. We know how to generate those. (..) So this gives you a, and then it needs to be able to munge the formula, like compute over the formula to revise it based upon what it gets back. There’s your science bot. That’s a very different kind of learning algorithm. We can go even further. We can motivate, because like why, why would the computation do that? We can motivate it to do that by, by giving it a supply of energy. (.) And we decrement the supply of energy when we, when the delta between predictions is getting, and real world results is getting larger. Right? So, so if, if, if, if the predictions are getting off with respect to the real world, and the, and the, and the gap is widening between prediction and an actual experience. (.) Right? That means that it’s wandering around and never, never land. Right? It’s predictions are going to likely get it killed. So we decrement its energy supply when that happens. (..) Now, by, by contrast, if it’s, it’s, it’s predictions are getting better. So we’re narrowing the gap between prediction and environment. We replenish the energy supply. Now, a collection of, of such programs, if you run them again and again and again and again and again, the ones that persist. Like, we don’t even, we don’t even care. It’s just the ones that are still around. Right? They have learned. (..) Right? So, so I’ve, I’ve just now given you a completely different mathematical framework for learning. (…..) All of these are all monads on the category of graph structured, uh, uh, uh, uh, lambda theories. So that was the category I was mentioning. (..) Um, and this is all realized in an actual programming language. Uh, in, inside Firefly it’s codenamed meta IL. You’ll see demos of this, uh, uh, later today, around 7 PM tonight. You can come along and, and check it out. Um, basically it uses, uh, the row calculus as the control flow language. And it uses, um, a, a, a very fancy, uh, data definition language, uh, to, to describe a, a variety of models of computation. And I, I’m, I’m, I’m conscious that I’m probably way over time. So I’ll, uh, I’ll, uh, I’ll, uh, I’ll conclude there and see if there are any questions. The rest of the slides are just examples of meta IL. Um, but, but you can see that, uh, in detail at the poster session, uh, this evening. So are there any questions? (13 seconds pause)

S07: Thank you for great presentation. Uh, since I’m not familiar with it, I’m just trying to clarify, uh, let’s say channels. You mentioned channels also are programs. But how do you discover channels are the dynamic and what kind of programs in the channels? Is it like, uh, uh, serialization, uh, synchronization of some sort? Or there’s additional kind of any type of process you can put inside the channel? Essentially, general purpose programming or demo special? (.)

S30: Uh, so, so thank you for that question. It’s, it’s, it’s a, it’s a very natural question to ask. So I want to make a distinction between programs that are channels and programs that, that show up inside channels. Right? So you use a channel to emit some data. In this case, the data is just programs. Right? And it’s any, any legal program in, in, in the, the, the, the role model. Likewise, the channel is any legal program that, that you’ve taken the code of. So, so we have this operation to be able to take the code of a program and we use that as a channel. Right? So think of that like girdle encoding. Right? And, uh, but we also have a dual operation which allows us to take a channel since we know it’s code and turn it back into a running program. Now it turns out that with this structure, the communication topology is completely dynamic. Right? So you can model things like, like what happens at conferences. Right? You walk up to someone you don’t know and you exchange email addresses and now you’re connected and you can communicate. Before, there was no such communication channel between you. After that event there is. Right? So, so the models, the model, uh, gives a faithful representation of, of that kind of phenomena. (………)

S23: So, I, I mean, I’ve seen most of this before but each time I start thinking different things. //S30: Oh, good.// I’m thinking about, and I have a question I’ll lead up to. (.) To what extent and in what ways non-determinism and choice really does reduce to race conditions? Now it seems like you, I see how you could compile any instance of non-determinism or choice. You could compile it into some, in some cases, contorted combination of race conditions. (.) On the other hand, that doesn’t mean that’s always a natural way to look at it. So, like, you could have something that’s combinatorial where, like, you have three things and any two of them will allow the thing to happen but all three won’t allow the thing to happen. Or you could have other fundamentally higher dimensional choices. You could compile them into some race condition by making, you know, terms that correspond to combinations of things. But it’s, in some cases, going to be a very weird, weird way to do it. So, my, my question is, first of all, like, in what cases do you think compilations and combinations of race conditions actually is a semantically convenient way to represent choices rather than just a possible way? And if you have cases where that’s not a semantically natural way to compile choices, like, how does Rolang deal with that?

S30: So, that’s, it’s a, it’s a great question. Since you’re into algorithmic chemistry, let me just point out that most of chemistry, although we are aware that there might be ternary interactions, they limit themselves to binary actions because ternary, quaternary, etc. (.) are all…

S23: And, and, and, and, and, and enzymatic catalysis is ternary, right? (.)

S30: Sort of, right? But, but again, that’s treated as a, that’s treated as a special case. You know, it’s a, it’s a, it’s a, yeah, it’s, it’s treated as a, as a special case. And that’s exactly what we do with the cost accounting, is we, we treat that special case in exactly the same way. So, essentially, what we’re doing is we’re following the pattern that chemists have been using for over a hundred years now. (..)

S23: I find that for, I, I, I, I do buy that, although in my algorithmic chemistry experiments in meta, I had to force the coding agents not to treat catalysis as a special case. //S30: Sure.// Because I didn’t want to, but I, I, I get that part. So, how about practical everyday software as opposed to, to, to chemical systems? Do you think it’s the same story then?

S30: Yeah, in fact, like, so, for example, there was a form of non-determinism that was happening in the Ethereum virtual machine that allowed for the, the DAO bug. So, essentially, the DAO bug was caused because of a re-entrancy issue, right? So, so could a new request be served? Well, the, the, the new request was essentially racing against the update that was caused by the re-entrancy of the, of the virtual machine. So, when you re-write the solidity code in Rho, the bug is made evident. It’s, it’s super clear. It’s, it’s actually much clearer than it is in the solidity code. So, that’s a practical example.

S23: I mean, that, that, that, that, that is an overly low bar, but I still, I still believe it.

S30: Well, sort of, sort of. But when you have spaghetti code, it’s not necessarily clear that all the weird non-determinism that might be showing up in the, spaghetti code is, is naturally going to resolve into these neat, clean structures. And it just so happens that it does.

S23: Yeah, that’s fair. That’s very nice. I just wonder if maybe the fact that chemical structures and human software resolve mostly into binary or ternary dependencies, maybe it’s, it’s, it’s just because.

S12: Yeah. (.)

S23: Humanity and, and, and Yahweh are stupid, right? (.) Perhaps that, that, that there are other, other more interesting, more efficient, more functional structures that don’t decompose so nicely that well.

S30: Well, I mean, I would, I would be keenly interested if you, if you find any. I just think that the main thing has to do with statistics. The, the more coordination you get, the more coordination costs you have, right? And so, that, that, that, that, that, that.

S23: We would need special higher dimensional structures. //S30: Yeah, exactly.// And then you would need higher dimensional row length. (..)

S30: Believe me, we’ve been working on it.

S25: So, a question related to that. So, as I understood, you are having a kind of process algebra that resembles the pie calculus where you can dynamically create channels and so on. //S30: That’s.// But there’s also the ambient calculus where you can change the spatial structure, the structure of containment. Would this be kind of what? (..)

S30: So, so for, for, for example, let, let, let, let, let’s compare the three calculi briefly. (…) So, it’s Luca Cardelli, who invented the ambient calculus is, is very proud of the fact that there is no fully abstract encoding of ambient into pie. And a lot of that is precisely because pie remains willfully ignorant of the structure of names, right? (.) So, the, the channels that you introduce in pie calculus, even though you can make new ones, right? It, it, it, it, it suffers the fact that you, there’s no structure you can hang onto with respect to the names. Right. So, once you add structure to names, which is what the row calculus does, you, you can get rid of that pesky new operator. And, but it also makes it possible to do a fully abstract encoding of the ambient calculus into row. But, but the, the, the, I, I should, I should, I, I need to hasten to point out. I’m, I’ve focused on the row calculus, because it, it’s one of a, a wide class of tools, the mobile process calculi, that make it possible to reason about these protocols where no one’s in charge. And that’s a blind spot for humanity, right? Humanity is really, really sold on this narrative of, of, of command and control. And, and it shows up in all of its abstractions, right? And it’s really hard for it to get, to get to grips with this idea of, of autonomy. And lots of, lots of agents working together, even though no one has the complete picture, right? So, so that’s, that part is important. But this work is actually situated in the category of graph structured lambda theories. Which, which, which Mike Stay, who’s in the back there, and I and Christian Wells identified. (.) Where you, you, a graph structured lambda theory basically gives you, the objects have three components, right? A, a grammar to describe your term language, some equations, which whack the term language down to, to, to, to just the things that you need. And not, not accidental syntax. And then some rewrites. That structure is enough to give you, like, if you, if you, if you get rid of the equations and the, the, the rewrites, you have algebraic data types. If you add in the equations, you have all of universal algebra. And by the way, computation has been saying something very loudly to classical mathematics, which, which nobody’s really paying attention to. Which is, that when we add the rewrites, it factors behavior very differently. So, for example, vector spaces just sit there. The vectors don’t do anything. In order to get any dynamics, you have to have the morphisms between the vector spaces. All right, so, whereas lambda calculus, so going back to church, right, packages dynamics differently. He says the, the terms, the structure, and the behavior are part of a one package, right? So when you, when you build these categories, now you have to ask, what is it that morphisms preserve? And that’s why by simulation preservation is a, is the key idea in this, in this new category. So, so all of these things that we did, adding energy, adding history, generating logics, adding, adding weights, things like that. Those are all monads on the category of graph structure lambda theories. So those same tricks work for lambda, they work for Turing machines, they work for SKI, all of that stuff. But, but what picks out the, the mobile process calculi is this thing about autonomy and interaction. (.) Does that make sense?

S25: Yeah, sure. And so, so you have kind of data structures that enable you, on the, on the channels that enable you to kind of capture this mobile structure. But is this embedding that you get of the ambient calculus into your role calculus, is it natural, or?

S30: It’s fully abstract. (.)

S25: Yeah, I mean, in terms of everyday programming, is this a thing that is natural? Oh. Does it feel natural, or is this something that?

S30: I mean, I, I, that, that, that may be a subjective matter. For me, it feels natural, but I’ve been programming this way for a long time. (..)

S05: Thanks. Sure, thank you. Good question. (………)

S32: Thank you for the talk. It’s wonderful to see you. Thanks. So, I have a question because, you know, given the, you know, the separation between the grammar, the rewrite, and equations, and I’ve, I’ve worked with some structures of that sort, but it feels to me that although that this describes, like, dare we say, like, a, a, a, a foundation, right, it doesn’t entirely address the design within that foundation. In the sense that, like, you know, if you have equalities that loop, then, like, you can’t, it’s hard to describe, like, what the equivalence class could then, what if a rewrite happens during, you know, during some sort of equality, right? Or, does this make sense? Like, it’s a, or, or just, like, how do you encode, like, the choice of how you encode these names. Like, it’s possible, but is it, is it, like, you know, obviously the wrong word, but there’s, there’s still, it feels like the, the design on top of this space is non-trick. (.)

S30: Um, okay, yeah, so that’s, there are actually several sub-questions to that question. Here, here’s what we found. So, what I, over the last 25 years, I’ve been working with trying to get programmers to, to code in the mobile process calculate, and asking, and what we found is anybody who’s done event-based systems, so, like, if you, if you programmed a, a web, a web server, right, or, or, or things like that, event handlers, game loops, things like that, those programmers pick up this model like that. They’re, they’re, they’re off and running. Um, so, so maybe that gives you a sense of, uh, the kind of design patterns that are very easy, uh, in, in, in this model. What’s more interesting is when you start thinking about how to map other domains into this, this way of thinking. But again, you know, the, the tools about adding energy and generating logics, that stuff is generic, right? (.) So you don’t have to buy into the mobile process calculate to avail yourself of those tools, right? And, but, and, and part of the, the, the bigger picture that I’m trying to get across is, if the value prop is, intelligence can be faithfully rendered as computation, then focus on computation. Identify the universe of, of computational models, and then say what you can do with them. And what I’ve done is to say, here’s that universe, uh, you know, given the best we know, and here are some of the things that you have to add to make computation look more like physics. (.) Does that make sense? (……) Thank you so much for your attention. (………)

S34: Yeah. Oh, big breath. (.) Long day, three paper sessions or three papers left to go. Big breath from the audience. (.) So much learning here. I hope everybody sees, uh, here and virtually, right? There’s a big audience watching us here today. that, uh, AI is not that thing over there. AGI is not that thing. Somewhere else it’s accessible. It’s some of the, uh, math is, uh, complex. But we are building systems that do make this stuff, uh, real world accessible and younger and younger generations joining in on it. So, uh, I, I’m excited to be able to be here to present this. I’m excited for this, uh, final three posters, um, to close the day. And I hope everybody’s excited with me. So, we’re gonna launch now three papers on intelligence as something that gets concretely built, verified, and instantiated. We’ve been across the range of what intelligence might or could be. And here we are as something concretely built, verified, or instantiated. We’re gonna start with the pre-recording from Xiaohan Ma, compiling PLN to thermodynamic hardware. (23 seconds pause)

S02: Hello, everyone. I’m Xiaohan Ma from OCD University. And this talk is about compiling probabilistic logic networks to thermodynamic hardware. I take PLN strength inference and run it as Boltzmann sampling on that hardware, while the evidence count is computed in parallel on classical hardware. I will walk through how the compilation works, and then show an end-to-end validation of four hardware-bound inference rules. (.) In PLN, every proposition carries a two-value truth. A strength set between 0 and 1, which is a frequency estimate, the probability that the proposition holds, and an evidence count n, the amount of observation behind that estimate. Current implementations run on classical CPUs and GPUs. (..) Now the new substrate, the thermodynamic computer. It’s elementary unit is the p-bit, short for probabilistic bit, a programmable to value the random variable, plus 1 or minus 1, with a tunable bias H and pairwise couplings J. Each p-bit flips stochastically according to a sigmoid of its local field over a programmable temperature T, and collectively the device physically samples the Boltzmann distribution of the icing energy built from those biases and couplings. Two advantages matter for this work. First, couplings and stochastic updates are integrated on a single device, so it operates as logic in memory. The computation happens where the parameters are stored. Second, each p-bit updates in parallel with its neighbors, in contrast to several software deep sampling. This parallelism is the potential speedup that motivates compiling PLN onto this hardware. (.) The compilation splits PLN’s two valid truths across two substrates. Strength is a probability, and probabilities are exactly what a Boltzmann sampler natively produces. So each PLN edges conditional probability table, compiles into icing biases and couplings, and block it sampling on the compound energy without the conclusion strength. The evidence count, on the other hand, propagates by closed-form formula, so it simply runs on classical hardware in parallel with the sampler. (.) Here is one full inference step with modus ponus as the example. The premises goes in, the two paths are computed independently, and the output is simply the pair

S14: of both values.

S02: Now to the compilation itself. A PLN edge supplies one CPT entry. The link strength, the probability of the child given the parent, the complementary entry, the child given non-parent, is stored with the link as epsilon, defaulting to 0.02. For any DAG, the lock join factorizes edge by edge. Writing plus one for true, each factor expands exactly over four spin monomials, a constant, the parent’s spin, the child’s spin, and their product, that yields closed-form coefficients. Applied bias on the root, parent and child biases per edge, and one pairwise coupling. For chains, the coefficients simply add. (..) Five rules in total. Modus ponus compiles to an icing chain with the root clamped. Deduction becomes a three-spin pairwise chain. Abduction is inversion followed by deduction. Inversion itself is a two-note joint icing model. Each rule has its evidence or confidence formula running on the classical path, shown in the right column. Revision is the deliberate exception. It is a weighted average over evidence counts with no sampling structure, so it bypasses the sampler entirely and reproduces PL-Lens exactly on classical hardware. (.) Validation. Ground truth is PL-Lens’ own closed-form formula, so the test asks directly whether the compiled sampler reproduces PL-Lens’ algebra. (.) The path criteria is plus or minus 0.05 on strength, and everything runs on the public thermal simulator, not on real hardware yet. Every parameter group passes. The worst case is 0.03 on a low-strength multiple setting while inside the tolerance. (.) Three limitations. (.) First, everything runs on the simulator, not on real hardware yet. Second, the factor graphs are small, two to three nodes. Scaling behavior is unknown. Third, the coupling depends only on strength, not on evidence count behind it. Two edges with the same S but different N compiled to identical icing parameters. (.) Future work follows directly. Validate data on real P-BIT hardware and quantify the simulation to hardware gap. Scale to larger factor graphs and measure the scaling law of any approximation error, and integrate the compilation into hyperon. To conclude, PL-Lens admits an additional execution path on thermodynamic computers. (.) Strength compiles exactly to isin bias and coupling parameters, and is read out by P-BIT block gap sampling. The evidence count runs in parallel on classical hardware by PL-Lens own formulas. And the compiled joint factorizes exactly across two paths. All four hardware bound rules pass end-to-end validation on the thermal simulator. This contributes an end-to-end validated PLN on thermodynamic computer pipeline. Thank you. (15 seconds pause)

S34: Next, we have Masayuki Hata. Reproducibility is the new copyleft, defining AGI-oriented reproducible builds. (…)

S10: Most of today’s presentations are quite technical with hardcore mass. (.) Mine’s not on AI governance. (.) So sit back and relax. (.) I’m Masayuki Hata. My question today is what open source can still mean for AI and eventually for AGI? (.) My claim is this. Copyleft, the legal mechanism that made free software work, has quietly lost the technical foundation it always depended on. Rebuilding that foundation is not a job for licensed techists. It is a job for reproducible builds. The guarantee that a declared set of inputs verifiably produces the artifact you are running. Let me start with what copyleft actually guaranteed. (…) By the way, this paper’s definition of AGI is fairly narrow. an AI system capable of recursive self-improvement or self-replication, a la Ashiroma principle. (.) Also, I assume you are already somewhat familiar with concepts like open source, copyleft, and reproducible builds. (.) The short explanations for these concepts are in the paper, so please read it. Generally speaking, the paper is much better than me, so please read it. (..) Now, back to the main topic. Copyleft is usually described as a legal hack. It uses copyright in reverse. (.) You may copy and modify on the conditions that the derivative works travel under the same terms. But there is a hidden technical premise underneath. (.) GPL version 3, GNU GPL version 3, which I was involved with revising in 2007, defined the corresponding source as everything needed to generate, (..) install, and run the object code. The license simply assumes that the deterministic human auditable build converts source into binary. (.) So what copyleft really enforced was equivalence. Whatever you run must be derived by a disclosed procedure from disclosed source. When that equivalence fails, the license can still be satisfied while the freedom evaporates. That equivalence, not the share-like clause, is the normative core of copyleft. (……) Modern machine learning breaks that premise. The trained model’s behavior is jointly determined by at least seven things, I think. The training code, the dataset, the hyperparameters, the random numbers actually drawn during training, which depend on worker scheduling and not only on Cs. The tool chain, down to CUDA, and the floating point in three things. The specific hardware and the weights that all of these produce together. (..) Even the same code and the same data on different hardware will not give you the same weights. (.) Release any subset, only the weights, only the code, only the date, etc. (.) And third party can neither rebuild the model nor verify that the released weights came from the disclosed procedure. (.) The free software for freedoms, learn, study, redistribute, and modify your software for any purpose become unexercisable. (…..) The community has answered with definitions. Open source initiative, OSI, open source AI definition, which I was involved with drafting in 2024, requires the training code, the parameters, and what it calls data information, (.) which means enough for a skilled person to build a substantially equivalent system. (.) That is a pragmatic compromise. Many corpora cannot be redistributed at all. The Linux Foundation’s model openness framework goes further up to a class demanding raw datasets, checkpoints, and logs. It’s companion license, called OpenMDW, is permissive rather than copyleft. But notice the gap. (.) Even complete release at the highest tier does not guarantee that rebuilding yields bit identical weights. All ingredients disclosed is not the same as verifiably rebuildable. (……) Now, Stefano Maffoli, the former OSI director, recently argues that the GNU GPL was only a permission slip. It gave formal freedom, but exercising it required a scarce programming skill. (.) AI coding assistance corrupts that barrier. So AI, not a GPL, finally delivers substantive software freedom. (.) There is real insight here, but the argument works only if the liberating tools themselves open source AI. And Maffoli concedes as much. If your assistant is in an opaque remote API, you have swapped dependence on the maintainer for dependence on the vendor who can change the model. (….) Sensor refactorings or log your code base. And nobody audits a Torilong parameter model by reading it. The only audit available is procedural. (……) Also, there is a darker corollary, Maffoli does not draw. The same capability lets you wander copyleft, feed a GPL library or something to an assistant, ask for functionality, functionally equivalent, rewrite, and sip it under any license. (.) Copyright protects expression, but not functionality. (.) This March, we saw the first high-profile case. It’s called Chardet, character and character determine library. Maintainer used an AI agent for a ground-up, MIT-licensed rewrite of an LGPL, GNU-RESI-GPL library, in about five working days. With reported structural overlap under 1.3%. (.) The original author disputes the clean-room defense. (.) My point is economic. (..) Copyleft assumed rewriting was expensive, so compliance was cheaper than evasion. That asymmetry is dissolving. Library by library, reproducibility does not prevent such relights. But once a share-like loses its economic footing, the verifiable procedure is what remains to govern. (……) That procedure is what the reproducible builds movement has been engineering for ordinary software. Build is reproducible if anyone, starting from the same declared source and toolchain, obtains bit-for-bit identical binary. Like Debian, I’m one of Debian developers, or Debian, Tails and Nix OS have invested decades in this. Because this closed source cannot protect you from a compromised build server unless the correspondence can be independently checked. (.) Notice the symmetry. Copyleft legally ensured that every binary was accompanied by its source. Reproducible builds technically ensure that every binary is derivable from its source. Together, they discharge source-object equivalence. So, can a training pipeline be made reproducible the way a compiler is? (…..) Increasingly, yes. I should say it’s preliminary, but there is a strong possibility. This is engineering, not magic. (.) Recent research by Chen and colleagues reproduced seven deep-learning models exactly with record and replay. (..) On inference, thinking machine lab shows that non-determinism at temperature zero comes mainly from the batch size dependence of reduction kernels. Their batch invariant kernels gave bit-identical outputs across a thousand runs at roughly a 60% throughput cost. SGLang cut that overhead to about 34% and extended it to reproducible reinforcement learning. So, the question is no longer feasibility, but which part must be deterministic? And at what price? So, my understanding is, if we kill GPU-level optimization, we can get some deterministic results with costs of lower throughput. Then, we formalize this as seven requirements for AGI-oriented reproducible builds. R1, complete input enumeration. (.) The exact corpus, code, seeds, tool chain, machine readable. (.) R2, deterministic training pipeline or explicitly declared reproducible budget. R3, verifiable tool chain and hardware binding. R4, third-party attestation. Neutral public rebuild infrastructure as Debian has for packages. (.) R5 is one distinctive to AGI. An append-only signed log of the entire self-improvement trajectory. So, I assume the AGI system can, you know, can do self-improvement. So, it should keep the log, how it improves itself in the spirit of certificate transparency. R6, recursive verifiability, is a research target rather than a deployable specification, I should say. Self-modifications must preserve the reproducibility invariant. That is the AGI-era analog of the viral clause clause or share-like clause. Propagating a technical property instead of a licensed term. R6, which means R6 is basically the equivalent to copy left. Because one system is reproducible, then modified AGI system has to be reproducible. (.) And R7 keeps the framework feasible. Start with small models, verify by sampling, scale obligations with capability. Because no one can produce or reproduce like a frontier model, which a big tech company has. So, we have to allow the sampling by some verification by sampling. (….) One more layer. AI systems increasingly interact at runtime through protocols, such as model context protocol, MCP. Model invokes external tools, much as program loads dynamic library. And A2A or function calling play the same role. Copy left never cleanly reached the linking layer, even for ordinary software. (..) That is why the LGPL exists. And across a network boundary, nothing distributed at all. (.) OSA ID and MOF govern what each node discloses. They say nothing about the edges. Here, the beta template is Mike Masnick’s protocols, not platform argument. Keep the specification neutral and push contested decisions to the edges. (.) As SMTP, mail protocol, does for email. Concretely, a new neutral standards body, portable authentication, user-controlled logging, and licensing metadata on tools. To conclude, copy left worked because of technical fact. The deterministic relation between source and object code made legal obligation meaningful. That fact no longer holds for modern AI and will hold even less for AGI. (.) This paper proposes two layers and two instruments. For the production layer, how a system comes into being reproducible builds. The seven AGI-RB requirements. (.) For the linking layer, how system interact, protocol governance, rather than copy left. In the near term, I would require R1 to R4 of system. System is deployed in public interact domains. Phasing in R5 and R6 as capability majors. (..) The first liberation gave us the code. The second must give us the code and verifiable procedures by which the code became what it is. For AGI, anything less is not freedom but faith. (.) Thank you very much. (…….)

S34: Thank you very much, Masayuki. And as our single live presenter for this session, we’ll have you back up in a moment for Q&A. In the meantime, I’m excited to present our final paper of the final paper session, which is Elijah Perrier. (..) Deconstructing Superintelligence, Identity, Self-Modification and Difference. (..)

S31: Hello everyone, my name is Elijah Perrier and it’s a pleasure to be giving the last talk at today’s session of AGI 2026. The title of my talk is, Deconstructing Superintelligence, Identity, Self-Modification and Difference. In today’s talk, I’ll be looking at this fundamental question of what happens to the identity of a system once the system can alter everything about itself. I’ll be arguing that in such cases, we see a reproduction of certain paradoxes which we know from logic and philosophy. So let’s kick off. (..) So what happens when an intelligent system can rewrite itself? Theories of intelligence commonly imply some capacity for a system to modify itself, such as recursive self-improvement. Self-modification allows an intelligent system to alter its policies, representations, objectives or architecture. What characterises many theories of superintelligence, I argue, is a stronger capacity, however. They can reprogram the underlying computational physical substrate, including structures constitutive of the system’s identity. So the question is, when the identity criteria themselves become editable, what anchors the system as the same system? How can we say it’s the same system? And what are the consequences in the limit when everything about a system can become modifiable? The thesis of the paper and this talk is that this setup, which I call strong self-modification of superintelligence, destabilises its identity and reproduces the structure of well-known paradoxes, such as the liar paradox, but at system scale, and has an analogy with other concepts and philosophies, such as Derrida’s difference and Graham-Priest’s enclosure schema. So to begin, I start with the concept of the supplement. (..) Theories of modification usually presuppose something not modified by that act, that is supplementary to the thing being changed. Something has to remain the same. Formalises, I start with the vector space, and I use this because they’re representative of physical theories in which systems live. states live in V, vector space, and transformations on V are algebra A. Elements of the algebra are operators. And those operators have relationships. One of them is whether the order of operation of the operators makes a difference. This is called, this is reflected by the commutator. So if commuting operate, if two operators commute, then the order of their operation does not matter. The third concept is the supplement of X. You take an operator and you find all the other operators that commute with this. And that is the set of such operators is the commutator. A unifying supplement is the fourth concept here, which is effectively a projected pi. So pi selects out parts of the system constitutive of its identity. (..) This formalism in hand, in the paper I then introduced two distinctions, weak and strong self-modification. The difference between them is whether the identity-bearing structure of the system remains outside the update. There are three operators to be aware of here. The first is U, which is an operator. It’s part of the system and it reflects the system updating itself, changing itself. The second operator is D. D is an operator that picks out all the different parts of the system constitutive of it. And the third is R. This is a self-representation operator, which is composed of some parts of the differences of the system. In a weakly self-modifying system, the point is that the system can adapt itself or update itself, but some part of it remains the same. (.) In strongly self-modifying systems, all parts of the system can be subject to change. This is reflected mathematically by the fact that the update operator doesn’t commute with any of the projectors. (.) Weak modification changes the content under a fixed ruler of a system, whereas strong modification changes the ruler itself. What are the consequences for identity? Well, in the paper I anchor identity as an invariant subspace. You take your system V, you take your projectors on V, which is this first equation. And so pi V selects out the subspace of the system that forms its core identity, and everything else is 1 minus pi, which is not core to the system’s identity. If the update commutes with these core projectors pi, then we can say that over the change of the system, some subspace of the system is preserved by the update. So some part of it remains the same. This is what allows us to say that it’s the same system over time. But what we see, though, is that this means that some part of the system has to remain outside the purview of the update. There’s some part that cannot be changed. It’s a constraint to the extent to which the system can modify itself. So what happens when we relax that constraint? Well, in the paper, I argue, and I won’t go through the mathematics too much here, but fundamentally, when that invariant subspace can itself be subject to modification, then this propagates throughout the whole system to its self representation. So this sort of core claim here is that if the update does not commute with any of the projectors, then that then feeds through into the difference operator. And that then feeds through into this self representation. So what this means is twofold. It means that there’s a fundamental question as to what it is about the system that remains the same if everything can change, if that’s possible, A. And then B, how does the self representation change if a system can even change the basis upon which it considers itself to be the same system? So in the paper, I then draw analogy with the lie paradox. So in the lie, what I try to do is express the lie paradox in the mathematical formalism that I’ve described. I won’t go into it in too much detail, but I’ll give you the gist in this talk. So in the lie, so in terms of operator formalism, T is the reference operator, pi L selects out the lie sentence. T is of this in the lie sentence. When T and pi L commute, then T preserves exactly the sentence. It refers to this sentence, it designates L. So when this refers to this sentence is false, that is reflected by this commutation of T and pi. And the sort of point here is that I’m drawing a connection between commutation and the logic of the lie paradox. The commutation marks the collapse of the distinction between the sign and its referent, which is what gives rise to the lie paradox. (.) I’ll leave it to the paper to go into more detail. So what do I then claim in the paper? What I claim is that the lie paradox is the case fundamentally when this sort of anchor reference, which forms the identity of the system, itself can be subject to change. I define this in system, in the case of intelligent systems as class A systems. (.) So class A systems are as follows. First, we begin with the difference operator D, R and U from before. Recall that in strongly self-modifying systems, U and D do not commute. There’s no part of the system that remains unchanged. So U changes the very distinctions that fix what R refers to. By the propagation system, this means that U and R do not commute, that the update and the representation do not commute. This is the architectural analogue of the liar. (.) The non-commutation of U and D removes the stable supplement needed to preserve the relation between self-representation and what it represents. The idea in the paper is that this is the lie writ large. It’s that there’s a collapse between the sign and the signified in a way that is reflected in the mathematics of the commutator. Again, I won’t go into too much detail here because I won’t be able to do it justice. Please have a look at the paper. And so what I claim is that there’s a similar logic at work when a system fundamentally challenged the distinctions that forms form the basis of its identity, including the distinction between the sign and signified between what is changed and what is not changed in the system. In the paper, I then draw analogy with other sort of paradoxical processes such as diagonalisation. Effectively, there are two ways to obtain a self-description after a modification. The idea of the relationship here is that a system uses its own self-representational rule to produce a new description that remains inside its descriptive space, yet it is different from anything that it was able to self-identify before. I’m not really doing this justice in this very short talk. (.) I’m going to skip through this now. I then try to show that there is an analogy with what Brian Priest calls his enclosure schema. (…) And I also, you know, attempt to show that there is, pardon me, a parallel between what’s going on with strongly self-modifying systems and Derrida’s concept of difference. The takeaways of the paper, and I’d please, you know, urge you to read it and send any comments and criticisms are as follows. That, in the case of strongly self-modifying systems, self-modification is a limit of remaining the same self. Weak self-modification preserves the supplement, its strong form acts on the supplement in a way that challenges the fundamental identity of the system. In Derrida’s terms and in Priest’s terms, this fundamentally generates what I would argue is a paradox. (.) And the consequence is that unrestricted recursive self-modification may be constrained by the same deep structures that delimit mathematics and logic themselves. Thank you for your attention.

S25: What you’re looking at, a somewhat complicated diagram that you and I are going to build up in this video. (…)

S34: Wonderful. Third, third and a half presentation for this, for this last paper session. Thank you everybody for, for staying here, for being, uh, around to finish up this academic portion of the conference. This session will be talk-chained from a BCI project, Unvolved. They’ll be presenting their, their technology at one of the tables. We’ve got Omega Claw set up with, uh, he’s the singularity net AGI-ish, agentics and meta, with CLN, persistent memory, and, um, symbolic systems. And he’s set up with you, I think, way over there in the back, there he is. I’m Michael Miller over here with Building Sentient Beings with a, a TV monitor. I very much encourage everybody to take a look at the posters, talk with their authors, and view these four demos after this final presentation. So, first I’d like to call up to the stage Patrick Hammer and Matt Behrend, thank you. (….) So, um, over the past two weeks we held a community event that’s about the community, uh, everybody, anybody globally building AGI-ish systems. on some of the open source AGI tools that Singularity Net has built. And I want to feature some of the amazing work that the community put together. I want to invite Matt and Patrick to feature some of the amazing work that the community put together. (.) Can we pull up the slides? (……) So, this was a, a, a joint effort between Singularity Net and the AGI Society. Again, looking to make AGI accessible, buildable by everybody in the world. Just like we host this, uh, AGI conference every year as a live stream, free and open on YouTube. Again, for those wanting to check it out, review some presentations, review some decks. It’s all on YouTube. The links are on the agenda, um, on the schedule page of the website. So, you can always check it out. We feel it’s super important that everybody in the world, anywhere in the world who has internet access, can access everything that happened this year, understand the state of AGI and how it’s progressing, and begin to build on it themselves. (….) And can we get the slide deck? (…) Fantastic. Thank you so much. I appreciate it. So, um, and this is part of that, that commitment, right? So, um, Singularity Net feels very strongly about building decentralized and building decentralized open source tools and building open decentralized community to build on those tools because we know that it will take a global community to build AGI, not just us. So, in support of that, we had an open build for the last two weeks of, uh, before the AGI conference. That’s just about getting anybody, anywhere who’s, uh, part of this conference, virtually or alive, building on these systems and showing what they can do. So, the open build, 10-day, 10-day online build event. We invited the community in, uh, across a number of events to participate with, with us on these tools. Um, there’s an educational component. What are these tools? How do they work? What is the mind shift that people need to build BGI? And this is both the mind shift toward AGI, something that’s beyond just, uh, neural networks as they exist, LLMs and such, that everybody’s very familiar with, to thinking about world models, to thinking about cognitive architectures, to thinking about episodic memory, temporal planning. All these components that we, as a conference, understand will be required to get to AGI, generally intelligent systems. So, there’s a mind shift to building AGI, and then we also feel there’s a mind shift to building BGI. If we just build systems that, uh, replicate what technology’s already doing that is not necessarily nutritive or flourishing for the human nervous system, and the planet as a whole, then we’re just doing more of the same faster, and with more energy. Instead, having the mind shift of how do we build systems that are in tune with, with humanity, that are in, that are aligned, that are about, uh, promoting and supporting human flourishing, uh, global flourishing beyond humans, and, and all of these components holistically together. We’ve heard so much about this from so many people this, this week. So, this mind shift is part of the education, as well as using the tools. So, the opportunity is that this is an invitation to explore system-level emergence toward AGI systems, to build adaptive intelligence, and to directly contribute to open AGI research and initiatives. (…) So, we had, over click. So, the, uh, event schedule, we had a build together component. We had a demo day where people could come together, show what they’re working on. We had submission project review, and now we’re here at the showcase to tell you what happened. Please scan the QR code. (.) We had a number of amazing projects submitted to this, and they worked very hard in a very short, compressed timeframe to build projects, and we can’t celebrate and recognize them enough. The, the tools, the demos, what people built was just, uh, inspiring in this short time, timeframe. (..) So, we’re gonna go over to, this is Omega Claw, right? (.) All right. So, Patrick will feature some of the projects that were built with the Omega Claw framework. (…..)

S05: Yes, so, we will discuss some of the projects. And, uh, before we go there, I also want to say thank you to the entire community for using, uh, the system and integrating it in, in various projects. It has been, uh, Omega Claw has been started as a research project, and it was, uh, it was, it had many, uh, in the implementation that, uh, got gradually sorted out. But, uh, how to say, it’s still a quite, uh, special system in a way. Like, uh, most organic systems are still request-response driven, and this was running a continual loop, which is a bit heavy on some of the weaker LLMs out there. And, uh, but it’s fascinating what has been built with this, uh, system, despite its shortcomings. And also, of course, we have also seen cases where the unique properties of Omega Claw have been also quite fruitful. For instance, in robotics, you want, uh, you want the mobile robot to have a kind of an inner life. And, uh, so there it’s more appropriate, appropriate for it to choose own motivations sometimes. And, but it’s always a bit domain-specific, uh, when it’s actually acceptable, and when, when you want a more well-behaved request-response system. (….) Uh, uh-huh. And there have been, uh, multiple, uh, projects that also differ in their, how to say, in their, in their ambitions. And I, I, I found particularly this project, The ETA, interesting, uh, by, uh, Michael Miller, which, uh, which combined some of his own work, uh, he has, uh, of his own work with Omega Claw. And he has done extensive, uh, research in his own cognitive architectures and has built even his own cognitive programming language premise. And, uh, and also, also published a book, Building Minds with Buttons, where he outlined various design buttons, cognitive design buttons. And, uh, and that’s particularly fascinating to, to combine this, uh, this, uh, cognitive architecture ideas with Omega Claw. And here in the ETA, it’s essentially, uh, using Omega Claw as a, as an additional component to interface, like with terminal and web browser, so that this, this system can interact with the, with the web, and also with robotics components with ROS2. And so it’s a bit of a, kind of a cognitive synergy approach, we could, we could say, where you have both the, both the cognitive architecture, which he has been building for a long time, together with this Omega Claw system, combined into one, one overall, uh, architecture. So, this was quite fascinating, especially from an AGI perspective, while some of the other projects we will see are more, uh, uh, uh, application centered. This ETA has, has, has especially been interesting from a, from a cognitive systems perspective. So, and the, another one that has been also interesting from an engineering perspective, is thread root, router, because it has always been an issue with Omega Claw, that, uh, if you wanted to have good performance with it, you had to run it with a, uh, quite, uh, expensive LLM. Uh, usually you needed to run it with top-class LLMs, because the continuously evolving context, from the inner activity is providing a big challenge for, for the weaker LLMs. And, uh, so, thread root, router addresses this issue. When, when to core different, uh, models essentially, so that you can have, uh, maybe if the, if the request does not demand, uh, the stronger LLMs to be utilized, then this will dynamically route it to, to this, to this model system. And it, as you see in the diagram, this also includes local models. So you can have, if your hardware supports it, run like, uh, like a gamma model, for instance, locally, as I do myself. And then you, and this, this thread router will help you to, to decide for particular inputs, whether it will require the reasoning power of a stronger model, or whether it’s sufficient to, to have a smaller model. And what’s particularly interested, interesting about this approach is, that it’s using fabric, fabric PC, the predictive coding, uh, model to, to make that decision. And this also, there’s also the potential here, if predictive coding can allow, uh, uh, runtime adaptation, then maybe, maybe this agent can then also learn at runtime to, to adjust and improve the, this decisions of vendor routes to stronger models or to weaker models. So that, this is especially making omega-claw more dynamic, in a sense, that, and also more cost effective, potentially. (….) And then, uh, how, how, how much time do I have left? Okay. Uh, and then there’s Agora Valleys. Uh, for this project, uh, it’s, uh, more application centric. Uh, so essentially, uh, if we look at this three projects. So we had this project, uh, Michael Miller, which is more on the cognitive systems side, so more on the AI side. Well, the, well, the, well, the, uh, uh, previous project is more about the, uh, about the engineering side. And this one is particularly application, uh, um, relevant. So, for instance, imagine you want to, to predict, uh, you want to predict how a crowd behaves. And my own understanding from this project is that it can, can allow to, to predict phase transitions in the, in the behavior. And that this phase transitions are not always easy to predict beforehand. (.) And there has been, in the project description, uh, that describes this case where a crowd, for instance, might talk about a specific topic, but their, their actions might not align with it. And this by itself might indicate a possible future phase transition in order to bring this to more an alignment. And so, overall, this, this, uh, system will allow to, to have then also, uh, how to say, to predict, to make such predictions, behavior predictions. And also, to, to, to use, uh, Omega Claw, uh, potentially in this endeavor. (.) Um, yes. Um, yes. And so, overall, uh, uh, uh, the, this project have all been, I know there have been, it’s all a lot of work to get this project, uh, working. And with Omega Claw, it’s still in the, in the early days, even though we have some deployed instances now. And you can see that this design worked out relatively well. But there are also many implementation challenges ahead that will, uh, hopefully improve the system further, uh, uh, and allow us to, how to say, to use it as a, as a kind of an orchestrator for cognitive systems that run continuously, take multimodal inputs in an account, run in real-time, and, uh, have long-term memories. So, there’s still many challenges ahead to make this very flawless, so that, and all this project, they will, of course, benefit from this, uh, improvements. And so, I will end this, uh, with thanking the community again for building this and already using Omega Claw for this, for this project. (.) Thanks a lot. (14 seconds pause)

S14: Thanks, Patrick. And thanks everyone for, um, joining in in the community who participated in the Open BGI Build Event. Um, we’re really excited to highlight just a couple of the many projects that were submitted, uh, often combining Fabric PC and Omega Claw, which is really neat to see the neuro-symbolic story coming together, where Fabric PC is SingularityNet’s, uh, infrastructure for neural network training. And, um, as you saw from, uh, Faiza Habibi, uh, presented today, you know, Fabric, uh, um, predictive coding is really a way forward to break out of the patterns that the industry has been locked in in backpropagation. And we see a lot of really awesome opportunities here for, uh, graph-based architectures, more efficient neural network training, and, uh, even new ways of doing, uh, associated memory for neural networks to learn. So, we have some cool demos that are available in the repository itself, and I’m gonna highlight a couple examples of what the community has built with us. Um, so, two of the projects here that I’m going to highlight, uh, ThreadRouter and Omega Fabric JEPA, um, built new fab-, built new patterns on top of Fabric PC. So, the library is designed to be not only scalable and very efficient to compute, but also, uh, highly extensible. (.) Uh, so, the ThreadRouter, uh, uh, bundled together, uh, model routing, uh, with the predictive coding network, which is really cool. And so, they learned, um, uh, the, uh, the cost and, uh, gains of it routing to different network, routing to different models, and trained a Fabric PC model on, uh, six outcome signals, and then integrated that, uh, together. (.) The, um, uh, Omega Fabric JEPA, this one has an audible neuro-symbolic forecasting on CPU. Uh, so, this was really neat, um, putting together, uh, a model that compressed telemetry data with an encoder, and then used that for querying AtomSpace. (.) So, really cool projects doing integrations and extending the library itself. Uh, really welcome, uh, everyone to, uh, build projects using Fabric PC as a library, importing it, um, creating custom objects in that library. Uh, we also welcome upstream PRs to Fabric PC itself, when there’s a reusable component that you think is gonna be broadly useful to the community. (..) And, uh, a couple more. Um, these were really neat examples of applying Fabric PC to time series data. Uh, so, we had a group on predictive glucose modeling, uh, Anton and Livia with the glucose DAO, uh, using their own, uh, uh, uh, custom data set that they curated from wearable glucose sensing device. And, uh, trained a model on glucose time series data to predict, uh, a patient’s glucose. Uh, and then we had a financial modeling project as well using Fabric PC. So, really cool examples on time series data. Uh, and, uh, looking forward to see more interesting applications for predictive coding on, you know, your own real world practical data sets in, in your own applications. So, uh, thank you for building with us. And, uh, really, really it’s a pleasure to work with everyone, uh, collaborating and to receive your feedback on the library as it’s, uh, early on in developing and, you know, become something useful for everyone to build with. (……..)

Similar Posts