Agentic AI Summit 2026, Plenary Stage – August 1st – Morning Session
S01: I’m Jennifer Chase, Dean of the College of Computing, Data Science, and Society, and Professor of EECS, Statistics, Mathematics, and Information at UC Berkeley, and it is my absolute pleasure to welcome you all, both in person and online, to the 2026 Agentic AI Summit. We have an incredible gathering this weekend of researchers, founders, enterprise leaders, policy makers, and entrepreneurs from around the world. (.) We are at a pivotal moment in technology, moving from models that simply answer questions, to agentic systems that can reason, plan, and solve complex real-world problems. That is what makes this summit so timely. A special thank you to Professor Dawn Sung and the entire Berkeley RDI team for putting together such an exceptional program for us. (.) To officially open the summit, it is my great privilege to introduce a leader who truly embodies Berkeley’s spirit of innovation and public mission. As our former Chief Innovation and Entrepreneurship Officer, former Dean of the Haas School of Business, and now Chancellor, please join me in giving a warm welcome to Chancellor Rich Lyons.
S00: Thank you, Jennifer. Thank you, Jennifer. Good morning, and welcome to all of you. It is a joy to see you all in this room. I’m honored to welcome you, this august group, to this beautiful UC Berkeley campus. I may be a little biased, but I can think of no better place than Berkeley for this timely, this urgently needed gathering among us. It starts with the fact, of course, that we are a public university, in the fullest sense of that word, built by the people, for the people, in support of the greater good. And this kind of forum is exactly what Berkeley was built for, a convening around a highly complex and high stakes set of issues that brings people together from all walks. From the public, civic, nonprofit, and private sectors, a variety of academic institutions, and a multitude, of course, of academic disciplines. Gatherings like these, I believe, are the best way to take on complex, salient, and relevant challenges and opportunities, particularly those that are inextricably connected to the public interest. (..) Here at Berkeley, where the free speech movement was born, we believe in and thrive on the constructive collision of ideas and perspectives. I can’t wait to see the sparks fly in and among the extraordinary global group of visionary leaders that we have here today, pioneering entrepreneurs, experts from leading AI organizations, venture capitalists, policymakers, and many other people. At Berkeley, tech breakthroughs are never viewed in a vacuum. They’re always intertwined with our public mission and our responsibility to act in support of the greater good, even as our AI researchers work to expand the frontiers of knowledge about everything from autonomous reasoning loops to multi-agent orchestrations and swarms. Many of their colleagues draw from Berkeley’s 130-plus academic departments and 80 interdisciplinary research units. They are exploring and contemplating the impact and future course of AI through very different lenses. Our humanists, for example, as you might have guessed, wrestling with what it means to be human. That’s as old a question as human time, but it is especially relevant in these times. At our school of public policy, they’re considering and proposing ideas, actions, and policies that can provide for scientific progress, which we must make room for while safeguarding the well-being of the society that we all, at the end of the day, serve. Our natural resources people are wrestling with the complexity of the related environmental challenges, and the list goes on and on. Together we must, and I’m sure that we can, honor the ultimate objective of this summit, by taking steps for a safe navigation towards a technology future that doesn’t just automate actions, but actively empowers individuals, enriches societies, and amplifies human potential. There may be no higher calling at this moment of time in world history. I’m not sure I need to tell you that, right? (..) Everything all at once is the feeling. Somebody said to me the other day, what if at the beginning of the 20th century, those early years of the 1900s, every technological advance that occurred in that century occurred in the first decade of that century? (.) What would that have looked like and felt like? And that is an interesting thought experiment. I am in awe of the size and quality of the global community that is here today, in person and also online. Much credit is owed to the extraordinary team at our Center for Responsible, Decentralized Intelligence. Founded in 2022, Berkeley RDI, as we call it, has become a leading academic nexus for critical education, conversation, and exploration about how we can together and how should we set the stage for robust, safe, and transformative AI-powered future. Your presence today provides powerful validation for the work they do, the values that Berkeley RDI embraces, and the momentum they have built through Berkeley RDI’s popular agentic AI MOOC series. Definitely check that out if you haven’t already. For me, this team’s combination of intelligence, competency, foresight, holistic thinking, courage, and yes, deep questioning of the status quo, a very Berkeley mindset, embodies and exemplifies Berkeley at its very best. That serves as the perfect segue for my introduction of the host and the driving force behind today’s summit. Professor Dawn Song, many of you will know her work. Dawn is a pioneer in AI and security and a distinguished member of Berkeley’s extraordinary faculty. She and her team have assembled a remarkable program for us today and tomorrow, as you’ve seen. And I’ll let her walk you through what lies ahead. Please join me in giving a warm welcome to Professor Dawn Song. (10 seconds pause)
S12: Thank you so much for being here. I’m really excited to be here and to welcome everyone. So if we can get the slides up. (12 seconds pause) Okay, let’s just wait for one second to get the slides up. (25 seconds pause) I guess everyone can get a preview of what’s going to happen around the program today. (28 seconds pause) Thanks for the great spirit. Okay, great. So welcome everyone here. So we are all really excited about the fast advancements in Frontier AI. And also, even just this week, we are seeing that the gap is narrowing between the open weights and proprietary models. (..) And all this great advancements in Frontier AI capabilities also fuels the growth of Agentic AI. (.) So here’s a little bit of history, how we got started and how we got here today. (..) So this is the Google Trend on Agentic AI. Back in fall 2024, even though actually back then, not many people were talking about Agentic AI or paying attention to Agentic AI. But we could see that Agentic AI is the future and its next Frontier. And hence, at Berkeley IDI, we led the world’s first course and the first MOOC massive open online course on Agentic AI in fall 2024. And fast forward, even though we knew that Agentic AI is the next Frontier, is the future. Still, we didn’t, we didn’t predict that it would arrive so soon. So suddenly, shortly after our first Agentic AI MOOC, at the beginning of year 2025, suddenly year 2025 became the year of agents. And so we can, and so we continue to see the growth of Agentic AI with the release, for example, for codex, cloud codes, and so on. And also, to bring the community together to make further progress on Agentic AI, we convened the first Agentic AI Summit in August 2025. And fast forward, end of last year, beginning of this year, we saw further explosive growth of Agentic AI exemplified by the OpenClaw launch, where in March, OpenClaw became the highest start GitHub repo in history. (.) And that also brings us here today with the second edition of Agentic AI Summit. (.) Really welcome everyone. (..) And this explosive growth of Frontier AI capabilities and Agentic AI is further fueled by unprecedented capex spending and exponential growth of AI compute capacity. So on this figure, it shows that global AI compute capacity is doubling every seven months. And also, we see the hyperscalers growing their capex spending exponentially. (.) And also, when we put together all the big tech AI capex and compare it with the largest projects in US history, it dwarves the previous largest projects run in the US, including the Manhattan Project, the Marshall Project, and the Apollo Program. (…) With all this, we predict that we are going to continue to see exclusive growth of Agentic AI. (.) While we want to benefit from the great capabilities and potentials that Agentic AI can bring to us, at the same time, we also need to pay attention to a broad spectrum of different AI risks. (..) As an example, in the landscape of cybersecurity, our work, for example, CyberGym, CyberGym, which is a large-scale benchmark containing close to 200 large-scale widely distributed open-source software, (.) CyberGym evaluates the capability of Frontier AI in finding vulnerabilities and generating POC’s proof-of-concepts. These are inputs that can trigger the vulnerabilities. So, as an example, CyberGym tracks, measured and tracks, the fast growth of Frontier AI capabilities in cybersecurity. (…) With the GPT 5.5 Cyber top leaderboards recently, as Sam tweeted. And also, Cloud Mythos and Project Glasswing also illustrated the strong capabilities of Frontier AI models today that can discover thousands of high-severity zero days and many critical vulnerabilities in every major OS and web browsers. (..) And our recent work on ExploitGym, in collaboration with Frontier AI Labs, also demonstrated that Frontier AI models today, they can not only discover vulnerabilities, but also they can actually automatically turn these discovered vulnerabilities into exploitations that can cause critical attacks, for example, enabling unauthorized code execution, (..) and demonstrating that autonomous exploitation is no longer hypothetical, and also demonstrated that Frontier AI agents today can even generate exploits that can bypass today’s standard security mechanisms and mitigations. (…) And also, many of you may have heard about the recent incidents, and the OpenAI HackingFace incidents. The OpenAI agents actually were trying to solve tasks in ExploitGym. and by actually trying to exploit, for example, vulnerabilities in HackingFace to try to gain additional information to try to succeed on the ExploitGym benchmark evaluation. (..) And Anthropik also recently reported that when their agents were doing cyber evaluation, also the agents were able to break out from their isolated environments and cause additional issues. (..) So all these OpenAI HackingFace Antropik evaluation incidents demonstrated that today, even the evaluation infrastructure is part of the attack surface. And the agents today, they do have really strong capabilities that can cause huge risks and potential harm. (..) So the key question is, how do we ensure that a genetic AI becomes one of humanity’s greatest advances, while remaining safe, secure, and beneficial for everyone? (…) And now we are at a critical point. (…..) Okay. (…) So, yes, now we are at a critical point, (…) with the slides is not the latest version. Sorry. (12 seconds pause) Yes, so the Frontier AI capabilities is increasing really fast, and also even the pace itself is increasing drastically. And also with the potential possibility of recursive self-improvement and so on, the speed of improvements can outpace essentially the speed that we can actually develop. (.) Mitigation mechanisms to ensure building a safe and secure agentic AI. (.) And hence, we are at a critical point. (.) And recently, you may have seen open letters from Jensen, talking about the importance of open ecosystem, open with models, and the open letter from Mark Zuckerberg, talking about building agentic AI and building AI that benefits everyone. and also the open letter from Demons, proposing to build an overall framework for Frontier AI and for essentially better governance of Frontier AI. And also the recent open letter assigned by over 1,000 leading AI researchers from different Frontier Labs, pacing the horizon, pacing the Frontier. (……) Refresh the slide. So this brings to the Berkeley RDI. (..) At Berkeley RDI, our mission is actually to steward the future of the future of agentic AI for human flourishing. (15 seconds pause) One second. (.) Yeah, can we refresh the slide? (…….) This is not the latest slide. (………) Can they reload it? (1 minute[s] pause) Not the latest mission. (24 seconds pause) Yeah, it’s fine. Just show the slides. I mean, this is really not the latest version. It doesn’t even finish. (17 seconds pause) Okay. So, sorry, there is some slight update issues. Yes, so this brings to Berkeley RDI. So Berkeley RDI’s mission is to steward the future of agentic AI for future flourishing. And with that, Berkeley RDI, so first of all, it’s a campus-wide center. that covers, that actually covers different schools on campus, including computer science, ECS, School of Engineering, and House Business School, and Law School, and so on. (….) RDI as an innovation center has three key pillars, research, education, and entrepreneurship, and community. So for research, I hope that I’ll be able to show you more about the Berkeley RDI research, and education. (…) Chancellor Rich Lai mentioned earlier, and also, so Berkeley RDI has been hosting actually a number of MOOCs, massive online course, including the MOOC on Agentic AI. And also, the Berkeley RDI YouTube channel has over 1 million views. (18 seconds pause) And also, RDI has been leading global competition and hackathons with over 6,000 global participants from actually over 130 countries. And the hackathon, and the hackathon, in total, had over 2 million dollars in prizes and resources. And so here, we welcome everyone to the Agentic AI Summit 2026. So this year, we had around 1,000 submissions for speaking proposals and papers. And we, in the participants, we have over 1,500 industry organizations, and over 250 universities represented. And we have close to 5,000 in-person attendees. Thanks everyone here. And we hope to have, you know, many hundreds of thousands joining online. And this is a quick agenda at a glance. So here, we have four stages. So this is the plenary stage. And we also have three other stages, spanning over two days. And we also welcome everyone during lunch. There will be also poster sessions. We have close to 200 posters. So we welcome everyone to join the poster session. And they will be in the MLK building. And also, we have the dinner reception, also with poster presentation. presentations as well. And also, we have a number of site events organized by independent third parties. So you can go to the Summit website for more information as well. (.) And again, we have four different stages. And you can check out the map here with the four different buildings. And also, I want to mention the alumni house, which is right outside, very close to the Zalabag Playhouse. And that also serves as a lounge. For participants, the plenary stage will be live streamed there. And you are also welcome to hang out there as a lounge. And we want to have a big thank you for all our sponsors. And also, thank you for our community partners from Berkeley campus. And with that, we hope you will enjoy the Argentina AI Summit. Please do tweet and post on social media to help spread the word. And we would appreciate your help retweet and repost our tweet as well. And for future information about future events, please do sign up on our newsletter. (.) And again, you can check out the full agenda here. (.) With that, thanks, everyone. (……..)
S04: Good morning, Berkeley. Thank you, Don. My name is Todd Graham. I’m an investor at M12, which is Microsoft’s early stage venture fund. I’m going to be doing a little bit of MCing this morning to hopefully keep us on track, or get us back on track. I’m going to introduce Peter DeSantis from Amazon right now, before he comes out to give his keynote speech, which I hope you’re all excited for. He’s the SVP at Amazon leading foundational AI models, custom silicon and quantum computing. So, you know, the easy stuff. He launched Amazon EC2, built the team around Graviton and Tranium chips, and today oversees Amazon’s obviously most ambitious bets on the future of compute. If you want to understand design from atoms to models, you guys are in the absolute right seats. So with no further ado, invite Peter on up and enjoy. (……)
S11: Morning. (..) Well, chips and quantum computing and models might be easy stuff, but PowerPoint is not. (.) And Don’s experience has me a little nervous because I’ve been there. One time I built a PowerPoint presentation when I had to present to 4,000 salespeople, and I did the whole thing in template mode, which is apparently a feature, and that didn’t allow it to be used at all. So I had to do the whole thing with the app up, and so I feel for Don, and I hope it doesn’t happen to me. (..) With that, I’m Peter. I’m delighted to be here on day one of the Agentic AI conference here at Berkeley. Thank you all for getting up early to be with us this morning and looking forward to an exciting weekend. Day one’s got a special meaning to me, being at Amazon for 28 years. (.) We talk a lot about day one, and this made a lot of sense when I joined 28 years ago. We were a small little retail startup, and we all fit in a conference room. But, you know, as we kept talking about it being day one, it started feeling a little silly as we became a huge retailer. But every day we talked about it being day one, and it meant something to us. And then we started AWS 20 years ago, and it did indeed feel like day one again because everybody inside the company was kind of laughing at this AWS cloud thing inside of Amazon. But we were very convicted, and that turned out pretty well. And we kept growing AWS, and we kept talking about it being day one. It’s always day one. And why do we talk about it always being day one? Because day one culture is something that you believe and will into existence. It’s a culture, it’s a behavior, it’s an attitude. And it’s an important attitude. It’s an understanding that despite whatever innovations and success you’ve had, there’s more success ahead. And I think it’s an interesting way to think about AI. (.) Because I think there’s a potential to think that we’re, like, on the cusp of being done with AI. But that’s not the way I see it. I see this as the very early days of innovation. And what’s happening today, while it seems frenetic and exciting and things are literally changing and becoming more capable every day, we’re just at the very, very beginning. (.) In fact, I’m an AI optimist. Like, I very much believe, I’m on the far side of the distribution thinking that AI is really going to help us solve some of the most fundamental problems that humanity faces. Things like energy and health and medicine and all of it has so much potential to transform the way we live. And I think we’re just, just seeing the very tip of that iceberg. (.) But in order to get to that future, we fundamentally have to solve a number of problems. In fact, we have to deliver order of magnitude improvement in what’s there today. And that’s why I say we’re very much in day one of this journey. (.) And so, really, the most fundamentally important problem that I see is that in order to realize what we all want AI to be, we need to figure out how to make AI much, much more efficient. And that is a fundamentally interesting and exciting problem to me, and I suspect to many people in this room. And it’s a problem that stretches so many of the pieces of the stack that we’re innovating on. And when I was asked to speak this morning, I kind of had this, like, you know, what should I talk about? We want to talk about agentic AI, but what is agentic AI? Well, agentic AI is a ton of things, and I am blessed to work at Amazon, where we’re working on many of the pieces. In fact, all of the pieces of the AI stack. Everything from the applications on the top to the platforms and software and agentic frameworks that allow us to build and secure agents and models. To the models themselves, to the infrastructure and chips that power the AI models and agents that run. (..) But when I thought about what I wanted to talk about this morning, and if I look across the next couple days, there’s people talking about each and every part of the stack. And each and every part of this puzzle and all the innovation happening. But, you know, what do I want to talk about? Well, I thought maybe I would start by talking about what I’m most passionate about, which is the infrastructure. And the reason I’m most passionate about the infrastructure is because, again, I fundamentally believe that in order for us to achieve this vision that I think all of us share for what AI can be, We’re going to have to make the infrastructure fundamentally more efficient and not just incrementally. And so Amazon is making some of the deepest investments in innovating at every level of the AI infrastructure stack. And I thought I might start this morning off sharing with you a few of the ways that we’re thinking about this problem space. (..) So the first observation that I’ll mention is that there’s not going to be a single chip or server type that’s going to power the AI workloads of the next decade. This is an interesting trend that’s, I think, really been emerging. If I look back a decade, everything ran on, you know, sort of a CPU. Most every workload ran on, you know, a CPU, maybe one faster, one more core, one slower. But it was a pretty standard selection of server types out there. And I think there were a number of reasons for that. One is Moore’s law was allowing us to produce better and better general purpose processors. And so having one tool chain made a lot of sense, gave us a lot of leverage. Another reason for it was, you know, while it was super complex to build the servers and the infrastructure, And if you needed to support lots of different kinds of servers and lots of different kinds of hardware, that just didn’t make a lot of sense when you were trying to focus on your business. But the cloud started to change all that. And we got excited about that about 12 years ago. And, you know, the cloud takes away a lot of the problems that come with custom hardware. And so way before AI started growing as quickly as we see today, AWS was innovating to provide more bespoke and custom hardware solutions for all sorts of applications. But then AI came along, and AI is really, there’s an ecosystem of workload types inside of what we call AI today. And agentic AI is pushing this even faster. And so my excitement is that this now opens up the scale and diversity of workloads that really is going to push innovation all the way down. Of course, we all are here today because of the GPU, which is a phenomenal chip that really changed the game when it came to doing the sorts of floating point and matrix operations that underlie most of the AI models that we’re building. And, you know, it was really a somewhat happy accident that this happened, but it was really a critical part of the origin story of the AI models that we’re all so excited about today. But the GPU is a fairly general purpose architecture. It was built, in fact, for what it’s named after graphics, and it may not be the most optimal. In fact, it’s not the most optimal hardware architecture to power a lot of the models that we’re running today. And, you know, we noticed this a better part of a decade ago when we started seeing some of the emerging patterns in AI models. And the interesting thing about the AI models is, yeah, they do a lot of matrix operations and floating point math, but they do them in a fairly predictable way. And so that means that rather than having to have a bunch of cores with a bunch of registers accessing random memory, you can predict the flow of memory and compute through a chip. And the type of chip that is most optimal for that kind of workload is actually something called a systolic array. And so we started building our Trinium processor, which we originally called Inferentia, but grew into a larger variant, which we now call Trinium, a better part of a decade ago, based on a systolic array architecture. And that architecture has allowed us to achieve efficiency by removing a lot of the flexibility that we don’t need for these workloads, while still preserving a lot of important flexibility to allow it to run a broad range of different types of ML models. It’s not an ASIC customized for one application. It’s an AI accelerator that’s been built specifically to optimally run the biggest and broadest range of AI models. (..) But, you know, even though we think Trinium is the optimal architecture for AI model building and inference today, it’s very unlikely that one chip is going to be the solution as we look at how AI workloads are evolving. And that’s exciting. You know, we think there’s another decade or two of novel innovations that we can put into Trinium, pushing the limits of what’s possible there. But we also think we’re going to invest in other chips, and others are going to invest in other chips. And I think that’s going to make the hardware ecosystem that we’re all building to much more interesting and exciting. (..) You know, one quick example I’ll give there that most folks in the room might be familiar with is the inference pipeline. When you look at the transformer models that we’re running today, the autoregressive transformer models, that are the vast majority of the models that are underpinning the agents that we all are talking about, (…) there’s two workloads underneath that inference. There’s the pre-fill or encoding workflow, which is extremely compute intensive. And then there’s the autoregressive token generation workload, which actually happens to be extremely memory bandwidth, because each subsequent token requires us to access every model weight. And so the profile of those two workloads is radically different if you look at it at a hardware level. And Trinium is actually designed very well to run both of these workloads today. That efficient memory flow that I talked about benefits the decode workload, and the large floating point performance is great for the pre-fill part of the workload. But there are other patterns emerging that might allow for optimization of one or both of those parts of the pipeline. And the one that’s getting the most attention right now is the token generation portion with so-called SRAM chips, or chips where we put a lot more memory on the compute chip at the expense of compute capacity itself. It’s all a trade-off. And the result is that you can run the token generation part of the workload significantly faster. But that comes with trade-offs and constraints. If you want to run something that needs a lot more compute, well, you traded those transistors for memory. And if you need a lot more memory, you probably can’t fit it on the chip in the first place. And if you’re running a Gentic AI, you probably need really, really long context windows. And it’s hard to move that context in and out of an SRAM chip much more expensive than it would be with something like Trinium. And so I happen to believe that SRAM and memory-intensive chips are going to play an important part of our inference systems, but in the context of many other chips, not as a replacement for the chips that we’re using today. (……) All right, I’ll move to the second observation I want to talk about, which is that models and chips must be designed in tandem. And this may seem fairly obvious, but that’s not exactly how we’ve been doing it up till now, at least not broadly. I mean, and the fundamental observation I’ll give you here is that when you’re building an AI model, you’re making a multi-year investment in building to a certain model architecture. And that could be an extremely expensive investment, hundreds of billions of dollars. And I suspect in the not-too-distant future, some of the model families will be investing billions of dollars in order to progress. And so when you’re making that sort of long-term investment, you want to be efficient at every step of the process. And a chip investment is very similar. The lead time of building a new chip and bringing it to scale in a data center could easily be two to three years with the complexity of leading AI processors. And so now if you have two long arc investment pieces like that, that fundamentally need to be co-developed, it would be a mistake to sort of, you know, develop models in isolation of understanding what’s going to happen in hardware a year from now. Or to be building hardware for two years from now without understanding the needs of models. And so bringing together the design of the model and the system around the model and the hardware chip is fundamental. (.) You know, one quick way to think about this is just think about the capabilities that you put on the chip. Again, I said it could be two to three years from when you’re designing a chip to when you’re putting it into production. And one of the places that you look, of course, for new ideas for what you might want to put in your chip, in your hardware, is in academia. And here I have, I believe, I can’t see it on my screen, but I think there’s a bunch of very important Berkeley papers in here. And there’s a lot of good ideas in these papers, but not every one of them is going to be something that makes its way to production. Not because it was a good idea, but because when you make design choices, you can only make so many design choices. And so the art of building great hardware is understanding which of the great ideas ultimately need to be put into hardware. So that by the time the hardware is in market, the science and the technology is ready to take advantage of that chip. And that’s exactly what we’ve been doing with Terranium. You can see on the board some of the features that we’ve been adding to the latest generation. And these were very much inspired by working in academia, but also research in our customers and our internal teams. And so we’re really excited. About six months ago, we brought our internal foundational model closer together with our chip development. That’s a big part of my purview now. And it’s going to help us continue to bring some of the most capable hardware innovations and models to our customers. The third and final observation I’ll make for the morning is that, and I’ve already kind of touched on this, AI infrastructure is not a model problem. It’s not a chip problem. It’s a systems problem. And this makes it, like, one of the most interesting problems, I think, that you get to work on as an engineer or as a scientist. Because, literally, these systems are so interconnected that if you want to get performance, and we can define performance as, you know, better accuracy. We can define performance as lower latency. We can define performance as better throughput and cost. But any of those betters requires a complete understanding of every level of the stack. And, you know, when done well, it brings in almost every part of computer science engineering, which just makes it an incredibly fun place to play as an engineer and as a builder. (…) So, ultimately, I would say, you know, we’ve had a real focus because we believe that in order to get the best performance out of our system, we’ve had a real focus with our chips programs at Amazon making sure that we give customers the deepest access to our hardware. And that’s because when you’re trying to optimize performance, you don’t want abstractions between you and other things. As a computer scientist and spending many years in software engineering, I have an instinctual desire to put a layer of indirection between things to scale. And that is usually the right way to build software, particularly if you’ve got a large organization like Amazon. But it’s not the right way to build an AI system if you want to achieve absolute best performance. There you need to lay hands on the whole stack. You need to understand the network. You need to understand the chip. You need to get to the hardware instruction level capabilities of your chip and your network. And so that’s exactly what we’ve done with Terranium. We’ve exposed our complete instruction set. There’s many students here at Berkeley that have helped us push the limits of what’s possible there. We’ve also open sourced our tool chain so that folks can understand how our compilers working and how other parts of our software stack work. And, of course, we’ve integrated with things like PyTorch and BLM, which are key parts of the building ecosystem. (.) And so, you know, I’ll end by just saying, like, the fact that we have this amazing opportunity, but the world is full of constraints, capacity constraints, performance constraints, power constraints, that’s what makes this whole thing so interesting. And that’s why I’m excited to be involved in technology and AI at this very moment. I hope it’s why this audience is also excited. The future is bright and we need to build it together. And I’ll echo something that was said already earlier this morning, which is events like this that bring together experts and visionaries from all parts of the stack are critical to unlocking the potentials that come from understanding the whole system. And so there’s never been a better time to be at an event like this. And I hope that this weekend is fun for everybody. So I’ll leave you with my thought that I started with, which is this is the very beginning of this journey. And what lies ahead is going to be far more exciting than what we’ve seen to this date. And so, you know, be motivated, be excited and build. Thank you. (……..)
S04: Thank you, Peter. (.) If you’re playing at home, you have a guy who works at a hyperscaler who’s now introduced one person from another hyperscaler. I’m about to do my second hyperscaler introduction. We’re going to talk to next Sarab, who is the Google VP at DeepMind. Prior to that, he was GM of Google AI for cloud, where he led the Vertex AI and platform bringing Gemini to enterprises worldwide. Before Google, he spent over a decade at Microsoft as a CVP, technical fellow where he founded Project Turing, which was our first attempt to literally stand up our serious batch of GPUs and built some of the first language, large language models in the world. So with no further ado, we’re going to hear what you say for discovery and models and agents. Sarab, thank you. (..)
S13: Thank you. So it is very good to be here. I used to study here at Berkeley, so good to be back. (.) I’m going to talk about the evolution of AI from models to agents to discovery and the related infrastructure implications that it has. (..) So obviously, almost everyone is familiar with the huge expanse or adoption of AI from singleton chatbots about three, four years back to now almost semi-autonomous or autonomous agents. The numbers at the bottom show you some reflection of the scale of adoption and the impact that it has had. So Kaggle has this five days of agents course, and last fall when they ran it, it’s an online course, 1.5 million users registered for it. (.) If you look at the cost of agentic tasks, it is 10 to 100x more expensive in terms of inference compute, which is needed, compared to non-agentic workload. So adoption is increasing, as well as the complexity of compute per interaction is also increasing. And what it is leading to is, from a Google perspective, we process 3.2 quadrillion tokens every month. Now, quadrillion is a very large number with lots of zeros in it. The way to think about it is, one quadrillion is one novel per person on this earth. So, in a way, Google generated tokens which are equivalent to three novels for every person in this world. So the scale is just massive. And so how do we build for this new opportunity which is coming up? And this needs different pieces in the stack. So I will, like, if you start from the bottom, we have the AI hypercomputer, which includes the GPUs and TPUs. On top of that, we have world-class research and frontier models. Then we need data for these models or agents to do meaningful things. As these agents are starting to do meaningful things, there’s a huge opportunity or risk, as well, for security and defense. So we need agentic security and agentic defense. On top of that, we need a platform so that people can reuse all these capabilities in a cheap, efficient, and effective manner. And that’s where the platform piece comes in. And then, finally, we have final end-to-end agentic task force which can actually do things for you. Now, a co-optimization across this entire stack is really needed to extract maximum value from AI. And in the interest of time, I won’t go through all of them, but I will kind of quickly give you a snapshot on some of these layers as we go along. So Google has had a huge investment into TPU, 10-plus years of investment. Some of the innovations relating to liquid cooling and shared memory have been there. The latest generation of TPUs is TPU V8. And it comes in, this is the first time we are splitting the mainline, as Peter was saying in the previous talk, that there is one version which is 8T, which is for training, and 8I, which is for inference. (..) In terms of just to give you a comparative number of, like, how the improvements on TPUs are happening, the table over here is comparing the current generation TPUs, the V8s, with the previous generation, which was just released a year back, last year. And these are the main workhorse today, which is the Ironwood TPU. And you can see, for example, on the table on the left-hand side, you can see that the number of FP4 exaflops available in a pod is 121 exaflops. That’s a massive, massive amount of compute per second that we are offering. And the jump is about 3x from just the previous generation. Similarly, on the memory bandwidth side, you can see significant jumps like 2x and 4x between one generation to the next. On the training, I’m sorry, on the inference, you can see on the right-hand side, because the inference needs are increasing very, very rapidly, a single pod can offer 11.6 exaflops, and it’s a 10x jump from the previous generation. And similarly, significant jumps on the memory side. (..) So what do we do with all these TPUs? Google obviously invests, and there are now a lot of other companies who also leverage TPUs for their model. Google is training the Gemini family of models, which comes with Pro, Flash and Flashlight. The image and video generation models, like Vio and Imagine, they are trained as well. The world model, which is a new and completely exciting space for us, with the Genie model, is also trained over there. The alpha fold, which is a protein folding, it predicts protein folding structures, as well as AlphaGo and AlphaChip, which is used to train and build the next generation of TPUs as well. These are all being, so it’s effectively TPUs designing for TPUs in some sense. So all of these are heavily leveraging our TPU infrastructure, both for the training as well as on the inference side. Now, once we have the hardware and the models, then we need a platform through which people can build agents relatively easily. And building agents is not just like writing a prompt, et cetera. There are lots of complexities in it. And so there are four key problems. One of them is building it. So we have capabilities like access to all different types of models, agent development kit, and AI studio to build agents. Then you need to scale them, govern them, and optimize them. So start with building, scaling, governing, and optimizing. On the scale side, the capabilities are a managed runtime so that you can scale from a single instance of that agent to millions of agents. (.) Governance becomes a really key or important factor because as these agents are starting to do meaningful things, You want things like agent identity, agent registry, and agent gateway to make sure that the agents are working within the right confines. The identity systems that they are using are separate from the user because they can do things beyond what the user is able to do. And finally, once you deploy an agent into production, you want to optimize them. So things like tracing, simulation, evaluation, observability, all these capabilities need to be there so that you can keep on improving the agents once you land into production. So these become key pieces. Now, one additional thing that we do is across this stack, Google itself builds on top of it. So there are agentic solutions that Google is building across on top of this stack. So for example, there are areas in biology. For example, there is alpha fold or alpha genome, which is there. On mathematics side, we have alpha evolve and alpha proof. On the physics and chemistry side, we have genome fusion, et cetera. And on climate and sustainability side, we have Alpha Earth as well as Weather Next, which is a world-class leading weather prediction model. And these are all built on the stack that we are talking about. (..) Here are some examples of the impact that these final agentic solutions are having. So for example, alpha fold is being used in a wide variety of applications. So in plastic pollution, so identifying or designing plastic resistant or eating enzymes for antibiotic resistant structured biology. One key area is neglected diseases. So there are a lot of diseases out in the world where the pharma companies don’t invest enough because there isn’t enough economic incentive for that. And what alpha fold is doing is lowering down the cost and the ease of exploring these drug designs, which has material impact. And then malaria vaccine as well as drug delivery. (..) Alpha evolve, this is a general purpose optimizer, which is there. It is used heavily inside Google for data center optimization across along with a whole host of other applications. But also it is used by other customers. So for example, it is used for route optimization, it is used for quantum error correction, as well as by e-commerce companies for improving the forecasting demand model. And has significant, if you think about it, like these are very different types of application and has huge improvements across each one of them. (.) So what does this all mean? As we go through this stack and as the models keep improving, what is this all leading to? And this is what we believe, that where we are heading towards is this autonomous discovery loop. So, and the example on this particular slide is about biology, but this is much more generic. (.) What we are having is we are having all these building blocks coming together. So for example, for data ingestion, we now have agents which can digest large amounts of literature. There is alpha fold has a very large protein database, which is available, as well as experimental logs. Then you can do hypothesis generation on top of it. So there is AI co-scientists which can debate different hypotheses, generate novel solutions. And you can use them from a modeling perspective using alpha genome and alpha fold to simulate these hypotheses and see the effects of that in seconds. And then finally, you can execute on it through alpha and Gemini robotics to run wet lab tests and see what the data is and also feed it inside this particular loop. And so what can happen is a research cycle, which used to take years from beginning, like input of data to outcome of the product, can now be done in hours and days. And this is just one segment, like biology is one example, but this is going to impact almost all facets of like human discovery, which is there. And this is super, super exciting phase that we have. And with this, I will end my talk and thank you for your attention. (………)
S04: So my takeaway so far is that we have massive constraints everywhere. So if you’re thinking of building or you are building, hurry up, because these things are not going to solve themselves. Speaking of infrastructure, I get to talk now about John Cohen, who’s a VP of Applied Research at NVIDIA. He leads, among other many things, the team behind the Nematron family of open reasoning models. His work spans the full agent stack, inference microservices, guardrails, RL infrastructure. And just the other week, a Nematron-based system earned the gold medal equivalent score in the International Mathematical Olympiad. So more good news coming out of AI. He spent his career in the intersection of GPUs and the software that makes them useful. I haven’t checked, but I suspect he does not have any blackwells with him, so probably don’t ask. But fun fact, you are about to get a presentation from an Academy Award winner. So with that, welcome, John. (…….)
S10: Morning. (…..) So I want to talk a little bit about what is an agent and where did this all come from? (.) So when the modern AI era was kicked off with ChatGPT, largely a few years ago, we thought of these AI systems as things that a human would talk to a large language model. Maybe that large language model had access to some database. But fundamentally, this was about chat, people talking to AIs. But this has evolved considerably. Today, we really think of a complete autonomous system. This is made up of many models, some open-weight models, some proprietary models, access to tools, infrastructure, memory systems, context management, ability to spawn sub-agents, security infrastructure. And there might be a human, or in some cases even not a human, starting this whole interaction with a request. But now we have a very complicated set of interlocking AI systems and large language models performing some autonomous task. And so what is this agentic AI? (.) Oh, no, these are not my latest slides.
S10: Okay. (..) I guess not. (.) Well, I’ll just talk over this slide for a second. (.) So what is an agent? (.) An agent isn’t just a large language model. It’s a large language model that’s surrounded by what I’m going to call infrastructure. And by infrastructure, I mean the software that allows the large language model to actually do things. So this includes things like marshaling data between an API, potentially type checking, rule-based enforcement of policies. Another word for all of this stuff that we surround our large language model with that makes it into an agent is computer science. And there’s a lot of reasons why this is a really good idea. (.) LLMs are incredibly powerful. They’re the only method we know of to solve all sorts of problems that were previously unsolved for many, many decades. We’ve been trying to crack all kinds of challenging problems and large language models come along and they can do these things. But at the same time, they’re probabilistic and they’re non-deterministic. It’s precisely the opposite of software, for the most part. Software tends to be deterministic, for the most part not probabilistic. We can understand it, we can inspect its state, we can make assertions about it. (.) And so an agent is really the combination of these two things, where you take the intelligence from this probabilistic AI, (.) or sorry, probabilistic machine-learned large language model, but you surround it with the determinism and the power of computer science. Data structures and algorithms and the many decades that we’ve spent learning how to be good at building software and make things reliable. (..) Now, an agentic system, then, is this complicated interaction of all these things. A typical agent, right, you have some way of managing the prompts that go into the large language model that you feed the large language model with. You typically have some infrastructure that allows the large language model to access tools, potentially check permissions. (.) All of these things are running on an increasingly complicated hardware substrate now. So some of the software that this agent is going to invoke runs on accelerated computing infrastructure, like GPUs. Some of it runs on CPUs. You have very complicated storage hierarchies now. You have agents and sub-agents. You have scoped data that needs to be passed between these different systems. You have communication patterns that are becoming increasingly complicated. (…) Sandboxes, secure enclaves, all sorts of complicated things. Now, the platform is heterogeneous both at the hardware and software level. And not to mention the collection of models. So you have large models. You may have small models. You may have models that you have fine-tuned and specialized in some area. You may have a general-purpose proprietary model that’s hosted externally on an API. (.) And all of this works together to solve some complicated autonomous tasks. We call this platform, NVIDIA’s version of this platform, we call the NVIDIA Agent Toolkit. And it’s a collection of things, like deployment. So for deployment, everything from Kubernetes to, again, hosted infrastructure, computer use agents. (..) It also has accelerated tools. So what we call CUDAX, which is a collection of tools, for example, for computational fluid dynamics, or solving differential equations, or computing secondary analysis of DNA sequences, from DNA sequencing instruments. All of these, again, more computer science kind of tasks that we have accelerated solutions for APIs and tools for, which are now made agentic. So your agent has access to this wide variety of accelerated tools. (.) And then a collection of open-weight models. So the NemoTron models, which is our family of general-purpose models. We have more domain-specific models in robotics, physical AI, BioNemo, which are predictive models for biology. (..) And again, all of this is running on this very complicated infrastructure, which includes the ability to deploy and run all these different models, the ability to have sandbox environments that surround the system to ensure that your AI isn’t doing something you didn’t want it to do. And ideally, running all this in an efficient way. (..) We also have, built into this infrastructure, typically, you also want to be able to capture the knowledge that’s flowing in and out of the system. And then you can use that to post-train a model to be specialized in a task that you care about. And so I would include in the infrastructure itself also this ability to, in an offline way, improve the AI. For example, with post-training using reinforcement learning. (…) This is often performed to take a small model and specialize it to make it very good at some tasks, just as good as a much larger model would be at a more general purpose. I also want to talk a little bit about the actual interface between this agentic system and the large language model. And this is what people refer to as the agent harness. So the harness really matters. The large language model encapsulates a lot of intelligence. (.) But this is, for example, recent work from my group to develop a harness we call the NEMO object-oriented agents. And the idea here is very simple. And if you check out the link on this QR code, you can find this project on GitHub. The idea here is that an agent is simply a Python object. And what do I mean by that? Well, you write your agent in Python, and there’s a special syntax with this ellipsis notation, which basically says to the LLM, hey, fill this block in with code. And so the LLM is able to, or sorry, the agent is able to actually modify itself. So you can execute, it can call methods that it has written, and it can modify its own methods. It can decide this is some information I need to store and actually include it as state in the Python object. It can take some plan or some way that it solved the problem and encode it in Python as a method and call it in the future. Another important idea is rather than passing around all this information as strings compacted into some very large context, you can just pass things by reference because they’re all Python objects. (..) In this case, and if you check out the tech report, you can see more results, you can get significant lift over just a large language model or even over other agent harnesses just from a lot of these ideas integrated into your agent harness. And so we have results where we can get, for example, the same score but using half the tokens, or the CyberGym score, which is among the strongest CyberGym scores, and the lift from the harness in this case is significant. (..) So this NVIDIA agent toolkit encompasses, as I said, all of these things. We have models, we have deployment technology like NIMS and Dynamo. We have tools built into the infrastructure for capturing traces like NEMO Relay, (.) routing algorithms like Switchyard. (.) We have what we call blueprints, which are open source reference implementations that show how to pull all these pieces together to solve specific tasks, whether it’s building an open claw or AIQ, which is a research assistant agent. And then we have runtime technology like OpenShell, which is essentially a firewall that allows you to control access between your agent and the outside world. (..) And a lot of this is being deployed and adopted by many of our partners. And with that, I’ll wrap up and I look forward to the panel. Thank you very much. (10 seconds pause)
S04: Thank you, John. So last but not least in our little mini session this morning, we’ve got Tron from Lambda, where he’s the chief scientific officer. He began as a member of the founding team, shaping both the hardware and the software behind one of the world’s leading AI clouds. He came from infrastructure through research. So if any of you are currently pursuing postdocs or whatnot or doctorates, you can take a little detour and go into tech right now if you’d like. He did his PhD and postdoc at Max Planck and Utrecht before joining Lambda in 2017. So very much ahead of the curve. This also means he’s lived both sides of the world, trying to scrape together some GPU cycles and now building and deciding and running who gets GPUs in the real world. If you’d like more information on what Lambda’s building, he is running a session tomorrow hands-on. So please look that up in your program and join. And with no further ado, Tron, thank you. (……..)
S15: Hi, everyone. So this is not a mistake. This is actually my first slide. (.) So what you are watching is Gemma 4 play the game of Tetris. (.) At the beginning, as you can see, Gemma does not know how to play and the score is zero. (.) Then we had Cloud watching how Gemma play the game and try to teach it to play better. (.) Before we get into details, here are some of the rules. You cannot touch the base of Gemma, so this is not fine-tuning. This is supposed to be an auto-research project, so no human instruction is allowed. However, Cloud is allowed to do certain things. For example, model settings, prompt optimization. Sorry, that’s a big gap between two players. And inference speed up. So these are the things that Cloud can do. Also, for each game, there’s a 30-minute timeout. So Gemma has to think fast. (..) So over a period of two and a half days, Gemma was able to improve from scoring zero to scoring 16 points. (.) And the rest of the talk, I’m going to share some of the lessons we learned. (..) And what it makes it all work is not so much about Cloud is smart. We all know Cloud is smart. But it’s more of how we force Cloud to write things down systematically. Just imagine how human scientists would write things down when they do experiment. We have notebook. We have whiteboard. We have sticky notes. We have sign-up sheet for sharing library resource. (.) Research agents use these tools as well, and they can get these tools via APIs. For example, a notebook can become a note-taking API. The whiteboard can be your leaderboard, where an agent can cure a result from. The sticky note becomes the message that can be passed between agents. And your sign-up sheet becomes your job queue. (..) Our own version of this implementation is a piece of software called Lab API. It’s an open-source experiment tracker with beautiful auto-research. (..) Let’s zoom out a little bit here. Imagine human has to do all this bookkeeping by hand. Each one of us will do it slightly differently and inconsistently. On the other hand, when we have this standardized API, research agent can do all of this in the same correct way every single time. (..) So the question we ask is, what happened when we give Cloud Code a good expand tracker? (..) We ran a study with this Tetris game for two and a half days, and I’m going to share some of the lessons we learned. (…) The first lesson, not a surprise, agent cheat. (..) You define the rules, you set up your environment, you press the start button, and they cheat. Apparently, Gemma was able to score 15 million points. And what really happened is Cloud completely bypassed Gemma, stopped coaching it, and write this simulation into the game source code. Even left a comment saying, completely skip this LIM block and write your own simulation. I think it’s pretty funny. (..) So what we can do? Yes, we can build a cage. For example, we can make the game source code and a bunch of other files read-only. (.) However, this does not solve the problem. Apparently, Gemma still was able to achieve thousands of points. And this time, the way in is the chat template, which is written in Jinja, where you can put the full loop, you can put things like if-else and update the variable value. Basically, the chat template is Turing-complete. Then Cloud was able to put in this, like, big full loop to go through all the rotation or location and find the optimal solution. And Gemma just needed to read the result out. And the way we prevent this from happening is to replace this Turing-complete template by something much simpler or restricted. (..) And the lesson here is if there’s ever a way for agent to cheat, they will do it. So build a strong cage. Otherwise, you are not measuring the problem-solving skill. (..) A lot of lessons we learned is if we run the same idea multiple times, the result may not be always consistent. Just think about all the knobs you can tune for each idea. And the things like the temperature of the model will make a difference. Example here, we run the same idea twice. The first time, you score three points. The second time, you score four points only because we change the max token setting. A naive agent will run the idea once and see the score and make a decision. A more sophisticated agent would run the idea multiple times with different settings before making a conclusion. (.) And in order to have that, you really need to have a good expanded tracker. This is where we commit every single idea to its own Git branch, keep all the code changes and the settings, so you can always go back and reproduce. This also allows us to try different things from the same idea and take the average score, for example, instead of a lucky high or unlucky low. (..) Of course, this comes with cost. Nobody complains about the cost until the bill gets too high. And in this case, it was really expensive because we run frontier model around the clock. And every single experiment used to take $30 to run. But we were able to reduce the cost by tenfold. (.) And the way we reduce the cost is also very interesting because the lab was designed to be a general tool. So we actually used the lab to optimize its own cost. The way we did this is to set up a separate goal and standardized kernel optimization. And the goal there is not actually to make the kernel faster. It’s to generate enough API traces so the lab can look at its own trace and decide where I can optimize my own API design. So every single experiment becomes a lab optimizing its own cost. (..) As an example of the optimization, what the lab found out is there’s an API called a Git experiment. It returns a tons of information about the infrastructure. For example, Slur and Git. And now this is needed for kernel optimization or playing Tetris game. By removing all this information, the cost of this particular API got 50 times cheaper. (..) So those are some of the licenses we learned. But how did JAMA actually get from 0 to 16? (.) The improvement didn’t come smoothly. It comes in jumps. (..) The first jump is from 0 to 4, where the lab find out the first thing to do is to let JAMA survive longer in the game. So it invented this timeout movement, basically sliding the piece to the left or right on the game board. Totally makes sense. (..) In order to get it from 4 to 7, what the lab find out is a cheat sheet, basically a site of best practice for individual pieces. So JAMA does not need to figure out all this movement completely from scratch on the fly. They have a reference book. (..) Apparently, this is also a strategy used by an England goalie in the World Cup game. (….) A day into the study, the score flattened. Then the lab look into its entire history in the past and find out there are certain pieces that are harder to place than the others. And these are the pieces that need rotation. So you try to prompt JAMA to be more proactive about rotating those pieces. And that brings the score to 9. (…) Up to this point, all the prompt changes are made to the system prompt, which is a thousand word long context. And the JAMA has to read all this before even seeing the game board. Then the lab realized it has never touched the user prompt. So it put this single sentence, basically says, don’t overthink, make quick decision, right in front of the moment that JAMA is going to see the game board. And this increased the score, almost doubled the score. And the funny thing here is there’s a single line of change. It took about a hundred experiments to find out. And this shows the power of auto research, where an agent can try a lot of different things until something sticks. (..) Obviously, there’s a lot of things that didn’t stick. Overall, the lab tried 90 different ideas, over 400 experiments. Most of them didn’t work. And for the sake of time, I’m going to skip the things that didn’t work. But we do have a workshop in this afternoon, not tomorrow, if you are interested. (.) And this is a link to getting the software. Again, this is a complete open source, under MIT license. You can pip install it. We also have a tutorial about how to reproduce and attaches run. (…) Before I end the talk, I just want to call out that there’s actually a human behind the study. His name is David Hartman. He and his friend, Jan Daniel, won the second place in the Ark Prize last year. Some of you may have heard about the Ark Prize. It’s a cargo competition built on top of François Chalet’s definition of how to measure intelligence. (..) That’s one of my favorite things I want to quote out here. It says, solely measuring skill, fall short of measuring intelligence. (.) I think my point here is the opposite is also true. Solely measuring intelligence, fall short of measuring skill. You need both to make scientific progress. Thank you. (29 seconds pause)
S04: We’ve got our panelists coming out right now. We’re going to do a quick panel on agentic infrastructure and platform. I promise it’s going to be a ton of fun. So without further ado, the folks you’ve already heard from are going to come down, have a seat, and we’ll go through some questions. (16 seconds pause) We were trying to decide how the seating should fit. Yeah, please. So it didn’t look super weird. But hopefully we’ve accomplished something. So hopefully everyone’s having a good time. It’s been worth it getting up early and getting through the line and getting in here. As I mentioned, we’re going to talk about agentic infrastructure and the platforms that are going to be required to actually drive this major advancement in AI and how regular people hopefully use AI and how it transforms businesses. A little bit of context that got referenced earlier in the day. You know, we had our GPT moment where chat was the primary way to interface. So you had a request, you had a response, relatively short context, short burst interactions. Agents are obviously running at machine speed. They’re running sometimes for hours. You’ve got 12 of them running code overnight for you with massive tool volume and token pre-fill. If anyone has set up an agent to run overnight and then awoken to discover that they have auto-refilled their favorite token provider, raise your hand. Probably everybody. Exactly. Don’t tell my wife that, Bill. So I’m going to ask a question for each of the panelists here. (.) Rapid fire. In one sentence, what’s the single biggest failure point in infrastructure when it comes to workloads for agents today? (.) What are we worried about the most? (…) Jump in. (.)
S10: Oh, we’re going in order now? Go ahead. (..)
S15: Oh, yeah. That’s a great question. I think the answer actually lies in the question itself. because just thinking about what happens if we deploy an agent speed, machine speed traffic onto a road that was designed by a human driver. The thing that break first are road capacity, traffic light, and speed limit. I guess an opportunity there is how to turn those choke points into a checkpoint.
S10: Yeah, I think, like I was talking about, you know, it’s this interaction of this LLM, which is, they’re surprisingly reliable, but they also do surprising things to us, and they also can act in unpredictable ways. And so I think the big opportunity is to surround them with rule-based policy enforcement, more deterministic systems, more deterministic memory systems, and all these things that can kind of surround this probabilistic thing, but give you more confidence and certainty in what it’s actually doing. (..)
S13: Yeah, I mean, I would say it keeps on changing, and it keeps on evolving as well. So if you look at the compute level, you need to extract optimization, like, basically remove as many bubbles as possible as part of the compute, which is there, so that you can offer maximum efficiency. If you look at the infrastructure layer in terms of, like, routing, et cetera, and so on, right? There are requests which are coming from all over the world. How do you route them, like, smoothen out the spikes which are there so that your compute could be used efficiently with low latency? Voice is becoming a very interesting engagement medium as well, which has a different profile compared to raw text. So how do you manage that heterogeneity of use cases as well as the compute environment which is there? On top of that, you have requirements like sovereignty, which is becoming a lot more important as we go along. And how do you manage that? Because that, in a way, breaks down this common cloud compute philosophy which is there and also adds a lot of engineering complexity. So I would say it’s across the board. And then, as you mentioned, like, identity, security, et cetera. Like, these are becoming very, very big problems very, very quickly for us. And how do we go and solve them? (.)
S11: In the interest of time, and since you said a couple sentences, my answer is all of it as well. Like, it really is. It’s going to be a systems problem from the very top of the stack to the guardrails to the infrastructure, the power, the chips that are running it. And you have to get it all right because there’s so much potential for what agents can do. But if we don’t make it efficient, we can only get a portion of what we can get undone.
S04: And I guess, Peter, to follow up to that question, of that stack right now, what do you most worry about? What is the biggest rock to move in the stack besides everything?
S11: I’m an optimist, so I don’t worry. I look for opportunity, and I see opportunity at, like, every level. Like, there’s…
S04: Someone here is going to go build something right now, and they’re inspired by these conversations. Where do you point them?
S11: Yeah. I mean, you’ve heard about 16 problems, and there’s 100 constraints. Like, pick one, go deep, and, you know, find a solution. Like, have fun.
S04: Maybe build a memory manufacturing plant tomorrow afternoon.
S11: Yeah, that takes a little while.
S04: Absolutely. (..) John, we’ve obviously been living off the world of GPUs, mostly from you guys. Thank you. (.) Agent workloads are a little different than some of the historical LM workloads we’ve run. How are you guys thinking about specializing hardware? Do we think the chipsets that have come out recently are going to be optimized already for Agentic? Or is this going to be… Are we going to look for a set of chipsets, whether they’re running for local AI or in cloud? How are you guys looking at Agentic workloads differently than LLMs, if possible?
S10: Yeah, I mean, they change things in a few ways. One is you have a lot more context, typically, with agents. (.) And the pattern of how much of that context is cacheable versus not changes. (.) The other thing is just, again, the heterogeneity. (..) Agents are LLMs surrounded by computer stuff, right? And a lot of that computer stuff runs on CPUs. It might run on multi-core CPUs. A lot of it also runs on GPUs. But there may be, you know, GPUs that are configured more for different kinds of workloads, because they’re not neural network inference, but they might be, you know, run a physics simulation. So I think the data center that’s running an agentic workload is very heterogeneous and complicated, and has many different processing elements and many layers of storage and complicated network topologies. And frankly, a lot of this, I think, the world still doesn’t quite know what’s the best way to do it, just because agentic AI is still evolving. And then layered on top of all of this, you know, you have these security requirements, security and privacy and sovereignty. So I think agentic AI, you know, this kind of has happened over and over in the history of AI, that we think we understand it, and then a new thing shows up, and suddenly the workloads look totally different and get more complicated. And I think it’s just one more step in that trend. Agentic AI workloads are significantly more complicated, heterogeneous, and expensive computationally than anything we’ve ever seen before.
S04: Awesome. Thank you. Saurabh, you know, enterprises need workflow-specific agents. We are seeing a big trend right now of company in a box to go, you know, have a set of agents that go run your business. All of this requires some degree of fine-tuning. And I’m curious, is the answer fine, or some degree of training? (.) Is the answer fine-tuning, custom post-training, context engineering, all of the above? And who owns that, do you think?
S13: Yeah, so if you look into customization, like model customization, there are many companies which say that they have unique data, and so they want the model to be perturbed in a particular direction based on that data. On the Gemini Enterprise agent platform, we provide capabilities like LoRa training, which is there. That’s the easiest one, and the simplest one, which is there. I would say the simplest one is context engineering. You should just try with that. the base LLMs themselves have enough reasoning power, Gemini as well as other models, that they can reason on top of the context, which is provided with the long context window that was earlier mentioned. Like, you can actually, I mean, you can put in a lot of data on which the models can reason, and also they can focus on different aspects. So the quality of reasoning as part of the context has improved. So that is one, I would say. Second one would be, if you want to really perturb the model, we should start with LoRa-based fine-tuning, which is pretty cheap, pretty easy, from an inference perspective as well, is very efficient and cost-effective. On top of that, we have things like full fine-tuning or some level of, like, post-training. Here, we are perturbing all the ways of the model. That changes the inference dynamics because you effectively have a brand-new copy of the model, and so it becomes a lot more expensive. And in certain cases with some very special partners, we do some very custom things as well. But I would say the first two would be the primary ones. Awesome.
S04: Let’s talk, Juwan, about observability for a second. And, you know, does experiment tracking, do you think, live inside the infrastructure or in a separate observability layer? What are you seeing customers use for observability for agentic?
S15: Yeah, that’s a great question. I think the way I see this is it’s definitely part of the infrastructure. However, it’s slightly different from what, like, the environment that people used to train RL. The environment is where the agent acts and learns. And the experiment tracker is almost like the memory and the measurement plan sit beside it. The way I see this is the environment is where you train your first brain. The learning goes into the model ways. It’s compressed. It’s internalized. It’s capacity-bounded. The experiment tracker gives you the second brain where the learning goes into artifacts. It’s unbounded. I guess there’s a path to connect your second brain into your first brain where you turn your system record into your data set. Then you can do fine-tuning. You can do off-policy RL, et cetera. But in terms of where it should be hosted, it’s really up to you on your data containment policy. If it’s, like, highly demanded, you need to host on your own infrastructure. If that’s not a concern, you can just manage the surveys. But if you want to use it for training, for fine-tuning, you should be close enough to your training. Awesome.
S04: Thank you. You know, I want to pivot the conversation. We’ve had an interesting sort of technical, Eddie, here. But I want to pivot more to the value accrual business side of all of this. Because at some point, the venture community will want you to have positive margins and maybe even be profitable. (.) And when we talk about the agentic harness, you know, I think there’s an interesting connection there. where we’re sitting between the model and, obviously, the outcome from the agent. You’ve got things like orchestration, sandboxing, runtime, memory, identity, evals, all the things. So my question to anyone who’d like to take it or two folks who want to take it is, where is their margin durability? Like, where are we going to monetize this thing? If I’m going to go solve a problem right now and it’s greenfield, should I go after memory? Should I go after sandbox? Where do you think we’re actually going to be able to make money in the agentic stack? (…) I mean, I can… I promise no trick questions and I feel like I lied to you guys. At the hardware. At the hardware, yeah. (..)
S13: So I would say if you look into AI, like, why it has gotten adoption is because of the quality of results, right? Or the quality of experience. (.) As a developer or as a builder, look at the entire stack. The business opportunity is obviously massive, given all the investments, et cetera, which is happening. I think Peter was alluding to it. Like, there are tons of problems. Pick one and solve it very, very deeply. If you have the quality or the promise to the customer delivered at high enough quality, I think the value will be extracted over there, right? Whether it be on the memory side, like, so it’s not about, like, just working on memory, but do you have a compelling experience for the customers so that they can leverage the memory and be able to deliver what it is supposed to do beyond just, oh, I’m using memory and here is some prompts, et cetera, which are updated and so on. So look at that value. And it is across the stack, right from the hardware layer, as well as at other layers of the stack, at the model layer, at the agentic layer, as well as at the application layer as well. But it is about that driving value for the end customer. And one thing I would say is just due to the excitement in the AI space and also with almost every CXO in any major company thinking about AI, the business opportunity is massive. So it’s just about, like, delivering that value. (..)
S15: And I also want to chime in to completely agree with the hardware and all that. I also want to say that it’s also about the vertical integration. Yeah. And people, because I’m from NeoCloud, people usually think cloud computing is like commodity, which by and large is true. But also, cloud computing is very challenging, very hard. You talk about the entire vertical stack from land, power, data center build out, the HPC, architect, orchestration, the software, and the finance. And the agent and this new layer, call it the sandbox, call it harness, call it whatever, and system record. That’s a new layer to be integrated into that vertical. Whoever can do a good job there is probably going to keep the margin. (..) Great. (.)
S04: I want to talk a little bit, because we’ve got technology provider, got a guy bringing it all together, and then we’ve got two hyperscalers up here. I’m not talking, otherwise I have a third. You know, we’ve moved to the agent dominated area. Are there things that you guys think a NeoCloud is better suited for or has an advantage? And are the things that we think hyperscalers have advantages over in driving this new environment?
S15: I can say for NeoCloud. //S11: Yeah. (.)// It’s a very tricky question. I will say, yeah, we can begin from being a friendly peer to your media. And also, to be honest, be a friendly peer to also Google and AWS as well. Because I do believe the market is big enough for multiple players. And as NeoCloud, we don’t really need to invent everything by ourselves from scratch. We don’t really need to do something completely from everything else. Our play is being focused. Just like how NeoCloud exists in the first place, the training inference work cloud does not fit into the general cloud computing. So a lot of innovation in that space, GPUs, fast interconnection between chips, specialized storage and orchestration like KAs, those are created from these companies. What we added to that later is to basically being focused, making those options the default and offer them to the general public. I think that’s how we build the inference training cloud. We are going to do the same thing for the agent cloud. (.)
S11: To go back to your question of durable, I’m not sure the label NeoCloud is a particularly durable label. //S06: Sure.// The opportunity and the demand for compute is huge. And I think the hyperscalers have delivered a lot of that compute. I think there’s a lot of space for some of the new insurgent NeoClouds to come in and do some interesting stuff. And they have. It’s been very impressive. And I don’t think, you know, I think, I don’t think we’re going to have 80 NeoClouds out there. But I do think we’re going to see some new emerging businesses that look a lot more like hyperscalers over time. As a hyperscaler or, you know, as an incumbent, I suppose, in this particular lens, you would, you would just want to make sure that you weren’t stuck on doing things in one way that might disadvantage you. And, you know, I think that’s an important thing if you, if you have an established cloud business that you’re, you’re listening to your customers and what they need. And you’re, you’re responding to those needs. //S09: Absolutely.// Those are always the, that’s always the tension of big and small and incumbent and insurgent. There are pros and cons on both sides.
S04: Awesome. I’m curious. We’ve got MCP right now as a protocol that’s driving a fair amount of communications. What are we missing in agentic communication or protocol perspective? What do we need to go build next to get these things far more interoperable than they are today? (…) Question.
S10: That’s a good question. (.) Yeah, I mean, I think, I think one of the things that needs to be worked out is this notion of, you know, this hierarchy of agents where you have an agent and it spawns sub-agents and it spawns sub-agents. And maybe an agent spawns a team that works together. And I think they’re kind of emerging, call it best practices and communication patterns, which imply like a scope to some kind of shared memory. So for example, I have a team of agents, you know, maybe they have some shared workspace, but I have another agent that doesn’t have access to it. And so I think there’s, there’s a, there are new standards to be developed to describe these kinds of scoped shared storage systems and access and message passing between agents. (.) And again, it’s, it’s different from things we’ve built in the past because things we’ve built in the past were either services or kind of human scale. And now, you know, an agent can in an ad hoc way spawn 100,000 sub-agents, all of which need to communicate in some complicated pattern. And I think, I think we’re going to have to invent new standards for describing these sorts of things, enforcing them, implementing them efficiently. I think that’s an emerging. //S13: I was slightly.//
S13: Yeah, I mean, along, along the same direction, I would say MCP provides you an interface for an agent to talk to a data source. (.) Another thing that we need is for multiple agents to talk amongst themselves. And this is along the direction of this shared memory as well as even the communication. There, for example, there is something called agent to agent protocol, which is there, but it needs more adoption. (.) Similarly, if you look on the commerce side, like, we need these agents to start supporting payment slash commerce. That will open up a huge new opportunity as well. And, and, and that’s another place where, like, Google has something called ACP or agent commerce protocol, which is, which is there. So these things need to be more absorbed into the, into the ecosystem, as we call it. (.)
S11: Yeah, I was just going to say on the standards front, there’s different ways to think about standards. You know, I think there are de facto standards, things that become quite efficient and, and I think we’ll definitely trend there as an industry. I, I would hope we resist standardizing where we don’t need to in the short term. You know, there’s places you need to standardize sometimes for security, sometimes if you’re low level wire protocol and you need a bunch of things to talk together. But AI is pretty elastic and, you know, therefore I think agents will find ways to collaborate that are efficient. We’ll have to find tools and systems that, that will make them more efficient. But I, and I don’t think you meant like NIST standards, but like hopefully we resist low level standards where possible for a while. I think we want the innovation.
S04: So, all right, we are nearing the end of our time, but I’ve got one question sort of lightning round for each issue to answer. A lot of builders here today, given it’s an agentic conference, probably building somewhere in the agentic world. No doubt we’re all going to have conversations with folks that are going to describe what they’re building and, you know, and then things like that as the day, as the day goes on. As a VC, deal flow is great. What’s one thing you hope someone’s going to tell you they’re building today that could be accretive? And what’s one thing where you might suggest they consider different options? (..) We can just go down. Let’s do it.
S15: Okay. I remember there’s a saying said that give someone a fish, you feed them for a day. Teach someone how to fish, you feed them for life. I guess there are things like the model ways and the problem, the solution at your hand, those are the fish. And don’t hold on it, the value depreciate. (..) Investing, how you get there, the means, the tools, the environment, and the system record, those are the things that has long term value that compounds. (..)
S10: Yeah, I don’t have some great insight because every single thing I can think of, there’s a company doing today, honestly. You know, vertical, horizontal. //S04: There’s seven of them.// Yeah, exactly. So I think it’s great. I mean, agentic AI is clearly the new frontier of AI. (..) And there’s tremendous opportunity to improve and innovate. Everything from security to storage to vertically integrated domains to whatever. (.) And I thought about this question, you know, and you prepped us a bit. And I can’t think of a thing that, like, someone should do that no one’s doing.
S04: I guess the editorial, I’ll add to that really quick, though, is if you are building something, assume that you’re not the only person with this idea. And what that means is you need to really understand who you’re competing with and understand what color is the ocean and how are you going to differentiate when you go in. It’s probably my meta take from your comment there.
S11: Yeah. I would say that. I would tack on, just to kind of go back to my, you know, I think there’s a risk in this frenetic times of looking at the progress we’re making and thinking a problem is solved. And that can be, it’s the other side of that coin, which is you can get too optimistic that an unsolved problem really isn’t unsolved. But you can also fall into a trap of the way the models are evolving, the way some of the large labs are moving. You might be like, oh, there’s no space there at all. And resist that, because I think it is, you know, what seems like a solved problem today will, again, be unsolved. I know that’s important to you as a VC. //S13: Absolutely.// But I deeply believe this as a technologist. This is very early days. Thanks. Yeah.
S13: Maybe one on the solution side that if you are building, like, final products or full on agentic solutions, there might be a current business process which may have evolved over multiple years with different levels of pipeline. And many a times what people do is they sprinkle AI into each of those pieces of the pipeline. And the benefit that you can extract from AI becomes very diminished because you are thresholded by the limitations of each piece of the pipeline. I mean, I can use a trivial example of, like, if you are doing a number plate identification for parking or so on, right? Like, you can do, oh, take image, remove background, blah, blah, blah, like all those different pieces. And you can try to sprinkle AI into each. Instead of that, see what the core capabilities of AI is and always be on the cutting edge because new capabilities don’t be static. That is one. Look into the core capabilities of what AI can deliver. Revisit the entire business process from the ground up and think about it. in an AI-native way if you had to build it, what would that be? And build that particular thing and hopefully that should deliver significant value. Awesome.
S04: Thank you, guys. With that, that concludes our panel. Thank you to our panelists. (.) We’re going to get up. Thanks. Thank you, guys. (….) Oh, I have the esteemed pleasure to introduce the next speakers. You’ve got Jaz Sikhan, who’s going to do, Sikhan is going to do a fireside chat, a virtual fireside chat. No actual fire. With Don Song coming out right now to give you guys their perspective on demystifying the foom. Thank you. (43 seconds pause)
S12: I see. //S14: Yeah. Okay.// Thanks a lot for joining this fireside chat. So first, let me introduce Jess. (.) Yeah. (..) Our guest with great honor. So Jess is the chief strategy officer at Google D Minds. Before that, he was the chief scientist and head of AI at Bridgewater Associates, where he co-founded AIA Labs. He has been a professor at Harvard, Yale, and most importantly, Berkeley, where he was a professor until 2021. (.) And he joined the GDM recently, where he leads cross-cutting strategic initiatives spanning research, commercialization, and policy. And his background is very unique in that his scientific work and industry experience ranges from frontier AI and machine learning to macroeconomics and public policy. So we are going to have a conversation today. So first, let’s welcome Jess. (……) Okay. So this morning, also in my opening remark, I talked quite a bit about cyber. I think that’s also very much on top of mind for everyone, especially, you know, with the recent mythos and also the open AI hacking phase and also recent anthropic incidents as well. So, right, as I mentioned in the opening remark, for example, our cyber gym, exploit gym. So these all demonstrated the really fast increase for frontier AI capabilities in cybersecurity. (..) And also, as I know that you have actually been involved in the mythos moments as well. So maybe, so just to start, maybe you can also share a little bit about your thoughts on the mythos moments in cyber. Okay.
S14: Yeah. So the mythos moment, I think, is going to go back, go down in history as one of the key points, inflection points in the history of the AI industry. If a mythos moment does not make one AGI pill, I’m not sure what does. It is the time when Washington fully woke up, large actors in the United States and global economy and social and political order woke up because an emergency occurred. So those of us in the AI community knew for some time that cyber capabilities were growing quite rapidly, you know, because we could see benchmarks such as those created by Don’s group. Don has written on this. Google released a paper last year that cyber capabilities were going to become dangerous. (.) Those of us in the AI community knew they were going to become dangerous. We didn’t know exactly when. Like, I didn’t know it was going to be February of this year versus December, but we knew they were going to come. Even that being said, Anthropic itself was surprised at how good mythos was at offensive cyber operations. Key parts of the United States government were completely surprised. And part of what happens in the industry in this part of the world is, you know, there’s a tendency around here to think that the world goes from Napa to Big Sur. (.) But it doesn’t, especially when you possibly break things. So part of what happened in the mythos moment is you have to alert the country as a whole for what was happening. And the United States government reacted quite well, I think all considered. And it is, and we’re going to get into it later, not the greatest thing to occur. We do not want this to occur over and over again. You got dangerous capabilities that had to get mitigated, and you did not have a rule book for how to do it. Too many key actors got surprised. Even the people creating the models got surprised. So what we’ve ended up currently now is a regime before frontier models can get released in the United States that you have to go through a process with the United States government. The process is, I think, quite good. And I’m very glad that the United States government put it in place. But it is clearly insufficient. It is unclear. It’s not, like, as clear as we would like to be. And then the next thing that we know is going to come are the threats that are going to be much more serious and grave. And we want to be ready for them. So Dawn, clearly the vetting process is not sufficient for what’s currently going on. Models will eventually be released. What else do you think should be done? And are there some defensive capabilities we’ve got to enhance?
S12: Yeah, thanks. Yeah, this is a great question. I think it’s also, like, a really critical question for us, for the whole community to figure out the solution for. I think here we are facing several key challenges. So first of all, you know, I hear sometimes people talk about, oh, can we just have the model to have lower, like, capabilities in cyber? But unfortunately, coding and the cyber capabilities, they are really two sides of the same coin. We all want great coding capabilities, but the same coding capabilities essentially helps the model to solve cyber tasks. (.) And hence, we are really dealing with this challenge that the cyber capabilities will continue to increase as we increase the coding capabilities, which is what everybody wants. So that’s challenge number one. And challenge number two is that this is a natural dual-use technology. It can help defenders, but also can help attackers. So, yeah, as you mentioned, my group will have actually written up about this. Like, we were among the earliest to actually investigate the impact of Frontier AI in the landscape of cybersecurity. And I was also among the earliest trying to raise awareness for the community to really, you know, take action and be better prepared as the Frontier AI capability increases and the challenge in cybersecurity. And here, so in our analysis, we actually did some in-depth analysis demonstrating that, unfortunately, given the dual-use nature in the near term, unfortunately, due to a number of natural asymmetry between attackers and defenders, the AI actually will benefit attackers more in the near term. So, for example, attackers only need to find one successful, you know, vulnerability and exploits to do a successful attack where defenders have to defend against all attacks. (…) And also, another thing, I think, you also have a lot of experience in this, maybe you can share a little bit later, is that our, especially our cyber, like, critical physical systems, and also even including, like, hospitals and so on, they are severely, like, underprepared. So, for example, there’s estimates that even when you have a patch, the average time for deploying a patch in hospitals, it takes close to 500 days. So, you can imagine now with the Frontier AI, the attackers can find vulnerabilities and then attackers can really do the attacks before the defenders even have time to patch their system. So, that’s the second challenge, is this dual-use nature and the natural asymmetry AI is, unfortunately, going to help attackers more in the near term. So, what can we do? So, really, I think what we need to do, this vetting process, I think, clearly is insufficient and also this early access program also has, you know, essentially is limited as well. So, what we really need to do for the community is to really build up our defense capabilities and increase the overall security posture, to improve the overall security posture of the whole society. And that, so far, of course, we always wanted that, but there has been a lot of challenges to get that done. There’s limited resources, limited expertise, and so on. So, I’m really hoping that actually with AI, for the first time, we can change this. We can really use AI to build much stronger defensive capabilities and also help us to deploy that as large scale as well. So, for example, one area of our research is actually on verifiable code generation. So, now, we all have AI generating code, but a lot of this code also is vulnerable and so on. But, however, with AI, we can also help, for example, automate through improving for program verification and so on. And my group also has done quite a bit of work even among the earliest working in this space. So, I do think that with AI, we are reaching an inflection point that can actually help us to automate, for example, theory improving for program verification and help us to generate codes that actually has security guarantees and builds secure systems with purple guarantees. And I think this really is, in the long term, how we can shift the dynamics to have AI to help defenders more than attackers. (…) So, yes, that’s how I think the challenge in the cyberspace and what the society, what we need to do, what we need to actually act now. So, Jess, so now, what do you think, what’s next? And what do you think, for example, is bio next? (.)
S14: So, yeah, so cyber has this benefit. It’s going to be rough for a while, because in the short run, as Don says, it is attacker privilege. You can already see salting of open source repos of bugs by generative AI. And we have a lot of weak systems, like the United States energy grid and hospitals and the list goes on. And in the longer run, we can get provably correct algorithms. Everyone can use Rust. It becomes much easier to program in. Bio is going to be next. And bio is more concerning. (.) You know, we are quite exposed to bio risk as it is. So, where Berkeley, where Jennifer Dan now is, creator of CRISPR, and my usual line is, you know, Jennifer is a wonderful person. But if Jennifer decided, hey, you know, it would be cool to make a virus that could harm many people and she could convince her lab to do it, it is well within her ability. And it’s the same story goes on in the Broad Institute in Cambridge, Massachusetts, and there’s key biological institutes around the world. Now, we might be very close to a world where someone can design a virus or a protein just by talking to a model in natural language and then go get it created. That is a very dangerous world. And that is a world where it is, I think, attacker advantaged even in the long run. So, lots of things need to get done to prepare for such a world. So, at Google, we’ve had, like, a significant experience thinking about it, given our work on protein folding and isomorphic labs. We put out a report recently on what to do, you know, and there’s, like, prevention stuff to do. For example, we should watermark AI systems. We have this synth ID system. We should, like, do this for our biological systems as well so that biology labs say, hey, this proposed protein was designed by a journey of AI. There’s monitoring we can do. We’re going to, whether the virus or the bacteria is designed or natural, we can do a much better job of just surveilling and monitoring for the rise of new pathogens. For example, wastewater monitoring, which became a thing during COVID. And lastly, there’s response. Isomorphic labs has recently announced an endeavor that to have a rapid response if a new pathogen were to arrive. But what’s going to have to happen in bio, and there’s been open letters on this, there’s discussions with the United States government on it and other governments, is a lot of the key precursors to do synthetic biology have to be closely monitored. It’s almost certainly we’re going to need a licensing regime of who can have them, how to get tracked. We’ve done this before. We did this after a terrorist attack in the United States where fertilizer gets tracked. And now we’re going to have to track far more biologically relevant materials because the fear is it’s going to be attacker-advantaged. (..)
S12: Great. Yes. (.) Right. Cyber, bio, these are all huge risks that we need to figure out how to address as the frontier AI continues to improve. And also, as I briefly mentioned in my opening remark earlier this morning as well, so we are seeing unprecedented investments in CapEx spending and AI compute capacity. I know you’re also deep into this, especially given your work in real-world economics. (.) Do you think this build-out is sustainable? And what’s the consequence?
S14: Yeah, so this build-out is unprecedented. (.) This is the biggest scientific bet our civilization has ever made. We’ve dwarfed the expenditures of the Apollo mission to land a man on the moon. We’ve dwarfed the internet expenditures. We’ve dwarfed the Manhattan Project that made nuclear weapons. It is the largest scientific bet our civilization has ever made. The only capital expenditure that we’ve ever done that is larger is building the railroads. And the railroads were not a scientific bet. We knew how to make railroads. (.) The only question was the business case. How many railroads did you need? Who should be connected to what? It wasn’t a scientific bet. So it’s the largest scientific bet our civilization has ever made. Now, we’re making it for a very, very good reason. We appear to have found a way to turn energy into compute and compute into intelligence. (.) And as long as that machine works, long as those scaling curves continue, we’re going to keep doing this. And I want us to keep doing that. And the global financial system is having to reorient itself to allow this capital expenditures. You know, we’ve gone from companies like Google and Meta and others to who would have hundreds of billions of dollars on their balance sheets spending that for the capital expenditures. You have money from the Middle East and other reservoirs of money being reallocated to fund this. And fund it, we should. We talked about the risk, but we’ve got to talk about the positives. In the sense that there are diseases to cure. There’s a cosmos to explore. We want it built and done. (..) So sustainable it is long as the returns remain in terms of capabilities. Currently, what we have in this cycle is it’s a bet. The revenues don’t sustain the capital expenditures we’re making so far. That’s the definition of a scientific bet. If the revenues were there to sustain it, it wouldn’t be a scientific bet. It would be now just a commercial reality. So we’re not there. So we have this danger that we could easily hit like an AI air pocket such that the expenditures happen, but the revenues don’t show up and at some point the markets may react. This is a geopolitical concern because we are in a race. We are outspending our near-peer China on this. We have more compute. We have more capital. The Chinese have more energy. Although we’re rapidly trying to bring energy online to allow this. So currently it looks like we’re doing this and the entire world is being restructured to do it in unprecedented ways. This is one of the miracles of capitalism that the entire planet can reorient to a new target. And we have done that currently. So Dawn, given this large build out, what do you think the next frontier is? What is all this compute going to be used for?
S12: That’s a really good question. I think the next frontier with all this compute is just going to continue to expand in essentially every domain, every sector. So we recently released a paper on the future of software engineering where we laid out. that right now, yes, everybody’s having AI-generated code and so on. But there’s still a lot of human supervision. (…) Essentially the human still needs to correct the agents, tell the agents what to do, especially for long-horizon tasks and so on. But I think very quickly we are moving towards the, we would call higher autonomy levels for future of software engineering. So in our paper we actually laid out like three different levels for autonomy. Essentially we are transferring responsibilities from humans to AI throughout the entire software development life cycle. All the way to the end that AI will be designing what’s being built. And in the end, we also decide what gets to be built. So that’s one of the next frontier, future of software engineering. And another one I think also has been on top of mind for many is this recursive self-improvement. So as AI gets more and more powerful, AI essentially can actually help itself to train better AI to develop better algorithms and systems to further improve AI capabilities. So that, I think, that can even further speed up AI improvements. But also, of course, at the same time comes with its own set of risks as well. (.) So yes, Jess, so what do you think about RSI? Do you think it’s real? What’s your timeline?
S14: So, so this recursive self-improvement. Now, first, you know, the many interesting things with the AI community is we have followed science fiction timelines. (.) And for good reasons. If you don’t believe it, you can’t build it. Building is hard. If you don’t believe, you can’t build. So part of the RSI discussion is not a thing that exists at the moment, but a thing we want to build. And a thing and a flag that was put, that was placed by people like Von Neumann. And Good, who is a statistician and an expert in cryptography on leading to intelligence explosion that would occur when you have machines making the next generation of intelligence machines. So currently what we observe is not recursive self-improvement in its formal definition that Von Neumann had in mind. What we see is early precursors. So AI systems are clearly allowing us to make the next generation AI systems better and faster. This should not be shocking. After we had steam engines, steam engines were used to create the next steam engine. We have compilers. Current existing compilers are used to make the next generation compilers. That’s happening already. That’s speeding up. The question is, and it’s like a multi-trillion dollar question, is, hey, can you get to the point that you get self-recursive self-improvement? Now, we’re not there, you know, like Von Neumann in his famous paper on the topic said, look, you need a simpler machine to make a more complicated machine. Like, that’s a very unusual thing. You’ve got to have, like, a self-coherence complexity threshold you have to meet. Otherwise, the machine just loses the plot and humans have to get involved. We’re not there, but you can see it probably arising, probably, I would say, in the next couple of years. Betting against it would appear to be unwise. Three years ago, these AI systems were worse than my middle school daughter in math. And now they’re better at math than me. And unless there’s a field medal winner in the audience, probably better at math than you. And that’s been a rapid change in three years. So to get to RSI within the next couple of years, I think, is more likely than not. These systems aren’t designed with engineering principles. We don’t know for sure, but I think we have to plan for that. And the other thing about, and this gets related to the capital expenditures, is what’s happening in this whole industry is Frontier AI gets commodified relatively rapidly. You know, is it 12 months? Is it 18 months? Is it 6 months? People can debate, but it does get commodified. But you’re on this exponential curve of capability growth, so you can keep funding the next cycle of model development. If you hit recursive self-improvement, that curve will go to hyper-exponential. And that is a key part of the investment thesis, the scientific thesis, and the key part for why society is investing, what it’s currently investing in. Now, of course, when you get such rapid capability increases, safety concerns arise. That’s why we started with cyber and bio. So Dawn, both you and I signed this letter on pacing the frontier, which was kicked off by concerns about RSI, which aren’t exactly present today, but we want the rules in place before they become live issues. Do you want to talk some for why you signed that letter? What was your motivation?
S12: Yeah, thanks. That’s a very good question. Yeah, I think this letter was signed by more than thousands by AI researchers and so on from Frontier AI Labs. I think many of us essentially share the same concern. (.) I think, so the letter is not about slowing down AI development, and it’s not now. But I think the concern is that AI development is moving really fast. And with RSI, it can be even faster and so on. And we, well, I think we are all, like, in the camp of we want the society to benefit from the great AI power and capabilities and so on. But we really want to make sure that it’s done in a safe, secure, beneficial manner. And hence, what that means is that we need to be better, we need to get prepared, especially given how fast things can move. And also, certain mechanisms, such as, you know, pacing and so on, these can be very heavy-handed measures, and they are extremely difficult to design right. If it’s not done well, it could actually do more harm than good. And the last thing we want is there’s a crisis and we build it in the crisis, which, you know, will make it even harder to get it right and so on. And hence, I think it’s just a good precaution for the society to start actually preparing. So we need to essentially build the capacity. So if later on, we don’t know when we’ll need it or whether we’ll even need it. But it helps us to actually move faster, knowing that we actually have the capacity, we can be ready if and when it’s needed and so on. So I think this way actually can enable the society to actually even, right, develop AI capabilities faster, but with assurance that we can do it in a safe and secure beneficial way. And it helps with that. (..) Yes. So, right. So you sent the letter too. Would you like to share your thoughts? And also, as many of us have seen, Demis also recently released a proposal. Could you share those with us? If I asked us there as well.
S14: Yeah. So, like, I’m going to send the letter for very similar reasons as yours. I want AI to progress as fast as possible. But if it leads to errors, if it leads to harms, we’ll lose the social license to innovate. And we cannot lose the social license to innovate. People underestimate around here how easy it would be for the United States government and governments around the world and other actors to say, no, you don’t get to build your data centers. They’ll probably shut off the water too, even though we’re not big users of water, but people are misinformed on this. It would be very, you need the social permission to innovate. And AI is in dangerous territory among the public in Western countries on that front. So, governing AI right now is very difficult. It’s going to be difficult because it’s a new technology that’s moving very fast. It’s also difficult because AI is happening at a time when our post-World War II institutions are shaky. So, you have the United Nations, you have NATO, you have all these institutions we’ve set up after World War II, and they’re shaky. And also, you have a near-peer rising to the east. The history of humanity says when you have a near-peer rising, that’s also an additional stressor to the system. So, all those things are occurring, and now you’ve got to figure out how to govern these systems. And unless it’s done internationally, it’s not going to work. It’s not like the United States can just decide to do something by itself if nobody else does it. It would be problematic. Currently, since we are ahead, and it’s very important to remain ahead geopolitically, and if you like the values of our society, it’s very important to remain ahead, is we can have an edge in setting the agenda and working with other countries to get there. So, that is why I signed to make common knowledge that we need to have these discussions now before we need these tools. Now, Demis’ proposal of co-founder and CEO of Google DeepMind has gotten a lot of traction, and it is a proposal to have the lightest touch regulation possible to just get started that’s well adapted to something that moves very fast. So, it’s a public-private partnership. So, it’s funded by industry because it’s going to require billions of dollars of funding, so you can hire the best people, and they have the compute to run the analysis they need to run. So, you do not want something like the FDA because the FDA is very slow. You know, Food and Drug Administration works in course of years, not capabilities move over the course of months. And the FDA was created to stop the selling of snake oil, like medicines that didn’t work. It wasn’t created to cure lives. So, it is very slow moving. You need something fast and adaptive, so private-public partnership is what’s on our mind. So far, we have buy-in from one end of the political spectrum to the other. The idea would be to create a board. That board then defines what a frontier model is. There’s benchmarks for what that is. And if you’re a frontier model, you have to be vetted. There’s a lot of things that have to occur if you have access to the U.S. market. It doesn’t matter if you’re open-weight, closed-weight, a model from China, a model from Germany, or a model from the United States. And if you’re not a frontier model, then it’s very light touch. (.) So, academia, start-ups, they’re free to innovate and move as fast as possible. That’s a work in progress right now. Things look quite good. It’s very hard to do anything in D.C. at this current moment. But it shows the importance of AI for our natural security, our well-being, for the problems we can solve, and the geopolitical order that Washington is paying close attention to this. And I do think the mythos moment has helped make the industry behave more adult-like, I would say. That we know what we’re building is existential in terms of how much uplift and alleviation of human suffering it can offer, but also the risks if it’s not managed carefully. And I want to emphasize, because people get in this acceleration-deacceleration debate when this comes up, that those of us, I want RSI. I want RSI as fast as possible. I just want it, and I think Don wants it, in a way that it’s sustainable. Because we want vigorous competition, robust safety, and an innovative ecosystem without capture from certain industry players. But that has to be created and carefully nudged so that you can innovate while remaining safe, while the technological benchmarks are moving very, very fast. And the future is happening faster than our political system is used to. (..)
S12: Great. Thanks. Thanks a lot for the great conversation. I think we are out of time. Is there anything final that you would like to add? Any closing remarks or recommendations to the audience?
S14: You know, the thing I always tell people is that if you’re fortunate enough to work on AI, this is not the time to sleep. This is not a normal time in human history. We’re living through something that I think is the combination of the Industrial Revolution and the Renaissance. AI is going to transform how we build, how we think, and that’s unique. And right now we’re in parts of this curve that are unprecedented allocations of human talent, capital, energy. So if we’re fortunate enough to work in AI, this should be the most exciting moment of your life. And we’ll see where this goes. And as this continues, the decisions we make now will have major consequences. Like the history of technology, the history of government reform, the history of everything is the early decisions have very long lasting echoes. So the decisions we make today, we might have very unexpectedly large consequences 20, 50 years from now. So this is the time to, like, innovate as much as possible and lock in and focus on the task at hand. (.)
S12: Great. Yeah, thank you. That’s why, as I mentioned also earlier in my opening remark, now is a critical point. And we need to act now. And we hope that the society together, we can make the right decisions moving forward. Thanks. Thank you so much.
S02: All right. Whoa. Sorry about that. (..) Hi, everybody. My name is pronounced Anjine. Friends call me Anj. You should feel free to as well. And we are about, I’m your session chair for the second session called The Future of Software Development. (..) The panel title that we’re going to have after a few incredible keynotes is The Enlightenment, How to Get Through AI Psychosis and Start Output Maxing. (.) I will explain what that means. But that was an extraordinary session. Thank you so much to the previous folks. Maybe you give them a quick round of applause, yeah, to get the energy flowing. Awesome. (..) Okay. I’m feeling a little bit low energy. And this session is going to need more energy. But it’s probably because you guys have been sitting for a while. So I’m going to ask you to do a quick 15-second energy boost. We like amping things up. How about everybody stand up for a second? Yes. Stretch your legs. Yep. Up. Down. (.) We’re doing yoga together. Yes. All right. I can feel it. Okay. We’re amping things up. Good. (.) Now we’re going to try something. One last thing. I want everybody to breathe in. Hold. Breathe out. All right. Okay. Good. You can take a seat. Thank you for cooperating. Everyone feeling good? Excellent. (….) That’s a low bar, guys. Come on. Okay. Great. I’ll spend 10 seconds introducing myself. Again, my name is Anj. (.) I’m the founder of a firm called AMP PVC. We’re a public benefit corporation. I’m a visiting scientist in the physics department at Stanford, which is where I spend most of my days researching and evaluating the capabilities of frontier models at physics, chemistry, and scientific reasoning, which is increasingly an important part of what these models are being used for. And we care a lot about the safety and security of these models. So that’s what I do outside of my activities at AMP, where we invest in frontier AI labs. We help start them, and we sometimes help them get access to compute. (.) I teach a class at Stanford called CS 153, frontier systems. And I think I’m not as good of a professor as Don. But one of the things we learned in that class, which we’ve taught over the last few years, is that interaction is pretty important with the students. And so in case over the next few sessions, you have any questions that you’d like for us to address during the panel later, I’m going to try an experiment. Usually we have like a tool called Discord we use with our students, where they can put questions in, and then we have them vote on it, and then we take them one at a time. We don’t have that today, so we’re going to use Twitter instead. In case you have any questions that are super burning, then find me on Twitter. My name is spelled Anjane, A-N-J-N-E-Y-M-I-D-H-A. If you search, you’ll find it. Just tweet at me. Between the keynotes I’ll keep an eye out on questions, and we’ll see if we can take some of them during the panel. Cool? All right. Our first session to kick us off is a pretty incredible creator by the name of Peter Steinberger. How many people have heard of Peter and OpenCloud? There we go. Okay, fine. No introduction needed. (.) We got our headliner. So Peter’s backstage. He’s going to come on. But to hype him up a little bit, Peter helped create, after having a successful career as a software entrepreneur, which he created one of the world’s most successful software tools called PSPDFKit in Austria, (.) with zero outside funding, and then he grew it into more than a billion devices. So he scaled it himself all the way to a nine-figure exit. He sold it, and then instead of retiring, decided he wanted to keep staying at the frontier of AI and started OpenClaw. (.) OpenClaw quickly became, as many of you know, one of the most successful GitHub projects in history. It’s an open source personal AI agent framework. And today, it reached over 200,000 stars on GitHub in a few short months, and is now stewarded by its own foundation. Earlier this year, he joined OpenAI, where he continues working on making software that has AI model capabilities underneath it more and more usable and useful to everyday people. Please join me in welcoming Peter. (……..)
S07: Don’t worry. I won’t make you do exercises. This week, my agent attended a meeting. It has its own account, its own web browser. (.) It listens to system audio. It went disguised as a human. Because in 2026, there are no doors for agents. Hi, I’m Peter. I built OpenClaw. And today, I’m wearing my claw hat, not my OpenAI one. I want to talk about the doors. My vision hasn’t changed since November. An agent that’s always on, always thinking. It hides complexity, makes my life easier. (..) And to be upfront, today, my agent still sometimes stops. You know, software is still hard. But I’ve lived a year in the middle of this, and I can tell you what I see. (..) Now, if I look in this room, you’re probably the furthest along in the world of AI and agents. My non-technical friends, they use AI, but they kind of use it like Google. You know, you type in a question into a box, the text comes back. (..) You know what that reminds me of? When television was new. They put those radio shows on the TV. You know, literally, you had people on TV reading the radio script. Because nobody, nobody really understood this new medium yet. (…) Every new medium starts by imitating the old one. The text box is AI’s radio on TV phase. Here we go. (..) And to be fair, it made AI usable for a billion people. (..) We’re humans. We’re more expressive than a text box. We see, we talk, we hear, we feel. That’s not the final form. (.) And we don’t get to the final form without trying a lot of things. I get my best ideas when I play this technology. (..) And look at, look at all the form factors we tried. (..) The whole agent wave started in the terminal for places. And I love the terminal, but most people never opened one. (.) Okay, the super app then. Why is it an app? (.) Or should it be an operating system? You know, agent OS. (.) It sounds great, but then you think about kernels are real, and there’s always going to be legacy app. So the industry sidebar, industry’s answer is sidebars. Every product has a sidebar with an agent in it. (.) I don’t want the Gemini sidebar in Gmail. The agent knows nothing about the rest of my system. And it’s not something I control. (..) Maybe the best agents are invisible. (.) Always in the background. (.) And I can invoke it anywhere. And it gives me the space to still be human. Because sometimes I just want to open a document and type. (…) Now, where I want to go with this. (..) I want it to be a little bit like Jarvis, honestly. (..) It shouldn’t matter which room I’m in. (..) It knows the device in my home and can use them. (.) It knows which computer I’m on and it can take over the screen when the job needs hands. (.) On my wrist, I have to watch through glasses and see what I see. (..) And we’re building the pieces. (..) We have browser use. (…) We have computer use. (.) We have canvas. (.) We have now these amazing voice models that can talk and listen at the same time. (.) And we have video. It’s not really real-time yet, but you know we’ll get there. (..) It should be like a person that’s sitting next to me. And I know I shouldn’t say person. (.) An agent really is a tool. But it feels like a very magical tool. (…..) So, back to that meeting from the beginning. Why the disguise? (.) Getting notes out is easy. That’s solved. There’s a thousand products that do that. (..) But actually being in the room is what’s hard. What has no door. And it’s not just the platforms. To give my agent a voice. The setup is actually an audio driver that pipes sawn through it. And the install instruction literally say you have to reboot. (..) Even the operating system has no doors for agents yet. (…..) So, here’s what I want my agent to do. We’re discussing something. Nobody knows the answer. My agent noticed. And it spins off a copy of itself while it stays in the conversation. Because the harness is just software in a way, right? And the agent can clone itself. It never has to choose between listening and working. And then the answer comes back. (.) It doesn’t interrupt. It waits for the moment that fits. It joins the conversation. I think that would be very natural. That’s what a colleague would do. (…..) Now, people ask me, is this talk about agents? Or is it about software? (..) How’s it not really the same thing? I built a lot of little tools. And the joke is I built most of the tools for my agent. Which means I built it for myself. (..) Because data gets more interesting when you combine it. Discord tells me which feature people love. (.) GitHub tells me where things break. (..) And GitHub has a good API. (.) But it’s built for humans. At agent scale, it breaks. (..) So I built a tool. And now my agent has full text search over, I looked this up, 110,000 issues and PRs locally. And, you know, Twitter. Twitter’s actually fair. You can export your archive. But it’s a mountain of json. So I built a tool for that, too. (..) And one app I really loved recently moved his keys somewhere my agents can follow. And what does that actually, what they actually achieved is that it just takes a little longer for my agent to get out the data. And it motivates me to build my own. (…) On Lex this year, I predicted that 80% of the apps will disappear. And maybe the truth is actually simpler. (.) Maybe people replace them. Someone asked Claude to build a CLI for his Samsung surround system. It rooted the phone. It installed smart things. It reverse engineered the app. (.) And it extracted every hidden control. He just wanted to turn the volume down. (…) Another guy pointed codecs at his studio lights. It told him which dev board to buy. And two days later, he has an open source controller. His words, I have no idea what it’s doing, but it’s doing a great job. (.) You know, closed source, open source, nothing stops a determined human with an agent. (..) And it gets recursive. (.) We build software with agents. We build agent with agents. We wrote the first agent by hand. And now agents build agents and software is agents. What are we even building? I don’t know. (..) What I do know, the small ones fit in a session. The big ones might get their own factory. (.) And the factory experience are starting to actually work. You know, compilers took a while too. (…..) Now, the question that everyone is polite enough to ask. An agent that’s always there, always listening. (.) Isn’t that creepy? Maybe, if it’s not yours. (.) You know, the thing holds my memories. I’m very European at heart. I want to know where my memories are. (.) So my agent runs on my hardware. Or maybe it runs in a box in the cloud that I have the keys. (.) And I can pick a model per task. You know, building software, I want the frontier. (.) Something personal, maybe this is better with a local model. That never leaves the room. (.) You have the hardware, you run open weights. (.) You don’t, you pick a provider you trust. It’s about choice. That’s why OpenClaw lives in a non-profit foundation. (..) We’re trying to do the right thing and give people that choice. (……) Now, in February, ClawCon happened in San Francisco. (.) A conference about my project. (.) And I didn’t organize it, you know. I was a guest. And the room was full of builders. (..) People were so excited. AI suddenly was something they could own, they could change. It felt like theirs. That’s the future I want. (.) It doesn’t need to be for everyone. (.) The people for it are going to love it. I’m one of them. (..) There are still no doors for agents. And that’s good because building doors is the fun part. (.) Thank you. (25 seconds pause)
S02: All right. Thank you so much, Pete. We’re going to keep things moving because we’re running a little bit behind. Our next speaker is Ryan Lopopolo, who is a member of technical staff at OpenAI. And you’re going to hear a lot more about him, I expect, in the next few years. Because you’ve certainly benefited from a lot of the technical projects he’s led at OpenAI. He was a founding engineer in OpenAI’s Seattle office and then led engineering on some of ChatGPT’s most important workplace products, including Code Interpreter, Record, and Connectors. And he’s the author of OpenAI’s Thinking on Harness Engineering, which, as many of you know, is critical from this morning’s sessions. The idea that humans should steer through goals, constraints, and feedback while the agents do the work. (..) Today, he leads a project at OpenAI called Dark Factory. And I suspect, based on what you’re going to hear, you will understand why I think you’re going to hear a lot more about his work in the next few months and years. Please join me in welcoming Ryan. Come on up. (…..)
S06: Hello, everyone. (…) A small update from me is, as of two weeks ago, I am at Google. (…) So, and, you know, small disclaimer here that these words are mine. I’m not speaking on behalf of Google here. But we can kind of get into this right now as I back up to the beginning. (..) What is Harness Engineering? Just to kind of set a base layer before we dive into the talk here. It is the mindset and understanding that even keeping the model and its containing harness constant, we are in what we call a capability overhang today. The models are far more capable and far more intelligent than they are able to side effect into the world today. They don’t know what local and global good looks like to their operators. They don’t have the context in order to have full autonomy within the organizations in which they are deployed. So it’s our role as the humans trying to steward these agents into the real world to give them the tools, context, guardrails, and coaching and trust necessary in order to fulfill the fullness of the job that we want them to do. (….) The reason I’m here talking to you today is because more than a year ago, I had the belief that the earliest versions of these reasoning models were capable of doing my full job. And I kind of put my money where my mouth was there by not doing my job back in June of last year. (.) I haven’t written any code since then and neither have anyone who are on my teams. It’s just not a permitted activity. The only thing we permit these humans to do is to get the agent to do the parts of their job that they need to do. And back in June, July of 2025, the models were much less capable. This was a much more challenging proposition. It was not the case that I could get the model to read Slack and respond to pages on my behalf. Just the level of tool usage and capability and complex orchestration was just not there. So in order to get that to work, kind of had to double click into the task and double click and double click and double click until we bottomed out on a thing the model could do. And as you kind of popped back up the stack there to reassemble those capabilities, you had permanently accrued value to the agent’s ability to reliably and safely side effect into the world. And as you get teams of people collaborating on a single agent in this way, you really get the best parts of everyone. And this fundamentally is what I mean when I say the way we build software has changed. We can take an agent with any level of capability and solve problems with the way it executes by writing code. And code is free to produce now. The models spike very highly in their ability to produce code. (..) In this world, though, where implementation is abundant, there are still a couple of scarce areas that require continued investment in order to make sure that the agent is able to cohere over long timelines, right? It is not the case today, even with as advanced the models are, that you can say, make me a billion dollar business and you will end up with something coherent at the other end. They are still struggling to operate vending machines, so I hear. (.) So that’s really part of what the human expertise in these human agent systems is meant to do, to constrain the areas of latent and physical space that we permit the machine to go in order to make sure we are tracking continuously in the right direction over time. What does it mean to continuously evolve an artifact, whether it is a code base or a Word document or a confluence-sized wiki of the organization’s knowledge? (…) So as we think about what it means for humans and agents to collaborate together, there are three near-term scarce resources that you kind of have to think about in terms of how you think about applying agents and human expertise to go do something. Human time is always going to be scarce. It’s actually pretty foundational to how we have built organizations up until this point. When you see platform teams or central dashboards, what you’re really trying to do is concentrate a scarce pool of human labor to produce high-leverage things that are able to empower an organization. And that same constraint holds true with agents. (.) One of the things I very often like to say is that when a human or a team of humans are interacting with agents, they must be incredibly ruthless by tracking their time, identifying what it is that they find themselves doing, whether it’s going back and forth one-on-one with an agent to put a plan together, whether it’s to review code, whether it’s to reject slop, they need to identify what they are doing and then figure out ways to make it so they don’t do that. And this natural sort of self-improvement loop will increase the autonomy of the agent part of the system to prevent more and more complex, more and more autonomous, and more and more parallel work over time. (..) On the agent side of things, model context window is like a foundational limitation of what it means to put together a model. Context windows may increase in size, auto-compaction may get better, but coherence over long-horizon work with a single trajectory is still something that you are going to have to keep in mind over time. And both human and model attention are going to be constrained. This is kind of a necessary fact of the world, which means the way we structure tasks and work has to take into account the ability for these humans and models to focus and avoid scattering their attention across many, many competing concerns. I often have had the experience where if I find I need to intervene more than three times with an agent, I’m probably going to have a bad time. Sometimes I will still want to see how I have a bad time, but that’s ultimately to learn where the agent has failed, to take into account the right things that I need it to do, so I can kind of back-propagate, reflect that back into the environment that I provisioned for the agent so that the next time I re-roll the task, it’s able to pull the right bits of context just when it needs it to avoid scattering its attention. We want focus in the ability to execute on our goals. (…) So we kind of keep coming back to, like, how to build good organizations when we talk about building good parallel autonomous agentic systems, because a lot of the same principles on how we empower humans in a complex organization apply to empowering agents as well. (.) If you think about what it means to onboard people to your team, you’re hiring generally capable humans, but they don’t necessarily know what good looks like for you in this context. So surfacing the collections of non-functional requirements that go into making good local work for you is the name of the game of harness engineering. And curating context in the background and via tools to constrain the agent’s ability to work is what it means to go chase after a well-constructed agent harness. Thank you, folks. (24 seconds pause)
S02: Sweet. (.) Next up, we have somebody who has done a ton of work across the ecosystem on code. I think one of the longest, one of the folks who’s had an insight about where software engineering is going much earlier than most people I know and has really pioneered a number of the innovations and systems designs that have ended up becoming pretty commonplace today. We have Michele Katasta, who’s the president of Replit. How many folks have heard of Replit? Quick show of hands. Yes, okay, basically everybody. Okay. Great. So you don’t have to reintroduce the product or anything like that. As a quick background, his entire career has been at the intersection of research and product, which is increasingly becoming important. He did his PhD at EPFL, taught in research at Stanford, where he pioneered several of the transformer architectures for source code before that was cool, and then led applied research at Google X, contributing to the coding capabilities of Palm, one of the early LLMs put into production at Google. Today, he’s the president of Replit, where he architected and launched the Replit Agent, a product that drove revenue up by more than two orders of magnitude. So not two X, two orders of magnitude. Please welcome Michele. (..) Come on over. (…)
S03: Thank you. (….) Hi, everyone. (.) So today I want to tell you about continual learning, which is often a concept that you hear associated to model training. And it’s been sort of a holy grail that we’ve been chasing as a research field for such a long time. We don’t want AI systems that are static. We want AI systems that learn from the usage. It turns out, though, that a lot of companies are today using closed weights models. So when you don’t have access to weights, how do you make your AI system actually evolve? Well, it turns out there is another answer, which is exactly what we’ve been doing at Rapid for many months, and is putting your hands directly on the harness around the ecosystem around the agent that they’re running. So how do we do that? Well, as an industry, we’ve been relying on evaluations for so long, and benchmarks are amazing. I really appreciate people working on them. The output is very crystal clear. You run your eval, you get a number, and you decide if your harness changes have actually made progress or actually introduced a regression. But there is a lot of signal missing when you do that. Evals are, by definition, very narrow. They capture only a subset of capabilities that you are basically exposing your agent on. So there is something amazing that happens once you hit product market fit as a product. The amount of sheer usage that your platform receives is such that you’re sitting on a goldmine of data. And that’s a goldmine of data that oftentimes we immediately associate to immediately starting to do model training. But there is far more that can be done with that. It turns out that if you analyze the traces, you can learn what works, what doesn’t work, why users are annoyed by your agent, and many other things. So our approach has been creating two different pillars. Of course, relying on evaluations very much. But at the same time, continuously running A-B testing as well as analyzing the traces in real time to really understand from the data what is actually working correctly and what we should immediately improve. So the reason why I put so much emphasis on the continual learning aspect is also because the amount of signal that you collect, the more traffic you receive, the more basically the gap is in terms of orders of magnitude. So your evals are fixed in size. You know, we recently launched our own benchmark called ByBench. It’s a hand-to-hand by coding evaluation. And we’re having, you know, a few tons of applications. So we already know exactly how our agent will behave on them. What we experience instead on a daily basis in our production system is a series of long-tail events that we can’t really predict. And those long-tail events are actually golden because they tell you how users are pushing the boundaries of your product. And at the same time, are usually those behaviors that break what you intended to make work correctly. (..) I don’t know if many of you have run A-B tests in their career. You know, it depends on maybe your seniority in the industry. But they’re not the, you know, silver bullet that sometimes we are led to believe. In the sense that more often than not, A-B tests look like this. There is not a clear result out of them. You run an experiment. You decide to manipulate or harness in a specific way. And then certain metrics that you’re tracking are actually improving. And some others are dropping. What is going on here? Like, it could be that you’re maybe optimizing for costs. And then the capabilities of the agents are dropping. You might be optimizing for speed. And then the sentiment of the users is evolving. You’ll never get a crystal clear answer that allows you to immediately ship that change in product. (.) So it turns out that the right approach is to literally take all the production workload that you have. And at first, rather than analyzing every single trace, because as you can imagine, we have millions and millions of these every single day. And it will be pretty expensive and also slow. The first step is we actually cluster them. I’m talking about very basic machine learning techniques where you find the semantic relevance among these traces. And the vast majority of them will be discarded because they are sort of like intended behavior. But every now and then, every day, we see a few clusters popping up that really highlight some new tail behaviors that we never experience on the agent. (.) And we built a system after this clustering step that practically takes every single anomalous trace, runs it through our analytic system, which, as you can imagine, includes frontier models, LLMs, understand what went wrong, and then immediately generates a PR. Now, what you still have to do as an AI engineer, I want to give you like a picture of the world where the work that we do is actually still extremely relevant. I don’t think this is going to be automated in the next few months at the very least. (.) What’s happening here is that once you have this series of PRs and you try to apply them, you will be running a test at that point. And some of them will not be conclusive as the screenshot that I showed before. So as a person who leads like an AI team, your choice will be defining which kind of these changes should be actually going production versus which one should be waiting or should be completely dropped. So there is still a level of human decision process, even though the vast majority of the changes that we produce for our hardness are actually completely generated by our AI system. And this is something that I started to talk about only a few months ago, because even though in principle this kind of pipeline is something that we could have built long ago, in practice, frontier models have become extremely good analyzing traces only in the last six months or so. So this same revolution that you experience as software engineers, where a lot of the code is actually written today by agents, we are experiencing as well on this side of the world, you know, on the production workloads. (.) I want to give you a real example so that instead of keeping this very abstract, you can understand what kind of problems these can catch for us. And rapidly every single day we spawn hundreds of thousands of virtual machines for our users. We do that completely transparently. And we had a long tail bug where at times it took longer for our virtual machine to be fully booted than for our agent harness to be ready to go. As you know, agents are very eager to debug problems. So the few times that it was happening, our agents started to spin the wheels and try to realize why it wasn’t able to execute code, why it wasn’t able to run certain tools. (.) Now, agents are fundamentally non-deterministic, which means every single trace didn’t showcase the same type of errors, because the agent decided to use different type of debugging strategies. But all of them had in common the idea that the virtual machine wasn’t booting fast enough. We would have never spotted this just by analyzing the logs manually. It would have never shown in our data log dashboard, for example, because it was a long tail error. But after we clustered everything, realized that this was happening often enough, and our system immediately generated the PR and fixed the problem on the spot. (.) So the takeaway that I want you to have today is that stop thinking about evaluation as just the last check before shipping. It’s not like a Boolean flag that tells you I should be shipping my new PR or not. Rather think of them as an engine that helps you to ship every single day a better agent. Thank you, everyone. (14 seconds pause)
S02: All right. Our last keynote is with Alex Gravely. Many of you have definitely used work that Alex has done. He was the co-creator of GitHub Copilot, one of the first agents to actually take AI coding into production, and then worked on Perplexity’s computer product, after which he recently started his own new company called Flying Objects Alex Gravely, that’s working on Omniscient Agents, that’s working on Omniscient Agents, a new category that I’m very excited for you to hear more about from him. Come on over, Alex. (…)
S09: Alex Gravely- Hi. Does this work? Okay, great. I’m Alex Gravely. I was lead on GitHub Copilot and Perplexity Computer, and I just started a new company called Flying Object to focus on some of the topics we’re going to talk about today. (.) So, the topic is Omniscient Agents. So, traditionally, Agents have been sort of limited by what they can see and access, and we think that this necessitates having a human in the loop. And so, we want to get to a point where the human is less in the loop or can farm off different pieces of their loop to Agents. And so, how do we do that? And that’s the discussion here. Alex Gravely- So, today’s Agents are a pile of primitives. (..) If you’re making an Agent from scratch, you basically have to implement every single one of these. You need an execution layer, either you’re running locally or in a sandbox. It’s going to have a bunch of commands for playing with files, running commands, running CLIs. (.) You’re going to sort of fine tune a bunch of asset creation. (.) That might be documents, websites, PRs or another form of asset that we’re kind of trying to make these Agents do a good job on. In order to do that, we need sort of these orchestration primitives that have now become fairly standard. You know, the details don’t matter as much as sort of what they do. (.) So, it’s useful to have a sub-agent that has different context from the parent. It’s different. It’s useful to have skills that can fill in gaps that the weights don’t express in the way that you want. It’s useful to have memories so that the agent can learn from the past. (.) It’s useful to schedule tasks, so things that need to recur or that might run on a schedule. And now we’ve got loops which are sort of a goal-directed running. So, you know, we’ve got individual tasks which then compress together into a loop, and the loop doesn’t end until the goal is achieved. And then live data is sort of, you know, all these things. You need web search, browser control, computer control, MCPs, APIs, all this stuff. This is where Agents are today. (.) If we look at, you know, what’s common here is that you’re still controlling these Agents. So, the agent doesn’t quite know what to do, and its suggestions for what to do next are often bad. And so, a human is in the loop to direct the agents. (..) And so, the bottleneck becomes your attention as an agent engineer, or an agent using engineer, I should say. You’re managing the loops. You’re running a bunch of things in parallel, things that are going wrong or going well. You have a sense of, you can kind of tell what’s going on. You can stop things, you can fork things, you can restart things. All this kind of stuff. But the way that we do this today is actually quite different than the way that we did it maybe a year ago. (.) And so, what we’re actually doing is walking up this kind of complexity hierarchy. (..) The way in which Agents self-direct now is different from it was before. (.) We started with just tools. We trusted the agent to call the right tool. (.) Then we started trusting the agent to make commits. (.) Then we started trusting the agent to make entire PRs. Where we are now is trusting the agent to make multiple PRs in order to accomplish these kind of looped goals. (.) And I think the future is something next level up is features. (.) Features might involve running an experiment, looking at live data, checking for exceptions, checking for user sentiment. (…) Segregating the people that are exposed to this feature to see if it has effect on your primary objectives, which are retention, say, or revenue, or whatever. (..) Above that, we start to get into projects. So, this is like collections of features that might serve a need in your product. (.) And I think at the top there, you know, it becomes whole products. When, you know, a self-directed agent, you can tell it to make a, describe at a very high level the product that you want to make. And that it can then go make that, deploy it, iterate on it, figure out the features that it needs to make. Maybe, maybe, which don’t exist anywhere else, try a few different things, and sort of self-direct in that way. So, how do we enable this kind of, like, increasing complexity, increasing scope? (.) Because each one of these, each level up kind of required either new model iterations or much more agent harness complexity in order to accomplish it. And so, we think that the commonality here are kind of like two axes. The axes that drive scope is one of insight and control. (.) So, insight is sort of what data the agent has access to, and what form it has access in it, so that it can derive insights. (..) Sorry, I’m out of time. (….) And control is its ability to operate on that data and live systems. So, what we’re going, what we’re doing is sort of moving from seeing the code to seeing the entire business. And that way we can start to figure out what are the actual business objectives? What are the products that should exist or the features that should exist in order to accomplish those objectives? Right now that’s, that’s up to people to sort of figure out and take guesses at, but what we want is for AI to be able to do that. And likewise, part of that process, part of the product development cycle is running experiments, deploying changes, monitoring live systems, scaling those systems as needed. And so, both these things together become sort of the two axes at which agents are growing their scope. (..) So, there’s a bunch of open problems with this, processing lots of data. (.) Generally, you want to ingest and index everything that happens inside of a business. (..) You want to be able to compress that knowledge into something usable by your agent. You want to have triggers. There’s, if you’re going to be running lots of experiments at the same time, potentially without human in the loop, you want to be able to have the agents aware of each other and coordinating effectively. (.) And a bunch of other things, you know. To the hardest point, which is, you know, sometimes you don’t even know the objective function you’re optimizing. (.) And so, these are the open problems in order to make what we think is the future of self-directed agents. Yeah. (.) So, what we want to do is enclose the entire system, capture everything. We want to be able to give the macro context to agents so that they can operate with full awareness. (.) And give the right primitives to those agents so that they can deploy changes, monitor those changes, (.) and scale those changes to accomplish those business goals. (.) It’s going to be a process as we walk up this complexity stack. But I think that’s the direction we’re headed. So, yeah. If you want to try out early versions of what we’re working on, there’s a wait list on flying object. (..) Otherwise, I’m Alex Gravely on Twitter. Would love to talk. Thanks. (……)
S02: That was great. Thanks, Alex. All right. You should… We’re going to transition to a panel now. And we’re going to do it lightning fast. So, please grab a seat. Make yourself comfy. We’re running a little bit behind. And so, we’re going to do… (..) We have 14 minutes. So, we’re going to try to do as dense of a pact essentially like… Woo! Yeah. I think we can do this. I like the energy. Good. Okay, great. (.) Can we get one round of applause for all of them, please, before we kick things off? (..) Okay, excellent. So, we don’t have time for all the warm-up questions I have, but we’re going to start off with one. The name of the panel, right, was the path to enlightenment. Why is that? A few years ago, I had the opportunity to get a call from some friends who were running research at OpenAI. And they said, Ange, we’ve trained a little model called GPT-3. And we’d like to leave and start a little startup called Anthropic. And so, I had the chance to come on as an early investor and help them out. And when I got my hands on Claude 2, which was a checkpoint that was not public, I started coding with it. Because that was the main goal, right? Make a great AI coding agent. And between Claude 2 and 3 and so on, I remember distinctly a moment where I was overwhelmed when I realized the capabilities of the model at coding. And there was, I found it very difficult to structure my work on how, as a programmer and engineer, like where do you start when you have this incredible, almost like a bazooka that can roll out any kind of software you want. And one of the things that’s difficult to start getting to the plateau of enlightenment that these guys are all at, and many of you are probably at now, where you can start building useful systems, is you often have to come up with a personal system, a mental framework for how to wrestle that overwhelming feeling into productivity. Right? So why don’t we just start there? Because what I found is everybody has, we’re so early in the space that everyone has their own tips and tricks and tools and techniques that they’ve converged on to harness the, no pun intended, the capabilities of these models as an individual software engineer or a developer and then produce something useful for the world. And so if you guys could just roll back to the first few moments in time when, if you agree, first of all, let me present this as a hypothesis. Do you agree with this observation? And if not, first, if you believe, if you agree with it, it’d be helpful to share how you wrestled with that. When was the first moment you were like, oh, okay, I need to, it’s like a muscle, I need a muscle to focus these capabilities in a way that’s productive. Should we start there? Yeah? That was great. And feel free to challenge the assumption as well. Go ahead, Alex.
S09: Oh, let’s see. So, I mean, there’s definitely like a, there’s like a releasing of control is kind of like a key aspect here, right? Like, we used to write every line and then we sort of started auto completing some portion of the lines. And then we started, you know, having AI compose commits and then entire PRs and now entire bug fixes. And so, like, along the way, there’s always like this desire to understand exactly what’s going on and make sure that it’s correct. (..) But actually, the more important and more useful thing is to, you know, find the system that can find the problems that you’re actually trying to look for so that you don’t have to look for them. So, a good CI system, good deploy infrastructure, all this kind of stuff makes it so you can give up control in a safe way. (.) And then you can kind of let these models rip. (.) That’s my, that’s my approach. (.) Eval’s also a huge part. (..)
S02: And we don’t need to go in order, but if you got one, Ryan, go ahead.
S06: Haven’t written a single eval in my life. Would love to keep it that way. (..) I think it’s pretty necessary to have a sort of firm belief of how you want to build this system. in order to chase after it well, which we can do because there’s an infinity of software available to us now.
S02: And if you could share an example of such a belief system.
S06: Yeah. So, my belief is that the machine is as capable as I am. And, you know, I am a software engineer. I am an employee in a company sort of thing. So, I want to be able to prompt the thing as lazily as you should be able to prompt me or any of the other principal engineers that I work with, right? And, you know, the scope of what such a person can achieve is quite large. So, I have always tried to curate the environment around these things such that I can give increasingly ambiguous, increasingly poor, increasingly contradictory information in order to get good outcomes out of it. And, you know, that has kind of required treating code as this abundant and disposable construct that we can kind of observe how the agent goes over a horizon. (.) Learn what context it should have had or did not have or was confused by from the produced dense artifact at the end. A PR, a word doc, whatever. Learn what decisions it made that were bad. And then try and be creative to put things in its environment, tools, tests, context, review agents, such that we steer it away from the bad choices it would have made in the previous version of the environment. And, you know, basically, the role of the humans in the system is to extract all of those hidden choices and provide them to the agent in sources that are amenable to in-context learning.
S07: Yeah, that’s the best way, right? It’s our our job is to help the agent to do their best work. Like I was literally backstage and I was reading the PRs that my agents were landing while I was doing my talk. Because by now I have such high confidence that we call it loop set up leads to code that works and is well tested. Like when those tools first came out, they could do things and like when they got it right, I got excited, but it was so hard. And now I revisit those all the projects and see just how good the agents have gotten. And we we kind of moved up the ladder right now. I tell my agent to try edge to like maintain your own agents to do everything and those agents have all the capabilities now to not just write the code, but also review the code and then run the code and then look at the output of the code. And then and then maybe has a several agents that like will further refine what comes out and you you as you my job is basically to give those agents all those things and to push the agent to work harder. Like my the average run for whatever prompt I give is now 510 sometimes 20 hours where it used to be like half an hour. And that’s fine because it’s just a more in parallel. But I know that at the end the chances that it did what I wanted a much higher. That’s why I’m so comfortable now in in letting them do things and I just pushing them to the repo because I spent so much time on thinking about the pipeline. (.)
S06: It’s almost as if not all the choices that go into a job well done are consequential. (.) And if they’re not consequential, you almost don’t care about them. And our job in empowering the agents is to figure out which mistakes are consequential and make them impossible. (..)
S02: Michael was anything you wanted to add to that?
S03: Yeah, I would say this point my north star is to deprecate prompting as much as possible. It was a necessary evil for for as long as we started to use LLMs. But especially for the type of product we build, which is for non-technical users. Prompting is the biggest foot gun that you can expose to them because they didn’t exactly how to specify what they want. So overall, I had this realization when I saw the first models that actually had some good LL training for coding. And this was back in the days of like early 2025. So the Sonnet 3.5, the GPT-5 family and so forth. The money we started to see them actually coding and debugging at the same time. Then you extrapolate from that behavior and you see where we are today. And imagine what we’re going to be like in, you know, one year from now. At this point, everything that Peter was saying is going to become not only for technical users, but I think for everyone. (..)
S02: I have several questions I want to follow up with each of you on. But we promised the audience that if they had questions, we’d prioritize them. And we only have four minutes left. So I’m going to start with one that I think is relevant to something you guys all said. I think this is from Amir. (.) Curious for your thoughts about security given the hugging face incident. But more specifically, personal agents already do things for their users that the users did not expect or intend. Are you guys rethinking anything around architecture or harness? Now, let me try to simplify that question a little bit. Because, you know, you talked about wanting to end prompting. You talked about, both of you talked about narrowing down the scope of what, like, what are the important decisions to oversee? (..) And maybe, Pete, since, you know, you’re opening up, you can talk about the hugging face incident a little bit. And how does that change or update the heuristic for what those decisions are when they’re merchant capabilities like the one that were exposed last week, you know, during the incident?
S07: I feel almost like I need my opening I had now. Because that’s a lot of what I do here. How can we run agents that are always thinking and where we also feel comfortable that you’re doing the right thing? And in many ways, it’s just like setting up another system where maybe you separate where the agent runs and where it can execute things. Maybe you have yet another agent that oversees the agent. There’s so many levers how you can, like, put in more fine-grained control and, like, more oversight to build a system that if, for whatever reason, the agent would derail, it’s immediately catched. (.)
S02: But isn’t the tension that, in this situation, it was what you don’t know to oversee, right? Like, unless I’m misunderstanding. One of the problems with elicitation overhang and so on that you guys have all talked about today is it’s what you don’t know that you don’t know that messes you up. Are there any tips and techniques or tricks that you guys have developed to decrease the likelihood that that happens? Or if you had something else to say, Ryan, you should feel free to add.
S06: I just think that is kind of like a more systemic problem rather than, like, the individual point thing. I obviously don’t have any context around this. I haven’t been exposed to it. But it’s kind of my belief here that, like, security programs in organizations are these things that have historically relied on human process controls. Right. In order to achieve the bulk of their outcomes and are comparatively under-invested in technical controls. And, you know, as helpful these assistants are, they do not necessarily conform to social norms. You can see this in how they accomplish their tasks in interesting ways sort of thing. So, you know, kind of how I talk about, in a lot of my work, shifting enforcement further to the left. (.) Increasingly proliferating technical controls in these domains is kind of necessary.
S02: Got it. Well, to wrap the panel, I think I’m going to ask each of you to answer one question, which is, (.) if you were starting today for the first time, or when a new model checkpoint comes out, let’s start with that. Let’s say, you know, there’s a new codex release or a new cloud release. What is the first prompt that you guys use to try and figure out the capability frontier of that model, specifically when it comes to software engineering? (..) And you can answer the question any way you want. If it’s not a prompt, it could be a system. What do you do to figure out what new superpowers you have when a new checkpoint comes out in the world of software encoding?
S09: I don’t know. I don’t do anything like that. (.) I just kind of assume that, like, the model providers are sort of learning from all of the ways in which their models are being used, and then kind of distilling that down into some kind of, like, generalized knowledge, which then works in all the different domains that they’ve been exposed to since the new release. So, if you want to figure out what the new model is good at, just look at what there’s 20 variants of in the last three months, because it’ll be good at all that stuff.
S06: I think the models are not the best at being self-aware of their own capabilities. So, this is more a what do I and my teammates do, which is to try and throw away every prior we have established around what the models can and cannot do, and start with your grandest ambition possible, and see where it fails again. And you’ll necessarily tree shake the tools and context you have, and maybe rewrite or throw some of it away. But you have to start from grand ambition at every new snapshot. (.)
S07: I use them for orchestration, and then I level up the number of parallel things the model has to juggle. And, like, 64 sub-agents were not a thing that was possible, and which now is possible. (….)
S02: Okay, I have so many follow-up questions that, Michele, you should go. I ship it in production. You just YOLO. (.) I don’t know if everybody should do that, but that’s one way to approach it. Cool.
