Video: AI agent sprawl: why you’ve lost track of your agents and how to regain visibility | Duration: 3616s | Summary: AI agent sprawl: why you’ve lost track of your agents and how to regain visibility | Chapters: Welcome and Introduction (0.175s), Agent Inventory Challenges (331.315s), Agent Management Gap (876.235s), Agent Inventory Management (1078.585s), Agent Building Demo (1507.795s), Agent Management Platform (2017.42s), Q&A Session (2877.415s), AI Security Risks (3262.16s), Closing Remarks (3581.455s)
Transcript for "AI agent sprawl: why you’ve lost track of your agents and how to regain visibility":
There we go. Alright. Good morning. Good afternoon. Good night. I guess if you're crazy and you're joining us and it's, weird hours, but, needless to say, welcome to everybody. Excited to have you here today, to discuss one of the biggest topics, I think, out there today. And, like, how do you manage agents, how do you manage inventory and risk and all the different things that come with those. So really excited to talk through this today. Let me get the screen up. So as we dive in, a quick bit of housekeeping. Of course, if you've never been to a webinar before, there is a chat function and there are q and a's. We will absolutely do our best while we're going through this to keep an eye and say hi in the chat, to answer questions. If you do have any questions, throw them in there. We've definitely got some time here at the end for Chris and I to tick through those, but we'll also do our best to keep an eye on them as we're going, and just hit them, you know, sort of in thread while we're going. So we will do do our best to do it. We're gonna set the stage a little bit. We've got a little bit of a product demo as we get into it. But, you know, sort of framing this conversation here today, you know, so you guys know where we are coming from. I will my name's Connor Jensen. I'm the field chief data officer here at Dataiku, based in, you can't tell from my windowless basement here, but, where the weather is actually nice in the Midwest at the moment outside of Chicago. Chris, do you wanna introduce yourself? Yeah. Pleased to be here. My name is Chris. I lead the Gen AI fuel specialist, based out of Amsterdam. So, really excited for today. Awesome. So if you aren't familiar with Dataiku, five seconds sort of, like, introduction to the company and to what we do. You know, our focus is really being on the platform for AI success, which requires all of these ingredients that we've seen. Platform is 13 years old. Both Chris and I have been here for quite some time. I've been here for a little over seven and using the platform for ten. I was actually a customer before I came here. Chris has been here almost as long as I have. So, we've been doing this for a minute. Around this sort of evolving space of data into data science, into machine learning, into generative AI, and now agentic stuff. And what we've seen throughout this journey, that we as a company and, obviously, you know, we as individuals have been on is it takes the combination of the right people, using the tools, using the data, using, you know, their expertise in being able to apply it to technology with the appropriate orchestration. If you can't get the tools and the data into the hands of the users, doesn't do anything for you. And then one of the big bogeymen that's out there that tends to sort of throw people off is the governance. Governance can sometimes trip us up and can sometimes help us go faster. But without all of these pieces together, we really see sort of, you know, different failure points in this AI world. Dataiku from a platform perspective is really, centered around those three different capabilities of enabling all of your people regardless of their different skill sets. So your, tech folks, your engineers, your data scientists, your AI, as well as your business people, marketing analysts, financial, folks, different, you know, functions there, bringing everybody into one platform to be able to work with the data, to be able to work with not just agentic AI, which we're going to speak a lot about here today, but the analytics, the modeling, the data science, and machine learning that goes under that, and really this whole ecosystem that together is what actually provides the power for your generative AI in your agentic systems that you're building. But doing that within whatever the enterprise architecture you have is, depending on where you are at. So that's the core focus of Dataiku. You'll see a little bit of the tools we're talking about today. And, again, if you don't have familiarity with it, you guys can certainly, you know, throw questions up. We'll, you know, answer them as we can. And then, you know, obviously, we're happy to answer other questions about that. But from a from a flow here today as we're talking about this idea of agentic sprawl, how are we going to think through this? The first piece is really setting the stage for us around, like, why this agentic world is really breaking sort of our notions of inventory and management of what we're doing from a sort of analytics and AI, perspective. Secondly is what does defensible inventory capture? Really starting to think through inventory is not just simply a list of names. What does it actually take to inventory properly what you're building from an agentic perspective? And then we will talk through what the product, enables for our customers to do that today. And then, as I said, we have plenty of time baked in to talk through any questions, we have to give some examples of things we've seen from our our customers and the like. So why is this a problem? Hopefully hopefully, most of you already have a notion of sort of what this problem looks like, but let's talk through sort of what we are seeing and what we're experiencing from our customers as they're really trying to move into this world of agentic AI. The trickiest starting point here is literally just knowing what the agents in your organization are. That it seems kind of funny to sort of say that as a as a challenge, but it is absolutely something that we are experiencing ourselves, that we're seeing in all of our customers because you have the the idea and the spread of agents. One, agents are easy to build. You know, one of the products that that both Chris and I have been talking about a lot lately is Dataiku's Co Build, which is, you know, built on top of things like Cloud Code or Genie or Cocoa or these other sort of prompt based building tools out there. But, really, over the last couple of years, there's been lots of questions and lots to talk about, like, how do you build agents? And I don't think that's actually the interesting challenge. I think building agents is a challenge and, you know, certainly something that organizations and people are learning to do. But building agents is going to be the easiest part of the puzzle. The hardest part is just figuring out what are all of those agents, what are they doing, who's managing the event, etcetera. You have every software company in the world now is, of course, an AI company. Every piece of software that you've ever bought now has embedded agents, embedded AI capabilities you don't even necessarily know are happening, under the hood. You have all of your different devices of, you know, are people using, you know, your organizational AI? Are you can you track them across, you know, where it's happening on a mobile device, on a laptop, in server settings, different things like that? You also have the contractors. Right? So these are all things that we're used to tracking, that we know what our apps are. We know who all the devices are. We know what contractors, but agents are all over the place. They're embedded in these SaaS apps. They're coming embedded in your devices. Perhaps contractors are using some different ones. But every team that has access to one of these things can be building and deploying its own agents, perhaps. New software you buy may have agents running in the background you don't know. This becomes a humongous risk across all different things because it's a very amorphous thing that we're trying to track. You know, even necessarily defining an agent actually can be a little bit of a challenge in and of itself. So this is it's a it's a unique challenge. And, you know, this if we're looking at it from an IT lens, you know, certainly for as an IT perspective, I was actually at a lunch with some some CISOs last week, and we were talking a lot about, you know, this this sort of challenge and and keeping on top of how they're approaching this. But from a security perspective, this is kind of a nightmare in and of itself because you now have these sort of nonhuman actors in your system doing stuff that you can't keep track of if you don't know they exist. And this is where this inventory question, I think you know, while inventory may seem simple and trite, I think the thing we're going to come back to over and over again in this is that you cannot manage, you cannot appropriately risk, you cannot securitize things that you don't know exist. At the heart of being able to manage and govern your agentic sprawl is simply knowing what you have. That's the first step, and it's already hard to do just from that. So when we think about this idea of inventorying, of keeping track of all this stuff, you know, typically, especially from a data perspective, we're able to really know what's being created. What are the data warehouses that are out there? What are the dashboards that are out there? What are the machine learning or data science models that are out there? What are the rags? Right? Even as we get into the generative AI perspective, you know, what are the different rag applications that we've built? What are the different gen AI models that we have provided? We have a, you know, a clear sense of those when we're able to control those from a central place. When we think about agents, though, it really starts to change that dynamic because part of the promise of agents is that anybody can build them. So any team that's out there that has access to our AI tools can build and ship agents and should be building and shipping agents. But are they doing so within your sort of IT typical SDLC perspective? Maybe, maybe not. If you're a product in a tech company, you might have that muscle. But outside of a tech organization or a tech company, it's very rare that all these individual teams across your company have that that experience, have that practice. And, honestly, even within a tech company, it's also not a given. Right? You know, to assume that everybody, you know, in your marketing team or your finance team or even in, you know, perhaps your sales teams know how to do this. And so you have enabled or are hopefully enabling people all over your organization to ship these agents, but now that means you have multiple sources that they're coming out from. Agents can hold their own credentials. This is a really interesting notion. You know, if you think about sort of what it means to be a sort of an account in in your Active Directory or your LDAP or, you know, how your organization manages accounts, Typically, you have two types of accounts. You have people, you have users, and you have service accounts. The service account is something that's, you know, built out there to run autonomously so that I can put something in production and if I'm there, you know, not there next week to run it, Chris can take care of it or or however we want that to happen. But those are typically there to just be part of automated systems. But with agents, a service account becomes something different because that agent is actually able to make decisions and change things. And so if you've got a service account that has API keys, that has access to everything, and you're putting an agent within it, they can now start to access tools, run models, change things without a human logging in, without being able to see that, oh, Connor went in and changed that. Right? When the service account goes and does things. One of the things that I've actually heard is companies actually starting to create a third type of account that is an agentic service account. That is a humongous change. Right? And, you know, twenty years of working in data, there have been two types of accounts for me when we're thinking from an IT perspective. It's a human error. It's a service account. And so introducing this third type of account is a is, I think, a huge step for us to start to think about that. But have you already taken that step? And then, again, have you now made it so that every single person who can deploy an agent knows how to deploy an agent as a service account, things like that? They also are persistent in autonomous ways. Right? You know, you can put an agent out there in the same way that a dashboard is sort of persistent. But a dashboard that's out there that nobody has looked at in five years is just an orphaned dashboard and it's not doing anything. Right? It's consuming resources and it's, you know, maybe annoying, but it's not doing anything. It's just sort of, like, out there as a ghost element. An autonomous agent that persists five years down the road that nobody knows is out there, that it's taking actions, that it's doing something, that's scary. Right? That actually starts to become something where if I built an agent that I deployed and, you know, I leave and IT didn't know that I had done that or Chris or whoever replaces me doesn't know who'd done that. How do you manage that you this thing that I created and put in place that nobody knows is there, especially if it's not attached to my user account? And one of the challenges, but also then one of the sort of flip sides of the agents need to be able to work across platforms, across systems to be effective, which means it's it's not an easy task. So it's sort of actually really challenging tasks right now is making it so that an agent can work across things. But, obviously, the MCP server, the idea of a two a, which I haven't seen too much in practice, but it's sort of you know, obviously, the the need is there and people are starting to work towards that, is that you can have a single workflow with multiple LLMs engaged, multiple data sources engaged, reaching out across different systems to access tools via MCP servers, etcetera. You know, you're now taking that that, you know, that unknown question of, is this you know, how is this agent? What does it do? How is it out there? This is really tricky. Like, one agent could actually have sub agents under it. Right? Actually, in fact, often has sub agents under it. So does that count as one agent? Does that count as five agents? Does that count as 20 agents if this agent's calling for it? So, again, this idea of it becomes really challenging to sort of think through all these questions of inventory and looking at this. The question at the heart of this of, like, how do I manage this? How do I wrap my arms as a CIO or as a CISO or, you know, even just sort of a business leader trying to understand what's happening? This has been the number one question that I have been asked by our customers, by, you know, potential customers, by friends and colleagues in the industry, you know, over the last, probably now, eighteen months, is how do I know how do I manage, understand what I have? And it's getting worse over that eighteen months, not better. The pace of the ability to deploy agents is simply increasing. The management aspect of it has not increased, at that same pace by by any means. You know, we sort of see like, I made the not such a joke joke about, you know, every software you've ever bought is now an AI software. But, you know, at the beginning of last year, while everything said it was AI, roughly 5% of, you know, actual enterprise software contained real AI. Again, real AI. Easier I I throw that around. I realize that that's not necessarily the easiest thing to define. But, you know, 5% of companies really had AI embedded in their software eighteen months ago. By the end of this year, it'll be 40% or more. Only 21% of organizations say that they maintain a real time inventory of active agents. We've definitely talked with, and I know organizations that have put these pieces in place. And you'd be hard pressed to get any of them to say that they have 100% confidence that even within these things they've created, that they actually have 100% of agents in there. But the fact that, you know, only one in five companies even goes to you know, is willing to say that they have anything in place is very indicative of sort of where we're at. And, you know, this idea of sort of how mature are we? You know, now that we're down to sort of 12%, only one in 10 companies really have something mature even though the vast majority of companies have a dedicated process for deploying or managing agents. So where is this gap coming from? The companies that say they really know what they're doing versus all the places that are sort of, you know, starting. So how do we sort of mature this? The this this gap will continue to grow. Agents are everywhere. Agents are easy to build. We have very low visibility into their performance, into behavior, obviously, nondetermic systems, all that sort of stuff. And it's very unclear who the owner of an agent is. Right? You know? So if an agent that was deployed is used in the wrong way or if I build it and deploy it and it does something, who's accountable? How do we do that? That's messy. That's very messy. And, again, if I don't even know what agents are out there, I can't even begin to ask some of these questions versus how do I actually create a trusted agent ecosystem? Do I have complete inventory and visibility of what's out there? How do I actually start evaluating the performance of an agent, not on uptime and SLAs and, you know, that it gave answers, but actually tying that evaluation to the business outcome that the agent is supposed to do, is supposed to solve. And then, finally, how do I actually start to govern and control all this? Because I can't govern what I don't know is there. So over to you, Chris. Talk about what inventory. looks like. Exactly. We can we can go to the first side. So, breaking that down. Right? So we've kind of discussed the core elements of of this challenge. And like Colin mentioned, it's a big challenge. But if we break it down into kind of fundamental principles, this is this is what we've, what we're what we're using as, as baseline internally. So each agent that you have fundamentally should have an owner. Right? Who actually owns, this agent. It's an accountable person. And oftentimes, that's an easy question to answer. The harder question becomes, if that person leaves, who is then the next person that owns this agent? And how do we make sure that to track that? Right? How do you make sure that these agents don't become ownerless, over time, let's say? Next, of course, is purpose. So what does this agent actually do, and who's who does it serve? Right? Who is who is its audience? Let's say, who is its consumer? Data access. So what does the actual agent have access to? And that will help with the final one, which we'll get to. Tools and permissioning. So is this a read only agent? Is this an agent that can actually write, and, affect production systems potentially? Does it have human in the loop, for those, write accesses? All of these things will collate into this final one, which is a risk there. Right? Based on that previous information, how how risky essentially is this agent, going to be. So if we go to the next slide, how does agent management kind of think about this? So we've got agent inventory. And, again, like we said, this is really the the foundational aspect, of of this challenge. So what we do is we scan and identify, the organization's agents across the different platforms, that, that those agents are are living in. This allows you to then establish an owner, and, of course, the map of purpose essentially. And we do this through automatic agent discovery. We do then automatic metadata enrichment, and we allow you to search, through all of those agents within within that specific inventory. We then go to the next slide. Once you kind of got those agents set up, we can start to think about monitoring. So at this point, you've got your agents. You've assigned owners. How do you actually know which agents are, are being used, which ones aren't being used? Is there a drift, at any point in time in terms of tool usage? I think for me, this is one of the most interesting concepts, actually. The idea of, not necessarily looking how the agent is answering, but what tools is it calling. And over time, this can really start to to to to show you interesting patterns in user behavior and what questions, they're asking asking the agent. For example, those questions could become more difficult questions, right, where the agent has to go through many more different sequences of tools. I also think tool ordering is also very interesting. Which tool does it call first, second, third? And this metadata, this information that lives on top of your your traces, highlighting that to you, and potentially, of course, feeding you alerts, over time as well. We also, of course, look into things like cost analysis, like cost per conversation, cost per message, all of these, very important factors, when it comes to monitoring. Next, and, last but certainly not least is risk management. Right? When we start thinking about these agents, we want to, make sure that we're managing that riskiness. Right? Different agents, again, like read only agents, post perhaps a a a a fairly small amount of risk, compared to, agents that can actually start to, look into and and adjust production systems. They can actually start to, I don't know, schedule maintenance, potentially approve, certain things internally even. And so we wanna make sure that you can, apply a risk framework, assess that risk first and foremost, apply that risk framework, and, of course, manage incidents, for these, different agents across your organization then as well. Alright. So next topic. So what can you actually do in Dataiku today? Dataiku, of course, the platform for AI success, and I hope that you're kind of gauging from these slides the the level of breadth that Dataiku, is offering, with with the platform. But, fundamentally, we allow you to go through that entire life cycle. Of course, agent management and inventory management of those agents is is an absolutely critical component. But Dataiku itself also allows you to, of course, build those agents. Right? Bringing in those subject matter experts, people that actually know what these agents should be doing within the business, and build these out, in either visual agents or structured visual agents, connect them to different data assets, different underlying tools, then, of course, orchestrating this. Right? Connecting this to the entire ecosystem of data within your organization, whether that's new modern legacy technologies like Snowflake and Databricks or, technologies, that you might have had for for, you know, five, ten, twenty years, in in some cases. And then, of course, overseeing that within within the platform. So the LMSH offering, an AI gateway on top of your, on top of your LMS connections. So if we go to the next slide, we can break this down in terms of, in terms of timeline. And I think this is really interesting, and I think it points to the philosophy that data goes have throughout this kind of, call it AI revolution, if you will, where each release, and suddenly in each year, we've actively been thinking around governance, first and foremost. And so that starts with the LLM mesh. Like I mentioned, this is the AI gateway for all LLM, interaction across data and across your organization potentially as well, where you can apply things like cost control at the LLM level and things like toxicity detection, various different guardrails that you may wanna apply. Then feeding into agents, and then how do you actually connect those agents, those ten, twenty, 30 different agents, and actually allow the business to start consuming those agents in a production environment. And that, of course, was the advent of, visual agents, structured visual agents, and agent connect. And just recently, we've released Cobalt. And and for me, this is one of the most exciting releases, that, that that I've been a part of in in the seven years that I've been here, where we're fundamentally shifting how, people are interacting with Dataiku. Right now, it's via an agentic chat interface, where you're using a chat interface to build out, an interpretable, reviewable workflow. Right? You're not generating thousands of lines of code. Instead, it's a very interpretable overview of the logic that the agent is actually creating on top of the data assets that at which you you've connected it into. And then finally, of course, agent management. Right? Coming very, very soon where we're we're releasing, a a platform that allows you to indeed look at these different agents, across your organization, not just inside of Dataiku, but also across all the other platforms, which you, which you have got connected into it, and enables you to do the things that we discussed in in those previous slides. So with that, I think it's time for a demo. So if you stop sharing, I think I should be able to take over here. Perfect. So I think that's coming through, alright. Alright. So, a short demo of, like, actually building these agents inside of Dataiku. In this case, we've got a factory reactor anomaly detection, project. It's actually a project that we've, we built, in conjunction, with, some of our customers. And so in this case, this is a very realistic project. And it might seem overwhelming to you at the start, but it actually isn't. Fundamentally, it's broken down into different zones at which you can see here. We won't dive into the detail of all of these different zones, but most Agentic projects are indeed comprised of all of these different assets. Right? You've got on the left hand side here, you've got, data preparation. Right? We're doing various different, data preparations. We're pulling in data from various different places, in this case, Snowflake. That can, of course, be any kind of, data that you've got sitting across your organization. Then we're building, machine learning models. Right? Real anomaly detection models, which we're building out in a visual way. And then finally, we're connecting all of those different systems into mesh into, agents, I should say. So we can see here these various different pink, agents, which we've, which we've connected into. Now recently, my philosophy around building agents has actually shifted quite a lot, and that's because I'm building a lot of agents. And, oftentimes, when you're building agents, it feels like you're kind of in the dark. Right? You you're you're adjusting the prompt, then you're chatting with it. You're seeing what works. Okay. That seemed to work. Then you go back. Maybe you just do a little bit more. All I found is that that's not really, a very scientific, way of of of building agents. And so instead, what what I've now started using is data user review functionality. And, again, like, I think this part of the platform is really, really exciting where, the first thing that I actually do is, like, sit down and discuss with my SMEs, with my stakeholders, what are the actual test cases that we've got? Like, what do we expect people to ask this agent? And this leads to really interesting concepts. Right? You've got this overarching idea for an agent, but there's also edge cases. Right? What if someone thinks that they can ask this to the agent? What should the agent behavior be then? Like, what's the intended purpose, of that agent? And So trying to gather this information is, is a challenge in and of itself, but it will pay dividends doing this at the start because it means that you're massively accelerating your your development cycle. You can optimize for something instead of, kind of navigating, in the dock. So we can see here the different test cases that, that, that we've gathered. We've got them reference answers. You don't have to have reference answers, but if you have, this is extremely powerful. Oftentimes, you can find this from, like, even Slack or email, or ticketing systems that you may have where people have been asking this question to, potentially other humans in the past. How have those humans answered that question? And then can you maybe port that into a test case for your agent? That's oftentimes kind of the workflow that we go through. And then expectations. So how do we expect the agent? Which tools do we expect the agent to actually use in this case, and throughout? We can then come over to settings. And under settings, we can actually start to apply different traits. So these are LLM as a judge traits, traits that we want the, the agent to actually, how we want the agent to behave and how we wanna measure, the agent. We can then come over to results, and I ran a small test there. So we can see one of these test cases that ran it actually failed. We can see the execution. So what actually happened, to this agent? We can see the test, the reference answer, the expectation, and we can see that this trade failed, the reference rate. So the agent just didn't consistently have a reference when it answered. So we weren't sure where that answer was actually coming from. We can also see the trajectory of the agent and then the underlying trades. Right? So you can actually see, for each trade, the LLM has to judge why did it pass, why did it fail. And then from here, you can, of course, involve your SMEs. Right? So in this case, an SME can come and provide reviews for each individual, agent run, so that you can use that as a reference point as well. And so oftentimes, what I'll do is I'll I'll I'll do a first set, of runs. I'll take a look myself and, usually, myself, I can make some adjustments. And those adjustments can be there's many parameters which you can start to optimize and adjust for. Right? There's there's obviously the prompt. There's the tool. There's the LLM, which grade of LLM are you gonna use? Are you maybe gonna bump it to to a to a a a a more powerful LLM so that it's better at instruction following? The data pool also has another set of parameters, which is the type of agent that you wanna choose. So do you wanna use a simple visual agent? I'm gonna show you what a simple visual agent looks like within the flow head. So we can see this maintenance agent, that we've got is connected to, to some underlying, structured and unstructured data, which we've connected to various different sets of tools. You can see this is a simple visual agent. By its very nature, it is very simple. Right? We have LLMs, and these LLMs can be any. Right? The LLM mesh is completely agnostic to the underlying, systems that you wanna use, the vertex models, anthropic models, OpenAI, oh, even some open source ones if, if you have maybe, models posted on premise, or, through your model providers. We then have instructions. So what do we want the agent to actually do? And then we've got those tools. And oftentimes, this is maybe the first thing that you build out because it is simple. Right? Then we can write down this prompt, and in in this case, the prompt isn't very complex. But oftentimes, what tends to happen is the agent doesn't behave in the right way, and so I add to the prompt. And then I go back, and then there's an edge case. And then, okay, I need to add that to the prompt as well. Quite quickly, what you realize is the prompt is actually becoming some form of workflow. Right? You're actually writing out, you know, if this, then that. Right? And, oftentimes, what happens is as these prompts grow, and I've seen some prompts that are, like, two, three pages long, when these prompts are growing, you need a larger LLM, a a more intelligent LLM to follow those instructions more consistently and more reliably. And that's certainly something that you can do, but there's obviously the cost drawback at that point. Right? Are you comfortable with every single time, a user asks a question that that is a significant cost expenditure, or not? And if the answer is no, and it it usually is, we can start to think about actually, breaking that prompt down. And so when you break a prompt down, you can go to a structured visual agent. So a structured visual agent, as you can see in this flow diagram, is essentially breaking that prompt down into, nondeterministic and deterministic steps. So if I come over to create a block, you can see these pink icons like agentic loop. Right? That's that simple visual agent. The LLM, the prompt, and the tools. A simple LLM request, so no agentic loop whatsoever. But then we've also got things like traditional routing, full loops, parallel execution, reflection, things like context compression as well, generating artifacts. So you can see there's a huge amount of blocks, logical blocks for you to choose from. And as you can see, that's exactly what I've done in this case. So, again, this is a maintenance, use case that we've got there. So the first thing that we wanna do when, when the user, asks a message is analyze what reactor user is talking about there. So we can see we've we've provided some instructions. And at that point, we know which reactor because the agent has identified. We can actually start deterministic new routing. Right? So if we can't identify a reactor, we go back to the user and we ask for clarification. If we can, we wanna move on to establishing a root cause. And then depending on that root cause, we wanna root it to a very various different, elements here. So quite quickly, you can start to see, like, okay. This logic that I've got within within this, agentic prompt, I can actually start breaking down, into, into a workflow like this. And what's really interesting is when you do that, you can actually start, potentially reducing the the models that you're using. In this case, you can see I'm using not even a reasoning model. I'm using something like GPT 4.1 mini, which is very cost effective. But oftentimes, there's more than enough than I need for these very simple tasks, that I'm asking for because I've decomposed, that structure completely. And so these are just some of the agents that you can build on Dataiku. Of course, you can build on code agents using LandGraph and your favorite, IDEs, as well. But from here, I'll hand it back over to Conor so, he can discuss, how we start to manage both these agents built inside of Dataiku, and outside of Dataiku. So I. will stop there. Alright. Thank you, Chris. I love seeing that. Again, you know, the I mentioned this a little bit sort of passing before was really feel like we were tackling two problems this year, maybe the last couple of years, but really, like, sort of, like, finally solving them. The first one is the the builder challenge. Right? Like, how how do we start finally start to fulfill this promise of, no. You don't have to go learn Python or lang chain, you know, or SQL or whatever to be able to work with data, to be able to work with AI. You know, that I think has been a a promise that has been growing over a while. You know, there are visual aspects of there are different things, but, you know, really, this sort of last couple of years has really seen that start to come to fruition through things like prompt based stuff in Cobuild or things like structured visual agents that that Chris showed here. But this solving the builder problem, creates the second problem of this inventory challenge of managing these things. And, you know, I'm personally super, you know, sort of most excited about this next piece piece that we're coming with, which is our agent management platform. Mentioned it before. This is Dataiku's newest, product suite. It is currently in early adoptions or beta with some of our our customers and going GA next quarter. So excited to see this really sort of, like, come out into the world. And, you know, hopefully, to really start to set the stage and help define for the industry what it really looks like to manage agents. So what does this allow us to do? The first and the most important question that I think we're trying to solve here is simply this inventorying question of, as a CIO or as a CISO or as a Head of AI who is reporting back to the board or somebody, how do I know what agents are in my organization and what are they doing across platforms. Now, across platforms is where everybody else has an answer until you get to that one of, No, I don't want to do this 15 different times in 15 different places. How do I look across platforms and know where this is at? So here you see, you know, our homepage, obviously, is a quick start of just simply, here's an overview of the agents that I have in my organization right now. But where does this come from? And so this this is where, you know, you really need to start because if we think back to the, we don't have to go back that far, but thinking about all the data governance or data catalog tools that have come along in the last few years, this idea of master data management and metadata management and creating data catalogs and stuff like that. Where does it always break down? It always breaks down at the manual effort that it takes to do these things. If somebody has to go in and register a table, has to go in and define the fields, has to go in and build that metadata, That will never happen. You can never have that sort of, you know, expectation. And that's where the agent world has been up until this point is for me to be able to manage agents, somebody has to have built and deployed it and then filled out a form or put it in my Google Sheet tracker or, you know, however we're managing these things versus doing it autonomously. And this is where we really sort of think through, how do you start to do this? So Getacoo in the back end, Christian went through all the LLM mesh and all the capabilities that we have there. We've always been an integrated platform that works across wherever you are on prem, in the cloud, whichever hyperscaler you're working with, wherever you have your data, and surfaced all of that. And we take that same sort of expertise and and approach to, we know, as cool as, you know, the things that Chris just showed on and how easy it is to build agents in Dataiku, that no one is gonna build all of their agents in Dataiku. I don't build all of my agents in Dataiku, and I don't expect that anybody ever will. And so that we have to build first from that perspective of you will be building agents and working with agents that come embedded in platforms like the ones you see on the screen here, Vertex, Snowflake, Databricks, Bedrock, Salesforce. Salesforce and Agentforce is a big one. Azure, Foundry. The list of the places that we're sort of continuing to work with is being driven by those beta customers, but we have a long list of these integrations that we've built and are continuing to build that we're working through to make sure that you can just simply plug that in. And that's the, you know, beauty of this is to say, okay. I want to add a new connection. I'm gonna come in here. I'm gonna select from these prebuilt connectors that we have. In addition to some of the things that are here, we're also working on how do you have connections for custom built ones. If you're managing your own, generative AI stack on prem, if you're doing it in your own, you know, your own, like, NVIDIA hardware or however you're building that, that there's the ability to, you know, work with via, like, an SDK to sort of register things like that there. But I'm gonna come in. I'm gonna select from these prebuilt connectors. I'm gonna say I wanna collect connect to Snowflake. I obviously will probably have multiple Snowflake instances, so I'm gonna call it. I'm gonna build, you know, a new one. I put that URL in there. And then once I hit connect for that, it's gonna ask me now I can say, how often do I do that? So once I've created that connection, it will go to those systems and pull everything in autonomously, but also on a recurring updated basis. This can then be driven by, of course, which platform are we talking about and how many people are working within it? Do I need this to run hourly? Is this a platform where I have a thousand people who are working in it that I really need to stay on top of? Is it something that is only being used by my AI team? And so, you know, we're gonna be working at their pace, and I can do it daily or weekly, etcetera. But I can turn these on and off. I can manage the frequency in different ways, and I can also go in and look at, you know, what's the history of those pulls. If there was a failure, why did that happen? Sometimes things happen with our codas. Sometimes we have agentic, you know, throttling, things like that. So I can very much keep on top of what's happening. And then I can look at, here's new agents that have come in that I need to make sure that I go and validate what's happening there. But we want to take that, you know, effort off of the people, off of your AI team, off of individual users to go in and register agents because your inventory will never be, you know, complete if you are requiring every person who's building an agent every time to go in there and manually register that. You know, speaking as a guilty offender, I'm terrible at managing stuff like that, as I would say most of us are. And so we just, you know, have to have that happen autonomously. And so working with all the biggest software companies out there to make sure that we're taking the connections that we already have into those systems and expanding it into managing those agents, which is in itself its own tricky task. As you talked about some of these various platforms up here, they don't necessarily have a single element internally that is an agent. And so, you know, for each one of these softwares that are going in there, we're working with them to define what are the different agentic systems that they have within to be able to show all those different things. And in some cases, you know, saying that we have one connection into something like Snowflake actually means five different connections into Snowflake to deal with the different systems under the hood there that do not. So once it comes in, now how do I sort of think about this? Obviously, I want to have, you know, that that big picture. What's out there? How many total agents do I have? How many of them are in production? How many of them are having issues or showing from an operational perspective? But also from a behavioral perspective. Right? So it's not just is the agent working, it's is the agent working properly? Those are two fundamentally different questions and incredibly important because I'd rather that an agent goes down and is not working than it is working or it is functioning, maybe, but it is not working properly and we're having, you know, erroneous behavior. So I need to be able to track those different things, and then I can actually put places things in place to do Drift and that. So let's dig into this inventory and actually sort of say what's here. So, obviously, I can go and I can search through what you know, how who's using these agents? Who's building them? What system are they in? Are they you know, if we have some governance on them, are they certified? And these are things that we are aware of said that this, yes, can be in production. This can be customer facing. This can be spread to different users, or we're working on it, or it's not certified at all, but we want to take whatever we have from a governance perspective, put that in place. We have a way to look at the risk level of these different things that are certified in in there. What is the risk level? How careful do I need to be with this agent? Who's building it, etcetera? If I go and dig into one of those, specific agents, I can really start to see first, I see the the monitoring, the operational perspective. Is this agent functioning? How often is it being called? How many tokens does it cost today? It's a lot of tokens. Hopefully, that's not true. What's the uptime? What's the response time? What do I see? What's my error rate? You know, these are sort of like almost standard IT ish, API metrics, the type of things that we've always wanted to know about the connections that we're creating. And I can see how that's going. I can look at stuff like average response and things like that. But I also need to be able to check into behavior, and this is where this really you know, the challenge of agents as opposed to just APIs really comes to the fore is that it can be functioning, it can be firing, but it can be giving the wrong answer. It could be drifting. And in this case, we see this. We have that there is a new topic. So this is being driven by the tools it's calling. Right? So as Krishna was talking about with that, structured visual agent flow or things like that, we're able to pull out from that agent what tools it has access to, what the flows of that agent are, and actually look at the behavior of it over time. And here we can see that over the last four quarters of this has been in production, we've seen a change in which topics are being called. But we also see that now this quarter, we have, 16% of our results are in an entirely new category. Right? So this then is a big drift. Maybe that's okay. Maybe we've added a new tool to the agent. Maybe, you know, there's something that's going on elsewhere in the company. And and so it doesn't tell me that there's necessarily a problem with the agent, but it does tell me that in the last quarter, the behavior of this agent has changed and I need to go look into that. Hopefully, we've started a new tool. And so now the fact that we have this new behavior is because we've updated the agent and we've made it better. But, you know, I can actually start to look at this and what I'm actually seeing is that, you know, I've got a router that's dropping things to where I'm having escalations happening. So this patch compliance may actually be behavior that's outside the intended, answers or the capabilities I built for this agent to date. And so this 16% of results that's happening may all be coming basically, it doesn't know what to do. It's an exception, and it's going back to an end user, which that is good in that if it's if the agent is not getting a question or is getting a question that doesn't have the ability to answer and it has, you know, a path to sort of send that off to the team to answer, great. That's the right sort of behavior that I wanted to happen, but I also the reason we have the agent is to hopefully take the load off of the team. So this is a problem that we need to go and look into and do some more of. Beyond simply that, though, I also wanna know the risk of this. You know, I'm not you know, this is a capacity planning agent. This is where when we talked about those, is this certified or not? These are some of the steps that there does actually become some manual stuff. So I can go in and create a risk assessment. This is how we can actually go through and certify, is this, you know, agent how does it sit? There's a questionnaire. It's gonna ask me, what is this doing? You know, this is is capacity planning. This is, you know, obviously, I'm looking at infrastructure stuff. This is not an autonomous decision making agent. This is input into my infra team to look at what they need to be able to do. So I can go through and say, okay, this is global. It's definitely, obviously, pulling stuff out of databases. It's probably gonna be looking at some external APIs in the file system. Does it work within customers? No. Good. Is it looking at, like, that? No. It's probably just really internal business data. Maybe we pull in some public stuff, but but it's not public only. So we'll say internal business data. No. It doesn't have health information. Does it take actions? No. What's the expected call volume? Moderate. You know? Well, no. This is input into human. So this helps then us define what the risks of this particular agent are. I'm not looking at any sensitive data on this. It's advisory. It's into mine or, actually, I should say, inputs into just human decision making, my infra team. But it does likely interact with other agents in it, and now I can go and say, okay. These are the suggested risks that I need to look at. So let's assume I'm happy with that. I go ahead and confirm and start the assessment, and this is now flagged this as a bronze one. This now is gonna end up in that pending bucket. Right? So I had, you know, my sort of, like, certified pending or not done. This is now my pending bucket. I've sent it off now to my risk management team to go through and validate these things and make sure that's there. So again, we're taking from, you know, we've centered this conversation here today around agentic sprawl. And the starting point I'm going to continue to try to come back to is, I can't manage or monitor or risk assess something that I don't know that it exists. So the first step is absolutely just purely identifying and doing so in an autonomous and comprehensive way across your organization what agents are then out there, and that then powers these other things of behavioral tracking, of monitoring, of risk assessment, and things like that. So, super excited again about the agent management platform. It's something that, again, is the number one question I've had from the organizations that we've talked to, especially, you know, sort of your, like, executives within IT or AI organizations is how we do this. So I'm really happy that this is coming and, you know, excited to sort of share this with you guys. I'm gonna go ahead and take this down. We've got about ten minutes left. I know there was a couple of questions early on around, sir, is will the, everything be shared? So, I'll do those. I can definitely answer one quick one. So, Fred, I'm I'm very impressed. I get a lot of questions about what state or country is up here on the wall here, but, properly recognized by Fred is that that is the Nurburgring Nordschleife. So, you know, golf flap on that one. I am a huge gear head. So, you know, among the various hobbies represented in the background here, but probably not as an you know? Well, I'm a nerd, but anyways. So what other questions? I see one there's the structured agent view is the most suitable option for building stateful agents. Yeah. Let me throw that, Chris. Like, what do you how do you wanna take that from? Yeah. Exactly. I'd say I'd say so. This is how I'm building, it's essentially a state driven agent. So, stage, for those of you who don't know, is essentially a way in which you can track, certain variables during agentic runs. So agents can interact with that state. They can update the state, pull from the state. And this also allows you to have different agentic runs pass over information to one another. So, if one agent updates the state, the other agent can then pull from that same state, if you let it, of course. And so structured visual agents allows you to do exactly that, work with, the state machines. You can, of course, if you want to and you're used to already, code line graph agents as well. Those work with, with state components. So you've got both options. No. Very, very cool. But, yes, great great question. The other question that I saw in there, from early on in in the conversation was the the definition of single source of truth, from Raymond. That one, I think, is a fun one, because I think that that's actually a much harder to answer question today than it was before. You know, the the core definition of single source of truth is obviously, you know, still the same of that within an organization, there is a single field that is, you know, the approved, managed, metadata, you know, master data fields for a specific, you know, piece of information. And that anytime you're trying to answer specific revenue questions, you know, headcount questions, whatever that that underlying question is, that it goes back to a single thing. With generative AI, that becomes really hard. It's actually one of the things that I've talked to a number of, you know, sort of startups or software vendors in this space creating back ends to help put in between sort of a a, an agentic system, whether that be sort of a a, like, a chat window and, you know, sort of prompt based agentic system or other things that are helping to, you know, bring this you know, our world of sort of, like, properly metadata and and managed data into those AI systems to reduce that risk of hallucination and and improper actions or improper responses that are coming in that. I I I think that is a interesting challenge to solve from a technical perspective to say, how do we make sure that we put into, behind whatever AI system we're talking about a layer that says, okay, if the question comes and it's related to this data, then it's going to come back to that same metadata and governed data to say that we're going to get the same response. If Chris and I and whoever else, Raymond, we all ask the data the same question, the AI the same question, we should be getting the same response, which isn't a challenge to solve. I personally feel that the harder challenge to solve in that is the thing that has persisted throughout the age of sort of trying to create single sources of truth is that in reality, if the three of us go off looking at the same data, we can come back with three different answers. Right? Even with three fields that are the same, I can ask the same question three subtly different ways. The source of truth may be the same. The answer may be different depending on how that's done. Let alone where we're talking about, you know, especially large scale enterprises where from mergers, acquisitions, things like that, transitions from systems to systems, having a single source of truth when you have six, twelve, 25, 59, I think is the highest number I've ever heard of, ERP systems. How do you create a single source of truth from stuff like that? Right? So the idea and the premise of a single source of truth, I think, is still the same. And the aim of moving towards having those single sources of truth very effectively represented within agent agentic and AI systems is tricky. I do think that this does in some way come back to where, ideally, the sprawl of agents and the ability for us to create more systems accessing that data helps to highlight more quickly where those challenges lie. But the actual problem solved still, you know, requires manual effort and real work to sort of go and read those things. What else? There's a couple other yeah. Go ahead, We're gonna. find one some notes. in chat, if there aren't any others, which is a good one. Is there a free trial available? So I will put the link in chat. Of course, there is. There's a fourteen day free trial, on the website. Third, to check it out. It's really easy to set up. Just go through the instructions. If you wanna do anything with agents, just connect your, API token. That's also really easy to do in the in the launchpad, and use the academy, academy.dataiku.com. I can, I can throw the link in there as well? That takes you through some of the onboarding and tutorials. I'm I'm very passionate about the platform, so I really do encourage all of you to, to try it out. Yeah. There's one other one that I that I actually wondered if this was gonna come up as we were talking about prepping for this today was the thing yesterday with, OpenAI sort of going out with the news about, you know, their sandbox models going rogue and accessing Hugging Face. I mean so I'm curious. Chris, what's what's I'll I'll let you what's your thoughts on this first? Man, I don't think the the dust has even settled on it yet. It's it's it's, it's scary. It's probably the the the the one word that comes to mind. I think most importantly, the thing that it's highlighted to me is, like, the sovereignty conversation again. Right? The fact that, this organization had no means to defend itself, other than going to an open weights model, which didn't have those guardrails and restrictions, which meant that it could fight, an attack, I think, opens up that topic of, like, sovereignty and do organizations need to start considering having even if it's not that the the the data center level of intelligence, you know, the the thousands of GPUs required to run, something like a a a Fable perhaps. But do organizations need to start considering having a small part of that intelligence being sovereign to them? Because it's a huge amount of risk exposure not having that. And we've seen with different regulations coming in left, right, center, The US apparently coming up with, with something about, open weights models very soon. I think that's gonna become a very, very big conversation even for someone like myself. Right? I'm using LLMs a lot. I am looking at GPUs, significant investments in in GPUs, and and owning some some level of that myself because I'm seeing that the subscriptions that I've got, are getting a little bit out of control. So investing that capital upfront, for that level of intelligence, which to be honest, like, the level of intelligence that you can get now, you know, we're talking like, GLM 5.2, for example. Like, that's more than enough 90% of all tasks that you're ever going to do, right, as a as an individual. So I think it's it's really interesting. Brings up, a lot of questions. I the the thing that I've really sort of felt as I sort of processed it over the last, you know, day, this this particular one really kinda came to, like, two conclusions about it at this point. Right? You know, the desk is not yet settled. But one of the things that, and I think this is in McKinsey's sort of state of AI report. I don't remember exactly what they call theirs, but, you know, the one that they published a couple months back. You know, they looked at sort of, like, what the companies who are really ahead on leveraging agentic and generative AI technology was and, like, what the patterns there were. And, you know, one of the things that was, I guess, unsurprising was that the companies who were the furthest ahead were also the ones who had had the most major, like, you know, blow ups. Right? Like, there is an element of risk in the things that we're doing, and we're trying to do this. Not to, like, excuse the social sort of, like, hand wave away that, you know, like, it's scary as heck. What does what what OpenAI is talking about there? But that, like, there is still a huge level of r and d and experimentation. Obviously, at at, you know, places like Anthropic and OpenAI, they're continuing to push the bounds on what the actual sort of, like, foundation models do. But within an organization, there is a certain amount of risk that is coming with continuing to build it and push. And even if you're not at the bleeding edge, you know, just, you know, you obviously, nobody's just setting these systems loose. And so this highlighted to me this sort of need for this governance and this awareness and this inventory because that it happened in a test environment where they were watching something and they knew that it happened is, like, a level of sort of, like, concern and should give everybody pause that's working on these systems. But they knew it happened because it was an agent that somebody had built, that somebody had sandboxed, and they were doing testing on. Worse is that there's probably stuff like this out there happening that people don't know about. And that's the part that is a bigger concern to me. I'm concerned about sort of, like, what OpenAI and Anthropic and, you know, sort of, like, the leading edge of the frontier models are doing. But all of the companies and organizations that are sort of stepping up and and taking swings and working with these things, those stats from the beginning of sort of, you know, 75% of companies are leveraging this stuff. Only 10% feel like they really know what they're doing. That's probably true. So how many of that 65% delta have something not to maybe the sort of, like, extent of what happened here at OpenAI, things like this happening in the background. And so it just to me highlights the need for continued investment across the spectrum, both at places like, the, you know, the the frontier model companies, but, you know, at organizations and other places in the tech ecosystem that the governance, the inventory, all of this needs to catch up. Because more incidents like that are out there already happening or bound to happen. And if we don't know that they happen, that's where, like, the shit really hits the fan. So, I guess with that, you know, I said we're coming right up at the top of the hour. Thank you so much for everybody that joined. It was a pleasure. I really appreciate, the participation, the questions, some some great questions in there. Thank you to Chris. Of course, I really appreciate you taking the time. Always a pleasure to spend time talking about fun and tricky topics, and I hope everybody has a great afternoon, evening, etcetera. Right? Cheers, Thanks very. much, everyone. Alright. Thanks, Bye bye. Chris. Bye, all.