Podcast

Securing Agentic AI: The OWASP Approach

Guest: Steve Wilson
Host: Mo Sadek. Technical Marketing Director, Alice
Episode #2
-
Feb 2026
Securing Agentic AI: The OWASP Approach
"You know who else is prone to hallucinations and prompt injection? Every employee you have."

Episode description

In this episode, Mo Sadek is joined by Steve Wilson (Chief AI and Product Officer at Exabeam, founder and co-chair of the OWASP GenAI Security Project) to explore how OWASP is shaping practical guidance for agentic AI security. They dig into prompt injection, guardrails, red teaming, and what responsible adoption can look like inside real organizations.

Meet the guest

Steve Wilson, Chief AI and Product Officer at Exabeam

Steve Wilson

Chief AI and Product Officer at Exabeam

Steve Wilson is Chief AI and Product Officer at Exabeam and founder and co-chair of the OWASP Gen AI Security Project, where he helps shape the standards for secure, production-ready AI. A recognized AI security leader, he’s a Google Cloud AI Innovation All-Star and author of The Developer’s Playbook for Large Language Model Security. He previously held leadership roles at Sun, Citrix, and Oracle and holds 11 patents.

Full transcript

Securing Agentic AI: The OWASP Approach

Curiouser & Curiouser, Episode 2 with Steve Wilson

A lightly edited transcript. Disfluencies and false starts have been cleaned up for readability. The substance is unchanged.

Steve Wilson: Technological change is accelerating right now. Compare the rate of change in the 90s, 2000s, or 2010s to the 2020s. These are exponential curves. So we're going to need to find things that can move at high speed, and that's always going to be tricky for the government. We've got to find ways to have those government agencies support programs that can run at higher speed. We've had really good collaboration with the likes of NIST and MITRE.

Mo: If AI has ever made you stop and think, "wait, what is happening?", you're not alone. I'm Mo, and I'm a security researcher asking the same questions. On Curiouser and Curiouser, we have open conversations with experts, researchers, and leaders working at the edge of this space, talking through how AI is taking shape, what's shifting, and how the people inside the work are thinking about it as it happens. So join us and listen in as the conversation takes shape.

Meet Steve Wilson

Mo: Welcome back to Curiouser and Curiouser. I'm really excited about today's episode. I think I say that every time, but we just get some of the best people, and today is no exception. We've got Steve Wilson, Chief AI and Product Officer at Exabeam, and, more excitingly to me, the founder and co-chair of the OWASP GenAI project. Great to have you here, Steve. I've probably ruined the intro, so I'll let you introduce yourself, because you have an amazing history.

Steve Wilson: Thanks a lot, Mo, and thanks for having me. I won't bore you with my whole life story, but for this purpose: I started my first AI company when I graduated from college in 1992. I've been doing AI projects on and off since then, while working on developer tools and large-scale cloud infrastructure. When ChatGPT came out, it was clear this was a next step in the evolution of AI, and that's when I put together the first version of what's now called the OWASP Top 10 for large language models. That became a kind of definitive document for how you approach security with these things. We had a couple hundred people contribute back in 2023, which I thought was amazing. The group has since grown to 22,000 people. We've put out 30-something white papers on different aspects of AI security, everything from red teaming to CISO guides, and most recently a new Top 10 for agentic applications.

In my day job at Exabeam, I run our product teams, and we use generative AI to build advanced cybersecurity agents that we include in our security operations platform. More recently, we're defining a new discipline we call agent behavioral analytics: how do I watch what the agents are doing and keep tabs on them, just like I try to do with my employees?

From 1992 to now: three decades of full-stack change

Mo: That's all pretty exciting. Just to confirm, did you say 1992? That's massive in terms of the shift in tech. From 1992 to now, you've lived through multiple tech shifts. What did it feel like seeing AI then, seeing it ten years ago, seeing it two years ago, and seeing it today?

Steve Wilson: In a nutshell: I start this AI company in 1992, and we're trying to build advanced neural networks on my Mac II with eight megabytes of RAM. We actually did some cool stuff, kept ourselves fed for a few years, and then the worldwide web happened and we said, this AI stuff isn't ready, we'd better go figure out the web. So we folded up the company. I got lucky and landed a job at Sun Microsystems working on Java to learn how the web worked. That was amazing.

Fast forward to 2023. I looked around and said, this AI stuff is amazing, and everything I'm doing right now is yesterday's news. So I quit my job and found a new one that would put me in a spot to do this AI stuff full time. I broke some of this down in the book I wrote for O'Reilly, but when I look at the difference between 1992 and 2024, that's 32 years, which is an easy computation for Moore's law. Things should be about 65,000 times as fast. You want to know the actual difference between the math coprocessor in my Mac II and an NVIDIA GPU? 143 million times. And to train a neural network, we take 10,000 of those and gang them together.

So there's a lot that changed. We invented transformers. We had to build the internet to give us a place to get enough data to train these things. We had to build large-scale computing infrastructure on top of GPUs. Thirty years of full-stack innovation went into that change. It's dramatically different, and the acceleration is crazy, because it keeps changing faster and faster.

How building software changed: waterfall, agile, and now agents

Mo: You touch on a couple of really good things, and one of the biggest is how fast infrastructure has changed. But people have changed too. Product in 1992 versus product now, who you're selling to, your ICPs, all of that changes how we build things. You come from a software engineering background and have seen some of the biggest changes in how people build. So what does that look like in your teams, and where are the challenges?

Steve Wilson: It's interesting to look at the cycle we're going through in software engineering. My dad was what I'd call a real engineer. He had a PhD in physics and got his first job at Hewlett-Packard when he got out of college, way before HP was making computers. He and his team built the first ultrasonic imaging machines, the kind you see at the hospital these days. So when he was teaching me to program as a kid, he was teaching me engineering. You plan it, write a specification, plan how it's going to work, think about it, and then you write the code. It was this incredible, disciplined development, what we'd call waterfall, that happens when you need stuff to work.

Then we made this big shift to agile and said, that's just too slow. If I'm going to keep up with the market, I need to iterate faster, so I'm not going to spec everything beforehand. That led to a set of goodness, but also a set of badness. A lot of teams stopped thinking about architecture, stopped planning in advance, and just made it up as they went along. Some parts of software really suffered.

What's crazy now is that while I've continued to be involved in building software, I hadn't written code in two decades. At some point, as you go down the management chain, you hang up your IDE and say, I'm a manager, I manage people who write code, or I help define products. 2025 was a watershed year for me. I discovered ChatGPT could write code that was not trivial. Then I got Cursor, then I got Claude Code, and all of a sudden my hobby projects went from a hundred lines of HTML to 50,000 lines of Python.

The amazing thing I learned, to tie it back, is that to make those agents really work, you have to tell them what to do, really exactly. You have to sit down and plan what you want, sketch out the architecture, maybe write a PRD. If you give all that as context to your bots, they can do real work for you. So it all comes around. The more things change, the more they stay the same.

Standing up a software team of agents

Mo: I ran into the exact same problem, and I've not been building software that long. I took about six months to sit down and build a project, and it was fantastic, until something went wrong about eight or nine hours in. I realized I wasn't using Git and wasn't making commits, and there was no way to salvage it. Hundreds of thousands of tokens completely gone, and hours of my time, which honestly would have taken months before. The efficiency is crazy.

Six months later I have all these agents orchestrated. Like you said, write a PRD: I have a little software engineering team alongside me, a project manager, a product person, agents testing the problem as we go and building unit tests, one making sure the microservices are properly structured. Every part of the apps I build now is really good, in my opinion. It's crazy how enabled we are, people who aren't even full-time software engineers, to go build our visions just because we understand how software is built.

We're seeing that in traditional software engineering too. Engineers are picking up these tools and running with them. It took me six months to figure out I needed to properly use code management; they're totally beyond that, using it for tab completions to write insane amounts of code alongside the code they're being guided to build. People are enabled at every level to do amazing things. But you're also the Chief AI Officer. How do you see that going into your organization? That role is new to the industry, spanning innovation to using AI responsibly in product, and I've seen the Chief AI Officer, product, and the chief security officer working really closely in tandem. How does that play out for you day to day?

The Chief AI Officer role: getting the whole org on board

Steve Wilson: When I joined Exabeam a couple of years ago, the first thing I did was push generative AI into our products. We built the first viable cybersecurity co-pilot, then the first family of cybersecurity agents to help run your SOC. At the same time, our CEO was saying, how do I get more leverage? I'm reading about companies transforming themselves with AI. We're putting it in our products, but we don't have a strategy for using it to run our own business. So he asked me to take on this additional role, because everyone knows I'm the AI cheerleader.

We had to start at the beginning. I got together with our CFO, our chief legal officer, and our CISO, none of whom were big AI users. Some were actively resistant. They said, this is scary, why would you use this? You lose your IP, they give you bad answers. So we did a lot of education, which led to first baby steps: building an acceptable use policy for these AI tools, having the CISO invest in that, and helping him understand that this wasn't only going to unlock the tools, it was going to let him make them safe. Without allowing people to use them, he's just fighting shadow AI. Instead, with an acceptable use policy, blessed and evaluated tools, and the CFO and chief legal officer evaluating the data retention and data use policies of the companies that build them, I now have those tools in-house, I can log the data coming out of them, and I can put that into compliance.

Along the way we selected a core AI platform, what's now called Google Gemini Enterprise. That gives everyone in the company standard chatbot functionality, but it also lets anybody build their own agents, which is unlocking a lot of creativity and pointing us to the next place this goes: every enterprise is going to have more agents than employees pretty soon.

Why OWASP, and why it moves fast

Mo: There's something in there that I think was the unblocker for you, and it helps explain where OWASP comes from. Coming from an application security and product security background, a lot of the friction was never technical, it was a language thing. There's a language barrier between the security person and the product person, and getting the thing done or communicating why it's risky, especially when it's also a revenue problem. When you can communicate in the same language, or find a common layer of technical communication between these groups, it's much easier to get things done.

I also tend to be conservative on risk, and I don't think every risk AI introduces is novel. But we need to be able to attribute them, to ask where the risk exists, what language describes it, and how we put the impacts in terms all the stakeholders understand. That's where OWASP is invaluable. So take us back, not to the 1990s, but to when you started the OWASP GenAI project and those first Top 10s. What were the guiding principles? We're all familiar with the web project, but for LLMs it's quite different.

Steve Wilson: The more things get different, the more they stay the same. I was working at an AppSec company whose CTO, Jeff Williams, wrote the original web Top 10, back when it was the only Top 10. Fast forward to 2023, and before RSA I was working with our marketing department asking, what are we going to say about AI? It was so new, but we knew everyone would want to talk about it. So we started writing things down. Then people knew I was interested, and they started sending me things, a little article about data poisoning, this or that. There was really a security story about what needs to happen for these things, because they are novel. Sure, it's software, and everything from the web Top 10 to what you need for cloud computing stacks up. It's not completely new. It's software, it runs on the web. But there were unique things.

So I put together a draft version of the Top 10 and went to Jeff and asked if it was worth doing. He said, you should go talk to the OWASP folks, and introduced me to the board. What everybody realized is there'd been research done on this, scattered across papers and university research, but there was no single place you could go to learn about the security risks of large language models. Going back to that agile discussion, we said, NIST or MITRE is going to take a year or two to put out guidance on something new like this. We're a bunch of open source nerds, we can work quick. So we said, we're going to do it in six weeks, three two-week sprints, and at the end of sprint three we're shipping it. It's time-boxed.

Tons of people read it, and it ballooned into a hunger for more information as this became the defining trend of the 2020s. The thing about OWASP that's really different from the big security organizations is that it can move so fast, but it's well respected enough. 2025 was interesting: I went to Davos for the World Economic Forum, spoke at the UN. When people want to talk about AI security, OWASP has become someone who has to be at the table.

The dependence on standards

Mo: Having that seat at the table is important. For me, a lot of this has become a quality issue too. Having a great product in the age of AI is about trust; there's a lot of data on the line. But when things move so fast, government organizations move pretty slowly, and OWASP fills that gap: we've got a project, a timeline, we'll get it out, we'll iterate out loud in the community. Even when there isn't a finished standard, there's guidance in draft that people can go try, come back, and say, this doesn't actually work like this, we should update it. When we were working on the agentic GenAI piece, it was very async, a lot of it happening in a Google Doc, people coming back and tweaking definitions.

One big thing: MITRE lost funding for part of last year. That caused a lot of fear and uncertainty, because MITRE isn't just a US thing; the world looks at these standards and they inform how a lot of people do things. So how do you feel about this dependence on the standards?

Steve Wilson: At the bottom of it, MITRE, CISA, and NIST are incredible resources that have delivered tons of value to the industry. At the same time, we have to evaluate the models for how we support them. They can be completely bottom-up, the way we built some of this at OWASP. But we also found there's a set of things we could do with more resources. Even in our own little OWASP group, we've gone out and gotten corporate sponsors who donate time and energy so we can do more, run events, spread the word. Government, in some ways, is the thing with the most money and influence, and the thing that could get baked into regulations. Nobody's going to write a law around the OWASP Top 10, and should they?

But what we know is that technological change is accelerating. Compare the rate of change in the 90s, 2000s, or 2010s to the 2020s; these are exponential curves. We're going to need things that can move at high speed, and that's always tricky for government. So we've got to find ways for government agencies to support programs that can run at higher speed. We've had really good collaboration with NIST and MITRE, which is great. Maybe that's the model going forward.

Mo: There's a lot of collaboration, but they need to be able to support these things. OWASP can move because of the community, but there isn't a lot of support from external places for these government-centric standards. They ask experts to come in and give opinions, but it takes a long time. The agility piece hurts a little, but by the time they come out, they're really well structured and provide a strong framework that something like OWASP can plug into: NIST says this, so here's how you'd do it, or MITRE recommends these things, here's how you use OWASP to think about your maturity. There's a nice synergy.

Going higher level, I think there's an over-indexing on the technology piece. People say, because we have standards and ways to implement, we should just implement. But there are a lot of misconceptions about AI safety and guardrails, and maybe we think too much about using AI without thinking about securing it, or we over-index on some of the security mechanisms meant to protect AI. So what are the biggest misconceptions about AI security and safety right now?

The biggest misconceptions about AI security

Steve Wilson: It's been an interesting shift over the last 24 months. Go back to 2023, early 2024, and a lot of CISOs took a stance of, nope, keeping this out. I'm going to my Zscaler, finding the category that says GenAI, and turning it off. Phew, solved that problem. Last year, in my job at Exabeam, we talked to a lot of CISOs, and they'd come to me and say, I can't do it anymore, I can't keep it out. I'm getting pressure from my boss and the board, they're going to roll this out, and I need a security strategy. Regulating it out of existence isn't going to hold water anymore; there's too much apparent benefit. There are good arguments about good ROI versus bad ROI, but people are going to do it now. It's not optional.

So they're looking for strategies, asking, do I need new tools? There's a new startup every five minutes saying they'll solve agentic security, and we've been through multiple generations of them. There was a whole first generation I'd call the guardrails tools; several of the founders worked on the first version of the Top 10 with me, went off and started these companies, built the first sets of guardrails, grew with some customers, and already sold to bigger cybersecurity companies. Now people are coming in behind saying, maybe that's not enough.

If your cybersecurity strategy for securing AI is to use AI to secure the AI, that's what I call the turtles all the way down problem. What's the best way to screen for prompt injection? You can't do it with a regular expression; it's not like looking for SQL injection. The only thing you land on is natural language processing, so you'll probably use a large language model. So the thing you're using to screen for prompt injection is another thing just like it, also vulnerable to it, and you stack those and hope it solves the problem. It doesn't. Whether it's OpenAI, Anthropic, or Google, they've all come out and said prompt injection and hallucinations are durable, endemic to the way these things are built. So get used to it and plan for it.

The analogy I draw: when people say, my God, if they're vulnerable to that, how can I trust them to do any work in my organization, all I can say is, you know who else is prone to hallucinations and prompt injection? Every employee you have. We just call it phishing, and compromised credentials, and malicious insiders. We have all sorts of names for it, but it's the same. Our humans can be deceived, our humans exfiltrate data, so we wrap them in security tools and build things like insider threat programs. The next thing we add to insider threat programs is our agents. We're going to have to learn from how we deal with very imperfect humans to deal with very imperfect, very fast AIs. That's the next step in this evolution.

Humans in the loop, red teaming, and defaults

Mo: It's funny how human-in-the-loop and humans on the front line have come full circle. You mentioned phishing: your organization's best defense is the humans, but they're also the weakest link. Same with AI. When you implement guardrails, the easiest way through them is often a human. So either you have a human in the loop watching, making sure drift isn't happening, with transparent systems reporting properly so that when a human does need to decide, it's easier and there's less fatigue reviewing all these outputs, or you're doing red teaming: actually testing, doing human business-logic fuzzing and penetration testing, putting your apps in real-world scenarios.

We need more of that, but there's been a lot of passing the buck, saying this can be done better outside the organization, and sometimes that's true and sometimes it's not. What I mean by outside is organizations saying, this is difficult, so let's just depend on default configurations, turn off the GenAI here, and that'll be good enough. But your defaults are almost never good enough. So making sure you're paying attention to default configurations and have a comprehensive program all the way around is super important, and that's one of the foundations OWASP helps you build. It always seems to come back to speed; the technology moves faster than we can anticipate. Just this past weekend, an agent came out and everyone was shocked, even though it had been around for months. As soon as it went viral, people panicked.

When we were looking at skills a couple of months ago, the first thought was, this is supply chain security. Two years ago, Cursor had an issue where you could pull skill files from outside your organization, they made calls to other places, and out of nowhere you had RCE. It's a supply chain problem, and it's been around. I had a little tool that would look at a skills file and tell me if there was anything weird in it. Suddenly that basic cleanliness became important again. So of these normal defenses we already have, what do you feel is most overlooked?

Supply chain, the App Store analogy, and staying on the edge safely

Steve Wilson: There's a lot to unpack. When you look at something like that agent, this has happened before and will happen again. Go all the way back to 2023 and you see the first examples where somebody shipped an agent toolkit and everybody started slapping together agents running around the web using their credentials to do who knows what. Everybody was excited for a few weeks, until they realized how terrifying it was, and it went back down.

First and foremost, you've got to take some personal or corporate responsibility for what you deploy. If I'm talking to consumers, we have to talk about what's safe on the internet, and consumers are not good at this; there are all sorts of ways they already get ripped off. But if you're talking to businesses, go back to first principles. You can't block out AI, but you can move at a measured pace where you allow things onto your network that you've done some inspection of. There are second-order bits, like letting in a tool that brings in other content, and we deal with that today, just maybe a little slower. In our software development shops, everybody rummages around GitHub and brings down stuff, and we built a whole supply chain security industry around managing that. We got that from Log4j and SolarWinds and all these other things.

That kind of agent is teaching us about a new layer of this that providers need to get better at. If consumers are going to use agents and grab skills, this needs to be managed more like the App Store on your phone. The Apple App Store is an amazing example of a supply chain that is mission-critical, full of innovative, weird, pretty open stuff, that's also locked down and pretty secure. So the providers of some of these things are going to have to find places where they can wrap it so the everyday user can safely get at the cool stuff.

For the people developing software and living at the edge, trying the latest stuff, just don't have your rose-colored glasses on while you do it. Do it in safe places, in sandboxes. Have a burner laptop that isn't full of your credentials while you figure out how this stuff works. That agent went from nothing, to something, to actively breached in about 72 hours. You weren't going to be noticeably behind the curve by taking a little more time to poke at it and see what's really under the covers.

Mo: There's this weird FOMO, as if the technology is going to disappear tomorrow just because it showed up today. Every new AI release in the last couple of years comes out, and within a week something bad happens, and we have to roll it back and say, hang on, this is a little too fast. Then we make a standard and say, here's how to deploy it safely, don't just try it, play with it in this sandbox. It's not sustainable. I've made a grim prediction that 2026 is the year we need a really big AI incident to set everything straight, our big breach moment. But I'd rather leave on a happier note, because there's a lot of potential here, and the last thing we want is to make people think they shouldn't adopt this. They should, really fast, because it'll benefit everyone. So as this space matures, and I don't just mean AI but the guidance we provide, what signals to you that something is ready for an organization to roll out at larger scale, rather than just adopting early and seeing what happens?

Knowing when AI is ready to deploy

Steve Wilson: The first thing is to understand what you want to get done. So many times I see people rolling out projects saying, it's neat technology, let's figure out where we can apply it and sort that out later. Then you get statistics like MIT saying 95% of the projects failed, and I think 90% of them probably didn't even have a goal when they started. People just played around and ran out of energy.

It's funny, people assume when they watch what I've done at OWASP, or read my book, that I'm the big caution guy, very measured and careful with the AI stuff. I'm not. I'm the world's biggest AI maximalist. I use this stuff every day. At Exabeam, we have so many big enterprise customers, and we built AI agents into our product, turned them on by default with access to versions of people's cybersecurity data. People assume I'd say don't do that, but what we had to do was ask, what are the use cases we can put these to that are safe and valid, with eyes open about what they're good and bad at?

So often these things look like they can do anything; you can build a demo that looks like anything is possible. These new AIs are great at things we could never do before: speaking English, speaking Japanese, writing Python, summarizing a document. That was science fiction a few years ago. But you know what they're terrible at? Arithmetic, following instructions, repeatable automation. They're awful at those, just like people. There's a reason people use calculators and spreadsheets instead of doing it in their heads, and a reason you have a calendar to remind you what to do.

The thing I see most is people trying to put these into places where you could just know that's probably not a good use case. On the other hand, if I optimize it to do what it's good at, and restrict it from even trying the things it's bad at, I can build some really great use cases. We built cybersecurity agents into our core platform, and customers tell me they're three to five times faster than they used to be, turning everyone in the SOC into a team leader rather than an individual contributor. That's value. But when people start saying, I'm building the autonomous SOC and automating all the humans out of the loop, I say, you're crazy. We're nowhere near that, it's not a good idea, and it won't be in the foreseeable future.

Where to find Steve and OWASP

Mo: I have some thoughts about the autonomous SOC and moving the analysts out, and about empowering your best people to do more while creating a training gap, but that's a story for another day. I know we're close to time. First, thank you so much for chatting today. You're a trove of information. Where can people find you? I know you'll be at RSA with us.

Steve Wilson: Hit me up on LinkedIn, I post about this stuff regularly, just search Steve Wilson Exabeam or Steve Wilson OWASP and you'll find me. Send a connection request. You can also look at my book, The Developer's Playbook for Large Language Model Security, from O'Reilly. And I'm virtualsteve on Twitter.

Mo: And since we're both part of the OWASP project, where can people find us there?

Steve Wilson: For my bit in particular, we've got genai.owasp.org. That's got everything we're doing on the Top 10 lists, agentic security, red teaming guides, all of it, and it's all free. It also has information if you want to join, contribute, or participate. We'd love to have you.

Mo: It's literally the same URL, genai.owasp.org, and that's where all of our agentic stuff is. We've also got a couple of Slack channels going, so anybody in the community can join. Steve, thank you again so much. It has been a pleasure.

Steve Wilson: Thanks for having me.

Mo: If this episode helped cut through the noise, like or subscribe so you don't miss what's next. Thanks for spending time with us. Until next time, stay curious.

Read Full Transcript

SOUNDBITES

Curiouser Soundbites: AI Is Not Magic. It's a Powerful, Imperfect Operator.

Blog

AI may feel magical, but it behaves like a fast, probabilistic operator, and security leaders must govern it accordingly.

Learn More

COMING UP

Black Hat USA 2026

Event

Alice @ Black Hat USA - Where AI systems are tested the hard way, before attackers do.

Learn More

GO DEEPER

Mitigating the Risks of Agentic AI

Webinar

As AI evolves from chatbots to autonomous agents, new security vulnerabilities are emerging. Explore the critical strategies needed to identify and manage the unique risks of agentic AI before they scale.

Learn More

Subscribe for new episodes

What’s New from Alice

Curiouser Soundbites: What a Former Google Cloud CISO Wants Leaders to Know About AI

blog
Jul 10, 2026
,
 
Jul 10, 2026
 -
5
 min read
Jul 10, 2026
 -
5
 min watch
July 10, 2026

Everyone's watching the flood of new AI vulnerabilities. Former Google Cloud CISO Phil Venables is watching something else, and it's the shift leaders can't afford to miss.

Learn More

Demystifying AI Red Teaming

whitepaper
Jun 25, 2026
,
 
Jun 25, 2026
 -
This is some text inside of a div block.
 min read
Jun 25, 2026
 -
This is some text inside of a div block.
 min watch
June 25, 2026

Your AI passed every check. That doesn't mean it's safe. Learn how to red team AI systems before adversaries find the gaps you missed.

Learn More