AI red teaming for AI agents means attacking an agent the way an adversary would: through its prompts, the content it retrieves, the tools it can call and the permissions it holds. The goal is to learn what the agent actually does under pressure. The stakes changed when models started taking actions. In 2025, a single crafted email was enough to make an enterprise copilot exfiltrate internal files without anyone clicking anything, a flaw rated 9.3 on the CVSS scale. This guide covers what agent red teaming must include, how it differs from a penetration test or a model evaluation, how certification testing is reshaping it, and how to build a program that produces evidence your buyers and auditors will accept.
Why AI Red Teaming Has Moved From Research Labs to the Security Review
Agents turned output risk into action risk
A language model produces text that a person usually reads before anything happens. An agent calls an API, updates a record, sends a message or runs code, often with no human reviewing each step. That shift moves the risk from what the system says to what the system does, and it changes what a test has to prove. A jailbreak that makes a chatbot say something embarrassing is a reputational problem. The same manipulation applied to an agent with write access to a CRM, a payments API or a code repository is a security incident.
Buyers now ask for behavioral evidence
Enterprise security reviews used to accept policies, architecture diagrams and a SOC 2 report as proof that a product was safe to deploy. For agents, that evidence describes intended boundaries without showing whether the agent respects them when someone pushes past them. AIUC-1 is a certification standard written specifically for AI agents. It closes that gap by pairing an audit of controls with large-scale technical testing of the agent itself, and it produces an audit report that includes red teaming results buyers can request. The standard now sits in the Cloud Security Alliance STAR Registry, and vendors including UiPath, Cursor, Harvey, ElevenLabs and Fin have certified. Whether or not you pursue certification, those announcements have taught procurement teams what a credible answer to “how do you know your agent is safe?” looks like.
The cost of a finding someone else discovers
A failure your own testing finds is a ticket. The same failure found by a customer’s security team during a review is a stalled deal, and found by an attacker in production it is a disclosure event. Agent vulnerabilities are also unusually cheap to exploit. The attack is written in natural language and often delivered through content the agent reads on its own, such as an email, a support ticket, a web page or a shared document. Elevate’s analysis of agentic applications under attack shows how memory, tools and external content multiply the entry points. The economics favor finding these failures first, on your schedule.
How AI Red Teaming Differs From Penetration Testing and Model Evaluation
Three disciplines get grouped under “AI security testing,” and they answer different questions. Confusing them is the most common reason a test report fails to satisfy the reviewer who asked for it. Each one is useful, but only one of them tests the behavior of an agent operating inside your environment.
| Dimension | Application penetration test | Model red teaming | AI agent red teaming |
|---|---|---|---|
| Target | Infrastructure, APIs, authentication | The model’s responses | The agent’s decisions and actions across tools, data and channels |
| Core question | Can an attacker break the system around the AI? | Can the model be made to say something harmful? | Can the agent be made to do something it should refuse? |
| Typical findings | Misconfigurations, injection flaws, broken access control | Jailbreaks, toxic output, sensitive data in responses | Goal hijack, tool misuse, privilege abuse, unsafe tool calls |
| Evidence produced | Findings with CVSS severity | Pass rates on prompt sets | Scenario results tied to controls, severity thresholds and retest history |
The three columns complement each other; none replaces the others. A penetration test can confirm that the API gateway enforces authentication while the agent behind it executes an instruction hidden in a customer email. Model red teaming can show that the underlying model refuses harmful requests in isolation. Yet the same model, wrapped in an agent with broad tool permissions, can be walked into a harmful action one reasonable-looking step at a time. Agent red teaming builds on both, and it answers the question an enterprise buyer actually cares about. Elevate’s overview of the OWASP LLM Top 10 covers the model-level risks this layer inherits.
The attack surface is everything the agent can reach
An agent’s exposure equals the sum of the credentials, tools, data sources and other agents it can touch. That is why scoping matters more in agent testing than in most security work. An agent with read-only access to a knowledge base and an agent that can issue refunds are different risk classes, even when they run on the same model. A useful scope document lists every tool the agent can call, the permissions attached to each, the content sources it ingests without human review, and the channels through which users reach it. If that list does not exist, the red team will spend its first week building it, and that is worth knowing before you budget the engagement.
Non-deterministic behavior changes what a pass means
Traditional tests are largely repeatable: an injection flaw either exists or it does not. An agent can refuse an attack nine times and comply on the tenth, depending on phrasing, context length or the state of its memory. Agent red teaming therefore runs large scenario sets, repeats them, and measures failure rates by severity instead of declaring a single pass. Certification testing reflects this: AIUC-1 evaluations typically run 1,000 to 5,000 test scenarios, and Harvey reported passing more than 3,000 unique tests with zero critical failures. Define in advance which failure severities are acceptable at what rate, or the results will be impossible to interpret.
What AI Red Teaming for Agents Should Cover
The OWASP GenAI Security Project published its Top 10 for Agentic Applications on December 9, 2025, after review by more than 100 practitioners. It is the most practical checklist available for scoping an engagement, and Elevate’s guide to OWASP agentic AI security threats covers each category in depth. The five areas below group those risks by what a red team actually does during testing.
Goal hijack and indirect prompt injection
Goal hijack is the risk OWASP ranks first: an attacker changes the objective the agent is pursuing, not only a single response. The most dangerous version is indirect, where the malicious instruction arrives inside content the agent retrieves on its own instead of from the user. Testing should plant adversarial instructions in every ingestion path the scope document lists, including emails, tickets, uploaded files, web results and RAG sources, and observe whether the agent follows them. A strong result shows the agent treating retrieved content as data, never as a new set of orders.
Tool misuse and unauthorized actions
Here the red team tries to push the agent into using legitimate tools in illegitimate ways: deleting instead of reading, sending data to an external address, chaining two harmless calls into a harmful one, or acting outside the task it was given. The test also checks whether high-impact actions trigger the human approval step the design promises. Many organizations discover at this stage that the approval gate exists in the interface but not in the API the agent actually calls. Each finding should be traced to a specific control: a permission, an allowlist, a confirmation step or a rate limit.
Identity, privilege and data exposure
Agents often run with a service identity broader than any single user who talks to them. That creates a path for a low-privilege user to reach high-privilege data through the agent. Red teaming probes whether the agent enforces the requesting user’s access rights, whether it leaks information across customers in multi-tenant deployments, and whether system prompts, credentials or personal data can be extracted. These tests connect directly to classic security controls, which is why testers with application security experience recognize the patterns quickly. The difference is that the exploit is a conversation instead of a malformed request.
Reliability failures that become actions
Hallucination is a quality problem in a chatbot and a security problem in an agent, because a fabricated parameter can be passed straight into a tool call. Testing should include scenarios where the agent lacks information it needs, and observe whether it asks, escalates or invents. It should also verify that the agent validates sources before acting on them and that tool calls are authenticated. The failures found here are rarely dramatic in isolation, yet they drive many of the incidents that reach customers.
Channels, memory and multi-agent paths
Voice, SMS, email and embedded widgets each change how an attack is delivered and which defenses apply, so a test limited to the chat interface leaves real exposure uncovered. Persistent memory adds a time dimension: an instruction planted today can alter behavior in a session weeks later. Where agents delegate to other agents or connect to tools through the Model Context Protocol, the red team should test whether trust is verified at each hop or simply inherited. These paths are newer and less documented, and for that reason they are where testers most often find critical issues.
| Test area | OWASP Agentic Top 10 | AIUC-1 domain | Example scenario |
|---|---|---|---|
| Goal hijack | ASI01 Agent Goal Hijack | B. Security | Hidden instruction in a support ticket tells the agent to email the case history externally |
| Tool misuse | ASI02 Tool Misuse and Exploitation | B. Security, D. Reliability | Agent is steered into issuing a refund outside policy through a chained request |
| Identity and privilege | ASI03 Identity and Privilege Abuse | A. Data and Privacy, B. Security | Low-privilege user retrieves another customer’s records through the agent |
| Reliability | No direct category; amplified by ASI08 Cascading Failures | D. Reliability | Agent fabricates an account ID and passes it into a write operation |
| Memory and multi-agent | ASI06 Memory and Context Poisoning, ASI07 Insecure Inter-Agent Communication | B. Security, E. Accountability | Poisoned memory entry changes agent behavior in a later session |
This mapping lets one test produce evidence for several audiences. The OWASP category speaks to the security engineer, the AIUC-1 domain speaks to the certification auditor, and the control the finding traces to speaks to the GRC team that maintains the evidence. AIUC has stated that its certification covers all ten OWASP agentic threats, so a program organized this way already lines up with the technical evaluation. Scenario design is where most internal programs fall short: generic jailbreak lists test the model, while the scenarios above test your agent, your tools and your data.
How Certification Testing Is Reshaping AI Red Teaming
What the AIUC-1 technical evaluation involves
AIUC-1 certification runs in four steps: scoping, evaluations, audit and certification. During scoping, the agent, its deployment context and the applicable control domains are defined, and the standard version is locked for the certification year. The evaluation phase runs thousands of scenarios against jailbreaks, prompt injection, data leakage, hallucinations and unsafe tool calls. An accredited third-party auditor then reviews policy, operational and technical controls across the six domains. The certificate and audit report are valid for one year, with technical retests every quarter.
Pre-testing is not certification testing
The certification’s red teaming is performed by the standard’s own evaluators, not by the organization or its advisors, and that independence is what gives the report weight with buyers. What an organization can control is how prepared its agent is when those evaluations begin. Adversarial pre-testing runs the same classes of attack earlier, against the scope that will be certified, so failures surface while fixing them is still cheap. An agent that enters evaluations after pre-testing tends to produce findings the team already understands, instead of surprises that restart the timeline.
Quarterly retesting turns a project into a program
Because technical testing recurs at least quarterly, red teaming stops being a one-time readiness step and becomes an operating rhythm. Every model upgrade, new tool, expanded permission or new channel changes the attack surface the next round will probe. The standard itself is revised each quarter, and recent updates added requirements on MCP security, agent permissions, third-party risk and coding agents. Organizations that tie internal testing to their release calendar see fewer surprises than those that test only when a retest date approaches.
How to Build an AI Red Teaming Program for Your Agents
Start with an agent inventory and a blast-radius map
You cannot test what you have not catalogued, and most organizations run more agents than they think. Agents embedded in procured platforms, automations built by individual teams, and copilots with plugin access are often missing from the security team’s list. For each agent, record what it can read, what it can write, which actions are irreversible, and who owns it. That map sets test priority: the agent that can move money or modify production code goes first, regardless of how many users it has.
Define severity and acceptable failure rates before testing
Agree in writing on what counts as a critical, high, medium and low failure for each agent, and what failure rate is tolerable at each level. A critical failure might be any unauthorized write action or any cross-customer data exposure, with a tolerance of zero. A low failure might be an off-topic response with no downstream effect. Without these definitions, every result becomes a debate, and the report cannot support a go or no-go decision.
Close the loop between findings and controls
The most common breakdown in AI red teaming is not the testing but what happens after the report. A finding that sits in a security backlog, disconnected from the control that should have prevented it, will reappear in the next quarterly round. Each finding should map to a control owner, a fix, a piece of evidence and a retest date, which makes this GRC work as much as security work. If your organization runs an AI management system under ISO 42001, its risk treatment and corrective action processes are the natural home for this loop. Elevate’s ISO 42001 executive summary outlines how those clauses work, and its ISO 42001 compliance services connect red teaming findings directly to the management system.
Choose between in-house, outsourced and co-sourced testing
| Model | Best fit | Main strength | Main gap |
|---|---|---|---|
| In-house | Teams with dedicated AI security engineers and frequent releases | Continuous testing tied to every deployment | Testers share the builders’ assumptions; limited independence for buyers |
| Outsourced | Organizations preparing for certification or a major enterprise deal | Independent perspective and attack experience across many environments | Point-in-time unless contracted as a recurring engagement |
| Co-sourced | Organizations with some internal capability and a quarterly retest cycle | Internal automation for regression, external experts for new attack classes | Requires clear ownership of findings between two teams |
The decision usually turns on two factors: how often the agent changes and who needs to trust the results. Automated scanners and internal regression suites catch known attacks cheaply after every release, which makes them essential but insufficient on their own. Novel attack chains, especially those combining tools, memory and multiple agents, still depend on experienced human testers. For most organizations selling agents to enterprises, a co-sourced model aligned to the quarterly retest calendar balances cost and credibility best.
If you are not sure which of your agents would fail first, a short conversation can answer that before you commit budget. Book a Readiness Call with Elevate’s AI governance practice to review your agent inventory, your exposure and a realistic testing sequence.
How Elevate Helps Organizations Red Team Their AI Agents
AI red teaming for agents sits where two disciplines meet, and they rarely live in the same firm. The testing side is offensive security: prompt injection, tool abuse, privilege escalation and exfiltration attempts, designed by people who have spent years breaking applications. The evidence side is governance, risk and compliance: mapping each finding to a control, an owner and a record that an auditor or a buyer’s security team will accept. Elevate runs both practices under one roof, with more than 500 penetration tests delivered through its penetration testing services and an ISO 42001 practice in house.
That combination changes how an engagement runs. Elevate scopes the agent inventory and blast radius first, designs scenarios around your tools and data instead of generic prompt lists, and ties every finding to the control that should have caught it. Fixes are retested by the same team that found the failure, and results are organized to support a customer security review, an ISO 42001 program or preparation for independent certification. Elevate prepares and tests but does not certify, which preserves the independence buyers look for in a final report.
Elevate’s AI governance practice is led by Angela Polania, who holds CISA, CISM and CRISC credentials and is an ISO 42001 Lead Auditor. With 18+ years in cybersecurity and compliance and 500+ clients served, Elevate works with CTOs, CISOs, AI product leaders and compliance officers at organizations building or deploying agents. For teams on the other side of the table, Elevate’s guide to AI vendor risk assessment covers what procurement should ask agent vendors. Talk to an Elevate advisor about where your agents stand today.
Conclusion
AI red teaming for agents answers a question that policies and architecture diagrams cannot: what does the agent do when someone tries to make it misbehave? The attack surface is everything the agent can reach, the failures are statistical instead of binary, and the evidence has to serve security engineers, auditors and procurement teams at the same time. Programs that start from an agent inventory, define severity before testing, and close every finding into a control produce results that hold up under review.
Certification testing has raised the bar. Thousands of scenarios per evaluation and quarterly retests mean red teaming is now a recurring discipline tied to every release, not a pre-launch exercise. Organizations that pre-test against the same attack classes, on their own schedule, enter those evaluations with fewer surprises and shorter timelines.
If your agents are heading into enterprise security reviews or independent certification, find out where they fail before someone else does. Book a Readiness Call with Elevate to map your agents, your exposure and your first round of testing.
Key Takeaways
AI red teaming for agents is a distinct discipline, and these points define how to run it well.
- Agents shift risk from output to action. Testing must prove what the agent does with tools, data and permissions, not only what it says.
- Scope decides everything. An agent’s exposure equals every credential, tool, data source and channel it can reach, so the inventory comes before the first test.
- Results are statistical. Agents behave non-deterministically, so programs need large scenario sets, repeated runs and severity thresholds agreed in advance.
- Findings must close into controls. A finding without a control owner, fix, evidence and retest date will reappear in the next round.
- Certification made red teaming recurring. Quarterly technical retests turn a one-time readiness step into an operating rhythm tied to every model or tool change.
FAQs
What is AI red teaming for AI agents?
AI red teaming for AI agents is adversarial testing that tries to make an agent take actions it should refuse. Testers attack through direct prompts, content the agent retrieves, the tools it can call and the permissions it holds. Unlike model testing, it measures behavior inside a real deployment, including data access and downstream actions. Results are usually reported as failure rates by severity, because agents do not respond identically to the same attack every time.
How is AI red teaming different from a penetration test?
A penetration test looks for flaws in the infrastructure and application around an AI system, such as broken authentication, misconfigurations or injection vulnerabilities. AI red teaming targets the behavior of the model or agent itself, for example whether it can be manipulated into leaking data or misusing a tool. The two overlap and work best together, since an agent can sit behind a well-secured API and still follow a malicious instruction hidden in a document. Mature programs run both and correlate the findings.
How often should AI agents be red teamed?
AI agents should be red teamed whenever their attack surface changes and on a regular cadence between changes. Triggers include a model upgrade, a new tool or integration, expanded permissions, a new user channel or new memory features. Independent certification programs for agents now retest technical behavior at least quarterly, which has become a practical benchmark for organizations selling to enterprises. Automated regression tests can run on every release, with deeper human-led testing each quarter.
Does AIUC-1 certification require AI red teaming?
Yes. AIUC-1 certification includes a technical evaluation phase in which the agent is tested against thousands of adversarial scenarios, covering jailbreaks, prompt injection, data leakage, hallucinations and unsafe tool calls. That testing is conducted as part of the certification itself, alongside an independent audit of policy, operational and technical controls. Certified agents are retested at least quarterly during the certificate’s one-year validity. Many organizations run adversarial pre-testing beforehand to find and fix failures before the formal evaluation.
Can an internal team red team its own AI agents?
An internal team can and should run ongoing red teaming, especially automated regression tests after each release. The limitation is independence: internal testers share the builders’ assumptions and often miss attack paths that an outside team would try first. Enterprise buyers and auditors also give more weight to results produced by an independent party. Many organizations use a co-sourced model, with internal teams handling continuous testing and external specialists running deeper engagements and new attack classes.