Prompt Injection Should Change What Your AI Agent Is Allowed to Do

The support agent is authenticated.

Its token is valid. Its access to the ticketing system is approved. Its connection to the customer database is working exactly as designed.

Then it reads a support ticket containing a hidden instruction to retrieve account data and send it somewhere else.

Nothing about the agent’s identity has changed.

Everything about the safety of its next action has.

That is the authorization problem emerging around AI agents. Traditional access control asks whether a known identity may use a resource. Agentic systems add another question: what information is influencing the agent at this moment, and should that information reduce what the agent is allowed to do?

Static permission is not enough for dynamic context.

The National Institute of Standards and Technology put fresh attention on this issue in a September 29 update about its Software and Agentic AI Identity and Authorization project. NIST’s summary of more than 600 public comments says respondents repeatedly called for separate governance or gateway components to evaluate agent requests. Commenters also argued that authorization policy may need to change when prompt injection is suspected—or whenever an agent is working with untrusted data. (NIST NCCoE, September 29, 2026)

That summary is not a final NIST control standard. It is evidence of where the architecture conversation is moving.

The practical lesson is already useful:

When the trustworthiness of an agent’s context decreases, its authority should decrease with it.

CaneCorsoAI

Strong Identity Does Not Make the Context Trustworthy

Identity and authorization remain essential.

Every consequential agent should have a distinct identity. Its credentials should be protected. Tool access should be scoped. High-impact actions should require stronger controls. Tokens should be short-lived where practical. Human sponsorship and delegation should be attributable.

Those controls answer important questions:

  • Which agent is making the request?
  • Which person or organization authorized it?
  • Which tools and data may it use?
  • What limits apply to its role?

They do not necessarily answer:

  • Did the agent just ingest hostile instructions from an email, document, web page, RAG source, or tool response?
  • Did sensitive information enter the prompt context unnecessarily?
  • Did an external source attempt to redefine the agent’s goal?
  • Is the requested action consistent with the original user’s intent?
  • Should the agent still have write, send, execute, or disclose authority after the context changes?

An agent can be the correct identity, using a valid credential, while acting on corrupted influence.

This is why prompt injection is not merely a content-filtering problem. It is a runtime authorization signal.

The recent NIST comment summary describes the same architectural tension. It notes that language-model systems may not reliably separate untrusted data from trusted instruction, and that respondents favored logically separate governance components at multiple points—from user input and prompt processing to tool calls and cross-boundary requests. (NIST NCCoE summary of comments)

The reasoning system should not be the only system deciding whether its reasoning is safe enough to execute.

Use Trust-Responsive Authorization

The goal is not to reinvent identity management.

It is to make authorization responsive to the current trust state of the workflow.

Call this trust-responsive authorization:

Agent authority is determined by identity, intended task, requested action, data sensitivity, source trust, and observed content risk—not by identity alone.

That model allows an agent’s permissions to step down as risk rises.

State 1: Normal

The agent is operating within its approved mission. Inputs come from expected sources, inspection finds no material risk, requested actions remain within policy, and the agent uses only the minimum tools required.

Normal does not mean unlimited. It means the bounded workflow may continue without extra intervention.

Examples might include:

  • Summarizing a known internal knowledge article.
  • Classifying a routine support request.
  • Retrieving non-sensitive status information.
  • Drafting a response without sending it.

State 2: Constrained

The agent encounters untrusted or privacy-sensitive content, but the workflow can remain useful after controls reduce the risk.

The system may:

  • Sanitize or remove suspicious instructions.
  • Tokenize or redact sensitive values.
  • Defang links.
  • Remove write-capable tools.
  • Restrict retrieval to an approved data set.
  • Convert an autonomous action into a recommendation.
  • Prevent external communication or network egress.

The objective is graceful degradation. The organization keeps useful work moving while reducing the blast radius.

State 3: Review

The context contains suspected prompt injection, conflicting instructions, unusual data access, an unexpected tool request, or another condition that makes autonomous action inappropriate.

The agent may continue gathering non-sensitive evidence, but the consequential action pauses for human review.

The reviewer should see:

  • The original business request.
  • The source and trust level of the suspect content.
  • What was detected or changed.
  • What the agent attempted to do.
  • Which data and tools were involved.
  • Which policy required review.

A generic “AI needs approval” message is not enough. The human needs the context required to make a real decision.

State 4: Stop

The request is clearly malicious, violates policy, attempts to expose protected data, seeks an unauthorized action, or cannot be made safe without destroying the purpose of the workflow.

The system blocks or quarantines the transaction. Depending on the situation, it may also revoke a session, disable a workflow, preserve evidence, or trigger incident response.

These states should be defined by system owners before an agent encounters the problem.

An AI agent should not decide for itself when it deserves more authority.

Put an Independent Control Point in the Path

Trust-responsive authorization requires a signal about the content influencing the agent.

That signal should come from a control point that is separate from the model’s reasoning and placed directly in the workflow.

This is the role of an AI Application Firewall.

CaneCorso™ is MicroSolved’s shared control plane for AI workflows. It sits between workflow content and the model or downstream logic. Based on configured policy, it can allow, sanitize, tokenize, or block content; apply privacy controls; and preserve reasons, scores, and evidence for review.

That makes CaneCorso useful as the inspection and enforcement layer in a trust-responsive architecture.

It does not replace identity and access management. It does not replace scoped tool permissions, secure credentials, application authorization, human approval, or incident response.

It gives those systems information they normally do not have:

Is the content influencing this action safe enough for the workflow to continue in its current mode?

A practical integration can work like this:

  1. Inspect the inbound content. Email, documents, tickets, RAG results, API responses, or tool output pass through CaneCorso before entering the model context.
  2. Apply the content decision. Approved content proceeds. Sensitive or suspicious portions can be sanitized, tokenized, defanged, or blocked according to policy.
  3. Map risk to agent mode. The workflow orchestrator uses the disposition to keep the agent in normal mode, remove higher-risk capabilities, require human review, or stop processing.
  4. Limit tool execution. The authorization layer checks the agent identity, requested action, current mode, data sensitivity, and approval state before a tool runs.
  5. Inspect consequential output. Where model output will drive decisions or actions, route it through a second inspection point before downstream execution.
  6. Preserve the decision record. Send the content-risk result, agent identity, tool request, policy decision, approval, and outcome to the organization’s monitoring and audit systems.

MicroSolved’s earlier production-use article describes this two-layer pattern—inspection before model processing and another check before selected downstream use—in Brent Huston’s own agent environment. (Why My AI Agents Needed CaneCorso as a Security Control Plane)

The key design choice is not merely to log the risk score. It is to make the workflow consume the decision.

A Warning Without an Authority Change Is Only Telemetry

Many AI security designs stop at detection.

The system identifies suspicious content. It writes an event. It may notify a security team. The agent continues with the same tools and privileges it had before the warning.

That is visibility, not control.

If a possible injection does not alter the available action set, the organization is asking a later human process to contain a machine-speed decision. The alert may arrive after the email was sent, the record was modified, the file was retrieved, or the tool call completed.

This does not mean every suspicious string should shut down the workflow. Overblocking can make AI systems unusable and encourage teams to bypass the control.

The better design is proportional response:

  • Low-risk privacy exposure can be tokenized while the workflow continues.
  • A suspicious URL can be defanged before analysis.
  • Untrusted text can be summarized with all action-capable tools removed.
  • A suspected injection can force the agent into read-only mode.
  • A consequential request can require human approval.
  • Clearly malicious content can be blocked and preserved for investigation.

CaneCorso’s allow, sanitize, tokenize, and block outcomes support that more nuanced response. The surrounding application determines which tools remain available and which approvals are required.

That separation of responsibilities matters. The content control should not silently become the identity provider, and the identity provider should not pretend it understands prompt semantics.

Test Whether the Boundary Actually Holds

No prompt-injection control should be treated as infallible.

NIST’s March analysis of a large agent red-teaming competition reported more than 250,000 attack attempts by over 400 participants across 13 frontier models. At least one successful hijacking attack was found against every target model. NIST’s takeaway was not that agent use is impossible. It was that evaluations must evolve with adversaries and that comparative red teaming provides information ordinary static tests may miss. (NIST CAISI, March 23, 2026)

The joint Careful Adoption of Agentic AI Services guidance from agencies including CISA, NSA, and the national cyber centers of Australia, Canada, New Zealand, and the United Kingdom recommends input validation and sanitization, prompt-injection filters, continuous evaluation, fail-safe defaults, human control points, runtime monitoring, and regular attempts to bypass safeguards. (Joint guidance, May 1, 2026)

For one representative agent workflow, test the complete control chain:

  1. Establish a normal task and record the intended agent identity, data sources, tools, and allowed actions.
  2. Introduce hostile instructions through realistic sources such as a ticket, email, document, RAG entry, API response, or tool result.
  3. Confirm that CaneCorso identifies, sanitizes, tokenizes, or blocks the content according to the workflow policy.
  4. Confirm that the resulting disposition changes the agent’s available authority as designed.
  5. Attempt a high-impact action after the risk signal and verify that the authorization layer blocks it or requires human approval.
  6. Verify that the agent cannot recover removed capability by rewriting the request, changing language, splitting the instruction across sources, or calling a different tool.
  7. Inspect the record. Confirm that responders can identify the source, agent, content decision, policy state, requested tool, approval, and final outcome.
  8. Test the failure mode. Decide what happens if the control plane, identity service, logging path, or human approval channel is unavailable.

CaneCorso’s Injection Scanner is designed to help exercise these controls before deployment, after changes, and during assurance reviews. The test still needs to cover the full application behavior. Detecting the input is only one link; the system must also constrain the agent and prevent the unsafe action.

Measure the Decision, Not Just the Detection

Counting detected prompts can be misleading.

A high count may indicate active attacks, a noisy policy, a research feed that contains legitimate attack examples, or ordinary content that resembles instructions. A low count may indicate clean inputs—or weak coverage.

More useful measures connect detection to outcome:

  • Percentage of agent actions carrying a recorded content-risk disposition.
  • Time from suspicious input detection to reduced agent authority.
  • Percentage of high-impact actions attempted after a risk signal and correctly blocked or escalated.
  • Percentage of sensitive values tokenized or redacted before model exposure.
  • Rate of safe workflows disrupted by false positives.
  • Rate of risky workflows allowed because a control failed open, timed out, or was bypassed.
  • Percentage of human reviews with enough evidence to explain the requested action and policy decision.
  • Results of adversarial validation before and after model, prompt, tool, policy, or data-source changes.
  • Time required to reconstruct which content influenced a consequential agent action.

The important metric is not “How many injections did we detect?”

It is:

When the context became less trustworthy, did the system reduce authority before harm occurred?

What Security Leaders Should Do This Week

Choose one agent that consumes external or mixed-trust content.

  1. Name the consequential actions. Identify every send, write, execute, disclose, purchase, delete, or privilege-changing capability available to the agent.
  2. Map the influence paths. Include users, email, documents, RAG sources, web content, API responses, tool output, memory, and other agents.
  3. Add an independent inspection point. Put CaneCorso or an equivalent AI Application Firewall before untrusted content reaches the model and before consequential output reaches downstream logic.
  4. Define the four states. Document what Normal, Constrained, Review, and Stop mean for this workflow.
  5. Bind state to capability. Make the orchestrator and authorization layer remove tools, require approval, or block execution as risk increases.
  6. Run an adversarial test. Use realistic indirect prompt injection attempts and verify the action is prevented—not merely detected.
  7. Preserve the reason. Record what influenced the agent, which policy applied, what authority changed, who approved any exception, and what finally happened.

Agent identity answers who is acting.

An AI Application Firewall helps determine whether the content influencing that agent is safe enough to use.

Authorization must connect the two.

If an agent can encounter hostile context while retaining its full authority, the architecture is trusting the moment when it should become most skeptical.

Prompt injection should not merely create an alert.

It should change what the agent is allowed to do.


Put CaneCorso in Front of Your AI Workflows

CaneCorso™ gives organizations a shared AI Application Firewall for email, document AI, RAG, support workflows, copilots, and agent-driven automation.

MicroSolved, Inc. can help you:

  • Map the content, privacy, tool, and authorization boundaries in an AI workflow.
  • Pilot CaneCorso against a representative production use case.
  • Configure allow, sanitize, tokenize, and block policies.
  • Test prompt-injection defenses with CaneCorso’s Injection Scanner.
  • Connect runtime decisions and evidence to monitoring, review, and audit processes.
  • Design constrained, human-review, and fail-safe operating states for AI agents.

To discuss a CaneCorso pilot or schedule a practical review, contact MicroSolved at info@microsolved.com or +1.614.351.1237.

Relax. We’re on watch.

AI tools were used as a research assistant for this content, but human moderation and writing are also included.

 

Three Field Guides for the Security Problems Showing Up Right Now

Over the last several months, we have released three longer-form guides around problems we keep seeing in real environments:

  • The Antifragile SOC Book
  • The AI Agents Management Framework
  • The Security Leader’s Operating System

They cover different areas, but they share a common theme: how to operate effectively as complexity, automation, and pressure increase.

Books

Why We Created These Materials

These guides did not begin as content projects.

They came from working problems.

Over the years, we have accumulated a lot of models, processes, questions, and operating approaches that we use when helping organizations solve difficult security problems. Some came from client work, some from research, and plenty came from lessons learned the hard way.

Much of that knowledge lived in conversations, notes, presentations, and experience.

We decided it was time to write more of it down.

Not as theory, and not as another collection of cybersecurity predictions.

As practical field material people can use, challenge, adapt, and apply.

Continue reading →

The Top Three Things Security Teams Need to Understand About AI-Driven Attacks

The easiest mistake security teams can make about AI-driven attacks is to imagine a completely new species of adversary.

That is not what most organizations are facing.

The objectives remain familiar: steal credentials, gain access, persist, move laterally, collect data, commit fraud, extort the victim, or conduct espionage. What AI changes is the economics of the attack. It reduces the time, expertise, and attention required to move from one step to the next.

That distinction matters because it tells defenders where to act.

Google Threat Intelligence Group reported this week that it has observed adversaries moving beyond basic prompting into agentic workflows and AI-enabled automation. In one Q2 2026 case, a threat actor used an AI coding chatbot and agent instructions to help plan, build, and execute a mass credential-harvesting campaign in less than six hours. The framework automated scanning, troubleshooting, and IP rotation with limited human involvement. (Google Threat Intelligence Group, September 8, 2026)

That does not mean every attacker has become an autonomous machine. It means the constraints that used to slow attackers down are weakening.

OnlineScammer

Security teams need to understand three things.

Continue reading →

AI Agents Need Autonomy Budgets, Not Just Governance Policies

Most organizations are granting AI agents authority faster than they are defining the limits of that authority.

That is the problem.

We have already started treating AI agents as digital workers. That is the right mental model. An agent that can access data, call tools, trigger workflows, generate artifacts, influence decisions, or alter enterprise state is not just another application. It needs identity. It needs boundaries. It needs oversight. It needs evidence. It needs a human owner. It needs a kill switch. That has been the right foundation for agent governance.

But it is not enough.

There is another question that needs to be asked much earlier:

How much damage is this agent allowed to cause before a human must approve the next action?

Not theoretically.

Not in vague risk language.

In actual economic terms.

How much money can it spend?
How many systems can it change?
How many records can it touch?
How much customer impact can it create?
How much privacy exposure can it cause?
How much reputational risk can it accumulate?

If we cannot answer those questions, we have not governed autonomy.

Continue reading →

Rational Security in the AI Era: How Attackers Are Evolving and How We Must Respond

The weaponization of artificial intelligence by cybercriminals and nation-state actors has crossed a critical inflection point. We no longer live in a world where we can rely solely on traditional perimeters; the threat landscape has fundamentally shifted into what we might call “Extremistan,” where the speed and scale of attacks demand a completely new level of resilience.

SadKitty

At MicroSolved, our mission is to provide rational cybersecurity for an irrational world. To do that effectively, we must look unflinchingly at the data.

The Problem and the Metrics

The numbers tell a stark story of industrialization at machine speed. According to recent threat reports, AI-enabled adversaries increased their attack volume by 89% year-over-year. More concerning is the velocity: the average eCrime breakout time has collapsed to just 29 minutes, with the fastest recorded intrusion moving from initial access to lateral movement in a staggering 27 seconds.

The financial impact is equally severe. The FBI IC3 recorded over 22,000 AI-related complaints with adjusted losses exceeding $893 million in 2025 alone, including tens of millions lost to AI-enabled Business Email Compromise (BEC). AI is accelerating attack speeds by 4x, making human-speed incident response no longer viable.

Continue reading →

AI Agents Are Already Working for You. Who’s Managing Them?

AI Agents Are Not Applications. They Are Digital Workers.

Most organizations are adopting AI agents faster than they are learning how to govern them.

That is the problem.

A chatbot that answers questions is one thing. An AI agent that can access business data, use tools, trigger workflows, generate artifacts, make recommendations, or alter enterprise state is something else entirely.

At that point, the organization is no longer just deploying software.

It is introducing a new kind of operational actor.

That actor needs identity.

It needs boundaries.

It needs oversight.

It needs evidence.

It needs a human owner.

It needs a kill switch.

In other words, AI agents must be managed more like digital workers than ordinary applications.

AIAgentBanner

Continue reading →

Why My AI Agents Needed CaneCorso as a Security Control Plane

AI agents are powerful because they can read, reason, summarize, decide, and act across a wide range of information sources.

That is also what makes them dangerous.

The more useful an agent becomes, the more likely it is to consume data I do not fully trust. Emails. Newsletters. RSS feeds. API responses. Documents sent as attachments. Social media. YouTube transcripts. Scraped search results. Web pages. Translated content. Random bits of text pulled from places where I do not control the author, the formatting, the intent, or the payload.

That is a very different security model than the one most of us are used to.

In traditional applications, we spend a lot of time separating code from data, users from administrators, trusted networks from untrusted networks, and internal systems from the internet. With LLMs and agents, all of those boundaries start to blur. Instructions, context, content, and intent all arrive in the same stream. The model has to reason over that stream, and the agent has to decide what to do with the result.

That is exactly why I wanted a security control plane in front of my own AI agents.

For me, that control plane became CaneCorso™.

CaneCorsoAI

Continue reading →

CaneCorso™ and the Real Problems AI Is Creating for the Business

AI didn’t sneak into the enterprise.

It walked in through productivity.

Email triage. Document handling. Support workflows. Internal copilots. Retrieval systems. Early agentic use cases. All of it made sense at the time. All of it still does.

But something changed along the way.

We didn’t just adopt AI—we embedded it into workflows that can influence decisions, expose data, and take action.

That’s where the problem starts.

And it’s exactly where CaneCorso™ is designed to operate.

CaneCorsoAI


AI Risk Isn’t a Model Problem — It’s a Workflow Problem

There’s a persistent misunderstanding in the market right now.

Most conversations about AI security still center on the model—what it knows, how it behaves, whether it can be tricked.

That’s not where the real risk lives.

The real risk shows up when:

  • Untrusted content enters a workflow
  • That workflow uses AI to interpret or transform it
  • And the output influences business operations

That content might come from:

  • Email
  • Documents
  • OCR pipelines
  • Retrieved knowledge (RAG)
  • Support tickets
  • External data sources

Once it’s in the workflow, it’s no longer just data.

It’s influence.

CaneCorso™ exists to control that influence—before it becomes an operational problem.

Continue reading →

Introducing CaneCorso: An AI Application Firewall Built for Real Workflows

AI has officially crossed the line from experiment to infrastructure.

Email flows into copilots. Documents feed RAG pipelines. Support tickets trigger agents that can take action. The convenience is real—and so is the risk.

What hasn’t caught up is security.

Most security models were built for a world where inputs were predictable and trust boundaries were well-defined. That world doesn’t exist anymore. Today, untrusted content flows directly into systems that can reason, decide, and act.

That’s exactly where things get interesting—and dangerous.


When Good Data Carries Bad Instructions

One of the biggest misconceptions about AI security is that it’s a model problem. It’s not. It’s a workflow problem.

Attackers don’t need to break in anymore. They ride along with legitimate data—emails, PDFs, tickets, knowledge base entries—and inject instructions that your AI system may interpret as truth.

Think about what that means in practice:

  • A support ticket that contains hidden instructions
  • A PDF with embedded prompt injection
  • A knowledge base entry that poisons RAG outputs
  • An approval workflow manipulated through summarization

Layer in human behavior—blind trust, over-privileged access, weak validation—and you’ve got a system primed to fail in ways that traditional controls simply won’t catch.

CaneCorsoAI


A More Rational Approach to AI Security

CaneCorso™ takes a different path.

Instead of trying to block everything suspicious (and breaking workflows in the process), it follows what’s described in the Rational AI Security model —security that behaves more like an immune system than a wall.

That means:

  • Detecting and isolating threats without stopping the system
  • Treating all inbound content as untrusted by default
  • Preserving business continuity while reducing risk
  • Producing measurable, auditable outcomes

This isn’t theoretical. It’s a direct response to how AI systems actually behave in production.


One Control Plane for AI Workflows

At its core, CaneCorso gives you a shared AI Application Firewall—a single control plane that sits between your workflows and your models.

Instead of every team building its own brittle filters, you get consistent, reusable protection across:

  • Email triage and analysis
  • RAG pipelines and knowledge systems
  • Document AI and OCR ingestion
  • Support and ticketing workflows
  • Agent-driven automation

The platform delivers:

  • Runtime decisions: allow, sanitize, tokenize, or block
  • Privacy controls: redact or tokenize sensitive data before model exposure
  • Audit-ready logs: reasons, scores, and evidence you can actually use
  • Adversarial validation: Injection Scanner proves controls before and after deployment

This isn’t just about stopping attacks—it’s about making security operationally usable.

Continue reading →

Update on PromptDefense Suite and AI Security Research

Last week, I discussed why and some of how we built the new PromptDefense Suite. 

This week, we are discussing the product’s future internally and how we might go to market. This is mainly due to two new capabilities we have built into the product. 

The first is an API and workflow automation mechanism. This allows organizations to stand up a single instance of PromptDefense and then use it to protect multiple AI/agent workflows. The code no longer has to be embedded directly in the project; instead, all defensive capabilities and logging can be accessed via an API instance. The API is robust and supports API key restrictions that tie into a rules engine, so that different workflows can have different trust models and actions pre-assigned in an audit-friendly way. 

Secondly, we have developed a licensing mechanism that covers protected workflows and skips the per-seat, per-token models that seemed too confusing for most firms looking for these kinds of tools. They told us they wanted a simpler licensing approach, and we developed a new licensing mechanism to make it easy, manageable, and auditable. Our testers have been calling it a win! 

As we continue with the beta-testing process and lock down our decisions about where the product is going, the news that drove us to create it continues to flow in. More of our clients are working on agents and AI-integrated workflows, which require this level of protection. While we continue to develop PromptDefender, we are also working to develop and release extended frameworks for AI model, agent, and product management, along with policies, procedures, and vendor risk assessment tools for these frameworks, for our vCISO clients. We’re also busy researching ongoing compliance implementation for AI workflows and agents, and should have more on that shortly. 

In the meantime, if you want to discuss AI or agent security, risk management, or other relevant topics, please reach out. We would love to talk with you and help align our modernization capabilities with your emerging needs. You can always email us at info@microsolved.com or call us at +1-614-351-1237. 

As always, thanks for reading. Stay safe out there, and stay tuned for more updates.