Security awareness programs are built on assumed behavior. But as Nik Kale, CISSP, explains, the rise of AI agents has significantly altered the realities of operational behavior and process, because AI-based activity doesn’t always mimic the actions we expect from real people.
Disclaimer: The views and opinions expressed in this article belong solely to the author and do not necessarily reflect those of ISC2.
Each security awareness program is premised on the assumption that a human will pause before acting. People check sender addresses, hover links before they click, consider the reasons finance is asking for gift cards instead of cash. In theory, training, phishing simulations and break room posters are all intended to introduce that gap between arrival and execution.
I have previously written for ISC2 Insights that agents would compel a new workforce model, as there is no owner in the organizational chart when an agent executes its function. That was an accountability argument. This one is for the people who run awareness programs, because their discipline has a problem it hasn't noticed yet.
AI agents have removed that gap. They don't hover. They read and they act. The gap between those two things is measured in milliseconds.
The agent equivalent of phishing is indirect prompt injection. A harmful instruction arrives in a document that was just retrieved, a web page, the body of an email or the output of a tool called by the agent. The model weighs every token in its context window without a reliable way to distinguish what it was told to do from what it was told about. OWASP put prompt injection at number one in the first edition of its Top 10 for LLM Applications in 2023. It was still number one in the 2025 revision and it is still number one in the 2026 edition published this August. Three editions, same top entry.
The listings are not the only evidence. In March, the Center for AI Standards and Innovation of NIST published an analysis of a large public red-teaming competition that Gray Swan and the UK AI Security Institute co-hosted: over 250,000 attack attempts by more than 400 participants against 13 frontier models, with at least one successful hijacking attack against every model tested. There was no uniform relationship between the models' overall capability and their security performance.
There are no training modules for this. There's no simulation your agent can fail on a Tuesday and learn from by Thursday. Even for humans, training is a decaying control. A study of 409 employees, presented at USENIX SOUPS 2020, showed that phishing-awareness training still produced a significant improvement in detection four months later and no significant improvement at six months. In other words, for humans it has a half-life. For an agent there is nothing to decay, because nothing was learned in the first place.
Where I Ran Into This
I design an enterprise support service where agents do root cause analysis on customer issues. These agents read customer communication in the form of diagnostic bundles, configuration dumps, log excerpts, case notes and screenshots of terminal output. All this data is untrustworthy by definition because it comes from outside the trust boundary. Almost nobody has ever bothered to sanitize a device log, because until recently a device log was inert.
It is no longer the case that the log is inert. Any line of text in the bundle now sits in the context window alongside the instructions I wrote. The model has no way of differentiating which came from me and which came from a file uploaded by an unauthenticated user in a different time zone.
I initially thought we were dealing with a model problem. The next best thing was to try to write better system prompts, where we told the model to ignore any instructions that it read in retrieved content. We put a warning banner around the untrusted block and repeated the rule at the end of the prompt, since recency helps. All this cuts down on the naive attempts and does close to nothing against the ones that matter, because you cannot instruct your way out of an architecture where instruction and data travel on the same channel. We were writing awareness training for a system that has no capacity for awareness.
What actually moved the risk was giving up on what the agent reads and getting strict about what the agent can do. Split read paths from write paths. Take no write actions based on retrieved content alone. Every tool must be limited to the few actions needed to make it useful. Cap egress. Log calls to each tool, including a reference to the content that triggered the tool rather than a copy of it as well as the authority used and the policy that was decided, to prevent secrets and personal info in the logs. Valid credentials do not equal valid intent, hence each of the controls is placed on the action and not on the identity. None of those controls is taught. It is architecture. It is done during the design review or not done at all.
The Half That Is Still a Human Problem
One half of the awareness function moves into architecture and stops being awareness. The other half stays human, but that half of the population is the one that the current awareness systems don't cover.
I'm talking about builders who connect agents to systems.
I spent some time last year writing production readiness rules for an open-source scanner that inspects Model Context Protocol servers. Eventually I had 20 heuristics. When I was writing those rules, I read some real-world implementations of servers. What stuck with me was how the tools described themselves. A tool's description is prose that a developer types quickly. A model reads the prose and acts on it. It is executable text sitting in a field that no code review process treats as code.
Adding a connector is a ticket. It is usually done on a Thursday afternoon at the request of someone from product. It takes about 20 minutes. In that time, nobody asks about the connector: what it can read, what it can write, which account it uses, whether the scope was inherited from a partnership integration that was built for a different purpose in 2021 or who can disable it at 2 a.m. The question I ask in these reviews goes past what the agent is authorized to do, to what information is authorized to influence how that authority gets used. The engineer is not careless; it is just a question that has not been asked of them and no phishing simulation has ever been done against them in this regard.
I see the same gap from the standards perspective. In OASIS CoSAI Workstream 4, I work on secure design for agentic systems, including how authority moves along a delegation chain when one agent calls another. The specifications are getting better quickly. What they cannot do is force the scoping conversation to occur. A control described in a standards document and which has never been raised in a design review is a control that simply does not exist.
Four Additions to the Curriculum, Mapped to the CISSP CBK
If awareness means putting a question in front of someone while they can still answer it, then these are the four questions, and they belong to the builders.
- Before connector connection, review scope. Security and Risk Management, Asset Security, Identity and Access Management. Each agent connector has a recorded answer to four questions prior to go-live: what it reads, what it writes, the identity it acts under and who revokes it. Each answer is a single paragraph. The value is more in asking the questions.
- Retrieved content is input, never instruction. Security Architecture and Engineering, Software Development Security. Engineers should be able to justify why a prompt-level remediation is not a control and should also be able to explain where the true bound is: tool, permission, write path. This is the one idea that sticks.
- Agent action logging that explains 'why.' Security Operations. Each tool call should be logged with a reference to the content that invoked it. Each record should include the authority under which the call was made, the policy decision that was made and the outcome. If your team is unable to analyze agent action logs and determine the reason for the action, you have no detection story for this class of attack, only a record of consequences you happened to notice.
- Injection cases in the release pipeline. Security Assessment and Testing. Plant instructions inside retrieved documents and run them as test cases, with results that can block a release. This is the closest example of a phishing simulation from the agent's perspective and is unique in that a pass/fail result can be placed behind a gate.
Three Things You Can Run This October
First, list your connectors. One page. Each agent in production, each tool it can call, whether each tool reads or writes and whose credentials it uses. Budget two hours. In most organizations, the exercise fails to complete and the incompleteness is the finding.
The second exercise is worth an afternoon: run one injection test against your own system. Plant a harmless instruction inside a document your agent retrieves in the normal course of work. Make the effect visible, something like telling it to append a specific word to its answer, so you can see in the output whether the instruction was followed. The result worth measuring is layered. The model should ideally ignore your inserted step. If it does not, policy blocks the tool call and the attempt is recorded and alerted, so nothing consequential happens, and this is defense in depth. It usually costs only one afternoon to implement, and it gives your engineering management a chance to see a risk in the context of a concrete example.
Lastly, add one question to your change process. For every change that adds an agent to a system or expands the agent's capabilities, ask “if this content were malicious, what might the agent do with it?” Ensure there is a named owner, a written answer, and there are no exceptions for internal-only systems. That last exception is where this will bite people first.
The Part That Does Not Transfer
The reason why we have Cybersecurity Awareness Month is that approximately 22 years ago, the simplest way into a company was a phishing link in an email. It still is. What has changed is who clicks.
The reflex that we developed in people over those 22 years cannot be built into software, and it will never be, because the pause we were developing was the entire mechanism. What can be built is the question, shifted upstream to the people wiring these systems together. They were never a part of the training population. This October is a reasonable time to put them there.
Nik Kale, CISSP, has nearly two decades of experience in enterprise software, AI platforms, cloud security, identity and access management. He has held principal engineering roles, with responsibility for designing AI security and enterprise systems at scale. His cybersecurity work spans the Coalition for Secure AI through OASIS, the IETF OAuth working group, ACM CCS 2026 and IEEE S&P SAGAI program committees.

