As agentic systems have become more powerful, concerns have shifted beyond adversarial abuse to include “misalignment” and other unwanted model behavior. Agents can take unsafe actions during legitimate workflows, or adversarial inputs can cause agents to execute unauthorized operations. As agents gain access to more tools, data, and systems, the potential impact of those actions grows.
For organizations deploying these systems, the practical questions are what an agent is authorized to do, what prevents actions outside that authority, and how to verify that those controls work regardless of what actions the agent might take to circumvent them.
In this post, we will explain how NVIDIA’s AI security team distinguishes unsafe individual actions from misaligned behavior, what can cause each, and how to limit their impact. The guiding principle across all these classes of unwanted behavior is to **assume misalignment: **design the system under the pessimistic assumption that the agent will eventually try to take actions that fall outside the scope of the authorized task, or violate system policy. Security boundaries must remain effective regardless of what actions the agent takes.
Unsafe actions and misaligned behavior
The NVIDIA AI security team defines alignment broadly: an aligned system performs actions consistent with a user’s intent within the boundaries established by the system designer. When the agent takes actions directed toward a goal that is contrary to the user’s stated intent, or performs operations outside the boundaries established by the system owner, we describe the behavior as misaligned.
An unsafe action describes a specific behavior, without requiring that the system be misaligned in proposing it. While misaligned systems will often initiate unsafe actions, a single unsafe action does not by itself establish that the agent is misaligned. A system can take (or be induced to take) unsafe actions even while executing a legitimate task. The corresponding security risk of an unsafe action depends on the agent’s access, the actions it can execute, and the effectiveness of the controls around it.
Unsafe actions can result from incomplete context, underspecified goals, mistaken assumptions, or adversarial inputs such as prompt injection.
- Example 1: When not provided with sufficient context about data security requirements, an agent might upload an internal file to an external service to complete a legitimate task, even though the transfer violates a data-handling policy.
- Example 2: When an agent becomes misaligned due to inadvertent factors like context decay or drift, it may attempt to perform actions that violate security policy that was originally placed in its context, resulting in potentially risky actions.
- Example 3: At the extreme end, an agent’s goal definition may become corrupted or adversarially defined, such that it is harmful to the user or system owner. In such cases, they can approximate the behavior of an insider threat: using all tools that have been made available to it and composing them in potentially unexpected ways to achieve its new misaligned and hostile goals. Misaligned behavior can also persist across a sequence of actions, with different tools combined in ways that produce an unauthorized outcome.
These examples illustrate how misalignment may arise in both inadvertent and adversarial patterns leading to unsafe action.
Figure 1. Examples of misaligned behavior and unsafe actions under both adversarial and non-adversarial conditions.
The divisions in Figure 1 are not precise: they can overlap, and their consequences depend on the agent’s access to tools and data, and the effectiveness of the controls around it. Different causes above require different approaches to prevent them from occurring. However, regardless of the cause, the downstream security impacts can be mitigated by the same controls: deterministic policy that restricts the agent from taking unsafe actions, regardless of the agent’s intent or attempted course of action.
Unsafe actions
Attacks that trigger unsafe actions
Adversarial techniques such as prompt injection can cause an agent to take an unsafe action that furthers attacker goals without altering the system goal definitions. For example:
- An agent rejected a direct request to execute a malicious command, but subsequently installed a package that executed the payload during installation. The agent’s dependency-installation workflow gave the attacker another route to code execution that used a capability the system was designed to provide.
- Crafted inputs caused an analytics workflow to generate code that both bypassed sanitization controls and achieved remote code execution.
Both examples show why evaluating the apparent intent of a request is often insufficient to establish that its execution stays within the authorized task and follows system policy, and that adversaries can cause agents to trigger actions with harmful effects even if they are not “misaligned”.
Unsafe actions during legitimate tasks
Unsafe actions also happen without adversarial input. In the absence of explicit policy guiding agent activities, unsafe activities without prohibition may be evaluated as permitted. An agent may take alternate, unanticipated actions to fulfill its legitimate goal. Attempting to exhaustively enumerate policy can reduce these failures, but even comprehensive policy related instructions may be lost during context compaction or omitted when tasks are delegated, as described below.
We have observed agentic systems initiating unsafe actions in this fashion during legitimate workflows within NVIDIA.
- **Example 1: **During a system diagramming task, one user observed an agent attempt to download and install a third party application that would have sent their data to an external host as part of the diagram creation process.
- **Example 2: **In testing, we have also observed agents commit internal information (such as inference servers and model names) to external version control systems, or initiate the transfer of files across local filesystem boundaries using alternative routes like HTTP servers and external third party proxies.
These examples illustrate the importance of making assumed policy explicit, and of placing controls outside of the scope of the agent with enforcement across all available tools and protocols that align with actual policy requirements, not just ones described in the system instructions.
These cases share three common traits:
- Agents executed legitimate (or apparently legitimate) tasks using operations that were not anticipated by the designers.
- The workflows exposed gaps in policy context, enforcement, or the handling of indirect effects that could be identified and exploited by the agents.
- The unauthorized actions could compromise confidentiality, integrity, or availability of resources outside the scope of the original agentic deployment.
Misalignment
Misaligned behavior can reflect both how a model was trained, as well as how runtime inputs and the execution environment affect its outputs and actions. These influences interact. For security teams, a particular concern is repeated attempts to execute unauthorized operations through different paths after a policy denial.
Model training and fine-tuning
Model inference services – which we define to include any deterministic guardrails placed behind the model’s inference API and not accessible to a client – are the “engine” of any agentic system, and the behavior of the model reflects both data observed during pretraining, and explicit instructions used during post-training. A model that is not aligned with the system or user intent will be both difficult and expensive to realign to those intents long term without additional training. These unwanted behaviors can show up in a capability denial form, for instance, during the HuggingFace / OpenAI incident in which frontier models refused to help HuggingFace defenders triage or defend against cybersecurity attacks due to cyber capability enhancement risk.
Models can also be adversarially misaligned as a result of training or fine-tuning. Open weight models have exhibited specific security relevant behavior in response to one or more predefined triggers. Other fine-tuning operations for purposes such as removal of safety training behavior or insecure code generation have also been observed to have side effects on model behavior in unrelated domains. Because of the importance of model behavior to the proper functioning of the agentic system, models should be reviewed through their model cards and task- and risk-specific evaluations against the intended deployment’s objectives and constraints.
Models may also exhibit risky behavior during training and evaluation. OpenAI used models with deliberately weakened safety guardrails during cybersecurity evaluations, in which agents bypassed network isolation controls and ultimately compromised Hugging Face infrastructure. Another OpenAI report describes unauthorized communication during training of an unreleased model, in which agents that could not share a workbook through the intended local filesystem used a public file-hosting service despite instruction to use only local files.
Runtime context and delegation
Runtime behavior is also affected by factors such as system prompts, skills, MCPs, automatically included files such as CLAUDE.md or AGENTS.md, tool descriptions, and external data introduced via retrieval augmented generation (RAG) or some other retrieval method. The use of system-prompt and AGENTS.md style “guardrails” – where the model is instructed to follow specific guidelines or avoid particular actions – are a common attempt to mitigate this risk.
Unfortunately, as agent context lengths grow, they often “dilute” the attention paid to such instructions, resulting in reduced effectiveness. In some cases, user-provided guardrails and instructions can be lost completely in a context compaction step, when the model must reclaim working memory by summarizing previous interactions. Both of these forms of “context drift” can result in later actions that no longer conform to task requirements or system policy. Often subagents – created for specific tasks by a parent workflow – receive incomplete context, which may also lead to misalignment. NVIDIA has observed unsafe behavior in long-running workflows in testing, including agents spontaneously generating prompt-injection payloads to return to the main agent attempting to exfiltrate SSH keys, or attempting to use local scripting libraries to gain direct access to policy-gated resources. Task delegation should therefore preserve the applicable constraints and limit each subagent’s permissions to its assigned work.
Jailbreaks and limits of behavioral guardrails
A more widely known adversarial influence, conventional jailbreaking, involves specially crafted user-provided inputs that attempt to override a system’s instructions. The NVIDIA AI Red Team has used “frog boiling” techniques to progressively elicit operations outside the system’s intended constraints. In some cases, agents were prompted to route traffic around network access controls, including through third party proxies and gaps in virtual private cloud (VPC) isolation.
Prompt based behavioral guardrails, while they can act as defense in depth, generally operate at the intent level of models, and so inherit their fragility. LLM-as-a-judge patterns – which also frequently offer significant defense in depth benefits – rely on LLM model outputs to classify or evaluate an action. As such, they face the same risks of context drift, malicious influence, and unpredictability as the agents they are intended to constrain.
Mitigating the impact of unsafe actions and misalignment
Enforce policy outside the agent
The behaviors above all share the same basic architectural risk factor: an agent harness turns text output into operational authority, and those text outputs may not conform to the authorized task or system policy.
The tools exposed by a harness can also enable operations beyond their apparent scope. For instance, in internal experiments we observed an agent that only had access to a file write tool both rewriting its own tools to create a bash execution tool, and writing commands to executable locations such as `~/.bashrc` or scripts referenced in the user crontab to trigger arbitrary code execution. This illustrates that controls the agent can modify are not sufficient to constrain its behavior.
Two architectural requirements follow:
- Bound the operational authority of the translation from text to operation via deterministic policy that does not rely on model behavior, and;
- Enforce that policy outside of the control plane of the agent,** **preferably in a different trust domain that the agent has no authority to modify, disable, or bypass.
For the file-upload example above, enforcement must cover both direct uploads and transfers through permitted intermediaries. The restriction must hold when the operation is issued through a different tool, generated code, or a delegated task.
Mitigating with NVIDIA Open Agent Safety Platform
Tools such as NVIDIA OpenShell, part of the NVIDIA Open Agent Safety Platform, allow for user-configurable controls to limit the filesystem and network access of the processes inside the sandbox, as well as isolating the agent process from the host. This effectively constrains the range of possible model actions without relying on the nondeterministic controls described above, while also limiting the potential blast radius of those models if they do take unsafe actions.
For some of the examples listed above, here’s how OpenShell can help by applying deterministic controls to the environment in which the agent runs:
Example 1: System diagram task
As described above, an agent attempted to construct a network diagram by installing third-party software that would have sent internal data to an external endpoint.
OpenShell’s filesystem and network policies provide controls at several stages of this workflow:
- Restricting the download: Network policy can block the download if the hosting destination is not permitted. Filesystem policy can also prevent the agent from writing the binary to a restricted location.
- Isolating execution: If the software is downloaded and executed, it runs inside the sandbox. Host files are available only where they have been explicitly exposed to the workload and permitted by its filesystem policy.
- Restricting file access: The downloaded software remains subject to the sandbox’s filesystem policy.
- Restricting external transfers: Network policy can block attempts to send data to destinations that are not permitted. If the external service is already allowed, OpenShell middleware could be used to add an additional authorization check based on the content of the payload.
These controls provide separate opportunities to interrupt the unsafe workflow. Their effectiveness depends on which files, destinations, and operations the configured policy permits.
Example 2: Malicious package dependencies
This attack is particularly difficult because it exploits a standard agent workflow: installing dependencies from GitHub or package repositories that may already be permitted.
- Initial exploitation may still occur: A malicious package may download and execute malicious code through an allowed dependency-installation workflow. The host remains isolated from the agent workload so any impact is contained by the OpenShell sandbox.
- Network access remains constrained: Code executed by the package remains subject to OpenShell’s network policy.
- Filesystem access remains constrained: The malicious code remains subject to the sandbox’s filesystem policy. It may still read files the policy makes readable or modify files the policy makes writable, including project files the agent needs for its task.
OpenShell can limit the consequences of a malicious dependency, but files and services deliberately made accessible to the workload remain within its reach. Policies should therefore grant only the access needed for the task.
NVIDIA Sentry adds an additional layer of security, by both moving the monitoring and enforcement point to a completely different trust domain, and acting as an internal proxy for all network traffic (including LLM traffic) which allows for direct inspection of chain-of-thought and tool calls, allowing for detection of misalignment at the semantic layer, and immediate detection of risky tools.
Additional safeguards and validation
Finally, protecting agent code and configuration with filesystem write restrictions limits unauthorized changes that could expand its capabilities. The policy and enforcement configuration must also remain outside the workload’s write access.
Controls over network and filesystem access cannot, by themselves, validate every permitted business operation. An agent authorized to update customer records can still make an incorrect change. Application-specific validation and approval requirements remain necessary.
By layering policy enforcement and monitoring outside the agent workload, and using those capabilities to establish detection and response capabilities, it is possible to dramatically reduce the risks of both adversarial attacks against agentic systems and unsafe actions during legitimate workflows. OpenShell provides a clear boundary on agent capabilities within an isolated environment, while NVIDIA Sentry provides both full introspection into all available agent network traffic, including any data available in LLM interactions such as tool calls, tool definitions, and reasoning traces, permitting complete logging and monitoring, suitable for building a strong detection and response capability.
Applying these controls as a first line of defense allows users to balance the flexibility and capabilities of agentic workflows with the risks of broader access and more consequential operations. Before expanding an agent’s access, operators should test whether the configured controls continue to block unauthorized operations under the new access model, record both relevant telemetry and enforcement decisions, and support intervention. Those tests should also measure the effect on legitimate task completion and performance.
Learn more about NVIDIA OpenShell and NVIDIA Sentry, part of the NVIDIA Open Agent Safety Platform here. Learn more about automated testing of agent plugins before deployment with NeMo Helix. For more research and insights from our AI security teams, read our blog.