The rise of the agentic enterprise
How autonomous workflows could reshape decisions, coordination and productivity�and where human oversight remains essential as AI moves from assistance to execution.
Read articleHow does your AI behave when someone tries to break it?
Conventional evaluation asks whether an AI system succeeds with representative inputs. Adversarial evaluation asks what happens when an actor deliberately manipulates the data, instructions, context or tools on which that success depends. Both are necessary: normal accuracy says little about the security of an exposed workflow.
Build the threat model around assets and actions, not only the model. Map training and fine-tuning data, retrieval sources, prompts, identities, tool permissions, outputs and external components. NIST�s 2025 taxonomy distinguishes evasion, poisoning, privacy and misuse attacks across the AI lifecycle, making clear that risk extends far beyond a malicious chat message.
Exercise realistic attack paths: instructions hidden in a document retrieved by the system; an email that attempts to redirect an agent; sensitive-data extraction through repeated queries; unsafe tool chaining; output that exploits downstream software; resource exhaustion; and a compromised supplier. Test whether controls limit the blast radius when the model follows the attacker.
Red teaming should be continuous and independent enough to challenge design assumptions. Preserve attack cases as regression tests, seed canary data, log tool calls and privilege changes, and rehearse containment with security, engineering and business owners. OWASP identifies excessive functionality, permissions and autonomy as root causes of damaging agent behaviour; prompt tuning alone cannot remove them.
Track attack success rate, unauthorised actions, data exposure, detection time, containment time and recovery quality. No defence is foolproof, so resilience matters as much as prevention. A trustworthy system is not one that never encounters hostile input; it is one whose boundaries remain enforceable when that input arrives.
Related macro
Articles
How autonomous workflows could reshape decisions, coordination and productivity�and where human oversight remains essential as AI moves from assistance to execution.
Read articleWhat separates companies that scale AI from those that accumulate experiments�and how operating models, economics and governance determine whether adoption creates measurable value.
Read articleFocus
Architecture becomes strategic when common capabilities are reusable across use cases rather than rebuilt around every new application.
Production introduces lifecycle, reliability and observability requirements that experimental environments are rarely designed to handle.
Strategic challenges
Executives must make investment and positioning decisions while technologies, economics and competitive implications continue to move.
Models, prompts, tools and autonomous actions introduce pathways that conventional application security may not fully address.
POV
A machine should gain decision authority only where its behaviour can be understood, tested and contained under real operating conditions.
Accuracy under normal conditions matters less when one uncontrolled failure can trigger actions the organisation cannot contain.
Strategic impact
Structured deployment, evaluation and monitoring allow teams to change models and configurations without losing visibility or control.
Economic modelling connects AI adoption to specific business drivers and makes the conditions behind expected returns explicit.
What we observe
We frequently see collections of use cases and technology initiatives without explicit choices about competitive or business priorities.
We often see advanced agents layered onto fragmented processes, weak integrations and decision rights that were never clearly defined.