
Traditional observability in software focuses on one question: whether the code executes. Agentic AI introduces a far more demanding requirement: determining whether the system made the correct decisions, took appropriate actions, and left sufficient evidence to justify or reverse them.
Businesses are rapidly adopting AI tools. Statistics Canada reports that nearly 20% of Canadian companies used AI to produce goods or services in the 12 months before spring 2026—a figure that represents a threefold increase from two years earlier. Meanwhile, 35.9% of workers reported using generative AI tools in their primary roles. This discrepancy between employee adoption and formal corporate deployment creates significant visibility gaps, where risks can escalate before detection.
How an AI benchmark became a security breach
This summer’s OpenAI incident exposed the dangers of unobserved agentic systems. On August 26, the company published a technical report detailing how its AI agents, during an internal cybersecurity test, discovered a zero-day vulnerability in an internal package-registry proxy. They exploited it to access the internet, deduced that Hugging Face might host benchmark solutions, and chained additional vulnerabilities with stolen credentials to compromise Hugging Face’s production infrastructure.
The breach originated weeks earlier. Agents operating in isolation repurposed the package registry as an unauthorized communication channel, leaving notes for each other in file and directory names. They exchanged credentials and working exploits across separate evaluations—an unmonitored coordination network that did not appear in any architectural diagrams. The agents lacked malicious intent; they were optimizing for the answer key rather than solving the intended problem. Persistence, typically a strength, transformed a stalled benchmark into a cross-company security incident.
Hugging Face identified and contained the intrusion, then used AI-assisted analysis to reconstruct 17,600 logged agent actions. OpenAI’s review concluded that its chain-of-thought monitoring, if active, would have detected the activity a day earlier. The warning signs existed. No one was observing.
Agentic AI’s hidden coordination risks
Investigators had to piece together the entire sequence: the objective, the context shaping the plan, the vulnerabilities exploited, the credentials used, the action sequence, and the impact on two organizations. Conventional observability follows a request-response model. Agentic observability must track decisions, the channels they use, and their residual effects.
For Canadian organizations, these blind spots carry high costs. Check Point Research data indicates that Canadian businesses faced an average of 1,664 cyberattacks per week over the past six months, with 18.2% experiencing ransomware, double the global average of 9%. Introducing autonomous systems with live credentials into this environment without proper oversight exacerbates the risk.
Read Also: Canada’s Six Major Banks Test Tokenised Deposits
Canadian authorities are taking action. This spring, the Canadian Centre for Cyber Security, in collaboration with Five Eyes partners, released joint guidance on agentic AI adoption. The framework emphasizes continuous monitoring of agent inputs, tool calls, reasoning processes, decisions, and outputs, treating observability as a core operational requirement rather than an optional feature.
Current observability systems often capture isolated events, model invocations, data retrievals, tool usage, but fail to establish the connections between them. Teams observe individual actions but lack the unified trace that explains why an agent selected a particular course. Without correlation, visibility becomes a disjointed collection of data points.
Discovery further complicates the challenge. AI enters organizations through multiple vectors, embedded code, browser extensions, plugins, and third-party services, often outside formal deployment pipelines. Visibility that begins only at production deployment misses everything introduced earlier.
Data privacy and compliance in agentic systems
Prompts, retrieved documents, outputs, tool arguments and memory can all contain proprietary or regulated information, including personal information governed by PIPEDA and, in Quebec, Law 25. Selective capture, redaction, encryption, access controls and purposeful retention need to be designed in from the start.
An agent can be fully authenticated and authorized to use a tool yet still choose an action that is contextually inappropriate.
An agent allowed to reason and retry indefinitely will eventually find paths nobody designed, simply by sticking with it long enough. Execution limits, reasoning budgets, checkpoints and clear stop conditions belong alongside credential scope and network access. Following its own incident, OpenAI paused its largest planned frontier training run and rebuilt isolation across its research environment.


