NVIDIA and Lakera AI Propose Unified Framework for Agentic System Safety
As artificial intelligence systems become more autonomous, their interaction with digital tools and data introduces significant new risks.
As artificial intelligence systems become more autonomous, their interaction with digital tools and data introduces significant new risks.
Researchers from NVIDIA and Lakera AI have collaborated on a paper proposing a unified framework for the safety and security of advanced "agentic" systems.
The proposed framework addresses the limitations of traditional security models in managing the novel threats posed by AI agents capable of real-world actions.
The core of the framework shifts the perspective from viewing safety as a static feature to considering safety and security as interconnected properties. These properties emerge from the dynamic interactions between AI models, their orchestration, the tools they use, and the data they access.
This holistic approach aims to identify and manage risks across the entire lifecycle of an agentic system, from development to deployment.
As artificial intelligence systems become more autonomous, their interaction with digital tools and data introduces significant new risks.
Researchers noted that conventional security assessment tools, such as the Common Vulnerability Scoring System (CVSS), are insufficient for addressing the unique risks in agentic AI. A minor security flaw at the component level could escalate into significant, system-wide user harm.
This new model introduces a comprehensive method for evaluating these complex systems. It provides a structured approach to understanding how localized hazards can compound and lead to unexpected, large-scale failures.
The framework is designed for enterprise-grade workflows, ensuring that as agents become more integrated into business processes, their actions remain aligned with safety and security policies.
The paper further explores the crucial phase of risk discovery, which utilizes an AI-driven red teaming process. Within a sandboxed environment, specialized "evaluator" AI agents are deployed to probe the primary agentic system for weaknesses.
These probes simulate various attack scenarios, from prompt injections to sophisticated attempts at tool misuse, to identify potential vulnerabilities before exploitation. This automated evaluation enables developers to detect and mitigate novel agentic risks, such as unintended control amplification or cascading action chains, in a controlled setting.
To support the advancement of this field, the researchers have released the Nemotron-AIQ Agentic Safety Dataset 1.0, which contains over 10,000 traces of agent behaviors during attack and defense simulations. This resource provides a valuable tool for the broader community to study and develop more robust safety measures for the next generation of agentic AI.
Based on reporting by Cyber Security News.
