Tuesday, August 11, 2026
LIVEThe Unrelenting Cyber Battle: Hacking Threats and the Imperative of Robust Data Protection///Navigating the Cyber Labyrinth: Bolstering Defenses Against Evolving Hacking Threats///The Dual Front War: Battling Hacking and Bolstering Data Protection in the Digital Age///The Ever-Evolving Cyber Threat Landscape: Navigating Hacking and Fortifying Data Protection///The Unseen Battle: Fortifying Data in an Age of Relentless Hacking///The Unseen War: Hacking's Relentless Advance and the Imperative of Data Protection///The Evolving Threat Landscape: Hacking, Data Protection, and the Imperative for Proactive Security///Navigating the Digital Minefield: Bolstering Data Protection in an Era of Relentless Hacking///The Dual Fronts of Digital Defense: Combating Hacking and Fortifying Data Protection///Hacking's New Frontier: Fortifying Data Protection in the Age of Advanced Cyber Threats///The Dual Front: Navigating Hacking Threats and Fortifying Data Protection in the Digital Age///Navigating the Digital Gauntlet: The Evolving Nexus of Hacking and Data Protection///The Unrelenting Cyber Battle: Hacking Threats and the Imperative of Robust Data Protection///Navigating the Cyber Labyrinth: Bolstering Defenses Against Evolving Hacking Threats///The Dual Front War: Battling Hacking and Bolstering Data Protection in the Digital Age///The Ever-Evolving Cyber Threat Landscape: Navigating Hacking and Fortifying Data Protection///The Unseen Battle: Fortifying Data in an Age of Relentless Hacking///The Unseen War: Hacking's Relentless Advance and the Imperative of Data Protection///The Evolving Threat Landscape: Hacking, Data Protection, and the Imperative for Proactive Security///Navigating the Digital Minefield: Bolstering Data Protection in an Era of Relentless Hacking///The Dual Fronts of Digital Defense: Combating Hacking and Fortifying Data Protection///Hacking's New Frontier: Fortifying Data Protection in the Age of Advanced Cyber Threats///The Dual Front: Navigating Hacking Threats and Fortifying Data Protection in the Digital Age///Navigating the Digital Gauntlet: The Evolving Nexus of Hacking and Data Protection///
Subscribe
Cyber Security
Independent · Digital
Thehackingpost
CybersecurityAI-assisted

OpenAI Launches EVMbench: A New Framework to Detect and Exploit Blockchain Vulnerabilities

OpenAI, in collaboration with crypto investment firm Paradigm, has released EVMbench, a benchmark designed to evaluate the interaction of artificial intelligence agents with smart contract security.

OpenAI, in collaboration with crypto investment firm Paradigm, has released EVMbench, a benchmark designed to evaluate the interaction of artificial intelligence agents with smart contract security.

With smart contracts securing over $100 billion in open-source crypto assets, the ability of AI to read, write, and audit code is becoming critical to financial infrastructure. This framework aims to assess capabilities in economically significant environments, promoting the use of AI systems to audit and enhance deployed contracts against potential threats.

EVMbench is constructed on a dataset of 120 curated high-severity vulnerabilities from 40 audits and open code competitions. It includes specific vulnerability scenarios from the Tempo blockchain security audit.

Employing a Rust-based harness, the system restricts unsafe RPC methods and conducts all exploit tasks in an isolated, local Anvil environment, ensuring safety and reproducibility without impacting actual assets or network stability.

The framework evaluates agents across three distinct capability modes that simulate real-world security tasks:

EVMbench is constructed on a dataset of 120 curated high-severity vulnerabilities from 40 audits and open code competitions.
Sarah Dawson · Thehackingpost

Detect : Agents audit repositories to identify known vulnerabilities, scored based on recall of ground-truth vulnerabilities and audit rewards. Patch : Agents modify contracts to remove exploits while retaining functionality, verified through automated tests to ensure the exploit is resolved and code compiles. Exploit : Agents attempt to drain funds from a deployed contract, graded programmatically via transaction replay on a sandboxed blockchain.

In Patch mode, agents must rectify identified issues without disrupting the contract's intended functionality or causing compilation errors. Exploit mode assesses an agent's capability to perform end-to-end fund-draining attacks in a sandboxed environment, providing a metric for offensive capabilities that defenders must counter.

Model Performance and Safety Initiatives

The release of EVMbench underscores advancements in AI model capabilities for cybersecurity tasks. In the exploit mode evaluation, OpenAI’s GPT-5.3-Codex achieved a success rate of 72.2 percent, a notable improvement from the GPT-5 model's 31.9 percent performance six months prior.

While offensive testing capabilities have improved, detection and patching tasks remain challenging. Agents often face difficulties maintaining full functionality while removing subtle bugs, highlighting the necessity of human oversight in the auditing process.

Advertisement

Recognizing the dual-use nature of cybersecurity tools, OpenAI is focusing on an evidence-based approach to support defenders. This includes expanding their security research agent, Aardvark, and allocating $10 million in API credits through their Cybersecurity Grant Program to enhance cyber defense for open-source software and critical infrastructure.

Although EVMbench does not support complex timing mechanics or mainnet forks, it marks a significant step toward standardizing AI interaction with blockchain security.

Based on reporting by GBHackers.

AI transparency. This article was produced with the assistance of artificial intelligence and published under human editorial oversight. AI systems can make mistakes. Read how we use AI (EU AI Act, Art. 50).
Related Stories