anthropic keynote whitepaper by zartis

The Agentic Revolution

Most organisations are making AI decisions based on capability they evaluated twelve to eighteen months ago. Based on the keynote by Ralph Ramos, Member of Technical Staff at Anthropic, this paper sets out why the model you evaluated and the model you are about to deploy are not the same product, and what that means for governance.

Introduction

At the 2026 Zartis AI Summit, Ralph Ramos, Member of Technical Staff at Anthropic, argued that most organisations are making AI decisions based on capability they evaluated twelve to eighteen months ago. The model they assessed and the model they are about to deploy are not the same product, and the gap between them, in terms of what AI can autonomously do, is larger than most decision-makers appreciate.

This paper is based on that keynote. It examines what changed inside AI capability in the space of eighteen months, why evaluation cycles built for incremental software releases break down against a categorical shift rather than an incremental one, and what is already happening inside engineering teams while most governance frameworks still assume a chatbot from eighteen months ago.

The central argument: organisations should not build governance around what a specific model version can do today, since that capability will not hold still. They should build it around principles, the categories of action an agent may take, the escalation path when it exceeds them, and the audit trail that makes its behaviour reviewable, which stay durable as capability moves on.

Ralph Ramos giving the keynote on The Agentic Revolution

Govern for the AI you'll deploy, not the AI you piloted

Get the full paper on why the evaluation cycle most organisations run is built for a product that no longer exists.

Closing the eighteen-month gap between the AI you evaluated and the AI you are about to deploy takes more than awareness of the problem.

Follow these steps for a clear path forward:

Design governance around principles and mechanisms, not around one model version's specific behaviour

Define the categories of action an agent is permitted to take, and the escalation path for when it exceeds them

Treat AI tools already in use across your organisation as a current state to manage, not a future risk to prevent

Make the sanctioned AI path easier to use than the unsanctioned one, rather than trying to restrict access outright

Prepare the wider organisation conceptually, not just technically, before agentic tools reach non-technical workflows

Survey your engineering leadership on what agentic tools are already in daily use, across sanctioned and unsanctioned channels, within the next 30 days

Download the whitepaper