Thesis: The enterprise return on Claude Code is not faster individual developers. It is the organisational capability you build when you treat the tool as a way to codify, govern and distribute how your whole engineering function works, and that capability is what survives contact with delivery metrics when raw velocity does not.
The Productivity Number Most Buyers Are Quoting Does Not Hold Up
In 2025, METR ran a randomised controlled trial on experienced open-source developers working in repositories they knew well, averaging five years on the codebase. Before they started, the developers predicted that AI tools would cut their task completion time by 24%. After finishing, they believed AI had made them roughly 20% faster. The measured result was the opposite: allowing AI tooling increased completion time by 19%.
That single finding should change how a Head of Engineering reads every Claude Code business case that lands on their desk. The per-developer speed story is contested even under controlled conditions, and the people doing the work are poor judges of their own gains.
It gets harder. Faros AI’s analysis of engineering telemetry through 2025 found that individual output did rise sharply with AI adoption: 21% more tasks completed and 98% more pull requests merged per developer. Organisational delivery metrics, meanwhile, stayed flat. Bugs per developer rose 54%, incidents per pull request rose 242%, and median pull request review time rose 441%. The work moved faster at the keyboard and then queued up everywhere else.
This is the gap the title points at. If you buy Claude Code to make developers individually quicker, the evidence says you may not get what you paid for, and even where you do, it evaporates before it reaches the business. The case for an enterprise rollout has to be built on something more durable.
Why “Developer Productivity” Is The Wrong Yardstick for an Enterprise Buyer
The 2025 DORA report, drawn from nearly 5,000 practitioners, reached a conclusion that reframes the whole question. AI, it found, acts as an amplifier: it magnifies the strengths of high-performing organisations and the dysfunctions of struggling ones in equal measure. With AI adoption among developers now around 90%, the tool is no longer the differentiator. What you have built around it is.
That is why time saved in writing code keeps reappearing as time spent on review, verification and rework. A faster first draft does not fix a slow review process, a brittle test suite, or undocumented standards that only three senior engineers carry in their heads. It loads more onto them. DORA’s own framing is that value stream management matters precisely because it reveals where AI productivity gains evaporate across the delivery lifecycle.
So the question a Head of Engineering should ask is not “how much faster is each engineer.” It is “what does my organisation now do that it could not do before.” That reframing points at three capabilities Claude Code can build at the organisational layer, none of which shows up in a per-seat productivity chart.
Capability One: Codifying Institutional Knowledge So It Stops Walking Out The Door
Most engineering organisations run on knowledge that lives in people. How we structure a service, which patterns we have banned and why, what “ready for review” means here, how we handle a database migration. That knowledge is expensive to transfer and it leaves when people do.
Claude Code turns that tacit knowledge into shared, version-controlled assets. CLAUDE.md memory files operate in a hierarchy: an organisation-wide file that loads in every session, a project file checked into each repository, and personal files for individual preferences. Custom slash commands capture repeatable workflows. Skills package domain-specific instructions that load on demand. Subagents handle delegated work in isolated context. Model Context Protocol servers connect the tool to your existing systems, and plugins bundle all of it for distribution across teams.
The payoff is not that one developer codes faster. It is that a new joiner, on day one, works inside the same standards, context and review expectations as your most senior engineer, because those expectations are now written down and enforced rather than absorbed over months. Consistency stops depending on who happens to pick up the ticket. That is an organisational capability, and it compounds.
Capability Two: Governance and Control That Let a Regulated Enterprise Actually Adopt This
For most enterprises, and certainly for any in financial services, healthcare or other regulated sectors, the blocker is never whether a tool is useful. It is whether it can be controlled. Claude Code is built to be governed centrally rather than configured per laptop.
Server-managed settings deliver policy from the administrator’s console and refresh during sessions, with no device management infrastructure required, and a settings hierarchy where managed values cannot be overridden by an individual user. Administrators set fine-grained permission rules over what the tool may do, restrict which models and which external integrations are available, and enforce an organisation-wide instruction file that cannot be excluded.
Deployment bends to your existing security posture rather than the reverse. Claude Code runs on the Claude API, on Amazon Bedrock, on Google Vertex AI and on Microsoft Foundry, inheriting each provider’s identity, compliance and data-residency controls. On the data question that compliance teams ask first: for commercial users on Team, Enterprise and API plans, Anthropic does not train its models on your code or prompts. Anthropic holds SOC 2 Type 2 and ISO 27001 certification, and for healthcare organisations a Business Associate Agreement combined with Zero Data Retention extends to Claude Code usage.
This is the unglamorous capability that determines whether anything else on this list is reachable. Without it, a regulated enterprise cannot move past the pilot.
Capability Three: Observability You Can Actually Manage
You cannot improve what you cannot see, and per-seat licence counts tell you nothing about whether AI is helping your organisation deliver. Claude Code exports full OpenTelemetry data: metrics, events and traces that flow into the monitoring stack you already run, whether that is Datadog, Grafana, Splunk or another backend.
Critically, the signals are attributable and aggregatable. Every event carries identity and organisation context, and usage can be tagged by team, department or cost centre. That means you can measure adoption and effect where it matters, at the level of a team and a value stream, rather than celebrating a vanity number of active seats. It is the instrumentation that lets you find DORA’s evaporating gains before they cost you, and to manage to the seven organisational capabilities DORA identifies as the real multipliers: a clear and communicated AI stance, healthy and accessible internal data, strong version control, working in small batches, a user-centric focus, and quality internal platforms.
What This Means for a Rollout
Read together, the evidence and the capabilities point to one conclusion: Claude Code in the enterprise is a capability programme, not a licence purchase. The organisations that get a return are the ones that invest in the surrounding system, the standards encoded in shared files, the guardrails enforced through managed settings, the enablement that teaches engineers to direct the tool well, and the measurement that ties it all back to delivery outcomes.
That is also where the gains are largest, because AI amplifies what is already there. A licence rolled out into an organisation with undocumented standards and a slow review process will amplify exactly those problems. The same licence rolled out into a deliberately built system raises the floor for every engineer at once.
Zartis works with engineering leaders to build that system rather than just supply the tool. As a Preferred Services Partner in the Claude Partner Network, we run Claude Code Workshops and AI-Enabled SDLC enablement that move teams from individual experimentation to governed organisational capability: codifying your standards into shared configuration, setting the permission and observability framework your security team will sign off on, and measuring the effect where it shows up in delivery, not in seat counts. If your teams are already using Claude Code, or about to, the question worth answering is not how fast it makes one developer. It is what capability you are building around it. Talk to us about a Claude Code Workshop and we will help you answer it.