This piece argues that the choice of delivery partner is a bigger determinant of whether a custom generative AI application reaches production than the choice of model, and that this choice can be tested against three concrete dimensions before a contract is signed: technical depth, delivery track record, and governance maturity. None of the three can be assessed from a demo. All three can be assessed from questions asked in procurement.
Why the failure pattern is not a technology problem
McKinsey’s March 2025 global survey found that 88% of organisations now use AI in at least one business function, up from 78% a year earlier, yet only 39% attribute any EBIT impact to it, and just 6% of respondents qualify as “high performers” capturing disproportionate value (McKinsey, The State of AI). Deloitte’s most recent enterprise AI survey puts a similar gap in different terms: 74% of organisations hope AI will grow revenue, but only 20% can currently demonstrate measurable impact, with a persistent skills gap cited as the top barrier to integration.
Root-cause analysis aggregated across enterprise AI pilots points at the same handful of blockers repeatedly: data readiness, integration complexity, change management, and unclear technical ownership of the system once it is live. None of those are model problems. They are delivery problems, which is exactly the category a development partner is supposed to solve and, per the MIT NANDA finding, often does not.
Dimension one: technical depth
A vendor that can produce a working prototype in a sales cycle has demonstrated that the underlying model works. It has not demonstrated that their engineering practice will hold up once the application is handling real production traffic, real data drift, and real edge cases. Three specific questions separate the two.
Ask for evaluation scores on a representative dataset, not a demo transcript.
Reference-free evaluation methods for retrieval-augmented generation, scoring faithfulness, answer relevancy and context relevance without hand-labelled ground truth, are now a mature academic and open-source discipline. The Ragas framework established the methodology; open-source tools such as DeepEval operationalise it as a repeatable test suite a delivery team can run in continuous integration, the same way a unit test suite runs against conventional code. A partner with real evaluation discipline can show you scores against a defined dataset. A partner without it can only show you that the demo worked on the day it was recorded.
Ask which evaluation metric they are using, and why.
This question matters more than it first appears. A 2026 study comparing four widely used evaluation libraries, Ragas, DeepEval, RAGChecker and Opik, on the same applied tasks found meaningful disagreement between them (arXiv:2607.07302). A vendor’s claim that “we run automated evaluation” is not, on its own, evidence of quality. The follow-up question, which metric, scored against what baseline, and what happens when the retrieval index or model version changes, is what separates a partner with real evaluation infrastructure from one repeating a phrase from a sales deck.
Ask how they red-team the system before it ships.
HarmBench established a standardised benchmark for automated adversarial testing and refusal robustness; open-source tools such as promptfoo now let a delivery team run that kind of red-teaming as part of a CI pipeline rather than as a one-off manual exercise before launch. A partner who cannot describe a concrete red-teaming process, what they test for, how often, and what fails the build, is not applying software engineering discipline to a system that behaves non-deterministically. A 2025 survey of software engineering practice specific to large language model systems is explicit on this point: conventional software testing rigour is necessary for LLM-based applications but not sufficient on its own (arXiv:2506.23762).
Dimension two: delivery track record
The second dimension is whether the partner has actually taken a comparable system into production and kept it there, not whether they can build one.
Ask for a system that has been live for six months or longer, not one that launched last quarter. A system that survives its first production month, when initial bugs surface, is a different proof point from one that has survived a data source changing format, an upstream model provider deprecating a version, or a spike in traffic the original design did not anticipate. Ask specifically how the partner handled data readiness and system integration on that project, since these are the two most cited reasons pilots stall before reaching production, well ahead of model choice.
It is also worth asking how the partner scopes the build itself. Andreessen Horowitz’s 2025 survey of enterprise CIOs found that generative AI spend has moved from experimental “innovation” budgets into permanent software line items, with fewer than a quarter of enterprises still funding it as an experiment, and that off-the-shelf, AI-native applications are increasingly covering the commodity use cases that used to justify a fully custom build (a16z, How 100 Enterprise CIOs Are Building and Buying Gen AI in 2025). A delivery partner worth hiring for a custom build should be able to tell you plainly when a custom application is not the right answer, and point you at an off-the-shelf alternative instead. A partner who proposes a full custom build for every problem, regardless of whether the differentiation justifies it, is optimising for their own billable scope rather than the client’s outcome.
McKinsey’s November 2025 survey adds a scoping dimension worth asking about directly: 23% of organisations are already scaling an agentic AI system in production and 39% are experimenting, meaning the target architecture for a “custom generative AI application” in 2026 is increasingly a multi-step agentic workflow rather than a single-turn chat or retrieval interface. A partner’s production track record should reflect that shift, not just static retrieval-augmented systems built two years ago.
Dimension three: governance maturity
The third dimension is whether the partner owns the system’s security and compliance posture after handover, or simply delivers code and exits.
Look for observability and tracing built into the delivery, not bolted on after an incident. Tools such as Langfuse combine tracing, prompt versioning and evaluation in a single, self-hostable platform, which matters directly for data governance: a self-hostable option means trace and prompt data does not have to leave the client’s own infrastructure to be inspected. Ask whether the partner’s proposed architecture includes programmable guardrails at the input, dialogue and retrieval layers, the pattern reference toolkits such as NVIDIA’s Guardrails implement, and whether those guardrails are tested as part of the same red-teaming process described above rather than treated as a separate afterthought.
A shared vocabulary is useful here. The OWASP GenAI Top 10 risk taxonomy, covering prompt injection, data leakage and adversarial manipulation among others, gives a buyer a checklist that does not depend on any single vendor’s terminology. Ask a prospective partner to walk through how their delivery process addresses each item on that list, specifically for the application being proposed, not in the abstract. A vendor that has genuinely built the discipline into their delivery process will answer in specifics. One that has not will answer in generalities.
Finally, ask who owns the system once it is in production: who is named, what their response commitment is, and whether that ownership is written into the contract or implied by goodwill. A partner that treats delivery as advise-and-build, rather than advise-then-disappear, will have an answer ready. A vendor operating as a prompt-wrapper shop, assembling a thin interface over a foundation model API with none of the evaluation, red-teaming or observability discipline described above, typically will not, because there is no delivery risk on their side of the contract to manage.
What this means for the build-versus-partner decision
None of this argues that every enterprise should partner rather than build. MIT Sloan Management Review’s build, boost, or buy framework is a reasonable starting point for that decision, and for genuinely commodity use cases, buying an off-the-shelf AI-native application is often the right call, as the a16z data above suggests enterprises are increasingly recognising. The argument here is narrower: where a custom application is the right call because the use case is genuinely differentiated, the evidence from MIT NANDA indicates that doing it with a specialist partner roughly doubles the odds of the project reaching measurable value compared with a purely internal build. That advantage is not automatic. It depends on which partner, assessed against the three dimensions above, not against how polished their demo was.
How Zartis fits this framework
Zartis works as an AI transformation partner across AI Development, AI Integration, AI Agents and AI Data Platform engagements: advising on architecture and governance, and then carrying delivery risk into production, rather than handing over a prototype and stepping back. That means evaluation and red-teaming discipline built into the development lifecycle from the first sprint, not added before a demo; data and system integration readiness assessed and addressed as part of the build, since it is one of the most cited reasons pilots stall; and governance, observability and named ownership carried through to production support, not left as an unresolved question at handover. Zartis engineering teams are equally set up to scope a multi-step agentic workflow as a static retrieval system, reflecting where custom application architecture is heading rather than where it stood two years ago.
The practical takeaway for a CTO or VP of Engineering evaluating generative AI development services: ask for evaluation scores against a defined dataset, ask for a production system that has survived six months of real traffic, and ask who owns the system after launch. A partner with a genuine engineering practice behind them will answer all three specifically.
References
- MIT NANDA / Project NANDA, “The GenAI Divide: State of AI in Business 2025”
- McKinsey & Company, “The State of AI: How Organizations Are Rewiring to Capture Value”
- McKinsey & Company, “The State of AI in 2025: Agents, Innovation, and Transformation”
- Deloitte, “State of Generative AI in the Enterprise”
- Andreessen Horowitz, “How 100 Enterprise CIOs Are Building and Buying Gen AI in 2025”
- MIT Sloan Management Review, “Buy, Boost, or Build? Choose Your Path to Generative AI”
- Es, Shahul et al., “Ragas: Automated Evaluation of Retrieval Augmented Generation”
- Mazur, Maksim et al., “Evaluating RAG Metrics in Applied Contexts: An Experiment, Its Findings and Its Limitations”
- Mazeika, Mantas et al., “HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal”
- “Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead”, arXiv:2506.23762