_Summary
Should we build our own AI for compliance, or buy AI solutions?
At some point, most organizations running ESG, regulatory, or due diligence workflows ask themselves this exact question, and the most common answer we hear back is some version of “we’re already managing AI internally.”
It’s a reasonable instinct. The technology looks accessible, open-source models are strong, and keeping everything in-house feels like the safer, more controlled path. The research on how these projects actually play out, across multiple independent studies published in the last twelve months, tells a more complicated story, and points to a small number of specific, recurring reasons internal AI projects in compliance stall.
The problem isn’t building an AI prototype. It’s turning it into a system that actually works.
Most companies don’t struggle to start an AI project. They struggle to finish one.
An internal team can build a chatbot, connect an LLM to company documents, automate a few tasks, or demonstrate an impressive proof of concept in weeks.
The difficult part comes afterwards:
- Can the system work reliably with real company data?
- Can it produce consistent results?
- Can users trust its answers?
- Can every important decision be traced and explained?
- Can it integrate with existing systems?
- Can it be monitored when the model, regulations or underlying data change?
- And, perhaps most importantly, can the company prove that the investment is actually creating measurable business value?
This is where the evidence is increasingly clear: the gap between experimenting with AI and operating AI at scale is enormous.
And that gap is one of the strongest reasons for companies to consider working with a specialized software and AI partner rather than trying to build everything themselves.
AI adoption in-house is high, but successful implementation is not
A 2025 study from MIT’s Project NANDA, The GenAI Divide: State of AI in Business 2025, analyzed more than 300 publicly disclosed AI implementations, alongside interviews and surveys of business leaders and employees.
Its most striking finding was that 95% of organizations in the study saw no measurable financial return from their GenAI pilots. Only around 5% were achieving significant value.
But the important point is what MIT says is behind that result.
The problem was not simply that the AI models were not powerful enough.
Instead, MIT identified a “learning gap”: systems that did not adapt sufficiently to the organization’s context, workflows and feedback. The research also highlights a major implementation gap between experimenting with AI and embedding it into real business processes.
That distinction matters.
The question isn’t:
“Can we make AI work?”
The question is:
“Can we make AI work reliably inside our organization?”
Those are very different engineering problems.

From prototype to production: where internal projects usually get stuck
Another 2025/26 data point tells a similar story.
Research from IDC and Lenovo reported that for every 33 AI proof-of-concepts launched by enterprises, only around four reach production.
In other words, a successful demonstration is not the same thing as a successful AI product.
A prototype can work with:
- a limited number of documents
- clean data
- carefully selected examples
- a small group of enthusiastic users
- manual checking
- engineers available to fix problems immediately
Production is different.
A production system has to cope with:
- inconsistent data
- thousands or millions of documents
- changing regulations
- different users and use cases
- security and permissions
- system integrations
- model updates
- unexpected inputs
- monitoring and logging
- performance requirements
- cost control
- human oversight
- audit requirements
This is the point at which many internal AI projects stop being an AI experiment and become a software engineering project.
And that is a much bigger undertaking than choosing an LLM and building an interface around it.
Why do internal AI projects fail?
There is no single reason.
The pattern emerging across research is that AI projects fail because organizations underestimate the system around the AI.
There are five recurring problems.
1. The prototype works: but the workflow doesn’t
AI is often introduced as a technology project:
“Let’s see what this model can do.”
But businesses don’t need models. They need outcomes.
A compliance team doesn’t need an impressive chatbot. It needs a faster and more reliable way to review evidence, identify gaps, prepare reports and document decisions.
That requires AI to be integrated into the actual workflow.
MIT’s research points directly to this problem: organizations that successfully capture value tend to move beyond generic AI tools and develop solutions that are deeply integrated into specific business processes.
This is one reason why many internal pilots remain pilots.
The technology works.
The business process around it has not been redesigned.
2. Good AI requires much more than a good model
It is tempting to think of an AI platform as:
User → AI model → Answer
In reality, a production-grade system often looks more like:
User → application → permissions → data → retrieval → multiple models/tools → validation → business rules → human review → audit trail
And all of those components need to work together.
For a regulatory or compliance use case, the system may need to:
- identify the relevant documents;
- retrieve the correct regulatory context;
- distinguish current information from outdated information;
- reason over several sources;
- generate an answer;
- identify uncertainty;
- provide evidence for the answer;
- apply business rules;
- route exceptions to a human;
- record what happened for future audit.
The LLM is only one component.
The real product is the system around it.
That system needs to be engineered, tested and maintained.

3. Internal teams often have the right domain-expertise: but not the right capacity for an AI platform
This is an important distinction. Compliance professionals understand the business better than an external software company ever will.
They know:
- which regulations matter;
- which documents are authoritative;
- where the current process breaks;
- which decisions require human judgment;
- what auditors ask for;
- what regulators expect.
That expertise is essential.
But it does not automatically translate into the ability to build and operate an AI system.
Moody’s 2026 research, based on a survey of 600 risk and compliance professionals, found that 41% identified a lack of internal expertise or skills as the biggest barrier to scaling AI. Other significant barriers included regulatory uncertainty, legacy-system integration and insufficient resources.
This is revealing.
The biggest problem isn’t necessarily that companies don’t understand AI.
It is that they don’t have enough people with the combination of:
AI engineering + software architecture + data engineering + security + AI evaluation + domain understanding + production operations.
Finding all of those capabilities internally is difficult.
And even when a company has them, those people have to maintain the system indefinitely.
4. AI needs governance and measurement from day one
Traditional software is relatively deterministic. If the same code receives the same input, we generally expect the same result.
AI systems are different. The model can change. The underlying data can change. The retrieved context can change. Prompts can change. External services can change. And the output can vary.
That means an enterprise AI system needs its own layer of evaluation and governance.
Organizations need to know:
- How accurate is it?
- When does it fail?
- Which types of cases are unreliable?
- Which model produced the answer?
- Which sources were used?
- What data did the system access?
- When was the result generated?
- Was a human involved?
- What happens when the AI is uncertain?
- Has performance changed since the last model update?
NIST’s AI Risk Management Framework explicitly treats AI risk management as something organizations should address across the design, development, deployment and use of AI systems. Its Generative AI Profile further emphasizes evaluation and trustworthiness considerations throughout the AI lifecycle.
And as AI becomes more autonomous, the problem becomes even harder.
Gartner forecasts that more than 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.
The lesson is not that AI agents don’t work.
It is that autonomy increases the amount of engineering and governance required around the model.
5. How much does it actually cost to run AI in-house?
Even organizations that clear the technical and domain hurdles run into a third one: economics. KPMG’s Q2 2026 survey found that nearly half of organizations have already questioned, delayed, scaled back, or paused an AI deployment once its expected costs began outweighing its value. Only around a third report full visibility into what their AI systems actually cost to run day to day, and those with full visibility are five times more likely to report established ROI than those without it.
The Moody’s compliance-specific data adds a matching piece: 41% of respondents cite a lack of internal expertise or skills as the primary barrier to scaling AI in risk and compliance, ahead of regulatory uncertainty, integration complexity, or budget itself.
Skilled AI engineers capable of building and maintaining a production-grade system are scarce and expensive, and that cost doesn’t end at launch, it recurs every time a model is upgraded, a regulation changes, or something breaks in production.
The second cost is usage itself. Most frontier models are priced per token, and token consumption for a workload involving large documents and iterative regulatory reasoning is genuinely difficult to forecast in advance. KPMG’s finding that access to lower-cost, high-fidelity models is now the fastest-rising influence on AI strategy, up 7 points in a single quarter, reflects exactly this pressure: organizations discovering that using AI and knowing what it costs and what it’s worth are two very different states to be in.
Common questions compliance leaders ask before deciding
“Isn’t building in-house cheaper in the long run, since we avoid ongoing vendor fees?”
Usually not, once the full cost is counted. The upfront build is the easy part to estimate. What’s harder to see coming is the recurring cost: retraining as regulations change, maintaining secure infrastructure, and the specialized engineering talent needed to sustain it, which per Moody’s own survey of compliance professionals is already the single most-cited barrier to scaling AI in this field.
“We have strong engineering talent. Isn’t that enough to build this ourselves?”
Engineering capability and sustained AI operations are different disciplines. MIT’s research points to a “learning gap” as the dominant failure mode for enterprise AI projects, meaning the initial build often isn’t where things break down. It’s the months and years of adaptation afterward, which is exactly where the KPMG data shows nearly half of organizations end up scaling back deployments once real costs and complexity appear.
“Our data is too sensitive to hand to an external platform, so shouldn’t we build it ourselves?”
That’s a legitimate concern, but it’s really an argument about where AI runs, not who builds it. A specialized platform can run entirely inside your own environment or an EU/Swiss data centre, with zero data leaving that perimeter, which addresses the sovereignty concern directly without requiring you to also build and maintain the AI system yourself.
“Couldn’t we just start with a small pilot and expand later?”
This is exactly the pattern the research warns about. IDC found that only around 4 of every 33 AI proof-of-concepts reach production. Gartner’s forecast for agentic AI cancellations describes precisely this trajectory: promising pilots that stall once real document volume, real regulatory nuance, and real audit scrutiny enter the picture.
“If we do decide to build, what’s the minimum we’d need beyond engineering talent?”
At least three things the research consistently flags as commonly missing: a mature AI governance model (only about one in five organizations has one, per Deloitte), a defined way to measure whether the system is actually working (a third of compliance teams surveyed by Moody’s aren’t measuring this at all), and full visibility into what the system costs to run, since KPMG found that visibility alone correlates with a five-times-higher rate of established ROI.
So why work with an AI software company and why partnering can work better than building alone
This is perhaps the most important evidence for the build-versus-buy decision.
MIT’s Project NANDA research found that external partnerships with customized, learning-capable AI systems reached deployment at roughly twice the rate of internally built tools — around 67% versus 33% in the study’s sample. The researchers caution that these figures are based on a limited sample and should not be interpreted as a universal success rate, but the direction is significant.
The implication is not:
“Companies should outsource AI.”
It is:
“Companies should not assume that building internally is automatically the lower-risk path.”
In many cases, the better approach is to combine internal domain expertise with external software expertise.
The company retains control over its data, business rules and decisions.
The software partner provides the technology layer needed to make those decisions scalable, measurable and maintainable.
Measurable business efficiency, predictable AI pricing, trustworthy and secure AI with Dydon AI
This is the layer we built Dydon AI around. Not one model handling everything, but several models and tech solutions combined for what each part of the job actually requires, running on infrastructure that stays inside a client’s own environment or Swiss/EU data centres.
It’s also why we don’t price on raw token consumption: we work with clients to identify where the value actually sits, then structure pricing around predictable business outcomes, so budgeting doesn’t depend on token-usage guessing in advance. And it’s why every answer routes through human validation before it reaches a report, not as a limitation, but because that’s where the market itself, per Moody’s own respondents, has already concluded the real trust lies.
The build-versus-buy decision isn’t really about whether your team is capable. It’s about whether the ongoing cost of building, orchestrating, measuring, and paying for a regulatory-grade AI system, indefinitely, is where you want that capability to sit.
Want to understand where AI could actually create value in your workflow?
Dydon AI offers an AI-potential analysis to identify where AI can deliver measurable efficiency, where human validation should remain in the process, and what an enterprise-ready implementation could look like.
Get your AI-potential analysis →
Sources
- Challapally, A. et al. (2025). The GenAI Divide: State of AI in Business 2025. MIT Media Lab, Project NANDA.
- Gartner, Inc. (2025). Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Press release, June 25, 2025.
- IDC / Lenovo (2025). AI proof-of-concept to production benchmark findings.
- KPMG International (2026). Global AI Pulse Q2 2026: From Deployment to Value Realization. June 2026.
- Moody’s, with We Live Context (2025). From Reactive to Proactive: How AI is Transforming Risk and Compliance.
- Moody’s (2026). Navigating the Shift: How Agentic AI is Reshaping Risk and Compliance.
- Deloitte (2026). State of AI governance maturity findings, cited via Dreamix RegTech industry analysis.