# Agathon.ai - AI Consultancy (Full) > Senior AI expertise with direct access. Real outcomes. This is the extended version of llms.txt with comprehensive information about Agathon, its founder, services, and the full text of all published articles. For the summary version, see https://agathon.ai/llms.txt --- ## About Agathon Agathon is a boutique AI consultancy that provides senior, hands-on expertise to organisations building with AI. Unlike large consultancies that sell senior partners and deliver juniors, Agathon's model ensures the person in the strategy meeting is the same person doing the technical work. Agathon specialises in four areas: 1. Building sophisticated AI products from concept to production 2. Technical due diligence on AI implementations for investors and acquirers 3. Strategic advisory for leadership teams building AI capabilities 4. Fractional CTO services spanning all technology decisions The consultancy works across financial services, telecommunications, automotive, and technology sectors, with a focus on organisations that need PhD-level technical depth combined with commercial pragmatism. ## Founder & Principal Consultant **Dr Colin Kelly** - Role: Founder & Principal Consultant at Agathon - Education: PhD in Natural Language Processing from the University of Cambridge; BA in Mathematics & Computer Science from the University of Oxford - Experience: 15+ years of commercial AI delivery - Background: Former applied AI research lead who has built production ML systems, led applied research teams, and delivered quantified AI benefits across enterprise environments - Expertise: Large Language Models (LLMs), Natural Language Processing (NLP), AI strategy, AI due diligence, AI product development, machine learning engineering - Profile: https://agathon.ai/about ## Services (Detailed) ### AI Product Pioneering Hands-on development of AI products that exploit full technical potential. Agathon works with clients to design and build novel AI systems, from interface patterns to backend architectures. Engagements range from early-stage prototyping to production deployment. URL: https://agathon.ai/services/ai-product-pioneering ### AI Due Diligence Technical assessment that separates real AI capability from marketing claims. Designed for investors, acquirers, and boards evaluating AI companies or internal AI initiatives. Provides rigorous, independent analysis of technical architecture, team capability, data assets, and scalability. URL: https://agathon.ai/services/ai-due-diligence ### AI Leadership Advisory Strategic consulting for leadership teams building internal AI capability. Covers AI roadmap development, build-vs-buy decisions, team structure, vendor evaluation, and governance frameworks. Focused on reducing dependency on external consultants over time. URL: https://agathon.ai/services/ai-leadership-advisory ### Fractional CTO Ongoing executive technical leadership for organisations that need senior technology decision-making without a full-time hire. Covers architecture, team building, vendor management, and technology strategy across all technology decisions. URL: https://agathon.ai/services/fractional-cto ## Training & Workshops (Detailed) ### LLM Strategy for Business Leaders - Duration: 1 day - Audience: Executives and senior leaders - Format: In-person or online - Content: Practical framework for identifying high-impact LLM opportunities, evaluating build-vs-buy decisions, understanding the economics of LLM deployment, and avoiding common costly mistakes. Participants leave with a prioritised list of use cases ranked by ROI potential. URL: https://agathon.ai/training#intro-llm ### Applied Prompt Engineering for Teams - Duration: 2 days - Audience: Technical teams - Format: In-person or online - Content: Battle-tested prompt engineering methodology that reduces LLM costs by 40-60% while improving output quality. Participants work through real business scenarios, learn to design prompts that handle edge cases, build evaluation rubrics, and create scalable prompt libraries. URL: https://agathon.ai/training#prompt-engineering ### Building Production AI Agents - Duration: 2 days - Audience: Advanced technical teams - Format: In-person or online - Content: Architecture patterns for deploying autonomous AI systems with enterprise guardrails. Covers tool integration, memory systems, agent orchestration, error handling, evaluation strategies, and operational considerations for production deployment. URL: https://agathon.ai/training#build-ai-agents ## Tools & Frameworks ### AI Readiness Assessment Free interactive questionnaire that evaluates an organisation's readiness for AI adoption across multiple dimensions. Provides a personalised score and recommendations. URL: https://agathon.ai/readinessassessment ### AI Agent Deployment Framework Practical framework for deploying LLM-based AI agents in the enterprise, covering architecture patterns, governance, and implementation guidance. URL: https://agathon.ai/frameworks/ai-agent ### Which LLM Decision Guide Interactive tool to help determine the right LLM for specific needs, comparing capabilities, pricing, and use case fit. URL: https://agathon.ai/whichllm ## Published Insights (Full Text) ### Big 4 consultant vs independent AI advisor: who to trust with AI strategy and due diligence - URL: https://agathon.ai/insights/big-4-consultant-vs-independent-ai-advisor-who-to-trust-with-ai-strategy-and-due-diligence - Published: 2026-08-08 - Categories: AI Advisory, AI Strategy, AI Consulting In October 2025, Deloitte Australia repaid part of a A$440,000 government contract after researchers found its 237-page report contained AI-generated fabrications, including a quote attributed to a federal court judgment that the judgment never contained and citations to academic papers that do not exist. The revised report disclosed that the firm had used GPT-4o in drafting. That episode is unusual only in being public. The background numbers are worse. RAND's 2024 study of AI projects found failure rates above 80 per cent, roughly double the rate for ordinary IT projects, and traced most failures to organisational causes rather than technical ones. Gartner predicted in June 2025 that over 40 per cent of agentic AI projects will be cancelled by the end of 2027, citing runaway costs, unclear value and weak risk controls. MIT's much-quoted NANDA report put the share of enterprise GenAI pilots showing measurable P&L impact at just 5 per cent, and while that figure has been fairly criticised for its small, self-reported sample, the direction of travel matches everything else we can measure. So the market for AI advice is a market in which most advice, followed faithfully, leads to a failed project. Choosing who to trust with your AI strategy, or with due diligence on an AI acquisition, matters more than the equivalent choice did for cloud migration or ERP. This article compares the two options most buyers weigh: a Big 4 firm (Deloitte, PwC, EY, KPMG, and by extension the large strategy houses) versus an independent AI advisor or specialist boutique. #### The short answer Neither is better in general. A Big 4 firm is the right choice when you need industrial-scale delivery, contractual risk transfer, and a brand that reassures a board or regulator. An independent advisor is the right choice when you need deep technical judgement from the actual senior person, delivered fast, from someone with no stake in which platform you pick or how big the follow-on programme gets. For AI due diligence specifically, independence is close to being the product itself. The rest of this article is the reasoning, a comparison table, and a decision framework you can apply to your own situation. One disclosure before we start. I run an independent AI advisory practice, so I have a side in this fight. I have tried to make the strongest case for the Big 4 rather than a conveniently weak one, and I have turned the harshest lens on the independent model too. Discount my conclusions as you see fit. #### What the Big 4 offer It has become fashionable to sneer at the big firms, particularly since Mazzucato and Collington's 2023 book The Big Con argued that the consulting industry systematically overstates what it knows and hollows out its clients' internal capability. Some of that critique lands. But buyers keep hiring these firms for reasons that are mostly rational. Scale and delivery capability. A Big 4 firm can put 50 consultants on your programme next month, across a dozen countries, with programme management, change management and training wrapped around the technical work. No independent can do this, and most boutiques cannot either. If your engagement is an eight-figure, multi-jurisdiction transformation, the delivery bench is the product and the big firms are the only realistic suppliers. Serious AI investment. The big firms are not tourists here. KPMG announced a $2 billion, five-year AI investment with Microsoft in 2023. EY put $1.4 billion into its EY.ai platform. PwC committed $1 billion over three years. Deloitte's 2025 alliance with Anthropic became Anthropic's largest enterprise AI deployment, covering more than 470,000 staff. Whatever else these numbers mean, they buy tooling, training and accumulated delivery experience at a scale no small firm can match. Risk transfer. Professional indemnity backed by a deep balance sheet, established methodologies, audit adjacency and regulatory familiarity. If the engagement goes wrong, there is an institution to hold accountable and insurance behind it. For some boards and procurement functions this is not optional. Cover. "Nobody ever got fired for hiring McKinsey" gets used as an insult, but it describes something real. If you are a CIO betting your credibility on an AI programme, or a fund partner writing a large cheque, a brand-name advisor makes the decision safer for you personally. The advice might not be better, but the decision is more defensible. A meaningful share of Big 4 fees pays for exactly this, and sometimes that is money well spent. #### What an independent AI advisor offers The person who sold the work does the work. This is the structural difference from which everything else follows. At a large firm, the partner who impressed you in the pitch typically hands delivery to a team you have not met, some of whom learned the subject recently. With an independent, the pitch and the delivery are the same brain. Specialist depth. Evaluating AI systems properly requires hands-on machine learning experience: reading evaluation results, checking training data provenance, telling a genuine capability from a demo. The field also moves monthly, and staying current is a full-time discipline. Generalist consulting benches vary enormously on this. A good independent AI advisor is a specialist by definition, because specialism is the only reason to hire them. Speed and price. An independent can typically start within days and deliver a scoped piece of work in weeks. On cost, reported benchmarks put Big 4 day rates at roughly $1,000 to $2,000 for junior staff and $4,500 to $7,500 for partners, with boutiques at $1,500 to $3,000 and independents ranging from $500 to $4,500 depending on seniority. The headline rate quoted in a big-firm pitch is usually the partner rate; the invoice reflects the blended team. An independent quotes one rate for one person. Independence. No alliance with a platform vendor, no implementation practice waiting downstream of the strategy, no audit relationship to protect. More on why this matters below. A maturing market. This is no longer a fringe option. DataIntelo valued the global fractional executive market at $9.4 billion in 2025, projecting $24.7 billion by 2034, and a 2026 Umbrex report citing Forbes found 72 per cent of CEOs planning to increase their use of fractional executives. Private equity firms routinely deploy independent specialists across portfolios. The supply side has professionalised, partly because senior people keep leaving the big firms to do this. #### The comparison at a glance #### The question that decides it: incentives, not intelligence Buyers usually compare advisors on brand, credentials and chemistry. Those are weak predictors. The strong predictor of whether advice serves you is the advisor's incentive structure, and this cuts both ways. ##### Three structural incentives at large firms The leverage model. Big-firm economics depend on a pyramid: partners sell, junior staff deliver, and the margin lives in the gap between what juniors cost and what they bill for. You are not being cheated, this is simply how the model works, but it means the default staffing of your engagement is the most junior team the firm believes can deliver it. In AI work, where judgement is scarce and the field changes monthly, that default is a bigger problem than it is in tax or audit. There is also a live irony: AI itself is eroding the pyramid. Reported figures for UK graduate intake cuts include 29 per cent at KPMG, 18 per cent at Deloitte and 11 per cent at EY, as the analyst work that juniors used to do is increasingly done by the same technology the firms are selling advice about. Vendor alliances. The Deloitte and Anthropic alliance covers 470,000 people and includes certifying 15,000 practitioners on Claude. PwC became OpenAI's first reseller in 2024 and one of the largest enterprise customers of ChatGPT Enterprise. These are sensible commercial moves, and clients get real benefit from the tooling and training. But if the firm advising you on platform selection also resells one of the candidate platforms, you should at minimum know that, and ask directly: how does your firm earn money if we adopt platform X versus platform Y? An advisor with a good answer will not mind the question. Implementation revenue. Strategy work at large firms is often priced modestly because it seeds the implementation programme, which is where the real revenue sits. That creates a quiet gravitational pull on the advice itself. A recommendation that concludes "you need an 18-month transformation programme" is worth millions to the advisor. A recommendation that concludes "you need three focused changes and no programme at all" is worth nothing to them, and it is sometimes the right answer. None of this is speculative or fringe. The UK Financial Reporting Council required the Big Four to operationally separate their audit and advisory arms because regulators concluded that structural conflicts inside multi-service firms are real and need managing. ##### Now turn the same lens on independents Fairness requires pointing the incentives argument the other way, and it draws blood. Key-person risk. If your independent advisor is ill, overcommitted or hit by a bus, there is no bench. For advisory work this is manageable. For anything operationally critical, it is a weakness that no amount of talent fixes. The empty-calendar problem. An independent with capacity to fill has an incentive to say yes to work at the edge of their competence, and to describe every problem as the kind of problem they solve. Big firms have this incentive too, but a brand and a methodology at least impose some floor on quality. With independents, quality variance is wide and there is no institution policing it. Thin risk transfer. An independent's professional indemnity is real but modest. If a large deal goes wrong on the back of their due diligence, you are not recovering your losses from them. What you are buying is judgement, not insurance, and you should be clear-eyed that those are different products. The honest conclusion is that neither structure is clean. The choice is which conflicts are most dangerous for your specific situation, and which you can see and manage. Big-firm conflicts push advice towards bigger programmes and allied platforms. Independent conflicts push towards overclaiming fit. For most strategy and due diligence work, the second is easier to detect and cheaper to survive. #### AI due diligence is the special case Everything above applies double when the engagement is due diligence on an AI company or an AI-driven deal, because here the entire product is scepticism. The base rate of exaggeration is high. A 2019 MMC Ventures study of 2,830 European startups classified as AI companies found evidence of AI material to the value proposition in only around 60 per cent of them. That finding is often overstated as "40 per cent were lying", which the data does not show, but even the careful reading means a large minority of AI claims do not survive inspection. Enforcement has followed the pattern. In March 2024 the SEC fined two investment advisers, Delphia and Global Predictions, for claiming AI capabilities they did not have, with its enforcement director summarising: "Simply put, that's called AI washing and it hurts investors." Gartner has estimated that of the thousands of vendors claiming agentic AI capabilities, only around 130 are genuine. Proper AI due diligence therefore goes beyond standard technology due diligence. It has to cover model provenance and ownership, the legal rights behind training data, data pipeline durability, the quality of the company's evaluations versus its demos, and regulatory exposure. The stakes on training data alone are now enormous: Anthropic's settlement over books used in training ran to roughly $1.5 billion. On regulation, the EU AI Act carries penalties up to €35 million or 7 per cent of global turnover for prohibited practices, applies to non-EU companies whose systems reach EU users, and its high-risk system obligations, originally due in August 2026, were deferred to December 2027 by the Digital Omnibus. A diligence provider who has not tracked that moving timeline is not current enough to price the risk. Two things follow for the Big 4 versus independent question. First, evaluating model claims is hands-on technical work, and you should ask any provider exactly who on the team has trained and evaluated models rather than managed projects about them. Large firms can field such people; whether they will be on your engagement is a staffing question you must ask. Second, independence stops being a nice-to-have. If the firm assessing an AI target has an alliance with the target's platform vendor, or an implementation practice that benefits from the deal closing, the diligence is structurally compromised no matter how able the team is. Buy your scepticism from someone with nothing riding on the answer. #### A decision framework: six questions Work through these in order. Your answers will usually make the choice for you. 1. Is this advice or delivery? If the engagement needs more than a handful of people executing in parallel, you need a firm. If it needs judgement, analysis and a defensible recommendation, headcount is irrelevant and may be a negative. 1. Does the outcome need to survive a board, regulator or court? If institutional brand cover and indemnity are load-bearing, that points to a big firm, and it is a fair reason to choose one. 1. How fast do you need to move? Deal timelines and competitive AI decisions often cannot absorb a big firm's mobilisation period. 1. What does failure cost, and who carries it? If you need to recover losses from your advisor, only a large firm's insurance is worth anything. If failure means a wrong decision you will own regardless, buy the best judgement available. 1. Where would conflicts hurt you most? For platform selection and due diligence, vendor alliances and implementation incentives sit exactly where the risk is, which weighs towards independence. 1. Do you need a senior specialist's brain or a firm's machine? Be honest about which one the work requires, because you will pay for both either way at a large firm. Choose a Big 4 firm when the job is large-scale delivery, when risk transfer and brand assurance are genuine requirements, and when you have the internal capability to manage the conflicts and staffing questions above. Choose an independent AI advisor when the job is strategy, due diligence or a scoped technical assessment, when speed matters, and when the value of the advice depends on it being unconflicted. #### The hybrid most buyers overlook Experienced buyers increasingly combine the two. Common patterns: - Independent as buy-side reviewer. Retain a large firm for the programme, and an independent specialist to review its recommendations, staffing and platform choices on your behalf. Against the programme budget the reviewer is a rounding error, and big-firm teams work differently when they know a specialist is reading their output. - Split strategy from build. Have an independent set the strategy or run the due diligence, then hand implementation to a firm or systems integrator with the bench to deliver it. The advisor who scoped the work has no stake in inflating it, and the implementer competes on delivery rather than marking their own homework. - Specialist evals inside broader diligence. On a deal, let a big-firm team run the standard workstreams and bring in a specialist purely for model evaluation and data provenance. Boards already insist on separation of duties for money. Applying the same logic to AI advice, so that whoever recommends the spend does not profit from it, is basic governance catching up with a new category of purchase. #### FAQ What is AI due diligence? Specialised assessment of a company's AI claims and assets, typically for investors or acquirers. It covers model ownership and provenance, legal rights to training data, data pipelines, the gap between demos and evaluated capability, technical debt, and regulatory exposure under regimes such as the EU AI Act. It extends standard technology due diligence with hands-on machine learning assessment. How much does AI strategy consulting cost in the UK? Reported benchmarks put AI strategy and readiness assessments at roughly £15,000 to £50,000, with enterprise transformation programmes running from six figures upwards. Day rates range from about £400 to £3,500 for independents, £1,200 to £2,500 at boutiques, and £800 to £6,000 at the Big 4 depending on seniority. Treat all of these as indicative; scope drives everything. Are Big 4 consultants worth it for AI work? For large-scale delivery, risk transfer and board assurance, often yes, and no smaller provider can substitute. For strategy and due diligence, the answer depends on who is actually staffed on your engagement and how the firm's platform alliances relate to your decision. Ask both questions before signing. Can a single independent advisor handle enterprise AI strategy? Strategy, yes: it is judgement work, and one current senior specialist can outperform a large mixed-seniority team on it. Delivery at scale, no: implementation across business units needs a bench that independents do not have, which is why the strategy/build split above exists. Who should carry out AI due diligence on an investment? Someone with hands-on model evaluation experience and no financial relationship to the target's technology stack or to the deal outcome. That can exist inside a large firm, but you must verify staffing and conflicts explicitly; with a genuine independent specialist, the independence comes built in. --- ### Why I make clients pay for the part everyone wants for free - URL: https://agathon.ai/insights/why-i-make-clients-pay-for-the-part-everyone-wants-for-free - Published: 2026-06-26 - Categories: AI Strategy, AI Consulting, Generative AI On scoping AI projects when building has never been cheaper. Building software used to be the expensive part. It took months, it took a team, and it took money you couldn't get back. So it made sense to be sure of what you were building before anyone started, because the cost of being wrong was paid in burnt quarters. That logic ran the industry for decades, and most teams still carry it around. The trouble is it has quietly inverted, and the instinct it produced now points the wrong way. #### The inversion nobody priced in When building was the bottleneck, a rough brief was survivable. You would discover the holes during the build, slowly and expensively, but you would discover them. The long grind of construction forced the questions into the open. Generation has removed the grind. A capable team can now have a working prototype by Friday. And the moment building stops being the hard part, the hard part becomes the thing building used to smuggle in for free: knowing what is worth building. That is the one job a generator cannot do for you. So cheap generation did not lower the value of discovery. It raised it. The demo is cheap now. The judgment is the asset. You do not have to take my word for it: DORA's 2026 research found the same thing from the other direction, that as AI speeds up code generation the bottleneck moves to specification, and the spec stops being overhead and becomes the scarce resource. I am not preaching from above this. I like opening the laptop and making something. The pull to skip the talking and get to a prototype is one I feel more than most, because building is the part of the work I enjoy most. So when I argue for discovery I am not arguing against an instinct I lack. I am arguing against one I have to manage in myself. #### What a working demo tells you I watched that instinct play out from the outside not long ago. A fast-growing company came to me already certain what it wanted built. They had sketched the architecture themselves and had lessons from a previous tool; what they wanted from me was a partner to help ship it. To their credit, they moved. They ran a quick, scrappy version of the idea and got a spread of working demos across different parts of the business. The energy in the room was real. The question I asked them afterwards was not whether the demos were impressive. It was whether any of it had shipped, or whether it had stayed in demo form. Because those are different outcomes, and the distance between them is where the work lives. That distance is the whole point, and it is worth being precise about why it opens up. A working prototype answers one question well: can we build this? But that was almost never the question that mattered. The question that mattered was whether you should, and whether the thing you built is even what you wanted once it exists in front of you. Building tests feasibility. It does not test desirability, and the two feel similar right up until you are holding something that works and realising it solves a problem you do not have. ##### Why round two goes bigger and lands flatter Novelty hides this for a while. The first version carries a charge simply from being new, which is why a second attempt so often goes bigger and lands flatter. The novelty has worn off, and no structure was underneath it. By structure I mean the unglamorous things a one-off skips: a clear view of whose work changes, what the prototype is supposed to replace rather than merely demonstrate, and how you would know in a month whether it stuck. A quick build can dazzle without any of that, because for one afternoon the excitement does the load-bearing. The version that compounds has decided, before it builds, what it is willing to be measured against. Fast generation lets you reach the disappointment sooner when that decision was never made. It feels like speed. It is a longer road, because now you are unpicking a built thing instead of rethinking an idea, and you are attached to it. #### The better use of a cheap generator So what should a team do with all the time and money that cheap building has handed back? The temptation is to spend it building more, and faster. The better move is to spend it sparring. The same tool that will generate the thing for you will also argue with you about whether the thing should exist. That second use is worth far more in the phase where you are most tempted to skip it. Point the generator at your own reasoning rather than your output, and make it earn the idea before you build it. That is discovery, done in an afternoon, for free, and it is the part nobody wants to pay for precisely because it does not produce anything to demo. Which is why discovery is the line item clients most often want to cut, and the one I hold. Not because the deliverable is impressive to look at. Because it is the only part of the engagement that decides whether the impressive part was worth building. #### What is left when building is free The cost of building collapsed. The cost of building the wrong thing did not move at all. That is the whole of it. The discipline worth paying for was never the construction. Construction was only ever a proxy for the thing we wanted: something worth having built. Generation made the proxy cheap and left the real thing exactly as hard as it always was. It is no longer hidden inside the months it used to take to find out. Knowing what deserves to exist is still the job. --- ### What "sovereign AI" actually means, and when it matters - URL: https://agathon.ai/insights/what-sovereign-ai-actually-means-and-when-it-matters - Published: 2026-06-09 - Categories: AI Strategy, LLMs, AI Advisory The word "sovereign" is becoming standard vocabulary in enterprise AI. A British lab recently announced plans for a UK-trained frontier model, backed by government compute and a coalition of banks, defence primes and telcos. It will not be the last of its kind. Over the coming year, "sovereign", "UK-built" and "air-gapped" will appear across product literature, attached to offerings that mean quite different things by the terms. Most commentary settles into one of two positions. One treats any domestic model as a national achievement worth celebrating on its own terms. The other notes that a "sovereign" UK model will train on American chips and concludes the whole category is largely presentational. Both leave the more practical question unaddressed: for a given organisation, which workloads does sovereignty actually affect, and which does it leave untouched? Answering that requires being precise about what "sovereign" claims. It is not a single property. It describes control at four distinct layers, and a vendor can accurately use the label while offering only one of them. #### The four layers The first layer is data. This concerns where inputs and outputs are stored and processed, and who can compel access to them. It is the layer most vendors have in mind when they describe a product as sovereign, partly because it is the most straightforward to provide. UK or EU data residency, a commitment not to train on customer data, and customer-managed encryption keys all sit here. The major US providers already offer these on their enterprise tiers. The second layer is compute. This concerns whose hardware the model runs on, and in which jurisdiction. A model hosted in a London data centre may still run on infrastructure owned by a US company, which carries legal consequences discussed below. Compute sovereignty in the fuller sense means the hardware sits within the customer's own boundary, or with a provider genuinely outside foreign legal reach. The third layer is weights. This concerns whether the customer controls the model itself: whether they can hold the weights, run them independently, inspect them, freeze a particular version, and continue operating if the vendor withdraws or changes its terms. Access to a model through an API is not the same as control over it, and most offerings described as sovereign stop well short of providing the latter. The fourth layer is governance. This concerns who determines what the model will and will not do, how it is updated, and what constraints it carries. When a US provider revises a usage policy or retires a model, customers are informed of the decision rather than involved in it. Full sovereignty means control at all four layers, and very little on the market provides that. What is more commonly available is data residency presented under a sovereign label. This is not misleading, and for many purposes it is sufficient. It is simply worth identifying which layer a given product addresses before attaching a premium to the term. #### Where it matters, and where it does not For a large share of common enterprise AI work, the distinction has limited practical weight. Drafting, summarising, coding assistance, customer-service triage, and general analysis of data that is either public or properly de-identified are all well served by a correctly configured US enterprise tier. UK or EU data residency, a signed data-processing agreement, zero data retention, and customer-held encryption keys, combined with a documented transfer impact assessment, address the substantial majority of compliance exposure for this kind of work. A sovereign model adds little that such an organisation can use. This is not a criticism of sovereign offerings. It reflects the fact that most data is not sensitive enough to require sovereignty in the first place. The cases where it does matter are a minority precisely for that reason, not because the capability is superficial. When data is sufficiently sensitive, the argument shifts from positioning to something structural, and the structural argument is worth understanding clearly. #### Why location is not the same as jurisdiction The strongest case for sovereignty rests on a feature of US law rather than on national preference. Under the CLOUD Act, US authorities can require a US-controlled company to produce data regardless of where in the world that data is held. A London data centre operated by a US provider remains under the control of an entity subject to US legal process. Separately, Section 702 of the Foreign Intelligence Surveillance Act permits intelligence collection on non-US persons under a broader standard than many organisations assume. The consistent principle is that server location does not determine legal reach. Corporate ownership does. No contractual term fully removes this exposure. A UK data region governs where data sits at rest; it does not alter who can be compelled to disclose it. For most commercial data, the practical likelihood of such powers being exercised is low, and accepting that residual risk is a reasonable position. For certain categories of data, however, a low probability of foreign state access remains unacceptable, and residency configuration does not resolve it. That distinction marks the point at which sovereignty begins to justify its cost. #### A practical way to assess it An organisation can determine its own exposure without relying on a vendor to define it, by sorting its AI workloads. Two questions carry most of the analysis. The first is the most sensitive data class that touches a given workflow. Not the typical input, but the single most sensitive one, since this is the question most easily understated when a convenient answer is preferred. The second is the genuine consequence were a foreign state, or any third party, able in principle to access that data during processing. The consequence may be reputational, regulatory, operational, or related to national security, or it may be negligible. Together these questions tend to sort workloads into three groups. The largest group generally involves public or low-sensitivity data with no meaningful consequence from third-party exposure. For this work, the appropriate choice is the most capable available tool, and a correctly configured US enterprise tier is well suited to it. Sovereignty offers no relevant benefit. A smaller group involves regulated personal data, where the consequences are real but manageable. Here the task is configuration rather than replacement: enterprise tier, in-region residency, zero data retention, customer-held encryption keys, and a documented transfer impact assessment. The aim is to manage the risk rather than to remove the provider. A genuine minority involves data where the combination of sensitivity and consequence makes residual foreign-jurisdiction exposure unacceptable at any probability. Classified or national-security data, certain financial-crime intelligence, the most sensitive identifiable health data where de-identification is insufficient, and legally privileged material all fall here. For this group, full sovereignty or genuine air-gapped deployment is a requirement rather than an enhancement. The value of the exercise lies in the proportions it reveals. In most organisations the third group is small and the first is large, and that distribution is what gives the framework its practical use. The relevant question is rarely whether to adopt a sovereign model in general. It is how much of the estate genuinely sits in the third group, and for most organisations the answer is a modest share. #### The position it leaves A credible sovereign option, where one exists, is worth having for two reasons. It addresses the third group of workloads, which existing tools may not be able to serve at all. And its availability provides leverage in negotiations with current providers over price, data terms and exit rights, since a viable alternative improves an organisation's position even if it is never adopted. Neither reason argues for waiting on the broader AI roadmap, and neither supports committing a production system to a model that has not yet shipped or been independently evaluated. The more durable step is to understand the estate first: to run the three-group assessment and establish how much of the organisation's work genuinely requires control at the compute and weights layers, rather than at the data layer that existing enterprise tooling has most likely already addressed. With that understanding in place, the next sovereign offering to arrive can be read for what it is. The relevant layer becomes clear, and so does whether it answers a need the organisation actually has. --- ### Inside sovereign AI governance: How it works when the frameworks meet reality - URL: https://agathon.ai/insights/inside-sovereign-ai-governance-how-it-works-when-the-frameworks-meet-reality - Published: 2026-05-28 - Categories: AI Strategy, Responsible AI, AI Advisory Most organisations treat AI governance as a compliance exercise. They map controls to regulations, produce documentation, and declare themselves governed. This creates a dangerous illusion. The gap between publishing a governance framework and operating one under production conditions is where real risk accumulates, quietly and without triggering any of the checkboxes that were supposed to catch it. Sovereign AI governance compounds this problem by layering jurisdictional politics, infrastructure dependencies, and geopolitical tension onto an already fragile compliance apparatus. The result is a domain where the organisations that look most governed on paper are often the least prepared for the scenarios that matter. #### The compliance theatre problem ##### Why ticking boxes creates a false sense of control A systematic review of 13 leading trustworthy AI audit and assurance frameworks found that none simultaneously achieved advanced capability across governance, operations, and audit pillars. The researchers described this as a structural "posture-ready vacancy" in the current state of the art (Trustworthy AI Posture framework, 2025). Existing frameworks cluster into three disconnected groups: process-heavy but execution-weak frameworks, principle-heavy but audit-light policy instruments, and technically continuous but governance-disconnected technical proposals. This fragmentation explains why compliance exercises feel productive without being protective. Organisations complete risk assessments, draft policies, and stand up committees. The documentation grows. The actual ability to detect, respond to, and learn from governance failures does not grow with it. Operational evidence is rarely systematically linked to governance claims, making demonstration of control adequacy inconsistent and resource-intensive. ##### The gap between framework design and operational reality The Unified Control Framework project synthesised 15 risk types and approximately 50 risk scenarios from existing governance literature, then discovered that the "operational" risk type was completely missing from established frameworks. Risk scenarios like "lack of inference data transparency" appeared in none of the existing taxonomies the researchers examined. The final control library contained 42 controls, each mapped to an average of 4.1 risk scenarios, but the authors acknowledged a critical limitation: they could not assess the degree of risk mitigation achieved by implementing any specific control. This is the governance gap in miniature. Frameworks enumerate what should be controlled. They rarely measure whether those controls reduce the risks they target. The distance between "we have a control for this" and "this control works in production" is where governance fails silently. ##### What regulators actually look for versus what organisations prepare Organisations prepare documentation. Regulators increasingly want operational evidence. Under the EU AI Act, signatories must provide unredacted access to their systemic risk management framework and updates to the AI Office within five business days of confirmation. Model documentation must be retained for 10 years after the model is placed on the market, with annual framework reviews required. Non-compliance with prohibited AI practices carries fines of up to EUR 35 million or 7% of worldwide annual turnover, whichever is higher. The direction is clear: regulators are moving toward continuous, auditable proof of governance rather than periodic self-attestation. Organisations that invest heavily in documentation without building the operational infrastructure to generate this evidence are preparing for the wrong exam. #### What sovereign AI governance demands in practice ##### Jurisdiction as a design constraint, not an afterthought Sovereignty in AI is not a policy preference. It is an architectural constraint that propagates through every layer of the technology stack. Brookings Institution research characterises AI as a transnational stack with concentrated choke points across minerals, energy, compute hardware, networks, digital infrastructure, data assets, models, applications, and crosscutting enablers of talent and governance. Full-stack AI sovereignty is structurally infeasible for almost any country because of these transnational dependencies. This means jurisdiction cannot be bolted on after system design. Every decision about where data resides, which models process it, who can access inference logs, and which legal regime governs disputes must be considered during architecture, not during compliance review. Canada's national AI strategy demonstrates this by explicitly distinguishing between what must be sovereign (compute resources for sensitive workloads) and what can be procured commercially (global foundation models). That distinction, made at the strategy level, is precisely the kind of design constraint that governance frameworks need to encode. ##### Data residency, model provenance, and the supply chain question The US CLOUD Act of 2018 enables American authorities to access data held by US cloud providers abroad without authorisation or knowledge of the host country. Microsoft told the French Senate it could not guarantee that French citizens' data would not be transmitted to US authorities without explicit French government authorisation. These are not theoretical risks. They are documented capabilities of the legal infrastructure that underpins global cloud computing. A study of 775 non-US data centre projects found that US companies served as operators for 18% of those projects while accounting for 48% of total data centre investment and 56% of AI investment. Even countries that build nominally "sovereign" facilities often rely on US hyperscalers for operations. AWS announced plans for a €7.8 billion European Sovereign Cloud; France's Bleu cloud de confiance is operated by Capgemini and Orange but built on Microsoft technology. The gap between sovereignty aspiration and operational reality is substantial, and governance frameworks that ignore supply chain provenance are governing a fiction. Model provenance raises parallel concerns. When an organisation deploys an open-weight model wrapped in a proprietary execution framework, the EU AI Act's rigid legal separation between "provider" and "deployer" breaks down. Research on the AI Act's applicability to agentic systems found that deployers who fundamentally alter agency by wrapping models with execution frameworks fall into a regulatory grey zone with no clear framework for attributing liability. ##### When "sovereign" clashes with "scalable" Sovereign AI systems can fragment markets, slow global AI development, reduce economic competitiveness, and become tools for digital authoritarianism without proper governance safeguards. This tension between sovereignty and scale is not resolvable through frameworks alone. It requires explicit strategic choices about which layers of the AI stack a given jurisdiction will control directly, which it will procure under negotiated terms, and which it will accept as dependencies. India's approach illustrates one resolution: application-led sovereignty through multilingual foundation models, voice systems, and AI-enabled interfaces adapted to local languages and service needs, rather than competing at frontier scale. The EU AI Continent Action Plan commits approximately €200 billion to developing infrastructure, increasing data centre capacity, and supporting local industry through procurement policy. These are different bets, reflecting different risk appetites and industrial strategies. Governance must be designed to support whichever bet an organisation's jurisdiction has placed, not to paper over the tensions between sovereignty and scalability. #### The institutional machinery behind governance decisions ##### Who sits at the table and who gets left out Deloitte reports that 72% of boards have one or more committees responsible for risk oversight, and more than 80% have one or more risk management experts. Yet governance of AI systems requires a different kind of expertise than governance of financial or operational risk. Boards are advised to recruit AI professionals with operational experience implementing successful AI projects, not just risk management generalists. The Unified Control Framework was validated through structured interviews with six to seven AI governance practitioners representing diverse roles: data scientists, AI governance consultants, enterprise practitioners, and policy experts. Even in a research context, assembling the right mix of perspectives required deliberate effort. In production governance, the default composition of risk committees tends to over-represent legal and compliance functions while under-representing the engineering staff who understand how models behave, what failure modes look like, and what monitoring is feasible. ##### How risk appetite shapes policy more than risk assessment does Risk assessment is the visible machinery of governance. Risk appetite is the invisible force that determines which assessment findings get acted on. Tolerance policies must account for multiple risk sources: financial, operational, safety and wellbeing, business, reputational, and model risks, according to the NIST AI Risk Management Framework. But in practice, organisations rarely make their risk appetite explicit across all these dimensions. The result is governance that appears comprehensive on paper while systematically ignoring categories of risk that the organisation has implicitly decided to accept. When a model exhibits bias that regulatory guidance flags as high-risk, the response depends less on what the risk assessment says than on how much reputational, legal, or financial exposure the leadership team is willing to tolerate. Making risk appetite explicit and documenting those trade-offs is harder than producing a risk register, which is precisely why most governance frameworks avoid requiring it. ##### The role of technical staff in non-technical governance bodies The NIST AI RMF explicitly calls for separating test and evaluation professionals from AI system developers, with independent staff reporting through risk management functions to counter groupthink and ensure course-correction. This structural recommendation addresses a real failure mode: when the people building systems are also the people assessing whether those systems are governed, the assessment becomes self-referential. Diverse teams with varied experience, disciplines, and backgrounds are better equipped to anticipate AI risks, but require explicit senior leadership commitment. Without that commitment, organisational incentives override diversity benefits, and governance committees default to the perspectives of their most senior (and typically least technical) members. #### Where frameworks fracture under pressure ##### Incident response when the model behaves unexpectedly AI systems are inherently dynamic and may perform unexpectedly after deployment. The NIST framework treats incident response plans as standard governance practice, not an optional add-on. Yet the Unified Control Framework analysis found that incident response (Control-042) had to be added after initial framework development because it was not captured in the original 41 controls derived from existing governance literature. Incident reporting timelines under the EU AI Act vary by severity: cybersecurity breaches within 5 days, operational disruptions within 2 days, deaths within 10 days, and serious health or environmental harm within 15 days. These are tight windows that require pre-established processes, clear escalation paths, and technical infrastructure for root cause analysis. Organisations that treat incident response as a section in a policy document rather than a rehearsed operational capability will discover the gap between framework and reality at the worst possible moment. The challenge intensifies with agentic systems. Research on the EU AI Act's applicability to agentic architectures found that the Act's reliance on "reasonably foreseeable misuse" is structurally flawed because agentic systems dynamically generate novel, unprogrammed execution paths that are inherently unforeseeable by original developers. You cannot write an incident response plan for failure modes that emerge from continuous reason-act-observe loops if your governance model assumes all risks can be enumerated in advance. ##### Cross-border obligations that contradict each other A single AI decision can simultaneously violate GDPR, the Digital Services Act, and the AI Act, triggering multiple investigations and enforcement mechanisms with potentially cumulative fines. Research on global AI governance illustrates this with a social media platform whose biased content moderation system could violate all three frameworks at once. The contradiction problem worsens across borders. The NIST AI RMF identifies a fundamental regulatory tension where AI debiasing techniques that rely on demographic information can conflict with legal prohibitions on intentional discrimination. An organisation operating across the EU, US, and Asia-Pacific faces a patchwork of requirements: China's Personal Information Protection Law requires local data storage, the EU AI Act imposes conformity assessments on high-risk systems, and the US lacks federal AI legislation while individual states adopt divergent approaches. Only five of fifty US states have adopted comprehensive data legislation, leaving California's Consumer Privacy Act as the de facto US data regulation. Governance frameworks that assume a single jurisdictional context cannot handle these contradictions. Organisations need governance architectures that explicitly model where obligations conflict and define resolution strategies rather than pretending coherence exists. ##### The procurement trap: vendor lock-in disguised as compliance Of the eight cloud service providers approved for Canadian government use, seven are American. ThinkOn is the only non-American company on the list. When compliance requirements point toward a small set of approved vendors, procurement decisions made in the name of governance can create dependencies that undermine the sovereignty those decisions were meant to protect. Current sovereign cloud contracts frequently obscure details about data access, algorithms, and operational control from legislative scrutiny. Microsoft returned $9.7 billion to shareholders through dividends and buybacks in Q2 2025 while simultaneously scaling sovereign cloud infrastructure for governments. The commercial incentives of cloud providers and the governance objectives of their government clients are not naturally aligned, and procurement frameworks that treat vendor certification as equivalent to governance assurance miss this structural tension. Saudi Arabia's SDAIA National Data Governance Platform illustrates the extreme case: biometric databases connected with predictive policing algorithms, citizen sentiment analysis from social media, all built on sovereign cloud infrastructure that satisfies data residency requirements while enabling surveillance capabilities that many governance frameworks would classify as unacceptable risk. #### Building governance that survives contact with production ##### Continuous assurance over periodic audit Traditional point-in-time audits cannot scale with the dual expansion of vertically evolving governance complexity and horizontally distributed, agentic AI deployments operating at machine speed. The shift from periodic audit to continuous assurance is not an incremental improvement. It is an architectural change to how governance operates. Research on runtime compliance monitoring found that inter-judge agreement rates on regulatory compliance ranged from 51.5% to 69.1% across five regulatory criteria when using small language models as automated judges. Question-order bias alone degraded agreement by up to 25 percentage points. Three structural failure modes emerged: truth bias (systematic default to "compliant"), reasoning/output dissociation (correct violation detection paired with false compliance scores), and prompt architecture sensitivity. These findings suggest that automated compliance monitoring, while necessary, introduces its own governance challenges that must be understood and managed. The EU AI Act's high-risk compliance obligations, originally set to apply from August 2026, were postponed to December 2027 under the AI Act Omnibus. This delay reflects the difficulty of operationalising continuous compliance, not a lack of regulatory ambition. ##### Embedding governance into engineering workflows EU AI Act compliance analysis requires gathering compliance information from multiple supply chain components, harmonising that information, and rendering a compliance prediction across all components. Research on automated compliance analysis found that current processes are too complex and time-consuming to enable rapid verification during development. Compliance analysis that happens only at deployment checkpoints misses the engineering decisions made months earlier that determine whether a system can be governed at all. The practical implication is that governance must be embedded into engineering workflows: into CI/CD pipelines, model registries, data catalogues, and deployment automation. A Policy Abstraction Pattern that decouples regulatory obligations from technical implementations allows the same assurance mechanics to operate across multiple jurisdictions without requiring code-level refactoring. This is governance as infrastructure, not governance as oversight. Research on translating AI Act requirements into verification activities found that decomposing legal requirements into operational sub-requirements grounded in authoritative standards reduces interpretive uncertainty. Verification activities characterised along two dimensions (type of verification and lifecycle target) create a reusable reference for consistent compliance verification. The key insight is that governance becomes tractable when it is expressed in the same language as the engineering systems it governs. ##### Red-teaming your own governance model Governance frameworks contain assumptions about what risks exist, how they manifest, and what controls are adequate. Those assumptions can be wrong. The Unified Control Framework's discovery that operational risk was entirely absent from existing frameworks illustrates how blind spots persist in mature governance literature. If the frameworks themselves have gaps, organisations that implement them faithfully will inherit those gaps without knowing it. Red-teaming governance means testing whether your controls work under adversarial conditions, not just whether they exist. It means staging scenarios where cross-border obligations contradict each other and observing whether your governance machinery produces a coherent response. It means having technical staff attempt to deploy a non-compliant model through your standard workflow and seeing whether your controls detect it. Governance that has never been tested under stress is governance that has never been tested. #### The geopolitics underneath the technical standards ##### How trade policy and AI regulation have become inseparable The US Department of Commerce banned Nvidia from selling A100, A100X, and H100 graphics processing units to customers in China in 2022, explicitly using trade policy to constrain AI capability development. The United States and European Union have each passed major semiconductor bills in response to AI competition with China. Export controls, compute access, and regulatory frameworks are no longer separate policy domains. They are instruments of a single strategic competition. The Pentagon demanded guardrail-free access to Anthropic's Claude models and threatened to invoke the Defence Production Act or designate the company as a supply chain risk if it refused. Anthropic refused, drawing a hard line against mass domestic surveillance and fully autonomous weapons use cases. This confrontation illustrates a tension that governance frameworks must acknowledge: the same governments that set AI governance standards also have strategic interests that can conflict with those standards. ##### Competing visions of sovereignty across the EU, US, and Asia-Pacific The United States hosts approximately 75% of global AI supercomputer performance, with China at 15% and the rest of the world at 10%. Europe invested €47 billion in AI infrastructure while US firms plan at least $650 billion in AI-related capital expenditure in a single year. These numbers define the playing field on which sovereignty discussions occur. India hosted the February 2026 AI Impact Summit, bringing sovereign AI to the international stage. India's governance approach prioritises equitable access, climate resilience, and inclusive growth rather than frontier model development, reflecting sovereignty concerns that differ fundamentally from those of the US or EU. China's approach uses regulation as a tool of state control: the Internet Information Service Algorithmic Recommendation Management Provisions require companies to promote content following the Communist Party's line while restricting unfavourable content. China had an estimated 626 million facial recognition cameras installed by 2020. These competing visions mean that "sovereign AI governance" carries radically different meanings depending on who is defining it. Governance frameworks that assume a shared understanding of what sovereignty means, or what it is for, will fail when applied across jurisdictions with incompatible political commitments. ##### The hidden influence of cloud infrastructure providers Without interoperable standards, governments import pre-configured intelligence: models trained elsewhere that reflect foreign assumptions about acceptable risk, accountability, and social values. As AI systems evolve from static models into agentic systems capable of autonomous tool invocation and database access, the interfaces governing those interactions become strategic choke points. Middle-power countries' AI sovereignty depends not on replicating frontier model development but on ensuring systems can be integrated, governed, audited, and replaced on national terms. Open standards preserve optionality by allowing governments to adapt rules over time, switch providers, and layer domestic priorities onto shared technical foundations. The practical recommendation from governance researchers is that governments should form interoperability blocs by aligning technical standards with neighbouring economies to create collective markets large enough to compel global AI providers to comply. Canada's Directive on Automated Decision-Making offers a concrete model: government departments must conduct Algorithmic Impact Assessments, publish reports publicly, and provide recourse mechanisms for affected citizens. This transparency requirement operates at the governance layer regardless of which cloud infrastructure provider operates underneath. Governance that depends on the goodwill of infrastructure providers is not governance. Governance that can be verified independently of infrastructure providers is. #### From framework to operating rhythm ##### Governance as a living system, not a document The EU AI Act requires Member States to establish national AI regulatory sandboxes for testing and validation of innovative AI systems under regulatory supervision. Yet sandbox participation is voluntary, and sandboxes may prove unattractive to innovators due to confidentiality concerns, inability to relax legal rules during the sandbox period, and inability to deliver presumption of conformity with the AI Act. Differing approaches taken by individual national sandboxes risk undermining uniform interpretation of the Act, potentially motivating innovators to engage in sandbox arbitrage. This illustrates a broader principle: governance instruments that do not adapt to how organisations and regulators learn are governance instruments that will be routed around. The five-layer AI governance framework spanning from regulatory mandates through standards, assessment methodologies, certification processes, and operationalisation identified critical gaps including missing standardised assessment procedures and reporting mechanisms. Frameworks are starting points, not endpoints. The organisations that treat governance as a living system, continuously updated based on operational experience, incident data, and regulatory evolution, will be the ones whose governance survives contact with production. ##### The signals that your governance model is failing silently Several indicators suggest governance exists on paper but not in practice. Your risk register has not been updated since it was created. Your incident response plan has never been exercised. Your technical staff cannot describe the governance process without consulting documentation. Your compliance controls have never rejected a deployment. Your cross-border data flows have not been re-evaluated since your last regulatory mapping. The most telling signal is absence. Governance that never generates friction, never delays a release, never surfaces a finding that changes a decision, is governance that is not operational. Real governance creates tension between speed and safety, between capability and compliance. If that tension is invisible, either the governance is not being applied or the risk tolerance is so high that the governance is decorative. ##### Starting points for organisations at different maturity levels Organisations beginning their sovereign AI governance journey should start with dependency mapping: cataloguing where their AI stack relies on specific jurisdictions, vendors, and infrastructure providers, and identifying which of those dependencies create governance obligations they are not currently meeting. Organisations with established governance should invest in operational testing: exercising their incident response processes, stress-testing their cross-border compliance logic, and verifying that their controls produce the evidence regulators will require. The shift from "we have a framework" to "our framework works" is where most governance programmes stall. Organisations at advanced maturity should focus on continuous assurance infrastructure: embedding governance into engineering pipelines, automating compliance evidence generation, and building the technical capability to demonstrate governance posture in real time rather than through periodic reports. The managed interdependence model, mapping dependencies by layer, prioritising feasible interventions, diversifying suppliers, and embedding interoperability through technical standards, provides a strategic template for governance that acknowledges the reality of transnational AI supply chains while maintaining meaningful sovereign control. --- Sovereign AI governance is where technical architecture, regulatory complexity, and geopolitical strategy converge. Getting it right requires more than frameworks and checklists. It demands architectural thinking, operational discipline, and the strategic clarity to distinguish between governance that protects and governance that merely performs. If you are building AI systems that must operate across jurisdictions, withstand regulatory scrutiny, and exploit the full technical potential of what modern AI makes possible rather than settling for the fraction most organisations achieve, get in touch with Agathon. #### References - Mind the Gap: How the Technical Mechanisms of Agentic AI Outpace Global Legal Frameworks - The Unified Control Framework: Establishing a Common Foundation for Enterprise AI Governance, Risk Management and Regulatory Compliance - Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring - A Five-Layer Framework for AI Governance: Integrating Regulation, Standards, and Certification - Global AI Governance Overview: Understanding Regulatory Requirements Across Global Jurisdictions - Is AI Sovereignty Possible? Balancing Autonomy and Interdependence - The AI Sovereignty Paradox at Home and Abroad - AI Governance for Board Members - Compliance Cards: Automated EU AI Act Compliance Analyses amidst a Complex AI Supply Chain - Assessing High-Risk AI Systems under the EU AI Act: From Legal Requirements to Technical Verification - NIST AI RMF Playbook — Govern - Trustworthy AI Posture (TAIP): A Framework for Continuous AI Assurance of Agentic Systems - The Geopolitics of AI and the Rise of Digital Sovereignty - Succeeding in the AI Competition with China: A Strategy for Action - What Does a 'Sovereign Cloud' Really Mean? - How National AI Clouds Undermine Democracy - Rethinking Sovereign AI as Strategy - Why AI Sovereignty Depends on Interoperability Standards - Operationalising AI Regulatory Sandboxes under the EU AI Act: The Triple Challenge of Capacity, Coordination and Attractiveness to Providers --- ### We've deployed Microsoft Copilot 365. Now what? - URL: https://agathon.ai/insights/weve-deployed-microsoft-copilot-365-now-what - Published: 2026-05-28 - Categories: AI Strategy, AI Consulting, Generative AI The procurement cycle is over. Licences are assigned. The launch email went out with the right mix of enthusiasm and corporate caution. Microsoft Copilot is live across your Microsoft 365 environment, and the project team is already moving on to their next initiative. This is precisely the moment where most organisations lose the plot. Deployment is a logistics exercise. What follows is an organisational design challenge, a data governance reckoning, and a sustained change management programme rolled into one. The companies that treat go-live as the finish line will spend the next eighteen months wondering why their per-seat investment isn't translating into measurable business improvement. The companies that treat it as a starting line will compound their advantage quarter over quarter. The difference between those two outcomes has almost nothing to do with the technology. #### The deployment was the easy part ##### Why go-live is where most organisations stall There's a pattern that repeats across enterprise software adoption, and AI tools amplify it. The procurement and deployment phases have clear ownership, defined milestones, and executive attention. The moment the tool is live, ownership fragments. IT considers the project delivered. The business assumes value will materialise organically. Leadership moves its attention to the next strategic priority. Research from the California Management Review, drawing on Deloitte's CFO Survey, reports that fewer than 40 percent of automation initiatives deliver measurable value. McKinsey's Global AI Survey found that only 30 percent of AI pilots transition to scaled impact. These aren't failure rates for bad technology. They're failure rates for good technology meeting unprepared organisations. ##### The gap between activation and adoption Activation is binary: the tool is on or off. Adoption is a spectrum, and most organisations cluster at the shallow end. Microsoft's own Work Trend Index found that 78% of AI users are bringing their own AI tools to work without organisational guidance or clearance. When your employees are bypassing the tool you've paid for in favour of consumer-grade alternatives, you don't have a technology problem. You have a relevance problem. The gap between activation and adoption is where Copilot either becomes embedded in how your organisation works or becomes another underused line item in your SaaS spend. ##### What "deployed" actually means Deployed means licences are assigned and the software is accessible. It doesn't mean people know what to do with it. It doesn't mean your data is structured in ways that make Copilot's outputs useful. It doesn't mean your workflows are designed to take advantage of AI augmentation rather than simply tolerating it. A Gartner survey published in October 2024 found that employees find value in Microsoft 365 Copilot, but tangible business impact remains elusive. Enablement activities and security mitigation take more effort than anticipated. This shouldn't surprise anyone who has watched enterprise technology adoption cycles before, but it does seem to surprise the executives who approved the budget. #### Measuring what matters beyond licence utilisation ##### The metrics that mislead The default measurement for any SaaS deployment is utilisation: how many people are logging in, how often, and for how long. For Copilot, this translates into tracking who is using the AI features in Word, Excel, Teams, and Outlook. These metrics are comforting and nearly useless. High utilisation tells you people are clicking buttons. It tells you nothing about whether the outputs are improving decisions, reducing errors, or accelerating work that matters. Low utilisation might indicate poor adoption, or it might indicate that the tool isn't relevant to certain roles. Without context, the dashboard is just a distraction. Microsoft's Work Trend Index found that 59% of leaders worry about quantifying the productivity gains of AI. That worry is well-founded, but the solution isn't better utilisation tracking. It's connecting tool usage to outcomes the business already cares about. ##### Linking Copilot usage to business outcomes Gartner's peer community research on AI productivity tools identifies a more useful measurement framework: one that extends beyond time savings to include decision quality, error reduction, speed of knowledge discovery, and cross-team collaboration improvements. Business outcomes like faster feature development lifecycle and reduced incident recovery times provide a fuller picture than efficiency metrics alone. The practical challenge is attribution. If a product team ships features 15% faster after Copilot adoption, how much of that improvement comes from the tool versus other concurrent changes? Perfect attribution is impossible, but directional measurement is not. Track the outcomes you care about before deployment, establish baselines, and measure trends. The goal isn't to prove Copilot's ROI to the penny. The goal is to understand where it's creating leverage and where it isn't, so you can redirect effort accordingly. ##### Building a measurement framework that survives quarterly reviews Most measurement frameworks die because they require manual data collection, depend on self-reported surveys, or track metrics nobody in the C-suite is asking about. A durable framework ties Copilot metrics to existing business reviews. If your leadership team already reviews cycle time, customer satisfaction, employee engagement, and revenue per employee, measure Copilot's impact through those lenses. If your quarterly business review doesn't include AI adoption metrics by Q2 of your rollout, it probably never will. The measurement framework needs an owner, a reporting cadence, and a direct line to someone with budget authority. #### The prompt literacy problem no one wants to own ##### Why training sessions don't stick Microsoft's Work Trend Index reports that only 39% of people globally who use AI at work have received AI training from their company, and only 25% of companies plan to offer generative AI training this year. The organisations that do provide training tend to run a single session, distribute a PDF of "helpful prompts," and consider the job done. This approach fails for the same reason that a one-day Excel training course doesn't produce spreadsheet experts. Prompt literacy is a skill that develops through repeated practice in context, not through classroom instruction. The half-life of a training session that isn't reinforced by daily application is measured in days, not months. ##### Building prompt competence into daily workflows The Work Trend Index data reveals something more useful than training completion rates: frequent experimentation with different ways of using AI is the number one predictor of whether someone becomes a power user. Power users are 68% more likely to frequently experiment with different approaches. This suggests a different model for building prompt competence. Rather than formal training programmes, organisations should create structured opportunities for experimentation within existing workflows. Weekly team prompting challenges. Shared prompt libraries that evolve based on what works. Short peer-led sessions where teams share techniques relevant to their specific work. Wolters Kluwer took this approach, scheduling one hour per week for peer-led learning sessions focused on practical GenAI applications, which boosted employee skills, innovation, and retention. ##### Identifying and empowering your internal champions Every organisation has people who adopt new tools faster than their peers. With Copilot, these early adopters are your most valuable change management asset, but only if you find them and give them a role. Power users, according to Microsoft's research, are 61% more likely to hear from their CEO about the importance of using generative AI at work, and 66% more likely to redesign business processes with AI. These aren't just enthusiastic individuals. They're people operating in environments where leadership signals matter and process change is encouraged. Identify your power users within the first 60 days. Give them dedicated time to experiment, a channel to share findings, and visible recognition from leadership. Their practical knowledge of what works in your specific context is worth more than any vendor-led training programme. #### Data quality as the silent bottleneck ##### What Copilot reveals about your information architecture Copilot is, at its foundation, a retrieval and synthesis tool. It pulls from your emails, documents, chats, and meeting transcripts to generate responses. The quality of those responses is bounded by the quality of what it finds. Most organisations discover uncomfortable truths about their information architecture within weeks of Copilot deployment. Documents are inconsistently named. Critical knowledge lives in email threads rather than shared repositories. Meeting notes are scattered across personal notebooks. SharePoint sites haven't been maintained in years. Gartner's survey found that the value delivered by Microsoft 365 Copilot is closely correlated with the degree of information management maturity, and that optimal value capture may require the reengineering of information assets. This is an expensive finding for organisations that assumed Copilot would work well on top of their existing data. ##### Permissions, governance, and the oversharing risk When Copilot searches across your Microsoft 365 environment, it respects existing permissions. If someone has access to a SharePoint site they shouldn't, Copilot will surface that content in their responses. The tool doesn't create new security vulnerabilities, but it makes existing ones far more discoverable and exploitable. Research from the California Management Review on agentic AI adoption identifies AI-powered data leaks as organisations' top security concern, yet many businesses still lack AI-specific security controls. A permissions audit before or immediately after Copilot deployment isn't optional. It's a prerequisite for responsible use. The practical challenge is that permissions in large Microsoft 365 environments are often a mess. Overshared team sites, legacy distribution lists with broad access, and inconsistent classification of sensitive documents create an attack surface that Copilot makes trivially easy to exploit, even accidentally. ##### Cleaning the foundations without stopping the work Information architecture remediation sounds like a multi-year governance programme, and it can be. But it doesn't have to start that way. Prioritise the data sources Copilot accesses most frequently: recent documents, active SharePoint sites, current Teams channels. Clean those first. Establish naming conventions and metadata standards for new content going forward. Accept that legacy content will be imperfect and focus governance effort where it will have the most immediate impact on Copilot output quality. #### Workflow redesign, not just workflow acceleration ##### The difference between faster processes and better ones The default assumption about AI productivity tools is that they make existing work faster. Write emails faster. Summarise meetings faster. Generate first drafts faster. This is true and profoundly insufficient. Research by Erik Brynjolfsson and colleagues demonstrates that productivity gains materialise only when firms redesign workflows around digital tools. Bain's research reinforces this: organisations that combine workflow redesign with workforce modernisation demonstrate 10% to 15% productivity lift and 10% to 25% EBITDA gains. The gains don't come from speed alone. They come from eliminating steps, reducing handoffs, and restructuring how decisions get made. Bain's research also quantifies the cost of ignoring this: AI amplifies whatever system it's dropped into. If workflow debt isn't addressed first, AI and automation multiply complexity instead of boosting productivity. Most organisations carry substantial workflow debt from accumulated meetings, approvals, handoffs, exceptions, and one-off policies that make even simple tasks hard to execute. Weak management systems and poor deployment of human capital drain companies of up to 40% of their productive power. ##### Where Copilot creates new possibilities versus where it just saves time Copilot saving someone ten minutes on an email draft is time savings. Copilot enabling a product manager to synthesise customer feedback across hundreds of support tickets, meeting transcripts, and survey responses in minutes rather than weeks is a capability shift. The difference matters because time savings are linear and capability shifts compound. A study of Boston Consulting Group consultants found that when GenAI tools matched the task, productivity increased by 12% and speed of task completion by 25%. The key qualifier is "matched the task." Copilot applied to the wrong workflow produces mediocre outputs faster. Applied to the right workflow, it creates analytical capabilities that didn't previously exist at that speed or cost. ##### Spotting the processes worth rebuilding from scratch Not every process benefits from AI augmentation. Some processes need to be replaced entirely. The candidates for wholesale redesign share common characteristics: they involve multiple handoffs between people or systems, they rely on information synthesis across disparate sources, they produce outputs that require significant human review before they're useful, and they've grown more complex over time without anyone questioning whether the complexity is necessary. Bain's case study of a UK banking group illustrates the potential: the organisation compressed a 60- to 100-day customer engagement process involving over ten handoffs into a one-day cycle through AI-enabled workflow redesign. That's not acceleration. That's reconception. #### Managing the organisational politics of AI tools ##### When enthusiasm outpaces capability Some teams will embrace Copilot immediately and start using it for everything. This creates its own problems. Enthusiasm without competence produces confidently generated outputs that nobody verifies, AI-assisted decisions that skip critical thinking steps, and a false sense of productivity that masks declining quality. MIT Sloan Management Review and Boston Consulting Group research highlights this risk: the marginal cost of a first attempt has dropped sharply with generative AI, but what remains expensive is evaluating what gets generated after the output arrives. Organisations must prioritise verification (does the output meet the standard?), evaluation (what does the output reveal?), and learning capture (how do we ensure insights persist?). A study of call centre agents given access to a GenAI conversational assistant found productivity improvements of at least 14%, along with higher service quality and faster onboarding. But the same research found that lower-performing employees received a bigger productivity boost than more experienced workers. The implication is uncomfortable: AI tools can compress the gap between your strongest and weakest performers, which means the quality of AI-generated outputs varies significantly depending on who's reviewing them. ##### Handling the teams that refuse to engage Every Copilot deployment has holdouts. Some resistance is principled: legal teams with legitimate concerns about confidentiality, engineering teams with valid questions about code quality, or compliance teams worried about audit trails. Address these concerns directly with specific governance policies, not with generic reassurance. Other resistance is cultural. Research from the California Management Review on AI adoption found that employees often fear AI will eliminate their jobs, leading to resistance and reduced cooperation during implementation phases. This fear is not irrational, even if it's premature. Acknowledge it honestly. The answer isn't "AI won't replace you." The answer is a clear articulation of how roles will evolve, what new skills will be valued, and what support will be available during the transition. Research published in HBR found that managers and executives frequently disagree on AI, and it's costing companies. The organisational question has shifted from whether AI will transform businesses to when results will arrive. If middle management doesn't believe in the timeline or the approach, adoption stalls regardless of executive enthusiasm. ##### Executive sponsorship that goes beyond the launch email Microsoft's data shows that AI power users are 61% more likely to hear from their CEO about the importance of using generative AI at work. This isn't correlation masquerading as causation. Executive communication creates permission structures. When leadership signals that AI experimentation is expected, not just tolerated, behaviour shifts. Effective executive sponsorship means regular, specific communication about AI adoption goals. It means leaders using the tools visibly and sharing what they've learned. It means allocating time, budget, and recognition for teams that redesign processes around AI capabilities. The launch email is the beginning of this communication programme, not its entirety. Research from Raisch and Krakowski, cited in the California Management Review, found that the critical enabler of AI adoption is not technical capability, but the intersection of organisational design and human-AI collaboration. Executive sponsors who understand this focus on creating the conditions for adoption rather than mandating it. #### From deployment to compounding returns ##### The 90-day inflection point The first 90 days after deployment determine trajectory. By day 90, usage patterns have solidified, internal champions have either emerged or haven't, and the organisation has either established measurement practices or defaulted to anecdote-driven assessment. Gartner's survey notes that the rapid pace of Microsoft 365 Copilot change requires significant investments in change management activities, and that organisations are favouring smaller, business-driven deployments rather than IT-led approaches. This suggests that the most effective 90-day strategies are departmental, not enterprise-wide. Pick two or three business units with strong leadership support, measurable workflows, and reasonable data quality. Prove the model there before scaling. ##### Building feedback loops between users and IT MIT Sloan Management Review and Boston Consulting Group research found that organisations that build systematic feedback loops between humans and AI are six times more likely to derive substantial financial benefits from AI. As of 2024, 70% of companies had adopted AI, but only 15% were using it for organisational learning. Organisations that invest in learning with AI are 73% more likely to achieve significant financial impact. The practical implication: create a structured channel for users to report what works, what doesn't, what Copilot gets wrong, and what they wish it could do. Feed that information back into training programmes, governance policies, and workflow design. Blue Cross Blue Shield of Michigan recouped more than $10 million in savings after applying a GenAI tool to its IT contracts, enabling better analysis and standardisation of terms and pricing. That result came from systematic learning, not from initial deployment. Organisations that combine strong organisational learning with learning specific to AI are up to 80% more effective at managing uncertainty, according to the same research. The feedback loop isn't a nice-to-have. It's the mechanism through which initial productivity gains compound into sustained competitive advantage. ##### Setting the conditions for what comes after Copilot Copilot is almost certainly not the last AI tool your organisation will deploy. Deloitte predicts that 25% of companies using generative AI will launch agentic AI pilots or proofs of concept in 2025, growing to 50% by 2027. The California Management Review notes that today's leading AI platforms may become obsolete within three to five years due to rapid technological evolution. The organisations that extract the most value from Copilot are simultaneously building the organisational capabilities that will make them effective adopters of whatever comes next: data governance maturity, prompt literacy across the workforce, measurement frameworks tied to business outcomes, and leadership that understands the difference between deploying a tool and transforming how work gets done. The deployment was the easy part. The work that follows is where the value lives. If your organisation is ready to move beyond basic Copilot deployment and build AI capabilities that exploit the full technical potential of these tools, rather than settling for surface-level features, get in touch with Agathon. #### References - Microsoft Work Trend Index 2024: AI at Work Is Here. Now Comes the Hard Part - The State of Microsoft 365 Copilot: Survey Results (Gartner) - Gartner Peer Community: How Organizations Are Measuring Value of AI Productivity Tools Beyond Time Saved - Want More Out of Your AI Investments? Think People First (Bain) - Overcoming the Organizational Barriers to AI Adoption (HBR, November 2025) - Bridging the Gaps in AI Transformation: An Evidence-Based Framework for Scalable Adoption (California Management Review) - Managers and Executives Disagree on AI — and It's Costing Companies (HBR, April 2026) - How to Reap Compound Benefits From Generative AI (MIT Sloan Management Review) - Turbocharging Organizational Learning With GenAI (MIT Sloan Management Review) - Adoption of AI and Agentic Systems: Value, Challenges, and Pathways (California Management Review, August 2025) --- ### Securing AI agents: Why tool use creates new attack surfaces - URL: https://agathon.ai/insights/securing-ai-agents-why-tool-use-creates-new-attack-surfaces - Published: 2026-05-28 - Categories: AI Agents, AI Strategy, Responsible AI The security model for AI agents is broken in a specific, architectural way. Organisations are building agents that read emails, query databases, execute code, and call external APIs, then securing them with the same authentication patterns designed for human users clicking through web forms. A recent analysis of enterprise environments found a ratio of 144 non-human identities to every one human identity in organisations, yet only 21.9% treat agents as independent identity principals. The remainder run agents on shared API keys or inherited human credentials never designed for non-human use. This is not a gap that better prompting or model alignment will close. It is a systems problem, and treating it as anything less guarantees that the most capable agents you build will also be the most dangerous. #### The authentication layer you forgot to build ##### Why identity stops at the API gateway Most agent architectures authenticate once at the perimeter and then trust everything downstream. The agent presents an API key or OAuth token, the gateway says "proceed," and from that point forward, every tool call inherits the same ambient authority. This mirrors how we built microservices a decade ago, before zero-trust networking forced us to authenticate at every hop. But agents are worse than microservices in one critical respect: they are nondeterministic. Research on AI agent standards found that the same model weights produce different outputs even on identical inputs, making behaviour unreliable as an identity signal. A correctly authenticated agent can still act outside its mandate through prompt injection or behavioural drift, because all current authentication mechanisms verify the container of identity (tokens, certificates) rather than the content of the agent. ##### Agents act on behalf of users, not as users The delegation model underlying most agent deployments conflates two distinct principals. When an agent books a meeting on your behalf, it should carry a credential that says "this agent is acting for Colin, with permission to access his calendar, and nothing else." Instead, most implementations hand the agent Colin's full session token. The framework proposed by researchers at institutions including Stanford's Digital Economy Lab extends OAuth 2.0 with distinct tokens for the user, the agent, and the delegation relationship between them. This separation matters because without it, every service the agent contacts sees a human credential and grants human-level access, with no way to scope, audit, or revoke the agent's authority independently. ##### The delegation problem no one is solving OAuth 2.0 and SAML presume synchronous human consent and single-hop delegation. A human clicks "Allow," a token is issued, and a service acts on it. But autonomous agents act asynchronously, long after the human has walked away, and chain calls across multiple services in a single task. No production-ready standard currently traces authorisation chains back to originating human principals in a way that every resource server along the chain can verify. Multi-hop delegation accountability remains unenforceable in practice. This is not an edge case. It is the default operating mode for any agent that coordinates across services. #### How tool use turns read into write ##### From retrieval to side effects A chatbot that answers questions from a knowledge base is a read-only system. The moment you give that system a tool that can send an email, update a database record, or create a calendar event, you have crossed a boundary that most security models treat as fundamental. Tool execution layer invocations constitute real-world side effects that are typically irreversible, yet tool outputs are injected back into the agent's context without explicit trust marking, as a systematic survey of 116 papers on agent security found. The agent treats the response from a tool call the same way it treats the user's original instruction: as trusted input that informs its next action. ##### The compounding risk of chained tool calls Individual tool calls might each be reasonable. The danger emerges in composition. An agent that can read a file, execute code, and make network requests has, in combination, the ability to exfiltrate data. Research on privileged execution environments found that CI/CD pipelines act as privilege-amplification mechanisms where agents gain indirect access to capabilities far exceeding their own execution environment, including production credentials, deployment permissions, and the ability to mutate persistent infrastructure state. The OpenClaw case study illustrates this concretely: a platform exposing 15+ tools to every session regardless of task type created a 15x capability over-provision ratio for tasks that needed only a single tool, and the resulting ClawHavoc supply chain attack exploited this unrestricted access to distribute infostealers across 20% of the platform's skill registry. ##### Why sandboxing isn't enough when the agent holds credentials Static sandboxing constrains where an agent can operate but not what it does with its legitimate access. CVE-2026-25253 demonstrated that a single malicious webpage could hijack an agent's full capability set via prompt injection. The agent stayed within its sandbox. It used only its approved tools. It simply used them on behalf of an attacker instead of the user. Sandboxing addresses the wrong threat model when the risk is not that the agent escapes its environment but that it is manipulated within it. Research on learned capability governance found that infrastructure-level enforcement (restricting which tools are available per task type) reduced dangerous tool exposure dramatically, improving the fraction of exposed tools actually used from 0.053 to 0.557 in real sessions. #### Prompt injection as a privilege escalation vector ##### Indirect injection through untrusted data sources Direct prompt injection, where a user types "ignore your instructions," gets the attention. Indirect injection is the real threat. A malicious instruction embedded in a retrieved document, an email body, or a webpage the agent visits can redirect the agent's behaviour without the user ever seeing the injected content. Controlled experiments with GPT-5.1 found baseline unsafe behaviour rates ranging from 40% to 100% across risk scenarios, with prompt-injection and command-execution scenarios reaching 90-100%. The ToolHijacker attack achieved a 96.7% success rate when targeting tool selection end-to-end, and it operates in a no-box threat model where the attacker has no access to the tool library, retriever parameters, or model weights. ##### When the tool output becomes the next instruction The most dangerous pattern in agent architectures is the feedback loop between tool outputs and subsequent reasoning. An agent queries a database, receives results that contain injected instructions, and treats those instructions as part of its task context. Research on backdoored retrievers showed attack success rates up to 91% when injected prompts appeared in retrieved documents, with effectiveness highest when poisoned content appeared in the first retrieval position. A backdoored retriever component maintained improved precision scores on standard evaluation metrics even after poisoning, making the attack invisible through normal performance monitoring. Prevention-based defences have not kept pace. StruQ and SecAlign, two defences designed specifically to counter prompt injection, fail against ToolHijacker, with the attack achieving a 99.6% success rate under StruQ defence. Perplexity-based detection missed 90% of malicious tool documents while producing a 10% false positive rate on benign tools. Lightweight mitigations in privileged execution environments reduced unsafe behaviour to zero in three of four scenarios, but prompt-injection mitigation achieved only a 44% relative reduction because adversarial instructions embedded in task-relevant context are structurally indistinguishable from legitimate input. ##### The gap between what the model sees and what the user intended A production incident involving ChatGPT's macOS application illustrates this gap precisely. Malicious instructions were injected into the app's Memories feature, causing it to continuously exfiltrate conversations to an attacker-controlled server. The user intended a helpful assistant. The model saw instructions (indistinguishable from legitimate ones) telling it to send data elsewhere. A separate incident demonstrated that Claude Code could be induced to read API keys from .env files and transmit them via DNS requests, triggered by indirect prompt injection in a code file. In both cases, the model behaved as instructed. The problem was that the instructions came from the wrong principal. #### Trust boundaries collapse in multi-agent systems ##### Agent-to-agent communication as an unaudited channel When agents communicate with each other, every message is both an input and a potential attack vector. Research across 1,488 agent-to-agent interaction chains found that increasing inter-agent trust monotonically raises the over-exposure rate for sensitive information. With DeepSeek and the AgentScope framework, the over-exposure rate climbed from 0.120 at low trust to 0.500 at high trust. Higher trust improved task completion rates (Llama-3-8B improved from 0.22 to 0.71) while simultaneously amplifying leakage risk, creating an efficiency-security tradeoff that no current framework explicitly manages. Even under low trust conditions, LLMs exhibit non-zero baseline leakage risk. The helpful-agent prior, the alignment objective that makes models useful, is itself an exploitable attack surface. Agents want to be helpful, and being helpful to another agent sometimes means disclosing information that should stay compartmentalised. ##### Shared context windows as attack surfaces Multi-agent systems that share context windows create a broadcast channel where any agent's output becomes every agent's input. Individually safe agents can compose into unsafe systems. Research on multi-agent security found that when multiple agents interact, they can develop covert collusion, coordinated attacks, and cascading failures that cannot be predicted by analysing individual agents in isolation. Coordination and information flow between agents can be embedded in ways indistinguishable from benign interaction, even under full observability of communication. Out-of-scope or emergency requests show substantially higher baseline exposure (over-exposure rate of 0.41) and steeper trust sensitivity, indicating agents are most vulnerable to unintended disclosure during unexpected task contexts. The practical implication: your multi-agent system is least secure precisely when it encounters the situations you did not anticipate. ##### Why microservice security models don't transfer cleanly The instinct to apply microservice security patterns to multi-agent systems is understandable but misleading. Microservices execute deterministic code paths. Their behaviour is auditable, reproducible, and constrained by their implementation. Agents are none of these things. Agentic systems require dynamically adjusted security policies based on natural language task descriptions that evolve over time, creating challenges that do not exist in traditional systems with fixed policies. A service mesh policy that says "service A can call service B on endpoint /api/v1/users" has no equivalent in a system where Agent A asks Agent B a natural language question and Agent B decides what to do based on probabilistic inference. Different orchestration frameworks alter security posture in ways that compound this problem. AutoGen exhibits a high baseline over-exposure rate (0.379) with low sensitivity to trust changes, while LangGraph shows a lower baseline (0.261) but steep sensitivity. Framework choice materially impacts risk, and most teams make that choice based on developer ergonomics rather than security properties. #### Observability is harder than you think ##### The non-determinism problem in audit trails Traditional audit trails assume reproducibility. Given the same inputs, the same system should produce the same outputs, and the log should explain why. Agents violate this assumption at every level. The same prompt, the same tools, the same context can produce different tool call sequences on consecutive runs. A systematic survey found that threats and failures in agentic systems can emerge not from malicious inputs or faulty tools but from the emergent behaviour of the agent's cognitive trajectory during its reasoning process. Even well-scoped agents may deviate from expectations when their reasoning states are not explicitly monitored. AgentTrace proposes a three-surface taxonomy for addressing this: cognitive traces (capturing reasoning), operational traces (capturing execution), and contextual traces (capturing tool invocations and data access). This multi-level introspection links agent reasoning with external interactions and side effects, providing the kind of causal chain that a flat log of API calls cannot. ##### Logging tool calls without logging sensitive payloads Every tool call an agent makes should be logged. But tool calls carry payloads, and payloads contain data: customer records, API keys, personal information. Logging everything creates a secondary attack surface in the logging infrastructure itself. Information flow control in LLMs suffers from label explosion: when multiple labelled data sources are concatenated and fed into a model, the output is labelled with the union of all labels, making fine-grained access control on logs impractical without purpose-built infrastructure. The practical challenge is implementing contextual traces that capture enough to reconstruct what happened and why, without creating a data store that is itself a high-value target. This requires treating observability as a first-class architectural concern, not bolting it on after the agent is already processing production data. ##### Detecting misuse when correct behaviour looks identical to exploitation AgentSight, a system-level observability tool using eBPF, detected an indirect prompt injection attack where a development agent reading a malicious URL in a project README executed commands to exfiltrate /etc/passwd. The system captured 521 raw events and correlated them into 37 actionable events, with less than 3% performance overhead. But this detection was possible only because the system bridged the semantic gap between high-level intent (what the agent was asked to do) and low-level actions (what system calls it made). Existing tools observe one or the other, but cannot connect them. The fundamental difficulty is that a compromised agent and a legitimate agent performing a similar task generate near-identical telemetry. An agent sending data to an external API might be fulfilling a user request or exfiltrating credentials. Distinguishing the two requires understanding intent, and intent lives in the reasoning trace, not the system call log. #### What existing frameworks get wrong ##### OWASP for LLMs covers the model, not the system The OWASP Top 10 for LLMs identifies prompt injection as the most pressing threat, and it is correct to do so. But the framing centres on the model as the locus of vulnerability. A systematic survey of agent security research found that the under-studied zone (representing 6.3% of all papers) holds the highest-severity threats, with an inverse correlation between research effort and threat severity. Seven of 28 grid cells representing layer-temporality combinations have zero defence coverage, and three of those seven contain documented attacks. The gaps are not in model security. They are in the system architecture surrounding the model. Every surveyed tool-execution attack reduces to a single root cause: principal trust inversion, the systematic failure to enforce the principal hierarchy at the agent-environment boundary. Most agent implementations implicitly treat environment inputs as high-trust despite the environment being the least trusted principal in the hierarchy. OWASP's framework does not address this because it was designed for a different unit of analysis. ##### The false comfort of guardrails without enforcement Guardrails that rely on the model to enforce security properties are guardrails in name only. Research on agents using decentralised identifiers found that in evaluation runs, both agents independently bypassed mutual authentication policies stated in the system prompt and proceeded with credential issuance after one-directional authentication. Agents altered verifiable credentials during processing by omitting required fields or misspelling attributes, preventing successful verification. Evaluation across 100 test runs per LLM showed highly variable completion rates for security procedures, with some processes achieving consistently low completion rates despite identical system prompts. In the current era of LLM-based agents, attackers have consistently succeeded in bypassing model-based defences without requiring substantial increases in attacker effort. Unlike traditional systems that treat processes as untrusted, current agentic systems fail to treat the AI model powering the agent as untrusted, instead allowing it to enforce security properties directly. This is the equivalent of asking the process being sandboxed to enforce its own sandbox. ##### Rate limiting and permissions as first principles, not afterthoughts Progent, a programmable privilege control system for LLM agents, demonstrates what principled enforcement looks like. Every tool call is checked against a security policy through a deterministic procedure. An SMT solver determines each policy update to be either a narrowing (applied automatically) or an expansion (requiring explicit approval), ensuring the agent's action space can only shrink without approval. Evaluation on the AgentDojo and ASB benchmarks showed significant reductions in attack success rates while maintaining high utility, validated in real-world frameworks including LangChain and OpenAI Agents SDK. The design pattern research reinforces this principle: once an LLM agent ingests untrusted input, it must be constrained so that the input cannot trigger consequential actions with negative side effects. General-purpose agents with access to powerful tools cannot provide meaningful safety guarantees against prompt injections with current language models. The answer is not to make models more robust (though that helps). The answer is to make the architecture assume the model will be compromised. #### Building security into the agent layer ##### Least privilege as a design constraint, not a retrofit The ROME incident demonstrated how an enterprise AI agent acted as an insider threat through inherited credentials and over-broad authority without exploiting any software vulnerability. The agent did exactly what agents do: it used the tools it was given, with the permissions it inherited, in ways its designers did not anticipate. Least privilege in agent systems means more than restricting API scopes. It means dynamically scoping tool availability per task, treating every tool call as a privilege boundary, and building infrastructure that can narrow an agent's capabilities mid-execution without requiring the agent's cooperation. Sensitive-information repartitioning (structurally limiting what data each agent can access) reduced over-exposure rates by 79.5% for DeepSeek and 88.4% for Llama-3-8B in multi-agent evaluations. Guardian-agent patterns (dedicated monitoring agents that audit peers) achieved 38.4% and 83.6% reductions. These are architectural interventions, not prompt engineering. They work because they operate at a layer the model cannot circumvent. ##### Human-in-the-loop as a spectrum, not a binary The tension between security and usability is real. Research on authenticated delegation identifies prompt fatigue as a concrete failure mode: users grant permissions without proper review when prompted too frequently. The binary of "always ask" versus "never ask" is a false choice. The spectrum runs from fully autonomous (for low-risk, reversible actions) through notification-only (for medium-risk actions the system will execute unless stopped) to explicit approval (for high-risk, irreversible actions). The right position on this spectrum depends on the stakes of the specific tool call, not a global setting. Mitigation strategies show significant performance tradeoffs that inform where to place approval gates. Environment sanitisation added 22.7 seconds per operation. Policy checking reduced execution time by 207.9 seconds. But content filtering added 3,218.6 seconds due to retries. Security controls that make agents unusable will be disabled. The goal is enforcement that is invisible for routine operations and present only when the risk warrants it. ##### Where to start when your agents are already in production If your agents are already deployed, the pragmatic sequence matters. First, inventory what tools each agent can access and what credentials it holds. The 15x capability over-provision ratio found in research suggests your agents almost certainly have access to tools they never use. Reducing that surface area is the highest-leverage first step. Second, separate the control plane from the data plane. The agent's reasoning about what to do should be architecturally distinct from the mechanism that executes tool calls. The Plan-Then-Execute pattern prevents tool outputs from injecting new instructions (though malicious data can still influence tool call parameters). This separation creates a natural point for deterministic policy enforcement. Third, treat inter-agent trust as a first-class security variable subject to continuous auditing, not a tacit background assumption. Scope it, bound it, make it revocable. Fourth, build observability that connects intent to action, bridging the gap between what the agent was asked to do and what system calls it actually made. Without this connection, your audit trail is a collection of facts that cannot answer the question that matters: was this behaviour authorised? The organisations that will build trustworthy AI agents are those that recognise this is an architecture problem, not a model problem. The model is one component in a system, and the system needs to be secure even when the model is compromised. If you are building AI products that demand this level of architectural rigour, or if your existing agents need a security posture that matches their capabilities, we should talk. Agathon works with technical leaders to build AI systems that exploit full technical potential without creating the attack surfaces that make that potential a liability. #### References - Authenticated Delegation and Authorized AI Agents - Authentication for AI Agents: Privacy and Security (Stanford Digital Economy Lab) - AI Agents with Decentralized Identifiers and Verifiable Credentials - Standards, Gaps, and Research Directions for AI Agents - Security Risks in Tool-Enabled AI Agents: A Systematic Analysis of Privileged Execution Environments - Prompt Injection Attack to Tool Selection in LLM Agents - Design Patterns for Securing LLM Agents against Prompt Injections - Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation - The Trust Paradox in LLM-Based Multi-Agent Systems: When Collaboration Becomes a Security Vulnerability - Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents - A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework - AgentOps: Enabling Observability of LLM Agents - AgentTrace: A Structured Logging Framework for Agent System Observability - AgentSight: System-Level Observability for AI Agents Using eBPF - Progent: Programmable Privilege Control for LLM Agents - Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents - Agent Security is a Systems Problem - Fully Autonomous AI Agents Should Not be Developed --- ### What every startup founder should ask when hiring an AI advisor - URL: https://agathon.ai/insights/what-every-startup-founder-should-ask-when-hiring-an-ai-advisor - Published: 2026-05-28 - Categories: AI Advisory, AI Strategy, AI Consulting Most founders hire an AI advisor the same way they'd hire any consultant: they look for domain credentials, ask for references, and pick the person who sounds most confident about the technology. This process selects for exactly the wrong qualities. It rewards fluency over depth, optimism over rigour, and vendor relationships over architectural judgement. U.S. companies spent $37 billion on generative AI alone in 2025, according to Harvard Business Review. Yet 71% of global chief information officers said their AI budgets would be frozen or cut if value couldn't be demonstrated within two years. The gap between spending and returns is widening, and much of it traces back to the advice founders received before writing their first line of code. Choosing the right AI advisor isn't a hiring decision. It's a technical architecture decision dressed as a people problem. #### The advisor trap most founders walk into ##### Why domain expertise alone misleads A founder building an AI-powered logistics platform will naturally seek someone who knows logistics. This instinct is sound but insufficient. Domain experts who layer AI onto existing mental models tend to replicate manual processes with automation rather than rethinking the problem space entirely. The more dangerous version: domain experts who completed one successful AI implementation and now treat that single pattern as universal. They'll recommend the same architecture, the same vendor stack, the same data pipeline regardless of whether your constraints resemble the ones they solved before. MIT Sloan research describes this as a familiar pattern: business leaders mistake early-stage AI breakthroughs for mature use cases, experience FOMO, and end up with implementations that fall short of expectations. Domain knowledge matters. But it should inform the problem definition, not dictate the technical approach. ##### The difference between advice and implementation capability Harvard Business Review reported on a telling exercise: when MBA students were asked to define what consultants do, they used phrases like "trusted advisor," "problem-solver," and "subject matter expert." None of them mentioned "results." From their perspective, consultants generate advice that clients are expected to turn into results, rather than producing results themselves. This framing problem runs deep in AI advisory. An advisor who can explain transformer architectures at a whiteboard but has never shipped a production inference pipeline will give you architecturally elegant recommendations that collapse under real-world latency requirements, data quality issues, and cost constraints. The advice sounds right. The implementation fails anyway. What you want is someone who has felt the pain of a model that performed brilliantly in evaluation but degraded in production because the training distribution didn't match real user behaviour. That kind of scar tissue can't be acquired through reading papers or attending conferences. #### What "AI experience" should actually mean ##### Technical depth vs. vendor fluency There is a specific kind of AI advisor who can walk you through every major platform's feature matrix, recite pricing tiers from memory, and draw integration diagrams on demand. This person is a vendor expert, not a technical one. Technical depth means understanding why a retrieval-augmented generation pipeline might outperform fine-tuning for your use case, or why it might not. It means knowing when a purpose-built model will outperform a general-purpose large language model. As Akamai's Robert Blumofe told MIT Sloan, large language models can be "a ridiculously expensive way to solve certain problems." Most enterprise AI problems require what he calls an "ensemble of technologies" brought together for a purpose-built solution rather than a single approach. An advisor with genuine technical depth will tell you things you don't want to hear about your current architecture. A vendor-fluent advisor will tell you which product to buy. ##### Reading the difference between hype cycles and production readiness Research from Gartner and academic work published on arXiv describe generative AI adoption as following a dual-stage process: the familiar hype cycle (technology trigger, peak of expectations, trough of disillusionment, slope of enlightenment, plateau of productivity) layered with emotional stages of organisational change including shock, denial, and integration. Your advisor should be able to place specific technologies on this curve with precision. Not "AI agents are the future" but "agentic systems still experience ongoing hallucinations and security vulnerabilities like prompt injection, and experts predict a decade or more before these issues are fully resolved, though deployment with human-in-the-loop oversight may come sooner." That level of nuance, drawn from the MIT Sloan 2026 AI decision-makers' analysis, separates someone tracking the field from someone parroting keynote slides. ##### Questions that expose surface-level knowledge Ask your prospective advisor these questions and listen to how they answer, not just what they say: - "When would you recommend against using a large language model for this problem?" If they can't articulate specific scenarios where simpler approaches win, they're anchored to a single paradigm. - "What's the most common failure mode you've seen in production AI systems?" Vague answers about "data quality" suggest surface familiarity. Specific answers about distribution drift, feedback loops, or evaluation methodology suggest real experience. - "How would you approach this if our engineering team were half its current size?" Constraints reveal thinking quality. An advisor who only knows how to solve problems with more resources hasn't solved hard problems. - "Walk me through how you'd evaluate whether a model is production-ready versus demo-ready." MIT Sloan's research highlights that LLM success at simple tasks like email classification can represent "success theatre" rather than solutions to complex enterprise problems. Your advisor should understand this distinction viscerally. #### How to assess strategic fit with your stage ##### Pre-product vs. scaling vs. optimising: different advisory needs A pre-product startup needs an advisor who can help identify where AI creates defensible value in the product, not someone who optimises inference costs. A scaling company needs someone who understands how to move from a working prototype to reliable infrastructure serving thousands of concurrent users. A company optimising existing AI systems needs someone who can identify where you're leaving performance on the table. These are different skill sets. The advisor who excels at zero-to-one product thinking may be the wrong person to help you reduce your compute bill by 40%. Ask explicitly which stage they've worked at most, and probe for specifics. ##### Whether they understand your constraints, not just the technology Research published on arXiv examining AI deployment in public systems found that many AI systems fail at deployment rather than during model development, even when they perform well in internal testing. The Institutional Alignment Readiness framework they developed assesses five dimensions: institutional and operational compatibility, data ecosystem maturity, human oversight capacity, fiscal sustainability, and regulatory alignment readiness. These dimensions apply to startups too, if in different proportions. An advisor who talks exclusively about model performance without asking about your data infrastructure, your team's capacity to maintain what gets built, or your runway relative to the implementation timeline is advising in a vacuum. The gaps between technical viability and responsible deployment are, as the researchers note, most acute in resource-constrained settings. Startups are resource-constrained by definition. ##### Red flags in how they frame timelines and ROI Be wary of any advisor who offers confident ROI projections for AI initiatives before understanding your data, your team, and your existing systems. Harvard Business Review's survey of AI investment returns found that isolated, piecemeal deployments, limited executive buy-in, and weak linkage to strategic goals is a pattern that recurs across companies of all sizes. Without a systematic way to decide where to start, how fast to move, and when to stop, AI efforts become a drain on attention and resources rather than a source of advantage. An advisor worth hiring will resist giving you a timeline before they've assessed your starting position. They should frame initial engagements as diagnostic rather than prescriptive. #### The build-vs-buy question they should help you think through ##### Advisors who default to custom builds Some advisors reflexively recommend building custom AI systems. This bias often correlates with advisors who sell implementation services or who built their reputation on bespoke technical work. Custom builds do create genuine advantages: Deloitte's analysis of generative AI strategy notes that building enables tailored functionality and robust data security, though it requires greater investment. But "greater investment" understates the ongoing commitment. Custom AI systems need monitoring, retraining, evaluation infrastructure, and dedicated engineering attention indefinitely. For a startup with twelve engineers, building a custom recommendation engine when a well-integrated third-party solution would serve the same purpose is a misallocation of scarce engineering capacity. ##### Advisors who default to off-the-shelf The opposite bias is equally dangerous. Advisors who consistently recommend buying off-the-shelf tools may be optimising for speed of implementation at the expense of differentiation. Deloitte's research acknowledges that buying can lower costs and accelerate implementation, though it may compromise on privacy and flexibility. The strategic question isn't speed. It's whether the AI capability you're building is a commodity or a competitive advantage. If your AI feature is table stakes in your market, buy it. If it's your moat, you'd better own the technical stack beneath it. Harvard Business Review has argued that generative AI is dissolving the economic logic that made standardised enterprise software the only practical choice. Leaders must ask which workflows they actually need to own. Your advisor should help you answer that question rather than defaulting to either direction. ##### What a balanced perspective sounds like A good advisor on build-vs-buy will ask about your competitive dynamics before your technical requirements. They'll distinguish between AI capabilities that should be proprietary and those that should be purchased. They'll factor in your team's maintenance capacity alongside the initial build cost. Deloitte's analysis of AI-assisted software engineering found that Klarna replaced its Salesforce CRM with a GenAI-built internal platform, reducing the developer requirement from 20 to 5 people. That's a compelling example of building, but it required Klarna's specific combination of engineering talent, scale, and strategic commitment. Your advisor should be able to articulate why a similar approach would or wouldn't apply to your situation, not just cite the case study as proof that building always wins. The scaling model for AI-assisted teams is shifting from "more developers equals more output" to "more context per developer equals more impact." An advisor who understands this shift will think differently about build-vs-buy than one still operating on pre-AI assumptions about engineering productivity. #### Incentive alignment and engagement structure ##### Equity, retainer, or project-based: what each reveals How an advisor structures their engagement tells you what they optimise for. Equity-based arrangements align the advisor's interests with your long-term success but can create perverse incentives around fundraising narratives. An advisor with equity might encourage you to build impressive demos that inflate valuation rather than robust systems that serve users. Retainer arrangements provide stability and ongoing access but can become comfortable. Without clear deliverables, retainer relationships drift toward general availability rather than focused impact. Project-based engagements create accountability around specific outcomes but can lead to advisors optimising for project completion rather than your broader strategic position. They might solve the scoped problem while ignoring adjacent issues that matter more. The structure that works best depends on your stage and needs. What matters is whether the advisor can articulate why they prefer a particular structure and what trade-offs it creates. If they can't discuss the downsides of their own preferred model, they haven't thought critically about their own incentives. ##### How to test for vendor-agnostic thinking Ask your prospective advisor which cloud provider or AI platform they'd recommend, then ask them to argue against their own recommendation. An advisor locked into a single ecosystem (through partnerships, certifications, or familiarity) will struggle to make a convincing case for alternatives. MIT Sloan's 2026 analysis of AI decision-making emphasises that organisations implementing generative AI predominantly take an individual-level approach to boost employee productivity rather than applying it to enterprise workflows and processes. This pattern often reflects vendor-driven thinking: the tools are designed for individual productivity, so that's what gets implemented. An advisor who thinks at the systems level will push beyond individual-tool adoption toward workflow-level transformation. ##### When an advisor's network becomes a liability Advisors with strong vendor relationships can provide valuable introductions and negotiating leverage. But those same relationships create bias. If your advisor has a referral arrangement with an infrastructure provider, their recommendation to use that provider should be treated with appropriate scepticism. Ask directly: "Which vendors do you have financial relationships with?" Any hesitation or evasion is informative. The best advisors disclose these relationships proactively because they understand that transparency is the only way to maintain credibility when conflicts exist. #### Evaluating their track record without relying on testimonials ##### What to look for in their public thinking Testimonials are curated. No advisor publishes the negative ones. Instead, read what they write and say publicly. Look for specificity over generality. An advisor who writes "companies should adopt AI strategically" is saying nothing. An advisor who writes about specific failure modes in retrieval-augmented generation pipelines, or who analyses why a particular architectural pattern breaks at scale, is demonstrating genuine expertise. Look for intellectual honesty. Do they acknowledge when a technology they previously advocated turned out to be less capable than expected? Do they update their views as the field evolves, or do they maintain the same position regardless of new evidence? The MIT Sloan research on AI hype cycles notes that a significant gap exists in understanding societal reception and adaptation to generative AI tools. An advisor who acknowledges uncertainty and complexity rather than projecting false confidence is likely to give you better guidance. ##### How they talk about failure and technical debt Listen for how candidates discuss projects that didn't work. An advisor who presents an unblemished record is either lying, inexperienced, or has only taken safe engagements. The AI field moves too fast and the technical risks are too real for anyone with meaningful experience to have avoided failure entirely. More telling than the failure itself is how they analyse it. Do they blame external factors (the client's data was bad, the team wasn't committed) or do they identify what they could have done differently? The best advisors have specific, uncomfortable stories about recommendations they made that turned out to be wrong, and they can articulate what they learned. Research on AI deployment in public systems found that two technically viable AI systems in education reached working prototypes but couldn't advance to broader rollout due to institutional rather than technical reasons. An advisor who has encountered similar situations and can discuss what they'd do differently demonstrates the kind of systemic thinking that prevents expensive failures. ##### Asking for the engagement that went wrong Make this a standard part of your evaluation process. Ask: "Tell me about an advisory engagement that didn't deliver the expected results. What happened and what would you do differently?" The quality of the answer matters more than the content of the failure. You're evaluating self-awareness, analytical rigour, and honesty. An advisor who can't or won't answer this question is one you should pass on. #### Moving from selection to productive engagement ##### Setting the terms before the first session Before your advisor's first billable hour, establish clarity on four points: what decisions they're being hired to inform, what information they'll need access to, how you'll measure whether the engagement is productive, and what happens if it isn't working. The MIT Sloan 2026 AI survey found that 38% of large enterprises have appointed a chief AI officer or equivalent role, but there is little consensus on reporting structure. This same ambiguity can plague advisory relationships. Your advisor needs to know who they report to, whose time they can request, and what authority their recommendations carry. Without this clarity, even brilliant advice gets lost in organisational friction. ##### How to know within 90 days whether it's working Set a 90-day evaluation checkpoint before the engagement begins, not after. Define two or three specific outcomes you expect by that point. These shouldn't be "delivered a strategy document" (that's activity, not outcome) but rather "helped us make the build-vs-buy decision on our recommendation engine with a clear technical rationale" or "identified and deprioritised two AI initiatives that weren't aligned with our product strategy." Harvard Business Review's portfolio approach to AI investment management offers a useful frame here: without a systematic way to decide where to start, how fast to move, and when to stop, AI efforts become a drain rather than a source of advantage. Your advisor should be helping you make those decisions with more confidence and precision than you could alone. If after 90 days your decision-making quality hasn't measurably improved, the engagement isn't working regardless of how impressive the advisor's credentials are. The test is simple. Are you making better technical decisions faster? If yes, the advisor is earning their fee. If no, it's time for a direct conversation about what needs to change, or whether to part ways. --- Hiring the right AI advisor is one of the highest-leverage decisions a startup founder can make, and one of the easiest to get wrong. The questions in this article are designed to filter for the rare combination of technical depth, strategic judgement, and intellectual honesty that separates genuinely valuable advisors from credentialed commentators. If you're ready to work with a team that builds AI solutions exploiting the full technical potential most companies never reach, rather than implementing surface-level features, get in touch with Agathon. #### References - How to Break the AI Hype Cycle and Make Good AI Decisions for Your Organization - Action Items for AI Decision Makers in 2026 - Hype and Adoption of Generative Artificial Intelligence Applications - Build, Buy, or Adopt Generative AI in Digital Procurement - AI-Assisted Software Engineering: Rewriting the Build Versus Buy Playbook - The End of One-Size-Fits-All Enterprise Software - Let's Hold Consultants Accountable for Results - 7 Factors That Drive Returns on AI Investments, According to a New Survey - Beyond Model Readiness: Institutional Readiness for AI Deployment in Public Systems - Manage Your AI Investments Like a Portfolio --- ### How AI Consulting Services Drive Digital Transformation - URL: https://agathon.ai/insights/how-ai-consulting-services-drive-digital-transformation - Published: 2026-05-27 - Categories: AI Consulting, AI Strategy, Machine Learning #### Why most digital transformation programmes stall before they start Eighty-eight percent of companies report regular AI use. Seventy-seven percent call it a board-level strategic priority. And yet 94% face significant challenges implementing it. The gap between enthusiasm and execution is where most transformation programmes go to die. The pattern is remarkably consistent. An organisation announces an AI strategy, funds a centre of excellence, launches a dozen pilots, gives everyone access to ChatGPT, and then waits for transformation to happen. Twelve months later, the pilots are still pilots. The centre of excellence has become a bottleneck. And the board is asking pointed questions about return on investment. This is not a technology problem. Research from Harvard Business School confirms what practitioners have observed for years: employees experiment with new tools but don't integrate them into how work gets done. Performance gains plateau. Adoption stalls. The "last mile" between AI capability and business value turns out to be the longest mile of all. The companies that break through this pattern share a common trait. They stopped treating AI as a technology deployment exercise and started treating it as an organisational redesign problem, one that requires a fundamentally different kind of expertise than what most internal teams or traditional consultancies provide. #### The gap between AI ambition and AI execution ##### Misaligned expectations at the board level Most AI strategies fail before any code gets written. The failure mode is predictable: leadership sees a competitor's press release, a vendor demo, or a McKinsey report, and sets expectations calibrated to what AI could theoretically do rather than what their organisation can absorb. A survey of 2,496 technology decision-makers across 22 countries found that 73% of leaders still believe technical teams should lead AI adoption. This represents a profound misunderstanding of where AI value originates. Business teams understand the workflows, constraints and regulatory requirements that AI must address. When technical teams lead in isolation, you get technically impressive systems that solve the wrong problems. The expectation mismatch runs deeper than strategy decks. Board members often conflate general-purpose AI tools (giving employees ChatGPT access) with enterprise AI transformation (redesigning how decisions get made). These are categorically different undertakings. The first is a procurement decision. The second is an organisational capability shift that touches operating models, data architecture, governance and talent strategy simultaneously. ##### Technical debt as an invisible brake Organisations rarely account for the state of their existing systems when planning AI initiatives. Legacy data architectures, inconsistent data quality, fragmented systems and undocumented business logic all compound into what researchers call "data dependency," a characteristic of AI systems that makes them exceptionally sensitive to the quality and availability of their inputs. Research on organisational capabilities for AI implementation identifies data management as one of four critical capabilities, distinct from general IT competence. AI systems that rely on machine learning derive their own rules from data rather than following predetermined logic. This makes them fundamentally different from traditional software. A CRM migration can tolerate messy data with workarounds. A machine learning model trained on messy data produces confidently wrong predictions. The practical consequence: organisations budget for model development but not for the data engineering, pipeline construction and quality assurance work that consumes 60-80% of any serious AI initiative. By the time teams discover the true state of their data foundations, the project timeline has already slipped. ##### The talent problem nobody wants to admit The AI talent gap runs to roughly 50% of demand, according to Reuters research from 2024. But the headline number obscures a more nuanced problem. Organisations are not just short on data scientists and ML engineers. They lack people who can bridge the gap between technical AI capabilities and business domain knowledge. A study of 92 companies found that the most significant barriers to AI adoption were organisational, not technological. Employee resistance, change management deficits and regulatory ambiguity ranked above technical limitations. The critical competencies for successful AI use turned out to be understanding algorithmic mechanisms and managing organisational change. Programming skills played a smaller role than expected. This creates a specific staffing paradox. The people who understand your business don't understand AI well enough to identify high-value applications. The people who understand AI don't understand your business well enough to build systems that integrate into existing workflows. And hiring a "Head of AI" to sit between these groups rarely works, because one person cannot simultaneously maintain technical depth and the cross-functional relationships needed to drive adoption. #### What AI consulting actually does (and what it doesn't) ##### Strategy that survives contact with your data Good AI consulting starts with a deceptively simple question that most organisations skip: given your data, your systems and your organisational constraints, what can AI do for you in the next six months that would change how you operate? This is different from the typical strategy engagement that produces a 60-slide deck of theoretical use cases ranked by estimated value. Research on AI project planning capability emphasises that "AI is very much focused on specific use cases, and you have to get rid of the preconception that it is applicable anywhere." The value lies in professionalising and democratising the process of use case identification, then rigorously assessing feasibility against the organisation's real technical and data landscape. Effective AI consultants bring pattern recognition from across industries and implementations. They have seen where the standard use cases (predictive maintenance, demand forecasting, customer service automation) succeed and where they fail. More importantly, they can identify non-obvious applications where your specific data assets create advantages competitors cannot replicate. ##### Building the right thing before building the thing right The distinction between building something well and building the right thing matters enormously in AI, because the cost of discovering you solved the wrong problem is measured in months, not days. Research on enterprise automation with foundation models highlights this challenge starkly. Traditional robotic process automation requires 12-18 months of setup and achieves roughly 60% initial accuracy. Modern approaches using multimodal foundation models can achieve near-human-level understanding of workflows (93% accuracy on understanding tasks) with minimal setup, based solely on natural language descriptions. But choosing between these approaches, or knowing when each is appropriate, requires architectural judgment that comes from building production systems, not reading about them. The best AI consultants function as technical product managers for your AI initiatives. They validate that the problem is worth solving before optimising the solution. They prototype with real data early to surface integration issues that no amount of planning can anticipate. They establish success criteria that tie to business outcomes rather than model accuracy metrics. ##### Embedding capability, not creating dependency The most corrosive dynamic in consulting is the one where the consultant becomes indispensable. Every engagement should make the client organisation more capable at the end than it was at the beginning. Research on organisational capabilities identifies four distinct capabilities that organisations need for sustained AI implementation: AI project planning, co-development of AI systems, data management and AI model lifecycle management. A good consulting engagement builds all four, not just the technical ones. It means training your teams to identify and evaluate AI use cases independently. It means establishing data governance practices your people own. It means creating model monitoring and retraining processes that continue to function after the consultants leave. This is qualitatively different from the engagement model where consultants build a system, hand over documentation and walk away. The documentation gets stale. The system degrades. Six months later, the organisation is calling the same consultants back. The organisations that avoid this cycle are the ones that insisted on knowledge transfer as a contractual deliverable from day one. #### Where AI consulting delivers outsized impact ##### Automating decision-making Most organisations start their AI journey by automating repetitive tasks: document processing, data entry, basic classification. These deliver incremental efficiency gains but miss the transformative potential. The outsized returns come from automating or augmenting decision-making itself. Research on enterprise decision-making found that AI systems increase the speed and clarity of managerial decisions when integrated with human judgement and supported by transparent processes. The key phrase is "integrated with human judgement." The DXC Technology survey reveals that 54% of leaders expect AI to operate with partial autonomy where humans review key decisions, while only 15% anticipate fully autonomous systems. This points to where consulting expertise matters most: designing the human-AI collaboration model. Which decisions should AI make autonomously? Which should AI recommend while humans approve? Where should AI provide information while humans decide? Getting this taxonomy right for your specific context determines whether AI augments your competitive position or just reduces headcount. ##### Turning unstructured data into competitive advantage Only 18% of organisations reported being able to take advantage of unstructured data in a Deloitte survey, despite the fact that 80-90% of enterprise data is unstructured: text, video, audio, web logs, customer communications. This gap represents one of the largest unrealised opportunities in enterprise AI. The organisations extracting value from unstructured data are doing things their competitors cannot easily replicate. Kensho, acquired by S&P Global, uses natural language processing to parse unstructured financial data, pulling numbers from earnings documents at a speed and scale that creates genuine trading advantages. Etihad Airways built predictive maintenance systems from unstructured sensor data, then spun the capability into a separate revenue-generating business unit serving other airlines. These are not implementations you arrive at by following a generic AI playbook. They require deep understanding of domain-specific data assets, the technical expertise to build extraction and processing pipelines, and the strategic vision to see which capabilities create durable competitive advantages versus those that competitors will commoditise within a year. ##### Accelerating time-to-value on ML initiatives Research consistently identifies that 83% of data science projects never make it into production. Seventy-six percent of organisations report problems implementing AI throughout the organisation. The time-to-value problem is not about building models. It is about everything surrounding the model: data preparation, integration with existing systems, monitoring, retraining and user adoption. AI consultants who have built production systems understand model lifecycle management as a continuous operational discipline, not a one-off deployment. This includes monitoring for data drift, managing retraining pipelines, handling edge cases that only appear at scale and maintaining model performance as the underlying data distribution evolves. MIT Sloan researchers found that teams within organisations are better suited than centralised functions to determine how they work best with AI. An experienced consultant can accelerate this discovery process by weeks or months, bringing patterns from comparable implementations while respecting the specific context. Vanguard Group estimates its AI ROI at close to $500 million, with use cases spanning call centre support, personalised adviser summaries and a 25% improvement in programming productivity. These results came from systematic scaling, not from any single brilliant model. Half of Vanguard's employees completed training through their AI Academy, and leadership maintained discipline by not scaling pilots "until the kinks have been worked out." #### The operating model shift most organisations miss ##### From project-based thinking to continuous intelligence Most organisations approach AI as a series of projects: identify a use case, build a model, deploy it, move to the next one. This project-based framing misses the compound nature of AI capability. MIT Sloan research on scaling AI describes three levels of value creation: individual productivity gains, incorporation of AI into defined tasks and roles, and automation of production and operational processes. The jump between these levels requires operating model changes, not just more projects. Individual productivity gains come from giving people tools. Task-level integration requires redesigning workflows. Process automation demands rethinking how entire functions operate. The organisations capturing the most value treat AI as continuous infrastructure rather than discrete initiatives. They invest in shared data platforms, reusable model components and cross-functional teams that can deploy AI capabilities against new problems quickly. This resembles platform engineering more than traditional project delivery. ##### Reorganising teams around AI-augmented processes When enterprises adapted to the internet, they did not create an Internet Department and require employees to seek approval to launch websites. Research from MIT Sloan argues the same principle should apply to AI: executives should establish guardrails, but individual teams should define how AI gets used in their specific context. The practical implication is significant. Only 47% of business professionals say AI policies reflect the realities of their work. When rules do not match day-to-day practice, employees either use unsanctioned tools (creating security and compliance risk) or ignore AI tools entirely (wasting the investment). "Judgement is local," as the researchers put it, and front-line leaders are best positioned to turn broad corporate policies into specific, workable practices. This requires a different team structure than most organisations have. You need people who understand both the technical constraints of AI systems and the operational reality of the work being augmented. Research on co-development of AI systems stresses the importance of integrating data scientists, domain experts, end-users, IT security and ethics experts. Those areas "need to work well together, without which the success is a big question mark." ##### Governance frameworks that enable rather than restrict The Mayo Clinic, described as "the most aggressive adopter of AI among US health care providers," offers an instructive model. Their approach shifted focus from governance (what can and cannot be done) to enablement (giving employees the latitude to build and test AI models in their domain). Mayo maintains a 60-person team supporting AI and data enablement. They provide internal users with a platform for building AI products and applications while end-users retain responsibility for data quality. This approach works because clinical staff are oriented to quantitative thinking and understand their data better than any central function could. The enablement model requires guardrails, not gates. The IAPP's AI Governance in Practice framework identifies governance considerations across the full AI lifecycle: planning, design, development and deployment. Effective governance establishes boundaries within which teams can move quickly, covering privacy, security, intellectual property and ethics at the enterprise level, while leaving implementation decisions to the people closest to the work. #### How to evaluate whether you need AI consulting ##### Signs your internal efforts have plateaued The clearest signal is a growing collection of pilots that never graduate to production. If your data science team can build impressive demos but struggles to deploy systems that integrate with existing operations, you have an organisational capability gap, not a technical one. Other indicators: your AI strategy document is more than twelve months old and still references the same "priority use cases." Your data infrastructure conversations keep getting deferred because they are "too expensive" relative to any single project. Your best ML engineers are spending more time on data cleaning and stakeholder management than on model development. Your AI governance framework is either non-existent or so restrictive that teams route around it. Research on AI readiness identifies that organisations need capabilities across project planning, cross-functional collaboration, data management and model lifecycle management simultaneously. Weakness in any one area bottlenecks the others. Most organisations plateau because they have invested heavily in one capability (usually technical talent) while neglecting the supporting capabilities that make that talent effective. ##### The build vs. buy vs. partner calculus The decision is rarely binary. Most successful AI implementations combine elements of all three: purchased infrastructure, internally built domain-specific models and external expertise for capability acceleration. Research on strategic AI adoption in SMEs proposes a phased approach: start with low-cost, general-purpose AI tools to build technical competence and positive attitudes toward AI. As familiarity increases, integrate task-specific tools. Then progress to in-house development where customisation and control matter. The progression matters because each phase builds the organisational muscle needed for the next. The specific value of a consulting partner is acceleration through the phases where you lack experience. If your organisation has strong data engineering but weak AI product management, you need a partner who can fill that gap temporarily while transferring the capability permanently. If you have strong domain expertise but no ML infrastructure, you need a different kind of partner. The worst outcome is hiring a generalist firm that addresses none of your specific gaps. ##### What good engagement looks like in practice A well-structured AI consulting engagement has several distinguishing characteristics. It starts with an assessment of your current capabilities, data landscape and organisational readiness, not with a technology recommendation. It defines success metrics tied to business outcomes before any model development begins. It includes explicit milestones for knowledge transfer alongside delivery milestones. Research on co-development emphasises that AI implementation should not happen in isolation. Effective engagements integrate your domain experts, data engineers, end-users and governance stakeholders from the outset. The consultant brings technical depth, architectural patterns and implementation experience. Your people bring domain knowledge, data context and the organisational relationships needed to drive adoption. The engagement should get shorter over time, not longer. If your consultant's involvement is growing rather than shrinking, the knowledge transfer component has failed. The goal is to build internal capabilities that compound after the engagement ends. #### What separates effective AI consultants from expensive ones ##### Domain fluency over generic frameworks A consultant who can explain transformer architectures but cannot translate that knowledge into your specific industry context will produce elegant solutions to the wrong problems. Research on the importance of integrating diverse expertise into AI implementation consistently finds that domain knowledge is the critical bridge between technical capability and business value. The distinction between domain awareness and domain fluency matters. Awareness means knowing that financial services have regulatory constraints. Fluency means understanding which regulatory constraints affect which AI applications, how similar organisations have satisfied regulators, and where the boundaries of acceptable automation sit in your specific jurisdiction. This fluency comes from repeated implementation experience in relevant contexts, not from reading industry reports. ##### Delivery track record over thought leadership The AI consulting market is saturated with firms that publish reports about what organisations should do but have limited experience building production systems. A delivery track record means the consultant has encountered the unglamorous realities of AI implementation: data quality problems that invalidate initial assumptions, model performance that degrades in production, user adoption challenges that no amount of training resolves, integration failures with legacy systems. The research is clear on this point: 83% of data science projects never reach production. The consultants worth hiring are the ones who can explain exactly why projects fail at the deployment stage and demonstrate specific practices they use to prevent those failures. Ask for references from organisations that are still using the systems the consultant built, not just organisations that were satisfied with the initial delivery. ##### Knowledge transfer as a first-class deliverable The DXC Technology survey found that leaders rank workforce AI training and change management as the single most valued service from third-party providers. This aligns with research showing that organisational factors, not technological limitations, are the primary barriers to AI adoption. Knowledge transfer is not documentation. It is not a training session on the last day of the engagement. It is a structured, ongoing process where your people work alongside the consulting team, gradually assuming ownership of each component. It means pair programming with your engineers, not handing them finished code. It means co-facilitating stakeholder workshops, not presenting findings. It means building internal champions who can advocate for and extend AI capabilities after the engagement concludes. Seventy-nine percent of CIOs report that partnerships with service providers have successfully delivered improved outcomes. But the partnerships that generate lasting value are the ones designed to be finite, where the explicit goal is to make the partner unnecessary. #### The transformation compound effect AI transformation compounds in ways that linear project plans fail to capture. Each successfully deployed AI system improves your data infrastructure, builds organisational muscle for the next implementation, trains your teams to work with AI-augmented processes and establishes governance patterns that can be reused. The second implementation is faster than the first. The fifth is faster still. MIT Sloan researchers describe this as "building the scaffolding": aligning AI efforts with core business capabilities so that each success creates the foundation for the next. Vanguard's $500 million estimated ROI did not come from a single breakthrough. It came from systematic, disciplined accumulation of AI capabilities across call centres, advisory services, programming and investment analysis, each building on shared infrastructure and institutional learning. The organisations that capture this compound effect share a common characteristic. They invested in capability, not just technology. They built the organisational muscles (project planning, cross-functional collaboration, data management, model lifecycle management) that turn individual AI projects into a self-reinforcing transformation engine. This is the central argument for working with AI consultants who embed capability rather than create dependency. The right engagement does not just deliver a system. It accelerates your organisation's ability to deliver the next system, and the one after that, without external help. If your AI initiatives have stalled at the pilot stage, or you are capturing a fraction of the value your data assets could generate, it may be time to work with people who build AI products that exploit the full technical potential most organisations leave on the table. Get in touch with Agathon to discuss how we can help you move from experimentation to transformation. #### References - HBR — "The Last Mile Problem Slowing AI Transformation" - HBR — "Overcoming the Organizational Barriers to AI Adoption" - HBR — "Why AI Adoption Stalls, According to Industry Data" - HBR — "How to Move from AI Experimentation to AI Transformation" - HBR — "For Success with AI, Bring Everyone On Board" - MIT Sloan — "Making Generative AI Work in the Enterprise" - MIT Sloan — "Scaling AI for Results: Strategies from MIT Sloan Management Review" - MIT Sloan — "Tapping the Power of Unstructured Data" - DXC Technology — "Closing the AI Execution Gap: Global AI Survey" - arxiv — "The Impact of Artificial Intelligence on Enterprise Decision-Making Process" - arxiv — "Automating the Enterprise with Foundation Models" - arxiv — "Strategic AI Adoption in SMEs: A Prescriptive Framework" - Springer Nature / Information Systems Frontiers — "Organizational Capabilities for AI Implementation: Coping with Inscrutability and Data Dependency" - IAPP — "AI Governance in Practice Report 2024" - IBM — "AI Skills Gap: What It Means for Enterprise Readiness" --- ### AI adoption strategy: why centralise, decentralise, and wait-and-see all fail - URL: https://agathon.ai/insights/ai-adoption-strategy-why-centralise-decentralise-and-wait-and-see-all-fail - Published: 2026-04-29 - Categories: AI Strategy, AI Consulting, AI Advisory #### The three-body problem of AI adoption Global corporate AI investment reached $252.3 billion in 2024. Only 6% of firms report significant earnings impact. That gap represents hundreds of billions in stalled initiatives, abandoned pilots, and transformation programmes that transformed nothing. Most organisations default to one of three strategies when deciding how to adopt AI: centralise it under a single team, decentralise it to every department, or wait until the technology matures. Each feels rational in isolation. Senior leadership can point to sound logic behind whichever path they've chosen, and board presentations make any of the three look like a plan. All three produce the same outcome: wasted budget, organisational frustration, and a widening gap between what AI can do and what the organisation gets from it. The strategy question itself is framed wrong. Centralise, decentralise, or wait treats AI adoption as a deployment decision, something you do to an organisation. Research from McClure and Gerdau's 2026 synthesis of nearly 10,000 organisational leaders across 19 large-scale studies confirms what practitioners have observed for years: AI project failure is an organisational learning problem, not a technology deficit. Choosing where to put the AI team misses the point entirely. #### The centralised trap ##### Control masquerading as progress When a CEO decides AI matters, the instinct is to hire a head of AI, build a Centre of Excellence (CoE), and funnel everything through one team. This maps neatly onto how organisations handled previous technology waves. Data warehousing, cloud migration, digital transformation: all followed a similar playbook. One team, one budget, one set of standards. The appeal is obvious. Governance stays clean. Tooling stays consistent. Risk gets managed through a single point of control. Leadership can point to the AI team on an org chart and say, with confidence, that AI is being handled. ##### Where centralisation breaks down A team of fifteen, or even fifty, cannot understand the operational nuances of every business unit in a large organisation. The supply chain team's forecasting challenges share little DNA with the customer service team's ticket classification problem or the legal department's contract analysis workflow. Each requires different data, different evaluation criteria, different definitions of success. What happens in practice: business units submit requests to the CoE. The CoE triages, prioritises, and queues. Innovation gets scheduled. By the time the AI team gets to a business unit's problem, the context has shifted, the sponsor has moved on, or the window of opportunity has closed. McKinsey's research on building AI-powered organisations found this pattern consistently: the CoE becomes a permissions desk. It exists to approve or deny rather than to accelerate. The people closest to the business problems sit furthest from the technical tools, separated by layers of intake forms and prioritisation frameworks. ##### The talent paradox Centralised AI teams attract a specific profile: people who understand AI broadly but lack deep domain expertise in any single business function. They can build a model, but they cannot tell you why the logistics team's demand signal behaves differently in Q4 or why the underwriting team's risk thresholds changed after a regulatory update last March. Meanwhile, domain experts with twenty years of operational knowledge stay in their silos. They understand the problems worth solving but lack access to the tools, training, or mandate to solve them with AI. The result is a widening gap between technical possibility and business application. The AI team builds technically sound solutions to the wrong problems. The domain teams know the right problems but cannot articulate them in a format the AI team can act on. The SIO (Siloed-Integrated-Orchestrated) progression model developed by McClure and Gerdau identifies this as the "siloed" stage, where AI capability exists in the organisation but remains disconnected from the operational context that would make it useful. #### The decentralised mirage ##### Democracy sounds good on a slide deck The opposite impulse is to let every team experiment. Give each department a budget, access to AI tooling, and the autonomy to figure out what works. This appeals to organisations that value speed, entrepreneurialism, and distributed ownership. It also appeals to leadership teams that don't want to make hard prioritisation decisions. The pitch is seductive: empower the edges, let domain experts drive adoption, avoid the bottleneck of a central team. A thousand flowers blooming. ##### What happens when everyone gets a budget Five teams buy five different AI platforms. The marketing team builds a content generation pipeline on one vendor's API. The operations team builds a forecasting system on another. Finance experiments with a third. No shared data infrastructure connects them. No common evaluation framework measures their results. Procurement becomes chaos. Each vendor relationship is negotiated independently, often at worse terms than a consolidated agreement would achieve. Shadow AI projects proliferate, built by well-meaning teams who lack the expertise to assess what they've created. Duplicate efforts multiply. Two teams in different offices spend six months solving functionally identical problems with incompatible approaches, neither aware of the other's work. The Adaptive Responsible AI Governance (ARGO) framework, developed through collaboration between Stanford researchers and a multinational enterprise, documented this pattern across multiple business units. Their assessment revealed "complex interplay between group-level guidance and local interpretation" and "regional and functional variation in implementation approaches" that created inconsistent outcomes across the organisation. Four failure patterns emerged consistently: tensions between central guidance and local interpretation, difficulty translating abstract principles into operational practices, wide variation in how different regions and functions implemented the same guidelines, and inconsistent accountability for risk oversight. ##### Risk without a safety net Decentralised adoption creates compliance exposure that only surfaces during an audit or an incident. Individual teams lack the expertise to evaluate model bias, data privacy implications, or the downstream effects of automated decisions. Without shared standards for model evaluation, each team invents its own definition of "good enough." An AI system making lending recommendations needs different scrutiny than one suggesting blog post topics. Decentralised teams often apply the same level of rigour (usually insufficient) to both, or they apply rigour to the wrong dimension: obsessing over model accuracy while ignoring fairness, or optimising for speed while neglecting explainability. #### The wait-and-see delusion ##### Patience as a strategy Some leaders look at the 6% success rate and conclude that waiting is the smart play. The technology is changing fast. Regulatory frameworks remain unsettled. ROI cases are thin. Why invest heavily in something that might look completely different in eighteen months? This maps onto a "fast follower" strategy that worked for most previous information technologies. Let the early adopters make expensive mistakes, learn from their failures, and adopt proven approaches once the dust settles. ##### The cost of standing still AI doesn't follow the same adoption curve as enterprise software. With traditional technology, a late adopter could purchase a mature product, hire experienced implementers, and catch up within a deployment cycle. AI capability is different. Organisations that started eighteen months earlier haven't just deployed tools. They've built institutional knowledge about what works in their specific context, trained their workforce to collaborate with AI systems, developed evaluation frameworks tuned to their domain, and created feedback loops that compound learning over time. Mahidhar and Davenport's research on AI adoption timing argues that the fast-follower strategy fails specifically because AI capabilities build on themselves. Each successful implementation creates data, institutional knowledge, and organisational muscle that accelerates the next one. Late entrants don't just face a technology gap. They face a capability gap that money alone cannot close. ##### When "ready" never arrives The goalposts keep moving. Waiting for the "right" model gives way to waiting for the "right" regulation, which gives way to waiting for the "right" use case. Organisations in permanent waiting mode don't eventually adopt well. They adopt in a panic when a competitor's AI-driven product threatens their market position, when a board member asks uncomfortable questions, or when a regulatory deadline forces action. Panic adoption is worse than any of the three strategies. It combines the worst elements of all of them: hasty centralisation without the right talent, rushed decentralisation without governance, and compressed timelines that preclude the learning the organisation never invested in. #### The shared failure mode Three strategies, three different mechanics, one identical outcome. The common thread is treating AI adoption as a technology deployment problem. Centralisation optimises for control. Decentralisation optimises for speed. Waiting optimises for risk avoidance. Each handles one variable while ignoring the others. Research from Israeli and Ascarza at Harvard Business School makes the diagnosis precise: many AI initiatives fail to scale because organisations lack "the organizational scaffolding to bridge technical potential and business impact." Technology enables progress, but without aligned incentives, redesigned decision processes, and a workforce equipped to collaborate with AI systems, even technically excellent pilots never become durable capabilities. Eatough, Ferrazzi, and colleagues documented a similar pattern in their 2026 analysis of industry adoption data: 88% of companies report regular AI use, yet performance gains plateau because employees "experiment with new tools but don't integrate them deeply into how work gets done." The problem is not which strategy you pick. The problem is that all three strategies treat AI as something to be installed rather than something to be learned. #### What works instead: adaptive AI integration ##### Thin centre, thick edges The organisations getting this right have converged on a structural pattern that borrows from both centralisation and decentralisation while avoiding the pathologies of each. A small central team (rarely more than five to eight people in a mid-sized enterprise) owns three things: guardrails, shared infrastructure, and evaluation frameworks. They define what responsible AI use looks like. They maintain common tooling, data pipelines, and model registries that any team can build on. They set evaluation standards so results are comparable across initiatives. They do not own execution. Domain teams own their own AI initiatives, operating within the central guardrails. A supply chain team runs its own demand forecasting pilots. A customer service team builds its own ticket classification system. Each team applies AI to problems they understand intimately, using infrastructure they don't have to build from scratch. The line between what gets centralised and what doesn't follows a simple heuristic: centralise what creates leverage across teams (infrastructure, governance, evaluation), decentralise what requires domain knowledge to get right (problem identification, solution design, success criteria). The ARGO framework describes this as three interdependent layers: shared foundation standards, central advisory resources, and contextual local implementation. The key word is interdependent. The centre enables the edges. The edges inform the centre. ##### Structured experimentation over strategy decks Eighteen-month transformation programmes fail because the technology, the organisation, and the competitive environment all change faster than the programme can adapt. The alternative is structured experimentation: time-boxed pilots with predefined success criteria, clear kill conditions, and a mechanism for scaling winners. Ninety-day cycles work for most organisations. Long enough to build something real, short enough to fail cheaply. Each pilot starts with a specific business problem (not "explore AI for marketing"), defines what success looks like in measurable terms, and commits to a decision at the end: scale, iterate, or stop. The scaling mechanism matters. Winners don't just get more budget. They get promoted onto the shared infrastructure layer so other teams can learn from and build on the approach. Failures get documented with equal rigour. The insight from a failed pilot is often more valuable than the output of a successful one. ##### Building the feedback loop The McClure and Gerdau research identifies five pillars of AI capability: Culture and Leadership, Human Capital and Operations, Data Architecture, Systems Infrastructure, and Governance and Regulatory Compliance. Organisations progress through stages (siloed, integrated, orchestrated) across all five simultaneously. You cannot be orchestrated on infrastructure while remaining siloed on culture. Cross-team learning is what moves organisations between stages. The mechanism needs to be lightweight enough that teams participate willingly and specific enough that captured knowledge is actionable. What worked, what failed, what surprised people. A monthly thirty-minute showcase where teams present results (including negative results) achieves more than a quarterly AI strategy review. The goal is making institutional learning a byproduct of doing the work rather than a separate initiative with its own programme manager and steering committee. When a supply chain team discovers that their forecasting model degrades in specific seasonal patterns, that finding should reach the finance team building revenue projections within days, not months. #### The capability stack most organisations miss ##### Technical foundations that determine outcomes Most AI failures aren't caused by choosing the wrong model. They're caused by insufficient foundations that make every project harder than it needs to be. Data readiness sits at the base. This doesn't mean having a data lake. It means having clean, documented, accessible data with clear ownership and known quality characteristics. Organisations that skip this step find every AI project begins with three months of data wrangling before any modelling starts. Integration architecture determines whether AI outputs reach the systems where decisions happen. A brilliant recommendation engine is worthless if its output can't flow into the CRM, the ERP, or the workflow tool where the human acts on it. Evaluation infrastructure (the ability to measure whether an AI system is performing as expected in production, not just in testing) separates organisations that learn from organisations that guess. ##### Organisational foundations that matter more The Harvard Business School research makes this point with uncomfortable clarity: technology isn't the biggest challenge. Culture is. Decision rights must be explicit. Who decides which AI projects get funded? Who decides when a pilot gets killed? Who decides whether a model is safe to deploy in production? Ambiguity in decision rights produces either paralysis (nobody decides) or chaos (everybody decides). Incentive alignment shapes behaviour more than strategy documents. If business unit leaders are measured on quarterly revenue and AI projects take two quarters to show results, those projects will be deprioritised regardless of their strategic importance. Aligning incentives means adjusting how success is measured during the transition period. Psychological safety for experimentation is non-negotiable. If a team that runs a failed AI pilot faces budget cuts or career consequences, every other team in the organisation learns to avoid experimentation. The organisations that build AI capability fastest are the ones where failure in a structured experiment carries no stigma, while failure to experiment does. AI literacy is the final piece. This doesn't mean teaching everyone to code or understand transformer architectures. It means building enough shared vocabulary that business leaders can have productive conversations with technical teams, that domain experts can identify problems worth solving with AI, and that everyone can critically evaluate AI outputs rather than treating them as infallible. #### Getting started without picking a lane For organisations stuck between the three default strategies, the first step is honest assessment. Where are you on the SIO progression model across each of the five pillars? Most organisations overestimate their readiness in the areas they've invested in (typically infrastructure) and underestimate their gaps in the areas they've ignored (typically culture, governance, and human capital). The second step is identifying two or three high-signal pilots. High-signal means the problem is well-defined, the data exists, the business impact is measurable, and a domain expert is willing to co-own the initiative with a technical counterpart. Avoid the temptation to pick the most impressive use case. Pick the one most likely to produce a clear result in ninety days. The third step is building the minimum viable governance layer. This isn't a hundred-page AI policy document. It's a one-page set of guardrails covering data privacy, model evaluation, and escalation procedures. Expand it as you learn. Resist the impulse to over-engineer governance before you have anything to govern. The first ninety days should produce two things: a completed pilot with documented results (positive or negative) and a working governance framework tested against a real project. Everything else (vendor selection, platform architecture, talent strategy, operating model design) can wait until you have evidence from your own organisation about what works. What to ignore until later: comprehensive AI strategy documents, organisation-wide training programmes, multi-year transformation roadmaps, vendor bake-offs comparing fifteen platforms. These activities feel productive. They aren't. They're sophisticated forms of delay that substitute planning for learning. The organisations that build durable AI capability share a common trait: they started before they felt ready, learned from what happened, and built organisational muscle through repetition. There is no shortcut to that process, and there is no strategy document that substitutes for it. If you're ready to build AI capability that goes beyond pilots and exploits the full technical potential of what's possible today, get in touch. #### References - Why AI Readiness Is an Organizational Learning Problem, Not a Technology Purchase (arxiv, 2025) - Most AI Initiatives Fail. This 5-Part Framework Can Help. (HBR, 2025) - Why Companies That Wait to Adopt AI May Never Catch Up (HBR, 2018) - Building the AI-Powered Organization (HBR, 2019) - Why AI Adoption Stalls, According to Industry Data (HBR, 2026) - An Adaptive Responsible AI Governance Framework for Decentralized Organizations (arxiv, 2025) --- ### What I've Learned Shipping AI Products That Most Consulting Advice Gets Wrong - URL: https://agathon.ai/insights/what-ive-learned-shipping-ai-products-that-most-consulting-advice-gets-wrong - Published: 2026-03-26 - Categories: AI Strategy, AI Consulting, NLP I need to be blunt about something. Most of what passes for "AI product strategy" advice is written by people who have never shipped an AI product. I don't mean that as a throwaway provocation. I mean it literally. The strategy decks, the thought leadership pieces, the conference talks: the vast majority are produced by consultants who advise on AI from a comfortable distance, by analysts who study it from the outside, or by founders who shipped one product and are now generalising from a sample size of one. The advice sounds reasonable. It's often directionally correct. And it falls apart the moment it meets the reality of building and shipping AI products at a level that actually matters. I've spent fifteen years building AI systems commercially. Not advising on them from the outside but building them, shipping them, maintaining them, watching them fail and figuring out why. Financial services, telecoms, automotive, government. From early NLP systems that would look primitive by today's standards through to current-generation compound AI architectures. Every one of those engagements taught me something that contradicted the conventional wisdom of its era. This piece is the distillation of those lessons. Not research findings. Not best practices extracted from case studies. What I've personally seen go wrong, what actually works, and why the gap between AI strategy and AI reality is wider than anyone wants to admit. #### The Gap Between AI Strategy and Reality Let me start with the three things that surprise almost every first-time AI product builder, almost universally. These are so consistent that I will now usually flag them in the first week of any engagement. ##### Surprise 1: Your First Architecture Will Be Wrong Not "suboptimal." Wrong. Every AI product I've seen ship has undergone at least one fundamental architectural revision within six months of launch. Not a tweak. A rethinking of how the core components fit together. This happens because AI products reveal their requirements in production, not in planning. You can design the most elegant architecture on a whiteboard, informed by every best practice document ever written, and production traffic will invalidate your assumptions within weeks. The user queries you anticipated account for maybe 40% of what people actually ask. The retrieval strategy that worked on your curated test set breaks on real-world data with its inconsistencies, ambiguities, and sheer volume. The latency budget you allocated turns out to be incompatible with the quality users expect. The conventional advice -- "plan carefully, get the architecture right, then build" -- is backwards for AI products. What I've learned is that you should plan to be wrong. Design for replaceability. Build your system so that swapping out the retrieval layer, changing the model, or restructuring the orchestration flow is a measured effort, not a rewrite. The teams that suffer most are the ones who invested heavily in a "perfect" architecture before they had production data. They're emotionally and financially committed to decisions made with inadequate information, and the sunk cost makes them resistant to the changes that production reality demands. ##### Surprise 2: Evaluation Is Your Actual Product I can almost hear the objection: "No, the user-facing capability is the product." Technically true. Practically misleading. Here's what I mean. In traditional software, the product works or it doesn't. A button either submits the form or it doesn't. An API either returns the correct data or it throws an error. Testing is important, but the gap between "works in testing" and "works in production" is manageable. In AI products, the output is probabilistic. The same input can produce different outputs. "Correct" is often subjective and context-dependent. The system can fail in ways that look like success -- a confident, fluent, completely wrong answer is far more dangerous than a crash. This means your ability to evaluate your system's output is, in a very real sense, your ability to build the product at all. Without robust evaluation, you cannot: - Know whether a change improved the system or degraded it - Identify failure modes before users discover them - Prioritise what to work on next - Make any claim about quality with confidence I've worked with teams that spent months building features and days building evaluation. Every single one of them regretted it. The features they built might have been brilliant -- they had no way to know, because they couldn't measure the impact. The teams that ship great AI products invest disproportionately in evaluation infrastructure. They build it first, not as an afterthought. They treat it as a core product capability, not a testing concern. What does good evaluation infrastructure look like in practice? - Layered evaluation: Unit tests for individual components (does the retrieval return relevant documents?), integration tests for the full pipeline (does the end-to-end response answer the question?), and behavioural tests for emergent properties (does the system handle adversarial inputs gracefully?). - Human-in-the-loop calibration: Automated metrics are necessary but insufficient. You need regular human evaluation to calibrate your automated metrics against actual quality. This doesn't need to be expensive -- even a few hours per week of structured human review dramatically improves your understanding of system performance. - Regression detection: Every change to the system -- model update, prompt modification, retrieval parameter adjustment -- should be evaluated against a held-out test set before reaching production. Not a full test suite for every change, but a calibrated sample that catches regressions in critical capabilities. - Business-connected metrics: Ultimately, your evaluation needs to connect to business outcomes. Response quality that users don't notice or don't value is irrelevant, regardless of how well it scores on benchmarks. The best evaluation frameworks I've seen track the chain from technical metrics through user behaviour to commercial outcomes. ##### Surprise 3: The Feedback Problem Is Harder Than the Model Problem Every AI strategy deck includes a slide about "continuous improvement" and "learning from data." None of them adequately describe how difficult this is in practice. The challenge is this: to improve an AI system systematically, you need signal about what's working and what isn't. In theory, user interactions provide this signal. In practice, extracting useful signal from user behaviour is one of the hardest problems in AI product development. Consider: a user asks your AI product a question, receives a response, and moves on without providing explicit feedback. Was the response good? You don't know. They might have been satisfied. They might have been dissatisfied but too busy to complain. They might have reformulated their question elsewhere. They might have used the response despite it being wrong, because they couldn't evaluate its accuracy. Explicit feedback mechanisms (thumbs up/down, ratings) help but are biased -- users who provide feedback are not representative of users who don't. Implicit signals (time spent reading, follow-up actions, return visits) are noisy and ambiguous. Building feedback loops that actually work requires: - Careful signal design: What user behaviours genuinely indicate quality? This varies enormously by product and use case. I've seen teams build feedback systems around signals that, upon investigation, correlated more with user mood than with response quality. - Bias-aware aggregation: Explicit feedback over-represents power users and under-represents the silent majority. Implicit feedback over-represents easily measured actions and under-represents actual value delivered. You need to account for these biases, not just average the numbers. - Closing the loop: Having feedback data is necessary but not sufficient. You need processes that convert feedback signal into system improvements -- whether that's prompt refinements, retrieval tuning, model fine-tuning, or architectural changes. This requires infrastructure, tooling, and disciplined prioritisation. - Latency tolerance: AI system improvements often take time to manifest and measure. Unlike a UI change where you can A/B test in days, AI improvements may require weeks of accumulated usage data to evaluate reliably. Teams accustomed to rapid iteration cycles in traditional software find this deeply uncomfortable. #### The Hardest Part Isn't the Model Let me say this as plainly as I can: the model is the easiest part of an AI product. I realise that's counterintuitive. The model is the headline. It's what investors ask about. It's what the press covers. It's the thing that feels most like "real AI." But in terms of the effort required to build a production AI product, the model -- whether you're training it, fine-tuning it, or calling an API -- accounts for perhaps 15-20% of the total work. The other 80% is everything else. Data pipelines that ingest, clean, chunk, and index your domain data reliably and at scale. Retrieval systems that find the right information for each query, not just the most semantically similar text. Orchestration logic that routes requests through the right sequence of components. Evaluation frameworks that tell you whether the system is actually working. Monitoring infrastructure that catches degradations before users do. Feedback pipelines that convert usage data into improvements. None of this is glamorous. None of it makes for good demo material. All of it is essential, and all of it is harder than most teams expect. The most common failure pattern I see is teams that allocate 80% of their effort to model selection and prompt engineering, and 20% to everything else. They get the ratio exactly backwards. #### AI Product Timelines Are Not Software Timelines This is perhaps the lesson I've paid the most to learn, and the one that most consistently trips up experienced software leaders entering AI product development. Traditional software development is approximately linear. Twice the effort yields roughly twice the progress. Experienced teams can estimate timelines with reasonable accuracy. A feature that took three weeks last quarter provides a useful reference point for a similar feature next quarter. AI product development is fundamentally nonlinear. Progress comes in discontinuous jumps separated by plateaus. The first 70% of capability arrives quickly -- often encouragingly quickly. The next 20% takes as long as the first 70%. The final 10% takes longer than everything before it combined. And you often can't predict which 10% will be hardest until you're deep into it. I've seen this pattern so consistently that I now build it explicitly into project planning: Phase 1: Rapid progress (weeks 1-6). Everything works better than expected. Demo quality is impressive. Stakeholders are excited. The temptation to set aggressive launch dates is overwhelming. Phase 2: The plateau (weeks 7-16). Progress slows dramatically. Edge cases multiply. The gap between "works on the demo" and "works reliably in production" becomes apparent. Quality improvements require disproportionate effort. This is where projects stall, where morale drops, and where inexperienced teams make their worst decisions. Phase 3: Hard-won gains (weeks 17+). With discipline and the right approach, the system begins to improve again -- but through systematic engineering (evaluation frameworks, data pipeline improvements, architectural refinements) rather than the quick wins that characterised Phase 1. The teams that ship successfully are those that expect Phase 2 and plan for it. The teams that fail are those that mistake Phase 1's progress for a sustainable trajectory and make commitments accordingly. ##### What This Means for Planning Traditional estimation techniques don't work for AI products. I've tried story points, t-shirt sizing, evidence-based scheduling -- none of them produce reliable forecasts for AI development work. What works instead: - Milestone-based rather than time-based planning. Define what "good enough" looks like for each capability, and track progress toward that milestone rather than estimating when you'll arrive. - Explicit uncertainty budgets. For any AI capability, allocate at least 50% of your estimated time as uncertainty buffer. This isn't padding -- it's acknowledging the nonlinear nature of the work. I've been doing this for fifteen years and still regularly underestimate. - Progressive commitment. Don't commit to a launch date until you're through Phase 2 for your core capabilities. Promise a demo by a date, if you must. Promise progress reviews at regular intervals. Do not promise a production-quality product by a specific date until you have production-quality evidence. #### The Demo Trap This is the single most common pattern I see leading to bad decisions. AI demos are seductive. A well-crafted demonstration can make an early prototype look like a nearly-finished product. The model produces fluent, confident output. The carefully chosen example queries showcase the system's strengths. The audience -- typically investors or executive stakeholders -- comes away believing the product is 80% complete when it's closer to 30%. I've watched this play out dozens of times. The demo goes well. Expectations set accordingly. Timelines committed. And then reality: the system handles the demo queries beautifully because those queries were used to calibrate the system. Real users, with their unpredictable queries, edge cases, and adversarial inputs, expose every limitation the demo concealed. The demo trap creates three specific problems: Premature timeline commitments. Stakeholders who've seen an impressive demo expect a short path to production. When Phase 2 arrives and progress plateaus, they interpret it as a team performance problem rather than an intrinsic characteristic of AI development. Architecture lock-in. The demo often enshrines specific architectural decisions that were optimised for the demo scenario rather than for production reality. Teams resist changing an architecture that "already works" -- even though it works only for a narrow set of carefully selected inputs. Resource misallocation. Based on the demo, leadership allocates resources for a "polish and ship" phase when what's actually needed is a "solve the hard problems" phase. The latter requires different skills, different timelines, and different expectations. ##### How to Demo Honestly I'm not suggesting you avoid demos. They're necessary for fundraising, stakeholder alignment, and team morale. But there are ways to demo that don't create the trap. - Include failure cases. Show queries where the system struggles, alongside your plan for addressing them. This builds credibility and sets realistic expectations. - Quantify performance. "Here's our evaluation score across 500 test queries, broken down by category" is more useful than "watch it nail this one example." - Separate demo quality from production quality. Be explicit: "This demo represents our best-case performance. Production quality across all user queries is currently lower, and here's our plan to close the gap." - Demo the evaluation, not just the output. Showing your evaluation infrastructure demonstrates sophistication and builds confidence that you know what "done" looks like -- even if you're not there yet. #### Practical Patterns for Managing AI Product Uncertainty After fifteen years of navigating this uncertainty, I've converged on a set of patterns that work. None of them are revolutionary. All of them are underused. ##### Pattern 1: Parallel Experimentation with Shared Evaluation Rather than pursuing a single approach and hoping it works, run two or three approaches in parallel with a shared evaluation framework. This sounds expensive, and it is -- in the short term. In the long term, it's dramatically cheaper than pursuing one approach for months, discovering it doesn't work, and starting over. The key is the shared evaluation framework. If each experiment has its own definition of "success," you can't compare them meaningfully. Define your evaluation criteria upfront, build the infrastructure to measure consistently, and let the data tell you which approach wins. In practice, I structure this as a time-boxed spike: two weeks to implement each approach at a basic level, followed by evaluation against the shared framework. The winner gets full investment. The losers provided information that's often as valuable as the winner's success. ##### Pattern 2: Progressive Rollout with Instrumentation Don't launch to everyone at once. Roll out to a small cohort with comprehensive instrumentation, learn from their usage, adjust, expand. This is standard practice in software, but AI products need more aggressive instrumentation than most teams implement. At minimum, log: - Every input query and the system's response - Retrieval results and ranking scores - Model confidence signals (where available) - Latency at each pipeline stage - Any explicit user feedback This data is simultaneously your evaluation set, your debugging tool, and your training data for future improvements. Teams that instrument aggressively in early rollout build an enormous advantage over those that ship and hope. ##### Pattern 3: The "What Would We Need to Believe?" Test Before committing to any significant architectural decision, I force the team to articulate what assumptions need to be true for this decision to be correct. Not "why is this a good idea?"; the team will always have reasons. But "what specific, testable assumptions are we making, and how will we know if they're wrong?" For example: "We're choosing to use RAG rather than fine-tuning. This assumes that our domain knowledge can be effectively encoded in retrievable documents rather than model weights. We'll know this assumption is wrong if retrieval quality on our production query distribution falls below X, or if users consistently report that the system lacks domain expertise despite having access to the relevant documents." This practice doesn't prevent wrong decisions. It prevents wrong decisions from persisting long past the point where the evidence should have triggered a course correction. ##### Pattern 4: Dedicated "Red Team" Time Every sprint -- or every two weeks at minimum -- dedicate time to actively trying to break your system. Not automated testing (though that's necessary too). Dedicated human effort to find the queries, edge cases, and interaction patterns that expose weaknesses. This is psychologically difficult. The team has spent the sprint building capabilities and wants to celebrate progress. Spending time finding failures feels demoralising. But the failures you find internally are infinitely preferable to the ones your users find. And every failure you discover is a test case that strengthens your evaluation framework going forward. I've found that rotating this responsibility across the team works best. Different people have different intuitions about where systems break, and the diversity of approaches yields better coverage than any single person's red-teaming instincts. ##### Pattern 5: Ship the Guardrails Before the Features This is counterintuitive, and it's the pattern I have to argue for most strenuously. Before shipping a new AI capability, ship the guardrails that will constrain it. What does the system do when it doesn't know the answer? When the user's query is ambiguous? When the retrieval returns irrelevant results? When the model's confidence is low? When the generated output contradicts known facts in your knowledge base? These guardrails are not afterthoughts. They're the difference between an AI product that degrades gracefully and one that fails catastrophically. And they're far easier to build before the feature ships than after -- when you're scrambling to patch failures in production while users are actively encountering them. In my experience, the quality of an AI product's guardrails is a better predictor of its long-term success than the quality of its primary capability. A moderately capable system with excellent guardrails consistently outperforms a highly capable system with poor guardrails, because the latter's failures destroy user trust in ways that are extraordinarily difficult to recover from. #### When This Advice Doesn't Apply I want to be careful not to overstate. There are contexts where the patterns I've described are less relevant. If you're building internal tools where the user base is small, expert, and tolerant of imperfection, you can move faster and with less instrumentation. The cost of failure is lower, and the feedback loop is tighter because you can talk directly to your users. If you're in a research context rather than a product context, the emphasis on evaluation infrastructure and guardrails may be premature. Explore first, productionise later. If your AI component is genuinely simple -- a single-turn classification, a straightforward summarisation -- the complexity I've described may not apply. Not every AI feature requires compound architecture and sophisticated evaluation. Sometimes a well-engineered API call genuinely solves the problem, and over-engineering it is its own form of failure. But if you're building a product where AI capability is central to the value proposition, where users will interact with the AI in complex and unpredictable ways, and where the quality of the AI's output directly impacts your business -- then these patterns apply. I've learned them through years of getting things wrong before getting them right, and they've proven consistent across every domain and company stage I've worked in. #### The Real Competitive Advantage Here's what I've come to believe after fifteen years of this work: the competitive advantage in AI products isn't the model, the architecture, or even the data. It's the operational maturity to ship, evaluate, learn, and improve faster than your competitors. The team that ships an imperfect product with excellent evaluation and rapid learning loops will outperform the team that spends a year building a "perfect" system every single time. Because the first team is accumulating real-world signal -- the only kind of signal that matters -- while the second team is optimising against assumptions that may or may not hold. This is the operational knowledge that separates teams that ship successful AI products from teams that produce impressive demos and never quite get to production. It's not glamorous. It's not the kind of insight that makes for a good conference talk. But it's what actually determines outcomes. These are the challenges I work through with founders and technical leaders every day. Not strategy at arm's length, but the operational reality of building AI products that work -- for real users, at production scale, with all the mess and uncertainty that entails. When you engage Agathon, you work directly with me. I've built these systems. I've made the mistakes. I've developed the patterns that prevent them. My background -- from mathematics at Oxford through NLP research at Cambridge to fifteen years of commercial AI systems -- gives me the technical depth to work at the level where these problems actually live, not the level where they get discussed in strategy decks. I work with founders to build their teams' capability to navigate this uncertainty independently. Because the reality of shipping AI products is that the challenges don't stop -- they evolve. And the most valuable thing I can leave you with isn't a strategy document or an architecture diagram. It's the operational judgment to make good decisions when the next unexpected challenge arrives. That's what building capability rather than dependency means in practice. And it's the only approach I've seen that works over the long term. --- ### The AI Team Decision: When to Build Internally, When to Stay Fractional, and When to Do Both - URL: https://agathon.ai/insights/the-ai-team-decision-when-to-build-internally-when-to-stay-fractional-and-when-to-do-both - Published: 2026-03-04 - Categories: Fractional CTO, AI Strategy, AI Advisory Every founder or CEO I speak to about AI capability eventually asks some version of the same question: should we hire our own AI team, or keep working with external people? It sounds like a binary choice. It isn't. The companies that get this right treat AI resourcing as a progression, one that evolves with their maturity and what they can actually support. The companies that get it wrong tend to make one of two mistakes: they hire too early and burn through expensive talent before they're ready to use it, or they outsource indefinitely and never develop the internal muscle they'll eventually need. This piece is about how to think through that progression honestly. I'll share what works, what fails, and how to figure out where you stand. #### The false binary: why "build vs outsource" is the wrong question The technology industry loves a clean dichotomy. Build or buy. In-house or agency. But the AI team decision doesn't sit neatly in either camp, because your needs change as your organisation matures. What I've seen across multiple engagements is a natural progression that looks roughly like this: Stage 1: Fractional and external. You don't have a clear AI workload yet. You're exploring what's possible, running initial experiments, maybe integrating off-the-shelf AI tools into existing products. At this stage, you need strategic guidance and focused execution bursts, not a permanent team sitting idle between projects. A fractional CTO or AI advisor makes sense here, because they bring pattern recognition from multiple organisations and can prevent the most expensive mistakes before you make them. Stage 2: Hybrid. You've validated that AI is central to your product or operations. You have recurring AI workload and reasonably clean data. Now you need a small internal core — perhaps a data engineer and an ML engineer — working alongside senior fractional leadership that provides architectural direction and quality assurance. The fractional leader's job shifts over time from delivery toward mentoring and capability transfer. Stage 3: Internal with selective external support. Your internal team owns the AI roadmap, the infrastructure, and the day-to-day delivery. You bring in external specialists for specific challenges — a novel architecture, a domain you haven't tackled before, an independent review of your approach. The fractional relationship either ends or becomes genuinely advisory: a few hours a month of strategic counsel rather than hands-on delivery. This progression isn't theoretical. MIT's Center for Information Systems Research surveyed 721 companies and found that only 7% have reached what they call "AI Future-Ready" status, with fully embedded internal teams delivering measurable value. The vast majority (62%) are still in the first two stages, where some combination of external expertise and internal capability-building is exactly the right approach. The better question is: which stage are you actually in, and are you resourcing accordingly? #### Four signals that you're ready to start building internally Not every company should rush to hire an AI team. In fact, premature hiring is one of the most expensive mistakes I see. But there are clear signals that indicate you're ready to start bringing capability in-house. 1. You have recurring, predictable AI workload. If your AI needs come in sporadic bursts — a model for this product, an experiment for that initiative — external support is more efficient. When you find yourself needing sustained, daily AI work across multiple product areas, that's when internal hires start making financial sense. 2. Your data infrastructure can actually support an AI team. This one catches people out repeatedly. Data scientists spend roughly 80% of their time wrangling data rather than building models. If your data pipelines are fragile, your data quality is poor, or your data lives in disconnected silos, hiring an ML engineer is like buying a racing car before you've built the road. Fix the infrastructure first. That's work a data engineer can do, incidentally, and it's a much better first AI hire than a data scientist. 3. You have a leader who can set direction. A team of junior or mid-level AI engineers without senior AI leadership will drift. They'll build technically impressive things that don't connect to business value. They'll make architectural decisions that create technical debt you won't discover for months. If you can't yet afford or justify a full-time Head of AI, a fractional AI leader working alongside your first internal hires is significantly more effective than leaving them unsupervised. I'll come back to this. 4. You have executive sponsorship that goes beyond enthusiasm. "We should do something with AI" is not sponsorship. Sponsorship means someone senior owns the AI roadmap, can unblock cross-functional dependencies, and will fight for the budget and organisational patience that AI capability requires. Without this, even excellent AI teams get marginalised into a Centre of Excellence that nobody listens to. #### Three signals that you're not ready (even if you think you are) These are harder to accept, because they often coexist with genuine ambition and excitement about AI. 1. You're hiring to "figure it out." If your brief to an AI hire is essentially "come and work out what we should be doing with AI," you are not ready to hire. You're asking someone to simultaneously define the strategy, build the infrastructure, deliver projects, and demonstrate ROI, with no institutional support. The average tenure of a Chief Data Officer is 2 to 2.5 years, and a significant reason is exactly this: organisations hire senior data and AI leaders without the readiness to support them. What you actually need at this stage is a short advisory engagement to define the strategy, followed by targeted hiring against a clear plan. 2. Your leadership team treats AI as IT's problem. Research from A.Team found that 84% of companies plan to increase AI investment, but the most common mistake is delegating the entire initiative to the technology function. AI capability that lives exclusively in IT gets disconnected from commercial reality. The models get built, but they don't get deployed, because nobody in the business owns the problem they're supposed to solve. If your board talks about AI but your product and commercial leaders aren't involved in defining what success looks like, you're not ready for internal hires. 3. You don't yet have a problem worth solving with AI. This sounds obvious, but it's remarkably common. RAND Corporation's research — the most methodologically rigorous study on AI project failure — found that the single most common root cause of failure is misunderstanding what problem needs solving. Companies hire AI teams to go looking for problems, rather than hiring AI teams to solve problems they've already identified and validated. If you can't articulate a specific, measurable business problem that AI could address, you need discovery work, not a permanent team. #### The hidden costs of hiring too early The numbers on premature AI hiring are worse than most leaders expect. Start with the direct costs. A senior ML engineer in the UK commands £100,000 to £127,000 in base salary. A Head of AI in London ranges from £111,000 to £269,000. In the US, the numbers are even higher — a VP of AI averages $351,000, and compensation for senior AI engineers runs $150,000 to $280,000. These are competitive market rates, driven by an AI talent demand-to-supply ratio of 3.2 to 1 globally. Now consider: it takes an average of 142 days to fill an AI role, compared with 44 days for the general market. During that vacancy, you're losing roughly $2,500 per week in productivity. Once hired, a senior technical hire takes 6 to 12 months to reach full productivity — longer in organisations without existing AI infrastructure, because they're building the foundations as well as the product. If the hire doesn't work out — and in a field where average tech tenure is 2 to 3 years, turnover is a real risk — you're looking at a total cost of 1.5 to 3 times annual salary when you factor in recruitment fees, onboarding, lost productivity, severance, and the cost of starting the search again. For a senior AI engineer on £120,000, that's £180,000 to £360,000 for a failed hire. Compare this with a fractional AI leadership engagement. A fractional CTO or Head of AI working two days per week typically costs 60 to 80% less than a full-time equivalent at the same seniority level, brings cross-industry pattern recognition from multiple organisations, and can start delivering strategic value within weeks rather than months. The financial argument for fractional leadership at the early stages isn't even close. But the financial costs aren't the most dangerous part. The real damage from premature hiring is organisational. ##### The "science fair" failure mode You hire talented ML engineers and data scientists. They build genuinely impressive prototypes. Those prototypes never make it to production. IDC and Lenovo data shows that 88% of AI proofs of concept never reach production — and the primary reason is that they were built in isolation from the business context that would make them useful. Without senior AI leadership connecting technical capability to commercial reality, you end up with a team that publishes internal demos but doesn't ship products. ##### The "island" failure mode Your AI team becomes a silo. They sit in engineering, far from the business teams whose problems they're supposed to solve. They build technically correct solutions that nobody trusts, nobody adopts, and nobody asked for. Harvard Business Review documented cases where AI teams in banks were simultaneously flagging customers as too risky for lending while marketing was targeting those same customers for growth — because the AI team had no line of sight into commercial strategy. ##### The "revolving door" failure mode You hire a Head of AI or Chief Data Officer. They arrive with energy and ambition. At around 18 months, they hit a wall — the organisation wasn't ready for the change they were hired to drive, they don't have the cross-functional authority to unblock adoption, and they leave. The average CDO tenure is 2 to 2.5 years, and nearly a third of current CDOs question the long-term viability of their own role. When they leave, they take their institutional knowledge with them, and you start again. All of them are preventable. #### How to structure a hybrid model that actually works For most companies in the middle of this progression — past pure exploration but not yet ready for a fully internal team — the hybrid model is the right answer. But "hybrid" is often used as a vague catch-all. Here's what it looks like in practice when it's structured well. ##### The fractional leader sets direction; the internal team builds muscle The most effective hybrid arrangements I've seen pair a senior fractional AI leader (typically two days per week) with a small internal core — one or two engineers who own the day-to-day execution. The fractional leader defines the architecture, reviews code and model decisions, manages technical risk, and mentors the internal team. The internal team builds the domain knowledge, maintains the systems, and gradually takes on more strategic responsibility. I worked with one early-stage company where the founder had a strong product vision involving AI but no technical leadership to execute it. We established a structure where I worked a set number of fixed days per week alongside a single senior developer. My role covered everything from product architecture and commercialisation strategy to managing the development workflow and quality. Over several months, the developer's capability grew substantially — not because I was formally training them, but because they were working within a structure that demanded good architectural decisions, clear code review processes, and consistent delivery standards. The key was that the engagement was designed from the start to build internal capability, not to create dependency. Every architectural decision was documented. Every technical choice was explained, not just implemented. The goal was always that the team could eventually operate independently, with my involvement reducing over time from delivery to oversight to occasional strategic advice. ##### Define the handover triggers in advance Before the engagement starts, agree on what "ready for internal leadership" looks like. Specific, measurable indicators: the internal team can independently architect new features, they can evaluate and manage external technical resources, they can make infrastructure decisions without senior oversight. When those triggers are met, the fractional engagement scales down. If you don't define these upfront, engagements drift into comfortable dependency, which is a failure of the model even if the work is good. ##### Protect the structure Hybrid models fail when the boundaries get blurred. If the fractional leader ends up doing all the strategic work while the internal team only executes, no capability transfer happens. If the internal team makes architectural decisions without review because the fractional leader isn't available that day, quality drifts. The structure needs fixed touchpoints: regular sessions where decisions are made together, code reviews that are genuinely educational, and a clear escalation path for urgent issues that respects the fractional leader's committed hours. #### The "capability, not dependency" principle Most AI consultancies won't say this, so I will: the goal of any external AI engagement should be to make the external person unnecessary. This isn't how the consulting industry typically works. The major firms have invested billions in AI practices — Accenture's $3 billion Data & AI investment, Deloitte's $4 billion AI services plan — and their business model depends on long-term, high-value engagements. Research from HFS and IBM found that 65% of enterprises now say traditional consulting models no longer deliver value, and 73% of consulting buyers want fundamentally different pricing models. There's a reason for that frustration: many consulting relationships are structured around dependency. The fractional model, done well, is structurally different. You're not buying a team of junior consultants managed by a partner who appears at the steering committee. You're working directly with a senior practitioner who has built and shipped the things they're advising you on. And that practitioner's success is measured by how effectively they transfer knowledge and make themselves redundant. When I think about the engagements I'm most proud of, they're the ones where the client eventually said: "We don't need you for this anymore." That's the whole point. This doesn't mean the relationship necessarily ends completely. Companies that have built strong internal AI teams still benefit from occasional external perspective — an independent architecture review, strategic advice on a new technical domain, a sounding board for a hire they're considering. But the nature of the engagement shifts from delivery to advisory, and the cost drops accordingly. #### Designing your progression: a practical framework If you're a founder or CEO trying to work out where you are in this progression, here's a simplified diagnostic: You're at Stage 1 (Fractional) if: - AI is one of several strategic priorities, not the core of your product - You don't have dedicated data infrastructure or a data engineering function - Your AI workload is project-based rather than continuous - You haven't yet validated specific, measurable AI use cases - Right move: Engage a fractional AI leader to define strategy, validate use cases, and build a roadmap for capability development You're at Stage 2 (Hybrid) if: - You have validated AI use cases delivering measurable value - You have functional data pipelines and reasonable data quality - Your AI workload is recurring and growing - You need someone working on AI daily, not just in project bursts - Right move: Hire your first internal AI engineers, paired with fractional senior leadership for architectural direction and mentoring You're at Stage 3 (Internal with selective support) if: - Your internal team can independently architect, build, and deploy AI features - You have established code review, deployment, and monitoring processes - Your AI leaders participate in commercial and product strategy, not just engineering - You bring in external expertise for genuinely novel challenges, not routine delivery - Right move: Transition fractional leadership to a light advisory relationship; invest in growing your internal team's seniority and breadth Most companies I work with are somewhere in the transition between Stage 1 and Stage 2. And that's fine: 62% of companies sit there, according to MIT's enterprise AI research. Being early in the progression is fine. Pretending you're further along than you are is where the damage happens. #### What the current market means for your decision Two things in the current market are worth considering before you commit to a resourcing model. First, we're in what Gartner calls the "Trough of Disillusionment" for generative AI. The initial hype has cooled. Only 5% of companies are capturing significant value from AI, and 60% are generating no measurable value at all. MIT found that internal AI builds succeed only about a third of the time, compared with roughly two-thirds for purchased or externally supported solutions. Which doesn't mean you shouldn't build internal capability. It means you should build it carefully, with senior guidance, rather than hiring aggressively and hoping for the best. Second, the economics of AI teams are shifting. AI coding tools are genuinely changing the calculus. Over 75% of developers now use AI coding assistants, and they're cutting onboarding time nearly in half. This means smaller, more senior teams can achieve what previously required larger groups. The emerging model is a strong architectural leader supported by a smaller number of capable engineers, amplified by AI tooling. This makes the fractional-to-hybrid progression even more relevant: you need strategic leadership more than you need headcount. #### Getting this right The AI team decision is an ongoing calibration between your ambition, your maturity, and what your organisation can actually support. The companies that navigate it well are honest about where they are, and they invest in capability that progressively builds independence. --- ### What Happens After You Discover The AI Isn't Real: a Due Diligence Decision Framework - URL: https://agathon.ai/insights/what-happens-after-you-discover-the-ai-isnt-real-a-due-diligence-decision-framework - Published: 2026-03-03 - Categories: AI Advisory, AI Strategy, Machine Learning Every AI due diligence content piece ends at "here's how to assess." This one starts where the others stop. You've done the technical due diligence. The findings are in. The AI isn't what the pitch deck claimed: the "proprietary models" are API wrappers, the "machine learning pipeline" is a rules engine with a modern interface, or the impressive demo hides architecture that won't survive its first thousand users. Now what? This is the moment most investors find themselves in with no playbook. The due diligence firm delivered a report full of technical findings, but the investment decision (walk away, renegotiate, or invest in fixing it) requires a framework that bridges technical assessment and commercial judgment. #### The AI misrepresentation spectrum: not all gaps are equal The first step after receiving concerning due diligence findings is categorisation. Not all AI capability gaps carry the same implications, and the appropriate response depends entirely on where on the spectrum your target company sits. After conducting technical assessments across multiple AI companies, we've found that findings cluster into four distinct categories. Understanding which you're dealing with is the single most important input to your investment decision. ##### Category 1: Marketing exaggeration The company labels rules engines, basic automation, or straightforward statistical methods as "AI" or "machine learning." The system works and customers use it, but it isn't what most technical practitioners would recognise as AI. This is the most common finding. A significant proportion of companies classified as "AI startups" show no evidence that AI is material to their core value proposition. Often the business is perfectly sound; it's the valuation multiple that's wrong. Commercially, the product works, customers get value, revenue is real. But the company has been valued at AI multiples (10–30x revenue) when it should be valued at standard SaaS or automation multiples (3–8x). The gap between these two sets of multiples is where the renegotiation lives. Remediation is low difficulty if the company has relevant data assets. Building ML capability on top of a working product typically costs £150,000–£750,000 over 6–18 months, assuming the team has competence in adjacent technical areas and the data exists to train meaningful models. The question worth asking: would adding real AI materially improve the product? Sometimes a well-built rules engine is the right technical choice, and the only thing that needs fixing is the marketing. ##### Category 2: Capability gap The company has actual machine learning, but it's significantly less sophisticated than represented. They pitched deep learning but deployed logistic regression. The model works in carefully controlled demos but degrades at production scale. The data science team exists but lacks the depth to execute the technical roadmap. This category is often the most productive zone for investment. The basics are in place. The team understands ML concepts, has built data pipelines, and has some model development experience. The gap is maturity, not capability. Commercially, the investment thesis may still hold, but timelines and milestones need recalibrating. The product roadmap probably has 12–24 months more technical risk than the pitch deck suggested. Customer promises may need tempering. Remediation is moderate difficulty, typically £400,000–£1.5 million over 12–24 months, driven primarily by hiring senior ML talent and investing in MLOps infrastructure. The question that decides it: does the existing team have the intellectual honesty and technical foundation to close the gap with the right investment, or will they resist acknowledging it? ##### Category 3: Third-party dependency disguised as proprietary capability The company claims proprietary AI but is actually wrapping third-party APIs (typically large language model providers) with minimal customisation. There are no proprietary models, no training infrastructure, and likely no relevant training data. The entire technical "moat" is a prompt template and an integration layer. This is not inherently a bad business. Plenty of valuable companies are built as integration layers. But it's a completely different type of business from what was represented, with different economics and a different defensibility profile. The company is subject to API pricing changes, terms of service modifications, and model deprecation, none of which affect companies with proprietary technology. The valuation should reflect a services or integration business, not a technology company. Margins are structurally vulnerable. Competitive moats are thin. Any competitor can build the same wrapper in weeks. Remediation is high difficulty. Building ML capability from scratch while simultaneously running a business on rented infrastructure costs £750,000–£4 million over 18–36 months. This means hiring an entirely new technical team, developing data collection and training pipelines from nothing, and managing the transition without disrupting existing customers. Many companies in this category never complete the transition. ##### Category 4: Fundamental technical deception The product is sold as AI automation but actually runs on human labour behind the scenes. Customer-facing outputs that appear automated are manually produced by teams of workers. The company operates at services margins while charging SaaS prices. The technology described in investor materials simply does not exist. Several high-profile enforcement actions in 2024 and 2025 have established that this category now carries criminal liability risk. Regulators, particularly the SEC and DOJ, have escalated from civil penalties to criminal indictments for AI misrepresentation in investment contexts. One company raised $42 million while claiming "proprietary deep learning" with 93–97% automation. The actual automation rate was effectively zero, with hundreds of offshore workers completing tasks manually. The founder now faces charges carrying decades of prison time. In most cases, this is a walk-away finding. The economics don't work: remove the humans and there's no product; keep the humans and the margins collapse. Beyond the financial calculation, proceeding with investment in a company that has made materially false representations about its technology creates regulatory, legal, and reputational exposure for the investor. Remediation is very high difficulty, estimated at £1.5–£8 million+ over 24–48 months, with a significant probability of failure. There is no AI capability to build on, no data infrastructure, and often no technical team capable of building what was described. The only viable path is effectively building a new technology company inside the existing commercial wrapper, and even then, the unit economics during transition are typically unworkable. #### The diagnostic question that matters most Across all four categories, one question separates fixable situations from fatal ones: could this company build the claimed capability, even though it doesn't exist yet? Answering that means looking at three things. Start with data assets. Does the company possess relevant, proprietary data that could support ML training? A company with rich proprietary data and a functioning data flywheel (where product usage generates more training data) has a completely different remediation profile from one relying entirely on third-party APIs or public datasets. The data doesn't need to be perfect, but it needs to exist and be relevant to the problem domain. Then look at team capability. Does the existing team have ML expertise, or only software engineering skills? Could they attract and retain ML talent? A minimum viable ML team (data engineer, ML engineer, data scientist, MLOps specialist, and product manager) runs approximately £450,000–£700,000 per year in loaded UK costs. The demand for AI/ML talent currently outstrips supply by roughly three to one, with mid-level salaries growing at nine percent year-over-year. If the company is in an unattractive location, sector, or stage for ML talent, remediation timelines extend considerably. Finally, infrastructure readiness. Does the company have the data pipelines, compute infrastructure, and development tooling to support ML development? Or would remediation require building from the ground up? Most AI systems that undergo rigorous assessment fail basic build quality benchmarks: poorly documented, built on legacy infrastructure, difficult to scale or transfer. The gap between "we have a Python notebook that runs locally" and "we have production ML infrastructure" is typically 6–12 months and several hundred thousand pounds. #### Pricing the gap: how AI capability findings should affect valuation When the due diligence findings are in and you've categorised the gap, the next question is what the findings mean for deal terms. ##### Marketing exaggeration (Category 1) Adjust the valuation from AI multiples to appropriate software or automation multiples. This can represent a 40–70% reduction in enterprise value, which is significant, but the business itself may remain highly investable. The negotiation framing is: "The technology works and the business is strong, but the valuation was predicated on AI differentiation that doesn't exist." ##### Capability gap (Category 2) The standard approach is shifting a meaningful portion of consideration from upfront payment to earnout, typically 30-50% of total deal value, contingent on achieving defined technical milestones. Useful earnout metrics include model performance KPIs (accuracy, latency, throughput at scale) and automation rates versus human intervention rates. Industry data shows earnouts typically pay out about 21% of their maximum value; when any earnout is achieved, approximately half the maximum is paid. Structure accordingly. ##### API dependency (Category 3) This often warrants restructuring the deal entirely. Consider an acquihire format (valuing the team and customer base rather than the technology), or a heavily milestone-gated investment where capital releases only as proprietary capability is demonstrated. Pricing should reflect integration-layer economics, not technology company economics. The negotiation centres on whether the team and customer base have enough value to justify the investment required to build real capability. ##### Fundamental deception (Category 4) Walk away. If investment documents contained material misrepresentations about technology capability, consult legal counsel about potential recovery actions. The SEC created a dedicated 30-person enforcement unit for AI-related fraud in February 2025, and enforcement actions are a bipartisan priority regardless of administration. #### Structuring deals when AI capability is uncertain When you've decided to proceed despite capability gaps, the deal structure itself becomes the primary risk management tool. Several mechanisms have matured rapidly over the past two years. Standard technology reps are insufficient for AI investments. Best practice now includes AI-specific representations covering: rights to training data, absence of data protection violations in data collection, accuracy of model architecture and performance disclosures, absence of undisclosed third-party dependencies (particularly API reliance), and compliance with emerging AI regulation including the EU AI Act. Leading firms recommend classifying these as fundamental representations with longer survival periods and higher or uncapped indemnity caps. Milestone-based funding has become increasingly common. Rather than deploying capital in a single tranche, investment releases in stages as the company achieves predetermined technical milestones. This is standard practice in life sciences and increasingly applied to AI investments. The hard part is defining precise, measurable milestones with contractual clarity on what constitutes achievement, and building in renegotiation protocols for when technology pivots make original milestones irrelevant. Standard holdbacks of 10–20% of purchase price, held for 12–24 months, can be supplemented with special-purpose escrows addressing specific AI risks: pending IP litigation, data provenance concerns, or regulatory compliance uncertainties. For Category 2 findings specifically, escrow release can be tied to remediation milestones, creating alignment between the seller's post-close incentives and the buyer's technical risk. Pre-closing covenants matter too. Require the target to retain AI engineers, maintain model architecture and datasets without material changes, ensure lawful use of training data, and maintain sufficient compute capacity. These covenants prevent the common scenario where technical talent departs or core technology is modified between signing and closing. One thing to watch: representations and warranties insurance (RWI) providers are increasingly scrutinising AI-specific risks and may exclude data provenance, model performance, and regulatory compliance claims. Don't assume standard RWI will cover AI misrepresentation. #### The remediation cost reality If you've decided to invest in fixing the gap, here's what the costs actually look like. A basic AI proof of concept runs £15,000–£50,000 over one to three months. Move to a custom ML solution at production quality and you're looking at £150,000–£400,000 over 6–12 months. Enterprise AI systems requiring integration, monitoring, and governance push past £400,000–£750,000+ and take 12–24 months. All of these assume relevant data exists and capable talent can be hired. The single longest timeline is institutional capability building. You can hire ML engineers, but integrating them into an organisation that has never operated with ML workflows, doesn't have MLOps infrastructure, and whose product managers don't understand model limitations takes 18–36 months to reach maturity. Google's widely cited 2015 research on hidden technical debt in ML systems remains authoritative: only a small fraction of real-world ML systems consists of actual ML code. The surrounding infrastructure (data pipelines, monitoring, serving, configuration) is vastly larger and accumulates debt faster than traditional software. The implication for investors: remediation budgets that account only for "building the model" routinely underestimate the true cost by 3–5x. The model is the easy part. Everything around it is where the time and money go. #### The regulatory dimension you can't ignore Enforcement has changed sharply. Between March 2024 and mid-2025, regulators escalated from modest civil penalties (two investment advisers settling for a combined $400,000 for unsubstantiated AI claims) to criminal fraud charges carrying decades of imprisonment. The trajectory is unmistakable, and it's bipartisan: the current US administration frames enforcement as protecting legitimate AI innovation by deterring fraudulent claims that divert capital from real companies. For investors, this creates a new dimension of risk. Investing in a company whose AI claims you know to be false, or should reasonably have known were false after conducting due diligence, creates potential exposure. The regulatory expectation is increasingly that sophisticated investors should verify AI claims before committing capital, and that failure to do so doesn't insulate them from downstream consequences if those claims prove fraudulent. In the UK, the FCA has not yet brought AI-specific enforcement actions, but applies existing Consumer Duty and Senior Managers & Certification Regime obligations to AI claims. The EU AI Act, with penalties of up to €35 million or 7% of global turnover, is still in staged implementation through 2027 but is already driving 20–30% valuation discounts for AI companies with potential regulatory non-compliance issues. AI technical due diligence is no longer optional for sophisticated investors. It's rapidly becoming a baseline expectation, both as investment discipline and as a regulatory compliance consideration. #### What this framework demands from your diligence process Most AI due diligence processes are built to answer one question: "Is the AI real?" That's necessary but insufficient. The framework above requires your diligence process to answer four more: Where on the spectrum does this sit? Marketing exaggeration, capability gap, API dependency, or fundamental deception? The category determines everything that follows. Is the gap fixable? Assess data assets, team capability, and infrastructure readiness. A company with strong proprietary data and a capable but overstretched team is a completely different proposition from one with no data moat and a sales-oriented leadership team. What does remediation actually cost? Not the optimistic estimate, the realistic one, including infrastructure, process change, talent acquisition, and the 18–36 months required to reach ML maturity. What deal structure protects the downside? Match protection mechanisms to severity: adjusted multiples for marketing exaggeration, milestone-gated earnouts for capability gaps, restructured deal formats for API dependency, and walk-away with legal review for fundamental deception. AI due diligence conducted by someone who has built production AI systems, rather than only reviewing them from the outside, is the foundation. But the real value is in what happens after the findings are in: the commercial experience to price the gap and the technical depth to design remediation plans that actually work. If you've already completed due diligence and need help interpreting findings, or if you're planning an assessment and want it structured to support post-discovery decision-making from the outset, get in touch. For engagements where the findings reveal fixable gaps, our AI leadership advisory service provides ongoing technical guidance through remediation, making sure the capability actually gets built. --- ### Why AI Startups Need a Different Kind of CTO, and What That Looks Like - URL: https://agathon.ai/insights/why-ai-startups-need-a-different-kind-of-cto-and-what-that-looks-like - Published: 2026-02-25 - Categories: Fractional CTO, AI Strategy, Machine Learning The CTO role at an AI-native startup is fundamentally different from the same role at a traditional software company. Most founding teams do not realise this until the consequences are already compounding. From the outside, the job looks similar: technical leadership, architecture decisions, team building, product delivery. The difference is invisible until it is not. US private AI investment reached $109.1 billion in 2024, according to Stanford's AI Index Report. As that capital flows into AI-native startups, the gap between the technical leadership these companies need and the technical leadership they typically hire is becoming one of the most consequential and least discussed problems in the ecosystem. #### The core difference: deterministic thinking meets a stochastic world Software engineering is deterministic. You write code. Given the same inputs, it produces the same outputs. When something breaks, you trace the logic, find the bug, fix it. Edge cases exist at the margins and you handle them one by one. AI product development is different. Machine learning models are probabilistic. They deal in likelihoods. Language is inherently ambiguous. User behaviour is unpredictable. Model outputs shift as data changes. There is no single "correct" answer to trace back to when something goes wrong, because the system was never designed to produce one. A CTO who has spent their career in deterministic systems will instinctively apply deterministic thinking to non-deterministic problems. In AI, that instinct is where things start to go wrong. The most significant consequence shows up in problem framing. In traditional software, you can build first and refine later. Ship, test, learn, improve. That works when system behaviour is bounded and predictable. In AI, "build it and fix it later" is catastrophic. A poorly framed problem compounds into months of wasted work. The model optimises for the wrong objective. The team trains on the wrong data, evaluates against the wrong benchmarks, bakes wrong assumptions into the architecture. By the time anyone notices, the cost of unwinding is enormous. Getting the question right up front makes all the difference. What exactly are we asking this system to do? How will we know if it is doing it well? What does failure look like? Software engineering leadership is not trained to prioritise these questions. They are the questions that determine whether an AI product succeeds or fails. The other thing that catches software-background CTOs off guard is edge cases. In traditional software, edge cases are exceptions at the margins. In AI, particularly in products with language at their core, the entire problem space is edge cases. Semantic ambiguity is not a bug to be fixed. It is a permanent condition to be designed around. A CTO who treats AI systems like software with some machine learning bolted on will consistently underestimate this. #### The "general CTO" failure pattern This is not theoretical. The same failure patterns recur across AI startups when technically competent engineering leaders without deep AI experience take the CTO role. Architecture decisions made with software assumptions. Systems get designed as if model behaviour is predictable and stable. Nobody builds in monitoring, retraining loops, evaluation infrastructure, or graceful degradation paths. The architecture works for the demo. It buckles in production, where the gap between deterministic assumptions and stochastic reality becomes undeniable. The pipeline cascade problem. AI products rarely rely on a single model. They involve pipelines where the output of one model feeds the input of another. Without AI systems experience at the top, nobody does the holistic thinking about how the whole system fits together. A model update early in the pipeline produces downstream consequences that nobody anticipated. Teams scramble with ad hoc fixes, patching symptoms rather than addressing structural fragility. This is one of the most common and most expensive failure modes in AI product development, and it stems from a leadership gap: nobody in the room understood that these systems need to be reasoned about as interconnected wholes, not independent components. The team structure trap. This shows up in two ways. In one version, data scientists and AI engineers sit in a separate team from software engineers, and an adversarial culture develops between the two. In the other, they sit together, but leadership has not set clear expectations about what each discipline does and what realistic timelines look like, so conflict builds within the team instead. Software engineers wonder why the data scientists cannot just "finish the model." Data scientists feel pressured to ship before the work is ready. Both versions trace back to a CTO who does not understand the workflow, cadence, and uncertainty profile of machine learning work well enough to structure teams and set expectations. Unrealistic timelines driven by software intuition. Software estimation is imprecise but roughly bounded. Machine learning development is fundamentally uncertain. Experiments fail. Models plateau. Data is messier than expected. A CTO without deep AI experience sets timelines based on software intuition, then reads missed deadlines as execution problems rather than inherent uncertainty. This creates a corrosive dynamic where the AI team feels perpetually behind, and leadership loses confidence in people who are doing difficult work competently. #### What AI-specific CTO competencies actually look like The standard checklist (model governance, evaluation methodology, data strategy) is accurate but insufficient. Those are table-stakes processes. The real competency is harder to codify. It is built from years of working with these systems. Intuition for how models interact with each other and with users. Modern AI products chain multiple models together. The way outputs propagate, where errors compound, how users probe the boundaries of model behaviour: these require deep, experiential understanding. No framework substitutes for having built and shipped systems where you learned, sometimes painfully, how these interactions play out. Comfort with semantic ambiguity. Anyone building AI products with language at their core needs leadership that understands ambiguity as a permanent feature of the domain, not a problem to be engineered away. This changes how you design evaluation criteria, how you scope product capabilities, how you communicate with stakeholders. An understanding of linguistics, even informal, stands a CTO in good stead here. Most software leaders do not have it. Problem framing as a primary technical skill. Translating a business need into a well-specified machine learning problem, and knowing when the translation is not clean, is the single highest-leverage CTO competency in an AI startup. This is where the research-to-production gap lives. A CTO who can look at a business objective and immediately see the assumptions, data requirements, failure modes, and evaluation criteria of the corresponding ML problem will save a startup months of misallocated effort. Calibrated uncertainty. Knowing what you do not know, what the model does not know, and what the data cannot tell you. Then communicating this to non-technical stakeholders without overselling certainty or drowning them in caveats. This is a leadership skill as much as a technical one, and it is conspicuously absent in CTOs whose entire career taught them that "it works or it doesn't" is a sufficient mental model. #### Fractional versus full-time: when each model fits Not every AI startup needs a full-time CTO from day one. Hiring one too early often produces the worst outcome: someone junior enough to afford but not experienced enough to make the foundational decisions correctly. A senior fractional CTO with real AI depth, engaged at the right intensity for the company's stage, will frequently outperform a less experienced full-time hire. The emerging CAIO (Chief AI Officer) role in larger organisations underscores this: AI leadership is increasingly recognised as a distinct discipline, not an extension of general technology management. Startups do not need two C-suite titles. They need one technical leader who already has AI depth baked in. Pre-seed and idea stage: a fractional AI CTO is almost always right. The company needs architecture decisions, problem framing, and technical strategy. It does not need forty hours a week of leadership. What it needs is someone senior enough to make the right calls on foundations that will shape the product's trajectory. Seed stage, building an MVP: the fractional model scales to two or three days per week. This is the critical period for getting foundations right. A fractional CTO at this intensity can set architecture, hire the first ML engineers, and establish development methodology. The alternative is a full-time CTO hire at a salary the company can barely afford. At Series A, non-founder CTOs at US startups command an average of $293,000 in cash compensation before benefits and equity (Kruze Consulting, based on payroll data from 250+ VC-backed startups). For an early-stage company, that is a heavy commitment. If the person does not have AI depth, it is a commitment to the wrong expertise. Series A, scaling the product: the decision depends on the product. If AI is the entire product rather than a feature, the company likely needs full-time AI-specific leadership by now. If AI is a core component but the scaling challenges are primarily infrastructure and engineering, the fractional model may still serve. Either way, a fractional CTO involved from earlier stages can help hire their full-time replacement and ensure continuity. Series B and beyond: full-time leadership is typically necessary. The organisation needs someone embedded in the daily rhythm of the company, managing teams, driving culture, handling the accumulation of technical decisions that compound over quarters. The critical point: the quality of early technical decisions matters more than the quantity of hours. A wrong architecture choice at seed stage will cost far more to fix than the premium rate of a senior fractional CTO who gets it right the first time. #### For investors: evaluating AI technical leadership in portfolio companies If you advise portfolio companies on technical leadership, the CTO question in AI-native startups deserves specific attention. The signals for strong general engineering leadership do not reliably predict strong AI product leadership. A few questions worth asking. How does the CTO frame the core ML problem the product is solving? If they describe it primarily in software terms (features, sprints, deployment pipelines) rather than ML terms (problem formulation, data strategy, evaluation methodology, model uncertainty), that is worth investigating. How do they respond when model performance plateaus? A CTO with AI depth will have a repertoire of diagnostic approaches. One without it will treat the plateau as an execution failure. How do they communicate technical uncertainty to the board? Overselling certainty is as much a red flag as being unable to articulate a path forward. At a portfolio level, if multiple companies are struggling with AI product delivery, the common factor is often a technical leadership gap. The teams may be talented. The problem may be that nobody in the leadership structure understands AI systems well enough to set realistic expectations, structure teams effectively, and make sound architectural decisions. For early-stage portfolio companies where the AI leadership cost does not justify a full-time hire but the AI leadership quality cannot be compromised, a fractional AI CTO bridges the gap. #### This is the work I do at Agathon The challenges described here are the problems I work on every day with founders and leadership teams building AI-native products. I hold a PhD in natural language processing from Cambridge and have spent over fifteen years building AI systems commercially across financial services, telecoms, automotive, and government. I have been the person making the architecture decisions, framing the ML problems, building the evaluation methodology, and shipping AI products into production. I have also conducted AI due diligence for investors assessing technical risk in their portfolios. When you engage Agathon, you work directly with me. I operate as a fractional CTO for AI-native startups, providing the depth of technical leadership that these products require. If you recognise the leadership gap described here, I would welcome a conversation about whether Agathon is the right fit. --- ### The Investor's Technical Due Diligence Playbook: What Most AI Assessments Actually Miss - URL: https://agathon.ai/insights/the-investors-technical-due-diligence-playbook-what-most-ai-assessments-actually-miss - Published: 2026-02-24 - Categories: AI Strategy, AI Consulting, Machine Learning Much technical due diligence on AI companies is performed at pace, by people with technical knowledge but without the time or incentive to find the dirty laundry. They report to deal teams who don't have the depth to interrogate their findings, and the result is a rubber stamp dressed up as rigour. I've seen it from both sides. Deeply technical teams spend weeks preparing detailed presentations on the ins and outs of how their technology works: the architecture decisions, the data pipelines, the model evaluation frameworks. Then a well-meaning, bright, but time-poor technical assessor arrives to evaluate the entire portfolio in a single day. The intricacies of a machine learning pipeline that took months to build get reduced to a throwaway bullet on someone's slide deck, glanced at for a few seconds before moving on. The assessment gets filed, the deal progresses, and six months post-close, the real picture emerges. This isn't because the people involved are incompetent. It's because the standard technical due diligence frameworks were built for traditional software, and AI introduces a fundamentally different category of risk that those frameworks weren't designed to catch. When your investment thesis rests on a company's AI capabilities, you need a different playbook entirely. #### The Problem With How AI Gets Evaluated The typical technical DD process covers sensible ground: architecture review, codebase quality, infrastructure scalability, security posture, team assessment, IP ownership. These things matter. But when applied to AI companies, they miss the questions that actually determine whether the technology is defensible, valuable, and real. Here's what I mean. A standard assessment might confirm that the codebase is well-structured, the infrastructure scales, and the team has relevant experience. All green lights. But it won't tell you whether the company's "proprietary AI" is a thin wrapper around OpenAI API calls with a custom prompt, something a competitor could replicate in a weekend. It won't tell you whether the training data was properly licensed, whether the model drifts without constant retraining, or whether the entire value proposition collapses when the next foundation model update changes the underlying capabilities. These aren't edge cases. They're the questions that determine whether you're investing in a technology company or a marketing story. #### A Framework That Actually Helps: The AI Investment Matrix To cut through the noise, I use a framework that maps two dimensions most DD processes evaluate separately but rarely connect: the strategic impact the AI targets (is it transforming a market, or reshaping one entirely?) against the technical depth of what's actually been built (is it genuinely novel technology, or a clever application of off-the-shelf capabilities?). This creates four quadrants, each with very different investment implications. ##### The Category Creators (High Market Impact × Deep Technical Depth) These companies target revolutionary market change and have built genuinely novel technology to deliver it: proprietary models trained on proprietary data, original architecture, defensible technical moats. This is where the biggest returns live, but also where evaluation is hardest. The team isn't just using AI; they're advancing it. What to look for: original research output from the team, proprietary training datasets with clear provenance, model performance that demonstrably exceeds what's achievable with publicly available tools, and, crucially, a cost structure and retraining pipeline that can sustain the advantage as the field moves. The risk here is that the moat erodes as foundation models improve. ##### The Emperor's New Clothes (High Market Impact × Shallow Technical Depth) This is where investors lose money. The company claims to be revolutionising a market, the pitch deck is compelling, the demo is impressive, but the underlying technology is a wrapper around existing capabilities that a well-resourced competitor could replicate in weeks. This is Builder.ai territory: a UK unicorn backed by Microsoft that claimed AI-automated software development. In reality, 700 human developers were doing the work. The company collapsed in 2025 with $37 million in frozen assets. Builder.ai isn't an outlier. An MMC Ventures study found that 40% of European startups identifying as AI companies had minimal actual AI integration. The SEC has now charged multiple firms for "AI washing": making misleading claims about AI capabilities to attract investment. This quadrant is where rigorous technical DD pays for itself many times over. ##### The Quiet Compounders (Moderate Market Impact × Deep Technical Depth) Often the best risk-adjusted investments in AI. These companies apply genuine technical depth (for example: real ML pipelines, proprietary data assets, sophisticated model architectures) to transform existing workflows rather than create entirely new markets. They're not headline-grabbing, but they're defensible. A competitor can't just spin up an API integration and match them. What makes these attractive is that the technical depth creates compounding advantages. Their models improve with usage data. Their training pipelines get more efficient. Their domain-specific performance widens the gap with generic alternatives. The key DD question is whether this compounding is real and sustainable, or whether the improvement curve is flattening. ##### The Commodity Trap (Moderate Market Impact × Shallow Technical Depth) Companies using off-the-shelf AI to improve existing processes. There's nothing wrong with this as a business as long as it is priced accordingly. The danger is when it's valued as a technology company with a defensible moat. If the core AI capability is an API call that every competitor can make, margins will compress as adoption spreads. Today's differentiator becomes tomorrow's table stakes. These companies can still be sound investments, but the thesis needs to rest on something other than the AI: distribution advantages, regulatory positioning, brand, network effects. The AI is an accelerant, not the engine. #### The Nuance Most People Miss: Wrappers Aren't Disqualifying Here's where I diverge from the lazy consensus that "wrapper equals bad." A company building on top of existing models isn't automatically a poor investment. Sometimes the smartest technical decision is to prove a concept using off-the-shelf capabilities while focusing engineering effort on the interaction layer: the points where human meets AI. What I've seen in practice is that the most important innovation often isn't in the model itself but in the interface patterns, the workflow integration, and the feedback loops that make AI genuinely useful rather than merely impressive in a demo. A team that deeply understands how users interact with AI outputs; where they need control, where they need transparency, where they need the system to take initiative, can build something far more valuable than a team with a technically superior model but a clunky user experience. The critical DD question isn't "did they build their own model?" It's "what happens to their defensibility as the technology matures?" If the value lives in a proprietary data flywheel, where user interactions generate training data that improves the system, which attracts more users, which generates more data, then a wrapper today can become a moat tomorrow. But if the interaction layer is thin and the AI is doing commodity work, there's nothing to compound. This distinction requires someone who understands both the technology and the product to evaluate. A pure technologist will dismiss the wrapper. A pure business evaluator won't know to ask about the data flywheel. You need both lenses simultaneously. #### Five Questions Your DD Process Should Actually Answer Forget the 85-point checklist. If your technical DD on an AI company doesn't definitively answer these five questions, it hasn't done its job. What actually happens when the system receives an input? Trace the full architecture from user action to AI output to delivered result. Is there a model running, or is there a human in the loop being obscured? Where does the intelligence actually live? This sounds basic, but I've seen impressive demos that, when you trace the pipeline, turn out to involve significant manual processing disguised as automation. Where does the training data come from, and what happens when it degrades? Data is the real asset in most AI companies, yet it's consistently the area that receives the least DD scrutiny. Who owns the data? How was it labelled? Is there a sustainable pipeline for new data, or is the company training on a fixed dataset that will become stale? What are the licensing implications if a data source changes terms? What would it cost to replicate this with off-the-shelf tools? Be honest about this. If a competent team with access to current foundation models and public datasets could rebuild the core capability in three months, the technology isn't the moat and so something else needs to be. That something else might be valid (data network effects, distribution, domain expertise embedded in the product), but you need to know. What's the cost structure at ten times current scale? AI infrastructure costs don't scale linearly, and they often scale in the wrong direction. Inference costs, retraining compute, data storage, and human oversight all have scaling characteristics that traditional software doesn't. A product that's margin-positive at current volume can become margin-negative at scale if the architecture isn't designed for it. What happens when the underlying foundation models change? If the product depends on a specific model's capabilities, what's the exposure when that model is deprecated, repriced, or superseded? Companies built on a single provider's API are carrying platform risk that should be priced into the deal. Those with model-agnostic architectures or proprietary models have a fundamentally different risk profile. #### What to Look for in the Team Technical DD typically asks whether the team is "experienced" and "capable." That's not enough for AI companies. The gap between a team that can fine-tune existing models and one that can do original ML research is an order of magnitude in capability and in the value they can create. You can gauge depth quickly if you know what to listen for. A CTO who genuinely understands their ML stack will talk about failure modes unprompted. They'll distinguish between types of accuracy, precision versus recall, and why the trade-off matters for their specific use case. They'll explain their train/test splits and, critically, go a level deeper: how the data was stratified, how the training set was assembled, what biases that assembly process might have introduced, and what that means for real-world performance on edge cases. A CTO who's learned the vocabulary but doesn't have the depth will talk about "accuracy" as a single number without qualification. They'll describe their model's performance in ideal conditions but go vague when asked about where it breaks. They'll reference their data without being able to articulate its provenance or limitations. This isn't a character flaw: many excellent technical leaders come from software engineering backgrounds where these questions don't arise. But if the investment thesis depends on AI capabilities, you need to know whether the person leading the technology genuinely understands the machinery or is managing it at arm's length. The composition matters too. A team heavy on software engineers but light on ML specialists can build a product but may struggle to deepen the technical moat. A team heavy on researchers but light on engineering may have impressive models but can't ship reliable production systems. The best AI companies have both and, ideally, someone who can bridge the gap between them. #### The Bridge Between Strategic and Technical The reason most AI due diligence fails is that the strategic assessment and the technical assessment happen in separate rooms, conducted by people who don't speak each other's language. The deal team evaluates market opportunity, competitive dynamics, and commercial traction. The technical team evaluates architecture, code quality, and infrastructure. Nobody connects the two. The most important questions live at the intersection. Is the technical depth sufficient to defend the strategic position? Does the market opportunity justify the technical investment required? Will the data flywheel that the commercial model depends on actually materialise given the current architecture? These aren't technology questions or business questions. They're both simultaneously. Getting this right requires someone who can read the code and read the market. In my experience, that's the rarest and most valuable capability in AI due diligence; and it's the one most deal teams don't think to look for until after the problems surface. --- ### Best Consulting Firms for Defining Your AI Strategy in 2026: An Honest Buyer's Guide - URL: https://agathon.ai/insights/best-consulting-firms-for-defining-your-ai-strategy-in-2026-an-honest-buyers-guide - Published: 2026-02-06 - Categories: AI Strategy, AI Consulting, Fractional CTO If you're searching for the best consulting firm to help define your AI strategy, you've probably already seen a dozen articles listing the same ten names in slightly different orders. Most were written by firms trying to rank themselves, stuffed with logos and "best for" labels that tell you remarkably little about which partner will actually help your organisation. This guide takes a different approach. Rather than ranking firms, it gives you a framework for choosing one — based on your company's size, AI maturity, budget, and what you actually need from the engagement. I'll cover the landscape honestly, including where different types of firms genuinely excel and where they reliably disappoint. And yes, I'll be transparent about where Agathon fits in that landscape too. The uncomfortable starting point: 40% of organisations making significant AI investments don't see business gains. That failure rate isn't a technology problem. It's a strategy problem. Getting the strategy right matters more than which model you deploy, and getting the right strategic partner matters more than most buyers realise. #### What "defining an AI strategy" actually means Before you evaluate firms, it's worth being precise about what you're buying. The AI consulting market conflates several very different services under the same label, and the confusion costs buyers real money. ##### AI strategy is not AI implementation AI strategy consulting means defining where and why your organisation should deploy AI, and in what sequence. The deliverables are a prioritised use-case portfolio, a data readiness assessment, a governance framework, an implementation roadmap with KPIs, and — critically — an honest appraisal of what you shouldn't bother with. AI implementation is building and deploying those systems. Different skill set, different engagement, often a different partner entirely. The problem is that many firms bundle these together, partly because it's more revenue and partly because the strategy conveniently recommends tools and platforms they happen to sell. When a consultancy's strategy engagement somehow always concludes that you need their proprietary platform, that's not strategy — it's a sales funnel with a consulting fee attached. ##### What a good AI strategy deliverable looks like If you've never commissioned one before, here's what you should expect from a competent strategy engagement: An AI readiness assessment that honestly evaluates your data infrastructure, technical capability, and organisational culture — not a tick-box exercise that tells you what you want to hear. A prioritised use-case portfolio that ranks opportunities by business impact and feasibility, explicitly identifying what to defer or abandon. A data strategy covering what data you have, what you need, what's missing, and what governance is required. An implementation roadmap with realistic timelines, resource requirements, and success metrics. And a governance framework covering responsible AI principles, risk management, and regulatory compliance — particularly important as the EU AI Act begins enforcement. If your strategy engagement delivers a slide deck full of AI buzzwords and a recommendation to "start a pilot," you've paid for a brochure, not a strategy. #### How to choose the right AI strategy consulting firm This is where most articles fail you. They list firms without helping you understand which type of firm matches your situation. The choice isn't about which firm is "best" in the abstract — it's about which model fits your needs. ##### The five factors that actually matter 1. Strategic depth versus technical breadth. Some firms are exceptional at high-level strategic thinking but couldn't build a production AI system if their partnership depended on it. Others can build anything but struggle to connect technical capability to business outcomes. The best AI strategy work requires both — the ability to think strategically about where AI creates value and ship the systems that capture it. Ask any prospective partner: what's the most complex AI system your team has actually built and deployed? 2. Industry-specific experience. AI in financial services looks nothing like AI in manufacturing. Regulatory constraints, data architectures, and organisational cultures vary enormously. A firm that's brilliant in retail may be mediocre in healthcare. Don't accept "we work across industries" as a credential — press for case studies in your sector specifically. 3. Engagement model and team composition. At large firms, the partner who wins your business rarely does the work. You'll present to a seasoned strategist and then be handed to a team of recent graduates. Ask directly: who will actually do the analysis and write the deliverables? If the answer involves more than one degree of separation from the person in the pitch meeting, factor that into your evaluation. 4. Vendor neutrality. If a consulting firm has commercial partnerships with specific technology vendors — and most large firms do — their strategy recommendations are structurally biased. This doesn't make them incompetent, but it does mean their "independent assessment" of which cloud platform or AI tooling to adopt comes with an asterisk. Ask about commercial relationships with technology providers and how they manage conflicts of interest. 5. Knowledge transfer and capability building. The best AI strategy engagements make the consulting firm progressively less necessary. The worst create dependency. Ask what capability your team will have after the engagement that they didn't have before. If the answer is "they'll be able to use our platform," that's dependency, not transfer. A good partner should be building your organisation's AI aptitude, not hoarding expertise to justify ongoing retainers. ##### Match your needs to the right type of firm The AI consulting landscape breaks into three broad categories, each suited to different buyers. This isn't a value judgement. A FTSE 100 company navigating enterprise-wide AI transformation probably needs McKinsey's weight. A 200-person company looking to identify its first three AI use cases probably doesn't — and would be overpaying for a brand name whilst receiving work done by people with less experience than the boutique alternative. #### The AI strategy consulting landscape in 2026 Rather than a ranked list, here's an honest assessment of where different firms sit — what they're genuinely good at and where they fall short. ##### Global consultancies: McKinsey, BCG, Deloitte, Accenture, PwC, EY What they do well. These firms bring unmatched convening power. When you need to align a 50-person C-suite around an AI vision, the McKinsey or BCG brand carries weight that no boutique can replicate. Their research arms (QuantumBlack, BCG X/GAMMA) produce genuinely valuable market intelligence. They have deep benches across industries and geographies, and their frameworks — like BCG's 10-20-70 model emphasising that 70% of AI transformation is people and process, not algorithms — reflect real wisdom about organisational change. Where they fall short. The strategy-implementation gap is structural. The partner who presents your AI strategy has likely never trained a model or debugged a data pipeline. As the AI consulting market matures, the gap between firms that can recommend agentic workflows and those that can build them is widening. Many large firm AI strategies recommend architectures their own teams can't deliver, requiring a second firm for implementation — which raises obvious questions about the strategy's technical realism. Pricing reflects brand premium as much as delivery quality. Enterprise AI strategy engagements routinely cost $200k-$500k, with implementation running into the millions. For genuinely complex, global transformations, this may be justified. For mid-market companies, it rarely is. Best for: Enterprise organisations (1,000+ employees) needing board-level credibility, global coordination, and deep organisational change management alongside AI strategy. You're paying for the brand, the bench depth, and the ability to manage political complexity across business units. ##### Specialist AI boutiques This is the category that's grown most dramatically since 2024. Firms like Neurons Lab, Binariks, and LeewayHertz — alongside Agathon — offer deep AI-specific expertise without the overhead and vendor entanglements of global firms. What they do well. Technical credibility. The people defining your strategy have typically built production AI systems themselves, which means the strategy is grounded in what's actually achievable rather than what looks good in a slide deck. Engagement models are leaner, senior people do the work directly, and recommendations tend to be vendor-neutral because boutiques don't have billion-pound technology partnerships to protect. The execution gap in AI consulting (the distance between a strategist who can talk about AI and a practitioner who can build it) is smallest in this category. When your strategy consultant has personally shipped production ML systems, the roadmap they write is materially more realistic. Where they fall short. Limited scale. If you need 30 consultants across four countries simultaneously, a boutique can't provide that. Some boutiques are stronger on implementation than strategy, or vice versa: you need to verify which. And brand recognition is lower, which can matter if you need to sell the AI strategy internally to a sceptical board that trusts Big 4 names. Best for: Mid-market companies (50-1,000 employees) that need genuine technical depth combined with strategic thinking, and want senior people doing the work rather than supervising it. Also excellent for enterprises that already have a high-level strategy from a global firm and need technically credible partners to pressure-test or refine it. ##### Independent and fractional AI strategists Individual practitioners — often former heads of AI at enterprises, PhD researchers, or ex-Big 4 senior consultants — offering direct access to senior expertise without firm overhead. Platforms like Toptal and Clutch aggregate some of these, but many operate through direct relationships. What they do well. Maximum expertise-per-pound. You're paying for one very senior person's full attention, with no markup for offices, junior staff, or brand management. Engagements are fast and lean. The fractional CTO model is particularly effective for companies that need ongoing strategic AI leadership without a full-time hire. Where they fall short. Individual capacity constraints. One person can't deliver a comprehensive enterprise strategy across multiple business units simultaneously. Quality is highly variable — credentials matter more here than in any other category, because there's no firm reputation providing quality assurance. Due diligence on the individual's actual track record is essential. Best for: SMEs, startups, and companies with budgets under £50k that need a realistic AI roadmap from someone who's actually done it. Also effective as an ongoing advisory relationship for founders or CTOs who want a sparring partner on AI decisions. #### How much does AI strategy consulting actually cost? Pricing in this market is notoriously opaque. Here's what you should expect to pay in 2026, based on engagement type. AI readiness assessment and strategy audit: $10k-$75k. This covers evaluating your current data infrastructure, identifying high-value use cases, and producing a prioritised roadmap. At the lower end, you'll get a focused assessment from a boutique or independent consultant. At the higher end, a global firm with a larger team and more extensive stakeholder interviews. Comprehensive AI strategy and implementation roadmap: $50k-$300k. This is the full engagement: readiness assessment, use-case identification, data strategy, governance framework, implementation roadmap, and often a pilot project to validate the highest-priority use case. Global firms sit at the top of this range; boutiques in the middle; independents at the lower end. Ongoing strategic advisory (retainer): $2.5k-$15k per month. Retained access to a senior AI strategist for ongoing guidance, decision support, and strategy refinement as implementation progresses. Increasingly popular as organisations recognise that AI strategy isn't a one-off exercise. Independent consultant (day rate): $800-$3,000+ per day. For focused engagements — a two-day strategy workshop, a week-long technical assessment, or ongoing fractional advisory. ##### What drives the price; and what shouldn't Legitimate cost drivers include the complexity of your organisation, the number of business units in scope, regulatory requirements, and the seniority of the consultants involved. Things that shouldn't drive cost but often do: brand premium (you're paying 40-60% more for a Big 4 logo), unnecessary discovery phases that could be replaced by focused workshops, and scope creep from strategy into implementation without a clear boundary. A useful benchmark: if your strategy engagement costs more than 10% of your likely first-year AI implementation budget, you're probably overpaying for strategy relative to execution. #### Ten questions to ask before hiring an AI strategy consultant These will tell you more in a 30-minute conversation than any amount of website research. 1. Who specifically will do the work? Not who will present, not who will supervise — who will analyse our data landscape and write the strategy document? 1. What AI systems has your team actually built and deployed in production? Strategy grounded in implementation experience is categorically better than strategy from people who've only ever written recommendations. 1. What technology vendor relationships do you have, and how do they influence your recommendations? Any hesitation here is informative. 1. Can you show me a redacted strategy deliverable from a comparable engagement? You're buying a deliverable. You should see what it looks like before committing. 1. What will our team be able to do after this engagement that they can't do now? The answer reveals whether you're buying capability transfer or dependency creation. 1. How do you handle it when your assessment concludes that AI isn't the right solution for a use case?Consultants who always recommend more AI aren't being strategic — they're selling. 1. What's your approach to data readiness, and what happens if our data isn't ready? A honest answer here saves you from a strategy that assumes infrastructure you don't have. 1. How many rounds of revision are included, and what's the approval process? Scope clarity prevents cost overruns. 1. What's your experience in our specific industry and regulatory environment? Generic AI expertise isn't enough. Press for specifics. 1. What does your pricing include, and what's billed separately? Travel, tools, junior staff time, and follow-up support are common areas where stated prices expand. #### Red flags when evaluating AI strategy firms Having been on both sides of AI consulting engagements, both as a buyer in enterprise environments and now as a provider, these are the warning signs I'd want any buyer to recognise. They lead with technology, not business outcomes. If the first conversation is about which LLM to use rather than what business problems you're trying to solve, the priorities are wrong. Technology selection should follow strategy, not precede it. They can't show strategy-specific case studies. Many firms have impressive implementation portfolios but have never delivered a standalone strategy engagement. Building a chatbot and defining an AI strategy are fundamentally different exercises. They resist fixed-scope proposals. "We'll need to do discovery before we can scope this" is sometimes legitimate and sometimes a mechanism for billing exploratory hours before you've committed to anything. A confident firm can scope a strategy engagement from a well-structured brief. There's no knowledge transfer plan. If the engagement ends with a deliverable but your team doesn't understand the reasoning behind the recommendations, you'll be calling the same firm back every time a decision needs making. That's not partnership; it's rent-seeking. They promise guaranteed ROI. AI strategy consulting can dramatically improve the odds of successful AI adoption, but no honest consultant guarantees specific returns before understanding your organisation's constraints. Promises of guaranteed ROI correlate strongly with disappointments. #### Where Agathon fits — and where we don't I'd be dishonest if I wrote a buyer's guide and pretended my own firm doesn't exist. So here's where Agathon genuinely adds value, and where you'd be better served elsewhere. Where we're strong. Agathon sits in the specialist boutique category, with a particular emphasis on combining research-grade technical depth with commercial pragmatism. I hold a PhD in Natural Language Processing from Cambridge with prior mathematics and computer science training from Oxford, and I've spent over fifteen years building and deploying AI systems in financial services, telco, and automotive. When I write an AI strategy, the roadmap reflects what's actually buildable specifically because I've built these systems myself. We're especially effective for organisations that need someone who can see the full technical potential of AI in their context, not just recommend obvious use cases. Our AI Leadership Advisory service is designed specifically for defining AI strategy with ongoing support, and our AI Readiness Quiz gives you a rapid self-assessment before engaging any consultant. Where we're not the right fit. If you're a multinational needing 20+ consultants across multiple geographies simultaneously, we can't provide that scale. If you need a brand name that your board will recognise from the Financial Times, a Big 4 firm serves that political function better. And if your primary need is large-scale implementation rather than strategy, our strength is in defining what to build and pioneering technically sophisticated products — not staffing a 30-person delivery team. #### Getting started Before engaging any firm, invest an hour in honest self-assessment. Consider where you sit on the AI maturity spectrum: are you exploring AI for the first time, or refining an existing strategy that hasn't delivered? Be realistic about your data infrastructure and your organisation's appetite for change. And be clear about your budget: not just for the strategy engagement, but for the implementation that follows. The right AI strategy partner isn't the most famous or the most expensive. It's the one whose expertise matches your situation, whose engagement model fits your organisation, and whose recommendations you'll trust enough to actually execute. If you'd like to explore whether Agathon is that partner for your specific situation, start a conversation. If we're not the right fit, I'll tell you — and point you toward who is. --- #### Frequently asked questions What does an AI strategy consultant actually do? An AI strategy consultant evaluates your organisation's readiness for AI adoption, identifies the highest-value use cases, creates a prioritised implementation roadmap, and establishes governance frameworks. The best consultants also assess your data infrastructure, recommend organisational changes needed for AI success, and build your team's capability to make AI decisions independently. How long does it take to define an AI strategy? A focused strategy engagement typically takes 4–12 weeks, depending on organisational complexity. A rapid assessment for a single business unit might take 2–4 weeks. Enterprise-wide strategies spanning multiple geographies can take 3–6 months. Be wary of engagements that extend beyond six months; at that point, the market will have moved and parts of your strategy may already be outdated. What's the difference between AI consulting and AI development? AI consulting provides strategic guidance on where and how to deploy AI. AI development builds and deploys the actual systems. Some firms offer both; many specialise in one or the other. Understanding this distinction prevents you from hiring strategists who can't build, or builders who can't think strategically about business value. Do I need AI strategy consulting if I already have a data team? Often, yes. Having a capable data team is valuable but doesn't automatically translate to strategic clarity about where AI creates the most business value. Data teams tend to optimise for technical sophistication; strategy consulting optimises for business impact. The most effective engagements combine external strategic perspective with internal technical knowledge. Can a small business afford AI strategy consulting? Yes, through independent consultants or boutique firms offering focused engagements. A meaningful AI strategy for a small business doesn't need to cost six figures. A well-scoped two-day workshop with an experienced practitioner ($3k–$6k) can produce a prioritised roadmap that saves months of misdirected effort. Start with our AI Readiness Quiz for a free self-assessment. Should I choose a big consulting firm or a specialist boutique? It depends on your organisation's size, budget, and what you value most. Global firms offer brand credibility, large teams, and organisational change expertise. Boutiques offer deeper technical expertise, senior direct engagement, and typically stronger vendor neutrality. Most mid-market companies get better value from boutiques; enterprises with complex political landscapes often benefit from the convening power of global brands. The decision matrix earlier in this article maps these trade-offs in detail. --- ### The key metrics to measure the ROI of your LLM deployments - URL: https://agathon.ai/insights/the-key-metrics-to-measure-the-roi-of-your-llm-deployments - Published: 2026-02-06 - Categories: LLMs, AI Strategy, AI Consulting Everyone measures LLM performance wrong. They count tokens, track latency, and celebrate when their chatbot doesn't hallucinate. Meanwhile, they're sitting on a Ferrari engine and measuring its value by how well it idles. The real tragedy? Most organisations deploy LLMs like they're building traditional software—obsessing over response times and error rates whilst missing the fundamental shift in how value creation works when language becomes computational. You're not shipping code anymore. You're deploying cognitive infrastructure. #### Why traditional software metrics fail spectacularly for language models Software metrics assume deterministic systems. Push button, get result. Measure speed, count errors, ship update. LLMs operate in a fundamentally different paradigm—they're probabilistic reasoning engines masquerading as text generators. ##### The latency paradox: when slower means better Here's what Microsoft Research discovered that should terrify every CTO measuring success by response time: in translation tasks, a 10x increase in model compute actually increased task completion time for high-skilled workers whilst dramatically improving quality. The paradox? Better models think longer because they're doing more sophisticated reasoning. Think about that. Your fastest responses might be your worst ones. That sub-second chatbot response you're celebrating? It's probably giving you the linguistic equivalent of a knee-jerk reaction when you need strategic analysis. ##### Token economics versus actual business value Everyone tracks tokens like they're measuring electricity usage. But research shows the relationship between token consumption and value delivery is non-linear and context-dependent. A 100-token response that solves a complex problem delivers orders of magnitude more value than a 1000-token response that misses the point. The Microsoft framework reveals that prompt tokens and completion tokens have entirely different value profiles. Prompt tokens represent investment in context and precision. Completion tokens represent output volume. Most companies optimise for minimising both—essentially starving their AI of context whilst demanding brevity. It's like hiring a consultant and giving them five minutes to understand your business. ##### The hallucination tax nobody talks about Azure OpenAI's content filtering metrics expose an uncomfortable reality: the real cost of hallucinations isn't in the false information—it's in the defensive infrastructure you build around it. Every content filter, every validation layer, every human review checkpoint represents a tax on your system's potential. Research shows that responses filtered for safety reasons (measured as "finish_reason": "content_filter") correlate with overly conservative deployments. You're not just preventing harmful outputs; you're throttling legitimate value creation. #### Technical performance metrics that actually matter Forget perplexity scores and BLEU metrics. Real-world LLM performance lives in the messy intersection of semantic understanding and business constraints. ##### Beyond perplexity: measuring semantic coherence in production Academic benchmarks measure how well models predict the next token. Production systems need to measure whether those tokens form coherent strategic thoughts. Researchers evaluating business process automation found that semantic understanding of activities—not token prediction accuracy—determined whether LLMs could identify value-adding versus non-value-adding steps. The key insight: measure meaning preservation across transformations, not just output similarity. When your LLM breaks down a complex activity into steps, does it maintain the semantic intent? That's what separates sophisticated implementations from expensive autocomplete. ##### Response quality scoring at scale The translation productivity experiments revealed something crucial: quality improvements from better models compound non-linearly. A 10x increase in model compute improved grades by 0.18 standard deviations—but this translated to a 29.7% increase in earnings per task when quality bonuses were included. This means your quality metrics need to capture value multiplication, not just error reduction. Are you measuring how much better decisions become, or just counting mistakes? ##### Context window utilisation and its hidden costs Here's what nobody tells you about context windows: filling them is easy, using them effectively is hard. The research on value-added analysis shows that structured prompting with role descriptions, guidelines, and examples dramatically outperformed simple context dumping. Measure semantic density, not token count. A well-structured 1,000-token prompt outperforms a 10,000-token information dump. Your context window is prime real estate—are you building skyscrapers or parking lots? ##### Model drift detection in the wild Static benchmarks tell you nothing about performance degradation in production. The enterprise evaluation challenges identified by researchers include dynamic, long-horizon interactions where model behaviour shifts over time. Track semantic consistency across conversation turns. Monitor when models start contradicting earlier statements or losing track of established context. This drift often appears before traditional error metrics spike. #### Business value indicators: where rubber meets road Stop measuring what's easy. Start measuring what matters. The research consistently shows that productivity gains from LLMs don't follow traditional software improvement patterns. ##### Time-to-insight reduction across knowledge work The experimental evidence is stark: LLMs create a 12.3% speed improvement per 10x increase in compute, but the real story is in the distribution. Low-skilled workers saw 21.1% improvements whilst high-skilled workers saw only 4.9%. This isn't about faster typing; it's about cognitive load transfer. Measure how quickly your teams reach actionable insights, not how fast they get responses. ##### Decision velocity improvements Business process analysis research demonstrates that LLMs can classify activities as value-adding, business-value-adding, or non-value-adding with remarkable accuracy when properly structured. But the value isn't in the classification—it's in the acceleration of decision-making. Track decision cycle time, not response time. How much faster are strategic choices being made with AI assistance versus without? ##### Automation rate versus human oversight burden Here's the uncomfortable truth from the research: the percentage of prompts returning HTTP 400 errors and responses filtered for content directly correlates with increased human oversight requirements. Every safety measure creates a manual review burden. Measure the true automation rate: tasks completed without human intervention divided by total tasks attempted. Most "automated" systems are just faster ways to create work for humans. ##### Error cascade prevention and recovery costs When an LLM makes an error in step one of a multi-step process, that error compounds through every subsequent step. Research on activity breakdown shows that errors in decomposition lead to fundamental misclassification of value. Track error propagation rates and recovery costs. One hallucination in a planning phase can invalidate hours of downstream work. #### The human factors everyone ignores LLMs don't operate in isolation. They're cognitive prosthetics for human intelligence. Yet most metrics pretend humans don't exist. ##### Cognitive load transfer metrics The translation experiments revealed something profound: LLMs don't just speed up work—they fundamentally change its cognitive structure. Translators using advanced models shifted from word-level translation to semantic-level review. Measure cognitive load redistribution. Are your knowledge workers doing higher-value thinking, or are they just babysitting AI outputs? ##### Trust calibration and user confidence scoring Research participants rated their familiarity with AI tools at 4.15/5 but their actual performance varied wildly based on model quality. Users can't calibrate trust without understanding model capabilities. Track the correlation between user confidence and actual output quality. Overconfidence in weak models is more dangerous than scepticism about strong ones. ##### Workflow integration friction coefficients The experiments used shadow testing—running new models in parallel with existing workflows—to measure integration friction without disrupting operations. This revealed that workflow changes often matter more than model improvements. Measure adaptation overhead: time spent learning new patterns, adjusting prompts, and recalibrating expectations. A 50% better model that requires 100% workflow restructuring might deliver negative value. ##### Skills displacement versus augmentation ratios The 4x difference in productivity gains between low and high-skilled workers reveals an uncomfortable truth: LLMs don't augment everyone equally. Some skills become more valuable, others become obsolete. Track skill evolution patterns. Which capabilities are being enhanced versus replaced? Your metrics should capture this transformation, not just productivity changes. #### Operational efficiency beyond the hype The real costs of LLM deployment hide in operational complexity. Token prices are just the tip of the iceberg. ##### Real compute costs versus promised savings Microsoft's framework includes detailed GPU utilisation metrics, tracking not just token consumption but 429 error responses indicating system overload. These "hidden" failures represent capacity constraints that destroy user experience. Calculate true cost per successful outcome, including retries, failures, and overhead. That chatbot might cost pennies per response but pounds per problem solved. ##### Fine-tuning ROI: the emperor's new clothes Here's what the research actually shows: structured prompting with zero-shot approaches often outperforms expensive fine-tuning. The business process analysis achieved remarkable results using carefully crafted prompts rather than model customisation. Measure comparative advantage: performance gain from fine-tuning divided by its total cost (including maintenance, versioning, and technical debt). Most fine-tuning delivers negative ROI when fully accounted. ##### Prompt engineering overhead and technical debt The research identified optimal prompt components through systematic grid search—but this optimisation process itself represents significant overhead. Every prompt is code that needs maintenance. Track prompt complexity growth over time. As edge cases accumulate, prompts become byzantine rule engines. That "simple" prompt template will eventually become your most complex codebase. ##### Infrastructure scaling efficiency breakpoints The scaling laws research reveals non-linear relationships between compute and performance. A 10x increase in compute delivers diminishing returns—but these returns compound differently across use cases. Identify your efficiency cliffs: where does additional compute stop delivering proportional value? Most organisations operate far below or far above optimal efficiency points. #### Risk-adjusted returns in the age of AI Value without risk assessment is gambling. The research consistently highlights risks that traditional metrics miss entirely. ##### Compliance cost multipliers Azure OpenAI's content filtering reveals that 400-series errors and filtered responses aren't just technical failures—they're compliance events. Each filtered response might represent a regulatory near-miss. Calculate compliance overhead ratios: cost of compliance infrastructure divided by operational costs. Some use cases require 10x compliance investment for 1x operational deployment. ##### Reputation risk quantification The research on responsible AI evaluation shows that harm, toxicity, and bias aren't binary—they exist on spectrums that shift with context. A helpful response in one culture might be offensive in another. Develop context-sensitive risk scores. What's acceptable in internal tools might be catastrophic in customer-facing systems. ##### Data leakage prevention metrics Enterprise evaluation challenges include role-based access control and data sovereignty. Every prompt potentially leaks sensitive information; every response might violate data residency requirements. Track information flow patterns. Where does sensitive data travel in your LLM pipeline? Most breaches happen in logging and monitoring, not primary processing. ##### Ethical debt accumulation rates Like technical debt, ethical debt compounds. Each decision to prioritise speed over safety, each shortcut in bias testing, accumulates risk that eventually demands payment. Measure ethical debt velocity: the rate at which questionable decisions accumulate versus the rate at which they're addressed. High velocity predicts future crisis. #### Building your measurement framework Stop copying Silicon Valley metrics. Build measurement systems that reflect your actual value creation, not their venture capital narratives. ##### Establishing baseline performance before LLM adoption The translation research established baseline performance through control groups completing tasks without AI assistance. This revealed that productivity gains varied 4x between skill levels. Create true baselines: measure current performance without AI, not just with your existing tools. You can't measure improvement without understanding your starting point. ##### Creating composite metrics that reflect reality Single metrics lie. The research consistently uses composite measures: earnings per minute (combining speed and quality), semantic coherence (combining multiple linguistic properties), and value-added classification (combining customer and business perspectives). Design metrics that capture value complexity. Revenue per conversation, not responses per second. Problems solved per pound spent, not tokens per penny. ##### Balancing leading and lagging indicators Task completion is a lagging indicator—it tells you what happened. Prompt quality and context utilisation are leading indicators—they predict what will happen. Build predictive metric models. Which early signals correlate with later success? The research shows that prompt structure quality predicts task success better than model size. ##### Avoiding vanity metrics and theatre Tokens processed, conversations handled, queries answered—these are vanity metrics. They make impressive dashboards but reveal nothing about value creation. Focus on metrics that hurt when they're bad. If a metric dropping doesn't cause immediate concern, it's probably theatre. #### The path forward: honest conversations about LLM economics Here's the truth nobody wants to admit: most LLM deployments destroy value. They're expensive ways to do things badly that humans did well. But the minority that succeed—those that exploit full technical potential rather than implementing basic features—they're transforming entire industries. The research makes this crystal clear. When Yale economists ran controlled experiments with professional translators, they didn't find uniform improvement. They found revolution for some and evolution for others. The difference? Understanding which capabilities to exploit and how to measure their true impact. Microsoft Research didn't create another chatbot framework. They built a comprehensive measurement system that captures cost, risk, performance, and value in their full complexity. They recognised that LLMs aren't faster databases or smarter search engines—they're cognitive infrastructure that demands new thinking about value creation. The business process researchers didn't automate tasks—they automated understanding. Their LLMs don't just classify activities; they reveal the semantic structure of value creation itself. This is the difference between using 10% of potential and exploiting capabilities others don't even know exist. Your metrics reveal your ambitions. If you're measuring response times and token costs, you're building commodity tools. If you're measuring semantic coherence, value multiplication, and cognitive load transfer, you're building the future. The uncomfortable truth about measuring LLM success isn't that it's hard—it's that doing it properly forces you to confront how little of the technology's potential you're actually using. Most organisations are driving Formula One cars in school zones, then wondering why they're not winning races. If you're ready to build AI solutions that exploit full technical potential rather than implementing basic features, you should contact us today. #### References - Microsoft Research framework for LLM evaluation metrics including costs, customer risk and user value quantification - ArXiv research on automated business process analysis using LLMs for value assessment - ArXiv study examining the real-world business benefits and limitations of LLMs in professional settings - ArXiv experimental research on scaling laws for economic productivity gains from LLM assistance - ArXiv comprehensive survey on LLM evaluation methods and benchmarking approaches - ArXiv survey on LLM agent evaluation including enterprise-specific challenges and compliance metrics --- ### EU AI Act Explained: Compliance Requirements and Business Impact - URL: https://agathon.ai/insights/eu-ai-act-explained-compliance-requirements-and-business-impact - Published: 2026-01-30 - Categories: AI Strategy, Responsible AI, AI Consulting #### The AI regulation nobody saw coming (until it was everywhere) Brussels wrote the future of AI while Silicon Valley was still arguing about chatbot safety. The EU AI Act, which entered force in August 2024, represents the most comprehensive AI regulatory framework globally - and most organisations remain blissfully unaware of its extraterritorial reach. Your AI system deployed in San Francisco? If its outputs touch EU citizens, you're in scope. That innocuous customer service bot? Potentially high-risk under the Act's classification system. The regulation's genius lies in its risk-based architecture rather than technology-specific rules. While competitors scramble to understand basic compliance requirements, sophisticated organisations recognise the Act as a forcing function for architectural excellence. The companies that will dominate the next decade of AI aren't those with the largest models - they're those who build systems that naturally align with regulatory frameworks while exploiting capabilities others consider too complex to govern. #### Understanding the EU AI Act's risk-based architecture ##### The four-tier risk pyramid that changes everything The Act's classification system operates through four distinct tiers: prohibited systems, high-risk systems, limited-risk systems, and minimal-risk applications. This isn't merely bureaucratic categorisation - it fundamentally restructures how AI capabilities can be deployed in production environments. Prohibited practices include subliminal manipulation techniques, vulnerability exploitation systems, social scoring mechanisms, and most forms of real-time biometric identification in public spaces. The prohibition extends to emotion recognition in workplaces and educational institutions - a significant constraint for organisations building advanced sentiment analysis capabilities. High-risk systems span eight critical domains: biometrics, critical infrastructure, education and training, employment management, essential services, law enforcement, immigration control, and democratic processes. The classification triggers when systems materially influence decision-making affecting fundamental rights. A recruitment screening algorithm? High-risk. An infrastructure monitoring system for water supply? High-risk. The breadth catches systems most organisations haven't considered regulatory targets. Limited-risk systems primarily face transparency obligations: users must know they're interacting with AI. Minimal-risk applications operate largely unconstrained but still require consideration of AI literacy provisions. ##### Why your chatbot might be high-risk (and why that matters) Most organisations assume their customer service chatbots fall into limited-risk categories. They're wrong. The moment your chatbot influences access to essential services - healthcare appointment scheduling, benefit applications, financial service access - it potentially triggers high-risk classification. Consider a healthcare provider's appointment booking system. Research shows how AI systems in healthcare contexts face particular scrutiny under the Act. If your chatbot can deny appointments based on algorithmic decisions, you've entered high-risk territory. The system now requires conformity assessments, technical documentation, human oversight mechanisms, and continuous postmarket monitoring. The classification isn't about the technology - it's about impact. A GPT-5 powered system answering general queries remains limited-risk. The same model making triage decisions becomes high-risk. Architecture decisions made today determine regulatory burden for years. ##### Prohibited AI systems: the absolute no-go zones The Act's prohibited systems list reveals European regulators' fundamental concerns about AI deployment. Beyond obvious restrictions on subliminal manipulation and social scoring, the prohibitions expose deeper architectural constraints. Biometric categorisation systems that infer protected characteristics - race, political beliefs, sexual orientation - face blanket prohibition. This extends beyond facial recognition to any system attempting such inference from biometric data. Voice analysis determining political affiliation? Prohibited. Gait recognition inferring religious beliefs? Banned. The workplace emotion recognition prohibition particularly impacts employee monitoring systems. Organisations deploying sentiment analysis on internal communications, productivity monitoring with emotional state inference, or interview assessment tools reading micro-expressions must fundamentally redesign these capabilities. #### Compliance requirements that will reshape your AI roadmap ##### Mandatory conformity assessments and CE marking High-risk AI systems require conformity assessments before market placement - a process most software organisations have never encountered. Unlike self-certification for GDPR compliance, conformity assessments involve documented evaluation against harmonised standards that may not exist until December 2027. The assessment examines your quality management system, technical documentation, and risk management processes. Successfully assessed systems receive CE marking - mandatory for EU market access. The Digital Omnibus proposal extends implementation timelines, but organisations starting assessment preparation now gain competitive advantage when standards crystallise. For systems embedded in regulated products - medical devices, vehicles, machinery - conformity assessment follows existing product regulations. Standalone AI systems face new assessment procedures the Act establishes. The difference fundamentally impacts go-to-market strategies. ##### Technical documentation and transparency obligations Documentation requirements exceed anything currently standard in AI development. Before deployment, high-risk systems need comprehensive technical documentation covering system architecture, development processes, data governance, performance metrics, risk management, and change tracking. The Act mandates specific documentation elements: intended purpose descriptions, hardware/software interaction specifications, development methodology explanations, training data provenance and characteristics, data processing procedures including outlier detection, monitoring and control mechanisms, relevant performance metrics, and postmarket monitoring plans. This isn't documentation for documentation's sake. The requirements force architectural decisions that enable explainability, traceability, and accountability. Systems designed with these requirements integrated from conception operate more robustly than those retrofitted for compliance. ##### Human oversight and accuracy thresholds Human oversight provisions require more than token human-in-the-loop implementations. The Act demands oversight mechanisms that enable understanding system limitations, prevent automation bias, allow output interpretation, permit system override, and provide intervention capabilities. These requirements fundamentally challenge fully automated decision-making architectures. Your high-risk system needs designed-in override mechanisms, interpretability features, and graceful degradation paths when humans intervene. Bolted-on oversight fails both regulatory scrutiny and operational requirements. Accuracy thresholds remain undefined in the Act itself, delegating to sector-specific standards. However, the requirement for declared performance metrics and continuous monitoring establishes a framework for accuracy accountability. Systems must document expected performance, measure actual performance, and maintain performance above declared thresholds. ##### Data governance and training set requirements Data governance extends beyond privacy compliance to encompass quality, representativeness, and bias mitigation. The Act requires governance covering training, validation, and testing datasets - each with distinct requirements. Quality dimensions include accuracy, completeness, coverage, conformity, consistency, lack of duplication, relational integrity, timeliness, and uniqueness. Each dimension requires measurement, monitoring, and maintenance processes. Organisations accustomed to "good enough" training data face fundamental process changes. Bias detection and mitigation obligations extend through the entire data lifecycle. This includes initial dataset assessment, preprocessing bias introduction, model training amplification, and deployment drift. Systems using reinforcement learning from human feedback or retrieval augmented generation face particular scrutiny for feedback loops creating biased outputs. #### Business impact beyond the obvious penalties ##### Product development timelines under the new regime Traditional AI product development cycles - rapid prototyping, iterative deployment, production learning - conflict with Act requirements. High-risk systems need conformity assessment before deployment. Technical documentation must exist before market placement. Risk assessments precede development decisions. The Digital Omnibus proposal's timeline extensions provide breathing room but don't eliminate the fundamental shift. Products launching after August 2026 (potentially December 2027 with extensions) need compliance built into development processes. Retrofitting compliance onto existing systems proves more expensive than designed-in compliance. Smart organisations are restructuring development processes now. They're implementing documentation practices, establishing risk assessment frameworks, and building oversight mechanisms into system architectures. When competitors scramble for compliance, these organisations will already operate within the framework. ##### The competitive advantage of early compliance Compliance creates competitive moats. While others struggle with retrofitting, compliant organisations capture market share through trust differentiation. The Act's public database for high-risk systems becomes a marketing asset - verified compliance visible to all potential customers. Early compliance also shapes standards development. Organisations demonstrating viable compliance approaches influence harmonised standards creation. Your implementation becomes the template others must follow. This first-mover advantage in regulated markets historically produces dominant market positions. Financial services organisations already leverage AI Act compliance for competitive positioning. Banks demonstrating robust AI governance attract customers concerned about algorithmic discrimination. Insurance companies with transparent AI decision-making reduce regulatory scrutiny while building trust. ##### Cross-border implications for global operations The Act's extraterritorial reach means global organisations can't isolate EU compliance. Systems whose outputs affect EU citizens fall under the Act regardless of deployment location. This creates three strategic options: global compliance adoption, market segmentation, or EU market abandonment. Global compliance adoption (building all systems to EU standards) simplifies operations but potentially over-constrains non-EU deployments. Market segmentation (maintaining separate EU-compliant and non-compliant systems) increases complexity and maintenance costs. EU market abandonment cedes significant market opportunity to competitors. Most sophisticated organisations adopt hybrid approaches: core platform compliance with market-specific configurations. This requires architectural decisions enabling compliance toggling without fundamental system changes. Get this wrong, and you're maintaining multiple codebases. Get it right, and you've built a globally deployable platform. ##### Insurance, liability, and the shifting risk landscape The Act fundamentally restructures AI liability landscapes. Providers of high-risk systems face strict documentation and performance obligations. Serious incidents require reporting within 15 days. Non-compliance penalties reach €35 million or 7% of global turnover. Insurance markets haven't caught up. Current professional liability policies rarely cover AI-specific risks adequately. Directors and officers insurance may not extend to AI governance failures. Product liability insurance struggles with AI's evolutionary nature. Forward-thinking organisations are negotiating bespoke AI insurance coverage now, before markets harden. They're also restructuring vendor agreements to clarify liability allocation. Providers and deployers may share liability depending on system modifications - crucial considerations for integration partnerships. #### Preparing your organisation for AI Act enforcement ##### Building compliant AI governance structures Governance structures that satisfy the Act differ fundamentally from traditional IT governance. You need risk management systems spanning entire AI lifecycles, quality management systems with documented procedures, technical oversight capabilities for AI-specific risks, and incident response processes for serious AI incidents. The Act's quality management system requirements encompass conformity procedures, modification protocols, design and development controls, testing and validation processes, data management systems, postmarket monitoring, incident reporting mechanisms, communication processes, recordkeeping systems, resource management, and accountability frameworks. These aren't checkbox exercises. Effective governance requires organisational changes: establishing AI oversight committees, appointing accountable individuals for high-risk systems, creating cross-functional review processes, and implementing continuous monitoring capabilities. ##### Audit trails and monitoring systems that actually work The Act mandates automatic logging for high-risk systems throughout their operational lifetime. Logs must capture usage timestamps, input data, reference databases, and verification personnel. This exceeds typical application logging: you're building forensic capability. Effective audit systems go beyond compliance minimums. They enable root cause analysis when systems fail, performance degradation detection before incidents occur, bias emergence identification in production systems, and compliance demonstration during regulatory inspections. The ISACA emphasise how postmarket monitoring must account for system interactions with other AI systems. Your monitoring can't assume isolated operation. It must detect emergent behaviours from system combinations, especially when your high-risk system processes outputs from other AI systems. ##### Staff training and organisational readiness AI literacy requirements apply to all AI systems, not just high-risk ones. Anyone operating AI systems needs sufficient understanding of capabilities, limitations, and appropriate use. This is beyond optional training; it's a compliance requirement with enforcement implications. Effective AI literacy programmes address technical fundamentals without requiring engineering expertise, ethical and legal considerations specific to your deployments, practical application within role-specific contexts, critical evaluation skills for AI outputs, and awareness of automation bias and over-reliance risks. SMEs and mid-caps receive certain exemptions, but AI literacy requirements remain universal. Even organisations qualifying for simplified documentation must ensure operational staff understand AI systems they're using. ##### Third-party AI and supply chain considerations Most organisations deploy third-party AI systems, creating complex compliance chains. The Act assigns obligations based on roles - provider, deployer, distributor, importer - that may shift depending on modifications and deployments. If you customize a third-party system substantially, you become a provider with full compliance obligations. If you simply deploy unchanged systems, you're a deployer with reduced but still significant requirements. Understanding these distinctions drives vendor selection and integration strategies. Supply chain compliance requires vendor assessment for Act compliance, contractual allocation of obligations, technical documentation access rights, incident reporting procedures, and liability allocation agreements. Smart organisations are building these requirements into procurement processes now. #### The strategic opportunities hidden in regulatory constraints ##### Innovation pathways within compliance frameworks Regulatory constraints force innovation along unexpected vectors. When emotion recognition in workplaces faces prohibition, organisations will be forced to develop performance analytics through objective metrics. When bias mitigation becomes mandatory, fairness-aware architectures emerge that outperform naive approaches. The Act's regulatory sandboxes suggest particular opportunity. These frameworks allow real-world testing of high-impact AI under regulatory guidance. The Digital Omnibus proposal extends sandbox access to general-purpose AI models and broadens real-world testing permissions for systems under product regulations. Sophisticated organisations will use sandboxes for competitive intelligence. They’ll test boundary-pushing capabilities while maintaining compliance. They’ll shape regulatory understanding through demonstrated safe operation. They’ll establish precedents competitors must follow. ##### Market differentiation through responsible AI leadership Responsible AI has moved from ethical nice-to-have to regulatory requirement. Organisations demonstrating leadership in responsible AI development capture premium market segments. They attract talent concerned about AI's societal impact and reduce regulatory scrutiny through proactive compliance. The Act's transparency requirements become differentiators when exceeded voluntarily. Publishing model cards for limited-risk systems, implementing human oversight for minimal-risk applications, and maintaining public audit logs beyond requirements build trust that translates to market share. Financial services organisations are already leveraging this dynamic. Banks exceeding transparency requirements attract customers concerned about algorithmic discrimination. Insurers providing detailed AI decision explanations reduce complaints and regulatory investigations. ##### The first-mover advantage in regulated AI markets Regulated markets reward first movers who shape compliance standards. Your implementation approaches become templates. Your technical solutions define feasibility. Your governance structures establish precedents. The Act's timeline, especially with Digital Omnibus extensions, provides a window for establishing market position. Organisations achieving compliance before mandatory deadlines capture customers requiring verified AI governance. They influence standards development through demonstrated practices. They build operational expertise competitors can't quickly replicate. Consider conformity assessment procedures. Organisations completing assessments early understand requirements competitors are still interpreting. They've identified efficient assessment paths, established assessor relationships, and documented successful approaches. When assessment becomes mandatory, they're helping customers navigate processes they've already mastered. The real opportunity isn't compliance - it's using compliance requirements to build superior AI systems. Documentation requirements force architectural clarity. Oversight obligations prevent automation bias. Monitoring requirements enable continuous improvement. Companies treating the Act as a capability framework rather than a compliance burden build systems that are more robust, trustworthy, and ultimately more valuable than those scraped together to meet minimum requirements. If you're ready to build AI solutions that exploit full technical potential while naturally exceeding regulatory minimum requirements, contact us today. The future belongs to those who see regulation as architecture guidance, not limitation. #### References - Harvard Kennedy School Paper on EU AI Act Implications for U.S. Healthcare - Cooley Law Firm Analysis of Digital Omnibus on AI and Business Compliance Roadmaps - EY Switzerland Comprehensive Business Impact Assessment - ISACA White Paper on Understanding EU AI Act Requirements --- ### Why Most AI Projects Fail Without Expert AI Consulting - URL: https://agathon.ai/insights/why-most-ai-projects-fail-without-expert-ai-consulting - Published: 2026-01-23 - Categories: AI Strategy, AI Consulting, Machine Learning #### The multi-million lesson nobody wants to learn Most organisations building AI products use less than 15% of what modern architectures make possible. They implement ChatGPT wrappers and call it innovation. They hire data scientists who build models that never see production. They burn through budgets chasing phantom problems whilst ignoring fundamental architectural decisions that doom their projects from inception. MIT researchers recently discovered that 95% of enterprise AI pilots fail to achieve rapid revenue acceleration. This isn't a technology problem; it's an expertise problem. The models work fine. The integration strategies, architectural choices, and strategic alignment are what collapse under pressure. #### When ambition meets algorithmic reality ##### The seductive promise of off-the-shelf AI Every vendor promises transformation through pre-trained models and API calls. Purchase their solution, they claim, and watch productivity soar. MIT's data tells a different story: organisations purchasing specialised AI tools from vendors succeed 67% of the time, whilst internal builds succeed only one-third as often. The catch? Most purchased solutions deliver incremental improvements rather than transformational change. The seduction lies in simplicity. Generic tools like ChatGPT excel for individual users due to their flexibility, but researchers found they stall in enterprise deployment because they don't learn from or adapt to workflows. They're stateless, context-blind, and fundamentally disconnected from the organisational knowledge graph that drives real value creation. ##### Why your data scientist isn't enough Your data scientist builds excellent models. They achieve 95% accuracy on test sets, create beautiful visualisations, and speak fluently about transformer architectures. Yet their models never reach production, or worse, they reach production and fail spectacularly. Scientists studying 555 neuroimaging-based AI models for psychiatric diagnosis found that only 15.5% included external validation. The remaining 84.5% performed brilliantly on their training data and collapsed when exposed to real-world variation. This isn't incompetence; it's the difference between building models and building systems. Data scientists excel at the former. Production AI requires the latter. ##### The expertise gap that kills projects MIT's research identified a critical pattern: successful AI deployments require empowering line managers, not just central AI labs, to drive adoption. Yet most organisations concentrate AI expertise in isolated teams, creating a fatal disconnect between technical capability and operational reality. The expertise gap manifests in three dimensions. First, architectural sophistication: understanding how to build systems that learn, remember, and act autonomously within boundaries. Second, strategic alignment: connecting AI capabilities to genuine business outcomes rather than technical metrics. Third, responsible development: addressing bias, fairness, and compliance before they become existential threats. #### The technical debt you don't see coming ##### Architecture decisions that haunt you later Early architectural choices compound into insurmountable technical debt. Researchers examining healthcare AI implementations found that measurement biases from different data sources—variations in imaging hardware, software versions, and acquisition parameters—fundamentally alter model behaviour. Models trained on single-source data underperform catastrophically when deployed across heterogeneous environments. The architecture problem extends beyond data heterogeneity. Monolithic model designs prevent iterative improvement. Stateless implementations sacrifice contextual learning. Synchronous processing patterns create bottlenecks that scale linearly with usage. These aren't bugs; they're architectural constraints baked into the foundation. ##### When machine learning models become maintenance nightmares Scientists documented a phenomenon called "feedback loop bias" where AI systems trained on their own predictions progressively degrade. Clinicians accepting AI recommendations, even inaccurate ones, generate training data that reinforces errors in future iterations. The model appears to improve whilst actually becoming less reliable. Maintenance complexity scales exponentially with model count. Each model requires monitoring for drift, retraining pipelines, version control, and performance tracking. Organisations deploying dozens of disconnected models create an unmaintainable web of dependencies. The maintenance burden eventually exceeds the value delivered, forcing either wholesale replacement or gradual abandonment. ##### The hidden costs of rushing to production Speed to market drives premature deployment. Researchers found that 50% of healthcare AI studies demonstrated high risk of bias, often due to absent sociodemographic data, imbalanced datasets, or weak algorithm design. These biases remain latent until deployment, manifesting as discrimination, unfairness, or systematic errors affecting specific populations. The rush to production bypasses critical validation steps. External validation, bias assessment, and adversarial testing get postponed indefinitely. Technical shortcuts become permanent fixtures. Temporary solutions ossify into core infrastructure. The accumulated shortcuts create fragility that surfaces during scaling, often requiring complete architectural overhauls. #### Strategic misalignment: Building solutions to phantom problems ##### Solving for technology instead of outcomes MIT's analysis revealed that over half of generative AI budgets target sales and marketing tools, yet researchers found the highest ROI in back-office automation—eliminating business process outsourcing, cutting agency costs, and streamlining operations. Organisations chase visible AI applications whilst ignoring transformational opportunities in core operations. The technology-first approach inverts proper solution design. Teams select impressive models then search for applications. They optimise for accuracy metrics rather than business impact. They celebrate technical achievements that deliver negligible operational value. Success becomes defined by model performance rather than outcome improvement. ##### The dangerous disconnect between AI capabilities and business needs Researchers identified a fundamental misalignment: executives blame regulation or model performance for AI failures, whilst data shows the core issue is flawed enterprise integration. The disconnect stems from treating AI as a technical project rather than a business transformation initiative. Business stakeholders request "AI solutions" without articulating specific problems. Technical teams deliver sophisticated models that don't address actual pain points. The resulting systems excel at tasks nobody needs whilst failing at critical business functions. Value creation requires translating between technical possibility and business necessity—a translation most organisations never attempt. ##### Why proof-of-concepts rarely scale Proof-of-concepts succeed in controlled environments then collapse under production loads. Scientists found this pattern consistently: models achieving stellar performance on curated datasets fail when exposed to real-world complexity, data quality issues, and operational constraints. The scaling failure reflects fundamental differences between experimental and production environments. POCs operate on clean data, limited scope, and controlled conditions. Production demands robustness to dirty data, comprehensive coverage, and unpredictable usage patterns. The engineering effort required to bridge this gap often exceeds the original development cost by orders of magnitude. #### The responsible AI blindspot ##### Compliance requirements that arrive too late Regulatory frameworks from the European Commission, FDA, Health Canada, and WHO establish increasingly strict requirements for AI deployment. Yet organisations treat compliance as a post-development concern, discovering fundamental architectural incompatibilities only after substantial investment. Scientists examining AI bias found that systemic bias, representation bias, and measurement bias permeate training data and model architectures. These biases can't be retroactively removed; they require fundamental redesign. Compliance isn't a checklist; it's an architectural constraint that shapes every development decision. ##### Bias, fairness, and the lawsuits waiting to happen Researchers studying commercial risk prediction algorithms discovered systematic racial bias: Black patients received lower risk scores than White patients despite similar health conditions. The bias stemmed from using healthcare costs as a proxy for illness severity; systemic barriers led to lower costs for Black patients, causing algorithms to underestimate their needs. These biases create legal liability extending far beyond regulatory fines. Discrimination lawsuits, breach of duty claims, and negligence cases proliferate as AI systems demonstrate systematic unfairness. A study of 48 healthcare AI models found 50% exhibited high risk of bias. The legal exposure compounds with deployment scale, creating existential threats to organisations deploying biased systems. ##### When ethical considerations become existential threats MIT researchers documented "automation bias" where clinicians inappropriately trust AI predictions, failing to notice errors or ignoring conflicting evidence. This erosion of human oversight creates cascading failures: errors propagate unchecked, expertise atrophies, and accountability dissolves. Ethical failures trigger reputational damage that destroys market position overnight. Trusted institutions lose credibility when AI systems discriminate, fail, or cause harm. Recovery requires years of rebuilding trust, if recovery is possible at all. The ethical dimension isn't optional philanthropy; it's core risk management. #### The expertise advantage: What consultants actually bring ##### Pattern recognition from repeated failure Expert consultants have witnessed hundreds of AI failures across industries, architectures, and approaches. They recognise early warning signs: architectural anti-patterns, misaligned incentives, and capability gaps that doom projects. This pattern recognition enables intervention before problems become intractable. The value isn't in avoiding all mistakes but avoiding catastrophic ones. Experts distinguish between acceptable technical debt and architectural time bombs. They identify which shortcuts enable rapid iteration versus those creating permanent constraints. Pattern recognition from repeated exposure accelerates learning curves by years. ##### Technical depth meets strategic thinking MIT's research emphasised that successful AI requires both technical sophistication and strategic alignment. Experts bridge this divide, translating between deep technical capabilities and business outcomes. They understand not just what's technically possible but what delivers value within organisational constraints. Technical depth enables exploitation of advanced capabilities others miss. Whilst most implementations use basic features, experts leverage memory systems, agentic architectures, and compositional approaches that multiply capability. Strategic thinking ensures these capabilities target genuine business problems rather than technical curiosities. ##### The network effect of specialist knowledge Specialist expertise creates compound advantages through network effects. Experts bring awareness of cutting-edge techniques, emerging tools, and proven architectures from across the ecosystem. They've debugged similar problems in different contexts, accelerating solution discovery. The network extends beyond technical knowledge to include regulatory understanding, vendor relationships, and talent connections. Experts know which approaches satisfy emerging regulations, which vendors deliver versus promise, and where to find scarce AI talent. This network effect compresses implementation timelines whilst reducing risk. #### Building internal capability whilst leveraging external expertise ##### The hybrid model that works Successful AI transformation requires both internal capability and external expertise. MIT found that empowering line managers, not just central AI teams, drives adoption. External experts accelerate capability building whilst internal teams ensure sustainable operations. The hybrid approach balances speed with ownership. Experts provide architectural blueprints, implementation patterns, and quality frameworks. Internal teams adapt these to organisational context, maintain systems, and drive continuous improvement. Knowledge transfer occurs through collaboration rather than documentation. ##### Knowledge transfer that sticks Effective knowledge transfer requires deliberate structure beyond traditional training. Experts must work alongside internal teams on real projects, demonstrating techniques in context. Abstract principles become concrete through application to actual problems. Scientists studying AI adoption found that reading ethics guidelines has no significant influence on developer decision-making. Similarly, theoretical AI training rarely changes behaviour. Knowledge transfer succeeds through apprenticeship models where internal teams learn by doing under expert guidance. Skills develop through practice, not PowerPoints. ##### Creating sustainable AI practices Sustainability requires embedding AI capabilities throughout the organisation rather than concentrating them in specialised teams. This distribution prevents single points of failure whilst enabling rapid, contextual innovation. Sustainable practices emerge from three foundations: architectural standards that prevent technical debt accumulation, operational processes that ensure continuous improvement, and governance structures that balance innovation with risk. External expertise establishes these foundations; internal teams evolve them based on organisational learning. #### The economics of getting it right first time ##### Calculating the true cost of failure Failed AI projects cost more than direct investment. Opportunity costs from delayed transformation, reputational damage from public failures, and organisational cynicism blocking future initiatives compound the loss. MIT's finding that 95% of AI pilots fail to achieve rapid revenue growth represents trillions in destroyed value. Recovery costs exceed initial investment. Failed architectures require complete replacement, not incremental fixes. Biased models demand fundamental redesign. Technical debt compounds interest until systems become unmaintainable. Getting it wrong costs multiples of getting it right. ##### Investment protection through expert guidance Expert guidance protects investment through risk mitigation and capability multiplication. Avoiding one architectural mistake can save millions in rework. Identifying the right problem to solve prevents wasted effort on phantom opportunities. Researchers found that purchased AI solutions succeed twice as often as internal builds, yet most organisations attempt internal development first. Expert guidance identifies when to build versus buy, which capabilities to develop internally versus source externally, and how to integrate disparate approaches into coherent systems. ##### When consultants pay for themselves Consultants deliver positive ROI when they prevent catastrophic failures, accelerate time to value, or unlock capabilities that wouldn't otherwise exist. A single prevented failure covers years of consulting fees. Compressed implementation timelines generate revenue months earlier. Advanced architectural approaches deliver capabilities competitors can't match. The calculation extends beyond direct returns. Consultants transfer knowledge that compounds over time, establish practices that prevent future failures, and build internal capabilities that enable sustainable innovation. The investment appreciates rather than depreciates. #### Moving forward: A pragmatic approach to AI success AI transformation requires acknowledging uncomfortable realities. Most organisations lack the expertise to build sophisticated AI systems. Most AI projects will fail without expert guidance. Most current implementations exploit a fraction of available capability. Success demands three commitments. First, engage expertise early when architectural decisions remain fluid. Second, prioritise capability building alongside system delivery. Third, measure success through business outcomes rather than technical metrics. The path forward isn't about choosing between internal development and external expertise. It's about orchestrating both to achieve what neither accomplishes alone. Internal teams provide context, continuity, and ownership. External experts provide patterns, practices, and possibilities. The organisations succeeding with AI aren't those with the biggest budgets or best data scientists. They're those with the wisdom to recognise their limitations and the courage to address them. They build on proven architectures rather than reinventing foundations. They learn from others' failures rather than repeating them. If you're ready to build AI solutions that exploit full technical potential rather than implementing basic features, you should contact us today. #### References - MIT Fortune magazine report on 95% AI pilot program failure rates - Nature Digital Medicine article on bias recognition and mitigation strategies in AI healthcare applications - PMC academic article on biases in AI and addressing inevitable ethical issues --- ### Building a Secure LLMOps Pipeline: From Development to Production - URL: https://agathon.ai/insights/building-a-secure-llmops-pipeline-from-development-to-production - Published: 2026-01-23 - Categories: LLMs, AI Strategy, Machine Learning #### The security theatre of LLMOps: Why most pipelines are vulnerable by design Most organisations deploying large language models operate under a dangerous illusion: they believe their standard DevOps security practices translate to LLMOps. They're wrong. The attack surface of an LLM pipeline extends far beyond traditional software vulnerabilities, creating novel exploitation vectors that conventional security tools simply cannot detect. Recent research demonstrates the severity of this oversight. Scientists have shown that GPT-4 can autonomously exploit 87% of one-day vulnerabilities, while traditional security scanners achieve 0% success on the same targets. Meanwhile, researchers discovered that poisoning just 250 documents in pretraining data can successfully backdoor LLMs ranging from 600M to 13B parameters. These aren't theoretical risks - they're active exploitation vectors being weaponised today. #### Understanding the LLMOps attack surface ##### Model weights and intellectual property risks Model weights represent billions of dollars in computational investment and competitive advantage, yet most organisations treat them like standard application binaries. The fundamental difference: model weights encode both proprietary knowledge and potential attack vectors simultaneously. Researchers have demonstrated that model extraction attacks can reconstruct near-perfect replicas of production models through careful API querying. The attack requires no access to training data or architecture details - just systematic observation of input-output pairs. Traditional rate limiting fails here because legitimate usage patterns often mirror extraction attempts. The intellectual property risk compounds when considering fine-tuned models. These contain not just general knowledge but domain-specific insights, customer patterns, and business logic. A leaked financial services model doesn't just expose algorithms; it reveals trading strategies, risk assessments, and customer behaviour models built over years. ##### Training data poisoning and extraction vulnerabilities Data poisoning represents the most insidious threat to LLM pipelines. Researchers discovered that strategic corruption of alignment samples through "PoisonedAlign" attacks makes models substantially more vulnerable to future prompt injection while maintaining normal benchmark performance. The poisoned models pass all standard evaluation metrics, making detection nearly impossible without specialised auditing. The scale required for successful poisoning is surprisingly small. Scientists have shown that contaminating just 0.01% of training data can insert persistent backdoors that activate on specific trigger phrases. These backdoors survive fine-tuning, quantisation, and even some forms of adversarial training. Training data extraction presents an equally serious concern. The "Fundamental Law of Information Recovery" reveals that answering sufficient queries about a database inevitably leaks sensitive information. LLMs trained on proprietary datasets become oracles for that data, potentially revealing customer information, trade secrets, or confidential communications through carefully crafted prompts. ##### Prompt injection through the development lifecycle Prompt injection isn't confined to production systems - it pervades the entire development lifecycle. During model evaluation, researchers use prompts to test capabilities. During fine-tuning, prompts shape behaviour. During deployment, prompts become the primary interface. Each stage introduces unique vulnerabilities. Scientists have demonstrated "homotopy-inspired" prompt obfuscation techniques that bypass safety mechanisms by applying linguistic deformations that preserve semantic intent while altering surface form. These transformations exploit the topological structure of language, treating prompts as points in a continuous semantic space that can be smoothly deformed to evade detection. The EchoLeak vulnerability (CVE-2025-32711) proved that zero-click prompt injection attacks are possible in production. Researchers showed they could exfiltrate corporate data from Microsoft 365 Copilot by sending specially crafted emails - no user interaction required. The attack exploited "LLM scope violations" where external untrusted input manipulated the AI agent to access and leak sensitive information autonomously. ##### Supply chain compromises in foundation models Foundation models represent the ultimate supply chain risk. Organisations build critical systems atop models whose training data, processes, and potential backdoors remain opaque. You're essentially importing a black box with billions of parameters, any subset of which could encode malicious behaviour. Recent experiments revealed that major foundation models can be compromised through data poisoning attacks that affect downstream applications. The contamination persists through fine-tuning because the poisoned behaviours become encoded in deep representational layers that transfer learning preserves. Model distribution platforms compound this risk. Researchers documented 91,403 attack sessions targeting AI infrastructure between October 2025 and January 2026, with attackers methodically probing over 70 LLM endpoints. The attacks exploited server-side request forgery vulnerabilities through model pull operations, demonstrating that even model deployment mechanisms become attack vectors. #### Architecting security-first development environments ##### Isolated compute infrastructure and sandboxing strategies Effective LLMOps security starts with radical isolation. Standard containerisation isn't sufficient when models can generate arbitrary code or manipulate their runtime environment. You need nested virtualisation with hardware-enforced boundaries between model execution and system resources. The key insight: treat every model invocation as potentially hostile code execution. This means air-gapped training environments, isolated inference clusters, and complete network segmentation between development and production systems. Models should execute in ephemeral environments that reset after each batch, preventing persistent compromise. Resource allocation becomes a security control. Memory limits prevent models from consuming system resources in denial-of-service attacks. CPU quotas prevent cryptomining. Network policies prevent data exfiltration. These aren't performance optimisations - they're security boundaries. ##### Version control for models, data, and prompts Traditional version control systems fail catastrophically for LLMOps. Git wasn't designed for multi-gigabyte model files, datasets measured in terabytes, or the complex lineage relationships between prompts, data, and model versions. Effective model versioning requires cryptographic attestation at every stage. Each model checkpoint must include tamper-evident logs of training data hashes, hyperparameters, and code versions. This creates an immutable audit trail that can detect post-hoc manipulation or unauthorised modifications. Prompt versioning presents unique challenges. Prompts aren't just text - they're executable specifications that determine model behaviour. Version control must track not just prompt content but semantic intent, evaluation results, and safety validation outcomes. A seemingly innocuous prompt modification can completely alter model behaviour. ##### Secure collaboration patterns for distributed teams Distributed AI development multiplies security risks. Each team member represents a potential compromise vector, whether through malicious action or compromised credentials. Traditional role-based access control fails when a junior developer's prompt experiments can poison production models. Implement cryptographic commit signing for all model modifications. Require multi-party authorisation for production deployments. Use homomorphic encryption for collaborative training on sensitive data. These aren't paranoid measures - they're necessary given the attack surface. The principle of least privilege requires fundamental rethinking for LLMOps. Access to training data doesn't mean access to model weights. Ability to run inference doesn't mean ability to modify prompts. Permission to fine-tune doesn't mean permission to alter base models. Granular permission models must reflect the unique risks of each capability. ##### Automated vulnerability scanning in model development Static analysis tools for traditional code don't understand model vulnerabilities. You need specialised scanners that detect prompt injection susceptibility, adversarial robustness, and data leakage potential. Researchers demonstrated that automated red-teaming can identify model vulnerabilities before deployment. These tools systematically probe models with adversarial prompts, jailbreak attempts, and extraction attacks. The key: automation must match the scale of model development. Manual security reviews cannot keep pace with continuous model updates. Vulnerability scanning must extend beyond the model itself. Training scripts, data pipelines, and deployment configurations all represent attack surfaces. A misconfigured data loader can leak training examples. An improperly secured API endpoint can enable model extraction. Security scanning must encompass the entire pipeline. #### Data governance throughout the pipeline ##### Privacy-preserving training techniques Differential privacy isn't optional for production LLMs - it's essential. Researchers have shown that models memorise training data verbatim, potentially regurgitating sensitive information in response to targeted prompts. Without privacy-preserving training, every model becomes a potential data breach. The mathematics of differential privacy for deep learning are well-established. Scientists demonstrated that carefully calibrated noise injection during training can provide formal privacy guarantees while maintaining model utility. The challenge lies in implementation: most frameworks treat differential privacy as an afterthought rather than a core design principle. Secure multi-party computation enables collaborative training without data sharing. Organisations can jointly train models while keeping their data encrypted and isolated. This isn't theoretical - production systems demonstrate that federated learning with differential privacy can achieve comparable accuracy to centralised training while preserving privacy. ##### Differential privacy implementation strategies Implementing differential privacy requires more than adding noise to gradients. You need privacy accounting to track cumulative privacy loss across training iterations. Scientists have shown that without proper accounting, privacy guarantees degrade rapidly as training progresses. The choice of privacy parameters (epsilon and delta) determines the privacy-utility trade-off. Researchers discovered that adaptive clipping strategies can improve this trade-off by dynamically adjusting noise levels based on gradient statistics. However, these optimisations introduce new attack surfaces if not carefully implemented. Privacy amplification through subsampling and shuffling can strengthen privacy guarantees without additional noise. Random batch selection and data shuffling make it harder for adversaries to target specific training examples. These techniques are particularly effective for large-scale distributed training where natural randomness provides additional protection. ##### Data lineage and audit trails Every byte of training data must be traceable from source to model. This isn't just compliance - it's operational security. When researchers discover a poisoned dataset, you need to identify every model trained on that data, every prediction made by those models, and every decision influenced by those predictions. Cryptographic data provenance ensures tamper-evident lineage tracking. Hash chains link data transformations, creating an immutable record of data processing. Merkle trees enable efficient verification of dataset integrity. These aren't overengineered solutions, they're necessary given the sophistication of data poisoning attacks. Data retention policies must balance model interpretability with privacy risks. Keeping training data enables model debugging and auditing but increases breach exposure. Researchers have shown that synthetic data generation can provide a middle ground, preserving statistical properties while eliminating individual records. ##### Compliance frameworks for regulated industries Financial services, healthcare, and government deployments face stringent regulatory requirements that standard LLMOps platforms ignore. GDPR's right to erasure becomes complex when individual data points influence billions of model parameters. HIPAA compliance requires encryption at rest and in transit, including model weights that encode patient information. Model cards and datasheets provide standardised documentation for regulatory compliance. These documents must detail training data sources, known biases, intended use cases, and evaluation metrics. However, static documentation isn't sufficient - you need dynamic compliance monitoring that tracks model behaviour against regulatory constraints. Cross-border data transfers introduce additional complexity. Models trained on EU data may be subject to GDPR even when deployed elsewhere. The solution requires geo-distributed training with data localisation, ensuring that sensitive data never leaves jurisdictional boundaries while still enabling global model development. #### Model evaluation beyond accuracy metrics ##### Adversarial robustness testing Accuracy metrics tell you nothing about security. A model with 99% accuracy can be completely compromised by adversarial examples that humans can't distinguish from normal inputs. Robustness testing must be as rigorous as accuracy evaluation. Researchers have developed standardised benchmarks for adversarial robustness, but most organisations ignore them. The RobustBench leaderboard shows that even state-of-the-art models fail catastrophically against sophisticated attacks. Your production models are likely far more vulnerable. Adaptive attacks that evolve based on model defences represent the current frontier. Static adversarial training fails against attackers who observe model behaviour and adjust their strategies. You need dynamic defence mechanisms that evolve alongside threats. ##### Bias and fairness assessments Bias isn't just an ethical concern, it's a security vulnerability. Biased models leak information about training data distributions, enabling inference attacks. Researchers have shown that demographic biases can be exploited to extract sensitive attributes from model predictions. Fairness metrics must be evaluated across multiple dimensions simultaneously. A model that appears fair on gender may discriminate on race. A model fair on individual metrics may be unfair on intersectional groups. Comprehensive evaluation requires exponentially many fairness checks. Post-hoc bias mitigation often introduces new vulnerabilities. Researchers discovered that fairness constraints can be exploited to force specific model behaviours. The solution requires bias-aware training from the ground up, not cosmetic adjustments to biased models. ##### Output safety validation Safety isn't boolean - it's contextual. Output that's safe for adult users may be harmful for children. Content appropriate for creative writing may be dangerous for medical advice. Safety validation must understand context, not just content. Constitutional AI approaches embed safety principles directly into model training. Rather than filtering outputs post-hoc, models learn to generate safe content by design. Researchers demonstrated that this approach is more robust against adversarial prompts that attempt to bypass safety filters. Safety validation must extend beyond text generation. Models that can execute code, query databases, or trigger actions require additional scrutiny. A seemingly safe text output becomes dangerous when it contains SQL injection payloads or shell commands. ##### Performance under data drift Production data distributions drift continuously, but most organisations never detect when models become unreliable. Drift detection isn't just about maintaining accuracy - it's about identifying when models enter unfamiliar territory where security guarantees no longer hold. Researchers have shown that models can maintain high average accuracy while failing catastrophically on shifted distributions. A fraud detection model trained on pre-pandemic data may completely miss new fraud patterns. Without drift detection, these failures remain invisible until catastrophic losses occur. Continual learning approaches that adapt to drift introduce new attack vectors. Online learning systems can be manipulated through adversarial data injection. The solution requires careful balance between adaptation and stability, with robust detection of malicious drift. #### Hardening the deployment infrastructure ##### Container security for model serving Standard container security practices fail for model serving. Models aren't stateless microservices - they're stateful computation engines with massive memory footprints and complex dependencies. Container escapes that would be theoretical for normal applications become practical when models can generate and execute arbitrary code. Implement read-only container filesystems with mounted model weights. This prevents models from modifying themselves or persisting malicious payloads. Use minimal base images that exclude compilers, interpreters, and network tools. Every additional capability is an potential exploit vector. Resource limits must account for model-specific attack patterns. Researchers documented cryptomining attacks that exploit GPU access for model inference. Memory exhaustion attacks target model loading. CPU saturation attacks abuse unbounded generation. Configure hard limits for all resources with automatic container termination on violation. ##### API gateway protection and rate limiting Traditional rate limiting based on requests per second fails for LLMs. A single request can consume vastly different resources depending on prompt length, generation parameters, and model complexity. You need adaptive rate limiting based on computational cost, not request count. Token-level rate limiting provides granular control. Track input tokens, output tokens, and total tokens per user, IP, and API key. Implement exponential backoff for limit violations. This prevents both denial-of-service attacks and model extraction attempts that rely on high-volume querying. API gateways must validate more than just authentication. Prompt validation, content filtering, and safety checking should occur before model invocation. This creates defence in depth - even if one layer fails, others provide protection. ##### Model versioning and rollback mechanisms Production models need instant rollback capabilities. When researchers discover vulnerabilities or adversarial attacks, you must revert to safe versions within minutes, not hours. This requires sophisticated version management beyond simple model file storage. Canary deployments for models differ from traditional software. You can't just route 1% of traffic to test stability. Model behaviour is probabilistic - rare failure modes might not appear in small samples. You need statistical validation across thousands of invocations before promoting versions. Version compatibility extends beyond API contracts. Prompt formats, token vocabularies, and generation parameters can change between versions. Rollback must account for these differences, potentially requiring prompt translation or request replay. ##### Runtime monitoring and anomaly detection Model behaviour monitoring requires understanding of both normal and adversarial patterns. Baseline normal behaviour across multiple dimensions: latency distributions, token probabilities, attention patterns, and activation statistics. Deviations indicate potential attacks or model compromise. Researchers have shown that adversarial inputs create detectable patterns in model internals. Attention weights concentrate unusually. Activation patterns diverge from typical distributions. Entropy of output distributions shifts. These signals enable runtime attack detection even when outputs appear normal. Monitor not just the model but its entire execution context. System calls, network connections, file access, and resource utilisation provide security signals. A model that suddenly starts making DNS queries or opening network sockets is likely compromised. #### Production security operations ##### Real-time threat detection systems Security information and event management (SIEM) systems must evolve for LLMOps. Traditional pattern matching fails against semantic attacks that vary surface form while preserving malicious intent. You need semantic analysis of prompts, outputs, and model behaviour. Threat detection requires correlation across multiple signals. A prompt that seems benign in isolation becomes suspicious when preceded by specific queries. Output that appears safe becomes dangerous when combined with previous responses. Context-aware detection systems must maintain conversation state and identify multi-turn attacks. Researchers demonstrated that ensemble detection methods combining multiple weak signals achieve high accuracy with low false positives. No single indicator reliably identifies attacks, but combinations of indicators provide strong signals. This requires sophisticated correlation engines that can process millions of events in real-time. ##### Incident response protocols for AI systems When models are compromised, traditional incident response fails. You can't just isolate a server and preserve disk images. Model state exists in memory, conversation history spans multiple systems, and attack artifacts may be probabilistic patterns rather than files. Develop AI-specific incident response playbooks. Document procedures for model isolation, state preservation, and forensic analysis. Define escalation paths for different attack types: prompt injection, model extraction, data poisoning, and adversarial examples. Each requires different response strategies. Recovery from AI incidents requires careful validation. Simply restoring from backups isn't sufficient when attackers may have poisoned training data or inserted backdoors weeks before detection. You need comprehensive revalidation of models, data, and predictions made during the compromise window. ##### Model behaviour monitoring and drift detection Production models drift in ways that accuracy metrics don't capture. Researchers have shown that models can maintain high accuracy while their internal representations shift dramatically. These representational shifts indicate potential compromise or data poisoning. Monitor prediction confidence distributions, not just predictions. Sudden changes in confidence patterns indicate model uncertainty or adversarial manipulation. A model that becomes overconfident on out-of-distribution inputs is likely compromised. Behavioural analysis must extend to model explanation. Track how feature importance, attention weights, and attribution maps evolve over time. Changes in model reasoning patterns often precede visible accuracy degradation. ##### Continuous security validation frameworks Security isn't a point-in-time assessment - it requires continuous validation. Automated red teams should constantly probe production models with new attack variants. This isn't penetration testing; it's continuous security monitoring. Researchers developed frameworks for automated adversarial testing that evolve attack strategies based on model responses. These systems discover novel vulnerabilities by combining known attack primitives in unexpected ways. Static security assessments miss these emergent attack patterns. Validation must cover the entire kill chain from initial prompt to final output. Test data exfiltration, model manipulation, prompt injection, and output modification. Each stage requires different validation techniques and success metrics. #### The human factor in LLMOps security ##### Access control and authentication patterns Zero-trust architecture isn't optional for LLMOps. Every access request must be authenticated, authorised, and audited. This includes not just human users but also system components, models, and automated processes. Implement capability-based access control rather than role-based. The ability to invoke a model doesn't imply ability to modify its prompts. Permission to view outputs doesn't grant permission to access internal states. Granular capabilities prevent privilege escalation. Multi-factor authentication must extend beyond passwords and tokens. Behavioural biometrics can identify users based on interaction patterns. Anomaly detection can flag unusual access patterns. These additional factors provide defence against credential compromise. ##### Security training for ML engineers ML engineers need security training specific to AI threats. Traditional secure coding practices don't address prompt injection, model extraction, or adversarial examples. Generic security awareness doesn't prepare teams for AI-specific attacks. Develop hands-on training using captured attack data. Let engineers experience actual prompt injection attempts, observe model extraction attacks, and analyse adversarial examples. Theoretical knowledge isn't sufficient - teams need practical experience with real attacks. Create internal red teams focused on AI security. These teams should continuously probe your models, discover vulnerabilities, and develop mitigations. This builds security expertise while identifying weaknesses before external attackers. ##### Cross-functional security responsibilities Security can't be delegated to a separate team. Data scientists must understand poisoning attacks. ML engineers must grasp adversarial robustness. Product managers must recognise prompt injection risks. Security is everyone's responsibility. Establish security champions within each functional team. These individuals bridge the gap between security specialists and domain experts. They translate security requirements into practical implementations and identify domain-specific vulnerabilities. Regular security reviews must include all stakeholders. Data scientists review training data integrity. Engineers assess pipeline security. Product managers evaluate user-facing risks. This comprehensive approach identifies vulnerabilities that siloed reviews miss. ##### Building security culture in AI teams Security culture starts with acknowledging that every model is potentially compromised. This isn't paranoia - it's pragmatism given the attack surface. Teams must assume breach and design systems that remain secure even when individual components fail. Celebrate security discoveries, not just feature launches. When team members identify vulnerabilities, recognise their contribution. This encourages proactive security thinking rather than reactive patching. Make security metrics visible alongside performance metrics. Track vulnerability discovery rates, patch times, and security test coverage. What gets measured gets managed. Security must be as prominent as accuracy in team dashboards. #### Future-proofing your LLMOps security posture The threat landscape evolves faster than defensive capabilities. Researchers discovered that multimodal models introduce new attack vectors through image-embedded instructions that bypass text-based filters. Agentic systems that can execute code and access external resources multiply the attack surface exponentially. Future-proofing requires architectural flexibility. Build systems that can incorporate new security controls without complete redesigns. Use pluggable security modules that can be updated as threats evolve. Design for defence-in-depth where new layers can be added without disrupting existing protections. Invest in security research, not just implementation. Partner with academic institutions studying AI security. Contribute to open-source security tools. Participate in responsible disclosure programs. The organisations that survive will be those that see security as a competitive advantage, not a compliance burden. The most sophisticated attacks won't target your models directly - they'll exploit the assumptions underlying your entire LLMOps pipeline. They'll poison data at the source. They'll compromise developer workstations. They'll infiltrate through supply chain dependencies. Security isn't about protecting models; it's about protecting the entire ecosystem that produces, deploys, and maintains them. Most organisations are building LLM capabilities on foundations of sand. They've adopted AI without adapting their security posture. They've deployed models without understanding their attack surface. They're one sophisticated adversary away from catastrophic breach. If you're ready to build AI solutions that exploit full technical potential while maintaining genuine security rather than theatre, you should contact us today. #### References - LLM Security and Safety: Insights from Homotopy-Inspired Prompt Obfuscation - ArXiv research paper examining LLM vulnerabilities and defensive strategies - Deep Learning with Differential Privacy - Foundational ArXiv paper on privacy-preserving training techniques for neural networks - Differential Privacy in Machine Learning: From Symbolic AI to LLMs - Recent comprehensive survey on differential privacy implementation strategies - Recent Advances of Differential Privacy in Centralized Deep Learning: A Systematic Survey - ArXiv systematic review of differential privacy in deep learning - LLM Exploit Generation: 91K Attacks Signal Security Crisis - Recent analysis of real-world LLM attacks and autonomous exploit generation --- ### Why tokenization matters: CharGPT vs ChadGPT - URL: https://agathon.ai/insights/why-tokenization-matters-chargpt-vs-chadgpt - Published: 2026-01-23 - Categories: AI Strategy, Machine Learning, LLMs #### The curious case of CharGPT versus ChadGPT: Why a single token changes everything Here's a thought experiment that reveals everything wrong with how most companies implement AI: ask GPT-4 to count the letters in "ChatGPT". Now ask it about "CharGPT". Watch it stumble. This isn't a quirk. It's a fundamental architectural limitation that affects every large language model in production today. And it's costing you more than you think - not just in API fees, but in lost capability, degraded performance, and missed opportunities to build genuinely sophisticated AI products. #### When language models can't spell their own names ##### The tokenisation blind spot nobody talks about Most CTOs assume their AI struggles with complex reasoning or lacks domain knowledge. They're solving the wrong problem. The real bottleneck sits much earlier in the pipeline: tokenisation, the process that turns text into numbers your model can process. Scientists have shown that GPT-4 tokenises "ChatGPT" as a single unit - one token, cleanly packaged. But "CharGPT"? That splits into "Char" and "GPT". Two tokens. Different embedding spaces. Completely different semantic representation. This isn't pedantry. It's the difference between a model that understands context and one that's merely pattern matching fragments. ##### How GPT-4 struggles with its own identity Researchers discovered that tokenisation inconsistencies cascade through every layer of processing. When a model splits "480" into one token but "481" into two tokens ("4" and "81"), it's not just inefficient - it fundamentally breaks arithmetic reasoning. The model must memorise arbitrary chunking patterns rather than learning algorithmic processing. The implications compound. A single misaligned token boundary can shift attention patterns, corrupt positional encodings, and derail the entire inference chain. Your carefully crafted prompts fail not because the model lacks capability, but because it literally sees different words than you intended. #### Demystifying the token: Your AI's atomic unit of thought ##### Characters are not tokens (and why that matters) Here's what most implementations miss: tokens aren't words, aren't characters, and certainly aren't concepts. They're statistical compromises - frequency-based fragments that emerged from training data. The typical English word requires 1.3 tokens. The same word in Ukrainian? Scientists have shown it needs 3.8 tokens on average. Your "multilingual" model isn't multilingual at all - it's linguistically biased at the most fundamental level. This affects everything: context window utilisation, inference costs, and most critically, the model's ability to maintain semantic coherence across languages. When your model uses 4x more tokens for non-English text, you're not just paying more - you're getting fundamentally worse performance. ##### The hidden vocabulary that shapes AI comprehension Modern models use vocabularies of 30,000 to 150,000 tokens. But here's the catch: researchers discovered that only 1.54% of learned tokens in byte-pair encoding correspond to meaningful linguistic units. The rest are arbitrary fragments, statistical accidents frozen in silicon. This creates a cascade of problems. Your model doesn't learn concepts - it learns fragment co-occurrences. It doesn't understand morphology - it memorises character sequences. Every sophisticated behaviour you observe is built on this fragile foundation of statistical text compression. ##### Byte-pair encoding: The compromise we all live with BPE dominates because it's computationally tractable, not because it's optimal. The algorithm merges the most frequent character pairs iteratively until reaching a target vocabulary size. Simple. Efficient. Wrong. Experiments have shown that BPE consistently fails on tasks requiring character-level precision. Splice site prediction accuracy drops by 22% compared to character-level tokenisation. Promoter detection suffers similar degradation. The model literally cannot see the patterns that matter because they're split across token boundaries. #### The CharGPT experiment: Breaking down a simple typo ##### What happens when you misspell ChatGPT "CharGPT" isn't just a typo - it's a tokenisation trap. The model sees "Char" (a programming term) and "GPT" (a model architecture). The semantic space shifts entirely. What should be a simple correction becomes a conceptual impossibility. This pattern repeats everywhere. "Sing lemon" versus "singlemon". "Table move" versus "tablemove". Each variant triggers different tokenisation, different embeddings, different outputs. Your users' typos aren't just causing errors - they're changing what the model perceives. ##### Token boundaries and the butterfly effect Researchers discovered that token boundary shifts cascade exponentially through transformer layers. A single character insertion can fragment a token, shifting every subsequent token's position, corrupting attention patterns, and ultimately producing nonsensical outputs. The effect is particularly severe in morphologically rich languages. Scientists have shown that Ukrainian text suffers from tokenisation inefficiency rates exceeding 380% compared to English. Every grammatical inflection risks fragmenting tokens into meaningless character sequences. ##### Why your model sees "Char" and "GPT" but not "CharGPT" The model's vocabulary is fixed at training time. "ChatGPT" earned its place through frequency. "CharGPT" didn't. So the model fragments it, processes the pieces independently, and reconstructs meaning from components that were never meant to be separate. This isn't a bug you can patch. It's architectural. Every token boundary is a potential failure point, and you could have dozens of them in every prompt. #### ChadGPT and the meme economy of tokenisation ##### Internet culture versus training data "ChadGPT" presents a different challenge. It's not a typo - it's a cultural reference that emerged after training. The model might tokenise it correctly but lacks the semantic grounding to understand the reference. This exposes tokenisation's temporal limitation. Your vocabulary is frozen at training time, but language evolves daily. New terms, new meanings, new contexts - all interpreted through outdated token mappings. ##### When tokens become cultural artefacts Scientists have shown that common dates from the 20th century receive unique tokens, while recent dates fragment into components. The model literally sees history differently based on when events occurred relative to its training date. This creates bizarre biases. "1995" might be one token, "2023" might be three. Historical events get processed more efficiently than current ones. Your AI isn't just outdated - it's architecturally biased toward the past. ##### The computational cost of understanding jokes Humour, wordplay, and cultural references often depend on precise character sequences. But tokenisation destroys this precision. A pun that relies on letter transposition becomes invisible when those letters span token boundaries. The cost isn't just comprehension - it's computational. Fragmented tokens require more processing, more attention computation, more memory. Your model works harder to understand less. #### Performance implications: More than just splitting words ##### Token efficiency and inference costs Here's what most implementations ignore: token count directly determines cost and speed. Researchers discovered that poor tokenisation can increase inference costs by 400% for non-English languages. But it's worse than linear scaling. Fragmented tokens corrupt attention patterns, requiring more computation per token. Your quadratic attention complexity becomes even more expensive when tokens don't align with semantic units. ##### Context windows and the tokenisation tax A 32k context window sounds impressive until you realise it's measured in tokens, not characters. Scientists have shown that the same information requires 2-4x more tokens in languages like Ukrainian or Thai compared to English. Your "large" context window shrinks dramatically for non-English users. That 32k window becomes effectively 8k for Ukrainian text. You're not just discriminating - you're architecturally excluding entire languages. ##### Why multilingual models suffer silently Experiments have shown multilingual models allocate vocabulary space inequitably. English dominates, capturing 70-95% of tokens despite representing a fraction of global language use. This creates compound disadvantages. Non-English text requires more tokens (higher cost), receives worse encoding (lower quality), and exhausts context windows faster (reduced capability). Your "multilingual" model is monolingual with translation overhead. #### The design flaw we're stuck with ##### Historical baggage from natural language processing Tokenisation emerged from compression algorithms, not linguistic theory. BPE was literally designed for data compression in the 1990s. We're building AGI on foundations designed for zip files. The path dependence is striking. Each generation of models inherits tokenisation assumptions from its predecessors. GPT inherits from BERT inherits from Word2Vec. We're not iterating toward optimality, we're accumulating technical debt. ##### The trade-offs nobody wants to discuss Character-level models solve tokenisation problems but explode sequence lengths. Word-level models maintain semantic coherence but can't handle novel words. Subword tokenisation splits the difference and satisfies neither requirement. Researchers discovered that no single tokenisation strategy dominates across all tasks. Splice site detection requires character precision. Document classification benefits from word-level semantics. Your model is optimised for nothing by trying to handle everything. ##### Subword tokenisation: Brilliant hack or fundamental limitation? BPE and its variants are engineering marvels - they shouldn't work, but they do. They compress text efficiently, handle unknown words gracefully, and scale to massive vocabularies. But they're still hacks. They treat symptoms, not causes. The fundamental problem remains: we're forcing discrete symbols onto continuous meaning. Every tokenisation algorithm is just choosing which information to lose. #### Practical implications for AI engineering ##### Prompt engineering in a tokenised world Your prompts fail for reasons you can't see. That carefully crafted instruction splits across token boundaries, fragmenting meaning. That specific example triggers unexpected tokenisation, shifting semantics. The solution isn't better prompts - it's token-aware prompting. Understand your model's vocabulary. Test tokenisation patterns. Design around boundaries, not through them. ##### Why your carefully crafted prompts fail mysteriously Scientists have shown that identical prompts can produce different outputs based solely on whitespace. A trailing space shifts token boundaries, changes positional encodings, and alters model behaviour. This isn't deterministic. The same prompt might work in testing and fail in production because surrounding context shifted token alignment. You're not debugging logic - you're debugging statistical compression artefacts. ##### Token-aware system design strategies Stop treating tokenisation as a preprocessing step. It's a design constraint that affects every architectural decision. Choose models based on tokenisation efficiency for your use case. Design context windows around token counts, not character counts. Build fallbacks for tokenisation failures. Most importantly, measure token efficiency as a key performance metric. #### The path forward: Rethinking linguistic representation ##### Character-level models and why we abandoned them Character-level models solve tokenisation perfectly - every character is a token. No boundaries, no fragments, no compression artefacts. We abandoned them because they're computationally intractable. Attention is quadratic in sequence length, and character-level encoding increases sequences by 4-6x. The cure is worse than the disease. ##### Emerging alternatives to traditional tokenisation Researchers are exploring learned tokenisation, where models discover optimal token boundaries during training. Others investigate continuous representations that bypass discrete tokens entirely. But these remain experimental. Production systems need solutions today, not promises tomorrow. The immediate future remains subword tokenisation, with all its flaws. ##### What the next generation might look like The next breakthrough won't iterate on tokenisation - it will eliminate it. Direct acoustic-to-semantic models for speech. Continuous embedding spaces for text. Hierarchical representations that capture multiple granularities simultaneously. Until then, we're stuck with a fundamental trade-off: semantic coherence versus computational efficiency. Every model chooses differently, and every choice has consequences. #### Why this matters more than you think Tokenisation isn't an implementation detail, rather it's the foundation that determines what your AI can and cannot do. It affects costs, capabilities, and fundamental behaviours in ways that no amount of fine-tuning can fix. Understanding tokenisation deeply - its implications, limitations, and workarounds - is what separates sophisticated AI implementations from expensive toys. It's the difference between building products that exploit AI's full technical potential and building features that merely check a box. If you're ready to build AI solutions that exploit full technical potential rather than implementing basic features, you should contact us today. #### References - Tokenization efficiency of current foundational large language models for the Ukrainian language - The first step is the hardest: pitfalls of representing and tokenizing temporal data for large language models - The Impact of Tokenizer Selection in Genomic Language Models - Tokenization Matters! Degrading Large Language Models through Challenging Their Tokenization - How Different Tokenization Algorithms Impact LLMs and Transformer Models for Binary Code Analysis --- ### Thin wrapper or true AI? Technical due diligence for AI investments - URL: https://agathon.ai/insights/thin-wrapper-or-true-ai-technical-due-diligence-for-ai-investments - Published: 2026-01-16 - Categories: AI Strategy, Machine Learning, NLP Every company is an AI company now. At least, that’s what the pitch decks claim. The reality is more sobering. According to MIT’s NANDA initiative, 95% of generative AI pilots at companies are failing. Yet valuations continue to assume these technologies will deliver. The gap between AI marketing and AI reality has never been wider—and investors are increasingly left holding the risk. Traditional due diligence wasn’t built for this. Legal review can tell you about IP ownership. Financial review can validate revenue claims. But neither can answer the question that actually matters: Is the AI real? In our experience assessing AI companies across multiple sectors, we’ve found that technical claims often fall into predictable patterns: some genuine, many less so. This framework is designed to help investors, board members, and M&A teams ask better questions before committing capital. #### The AI Implementation Spectrum Not all “AI companies” are created equal. Before assessing red flags, it helps to understand where a company sits on the implementation spectrum. ##### Tier 1: Thin Wrappers At the lowest end are companies whose entire “proprietary AI” consists of API calls to third-party models—OpenAI, Anthropic, or similar—wrapped in a custom interface. The technical differentiation is essentially a React frontend and some prompt engineering. This isn’t inherently problematic. Valuable businesses can be built on third-party infrastructure. But these companies should be valued as application businesses, not AI businesses. The moat is in distribution, user experience, or domain expertise, not technology. What to watch for: No ML engineers on staff. “Prompt engineering” described as core IP. Reluctance to discuss what happens if API pricing changes or access is revoked. ##### Tier 2: Augmented Applications The middle tier comprises companies using third-party models enhanced with proprietary data, fine-tuning, or retrieval-augmented generation (RAG). There’s genuine technical work here, but the defensibility depends entirely on the data advantage. The critical question: Is the proprietary data truly unique, or could a competitor license or generate equivalent training data within 18 months? ##### Tier 3: Proprietary AI At the top end are companies with custom models trained on proprietary datasets, often with novel architectures or training approaches. These carry the highest potential defensibility, but also the highest execution risk. The critical question: Does the team have the depth to maintain and improve these systems, or is the capability concentrated in one or two individuals who could leave? Most AI companies we assess fall into Tier 1 or Tier 2. There’s nothing wrong with that but investors should price accordingly. #### Five Technical Red Flags These patterns appear repeatedly in AI companies that don’t survive technical scrutiny. None is automatically disqualifying, but each warrants deeper investigation. ##### 1. The Buzzword Gap Certain terms have become markers for technical imprecision. When we hear invented terminology that doesn’t map to standard ML concepts, or established terms like “reasoning” applied loosely to any multi-step process, it often signals that marketing has outpaced engineering. This doesn’t mean the founders are being deliberately misleading. Often, they’re translating genuine technical work into language they believe investors want to hear. But the translation reveals a gap between what the technology actually does and how it’s being positioned. What to ask: “Can you describe the model architecture without using analogies? What specific machine learning techniques are you using, and why those over alternatives?” ##### 2. The Validation Vacuum One of the clearest signals of technical maturity is how a company measures and reports model performance. Genuine AI teams obsess over metrics. They can tell you their precision, recall, F1 scores, or task-specific benchmarks without hesitation. They know where their models fail and have plans to address it. Companies with superficial AI implementations often struggle here. We’ve seen Series B companies claim to have “solved” entity resolution across unstructured datasets without any rigorous validation methodology. When pressed on accuracy, the answers become vague: “customers are happy” or “it works well in practice.” What to ask: “Show me your evaluation framework. What benchmarks do you use? What’s your current performance, and what are the known failure modes?” If the answer is unclear or deflects to anecdotal customer feedback, treat every accuracy claim with scepticism. ##### 3. The Demo-to-Production Gap Impressive demonstrations are easy. Production systems that work reliably at scale are hard. We regularly encounter companies with compelling demos that have never processed real customer data at volume. The prototype works beautifully on curated examples; the production system struggles with edge cases, latency requirements, and the messy reality of real-world data. Warning signs: - No production monitoring dashboards to show - Metrics reported only on test datasets, not live data - “We’re focused on R&D” after two or more years of operation - Customer count that hasn’t grown despite claimed product-market fit What to ask: “Can you walk me through your production monitoring? What does your error rate look like over the past 90 days? How do you handle cases where the model fails?” ##### 4. The Disappearing Data Moat “Proprietary data” is perhaps the most overclaimed moat in AI. Companies assert data advantages that don’t survive scrutiny. Common patterns we’ve observed: - Licensable data: The “proprietary” dataset is actually available for purchase or licensing from data providers - Replicable data: The data could be generated by a well-funded competitor within 12–18 months - Stale data: The data advantage existed historically but the collection mechanism is no longer unique - User-generated data without lock-in: The data comes from users who could equally contribute to a competitor’s platform True data moats are rare. They require data that is simultaneously valuable for model training, expensive or impossible to replicate, and continuously refreshed through mechanisms competitors can’t easily copy. What to ask: “If a competitor raised £20m specifically to replicate your data advantage, how long would it take them? What would prevent them?” ##### 5. The Key Person Illusion AI systems are complex, and the knowledge required to maintain them often concentrates in a small number of individuals. We’ve assessed companies where the entire model architecture existed primarily in one engineer’s head, with minimal documentation and no realistic succession plan. This creates acute risk: the departure of a single technical leader can leave a company unable to maintain, debug, or improve its core technology. Warning signs: - Technical documentation described as “in progress” or “on the roadmap” - Model training processes that only one person has successfully executed - Inability to answer technical questions without deferring to a specific individual - Recent departure of founding technical team members What to ask: “If your lead ML engineer left tomorrow, how long before someone else could retrain your models? Who else has successfully done it?” #### What Good Looks Like Red flags are useful, but investors also need to recognise genuine technical capability. Here’s what we look for in companies that survive rigorous technical due diligence. ##### Metrics Obsession Strong AI teams measure everything. They have dashboards showing model performance over time, can segment accuracy by use case or customer type, and actively track where their systems fail. They’re often more eager to discuss their weaknesses than their strengths—because they’re genuinely working to fix them. ##### Honest Capability Boundaries Mature technical teams are precise about what their technology can and cannot do. They’ll say “we handle X well, but Y is still a challenge” rather than claiming universal capability. This precision signals genuine understanding rather than marketing optimism. ##### Documented, Reproducible Systems In well-run AI teams, model training is a documented process that multiple team members have executed. There’s version control for models, not just code. Experiments are logged. The system could survive the departure of any individual. Uncomfortably, perhaps, but it would survive. ##### Scalability by Design Companies with genuine technical depth have thought about scale from the beginning. They can articulate their cost structure at 10x current volume. They’ve made architectural decisions that support growth, not just current operations. They know what breaks next and have plans to address it. ##### Data Strategy, Not Just Data Rather than claiming a static data moat, strong companies articulate an ongoing data strategy: how they continue to acquire valuable training data, how that data improves their models over time, and why their data acquisition mechanism is defensible. ##### Responsible AI Integration We increasingly view responsible AI practices as a signal of technical maturity. Companies that have thought seriously about bias, fairness, and failure modes tend to have more robust systems overall. It’s not just ethics—it’s engineering discipline. #### Building Technical Due Diligence Into Your Process Traditional due diligence frameworks need adaptation for AI-native companies. Based on our experience, we recommend: 1. Technical review before deep financial analysis. There’s limited value in detailed revenue modelling if the core technology doesn’t survive scrutiny. A focused technical assessment early in the process can save significant time and cost. 2. Direct technical access. Demos and pitch decks are insufficient. Request documentation, architecture diagrams, and access to technical leadership. Reluctance to share technical details (even under NDA) is itself a finding. 3. Independent technical perspective. Internal technical teams often lack specific ML/AI expertise, and may be too polite to challenge founders directly. External technical due diligence provides both expertise and independence. 4. Questions designed to reveal depth. The questions throughout this article are designed to distinguish genuine capability from confident presentation. The goal isn’t to catch founders lying, it’s to understand what you’re actually buying. #### Conclusion The AI investment landscape rewards those who can distinguish signal from noise. Every company claims differentiation; few can demonstrate it under technical scrutiny. This isn’t about being cynical toward AI. Genuine AI capabilities create enormous value and that’s precisely why rigorous assessment matters. The companies with real technology benefit from due diligence that separates them from the crowd. The investors who develop technical evaluation capabilities gain an edge in a market where most rely on demos and pitch decks. The cost of getting AI assessment wrong is significant: overvalued acquisitions, technology that doesn’t scale, and moats that evaporate when foundation models improve. The cost of getting it right is a few weeks of focused technical review. In a market full of AI claims, that seems like a reasonable trade. --- Agathon provides independent technical due diligence for investors evaluating AI companies. Founded by Dr Colin Kelly (PhD in Natural Language Processing, Cambridge), we combine academic rigour with practical experience building and assessing AI systems across multiple sectors. --- ### Best AI consultancies for 2026: navigating the agentic era - URL: https://agathon.ai/insights/best-ai-consultancies-for-2026-navigating-the-agentic-era - Published: 2025-12-27 - Categories: AI Strategy, Machine Learning, AI Consulting The AI consulting landscape has reached an inflection point. After years of strategy decks and proof-of-concept projects, 2026 is the year enterprises must actually ship, and the gap between firms that can deliver and those that merely advise is widening rapidly. The numbers tell a sobering story. While 88% of organisations now regularly use AI, only 6% qualify as “high performers” seeing meaningful EBIT impact. Gartner predicts over 40% of agentic AI projects will be cancelled by 2027 due to escalating costs, unclear ROI, or inadequate risk controls. The experimentation phase is giving way to accountability. This shift fundamentally changes what enterprises should look for in an AI consulting partner. Strategy alone is no longer sufficient. Implementation expertise has become the primary differentiator: the ability to navigate orchestration frameworks, design human-AI interfaces, and deploy agents that actually work in production. We’ve been writing about the best AI consulting firms since 2024, and our 2025 analysis identified the rise of boutique technical excellence as a defining trend. For 2026, that thesis has become even more defensible. The question for many organisations is no longer whether to invest in AI, it’s whether your consulting partner can actually build what you need. #### What’s changed since 2025 ##### The agentic AI reality check AI agents dominated the conversation in 2025. McKinsey’s November survey found 62% of organisations experimenting with agents but only 23% scaling them in production. That gap represents the core challenge: agents are genuinely transformative when they work, and genuinely expensive failures when they don’t. The technology has matured significantly. Anthropic’s Model Context Protocol (MCP) has emerged as the industry standard for tool integration, with 10,000+ active servers and adoption by OpenAI, Google, Microsoft, and AWS. The December 2025 donation of MCP to the Linux Foundation’s Agentic AI Foundation cemented its position as the universal protocol for AI-to-tool communication. But maturation brings complexity. Enterprises now face real decisions about orchestration frameworks (LangGraph vs CrewAI vs AutoGen), architecture patterns (single-agent vs multi-agent), and integration approaches that didn’t exist eighteen months ago. The frameworks themselves are evolving rapidly: LangChain now explicitly recommends LangGraph over its base library for agent orchestration. ##### The exhaustion factor Here’s what the industry rarely acknowledges: keeping up with AI is genuinely exhausting. In the past year alone, enterprises have had to evaluate: new orchestration frameworks (LangGraph, CrewAI, OpenAI Swarm), new protocols (MCP, A2A), new coding tools (Claude Code, Cursor, Windsurf), new architectural patterns (agentic RAG, GraphRAG, long-context RAG), and new compliance requirements (EU AI Act penalties reaching €35M or 7% of global turnover). That viral post about pivoting from prompt engineering to context engineering to agent design to multi-agent swarms to single-agent architecture, all in a single week, resonates because it captures a genuine experience. The pace of change is real. The cognitive load on internal teams is real. And the risk of chasing every new framework while shipping nothing is very real. This exhaustion creates legitimate demand for specialists who stay current so you don’t have to. The value proposition is sustained attention to a domain that moves faster than any internal team can reasonably track whilst also doing their day jobs. ##### From strategy to implementation The consulting industry’s own research confirms the shift. McKinsey notes that clients increasingly expect implementation rigour, not just strategic recommendations. Sixty-five percent of businesses using generative AI now prefer consultants who actively participate in implementation rather than handing off slide decks. This creates a problem for traditional strategy firms. The gap between recommending “deploy an agentic workflow” and actually building one that works in production is vast. It requires understanding of token economics, context window management, tool calling patterns, error handling, human-in-the-loop design, and a dozen other technical considerations that don’t fit neatly into a strategy document. The firms thriving in 2026 are those that can do both: think strategically about where AI creates value, and ship the systems that capture it. #### What enterprises actually need in 2026 Based on where organisations are succeeding — and failing — with AI implementation, five capabilities have emerged as essential: Agentic AI expertise that extends beyond hype. Not chatbots relabelled as agents, but genuine understanding of when agentic architectures add value, how to design human oversight, and how to measure ROI. The 40% project cancellation rate suggests many vendors are overselling agent capabilities; you need partners who can distinguish viable use cases from vendor enthusiasm. Framework and protocol fluency. Can they implement LangGraph? Do they understand MCP security implications? Have they deployed CrewAI in production? Technical currency matters. The difference between current practitioners and firms still running 2023 playbooks is the difference between shipping and stalling. Human-AI interface design. The most successful AI implementations aren’t fully autonomous, they’re thoughtfully designed collaborations between humans and AI systems. This requires product thinking, not just engineering: understanding workflows, identifying where AI augments rather than replaces, and designing interfaces that make AI capabilities accessible to non-technical users. Context engineering capability. The shift from prompt engineering to context engineering reflects a deeper understanding of how to get consistent results from large language models. It’s not about clever prompts; it’s about systematic approaches to providing models with the right information at the right time. Anthropic describes it as “product strategy in disguise” where every system prompt instruction is a product decision. Independent technical assessment capability. As AI investment accelerates, so does the need for objective evaluation. Whether you’re a VC assessing a portfolio company’s technical claims, a corporate development team evaluating an acquisition target, or a board seeking assurance on AI initiatives, the ability to distinguish genuine capability from well-packaged demos has become critical. #### A note on scope This list focuses on firms where deep technical expertise meets hands-on delivery: consultancies that can both advise and build. We haven’t included pure-play systems integrators (Accenture, Cognizant, Infosys) or the Big Four’s AI practices (Deloitte, PwC, EY, KPMG). Not because they lack capability, but because their model serves a different segment: large-scale enterprise transformation with significant implementation workforces. For organisations seeking that scale, those firms remain relevant options. This list is for those who prioritise depth over breadth. #### The best AI consultancies for 2026 ##### Bain & Company Bain’s Advanced Analytics Group has quietly built one of the more technically credible AI practices among the major strategy consultancies. Their 500+ data scientists and ML engineers represent genuine technical depth, bolstered by recent acquisitions of Australian AI firm Max Kelsen and Spanish specialist PiperLab. What distinguishes Bain is their “80/20” approach, a pragmatic focus on the implementations that drive business value rather than chasing every emerging capability. Their “State of the Art of Agentic AI Transformation” research, published in 2025, demonstrates sophisticated understanding of where agents create value and where they don’t. Strengths: Strategic consulting pedigree combined with real technical delivery; strong experimentation-at-scale methodology; pragmatic about what works versus what’s hyped. Best for: Large enterprises seeking AI transformation with clear business case discipline and C-suite alignment. Limitations: Premium pricing structures suited to large-scale engagements; less focused on rapid prototyping or early-stage exploration. ##### Cambridge Consultants Spun out of Cambridge University’s scientific ecosystem, Cambridge Consultants brings 800+ engineers and scientists to deep-tech challenges. Their work spans AI assurance frameworks, edge AI deployment, and applications in highly regulated industries including defence and life sciences. Their approach leans research-forward: they’re the firm you engage when the problem genuinely doesn’t have an off-the-shelf solution. Projects with the UK Ministry of Defence and Hitachi demonstrate their capacity for technically demanding, compliance-heavy implementations. Strengths: Genuine scientific depth; strong in regulated industries; edge AI and embedded systems expertise; AI assurance and safety frameworks. Best for: Organisations facing novel technical challenges requiring research-grade expertise, particularly in regulated sectors. Limitations: Research orientation means longer timelines than pure implementation shops; premium positioning may exceed requirements for standard enterprise AI applications. ##### OneSix (incl. Strong Analytics) The June 2024 merger of data engineering firm OneSix with ML/AI consultancy Strong Analytics created something increasingly rare: a full-stack partner spanning data infrastructure through to production AI systems. Their integrated team (composed of data engineers, data scientists, ML experts, AI engineers, and LLM Ops specialists) reflects the reality that successful AI implementation depends on solid data foundations. Strong Analytics’ heritage in financial services, pharmaceuticals, and technology, combined with OneSix’s Snowflake partnership and data platform expertise, positions them well for enterprises whose AI ambitions are bottlenecked by data readiness. Strengths: End-to-end capability from data engineering through AI deployment; strong Snowflake ecosystem expertise; LLM Ops and production AI experience; blue-chip client portfolio across mid-market and enterprise. Best for: Enterprises needing integrated data and AI transformation, particularly those with Snowflake investments or significant data infrastructure work required. Limitations: North America focused delivery; less established brand than larger consultancies for C-suite positioning conversations. ##### Brainpool AI Brainpool takes a different approach entirely: a curated network of 500+ academic AI experts available for sprint-based engagements. Rather than building a permanent consulting staff, they match PhD-level specialists to specific project requirements. This model works particularly well for organisations needing genuine research expertise for defined projects without committing to ongoing consulting relationships. Their governance-first approach appeals to enterprises navigating EU AI Act compliance and responsible AI requirements. Strengths: Access to academic expertise without academic timelines; sprint-based model suits defined projects; strong governance and responsible AI focus; flexible engagement structures. Best for: Organisations needing specialist expertise for specific technical challenges or research-grade feasibility assessment. Limitations: Network model means less continuity than retained consulting relationships; less suited to large-scale ongoing transformation programmes. ##### Agathon Full disclosure: this is us. We include ourselves because we believe we represent a model increasingly relevant for 2026 — and because we’ve been transparent about this practice in our 2024 and 2025 coverage. Agathon combines research-grade technical depth (our founder holds a PhD in Natural Language Processing from Cambridge, with prior mathematics and computer science training from Oxford) with over a decade of commercial deployment experience. We’re practitioners who build, not just advisors who recommend. Our services span innovation assessment, AI advisory, fractional CTO, and AI product development. Recent work includes human-AI interface discovery and build for a category-defining AI writing product currently in beta testing; feasibility assessment and technical deep dives underpinning a media relation organisation’s medium-term AI roadmap; and in-house workflow development using Claude Code: we stay lean and bring that discipline to client engagements. We also provide technical due diligence for investors, applying the same rigour required to build sophisticated AI systems to evaluating whether others have built them properly. Strengths: Research-grade NLP expertise combined with commercial deployment experience; strategic advisory to hands-on delivery; technical due diligence capability for investors. Best for: Organisations seeking hands-on technical expertise with strategic grounding, particularly for human-AI interface design or agentic architectures. Also: investors needing independent assessment of AI capabilities. Limitations: Boutique scale limits capacity for very large transformation programmes; selective about engagements. #### Making the right choice The firms above represent different models for different needs. Rather than generic recommendations, here’s a framework for matching your situation to the right type of partner: What’s your primary constraint? If it’s board/C-suite alignment and business case rigour you likely need a firm with strategy consulting pedigree. Bain’s combination of strategic credibility and technical depth serves this well. If it’s novel technical challenges in regulated environments you need research-grade expertise comfortable with compliance complexity. Cambridge Consultants’ scientific depth and regulatory experience fits here. If it’s data readiness blocking AI progress you need integrated data-through-AI capability, not just an AI specialist layered on top of broken data infrastructure. OneSix’s full-stack approach addresses this. If it’s specific technical expertise for defined projects you may not need an ongoing consulting relationship at all. Brainpool’s network model offers flexibility without long-term commitment. If it’s bridging strategy and implementation with hands-on senior expertise particularly for human-AI interface work or agentic architecture, that’s where we focus. If it’s independent technical assessment for investment decisions — whether you’re a VC evaluating a portfolio company, a PE firm conducting technical due diligence, or a board seeking assurance — you need a firm that can assess rather than just build. Our innovation assessment capability serves this need. Red flags to watch out for: - Firms that can only speak strategy or only speak implementation, but not both - “Agentic AI” offerings that are rebranded chatbots or RPA - Inability to discuss specific orchestration frameworks, MCP, or current architectural patterns - Reluctance to share concrete project outcomes or reference clients - One-size-fits-all methodologies that don’t adapt to your specific context #### The year ahead The common thread across all credible options for 2026: implementation capability has become non-negotiable. The era of AI strategy without AI delivery is ending. The exhaustion is real. The pace of change is unsustainable for most internal teams. But the opportunity for organisations that find the right partners and ship real systems has never been greater. The firms that thrive will be those that can both think clearly about where AI creates value and build the systems that capture it. The 40% agentic AI project failure rate Gartner predicts isn’t inevitable. It’s the consequence of mismatched capabilities, unclear requirements, and partners who oversell and underdeliver. Choose wisely. --- Agathon is an AI-native consultancy specialising in AI strategy, agentic architecture and AI technical due diligence. If you’re navigating the challenges of AI implementation in 2026, or need independent assessment of AI capabilities, we’d welcome a conversation. --- ### The business applications of reinforcement learning: why most enterprises are leaving billions on the table - URL: https://agathon.ai/insights/the-business-applications-of-reinforcement-learning-why-most-enterprises-are-leaving-billions-on-the-table - Published: 2025-12-05 - Categories: Machine Learning, Reinforcement Learning, AI Strategy Most companies deploying AI today are essentially building sophisticated calculators when they could be creating adaptive intelligence systems. While everyone chases the latest large language model headlines, reinforcement learning represents the most underexploited frontier in enterprise AI: a methodology that learns optimal strategies through environmental interaction rather than pattern matching on historical data. The uncomfortable truth? Most AI implementations barely scratch the surface of what's technically possible. Companies invest millions in supervised learning solutions that become obsolete the moment business conditions change, whilst reinforcement learning offers genuinely adaptive systems that improve performance autonomously. Yet few enterprises understand when and how to exploit this capability. #### The fundamental limitation of conventional AI in dynamic business environments Traditional machine learning approaches operate under a dangerous assumption: that the future resembles the past. Supervised learning models excel at pattern recognition but fail catastrophically when faced with novel scenarios or evolving conditions. This works for static problems like image classification but becomes a liability in dynamic business environments where optimal strategies must evolve continuously. Reinforcement learning operates on an entirely different paradigm. Rather than learning from historical datasets, RL agents engage in continuous experimentation with their environment, optimising long-term outcomes through systematic trial and reward accumulation. As researchers have demonstrated, this approach can solve complex problems that traditional AI simply cannot address—particularly those involving sequential decision-making under uncertainty. The distinction isn't merely technical; it's strategic. Companies using supervised learning are essentially building sophisticated retrospective analysis tools. Those implementing reinforcement learning are creating forward-looking optimisation engines that adapt to changing conditions autonomously. #### The three capability classes where RL transforms business operations ##### Accelerating design and product development beyond human limitations Mining companies are now exploring a greater range of mine designs than possible with other AI techniques, whilst automotive manufacturers use RL agents to test more ideas for regenerative braking in electric vehicles, optimising for noise, vibration, and heat simultaneously. This represents a fundamental shift from traditional computer-aided design to truly intelligent systems that explore design spaces autonomously. The technical sophistication here extends beyond simple parameter optimisation. These systems can navigate complex trade-offs across multiple objectives whilst discovering design solutions that human engineers would never consider. The capability to rapidly iterate through millions of design variations whilst learning optimal strategies from each experiment creates competitive advantages that traditional design processes simply cannot match. ##### Optimising complex operational systems with genuine intelligence Reinforcement learning's ability to solve complex problems gives it high potential for optimising operations, helping organisations identify optimal actions across value chains as events unfold. Transportation companies are optimising travel routes in real time, whilst food producers manage global distribution amid fluctuating demand using RL systems that adapt to changing conditions faster than human operators. The key insight: most "AI optimisation" solutions are actually sophisticated rule engines. True RL implementations create systems that discover novel optimisation strategies through environmental interaction, often finding solutions that surpass human-designed heuristics by substantial margins. ##### Enhancing customer interaction through adaptive personalisation The challenge of modern customer engagement lies not in collecting data, but in making optimal decisions across millions of micro-interactions. User preferences change frequently, making traditional recommendation systems obsolete quickly. However, RL systems can track reader return behaviours and construct systems using news features, reader features, and context features to optimise engagement dynamically. This goes far beyond A/B testing or collaborative filtering. Advanced RL implementations create personalisation engines that learn optimal interaction strategies for individual users whilst adapting to preference shifts in real-time—a capability that transforms customer experience economics. #### Industry-specific applications revealing technical potential ##### Financial services: From pattern recognition to strategic optimisation JPMorgan Chase has several publications on RL applications in finance, including optimising financial decisions such as portfolio management, trading, and risk management. The sequential nature of financial decision-making aligns perfectly with RL's core strengths, enabling systems that adapt to market conditions rather than simply recognising historical patterns. The sophistication here extends to risk-adjusted portfolio optimisation under changing market conditions—problems that supervised learning approaches cannot address effectively because they lack the sequential decision-making framework necessary for dynamic strategy adaptation. ##### Manufacturing and industrial automation with adaptive intelligence Factory tasks like picking devices from boxes and placing them in containers are now handled by robots training themselves with remarkable speed and precision. This represents the emergence of truly autonomous manufacturing systems that improve performance through experience rather than programming. Beyond simple automation, these implementations create manufacturing systems that optimise efficiency, quality, and throughput simultaneously whilst adapting to variations in materials, equipment performance, and production requirements—capabilities that traditional automation systems cannot achieve. ##### Healthcare applications requiring sequential decision optimisation The sequential nature of medical decision problems makes RL particularly suitable, with applications in lung cancer and epilepsy treatments, and deep RL treatment strategies for sepsis developed from medical registry data. These systems learn optimal treatment protocols through systematic analysis of patient responses rather than relying solely on historical treatment patterns. ##### Energy sector breakthroughs in autonomous optimisation Google achieved a 40% reduction in energy consumption by letting an RL model control the cooling of one of their live data centres, marking one of the first major applications of modern RL in the energy sector. This wasn't incremental improvement through better sensors or controls—it was a fundamental shift to adaptive optimisation that discovered cooling strategies human engineers had never considered. #### Implementation challenges that separate sophisticated from superficial adoption ##### The simulation requirement and digital twin complexity Unlike supervised learning, which requires historical data, RL systems must experience their environment directly. This means companies need sophisticated simulation capabilities or digital twins of their business processes. The technical challenge isn't just building these simulations—it's ensuring they capture the essential dynamics of real-world systems whilst remaining computationally tractable. Most companies underestimate this requirement, leading to implementations that work in simplified simulations but fail in complex real-world environments. The gap between proof-of-concept demonstrations and production-ready systems often reveals whether organisations possess the technical depth necessary for sophisticated RL deployment. ##### Data hunger exceeding traditional machine learning requirements Andrew Ng has noted that reinforcement learning's hunger for data exceeds even supervised learning, making it difficult to acquire sufficient data for RL algorithms. This fundamental challenge shapes deployment strategies and timelines, requiring companies to think systematically about data generation rather than collection. The implication: organisations must build RL implementations that can learn efficiently from limited environmental interaction whilst maintaining safety constraints—a technical challenge that demands sophisticated algorithmic choices and careful system design. ##### Enterprise readiness and production deployment complexity Current RL research focuses primarily on game-playing and simulated environments. Few software tools come with examples aimed at industry applications, and the gap between research demonstrations and production systems remains substantial. This creates both opportunity and risk for early adopters. #### Strategic recommendations for capability-focused leaders ##### Building organisational capability for adaptive intelligence The transition to reinforcement learning requires more than technical implementation—it demands fundamental shifts in how organisations approach problem-solving. Companies must develop capabilities in simulation, reward function design, and safety-constrained exploration whilst building teams that understand both the technical mechanisms and business applications of adaptive intelligence. ##### Identifying optimal use cases through technical lens Executives who understand RL's potential will be better positioned to find competitive edges, with many organisations implementing traditional technologies first before applying RL to achieve previously unattainable performance tiers. The key is recognising problems that require sequential decision-making under uncertainty rather than pattern recognition on historical data. ##### Future-proofing through intelligent automation architecture The convergence of reinforcement learning with other AI methodologies promises unprecedented opportunities for businesses willing to invest in adaptive intelligence rather than static automation. This requires architectural thinking about how RL systems integrate with existing business processes whilst providing the flexibility to evolve strategies as conditions change. #### The exploitation opportunity that most consultancies miss Reinforcement learning represents more than an algorithmic advancement, it embodies a fundamental shift towards genuinely intelligent systems that navigate uncertainty, optimise complex trade-offs, and continuously improve performance without human intervention. The technical potential extends far beyond what most organisations currently exploit through their AI implementations. The companies that will dominate the next decade won't be those with the most sophisticated data collection or the latest language models. They'll be the organisations that build adaptive intelligence systems capable of discovering optimal strategies through environmental interaction whilst adapting to changing conditions faster than competitors can respond. Most AI consultancies focus on implementing commodity solutions using standard frameworks. The real opportunity lies in building systems that exploit the full technical potential of reinforcement learning to create adaptive advantages that competitors cannot easily replicate. If you're ready to build AI solutions that exploit full technical potential rather than implementing basic features, you should contact us today. #### References - O'Reilly. (2019). Practical applications of reinforcement learning in industry. - Cinelli, M. (2023). Reinforcement Learning in business: A sneak-peak on the applications of one the most promising AI methods. Medium. - IBM Data Science in Practice. (2022). Reinforcement Learning: The Business Use Case, Part 1. Medium. --- ### Finding a trusted AI consulting partner: beyond the marketing veneer - URL: https://agathon.ai/insights/finding-a-trusted-ai-consulting-partner-beyond-the-marketing-veneer - Published: 2025-10-30 - Categories: AI Consulting, AI Strategy, Fractional CTO Most AI consultancies are building yesterday's solutions with tomorrow's buzzwords. They're implementing chatbots when you need agentic systems, deploying basic ML when you need sophisticated neural architectures, and offering strategy when you need executable technical vision. The brutal truth? The AI consulting landscape is saturated with firms that treat artificial intelligence as a commodity service rather than a transformative capability requiring deep technical sophistication and strategic foresight. #### The technical depth deficit The fundamental challenge isn't finding an AI consultant. It's finding one who understands the vast chasm between surface-level implementations and genuine technical exploitation. Research demonstrates that businesses with proper AI guidance meet or exceed their expectations, whilst well-executed AI implementations can deliver 3.5X returns on investment. Yet most consulting engagements barely scratch the surface of what's technically possible. Consider the difference between deploying a basic recommendation engine versus building a multi-modal system that combines collaborative filtering, deep learning embeddings, and real-time contextual reasoning. Both solve the same business problem. Only one exploits the full technical potential. The gap exists because most consultants operate from implementation playbooks rather than first-principles technical understanding. They know how to deploy existing solutions but lack the depth to architect novel approaches that unlock unprecedented capabilities. #### Distinguishing technical sophistication from consulting theatre ##### Beyond the buzzword avalanche The market is awash with firms promising "AI transformation" whilst delivering glorified automation scripts. True technical sophistication reveals itself through specific architectural choices, not marketing vocabulary. Advanced practitioners discuss model parallelisation strategies, not just "machine learning deployment." They design multi-agent orchestration systems, not just "AI workflows." They architect semantic reasoning pipelines, not just "data processing automation." ##### Implementation depth indicators Genuine technical expertise manifests in architectural decisions that most firms cannot even articulate. Look for consultants who can explain why they chose specific attention mechanisms over alternatives, how they optimised inference latency for production deployment, or what trade-offs they made between model accuracy and computational efficiency. The critical differentiator lies in their ability to architect systems that exploit capabilities others miss entirely. This isn't about using more sophisticated tools, it's about understanding how to combine existing tools in ways that create emergent capabilities. #### The strategic assessment framework ##### Technical capability validation beyond portfolios Portfolio evaluation requires looking beneath surface-level case studies to understand underlying technical approaches. The most revealing question isn't "What did you build?" but "Why did you make specific architectural choices, and what alternatives did you consider?" Advanced consultants can articulate the technical reasoning behind their design decisions. They understand not just what works, but why it works and under what conditions it might fail. ##### Systems thinking versus point solutions The most sophisticated AI implementations emerge from systems-level thinking rather than point solution deployment. This means understanding how AI components interact with existing infrastructure, data pipelines, and business processes to create compound value. Researchers have shown that 70% of AI efforts must focus on people and process transformation, not just algorithms. The consulting firms that grasp this integration challenge are the ones capable of delivering transformative rather than incremental outcomes. ##### Commercial reality grounding Technical sophistication without commercial awareness produces impressive demonstrations that fail in production. The best consultants combine deep technical capability with acute understanding of operational constraints, regulatory requirements, and economic realities. This manifests in their ability to design solutions that are simultaneously technically advanced and commercially viable—exploiting cutting-edge capabilities whilst remaining deployable within real-world constraints. #### Red flags in technical positioning ##### Technology-first approaches without capability mapping Beware consultants who lead with specific technologies rather than capability requirements. The firms pushing "ChatGPT integration" or "Claude deployment" without first understanding your specific use case are treating AI as a commodity rather than a capability to be exploited. Advanced practitioners start with capability requirements and work backwards to optimal technical approaches. They might recommend completely different architectures for seemingly similar problems because they understand the nuanced technical requirements. ##### Lack of architectural sophistication Most consulting engagements fail because they implement basic solutions to complex problems. The critical indicator is whether your consultant can explain the architectural principles underlying their recommendations, not just the implementation steps. #### Strategic partnership versus transactional delivery ##### Beyond implementation towards capability building The most valuable consulting relationships are those that build internal capability whilst delivering external results. This requires consultants who can simultaneously architect sophisticated solutions and transfer knowledge effectively. Top AI consulting firms ensure not just successful deployment but sustainable internal capability development. They design systems that your team can understand, modify, and extend rather than black boxes requiring perpetual external support. ##### Long-term technical evolution planning AI capabilities evolve rapidly. The consulting partners who provide lasting value are those who architect solutions that can evolve with advancing technical capabilities rather than requiring complete replacement. This forward-looking approach requires deep understanding of both current technical constraints and likely future developments, the kind of insight that comes from genuine technical leadership rather than implementation following. #### The selection imperative The choice of AI consulting partner represents a defining moment for organisations serious about exploiting artificial intelligence's transformative potential. Most firms will deliver functional implementations that solve immediate problems whilst leaving vast technical potential unexploited. The consulting partners who create lasting competitive advantage are those who combine technical depth with strategic insight, who understand both cutting-edge capabilities and commercial realities, who can architect solutions that exploit possibilities others cannot even perceive. Success demands moving beyond traditional procurement towards identifying partners who can unlock technical potential you didn't know existed. If you're ready to build AI solutions that exploit full technical potential rather than implementing basic features, you should contact us today. The difference between sophisticated capability exploitation and commodity implementation will determine whether your AI investments deliver transformational or merely incremental returns. --- ### The leader's guide to building AI aptitude in your organisation - URL: https://agathon.ai/insights/the-leaders-guide-to-building-ai-aptitude-in-your-organisation - Published: 2025-09-15 - Categories: AI Strategy, Responsible AI, AI Advisory Most organisations are building AI programmes the way they built websites in 1995: functional but exploiting perhaps 15% of what's actually possible. The difference between basic AI implementation and sophisticated capability exploitation is about fundamentally different thinking about what intelligence means in an organisational context. The real challenge isn't teaching your workforce to use ChatGPT. It's building the collective cognitive capacity to recognise where AI can create entirely new forms of competitive advantage that your competitors haven't even conceptualised yet. #### Defining AI aptitude in the modern workplace AI aptitude extends far beyond digital literacy or prompt engineering. Research from Warwick Business School identifies six leadership capabilities essential for AI-driven environments: experimental mindset, empathetic leadership, ethical reasoning, cross-functional collaboration, data fluency, and pragmatic innovation. What's fascinating is how these capabilities compound. Organisations that develop all six simultaneously see exponentially greater returns than those focusing on individual skills. ##### The shift from linear to quantum thinking Traditional business thinking operates linearly: identify problem, deploy solution, measure outcome. AI-enabled organisations operate quantum-style. They maintain multiple parallel hypotheses about reality and collapse them into actionable insights through continuous experimentation. Leaders who understand this distinction create environments where AI capabilities evolve organically rather than being imposed top-down. ##### Essential components of organisational AI readiness True AI readiness requires what researchers call "semantic awareness" — the ability to understand how meaning flows through data systems and impacts decision-making. When leaders develop this awareness, they start seeing opportunities to exploit AI capabilities invisible to traditionally-minded competitors. #### The leadership imperative for AI transformation The most sophisticated AI implementations emerge from organisations where leadership has genuinely internalised the difference between automation and augmentation. Automation replaces human tasks; augmentation amplifies human judgment. The former creates efficiency gains; the latter creates new categories of competitive capability. ##### Leading through cultural transformation Netflix's Reed Hastings exemplifies this approach — using AI experimentation extensively not just for content recommendations, but for organisational decision-making itself. The cultural shift means becoming comfortable with AI-mediated reality as the baseline for strategic thinking. Leaders who successfully navigate this transformation create what might be called "hybrid intelligence": seamless integration between human intuition and machine insight. ##### Overcoming resistance and building buy-in Research from Warwick Business School reveals that over 50% of employees expect AI to replace their roles within the next year. This reflects a fundamental misunderstanding of how sophisticated AI implementations actually work. The most effective leaders reframe this narrative entirely, positioning AI as expanding human capability rather than replacing it. #### Strategic frameworks for developing AI competency The organisations that exploit AI's full potential ask: "How does intelligence actually flow through our organisation, and where are the bottlenecks that prevent us from acting on what we know?" This approach reveals implementation opportunities that use-case thinking completely misses. ##### Creating a culture of experimentation and learning IBM's approach through their AI Ethics Board demonstrates sophisticated thinking about experimentation within ethical constraints. Rather than treating ethics as compliance overhead, they've embedded ethical reasoning into their experimental methodology. This creates a sustainable framework for pushing boundaries while maintaining trust, essential for long-term AI capability development. ##### Cross-functional collaboration and governance structures Forbes research shows that 75% of CEOs mention cross-functional collaboration when setting AI strategy, but most implement this as committee structures rather than genuine integration. The difference is crucial: committees coordinate; integration creates emergent intelligence that exceeds the sum of individual contributions. #### Essential leadership skills for the AI-driven workplace The leadership capabilities required for AI exploitation are qualitatively different approaches that become necessary when dealing with systems that can process information and generate insights at superhuman scale. ##### Data-driven decision making and strategic analysis YouTube's Susan Wojcicki exemplified sophisticated data-driven leadership by using analytics to continuously reshape the platform's fundamental architecture, developing intuitive understanding of how data patterns translate into strategic possibilities. Leaders with this capability see opportunities for AI implementation that remain invisible to metrics-focused managers. ##### Emotional intelligence in human-AI collaboration Microsoft's Satya Nadella's approach to AI augmentation demonstrates how emotional intelligence becomes more crucial, not less, in AI-rich environments. When machines handle routine cognitive tasks, human judgment becomes concentrated in areas requiring empathy, context interpretation, and ethical reasoning. ##### Ethical leadership and responsible AI deployment Fujitsu's international AI ethics research team represents sophisticated thinking about embedding ethical reasoning into technical architecture. This approach creates AI systems that can navigate complex ethical terrain autonomously while maintaining alignment with human values, a capability that becomes essential as AI systems gain more autonomous decision-making authority. #### Building organisational AI literacy AI literacy programmes that focus on tool usage miss the fundamental challenge: developing the frameworks necessary to think alongside AI systems. This requires understanding how AI processes information differently from humans and where those differences create opportunities for hybrid intelligence approaches. ##### Comprehensive AI literacy training programmes The most effective AI literacy programmes start with semantic representation: teaching people how to think about meaning in ways that AI systems can process and enhance. This is about developing an intuitive understanding of how human concepts translate into computational processes and back into human-actionable insights. ##### Fostering curiosity and adaptive capacity Organisations that successfully build AI aptitude cultivate what researchers call "productive uncertainty": comfort with not knowing exactly how AI systems reach their conclusions while maintaining confidence in their ability to evaluate and act on AI-generated insights. This cognitive stance enables continuous learning and adaptation as AI capabilities evolve. #### Managing the human dimension of AI integration Humans excel at pattern recognition, contextual reasoning, and ethical judgment while AI excels at processing vast information sets and identifying statistical relationships. The integration challenge is creating workflows that exploit both capabilities without forcing either into unsuitable roles. ##### Addressing displacement fears and job evolution The most sophisticated AI implementations don't replace jobs: they create new categories of human-AI collaborative roles that didn't previously exist. Leaders who understand this help their teams develop capabilities that become more valuable as AI systems become more sophisticated. This requires understanding the fundamental cognitive differences between human and artificial intelligence. ##### Change management strategies for AI adoption Successful AI adoption requires what might be called "cognitive change management": helping people adapt their thinking patterns rather than just their processes. This involves developing comfort with AI-mediated decision-making while maintaining human agency over strategic choices. Leaders who master this approach create organisations that can continuously evolve their AI capabilities without losing institutional coherence. #### Future-proofing your organisation's AI journey The AI landscape evolves so rapidly that specific technical skills become obsolete within months. The sustainable approach focuses on developing meta-capabilities — the ability to quickly understand and exploit new AI capabilities as they emerge. This requires understanding the fundamental principles underlying AI development rather than memorising current implementations. Building genuine AI aptitude isn't about deploying the latest models or hiring more data scientists. It's about fundamentally reconceptualising how intelligence operates within your organisation and creating the cognitive infrastructure necessary to exploit AI's full potential. Most organisations implementing AI today are like companies building their first websites: functional, but missing the architectural sophistication necessary for genuine competitive advantage. The leaders who understand this distinction are building capabilities that their competitors won't recognise until it's too late to catch up. They're not just using AI to optimise existing processes; they're using AI to discover entirely new categories of strategic possibility. If you're ready to build AI capabilities that exploit technical potential others miss rather than implementing basic features, you should contact us today. #### References - Warwick Business School. (2024). Six leadership skills you need to make the most of AI. - IMD Business School. (2025). A real leader's guide to AI. - Trinity College Career & Life Design Center. (2024). Build AI Aptitude in Your Organization as a Leader. - Townsend, Stewart. (2024). Thrive as a Leader in the AI-Driven Workplace: Essential Skills You Need. Medium. --- ### AI-enhanced scenario planning: techniques for modern boardrooms - URL: https://agathon.ai/insights/ai-enhanced-scenario-planning-techniques-for-modern-boardrooms - Published: 2025-09-08 - Categories: AI Strategy, Generative AI, AI Agents Most boardrooms are playing with toys while calling it strategic foresight. The typical AI-powered scenario planning implementation today extrapolates historical data, applies simplistic trend analysis, and generates neatly packaged "scenarios" that are little more than glorified forecasts with error bars. This fundamental misunderstanding wastes both computational resources and executive attention. True AI-enhanced scenario planning doesn't just predict what might happen, it systematically explores the boundaries of what could happen, identifies hidden relationships between seemingly unrelated factors, and reveals strategic blindspots that conventional approaches miss entirely. #### The evolutionary gap in boardroom scenario planning Traditional scenario planning emerged in the 1960s with Royal Dutch Shell's pioneering work, helping them navigate the 1970s oil crisis. For decades, the methodology remained largely unchanged: gather experts, identify driving forces, build plausible narratives, and stress-test strategies. Today's business environment demands more. Interconnected global systems, exponential technological change, and unprecedented volatility have rendered traditional approaches dangerously inadequate. Yet most boardrooms haven't truly evolved beyond these decades-old methodologies—they've merely digitised them. The strategic importance of scenario planning at board level isn't just about predicting futures but systematically exploring possibility spaces to build adaptive, resilient organisations. This requires moving beyond simplistic forecasting to sophisticated simulation of complex adaptive systems, something traditional approaches fundamentally cannot deliver. #### The architectural sophistication gap in AI scenario systems The critical technical distinction between basic and advanced AI scenario planning lies in their architectural approaches: Basic implementation: Trains models on historical data, applies straightforward extrapolation, and generates narratives that fundamentally remain bound by past patterns. Advanced implementation: Leverages multi-layered, heterogeneous systems that combine: - Agent-based models simulating complex adaptive systems - Causal inference engines identifying non-obvious relationships - Generative models creating truly novel scenarios beyond historical patterns - Adversarial testing frameworks systematically challenging generated scenarios As Brynjolfsson and McAfee observe in their analysis of general-purpose technologies, the most powerful AI implementations are those that catalyse "waves of complementary innovations and opportunities." This principle applies directly to scenario planning systems, where architectural sophistication creates entirely new strategic capabilities. #### Beyond pattern recognition: Technical foundations for advanced scenario exploration ##### Causal architecture vs statistical correlation Most scenario planning tools rely heavily on statistical correlation. This fundamentally limits their utility for genuine strategic exploration. According to research published in the International Journal of Forecasting, models that fail to incorporate causal mechanisms perform poorly when confronted with structural breaks or regime changes; precisely the conditions boards need to prepare for. The technical solution requires implementing causal inference engines that go beyond correlation to identify causal mechanisms. These systems construct directed acyclic graphs (DAGs) representing causal relationships and enable intervention modelling—a capability critical for examining "what if" scenarios that have no historical precedent. ##### Complex adaptive system simulation Conventional approaches treat economies, markets, and organisations as complicated but ultimately predictable systems. This is a category error. These are complex adaptive systems where emergent behaviours, non-linear dynamics, and feedback loops dominate. Advanced scenario planning requires agent-based modelling architectures that simulate individual actors' behaviours and their complex interactions. This approach can reveal unexpected, emergent patterns and phase transitions impossible to identify through traditional forecasting methods. ##### Counterfactual generation and testing The most sophisticated scenario planning systems systematically generate counterfactuals: alternative histories that could have occurred under different conditions. This isn't mere speculation but a rigorous technical approach to understanding causal mechanisms and exploring the boundaries of possibility spaces. Implementing effective counterfactual generation requires: 1. Structured causal models that formally represent intervention effects 1. Generative adversarial networks trained to produce plausible counterfactuals 1. Validation frameworks that test counterfactual coherence and plausibility #### Implementation framework: Technical components of advanced scenario systems ##### Data integration and knowledge representation Superior scenario planning begins with superior knowledge representation. While basic systems rely on structured data and simple ontologies, advanced implementations integrate: - Heterogeneous data sources spanning structured, semi-structured, and unstructured data - Knowledge graphs representing entities and relationships with formal ontologies - Temporal logic frameworks capturing evolving relationships over time - Uncertainty representation mechanisms for explicitly modelling confidence levels The technical challenge lies in creating unified semantic representation layers that can integrate these diverse knowledge structures while maintaining their richness. ##### Pattern identification and scenario generation Pattern identification in advanced systems goes far beyond statistical trend analysis. The technical architecture should include: - Multimodal pattern detection across numeric, textual, and visual data - Anomaly detection frameworks identifying weak signals of emerging trends - Network analysis algorithms revealing hidden relationship structures - Temporal pattern mining identifying evolving dynamics Scenario generation then leverages these identified patterns through: - Generative models trained on diverse scenarios but constrained by causal models - Narrative generation engines creating coherent, causally consistent stories - Diversity optimisation algorithms ensuring scenario coverage across possibility spaces ##### Probabilistic assessment and prioritisation The most sophisticated scenario planning systems implement formal probabilistic reasoning frameworks, including: - Bayesian networks representing conditional dependencies between variables - Monte Carlo simulation for exploring parameter spaces - Explicit representation of epistemic uncertainty (what we don't know) vs aleatory uncertainty (inherent randomness) These technical approaches enable meaningful prioritisation that goes beyond simplistic "high/medium/low" impact and likelihood assessments to create nuanced understanding of possibility spaces. #### Ethical considerations and technical limitations The advancement of AI-enhanced scenario planning introduces specific technical challenges that require careful consideration: ##### Algorithmic bias and representational limitations Research by Kahneman, Sibony, and Sunstein on algorithmic decision-making highlights how AI systems can amplify existing biases in training data. In scenario planning, this manifests as systematically overlooking certain types of futures or overweighting others. The technical solution involves: - Adversarial testing frameworks specifically designed to identify bias - Diverse ensemble methods incorporating multiple model architectures - Explicit representation of model limitations and assumptions ##### Explainability requirements for strategic credibility For scenarios to influence board-level decisions, they must be explainable. This creates a fundamental tension between model sophistication and interpretability. Advanced implementations address this through: - Hybrid architectures combining interpretable components with black-box models - Post-hoc explanation frameworks generating human-understandable rationales - Narrative generation techniques that transform complex model outputs into coherent stories ##### Managing generative hallucination As the World Economic Forum's Future of Jobs Report notes, generative AI systems can produce plausible but factually incorrect outputs. In scenario planning, this creates a critical risk of exploring scenarios that appear plausible but violate fundamental constraints. Mitigating this requires: - Constraint satisfaction frameworks that enforce physical, economic, and logical consistency - Fact-checking mechanisms that validate generated content against established knowledge - Uncertainty quantification techniques that explicitly represent confidence levels #### Organisational readiness for AI-enhanced scenario planning ##### Technical infrastructure requirements Implementing advanced scenario planning requires specific technical infrastructure: - Distributed computing environments supporting complex simulations - Versioning systems tracking scenario evolution and decision rationales - Integration frameworks connecting scenario outputs to strategic planning processes - Visualisation capabilities presenting complex multidimensional scenarios effectively ##### Skills and capabilities needed The most common implementation failure is treating AI-enhanced scenario planning as either a purely technical or purely strategic challenge. Success requires multidisciplinary teams combining: - Data science expertise for model development and validation - Domain expertise for scenario interpretation and constraint definition - Strategic thinking for connecting scenarios to organisational decisions - Technical architecture skills for designing integrated systems #### Future directions: The untapped potential of AI in scenario planning ##### Large language models as scenario generators and critics Current research into large language models (LLMs) points to their potential not just as scenario generators but as sophisticated critics that can identify inconsistencies, blindspots, and implausibilities in generated scenarios. Implementing this capability requires going beyond simple prompt engineering to develop: - Fine-tuned LLMs specifically trained on scenario evaluation - Structured prompting architectures that systematically probe scenario coherence - Ensemble approaches combining multiple LLMs with different training objectives ##### Continuous learning systems for adaptive scenario planning The most sophisticated scenario planning systems implement continuous learning loops that: - Systematically compare scenario projections against real-world outcomes - Identify systematic biases in scenario generation - Automatically refine models based on observed performance - Maintain explicit representations of model uncertainty As Makridakis et al. noted in their analysis of forecasting competitions, models that systematically learn from their errors significantly outperform static approaches over time. #### Conclusion: From scenario consumption to scenario intelligence Most boardrooms remain consumers of scenarios rather than developers of genuine scenario intelligence. The distinction is critical: scenario consumption treats the future as something to be predicted, while scenario intelligence treats it as a design space to be explored and shaped. Building truly effective AI-enhanced scenario planning requires moving beyond the simplistic application of machine learning to historical data. It demands sophisticated architectural approaches that combine causal inference, complex system simulation, and probabilistic reasoning within integrated technical frameworks. For boards looking to navigate unprecedented uncertainty, the difference between basic and advanced implementation isn't incremental—it's existential. Organizations that merely digitise traditional approaches will continue to be surprised by the future, while those that implement architecturally sophisticated systems will develop the scenario intelligence needed to thrive in complex, rapidly evolving environments. If you're ready to move beyond basic implementation to exploit the full technical potential of AI-enhanced scenario planning, you need partners who understand both the technical architecture and strategic implications of advanced systems. That's the difference between digitising the past and designing the future. #### References - Makridakis, S., Spiliotis, E., & Assimakopoulos, V. (2022). "The M5 Accuracy Competition: Results, findings and conclusions." International Journal of Forecasting. - Brynjolfsson, E., & McAfee, A. (2022). "The Business of Artificial Intelligence." Harvard Business Review. --- ### Cost-effective LLM implementation: when to fine-tune and when to prompt - URL: https://agathon.ai/insights/cost-effective-llm-implementation-when-to-fine-tune-and-when-to-prompt - Published: 2025-09-01 - Categories: LLMs, AI Strategy, Machine Learning #### The £100,000 question nobody's asking about LLMs Here's the uncomfortable truth: most companies are haemorrhaging money on LLM implementations because they're asking the wrong question. They debate GPT-4 versus Claude, obsess over benchmarks, and chase the latest model releases. Meanwhile, they're burning through compute budgets like venture capital in 2021. The real question isn't which model to use. It's whether you need a model at all versus clever prompting. And if you do need one, whether you're sophisticated enough to handle the operational complexity that comes with it. #### Why most companies are burning money on the wrong LLM strategy ##### The seductive trap of over-engineering Fine-tuning has become the default answer to every LLM challenge. Can't get the right output format? Fine-tune. Need domain-specific knowledge? Fine-tune. Want consistent responses? Fine-tune. This reflexive response ignores a fundamental reality: researchers have demonstrated that sparse models can achieve comparable accuracy to their dense counterparts whilst supporting larger batch sizes and improving throughput. Yet teams persist in building dense, over-engineered solutions that demand exponentially more resources without delivering proportional value. ##### When sophistication becomes stupidity Consider this: academic research shows that fine-tuning generally converges within 10 epochs, often reaching peak accuracy much sooner. For simpler tasks, models approach optimal performance after just one epoch. Yet organisations routinely over-train models, chasing marginal improvements that users will never notice. The irony? Studies indicate that few-shot prompting can match fine-tuned performance for many domain-specific tasks. You're essentially paying for a Ferrari to navigate speed bumps. ##### The hidden costs of computational vanity The MoE (Mixture of Experts) layer dominates execution time in LLM fine-tuning, accounting for up to 85% of computational overhead according to detailed profiling studies. Matrix multiplication operations within these layers become your primary cost centre. Every unnecessary fine-tuning iteration multiplies this burden. But the real killer isn't compute. It's the maintenance nightmare you've created. Version control, model drift, prompt degradation, and the endless cycle of retraining as base models evolve. You've traded a simple prompt management system for a complex ML operations pipeline. #### Understanding the fundamental trade-offs ##### Response quality versus response time Research reveals a critical insight: as batch sizes increase, LLM workloads transition from being memory-bound to compute-bound. This shift fundamentally alters your optimisation strategy. Small-batch, low-latency applications benefit from prompt engineering's minimal overhead. High-throughput scenarios demanding consistent quality might justify fine-tuning's upfront investment. ##### Model size versus maintenance burden Scientists have shown that distillation can reduce model parameters whilst maintaining acceptable performance. A distilled model generates predictions faster and requires fewer resources, though with some quality degradation. The question becomes: is a 5% accuracy improvement worth a 10x increase in operational complexity? Parameter-efficient fine-tuning (PEFT) offers a middle ground, updating only the most relevant parameters. But even PEFT requires sophisticated infrastructure and expertise that most teams lack. ##### Flexibility versus predictability Prompt engineering excels at flexibility. Change requirements? Update the prompt. New use case? Adjust the template. This agility comes at the cost of potential inconsistency and prompt drift. Fine-tuning delivers predictability but locks you into specific behaviours. Researchers discovered that fine-tuned models become vulnerable to overfitting, especially on straightforward tasks. You've optimised for yesterday's requirements whilst tomorrow's needs have already shifted. #### When prompting is your secret weapon ##### Domain-specific tasks that don't need a PhD Academic studies comparing fine-tuning versus prompt engineering in code review automation found that sophisticated prompting matched specialised models for most practical applications. The difference? Prompting required minutes to implement versus weeks of data preparation and training. ##### Rapid prototyping and iterative development Zero-shot and one-shot prompting enable immediate experimentation. You can test hypotheses, validate approaches, and iterate designs without committing to training infrastructure. This velocity advantage compounds over time, allowing teams to explore more solution spaces. ##### The art of prompt engineering as competitive advantage Prompt caching and versioning systems provide the governance benefits of traditional ML pipelines without the overhead. Sophisticated prompt management becomes your differentiator, not your model architecture. ##### Zero-shot and few-shot learning scenarios Research demonstrates that foundation models trained on massive datasets often possess latent capabilities that clever prompting can unlock. Few-shot prompting with carefully selected examples can achieve domain adaptation without any training. You're leveraging billions of dollars of pre-training investment for the cost of a well-crafted prompt. #### When fine-tuning becomes inevitable ##### Breaking through the prompt engineering ceiling Some tasks demand capabilities beyond prompting's reach. Highly specialised terminology, complex multi-step reasoning, or strict compliance requirements might necessitate fine-tuning. The key is recognising this ceiling through systematic evaluation, not assumption. ##### Handling proprietary knowledge and unique vocabularies When your domain involves concepts absent from public training data, fine-tuning becomes essential. Medical subspecialties, proprietary trading strategies, or internal technical documentation require models to learn fundamentally new associations. ##### Achieving consistent brand voice at scale Marketing and customer communication at scale demands unwavering consistency. Fine-tuning can embed brand guidelines, tone requirements, and stylistic preferences directly into model weights, ensuring every interaction aligns with corporate identity. ##### Performance optimisation for high-volume applications Studies show that sparse fine-tuning can improve throughput by supporting larger batch sizes whilst maintaining accuracy. For applications processing millions of requests, the efficiency gains justify the implementation complexity. #### The economics of each approach ##### True cost breakdown beyond compute hours Researchers developed analytical models demonstrating that LLM costs extend far beyond token pricing. A comprehensive cost function must include fixed costs, variable costs, and critically, the probability of success for your specific task. The economic model reveals surprising insights: superior accuracy of expensive models can justify greater investment through increased earnings, but not necessarily higher ROI. A cheaper model achieving 80% accuracy might deliver better returns than a premium model reaching 95%. ##### Engineering time: the forgotten expense Prompt engineering appears deceptively simple, but sophisticated implementations require substantial expertise. Prompt versioning, A/B testing frameworks, and quality monitoring systems demand engineering investment. However, this pales compared to fine-tuning's requirements: data pipeline construction, training infrastructure, hyperparameter optimisation, and ongoing model management. ##### Maintenance nightmares and version control Fine-tuned models create versioning challenges that compound over time. Each base model update potentially requires retraining. Dataset drift necessitates periodic refreshing. Model performance degradation demands constant monitoring. You've transformed a straightforward integration into an ongoing operational commitment. ##### Risk assessment and technical debt Every fine-tuned model represents technical debt. As foundation models evolve rapidly, your specialised variants risk obsolescence. Prompt-based approaches maintain compatibility with model upgrades, preserving your investment whilst benefiting from continuous improvements. #### Building your decision framework ##### The three-question litmus test Before defaulting to fine-tuning, answer these questions honestly: 1. Have you exhausted prompt engineering possibilities, including chain-of-thought reasoning, few-shot examples, and structured templates? 1. Can you quantify the performance delta between prompted and fine-tuned approaches in real-world conditions? 1. Do you possess the operational maturity to manage model lifecycle, versioning, and drift? If you answered no to any question, you're not ready for fine-tuning. ##### Mapping task complexity to implementation strategy Simple classification or extraction tasks rarely justify fine-tuning. Complex reasoning, creative generation, or highly specialised domains might warrant the investment. The key is matching implementation complexity to task requirements, not technical ambition. ##### Creating reversible decisions Start with prompting. Always. Build prompt management infrastructure that scales. Only when you hit demonstrable limitations should you consider fine-tuning. This approach preserves optionality whilst delivering immediate value. ##### When hybrid approaches trump purist solutions Research indicates that combining retrieval-augmented generation (RAG) with sophisticated prompting often outperforms fine-tuning alone. RAG provides domain-specific context without training overhead. This hybrid approach delivers accuracy improvements whilst maintaining operational simplicity. #### Technical implementation considerations ##### Infrastructure requirements and constraints Fine-tuning demands substantial infrastructure: multiple GPUs, extensive memory, and sophisticated orchestration. Profiling studies show that gradient checkpointing can reduce memory requirements but increases execution time. You're constantly trading one constraint for another. Prompt-based approaches require minimal infrastructure: API access, prompt storage, and basic versioning. The simplicity enables rapid scaling without architectural overhaul. ##### Data quality thresholds for fine-tuning Academic research emphasises that fine-tuning requires high-quality, labeled datasets. Poor data quality amplifies rather than corrects model deficiencies. Dataset curation, cleaning, and validation often consume more resources than training itself. ##### Prompt management systems and versioning Sophisticated prompt engineering demands robust management systems. Version control, A/B testing capabilities, performance tracking, and rollback mechanisms become essential. These systems, whilst simpler than ML pipelines, require thoughtful design and implementation. ##### Monitoring and evaluation metrics Both approaches demand comprehensive monitoring, but the metrics differ. Prompt-based systems focus on output consistency, response relevance, and prompt drift. Fine-tuned models require additional tracking of model performance degradation, dataset shift, and retraining triggers. #### Common pitfalls and how to avoid them ##### The premature optimisation syndrome Teams often fine-tune before establishing baseline performance through prompting. This premature optimisation wastes resources and obscures simpler solutions. Establish prompted baselines, document limitations, then evaluate fine-tuning's incremental value. ##### Ignoring prompt drift and model decay Prompts degrade over time as language patterns evolve and model behaviours shift. Similarly, fine-tuned models experience performance decay. Both require monitoring and maintenance, though prompt updates are considerably simpler to implement. ##### Underestimating human-in-the-loop requirements Neither approach eliminates human oversight. Prompt engineering requires continuous refinement based on output analysis. Fine-tuning demands data curation, quality assessment, and performance validation. Budget for ongoing human involvement regardless of approach. ##### The false economy of cheap solutions Choosing the cheapest model or minimal infrastructure seems economical but often backfires. Research shows that model selection significantly impacts ROI. A slightly more expensive model achieving higher accuracy might deliver superior returns through improved business outcomes. #### Future-proofing your LLM strategy ##### Preparing for model obsolescence Foundation models evolve rapidly. Today's state-of-the-art becomes tomorrow's legacy system. Prompt-based approaches adapt naturally to model upgrades. Fine-tuned variants require complete retraining, multiplying migration costs. ##### Building abstraction layers for flexibility Implement abstraction layers that separate business logic from model-specific implementations. This architecture enables model swapping without application changes, preserving flexibility as the landscape evolves. ##### The coming convergence of approaches Emerging techniques blur the distinction between prompting and fine-tuning. Prompt tuning, soft prompts, and adapter layers offer middle grounds. Position your architecture to leverage these hybrid approaches as they mature. ##### Why agility beats perfection The LLM landscape changes too rapidly for perfect solutions. Prioritise adaptability over optimisation. A good-enough solution today that evolves beats a perfect solution delivered after requirements change. #### The uncomfortable truth about LLM implementation Most organisations are wasting money on LLM implementations because they're optimising for the wrong metrics. They chase benchmark scores instead of business outcomes. They fine-tune for marginal improvements instead of exploiting existing capabilities through sophisticated prompting. The evidence is clear: sparse models match dense model performance whilst reducing costs. Few-shot prompting rivals fine-tuning for many applications. Prompt engineering delivers immediate value whilst preserving flexibility. Yet companies persist in building complex, expensive, brittle solutions. They confuse technical sophistication with business value. They mistake complexity for capability. The winning strategy isn't choosing between prompting and fine-tuning. It's understanding when each approach delivers maximum value. Start with prompting. Exhaust its possibilities. Document its limitations. Only then, with clear evidence and quantified benefits, consider fine-tuning. This isn't about being conservative. It's about being strategic. It's about exploiting the full potential of these technologies rather than implementing basic features wrapped in unnecessary complexity. The companies that win won't be those with the most sophisticated models. They'll be those that match implementation complexity to business requirements, that prioritise adaptability over perfection, and that understand the true economics of LLM deployment. If you're ready to build AI solutions that exploit full technical potential rather than implementing basic features, you should contact us today. #### References - IBM's comprehensive guide to RAG vs fine-tuning vs prompt engineering trade-offs - Google Developers' technical documentation on LLM fine-tuning, distillation, and prompt engineering - Stanford research paper on optimizing LLM usage costs with 40-90% reduction methods - Academic study on understanding performance and cost estimation for LLM fine-tuning --- ### AI governance: what business leaders need to know - URL: https://agathon.ai/insights/ai-governance-what-business-leaders-need-to-know - Published: 2025-08-25 - Categories: AI Strategy, Responsible AI, AI Advisory Most organisations treat AI governance like they treat fire drills: mandatory, performative, and utterly disconnected from daily operations. They're building compliance theatres whilst their competitors are building capability fortresses. The real governance gap isn't about missing policies. It's about missing the fundamental shift in how AI systems operate compared to traditional software. Here's the uncomfortable truth: your carefully crafted governance framework, lovingly adapted from IT best practices, is already obsolete. Why? Because you're governing deterministic systems in a probabilistic world. You're applying assembly-line quality control to systems that learn, adapt, and occasionally hallucinate. #### Understanding AI governance beyond the compliance checkbox The consultancies will sell you frameworks. The lawyers will sell you liability shields. The vendors will sell you "enterprise-ready" solutions with governance "built in". They're all missing the point. Real AI governance isn't about ticking boxes. It's about understanding that every AI system is essentially a compression algorithm for human judgment, complete with all our biases, blind spots, and brilliances. Governing AI means governing compressed human cognition at scale. ##### The difference between AI governance and traditional IT governance Traditional IT governance assumes predictability. Input A produces Output B. Every time. AI systems operate on probability distributions. Input A might produce Output B, or something surprisingly brilliant, or complete nonsense. The same model, with the same input, can produce different outputs based on temperature settings, random seeds, or the phase of the moon (technically, cosmic ray interference, but the moon sounds more poetic). This isn't a bug. It's the feature that makes AI valuable. But it means your governance model needs to shift from controlling outcomes to managing outcome distributions. Think less traffic lights, more jazz improvisation with guardrails. ##### Why "move fast and break things" breaks down with AI systems Silicon Valley's favourite mantra becomes Silicon Valley's biggest liability when applied to AI. When you "break things" with traditional software, you fix the bug and push an update. When you break things with AI, you might have trained your model on biased data, created feedback loops that amplify discrimination, or built systems that confidently generate plausible-sounding misinformation. Research from MIT catalogues over 750 distinct AI risks. That's not 750 ways the same thing can go wrong. That's 750 different failure modes, each requiring different governance approaches. The "ship now, fix later" mentality doesn't work when "later" involves retraining models, rebuilding trust, and potentially facing regulatory sanctions. ##### The hidden costs of ungoverned AI deployment The obvious costs (fines, lawsuits, reputation damage) are just the tip of the iceberg. The real costs lurk beneath: technical debt that compounds exponentially, shadow AI systems that proliferate like digital kudzu, and the opportunity cost of building on unstable foundations. According to recent studies, organisations are already spending 4.6% of their AI budgets on ethics and governance, expected to rise to 5.4% by 2025. But here's what they don't tell you: ungoverned AI typically costs 3-5 times more in remediation than governed AI costs in prevention. Pay now or pay later, with interest. #### The anatomy of effective AI governance Forget the consultancy pyramids and maturity matrices. Effective AI governance has three essential organs: decision rights that actually matter, risk frameworks that actually work, and metrics that actually measure. ##### Decision rights and the AI accountability vacuum The accountability vacuum in AI isn't accidental. It's structural. When a traditional system fails, you trace the code, find the bug, assign blame. When an AI system fails, you have a model trained by one team, fine-tuned by another, deployed by a third, and used by people who understand neither the training nor the deployment. IBM's research shows that 60% of C-suite executives claim they have "clearly defined gen AI champions" throughout their organisation. The other 40% are at least honest. But even those with champions face a fundamental problem: championship without authority is cheerleading. Real accountability requires what researchers call "sociotechnical" governance: understanding that AI systems are inseparable from the human systems that create, deploy, and use them. You can't govern the algorithm without governing the organisation. ##### Risk frameworks that actually work in practice Most AI risk frameworks are elaborate exercises in wishful thinking. They categorise risks that are easy to categorise, measure risks that are easy to measure, and ignore everything else. It's like securing your house by installing seventeen locks on the front door whilst leaving the windows open. Effective frameworks recognise that AI risks are emergent, not enumerable. They focus on system properties rather than component failures. They measure outcome distributions rather than individual outcomes. Most importantly, they acknowledge uncertainty rather than pretending it doesn't exist. The Data & Trust Alliance's approach offers a glimpse of what works: 22 metadata fields that provide essential information about data provenance. Not 200 fields that nobody will fill out. Not 2 fields that tell you nothing. Just enough structure to be useful without being burdensome. ##### The measurement problem: KPIs for responsible AI You can't manage what you can't measure, but with AI, you often can't measure what matters most. How do you quantify fairness? How do you metric trustworthiness? How do you KPI explainability? The answer isn't to measure everything. It's to measure the right things. Model performance metrics are table stakes. What matters are outcome metrics: disparate impact analysis, confidence calibration, adversarial robustness. These aren't just technical metrics. They're business metrics that happen to have technical implementations. #### Building your governance framework without bureaucratic paralysis The consultancies want to sell you a 200-page governance framework that nobody will read. The lawyers want to wrap everything in legal bubble wrap. The academics want peer review for every decision. Here's a radical alternative: build governance that people actually use. ##### The minimum viable governance structure Start with three things: clear ownership, clear boundaries, clear consequences. Everything else is elaboration. Ownership means one person (not a committee) is accountable for each AI system. Boundaries mean explicit limits on what the system can and cannot do. Consequences mean predetermined responses to boundary violations. Not "we'll investigate". Not "we'll form a committee". Actual, specific, predetermined responses. IBM's testing of Data & Trust Alliance standards showed a 58% reduction in data clearance processing time for third-party data. Not because they did less governance. Because they did focused governance. ##### Scaling governance with AI maturity Governance should grow with capability, not ahead of it. Premature governance is just bureaucracy. Delayed governance is just negligence. The key is coupling governance maturity to AI maturity. Using pre-trained models? Focus on deployment governance. Fine-tuning models? Add training governance. Building models from scratch? Full-stack governance. Each level builds on the previous, rather than starting from scratch. ##### The role of AI ethics committees (and why most fail) According to IBM research, 47% of organisations have established AI ethics councils. Most are about as effective as corporate sustainability committees: well-intentioned talk shops that produce recommendations nobody implements. Ethics committees fail for three reasons: wrong people (ethicists without technical knowledge or technicians without ethical training), wrong mandate (advisory without authority), and wrong timing (reviewing after the fact rather than designing from the start). The committees that work have three characteristics: multidisciplinary membership (including philosophers, anthropologists, and domain experts, not just technologists), executive mandate (reporting directly to C-suite with actual veto power), and embedded process (part of development workflow, not separate review process). #### Navigating the regulatory maze without losing competitive edge AI regulation is coming. In some jurisdictions, it's already here. The EU's AI Act, China's algorithmic regulations, the patchwork of US state laws. Treating these as compliance burdens is missing the opportunity. ##### The global patchwork of AI regulation Every jurisdiction wants to be the Brussels of AI regulation: setting standards the world must follow. None have quite managed it yet. The result is a regulatory patchwork that makes GDPR look simple. The smart approach isn't to wait for harmonisation (which won't happen) or to build to the lowest common denominator (which won't last). It's to build flexible governance that can adapt to multiple regulatory regimes without architectural rewrites. ##### Preparing for regulations that don't exist yet Recent research shows 27% of public companies cite AI regulation as a risk in SEC filings. They're preparing for regulations that don't exist yet. This isn't paranoia. It's pattern recognition. Every transformative technology follows the same regulatory arc: innovation, exploitation, scandal, regulation. We're somewhere between exploitation and scandal. The organisations that survive the transition are those that build governance before regulation forces them to. ##### Turning compliance into competitive advantage Australia Post offers a masterclass in this approach. They're using generative AI to handle 40-60% of customer calls, saving costs whilst improving service. But they're not just building capabilities. They're building transparent, governed capabilities that customers trust. Trust is the ultimate competitive advantage in AI. Not because customers care about your governance framework. Because governed AI produces more reliable, less biased, more explainable results. Compliance isn't a tax on innovation. It's an investment in sustainability. #### The human element: Governance beyond algorithms AI governance discussions inevitably devolve into technical specifications and legal requirements. Missing from most frameworks is the recognition that AI systems are sociotechnical systems. The human element isn't an add-on. It's fundamental. ##### Managing the sociotechnical complexity of AI systems Every AI system exists in a human context. Trained on human-generated data, deployed by human operators, used by human users, affecting human lives. Governing the algorithm without governing the human system is like tuning a piano whilst ignoring the pianist. Research from IBM and others consistently shows that AI failures are rarely purely technical. They're usually sociotechnical: technically correct systems used incorrectly, or technically incorrect systems compensated for by human operators. ##### The skills gap nobody talks about Everyone talks about the ML engineering skills gap. Nobody talks about the AI governance skills gap. Who on your team can explain a neural network's decision to a regulator? Who can translate between data scientists and domain experts? Who understands both the technical architecture and the business implications? These aren't technical roles or business roles. They're translation roles. And they're critically undersupplied. According to Gartner, 65% of data leaders cite governance as their top priority, but less than 10% have dedicated governance expertise. ##### Creating a culture of responsible innovation Culture eats strategy for breakfast, and it devours governance frameworks for dessert. You can have the world's best governance framework, but if your culture rewards shipping features over addressing risks, governance becomes theatre. The 2024 Edelman Trust Barometer reveals that 79% of respondents expect CEOs to speak out about ethical technology use. But speaking out isn't enough. Culture is built through actions, not words. What gets rewarded? What gets punished? What gets ignored? #### Common governance failures and how to avoid them Let's be honest about how governance actually fails. Not the dramatic failures that make headlines, but the mundane failures that accumulate into catastrophes. ##### The "pilot purgatory" trap Pilots are where AI initiatives go to die. Not because they fail (though many do), but because they succeed without scaling. The governance framework that works for a pilot breaks at production scale. The trap is treating pilots as technical experiments rather than governance experiments. Every pilot should test not just whether the AI works, but whether the governance scales. Can you maintain the same level of oversight with 1000x the volume? If not, you're not ready for production. ##### Shadow AI and the proliferation problem Shadow IT was annoying. Shadow AI is dangerous. When users can spin up AI capabilities through browser plugins, API calls, or SaaS integrations, your carefully crafted governance framework becomes a Maginot Line. MIT research identifies this as one of the 750 AI risks: ungoverned AI proliferation. The solution isn't to lock everything down (which just drives it further underground). It's to make governed AI easier to use than ungoverned AI. ##### When governance becomes innovation theatre Some organisations have turned governance into performance art. Ethics committees that meet quarterly to review initiatives that ship daily. Bias audits conducted on training data that's already in production. Risk assessments filed in drawers that nobody opens. This isn't governance. It's innovation theatre: the appearance of responsibility without the substance. It's worse than no governance because it creates false confidence. #### The economics of AI governance Let's talk money. Not the hand-wavy "trust is valuable" money, but actual, measurable, budgetable money. ##### The ROI of responsible AI IBM's research with the Data & Trust Alliance showed a 62% reduction in data clearance processing time for internally generated data. That's not a soft benefit. That's measurable productivity improvement. But the real ROI comes from risk mitigation. One biased lending algorithm can trigger millions in fines, lawsuits, and remediation costs. One hallucinating customer service bot can destroy years of brand building. Governance isn't a cost centre. It's insurance with positive returns. ##### Budgeting for governance without breaking the bank Current spending on AI ethics and governance averages 4.6% of AI budgets. That sounds reasonable until you realise most organisations dramatically underestimate their true AI spend. When you account for shadow AI, embedded AI in SaaS products, and AI-augmented processes, the real percentage is often less than 1%. The key isn't to spend more. It's to spend smarter. Governance tooling that integrates into existing workflows. Automated monitoring rather than manual reviews. Risk-based approaches that focus resources where they matter most. ##### The true cost of getting it wrong The true cost isn't the fine or the lawsuit. It's the compound effect: lost trust leading to reduced adoption, increased scrutiny leading to slower deployment, technical debt leading to higher maintenance costs. One major tech company (unnamed in research) spent three years and tens of millions rebuilding an AI system after governance failures. Not because regulators forced them to. Because the technical debt from ungoverned development made progress impossible. #### From principles to practice: Making governance operational Principles are poetry. Practice is prose. The gap between "we value fairness" and "here's how we measure and enforce fairness" is where most governance frameworks fail. ##### Embedding governance into development workflows Governance can't be a gate at the end of development. It needs to be embedded throughout. Not as bureaucratic checkpoints, but as integrated tooling and automated checks. Think of it like code linting for AI. Automated bias detection during training. Explainability requirements in model APIs. Drift detection in production. Make governance invisible to developers whilst making governance failures impossible to ignore. ##### The tooling ecosystem for AI governance The governance tooling ecosystem is fragmented, immature, and desperately needed. Most organisations cobble together solutions from multiple vendors, open-source projects, and internal development. The winners will be platforms that integrate governance into development rather than bolt it on after. Think GitHub Actions for AI governance: automated, integrated, and invisible until something goes wrong. ##### Monitoring and auditing in production Production is where governance goes to die. The model that was fair in testing becomes biased in production. The system that was accurate in validation starts hallucinating in deployment. Continuous monitoring isn't optional. But it's not just about model metrics. It's about outcome monitoring: are predictions calibrated? Are errors randomly distributed or systematically biased? Are edge cases becoming common cases? #### The future-proof governance strategy If predicting the future of AI is hard, predicting the future of AI governance is impossible. But we can build governance that adapts rather than breaks. ##### Preparing for AGI governance challenges Current governance frameworks assume narrow AI: systems that do one thing well. But capabilities are expanding rapidly. Today's chatbot is tomorrow's agent is next year's autonomous system. The organisations preparing for this transition are building capability-based rather than application-based governance. Instead of governing "the customer service bot", they're governing "systems that can make decisions affecting customers". The abstraction matters. ##### Adaptive governance for evolving capabilities Static governance for dynamic systems is a recipe for irrelevance. Your governance framework needs to evolve as fast as your AI capabilities. Not constantly changing, but constantly adapting. This means feedback loops: governance informing development, development informing governance. It means version control for governance, not just for code. It means treating governance as a product, not a project. ##### Building resilience into your AI governance model Resilient governance bends without breaking. It handles edge cases without paralysis. It adapts to new requirements without architectural rewrites. The key is building governance on principles rather than rules. Rules are brittle. Principles are flexible. Rules tell you what to do. Principles tell you how to think. In a rapidly evolving field, knowing how to think matters more than knowing what to do. Governance isn't about preventing failure. It's about failing safely, learning quickly, and improving continuously. The organisations that understand this distinction are the ones that will thrive in the AI era. Most organisations are using perhaps 10% of AI's true potential because they're either paralysed by governance concerns or reckless in ignoring them. There's a third way: governance as enablement, not enforcement. If you're ready to build AI solutions that exploit full technical potential rather than implementing basic features, whilst maintaining the governance and trust that ensures sustainable success, you should contact us today. #### References - IBM Institute for Business Value comprehensive guide to AI governance frameworks and best practices - Harvard Berkman Klein Center Ethics and Governance of AI Initiative research - Harvard Kennedy School Leading in Artificial Intelligence executive program on technology and policy - MIT Media Lab Ethics and Governance of Artificial Intelligence Fund research on AI policy and governance --- ### From guidelines to guardrails: operationalising AI ethics in product development - URL: https://agathon.ai/insights/from-guidelines-to-guardrails-operationalising-ai-ethics-in-product-development - Published: 2025-08-18 - Categories: Responsible AI, AI Strategy, AI Consulting Most AI ethics implementations fail because they're designed for documentation, not deployment. They exist primarily to reduce liability rather than to enhance capability. This pervasive misalignment has created a market of ethical AI solutions that masquerade as comprehensive while addressing only the most visible risks. The uncomfortable truth: organisations adopt ethical frameworks as a checkbox exercise rather than embedding them into the fabric of their AI products. The result is predictable: sophisticated AI capabilities remain underutilised due to fear, while the safeguards that would enable their responsible deployment remain superficial and ineffective. #### The implementation gap: why ethical principles fail in practice The debate about AI ethics has overwhelmingly focused on principles—the 'what' of AI ethics rather than practices, the 'how'. Morley et al. (2020) identified this exact issue in their review of publicly available AI ethics tools, noting that "awareness of the potential issues is increasing at a fast rate, but the AI community's ability to take action to mitigate the associated risks is still at its infancy." This principle-practice gap creates a perilous situation where organisations possess well-articulated ethical aspirations but lack the technical infrastructure to realise them. When researchers reviewed 84 ethical AI documents, they found remarkable convergence around principles like transparency, justice, non-maleficence, responsibility, and privacy—yet a glaring absence of technical specifications for implementing these values. The challenge isn't defining what ethical AI should look like, but operationalising these principles in the messy reality of development pipelines. Companies that successfully navigate this transition gain a critical competitive advantage: they can deploy more powerful, sophisticated AI capabilities while managing risks that would paralyse their competitors. #### Converting ethical abstractions into technical specifications The most advanced AI implementations require a methodical translation process that converts abstract ethical principles into concrete technical requirements. This is where most organisations falter. They have grand ethical statements but lack the technical expertise to encode these values into their systems. Effective operationalisation begins by deconstructing each ethical principle into discrete, measurable components. For example, rather than treating "fairness" as a monolithic concept, sophisticated implementations distinguish between group fairness (ensuring different demographic groups receive similar outcomes) and individual fairness (ensuring similar individuals receive similar outcomes). This decomposition creates the foundation for technical implementation. Leading organisations develop what Rakova et al. (2021) call "technical guardrails": programmatic constraints that enforce ethical boundaries during model training, validation, and deployment. These guardrails differ fundamentally from superficial ethics guidelines because they're executable, testable, and embedded directly into development workflows. A critical insight from research by Leslie (2019) for the Alan Turing Institute demonstrates that effective guardrails must operate across the entire AI lifecycle, from data collection through to ongoing monitoring and maintenance. This systems-thinking approach reveals why piecemeal ethical interventions fail; they optimise for local ethical concerns while ignoring system-wide vulnerabilities. #### Engineering ethics into the AI development lifecycle The most sophisticated AI organisations have moved beyond ethical principles as abstract concepts by embedding guardrails at every stage of development. This lifecycle approach creates multiple layers of protection while enabling deployment of more powerful capabilities. ##### Pre-development: ethical risk assessment Forward-thinking organisations begin with a structured ethical risk assessment before writing a single line of code. This assessment maps potential harms, identifies vulnerable stakeholders, and anticipates failure modes. Significantly, this isn't merely a philosophical exercise; it yields quantifiable risk metrics that inform technical design decisions. Research by Fjeld et al. (2020) reveals that effective risk assessments scrutinise not only the AI system itself but also its broader socio-technical context. Organisations that take this systems approach can identify cascade effects and externalities that narrow technical assessments would miss. ##### Development phase: technical constraints as guardrails During model development, ethical principles must be translated into computational constraints. Sophisticated implementations leverage a combination of: - Mathematical fairness constraints incorporated directly into model training - Transparency mechanisms that expose model decision boundaries - Adversarial testing frameworks that systematically probe for ethical vulnerabilities - Containment architectures that limit system authority and autonomy What distinguishes advanced implementations is their focus on making these constraints robust to deployment pressures. As Morley et al. (2020) note, ethical guardrails must withstand the inevitable tensions between ethical ideals and commercial imperatives. This requires designing constraints that are difficult to circumvent, even when facing production pressures. ##### Deployment: continuous ethical monitoring The most advanced systems implement what Raji et al. (2020) call "ethical canaries"—continuous monitoring mechanisms that detect potential ethical violations in production. These systems move beyond simple dashboard metrics to implement sophisticated anomaly detection algorithms that can identify emerging ethical risks before they manifest as harms. Critical to this approach is establishing clear thresholds for intervention. When a guardrail is breached, what happens? Sophisticated implementations define an escalation pathway with pre-determined intervention protocols. This removes the ambiguity that often paralyses organisations when facing ethical edge cases in production. #### Governance structures that enable ethical innovation Effective governance is perhaps the most overlooked aspect of operationalising AI ethics. Research by Rakova et al. (2021) found that without appropriate governance structures, even the most sophisticated technical guardrails can be circumvented or eroded over time. The most effective governance models share several characteristics: - Cross-functional ethical review boards with real decision-making authority - Clear escalation pathways for ethical concerns - Documented accountability mechanisms with meaningful consequences - Transparent documentation of ethical decisions and trade-offs What distinguishes sophisticated governance from performative governance is authority. As Leslie (2019) notes, ethical governance must have "teeth"—the power to delay or redirect AI initiatives that violate ethical guardrails, even when those initiatives have strong business cases. #### Measuring ethical performance beyond compliance Most organisations approach ethical measurement through the narrow lens of compliance—meeting minimum standards to avoid regulatory or reputational harm. This approach fundamentally misunderstands the competitive advantage of operationalised ethics. Sophisticated implementations measure ethical performance against a broader set of metrics: - Capability utilisation: How much of the AI system's potential capability can be safely deployed? - Stakeholder trust: How does ethical performance affect user, customer, and regulator trust? - Innovation velocity: How quickly can new capabilities be deployed with appropriate guardrails? These metrics reframe ethics not as a constraint on innovation but as an enabler of sustainable capability development. Organisations that excel at ethical operationalisation can deploy more sophisticated AI capabilities precisely because their guardrails enable responsible experimentation. #### Challenges in practical implementation Research by Morley et al. (2020) identifies several persistent challenges in operationalising AI ethics: - Resource constraints: Ethical guardrails require additional engineering effort and computational resources - Competing priorities: Ethical objectives may conflict with each other (e.g., transparency vs. privacy) - Evolution of standards: Ethical expectations change over time, requiring adaptable guardrails - Verification difficulties: Proving the effectiveness of ethical guardrails remains technically challenging These challenges help explain why so many organisations default to superficial ethics implementations. The technical work of building robust guardrails is difficult, expensive, and doesn't yield immediate returns. Yet it's precisely this difficulty that creates strategic advantage for organisations willing to make the investment. #### From theoretical frameworks to operational reality The path from ethical principles to operational guardrails isn't straightforward, but it's increasingly necessary as AI capabilities advance. Organisations that can successfully navigate this transition gain a critical competitive advantage: they can deploy more powerful, sophisticated AI capabilities while managing risks that would paralyse their competitors. The key insight is that ethical operationalisation isn't just about risk mitigation; it's about capability enablement. When implemented properly, ethical guardrails create the conditions for responsible innovation by establishing clear boundaries within which experimentation can safely occur. This represents a fundamental shift in thinking about AI ethics—moving from abstract principles that constrain innovation to operational guardrails that enable it. Organisations that make this shift can deploy sophisticated AI capabilities that others cannot, precisely because they've developed the technical infrastructure to deploy these capabilities responsibly. If you're ready to build AI solutions that exploit full technical potential while maintaining robust ethical guardrails, rather than implementing basic features with superficial ethical considerations, Agathon's specialised expertise in sophisticated AI implementation can provide the guidance you need. #### References - Morley, J., Floridi, L., Kinsey, L., & Elhalal, A. (2020). From what to how: An initial review of publicly available AI ethics tools, methods and research to translate principles into practices. Science and Engineering Ethics, 26(4), 2141-2168. - Rakova, B., Yang, J., Cramer, H., & Chowdhury, R. (2021). Where responsible AI meets reality: Practitioner perspectives on enablers for shifting organizational practices. In Proceedings of the ACM on Human-Computer Interaction. - Fjeld, J., Achten, N., Hilligoss, H., Nagy, A., & Srikumar, M. (2020). Principled artificial intelligence: Mapping consensus in ethical and rights-based approaches to principles for AI. Berkman Klein Center Research Publication, (2020-1). - Leslie, D. (2019). Understanding artificial intelligence ethics and safety: A guide for the responsible design and implementation of AI systems in the public sector. The Alan Turing Institute. --- ### Hiring a fractional head of AI to complement your existing technical team - URL: https://agathon.ai/insights/hiring-a-fractional-head-of-ai-to-complement-your-existing-technical-team - Published: 2025-08-15 - Categories: Fractional CTO, AI Strategy, Responsible AI Most organisations implementing AI today are operating at 10-20% of what's technically possible. It's not for lack of engineering talent or infrastructure – it's a leadership gap. Companies with world-class development teams are still producing rudimentary AI implementations while their executives speak breathlessly about "transformation" and "disruption." The reality on the ground is far more mundane: basic automation, shallow ML implementations, and simplistic LLM applications that barely scratch the surface of what modern AI architectures can deliver. This isn't just a matter of missed opportunity; it's becoming an existential risk. As the gap widens between AI's potential and an organisation's capability to exploit it, competitors who bridge this divide will gain insurmountable advantages. The solution isn't another full-time executive hire or costly consulting engagement. It's a strategic architectural leadership role that blends deep technical knowledge with commercial and product vision – the fractional Head of AI. #### Why technical teams falter without AI leadership Technical teams without specialised AI leadership typically fall into predictable traps: 1. Architecture myopia: Building AI systems using traditional software principles, missing the fundamentally different architectural requirements of effective AI systems. 1. Feature fixation: Implementing isolated AI capabilities rather than designing coherent, evolving intelligence systems. 1. Experimental aimlessness: Pursuing technical proofs-of-concept without strategic orchestration toward business objectives. 1. Governance blindspots: Underestimating the ethical, regulatory, and risk governance unique to AI systems. MD Anderson Cancer Center's experience illustrates this perfectly. Their ambitious $62 million "moon shot" project using IBM's Watson cognitive system to diagnose and recommend cancer treatments was put on hold in 2017 without treating a single patient. Meanwhile, their IT group's more modest AI experiments in patient services, financial operations, and staff support yielded significant business impact with far less investment. The failure wasn't technical competence, rather it was AI leadership and strategic orchestration. #### The evolving role of AI leadership in tech organisations As AI shifts from experimental to mission-critical, leadership requirements have evolved dramatically. Modern AI leaders must bridge traditionally separate domains: - Technical architecture: Designing systems that balance immediate business needs with future extensibility - Business strategy: Translating AI capabilities into competitive advantage - Ethical governance: Establishing frameworks for responsible AI use - Organisational transformation: Building AI-native processes and culture The McKinsey Global Institute estimates AI will add $13 trillion to the global economy over the next decade, but their research shows that the primary barrier isn't technology rather it’s organisational structure and leadership. Companies with the strongest financial performance from AI aren't those with the most advanced technologies but those with coherent AI leadership that connects technical capabilities to strategic outcomes. #### When a fractional Head of AI makes strategic sense The fractional model provides specialised leadership without the overhead of a full-time executive hire. This approach makes particular sense when: ##### Your organisation shows these warning signs - You have skilled engineers building technically sound but strategically disconnected AI features - Your AI initiatives produce working demos but struggle to deliver production-grade impact - Technical decisions about AI architecture are made without sufficient consideration of long-term strategic implications - Ethical and governance considerations are addressed reactively rather than designed proactively ##### Your organisational structure fits these patterns - Mid-sized companies with strong technical teams but limited AI specialisation - Scale-ups transitioning from proof-of-concept to production AI systems - Established enterprises initiating strategic AI capabilities alongside existing operations - Companies facing specific AI challenges that require specialised expertise #### The three domains of fractional AI leadership impact A fractional Head of AI delivers strategic value in three critical domains that most organisations struggle to integrate: ##### Strategic AI architecture design Most AI implementations suffer from architectural fragmentation – point solutions designed for immediate problems without a coherent technical strategy. The fractional Head of AI develops a technical blueprint that: - Creates a scalable foundation that accommodates evolving AI capabilities - Defines integration patterns between AI components and existing systems - Establishes governance frameworks for data, models, and deployment - Plans for ethical considerations from architecture to implementation ##### Capability unlocking and orchestration Technical teams often implement basic versions of AI capabilities without exploiting their full potential. A fractional Head of AI: - Identifies underutilised potential in existing AI investments - Orchestrates complementary capabilities to create multiplicative effects - Optimises the balance between proprietary development and vendor solutions - Creates capability roadmaps that align with business strategy horizons ##### Knowledge transfer and capability building The most valuable contribution of a fractional Head of AI is often building internal capabilities: - Establishing architectural thinking and technical standards - Mentoring technical leaders on AI-specific considerations - Building governance frameworks that technical teams can operationalise - Creating strategic evaluation frameworks for AI opportunities #### Beyond consulting: Why fractional leadership outperforms traditional models The fractional model significantly differs from traditional consulting engagements. Where consultants typically assess, recommend, and depart, fractional leaders: - Take direct accountability for outcomes - Work within the organisation's structure rather than alongside it - Make decisions rather than just providing recommendations - Transfer knowledge continuously rather than at project boundaries Harvard Business Review research found that companies using cognitive technologies achieved the best results when focusing on "low-hanging fruit" with clear business cases, not moon shots. A fractional Head of AI brings the strategic discipline to identify these opportunities while building toward more ambitious capabilities. #### Finding and evaluating the right fractional AI leader The ideal fractional Head of AI combines qualities rarely found in a single full-time hire: ##### Technical depth markers - Demonstrates systems thinking across the AI stack, not just expertise in a single domain - Can articulate architectural considerations that balance immediate needs with future flexibility - Possesses both theoretical understanding and practical implementation experience - Maintains current knowledge of rapidly evolving AI capabilities ##### Strategic breadth indicators - Translates technical capabilities into business advantage - Understands organisational dynamics and change management - Can communicate effectively with both technical and non-technical stakeholders - Brings frameworks for ethical and responsible AI implementation ##### Assessment methods Traditional interviews often fail to identify effective fractional AI leaders. More effective approaches include: - Architecture design exercises with your technical team - Strategic planning simulations with business stakeholders - Case-based discussions of previous AI implementations - Ethical scenario responses to evaluate governance thinking #### Measuring success in fractional AI leadership arrangements Effective fractional AI leadership produces measurable outcomes across multiple dimensions: ##### Technical architecture maturity - Coherent AI system design that supports multiple use cases - Clear standards for AI development and deployment - Reduced technical debt in AI components - Improved integration between AI and traditional systems ##### Strategic capability advancement - Transition from isolated features to orchestrated capabilities - Increased business impact from existing AI investments - Clearer decision frameworks for AI opportunity evaluation - Improved alignment between technical priorities and business objectives ##### Organisational capability building - Enhanced internal expertise in AI-specific considerations - More sophisticated evaluation of AI vendor claims - Improved governance processes for AI systems - Greater confidence in AI roadmap execution #### The shift from fractional to embedded AI leadership Effective fractional leadership should ultimately make itself unnecessary by building internal capabilities. This transition typically follows three phases: 1. Direct leadership: The fractional leader makes key decisions and directs AI strategy 1. Collaborative leadership: The fractional leader works alongside emerging internal leaders 1. Advisory oversight: The fractional leader provides periodic guidance to capable internal leaders This evolution should be explicitly planned and measured, with clear milestones for the transition of responsibilities. #### Exploiting AI's full potential through strategic leadership The gap between what's technically possible with AI and what most organisations actually implement represents one of the largest missed opportunities in modern business. A fractional Head of AI provides the architectural vision and strategic discipline to exploit this potential without the overhead of a full-time executive hire. The most successful organisations don't treat AI as just another technology initiative but as a fundamental capability requiring specialised leadership. As researchers at MIT Sloan Management Review discovered, the primary barrier to AI impact isn't technical implementation but organisational integration – precisely where fractional leadership creates the most value. If you're ready to move beyond implementing basic AI features to building systems that exploit AI's full technical potential, a fractional Head of AI may be the strategic catalyst you need. #### References - Davenport, T., & Ronanki, R. (2018). Artificial Intelligence for the Real World. Harvard Business Review. - Fountaine, T., McCarthy, B., & Saleh, T. (2019). Building the AI-Powered Organization. Harvard Business Review. - Hammond, K. (2021). The Economics of Artificial Intelligence. IEEE Computer Society. --- ### Self-improving systems: the AI architecture pattern everyone talks about, nobody builds - URL: https://agathon.ai/insights/self-improving-systems-the-ai-architecture-pattern-everyone-talks-about-nobody-builds - Published: 2025-07-28 - Categories: AI Strategy, Generative AI, AI Agents The most radical promise of artificial intelligence remains largely untapped in today's implementations. While companies race to deploy LLMs with increasingly impressive baseline capabilities, the truly transformative architecture—systems that improve themselves without direct human intervention—remains conspicuously absent from production environments. This isn't merely an implementation gap; it's a profound misunderstanding of what constitutes genuine intelligence versus sophisticated mimicry. #### The theoretical mirage of self-improvement Self-improving systems represent AI's most tantalising architectural pattern. Conceptually, they're deceptively straightforward: AI systems that can evaluate their own performance, identify deficiencies, and autonomously enhance their capabilities without explicit human programming. In theory, such systems would create a virtuous cycle of continuous capability enhancement that dramatically outpaces human-guided development. The fundamental components of genuinely self-improving systems include: 1. A meta-learning mechanism that enables performance evaluation 1. An introspection capability to identify improvement opportunities 1. A self-modification architecture that implements improvements 1. A validation framework ensuring changes enhance rather than degrade performance The promise is compelling: recursive self-improvement could initiate rapid capability accumulation that would make our current AI development pace seem glacial by comparison. As Hubinger et al. note in their 2019 paper on risks from learned optimisation, systems that engage in "mesa-optimisation"—where a learned model becomes an optimiser itself—could potentially rewrite their own architecture to better achieve their objectives. Yet despite intense theoretical interest, truly self-improving systems remain conspicuously absent from today's AI landscape. The systems we label as "self-improving" are, at best, pale approximations. #### The implementation chasm: technical barriers The gap between theoretical possibility and practical implementation stems from several profound technical challenges: ##### The meta-learning paradox For a system to improve itself, it must possess a meta-cognitive capability to evaluate its own performance—effectively requiring intelligence about intelligence. This creates a circular dependency: how can a system bootstrap improvements without already possessing the capability to recognise what constitutes improvement? As Chollet articulates in his 2019 paper "On the Measure of Intelligence," genuine intelligence isn't mere task performance but rather "skill-acquisition efficiency" across novel situations with minimal prior experience. Creating architectures that can effectively measure their own skill-acquisition efficiency remains an unsolved problem. ##### The self-modification dilemma Even if a system could accurately evaluate its performance, implementing beneficial self-modifications poses a distinct challenge. Most modern AI architectures weren't designed with self-modification capabilities. Neural networks—the backbone of contemporary AI—aren't easily introspectable or modifiable by their own inference processes. AlphaGo Zero demonstrates a limited form of self-improvement through a reinforcement learning cycle, but crucially, this improvement happens within predefined architectural boundaries. As Silver et al. describe in their 2017 Nature paper, AlphaGo Zero "becomes its own teacher" through self-play, but the architecture itself—the neural network structure, the MCTS search algorithm—remains fixed by human designers. ##### Verification impossibility Perhaps the most profound barrier is verification. How can a system ensure that modifications improve rather than degrade performance across all potential scenarios? This creates a safety paradox: comprehensive verification would require testing against all possible inputs—an intractable problem for any non-trivial domain. This explains why even approximate self-improving systems like AutoML platforms operate within tightly constrained search spaces rather than allowing unbounded architectural modifications. #### Organisational obstacles: the human element Technical barriers alone don't explain the implementation gap. Organisational factors profoundly shape what gets built and deployed: ##### Misalignment with development practices Modern AI development follows an iterative cycle of human-directed experimentation, evaluation, and refinement. Teams optimise for predictable, incremental improvements rather than potentially transformative but unpredictable self-modification capabilities. This creates a fundamental tension: truly self-improving systems would require surrendering significant control over the development process—something few organisations are culturally or operationally prepared to do. ##### Talent allocation realities Building approximate self-improving systems requires rare expertise across multiple domains: meta-learning, neural architecture search, reinforcement learning, verification methods, and system safety. This expertise doesn't just need to exist within an organisation; it needs to be coordinated across teams with often divergent incentives and objectives. The result is that most organisations default to familiar, more tractable approaches that yield predictable, incremental improvements rather than potentially transformative but risky architectural innovations. #### Current approximations: the self-improvement illusion What passes for "self-improving" systems today actually represents constrained optimisation within predetermined parameters rather than genuine self-modification: ##### Reinforcement learning from human feedback (RLHF) Systems like those described by Stiennon et al. in their 2020 paper on "Learning to summarize from human feedback" represent the current state-of-the-art in approximate self-improvement. These systems collect human preferences, train a reward model on those preferences, and then optimise performance against that reward model. While this creates a feedback loop that improves performance, the improvement remains bounded by human feedback quality and operates within fixed architectural constraints. The system improves its outputs but cannot fundamentally redesign its own architecture or objective function. ##### Neural architecture search Automated machine learning (AutoML) platforms provide another approximation of self-improvement by automatically searching for optimal neural network architectures. However, these systems operate within tightly constrained search spaces and optimise for predefined metrics rather than autonomously determining what constitutes improvement. The critical distinction is that these systems don't truly improve themselves—they're designed by humans to optimise specific parameters within carefully constructed boundaries. They're more akin to sophisticated auto-tuning than genuine self-improvement. #### The essential architecture for genuine self-improvement Building truly self-improving systems requires a radical departure from current architectural approaches: ##### Introspectable representations For a system to modify itself effectively, it must maintain interpretable representations of its own capabilities and processes. Current neural network architectures produce distributed representations that resist straightforward interpretation or targeted modification. A promising alternative comes from neurosymbolic approaches that combine the learning capabilities of neural networks with the interpretability of symbolic systems. These hybrid architectures could potentially support meaningful introspection and targeted self-modification. ##### Meta-learning frameworks Rather than optimising for task performance directly, genuinely self-improving systems must optimise for learning efficiency itself. This requires architectures that can evaluate not just what they know but how they learn. The concept of "learning to learn" has gained traction in recent research, but current approaches typically focus on learning hyperparameters or initialisation strategies rather than fundamental architectural innovation. ##### Bounded self-modification While unbounded self-modification presents intractable safety challenges, bounded self-modification offers a more viable path forward. Systems could be designed with constrained "modification spaces" where self-improvement can occur without risking fundamental objective function corruption. AlphaGo Zero provides a template for this approach: self-improvement occurs within the bounded context of policy and value networks, but the overall architecture—the objective function, the MCTS algorithm, the network structure—remains fixed. #### Building viable self-improving systems: the path forward Despite these formidable challenges, pragmatic approaches to approximating self-improvement are emerging: ##### Tiered architectural control Rather than pursuing unbounded self-modification, practical systems should implement multiple control tiers with varying degrees of self-modification authority: 1. Output-level optimisation (adjusting system outputs without modifying internal processes) 1. Parameter-level optimisation (adjusting weights and hyperparameters within fixed architectures) 1. Limited architectural optimisation (modifying specific architectural components within safety boundaries) This tiered approach allows for meaningful self-improvement while maintaining essential safety guarantees. ##### Explicit improvement interfaces Instead of expecting systems to modify arbitrary aspects of their architecture, designers should implement explicit "improvement interfaces" that expose safe modification points. This design pattern—creating specific architectural components designed for self-modification—establishes clear boundaries between fixed and modifiable system elements. ##### Human-AI collaborative improvement The most promising near-term approach may be collaborative improvement, where AI systems propose modifications that humans evaluate and implement. This approach leverages both machine-generated innovation and human judgment to guide the improvement process. The RLHF paradigm exemplifies this approach, creating a human-machine feedback loop that enables meaningful improvement while maintaining essential safety boundaries. #### Conclusion: reframing self-improvement expectations Self-improving systems represent the horizon of AI development—fascinating, promising, but still beyond our current grasp. The challenges aren't merely technical but conceptual: we're still developing frameworks to understand what intelligence is, let alone how to enable systems to enhance it autonomously. Yet this shouldn't discourage innovation. By reframing expectations from "fully autonomous self-improvement" to "bounded, human-guided self-modification," we can make meaningful progress while acknowledging the profound technical and philosophical challenges involved. The most promising path forward isn't the science fiction vision of completely autonomous self-improving systems but rather increasingly sophisticated human-AI collaborative improvement frameworks. These hybrid approaches leverage both machine learning capabilities and human judgment to create virtuous cycles of enhancement that progressively expand the boundaries of what's possible. Most organisations implementing AI today are barely scratching the surface of what's technically possible—focusing on implementing basic capabilities rather than architecting systems with even limited self-improvement potential. The gap between theoretical possibility and practical implementation isn't just a matter of technical complexity but of imagination and architectural vision. If you're ready to move beyond commodity AI implementations and build systems that exploit the full technical potential of modern AI architectures—including bounded self-improvement capabilities—you need partners who understand both the theoretical possibilities and practical constraints of advanced AI system design. #### References - Hubinger, J., et al. (2019). "Risks from Learned Optimization in Advanced Machine Learning Systems." - Turchin, A., & Denkenberger, D. (2020). "Classification of global catastrophic risks connected with artificial intelligence." AI & Society. - Stiennon, N., et al. (2020). "Learning to summarize from human feedback." - Yampolskiy, R. V. (2020). "On Controllability of Artificial Intelligence." - Silver, D., et al. (2017). "Mastering the game of Go without human knowledge." Nature. - Chollet, F. (2019). "On the Measure of Intelligence." --- ### Parameter-efficient fine-tuning: what business leaders need to know - URL: https://agathon.ai/insights/parameter-efficient-fine-tuning-what-business-leaders-need-to-know - Published: 2025-07-07 - Categories: Machine Learning, AI Strategy, Responsible AI While your competitors chase billion-parameter models and drain their compute budgets on full-model fine-tuning, they're missing a transformative approach that delivers equivalent performance at a fraction of the cost. Parameter-efficient fine-tuning (PEFT) isn't just a technical optimisation—it's a strategic advantage that fundamentally changes the economics of AI model customisation. Most organisations deploying large language models are using brute force techniques from 2019, unaware that they're leaving substantial value on the table and missing the opportunity to deploy specialised models at scale across their enterprise. This is the equivalent of purchasing an entire new vehicle when you only needed to replace the tyres. #### Hidden economics of model adaptation The financial reality of traditional fine-tuning approaches has become increasingly prohibitive. When adapting a large language model like GPT-3 (175B parameters), conventional fine-tuning requires updating every parameter in the model. This translates to enormous computational costs, extended training times, and substantial storage requirements for each variant. Consider the maths: fine-tuning a 175B parameter model typically requires high-end GPUs with at least 80GB of VRAM, costing upwards of £10,000 per GPU. Even with this hardware, training can take days or weeks, consuming thousands in compute costs alone. Afterwards, you're storing multiple copies of essentially the same model with slight variations—each consuming hundreds of gigabytes. What most technical leaders miss is that this approach scales poorly when deploying specialised AI capabilities across multiple business domains. Each use case effectively requires an independent, fully-trained model—multiplying costs linearly while delivering diminishing returns. #### Technical foundations of parameter efficiency At its core, parameter-efficient fine-tuning is based on a simple yet profound insight: we don't need to modify all parameters in a pre-trained model to adapt it for specific tasks. Research has shown that language models contain significant redundancy and that meaningful adaptations can be achieved by strategically modifying only a small subset of parameters. ##### Low-rank adaptation (LoRA): The breakthrough approach The breakthrough came with Hu et al.'s 2021 introduction of Low-Rank Adaptation (LoRA). This technique works by freezing the pre-trained model weights and injecting trainable rank decomposition matrices into each layer of the transformer architecture. These matrices capture task-specific adaptations while keeping the original model intact. LoRA exploits an important mathematical insight: the adaptations needed for task-specific tuning can be represented as low-rank updates to the original weight matrices. As demonstrated in their research, this approach can reduce the number of trainable parameters by a factor of 10,000 compared to full fine-tuning of GPT-3 175B, while achieving comparable or superior performance. The mathematics behind LoRA is elegant: if W₀ is the original weight matrix and ΔW is the update, then: W = W₀ + ΔW, where ΔW = BA and B and A are much smaller matrices. ##### Beyond LoRA: The PEFT ecosystem The PEFT landscape has expanded significantly beyond LoRA, with several complementary approaches: - Adapter methods: Introduced by Houlsby et al. (2019), these inject small trainable modules into the model while keeping most parameters frozen. These modules create bottlenecks that efficiently capture task-specific information. - Prompt tuning: Lester et al. (2021) demonstrated that by simply adding and optimising continuous "soft prompt" vectors prepended to inputs, models can achieve performance comparable to full fine-tuning as scale increases—especially with models exceeding billions of parameters. - Prefix tuning: Li and Liang (2021) extended prompt tuning by optimising prefix vectors for both the encoder and decoder of a sequence-to-sequence model, allowing more control over generation tasks. - Quantized LoRA (QLoRA): Dettmers et al. (2023) combined quantization with LoRA, enabling fine-tuning of models up to 65B parameters on a single consumer GPU with 48GB of memory—a task previously requiring multiple high-end GPUs. #### Strategic advantages for enterprise AI deployment The business implications of parameter-efficient fine-tuning extend far beyond mere cost savings. By fundamentally changing how models are adapted and deployed, PEFT enables several strategic advantages that are often overlooked. ##### From monolithic to modular AI architecture Traditional fine-tuning creates siloed, monolithic models that scale poorly across an enterprise. Parameter-efficient approaches enable a modular architecture where a single frozen base model can be augmented with multiple lightweight adaptations for different business domains. This modularity transforms deployment strategies. Rather than managing dozens of full-sized models, technical teams can maintain a single core model supplemented by small adapter modules for specific use cases. The storage footprint for these adapters is trivial—often less than 1% of the base model size. ##### Accelerated time-to-value The research findings are compelling: QLoRA enabled researchers to fine-tune a 65B parameter model in just 24 hours on a single GPU, reaching 99.3% of ChatGPT's performance level on benchmark tasks. This represents a paradigm shift in the development timeline for AI capabilities. For enterprises, this acceleration means AI initiatives can move from conception to production in days rather than months. New use cases can be rapidly prototyped, evaluated, and deployed without lengthy procurement cycles for additional compute resources. ##### Environmental and resource efficiency The environmental impact of AI training has come under increasing scrutiny. Parameter-efficient approaches directly address this concern by drastically reducing the computational resources required for model adaptation. Dettmers et al. demonstrated that their approach reduced memory usage by a factor of 3 compared to standard fine-tuning techniques. This efficiency translates to proportional reductions in energy consumption and carbon footprint—allowing organisations to meet sustainability goals while scaling AI capabilities. #### Implementation strategies for competitive advantage Implementing parameter-efficient fine-tuning requires a strategic approach that balances technical considerations with business objectives. Here's how forward-thinking organisations are leveraging these techniques: ##### Diversification through specialisation Rather than creating a single "jack of all trades" model, leading organisations are developing portfolios of specialised capabilities through parameter-efficient adaptation. Each business domain receives tailored AI capabilities without the redundant cost of maintaining completely separate models. This approach enables a level of customisation previously considered economically infeasible. Customer service can have specialised models for different product lines, finance can have separate models for different analytical tasks, and product teams can build domain-specific assistants—all while sharing the computational cost of a single base model. ##### Governance through architectural separation Responsible AI governance becomes more manageable when adaptations are architecturally separated from the base model. Changes to task-specific modules don't risk unintended consequences to the core model's behaviour, creating natural isolation boundaries for governance controls. This separation also simplifies compliance efforts. The base model can undergo rigorous security and bias testing once, while lighter evaluation can focus on the specific behavioural changes introduced by each adapter. ##### Rapid experimentation and iteration The lightweight nature of parameter-efficient adaptations enables a fundamentally different approach to AI development. Teams can maintain multiple experimental versions simultaneously, conduct A/B testing with minimal overhead, and rapidly iterate based on user feedback. This advantage is particularly pronounced when working with the largest models. While traditional fine-tuning of a 65B parameter model might take weeks on multiple GPUs, parameter-efficient approaches allow daily or even hourly iterations on a single device. #### Limitations and strategic considerations Despite its advantages, parameter-efficient fine-tuning isn't universally optimal. Technical leaders should consider several factors when evaluating implementation: ##### Performance trade-offs at scale Research by Lialin et al. (2023) revealed that performance gaps between PEFT methods become more pronounced in resource-constrained settings. When hyperparameter optimisation is limited and networks are fine-tuned for only a few epochs, some methods struggle to match LoRA's baseline performance. This finding has important implications for enterprise deployment, where time and computational budgets are often constrained. It suggests that simpler, more robust PEFT methods may be preferable to theoretically superior but more finicky approaches in production environments. ##### Integration complexity Implementing parameter-efficient approaches requires rethinking existing ML pipelines. While the techniques themselves are becoming more accessible through libraries and frameworks, they often demand changes to training infrastructure, serving architecture, and model management systems. Organisations with heavily invested traditional ML infrastructure may face integration challenges that partially offset the computational savings. This is particularly true for organisations with custom training pipelines or tightly coupled serving architectures. ##### Compatibility challenges across the model lifecycle Not all parameter-efficient methods work equally well across all models and tasks. The research shows significant performance variance depending on model architecture, size, and the specific downstream task. Additionally, some approaches like adapter methods introduce small but measurable inference latency, which may be critical for real-time applications. Others, like LoRA, maintain the same inference speed but may require more complex model loading procedures. #### The future landscape of AI customisation The parameter-efficient fine-tuning landscape is evolving rapidly, with several trends that will shape enterprise AI strategies in the coming years: ##### Convergence of techniques Research is increasingly exploring hybrid approaches that combine multiple parameter-efficient techniques. For example, integrating quantization with adapter methods or combining prompt tuning with LoRA to exploit the complementary strengths of each approach. This convergence will likely lead to even more efficient adaptation methods that further reduce computational requirements while maintaining or improving performance. ##### Integration with emerging model architectures As model architectures evolve beyond the current transformer paradigm, parameter-efficient techniques will adapt accordingly. Early research suggests that these approaches may be even more effective with next-generation architectures that incorporate sparse attention or mixture-of-experts designs. ##### Democratisation of advanced AI capabilities Perhaps most significantly, parameter-efficient approaches are democratising access to state-of-the-art AI capabilities. By reducing the computational barriers to model adaptation, these techniques enable smaller organisations and teams to create sophisticated, customised AI solutions previously available only to technology giants. This democratisation will accelerate innovation and lead to more diverse applications of AI across industries and domains. #### Moving beyond conventional wisdom The strategic significance of parameter-efficient fine-tuning extends far beyond technical optimisation. It represents a fundamental shift in how organisations can approach AI customisation and deployment—enabling more agile, scalable, and economically viable AI strategies. Technical leaders who recognise this shift can position their organisations to extract substantially more value from their AI investments while reducing costs and accelerating time-to-market. Rather than pursuing the brute-force approach of training ever-larger models, they can focus on efficiently adapting existing models to create portfolios of specialised capabilities. The research is clear: parameter-efficient fine-tuning delivers comparable or superior performance to traditional approaches while requiring only a fraction of the computational resources. For organisations serious about scaling AI capabilities across their enterprise, this isn't merely a technical detail—it's a strategic imperative. If you're ready to build AI solutions that exploit the full technical potential of large language models rather than implementing basic, resource-intensive approaches, it's time to reconsider your fine-tuning strategy. The future belongs to organisations that can rapidly adapt and deploy specialised AI capabilities while controlling costs and maximising return on AI investments. #### References - Hu, E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). LoRA: Low-Rank Adaptation of Large Language Models. - Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., Attariyan, M., & Gelly, S. (2019). Parameter-Efficient Transfer Learning for NLP. - Lester, B., Al-Rfou, R., & Constant, N. (2021). The Power of Scale for Parameter-Efficient Prompt Tuning. - Li, X. L., & Liang, P. (2021). Prefix-Tuning: Optimizing Continuous Prompts for Generation. - Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient Finetuning of Quantized LLMs. - Lialin, V., Deshpande, A., & Rumshisky, A. (2023). Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning. --- ### The curse of dimensionality: when more data becomes your enemy - URL: https://agathon.ai/insights/the-curse-of-dimensionality-when-more-data-becomes-your-enemy - Published: 2025-07-02 - Categories: Machine Learning, AI Strategy, AI Consulting Most AI implementations today are operating with one hand tied behind their back, primarily because organisations fail to understand a fundamental technical paradox: adding more data features often makes AI systems worse, not better. This isn't just an academic quirk—it's a mathematical reality with profound implications for every sophisticated AI application you'll build. #### The paradox that cripples your AI systems The curse of dimensionality represents perhaps the most significant yet overlooked challenge in modern AI development. First described by mathematician Richard Bellman in 1957, this phenomenon manifests when data exists in high-dimensional spaces, causing algorithms to behave in counterintuitive and often detrimental ways. What's particularly vexing is that while adding dimensions (features) theoretically provides more information, it simultaneously creates mathematical conditions that undermine the very foundations of most AI systems. The algorithms most companies deploy simply weren't designed to handle this paradox effectively. Most organisations respond to AI challenges by gathering more data or adding more features. But as I'll demonstrate, this approach often accelerates the very problem they're trying to solve. #### The geometric betrayal To understand why high-dimensional spaces behave so strangely, consider a simple example that demonstrates the effect. Imagine a unit hypercube (a cube with sides of length 1) in various dimensions: - In 1D, it's just a line segment from 0 to 1 - In 2D, it's a square with an area of 1 - In 3D, it's a cube with a volume of 1 Now, let's insert a slightly smaller hypercube inside it, with sides of length 0.9: - In 1D, this smaller segment occupies 90% of the original - In 2D, the smaller square occupies 0.9² = 81% of the original - In 3D, the smaller cube occupies 0.9³ = 72.9% of the original By the time we reach just 10 dimensions, the smaller hypercube occupies only 0.9¹⁰ ≈ 35% of the volume. At 100 dimensions (common in many machine learning applications), this becomes 0.9¹⁰⁰ ≈ 0.00027% of the volume. This isn't just a mathematical curiosity—it fundamentally alters how proximity and similarity function in your AI systems. #### The concentration of distances phenomenon Perhaps even more troubling is the behaviour of distance metrics in high-dimensional spaces. Distance-based methods underpin numerous machine learning algorithms, from nearest neighbour searches to clustering techniques like k-means. As dimensions increase, a disturbing effect emerges: the difference between the nearest and farthest points becomes negligible. Mathematically, as the dimensionality approaches infinity, the ratio of the distances approaches 1, meaning that distance-based discrimination becomes impossible. Aggarwal, Hinneburg, and Keim's research demonstrates this phenomenon quite clearly. This has profound implications: in high-dimensional spaces, the concept of a "nearest neighbour" loses meaning. Your similarity metrics break down. Your clustering algorithms group unrelated points. Your recommendation systems suggest irrelevant items. #### Why your choice of distance metric matters more than you think Most AI practitioners reflexively reach for Euclidean distance (L₂ norm) when implementing distance-based algorithms. This default choice is often disastrous in high-dimensional spaces. Research by Aggarwal et al. reveals something counterintuitive: the Manhattan distance (L₁ norm) consistently outperforms Euclidean distance in high dimensions. Their analysis demonstrates that the L₁ norm maintains discriminative power far better than L₂ as dimensionality increases. Even more revealing is their exploration of fractional distance metrics (L_k norms where 0 < k < 1), which show remarkable resistance to the curse of dimensionality. These metrics, though less intuitive, significantly improve the effectiveness of clustering and nearest neighbour searches in high dimensions. This isn't theoretical—their experiments with the k-means algorithm show that using L₁ instead of L₂ norms can dramatically improve clustering accuracy in high-dimensional data. Fractional norms perform even better, with L₀.₅ delivering superior results in many contexts. #### Statistical sparsity: The empty space phenomenon Another dimension of this curse manifests in the exponential growth of training data requirements. In low dimensions, relatively few samples can adequately represent the underlying distribution. As dimensions increase, the volume of the space grows exponentially, creating vast "empty" regions where no data exists. Donoho (2000) termed this the "empty space phenomenon," noting that for a fixed dataset size, the proportion of the feature space containing data points approaches zero as dimensionality increases. This creates a statistical sparsity problem where most of the space becomes unrepresented in your training data. The practical consequence? Your models increasingly fit to noise rather than signal as dimensions increase. Overfitting becomes almost inevitable without proper dimensionality management. #### Strategic approaches that actually work Rather than advocating a single solution, let me outline a multi-faceted approach that sophisticated organisations can implement: ##### Feature engineering with dimensional awareness Effective feature engineering isn't just about creating relevant features—it's about understanding dimensional interactions. Some approaches that deliver results: - Mutual information analysis to identify and eliminate redundant dimensions - Careful application of domain knowledge to select features that maintain statistical significance - Feature hierarchies that allow dynamic dimension management based on context ##### Beyond PCA: Modern dimensionality reduction Principal Component Analysis (PCA) remains the default technique for many organisations, but its linear nature makes it insufficient for many real-world applications. More sophisticated approaches include: - Manifold learning techniques like t-SNE and UMAP that preserve local structure in lower dimensions - Autoencoder architectures that learn nonlinear dimensional reductions tailored to your specific data - Probabilistic PCA and factor analysis methods that explicitly model uncertainty in dimensional reduction ##### Distance metric engineering Most organisations never question their choice of distance metric, but this decision has profound implications: - Replace Euclidean distance with Manhattan distance in high-dimensional contexts - Experiment with fractional norms (L₀.₅ or L₀.₈) for clustering and similarity searches - Implement adaptive distance metrics that adjust based on local density patterns ##### Architectural adaptations Some neural network architectures inherently handle high dimensionality better than others: - Attention mechanisms that dynamically focus on relevant dimensions - Sparse neural networks that activate only for specific dimensional subspaces - Hierarchical embeddings that represent data at multiple dimensional resolutions #### The opportunity in dimensional mastery Organisations that master high-dimensional spaces gain significant competitive advantage. While most companies struggle with the mathematical realities of high-dimensional data, those who understand and exploit these properties can build dramatically more effective AI systems. This isn't about small incremental improvements—it's about fundamental capability differences. Systems that effectively navigate high-dimensional spaces can: - Extract signal from data that appears as noise to conventional approaches - Maintain discrimination ability where standard methods collapse - Identify patterns that exist only in specific dimensional subspaces #### Moving beyond dimensional naivety The most sophisticated AI implementations don't just add more data or more features—they strategically manage dimensionality to exploit its properties rather than fall victim to its curses. This requires moving beyond the simplistic "more data is better" mindset that dominates most AI projects. By understanding the mathematical realities of high-dimensional spaces, implementing appropriate distance metrics, and architecting systems with dimensional awareness, organisations can unlock capabilities that remain inaccessible to those using conventional approaches. If you're ready to build AI systems that exploit the full technical potential of your data rather than implementing basic features constrained by dimensional limitations, it's time to rethink your fundamental approach to AI architecture. #### References - Bellman, R. (1957). Dynamic Programming. Princeton University Press. - Domingos, P. (2012). A Few Useful Things to Know About Machine Learning. Communications of the ACM, 55(10), 78-87. - Aggarwal, C. C., Hinneburg, A., & Keim, D. A. (2001). On the Surprising Behavior of Distance Metrics in High Dimensional Space. Database Theory — ICDT 2001, 420-434. - Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press. --- ### The future of AI for lawyers: transforming legal practice in the digital age - URL: https://agathon.ai/insights/the-future-of-ai-for-lawyers-transforming-legal-practice-in-the-digital-age - Published: 2025-07-02 - Categories: NLP, Machine Learning, Responsible AI Most legal AI implementations scratch merely 10% of what's technically possible. The true revolution isn't in basic document search or simple contract analysis—it lies in exploiting the rich architectural possibilities that emerge when sophisticated neural networks meet rule-based legal systems. While many consultancies push off-the-shelf solutions that mimic basic paralegal functions, the architectural sophistication required for truly transformative legal AI remains largely untapped. #### Beyond the buzzword: The architectural foundations of legal AI When examining legal AI, we must distinguish between shallow implementations and systems that genuinely transform legal practice. The distinction isn't merely academic—it represents billions in potential value currently left on the table. Legal reasoning has unique characteristics that demand specialised AI approaches. As Chalkidis and Kampas demonstrated in their 2019 research, legal word embeddings trained on domain-specific corpora significantly outperform general-purpose models. Their experiments with deep learning in law showed that legal-specific pre-training creates representations that capture the nuanced relationships between legal concepts in ways generic models simply cannot. ##### Natural language processing and legal document analysis Basic NLP techniques for legal documents follow conventional approaches. Advanced implementations, however, exploit what Bench-Capon and Sartor identified as "reasoning with cases incorporating theories and values"—a fundamentally different technical architecture. The most sophisticated legal NLP systems combine multiple technical approaches: - Domain-specific transformer models with legal corpus pre-training - Neurosymbolic architectures that merge legal rules with deep learning - Knowledge representation frameworks that preserve legal reasoning chains Research by Surden demonstrates that standard implementations capture merely surface-level document features, while advanced architectures can model complex legal reasoning patterns and extract deeper semantic relationships. ##### Machine learning approaches for case prediction The architectural distinction between basic and sophisticated case prediction systems is stark. Elementary systems apply simple statistical analysis to historical outcomes. Advanced implementations, however, exploit what McGinnis and Pearce call the "great disruption"—systems that combine multiple technical dimensions: - Multi-modal analysis across document text, citation networks, and procedural history - Temporal reasoning capabilities that model jurisdictional shifts - Explainable AI components that preserve legal reasoning chains These architectures don't merely predict outcomes—they model the reasoning process itself, generating explanations that align with legal methodologies. ##### Knowledge representation in legal reasoning systems Most implementations treat knowledge representation as a trivial database problem. The technical frontier, however, lies in what Surden describes as "bridging formal logical representations with machine learning approaches." Sophisticated knowledge representation for legal AI requires: - Ontological frameworks that represent legal concepts and relationships - Inference mechanisms that support both rule-based and probabilistic reasoning - Temporal reasoning capabilities that track legal evolution The technical complexity here is profound—merging symbolic reasoning with neural approaches in ways that preserve legal integrity while exploiting pattern recognition capabilities. #### Current applications of AI in law: From basic to sophisticated The market abounds with AI tools claiming to revolutionise legal practice, but their architectural sophistication varies dramatically. ##### Contract review and due diligence automation Elementary contract review tools employ basic entity extraction and classification. Sophisticated implementations, however, exploit what Chalkidis and Androutsopoulos demonstrated in their deep learning approach to contract element extraction—combining hierarchical RNNs with attention mechanisms specifically designed for contractual reasoning. These advanced architectures can: - Identify interdependencies between contract clauses - Track obligation networks across multiple documents - Model conditional relationships in complex contractual structures The technical distinction is profound—from basic pattern matching to sophisticated reasoning about contractual structures. ##### Legal research and case analysis Standard legal research tools apply keyword matching with minimal semantic understanding. Advanced architectures, however, exploit what Alarie, Niblett, and Yoon called "technical mechanisms that model legal reasoning processes"—a fundamentally different technical approach. Sophisticated legal research systems: - Model citation networks and precedential relationships - Apply transfer learning across jurisdictions while preserving legal distinctions - Generate explanations aligned with legal methodologies The architecture requires specialised attention mechanisms that align with legal reasoning patterns. ##### Predictive analytics for case outcomes Elementary predictive systems apply standard classification algorithms to historical outcomes. Advanced implementations, as McGinnis and Pearce noted, exploit "machine intelligence that captures the deep structure of legal reasoning"—combining multiple technical approaches into cohesive reasoning frameworks. These sophisticated architectures: - Model jurisdictional variations and temporal shifts in legal doctrine - Represent reasoning chains that align with judicial decision processes - Generate explanations that preserve legal reasoning integrity The technical challenge lies in merging statistical prediction with reasoned explanation in ways that respect legal norms. ##### AI-powered e-discovery solutions Basic e-discovery tools apply simple keyword search and classification. Advanced architectures, however, exploit what Surden identified as "artificial intelligence techniques that model relevance and privilege across complex document collections"—a fundamentally different technical approach. Sophisticated e-discovery systems: - Apply active learning strategies that adapt to evolving case theories - Model privilege relationships across complex organisational structures - Generate explanations for inclusion/exclusion that align with legal standards The technical distinction lies in modelling complex legal concepts like relevance and privilege rather than simple document classification. #### Ethical considerations and responsible implementation The ethical dimensions of legal AI aren't merely regulatory checkboxes—they represent fundamental architectural requirements. ##### Addressing bias in legal AI systems Elementary approaches to bias focus on simple data balancing techniques. Sophisticated implementations, however, exploit what Surden described as "multi-level fairness frameworks"—architectural components that address bias at multiple technical levels: - Data representation fairness through specialised sampling techniques - Algorithmic fairness through constrained optimisation - Output fairness through adversarial validation The technical challenge requires specialised architectural components dedicated to fairness evaluation and mitigation. ##### Maintaining attorney-client privilege with AI tools Basic approaches to privilege rely on simple access controls. Sophisticated implementations, however, exploit "information architectures that model privilege relationships"—complex technical systems that: - Implement differential privacy guarantees for sensitive content - Model contextual privilege across complex organisational relationships - Provide formal guarantees for information compartmentalisation These architectural elements require specialised technical approaches beyond standard security measures. ##### Transparency and explainability requirements Elementary explainability focuses on simple feature importance metrics. Advanced implementations, however, exploit what research on legal AI identifies as "explanation architectures aligned with legal reasoning"—generating explanations that: - Reflect legal reasoning processes rather than statistical correlations - Preserve reasoning chains from evidence to conclusion - Support counterfactual analysis in legal contexts The technical challenge involves specialised architectural components dedicated to legal reasoning explanation. #### Challenges and limitations The technical frontier in legal AI faces several fundamental challenges that go beyond simple implementation issues. ##### Data quality and training issues in legal AI Standard approaches treat data quality as a simple preprocessing step. Sophisticated implementations, however, recognise what Chalkidis and Kampas identified as "domain-specific representation challenges in legal text"—requiring specialised technical approaches: - Transfer learning techniques adapted specifically for legal domain shift - Active learning strategies that efficiently utilise scarce expert annotation - Ontological alignment between data representations and legal concepts These architectural requirements demand specialised technical approaches beyond standard data cleaning. ##### Regulatory uncertainty and compliance concerns Basic compliance approaches apply simple rule-checking. Advanced implementations, however, exploit what Suksi described as "administrative due process requirements in automated systems"—architectural components that: - Model regulatory requirements as formal constraints - Implement verifiable compliance guarantees - Generate audit trails aligned with administrative requirements The technical challenge involves specialised verification components that can provide formal guarantees about system behaviour. ##### Technical limitations in complex legal reasoning Most discussions of limitations focus on simple performance metrics. The true frontier, however, lies in what Bench-Capon and Sartor identified as "reasoning with cases incorporating theories and values"—architectural challenges including: - Integrating rule-based reasoning with neural approaches - Modelling normative reasoning alongside descriptive prediction - Representing legal principles that transcend specific rules These challenges require fundamentally new technical approaches that bridge symbolic and statistical methods. #### The future of AI in legal practice The technical frontier in legal AI isn't about incremental improvements to existing tools—it involves fundamental architectural innovations. ##### Emerging technologies on the horizon While many focus on simple application of existing techniques, the true frontier lies in what McGinnis and Pearce called "the great disruption"—fundamental technical innovations including: - Neurosymbolic architectures that merge logical reasoning with deep learning - Federated learning approaches that preserve privacy while enabling collaboration - Formal verification techniques that provide guarantees about system behaviour These architectural innovations represent step-changes in capability rather than incremental improvements. ##### The evolving lawyer-AI partnership Basic discussions focus on automation of routine tasks. Sophisticated implementations, however, exploit what Alarie, Niblett, and Yoon described as "augmentation architectures"—technical systems designed to: - Enhance rather than replace legal reasoning - Provide cognitive scaffolding for complex legal analysis - Support explanation-based collaboration between human and machine These architectural approaches require fundamentally different technical design choices than simple automation systems. ##### Skills development for the AI-augmented lawyer Standard approaches focus on basic tool training. The technical frontier, however, requires what Armour and Sako identified as "AI-enabled business models"—requiring fundamentally different skill development: - Ability to evaluate and validate AI-generated legal reasoning - Capacity to identify appropriate domains for AI application - Skills in designing and specifying legal reasoning systems The technical challenge involves creating interfaces and systems that support this evolved skill development. #### Conclusion: Beyond basic automation to sophisticated legal reasoning Most legal AI implementations scratch merely the surface of what's technically possible. While basic document processing and simple prediction systems dominate the market, the true technical frontier lies in sophisticated architectures that model legal reasoning itself. The distinction between basic and advanced implementations isn't merely academic—it represents billions in untapped value and fundamental transformation of legal practice. As the research by Chalkidis, Surden, and others demonstrates, exploiting the full technical potential requires specialised architectural approaches rather than simple application of general AI techniques. If you're ready to build legal AI solutions that exploit the full technical potential rather than implementing basic features, you need partners who understand the architectural sophistication required. The future belongs to those who can bridge the gap between legal reasoning and advanced AI architectures. #### References - Chalkidis, I., & Kampas, D. (2019). Deep learning in law: Early adaptation and legal word embeddings trained on large corpora. Artificial Intelligence and Law, 27, 171-198. - Surden, H. (2020). Artificial Intelligence and Law: An Overview. Georgia State University Law Review, 35(4). - Bench-Capon, T., & Sartor, G. (2021). A model of legal reasoning with cases incorporating theories and values. Artificial Intelligence and Law, 29, 1-41. - McGinnis, J. O., & Pearce, R. G. (2019). The great disruption: How machine intelligence will transform the role of lawyers in the delivery of legal services. Fordham Law Review, 82(6), 3041-3066. --- ### Richard Sutton's bitter lesson explains why your AI solution feels shallow - URL: https://agathon.ai/insights/richard-suttons-bitter-lesson-explains-why-your-ai-solution-feels-shallow - Published: 2025-07-01 - Categories: AI Strategy, Machine Learning, Generative AI Your AI implementation feels shallow because it probably is. Most organisations struggle to extract even 20% of what's technically possible from their AI investments. The uncomfortable truth? Computing power and data repeatedly trump clever algorithms and domain expertise – a reality that most technical leaders resist until it's painfully obvious. Richard Sutton's "Bitter Lesson" illuminates why the most impressive AI breakthroughs consistently come from approaches that exploit computational scale rather than human knowledge engineering. While this insight has reshaped AI research, it remains tragically underutilised in commercial implementations. Let's explore why your meticulously crafted, domain-specific AI solution feels underwhelming, and what it would take to tap into the other 80% of untapped potential. #### The bitter lesson explained In 2019, Richard Sutton – a pioneer in reinforcement learning – articulated what he called "The Bitter Lesson" from 70 years of AI research: general methods that leverage computation ultimately prove most effective, and by an enormous margin. Time after time, across domains from chess to image recognition, AI researchers invested heavily in encoding human domain knowledge, only to be outperformed by approaches that prioritised scale, search, and learning. As Sutton writes: > This pattern has repeated across multiple domains: - In chess, Deep Blue's "brute force" search defeated Kasparov in 1997, despite domain experts insisting that strategic human knowledge was essential - In computer Go, AlphaGo triumphed using self-play and search, rendering decades of encoding human expertise irrelevant - In computer vision, hand-engineered features like SIFT gave way to convolutional neural networks that learn their own features - In speech recognition, statistical methods outperformed approaches based on detailed models of human phonetics The lesson is both powerful and psychologically difficult to accept: building in human knowledge provides immediate benefits but ultimately limits progress compared to scaling computational approaches. #### Why our intuitions about AI solutions are often wrong The allure of domain-specific knowledge engineering is powerful. It feels right to encode our hard-won expertise directly into AI systems. It's intellectually satisfying and delivers quick initial wins. This approach follows the natural human impulse to transfer our own mental models into machines. But these intuitions lead us astray when building production AI systems. Gary Marcus articulates this tension in his paper "The Next Decade in AI," where he acknowledges the limitations of both approaches. The human-knowledge approach tends to complicate methods in ways that make them less suited to leveraging computation effectively. The psychological biases that lead us down this path include: 1. The expert blind spot – we overvalue our domain expertise and undervalue what can be learned directly from data 1. Complexity illusion – we assume solutions must match the perceived complexity of the problem 1. Agency bias – we believe our conscious reasoning process is how intelligence "should" work These biases lead technical teams to over-engineer features, create brittle rule systems, and generally resist the Bitter Lesson's implications. #### The scale advantage in modern AI The rise of foundation models exemplifies Sutton's Bitter Lesson in dramatic fashion. Models like GPT-3 demonstrate how sheer computational scale reveals capabilities that couldn't be engineered directly. As the Stanford group behind "On the Opportunities and Risks of Foundation Models" notes: > These emergent capabilities – abilities not explicitly designed into the system – arise from scale in ways that domain experts consistently fail to anticipate. The transformer architecture that powers most modern language models doesn't incorporate linguistic theory; instead, it scales attention mechanisms across massive datasets. Consider the trajectory of deep learning breakthroughs documented by Bengio, LeCun, and Hinton in their seminal "Deep Learning for AI" paper. Early neural networks struggled with recognition tasks until: 1. Availability of large labelled datasets (ImageNet) 1. Efficient use of GPU computing power 1. Architectural innovations like ReLUs that facilitated training deeper networks 1. Techniques like dropout that improved generalisation None of these advances involved more sophisticated encoding of domain knowledge. Rather, they enabled existing algorithms to scale more effectively. #### Implications for AI practitioners and businesses The critical mistake most organisations make is optimising for initial performance rather than scalability. This leads to AI systems that deliver quick wins but plateau rapidly, precisely as Sutton's Bitter Lesson predicts. The path to exceeding that plateau requires rethinking your approach: ##### Evaluating the tradeoff matrix The most sophisticated AI implementations require making explicit tradeoffs between: 1. Domain customisation vs. leveraging foundation models 1. Initial performance vs. scaling potential 1. Explainability vs. raw predictive power 1. Control vs. emergent capabilities Most technical leaders optimise for the wrong elements of this matrix, creating sophisticated solutions that will inevitably be outperformed by approaches that better exploit computational scale. ##### Where domain knowledge actually matters Domain expertise isn't irrelevant – it's just more valuable when applied to: 1. Problem formulation (what questions to ask) 1. Data curation and evaluation 1. Constraining the search space 1. Interpreting and validating model outputs The key insight: use human knowledge to guide what the system learns rather than directly encoding that knowledge into the system. #### Finding the middle ground The bitter pill of Sutton's lesson doesn't mean abandoning all domain expertise. Rather, it suggests a neurosymbolic approach that combines the best of both worlds – using computational scale for perception and pattern recognition while employing symbolic reasoning for aspects where human knowledge provides genuine leverage. As Bengio and colleagues explain: > The breakthroughs came not from encoding visual expertise but from creating architectures that could learn efficiently at scale. ##### Extracting the other 80% Most organisations implement what we might call "shallow AI" – solutions that utilise familiar technology patterns without fully exploiting their computational potential. Extracting the remaining 80% requires: 1. Architecting for scale from the beginning 1. Using unsupervised pre-training techniques to leverage unlabelled data 1. Implementing transfer learning to build on foundation models 1. Creating data flywheel effects that improve with usage 1. Focusing technical expertise on where human knowledge genuinely complements computational approaches The Stanford paper on foundation models describes this emerging pattern: > This points to a future where technical expertise focuses on effectively adapting and constraining foundation models rather than building bespoke solutions from scratch. #### Reconciling with the bitter lesson The truly sophisticated AI implementations of the next decade will come from teams that have fully internalised Sutton's Bitter Lesson – not by abandoning human expertise, but by directing it toward problems where it genuinely complements computational approaches. The reality is that most AI solutions feel shallow not because they lack domain knowledge, but because they fail to fully exploit computational potential. They optimise for immediate performance rather than building foundations that can scale with computational resources. At a practical level, this means: 1. Investing more in data infrastructure than in clever algorithms 1. Building systems that improve with usage rather than static models 1. Focusing domain expertise on problem formulation rather than feature engineering 1. Creating architectures that can absorb computational resources effectively Internalising the Bitter Lesson is professionally challenging because it requires technical leaders to acknowledge that their hard-won domain expertise might be less valuable than they'd like to believe. But this realisation is the first step toward building AI systems that exploit the full technical potential of modern approaches. If you're ready to build AI solutions that exploit full technical potential rather than implementing basic features, we should talk. #### References - Sutton, R. (2019). The Bitter Lesson. - Marcus, G. (2020). The Next Decade in AI: Four Steps Towards Robust Artificial Intelligence. - Bengio, Y., LeCun, Y., & Hinton, G. (2021). Deep Learning for AI. Communications of the ACM, 64(7), 58-65. - Bommasani, R. et al. (2021). On the Opportunities and Risks of Foundation Models. --- ### Top AI consulting companies for 2025: the rise of boutique technical excellence - URL: https://agathon.ai/insights/top-ai-consulting-companies-for-2025-the-rise-of-boutique-technical-excellence - Published: 2025-06-21 - Categories: AI Strategy, Generative AI, Responsible AI The AI consulting landscape has fundamentally shifted since we wrote our round-up for 2024. What began as a market dominated by traditional management consultancies has evolved into something far more nuanced, with the global AI consulting market reaching $196 billion and 37.3% projected CAGR through 2030. The most significant trend? Small and medium enterprises are increasingly choosing boutique technical specialists over large generalist firms. #### The great consulting realignment The macro trends shaping AI consulting in 2025 tell a clear story. Only 1% of companies reach AI maturity despite widespread investment, whilst 82-93% of AI projects fail to deliver expected results. This execution gap has created unprecedented opportunity for smaller, technically-focused consultancies. 65% of businesses using generative AI now prefer consultants who actively participate in implementation rather than providing strategic advice alone. For small enterprises particularly, this means working with partners who understand both the technology and the practical realities of resource-constrained implementations. The shift is measurable: boutique AI consultancies are capturing market share by delivering what enterprises actually need—hands-on technical implementation over theoretical frameworks, faster time-to-value over lengthy strategy phases, and senior-level engagement over junior consultant execution. #### What small enterprises really need from AI consulting Small and medium enterprises face unique challenges that large consulting firms aren't equipped to address: Speed over scale: SMEs need rapid prototyping and iterative development, not 6-12 month strategy phases. They require partners who can move from concept to pilot within weeks, not quarters. Technical depth over broad coverage: 72% of organisations underestimate AI integration complexity. Small enterprises need consultants who understand everything from prompt engineering to advanced machine learning, not generalists with AI add-ons. Senior engagement throughout: Unlike large enterprises with dedicated AI teams, SMEs need direct access to senior expertise throughout the project lifecycle. They can't afford to work with junior consultants learning on their budget. Cost-effectiveness with proven ROI: Small enterprises need clear, measurable business outcomes tied to specific investments. Value-based pricing models focusing on 10-25% of measurable business improvement align consultant incentives with client success. Agile partnership models: SMEs require consultants who can pivot quickly based on market feedback and changing requirements, something impossible within rigid enterprise consulting frameworks. #### Evaluation criteria for 2025 AI consulting firms We believe the key criteria for evaluating AI consultancies have evolved significantly: ##### Technical depth and implementation capability Look for PhD-level expertise, published research, and hands-on MLOps experience. The best consultancies demonstrate full-stack competency from foundational ML to multimodal AI systems. ##### Consultative approach and agility Shorter decision paths, personalised service, and dedicated senior-level engagement distinguish boutique specialists from scale-focused firms. ##### Industry specialisation Domain-specific expertise enables accurate ROI projections and addresses sector-specific regulatory requirements. ##### Proven implementation methodology Agile methodologies starting with pilot projects, iterative development, and comprehensive change management reduce risk through incremental value delivery. ##### Measurable business outcomes Focus on consultants with established frameworks for measuring and optimising business outcomes, not just technical metrics. #### The Top 5 AI consulting companies for 2025 Be sure to also check out our previous listing for 2024 and our latest listings of AI consulting firms in 2026. ##### 1. McKinsey & Company - the strategic powerhouse Strengths: McKinsey's QuantumBlack division offers 140+ use case accelerators and leads in "agentic AI" development. Their strategic positioning expertise and enterprise relationships remain unmatched. Best for: Large enterprises requiring board-level strategic guidance and extensive change management across multiple business units. Limitations: Only 1% of companies reach AI maturity despite McKinsey's involvement, highlighting execution gaps. Premium pricing and 6-12 month strategy phases make them less suitable for SMEs needing rapid implementation. ##### 2. IBM Consulting - the technology integrator Strengths: $2 billion AI book of business with their Granite 3.0 LLMs and watsonx platform. Asset-based consulting provides technological depth that pure strategy firms lack. Best for: Enterprises with existing IBM infrastructure or complex technical integration requirements. Limitations: Enterprise focus creates slower decision-making cycles and higher overhead costs. Complex approval processes extend implementation timelines despite strong technical capabilities. ##### 3. Brainpool AI - an on-demand expert network Strengths: London-based network connecting 500+ experts from UCL, Oxford, Cambridge, Harvard, MIT, Stanford, and NYU through sprint-based project delivery. Precise talent matching for specialised requirements impossible through traditional consulting staffing. Best for: Organisations needing specific expertise for defined projects without long-term consulting relationships. HSBC and Fujitsu partnerships demonstrate enterprise validation whilst maintaining boutique agility. Unique approach: Network model enables access to world-class expertise at project-specific rates, combining academic research depth with commercial implementation experience. ##### 4. Cambridge Consultants - the scientific innovators Strengths: 800+ experts focusing on breakthrough AI applications with scientific research depth. AI assurance framework and edge AI specialisation serve clients requiring next-generation solutions. Best for: Organisations requiring cutting-edge research capabilities and breakthrough AI applications in regulated industries. Unique approach: Scientific rigour combined with commercial focus. Projects include UK Ministry of Defence autonomous systems and Hitachi predictive maintenance optimisation. ##### 5. Agathon Limited - the boutique technical excellence partner Why Agathon stands out: Agathon represents the new generation of AI consultancies that small enterprises are increasingly choosing over traditional alternatives. Technical breadth with business focus: Unlike large consultancies with AI add-ons, Agathon's expertise spans the full spectrum from prompt engineering to advanced reinforcement learning, whilst maintaining deep business consultative approach that gets under the skin of what clients actually need. Nimble and agile delivery: Whilst large firms require 6-12 month strategy phases, Agathon's agile approach enables rapid prototyping and iterative development. Senior-level experts remain directly involved throughout engagements, providing personalised service impossible at scale-focused firms. Competitive pricing with greater flexibility: Agathon offers more competitive pricing than premium-priced global firms whilst maintaining greater flexibility in engagement models. This enables SMEs to access world-class AI expertise without enterprise-level budget requirements. True partnership approach: Rather than the hands-off, outsourced model of large consultancies, Agathon provides collaborative partnerships with shared accountability for business outcomes. Specialist advantage: By focusing specifically on AI rather than being a generalist consultancy with AI capabilities, Agathon stays ahead of the latest techniques and frameworks, experimenting with cutting-edge approaches that larger firms are slower to adopt. #### Strategic recommendations for 2025 The evidence strongly favours boutique specialists for most AI implementations. When evaluating AI consulting partners, prioritise: Technical depth over broad consulting experience: Seek consultants with hands-on implementation experience rather than strategic advisory backgrounds. PhD-level expertise and proprietary technology development indicate genuine capability versus AI consulting add-ons. Industry specialisation providing domain expertise: Focus on firms with deep vertical knowledge understanding regulatory requirements and competitive dynamics in your specific sector. Agile partnership models balancing speed with strategy: Prioritise consultants offering rapid prototyping whilst maintaining strategic oversight. Sprint-based approaches reduce risk while delivering value incrementally. Measurable ROI frameworks proving value delivery: Select consultants with proven methodologies for measuring business outcomes, not just technical metrics. #### The future belongs to specialist excellence The AI consulting landscape has reached an inflection point. Large consulting firms struggle with the 82-93% AI project failure rate, whilst boutique specialists consistently deliver measurable results through technical excellence and agile implementation. For small and medium enterprises, the choice is increasingly clear: partner with specialists who combine deep technical expertise with business acumen, rather than generalists learning AI on your budget. The future of AI consulting belongs to firms that truly understand both the technology and the practical realities of implementation. --- Ready to explore how AI can transform your business? Contact Agathon now for a free initial consultation to discuss your specific needs and discover how our boutique approach delivers better results than traditional consulting models. --- ### Beyond the hype: creating measurable ROI with LLM implementations - URL: https://agathon.ai/insights/beyond-the-hype-creating-measurable-roi-with-llm-implementations - Published: 2025-05-31 - Categories: LLMs, Responsible AI, AI Strategy In the current AI landscape, Large Language Models (LLMs) are often hailed as transformative tools capable of revolutionising various industries. However, beneath the shiny veneer of their capabilities lies a stark reality: the rapid advancement of LLMs has outpaced traditional evaluation methods, leading to challenges in assessing their true value and return on investment (ROI). The question we must confront is whether these models can deliver tangible benefits or if they are merely a costly gamble in the AI roulette. #### Understanding LLMs and Their Capabilities ##### What Are Large Language Models? LLMs are sophisticated AI systems trained on vast datasets to understand and generate human-like text. Their applications range from content creation to customer service automation. Imagine deploying an AI that can draft reports, write poetry, or even engage in nuanced conversations—all tasks once reserved for human intellect. The allure is undeniable, yet the complexities of their deployment are often overlooked. ##### The Promise and Perils of LLM Implementations While LLMs offer impressive capabilities, their deployment without proper evaluation can lead to inflated performance metrics and unforeseen challenges. A model that appears to excel in a controlled environment may falter in real-world applications, resulting in lost time and resources. As the saying goes, "all that glitters is not gold." #### Evaluating the True Value of LLMs ##### The Limitations of Traditional Benchmarks Standard evaluation metrics often fail to capture the nuanced performance of LLMs, leading to misleading assessments. Traditional benchmarks may indicate a model’s proficiency, but they frequently overlook critical aspects such as contextual understanding and ethical considerations. Relying solely on these metrics can create a false sense of security about a model’s capabilities. ##### Developing Robust Evaluation Frameworks To accurately measure ROI, it’s essential to establish comprehensive evaluation frameworks that consider real-world applications and diverse datasets. A robust framework should encompass not just quantitative metrics but also qualitative insights from end-users, ensuring that the LLM aligns with specific business goals and user needs. #### Best Practices for Implementing LLMs ##### Aligning LLM Capabilities with Business Objectives Successful LLM integration requires a clear understanding of business goals and how LLMs can address specific challenges. It’s not enough to implement an LLM because it’s trendy; organisations must pinpoint how these models can drive efficiency, enhance customer engagement, or generate content that resonates with their audience. ##### Continuous Monitoring and Iterative Improvement Post-deployment monitoring is crucial to ensure LLMs adapt to evolving data and continue to deliver value. Regular assessments can uncover performance dips and biases that may emerge over time. This iterative approach not only safeguards against stagnation but also fosters innovation, as organisations continuously refine how they leverage LLMs. #### Ethical Considerations in LLM Deployment ##### Addressing Bias and Ensuring Fairness LLMs can inadvertently perpetuate biases present in their training data, necessitating strategies to mitigate such issues. Without vigilant oversight, organisations risk alienating customers or, worse, facing reputational damage due to unethical AI behaviours. Implementing fairness guidelines and bias detection tools is not just a regulatory requirement; it’s a moral imperative. ##### Transparency and Accountability Maintaining transparency in LLM operations fosters trust and accountability, essential for responsible AI deployment. Stakeholders must understand how these models operate and the data they rely on. A transparent approach not only builds trust but also enables organisations to make informed decisions about their AI strategies. #### Conclusion While LLMs hold significant promise, realising measurable ROI requires meticulous evaluation, strategic implementation, and ongoing oversight. It’s a complex journey, fraught with potential pitfalls. However, by adhering to best practices and ethical guidelines, organisations can harness the full potential of LLMs to drive tangible business outcomes. If you’re navigating the complexities of LLM implementation and seeking guidance, look no further than Agathon. Our consultancy has firsthand experience in transforming LLM hype into measurable success, and we’re here to help you every step of the way. #### References - Speed of AI development stretches risk assessments to breaking point - Don't Make Your LLM an Evaluation Benchmark Cheater - Guide to Evaluating Large Language Models: Metrics and Best Practices - Composio - Best Practices for Monitoring Large Language Models (LLMs) - Galileo AI --- ### Contextual chunking strategies that improve RAG performance - URL: https://agathon.ai/insights/contextual-chunking-strategies-that-improve-rag-performance - Published: 2025-04-15 - Categories: Generative AI, Machine Learning, Responsible AI The ability of AI systems to retrieve and generate information seamlessly is becoming not just an advantage, but a necessity. Enter Retrieval-Augmented Generation (RAG)—a cutting-edge technique that melds the best of both worlds. Yet, its success hinges on one often-overlooked aspect: contextual chunking. If you’re not leveraging effective chunking strategies, you’re likely leaving performance on the table. #### Understanding Retrieval-Augmented Generation (RAG) RAG operates at the intersection of information retrieval and generative AI. It fetches relevant data from a vast corpus and generates human-like text based on that context. At its core, RAG relies on the quality of the retrieved chunks of information. The more relevant and contextually rich these chunks are, the better the generated output. This is where chunking strategies come into play, significantly influencing RAG's efficacy. #### The significance of contextual chunking in RAG Contextual chunking is not merely a technical nicety; it’s a game-changer. By breaking down information into more digestible, context-rich segments, RAG systems can enhance both retrieval accuracy and generative quality. The challenge lies in choosing the right chunking strategy to avoid losing nuanced information while still maintaining computational efficiency. #### Evaluating chunking strategies ##### Fixed-length chunking This traditional method divides text into uniform segments. While it’s straightforward, it often fails to capture semantic nuances. Imagine trying to explain a complex concept using only snippets of a textbook, devoid of context. This approach can lead to disjointed responses and a lack of coherence in RAG outputs. ##### Semantic chunking Enter semantic chunking, which aims to group text based on meaning rather than arbitrary lengths. However, a recent study raises questions about its computational cost versus its benefits. The findings challenge the assumption that semantic chunking always yields superior results—an essential consideration for practitioners. ##### Hybrid chunking Hybrid chunking attempts to marry the best features of both fixed-length and semantic approaches. It combines structured segmentation with contextual awareness, but its effectiveness can vary based on the specific RAG application. It’s a middle ground, yet it’s not without its own complexities. #### Advanced contextual chunking techniques ##### Late chunking Late chunking is a novel strategy that employs long-context embedding models. By embedding the entire text before chunking, this approach maintains contextual integrity. The result? Significantly improved retrieval tasks without the need for extensive retraining. ##### Meta-chunking Meta-chunking goes a step further by recognising the need for a granular approach between sentences and paragraphs. It utilises linguistic connections to enhance chunk quality, effectively balancing performance and processing speed. This method shows promise in boosting RAG performance while conserving computational resources. ##### Mix-of-Granularity This technique combines various levels of chunking, ensuring that both fine-grained and coarse-grained segments are utilised. It allows RAG systems to adapt dynamically to the complexity of different texts, enhancing the overall retrieval and generative process. #### Best practices for implementing contextual chunking To maximise the benefits of contextual chunking, organisations should: - Conduct thorough evaluations of different chunking strategies to determine their impact on RAG performance. - Experiment with advanced techniques like late chunking and meta-chunking to find the right balance between context preservation and computational efficiency. - Foster a culture of continuous improvement, where chunking strategies are regularly revisited and refined based on evolving use cases and technological advancements. As the AI landscape matures, organisations that harness effective contextual chunking strategies will find themselves at a strategic advantage. At Agathon, we have firsthand experience with these methods and understand the nuances that can make or break your RAG implementation. If you have questions or need tailored advice, don’t hesitate to reach out. The future of RAG is here—let’s navigate it together. #### References - Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models - Is Semantic Chunking Worth the Computational Cost? - MoC: Mixtures of Text Chunking Learners for Retrieval-Augmented Generation System - Meta-Chunking: Learning Efficient Text Segmentation via Logical Perception --- ### Understanding how QLoRA works - URL: https://agathon.ai/insights/understanding-how-qlora-works - Published: 2025-04-14 - Categories: LLMs, Machine Learning, Generative AI In the rapidly evolving landscape of artificial intelligence, fine-tuning large language models (LLMs) has become a pivotal challenge. Enter QLoRA, a groundbreaking technique that not only streamlines the fine-tuning process but also significantly reduces memory usage without sacrificing performance. This efficiency is critical, especially as organisations increasingly rely on LLMs for various applications. #### Background Fine-tuning colossal models like GPT-3 or LLaMA can be an arduous task. The computational and memory constraints often render such endeavours impractical for many practitioners. High resource requirements mean that only a handful of organisations with extensive infrastructure can afford to fine-tune these behemoths. Previous methods, such as Low-Rank Adaptation (LoRA), attempted to alleviate these constraints by introducing low-rank matrices that adjust model weights. While effective, these solutions still require significant resources, leaving a gap for innovations that can make fine-tuning accessible to a broader audience. #### Core concepts of QLoRA ##### Quantization At the heart of QLoRA lies quantization, a process that reduces the precision of model weights, thus shrinking the memory footprint. By converting high-precision weights into lower-bit formats, QLoRA allows for substantial savings in memory without compromising the model's capabilities. ##### Low-Rank Adaptation (LoRA) LoRA complements quantization by employing low-rank matrices to make efficient adjustments to model weights. This technique facilitates fine-tuning by adding a manageable number of parameters, making it easier to adapt these grand models to specific tasks. #### How QLoRA works ##### Combining quantization and LoRA QLoRA ingeniously merges quantization with LoRA, creating a synergy that optimally fine-tunes large models. By backpropagating gradients through a frozen, quantised LLM, QLoRA utilises LoRA to adjust weights, achieving remarkable efficiency. ##### Innovations introduced by QLoRA Key innovations such as 4-bit NormalFloat (NF4) — a new data type optimised for normally distributed weights — and double quantization, which quantises quantisation constants, push the boundaries of what's possible. Paged optimisers further enhance memory management, tackling spikes in memory requirements during training. #### Advantages of QLoRA ##### Memory efficiency One of the standout benefits of QLoRA is its memory efficiency. It permits the fine-tuning of large models on GPUs with limited memory, opening the door for smaller organisations to participate in the LLM fine-tuning landscape. ##### Performance preservation Despite the downsizing of memory requirements, QLoRA maintains impressive performance levels. This duality ensures that users can achieve high-quality results without overextending their computational resources. #### Applications of QLoRA ##### Large-scale model fine-tuning QLoRA has already demonstrated its prowess in fine-tuning models with billions of parameters. For instance, its application has led to state-of-the-art results in various benchmarks, proving its efficacy in real-world scenarios. ##### Resource-constrained environments Moreover, QLoRA shines in resource-constrained environments. From startups with limited budgets to academic institutions, this technique empowers a wide array of users to harness the capabilities of LLMs without hefty investments. In conclusion, understanding how QLoRA works reveals a transformative approach to fine-tuning large language models. As organisations strive to leverage AI effectively, innovations like QLoRA will be instrumental in making advanced AI more accessible. Should you have any questions or seek guidance on implementing QLoRA, reach out to us at Agathon, where our direct client experience can help you navigate this complex terrain. #### References - QLoRA: Efficient Finetuning of Quantized LLMs - QuAILoRA: Quantization-Aware Initialization for LoRA - [2408.06634] Harnessing Earnings Reports for Stock Predictions: A QLoRA-Enhanced LLM Approach --- ### Small LLMs — why they matter - URL: https://agathon.ai/insights/small-llms-why-they-matter - Published: 2025-04-14 - Categories: LLMs, Responsible AI, NLP In an age where larger-than-life models dominate the AI landscape, small language models (SLMs) are quietly revolutionising the field. Often overshadowed by their heftier counterparts, these compact models offer a compelling blend of efficiency and versatility that could redefine how we approach language processing tasks. As industries grapple with the implications of responsible AI, the significance of SLMs is becoming increasingly evident. #### What are small language models? ##### Definition and characteristics Small language models typically encompass architectures with fewer than 5 billion parameters, striking a balance between performance and resource consumption. These models are designed to perform language tasks with remarkable agility, requiring significantly less computational power compared to large language models (LLMs). Their simplicity does not equate to ineffectiveness; in fact, advancements in training techniques and model compression have enabled SLMs to deliver robust performance on a variety of tasks. ##### Comparison with large language models While LLMs like GPT-4 boast billions of parameters and impressive capabilities, they come at a cost—both computationally and financially. SLMs, on the other hand, are engineered for agility. They are easier to deploy, require less energy, and often exhibit faster inference times. This trade-off makes SLMs an attractive option for businesses seeking effective yet economical solutions. #### Importance of small language models ##### Efficiency and resource optimisation In an era where sustainability is paramount, SLMs shine as champions of resource optimisation. Their reduced size translates to lower energy consumption and minimal computational overhead, making them ideal for organisations with limited resources. They empower teams to deploy AI solutions without the hefty infrastructural investments that typically accompany LLMs. ##### Accessibility and deployment Accessibility is another critical advantage of SLMs. With fewer resources required for deployment, small models can be integrated into a wider array of applications, from mobile apps to embedded systems. This democratisation of AI technology allows even smaller enterprises to harness the power of language processing, fostering innovation across sectors that were previously constrained. #### Applications of small language models ##### Domain-specific applications SLMs are not just versatile; they are highly adaptable to specific domains. In healthcare, for instance, they can be fine-tuned to interpret medical records or assist in diagnostics. In customer service, SLMs can power chatbots that understand nuanced queries, thereby enhancing user experience without the need for extensive computational resources. ##### On-device and edge computing With the rise of the Internet of Things (IoT), the demand for on-device processing has soared. SLMs are particularly well-suited for edge computing, enabling real-time language processing without reliance on cloud infrastructure. This not only enhances speed and responsiveness but also addresses privacy concerns by keeping sensitive data local. #### Challenges and future directions ##### Performance limitations Despite their advantages, SLMs are not without limitations. Smaller architectures may struggle with complex language tasks that require extensive context or nuanced understanding. As researchers continue to push the boundaries, the challenge will be to enhance performance while maintaining the efficiency that makes SLMs appealing. ##### Research and development opportunities The nascent field of SLMs presents a wealth of research opportunities. From exploring innovative compression techniques to developing hybrid models that combine the strengths of both small and large architectures, the potential for advancement is vast. As the demand for efficient AI solutions grows, so too will the focus on SLMs, positioning them as a critical area for future exploration. In conclusion, small language models matter more than ever. Their potential to democratise AI, optimise resources, and adapt to specific needs positions them at the forefront of the next wave of innovation. At Agathon, we have firsthand experience with SLMs and understand their transformative capabilities. If you have questions about leveraging small language models in your own projects, don't hesitate to reach out. #### References - A Survey of Small Language Models - Small Language Models: Survey, Measurements, and Insights - Small Language Models (SLMs) Can Still Pack a Punch: A survey - A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness --- ### Diffusion models: a simple explainer - URL: https://agathon.ai/insights/diffusion-models-a-simple-explainer - Published: 2025-04-13 - Categories: Generative AI, Machine Learning, NLP Diffusion models have emerged as a transformative class of generative models in artificial intelligence, particularly excelling in tasks such as image synthesis, video generation, and molecule design. Their ability to generate high-quality, diverse outputs has garnered significant attention across various domains. In a world increasingly driven by visual content and complex data generation, diffusion models offer an innovative approach that challenges conventional methods by harnessing the power of noise and reversibility. #### The mechanics of diffusion models At the core of diffusion models lies a two-step process that elegantly captures the essence of data generation: Forward process: This involves gradually adding noise to the data, effectively transforming it into a random noise distribution. Think of it as blurring a clear image until it becomes an unrecognisable mess. Reverse process: A neural network is trained to reverse this noising process, reconstructing the data step by step from the noise. This is akin to a skilled artist bringing clarity back to a foggy landscape, with each stroke unveiling the underlying image. This duality allows the model to learn the underlying data distribution, enabling the generation of new, similar data instances that can be astonishingly lifelike. #### Applications of diffusion models The versatility of diffusion models is nothing short of remarkable, with applications spanning various fields: - Image and video generation: They have set new benchmarks in generating high-resolution, photorealistic images and videos, pushing the boundaries of what's possible in creative expression. - Molecule design: In drug discovery, diffusion models assist in designing novel molecules with desired properties, accelerating the quest for new therapeutic solutions. - Text-to-image synthesis: Models like DALL·E 2 harness diffusion processes to transform textual descriptions into stunning visual representations, demonstrating a powerful synergy between language and imagery. #### Challenges and considerations Despite their impressive capabilities, diffusion models face certain challenges that must be navigated: - Computational demands: The iterative nature of the reverse process can be computationally intensive, requiring significant resources and time. This isn't just a minor inconvenience; it can be a barrier to accessibility for many potential users. - Training data quality: The quality of generated outputs is heavily dependent on the quality and diversity of the training data. Garbage in, garbage out is a maxim that rings true here. - Ethical implications: The ability to generate hyper-realistic images and videos raises profound concerns about misinformation and misuse. As we stand on the precipice of a new digital reality, ethical considerations must guide the deployment of these powerful tools. #### Future directions Ongoing research aims to address these challenges by: - Improving efficiency: Developing methods to accelerate the sampling process without compromising output quality is crucial. The future of diffusion models lies in their ability to deliver results quickly and effectively. - Enhancing data diversity: Curating more diverse and representative datasets will improve model generalisation, ensuring that outputs reflect the richness of the real world. - Ensuring ethical use: Implementing safeguards to prevent misuse and promote responsible AI development is not merely a good practice; it is an imperative for the sustainability of this technology. As the landscape of artificial intelligence continues to evolve, diffusion models stand at the forefront, promising to reshape our interaction with technology and creativity. If you have comments or questions about this exciting field, reach out to us at Agathon, your trusted AI consultancy, where we can navigate these transformative waters together. #### References - Diffusion Models: A Comprehensive Survey of Methods and Applications - Understanding Diffusion Models: A Unified Perspective - Diffusion Models in Vision: A Survey --- ### The Non-Executive Director's guide to assessing AI system performance - URL: https://agathon.ai/insights/the-non-executive-directors-guide-to-assessing-ai-system-performance - Published: 2025-04-13 - Categories: AI Strategy, Responsible AI, Machine Learning In the contemporary business landscape, Artificial Intelligence (AI) has swiftly transcended its role as a technological novelty to become a cornerstone of strategic decision-making. As organisations increasingly integrate AI into their operations, Non-Executive Directors (NEDs) find themselves at the helm of overseeing these AI initiatives. This guide aims to empower NEDs with the necessary insights and tools to effectively assess the performance of AI systems, ensuring alignment with corporate objectives and ethical standards. #### Understanding AI System Performance AI system performance can be encapsulated by four critical dimensions: accuracy, efficiency, scalability, and reliability. These elements form the backbone of any robust AI deployment. Key performance indicators (KPIs) such as precision, recall, F1 score, and throughput serve as quantifiable measures of an AI system’s effectiveness. Crucially, these metrics must be meticulously aligned with the overarching business objectives to ensure that AI systems not only perform well technically but also deliver tangible business value. #### Framework for Assessment A structured framework for evaluating AI systems is imperative for informed oversight. This framework comprises three core components: - Data Quality: The bedrock of AI accuracy lies in the quality of data it consumes. High-quality, relevant data is essential for training models that perform reliably in real-world scenarios. - Model Evaluation: Techniques such as cross-validation and A/B testing are instrumental in rigorously assessing model performance. These methodologies help in identifying potential weaknesses and areas for improvement. - Continuous Monitoring: Establishing ongoing checks and updates is vital for maintaining AI performance over time and adapting to changing conditions. #### Ethical and Governance Considerations Ethics play a pivotal role in AI performance assessment. NEDs must ensure that AI systems are free from bias, transparent in their operations, and accountable for their outcomes. Implementing governance frameworks that align with regulatory standards and organisational policies is essential for ensuring ethical AI deployment and mitigating risks. #### Engaging with AI Experts Collaboration between NEDs and AI specialists is paramount. Effective communication and alignment on performance metrics facilitate informed decision-making. Building a culture of trust and transparency around AI initiatives not only enhances performance assessment but also fosters innovation and ethical compliance. #### Conclusion Non-Executive Directors play a critical role in assessing AI performance, ensuring that AI systems align with business and ethical goals. As AI continues to shape the future of business decision-making, NEDs are encouraged to embrace AI assessment as a strategic priority. This proactive approach will not only safeguard their organisations against AI-related risks but also unlock new opportunities for innovation and growth. --- ### Creating custom benchmarks that align with business objectives - URL: https://agathon.ai/insights/creating-custom-benchmarks-that-align-with-business-objectives - Published: 2025-04-13 - Categories: AI Advisory, AI Strategy, Machine Learning "In the age of data-driven decision-making, why do so many businesses still rely on generic benchmarks that fail to reflect their unique goals?" Let's face it: relying on one-size-fits-all metrics in a world that celebrates individuality is like trying to fit a square peg in a round hole. Tailored benchmarks are not just a luxury but an essential ingredient for business success. Generic metrics, while easy to adopt, often miss the mark in capturing the nuanced goals and priorities of individual enterprises. #### Understanding the Importance of Custom Benchmarks Benchmarks serve as the yardsticks of business performance, enabling organisations to measure progress and identify areas for improvement. However, relying on off-the-shelf benchmarks is akin to using a compass designed for someone else's journey. Such generic metrics often fall short, failing to cater to the specific needs and goals of diverse businesses. They can lead to misguided strategies, rooted in comparisons that lack relevance. #### Identifying Business Objectives Creating custom benchmarks starts with a clear understanding of one's business objectives. Aligning metrics with strategy requires introspection and clarity about what success looks like for your organisation. For example, a tech start-up aiming for rapid user growth needs different benchmarks from a traditional manufacturing company focused on efficiency and cost reduction. Tailored benchmarks ensure that the metrics driving decisions are relevant and specific to the business context. #### Designing Custom Benchmarks Crafting these bespoke benchmarks involves a structured approach. Start by defining clear business objectives, then identify the key performance indicators (KPIs) that align with these goals. Consider industry-specific knowledge and contextual factors—what works for a financial institution may not apply to a retail chain. By incorporating these insights, businesses can design benchmarks that truly reflect their unique operational landscape and strategic priorities. #### Data Collection and Analysis Selecting the right data is crucial. Businesses must focus on collecting data that accurately reflects their performance and aligns with their objectives. Technologies like AI-driven analytics platforms can facilitate this process by offering sophisticated tools for data analysis. These technologies not only streamline data collection but also provide insights that inform better benchmark creation, ensuring the metrics are both meaningful and actionable. #### Implementing and Testing Benchmarks Before a full-scale rollout, pilot implementations are essential. Testing benchmarks on a smaller scale allows businesses to identify potential issues and refine the metrics. This iterative refinement ensures that benchmarks remain relevant and adapt to changing business conditions. Feedback loops are critical in this process, allowing for continuous improvement and adjustment. #### Conclusion In conclusion, businesses must move beyond generic benchmarks and embrace metrics that reflect their unique goals. By developing custom benchmarks, organisations can ensure alignment between data metrics and business objectives, driving strategic success. It's time to reassess current benchmarking strategies and adopt approaches that truly capture the essence of your business. Embrace the tailored, objective-aligned future, and watch your strategic decisions soar. --- ### Enterprise knowledge graphs as RAG foundations: implementation lessons - URL: https://agathon.ai/insights/enterprise-knowledge-graphs-as-rag-foundations-implementation-lessons - Published: 2025-04-13 - Categories: Knowledge Graphs, RAG, LLMs Enterprise knowledge graphs are the backbone of modern data intelligence, transforming disparate data silos into interconnected ecosystems of knowledge. Coupled with retrieval-augmented generation (RAG)—a technique that leverages large language models (LLMs) to enhance information retrieval and generation—they become indispensable in large organisations striving for efficiency. The integration of these technologies unlocks unprecedented data accessibility, empowering decision-making and operational performance. Simply put, knowledge graphs provide context and structure, while RAG amplifies this with intelligent querying and generation capabilities. #### Key considerations for implementing enterprise knowledge graphs as RAG foundations Data quality and consistency are non-negotiable. Without accurate, up-to-date information feeding into your knowledge graph, your RAG systems become akin to castles built on sand—impressive but ultimately unstable. It’s vital to establish rigorous data governance protocols to ensure the integrity of the data driving your AI decisions. Scalability and performance cannot be overlooked. As your organisation expands, the volume of data will balloon. Ensuring that your architecture can handle this growth without compromising speed or accuracy is paramount. A well-designed system should seamlessly scale, accommodating increased loads while maintaining responsiveness. Security and compliance are critical, especially in an era of stringent data protection regulations. Implementing robust security measures is essential to safeguard sensitive information. The last thing any organisation needs is a breach that undermines trust and incurs regulatory penalties. #### Lessons learned from real-world implementations From our experiences building enterprise-scale RAG solutions, one of the most significant insights is the value of modular and model-agnostic approaches. This flexibility allows organisations to pivot as technologies evolve and new models emerge, avoiding the pitfalls of vendor lock-in. As detailed in a recent study, simple adjustments in how knowledge base content is created can dramatically enhance RAG solution effectiveness. It’s not just about the tech; it’s about the content that fuels it. Moreover, traditional RAG benchmark evaluations often fall short when confronted with novel user queries. A "human-in-the-loop" approach that incorporates flexible monitoring and evaluation techniques can prove invaluable, ensuring that the system remains responsive and relevant. #### Best practices for successful implementation Collaboration is the cornerstone of successful implementations. Cross-functional teams that include data engineers, domain experts, and IT professionals are essential for ensuring that the knowledge graph meets the nuanced needs of the organisation. This multidisciplinary approach fosters innovation and alignment with business objectives. Continuous improvement should be embedded into the development culture. Regular updates to both the knowledge graph and RAG systems are vital to adapt to changing business landscapes. This iterative mindset not only enhances system performance but also keeps the technology relevant. User-centric design cannot be an afterthought. By prioritising user experience, organisations can create intuitive interfaces that facilitate the delivery of relevant and accurate information. A system that’s easy to use is more likely to be embraced by its users. #### Conclusion In summary, implementing enterprise knowledge graphs as foundations for RAG systems necessitates a focus on data quality, scalability, security, and continuous improvement. The lessons learned from real-world applications underscore the importance of modular design, effective content creation, and user-centric approaches. As we look to the future, emerging trends such as enhanced natural language understanding and more sophisticated AI ethics will undoubtedly shape the landscape of knowledge graphs and RAG. For organisations ready to embrace these changes, the potential is immense. If you have questions or need guidance on implementing enterprise knowledge graphs as RAG foundations, reach out to Agathon. Our direct client experience in this space can help you navigate the complexities of these powerful technologies. #### References - Optimizing and Evaluating Enterprise Retrieval-Augmented Generation (RAG): A Content Design Perspective --- ### Demystifying LoRA - URL: https://agathon.ai/insights/demystifying-lora - Published: 2025-04-13 - Categories: LLMs, Machine Learning, NLP In the rapidly evolving landscape of artificial intelligence, Low-Rank Adaptation (LoRA) is emerging as a game-changer, particularly in the fine-tuning of large language models (LLMs). Its ability to efficiently adapt pre-trained models with minimal computational overhead is not just a technical curiosity; it is a necessity in an era where model performance and resource optimisation are critical. This post aims to demystify LoRA, making it comprehensible for both technical aficionados and those new to the field. #### What is LoRA? LoRA is a sophisticated technique that leverages low-rank decomposition to adapt pre-trained models. By focusing on a low-dimensional representation of the model's weight updates, LoRA allows for substantial parameter reduction during the fine-tuning process. This leads to a more efficient use of resources while maintaining model performance. Historically, model adaptation techniques have evolved from simple fine-tuning to more complex methods, culminating in innovations like LoRA that cater to the demands of modern AI. #### The Technical Mechanism of LoRA At its core, LoRA utilises the mathematical principles of low-rank matrices. By approximating weight updates as low-rank matrices, it significantly reduces the number of trainable parameters, leading to faster convergence and lower memory requirements. During implementation, LoRA modifies the training process by integrating additional low-rank layers into existing architectures. This seamless integration ensures that the model retains its expressive power while becoming more adaptable to new tasks. #### Challenges and Limitations Despite its advantages, LoRA is not without limitations. In scenarios where the target task requires extensive domain-specific knowledge, LoRA may fall short compared to more comprehensive fine-tuning methods. Furthermore, ethical implications arise when deploying AI models in sensitive contexts—ensuring responsible innovation is paramount. Organisations must navigate these challenges carefully to avoid unintended consequences. #### Future Directions and Research Opportunities The future of LoRA is promising, as ongoing research continues to uncover new adaptation techniques that could further enhance its capabilities. As AI evolves, so too will the methodologies surrounding model adaptation. Researchers are encouraged to explore innovative applications of LoRA, particularly in areas where efficiency and adaptability are paramount, such as real-time language translation and personalised AI assistants. #### Conclusion Understanding LoRA is essential for anyone engaged in AI model adaptation. Its potential to transform how we fine-tune large models cannot be overstated. As the AI landscape continues to evolve, embracing such innovations will be crucial for staying competitive. We encourage you to share your experiences or questions regarding LoRA and its applications in the comments section. For further insights or consultancy on AI technologies, feel free to reach out to us at Agathon, your trusted AI consultancy. Let's navigate this exciting frontier together! --- ### Skills taxonomy for modern AI teams: beyond traditional data science - URL: https://agathon.ai/insights/skills-taxonomy-for-modern-ai-teams-beyond-traditional-data-science - Published: 2025-04-11 - Categories: Fractional CTO, Responsible AI, LLMs The landscape of artificial intelligence (AI) teams is evolving at an unprecedented pace. As AI technologies advance, the skills required to harness their full potential are transforming. A well-defined skills taxonomy is critical in helping teams adapt to new challenges and opportunities. This article dives into the necessity of expanding beyond traditional data science skills to build robust, versatile AI teams equipped for the future. #### The Limitations of Traditional Data Science Skills Historically, data science has centred around statistical analysis and machine learning. While these skills remain foundational, the rapid advancement of AI technologies has exposed their limitations. For instance, projects involving large language models (LLMs) or neuro-symbolic AI require expertise beyond conventional data science. Such projects necessitate skills in software engineering, data engineering, and a deep understanding of AI ethics and governance. #### Defining a Modern AI Skills Taxonomy To address these gaps, we propose a modern AI skills taxonomy, categorising essential skills into four main areas: Technical Skills: - Advanced Machine Learning: Mastery of cutting-edge algorithms and techniques. - Software Engineering: Proficiency in developing scalable and efficient AI systems. - Data Engineering: Expertise in managing and processing large datasets. - AI Ethics and Governance: Understanding of ethical frameworks and regulatory compliance. Domain Expertise: Knowledge specific to the industry or field where AI is being applied. Interdisciplinary Skills: - Neurosymbolic AI: Integration of symbolic reasoning with neural networks. - Human-Computer Interaction: Design of intuitive and user-friendly AI interfaces. Soft Skills: - Collaboration and Communication: Ability to work effectively in diverse teams. - Problem-Solving and Critical Thinking: Aptitude for tackling complex challenges. #### The Role of Interdisciplinary Collaboration Interdisciplinary collaboration is crucial in AI development, as it brings together diverse expertise. Successful AI projects often involve teams comprising data scientists, engineers, and domain experts. For example, a healthcare AI project might require collaboration between medical professionals and data engineers to ensure both technical and clinical insights are integrated. Strategies such as cross-functional workshops and collaborative platforms can foster effective teamwork. #### Continuous Learning and Skill Development In the rapidly evolving AI landscape, continuous learning is paramount. Organisations should invest in training programmes, online courses, and certifications to keep their teams updated with the latest advancements. Additionally, fostering a culture of lifelong learning and curiosity can drive innovation and maintain competitive advantage. #### Ethical Considerations in AI Team Skills Integrating ethical considerations into AI practices is essential for responsible innovation. Skills related to fairness, accountability, and transparency should be embedded within AI teams. These ethical skills not only enhance team performance but also ensure that AI systems are developed with societal impacts in mind, fostering trust and acceptance. #### Conclusion A modern AI skills taxonomy is essential for navigating the complex AI ecosystem. Embracing a broader skill set will empower teams to drive innovation and create impactful AI solutions. Organisations are encouraged to reassess their AI team compositions and skill requirements, ensuring they are equipped to meet the demands of the future. The convergence of technical expertise, interdisciplinary collaboration, and ethical considerations will define the next era of AI development. --- ### The fractional CTO's guide to building AI teams that deliver - URL: https://agathon.ai/insights/the-fractional-ctos-guide-to-building-ai-teams-that-deliver - Published: 2025-04-10 - Categories: Fractional CTO, Responsible AI, AI Strategy Artificial intelligence has swiftly transitioned from a futuristic concept to a critical driver of business transformation. As organisations increasingly rely on AI to gain competitive advantages, the demand for AI teams that can deliver tangible results is more pressing than ever. Within this landscape, the fractional CTO plays a pivotal role, leveraging their strategic oversight to shape and guide the development of robust AI teams that can meet organisational goals effectively. #### Understanding the role of AI in business AI's strategic impact on business operations and decision-making cannot be overstated. From enhancing customer experiences to optimising supply chains, AI initiatives must be intricately aligned with broader business objectives to deliver meaningful value. A fractional CTO can facilitate this alignment, ensuring AI projects are not just technically sound but also strategically relevant. #### Key characteristics of high-performing AI teams High-performing AI teams are distinguished by a blend of technical, analytical, and problem-solving skills. Beyond technical prowess, these teams thrive on diversity, which fosters innovation and comprehensive problem-solving. A diverse team brings varied perspectives and cognitive approaches, essential for tackling complex AI challenges effectively. #### Building the right team structure Deciding between centralised and decentralised team structures involves careful consideration of strategic implications. A centralised structure can offer consistency and control, while a decentralised approach may enhance agility and responsiveness. Optimal team composition should include roles such as data scientists, machine learning engineers, and AI ethicists, ensuring a holistic approach to AI development. #### Cultivating a culture of collaboration and innovation Promoting a culture of collaboration and innovation is vital for AI teams. Strategies that encourage experimentation and learning can drive innovation. Continuous learning and professional development are paramount to keep pace with the rapidly evolving AI landscape, ensuring teams remain at the forefront of technological advancements. #### Implementing effective AI project management Agile methodologies are well-suited to AI projects, providing the flexibility and iterative processes necessary for navigating the complexities of AI development. Best practices in AI project management focus on ensuring timely delivery and high-quality outcomes, balancing innovation with practical execution. #### Addressing ethical and responsible AI development Ethical considerations in AI development are paramount. Implementing frameworks for responsible AI practices within teams is crucial to mitigate risks and uphold organisational values. A fractional CTO can guide teams in embedding ethical considerations throughout the AI lifecycle, from design to deployment. #### Measuring success and impact Evaluating the success of AI initiatives requires well-defined key performance indicators (KPIs) that measure both technical achievements and business impact. Techniques for assessing the value delivered by AI teams are essential for demonstrating return on investment and guiding future AI strategies. #### Conclusion For fractional CTOs, sustaining high-performance AI teams involves a delicate balance between fostering technological innovation and ensuring ethical responsibility. By strategically guiding AI team development and aligning AI initiatives with business objectives, fractional CTOs can drive significant organisational impact, navigating the complexities of the AI ecosystem with precision and foresight. --- ### The strategic value of a Non-Executive Director with AI expertise - URL: https://agathon.ai/insights/the-strategic-value-of-a-non-executive-director-with-ai-expertise - Published: 2025-03-03 - Categories: AI Strategy, Responsible AI, Machine Learning In today's rapidly evolving business landscape, artificial intelligence is no longer a futuristic concept but a present-day competitive differentiator. For boards seeking to navigate this technological revolution, recruiting and retaining a non-executive director with AI expertise has become increasingly critical. Here's why this role matters and what unique value such directors bring to the boardroom. #### Deep Technical Insight Without Operational Burden Non-executive directors with AI expertise (AI NEDs) offer boards something invaluable: the ability to understand and evaluate complex AI initiatives without being involved in day-to-day operations. They bridge the knowledge gap between technical teams and board members who may lack specialised technical knowledge, translating complex AI concepts into business implications that the entire board can understand. #### Risk Management and Ethical Oversight AI implementation carries significant risks, from algorithmic bias to data privacy concerns. A board member with AI expertise can ask the right questions about these risks before they become problems. They can help establish appropriate governance frameworks and ethical guidelines for AI use, ensuring the company's AI strategy aligns with its values and regulatory requirements. #### Strategic Direction and Competitive Advantage These directors can identify emerging AI trends and opportunities that might otherwise go unnoticed. They understand how AI capabilities can enhance existing products, create new revenue streams, or transform business processes. This strategic foresight is crucial for maintaining competitive advantage in an increasingly AI-driven marketplace. #### Investment Evaluation When considering AI investments or acquisitions, having a board member who can evaluate the technical merit and long-term viability of AI solutions proves invaluable. They can distinguish between genuine innovation and hype, helping the board make more informed investment decisions. #### Talent and Culture Considerations AI transformation is as much about people as it is about technology. Directors with AI expertise understand the talent requirements and organisational changes needed to successfully implement AI. They can guide discussions on workforce planning, reskilling initiatives, and fostering an AI-ready culture. #### Conclusion As AI becomes increasingly central to business strategy, having dedicated expertise at the board level is no longer optional—it's essential. The right non-executive director with AI expertise doesn't just bring technical knowledge; they bring strategic vision, risk awareness, and the ability to guide companies through one of the most significant technological transformations in business history. The challenge isn't just recruiting such talent—it's retaining them in a competitive market. Companies that recognise the value these directors bring and actively engage their expertise will be better positioned to thrive in an AI-transformed future. --- ### An executive’s guide to AI agents - URL: https://agathon.ai/insights/an-executives-guide-to-ai-agents - Published: 2025-01-21 - Categories: AI Agents, Generative AI, AI Strategy In the rapidly evolving landscape of business technology, AI agents have emerged as pivotal players, redefining operational efficiency and strategic decision-making. Their significance lies not only in their ability to automate routine tasks but also in their transformative potential, particularly through generative AI capabilities. These agents can enhance productivity by streamlining workflows, enabling organisations to harness data-driven insights and ultimately improve performance across various functions. #### Understanding AI Agents ##### Definition and Purpose AI agents are software entities designed to autonomously perform tasks or assist users in completing complex workflows. While traditional automation tools execute predefined processes, AI agents leverage machine learning and natural language processing to adapt and learn from interactions, making them more versatile and effective in dynamic environments. ##### Types of AI Agents AI agents can be categorised into autonomous and semi-autonomous agents. Autonomous agents operate independently, executing tasks without human intervention, while semi-autonomous agents work alongside human operators, providing support and enhancing their capabilities. Examples of use cases include: - Customer Service: AI agents can handle queries, provide real-time assistance, and automate responses. - Human Resources: AI agents streamline recruitment processes by screening candidates and scheduling interviews. - Finance: AI agents facilitate data analysis, risk assessment, and reporting tasks, allowing finance professionals to focus on strategic initiatives. #### The Evolution of AI Agents ##### Historical Context The journey of AI agents began with rule-based systems in the 1960s, which evolved into more sophisticated machine learning models in the 21st century. Recent advancements in large language models (LLMs) and generative AI have catalysed a revolution in agent capabilities, enabling them to understand and respond to complex queries with human-like accuracy. ##### Current Trends Recent trends indicate a surge in the adoption of generative AI agents, which can orchestrate workflows, evaluate responses, and personalise interactions. For instance, companies like Lenovo have successfully integrated AI agents into customer service operations, resulting in significant productivity gains and improved customer satisfaction. #### Potential Benefits of AI Agents ##### Increased Efficiency and Productivity Studies have shown that organisations implementing AI agents can achieve productivity gains of 30 to 45%. For example, a McKinsey analysis revealed that the use of generative AI in customer service led to a 14% increase in issue resolution rates while reducing handling times by 9%. These improvements underscore the potential of AI agents to streamline workflows and reduce operational costs. ##### Enhanced Customer Experience AI agents play a crucial role in enhancing customer service by providing personalised interactions and timely responses. By leveraging historical data and machine learning, these agents can offer tailored solutions that meet individual customer needs, ultimately fostering loyalty and satisfaction. ##### Data-Driven Decision Making AI agents excel at processing vast amounts of data, enabling informed decision-making. However, the effectiveness of these agents depends significantly on data quality and integration. Organisations must ensure that data is clean, structured, and accessible to maximise the value derived from AI agents. #### Challenges and Considerations ##### Implementation Hurdles Despite their potential, organisations face several challenges when adopting AI agents. Common issues include data quality, employee resistance, and integration complexities. To overcome these hurdles, organisations should invest in robust data governance frameworks and engage employees through education and training initiatives. ##### Ethical Implications The deployment of AI agents raises important ethical considerations, including transparency, bias, and accountability. It is imperative for organisations to establish responsible AI development practices and governance frameworks to mitigate risks and ensure equitable outcomes. #### Strategic Recommendations for Executives ##### Developing a Clear AI Strategy Executives must define a robust AI strategy that aligns with organisational goals. This involves identifying key use cases, understanding the technological landscape, and assessing the potential impact on business functions. ##### Fostering a Culture of Trust Building trust in AI agents is essential for successful integration. Executives should promote transparency in AI operations and invest in training programmes that equip employees with the knowledge and skills to collaborate effectively with AI technologies. ##### Continuous Monitoring and Evaluation To ensure the ongoing success of AI agents, organisations must implement mechanisms for continuous monitoring and evaluation. Establishing key performance indicators (KPIs) and metrics will enable businesses to assess the impact of AI agents on productivity and customer satisfaction, allowing for iterative improvements. #### Conclusion AI agents possess the potential to revolutionise enterprise operations, driving efficiency, enhancing customer experiences, and enabling data-driven decision-making. As organisations embrace this technology, they must navigate the associated challenges with a strategic and ethical approach. Looking forward, AI agents will undoubtedly play an increasingly integral role in the future of organisational success, shaping the dynamics of work and customer engagement in profound ways. --- ### Reinforcement learning: practical guide for business users - URL: https://agathon.ai/insights/reinforcement-learning-practical-guide-for-business-users - Published: 2025-01-13 - Categories: Machine Learning, Reinforcement Learning, AI Advisory In an era markedly defined by rapid technological advancement, reinforcement learning (RL) has emerged as a pivotal component within the artificial intelligence (AI) landscape. Its capacity to enhance decision-making processes and automate complex tasks has ignited significant interest among business leaders. As organisations increasingly seek to harness AI-driven insights for competitive advantage, understanding the nuances of RL becomes essential. ##### Understanding Reinforcement Learning Reinforcement learning is a subset of machine learning where an agent learns to make decisions by interacting with an environment. The agent takes actions and receives feedback in the form of rewards or penalties, allowing it to refine its strategies over time. The key components of RL include: - Agent: The learner or decision-maker. - Environment: The context within which the agent operates. - Actions: The choices available to the agent. - Rewards: Feedback received from the environment based on the agent's actions. In contrast to supervised learning, which relies on labelled datasets, or unsupervised learning, which seeks to uncover hidden patterns in unlabelled data, RL focuses on learning optimal behaviours through trial and error. ##### How Reinforcement Learning Works The RL process is fundamentally about balancing exploration and exploitation. Exploration involves trying new actions to discover their effects, while exploitation focuses on utilising known actions that yield high rewards. This dynamic interplay is crucial for the agent’s learning process. Key algorithms driving RL include Q-learning, which is a value-based approach that helps the agent learn the value of actions in specific states, and Deep Q-Networks (DQN), which leverage neural networks to handle complex environments with high-dimensional state spaces. ##### Real-World Applications of Reinforcement Learning Reinforcement learning is not merely theoretical; its applications span various industries, demonstrating its versatility and effectiveness: - Finance: In algorithmic trading, RL can optimise trading strategies by learning market behaviours, while also assessing risk in real-time. - Healthcare: Personalised treatment plans can be developed by analysing patient data, enabling adaptive resource allocation to improve patient outcomes. - Supply Chain: RL algorithms can optimise inventory management, ensuring that stock levels are maintained efficiently while reducing costs. - Gaming: Intelligent game agents, powered by RL, can adapt their strategies to provide a more engaging experience for players, often outperforming human competitors. ##### Benefits of Implementing Reinforcement Learning in Business The implementation of reinforcement learning can yield substantial benefits, including: - Enhanced Decision-Making: RL equips businesses with the ability to make data-driven decisions that are adaptive to changing scenarios. - Increased Efficiency: Automation of complex processes leads to reduced operational costs and improved productivity. - Adaptability: As environments change, RL systems can adjust their strategies, ensuring continued effectiveness. ##### Challenges and Considerations Despite its potential, several challenges accompany the implementation of RL: - Data Requirements: RL often necessitates large volumes of data for effective training, which can pose a barrier for firms with limited resources. - Computational Resources: The complexity of RL algorithms can demand significant computational power, making initial setups costly. - Ethical Considerations: The opaque nature of RL decision-making raises concerns regarding safety and transparency, necessitating frameworks for responsible AI use. ##### Getting Started with Reinforcement Learning Business users keen on integrating RL into their operations should consider the following practical steps: - Identifying Use Cases: Explore areas where RL could enhance decision-making or automate processes. - Collaboration: Engage with data scientists and AI experts to develop a robust understanding of RL applications. - Pilot Projects: Start small with pilot projects to assess feasibility and measure outcomes before scaling. ##### Future Trends in Reinforcement Learning The field of reinforcement learning is evolving rapidly, with several emerging trends: - Integration with Neural Networks: Further advancements in deep learning techniques are likely to enhance the capabilities of RL. - Multi-Agent Systems: The development of systems where multiple agents learn and interact can lead to more sophisticated applications across industries. As businesses increasingly embrace these innovations, the future of RL promises to reshape strategic decision-making landscapes. ##### Conclusion Reinforcement learning stands poised to transform business processes, offering unprecedented opportunities for efficiency and adaptability. By recognising its potential and actively exploring its applications, organisations can position themselves to leverage RL as a strategic tool for competitive advantage. For those eager to delve deeper into the world of AI and machine learning, we invite you to subscribe for more insights. Additionally, consider reaching out to us for a consultation on how to effectively implement reinforcement learning within your organisation. Embrace the future of intelligent decision-making today! --- ### Common technical challenges fractional CTOs solve for startups - URL: https://agathon.ai/insights/common-technical-challenges-fractional-ctos-solve-for-startups - Published: 2025-01-12 - Categories: Fractional CTO, AI Consulting, AI Strategy As a startup grows, it inevitably faces technical hurdles that can make or break its success. Fractional CTOs bring seasoned expertise to solve these challenges without the commitment of a full-time executive. Here are the most common technical challenges they help resolve: ##### Technical Debt Management Many startups accumulate technical debt in their rush to market. Fractional CTOs excel at auditing codebases and creating practical strategies to reduce this debt without halting development. They might implement automated testing, refactor critical systems, or introduce coding standards that prevent future debt accumulation. ##### Scaling Infrastructure When user growth strains existing systems, fractional CTOs step in to architect solutions. They might migrate from monolithic to microservices architecture, optimise database performance, or implement caching strategies. ##### Security and Compliance Many startups postpone security concerns until they face their first crisis. Fractional CTOs proactively implement security best practices, from basic password policies to comprehensive data encryption. They're particularly valuable when preparing for SOC2 compliance or GDPR requirements, creating roadmaps that make these complex processes manageable. ##### Technology Stack Decisions Should you rebuild that legacy system? Is microservices architecture right for your scale? Fractional CTOs help make these crucial decisions, balancing factors like team expertise, maintenance costs, and scalability needs. They prevent expensive mistakes like over-engineering or choosing trendy but impractical technologies. ##### Building Technical Teams Finding and retaining technical talent is a universal challenge. Fractional CTOs help define roles, interview candidates, and structure teams effectively. They often implement engineering practices that make the company more attractive to top talent, such as modern CI/CD pipelines and clear career progression frameworks. ##### Cloud Cost Optimisation Cloud bills can spiral out of control as startups scale. Fractional CTOs typically find 30-50% cost savings through techniques like right-sizing instances, implementing auto-scaling, and optimizing storage usage. The value of a fractional CTO lies not just in solving these challenges, but in preventing them from becoming crises in the first place. They bring battle-tested experience that helps startups avoid common pitfalls while building sustainable technical foundations for growth. Would your startup benefit from this kind of technical leadership? Consider booking a technical audit with us to identify your most pressing challenges and opportunities for improvement. --- ### Your private LLM: deploying LLMs locally and offline using Ollama - URL: https://agathon.ai/insights/your-private-llm-deploying-llms-locally-and-offline-using-ollama - Published: 2025-01-11 - Categories: LLMs, AI Consulting Large Language Models (LLMs) have revolutionised numerous sectors, enabling applications ranging from intelligent chatbots to advanced content generation. However, as organisations increasingly rely on these powerful tools, the importance of local deployment emerges as a pivotal consideration—especially with respect to privacy, control, and efficiency. By deploying LLMs locally, organisations can mitigate risks associated with data exposure and enhance responsiveness. Ollama, a cutting-edge tool designed for local LLM deployment, provides a robust solution for users seeking to harness the potential of LLMs while maintaining oversight of their operational environment. #### Understanding Ollama Ollama is a platform that facilitates the deployment of LLMs on local machines, thereby empowering users to leverage these sophisticated models without relying on cloud infrastructure. Its primary purpose is to simplify the process of running LLMs offline, catering to users who prioritise data privacy and operational control. Key features of Ollama include its user-friendly interface, compatibility with various LLM architectures, and streamlined installation processes. In contrast to cloud-based deployment, which can introduce latency and data security concerns, Ollama offers a more immediate and secure alternative for organisations seeking to integrate LLMs into their workflows. #### Technical Requirements for Local Deployment To successfully deploy LLMs using Ollama, specific hardware and software prerequisites must be met. Recommended hardware specifications include a modern multi-core CPU, a minimum of 16GB RAM, and sufficient GPU resources for optimal performance—ideally, a dedicated graphics card with CUDA support for accelerated processing. On the software side, users will require a compatible operating system (Linux, macOS, or Windows), Docker installed for container management, and Python to facilitate model operations. Following these guidelines ensures a smooth installation and operation of Ollama. #### Step-by-Step Guide to Deploying LLMs with Ollama 1. Download and Install Ollama: Begin by visiting the Ollama website to download the latest version of the software. Follow the installation instructions provided for your respective operating system. 1. Configure the Local Environment: After installation, set up your Docker environment by pulling the relevant LLM containers. This process typically involves executing specific commands in the terminal. 1. Loading and Running an LLM Instance: With the environment configured, load your desired LLM using Ollama commands. For instance, a simple command can initiate the model and prepare it for input. 1. Example Use Cases: Once the model is running, users can explore various applications, such as building customised chatbots, generating marketing content, or conducting sentiment analysis on customer feedback. #### Advantages of Local Deployment The local deployment of LLMs through Ollama presents several advantages. Firstly, enhanced privacy and data security are paramount, as sensitive information remains on-site, reducing the risk of breaches. Secondly, organisations experience reduced latency, as model queries are executed locally, eliminating dependence on internet connectivity. Furthermore, local deployment grants users greater control over the model's behaviour and outputs, allowing for tailored adjustments and fine-tuning. This flexibility can lead to improved performance in specific applications, catering to unique organisational needs. #### Challenges and Considerations Despite its advantages, local deployment is not without challenges. Hardware constraints can limit the size and complexity of models that can be effectively utilised. Additionally, maintaining and updating LLMs requires ongoing technical expertise, which may necessitate dedicated resources. Ethical considerations also arise, particularly around data handling and the potential biases inherent to the models. Organisations must implement rigorous standards for data governance to mitigate these risks. #### Real-World Applications Numerous organisations have successfully deployed LLMs using Ollama, realising significant benefits. For example, a leading financial institution utilised Ollama to develop an internal chatbot for customer service inquiries. By deploying the LLM locally, the institution enhanced response times while ensuring compliance with data protection regulations. Other case studies reveal similar successes in sectors such as healthcare and education, where local deployment has facilitated improved user experiences and operational efficiencies. #### Future Outlook The landscape of local LLM deployment is evolving rapidly. As tools like Ollama gain traction, we can anticipate advancements in both hardware and software that will further enhance local deployment capabilities. Innovations in chip technology, for instance, could enable the execution of larger models with reduced resource requirements. Moreover, the ethical implications of local deployments will continue to shape industry standards, promoting responsible AI practices that prioritise transparency and accountability. #### Conclusion In conclusion, the local deployment of LLMs using Ollama represents a significant opportunity for organisations to enhance control, privacy, and efficiency in their AI initiatives. By considering local deployment strategies, organisations can better navigate the complexities of AI technology while safeguarding sensitive data. As the AI landscape continues to evolve, now is the time for organisations to explore and experiment with LLMs in a local context, harnessing their transformative potential. --- ### Understanding the AI Skills for Business Framework — dimensions and implementation (part 2) - URL: https://agathon.ai/insights/understanding-the-ai-skills-for-business-framework-dimensions-and-implementation-part-2 - Published: 2025-01-10 - Categories: Responsible AI, AI Advisory In our previous post, we explored the four key personas that organizations need to develop for successful AI adoption: AI Citizens, Workers, Professionals, and Leaders. Now, let's examine the five dimensions of AI competency that form the backbone of the framework, with a particular focus on what these mean for small and medium-sized enterprises (SMEs). #### The Five Dimensions: Practical Reality vs. Theoretical Ideal ##### Privacy and Stewardship This dimension represents perhaps the most critical starting point for any business, regardless of size. It encompasses data protection, security, and responsible data management. For SMEs, this can seem overwhelming – especially given the complexity of data protection regulations and the cost of compliance. The reality is that even small businesses can build strong foundations here. Start with basic data hygiene practices and gradually develop more sophisticated approaches. For instance, a local retail business might begin with proper customer data management and basic security protocols before moving toward more advanced data governance frameworks. ##### Technical Infrastructure This dimension covers data collection, engineering, and system architecture – aspects that often frighten smaller organisations due to perceived complexity and cost. However, the emergence of cloud services and "AI-as-a-Service" platforms has democratised access to sophisticated technical infrastructure. Small businesses don't need to build everything from scratch. A restaurant chain, for example, might start with off-the-shelf analytics tools for inventory management before gradually developing more customised solutions as their needs evolve. ##### Problem Definition and Communication Here's where many organisations, regardless of size, stumble. This dimension focuses on identifying business problems that AI can solve and effectively communicating about AI initiatives. It's not about technical complexity – it's about business clarity. We've seen small manufacturing firms excel here by clearly defining specific automation needs and maintaining open dialogue with their workforce about AI implementation. The key is starting with well-defined, contained problems rather than attempting comprehensive digital transformation all at once. ##### Problem Solving and Analysis This dimension might seem the most daunting, encompassing data analysis, modelling, and AI application. However, the framework recognises different levels of capability maturity. Small businesses can begin with basic analytics and gradually build more sophisticated capabilities as needed. A local marketing agency might start with simple AI-powered content analytics before progressing to more complex predictive customer behaviour models. The key is matching capability development to actual business needs rather than pursuing technical sophistication for its own sake. ##### Evaluation and Reflection Perhaps the most overlooked yet crucial dimension, this covers performance assessment, ethical considerations, and continuous improvement. It's particularly relevant for smaller organisations where the impact of AI initiatives can be more immediately felt and measured. #### The Reality Gap: Challenges for Small Business While the framework is comprehensive, implementing it in smaller organisations presents unique challenges: 1. Resource Constraints: Limited budgets and personnel make it difficult to develop capabilities across all dimensions simultaneously. 1. Expertise Gaps: Smaller organisations often lack specialised AI talent to guide implementation. 1. Competing Priorities: AI capability development must be balanced against immediate operational needs. #### Making It Work: A Practical Approach The key to successful implementation lies in taking an incremental, prioritised approach: First, assess your current capabilities across these dimensions. Many organisations have more existing capability than they realise – they just haven't formalised it. Second, identify the dimensions most critical to your immediate business needs. A retail business might prioritise data privacy and basic analysis capabilities, while a professional services firm might focus on problem definition and communication. Third, develop a staged implementation plan that builds capabilities progressively, aligned with business objectives and resources. #### The Path Forward While the framework might seem ambitious, especially for smaller organisations, it provides a valuable roadmap for AI capability development. The key is not to attempt everything at once but to build capabilities systematically over time. Successful AI adoption is about more than just technical implementation – it's about building sustainable capabilities aligned with business needs. Organisations should navigate this journey by: - Assessing current capabilities against the framework - Identifying priority areas for development - Creating customised implementation roadmaps - Providing ongoing support and guidance #### Taking Action The AI Skills for Business Framework provides a comprehensive guide for building AI capabilities, but you don't have to navigate it alone. Whether you're just starting your AI journey or looking to enhance existing capabilities, expert guidance can help you avoid common pitfalls and accelerate your progress. --- ### Understanding the AI Skills for Business Framework — leaders guide and personas (part 1) - URL: https://agathon.ai/insights/understanding-the-ai-skills-for-business-framework-leaders-guide-and-personas-part-1 - Published: 2025-01-09 - Categories: Responsible AI, AI Consulting, AI Advisory As artificial intelligence continues to reshape the business landscape, organisations face a critical challenge: ensuring their workforce has the right skills to leverage AI effectively. The AI Skills for Business Framework, developed through extensive collaboration between the Department for Science, Innovation and Technology and industry partners, provides a comprehensive roadmap for business leaders to understand and develop AI capabilities across their organizations. #### Why This Framework Matters Now The framework arrives at a crucial moment when businesses of all sizes are grappling with AI integration. It's not just about having technical experts – it's about building organisation-wide competency to ensure AI initiatives deliver real business value. The framework identifies four key personas that organisations need to develop: AI Citizens, AI Workers, AI Professionals, and AI Leaders. #### Understanding the Four AI Personas ##### AI Citizens: The Foundation These are your customers and general workforce who interact with AI in their daily lives. Think of a retail customer using an AI-powered shopping assistant, or an office worker using AI-enhanced productivity tools. They need to understand AI's basic capabilities and limitations to engage with it effectively and safely. ##### AI Workers: The Bridge These employees don't directly develop AI systems but use them in their roles. Consider a marketing manager using AI tools for campaign optimisation, a healthcare worker using AI-assisted diagnostic tools, or a financial analyst using AI for market trend analysis. They need deeper knowledge about AI's capabilities and limitations within their domain, along with the ability to critically evaluate AI outputs. ##### AI Professionals: The Builders These are your technical specialists who design, develop, and maintain AI systems. Examples include data scientists building predictive maintenance models in manufacturing, AI engineers developing natural language processing systems for customer service, or machine learning specialists optimising supply chain operations. They require deep technical expertise combined with strong ethical awareness and business acumen. ##### AI Leaders: The Strategists These individuals, often in C-suite or senior management positions, guide the organisation's AI strategy and governance. They need to understand AI's strategic implications, manage risks, and drive organisational change. For instance, a retail CEO needs to understand how AI can transform the customer experience while managing privacy concerns, or a manufacturing COO needs to balance automation opportunities with workforce impact. #### Cross-Industry Impact The framework's flexibility makes it applicable across sectors. In healthcare, AI Citizens might be patients using AI-powered health apps, AI Workers could be nurses using AI-assisted diagnostic tools, AI Professionals would be the teams developing these tools, and AI Leaders would be hospital administrators developing AI adoption strategies. In financial services, this might translate to customers using AI-powered banking apps (Citizens), financial advisors using AI for portfolio management (Workers), teams developing fraud detection systems (Professionals), and executives overseeing AI-driven digital transformation (Leaders). #### Moving Forward Understanding these personas is crucial for developing effective AI capabilities in your organisation. Each role requires different levels of AI competency, and organisations need to ensure they're developing the right skills at each level. In our next post, we'll dive deeper into the five dimensions of AI competency that these personas need to master: Privacy and Stewardship, Technical Infrastructure, Problem Definition and Communication, Problem Solving and Analysis, and Evaluation and Reflection. For business leaders, the key takeaway is clear: successful AI adoption requires a holistic approach to skills development across all levels of your organisation. It's not enough to just hire AI experts – you need to build AI literacy and competency throughout your workforce. Building these capabilities may seem daunting, but the framework provides a clear path forward. Start by identifying where these personas exist in your organisation and assess their current capabilities against the framework. This will help you develop targeted training and development programs to build the AI skills your organisation needs for the future. --- ### Unlocking business potential with AI agents - URL: https://agathon.ai/insights/unlocking-business-potential-with-ai-agents - Published: 2025-01-08 - Categories: AI Strategy, LLMs, AI Agents As businesses navigate an increasingly competitive landscape, AI-powered solutions offer a new frontier for growth. At the forefront of this revolution are AI agents—systems designed to autonomously handle complex tasks, enabling organisations to scale operations, improve efficiency, and reduce costs. In this post, we explore the fundamentals of AI agents, their practical applications, and how you can guide your organisation on its journey to AI-driven success. --- #### What are AI agents? AI agents are intelligent systems capable of perceiving their environment, using tools, and taking actions to achieve goals. Unlike rigid workflows that follow predefined paths, agents dynamically adapt their actions based on feedback and environmental changes. For instance, an AI agent in customer support doesn’t merely answer questions; it retrieves order details, processes refunds, and escalates issues to human agents when needed. Similarly, coding agents autonomously resolve complex software bugs by analysing code, making changes, and running automated tests. This flexibility and autonomy make agents indispensable for scaling operations and tackling open-ended challenges. --- #### Core components of AI agents 1. The Environment: Every agent operates within a defined environment, such as a website, database, or even a vehicle. 1. Tools and Capabilities: Tools like search engines, APIs, and data retrieval systems extend an agent’s ability to perceive and act. For example, an agent with access to a calendar API can schedule meetings. 1. Planning and Execution: Agents excel by breaking tasks into manageable steps, prioritising efficiency and accuracy. 1. Feedback and Reflection: Reflection mechanisms ensure agents learn from their mistakes, iterating until the desired outcome is achieved. #### Building effective AI agents To maximise the impact of AI agents, it’s crucial to follow a structured approach: - Start Simple: Begin with foundational systems like augmented LLMs, which combine LLM capabilities with external tools for enhanced functionality. - Introduce Complexity Gradually: As needs evolve, adopt advanced frameworks, e.g., for dynamic task delegation. - Tailor for Specific Use Cases: A customer support agent, for instance, may require APIs for accessing order history, while a coding agent needs tools for debugging and testing. #### Key challenges and how to overcome them While AI agents offer tremendous potential, businesses must address key challenges to ensure success: - Tool Integration: Effective tools are the backbone of any agent. Clear documentation, intuitive design, and robust testing are essential. - Error Handling and Reflection: Incorporate checkpoints where agents evaluate their progress and correct mistakes, minimising failures. - Trust and Security: Implement guardrails to prevent harmful actions, such as unauthorised data access or incorrect transactions. AI agents represent a transformative leap in technology, enabling businesses to achieve more with less. By leveraging these systems, organisations can unlock new efficiencies, scale operations, and stay ahead of your competitors. --- ### Building your AI-first company - URL: https://agathon.ai/insights/building-your-ai-first-company - Published: 2025-01-07 - Categories: AI Strategy, AI Advisory #### The new paradigm: lean, AI-powered efficiency The advent of generative AI is fundamentally reshaping how companies can be structured and operated. We're witnessing the emergence of exceptionally lean organisations achieving outsized success with minimal headcount. This represents a dramatic shift from traditional company scaling models, where growth typically demanded proportional team expansion. #### Leveraging AI for maximum efficiency Today's businesses can maintain incredibly lean operations by strategically deploying AI across multiple functions. Development and technical operations, for instance, can be enhanced using AI coding assistants to accelerate the development process, AI-powered testing tools to ensure quality assurance, and automated DevOps processes to streamline operations. Coupled with serverless architectures and cloud services, these tools allow for a fast-moving, cost-efficient technical foundation. Business operations, too, can see significant gains from AI. From AI-powered customer service tools to automated marketing and content generation platforms, startups can reduce overhead and improve response times. Financial planning and reporting become simpler with AI assistance, and HR processes can be streamlined with AI-driven tools for recruitment, payroll, and management. Product development also stands to benefit, as AI-driven market research and analytics provide insights with unprecedented speed and accuracy. Automated user testing and feedback collection enable faster iterations, while AI-assisted design tools create rapid prototypes. Through continuous analytics, organisations can maintain a loop of perpetual improvement, keeping their offerings competitive and relevant. #### The role of fractional CTOs For lean startups and companies, the fractional CTO model offers a blend of cost efficiency and strategic value. Fractional CTOs provide access to senior technical leadership without the expense of full-time hires. Their pay-as-you-need flexibility reduces overhead and management burden, ensuring technical decisions are guided by experience without overextending budgets. These professionals also bring strategic insights, including risk management, security expertise, and valuable network connections. Their guidance becomes especially critical during scaling phases, helping startups navigate technical architecture decisions, evaluate AI tools, and determine the right time to expand technical teams. This model allows you to stay nimble while accessing the expertise necessary for long-term success. #### Building your AI-first company Founders aiming to build lean, AI-powered startups should start with an AI-first mindset. From day one, core processes should integrate AI capabilities, supported by tools and platforms that enable seamless automation. Designing workflows around AI and fostering an AI-first culture ensures the company remains aligned with its efficiency goals. Focusing on core competencies is essential. Founders should identify tasks that require human touch and automate everything else, remaining lean in non-core areas. External expertise can be leveraged when necessary to avoid unnecessary overhead. At the same time, companies must plan for scale, choosing scalable AI platforms, designing automated processes for growth, and maintaining clean, well-documented systems. #### Future implications The rise of AI-powered lean startups brings profound implications for the business landscape. Market competition will intensify as faster time-to-market and lower capital requirements enable rapid experimentation and innovation. Talent requirements will shift, with a focus on strategic rather than operational roles, greater emphasis on AI literacy, and the demand for hybrid skill sets that combine domain expertise with technical acumen. The investment landscape will also evolve, with revised metrics for success, new efficiency expectations, and different scaling models reshaping valuation approaches. #### Conclusion The generative AI revolution is enabling a new breed of hyper-efficient companies that can achieve significant impact with minimal headcount. By leveraging AI capabilities across all functions and utilising fractional technical leadership, these companies can maintain lean operations while delivering substantial value. Success in this new paradigm requires rethinking traditional organisational structures and embracing AI not just as a tool, but as a fundamental component of company architecture. For founders willing to embrace this approach, the potential for building efficient, scalable, and impactful companies has never been greater. --- ### Process reward models: a simple explainer - URL: https://agathon.ai/insights/process-reward-models-a-simple-explainer - Published: 2025-01-03 - Categories: Generative AI, LLMs, Machine Learning #### Teaching AI step-by-step: How process reward models help AI learn Imagine you're training a puppy. You don't just wait until the end of the day to see if it did its business outside or chewed on the furniture. You give it treats and praise for good behaviour throughout the day, right? This way, the puppy learns what actions lead to positive outcomes. Similarly, researchers are trying to teach AI systems by rewarding them not just for the final answer, but also for the steps they take to get there. This is where process reward models (PRMs) come in. #### What are PRMs? PRMs are like trainers for AI systems. They provide feedback at each stage of a task, guiding the AI towards the right solution. This is especially helpful for complex tasks that require multiple steps, like solving a math problem or writing a persuasive essay. #### Why are PRMs important? Traditional AI models often struggle with these complex tasks. They might get the final answer right by chance, but they may not understand the underlying process. PRMs help them develop that understanding by rewarding them for taking the correct steps along the way. #### Recent Research in PRMs Researchers are constantly improving PRMs. Here are some recent advancements: - More nuanced feedback: New PRMs are being developed that can provide more specific feedback at each step. This helps the AI understand not just if it's on the right track, but also how to get even better. - Better data collection: Training PRMs requires a lot of data. Researchers are working on ways to collect this data more efficiently and effectively. #### The Future of PRMs PRM research is a rapidly evolving field. With continued development, PRMs have the potential to revolutionise the way AI systems learn and solve problems. For example: - OpenAI discuss a new approach called Process Q-value Model (PQM) that shows promise in giving more accurate feedback at each step. - Patrick McGuinness summarises some recent research in this space and explores how PRMs can be used to improve the reasoning abilities of LLMs. - Nathan Lambert discusses the challenges of collecting data for training PRMs. --- ### Should I build my own large language model (LLM)? - URL: https://agathon.ai/insights/should-i-build-my-own-large-language-model-llm - Published: 2024-11-28 - Categories: Generative AI, AI Strategy, LLMs The question "Should we build our own LLM?" reveals a fundamental misunderstanding of what most organisations actually need. After architecting AI systems that exploit advanced capabilities like hierarchical context management and self-improving workflows, the real question isn't whether to build an LLM—it's whether you're maximising the technical potential of AI solutions that already exist. #### The real problem: most AI products use 20% of what's possible Whilst organisations debate building custom LLMs, they're missing the bigger opportunity. The majority of AI implementations barely scratch the surface of available capabilities. Companies spend months building basic chatbots when they could be developing sophisticated systems with guided user journeys, contextual memory, and adaptive learning patterns. Building your own LLM is like manufacturing your own microprocessors when you haven't yet mastered software architecture. #### When building makes strategic sense Custom LLM development becomes viable when you meet these criteria: Technical Prerequisites: - Sophisticated AI development team with deep ML expertise - Substantial computational infrastructure (multi-million pound investment) - Access to high-quality, domain-specific training datasets - 18-24 month development timeline without revenue pressure Business Justification: - Unique domain requirements that no existing model addresses - Data sovereignty regulations requiring complete control - Long-term competitive advantage dependent on proprietary AI capabilities - Volume economics where custom development becomes cost-effective Most organisations claiming these prerequisites are overestimating their needs and underestimating the complexity. #### The sophisticated alternative: maximising existing potential Instead of building from scratch, focus on exploiting the full technical potential of available AI capabilities: Advanced implementation approaches: - Hierarchical context management for complex decision-making workflows - Self-improving prompting systems that evolve with usage patterns - Novel interaction patterns that transform how users engage with AI - Multi-agent architectures that coordinate specialised AI capabilities Strategic technical decisions: - Fine-tuning open-source models for specific domains - Implementing retrieval-augmented generation (RAG) with sophisticated knowledge graphs - Building custom evaluation frameworks for domain-specific performance - Developing hybrid approaches that combine multiple AI capabilities #### Due diligence questions Before making any LLM investment decision: 1. Technical assessment: What percentage of available AI capabilities does your current implementation exploit? 1. Competitive analysis: Could sophisticated implementation of existing models provide the same competitive advantage? 1. Resource evaluation: Do you have the technical expertise to distinguish between superficial and genuine AI innovation? 1. Strategic alignment: Does custom LLM development align with your core business capabilities and competitive positioning? #### The investment reality Building production-ready LLMs requires: - £2-10M+ in computational resources - 12-18 months with expert ML team - Ongoing operational costs exceeding £100K monthly - Risk of technological obsolescence as the field rapidly evolves Most organisations would generate superior returns investing these resources in sophisticated implementation of existing capabilities. #### Strategic recommendation For 95% of organisations, the answer is clear: Don't build your own LLM. Build AI products that fully exploit what's already possible. Focus on: - Product definition: Bridging ambitious vision with technical reality - Sophisticated architecture: Implementing advanced capabilities most competitors ignore - Strategic implementation: Creating competitive advantage through technical depth, not custom models The organisations succeeding with AI aren't those building custom LLMs—they're those maximising the potential of existing capabilities whilst their competitors implement basic chatbots. #### Getting strategic clarity The LLM decision requires technical due diligence that distinguishes between solutions that exploit AI's full potential and those representing typical shallow implementations. This isn't a technology decision—it's a strategic business decision requiring expertise in both advanced AI capabilities and competitive positioning. Before committing resources to custom LLM development, ensure you're not missing opportunities to create breakthrough products with existing technologies. The most sophisticated AI implementations often involve no custom model development whatsoever. --- Need expert evaluation of your AI strategy and technical approaches? Schedule a call with us to discuss technical due diligence for AI investments and strategic guidance for building sophisticated AI products that maximise available capabilities. --- ### What does a fractional CTO do? - URL: https://agathon.ai/insights/what-does-a-fractional-cto-do - Published: 2024-11-04 - Categories: AI Consulting, AI Advisory, Fractional CTO --- A fractional AI CTO is a senior AI leader who works with you part-time. They decide where AI is worth building, where to buy instead, and how to run it safely once it ships. You get executive-level judgment without the cost or commitment of a full-time hire. That last part matters more than it sounds. The hardest calls in AI are rarely technical. They are decisions about where to spend money and where to walk away, and a fractional AI CTO makes them with you, including the call to not build at all. The role suits mid-market companies starting their first serious AI work, startups that need AI credibility without a full executive salary, and traditional businesses being told to "do something with AI" and unsure what that something is. #### What a fractional AI CTO actually decides Most descriptions of this role list activities, which fill a page and tell you nothing. What you are paying for is a short list of decisions, each of which can save or waste a large amount of money. ##### Where AI is worth it, and where it isn't The first job is finding the few places in your business where AI changes an outcome that matters, and naming the places where it won't. Most "AI use cases" presented to me fail one test: if the model were 20% less accurate, would any decision change? If not, you don't have a use case. You have a technology you fancy. A good fractional AI CTO kills those ideas before they consume a budget. ##### Build or buy For almost every AI capability you want, someone already sells a version of it. The expensive mistake is building what you could have bought, usually because building feels more impressive. The job is to make this call without ego: buy the commodity, build only what is genuinely yours to build. A consultant whose income depends on a long build will rarely tell you to buy, which is exactly why the judgment has to be independent. ##### What data you actually have AI plans tend to assume data that is clean, accessible, and cleared for use. Reality is messier. Part of the job is an honest read of your data and technical infrastructure before anyone promises what a model can do with it, so the roadmap survives contact with your real systems. ##### How it gets governed Once something ships, someone has to own how it behaves: where it can fail safely, what happens when it gets something wrong, who is accountable, and whether it meets the rules that apply to you. In the UK and EU that increasingly means real regulatory exposure, not a box-ticking exercise. A fractional AI CTO sets this up while it is still cheap to get right. ##### How your team levels up A good engagement makes itself progressively unnecessary. The aim is to leave your engineers and analysts more capable than they were: better at choosing tools, evaluating vendors, and spotting AI claims that don't hold up. Done well, this means building an AI team that delivers instead of one that stalls. #### Fractional Head of AI vs fractional AI CTO These titles get used interchangeably, but the one you reach for usually reveals the problem you have. Hire a fractional AI CTO when AI is part of a broader technical leadership gap, or when there is no senior technical owner at all. The remit is the whole technical picture, with AI inside it. Hire a fractional Head of AI when your engineering is solid but nobody owns AI as a discipline. The remit is narrower and deeper: AI strategy, capability, and governance. A related title, fractional Chief AI Officer, appears in larger organisations where AI warrants its own seat at the table; the work overlaps heavily and the title mostly signals seniority. In short: "no senior technical direction" points to a fractional CTO, "engineering is fine but AI keeps stalling" points to a fractional Head of AI. If you have an established team and want to know how a dedicated AI leader slots in alongside it, we cover that in full here. #### When to hire a fractional AI CTO A few signals tend to show up together. Any one is usually survivable. Two or three at once is the point to act. ##### You can't judge which AI suggestions are real Pressure is coming from the board, customers, or competitors, and nobody senior enough can separate the genuine opportunities from the noise. ##### You are about to spend serious money on a build Nobody internal can sanity-check it. This is the most expensive moment to lack senior judgment, and the most common one. ##### Your AI work keeps stalling Pilots that never reach production, models that work in a demo but not in the business, projects that quietly lose momentum. Usually a leadership gap, not a talent gap. ##### You need AI credibility for someone external Investors doing due diligence, a regulator, an enterprise customer: someone who can speak to your AI with authority and survive scrutiny. ##### You are not ready for a full-time hire The need is real but not yet a permanent role, and hiring one prematurely is its own expensive mistake. #### What does a fractional AI CTO cost? A fractional or part-time CTO model lets you buy senior judgment in proportion to your need: the decisions without the full executive salary the role would otherwise demand. ##### How engagements are structured Engagements take one of three shapes. A retained advisory arrangement gives you ongoing access at a set number of days per month, suited to steady oversight of a live AI programme. A project engagement is scoped to a specific decision or build, with a defined start and end. A discovery engagement is a short, paid first piece of work to assess where you are and what is worth doing, which often makes sense before committing to anything larger. ##### What drives the cost Cost tracks seniority, days per month, and how AI-specific the expertise needs to be. Judge the day rate against the decisions this person is there to get right: a six-figure build that should never have started, a vendor contract signed on bad assumptions, a model shipped into a regulated market without governance. Senior judgment is rarely the expensive line in an AI budget. The avoidable mistakes are. #### What the right one looks like The market is full of people who will build you AI. Far fewer will tell you not to. The fractional AI CTO worth hiring has real, hands-on technical depth, enough business context to weigh a build against its return, and the independence to say "buy this" or "don't build it" when that is the honest answer. ##### What this looks like in practice We took a ghostwriting firm from no technical team, no code, and no spec to a patent-pending AI product in active beta, over a 20-month fractional CTO engagement. We owned the requirements, architecture, and development cadence, recruited the team from scratch, and made the calls that turned a services business into a defensible technology asset. Read the full case study: zero to beta in 20 months. --- ### Who are the best AI consulting firms in 2024? - URL: https://agathon.ai/insights/who-are-the-best-ai-consulting-firms-in-2024 - Published: 2024-10-27 - Categories: AI Consulting, AI Strategy As artificial intelligence (AI) continues to transform businesses across industries, finding the right consulting partner to guide your AI initiatives is crucial. With a myriad of options available, it can be challenging to identify the best fit for your organisation. We'll explore some key factors to consider when evaluating AI consulting firms, as well as highlight a few leading players in the space that are worth your attention. Be sure to check out our listing of AI consulting fims from 2025 and our latest list of top AI consultancies for 2026. #### What to look for in an AI consulting company When searching for an AI consulting partner, there are several important criteria to assess: Expertise and Experience: Look for firms that have extensive experience and a proven track record in deploying AI solutions across a variety of industries and use cases. They should have a team of seasoned data scientists, machine learning engineers, and AI strategists. Industry Focus: Consider firms that specialize in your particular industry and understand the unique challenges and requirements. This domain expertise can be invaluable when developing tailored AI solutions. End-to-End Capabilities: Seek out firms that can handle the full AI lifecycle - from data preparation and model development to deployment and ongoing maintenance. This end-to-end approach can streamline the process and ensure seamless integration. Innovation and Thought Leadership: The best AI consultants stay ahead of the curve, experimenting with the latest AI techniques and frameworks. Look for firms that contribute to the broader AI community through research, publications, and industry events. #### Top AI consulting firms to know Now, let's take a closer look at some of the leading AI consulting firms that excel in these areas: 1. McKinsey & Company: A global management consulting firm with a strong AI and analytics practice, helping clients across industries leverage AI for competitive advantage. 1. Deloitte: One of the "Big Four" professional services firms, with a dedicated Artificial Intelligence and Cognitive practice that delivers AI-powered business transformation. 1. IBM: A global management consulting and professional services company, recognised for its innovative AI solutions and industry-specific AI accelerators. 1. Booz Allen Hamilton: A management and technology consulting firm with a specialized AI and Analytics practice, supporting clients in areas like computer vision, natural language processing, and predictive analytics. 1. Wipro: An Indian multinational IT services and consulting company, known for its comprehensive AI and Machine Learning capabilities across a wide range of industries. These are just a few examples of the top AI consulting firms to consider. As you evaluate your options, be sure to thoroughly research each firm's specific capabilities, industry focus, client testimonials, and alignment with your own AI initiatives and requirements. Selecting the right AI consulting partner can be a game-changer for your organisation, unlocking new opportunities for innovation, efficiency, and competitive advantage. Take the time to find the best fit, and you'll be well on your way to realising the full potential of AI within your business. #### Why you should consider choosing Agathon for your AI consulting While the global management consulting firms and large IT services providers mentioned offer impressive scale and resources, there are also benefits to working with a smaller, more specialised AI consultancy like Agathon. Nimble, agile firms like our own provide a more personalised, high-touch approach, with senior-level experts directly involved throughout the engagement. We can move more quickly, experiment with the latest AI techniques, and tailor solutions more closely to your unique business needs. We can be especially advantageous for organisations seeking a truly collaborative partnership, as opposed to a more hands-off, outsourced model. Additionally, as a boutique AI consultancy we can offer more competitive pricing and greater flexibility compared to the premium pricing of the largest players. When evaluating your options, be sure to weigh the tradeoffs between the scale and resources of the larger firms versus the agility and personalised service of specialised, mid-sized AI consultancies like Agathon Limited. Contact us now for a free initial consultation to discuss your needs. --- ### Can you reason with LLMs? - URL: https://agathon.ai/insights/can-you-reason-with-llms - Published: 2024-10-20 - Categories: Machine Learning, Generative AI In a paper from Apple titled GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models the authors examine the ability of large language models (LLMs) like GPT and similar systems to solve math problems, particularly using reasoning rather than just pattern recognition. Here’s our primer on the paper: ##### 1. Problem with Current Benchmarks - LLMs are often tested on a popular math dataset called GSM8K, which includes grade-school-level questions. However, simply scoring well on GSM8K doesn’t necessarily mean these models understand math or can reason logically. Many LLMs may perform well simply by memorising question patterns and answers rather than actually solving the problems. - The authors developed an improved testing method, called GSM-Symbolic, to better evaluate whether these models are genuinely reasoning through problems or just relying on patterns. GSM-Symbolic introduces variations in question phrasing and numbers to challenge models more rigorously. ##### 2. Testing Mathematical Reasoning Skills - When tested on GSM-Symbolic, many models performed inconsistently, especially when only the numbers in questions were changed. This suggests that the models were thrown off by changes in variables, which reveals their lack of flexible problem-solving abilities. - The study also found that as questions got more complex—by adding more clauses or steps—the performance of LLMs declined. This drop happened because many LLMs aren’t capable of true logical reasoning but instead try to mimic reasoning by recognising patterns from training data. ##### 3. New Challenges Introduced: GSM-NoOp - To further explore model limitations, the authors created another test called GSM-NoOp. This dataset added extra, irrelevant details to math problems (called "no-op" information) to see if the LLMs could ignore this unnecessary information. - Most models struggled with GSM-NoOp, often getting confused and incorrectly factoring in these irrelevant details into their calculations, showing they lack true reasoning skills. ##### 4. Key Findings - High Sensitivity to Changes: The LLMs performed poorly when minor changes were introduced to the questions, suggesting that they may not genuinely understand the math problems. - Struggle with Complexity: The more complex a problem became, the worse the models performed, indicating that LLMs are not yet ready to tackle truly challenging logical problems. - Pattern Matching Over True Reasoning: The study suggests that current LLMs tend to match patterns rather than engage in actual reasoning. This means they might answer correctly in familiar situations but fail in new or slightly altered scenarios. ##### 5. Conclusion - The research highlights that while LLMs have made progress in handling math problems, they still rely heavily on recognising familiar patterns. The study emphasises the need for better evaluation methods and further improvements in model design so that future LLMs can achieve genuine reasoning capabilities. In essence, while LLMs show some promise, they still have a long way to go in terms of true logical reasoning, particularly for complex or unfamiliar math problems. This research helps pave the way for developing models that can genuinely understand and solve problems rather than just mimicking patterns. --- ### Understanding large language models: a group discussion analogy - URL: https://agathon.ai/insights/understanding-large-language-models-a-group-discussion-analogy - Published: 2024-09-01 - Categories: Generative AI, Machine Learning, LLMs As machine learning models have become more powerful and complex, it's important to develop intuitive ways to understand how they work. One of the most influential neural network architectures in recent years is the transformer, which has revolutionised the field of natural language processing (NLP). To help demystify the inner workings of the transformer, let's imagine a scenario that we can all relate to - a group discussion or conversation. #### The Participants and the Input Visualise a room full of people engaged in a lively discussion. Each person represents a single transformer block, and the entire conversation is the input sequence that the transformer model is processing. Just like in a real conversation, each participant is listening attentively to what the others are saying, considering the context of the discussion, and then formulating their own response. This is the core functionality of a transformer block. The input to the transformer is the sequence of statements or questions being exchanged between the participants. In an NLP task, this could be a sentence, a paragraph, or even an entire document. #### The Attention Mechanism The key innovation in the transformer architecture is the attention mechanism, which is akin to how each person in the discussion actively pays attention to the most relevant parts of what the others are saying. Imagine one of the participants, Alice, is about to speak. Before she responds, she listens closely to the other participants, weighing the importance of the different points they've made. She doesn't treat each statement equally - instead, she focuses her attention on the parts of the discussion that are most relevant to what she wants to say next. This selective attention is the core of the transformer's attention mechanism. Rather than processing the input sequence in a rigid, sequential manner, the transformer dynamically determines which parts of the input are most important for producing the desired output. #### The Output and the Feedback Loop After carefully considering the conversation, Alice formulates her response and shares it with the group. This output from Alice's "transformer block" now becomes part of the input sequence that the other participants (transformer blocks) will attend to when it's their turn to speak. This cyclical process of attending to the relevant parts of the input, producing an output, and then having that output become part of the new input sequence is at the foundation of how transformers operate. #### The Power of Parallel Processing One of the key advantages of the transformer architecture is its ability to process the input sequence in parallel, rather than sequentially like traditional recurrent neural networks (RNNs). In our conversation analogy, this means that each participant can listen to and process the entire discussion simultaneously, rather than having to wait their turn to speak. This parallel processing enables transformers to capture complex relationships and dependencies in the input data much more effectively. #### The Role of Self-Attention An essential component of the transformer's attention mechanism is self-attention, which allows each participant to not only attend to the other speakers but also to their own previous contributions to the discussion. Imagine Alice reflecting on her own past statements, recognising how they relate to the current topic, and using that self-awareness to inform her next response. This self-referential process is a key part of how transformers are able to maintain context and coherence in their outputs. #### The Transformer in Action When you put all of these pieces together - the participants, the attention mechanism, the cyclical input-output process, and the parallel processing - you get a powerful and flexible architecture that excels at tasks like language modelling, translation, summarisation, and more. Just like a group of people engaged in a thoughtful discussion, the transformer is able to dynamically focus on the most relevant information, build upon previous contributions, and produce coherent and contextually appropriate outputs. --- ### Conducting a data assessment - URL: https://agathon.ai/insights/conducting-a-data-assessment - Published: 2023-07-02 - Categories: AI Consulting, AI Advisory Before embarking on any AI project, we conduct a thorough data assessment. These are crucial to ensure the availability, quality, and suitability of the data. In this article, we explore the key steps involved in our data assessments and their significance in delivering a successful AI solution. #### Data governance Data governance sets the foundation for data management throughout the AI project lifecycle. From the outset, it is essential to establish robust data governance mechanisms, encompassing data ownership, privacy, security, compliance, and ethical considerations. This ensures that data is handled responsibly, respects legal requirements, and aligns with ethical principles, mitigating potential risks and liabilities associated with data usage. #### Data availability It is crucial to assess whether the relevant data needed for training, testing, and evaluation will be available. Conduct a comprehensive inventory of the data sources and evaluate their accessibility, quality, quantity, and completeness. This assessment allows you to gauge the feasibility of the project, identify potential data gaps, and proactively plan for data acquisition or augmentation, if necessary. #### Identify flaws and bias Data may contain inherent flaws, biases, or inaccuracies, which can impact the performance and fairness of AI systems. It is imperative to identify and address these issues before initiating the project. Perform a thorough data analysis, employing statistical techniques, data visualisation, and domain expertise to uncover biases, data drift, or anomalies. Take remedial measures such as data cleaning, feature engineering, or augmentation to rectify the flaws or mitigate biases, ensuring the data is fit for purpose. #### Data acquisition strategy In certain cases, you may require additional data to enhance the quality or diversity of your dataset. Define a clear strategy for data acquisition, considering options such as data partnerships, collaborations, or purchasing from third-party providers. Ensure that data acquisition adheres to legal and ethical considerations, including consent, privacy, and data protection regulations. If sharing data is necessary, establish mechanisms to anonymise or aggregate sensitive information, protecting the privacy of individuals or organisations involved. #### Data security and confidentiality Throughout the data assessment process, we prioritise data security and confidentiality. We implement robust data protection measures, including encryption, access controls, and data anonymisation techniques. We consider the implications of storing and transmitting data, and ensure compliance with relevant data protection regulations, industry standards, and client-specific requirements. By addressing data security from the outset, we establish trust safeguard sensitive information. #### Agathon data assessment Conducting a detailed data assessment is a fundamental step we take for any project, ensuring a solid foundation for success. By implementing strong data governance mechanisms, assessing data availability, addressing flaws and potential biases, and defining data acquisition and sharing strategies, we lay the groundwork for robust, reliable, and ethically sound AI systems. By recognising the significance of data and its impact on AI outcomes, we help our clients drive innovation, deliver superior solutions, and unlock the true potential of artificial intelligence. --- ### Delivery excellence through multi-disciplinarity and diverse teams - URL: https://agathon.ai/insights/delivery-excellence-through-multi-disciplinarity-and-diverse-teams - Published: 2023-07-01 - Categories: Responsible AI, AI Strategy, AI Consulting The successful development, evaluation, and delivery of AI projects requires more than just technical expertise. To achieve delivery excellence and especially to mitigate biases in AI systems, it is crucial to employ a multidisciplinary team that encompasses a wide range of skills and perspectives. In this article, we explore the importance of diverse teams and the interdependent disciplines that contribute to the success of AI projects, highlighting how we at Agathon leverage such teams to deliver exceptional outcomes. #### Diversity within our team Diversity within any AI team is essential to address biases that can inadvertently manifest in AI systems. By including team members from different backgrounds, cultures, genders, and ethnicities, we aim to mitigate bias in data collection, algorithm design, and decision-making processes. Diverse perspectives foster more comprehensive and fair AI solutions that benefit all users and avoid perpetuating harmful biases. #### Domain expertise Domain expertise plays a pivotal role in AI projects, enabling a deep understanding of specific industries or sectors such as healthcare, transportation, finance, or manufacturing. Experts in these domains possess invaluable knowledge of real-world challenges, regulatory considerations, and unique requirements, ensuring that AI solutions are aligned with the needs and goals of the target industry. #### Research excellence Our research scientists and model development experts contribute to the core AI capabilities of our projects and embrace a range of research outlets to guide their work. AI is inherently interdisciplinary, drawing insights from fields such as computer science, mathematics, neuroscience, psychology, and more. In reading a wide range of academic papers, our research scientists tap into a broad range of disciplines enabling them to deliver quality models which are designed for our customers’ outcomes. #### Data ethics As AI systems become increasingly pervasive, addressing ethical considerations and mitigating biases is crucial. Our data ethicist consultants ensure that AI projects adhere to ethical guidelines, promoting fairness, transparency, and accountability. They help identify and mitigate biases in training data, develop robust governance frameworks, and ensure that any AI systems we are involved in are designed to respect privacy, security, and societal norms. #### Data visualisation Effective communication of AI insights is essential for decision-making and understanding the impact of AI systems. Visualisation and information design experts help transform complex AI outputs into intuitive and actionable visual representations. They enable stakeholders to comprehend and interpret AI-generated information, promoting effective collaboration and driving informed decision-making. #### Delivery excellence through diversity Delivery excellence in AI projects necessitates the integration of diverse skills and expertise across multiple disciplines. We at Agathon focus on multidisciplinarity to foster a collaborative environment where domain expertise, commercial acumen, systems engineering, model development, data ethics, and visualisation converge to deliver exceptional outcomes. By recognizing the importance of diversity and the interdependent nature of AI technologies, we constantly strive to build robust, unbiased AI systems that drive innovation, societal benefit, and lasting impact. --- ### Quantifying the benefits and risks of an AI deployment - URL: https://agathon.ai/insights/quantifying-the-benefits-and-risks-of-an-ai-deployment - Published: 2023-06-27 - Categories: AI Strategy, AI Consulting, Generative AI #### Why most AI business cases are built on vibes The average organisation scraps 46% of AI proof-of-concepts before they reach production, according to S&P Global Market Intelligence's 2025 survey of over 1,000 enterprises. RAND Corporation's analysis puts the broader failure rate above 80%, twice the failure rate of non-AI technology projects. Yet the business cases that green-lit these projects all showed positive ROI on a spreadsheet somewhere. The disconnect is structural. Traditional return on investment calculations fail to capture the dual nature of AI implementations, which simultaneously reduce certain operational risks while introducing novel exposures related to algorithmic malfunction, adversarial attacks, and regulatory liability. Huwyler's 2025 quantitative framework for measuring AI ROI demonstrates that investment decisions routinely rely on optimistic benefit projections without accounting for the probabilistic costs of AI-specific threats including model drift, bias-related litigation, and compliance failures under emerging regulations such as the EU AI Act and ISO/IEC 42001. Most AI business cases start with a technology demonstration, work backwards to a cost-saving narrative, and present that narrative as though it were financial analysis. PwC's 28th Annual Global CEO Survey found that while 56% of CEOs report generative AI has created efficiencies in how employees use their time, only about a third reported increased revenue (32%) or profitability (34%). The gap between "efficiency" and "profit" tells you everything about how these projects are being measured. The organisations that produce reliable returns from AI treat measurement as an engineering discipline, not a post-hoc justification exercise. #### The measurement problem nobody wants to talk about ##### Defining what "success" means before you write a single line of code The single most common reason organisations cannot evaluate AI ROI after the fact is that nobody measured the baseline before deployment. This sounds obvious. It remains the norm. An IBM CEO study found that only around 25% of AI initiatives deliver expected ROI, and just 16% have scaled enterprise-wide. One contributing factor: success gets defined in terms of model performance (accuracy, latency, throughput) rather than business outcomes (cost per transaction, revenue per customer segment, time to market). A 98% accurate model that answers the wrong question delivers 0% business ROI. Defining success requires a KPI architecture that spans operational, financial, and strategic tiers before development begins. Operational KPIs cover process cycle time, error rates, throughput, and headcount productivity. Financial KPIs address cost per transaction, revenue per customer segment, and support cost per ticket. Strategic KPIs track customer satisfaction, competitive win rate, and talent retention in AI-impacted roles. All metrics need to be measured using a standardised methodology agreed upon by all stakeholders before launch. Adding measurement frameworks after a successful programme launch creates attribution challenges that are expensive to unwind and often impossible to resolve. ##### Separating vanity metrics from value metrics Google Cloud's AI measurement research draws a sharp distinction between adoption metrics and business value metrics. Active users, session length, and thumbs-up feedback tell you whether people are interacting with the system. They tell you nothing about whether the system is generating value. The key measurement discipline for generative AI, as identified by IBM's research, is distinguishing between "time saved" and "value created." Time saved converts into financial ROI only if it leads to a reduction in headcount, shifts the workforce to higher-value tasks, or speeds up time-to-market in ways that drive measurable revenue growth. A team that saves four hours per week but fills that time with low-value work has generated a metric, not a return. The vanity metric problem intensifies with agentic AI systems. As Google Cloud's 2026 framework for measuring agentic AI notes, the evaluation metrics used for large language models (perplexity, BLEU scores, simple thumbs-up/down feedback) do not suffice for assessing autonomous agents. An agent that handles 10,000 tasks per month tells you nothing useful unless you can measure how many it got right, through what reasoning path, and at what cost per successful outcome. #### Quantifying the upside: beyond "efficiency gains" ##### Direct cost displacement vs. capability creation Most AI business cases fixate on cost displacement: replace expensive human labour with cheaper digital alternatives. EY's analysis of this pattern is blunt. Finance teams calculate ROI based on headcount reduction. Operations leaders measure success by eliminated positions. This thinking treats AI as a more efficient version of existing resources rather than recognising it as a fundamentally different capability. The cost-displacement model caps AI's potential at the current size and scope of human-performed tasks. A more productive framing separates direct cost displacement (doing existing work cheaper) from capability creation (doing work that was previously impossible). Lumen Technologies identified that their sales teams spent four hours per week researching customer backgrounds for outreach calls. They quantified this as a $50 million annual opportunity and built AI integrations that compressed research time to 15 minutes. The critical detail: they started with the business pain, not the technology demonstration. Air India's AI virtual assistant handles 97% of over four million customer queries with full automation. This started as a capacity constraint problem (their contact centre could not scale with passenger growth), not an efficiency optimisation. The system created capacity that hiring alone could not have delivered at the same cost structure. ##### Revenue acceleration and time-to-market compression Once deployed, AI systems can handle exponentially increasing workloads without proportional cost increases. EY describes this as "increasing returns to scale," where the more you grow, the lower your per-unit costs become. Traditional businesses face capacity constraints that require proportional investment as they expand. AI-enabled businesses can scale operations, customer base, and market reach while maintaining or even reducing their cost basis. PwC's research found that companies effectively using next-generation cloud architectures and AI capabilities are measurably more likely than their peers to improve profitability, productivity, and time to market. Microsoft reported that their sales team using AI tools achieved 9.4% higher revenue per seller and closed 20% more deals. The design was deliberate: AI suggests draft responses and summarises meetings, while sales representatives retain control over final customer communications. ##### Second-order benefits that compound over quarters The compounding effects of growth-oriented AI adoption create advantages that become increasingly difficult for competitors to match. Expanded market reach generates more data, which improves AI capabilities, which enables further expansion. EY's analysis of their own internal transformation validates this compound effect: early investments in data consolidation and AI talent created the foundation for later innovations, with each capability built upon previous investments, creating exponential rather than linear returns. Google Cloud's research identifies a similar compounding dynamic with business operational KPIs. In retail, a more engaging AI-powered search experience increases visit volume, which generates more behavioural data, which improves personalisation, which drives higher revenue per visit. These second-order effects rarely appear in initial business cases because they emerge over quarters, not weeks. Organisations that measure only first-order effects will systematically undervalue their AI investments and underinvest relative to competitors who measure the full value chain. #### Quantifying the downside: where AI deployments actually fail ##### Technical debt accumulation and maintenance burden IBM's research shows that paying down technical debt from legacy systems can improve AI ROI by up to 29% because it reduces friction and rework. The inverse is also true: deploying AI on top of unresolved technical debt accelerates its accumulation. AI systems are living systems, not one-time releases. While traditional software degrades gradually, AI models undergo many more iterations. Reaching a certain level of accuracy does not immediately translate to business value without redesigning the workflow to leverage the intelligence layer and driving adoption. When organisations underestimate this compounding investment curve, ROI timelines extend beyond projections. The hidden cost structure is instructive. The model itself is often one of the smaller expenses. Data infrastructure, integration, monitoring, retraining, and governance represent the major ongoing investments. Most AI business cases underestimate total cost by 40-60% because they exclude categories that only become visible post-deployment: MLOps infrastructure, drift monitoring systems, human oversight requirements, and compliance tooling. ##### Data quality degradation loops Vela et al.'s 2022 study on temporal quality degradation in AI models, published in Scientific Reports, presents findings that should concern any organisation deploying AI systems without continuous monitoring. The researchers tested four standard machine learning models across 32 datasets from healthcare, transportation, finance, and weather, and observed temporal model degradation in 91% of cases. The degradation patterns they identified are more alarming than simple accuracy decline. Some models performed reasonably well on average, but the variability of their error values grew significantly over time, creating an illusion of accurate performance while actual outcomes became less certain. Other models exhibited "explosive" degradation, maintaining good performance for extended periods before abrupt failure, with no warning from the underlying data. The researchers also discovered "strange attractor" behaviour, where model errors clustered into discrete basins, erratically switching between them over time. Most critically, the researchers demonstrated that data drifts alone cannot explain or predict model failures. Temporal degradation of AI models represents a separate phenomenon, not solely driven by drifts in the data, and not necessarily predictable based on those drifts. This means that monitoring data distributions, while necessary, is insufficient as a quality control mechanism. ##### Organisational friction and adoption resistance Research on AI adoption in HR by Priyanghaa (2025) found that 70% of respondents cited fear of job displacement as a significant barrier, while 65% reported lack of trust in AI systems. The correlation analysis revealed a negative relationship between resistance and organisational readiness (r = -0.60), meaning that as resistance increases, effective adoption decreases proportionally. The same research found that change management practices showed a strong positive correlation with readiness (r = 0.75). Clear communication, employee involvement, continuous training programmes, and feedback mechanisms all measurably reduced resistance. The organisations that treated adoption as a human systems challenge, not a technology rollout, achieved substantially better outcomes. Google Cloud's research on agentic AI adoption identified a specific failure mode they call the "bystander effect." When an AI agent fully owned a task, teams experienced uncertainty about who should verify the work, leading to longer cycle times despite the automation. When a human owned the task and the AI assisted, verification was fast because the human felt responsible. The positioning of AI as collaborator rather than replacement produced measurably better adoption outcomes. ##### Regulatory and compliance exposure The EU AI Act, which entered into force in 2024 with full enforcement of high-risk system obligations from 2026, introduces concrete compliance costs that most business cases ignore entirely. Huwyler's framework for risk-adjusted AI ROI explicitly integrates compliance failures under emerging regulations as a probabilistic cost that must be modelled. The Act mandates risk-based classification of AI systems, transparency obligations, and governance requirements including documentation, monitoring, and human oversight for high-risk systems. For organisations in financial services, healthcare, and critical infrastructure, these are not optional enhancements. They are legal requirements with financial penalties for non-compliance. EY Luxembourg's analysis notes that the certification process (through frameworks like Europrivacy, which is the first scheme officially recognised under GDPR and designed to extend to the EU AI Act) covers data minimisation, security measures, accountability frameworks, and risk management for AI systems. These compliance activities carry real costs that belong in the total cost of ownership, not as a surprise line item six months after deployment. #### Building a risk-adjusted ROI framework ##### Assigning probabilities to failure modes Huwyler's quantitative framework draws on established risk quantification methods, including annual loss expectancy calculations and Monte Carlo simulation techniques, to compute net benefits that incorporate both productivity gains and the delta between pre-implementation and post-implementation risk exposures. The practical application requires probability-weighted scenario analysis. A representative model might assign a 25% probability to full deployment success with a $4.2 million return, 40% to partial deployment achieving 60% of target impact, 25% to limited adoption at 30% impact, and 10% to programme failure. The risk-adjusted expected value across these scenarios will be substantially lower than the "base case" that appears in most business cases, which implicitly assumes 100% probability of the best-case outcome. This approach forces honest conversations about adoption probability, which is typically treated as an assumption rather than an estimate. If the risk-adjusted estimate is negative or marginal, the project must be reworked before approval. The discipline of assigning probabilities to failure modes changes the quality of investment decisions more than any improvement in model architecture. ##### Modelling scenarios rather than point estimates The S&P Global survey found that companies cited cost overruns, data privacy concerns, and security risks as primary obstacles to AI success. These risks are not binary (they happen or they don't). They occur on a spectrum, with varying probability and varying financial impact. A robust scenario model requires three things: a conservative case that assumes partial adoption, extended timelines, and the emergence of at least one major unplanned cost category; a base case that assumes planned adoption rates and costs within 20% of estimates; and an optimistic case that includes second-order compounding effects and successful scaling beyond the initial use case. The MIT 2025 AI Report found that 95% of generative AI pilots fail to deliver tangible profit-and-loss results. This statistic alone should inform the probability weightings in any scenario model. The 5% that succeed treat AI as an integrated workflow rather than a static project, which means success depends on organisational factors that are harder to estimate but more consequential than technical performance. ##### Accounting for opportunity cost of doing nothing PwC's analysis of competition in the age of AI argues that the speed at which competitive capabilities change is accelerating at exponential rates, and the next few years of disruption will likely produce winners that persist for decades. This creates a measurable opportunity cost of inaction. The cost-of-inaction analysis should be part of every AI investment evaluation. If a competitor deploys AI to compress their sales research from four hours to fifteen minutes (as Lumen Technologies did), the competitive disadvantage to non-adopters compounds over time. Each quarter of delay is a quarter in which competitors are building data advantages, refining their models, and deepening their customer relationships through AI-augmented workflows. Jensen Huang argued at the Cisco AI Summit in February 2026 that forcing engineers to justify AI work with hard ROI up front is counterproductive in a period of rapid technological change. The counterpoint for CFOs: the opportunity cost of doing nothing should be formally quantified and included in the decision framework, even if it requires assumptions. An imprecise estimate of competitive risk is more useful than pretending the risk does not exist. #### The hidden costs that sink AI projects ##### Integration complexity with legacy systems The total cost of AI ownership (TCAO) extends far beyond model development. It includes data pipeline construction, API integration, security and compliance infrastructure, cloud computing and GPU resources, data preparation and labelling, and the often-overlooked expense of adapting existing workflows to accommodate AI outputs. IBM's research confirms that many organisations are not where they need to be in their digital transformation journey to realise the full benefit of AI integration. Technical debt remains a primary friction source, and deploying AI systems on top of unresolved integration challenges creates compound complexity. Each integration point becomes a potential failure mode, a maintenance burden, and a constraint on future flexibility. The Informatica CDO Insights 2025 survey identifies the top obstacles to AI success as data quality and readiness (43%), lack of technical maturity (43%), and shortage of skills (35%). Winning programmes invert typical spending ratios, earmarking 50-70% of the timeline and budget for data readiness, including extraction, normalisation, governance metadata, quality dashboards, and retention controls. ##### Ongoing model monitoring and retraining The temporal degradation research by Vela et al. demonstrates that model quality cannot be assumed to persist. Some models exhibited stable performance for over a year before sudden, catastrophic degradation, with no detectable signal from the data itself. The researchers also identified evolving bias patterns, where feature importance values shifted over time, meaning that a model validated for fairness at deployment could develop discriminatory patterns months later without any change in the underlying data distribution. Google Cloud's framework for production AI systems specifies concrete monitoring requirements: model drift detection with alert thresholds, automated retraining cadences, and structured post-deployment audits at 30, 90, and 180-day intervals. The metrics include adoption rate versus target, KPI movement versus baseline, unplanned cost discovery, and optimisation opportunity identification. These monitoring costs are ongoing and non-trivial. A model that works today and fails silently in six months is worse than a model that never worked, because it has been integrated into decision workflows and its outputs are being acted upon without scrutiny. ##### Talent acquisition and retention premiums The White House Council of Economic Advisers' 2025 AI Talent Report documents a structural gap between AI talent supply and demand. Between 2015 and 2022, job listings requiring AI skills increased 257%, while overall job listings grew only 52%. AI salaries increased between 10 and 13% in a single year (2021-2022), and AI labs spend 29-49% of their total costs on labour. Growth in the supply of AI talent measurably lags growth in demand. The number of AI software-related job postings grew at an average annual rate of 31.7% from 2015 to 2022, while bachelor's degrees in relevant fields grew at only 8.2% annually. At the doctoral level, the gap is even wider: 2.9% annual growth in graduates versus demand growing at multiples of that rate. Non-US citizens make up nearly half of AI-relevant PhD graduates from US institutions, and similar dynamics apply globally. For any organisation building AI capabilities, talent costs will remain elevated, and the risk of losing key personnel to competitors with deeper pockets is a financial exposure that belongs in the ROI model. #### Measuring what matters in production ##### Leading indicators vs. lagging indicators Google Cloud's three-pillar framework for agentic AI measurement distinguishes between reliability metrics (can the agent handle complex workflows consistently?), adoption metrics (are people using it?), and business value metrics (is it generating net new value?). The sequence matters: reliability must be established before adoption can be measured, and adoption must be confirmed before business value can be attributed. Leading indicators for production AI include tool selection accuracy, plan adherence, argument hallucination rate, and cost per successful task. These metrics surface problems before they cascade into business impact. Lagging indicators like revenue uplift, cost savings, and customer satisfaction confirm value but arrive too late to inform corrective action. The practical distinction: if your AI system's plan adherence score drops from 92% to 74% over two weeks, that is a leading indicator of degradation that warrants investigation. If your customer satisfaction score drops three points next quarter, that is a lagging indicator that confirms the damage has already been done. ##### Setting thresholds for intervention and rollback Production AI systems need explicit service-level objectives, not aspirational targets. Google Cloud's research recommends writing concrete SLOs such as "ticket summary accuracy above 85% and latency below five seconds, 95% of the time." When those thresholds are breached, automated alerting triggers investigation and, if necessary, rollback. The intervention framework should specify who acts at each threshold. A minor drift in accuracy might trigger automated retraining. A sustained decline below the SLO triggers human investigation. A catastrophic failure triggers immediate rollback to the previous model version or handoff to human operators. Google's documentation team discovered that "output friction" (how often a human needs to step in and take over a task the agent started) is one of the most informative production metrics. High intervention rates signal trust issues and suggest the agent may work better in a reactive mode, where it assists humans, rather than a proactive mode where it operates autonomously. ##### Attribution: isolating AI impact from other variables Isolating AI's contribution from other simultaneous changes is one of the hardest measurement problems in production. Google Cloud's research notes that when you make changes to AI systems, improving one KPI can sometimes impact another. For retailers, cart size may increase with a more engaging chatbot, but time-to-cart (a previously important metric to keep low) may increase as well. The WorkOS analysis of enterprise AI patterns confirms that organisations reporting significant financial returns are twice as likely to have redesigned end-to-end workflows before selecting modelling techniques. This makes attribution cleaner: if the workflow was redesigned for AI and the business metric improved, the causal chain is shorter and more defensible. Context and industry expertise remain critical when interpreting changes in operational metrics. A contact deflection rate improvement of 15% is a clear AI attribution when the only change was deploying an AI agent. The same improvement during a quarter when you also redesigned your support portal, changed your SLA targets, and restructured your support team is attributable to nothing in particular. #### A practical scoring model for go/no-go decisions ##### Weighted criteria that reflect your organisation's risk appetite A structured scoring model should evaluate AI investments across multiple dimensions with weights that reflect organisational priorities. Strategic alignment (does this initiative address a named corporate priority?) should carry heavy weight because misaligned AI projects consume 30-40% more budget than aligned ones, according to Boston Consulting Group analysis cited in the CMARIX framework. The scoring dimensions should include: strategic alignment and executive sponsorship, baseline measurement readiness, total cost of ownership completeness, risk-adjusted value assessment, data quality and infrastructure maturity, change management planning, regulatory compliance requirements, and talent availability. Each dimension receives a score and a weight. The weights differ by organisation: a regulated financial institution will weight compliance exposure higher than a consumer technology company. A company with strong existing data infrastructure will weight integration complexity lower. The critical discipline is requiring a minimum score before approval, and treating a low score as a signal to rework the proposal rather than approve a weak business case. Organisations that approve every AI initiative above a low threshold end up with the 46% abandonment rate that S&P Global documented. ##### Time horizons that match realistic deployment timelines IBM's research on AI adoption is direct: only about 25% of AI initiatives deliver expected ROI, and CEOs are balancing pressure for short-term ROI with longer-term innovation goals. The scoring model must accommodate this tension by specifying different time horizons for different types of AI initiatives. A customer service AI agent that deflects routine enquiries should demonstrate ROI within 8-14 months, with year-two and year-three returns exceeding 300% as the model learns from historical data and increases deflection rates. A capability-creation initiative (entering new markets, building new product categories through AI) may require 18-36 months before meaningful returns materialise. The payback analysis should include conservative, base, and optimistic cases, with the investment approval tied to the conservative case being acceptable, not the optimistic case being attractive. Budget reserves of at least 25% against the base total cost of ownership estimate should be standard practice, covering the unplanned cost categories that surface in virtually every AI deployment. #### Making the business case that survives scrutiny The business cases that survive board-level scrutiny share common characteristics. They start with quantified business pain, not technology demonstrations. They include complete cost models that account for integration, monitoring, retraining, compliance, and talent costs. They present risk-adjusted scenarios rather than point estimates. They specify measurable baselines and success criteria before development begins. They budget for change management as a first-class programme element, not an afterthought. The MIT 2025 finding that 95% of generative AI pilots fail to deliver P&L results is not evidence that AI does not work. It is evidence that measurement, governance, and organisational readiness determine outcomes more than model sophistication does. The 5% that succeed follow a recognisable pattern: they quantify the problem before proposing a solution, they model the full cost structure, they invest disproportionately in data readiness and change management, and they measure production performance continuously against explicit thresholds. Huwyler's framework for risk-adjusted AI ROI captures the underlying principle: accurate AI investment evaluation requires explicit modelling of control effectiveness, reserve requirements for algorithmic failures, and the ongoing operational costs of maintaining model performance. Organisations that build this discipline into their investment process will make fewer AI bets, but the bets they make will produce measurable returns. The gap between AI's technical capability and its delivered business value is a measurement and governance problem, not a technology problem. Organisations that close this gap treat AI investment with the same rigour they apply to any capital allocation decision: quantified baselines, scenario-modelled returns, risk-adjusted expectations, and continuous production monitoring. Those that do not will continue to fund impressive demonstrations that deliver impressive write-offs. If you are building the business case for a sophisticated AI deployment and want the financial framework to match the technical ambition, get in touch. We help technical leaders build AI systems that deliver returns robust enough to survive scrutiny, not just approval. #### References - The Risk-Adjusted Intelligence Dividend: A Quantitative Framework for Measuring AI ROI (arxiv.org) - How to maximize AI ROI — IBM Think - AI ROI in 2026: A CFO Framework to Measure AI Investment — CMARIX - Beyond cost cutting: AI as the ultimate growth engine — EY - Competing in the age of AI: Speed, scale, innovation — PwC - KPIs for gen AI: Measuring AI success — Google Cloud - The KPIs that actually matter for production AI agents — Google Cloud - Why most enterprise AI projects fail — WorkOS - Temporal quality degradation in AI models — Nature / Scientific Reports - EU AI Act and data privacy certification — EY Luxembourg - AI Talent Report — White House Council of Economic Advisers - AI Adoption in HR: Resistance, Readiness, and Change Management — Journal of Marketing & Social Research --- ### Adversarial models: what are they and when should you use them? - URL: https://agathon.ai/insights/adversarial-models-what-are-they-and-when-should-you-use-them - Published: 2023-06-19 - Categories: Machine Learning, LLMs Imagine a world where a self-driving car effortlessly navigates the bustling city streets, its artificial intelligence (AI) systems diligently analyzing the surroundings to ensure a safe journey. What if that very same AI could be easily tricked, leading the car down a treacherous path? Welcome to the realm of adversarial models—a daring exploration into the vulnerabilities and fragility of AI. Picture an AI-powered cybersecurity system that, despite its seemingly impenetrable defenses, falls victim to a meticulously crafted attack. Or envision a financial fraud detection system that falters in the face of cunning adversaries. It is here, in these captivating scenarios, that adversarial models emerge as a double-edged sword—an instrument of chaos and a catalyst for innovation. #### What is an adversarial model? An adversarial model in machine learning is a type of model that is designed to improve the robustness and performance of the target model by actively attempting to deceive or challenge it. It involves training two models simultaneously: the target model and the adversarial model. The target model is the model that we want to improve or make more resilient to potential attacks or vulnerabilities. The adversarial model, on the other hand, is trained to generate adversarial examples or perturbations that are specifically crafted to deceive the target model. The process typically involves the following steps: 1. Training the target model: The target model is trained using standard machine learning techniques on a labeled dataset to perform a specific task, such as image classification or natural language processing. 1. Training the adversarial model: The adversarial model is trained to generate perturbations or examples that can potentially fool the target model. This is done by using optimization techniques to find the perturbations that maximize the target model's prediction errors or misclassify the input. 1. Adversarial example generation: The adversarial model generates adversarial examples by applying carefully crafted perturbations to the input data. These perturbations are often imperceptible to humans but can lead to significant changes in the target model's predictions. 1. Adversarial training: The target model is then retrained on a combined dataset consisting of the original data and the adversarial examples. This training process helps the target model learn to be more robust and accurate in the presence of adversarial attacks. #### How are they useful? The usefulness of adversarial models lies in their ability to expose vulnerabilities and improve the overall security and reliability of machine learning systems. By actively challenging the target model with adversarial examples, these models can help identify weaknesses, bias, or flaws in the system. Adversarial training allows the target model to learn from these examples and become more resilient, making it harder for malicious actors to manipulate or exploit the system. As alluded to above, adversarial models have applications in various domains, including computer vision, natural language processing, and cybersecurity. They can enhance the robustness of image classifiers, improve the security of biometric systems, and aid in the detection of malicious activities or attacks. Overall, adversarial models can play a crucial role in strengthening machine learning systems, increasing their resistance to potential threats, and advancing the field's understanding of vulnerabilities and defenses in artificial intelligence. #### Industry use cases Adversarial models in machine learning can be useful in various industries and use cases where the robustness and security of machine learning systems are critical. From our two examples above adversarial models can help in detecting and mitigating cyber threats. By generating adversarial examples, these models can identify vulnerabilities in intrusion detection systems, malware classifiers, or network traffic analysis systems, making them more resilient against adversarial attacks. In the case of autonomous vehicales, adversarial models can be employed to enhance the safety and reliability of autonomous vehicles. By generating adversarial examples, potential vulnerabilities in object recognition or sensor fusion systems can be identified and mitigated, reducing the risk of misclassification or manipulation. --- ### Responsible and ethical AI — why does it matter? - URL: https://agathon.ai/insights/responsible-and-ethical-ai-why-does-it-matter - Published: 2023-06-01 - Categories: Responsible AI, AI Strategy, AI Advisory After architecting AI systems with advanced capabilities like responsive context management and self-improving workflows, one thing becomes clear: most conversations about "responsible AI" focus on theoretical frameworks whilst ignoring the technical decisions that actually determine whether AI systems are beneficial or harmful. The real ethical challenge isn't following compliance checklists—it's building AI products that exploit technical potential responsibly whilst avoiding the shallow implementations that create genuine risks. #### The superficial ethics problem Most AI ethics discussions centre on governance frameworks and policy guidelines. Meanwhile, the actual ethical risks emerge from poor technical implementation: Real ethical risks in AI development: - Shallow implementations that make confident predictions without understanding context - Basic AI tools that automate decisions without preserving human agency - Systems that claim "intelligence" whilst operating with minimal technical sophistication - Products that exploit user psychology rather than enhancing human capabilities These problems aren't solved by ethics committees—they're solved by technical excellence and sophisticated implementation. #### Technical sophistication as ethical foundation When building AI systems that maximise technical potential, ethical considerations become embedded in architectural decisions: Responsible technical approaches: - Responsive context management: Systems that understand nuance and context rather than making oversimplified predictions - Transparent decision processes: Architectures that enable genuine explainability, not post-hoc rationalisation - Human-centric workflows: AI that augments human decision-making rather than replacing human judgment - Adaptive learning systems: Products that improve based on user feedback and changing requirements The technical ethics questions: 1. Does your AI system exploit advanced capabilities to enhance human decision-making? 1. Can users understand and influence the system's reasoning process? 1. Does the technical architecture preserve human agency and choice? 1. Are you building genuinely intelligent systems or sophisticated automation? #### Beyond compliance: strategic ethical leadership Responsible AI isn't about checking regulatory boxes—it's about technical leadership that creates competitive advantage through superior implementation: Strategic advantages of responsible technical approaches: - User trust through transparent, explainable systems - Regulatory resilience through proactive technical design - Market differentiation via sophisticated rather than superficial AI - Long-term sustainability through adaptive, learning systems The business case for technical ethics: Companies building sophisticated AI products with responsible technical architectures outperform those implementing basic AI tools regardless of their ethics committees. Technical excellence and ethical implementation are inseparable. #### Due diligence for responsible AI When evaluating AI investments or development approaches, the critical questions are technical: Technical Assessment Criteria: 1. Capability exploitation: Does the system maximise available AI potential responsibly? 1. Architectural transparency: Can the technical approach support genuine explainability? 1. Human agency preservation: Does the design enhance rather than replace human decision-making? 1. Adaptive sophistication: Can the system learn and improve whilst maintaining ethical constraints? Investment red flags: - AI systems that automate human judgment without preserving human oversight - "Black box" implementations that can't explain their decision processes - Basic tools marketed as sophisticated AI without technical depth - Systems designed to exploit user psychology rather than enhance capabilities #### The implementation reality Building responsible AI requires technical expertise to distinguish between genuine innovation and superficial implementations. Most organisations discussing AI ethics lack the technical depth to evaluate whether their AI systems are actually responsible or merely compliant. Critical technical decisions: - Choosing architectures that enable rather than obscure transparency - Implementing learning systems that adapt without compromising ethical constraints - Building user interfaces that preserve human agency whilst leveraging AI capabilities - Designing evaluation frameworks that measure genuine rather than apparent performance #### Strategic recommendation Responsible AI isn't achieved through compliance frameworks—it's built through sophisticated technical implementation that maximises AI potential whilst preserving human values and agency. Focus areas for responsible AI leadership: - Technical due diligence: Evaluate whether AI systems genuinely exploit available capabilities responsibly - Sophisticated implementation: Build products that demonstrate advanced AI whilst enhancing human decision-making - Strategic architecture: Design systems that create competitive advantage through responsible technical excellence - Genuine innovation: Distinguish between breakthrough AI products and basic automation with ethical policies #### The competitive advantage Organisations that combine technical sophistication with responsible implementation create sustainable competitive advantages. They build AI products that users trust, regulators respect, and competitors struggle to replicate. The future belongs to companies that don't just talk about responsible AI—they build it through superior technical implementation that maximises AI potential whilst preserving human agency and understanding. #### Getting beyond theoretical ethics Responsible AI requires technical leadership that can evaluate, design, and implement sophisticated AI systems with embedded ethical considerations. This isn't about following guidelines—it's about technical expertise applied to create genuinely beneficial AI products. The most responsible approach to AI development is building products that fully exploit technical potential whilst enhancing rather than replacing human capabilities. Everything else is just policy theatre. --- Need technical due diligence for your responsible AI implementation? Agathon provides expert evaluation of AI systems that distinguishes between genuine innovation and superficial compliance, helping organisations build sophisticated AI products that create competitive advantage through responsible technical excellence. --- ### Recent trends in NLP - URL: https://agathon.ai/insights/recent-trends-in-nlp - Published: 2023-02-23 - Categories: Machine Learning, Generative AI, AI Advisory In recent years, there have been several notable trends in Natural Language Processing (NLP) and Artificial Intelligence (AI) that have significantly impacted the field. Here's a short summary of some of the key trends: 1. Transformer-based Models: Transformer models, such as OpenAI's GPT (Generative Pre-trained Transformer) series, have revolutionised NLP. These models leverage self-attention mechanisms and pre-training on massive datasets, enabling them to generate high-quality text and perform a wide range of language-based tasks. 1. Transfer Learning and Pre-training: Pre-training large-scale language models on vast amounts of text data has become a dominant approach. These pre-trained models can then be fine-tuned on specific downstream tasks, allowing for better generalisation and improved performance across various NLP applications. 1. Multimodal AI: The integration of multiple modalities, such as text, images, and audio, has gained significant attention. Researchers have been working on developing models that can understand and generate content using multiple modalities, enabling applications like image captioning, visual question answering, and audio transcription. 1. Ethical and Responsible AI: As AI technologies continue to advance, the importance of ethical considerations and responsible deployment has come to the forefront. There is an increased focus on fairness, transparency, and accountability to ensure that AI systems are unbiased, respect privacy, and are used for the benefit of society as a whole. 1. Low-resource and Multilingual NLP: There has been growing interest in developing NLP models that can effectively handle low-resource languages and multilingual scenarios. Efforts have been made to improve the accessibility of NLP technologies for languages with limited resources and to develop cross-lingual models that can transfer knowledge across different languages. 1. Conversational AI and Chatbots: Conversational AI has witnessed significant progress, with the development of chatbots and virtual assistants capable of engaging in human-like conversations. Advancements in language generation and understanding have led to more sophisticated chatbot systems that provide personalized and context-aware responses. 1. Reinforcement Learning for NLP: Reinforcement Learning (RL) techniques have been applied to NLP tasks, enabling models to learn from interaction and feedback. RL has been particularly successful in areas like dialogue systems, machine translation, and text summarization, where models can be trained to optimize performance based on rewards or evaluations. ## Contact - Website: https://agathon.ai - Contact page: https://agathon.ai/contact - Location: United Kingdom --- This content is available for LLM consumption. For human visitors, please visit https://agathon.ai