Key takeaways
- An AI Centre of Excellence should be a small, accountable operating capability that sets guardrails, enables delivery teams, and measures outcomes—not a permanent approval bottleneck or a standalone innovation lab.
- For industrial enterprises, the charter must distinguish read-only decision support from workflows that create records or execute actions, while preserving ISA-95 boundaries between enterprise systems, manufacturing operations, and control.
- A defensible setup combines an executive-sponsored portfolio process with an AI delivery lifecycle that maps use cases, measures performance and risk, manages incidents, and retains evidence.
- The first 90 days should produce a charter, a prioritized pilot portfolio, a reference architecture, baseline measures, minimum controls, reusable delivery templates, and an explicit path from centralized enablement to federated adoption.
The objective is scaled value with controlled risk—not an “AI team”
Industrial enterprises are not short of AI ideas. They are short of a repeatable way to decide which ideas deserve scarce engineering, data, cybersecurity, and change-management capacity; how to make them safe enough to deploy; and how to prove that they improved an operating outcome.
That gap is visible in current enterprise research. McKinsey’s 2025 survey found that 88% of respondents reported regular AI use in at least one business function, but only about one-third said their companies had begun scaling AI programs across the enterprise. The same survey reported that 23% were scaling an agentic AI system somewhere in the enterprise and 39% reported experimentation. These are survey-reported indicators, not manufacturing benchmarks, but they describe a familiar industrial pattern: broad activity, uneven governance, and limited repeatability beyond the first few use cases. [2]
An AI Centre of Excellence (AI CoE) is the operating capability that closes this gap. It does not own every use case forever. It gives the enterprise one accountable way to select work, establish guardrails, provide reusable platforms and methods, assure delivery quality, and measure benefits. Microsoft defines an AI CoE as an internal team of experts that drives valuable outcomes and helps prevent fragmented or ungoverned AI adoption. [1]
The practical test: If a proposed AI CoE cannot explain who owns the business decision, what evidence supports an AI output, which systems it may read or write, how performance is measured, and how work continues when the AI is unavailable, it is not yet an operating model.
For a manufacturer, EPC contractor, or asset-intensive enterprise, that test matters more than the choice of model provider. An assistant that summarizes maintenance history is different from a workflow that creates a draft work order, and both are very different from a system that can alter an operating parameter. The CoE must make those distinctions explicit before adoption becomes widespread.
Start with a charter that makes the CoE accountable
The CoE charter is a short, executive-approved document—not a presentation deck. It should define the outcomes the CoE is responsible for, the decisions it can make, the authority it does not have, the funding model, and the expected relationship with business and site teams.
A useful charter answers five questions in plain language. Why is the enterprise investing in AI? Which outcomes matter, such as engineering cycle time, quality investigation time, schedule adherence, energy intensity, safety administration effort, or knowledge-retrieval accuracy? Who is accountable for business value and risk? What is the boundary between experimentation, production decision support, and production workflow automation? How will the CoE enable teams rather than becoming a queue of approvals?
| Charter element | Practical decision | Evidence of completion |
|---|---|---|
| Executive mandate | Name an accountable executive sponsor and a cross-functional steering group with a defined decision cadence. | Signed charter, meeting cadence, escalation route, and approved budget or capacity model. |
| Value thesis | Specify the operating outcomes the portfolio may target; do not use “innovation” as the only success criterion. | Outcome hierarchy and baseline measures for each selected use case. |
| Scope and exclusions | State which sites, functions, data classes, AI capabilities, and action types are initially in scope—and which are prohibited. | A published boundary statement and an exception process. |
| Decision rights | Separate portfolio prioritization, architecture review, risk acceptance, business ownership, release approval, and incident response. | A RACI or equivalent decision-rights map. |
| Funding and capacity | Decide whether the CoE funds shared capabilities, pilots, or neither; require business sponsors to own benefits. | Chargeback/showback logic and named business owners. |
| Exit and evolution | Define when a capability moves from CoE stewardship into a product, platform, site, or function team. | Transition criteria, support model, and ownership handover checklist. |
This does not make the CoE bureaucratic. It makes it possible to say “yes” quickly to a bounded, measurable use case and “not yet” to an initiative that lacks an owner, data authority, safe failure mode, or credible way to prove value.
Choose the operating model for the maturity you have—not the org chart you wish you had
Early in the journey, a centralized team usually has an advantage: it consolidates scarce skills, creates common controls, and prevents every function from buying disconnected tooling. Microsoft similarly recommends centralization at the outset, then an evolution toward an advisory model once governance and delivery practices are embedded in platform operations. [1]
Centralization is not a permanent destination. A mature CoE should retain ownership of enterprise-wide policy, platform patterns, evaluation methods, and portfolio visibility while delivery ownership moves closer to the process owner and site. The right model is often federated enablement: a small central hub and designated domain or site leads who use the common standards in real work.
| Model | Best fit | Strength | Failure mode to avoid |
|---|---|---|---|
| Centralized CoE | First 6–12 months; limited AI capability; high regulatory, IP, or OT risk. | Fast standardization and clear accountability. | The CoE becomes the bottleneck for every minor request. |
| Federated hub-and-spoke | Multiple functions or plants have capable product, data, or engineering teams. | Standards remain common while delivery stays near the work. | Local teams duplicate tools or bypass shared controls. |
| Advisory / platform-enabled | Common platform, guardrails, and delivery patterns are established. | Business teams move quickly with policy built into the path to production. | The CoE loses visibility of new use cases and risk exposure. |
The model can vary by capability. For example, model and application evaluation may remain centralized, while maintenance, quality, and engineering teams own their respective backlogs. The design principle is consistent: centralize what must be common; decentralize what must be close to the process.
Build a small, multidisciplinary founding team
An AI CoE is not synonymous with a data-science group. Industrial use cases combine process knowledge, data ownership, information security, plant operations, legal or compliance requirements, product delivery, and adoption. Microsoft’s guidance explicitly calls for business leaders, AI technical experts, data scientists, ML engineers, governance experts, security specialists, and AI operations professionals. [1]
A lean founding team can begin with five accountable roles. One person may cover more than one role at the start, but the responsibilities must be visible.
| Role | Accountable for | Typical industrial counterpart |
|---|---|---|
| Executive sponsor and steering group | Investment decisions, cross-functional conflict resolution, and business accountability. | COO, CIO, CTO, Chief Digital Officer, business-unit or site leaders. |
| CoE lead | Charter execution, portfolio health, dependencies, reporting, and operating-model evolution. | Head of digital, enterprise architect, or transformation leader. |
| Domain product owner | Use-case value, user workflow, acceptance criteria, adoption, and benefit realization. | Maintenance, quality, engineering, production-planning, or EHS leader. |
| Data, platform, and AI engineering lead | Source integration, data contracts, evaluation, deployment architecture, observability, and lifecycle management. | Data engineering, OT/IT architecture, MLOps, application engineering. |
| Risk, security, and assurance lead | Data classification, access model, supplier review, safety and quality boundaries, release controls, monitoring, and incident response. | CISO/OT security, quality, privacy, legal, and operational-risk representatives. |
The team needs regular access to operator, engineer, planner, quality, and maintenance expertise. Those people are not merely “stakeholders.” They decide whether a retrieval result is applicable, whether a recommendation fits the real workflow, and whether the measured benefit is genuine.
Create an intake process that prioritizes evidence, not enthusiasm
A good intake process has two goals: turn a loosely described business problem into a testable use case, and reject unsuitable work early without wasting political capital. Each request should be recorded in a small, consistent brief that names the process owner, users, decision or task, source systems, data authority, potential action, expected value, risks, and baseline.
The first prioritization question should be: What decision, handoff, or repetitive task will change? “Build an industrial copilot” is not a use case. “Help reliability engineers find released maintenance instructions and prior failure evidence for the correct asset before they plan a job” is a use case. It identifies the user, the decision, the evidence, and an observable outcome.
| Prioritization dimension | Questions to answer before a pilot | Red flag |
|---|---|---|
| Business value | Which operating metric can move? What is the baseline? Who owns the benefit? | Savings exist only as a generic percentage or vendor claim. |
| Technical feasibility | Are the needed data, identifiers, interfaces, and environments available? | The team cannot identify an authoritative source or a usable integration path. |
| Risk and control | Could an incorrect output affect safety, quality release, contractual commitments, or sensitive information? | The proposed action has no safe human review or fallback. |
| Adoption readiness | Which role will use it in which workflow, and what changes for them? | No named users or process owner are willing to test it. |
| Reusability | Can a connector, evaluation set, policy, or workflow pattern serve another use case? | The pilot requires a one-off exception to every standard. |
Score these dimensions, but do not pretend that a score creates certainty. A steering group should see the assumptions and decide deliberately. The portfolio should contain a mix of near-term, low-risk use cases that prove the delivery method and a small number of strategically important use cases that build reusable data or workflow capabilities.
Define industrial architecture boundaries before granting access
In industrial AI, architecture is a governance control. ISA-95, also known as IEC 62264, addresses integration between enterprise business systems and manufacturing control systems. It emphasizes the interface between Level 3 manufacturing operations management and Level 4 enterprise systems while maintaining distinct functional boundaries between enterprise systems and Levels 0–3 control and operations. [5]
An AI CoE should use that boundary to distinguish an information path from an action path. The information path may retrieve approved data from enterprise, operations, engineering, quality, and maintenance systems through governed connectors. The action path should enter a target system only through an explicit API, service identity, validation rule, native approval workflow, and audit record. Neither a chat interface nor a fluent answer should create an implied right to act.
| Capability tier | Example | Default CoE control |
|---|---|---|
| Observe and retrieve | Search released procedures, summarize a downtime timeline, classify incoming documents. | Read-only identities, authoritative-source metadata, permissions filtering, citations, logging, and evaluation. |
| Recommend and prepare | Suggest investigation questions; draft a CMMS work order, CAPA response, or engineering-change request. | Clear uncertainty, human review, editable draft, policy validation, and destination-system approval. |
| Execute a bounded transaction | Send an approved notification or create a pre-authorized low-risk ticket. | Allowlisted API, narrow parameters and asset scope, separation of duty, idempotency, monitoring, and rollback or manual fallback. |
| Alter plant control or safety state | Change a process setpoint, alarm threshold, interlock, or safety function. | Outside the default CoE scope. Requires separate engineered control, safety, cybersecurity, and site authorization processes. |
This is not an argument against automation. It is an argument for matching automation to evidence, risk, and controllability. A successful CoE creates a path to production that is fast for low-risk work and deliberately rigorous for high-consequence work.
Operate the AI lifecycle as a managed system
The NIST AI Risk Management Framework is voluntary guidance intended to help organizations incorporate trustworthiness considerations into the design, development, use, and evaluation of AI systems. [3] ISO/IEC 42001:2023 specifies requirements for establishing, implementing, maintaining, and continually improving an artificial intelligence management system. [4] Together, they are useful anchors for a CoE because they turn “responsible AI” from a value statement into repeatable work.
NIST organizes activities around Govern, Map, Measure, and Manage. The CoE can translate those functions into its delivery lifecycle without treating a framework as a substitute for engineering judgment.
| Lifecycle gate | Practical CoE activity | Required record |
|---|---|---|
| Govern | Assign accountable owners; define policies, risk appetite, roles, vendor requirements, and training expectations. | Charter, control catalogue, decision-rights map, supplier and data-use conditions. |
| Map | Describe intended users, operating context, data sources, decisions, affected parties, foreseeable failure modes, and escalation path. | Use-case brief, system context, data map, risk assessment, and acceptance criteria. |
| Measure | Test quality, groundedness, safety controls, access isolation, latency, cost, user usefulness, and relevant task performance against a documented baseline. | Evaluation dataset, test results, threshold decisions, residual-risk statement, and pilot report. |
| Manage | Release in stages, monitor performance and drift, handle incidents, retire or change the system safely, and report outcomes. | Release record, monitoring plan, incident log, change history, and periodic review. |
For generative AI, the evaluation set should include the difficult cases, not only representative demos: ambiguous requests, stale documents, near-duplicate revisions, permission boundaries, missing data, unusual abbreviations, conflicting instructions in source material, and cases where the correct response is “insufficient evidence.” NIST’s Generative AI Profile is particularly useful here because it addresses risks that are unique to or exacerbated by generative AI and suggests actions under the same four functions. [6]
Make reusable assets the CoE’s primary product
A CoE becomes scalable when each successful use case leaves the next team with a better starting point. The durable products are not only deployed applications. They are templates, connectors, policy checks, evaluation harnesses, data contracts, reference architectures, prompt and tool-design patterns, and adoption playbooks.
Examples of reusable industrial assets include a standard connector pattern for historian or CMMS data; a source-authority and revision metadata schema; a permission-aware retrieval pattern; a policy-controlled way to draft—not automatically approve—transactions; a red-team test set for plant and engineering documents; and a benefit-tracking template that links a pilot to a real operating baseline.
This product mindset has an important consequence: the CoE should publish its standards as practical paths, not PDF barriers. A delivery team should be able to discover which data classifications are permitted, which evaluation evidence is required, which API action types are allowed, and who can review an exception. The process must be stronger than shadow AI while still being easier to use than an unofficial workaround.
Measure value, adoption, and control performance together
A pilot that produces a polished demo but no changed workflow is not scale-ready. Equally, a pilot that reports time saved while hiding unreliable outputs or unsafe access patterns is not an enterprise success. The scorecard should contain operational, adoption, technical, and control measures.
| Measure family | Examples | Why it matters |
|---|---|---|
| Operating outcome | Investigation cycle time, engineering search time, schedule-replan time, work-order preparation time, rework, defect escape, energy variance. | Connects the use case to a process outcome rather than model activity. |
| Adoption and workflow | Eligible-user adoption, repeat use, completion rate, override or rejection reasons, training completion. | Reveals whether the workflow is trusted and usable. |
| Quality and reliability | Grounded-answer rate, task accuracy, citation completeness, revision correctness, latency, cost per completed task, availability. | Shows whether the system is fit for the decision it supports. |
| Risk and control | Access-control failures, policy violations, unsafe suggestions, incident response time, data-retention compliance, human-approval rate. | Ensures that efficiency does not conceal a growing control gap. |
The metric definition must come before the launch. For example, “save engineer time” becomes measurable only when the team records the existing time-to-answer, the eligible population, what counts as a completed answer, the review effort required, and any impact on quality or rework. Baselines should be held by the business owner, not invented after the pilot has succeeded.
A practical first 90-day setup plan
The first 90 days should not attempt enterprise-wide transformation. The objective is to create a credible operating baseline, prove the decision process, and deliver one or two bounded use cases with reusable assets.
| Timing | Primary outcome | Work that should be complete |
|---|---|---|
| Days 1–30: establish authority | The CoE has a mandate and a shared language for value and risk. | Executive sponsor and steering group named; charter approved; initial team assigned; intake brief and triage criteria published; initial policy boundaries set; candidate use cases baselined. |
| Days 31–60: design the path to production | Selected pilots have owners, architecture, controls, and evaluation plans. | Prioritized portfolio agreed; data authority and access maps completed; reference architecture and identity model reviewed; lifecycle gates defined; evaluation set and success thresholds drafted; user workflow and fallback documented. |
| Days 61–90: prove and operationalize | One or two constrained pilots produce evidence, not only demonstrations. | Users test in a controlled environment; results are measured against the baseline; incidents and failure modes are reviewed; reusable assets are published; scale, iterate, hold, or stop decisions are documented. |
At day 90, the steering group should make explicit decisions. Which pilots will proceed, under what residual risk and funding? Which assumptions were disproved? Which controls or platform capabilities need investment? Which assets are ready for reuse? And which ideas should be stopped rather than kept alive as permanent experiments?
Common failure modes—and the corrective action
The most common failure is treating the CoE as a centralized “innovation theatre” function. The corrective action is to attach each initiative to a business owner, a real workflow, a baseline, and a release decision. Another failure is treating the CoE as a compliance checkpoint that sees work only at the end. The corrective action is to publish clear patterns and bring risk, security, data, and operations representatives into use-case discovery.
A third failure is allowing the technology stack to dictate the portfolio. A platform can enable delivery, but it should not determine which operational problem matters. Finally, industrial enterprises must avoid collapsing business and plant-control boundaries because an AI interface makes integration appear easy. The CoE should be the team that asks what happens when the output is wrong, stale, unavailable, or maliciously influenced—and ensures the workflow can fail safely.
Frequently asked questions
References
[1] Microsoft Learn, “Establish an AI Center of Excellence,” https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ai/center-of-excellence
[2] McKinsey & Company, “The state of AI in 2025: Agents, innovation, and transformation,” https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
[3] U.S. National Institute of Standards and Technology, “AI Risk Management Framework,” https://www.nist.gov/itl/ai-risk-management-framework
[4] International Organization for Standardization, “ISO/IEC 42001:2023 — Artificial intelligence management system,” https://www.iso.org/standard/42001
[5] International Society of Automation, “ISA-95 Series of Standards: Enterprise-Control System Integration,” https://www.isa.org/standards-and-publications/isa-standards/isa-95-standard
[6] U.S. National Institute of Standards and Technology, “Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile,” https://doi.org/10.6028/NIST.AI.600-1




