Seventeen organizations have each decided the world needs their own way of scoring how "AI-mature" a business is. Gartner has one. So does McKinsey, BCG, PwC, IBM, AWS, MIT, the OECD, and the government of Singapore. Some of these are rigorous, evidence-based diagnostics built on survey data. Others are thinly documented marketing artifacts attached to a consulting pitch or a cloud contract.
This is a walk through all seventeen: what each one measures, what it's good for, and where it falls apart. Then a comparison across all of them, including the dimension most of them get wrong, and a look at where a much simpler self-assessment fits among a set of tools mostly built for organizations ten times the size of the ones reading this.
Two numbers are worth holding onto going in. AI adoption could add roughly $13 trillion to global economic output by 2030, about 1.2% of additional annual productivity growth, more than double what the steam engine or industrial robots added in their respective eras. And most estimates put the share of AI pilots that never make it past a proof of concept somewhere between 70% and 95%. Maturity frameworks exist to close that gap between the upside and the failure rate. Whether any given one does that for your business depends entirely on which one you pick.
These ten are built for organizations with multiple business units, dedicated data and platform teams, and enough budget to run a formal audit. They tend to be the most rigorously researched of the seventeen, and also the least usable without outside help.
Gartner's model evaluates a business across seven capability pillars, strategy, governance, data, product, engineering, operating models, and culture, through a guided questionnaire that outputs a visual heat map of where current capability sits against a target. It maps organizations to one of five stages, from Awareness to Transformational, and, notably, pushes users to map out process friction and manual workarounds before recommending any automation, on the theory that automating a broken process just makes the mess move faster.
McKinsey's assessment, built from its "Rewired" research and QuantumBlack's client work, evaluates six capabilities: a transformation roadmap tied to real value, a bench of specialist talent, a fast-moving operating model, a flexible distributed technology environment, data embedded across the business, and a genuine adoption and scaling motion. Getting there involves auditing budgets, governance policies, and org charts, followed by structured interviews with six to ten senior leaders. McKinsey's own research is blunt about the state of play: almost every company is investing in AI, but only about 1% consider themselves fully mature.
BCG runs two related frameworks. The first is the 10-20-70 rule: in BCG's research, only 10% of AI value creation comes from the algorithm itself, 20% from technical infrastructure and data, and 70% from people, process redesign, and change management. It sorts companies into four brackets, from Stagnating to Future-Built, based on capability across 53 digital and AI dimensions.
The second is Deploy-Reshape-Invent, three value plays that sequence how a business should expect returns:
BCG's own research suggests that for most companies, over 90% of eventual AI value comes from Reshape and Invent, not Deploy.
PwC's model produces a precise 0 to 100% score across twelve operational domains, mapped to five bands: AI-Vulnerable (0-29%), AI-Aware (30-48%), AI-Developing (49-66%), AI-Advanced (67-83%), and AI-Driven Leader (84-100%). The score comes from senior leader interviews benchmarked against a database of more than 500 organizations, and it's specifically designed to surface the gap between what a CEO and a CTO each believe about the company's own readiness, which PwC reports is often a double-digit-point spread.
MITRE's model spans six pillars (ethical and responsible use, strategy and resources, organization, technology enablers, data, and performance) broken into twenty specific metrics, scored through a free multiple-choice questionnaire called the Organizational Assessment Tool. It maps to five levels, Initial through Optimized, and deliberately doesn't require every organization to hit Level 5 on every dimension. A defense contractor and a mid-market retailer can set very different target profiles and both be "done."
CMMI Institute, the standards body behind decades of software process maturity work, launched AIM in mid-2026, extending its existing 31 practice areas with AI-specific guidance across eight domains: data, development, people, safety, security, services, suppliers, and virtual. Getting scored requires a formal appraisal from a certified Lead Appraiser, and the model includes crosswalks that map directly onto ISO/IEC 42001, ISO 23053, ISO 23894, and ISO 31000, letting one appraisal double as evidence for multiple compliance regimes at once.
IBM's original AI Ladder (Collect, Organize, Analyze, Infuse) has evolved into a generative-AI-specific model that scores seven dimensions on a 0-3 scale, rolling up into Silver, Gold, or Platinum. For generative AI specifically, IBM lays out five phases, from simply consuming an off-the-shelf model through to building and running custom models securely across environments, with explicit guidance on things like Model Context Protocol integration and real-time bias detection along the way.
AWS extends its existing Cloud Adoption Framework into four levels: Envision, Experiment, Launch, and Scale, evaluated across five pillars (business, people, governance, platform, operations). It's explicitly built to move a team from mapping theoretical use cases to running production workloads with SLAs, and eventually to a self-service marketplace of reusable internal components.
Accenture's research sorts companies into four groups rather than a strict ladder: AI Experimenters, Builders, Innovators, and Achievers, based on foundational versus differentiating capability. In its widely cited study, 63% of companies land in the Experimenter group, with a maturity score around 29 out of 100, while a smaller group of Achievers scores roughly double that and correlates with meaningfully higher revenue growth.
Built by MIT's Center for Information Systems Research from a survey of 721 companies plus follow-up executive interviews, this model maps four stages: Experiment and Prepare, Build Pilots and Capabilities, Industrialize, and Become AI Future Ready. The distribution is telling: 28% of companies sit at Stage 1, 34% at Stage 2, and only around 7% ever reach the top stage. Financial performance tracks the stages closely, with Stage 1 and 2 companies underperforming their industry averages and Stage 3 and 4 companies well ahead of them.
Smaller businesses and public bodies run into a different problem than enterprises: not too little rigor, but frameworks that assume resources they don't have. These four were built with that constraint in mind.
Built under the EU's Digital Europe Programme and delivered free through the European Digital Innovation Hubs network, the DMA evaluates six dimensions, including digital strategy, human-centric digitalization, data management, and AI and automation. Rather than a self-scored questionnaire, an EDIH representative runs a structured interview and maps the business's actual practices to predefined behavioral indicators, producing a 0-100 score benchmarked against similar-sized regional and EU peers. Completing it can also unlock direct digitization grants for micro and small enterprises.
Built by AI Singapore, AIRI scores organizations across five pillars and twelve dimensions on a 1-4 scale, landing them in one of four bands: AI Unaware, AI Aware, AI Ready, and AI Competent. It's designed to be run without outside consultants, and results map directly to Singapore's state-funded training and engineering support programs, closing what the index's own research calls the SME "execution gap."
An academic model built for industrial SMEs, DAMA-AHP scores 66 distinct elements across six dimensions, from people and expertise to production processes. What sets it apart is the Analytic Hierarchy Process it borrows from decision science, which lets a company weight each dimension and element according to its own strategic priorities rather than accepting a fixed, generic weighting.
Less a maturity ladder than a classification lens, the OECD's framework maps an AI system's characteristics against five values-based principles: inclusive growth, human-centered fairness, transparency, robustness and safety, and accountability. It's built for policymakers and risk assessors as much as business leaders, tracing data handling, traceability, and human oversight across an AI system's full lifecycle.
Forrester's model checks whether the baseline structures, initial setup, technical tooling, organizational preparation, are in place before a company launches a large-scale AI initiative. It functions less as an ongoing maturity ladder and more as a one-time, pre-project readiness check, which is exactly what it's good for: catching a gap before it becomes an expensive mistake. Much of Forrester's own methodology sits behind paid research, so public documentation on it is thinner than most of the frameworks above.
Salesforce's model tracks the maturity of an organization's responsible-AI practices specifically, across four stages: Ad hoc, Organized and repeatable, Managed and sustainable, and Optimized and innovative. It's narrower by design than most of the frameworks here, focused entirely on ethical design, bias review, and workforce alignment around responsible AI use rather than technical or financial capability.
Element AI's model, from a whitepaper published before the company was acquired by ServiceNow in 2020, evaluates five dimensions, strategy, data, technology, people, and governance, across five levels from Exploring to Transforming. It's worth knowing this one is a legacy artifact: with Element AI no longer operating independently, nobody is actively updating it for multi-agent systems the way an active vendor would.
Line them up and three clusters emerge, and which cluster a framework sits in tells you more about whether it'll work for your business than any individual feature does.
What almost none of them share is a scoring philosophy that guards against the most common way these assessments go wrong: a simple average across dimensions. If a business scores well on strategy and technology but poorly on governance and data quality, a plain average can still produce a healthy-looking overall number, one that clears the business to scale. Several of the more rigorous frameworks here, McKinsey's among them, instead anchor the overall score to the weakest dimensions rather than the mean, on the logic that a business with brilliant algorithms and no data governance isn't moderately mature, it's one incident away from a real problem. It's a principle worth borrowing even if you're not running any of these formally: don't let a strong score in one column paper over a failing one in another.
A fair question, especially for a team that's already run a CMMI or TOGAF assessment before: why does AI need its own maturity model at all? The honest answer is that AI systems behave in ways traditional IT governance was never built to catch.
That last row is the one worth sitting with longest. A traditional IT system either works or it doesn't, and when it fails, you can usually trace exactly why. An AI system can produce a confident, plausible, wrong answer with no error message at all, which is why every serious framework in this list treats a defined human-in-the-loop checkpoint as a maturity requirement, not an optional extra.
Enterprise maturity models spend most of their attention on scaling risk. For a smaller business, the more immediate risk is continuity. It's a familiar pattern: a business "adopts AI," but what that means is one or two employees using consumer AI tools in browser tabs, with no shared documentation of how those tools are configured or what they're being trusted to do. If that person leaves, the knowledge usually leaves with them, and if there's no defined process for catching a wrong or questionable AI output before it reaches a client, nobody's checking until something has already gone out the door. This both a capability gap as well as a business continuity risk wearing an AI-shaped costume.
This is exactly the gap the SME-specific frameworks above are built to close, and it's a large part of why we built the process redesign work we do at Orbflo around documentation and handoff points first, rather than treating a slick AI tool as evidence that the underlying process is sound.
Every framework above was built to evaluate businesses using AI as a tool a person prompts. The frontier that's already arriving is businesses running networks of autonomous agents that coordinate entire processes with far less human prompting in the loop, and most of these seventeen models haven't caught up to it yet. A handful, mainly AWS's and Microsoft's newer Agentic Adoption Model, have started building in three specific checks worth watching for regardless of which framework you use: whether your systems support modern integration standards like the Model Context Protocol so agents can connect to your tools and databases directly, whether your data can be delivered as a clean, low-latency stream rather than an overnight batch job, and whether governance checks (bias detection, drift monitoring, prompt-injection defenses) are built directly into your deployment pipeline instead of happening as a manual review after the fact. A framework that doesn't ask any of these three questions yet is evaluating yesterday's version of the problem.
The honest starting point is scale, not preference. A large, regulated enterprise chasing audit-grade compliance has real reasons to reach for CMMI AIM or PwC's assessment. A company mainly trying to fix its change-management and workforce problem is better served by BCG's 10-20-70 rule or McKinsey's approach. A cloud-native engineering team is going to get more out of AWS's or IBM's model than out of anything built for a boardroom. And a small or mid-sized business without a transformation budget is generally better off with something built for that constraint, like Singapore's AIRI or the EU's DMA tool, assuming it operates in one of those regions.
Outside those two regions, there isn't an obvious equivalent, which is part of why we built the AI Operating System Scorecard the way we did: a short, self-administered diagnostic across nine dimensions rather than a twelve-domain enterprise audit. It doesn't compete with McKinsey's or PwC's depth, and it isn't trying to. It's closer in spirit to AIRI or the DMA tool: something a business without a dedicated transformation team can run itself, in an afternoon, to get a rough read on where the real gaps are before deciding whether a heavier framework is worth the investment.
Whichever one you use, two habits carry across all seventeen. Score conservatively rather than averaging, so a strong result in one column can't quietly cover for a weak one in another.
If you want a starting number
The AI Operating System Scorecard takes about the same amount of time as reading one section of this article, and gives you a baseline read on decision authority, data readiness, and process clarity, three of the dimensions every framework above treats as foundational, before you decide whether a heavier assessment is worth running.

No, and treating any of these seventeen as universal is the fastest way to waste the exercise. Enterprise frameworks assume a transformation budget and specialist staff most companies don't have. Hyperscaler models assume you're building on that specific cloud. SME frameworks assume a size and geography that a mid-sized enterprise has usually outgrown. Match the framework to your actual constraints before you match it to its reputation.
Technically, yes, but it's rarely a good use of the effort. These models are built around weeks of leadership interviews, formal document audits, and in PwC's case, an actual advisory engagement to complete properly. A smaller business is almost always better served starting with a lighter, self-administered diagnostic and reserving a heavier framework for after it has a specific, well-scoped problem worth that level of rigor.
Averaging dimension scores instead of anchoring to the weakest one. A business can look mature on paper, strong strategy, strong technology, while its data governance or human-review process is genuinely broken, and an averaged score will hide that. Most of the frameworks that have caused real damage when followed too literally are the ones used this way. The fix costs nothing: score conservatively, and fix the weakest dimension before scaling anything built on top of it.
The AI Operating System Scorecard is a diagnostic tool that measures whether your business is structurally built to make AI compound, across nine dimensions including how decisions get made, how clearly your processes are defined and how your team is using and integrating AI.
The output is a clear view of where your biggest leverage gaps are and where to focus first.
One practical AI operating-system insight bi-weekly.
No fluff, no spam.