The AI readiness assessment tool you choose depends on the level you need to assess. Microsoft assessment and Cisco index offer free organizational self-assessments. AWS CAF-AI docs and Google framework offer free frameworks or security-focused documentation. These can establish an organizational baseline in 1-2 hours, but they do not evaluate codebase health, test suite quality, or review capacity. Academic research identifies those factors as the primary variables shaping AI coding tool outcomes.
TL;DR
The tools below span free vendor questionnaires to strategy-firm engagements above $500,000. I compared their published methodologies against DORA research, peer-reviewed studies, and practitioner accounts. The 11 tools marked not engineering-specific operate at the company level and skip codebase context, technical debt baselines, and workflow maturity, which are the factors that determine whether AI coding output becomes maintainable shipped code.
Engineering leaders keep hitting the same wall: the organization scores "scaling" on a maturity questionnaire, the pilot launches, and six months later the initiative stalls. McKinsey's State of AI 2025 survey reports that 88% of respondents use AI regularly in at least one business function, yet less than one in five say their organizations track KPIs for gen AI systems.
I spent time with the methodology documents, scoring rubrics, and pricing structures behind 17 named assessment tools across five categories. The goal was to identify which tools belong in an evaluation and which engineering-readiness factors each method omits. The Augment Code AI readiness framework is one of two tools in this 17-tool set built specifically for engineering teams, and it pairs with Augment Cosmos, the unified cloud agents platform that gives engineering teams shared context and memory across the software development lifecycle. I include the readiness framework in the comparison below alongside the research-firm, vendor, and consulting options. I used it to evaluate a rollout plan for a brownfield legacy repository, and it surfaced delivery workflows, review load, and codebase constraints because the diagnostic evaluates readiness below the company-level strategy and governance layer.
The Full Field: Seventeen Named Tools Across Five Categories
The tools I found split into research-firm diagnostics, free vendor self-assessments, consulting engagements, government and academic models, and a new engineering-specific category that emerged in 2026. In this set, 11 are not engineering-specific, four are partial engineering assessments, and two are explicitly engineering-specific.
| Tool | Vendor | Format | Access | Engineering-specific? |
|---|---|---|---|---|
| AI Maturity Model & AI Roadmap Toolkit | Gartner | Interactive tool + framework | Client-only | No |
| CIO Playbook: AI Readiness Assessment | IDC | Online self-assessment | Free | No |
| AI MaturityScape Framework Guide | IDC | Downloadable guide | Gated | No |
| AI Readiness Assessment | Microsoft | 45-minute multiple-choice online assessment | Free | No |
| AI Readiness Index | Cisco | Self-assessment + annual report | Free | No |
| M365 Copilot Assessment | Cloudiway | Automated tenant audit | Commercial | No |
| Cloud Adoption Framework for AI (CAF-AI) | AWS | Whitepaper / documentation | Free | Partial |
| AI Adoption Framework | Google Cloud | Whitepaper | Free | No |
| SAIF Risk Self-Assessment | Google Cloud | Security self-assessment | Free | Partial |
| CARE Score | CloudBees | Self-assessment scoring | Public | Yes |
| AI Readiness Assessment Framework | Augment Code | Diagnostic guide | Free | Yes |
| AI Readiness Assessment | Winder.AI | Consulting-led | Paid | Partial |
| GenAI Readiness and Adoption | Deloitte | Consulting engagement | Paid | No |
| AI Maturity Model | MITRE | Organizational self-assessment | Free | No |
| Enterprise AI Maturity Model | MIT CISR | Research-based model | Public | No |
| AI Adoption Maturity Model | Accenture + CMU SEI | Assessment + benchmarking | Launched 2026 | Partial |
| 2026 Agentic AI Readiness Index | Fivetran | Survey/index report | Published | No |
A few of these tools deserve direct references beyond the table. IDC's AI MaturityScape framework guide is the gated downloadable companion to the free CIO Playbook. The Cloudiway M365 Copilot assessment is a commercial tenant audit for Microsoft 365 environments rather than a general readiness questionnaire. MIT Sloan CISR's research-based Enterprise AI Maturity Model sits alongside MITRE as one of the two most vendor-neutral academic entries.
Two entries deserve flags. The Fivetran index is a vendor-sponsored survey. The Accenture and Carnegie Mellon SEI model is the only enterprise framework in this review that names an explicit engineering dimension among its eight core dimensions.
Free Self-Assessments: Useful Baselines With Built-In Bias
The free vendor tools give you an organizational baseline in an afternoon. Each one also tilts toward its vendor's commercial interests. I'd still run one or two of them; read the results knowing who wrote the questions.
Microsoft AI Readiness Assessment

Microsoft's free tool at Microsoft Learn assessment asks 45 questions grouped into seven pillars: Business Strategy, AI Governance & Security, Data Foundations, AI Strategy & Experience, Organization & Culture, Infrastructure for AI, and Model Management. It outputs a Microsoft maturity score across five maturity stages, from exploring through realizing. The tool may require Microsoft tenant permissions, and Microsoft does not publish a public scoring methodology.
Cisco AI Readiness Index

Cisco's Cisco assessment tool benchmarks you against a double-blind survey of 8,161 business leaders at 500+ employee organizations across 30 markets. The Cisco 2025 index scores 49 indicators across six weighted pillars: Infrastructure at 25%, Data at 20%, Strategy at 15%, Governance at 15%, Talent at 15%, and Culture at 10%. Indicators earn 25-50% credit for partial deployment and 100% for full deployment. Organizations land in one of four bands: Pacesetters (86-100), Chasers (61-85), Followers (31-60), and Laggards (0-30). In 2025, only 13% of organizations qualified as Pacesetters while 48% landed as Followers.
The bias sits in the weighting. Infrastructure carries the heaviest weight at 25%, Cisco sells infrastructure, and Cisco publishes a separate Infrastructure Focus report comparing Pacesetters against everyone else on infrastructure metrics. Culture gets the lowest explicit weight at 10%, even though practitioner research suggests it predicts adoption failure.
IDC CIO Playbook and MITRE AI Maturity Model

IDC's CIO Playbook assessment (September 2025) is free and adds regional and industry peer benchmarking, which the other free tools mostly lack. MITRE's AI Maturity Model covers strategy and governance, data and infrastructure, AI development and operations, workforce and culture, and risk and assurance. MITRE carries no vendor sales agenda, which makes it my preferred neutral option, though it leans less action-oriented than the commercial tools.
AWS and Google Offer Frameworks

AWS offers multiple scored public AI readiness assessments via AWS Marketplace, while Google does not appear to offer an analogous scored organizational AI readiness assessment. AWS's AWS CAF-AI is explicitly a "mental model," a documentation set spanning Business, Governance, Operations, and Security perspectives. AWS also states outright that it provides "prescriptive guidance" for implementing AI "on AWS." Google's AI Adoption Framework is a whitepaper. Its Data and AI Strategy Assessment runs through a Google sales assessment rather than as a standalone public tool. Use Google's SAIF Risk Self-Assessment when your concern is AI security posture against Google's Secure AI Framework (SAIF).
Paid Assessments: What $15,000 to $500,000 Buys
Consulting-led assessments buy depth, named owners, and a roadmap that free questionnaires can't produce. Prices vary by an order of magnitude. Published pricing estimates put free vendor self-assessments at 1-2 hours of effort. Independent mid-market assessments run roughly $15,000-$75,000 for 2-4 weeks. Enterprise practitioner assessments run $40,000-$120,000 for 4-8 weeks. Big Four or strategy-firm engagements run $100,000-$500,000+. AI-native sprints run $75,000-$250,000 over 90 days.
| Tier | Cost | Scope | Typical output |
|---|---|---|---|
| Free vendor self-assessment | $0 | 1-2 hours | Score, maturity stage, benchmark |
| Independent / mid-market | $15,000-$75,000 | 2-4 weeks, 1-2 consultants | Gap assessment, initial roadmap |
| Enterprise practitioner | $40,000-$120,000 | 4-8 weeks | Buildable roadmap for $1B+ enterprises |
| Big Four / strategy firm | $100,000-$500,000+ | Multi-week | Board-level report, multi-year program proposal |
| AI-native sprint | $75,000-$250,000 | 90 days, assessment + implementation | Combined engagement |
Gartner's Gartner toolkit sits behind a client relationship. It uses a seven-question survey rated Level 1 through Level 5 across five maturity stages. Gartner's Q4 2024 benchmark of 432 respondents reported that high-maturity organizations averaged 4.2-4.5 scores versus 1.6-2.2 for low-maturity ones. Deloitte's GenAI assessment requires a full consulting engagement. Winder.AI describes its 2026 engagement as the most engineering-adjacent of the consulting options. It stays model- and platform-agnostic and covers agentic systems, the EU AI Act enforcement timeline, and self-hosted open-source models.
One caution applies across the paid tier. CMU SEI notes that available AI maturity guidance faces limited evidence of real-world value and difficulty staying relevant as the technology evolves. Ask what deliverable you'd act on before you sign, and whether the firm has any incentive to find you unready.
Seven Gaps: What Every Enterprise Assessment Misses for Engineering Teams
Enterprise frameworks usually score strategy, data, governance, talent, and culture. AI coding rollouts need seven additional engineering checks. I could not find these dimensions in the frameworks above, and research links each one to outcomes.
Gap 1: Codebase context readiness
Codebase context readiness asks whether your codebase is structurally legible to AI tools. Most AI tools treat legacy code as isolated snippets and miss the context that explains why systems evolved into their current state, an observation reinforced by our own work on AI-powered legacy code refactoring. I tested this framing on a legacy-codebase modernization review. The relevant check was whether a tool could analyze repository-level context rather than isolated snippets. The Context Engine processes entire codebases across 400,000+ files through semantic dependency graph analysis. Augment has also published benchmark results for related products, including Context Engine MCP and Augment Code Review.
Gap 2: Documentation health and knowledge silos
DORA research found that giving teams AI tools with direct access to internal data (internal codebases, architectural diagrams, wikis, documentation, style guides) is a statistically significant multiplier for individual effectiveness and code quality. Yet DORA's 2024 report found a 25% increase in AI adoption associated with only a 7.5% increase in documentation quality, so documentation lags adoption rather than leading it. Practitioner Pete Hodgson's analysis describes what goes wrong: a coding agent produces the "wrong" solution because "it simply doesn't know the conventions that it should be adhering to, and doesn't have enough context about the existing codebase and architecture."
Gap 3: Technical debt baseline
A technical debt baseline shows your debt position before AI multiplies it. A large-scale empirical study found that AI coding assistants introduce more issues than they fix for runtime bugs and security issues. That baseline matters because AI-generated code can work, pass tests, and look reasonable while carrying design flaws that only surface later, especially when the intent behind the code is fuzzy.
Gap 4: Workflow maturity and delivery pipeline readiness
Workflow maturity and delivery pipeline readiness matter because AI output creates verification work inside existing version control and review processes. A productivity study on brownfield projects quantified one hidden cost: a "verification tax" of an estimated 4.3 minutes per suggestion spent checking AI output against existing code. Teams should identify workflow bottlenecks before asking whether AI reduces that friction, which is why the readiness framework we built leads with delivery friction rather than model selection. I used that workflow-bottleneck lens on a review-heavy brownfield rollout plan, and the exercise put review capacity at the center.
This gap also explains why individual adoption can diverge from organizational results. Augment Cosmos is the unified cloud agents platform that runs agents in the cloud with shared context and memory that compound across the team and the software development lifecycle. It exposes three primitives (Environments, Experts, and Sessions), so an agent workflow one engineer builds becomes an auditable, replayable capability the whole team can draw on. Building Cosmos surfaced four compounding problems that appear when engineers adopt agents without a unifying system:
- Setups fragment.
- Expertise gets trapped in one engineer's config.
- No quality signal shows which agent setups work across teams.
- The review bottleneck worsens because humans get pulled in only at the final PR.
The named public questionnaires in this review do not ask about those four failure modes. I tested the Cosmos Sessions model on a workflow where one engineer's billing-service prompt needed to become reusable team practice. Auditable, replayable sessions can turn one-off prompts into reusable workflows, though they do not by themselves prove which setup improves throughput. I also tested the Cosmos human-checkpoint model on a Linear-ticket-to-PR workflow. It moved review earlier than the final PR because teams can set policies for where human judgment is required before agents finish the work.
Gap 5: Greenfield versus brownfield codebase maturity
Academic research identifies codebase maturity as one of several key variables shaping AI coding outcomes. One arXiv paper notes that greenfield projects "dominate the positive-result studies" with minimal constraints and maximal AI use. The same paper found that brownfield projects with established conventions, test suites, and architectural patterns impose the verification overhead that the METR study measures as the primary cost. A coding readiness model documents the universal pattern: "fast initial satisfaction followed by systemic frustration..." Teams needed "a systematic approach to making the AI effective within a real codebase over time."
Run your software agents at scale
Cosmos gives your agents the context, tools, and feedback loops they need to get better with every workflow.

Gap 6: Test suite health and review capacity
AI changes the economics of testing in a way standard enterprise questionnaires rarely measure. The readiness question is whether generated tests increase useful regression protection or increase maintenance load. Teams can drown in low-value tests that make every refactor expensive, especially when review capacity is already the bottleneck.
Gap 7: Trust and the leadership-engineer perception gap
Organizational readiness assessments often include leaders as well as employees and end users, and recent AI-readiness surveys suggest leaders often overestimate readiness. DORA's 2024 report found that 39% of respondents report little to no trust in AI-generated code. When the CTO fills out a questionnaire, the answer captures executive belief. Developers still need enough trust to review, merge, and maintain changes that AI generates.
The Emerging Codebase-Readiness Category
The tools below assess whether codebases can support AI-assisted development. No single tool in this set yet combines codebase health scoring, AI-context readiness, and developer skills assessment in one framework.
- CodeScene AI-Ready Code / CodeHealth™: CodeScene's vendor-reported whitepaper (March 2026 edition) reported that "AI-generated changes fail significantly more often in unhealthy code, with defect risk rising by at least 60%." Its MCP server exposes Code Health analysis to safeguard AI-generated code.
- Galarza readiness assessment: Scores repositories 0-100 across eight dimensions, including test foundation, architecture clarity, type safety, and documentation, with bands from Agent-Ready (80+) down to Foundation (<40), plus a prioritized roadmap.
- @kodus/agent-readiness: The GitHub repository describes @kodus/agent-readiness as the open-source alternative to Factory.ai's Agent Readiness. The project uses TypeScript/Bun and carries an MIT license.
- CloudBees CARE Score: A self-assessment covering governance, cost visibility, productivity measurement, and pipeline visibility. For calibration, CloudBees reports that organizations self-assessed at an average of 83.6 out of 100, which CloudBees contrasts against operational reality.
- DX AI Measurement Framework™: DX describes its framework as measuring adoption and results across utilization, impact, and ROI. DX cites research spanning 435 companies and 135,000 developers. DX's vendor-reported longitudinal data says that across 400+ companies, AI tool usage rose an average of 65% while median PR throughput rose 7.76%, and "for most organizations, a 5-15% throughput gain is what the current generation of AI coding tools is delivering."
- Augment Code AI Readiness Assessment Framework: A free diagnostic guide (July 2026) that evaluates data infrastructure, tooling integration, team skills, governance, and strategy, and explicitly covers delivery workflows, review load, and codebase constraints. Its scope matches the seven gaps above more closely than any enterprise framework I reviewed. Apply the same skeptical lens you'd apply to any vendor-published diagnostic: mapped against the 17-tool matrix, it was the only vendor diagnostic in this set that explicitly covered delivery workflows, review load, and codebase constraints, and it evaluates the engineering system as well as the organizational program.
The DX throughput numbers make measurement a readiness requirement from day one. DX also reports customer examples it measured. A Booking.com measurement across 3,500 engineers produced a 65% increase in AI adoption through data-driven measurement. An Indeed training rollout moved adoption from roughly 25% to 97% across 2,000+ engineers after redesigning training around hands-on learning. The gap between adoption and throughput also shows why shared systems matter more than individual tooling. Agents that carry context and learned corrections across a team can compound in ways per-developer setups can't. I tested the Cosmos session-promotion pattern on a team rollout plan where AI usage rose faster than PR throughput. The mechanism explained why adoption alone underdelivers: corrections need a path into reusable team workflows rather than remaining in per-developer configs.
How to Choose: Match the Tool to the Decision
Let the decision define the assessment you choose. My recommendations after comparing the methodologies:
- Choose a free vendor self-assessment (Cisco, Microsoft, IDC) if you need a board-ready organizational baseline in under two hours and can discount each vendor's product-line bias.
- Choose MITRE if you want a vendor-neutral organizational model with no sales pipeline attached.
- Choose the Accenture + CMU SEI model if you need enterprise benchmarking that at least names Engineering as a dimension, keeping in mind SEI's caveat that all available AI maturity guidance faces limited evidence of real-world value.
- Choose a consulting engagement ($15K-$500K+) only when you need named owners, sequencing, and budget bands, and demand engineering roadmaps as the deliverable.
- Choose a codebase-readiness tool (CodeScene, Galarza, @kodus/agent-readiness) plus a measurement framework (DX) if the decision is whether and how to roll out AI coding tools. This pairing covers the seven gaps enterprise assessments miss.
- Choose the Augment Code engineering framework if you need to link organizational readiness scores to repository, workflow, and review-capacity checks. I used that approach after an org-level baseline, and it translated readiness into the codebase, workflow, and review-capacity checks engineering teams can change.
The sequence matters as much as the tool choice. Start with a free organizational assessment when the audience is the board or executive team, then move immediately to repository-level checks before approving a coding-assistant production rollout. A Cisco or Microsoft score can tell you whether strategy, governance, and infrastructure are ready enough to examine in more detail. A codebase-readiness review evaluates whether a specific brownfield system has enough documentation, test coverage, and review capacity to absorb AI output. Without those conditions, AI output can create more technical debt. Pair that with DX-style measurement, because throughput matters more than adoption alone. For engineering leaders, readiness depends on the codebase, reviewers, and workflow turning AI output into maintainable shipped code, and that requires both organizational and engineering levels of assessment, measured before rollout.
Whatever you pick, replace generic surveys with targeted developer friction interviews. The EPAM engineering playbook recommends asking, "Where is AI wasting your time today?" That vendor-published playbook also notes that the most impactful blockers often come from low-level constraints like "slow IDEs, corporate proxies blocking APIs, or restrictive security defaults." The named public questionnaires in this review do not ask about those low-level blockers; conversations do.
Run the Org-Level Assessment, Then Audit the Level It Skips
After an assessment, an engineering leader can face a mismatch between the organizational score and the repository reality. Your organization can score as a Pacesetter while your 10-year-old monorepo, thin documentation, and overloaded reviewers guarantee the frustration cycle that the research documents. Most standard questionnaires measure organizational readiness and leave the engineering checks for a separate audit.
Run a free organizational assessment for the baseline, then spend one week where enterprise tools cannot see: score a critical repository with a codebase-readiness tool, interview five developers about where AI wastes time, and baseline DORA metrics before rollout. The gap between a board-level Pacesetter score and the seven engineering checks above defines the work that Augment Cosmos targets. Cosmos gives engineering teams a shared cloud agents platform where context, memory, and reusable workflows compound rather than staying trapped in individual configs. Organization-wide change requires shared systems beyond individual adoption.
FAQ
Related
Written by

Molisha Shah
GTM
Molisha is an early GTM and Customer Champion at Augment Code, where she focuses on helping developers understand and adopt modern AI coding practices. She writes about clean code principles, agentic development environments, and how teams are restructuring their workflows around AI agents. She holds a degree in Business and Cognitive Science from UC Berkeley.