As AI agents move from prototype to production, a fundamental question faces every software company: is our product actually operable by autonomous systems? The Agentability Project has audited 100 leading SaaS products to answer this question, measuring performance across eight Agent Factors Engineering (AFE) principles. The results reveal an industry at an inflection point, with Expensify serving as a representative case study of both progress and persistent gaps.
The Agentability Framework
Agentability measures how effectively AI agents can operate software, scored from 0 to 100. The framework evaluates eight principles:
- Machine readability: Whether interfaces expose structured, parseable data rather than requiring visual interpretation
- Transparency: How clearly a product documents its state, capabilities, and API surface for programmatic access
- Shadow-UI avoidance: The absence of hidden modal dialogs, overlays, and dynamic content that block agent progress
- Defaults: Whether sensible preset values reduce the decisions agents must make
- Control: The availability of APIs, webhooks, and programmatic interfaces
- Chunking: How well complex workflows decompose into discrete, agent-friendly steps
- Status: Visibility into operation state and progress for agent monitoring
- Clean handoffs: Clear signals when human intervention is required
These principles emerged from observing where AI agents consistently fail when attempting to automate real business workflows. Products that score well enable reliable automation; those that don't force brittle workarounds or human supervision.
The State of SaaS Agentability
Across 100 audited products, the average agentability score sits at 38.3 out of 100—a grade that signals widespread unreadiness for the agent-driven workflows already emerging in enterprise environments. The distribution tells a more nuanced story:
- 22 products qualify as Agent-Ready (scoring 45 or above)
- 54 fall into the Developing tier (35-44)
- 17 are categorized as Lagging (20-34)
- 7 products score below 20, effectively Agent-Blind
This distribution suggests most vendors recognize the importance of programmatic access—few products are completely hostile to automation. But recognition hasn't translated into comprehensive implementation of agent-friendly design patterns.
Two Critical Blind Spots
83 of 100 products score zero on transparency, the principle measuring whether software clearly documents its capabilities and state for programmatic consumers.
Transparency represents the most striking failure mode across the audit. Having an API isn't sufficient if agents can't discover what operations are available, what parameters are required, or what state the system is currently in. Most products assume a human user who can explore the interface, read documentation separately, and build mental models through trial and error. Agents require explicit, machine-readable capability declarations—essentially, a product that can explain itself to autonomous systems.
The implications extend beyond technical convenience. When agents can't determine available actions through introspection, they must rely on outdated documentation, brittle hardcoded assumptions, or expensive LLM reasoning to guess at functionality. Each of these fallbacks introduces failure modes that make automation fragile and expensive to maintain.
80 of 100 products score under 20 on shadow-UI avoidance, indicating pervasive use of modal dialogs, pop-ups, and dynamic overlays that block agent workflows.
Shadow-UI elements—modals, tooltips, interstitials, and dynamic overlays—are optimized for human attention but catastrophic for agent operation. A human user immediately recognizes a modal dialog requiring dismissal; an agent may not detect it exists, may lack the capability to interact with it, or may misinterpret its contents as part of the underlying page.
This pattern proliferates because it serves legitimate product goals: progressive disclosure, user onboarding, upsell prompts, and contextual help. But each instance creates a potential blocking point for automation. The low scores suggest most product teams haven't yet considered shadow-UI through an agent-operation lens.
Where Products Excel
The data isn't uniformly discouraging. Machine readability scores average 71 across the audit—a strong signal that most products expose structured data through APIs, webhooks, or queryable interfaces. This foundation is essential; without it, agents would be limited to screen-scraping and computer vision, approaches that are both brittle and expensive.
Control scores average 48, indicating that roughly half of products provide meaningful programmatic access to core functionality. This suggests the basic infrastructure for agent operation exists in many cases; the gaps lie more in discoverability, reliability, and handling edge cases.
What This Means for Product Teams
For teams evaluating their own agentability, the audit data points to three high-impact areas:
Invest in Machine-Readable Capability Documentation
Traditional API documentation serves developers who read English, understand business context, and can experiment interactively. Agents require structured schemas that declare available operations, required parameters, and expected state transitions. OpenAPI specifications are a starting point, but comprehensive agentability requires going further: publishing capability catalogs, state machines, and constraint definitions in formats agents can parse and reason about.
Audit and Eliminate Blocking Shadow-UI
Conduct an inventory of modal dialogs, overlays, and interstitials in critical workflows. For each, ask: Does this block progress if not handled? Can it be detected and dismissed programmatically? Better yet, can it be eliminated entirely for API-driven sessions? Consider implementing a shadow-UI-free mode for authenticated agent access, where all such elements are suppressed or converted to API-layer signals.
Design Workflows for Decomposition
Complex multi-step processes need clean break points where agents can pause, verify state, and handle errors. Rather than requiring 12 sequential operations to complete a task, expose intermediate checkpoints with clear status signals. This improves chunking scores while making failures recoverable and debugging tractable.
Measuring Your Position
The patterns visible in Expensify's audit mirror those across the Agentability Index: strong foundational machine readability undermined by transparency and shadow-UI gaps that make reliable automation difficult. Every product team shipping APIs today should understand where they stand on these dimensions.
The shift toward agent-driven workflows isn't speculative. Enterprise customers are already deploying AI systems that attempt to operate software autonomously, and they're discovering which products enable reliable automation versus which require constant human supervision. Agentability is becoming a competitive differentiator and, for some buyers, a procurement requirement.
Product teams unsure of their current position can run a free audit to benchmark against the principles and identify the highest-impact improvements. The gap between average performance and agent-ready products is significant but closable—particularly for teams that already have API infrastructure and need primarily to address transparency and shadow-UI concerns.
The next generation of software competition won't just be about features humans can access, but about operations agents can reliably execute. Understanding agentability today means being prepared for the automation expectations of tomorrow.
How agent-ready is your product?
Run a free agentability audit and get a scored, prioritized fix list in minutes.
Run a free audit