Making a Website Usable by AI Agents: What Kondevs Did, and What the Data Says Most Sites Get Wrong

According to the Arobis AI Readiness Report (September 2026), 19.2% of sampled websites returned no usable page when fetched as GPTBot, despite only 3.8% explicitly disallowing it in robots.txt. Read that again. Nearly one in five sites is invisible to AI agents for reasons the site owners probably don't know about. The gap between what robots.txt says and what actually happens when a bot tries to load a page is, for most enterprises, an undiagnosed blind spot.

Kondevs published a detailed account of how they made their own site usable by AI agents. The piece is worth reading on its own terms, but the broader question matters more: what does "AI-agent-ready" actually mean for an enterprise website, and where does the industry stand?

The short answer: worse than most teams assume. And the fixes are less glamorous than the hype suggests.

The Three-Layer Problem Nobody Talks About in One Breath

"AI readiness" for a website is not a single checkbox. It breaks into at least three distinct layers, each with its own failure modes. First: can bots actually fetch your pages? Second: do those pages clearly describe what your business does? Third: is key information exposed in machine-readable formats like JSON-LD structured data?

Most conversations about AI discoverability jump straight to the third layer (schema markup, llms.txt files, MCP servers) while ignoring the first two. That's like debating the color of the front door on a house with no foundation.

A GEO audit of 107 small-business sites found that 64.5% did not plainly state what the business does near the top of the homepage. Almost two-thirds. And 27.1% had no JSON-LD structured data on any crawled page. These are not edge cases; they're the norm. If a human visitor struggles to figure out what you sell, an AI agent has no chance.

The Kondevs approach tackled all three layers, which is what makes it instructive. But the real lesson isn't the specific implementation. It's the sequencing: fix fetchability and content clarity before reaching for protocol-level tooling.

llms.txt, MCP, WebMCP: Promising, Unproven, and Draft

The tooling conversation around AI agent readiness has gotten noisy. Three names come up constantly: llms.txt, MCP (Model Context Protocol), and WebMCP. Each solves a different problem, and each sits at a different maturity level.

llms.txt is a lightweight publishing convention intended to point language models to useful site content. Reported adoption ranges wildly: 17.1% in a 1,500-domain scan (AIOScan, August 2026), 22.2% among 455 reachable small-business sites (CitedOn, August 2026), and 48.4% in another sample (Arobis). The variance alone should make you cautious about treating it as settled infrastructure. One study (summarized in the AIOScan report) found that Ahrefs observed no requests to llms.txt on 97% of domains hosting it. Publishing the file is cheap. Expecting it to change outcomes is, at this point, speculative.

MCP is more interesting architecturally. It's a client-server protocol for connecting AI applications to external tools and data, with typed inputs, bounded permissions, and explicit security guidance (input validation, rate limits, confirmation for sensitive actions). It's not a website discovery format; it's an integration protocol. That distinction matters. If you have a high-value workflow worth exposing to agents (booking, quoting, account lookup), MCP is the right conversation. If you're trying to make your "About Us" page findable, it's overkill.

WebMCP is the earliest of the three. The spec (updated October 2026) describes a draft browser API for exposing page functions as tools to compatible in-browser agents. "Draft" is the operative word. Building a critical dependency on it today creates rework risk if the spec moves, and browser support remains limited.

What Kondevs Got Right: Governance Before Glamour

The Kondevs write-up is notable for what it doesn't do. It doesn't treat AI agent readiness as a marketing stunt. It treats it as an infrastructure question: what needs to be fetchable, what needs to be structured, what needs to be secured.

That framing aligns with the MCP specification's own security guidance, which emphasizes that tool interfaces should validate inputs, enforce access controls and rate limits, sanitize outputs, and require confirmation for sensitive actions. In other words, enabling agents to interact with your systems is an integration architecture problem, not a content marketing problem. The same governance instincts that apply to API management (ownership, observability, fallback logic) apply here.

For integration architects and IT leaders in regulated environments (and with EU AI Act transparency obligations applying from 2 August 2026, that's a growing category), the question isn't whether to make systems agent-accessible. It's how to do it without creating a new class of ungoverned interfaces.

The Practical Sequence for Enterprise Teams

Based on the audit data and the Kondevs experience, a defensible order of operations looks like this. Start by testing whether AI crawlers can actually fetch and render your key pages: not just checking robots.txt, but verifying server responses and rendered output. The Arobis data shows that the gap between policy and reality is wide enough to matter commercially.

Then fix content clarity. If your homepage doesn't state what you do in the first scroll, structured data won't save you. After that, implement JSON-LD (Organization and service-level schema first, FAQPage if relevant). The Arobis report found 76.7% of sampled sites used JSON-LD, but only 63.6% declared an Organization entity. That gap is low-effort to close.

llms.txt? Publish it, monitor server logs for requests, and don't expect miracles. MCP? Only when you have a real integration problem it solves, with security controls designed before the demo. WebMCP? Watch, but don't build on it yet.

The Kondevs article is worth reading because it documents someone actually doing the work rather than theorizing about it. In a category full of protocol hype and adoption statistics that contradict each other across every sample, that kind of operational honesty is rarer than it should be. The site that loads, explains itself clearly, and exposes structured information with proper controls will outperform the site with a perfect llms.txt file that nobody requests. Reliability, as usual, is the feature that only gets noticed when it disappears.

Frequently asked questions

How do we test whether AI crawlers can actually fetch our key pages, beyond checking robots.txt?

Verify server responses and rendered output when fetched as GPTBot or similar user agents. The Arobis AI Readiness Report found 19.2% of sites returned no usable page to GPTBot despite only 3.8% disallowing it, so robots.txt alone is not a reliable indicator of actual fetchability.

Should we publish an llms.txt file, and how do we know if it's working?

Publishing llms.txt is low-cost and worth trying, but monitor server logs for actual requests before investing further. One study found Ahrefs observed no requests to llms.txt on 97% of domains hosting it, and adoption figures vary wildly across samples (17% to 48%), so treat it as a controlled experiment rather than proven infrastructure.

When does it make sense to build an MCP server, and what security controls are needed?

MCP is appropriate when you have a real, high-value workflow to expose to AI agents (booking, quoting, account lookup). The MCP specification emphasizes input validation, access controls, rate limits, output sanitization, and confirmation for sensitive actions. It is an integration protocol, not a website discovery format.

What structured data should we implement first to reduce ambiguity about our company?

Start with JSON-LD Organization schema, then add service-level schema and FAQPage if relevant. The Arobis report found 76.7% of sites used JSON-LD but only 63.6% declared an Organization entity, making that gap a low-effort fix with clear value for AI comprehension.

Is WebMCP ready for production use?

No. WebMCP is a draft browser API proposal (spec updated October 2026) for exposing page functions as tools to in-browser agents. Building a critical dependency on it creates rework risk if the spec changes, and browser support remains limited. Watch the standard but don't build on it yet.

Sources

  1. How Many Websites Block AI Crawlers? We Scanned 1,500 - AIOScan, aioscan.com
  2. How many websites block AI crawlers? We audited 107, see-geo.com
  3. llms.txt and AI-Crawler Rules Across 500 Small-Business ..., citedon.com
  4. AI Discovery File Adoption Research - Q3 2026, ai-visibility.org.uk
  5. AI Readiness Report 2026: Website Benchmarks - Arobis AI, arobis.ai
  6. AI Discovery File Adoption Research - Q1 2026 - AI Visibility, ai-visibility.org.uk
  7. AI Search Readiness Study 2026: Data on 419 US Sites, christopholivierconsulting.com
  8. AI Discovery File Adoption Research - Q2 2026, ai-visibility.org.uk
  9. 71 Agentic Web Statistics for 2026, marketsplash.com
  10. www.growthract.com › insights › blogs2026 AI Search Readiness Benchmark: 100 B2B SaaS Websites, growthract.com
  11. Frequently Asked Questions, aeo-expert.nl
  12. www.context.dev › blog › architecting-real-time-web-browsingArchitecting Real-Time Web Browsing for AI Agents: Tool Calling..., context.dev
  13. What Is WebMCP and Should Your Docs or Website Implement It?, documentation.ai
  14. 2026 AI Agent Search Protocols: How MCP, Semantic Markdown, and llms.txt Power Autonomous Web Retrieval, dev.to
  15. Building a Model Context Protocol (MCP) Server for Live Web ..., context.dev
  16. Building AI Agents That Can Safely Work With Live Web Data Using MCP, sitepoint.com
  17. The AI Search Glossary: A definition of 300+ terms, tryprofound.com
  18. github.com › mcp › brightdataMCP Registry | Brightdata · GitHub, github.com
  19. Top 15 Web Search MCPs to Connect Your AI To - Nimble Way, nimbleway.com
  20. Automating the Web with MCP: Infra that Doesn't Break - InfoQ, infoq.com

Related concepts & services

Key terms: AI Agent, Model Context Protocol (MCP)

Related articles

Case Studies

How we made kondevs.com usable by AI agents

What an AI agent can read, call and publish on kondevs.com, how it is built on Cloudflare, and the rules that kept it honest: read-only by default, one table per contract, errors and limits a machine can act on.

12 min read

Ready to streamline your integration?

Tell us the objective, and we'll tell you honestly how we would approach it.

Talk to an integration architect