A buyer who asks an AI assistant "which consultancies in Europe can take over the integration of our SAP landscape?" gets a composed answer, not a page of links. The assistant names a few firms, describes each in its own words and may cite a source or two. If your company is missing, the buyer may never reach your website. If it is there but described wrongly, with an old name or a service you no longer offer, the buyer reads that first.
This article covers how a company becomes findable, readable and quotable in answers from ChatGPT, Claude, Perplexity, Microsoft Copilot, Google's AI Overviews and similar tools: what it controls, what it does not, and what we did on kondevs.com. Our case study How we made kondevs.com usable by AI agents covers the other side: agents that call tools on the site.
Classic search ranks pages and shows links, and the user reads your page in your words. An answer from a large language model is composed, so it may name your company without sending anyone to your site. It cites sources or it does not, and a mention without a link leaves no trace in your analytics. And it describes you in the model's words: if your pages are vague or contradict each other, the description comes from whatever else the model has.
Behind each answer are three paths, and they are easy to confuse.
- Training. A model learns from text collected before its training cutoff; what it knows about you changes only with a new model. Crawlers such as OpenAI's GPTBot and Anthropic's ClaudeBot collect content that may be used for training.
- Search indexes. Assistants with web search look up candidate pages in an index. OpenAI states that sites which block its OAI-SearchBot are not shown in ChatGPT search answers. Google states that a page must be indexed and eligible to be shown with a snippet to appear as a supporting link in AI Overviews or AI Mode, with no additional requirements.
- Retrieval at answer time. The assistant fetches a page while answering and grounds its answer in it, a form of retrieval-augmented generation. ChatGPT-User, Claude-User and Perplexity-User make such fetches on a user's behalf.
Each path has its own crawlers, its own controls and its own pace, so no single measure fixes all three.
Access: can the machines reach the page?
- Crawler rules. robots.txt, standardised as the Robots Exclusion Protocol in RFC 9309, says which crawlers may fetch which paths. It is a request, not access control, and a crawler with a group of its own ignores the rules written for
*. - Usage preferences. Cloudflare's Content Signals Policy of September 2025 adds three preferences to robots.txt:
search,ai-input(use in AI answers at query time) andai-train. Cloudflare calls them preferences, not technical protection. If a CDN manages your robots.txt, read the file it actually serves. - Text in the HTML. A page whose text arrives in the first HTML response can be read by any fetcher; one that builds its text in the browser depends on the reader running scripts.
- A true sitemap and change notices. A sitemap with every real page and honest dates, plus IndexNow for changes. The IndexNow site lists Amazon, Bing, Naver, Seznam.cz, Yandex and Yep (not Google), shares a submission among them and does not guarantee indexing.
Understanding: can a model read the page correctly?
- Content without layout noise. Menus, banners and repeated footers dilute the text a model receives. Real headings, lists and tables help; a plain markdown version helps more.
- Answer-first structure. Lead with the answer, then the detail. A definition that starts "X is..." can be quoted on its own; key takeaways at the top and an FAQ at the end put the short version in a predictable place. A glossary fixes your vocabulary.
- Consistent entity facts. Legal name, former names, founding year, address, registration and VAT numbers, services and people, the same everywhere. A contradiction becomes the model's problem to resolve, and it may resolve it wrongly.
- Structured data. schema.org JSON-LD states the facts in machine form:
OrganizationwithlegalName,vatIDandsameAs,Person,Article,FAQPage,BreadcrumbList,DefinedTerm. Google says its AI features need no special schema and that structured data should match the visible text, so treat markup as a machine-readable copy of the page. Google stopped showing FAQ rich results in May 2026; we keepFAQPagebecause it states each question and answer explicitly for any parser.
Trust and reach: does anyone else confirm what you say?
- Sources and named authors. Linked sources show where a claim comes from; a real author with a profile gives it an owner.
- Consistent profiles elsewhere. Directories, vendor partner listings, LinkedIn, Crunchbase and Wikidata, each with the same facts as your site and linked from your
Organization'ssameAs. - Mentions by others. Articles, talks, partner pages and client references that name you.
This layer is mostly off your site, slow and outside your control. It matters most for comparative questions ("which firms do X in Y?"), because an assistant needs sources other than your own site to compare you with anyone.
We checked every surface below on the live site with plain HTTP requests on 11 October 2026.
Access
- robots.txt has one group for all crawlers and one that names twelve AI user agents, among them GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot and Google-Extended. Both carry the same rules and say
Content-Signal: search=yes, ai-input=yes, ai-train=yes. Only the publishing API is closed. We want to be indexed, quoted and learned from; a company with paid content may decide otherwise. We also check that the file served is the one we wrote, and that a request with an AI crawler's user agent gets the page, not a challenge. - Server-side rendering. Every page is rendered on Cloudflare's edge, so its text is in the first HTML response (why we chose Cloudflare Workers).
- The sitemap lists every page and article (53 addresses on 11 October 2026), with each article's modification date, and pairs each English page with its German counterpart.
- IndexNow. Every article published, updated or removed through our publishing API sends its address and the Content Hub's to IndexNow. In October 2026 we also submitted the whole sitemap.
- Full-text feeds. RSS and JSON Feed carry the newest 50 articles in full; the German Content Hub has its own pair.
Understanding
- Markdown for every page. A request with
Accept: text/markdownreturns a short header (title, description, canonical address, language) and the main content, without menus, footers or banners. - llms.txt follows the format Jeremy Howard proposed in September 2024: a short summary, "When to use KONDEVS" (which page answers which need), "Key facts" (registered details, founding year, services, platforms, named projects, contact) and the main links. llms-full.txt is the whole site in one file, generated from the live content and kept under about 40,000 tokens (about 38,000 on 11 October 2026).
- One JSON-LD graph per page. The
Organizationnode has the same identifier on every page and holds the registered name, registration and VAT numbers, address, founding year, former names, founder andsameAs. Authors arePersonnodes tied to their profile pages. Articles addArticle,BreadcrumbList,FAQPageand, when they list sources,citation; service pages addService; the glossary is aDefinedTermSet. Nothing is marked up that the page does not show, and the registered name, registration and VAT numbers come from one record in the code that also feeds the contact page, the footer and the privacy policy. - Takeaways, FAQs and a glossary. Every Content Hub article ends with an FAQ and the English ones open with key takeaways; each service page ends with an FAQ. The glossary defines 22 terms, each opening with "What is ...?" and a one-sentence answer.
- English and German. German pages mirror the English ones under /de/, each page names its counterpart with
hreflang, and markdown answers in the page's language. An assistant serving a German buyer finds a German page instead of translating an English one.
Trust and reach
- Authors and sources. Each article names its author, with a profile under /authors/ and links to external profiles. Articles that rely on outside claims end with a Sources list, which becomes
citation. - Off the site, the work is the same as for any company: directory profiles with the same facts as the site, linked from
sameAs, and a Wikidata entry once independent sources can support it.
A direct path for assistants that call tools
Assistants that can call tools need no web search for our content: the read-only MCP server at POST /mcp offers search_articles and get_article, which returns an article with its key takeaways, FAQ and sources, and llms.txt asks them to cite the article's address.
Visibility in AI answers is easy to measure badly: one question, one run, one screenshot. Our method is built to avoid that.
- A fixed set of questions. Brand questions (the company name, the domain) and buyer questions phrased as a buyer would ask them, without our name: by platform, by region, by problem. The wording never changes between rounds, so a different answer reflects a different world, not a different question.
- Fixed tools. Each question goes through a standard web search and an AI-assisted search, because assistants that browse build their answers from search results. We record whether kondevs.com appears, where, and whether its title and description are current.
- Entity checks. Wikidata, queried through its API by name and by official website, and the directories buyers consult.
- Indexability checks. The robots header and meta tag on sample pages, the robots.txt served and the sitemap's size. A measurement means little if the site has blocked itself.
- A dated baseline and a schedule. The baseline was taken on 29 September 2026 and repeated on 8 October 2026; the next round is planned for 15 October 2026. Each round records the date, the exact query, the result and the links.
This article describes the method, not our results. What nobody can measure precisely:
- Answers vary by day, user, location and model version. One run is an anecdote.
- There is no complete record. No assistant publishes every answer that named a company, and a mention without a link never reaches your analytics.
- First-party reports are partial. Google Search Console counts AI Overviews and AI Mode inside the "Web" search type, not separately. Bing Webmaster Tools' AI Performance report, in public preview since February 2026, counts citations in Copilot and Bing's AI summaries, not visits.
- Training data is a black box. You cannot see whether a page is in a training set, or when a model will reflect a change.
- Cause is hard to prove. A change in answers may come from a new model version or a competitor's page rather than from your work.
When you ask an assistant directly, use a fresh session without history, note the model and the date, keep the full answer and ask more than once.
- Write down your facts once. Legal name, former names, founding year, address, registration and VAT numbers, services, people and contacts, in one register that every page, profile and markup copies from. Remove contradictions first.
- Decide your crawler policy, then check what is served. Choose per crawler and purpose: search, answers, training. Read the robots.txt your CDN actually serves.
- Make the text readable without a browser. Fetch your key pages with a plain HTTP client. If the service descriptions are missing, fix rendering before touching markup.
- Rewrite the key pages answer-first. Each service page should say what it is, who it is for and what the client gets in its first lines, and end with the questions buyers actually ask.
- Add structured data that matches the page.
Organization,Person,Article,FAQPageandBreadcrumbListin one graph with stable identifiers. Never mark up what the page does not show. - Publish a short llms.txt and keep the index fresh. Key facts, when to use you, the main links; a sitemap with true dates; IndexNow; the site verified in Google Search Console and Bing Webmaster Tools.
- Build the off-site footprint, and measure. Directory profiles with the same facts, Wikidata when independent references exist, all linked from
sameAs. Take the baseline before you start, so you can see what moved.
Most of this is not marketing work. It is one set of facts published consistently through many interfaces: pages, markup, files, feeds and tools. We solve the same problem in enterprise landscapes, where an ERP, a CRM and a partner portal must tell the same story about the same customer. If you are working out what your website and your systems should expose to AI assistants, and in which order, talk to an integration architect. Tell us the objective, and we will tell you honestly how we would approach it.


