Your customer asks AI a question. The service searches the web, retrieves a few passages, and writes an answer before they ever visit your site.

When I publish an article on Finding Direction, one Markdown file supplies the article page, Thinking archive, RSS feed, sitemap, and related-content links. The file holds the title, summary, date, status, topics, body, source links, and image descriptions. Each published form points back to the same work.

AI search enters through a smaller opening. It can retrieve one paragraph from the article, combine it with passages from other sites, and present the result inside an answer. The reader encounters the idea before the page that gives it context.

The website remains the place where the complete thought, evidence, and authorship stay together. The publishing work now includes making those parts clear when a search service retrieves only a piece.

Weathered headstones and iron grave enclosures beneath trees in an old cemetery.
[ OLD CEMETERY / ARINA KRASNIKOVA / PEXELS ↗ ]

The Answer Comes First

Traditional search usually returns a ranked set of pages and leaves the choice to the person searching. AI search adds several steps after retrieval. Google describes query fan-out, where a model issues related searches in parallel, retrieves relevant sources, and uses them to generate an answer with supporting links.

QUESTIONRELATED SEARCHESFIND PAGESRETRIEVE PASSAGESWRITE ANSWERCITE

Search remains part of this process. Google says its existing technical requirements and quality signals apply to AI Overviews and AI Mode. OpenAI says any public site can appear in ChatGPT search and recommends allowing OAI-SearchBot so its responses can include summaries, snippets, citations, and links.

The additional selection happens inside the found page. A search service chooses the passage that supports its answer. That passage needs enough information to identify its subject, author, date, and source after it has been separated from the surrounding page.

[ FIELD NOTE / DISCOVERY ]

As AI changes how people find information, websites need to make their work easy to retrieve, identify, verify, and trace back to its source.

Give the Work an Address

The Markdown files behind Finding Direction already form a simple knowledge store. Calling it a store sounds grander than the implementation. It is a directory of source files with consistent fields and a build that knows what to do with them.

A knowledge store can also live in a content management application, relational database, or graph. Its value comes from keeping the work and the information that identifies it in one maintained source.

For an article, that source includes several kinds of information:

  • Identity: a stable ID, title, type, canonical URL, and any previous URLs.
  • Authorship: the person or organization responsible for the work.
  • Time: when the work was created, published, modified, or superseded.
  • Meaning: the summary, topics, sections, and relationships to other work.
  • Evidence: citations, image credits, source records, and links that support the claims.
  • State: whether the work is a draft, scheduled, published, private, archived, or removed.

The dates describe different events. Editing a typo today does not change when an article was first published. A policy may have a publication date, a later effective date, and another date when it was replaced. Keeping those events separate helps a search service select the current fact without rewriting the history around it.

A simplified record for this article looks like this:

id: article:is-seo-dead
type: Article
canonicalUrl: https://paulbout.in/thinking/is-seo-dead
title: Is SEO dead?
author: person:paul-boutin
createdAt: 2026-08-29
modifiedAt: 2026-08-29
status: draft
topics:
  - knowledge-systems
  - web-development
sources:
  - id: source:google-ai-search-guide
    supports: query-fan-out
relationships:
  - type: about
    target: concept:ai-search

The build uses these fields to create metadata, feeds, archive entries, related links, and machine-readable descriptions while the public page displays only what readers need.

[ FIELD NOTE / ADDRESS ]

When a passage leaves the page, its title, author, date, and source need to go with it.

Keep the Original

Search services build indexes from published pages. Internal AI applications often do the same with their own documents. They extract the text, divide it into chunks, create embeddings, and store those vectors with selected metadata.

SOURCEPUBLISHEXTRACTCHUNKEMBEDSEARCH INDEX

An embedding helps a service find passages with similar meaning. The accompanying metadata identifies what each passage describes. A chunk that contains “Coverage begins after 90 days” needs the product, jurisdiction, effective date, section heading, and source URL that make the sentence useful.

The index should be disposable. Rebuilding it from the maintained source preserves every fact, citation, relationship, and revision state. When an article changes, the publishing workflow regenerates the page and sends the updated source through indexing again.

Every derived chunk retains the record ID, canonical URL, title, section heading, author, version, modified date, and source references. Those fields let retrieval filter old material and return enough context for a person to inspect the answer.

The same structure helps with work beyond public search. A support agent can retrieve the current return policy and its effective date. A product comparison can connect a specification to the exact product version. A writing tool can distinguish my published position from a source I cited while developing it.

Publish From One Source

The source only helps discovery when the site publishes clear, accessible forms of it. Each form has a different job.

HTML carries the complete work. Put the meaningful content in the server response whenever possible. Use one descriptive h1, ordered section headings, ordinary links, useful image descriptions, and paragraphs that identify their subject. Keep navigation, cookie banners, and interface controls subordinate to the article in the document structure.

JSON-LD describes the work. JSON-LD is a W3C standard for expressing linked data in JSON. The Schema.org vocabulary provides common types for an article, person, organization, product, event, and other published material. For an article, fields such as headline, author, datePublished, dateModified, mainEntityOfPage, about, and citation connect the visible page to its identity and evidence. Google recommends JSON-LD for its supported structured-data features because it is easier to maintain at scale. Its generative-search guidance also makes clear that structured data is not a separate ranking shortcut for AI answers.

Sitemaps and feeds announce change. An XML sitemap lists the canonical URLs a crawler should discover and can include an accurate lastmod value. RSS or Atom publishes a chronological stream of new work. Sites with frequent changes can use IndexNow to notify participating search services when a URL has been added, updated, or deleted. The notification requests another crawl; the search service decides what to index.

APIs expose exact records. Documentation, catalogs, directories, and public datasets often benefit from a read-only JSON response with stable IDs, typed fields, timestamps, pagination, and explicit versions. That response contains the same facts as the public page. Generating it from the maintained source prevents a separate “AI version” from drifting away from what people read.

See What AI Finds

Crawler access determines which services retrieve the work. A broad rule for “AI bots” misses the differences between search, user-requested access, model training, and browser interaction.

OpenAI uses OAI-SearchBot for discovery in ChatGPT search and GPTBot for content that could be used in model training. Its publisher controls allow search discovery while declining training. Perplexity similarly documents PerplexityBot for its search index and a separate Perplexity-User client that fetches pages in response to a person’s request.

Express the intended access through robots.txt, page-level noindex directives, authentication, and firewall rules. Verify known crawlers with the IP ranges or reverse-DNS process published by the provider instead of trusting the user-agent name alone. Keep private and authenticated material outside every public path.

Treat llms.txt as an optional experiment. Google says it ignores llms.txt for Search, including its generative features. If another service uses the file, generate it from the same source as the rest of the site.

Retrieval needs its own tests. Start with a small collection of questions the site answers. For each one, record the page and passage that contain the evidence. Run those questions against the site’s search or vector index and check whether the correct passage ranks near the top with its title, date, author, and URL attached.

Public AI search changes as models, indexes, and competing sources change, so those results are observations rather than deterministic tests. Asking the same questions periodically reveals blocked crawlers, stale dates, duplicate URLs, weak passage identity, and old material outranking the current source.

Server logs show which approved crawlers reached each page and what response they received. OpenAI adds utm_source=chatgpt.com to referral links from ChatGPT. Google is rolling out a generative AI performance report in Search Console. These signals show different parts of discovery: crawling, citation, impression, and the visit that follows.

[ FIELD NOTE / MEASUREMENT ]

A citation confirms that a source was selected. Understanding the complete work still requires reading it.

Finding Direction holds the complete article: the argument, evidence, sources, and connections to other work. When AI retrieves one passage from this site, I want the reader to know who wrote it, when it was current, what supports it, and where to read the rest.