mudpie

original analysis · 4 min read

All 70 YC homepages linking llms.txt had meta descriptions

All 70 homepages advertising llms.txt also expose meta descriptions; 67 have canonical links. A fresh, uncapped comparison with ordinary metadata and on-origin versus offsite discovery.

Published · Updated

Every one of the 70 YC homepages advertising llms.txt in this scan also had a meta description. Sixty-seven had a canonical link, 68 had an H1 or H2, and 63 contained parseable JSON-LD with a type.

That is the most useful pattern I found in the agent-discovery data: these homepages expose the extra link alongside familiar HTML metadata.

I collected fresh homepage responses for the public website listings across nine YC cohorts. Of 1,586 company rows, 1,515 yielded parsed HTML. On that common denominator, descriptions appeared on 95.2% of homepages, while advertised llms.txt links appeared on 4.6% and Markdown alternates on 2.6%.

Signal in the homepage response Companies Share of 1,515
Nonempty meta description 1,443 95.2%
Nonempty H1 or H2 1,316 86.9%
Canonical link 1,003 66.2%
Parsed JSON-LD with a type 767 50.6%
Advertised llms.txt link 70 4.6%
Advertised Markdown alternate 40 2.6%

The Markdown subset tells a similar story. All 40 homepages advertising an alternate also had descriptions and headings. Thirty-nine had canonical links; 38 had typed JSON-LD. These counts describe co-occurrence in the same response. They cannot tell us which feature came first or whether one caused the other.

This is a different question from whether a file exists. Installmap’s September census requests the file directly and reports a different, broader YC population. Our scan asks whether the homepage advertises a link. A client can also discover a file by convention or another page. Neither our 4.6% nor a file-presence percentage measures actual agent use, and subtracting the two studies would not establish a discovery gap.

Looking at company pages makes the arrangement easier to understand. OpenTag’s homepage, linked from its YC profile, advertised pricing, docs, and llms.txt, alongside all four metadata signals. StarSling’s homepage, from its YC entry, advertised a Markdown alternate and two llms.txt destinations: one on the main origin and one on its docs subdomain.

That second example matters. A public reference collection can span multiple hosts even when discovery begins at one homepage. Across this scan, 68 companies advertised a same-origin llms.txt link and five advertised an other-origin one; three appeared in both groups. All 40 Markdown-alternate companies advertised a same-origin alternate.

The geography of ordinary docs links was different. There were 390 docs/developer keyword matches across all origins, but only 188 with a same-origin match. For that surface, a same-origin filter drops 202 companies. Applying one boundary to every reference type would hide how differently the links are distributed.

My initial question was how often agent-specific discovery appeared compared with ordinary SEO basics. In these homepage responses, it is a small additional layer, and the companies advertising it frequently expose the familiar metadata too. That is a more useful starting point than assuming a new file makes the rest of the page irrelevant.

For a founder, I would use the examples as a publishing checklist: make the homepage identify the product, give buyers clear reference destinations, and make any alternate point to material you maintain. The data supports checking what is advertised and how it fits together. It does not measure how well those references answer a buyer’s question.

For someone building an evaluator, keep each observation separate. A canonical link, typed JSON-LD, an advertised alternate, and a successful retrieval are different pieces of evidence. Combining them into one readiness number would hide the very pattern that makes this sample interesting.

Methods: Fresh credential-free homepage GETs ran September 19, 2026, 16:14–16:20 UTC, using robots policy and concurrency three. All 1,586 rows remain accounted for: 1,584 valid public hosts, one missing URL, and one single-label host. HTML signals use 1,515 parsed responses; 54 no-status rows plus 17 non-2xx rows remain unknown. We counted uncapped eligible anchor/link matches without following their destinations. Published URLs omit query/hash and private or credential targets. Metadata is raw HTTP HTML, not rendered visibility. Unadvertised files, runtime WebMCP, agent use, task completion, and ranking effects were not measured. This is a separate observation from the earlier HTTP pass.

Sources: YC company directory, Spring 2025 roster, Summer 2026 roster, the nine-cohort follow-up dataset, and the parser, guarded collector and recount methods. Raw historical pages are not distributed; the methods distinguish recounting retained observations from making a new collection.

About the author

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.

First1000 ↗ · X ↗