Sign in

Your website, as clean text a language model can actually read.

Point this at a URL and get every page back as Markdown, in one ZIP — main content only, no navigation, no cookie banners, no markup burning tokens without carrying meaning.

Sign in with Google. Nothing to install, nothing to configure.

Main content, not the whole page

Each page runs through a content filter that drops navigation and repeated chrome. That's the difference between a corpus you can search and one where every document matches the phrase “Book Now”.

One file per page, provenance intact

Every file opens with the URL it came from, so nothing loses its source on the way into a retrieval pipeline.

The token arithmetic is the argument

English runs near 1.3 tokens per word. A 662,000-word site is about 880,000 tokens of clean text — as raw HTML it would be several times that, nearly all of it markup.

Ready for whatever comes next

Ground a support assistant in your own answers, audit content at a scale nobody can read, brief translators with real terminology, or move platforms without losing the writing.

Every crawl counts the words as it goes, so you get the size of the site alongside the content — total, per folder and per page. Useful the day somebody asks what it would cost to translate.