The open web, read page by page
Ordinary web pages, read and kept as pages. The collection everything else is filed out of — where a page belongs to no speciality shelf, it belongs here.
- what kind of page declared
- 6
- per record
- 19.9 KB
How it divides
1 dimension · press a band to read it- General28,358,967 · 60.3%
- Technical6,590,523 · 14.0%
- Academic5,329,289 · 11.3%
- Business2,066,021 · 4.4%
- Government1,866,861 · 4.0%
- Remaining 3 values2,852,899 · 6.1%
28,358,967 passages are filed under General — 60.3% of those that declare it.
47,064,560 of 47,858,601 declared · 98.3%. The rest carry no value for this field, so these shares are of the declared set and do not add up to the collection.
What you can do with it
Ask the web a question, not a search engine — This is read into our own index rather than fetched from somebody else's results, so a question can be asked of the page itself.
Documents are read too — PDFs, Office files and slide decks are converted and read page by page, rather than returned as a filename to guess from.
Follow it back — Every page keeps the address it came from, so an answer can be checked where it was published.
Where it stops
It is not the whole web, and no index is. It holds what has been read so far.
A page behind a login or a paywall is not in here. Where a site refuses the crawler, that refusal is respected rather than worked around.
Freshness varies by page. Something that changes hourly and something that changed once in 2019 are not re-read on the same schedule.