Not a pile of web pages.A set oflibraries.
Most of what matters isn't really a web page — it's a judgment, a standard, a clinical trial, a filing. askFinz keeps each of those as the thing it is, so an answer can point at the right one.
index/ ├── research/ papers · data ├── news/ thousands of outlets ├── law/ acts · judgments ├── standards/ the specifications ├── patents/ grants · applications ├── medicine/ trials · clinical ├── markets/ filings · earnings ├── reference/ encyclopedic ├── code/ repos · docs ├── books/ long-form · manuals ├── courses/ syllabuses ├── products/ matched · reviewed ├── jobs/ posts · tenders ├── places/ travel · property └── media/ film · music · video
Read from the running system.
What's held, and what it's for.
The index keeps 44 separate collections. They are grouped below by subject — a few related collections sit under one heading here, because “papers, preprints and the datasets behind them” is more useful to read than three entries that mean the same thing to you.
Examples are illustrative — written to show what each collection is for, not generated live. They describe how results behave rather than asserting any specific finding, case or figure.
We don't scrape a worse copy of something published properly.
A patent is not a web page. Neither is a standard, a clinical trial or a court judgment. Where an organisation publishes its material properly, we read it from them — with the identifiers, dates and status that make it citable — instead of collecting whatever version happens to be lying around on the open web.
Specifications
Read from the standards bodies themselves, so a requirement can be quoted with the document and revision it came from.
Patents & trademarks
Taken from the offices that grant them, keeping the filing office, dates and status rather than a summary of a summary.
Clinical research
Registered studies with their sponsor, phase and status — the trial itself, not an article written about it.
Reading from the source is how these collections get as complete as they are — and it is the part a crawler alone cannot reproduce, no matter how much of the web it reads.
Quality is decided on the way in — and revisited afterwards.
Reading widely is easy; reading widely without filling up on junk is the hard part. Pages are reviewed as they arrive, anything borderline gets a second closer read, and what fails is rejected rather than kept and filtered later. These are live figures.
It also improves what is already held. As the software that reads and files pages gets better, it goes back over material indexed months ago and re-checks it — so a page read early benefits from every improvement made since, without anyone asking for it and without the page being fetched again. The index gets better with age rather than staler.
Read live from the running system. Shown as rates rather than totals — the index is still growing, and a rate tells you how it is built rather than how far along it is.
Fetched once, when it was first read.
reading…
- Read and storedthe page enters the index
- Cleaner textmenus, banners and filler stripped more accurately
- Structure kepttables and headings survive as tables and headings
- Filed by kindrecognised as the sort of thing it is, not a loose page
Nothing above sends a single new request to the publisher. Every improvement is made against the copy we already hold — which is why a page read early keeps getting better, and why the index improves with age rather than going stale.
Filed by what it is
A court judgment is kept as a judgment, with its jurisdiction and the cases around it — not as a page that happens to contain legal words. That is what lets an answer cite the right thing.
Read from the source
Where an organisation publishes its own material properly, we read it from there rather than scraping a worse copy off the open web.
Growing where it's used
Collections expand toward what people actually ask about. Because the index is ours, adding a new one is a decision we can simply make.
The shelves grow where people look.
Point us at yours.
Collections expand toward what readers actually ask about, and because the index is ours, adding to one is a decision we can simply make. If a source you rely on isn't here, tell us — and if the site is yours, you can have it read directly with priority rather than waiting your turn.