Skip to main content
Private access is opening up — request an invite
askFinz
Collections

Not a pile of web pages.A set oflibraries.

Most of what matters isn't really a web page — it's a judgment, a standard, a clinical trial, a filing. askFinz keeps each of those as the thing it is, so an answer can point at the right one.

5.2 KB
Per stored page
everything needed to find and rank it
1.7
Passages per page
each one separately searchable
48
Collections
each filed as the kind of thing it is
44 collectionsindex/
index/
 ├── research/     papers · data
 ├── news/         thousands of outlets
 ├── law/          acts · judgments
 ├── standards/    the specifications
 ├── patents/      grants · applications
 ├── medicine/     trials · clinical
 ├── markets/      filings · earnings
 ├── reference/    encyclopedic
 ├── code/         repos · docs
 ├── books/        long-form · manuals
 ├── courses/      syllabuses
 ├── products/     matched · reviewed
 ├── jobs/         posts · tenders
 ├── places/       travel · property
 └── media/        film · music · video
filed by what it is20 stores
Right now

Read from the running system.

Connecting…
368
Document formats
PDF, Office and more, converted to readable text
Full text
How documents are read
page by page, not skimmed or linked to
Open access
Scholarly sources
read from the publishers and repositories themselves
Continuous
Refresh
read as it changes, not rebuilt on a schedule
What's held

What's held, and what it's for.

The index keeps 44 separate collections. They are grouped below by subject — a few related collections sit under one heading here, because “papers, preprints and the datasets behind them” is more useful to read than three entries that mean the same thing to you.

Examples are illustrative — written to show what each collection is for, not generated live. They describe how results behave rather than asserting any specific finding, case or figure.

Where it comes from

We don't scrape a worse copy of something published properly.

A patent is not a web page. Neither is a standard, a clinical trial or a court judgment. Where an organisation publishes its material properly, we read it from them — with the identifiers, dates and status that make it citable — instead of collecting whatever version happens to be lying around on the open web.

Specifications

Read from the standards bodies themselves, so a requirement can be quoted with the document and revision it came from.

Patents & trademarks

Taken from the offices that grant them, keeping the filing office, dates and status rather than a summary of a summary.

Clinical research

Registered studies with their sponsor, phase and status — the trial itself, not an article written about it.

Reading from the source is how these collections get as complete as they are — and it is the part a crawler alone cannot reproduce, no matter how much of the web it reads.

Quality is decided on the way in — and revisited afterwards.

Reading widely is easy; reading widely without filling up on junk is the hard part. Pages are reviewed as they arrive, anything borderline gets a second closer read, and what fails is rejected rather than kept and filtered later. These are live figures.

It also improves what is already held. As the software that reads and files pages gets better, it goes back over material indexed months ago and re-checks it — so a page read early benefits from every improvement made since, without anyone asking for it and without the page being fetched again. The index gets better with age rather than staler.

137
Pages per domain
we read a site through, not just its front page
20%
Added to the skip list
of pages audited, the share confirmed as junk and never retried
77%
Junk caught on deep check
of the borderline pages sent for a second, closer read
Every page
Checked before it is kept
quality is decided on the way in, not on the way out

Read live from the running system. Shown as rates rather than totals — the index is still growing, and a rate tells you how it is built rather than how far along it is.

One page, read onceimproved in place
The publisher's page

Fetched once, when it was first read.

reading…

Re-checked from our own copy, whenever the software improves
  1. Read and storedthe page enters the index
  2. Cleaner textmenus, banners and filler stripped more accurately
  3. Structure kepttables and headings survive as tables and headings
  4. Filed by kindrecognised as the sort of thing it is, not a loose page

Nothing above sends a single new request to the publisher. Every improvement is made against the copy we already hold — which is why a page read early keeps getting better, and why the index improves with age rather than going stale.

Why it's built this way

Filed by what it is

A court judgment is kept as a judgment, with its jurisdiction and the cases around it — not as a page that happens to contain legal words. That is what lets an answer cite the right thing.

Read from the source

Where an organisation publishes its own material properly, we read it from there rather than scraping a worse copy off the open web.

Growing where it's used

Collections expand toward what people actually ask about. Because the index is ours, adding a new one is a decision we can simply make.

Something missing?

The shelves grow where people look.
Point us at yours.

Collections expand toward what readers actually ask about, and because the index is ours, adding to one is a decision we can simply make. If a source you rely on isn't here, tell us — and if the site is yours, you can have it read directly with priority rather than waiting your turn.