Methodology

How the catalogue is made

Every document here was published by a public body on its own website. This page says how we find it, what we keep, how we read it and how often each part is refreshed.

Where a document comes from

From the body’s own website: the pages where it publishes its board and committee papers, its governance documents and its registers. We find those pages by reading the site’s own sitemap and navigation rather than guessing at addresses, and we stay on the body’s own site.

Official statistics come from the publisher of the statistic, NHS England, and regulator ratings from the Care Quality Commission, not from anything a body says about itself.

How it is collected and archived

When a new document appears we fetch the file itself and store the bytes in our own archive, with a fingerprint of the file, the address we took it from and the day it arrived. A link on its own is not counted as collected.

We keep every original. Bodies reorganise their websites and old papers disappear; ours stay, and the record still links to the body’s address so the provenance is never rewritten. Where a body removed a paper before we collected it, we look for it in the UK Government Web Archive, run by The National Archives, and then the Internet Archive, and record which archive our copy came from.

Some sites refuse automated requests. We retry through an ordinary browser, and where a site still refuses we ask the body for access rather than disguise ourselves. Every document a collector cannot get opens a ticket that a person reviews.

How it is read

In full. The text of the whole document is extracted, and scanned documents are read by optical character recognition. A board pack is split into the separate papers it contains, and each paper is categorised from what its body says, not from its title, because titles are the part of a PDF most often mangled.

Every value we take from a paper keeps its place in the document and the sentence around it, so what you are shown is the passage itself rather than a claim about it. We do not write summaries into the record.

Dates are British and read from the document: the date shown for a paper is the meeting it went to, not the day we happened to collect it.

How often each part refreshes

Board papers and other documents
Every day. Each body's pages are checked for new documents, new ones are collected and read, and papers a body is due to publish ahead of a meeting are looked for first.
Who sits on each board
Every day, checked against the body's own website.
NHS performance statistics
Checked every day against NHS England's statistics calendar and loaded on the day each release is published: A&E, referral to treatment, cancer and diagnostic waiting times, and the NHS Oversight Framework.
Care Quality Commission ratings
Every week.
This site
Each day's collection is checked on a separate copy of the database and released here, normally the same day.

What a gap means

When a profile holds fewer papers than you expected, or a document you know exists is not listed, the gap is ours until it is shown to be the body’s. Most gaps are pages we have not yet reached, a site that refused us, or a document published somewhere we did not look.

Every failure to collect is recorded as a ticket for a person to review, and when one is answered the fix goes back into the collector and is re-run across every body, not just the one that raised it. If you know where a missing document is, tell us.

Where to read more

What is held today is on the about page. What you may do with it is in the data licence.