Methodology
How do you know?
This is the machinery behind every Red Line Review verdict. Every determined small-site application across all 33 London boroughs, over a rolling three-year window, coded into a 10-category refusal taxonomy and back-tested before publication — and published in full, because a judgement is only worth something if you can check the evidence under it.
Dataset scope
The London Small Sites Planning Dataset covers every small-site planning application across all 33 London boroughs, 2022–2026, a rolling window that captures the current policy equilibrium. The overall refusal rate is 43.6%, ranging from 15.2% refused in the most permissive borough (Kensington & Chelsea) to 70.5% in the strictest (Havering), a 55-point spread in raw refusal rates. The mix-adjusted spread, holding site type, transport access, conservation status and unit count constant, is narrower at around 43 points.
Every application is classified along nine axes: decision outcome, site type (conversion, demolish-rebuild, end-terrace, mid-terrace, backland, infill, extension), area within the borough, conservation area status, PTAL accessibility level, density, determination time, decision route (committee vs delegated), and case officer.
Sources and ingestion
Primary feed: the Mayor of London’s Planning London Datahub (PLD), queried via its Elasticsearch guest endpoint. Council back-office systems push application validations, decisions, and metadata to the PLD on a daily cadence under the London Planning Data Standard. Perfect Scale’s ingest pipeline pulls each borough by canonical local-planning-authority name, paginating through the full validation history within the analysis window.
The PLD feed is then augmented by direct harvest of each council’s public planning register for Decision Notices, Officer Reports and Committee minutes. This requires ten different scrapers because London boroughs use ten different portal systems (Idox, Northgate, Council Direct, Arcus/Salesforce, Agile/IEG4, Tascomi/PlaceHub, RBKC Atlas, NECSWS SPA, Aurora and one bespoke build). Each scraper is portal-aware, rate-limited, and resumable.
NLP extraction and the refusal-reason taxonomy
Decision Notices and Officer Reports are parsed for structured refusal reasons. Each numbered reason is extracted as raw text, then classified into a 10-category taxonomy currently in production: Design & Character (DES), Neighbour Amenity (AMN), Daylight/Sunlight (DLT), Space Standards (SPC), Heritage & Conservation (HER), Transport & Parking (TRN), Policy Non-compliance (POL), Infrastructure & Sustainability (INF), Flood Risk (FLD), and Other (OTH). Three further categories are being introduced in the next quarterly refresh: Insufficient Information (INS), Permitted Development Non-compliance (PDD), and Loss of Use / Community (LUC). They disaggregate patterns previously bundled under OTH. Categories are not mutually exclusive: a single reason can match both Design and Amenity, which is the most common pairing across the dataset.
Each reason carries a primary and secondary code, the source document, and the original text so any classification can be audited back to the council’s own words. Total reasons classified to date: over 18,000 across more than 10,000 source documents.
Decision routes and officer overturns
Every decision is tagged with its route, committee or delegated, and where an officer report exists, the officer’s recommendation is captured separately from the committee’s outcome. This lets the dataset surface overturn rates: how often a committee diverges from the officer it asked to write the report, broken down by borough, site type and area. The political-noise signal is one of the few in planning data that requires NLP-extracted document reading rather than transactional records alone, and it is also one of the most asked-about findings by developers preparing for a marginal application.
Modelling and back-testing
Two measures sit on the dataset. The first is a density model that estimates the approvable unit count for a given site type, area, and conservation context. Across the borough dashboards, in back-tests against decided schemes the density estimate lands within ±1 unit of the actual approved scheme 73% of the time, and within ±2 units 90% of the time, with no systematic bias in either direction. Every borough dashboard publishes the model’s back-test before going live.
The second is a cohort benchmark that reads a scheme against the empirical approval record for its area, site type, density, conservation status and PTAL band. It is surfaced only as a descriptive comparison (“schemes in this cohort have approved at 41%, against a borough average of 56%”) — never as a probability or point estimate for an individual application. Historical rates are history; the possibility of policy change is stated, not buried
Back-test sample sizes and confidence intervals for both measures are documented in the underlying borough analyses. A consolidated back-test summary, with per-cell n and CI, is being added to this page in the next quarterly refresh.
Evidence tiers
Every finding in every Perfect Scale output carries an evidence tier: Robust, Indicative, Suggestive, or Anecdotal. The tier reflects sample size, the appropriateness of any statistical test (Mann-Whitney U, chi-squared, Spearman correlation), and whether the cohort meets a minimum-n gate. Robust cells hold up 37% better than Anecdotal cells in the dashboard back-tests. The four tiers exist so readers know how much weight to put on each claim. A number drawn from 200 applications is not the same as one drawn from a handful.
How we stay honest
An independent measure is only worth something if it can be checked, and if it admits when it is wrong. So the method is published in full, every finding is graded by how much evidence sits behind it, and where the data is thin or a claim is weak we say so plainly rather than dress it up. Every number traces back to the councils’ own decisions; nothing rests on a figure you cannot follow to its source. When we test the models, we report what holds up and what does not. We would always rather state a limit than overstate a finding. That is the whole basis on which an evidence standard earns trust from every side.
Limits and cadence
The dataset is London-only by design, focused on schemes of 1–9 units. Conditions, officer intelligence, and refusal-reason extraction depend on harvested Decision Notices and Officer Reports. Coverage varies between boroughs because portals vary in accessibility (Westminster currently caps harvest at ~23% on its POST-blocked Idox; most boroughs sit at 70%+). Where coverage is below 50%, the relevant findings are marked Indicative at best and explicitly footnoted in every report.
Refresh cadence is quarterly: each borough is re-pulled from the PLD, re-harvested for any new refused applications, NLP-classified, and the workbook regenerated. The full pipeline is documented and resumable; any stage can be re-run independently. Borough-level data-window-end dates are published on every dashboard footer.
See the data for yourself.
Every London borough has a free public dashboard: approval rates, refusal patterns and quarterly trend, no sign-up. And when you’re weighing a specific site, a Red Line Review puts this evidence behind one verdict.