Why we need better definitions for SDSs and Local Plan evidence, and how these definitions can help us digitise from the outset.
Every plan-maker preparing a Spatial Development Strategy over the next three years will run into the same question within the first few weeks of starting: what does a proportionate evidence base for an SDS actually look like? They’ll very quickly run into a secondary question of equal importance: how should the SDS evidence relate to the Local Plans sitting beneath it? These are essential, foundational questions that have yet to be answered fully.
The draft NPPF is clear that evidence should be relevant, proportionate and up to date, and that plan-makers should draw on existing evidence and work jointly before commissioning anything new. What the draft NPPF does not do is specify what an evidence base should contain, how it should be structured, and what specific parts should be shared between SDSs and their constituent Local Plans. That gap is where confusion, duplication, delay and, potentially, future inconsistency creep in. That’s why we’re working hard as part of PropTech 6 to address this specific challenge.
Our thinking rests on two simple ideas. First, that we need clearer, shared definitions that help us understand the different types of data that feed an evidence base. Second, that both SDSs and Local Plans organise their evidence around the same set of standardised themes. This blog explores the first of these two ideas, showing how improved definitions can help us understand how we can move towards digital, and more automated, evidence.
An Essential Definition: Raw Data vs Derived Evidence
In our view much of the confusion around evidence bases comes from limited terminology and joint understanding of what we mean by terms like “evidence” or “proportionate” in the first place. Let’s take “evidence”: invariably we treat “evidence” as a single undifferentiated category, usually comprising a collection of PDF reports. However, in practice, an evidence base is built from a combination of lots of different types of data. Some of this data is easy to obtain, while some of it is highly processed, with lots of professional judgement applied. Some evidence has historically been consolidated in PDF reports, while other evidence lends itself more immediately to being digital and interactive. These are important distinctions that we think it is useful to have definitions for.
At a minimum we see the following different types of data all feeding into a plan:
Raw Data
Raw Data is information in its rawest usable form, drawn from a defined source, with no interpretive process applied. This type of evidence can be further sub-divided into additional types:
- National datasets: data published by national bodies (ONS, OS, MHCLG, DfT, Environment Agency, Met Office/UKCP18) that are consistent in structure and coverage across the country. Again, it’s important to state that this should be raw data so, as an example, the DfT’s Connectivity Tool, despite being national, wouldn’t fit into this category (because it is derived output using lots of other raw data).
- Local datasets: genuinely novel data held within a particular local authority or partner system, such as planning application registers or local housing land availability data. Again, however this is data in its rawest form. The important point about this category is that the data is in a local system that is collected based on repeatable business as usual processes. It is not a specific one-off commissioned piece of data – that’s the category below.
- Primary data: Primary data is data generated for a specific purpose (often a study), such as a commissioned household survey, a traffic count, or a stakeholder consultation exercise. This might be necessary data collected locally to enhance the basic picture. Again, because we’re looking at raw data here, this is data in its unprocessed form. The important point about this type of data is that is often a one-off commission, not a repeatable extract from a system.
Derived Evidence
Now we get to one of the most important distinctions. Derived Evidence is the name we give to processed data. This naturally includes the big PDF reports – what people most commonly think of when they talk about “the evidence base” in a plan-making context. These are usually sizable studies that get commissioned that consultants deliver. Importantly, Derived Evidence is not raw data. Most commonly it combines different sets of raw data and applies heavy processing in order to reach a new output.
Whilst the exact process for each piece of Derived Evidence is slightly different, the overall structure of the process is broadly the same. Each piece of Derived Evidence takes Raw Data (typically a combination of national, local and primary sources) and applies a defined analytical method (a methodology, a set of assumptions) to produce new outputs: datasets, zones, maps and recommendations that didn’t exist in the source data itself. A Housing Market Area assessment, a Functional Economic Area study and a Strategic Flood Risk Assessment are all examples of Derived Evidence, each taking multiple raw inputs through a process to reach a defensible output. As stated above, some national bodies also produce Derived Evidence – the DfT’s Connectivity Tool or DESNZ’s Heat Network Zoning model are all examples.
Consequences
The lack of shared language or understanding of how evidence bases are built up can quickly lead to a range of issues. Some of the commonly observed challenges are set out below (this is not an exhaustive list, but it covers some of the patterns we see most often).
“Local” data maintenance
In our experience, when we’ve conducted data audits with various authorities, we’re almost always told “we have a lot of local data.” Once the audit is complete, however, it’s very common to find that 80–90% of it is, in its rawest form, national data: an ONS table, an OS layer or an Environment Agency dataset that’s been downloaded, reshaped and absorbed into local systems over time. That’s usually a real surprise to the client and one that has important efficiency implications. Common efficiency implications include:
- We often find that local processes are being required to enable councils to work with national datasets. Often the reason for this is the complexity of handling large data, which necessitates local slicing and dicing that can often absorb a significant amount of resource and time.
- Another common issue is the local processing or maintenance of Derived Evidence. This could simply be managing layers that have been produced through consultancy processes, while in the most extreme cases, this work might include extracting zones, boundaries or figures from PDF reports (because the Derived Evidence was never commissioned with digital outputs in mind).
Recommissioned rather than refreshed
When Derived Evidence isn’t built with a preserved method and data trail, it is much harder to update. Often to facilitate an update, local authorities are pushed to start again from scratch. If a dataset changes, if a plan comes up for review, or if a neighbouring authority needs the same analysis for their own area, rather than re-running an established process against new inputs, the whole piece of evidence is recommissioned, often at full cost. Consultants are naturally happy to oblige. Genuine transferability would mean the method and the associated code are delivered in an open format, but again this requirement would need to be factored in at the commissioning stage.
Loss of primary data
Surveys, counts or consultation exercises are commonly commissioned specifically to inform a piece of Derived Evidence. In practice, this data is sometimes not retained separately and is instead folded straight into the analysis, reported on only in its processed form. The raw responses, counts or survey returns may never be handed over as a standalone asset. That means a council that commissioned and paid for the primary data collection can’t return to it later to ask a new question, validate a specific conclusion, or feed it into a different piece of analysis.
The Benefits of Definitions
Taken together, for us these three patterns point to a common underlying solution – if we can create clearer definitions of what’s raw and what’s derived we can ensure that:
- raw data is preserved and shared as a standalone asset, and
- that derived evidence processes are made more transparent and transferable.
If we can systematically do this, it will make it much easier to share evidence between tiers of authority within the same geography, but also across geographies. Therefore, we think getting the definitions right, and making sure they are clearly understood in the commissioning process, is a useful first step to helping make evidence gathering more efficient overall.
Why This Has to Be Digital First
Beyond the definitions, if Derived Evidence sits as static PDFs, and raw data sits as spreadsheets or sliced national datasets held within different team folders (or even different organisational GIS systems) there is no easy way for an SDS team to consolidate a clear understanding of the current state of evidence without significant manual effort. It becomes even more challenging for anyone to demonstrate a clean audit trail from the raw data to a specific decision. As SDSs and Local Plan evidence becomes populated in parallel, it becomes even more important to maintain clear lines of sight (ideally search, discoverability and interoperability) across the collective evidence.
A commitment to a digital-first evidence base is a foundational decision to help move towards a more manageable situation.
In practice, getting to a fully digital evidence base is likely to be a journey rather than an immediate destination for most authorities. However, the following key steps are likely to go a long way to helping to realise digital ambitions over time:
- Core data, held once: As we’ve seen, what’s described as local data is often 80–90% derived from national sources. There’s no good reason for every authority to separately download, reshape and maintain its own copy of national datasets other than the historic complexity in handling data of this size. Ideally, this core data would exist as a single source that all stakeholders can access across organisations and systems. National data is freely available in the first place, and keeping it current can nowadays be fully automated. Dealing with the data size can be easily addressed with modern web-based systems. With the need to maintain local versions of national data removed, efficiency and visibility benefits for everyone working on the plan can be quickly realised, which can underpin the business case for a digital-first approach.
- True local data, properly categorised: Once the national datasets are stripped away, it becomes much easier to see the genuinely valuable data held within local systems. Data most relevant to the plan can then be prioritised. Extraction from systems can largely be automated or streamlined, with many suppliers providing APIs or standard extract formats. Some more complex data remains in this category, for example up-to-date consolidation or housing and infrastructure. But consolidation of these data can be clearly prioritised for digitisation in a joined-up way using SDS-level governance to ensure standard formats, interoperability and visibility so that benefits are seen by all stakeholders in the plan area.
- Replicable processes, delivered as code: The final step to full digitisation then needs to progress from raw data alone to Derived Evidence. This is likely to be the hardest category to move into the digital sphere, however there is lots of precedent regarding how to do this. Universities in particular have spent the last decade being pushed toward reproducible research, with methods, data and code expected to be published alongside conclusions. Consultants have yet faced no equivalent expectations. However, if Derived Evidence is to be genuinely transferable rather than a one-off report, the process that produced it needs to be understood well enough to be shared, inspected and re-run. This is largely a knowledge and commissioning issue, but could become a real possibility once the first two steps are addressed.
Getting the definitions right won’t solve the consolidation of proportionate evidence on its own, but it will help to remove the confusion that currently sits behind many other problems such as duplication, cost or lack of interoperability. Once a team can say with precision whether something is raw data (national, local or primary) or derived evidence, it starts to become much clearer how they should go about gathering and maintaining that data. Only by treating each bit of data in the most appropriate way for its individual category, can we enhance the efficiency (and reduce future costs) of data consolidation.
But these categories are also essential for helping to understand proportionality. National data might get you 80% of the way towards exactly what you need. This in itself might be perfectly proportionate in certain cases. If you then do need to go beyond this, a specific bit of local system data might achieve the next 10%. Finally, if you do need to run a methodology to deliver the full piece of Derived Evidence, wouldn’t it be great if the “recipe” also existed for this in a way that could easily be re-used?
Whilst we see better definitions of data as a necessary first step to understanding the different parts of the evidence base, definitions are only half the picture. In our next blog we’ll turn to the second idea underpinning our PropTech 6 work: a standardised set of themes that both SDSs and Local Plans can use to organise their evidence.
If you’re developing an SDS or Local Plan evidence base and want to talk through how these ideas might apply to your own evidence gathering, we’d like to hear from you. Get in touch at info@cityscience.com to discuss our PropTech 6 work and how a digital-first approach could shape your evidence base.
