Analytic Methods Policy
1: Purpose
This policy outlines the requirement for all research conducted within the OpenSAFELY Service to comply with the permitted analytical methods of NHS England. It contains the following sections for reference:
- Purpose
- Scope
- Background
- Policy requirements
- How this policy will be updated
2: Scope
This policy outlines the analytic methods that are currently supported and not supported within OpenSAFELY due to resource for compute and output checking. The policy is approved by NHS England, and managed in collaboration with the OpenSAFELY team.
3: Background
The development of this policy takes into account a variety of constraints, such as: the available and funded OpenSAFELY compute; the capacity within the output-checking service, which oversees statistical disclosure controls (SDC), to satisfy contractual, legal and ethical obligations; the unknown privacy risks associated with some aspects of some analysis methods (eg generative and risk prediction models).
In addition, we must ensure we bring the public, profession and other stakeholders with us if and when there are analytic approaches that have, or are perceived to have, additional challenges around interpretability, explainability, safety, efficacy, cost-effectiveness1 and privacy (such as machine learning / AI models) are being run against GP and NHS data; maintaining the trust of the public and profession is a key priority for teams running the OpenSAFELY service.
4: Policy requirements
The OpenSAFELY output-checking service supports most established analytical methods. The following table is an indicative and non-exhaustive list:
| SDC-supported methods |
| Descriptive Statistics (counts, proportions, means, quantiles, etc.) |
| Basic Statistical Tests (e.g., t-tests, Chi-squared tests) |
| Generalized Linear Models (GLM), including linear, logistic and Poisson regression |
| Survival analyses (e.g., Kaplan-Meier estimates, Cox regression, parametric survival modelling) |
| Traditional Time Series Analysis (e.g., ARIMA, Exponential Smoothing, STL) |
| Visualisation of non-patient-level summary statistics (e.g., histograms, time-series, forest plots) and data smoothing for visualisation. |
| Structural Equation Models |
The following table lists the analytic methods currently not-supported by the OpenSAFELY output-checking service which oversees statistical disclosure controls (SDC).
Your OpenSAFELY study must not use any libraries, scripts, or run any code that uses these analytic methods:
| Not SDC-supported methods |
| Neural Networks (including deep NNs, CNNs, RNNs, ANNs) |
| Support Vector Machines (SVM) |
| Random Forests |
| Gradient Boosting Machines |
| Unsupervised Clustering Algorithms (e.g., k-means, PCA) |
| Advanced Time Series Models (e.g., LSTM, GRU, TCN) |
| Natural Language Processing Models (e.g., BERT, GPT) |
| Reinforcement Learning Models |
| Adversarial Learning Models |
| Large Language Models |
| Generative AI Models |
Your responsibility as OpenSAFELY users (researchers, analysts, data scientists)
- Study lead(s) must familiarise themself with the not SDC-supported methods.
- The study lead(s) must ensure that outputs used and requested for release comply with this policy.
Important notes:
- Users are not permitted to run methods on the Non SDC-Supported list even if they are not planning to release the model for external querying or inspection, without exceptional prior permission (see below).
- It might be acceptable for a Non SDC-Supported method to be used for an internal project analysis need (for example, using such models for propensity matching, inverse propensity weighting) with no requirement to release the outputs, model or performance statistics (see Appendix A for exception process).
- Currently, we do not support the release of models using Non SDC-Supported methods because (not exhaustive list):
- It is extremely labour intensive to evaluate large volumes of potentially disclosive data for release outside a TRE/SDE; and for some types of model that might be proposed for export there are no currently recognised means to evaluate disclosivity risk in any setting.
- Development of models that cannot be used as intended (e.g., for implementing risk prediction tools outside of OpenSAFELY) will utilise compute and resources unnecessarily.
- If you are planning to conduct analyses to produce a risk prediction model using the SDC-Supported methods, you must detail this on your application form “Study Information” section, so that NHS England and others can consider any regulatory or liability issues; in particular you must specify if you intend to implement the model in clinical practice. However, any risk prediction model files generated (using the SDC-Supported methods) may still not be approved for release if the model files are unreasonably large such that output-checkers cannot effectively assess for disclosivity risks.
- It is the responsibility of users to ensure outputs requested for release are clearly explained to output-checkers; only the minimal required; clearly in line with the project purpose; in line with this policy.
5: How this policy will be updated
As OpenSAFELY’s features and resources change over time, it is possible that various methods on the not SDC-supported list will be supported. To help us prioritise how the platform develops, email us (team@opensafely.org, subject: analytic methods suggestions) and share the specific method(s) of interest you would like us to support, including:
- Why is the method necessary for a future your project?
- Can the analysis be conducted using supported methods?
- What is the intended benefit of using the method?
- What have you considered to be the risks of using the method?
- How do you intend to mitigate these risks? Specifically, cite how you think the statistical disclosure control process could work.
We encourage users to also talk with their co-pilots as part of this feedback.
Requests for non SDC-supported methods will be logged and reviewed on a quarterly basis by NHS England, the OpenSAFELY team and shared with our governance board to determine how this policy will evolve over time2.
Appendix 1
Prior approval must be obtained before running a non SDC-Supported method for your project. The request should describe:
- The Non SDC-Supported method you wish to use.
- Why is this method necessary and/or beneficial for your project, and why is an SDC-supported method not sufficient?
- If approved to use the method, what outputs (if any) do you intend to request for release?
- the fitted model;
- outputs derived from the method (e.g. performance metrics, calibration measures, etc.);
- downstream analyses that use the method internally.
- Describe any privacy or disclosure risks that you have identified with the outputs you intend to request for release.
- How could these risks be mitigated? Specifically, cite how you think the statistical disclosure control process could work.
- Will you use this work to develop, validate or implement a risk prediction model for use outside OpenSAFELY (for example in clinical practice)?
If you know which non SDC-Supported methods you would like to use at the application stage, please answer the above questions in your application, otherwise e-mail team@opensafely.org with your request, detailing responses to the above questions were appropriate.
The OpenSAFELY team will carry out a focused assessment on whether the requested method can be practically supported within the service, taking account of compute requirements, operational feasibility and whether the requested outputs can be assessed through the output-checking service.
Your full request, together with the OpenSAFELY assessment, will be shared with NHS England, who will make the final decision. Approval to use a non SDC-Supported method does not automatically imply that any requested outputs will be approved for release.
Last updated: 24th July 2026