Suicide risk prediction models show promise—but not a ready-made answer

A systematic review of 167 models finds that accurate, independently tested prediction of self-harm and suicide remains difficult, with only a small number of models showing promising results in external validation.

·

·

5 min read

Suicide risk prediction models show promise—but not a ready-made answer — AI-generated editorial image

⏱ 5 min read

A large systematic review suggests that statistical models may help identify patterns associated with self-harm and suicide—but it also shows why these tools cannot currently provide a simple or definitive answer about an individual’s risk.

Seyedsalehi and colleagues examined 91 articles covering 167 multivariable prediction models. The models were designed to predict self-harm, suicide, or both. The review, published in 2025, included evidence identified through searches of major health and social care databases, with an updated search for external validations completed in October 2024.

What the review found

The evidence base was uneven. Of the 167 models, 76 predicted self-harm, 51 predicted suicide and 40 predicted the combined outcome. Yet only 14 models—8%—had been tested in an external dataset. Just 28 models, or 17%, were described in enough detail for another research team to attempt validation.

External validation matters because a model can appear to perform well in the data used to create it but work less effectively in a different population, health system or setting. More than 60% of the models and validations used data from the United States, while 72% relied on routine sources such as electronic health records or administrative databases.

The models varied considerably in complexity. The final models used between two and 8,071 predictor parameters, with a median of 13. Candidate parameters considered during development ranged from nine to more than 89,000.

Performance often weakened outside the original dataset

The review used the C-index to describe discrimination: how well a model distinguishes between people who do and do not go on to experience the outcome being predicted. A score of 0.5 is equivalent to a coin toss, while scores closer to 1 indicate better discrimination.

Among development studies, C-indices ranged from 0.61 to 0.97, with a median of 0.82. In external validation studies, they ranged from 0.60 to 0.86, with a median of 0.81.

For self-harm models, the median C-index fell from 0.85 during development to 0.73 during external validation. For suicide models, it fell from 0.82 to 0.76. Models predicting the combined outcome showed a rise from 0.79 to 0.85, although the review cautioned that the overall evidence was limited.

These figures do not show whether a model is clinically useful on their own. A tool can rank people by estimated risk without accurately predicting what will happen to a particular person. The review also found that calibration—the agreement between predicted and observed risk—was assessed infrequently: in only 15 development studies and nine of 29 external validations, covering six models in total.

Two model families stood out

The review identified five models with good predictive performance in external datasets. Two model families, OxMIS and the Simon models, showed both adequate discrimination and calibration in external validation.

OxMIS is a freely available, web-based, 17-item model intended to estimate one-year suicide risk among people with severe mental illness. It uses sociodemographic and clinical risk factors. In its original development paper, OxMIS reported a sensitivity of 55% and specificity of 75%. Its positive predictive value was 2%, while its negative predictive value was 99%.

The Simon models estimate 90-day risk of suicide attempt and suicide death after mental health specialty and general medical visits. They use 313 demographic and clinical characteristics from electronic health records. Across the four models, the original paper reported sensitivity ranging from 7.0% to 48.1%, specificity from 95.0% to 95.2%, positive predictive values from 0.26% to 5.4%, and negative predictive values from 99.6% to 99.9%.

The review rated OxMIS as the only model whose external validations were at low risk of bias. All model development studies were judged to have a high risk of bias, as were all but two external validations.

Why caution is still needed

The most frequently reported problems were incomplete or inappropriate evaluation of performance, seen in 92% of studies; insufficient sample sizes, seen in 77%; inappropriate handling of missing data, seen in 66%; and failure to account for overfitting and optimism in performance estimates, seen in 63%.

More data and more predictors did not automatically produce better results. High-dimensional models had a median C-index of 0.82, the same as low-dimensional models. Models based on routine data had a median C-index of 0.84, compared with 0.81 for models built from prospectively collected data.

The review also highlights an important limitation in the underlying evidence: suicide attempts and non-suicidal self-injury were not distinguished. The analysis used the NICE definition of self-harm—intentional self-injury or self-poisoning regardless of intent—but the authors note that these are distinct constructs.

What this means for care

Only 11 of the 167 models, or 7%, were available to clinicians as a tool for calculating suicide risk. The review concluded that accurately predicting suicide remains a substantial challenge and that apparently promising findings should be interpreted with realistic scepticism.

The findings do not support dismissing all prediction models. The authors said blanket criticisms of their predictive performance are not evidence-based, pointing to the five models that performed well in external datasets. At the same time, no single model is likely to be sufficient for accurately assessing suicide risk, and the clinical usefulness of the tools remains uncertain.

NICE self-harm guidance and NHS England suicide prevention guidance currently advise against using risk prediction tools. The review authors suggest that these findings should be considered when clinical guidance is updated in the future.

For now, the central message is one of proportion: statistical models may contribute to research and, potentially, to carefully evaluated clinical systems, but they should not be treated as definitive judgments about an individual. Their results need to be understood alongside the limits of the evidence and the circumstances of the person being assessed.

AI tools were used to assist with the preparation of this article.

Leave a Reply

Your email address will not be published. Required fields are marked *