SurPyval - Survival Analysis in Python

SurPyval is a Python package for survival analysis: the statistics of how long things last. The “thing” can be a bearing, a patient, a loan, a recession or a marathon runner, and the “time” can be hours, cycles, kilometres or dollars. What makes these problems different from ordinary statistics is that the data is almost never complete. Some units have not failed yet when the study ends (censoring), some were only inspected now and then, and some never reached your data set at all because of how it was collected (truncation). SurPyval is built around handling every one of those situations correctly, in any combination, with one consistent interface.

The package grew out of reliability engineering, so it has the tools an engineer expects (probability plots, offsets, B-lives, accelerated life tests, burn-in and warranty calculations, repairable systems, degradation), alongside the tools of biostatistics and econometrics (Kaplan-Meier, Cox regression, competing risks, frailty, predictive metrics).

Installation

SurPyval needs Python 3.11 or later and installs with pip:

$ pip install surpyval

To work on SurPyval itself, see Contributing.

Where to start

  • New to SurPyval? Read the Quickstart. In ten minutes it fits a first model, shows why censoring matters, and gives a small example of every area of the package.

  • New to survival data, or not sure whether your data is censored or truncated? Read Types of Data. Getting this right matters more than any choice of model.

  • Have data in a spreadsheet, a list of failures and survivors, or a maintenance log? Data Wrangling Examples shows how to turn it into the x, c, n, t arrays every fitter takes, and Conventions defines those arrays exactly.

How the documentation is organised

The pages are in five groups, in the menu on the left:

  • Quickstart & Intro: the pages above, plus Handy References - Aide-mémoire, a one-page summary of the functions of a distribution and how they relate.

  • Survival Analysis: the theory pages. Each explains a family of methods from first principles: the problem it solves, the mathematics, how to read the results, what can go wrong and when to use something else. They are written so that you can learn the methods from them.

  • SurPyval Modelling: the how-to pages. Each shows how to do that analysis in SurPyval, with worked examples whose output is computed when the documentation is built, so it always matches the current version. It also includes Example Applications, complete examples from several fields. Each theory page links to its how-to page and the other way round.

  • SurPyval: the API reference (API), with every class, function and argument, and the Changelog.

  • Community Guidelines: where to get help, how to report a problem and how to contribute.

What SurPyval covers

Every model is created the same way: a fitter (surv.Weibull, surv.KaplanMeier, surv.CoxPH, …) has a fit() method that takes your data and returns a fitted model, which you then ask for survival probabilities, hazards, quantiles, confidence bounds, plots and so on.

Area

What is in it

Theory / how-to

Non-parametric estimation

Kaplan-Meier, Nelson-Aalen, Fleming-Harrington and Turnbull estimators with confidence bounds and bands; the log-rank test; restricted mean survival time; success-run (zero-failure) testing.

Non-Parametric Estimation / Non-Parametric SurPyval Modelling

Parametric distributions

Continuous lifetime distributions (Weibull, Exponential, Gamma, LogNormal, LogLogistic, Gumbel, Normal, Logistic, Exponentiated Weibull, Beta, Rayleigh, Uniform, …) and discrete ones (Geometric, Poisson, Binomial, Negative Binomial, Beta-Geometric, discrete Weibull), fitted by any of the five estimation methods in the table below. Offsets (the three-parameter Weibull), limited failure populations and zero-inflation; mixture models; your own distribution from a cumulative hazard; the Royston-Parmar flexible spline model; automatic choice of distribution with fit_best.

Parametric Estimation / Parametric SurPyval Modelling

Regression

Proportional hazards (Cox, with strata and time-varying covariates, and parametric), accelerated failure time, accelerated life with life-stress relationships (power law, Eyring, exponential/Arrhenius and others), proportional odds, additive hazards, Buckley-James, shared frailty; Brier score and time-dependent AUC for validating predictions.

Regression Analysis / Regression Modelling with SurPyval

Competing risks

Items that can fail from one of several causes: cumulative incidence (Aalen-Johansen), parametric cause-specific models, cause-specific Cox regression, Fine-Gray subdistribution regression and Gray’s test.

Competing Risks Analysis / Competing Risks SurPyval Modelling

Recurrent events

Repairable items that fail again and again: the mean cumulative function, homogeneous and non-homogeneous Poisson processes (Crow-AMSAA, Duane, Cox-Lewis), imperfect repair (Kijima, G1, ARA, ARI), trend tests, and proportional-intensity regression.

Recurrent Event Analysis / Recurrent Event Modelling with SurPyval; Recurrent Event Regression Analysis / Recurrent Event Regression Modelling with SurPyval

Degradation

Measurements that drift towards a failure threshold: general path models, Wiener and Gamma processes, destructive degradation, remaining useful life of a monitored unit, and accelerated and step-stress degradation tests.

Degradation Analysis / Degradation Modelling with SurPyval

Multivariate

Dependent lifetimes joined by a copula (Clayton, Frank, Gumbel, Gaussian, independence), with censored and truncated margins.

Multivariate Analysis / Multivariate Modelling with SurPyval

Almost every fitted model can be saved to a dictionary or a JSON file and restored with surv.from_dict or surv.from_json (see Conventions for what a restored model keeps, and the few models that cannot be saved). A survival tree and a random survival forest are available as beta-stage models in surpyval.beta.ml (see Machine Learning (beta)).

Estimating a single distribution

SurPyval is unusual in how many ways it can estimate the parameters of a distribution. Most packages offer maximum likelihood, and perhaps probability plotting; SurPyval grew out of reproducing the probability plotting of engineering practice and found along the way that there are many ways to estimate parameters, each with its own strengths. The methods, and the data each can use, are:

SurPyval Modelling Methods

Method

Para/Non-Para

Observed

Censored

Truncated

Maximum Likelihood (MLE)

Parametric

Yes

Yes: left, right and interval

Yes: left and right, a different value for each observation if needed

Maximum Product Spacing (MPS)

Parametric

Yes

Left and right; not interval

Yes, but one common tl and/or tr for all observations

Probability Plotting (MPP)

Parametric

Yes

Right; left and interval only with heuristic="Turnbull"

Left (with the Kaplan-Meier, Nelson-Aalen, Fleming-Harrington or Turnbull heuristic); right only with heuristic="Turnbull"

Mean Square Error (MSE)

Parametric

Yes

Yes: left, right and interval

No

Method of Moments (MOM)

Parametric

Yes

No

No

Kaplan-Meier

Non-Parametric

Yes

Right only

Left only

Nelson-Aalen

Non-Parametric

Yes

Right only

Left only

Fleming-Harrington

Non-Parametric

Yes

Right only

Left only

Turnbull

Non-Parametric

Yes

Yes: left, right and interval

Yes: left and right

Maximum likelihood is the default (how="MLE") and, with Turnbull for a non-parametric view, handles every combination of data. The other methods have their own uses: probability plotting is the traditional engineering method and gives the familiar straight-line plot; maximum product spacing is particularly good for offset distributions and distributions with a finite bound, where the likelihood can be unbounded; and the method of moments is quick for complete data. Some distributions do not support every method (probability plotting needs a distribution that can be drawn as a straight line, and discrete distributions cannot use MPS). Whenever a method cannot use your data, SurPyval raises an error saying so rather than silently ignoring part of it. The methods are explained in Parametric Estimation and Non-Parametric Estimation, and Types of Data has the same table split by type of censoring and truncation.

A word of encouragement

Becoming a competent survival analyst depends on a strong understanding of censoring, truncation and observations, together with a solid understanding of the distributions used to describe lifetimes. Recognising that a real situation is censored or truncated is what keeps an analysis from going wrong, and it can be surprisingly difficult; Types of Data is the place to build that skill. A good understanding of the distributions lets you reason about the process that generated your data, and so choose an appropriate model, if any, for your problem. Survival analysis is an extremely powerful, and thoroughly interesting, tool, so don’t give up; or if you do give up, do the survival statistics on it.

Contents:

Survival Analysis

SurPyval Modelling

Indices and tables