Utilities

The data containers the fitters build from your arrays, the functions that convert between SurPyval’s data formats, and the abstract base classes every model derives from. The formats themselves – xcnt (times, censoring flags, counts, truncation), xrd (times, numbers at risk, deaths) and fsli (lists of failures, suspensions, left- and interval-censored values) – are defined in Conventions, and Data Wrangling Examples shows them in use.

Data Classes

SurpyvalData holds univariate xcnt data, split by censoring type; every univariate fitter builds one, and fit_from_surpyval_data accepts one directly. RecurrentEventData holds recurrent-event data (each row an event of an item); build it with handle_xicn.

class surpyval.utils.surpyval_data.SurpyvalData(x: ArrayLike | None = None, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, xl: ArrayLike | None = None, xr: ArrayLike | None = None, tl: ArrayLike | Number | None = None, tr: ArrayLike | Number | None = None, Z: ArrayLike | None = None, group_and_sort: bool = True, handle: bool = True)

Bases: object

Validated univariate survival data in the xcnt format, split by type of observation for the likelihoods.

Every univariate fitter builds one from its arrays (with xcnt_handler()); parametric fitters also accept one directly through fit_from_surpyval_data. Besides x, c, n and t (and tl/tr, the columns of t), it holds the observed (x_o, n_o), right censored (x_r, n_r), left censored (x_l, n_l) and interval (x_il, x_ir, n_i) rows separately. A censored row with a finite truncation bound on its censored side is stored as an interval ending at that bound, so the likelihood uses the probability of the part of the window it could have fallen in.

Examples

>>> from surpyval import SurpyvalData
>>> inf = float("inf")
>>> data = SurpyvalData(x=[1, 2, 3, 4], c=[0, 1, 0, 1], tr=[9, 3, 9, inf])
>>> data.x_o, data.x_r
(array([1., 3.]), array([4.]))
>>> data.x_il, data.x_ir
(array([2.]), array([3.]))

Initialize a SurpyvalData instance for survival analysis.

Validates, sorts, and stores survival data in the xcnt format. Supports uncensored, right/left/interval-censored, and truncated observations. Can convert to xrd format and select subsets by censoring type.

Parameters:
  • x (array-like, optional) – The primary data array of failure/event times. When c is 2 the corresponding x entry is a 2-element array [left, right].

  • c (array-like, optional) – Censoring flags for each value in x: * 0 = uncensored * 1 = right censored * -1 = left censored * 2 = interval censored

  • n (array-like, optional) – Number of occurrences for each value in x (positive whole numbers).

  • t (array-like, optional) – 2D array of truncation bounds [left, right] for each value in x.

  • xl (array-like, optional) – Left interval bounds for interval censored data. Cannot be used with ‘x’. Must be paired with ‘xr’.

  • xr (array-like, optional) – Right interval bounds for interval censored data. Cannot be used with ‘x’. Must be paired with ‘xl’.

  • tl (array-like or scalar, optional) – Left truncation bounds. Cannot be used with ‘t’. May be given alone (the right bound is then infinite) or with ‘tr’.

  • tr (array-like or scalar, optional) – Right truncation bounds. Cannot be used with ‘t’. May be given alone (the left bound is then minus infinity) or with ‘tl’.

  • Z (array-like, optional) – Covariates, one row per value of x (see also add_covariates()). Giving Z turns group_and_sort off, so the rows stay aligned with x.

  • group_and_sort (bool, default=True) – Whether to group and sort the data. Set False when using covariates to maintain data order.

  • handle (bool, default=True) – Whether to validate and process the input data. Set False for pre-validated data: x, c, n and a two-column t exactly as xcnt_handler() returns them.

Examples

Basic usage with uncensored data: >>> x = np.array([1, 2, 3]) >>> data = SurpyvalData(x)

Right censored data: >>> x = np.array([1, 2, 3]) >>> c = np.array([0, 1, 1]) # 2 and 3 are censored >>> data = SurpyvalData(x, c)

Interval censored data: >>> xl = np.array([1, 2, 3]) >>> xr = np.array([2, 3, 4]) >>> data = SurpyvalData(xl=xl, xr=xr)

Interval Censored with nested 2 arrays: >>> x = [1, 2, [2, 5], 3, 6] >>> c = [0, 1, 2, 0, 0] >>> data = SurpyvalData(x=x, c=c)

With truncation: >>> x = np.array([1, 2, 3]) >>> t = np.array([[0, 5], [0, 5], [0, 5]]) >>> data = SurpyvalData(x, t=t)

add_covariates(Z: ArrayLike) → None

Method to add covariates to the data. The covariates are stored in the Z attribute of the object. When doing regression survival analysis this method allows for the covariates to be added to the data in a consistent manner that also allows for the data to be converted to be passed to the fitters.

Parameters:

Z (numpy.ndarray) – The covariate array.

classmethod from_json(source: str | Path) → SurpyvalData

Create SurpyvalData instance from JSON text or a file.

Parameters:

source (str | Path) – Pass a pathlib.Path to load from a file; a str is always parsed as JSON text, so from_json("data.json") fails even though to_json("data.json") writes that file. The data is not re-validated.

Returns:

New instance created from JSON data

Return type:

SurpyvalData

Examples

>>> from surpyval import SurpyvalData
>>> data = SurpyvalData(x=[1, 2, 2, 5], c=[0, 0, 0, 1])
>>> restored = SurpyvalData.from_json(data.to_json())
>>> restored.x, restored.n
(array([1., 2., 5.]), array([1, 2, 1]))
to_json(filepath: str | Path | None = None) → str | None

Serialize SurpyvalData to JSON format.

The JSON is strict: the infinite truncation bounds (and any other non-finite value) are written as null and recorded under "non_finite", which from_json() reads back (see Saving and Loading Models).

Parameters:

filepath (str | Path, optional) – If provided, saves JSON to this file path

Returns:

JSON string if no filepath provided, None if saved to file

Return type:

str | None

to_xrd(estimator: str = 'Nelson-Aalen') → tuple

Converts the data into the xrd format. Observed and right censored data without right truncation is converted exactly with xcnt_to_xrd(). If the data has right truncated observations or left or interval censored observations, the Turnbull estimator is fitted and its (possibly fractional) numbers at risk and deaths are returned; the estimator parameter is used only in that case.

Parameters:

estimator (str, optional) – The method for estimation if data requires the use of the Turnbull estimator to convert to xrd, defaults to “Nelson-Aalen”.

Returns:

The xrd data, (x, r, d).

Return type:

tuple

class surpyval.utils.recurrent_event_data.RecurrentEventData(x: ArrayLike, i: ArrayLike, c: ArrayLike, n: ArrayLike, e: ArrayLike | None = None, tl: ArrayLike | None = None, tr: ArrayLike | None = None)

Bases: object

A class to handle and manipulate recurrent event data. Recurrent events are those that can occur more than once for each subject or item.

The recurrent fitters build one from their x, i, c, n arrays with surpyval.handle_xicn, which validates the input; build one that way to pass to a fitter’s fit_from_recurrent_data. The constructor itself does no validation.

Examples

>>> import numpy as np
>>> from surpyval import RecurrentEventData
>>> x = np.array([1, 2, 3, 4, 5, 1, 2, 3, 4, 5])
>>> c = np.array([0, 0, 1, 1, 1, 0, 0, 0, 0, 1])
>>> n = np.array([1, 1, 1, 1, 1, 1, 1, 1, 1, 1])
>>> i = np.array([1, 1, 1, 1, 1, 2, 2, 2, 2, 2])
>>> data = RecurrentEventData(x, i, c, n)
>>> data.to_xrd()
(array([1, 2, 3, 4, 5]), array([2, 2, 2, 2, 2]), array([2, 2, 1, 1, 0]))
>>> data[0:2]
RecurrentEventData(
    x=[1 2],
    i=[1 1],
    c=[0 0],
    n=[1 1]
)
>>> data.get_times_to_first_events()
SurpyvalData(
    x=array([1.]),
    c=array([0]),
    n=array([2]),
    t=array([[-inf,  inf]])
)
>>> data.get_interarrival_times()
array([1, 1, 1, 1, 1, 1, 1, 1, 1, 1])
property event_types: list

The distinct event types (marks) present in the data, excluding the None mark used for censored / end-of-observation rows, in the order of ordered_labels (mixed str and int marks are ordered too). Returns an empty list when the data carries no marks.

get_events_for_item(item: Any) → tuple[NDArray, NDArray, NDArray]

Get all events for a specific item or subject.

Parameters:

item (int or str) – The id of the item or subject.

Returns:

A tuple containing event times, censoring information and frequencies for the specified item.

Return type:

tuple

get_interarrival_times() → NDArray

Finds the interarrival times between events for each item. The class assumes that the time of the event is cumulative, sometimes it is necessary to know the interarrival times of the events. This method returns the interarrival times for each item. It is aligned with the items attribute.

Returns:

An array of interarrival times.

Return type:

numpy.ndarray

get_previous_x(min_x: float = 0) → NDArray

Finds the previous event time for each event. This is useful for calculating the time since the last event. This method returns the previous event time for each event. It is aligned with the items attribute.

Parameters:

min_x (float, optional) – Fallback minimum for the first event of each item. The item’s left truncation bound is used instead when it is greater, so a delayed-entry item’s first interval begins at its entry time.

Returns:

An array of previous event times.

Return type:

numpy.ndarray

get_right_truncation_close() → tuple[NDArray, NDArray, NDArray]

Per-item integration bounds for the NHPP likelihood’s right window-close.

The NHPP integral runs from each item’s entry time (its left truncation tl, handled by get_previous_x()) to the time its observation window closes. Historically that close was only known from an explicit right-censoring (c=1) row, so the integral stopped at the item’s last recorded time. When an item instead carries a finite right-truncation time tr the window closes there, and the integral must be extended from the last in-window time out to tr.

Returns three aligned arrays (x_last, x_close, rep_idx) with one entry per item whose tr is finite: x_last is the item’s last in-window time (its last event or right-censoring row), x_close is its tr, and rep_idx is a representative row index for the item (so per-item covariates Z can be gathered). Items with the default tr = inf are omitted, so untruncated data yields empty arrays and contributes nothing to the integral. Adding cif(x_last) - cif(x_close) to the log-likelihood therefore extends the telescoped integral to tr (and is exactly zero when a c=1 row already sits at tr, so the two ways of closing the window never double-count).

Returns:

(x_last, x_close, rep_idx) as described above.

Return type:

tuple of numpy.ndarray

get_times_to_first_events() → SurpyvalData

Get the times to the first events for each item or subject. In the estimation of recurrent or renewal events it can be helpful to know the distribution of the times to the first event per item. This method returns the times to the first events for each item. It is aligned with the items attribute.

Returns:

A transformed dataset containing times to the first events.

Return type:

SurpyvalData

item_observation_windows() → tuple[NDArray, NDArray]

The (entry, exit) observation window of each item, aligned with items.

The entry is the item’s left-truncation bound tl (delayed entry); with no truncation it is -inf, so the item is at risk from the start. The exit is the item’s last recorded time (its last event or its end-of-observation c=1 row) or, when the item carries a finite right-truncation time tr, that tr: observation of the item ends there, exactly as if it had an end-of-observation row at tr. This is the same window-close the NHPP likelihoods integrate to (see get_right_truncation_close()).

Returns:

(entry, exit), one value per item.

Return type:

tuple of numpy.ndarray

item_rows() → tuple[NDArray, NDArray]

The first row of each item and the item of each row, as positions in items (which is sorted).

Returns:

(first, inverse): first[k] is the index of item items[k]’s first row, and inverse[j] the position in items of row j’s item.

Return type:

tuple of numpy.ndarray

Examples

>>> from surpyval import handle_xicn
>>> data = handle_xicn([1, 2, 3, 4], i=[2, 1, 2, 1])
>>> data.items, data.i
([1, 2], array([1, 1, 2, 2]))
>>> data.item_rows()
(array([0, 2]), array([0, 0, 1, 1]))
split_for_nhpp_likelihood() → dict

Split the data into the pieces every NHPP-style likelihood needs: per-censoring-type event and previous-event times, the right-truncation window close, and the row masks (so regression variants can gather covariates the same way). This preamble used to be copy-pasted between the plain and proportional-intensity NHPP fitters and drifted (#288); it now lives here once (#296).

to_cause_specific_xrd(cause: Any) → tuple[NDArray, NDArray, NDArray]

Convert the recurrent event data to xrd format for a single event type (cause). The at-risk set r is shared across all causes (an item remains at risk for every cause until it leaves observation); only the event count d is restricted to the requested cause.

Parameters:

cause (object) – The event type to compute the cause-specific counts for. Must be one of self.event_types.

Returns:

A tuple (x_unique, r, d_cause) where d_cause counts only events of the requested cause and r is the shared at-risk set.

Return type:

tuple

to_xrd(estimator: str = 'Nelson-Aalen') → tuple[NDArray, NDArray, NDArray]

Convert the recurrent event data to xrd format.

Parameters:

estimator (str, optional) – The estimator to use, defaults to “Nelson-Aalen”.

Returns:

A tuple containing unique event times, risk set sizes, and the event counts.

Return type:

tuple

Data Wrangling Utilities

Validate data in one format and convert it to another. Each of these is importable directly from surpyval.

surpyval.utils.xcnt_handler(x: ArrayLike | None = None, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, xl: ArrayLike | None = None, xr: ArrayLike | None = None, tl: ArrayLike | Number | None = None, tr: ArrayLike | Number | None = None, group_and_sort: bool = True) → tuple[NDArray, NDArray, NDArray, NDArray]

Main handler that ensures any input to a surpyval fitter meets the requirements to be used in one of the parametric or nonparametric fitters.

It converts the inputs to numpy arrays and checks them: x is not empty and has no NaN (a scalar is one observation); c holds only -1, 0 and 1 (and 2 for a two-column x); n holds positive whole numbers; the truncation bounds have no NaN and the window of each row has its left bound below its right bound. For two-column x, a row with equal values is not an interval, a row with different values must be flagged 2 (when c is given), and an infinite end turns the row into one-sided censoring: [v, inf] becomes right censored at v and [-inf, v] left censored at v. If no interval is left, x is returned with one column.

Each row must then fit its truncation window (tl, tr]: a single value lies strictly above tl and at or below tr – strictly below tr if it is right censored – whether x has one column or two; an interval [xl, xr] needs tl <= xl and xr <= tr. Identical rows are then merged (their counts summed) and the rows sorted with xcnt_sort().

Parameters:
  • x (array) – array of values of variable for which observations were made.

  • c (array, optional (default: None)) – array of censoring values (-1, 0, 1, 2) corresponding to x

  • n (array, optional (default: None)) – array of count of observations at each x and with censoring c

  • t (array, optional (default: None)) – array of values with shape (?, 2) with the left and right value of truncation

  • xl (array or scalar, optional (default: None)) – array of the values of the left interval of interval censored data. Cannot be used with ‘x’ parameter, must be used with the ‘xr’ parameter

  • xr (array or scalar, optional (default: None)) – array of the values of the right interval of interval censored data. Cannot be used with ‘x’ parameter, must be used with the ‘xl’ parameter

  • tl (array or scalar, optional (default: None)) – array of values of the left value of truncation. If scalar, all values will be treated as left truncated by that value cannot be used with ‘t’ parameter but can be used with the ‘tr’ parameter

  • tr (array or scalar, optional (default: None)) – array of values of the right value of truncation. If scalar, all values will be treated as right truncated by that value cannot be used with ‘t’ parameter but can be used with the ‘tl’ parameter

  • group_and_sort (bool, optional (default: True)) – whether to group and sort the data. If False, the data will be returned in the order it was entered. This is useful for when validating survival data for which you also have covariates.

Returns:

  • x (array) – sorted array of values of variable for which observations were made.

  • c (array) – array of censoring values (-1, 0, 1, 2) corresponding to output array x. If c was None, every row is observed (0), except that the rows of a two-column x with different values are interval censored (2).

  • n (array) – array of count of observations at output array x and with censoring c. If n was None, count array assumed to be all one observation.

  • t (array) – array of truncation values of observations at output array x and with censoring c.

Raises:

ValueError – If the inputs break any of the rules above: for example x and xl/xr both given, empty x, arrays of different lengths, a NaN in x or in a truncation bound, an unknown censoring flag, a count that is not a positive whole number, a value at or below its own left truncation, or a right censored value at its own right truncation.

Examples

>>> from surpyval import xcnt_handler
>>> x = [1, 2, 3, 4, 5]
>>> c = [0, 0, 1, 1, 1]
>>> n = [1, 1, 1, 1, 1]
>>> t = [[0, 6], [0, 6], [0, 6], [0, 6], [0, 6]]
>>> xcnt_handler(x, c, n, t)
(array([1., 2., 3., 4., 5.]),
array([0, 0, 1, 1, 1]),
array([1, 1, 1, 1, 1]),
array([[0., 6.],
        [0., 6.],
        [0., 6.],
        [0., 6.],
        [0., 6.]]))
>>> xcnt_handler(x, c, n, tl=0, tr=6)
(array([1., 2., 3., 4., 5.]),
array([0, 0, 1, 1, 1]),
array([1, 1, 1, 1, 1]),
array([[0., 6.],
        [0., 6.],
        [0., 6.],
        [0., 6.],
        [0., 6.]]))
>>> xl = [1, 2, 3, 4, 5]
>>> xr = [2, 3, 4, 5, 6]
>>> xcnt_handler(xl=xl, xr=xr)
(array([[1., 2.],
        [2., 3.],
        [3., 4.],
        [4., 5.],
        [5., 6.]]),
array([2, 2, 2, 2, 2]),
array([1, 1, 1, 1, 1]),
array([[-inf,  inf],
        [-inf,  inf],
        [-inf,  inf],
        [-inf,  inf],
        [-inf,  inf]]))
surpyval.utils.fsli_handler(f: ArrayLike | None = None, s: ArrayLike | None = None, l: ArrayLike | None = None, i: ArrayLike | None = None) → tuple[NDArray, NDArray, NDArray, NDArray]

Validate data in the fsli format: separate lists of failures, suspensions (right censored), left censored values and intervals. Any combination may be given, but at least one must hold data. Each is returned as a float array.

Parameters:
  • f (array-like, optional (default: None)) – array of values for which the failure/death was observed

  • s (array-like, optional (default: None)) – array of right censored observation values

  • l (array-like, optional (default: None)) – array of left censored observation values

  • i (array-like, optional (default: None)) – array of [lower, upper] pairs, one per interval censored observation, with lower < upper

Raises:

ValueError – If no data is given, if f, s or l is not one-dimensional, if i is not of shape (k, 2), if any value is NaN, or if an interval’s lower value is not below its upper value.

Returns:

  • f (array) – array of values for which the failure/death was observed that have been checked for correctness

  • s (array) – array of right censored observation values that have been checked for correctness

  • l (array) – array of left censored observation values that have been checked for correctness

  • i (array) – array of interval censored data that have been checked for correctness

Examples

>>> from surpyval import fsli_handler
>>> f = [1, 2, 3, 4, 5, 6]
>>> s = [1, 2, 3]
>>> l = [4, 5, 6]
>>> i = [[1, 2], [3, 4]]
>>> fsli_handler(f, s, l, i)
(array([1., 2., 3., 4., 5., 6.]),
array([1., 2., 3.]),
array([4., 5., 6.]),
array([[1., 2.],
        [3., 4.]]))
surpyval.utils.xrd_handler(x: ArrayLike, r: ArrayLike, d: ArrayLike) → tuple[NDArray, NDArray, NDArray]

Takes a combination of ‘x’, ‘r’, and ‘d’ arrays and ensures that the data is feasible: the arrays are one-dimensional and the same length, x has no NaN and no repeated time, r and d hold whole numbers (integer-valued floats such as [5.0, 4.0] are accepted), every r is at least one, no d is negative and no d exceeds its r.

xrd data lists each distinct time once, with the number at risk and the number of deaths at that time. Each (x, r, d) row is therefore self-contained, so rows given out of order are sorted by x (carrying their r and d with them) rather than refused. A repeated time is refused: there is no unambiguous way to merge two rows that each claim to be the risk set at that time.

Does not check that r decreases, as it can grow when there is left truncation (late entry).

Parameters:
  • x (array) – array of values of variable for which observations were made.

  • r (array) – array of at risk items at each value of x

  • d (array) – array of failures / deaths at each value of x

Returns:

  • x (array) – array of values of variable for which observations were made, in increasing order.

  • r (array) – array of at risk items at each value of x

  • d (array) – array of failures / deaths at each value of x

Raises:

ValueError – If any of the rules above is broken.

Examples

>>> from surpyval import xrd_handler
>>> x = [1, 2, 3, 4, 5]
>>> r = [5, 4, 3, 2, 1]
>>> d = [1, 1, 1, 1, 1]
>>> x, r, d = xrd_handler(x, r, d)
>>> x
array([1., 2., 3., 4., 5.])
>>> r
array([5, 4, 3, 2, 1])
>>> d
array([1, 1, 1, 1, 1])

Rows out of order are sorted together:

>>> xrd_handler([3, 1, 2], [2, 5, 4], [1, 1, 1])
(array([1., 2., 3.]), array([5, 4, 2]), array([1, 1, 1]))
surpyval.utils.recurrent_utils.handle_xicn(x: ArrayLike, i: ArrayLike | None = None, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, tl: ArrayLike | None = None, tr: ArrayLike | None = None, Z: ArrayLike | dict | None = None, as_recurrent_data: Literal[True] = True, windows: dict | None = None, e: ArrayLike | None = None) → RecurrentEventData
surpyval.utils.recurrent_utils.handle_xicn(x: ArrayLike, i: ArrayLike | None = None, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, tl: ArrayLike | None = None, tr: ArrayLike | None = None, Z: ArrayLike | dict | None = None, *, as_recurrent_data: Literal[False], windows: dict | None = None, e: ArrayLike | None = None) → tuple[NDArray, NDArray, NDArray, NDArray]

Validate recurrent-event data given as arrays and assemble it into a RecurrentEventData, the object every recurrent fitter’s fit_from_recurrent_data takes.

Each row is one event (or the right-censored end of observation) of the item named in i, at time x measured from the start of that item’s life.

Parameters:
  • x (array like) – The event times (a 2-D [left, right] row for an interval-censored count).

  • i (array like, optional) – The item each row belongs to. Defaults to one item.

  • c (array like, optional) – Censoring flags: 0 an observed event, 1 the right-censored end of observation, -1 left-censored and 2 interval-censored counts. Defaults to all observed. Rows are sorted by item and time; an end-of-observation row at the same time as an event goes after it, whatever the input order.

  • n (array like, optional) – The number of events in each row. Defaults to 1.

  • t (array like, optional) – (N, 2) truncation bounds per row. Use tl / tr for per-item bounds instead.

  • tl (array like or scalar, optional) – Left-truncation (start of observation) and right-truncation (end of observation) times. An item with both a c=1 row and a finite tr must have them at the same time.

  • tr (array like or scalar, optional) – Left-truncation (start of observation) and right-truncation (end of observation) times. An item with both a c=1 row and a finite tr must have them at the same time.

  • Z (array like or dict, optional) – Covariates: one row per row of x (the same on every row of an item: covariates are per item), or a {item: covariates} mapping applied to every row of that item.

  • as_recurrent_data (bool, optional) – If True (the default) return a RecurrentEventData; otherwise return the validated (x, i, c, n) arrays.

  • windows (dict, optional) – Gapped observation: {item: [(start, end), ...]}. Every row must then be an observed event; the windows supply the censoring rows. Not combinable with t/tl/tr, Z or e.

  • e (array like, optional) – The event type (mark) of each row, for the cause-specific models; a missing value marks a row with no cause (such as the censoring row).

Returns:

The assembled data, or (x, i, c, n).

Return type:

RecurrentEventData or tuple

Examples

>>> from surpyval import handle_xicn
>>> data = handle_xicn([2, 5, 8, 3, 9], i=[1, 1, 1, 2, 2],
...                    c=[0, 0, 1, 0, 1])
>>> data.items
[1, 2]
surpyval.utils.fs_to_xcnt(f: ArrayLike | None = None, s: ArrayLike | None = None) → tuple[NDArray, NDArray, NDArray, NDArray]

Convert failure (f) and suspension (right-censored, s) times to the xcnt format, counting repeated times; see fsl_to_xcnt().

Parameters:
  • f (array like, optional) – Observed failure times.

  • s (array like, optional) – Right-censored (suspension) times.

Returns:

x, c, n, t – The distinct times, censoring flags (0 or 1), counts and (untruncated) truncation bounds, sorted.

Return type:

arrays

Examples

>>> from surpyval import fs_to_xcnt
>>> x, c, n, t = fs_to_xcnt([1, 3, 3, 7], [5, 9])
>>> x
array([1., 3., 5., 7., 9.])
>>> c, n
(array([0, 0, 1, 0, 1]), array([1, 2, 1, 1, 1]))
surpyval.utils.fsl_to_xcnt(f: ArrayLike | None = None, s: ArrayLike | None = None, l: ArrayLike | None = None) → tuple[NDArray, NDArray, NDArray, NDArray]

Convert failure (f), suspension (right-censored, s) and left-censored (l) times to the xcnt format, counting repeated times.

Parameters:
  • f (array like, optional) – Observed failure times.

  • s (array like, optional) – Right-censored (suspension) times.

  • l (array like, optional) – Left-censored times.

Returns:

x, c, n, t – The distinct times, censoring flags, counts and (untruncated) truncation bounds, sorted.

Return type:

arrays

Examples

>>> from surpyval import fsl_to_xcnt
>>> x, c, n, t = fsl_to_xcnt([4, 6], [8], [2])
>>> x, c, n
(array([2, 4, 6, 8]), array([-1,  0,  0,  1]), array([1, 1, 1, 1]))
surpyval.utils.fsli_to_xcnt(f: ArrayLike | None = None, s: ArrayLike | None = None, l: ArrayLike | None = None, i: ArrayLike | None = None) → tuple[NDArray, NDArray, NDArray, NDArray]

Converts the fsli format to the xcnt format, so that the data can be passed to one of the parametric or nonparametric fitters. The inputs are validated with fsli_handler(). Repeated values are counted in n. When there are intervals, x is returned with two columns, the other rows repeating their value.

Parameters:
  • f (array) – array of values for which the failure/death was observed

  • s (array) – array of right censored observation values

  • l (array) – array of left censored observation values

  • i (array) – array of [lower, upper] pairs of interval censored data

Returns:

  • x (array) – sorted array of values of variable for which observations were made.

  • c (array) – array of censoring values (-1, 0, 1, 2) corresponding to output array x.

  • n (array) – array of count of observations at output array x and with censoring c.

  • t (ndarray) – ndarray of truncation values of observations at output array x and with censoring c.

Examples

>>> from surpyval import fsli_to_xcnt
>>> f = [1, 4, 5]
>>> s = [2, 3]
>>> l = []
>>> i = []
>>> x, c, n, t = fsli_to_xcnt(f, s, l, i)
>>> x
array([1., 2., 3., 4., 5.])
>>> c
array([0, 1, 1, 0, 0])
>>> n
array([1, 1, 1, 1, 1])
>>> t
array([[-inf,  inf],
       [-inf,  inf],
       [-inf,  inf],
       [-inf,  inf],
       [-inf,  inf]])
surpyval.utils.xcn_to_fs(x: ArrayLike, c: ArrayLike | None = None, n: ArrayLike | None = None) → tuple[NDArray, NDArray]

Convert observed and right-censored xcn data to the fs format: one array of failure times and one of suspension (right-censored) times, each time repeated by its count.

Parameters:
  • x (array like) – The times.

  • c (array like, optional) – Censoring flags: 0 observed, 1 right-censored. Other values are dropped. Defaults to all observed.

  • n (array like, optional) – The count at each time, a whole number. Defaults to 1.

Returns:

f, s – The failure times and the suspension times.

Return type:

arrays

Notes

A two-column x (as xcnt_handler() returns for interval data) is accepted: its interval rows are dropped like any other non-0/1 flag, and every other row must have equal ends.

Raises:

ValueError – If c or n is not the same length as x, a count is not a non-negative whole number, or an observed or right censored row of a two-column x has different ends.

Examples

>>> from surpyval import xcn_to_fs
>>> xcn_to_fs([1, 2, 5], [0, 1, 0], [2, 1, 1])
(array([1, 1, 5]), array([2]))
surpyval.utils.fs_to_xrd(f: ArrayLike, s: ArrayLike) → tuple[NDArray, NDArray, NDArray]

Converts the fs format to the xrd format.

Parameters:
  • f (array) – array of values for which the failure/death was observed

  • s (array) – array of right censored observation values

Returns:

  • x (array) – sorted array of values of variable for which observations were made.

  • r (array) – array of count of units/people at risk at time x (including if it had an event at ‘x’).

  • d (array) – array of the count of failures/deaths at each time x.

Examples

>>> from surpyval import fs_to_xrd
>>> f = [1, 4, 5]
>>> s = [2, 3]
>>> x, r, d = fs_to_xrd(f, s)
>>> x
array([1., 2., 3., 4., 5.])
>>> r
array([5, 4, 3, 2, 1])
>>> d
array([1, 0, 0, 1, 1])
surpyval.utils.xcnt_to_xrd(x: ArrayLike, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, **kwargs: Any) → tuple[NDArray, NDArray, NDArray]

Converts the xcnt format to the xrd format: the distinct times, the number at risk at each and the number of deaths at each. The data is validated with xcnt_handler() first. Only observed and right censored rows without right truncation can be converted; left truncation is allowed and sets when each item enters the risk set, under the (entry, exit] convention.

Parameters:
  • x (array) – array of values of variable for which observations were made.

  • c (array, optional (default: None)) – array of censoring values (0 or 1) corresponding to x. If None, an array of 0s is created corresponding to each x.

  • n (array, optional (default: None)) – array of count of observations at each x and with censoring c. If None, an array of ones is created.

  • t (array, optional (default: None)) – array of shape (?, 2) of truncation bounds; the right bounds must be infinite.

  • kwargs (keywords for truncation, tl and tr, used in place of) – t as in xcnt_handler()

Returns:

  • x (array) – sorted array of values of variable for which observations were made.

  • r (array) – array of count of units/people at risk at time x (including if it had an event at ‘x’).

  • d (array) – array of the count of failures/deaths at each time x.

Raises:

ValueError – If any row is left (-1) or interval (2) censored, or right truncated.

Examples

>>> x = np.array([1, 2, 3, 4, 5])
>>> c = np.array([0, 1, 1, 0, 0])
>>> n = np.array([1, 1, 1, 1, 1])
>>> x, r, d = xcnt_to_xrd(x, c, n)
>>> x
array([1., 2., 3., 4., 5.])
>>> r
array([5, 4, 3, 2, 1])
>>> d
array([1, 0, 0, 1, 1])
>>> # Using left truncated data: under the (entry, exit] convention a
>>> # subject entering exactly at an event time is not at risk for it.
>>> x = np.array([1, 2, 3, 4, 5])
>>> tl = np.array([0, 1, 2, 3, 4])
>>> x, r, d = xcnt_to_xrd(x, tl=tl)
>>> x
array([1., 2., 3., 4., 5.])
>>> r
array([1, 1, 1, 1, 1])
>>> d
array([1, 1, 1, 1, 1])
surpyval.utils.xrd_to_xcnt(x: ArrayLike, r: ArrayLike, d: ArrayLike) → tuple[NDArray, NDArray, NDArray, NDArray]

Converts the xrd format to the xcnt format. Each death becomes an observed row, and the items that leave the risk set without dying between x[j] and x[j + 1], r[j] - d[j] - r[j + 1], become right censored rows at x[j] (after the last time, r - d). The input is validated with xrd_handler(), so the times must be distinct; rows out of order are sorted first. The result has no truncation and no left or interval censoring.

Note: left truncation cannot be recovered from the xrd format because the at-risk count r collapses per-subject truncation times into a single scalar. Use xcnt format directly when left truncation is present.

Parameters:
  • x (array) – array of values of variable for which observations were made.

  • r (array) – array of at risk items at each value of x

  • d (array) – array of failures / deaths at each value of x

Returns:

  • x (array) – array of values of variable for which observations were made.

  • c (array) – array of censoring values (0 or 1) corresponding to x

  • n (array) – array of count of observations at each x and with censoring c

  • t (array) – array of values with shape (?, 2) with the left and right value of truncation (all [-inf, inf])

Raises:

ValueError – If xrd_handler() refuses the data, or if the risk set grows from one time to the next (late entry), which the xcnt output cannot represent.

Examples

>>> x = np.array([1, 2, 3, 4, 5])
>>> r = np.array([5, 4, 3, 2, 1])
>>> d = np.array([1, 0, 0, 1, 1])
>>> x, c, n, t = xrd_to_xcnt(x, r, d)
>>> x
array([1., 2., 3., 4., 5.])
>>> c
array([0, 1, 1, 0, 0])
>>> n
array([1, 1, 1, 1, 1])
>>> t
array([[-inf,  inf],
       [-inf,  inf],
       [-inf,  inf],
       [-inf,  inf],
       [-inf,  inf]])
surpyval.utils.round_sig(points: NDArray, sig: int = 2) → list

Round each value to sig significant figures (used for the tick labels of probability plots).

Parameters:
  • points (array or scalar) – The values to round. 0 (which has no leading digit) and the non-finite values are returned as they are.

  • sig (int, optional) – The number of significant figures. Defaults to 2.

Returns:

The rounded values: a list for an array, a scalar for a scalar.

Return type:

list or scalar

Examples

>>> import numpy as np
>>> from surpyval import round_sig
>>> round_sig(np.array([1234.5, 0.012345, -0.5678, 0.0]), 2)
[np.float64(1200.0), np.float64(0.012), np.float64(-0.57), np.float64(0.0)]
>>> round_sig(0)
np.int64(0)

Random draws follow one seeding rule (see Conventions); SurPyval turns every random_state or seed argument into a generator with:

surpyval.utils.rng.as_generator(random_state: Any = None) → Generator

A numpy Generator for random_state.

Parameters:

random_state (None, int, array_like of ints, numpy.random.SeedSequence,) – numpy.random.BitGenerator or numpy.random.Generator None seeds the generator from numpy’s global RNG, so np.random.seed makes the draw reproducible (and advances the global stream by one draw). Anything else is passed to numpy.random.default_rng: a seed gives a reproducible stream independent of the global state, and a Generator is used as is.

Return type:

numpy.random.Generator

Examples

>>> import numpy as np
>>> from surpyval.utils.rng import as_generator
>>> np.random.seed(0); a = as_generator().uniform(size=3)
>>> np.random.seed(0); b = as_generator().uniform(size=3)
>>> bool(np.array_equal(a, b))
True
>>> bool(np.array_equal(as_generator(1).uniform(size=3),
...                     np.random.default_rng(1).uniform(size=3)))
True

Lower-level helpers in surpyval.utils, used by the handlers above:

surpyval.utils.xcnt_sort(x: NDArray, c: NDArray, n: NDArray, t: NDArray) → tuple[NDArray, NDArray, NDArray, NDArray]

Sort xcnt arrays by x (the interval midpoint for 2-D x), breaking ties by the lower truncation bound and then by the censoring flag (so at a tied time left-censored rows come first, then observed, then right-censored, then interval-censored).

Returns:

x, c, n, t – The same arrays, reordered together.

Return type:

arrays

surpyval.utils.group_xcnt(x: NDArray, c: NDArray, n: NDArray, t: NDArray) → tuple[NDArray, NDArray, NDArray, NDArray]

Collapse identical (x, c, t) rows, summing their counts.

Takes and returns the four xcnt arrays (t of shape (k, 2)); each group is represented by its first row, in order of first appearance. Rows containing NaN are never merged.

It sorts and counts with np.bincount: a walk over every observation in Python (a triple-nested defaultdict) is O(N) but with a very large constant – roughly 13 microseconds per observation, 94% of a 50,000-point Normal fit and two seconds at 100,000 points.

The group order is x-major, which is subtler than it looks: outer by first appearance of x, then of (x, c) within it, then of the full key. xcnt_sort runs immediately after this and re-sorts on c, t.min(axis=1) and x – but it is a stable sort, so rows tying on all three of those keep whatever order arrived. Rows sharing an x and c with different tr but an equal t.min() are exactly such a tie, so a plain sorted np.unique would silently reorder them; the lexsort below keeps the nesting.

When nothing needs grouping the inputs are returned as they are, rather than copied. Callers inside the package sort immediately afterwards, which copies, so nothing aliases in practice.

surpyval.utils.coerce_xcnt_x(x: ArrayLike) → NDArray

Coerce the x variable of xcnt-format data into a float numpy array.

Accepts a scalar (a single observation), a 1D array of event values, or a 2D array / list-of-pairs of [left, right] interval bounds. In a list (or tuple), any element that is itself a list, tuple or array is an interval row, and the scalar elements are rows with equal ends. Validates dimensionality, the absence of NaNs and the interval ordering (left <= right). Shared by the univariate (xcnt_handler) and recurrent (handle_xicn) handlers. Durations and dates are refused (see refuse_time_values).

surpyval.utils.format_truncation(t: ArrayLike | None, tl: ArrayLike | Number | None, tr: ArrayLike | Number | None, n_rows: int) → NDArray

Build the (n_rows, 2) truncation array from either a t matrix or separate tl/tr bounds (scalars, including 0-d arrays, broadcast to all rows). The default window is the whole real line [-inf, inf]. NaN bounds are refused. Shared by xcnt_handler and handle_xicn.

surpyval.utils.is_missing_event(value: Any) → bool

Whether a competing-risks event value marks no attributed cause – a censored observation. That is Python None or any missing value (NaN, pandas NA).

surpyval.utils.missing_events(values: NDArray) → NDArray

is_missing_event() of each element of a 1-D object array.

A per-element loop over is_missing_event is a noticeable part of a competing-risks fit at 1e5 rows (#515). pandas.isna of an object array checks each element as isna checks a scalar, and so agrees with is_missing_event for every label type – except an object that is array-like without being a numpy scalar, which the scalar isna converts to an array first. Such labels take the loop.

surpyval.utils.resolve_cr_censoring(e: ArrayLike, c: ArrayLike | None) → tuple[NDArray, NDArray]

Canonicalise a competing-risks event vector and its censoring flag.

A missing event value (None, NaN or pandas NA) marks a right-censored observation with no attributed cause; such values are canonicalised to None. When c is not supplied it is derived from the events – a missing event is censored (c = 1), an event present is observed (c = 0) – so competing-risks data may be given as (x, e) alone, without a separate censoring array.

Returns the canonicalised event array (object dtype, None for censored) and the censoring flag (unchanged if supplied, otherwise derived).

lifelines, scikit-survival and R’s cmprsk code competing-risks data as one integer column with 0 for a censored row. Such data read here as a cause called 0 and no censoring at all, and every incidence is wrong, silently (#486). So where c is not given, no label is missing and the labels are numbers including 0, this warns, saying how to convert the data; passing c says which rows are censored and silences it.

The functions above documented as surpyval.utils.<name> are what that namespace promises (its __all__). It also holds the input validators the fitters call (check_* and validate_* functions such as validate_coxph and validate_cr_inputs, and optional_column, validate_1d and validate_float_array). They are internal: their behaviour is covered by the fitters’ own documentation, and they may change without notice.

Abstract Base Classes

Every model derives from one of these, exported from surpyval; they define the minimal interface a model provides and are useful for isinstance checks and for writing new model types.

class surpyval.distribution.Distribution

Bases: ABC

Root abstract base class that every surpyval model inherits from.

The contract shared by all models – parametric, nonparametric, mixtures, compositions and degenerate models – is the survival interface:

  • sf() the survival (reliability) function

  • ff() the cumulative distribution (failure) function

Hf() is provided as a default derived from sf() and may be overridden by models with a closed form. Statistical summaries such as moment, entropy, random and to_dict are not part of this minimal contract; the ParametricDistribution and NonParametricDistribution subclasses add the ones appropriate to their model family.

Examples

It is abstract; every fitted model is one, whatever its family:

>>> import numpy as np
>>> from surpyval import KaplanMeier, Weibull
>>> from surpyval.distribution import Distribution
>>> x = np.array([1, 2, 3, 4, 5])
>>> models = [Weibull.fit(x), KaplanMeier.fit(x)]
>>> all(isinstance(m, Distribution) for m in models)
True
>>> [m.sf([2.5]).round(4) for m in models]
[array([0.609]), array([0.6])]
Hf(x: ArrayLike, *args: Any, **kwargs: Any) → ArrayLike

The cumulative hazard, -log sf(x) unless overridden.

abstractmethod ff(x: ArrayLike, *args: Any, **kwargs: Any) → ArrayLike

The failure (cumulative distribution) function at x.

abstractmethod sf(x: ArrayLike, *args: Any, **kwargs: Any) → ArrayLike

The survival (reliability) function at x.

class surpyval.distribution.ParametricDistribution

Bases: Distribution

A fully specified parametric model. In addition to the survival interface it supports random sampling and the standard statistical summaries (moments, entropy) and can be serialised with to_dict.

Examples

It is abstract; a model from a parametric fitter is one:

>>> from surpyval import Weibull
>>> from surpyval.distribution import ParametricDistribution
>>> model = Weibull.from_params([10, 2])
>>> isinstance(model, ParametricDistribution)
True
>>> model.sf([5, 10]).round(4)
array([0.7788, 0.3679])
>>> round(float(model.moment(1)), 4)
8.8623
Hf(x: ArrayLike, *args: Any, **kwargs: Any) → ArrayLike

The cumulative hazard, -log sf(x) unless overridden.

abstractmethod entropy(*args: Any, **kwargs: Any) → ArrayLike

The differential entropy.

abstractmethod ff(x: ArrayLike, *args: Any, **kwargs: Any) → ArrayLike

The failure (cumulative distribution) function at x.

abstractmethod moment(n: int, *args: Any, **kwargs: Any) → ArrayLike

The n-th raw moment.

abstractmethod random(size: int | tuple[int, ...], *args: Any, **kwargs: Any) → ArrayLike

Draw random samples from the model. random_state (an int or a numpy.random.Generator) gives a draw of its own; None draws from numpy’s global stream, so np.random.seed reproduces it (Conventions, “Random draws and seeds”).

abstractmethod sf(x: ArrayLike, *args: Any, **kwargs: Any) → ArrayLike

The survival (reliability) function at x.

abstractmethod to_dict() → dict

Serialise the model to a plain dictionary.

class surpyval.distribution.NonParametricDistribution

Bases: Distribution

An empirical model produced by a nonparametric estimator (Kaplan-Meier, Nelson-Aalen, Fleming-Harrington or Turnbull). Adds random sampling from the fitted estimate to the survival interface.

Examples

It is abstract; a model from a non-parametric fitter is one:

>>> import numpy as np
>>> from surpyval import KaplanMeier
>>> from surpyval.distribution import NonParametricDistribution
>>> model = KaplanMeier.fit(np.array([1, 2, 3, 4, 5]))
>>> isinstance(model, NonParametricDistribution)
True
>>> model.sf([2.5, 4]).round(4)
array([0.6, 0.2])
Hf(x: ArrayLike, *args: Any, **kwargs: Any) → ArrayLike

The cumulative hazard, -log sf(x) unless overridden.

abstractmethod ff(x: ArrayLike, *args: Any, **kwargs: Any) → ArrayLike

The failure (cumulative distribution) function at x.

abstractmethod random(size: int, *args: Any, **kwargs: Any) → ArrayLike

Draw random samples from the fitted estimate, with random_state as for ParametricDistribution.random().

abstractmethod sf(x: ArrayLike, *args: Any, **kwargs: Any) → ArrayLike

The survival (reliability) function at x.

class surpyval.distribution.MultivariateDistribution

Bases: ABC

A jointly-specified model of several correlated event-time series.

Unlike Distribution, whose functions take a single random variable, the multivariate interface is evaluated at a point in several dimensions – x is array-like with one column per series. Concrete implementations (e.g. copula models) glue together ordinary univariate surpyval margins with a dependence structure, so the contract is the joint survival interface plus sampling:

  • cdf() the joint cumulative distribution P(X_1<=x_1, ...)

  • sf() the joint survival P(X_1>x_1, ...)

  • pdf() the joint density

  • random() draw correlated samples (one row per realisation)

Examples

It is abstract; a copula model is one:

>>> from surpyval import Weibull
>>> from surpyval.distribution import MultivariateDistribution
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> model = Clayton.from_params([2.0], margins)
>>> isinstance(model, MultivariateDistribution)
True
>>> model.sf([[5, 15], [10, 20]]).round(4)
array([0.624 , 0.2354])
abstractmethod cdf(x: ArrayLike, *args: Any, **kwargs: Any) → ArrayLike

The joint CDF at each row of x.

abstractmethod pdf(x: ArrayLike, *args: Any, **kwargs: Any) → ArrayLike

The joint density at each row of x.

abstractmethod random(size: int | tuple[int, ...], *args: Any, **kwargs: Any) → ArrayLike

Draw correlated samples, one row per realisation, with random_state as for ParametricDistribution.random().

abstractmethod sf(x: ArrayLike, *args: Any, **kwargs: Any) → ArrayLike

The joint survival function at each row of x.