Utilities
The data containers the fitters build from your arrays, the functions
that convert between SurPyval’s data formats, and the abstract base
classes every model derives from. The formats themselves – xcnt
(times, censoring flags, counts, truncation), xrd (times, numbers at
risk, deaths) and fsli (lists of failures, suspensions, left- and
interval-censored values) – are defined in Conventions, and
Data Wrangling Examples shows them in use.
Data Classes
SurpyvalData holds univariate xcnt data, split by censoring type;
every univariate fitter builds one, and fit_from_surpyval_data
accepts one directly. RecurrentEventData holds recurrent-event data
(each row an event of an item); build it with handle_xicn.
- class surpyval.utils.surpyval_data.SurpyvalData(x: ArrayLike | None = None, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, xl: ArrayLike | None = None, xr: ArrayLike | None = None, tl: ArrayLike | Number | None = None, tr: ArrayLike | Number | None = None, Z: ArrayLike | None = None, group_and_sort: bool = True, handle: bool = True)
Bases:
objectValidated univariate survival data in the
xcntformat, split by type of observation for the likelihoods.Every univariate fitter builds one from its arrays (with
xcnt_handler()); parametric fitters also accept one directly throughfit_from_surpyval_data. Besidesx,c,nandt(andtl/tr, the columns oft), it holds the observed (x_o,n_o), right censored (x_r,n_r), left censored (x_l,n_l) and interval (x_il,x_ir,n_i) rows separately. A censored row with a finite truncation bound on its censored side is stored as an interval ending at that bound, so the likelihood uses the probability of the part of the window it could have fallen in.Examples
>>> from surpyval import SurpyvalData >>> inf = float("inf") >>> data = SurpyvalData(x=[1, 2, 3, 4], c=[0, 1, 0, 1], tr=[9, 3, 9, inf]) >>> data.x_o, data.x_r (array([1., 3.]), array([4.])) >>> data.x_il, data.x_ir (array([2.]), array([3.]))
Initialize a SurpyvalData instance for survival analysis.
Validates, sorts, and stores survival data in the xcnt format. Supports uncensored, right/left/interval-censored, and truncated observations. Can convert to xrd format and select subsets by censoring type.
- Parameters:
x (array-like, optional) – The primary data array of failure/event times. When c is 2 the corresponding x entry is a 2-element array [left, right].
c (array-like, optional) – Censoring flags for each value in x: * 0 = uncensored * 1 = right censored * -1 = left censored * 2 = interval censored
n (array-like, optional) – Number of occurrences for each value in x (positive whole numbers).
t (array-like, optional) – 2D array of truncation bounds [left, right] for each value in x.
xl (array-like, optional) – Left interval bounds for interval censored data. Cannot be used with ‘x’. Must be paired with ‘xr’.
xr (array-like, optional) – Right interval bounds for interval censored data. Cannot be used with ‘x’. Must be paired with ‘xl’.
tl (array-like or scalar, optional) – Left truncation bounds. Cannot be used with ‘t’. May be given alone (the right bound is then infinite) or with ‘tr’.
tr (array-like or scalar, optional) – Right truncation bounds. Cannot be used with ‘t’. May be given alone (the left bound is then minus infinity) or with ‘tl’.
Z (array-like, optional) – Covariates, one row per value of x (see also
add_covariates()). GivingZturnsgroup_and_sortoff, so the rows stay aligned with x.group_and_sort (bool, default=True) – Whether to group and sort the data. Set False when using covariates to maintain data order.
handle (bool, default=True) – Whether to validate and process the input data. Set False for pre-validated data:
x,c,nand a two-columntexactly asxcnt_handler()returns them.
Examples
Basic usage with uncensored data: >>> x = np.array([1, 2, 3]) >>> data = SurpyvalData(x)
Right censored data: >>> x = np.array([1, 2, 3]) >>> c = np.array([0, 1, 1]) # 2 and 3 are censored >>> data = SurpyvalData(x, c)
Interval censored data: >>> xl = np.array([1, 2, 3]) >>> xr = np.array([2, 3, 4]) >>> data = SurpyvalData(xl=xl, xr=xr)
Interval Censored with nested 2 arrays: >>> x = [1, 2, [2, 5], 3, 6] >>> c = [0, 1, 2, 0, 0] >>> data = SurpyvalData(x=x, c=c)
With truncation: >>> x = np.array([1, 2, 3]) >>> t = np.array([[0, 5], [0, 5], [0, 5]]) >>> data = SurpyvalData(x, t=t)
- add_covariates(Z: ArrayLike) None
Method to add covariates to the data. The covariates are stored in the Z attribute of the object. When doing regression survival analysis this method allows for the covariates to be added to the data in a consistent manner that also allows for the data to be converted to be passed to the fitters.
- Parameters:
Z (numpy.ndarray) – The covariate array.
- classmethod from_json(source: str | Path) SurpyvalData
Create SurpyvalData instance from JSON text or a file.
- Parameters:
source (str | Path) – Pass a
pathlib.Pathto load from a file; astris always parsed as JSON text, sofrom_json("data.json")fails even thoughto_json("data.json")writes that file. The data is not re-validated.- Returns:
New instance created from JSON data
- Return type:
Examples
>>> from surpyval import SurpyvalData >>> data = SurpyvalData(x=[1, 2, 2, 5], c=[0, 0, 0, 1]) >>> restored = SurpyvalData.from_json(data.to_json()) >>> restored.x, restored.n (array([1., 2., 5.]), array([1, 2, 1]))
- to_json(filepath: str | Path | None = None) str | None
Serialize SurpyvalData to JSON format.
The JSON is strict: the infinite truncation bounds (and any other non-finite value) are written as
nulland recorded under"non_finite", whichfrom_json()reads back (see Saving and Loading Models).- Parameters:
filepath (str | Path, optional) – If provided, saves JSON to this file path
- Returns:
JSON string if no filepath provided, None if saved to file
- Return type:
str | None
- to_xrd(estimator: str = 'Nelson-Aalen') tuple
Converts the data into the xrd format. Observed and right censored data without right truncation is converted exactly with
xcnt_to_xrd(). If the data has right truncated observations or left or interval censored observations, the Turnbull estimator is fitted and its (possibly fractional) numbers at risk and deaths are returned; theestimatorparameter is used only in that case.- Parameters:
estimator (str, optional) – The method for estimation if data requires the use of the Turnbull estimator to convert to xrd, defaults to “Nelson-Aalen”.
- Returns:
The xrd data,
(x, r, d).- Return type:
tuple
- class surpyval.utils.recurrent_event_data.RecurrentEventData(x: ArrayLike, i: ArrayLike, c: ArrayLike, n: ArrayLike, e: ArrayLike | None = None, tl: ArrayLike | None = None, tr: ArrayLike | None = None)
Bases:
objectA class to handle and manipulate recurrent event data. Recurrent events are those that can occur more than once for each subject or item.
The recurrent fitters build one from their
x,i,c,narrays withsurpyval.handle_xicn, which validates the input; build one that way to pass to a fitter’sfit_from_recurrent_data. The constructor itself does no validation.Examples
>>> import numpy as np >>> from surpyval import RecurrentEventData >>> x = np.array([1, 2, 3, 4, 5, 1, 2, 3, 4, 5]) >>> c = np.array([0, 0, 1, 1, 1, 0, 0, 0, 0, 1]) >>> n = np.array([1, 1, 1, 1, 1, 1, 1, 1, 1, 1]) >>> i = np.array([1, 1, 1, 1, 1, 2, 2, 2, 2, 2]) >>> data = RecurrentEventData(x, i, c, n) >>> data.to_xrd() (array([1, 2, 3, 4, 5]), array([2, 2, 2, 2, 2]), array([2, 2, 1, 1, 0])) >>> data[0:2] RecurrentEventData( x=[1 2], i=[1 1], c=[0 0], n=[1 1] ) >>> data.get_times_to_first_events() SurpyvalData( x=array([1.]), c=array([0]), n=array([2]), t=array([[-inf, inf]]) ) >>> data.get_interarrival_times() array([1, 1, 1, 1, 1, 1, 1, 1, 1, 1])
- property event_types: list
The distinct event types (marks) present in the data, excluding the
Nonemark used for censored / end-of-observation rows, in the order ofordered_labels(mixedstrandintmarks are ordered too). Returns an empty list when the data carries no marks.
- get_events_for_item(item: Any) tuple[NDArray, NDArray, NDArray]
Get all events for a specific item or subject.
- Parameters:
item (int or str) – The id of the item or subject.
- Returns:
A tuple containing event times, censoring information and frequencies for the specified item.
- Return type:
tuple
- get_interarrival_times() NDArray
Finds the interarrival times between events for each item. The class assumes that the time of the event is cumulative, sometimes it is necessary to know the interarrival times of the events. This method returns the interarrival times for each item. It is aligned with the items attribute.
- Returns:
An array of interarrival times.
- Return type:
numpy.ndarray
- get_previous_x(min_x: float = 0) NDArray
Finds the previous event time for each event. This is useful for calculating the time since the last event. This method returns the previous event time for each event. It is aligned with the items attribute.
- Parameters:
min_x (float, optional) – Fallback minimum for the first event of each item. The item’s left truncation bound is used instead when it is greater, so a delayed-entry item’s first interval begins at its entry time.
- Returns:
An array of previous event times.
- Return type:
numpy.ndarray
- get_right_truncation_close() tuple[NDArray, NDArray, NDArray]
Per-item integration bounds for the NHPP likelihood’s right window-close.
The NHPP integral runs from each item’s entry time (its left truncation
tl, handled byget_previous_x()) to the time its observation window closes. Historically that close was only known from an explicit right-censoring (c=1) row, so the integral stopped at the item’s last recorded time. When an item instead carries a finite right-truncation timetrthe window closes there, and the integral must be extended from the last in-window time out totr.Returns three aligned arrays
(x_last, x_close, rep_idx)with one entry per item whosetris finite:x_lastis the item’s last in-window time (its last event or right-censoring row),x_closeis itstr, andrep_idxis a representative row index for the item (so per-item covariatesZcan be gathered). Items with the defaulttr = infare omitted, so untruncated data yields empty arrays and contributes nothing to the integral. Addingcif(x_last) - cif(x_close)to the log-likelihood therefore extends the telescoped integral totr(and is exactly zero when ac=1row already sits attr, so the two ways of closing the window never double-count).- Returns:
(x_last, x_close, rep_idx)as described above.- Return type:
tuple of numpy.ndarray
- get_times_to_first_events() SurpyvalData
Get the times to the first events for each item or subject. In the estimation of recurrent or renewal events it can be helpful to know the distribution of the times to the first event per item. This method returns the times to the first events for each item. It is aligned with the items attribute.
- Returns:
A transformed dataset containing times to the first events.
- Return type:
- item_observation_windows() tuple[NDArray, NDArray]
The
(entry, exit)observation window of each item, aligned withitems.The entry is the item’s left-truncation bound
tl(delayed entry); with no truncation it is -inf, so the item is at risk from the start. The exit is the item’s last recorded time (its last event or its end-of-observationc=1row) or, when the item carries a finite right-truncation timetr, thattr: observation of the item ends there, exactly as if it had an end-of-observation row attr. This is the same window-close the NHPP likelihoods integrate to (seeget_right_truncation_close()).- Returns:
(entry, exit), one value per item.- Return type:
tuple of numpy.ndarray
- item_rows() tuple[NDArray, NDArray]
The first row of each item and the item of each row, as positions in
items(which is sorted).- Returns:
(first, inverse):first[k]is the index of itemitems[k]’s first row, andinverse[j]the position initemsof rowj’s item.- Return type:
tuple of numpy.ndarray
Examples
>>> from surpyval import handle_xicn >>> data = handle_xicn([1, 2, 3, 4], i=[2, 1, 2, 1]) >>> data.items, data.i ([1, 2], array([1, 1, 2, 2])) >>> data.item_rows() (array([0, 2]), array([0, 0, 1, 1]))
- split_for_nhpp_likelihood() dict
Split the data into the pieces every NHPP-style likelihood needs: per-censoring-type event and previous-event times, the right-truncation window close, and the row masks (so regression variants can gather covariates the same way). This preamble used to be copy-pasted between the plain and proportional-intensity NHPP fitters and drifted (#288); it now lives here once (#296).
- to_cause_specific_xrd(cause: Any) tuple[NDArray, NDArray, NDArray]
Convert the recurrent event data to xrd format for a single event type (cause). The at-risk set
ris shared across all causes (an item remains at risk for every cause until it leaves observation); only the event countdis restricted to the requested cause.- Parameters:
cause (object) – The event type to compute the cause-specific counts for. Must be one of
self.event_types.- Returns:
A tuple
(x_unique, r, d_cause)whered_causecounts only events of the requested cause andris the shared at-risk set.- Return type:
tuple
- to_xrd(estimator: str = 'Nelson-Aalen') tuple[NDArray, NDArray, NDArray]
Convert the recurrent event data to xrd format.
- Parameters:
estimator (str, optional) – The estimator to use, defaults to “Nelson-Aalen”.
- Returns:
A tuple containing unique event times, risk set sizes, and the event counts.
- Return type:
tuple
Data Wrangling Utilities
Validate data in one format and convert it to another. Each of these is
importable directly from surpyval.
- surpyval.utils.xcnt_handler(x: ArrayLike | None = None, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, xl: ArrayLike | None = None, xr: ArrayLike | None = None, tl: ArrayLike | Number | None = None, tr: ArrayLike | Number | None = None, group_and_sort: bool = True) tuple[NDArray, NDArray, NDArray, NDArray]
Main handler that ensures any input to a surpyval fitter meets the requirements to be used in one of the parametric or nonparametric fitters.
It converts the inputs to numpy arrays and checks them:
xis not empty and has no NaN (a scalar is one observation);cholds only -1, 0 and 1 (and 2 for a two-columnx);nholds positive whole numbers; the truncation bounds have no NaN and the window of each row has its left bound below its right bound. For two-columnx, a row with equal values is not an interval, a row with different values must be flagged 2 (whencis given), and an infinite end turns the row into one-sided censoring:[v, inf]becomes right censored atvand[-inf, v]left censored atv. If no interval is left,xis returned with one column.Each row must then fit its truncation window
(tl, tr]: a single value lies strictly abovetland at or belowtr– strictly belowtrif it is right censored – whetherxhas one column or two; an interval[xl, xr]needstl <= xlandxr <= tr. Identical rows are then merged (their counts summed) and the rows sorted withxcnt_sort().- Parameters:
x (array) – array of values of variable for which observations were made.
c (array, optional (default: None)) – array of censoring values (-1, 0, 1, 2) corresponding to x
n (array, optional (default: None)) – array of count of observations at each x and with censoring c
t (array, optional (default: None)) – array of values with shape (?, 2) with the left and right value of truncation
xl (array or scalar, optional (default: None)) – array of the values of the left interval of interval censored data. Cannot be used with ‘x’ parameter, must be used with the ‘xr’ parameter
xr (array or scalar, optional (default: None)) – array of the values of the right interval of interval censored data. Cannot be used with ‘x’ parameter, must be used with the ‘xl’ parameter
tl (array or scalar, optional (default: None)) – array of values of the left value of truncation. If scalar, all values will be treated as left truncated by that value cannot be used with ‘t’ parameter but can be used with the ‘tr’ parameter
tr (array or scalar, optional (default: None)) – array of values of the right value of truncation. If scalar, all values will be treated as right truncated by that value cannot be used with ‘t’ parameter but can be used with the ‘tl’ parameter
group_and_sort (bool, optional (default: True)) – whether to group and sort the data. If False, the data will be returned in the order it was entered. This is useful for when validating survival data for which you also have covariates.
- Returns:
x (array) – sorted array of values of variable for which observations were made.
c (array) – array of censoring values (-1, 0, 1, 2) corresponding to output array x. If c was None, every row is observed (0), except that the rows of a two-column x with different values are interval censored (2).
n (array) – array of count of observations at output array x and with censoring c. If n was None, count array assumed to be all one observation.
t (array) – array of truncation values of observations at output array x and with censoring c.
- Raises:
ValueError – If the inputs break any of the rules above: for example
xandxl/xrboth given, emptyx, arrays of different lengths, a NaN inxor in a truncation bound, an unknown censoring flag, a count that is not a positive whole number, a value at or below its own left truncation, or a right censored value at its own right truncation.
Examples
>>> from surpyval import xcnt_handler >>> x = [1, 2, 3, 4, 5] >>> c = [0, 0, 1, 1, 1] >>> n = [1, 1, 1, 1, 1] >>> t = [[0, 6], [0, 6], [0, 6], [0, 6], [0, 6]] >>> xcnt_handler(x, c, n, t) (array([1., 2., 3., 4., 5.]), array([0, 0, 1, 1, 1]), array([1, 1, 1, 1, 1]), array([[0., 6.], [0., 6.], [0., 6.], [0., 6.], [0., 6.]])) >>> xcnt_handler(x, c, n, tl=0, tr=6) (array([1., 2., 3., 4., 5.]), array([0, 0, 1, 1, 1]), array([1, 1, 1, 1, 1]), array([[0., 6.], [0., 6.], [0., 6.], [0., 6.], [0., 6.]])) >>> xl = [1, 2, 3, 4, 5] >>> xr = [2, 3, 4, 5, 6] >>> xcnt_handler(xl=xl, xr=xr) (array([[1., 2.], [2., 3.], [3., 4.], [4., 5.], [5., 6.]]), array([2, 2, 2, 2, 2]), array([1, 1, 1, 1, 1]), array([[-inf, inf], [-inf, inf], [-inf, inf], [-inf, inf], [-inf, inf]]))
- surpyval.utils.fsli_handler(f: ArrayLike | None = None, s: ArrayLike | None = None, l: ArrayLike | None = None, i: ArrayLike | None = None) tuple[NDArray, NDArray, NDArray, NDArray]
Validate data in the
fsliformat: separate lists of failures, suspensions (right censored), left censored values and intervals. Any combination may be given, but at least one must hold data. Each is returned as a float array.- Parameters:
f (array-like, optional (default: None)) – array of values for which the failure/death was observed
s (array-like, optional (default: None)) – array of right censored observation values
l (array-like, optional (default: None)) – array of left censored observation values
i (array-like, optional (default: None)) – array of
[lower, upper]pairs, one per interval censored observation, withlower < upper
- Raises:
ValueError – If no data is given, if
f,sorlis not one-dimensional, ifiis not of shape(k, 2), if any value is NaN, or if an interval’s lower value is not below its upper value.- Returns:
f (array) – array of values for which the failure/death was observed that have been checked for correctness
s (array) – array of right censored observation values that have been checked for correctness
l (array) – array of left censored observation values that have been checked for correctness
i (array) – array of interval censored data that have been checked for correctness
Examples
>>> from surpyval import fsli_handler >>> f = [1, 2, 3, 4, 5, 6] >>> s = [1, 2, 3] >>> l = [4, 5, 6] >>> i = [[1, 2], [3, 4]] >>> fsli_handler(f, s, l, i) (array([1., 2., 3., 4., 5., 6.]), array([1., 2., 3.]), array([4., 5., 6.]), array([[1., 2.], [3., 4.]]))
- surpyval.utils.xrd_handler(x: ArrayLike, r: ArrayLike, d: ArrayLike) tuple[NDArray, NDArray, NDArray]
Takes a combination of ‘x’, ‘r’, and ‘d’ arrays and ensures that the data is feasible: the arrays are one-dimensional and the same length,
xhas no NaN and no repeated time,randdhold whole numbers (integer-valued floats such as[5.0, 4.0]are accepted), everyris at least one, nodis negative and nodexceeds itsr.xrd data lists each distinct time once, with the number at risk and the number of deaths at that time. Each
(x, r, d)row is therefore self-contained, so rows given out of order are sorted byx(carrying theirranddwith them) rather than refused. A repeated time is refused: there is no unambiguous way to merge two rows that each claim to be the risk set at that time.Does not check that
rdecreases, as it can grow when there is left truncation (late entry).- Parameters:
x (array) – array of values of variable for which observations were made.
r (array) – array of at risk items at each value of x
d (array) – array of failures / deaths at each value of x
- Returns:
x (array) – array of values of variable for which observations were made, in increasing order.
r (array) – array of at risk items at each value of x
d (array) – array of failures / deaths at each value of x
- Raises:
ValueError – If any of the rules above is broken.
Examples
>>> from surpyval import xrd_handler >>> x = [1, 2, 3, 4, 5] >>> r = [5, 4, 3, 2, 1] >>> d = [1, 1, 1, 1, 1] >>> x, r, d = xrd_handler(x, r, d) >>> x array([1., 2., 3., 4., 5.]) >>> r array([5, 4, 3, 2, 1]) >>> d array([1, 1, 1, 1, 1])
Rows out of order are sorted together:
>>> xrd_handler([3, 1, 2], [2, 5, 4], [1, 1, 1]) (array([1., 2., 3.]), array([5, 4, 2]), array([1, 1, 1]))
- surpyval.utils.recurrent_utils.handle_xicn(x: ArrayLike, i: ArrayLike | None = None, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, tl: ArrayLike | None = None, tr: ArrayLike | None = None, Z: ArrayLike | dict | None = None, as_recurrent_data: Literal[True] = True, windows: dict | None = None, e: ArrayLike | None = None) RecurrentEventData
- surpyval.utils.recurrent_utils.handle_xicn(x: ArrayLike, i: ArrayLike | None = None, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, tl: ArrayLike | None = None, tr: ArrayLike | None = None, Z: ArrayLike | dict | None = None, *, as_recurrent_data: Literal[False], windows: dict | None = None, e: ArrayLike | None = None) tuple[NDArray, NDArray, NDArray, NDArray]
Validate recurrent-event data given as arrays and assemble it into a
RecurrentEventData, the object every recurrent fitter’sfit_from_recurrent_datatakes.Each row is one event (or the right-censored end of observation) of the item named in
i, at timexmeasured from the start of that item’s life.- Parameters:
x (array like) – The event times (a 2-D
[left, right]row for an interval-censored count).i (array like, optional) – The item each row belongs to. Defaults to one item.
c (array like, optional) – Censoring flags: 0 an observed event, 1 the right-censored end of observation, -1 left-censored and 2 interval-censored counts. Defaults to all observed. Rows are sorted by item and time; an end-of-observation row at the same time as an event goes after it, whatever the input order.
n (array like, optional) – The number of events in each row. Defaults to 1.
t (array like, optional) – (N, 2) truncation bounds per row. Use
tl/trfor per-item bounds instead.tl (array like or scalar, optional) – Left-truncation (start of observation) and right-truncation (end of observation) times. An item with both a
c=1row and a finitetrmust have them at the same time.tr (array like or scalar, optional) – Left-truncation (start of observation) and right-truncation (end of observation) times. An item with both a
c=1row and a finitetrmust have them at the same time.Z (array like or dict, optional) – Covariates: one row per row of
x(the same on every row of an item: covariates are per item), or a{item: covariates}mapping applied to every row of that item.as_recurrent_data (bool, optional) – If
True(the default) return aRecurrentEventData; otherwise return the validated(x, i, c, n)arrays.windows (dict, optional) – Gapped observation:
{item: [(start, end), ...]}. Every row must then be an observed event; the windows supply the censoring rows. Not combinable witht/tl/tr,Zore.e (array like, optional) – The event type (mark) of each row, for the cause-specific models; a missing value marks a row with no cause (such as the censoring row).
- Returns:
The assembled data, or
(x, i, c, n).- Return type:
RecurrentEventData or tuple
Examples
>>> from surpyval import handle_xicn >>> data = handle_xicn([2, 5, 8, 3, 9], i=[1, 1, 1, 2, 2], ... c=[0, 0, 1, 0, 1]) >>> data.items [1, 2]
- surpyval.utils.fs_to_xcnt(f: ArrayLike | None = None, s: ArrayLike | None = None) tuple[NDArray, NDArray, NDArray, NDArray]
Convert failure (
f) and suspension (right-censored,s) times to thexcntformat, counting repeated times; seefsl_to_xcnt().- Parameters:
f (array like, optional) – Observed failure times.
s (array like, optional) – Right-censored (suspension) times.
- Returns:
x, c, n, t – The distinct times, censoring flags (0 or 1), counts and (untruncated) truncation bounds, sorted.
- Return type:
arrays
Examples
>>> from surpyval import fs_to_xcnt >>> x, c, n, t = fs_to_xcnt([1, 3, 3, 7], [5, 9]) >>> x array([1., 3., 5., 7., 9.]) >>> c, n (array([0, 0, 1, 0, 1]), array([1, 2, 1, 1, 1]))
- surpyval.utils.fsl_to_xcnt(f: ArrayLike | None = None, s: ArrayLike | None = None, l: ArrayLike | None = None) tuple[NDArray, NDArray, NDArray, NDArray]
Convert failure (
f), suspension (right-censored,s) and left-censored (l) times to thexcntformat, counting repeated times.- Parameters:
f (array like, optional) – Observed failure times.
s (array like, optional) – Right-censored (suspension) times.
l (array like, optional) – Left-censored times.
- Returns:
x, c, n, t – The distinct times, censoring flags, counts and (untruncated) truncation bounds, sorted.
- Return type:
arrays
Examples
>>> from surpyval import fsl_to_xcnt >>> x, c, n, t = fsl_to_xcnt([4, 6], [8], [2]) >>> x, c, n (array([2, 4, 6, 8]), array([-1, 0, 0, 1]), array([1, 1, 1, 1]))
- surpyval.utils.fsli_to_xcnt(f: ArrayLike | None = None, s: ArrayLike | None = None, l: ArrayLike | None = None, i: ArrayLike | None = None) tuple[NDArray, NDArray, NDArray, NDArray]
Converts the fsli format to the xcnt format, so that the data can be passed to one of the parametric or nonparametric fitters. The inputs are validated with
fsli_handler(). Repeated values are counted inn. When there are intervals,xis returned with two columns, the other rows repeating their value.- Parameters:
f (array) – array of values for which the failure/death was observed
s (array) – array of right censored observation values
l (array) – array of left censored observation values
i (array) – array of
[lower, upper]pairs of interval censored data
- Returns:
x (array) – sorted array of values of variable for which observations were made.
c (array) – array of censoring values (-1, 0, 1, 2) corresponding to output array x.
n (array) – array of count of observations at output array x and with censoring c.
t (ndarray) – ndarray of truncation values of observations at output array x and with censoring c.
Examples
>>> from surpyval import fsli_to_xcnt >>> f = [1, 4, 5] >>> s = [2, 3] >>> l = [] >>> i = [] >>> x, c, n, t = fsli_to_xcnt(f, s, l, i) >>> x array([1., 2., 3., 4., 5.]) >>> c array([0, 1, 1, 0, 0]) >>> n array([1, 1, 1, 1, 1]) >>> t array([[-inf, inf], [-inf, inf], [-inf, inf], [-inf, inf], [-inf, inf]])
- surpyval.utils.xcn_to_fs(x: ArrayLike, c: ArrayLike | None = None, n: ArrayLike | None = None) tuple[NDArray, NDArray]
Convert observed and right-censored
xcndata to thefsformat: one array of failure times and one of suspension (right-censored) times, each time repeated by its count.- Parameters:
x (array like) – The times.
c (array like, optional) – Censoring flags: 0 observed, 1 right-censored. Other values are dropped. Defaults to all observed.
n (array like, optional) – The count at each time, a whole number. Defaults to 1.
- Returns:
f, s – The failure times and the suspension times.
- Return type:
arrays
Notes
A two-column
x(asxcnt_handler()returns for interval data) is accepted: its interval rows are dropped like any other non-0/1 flag, and every other row must have equal ends.- Raises:
ValueError – If
cornis not the same length asx, a count is not a non-negative whole number, or an observed or right censored row of a two-columnxhas different ends.
Examples
>>> from surpyval import xcn_to_fs >>> xcn_to_fs([1, 2, 5], [0, 1, 0], [2, 1, 1]) (array([1, 1, 5]), array([2]))
- surpyval.utils.fs_to_xrd(f: ArrayLike, s: ArrayLike) tuple[NDArray, NDArray, NDArray]
Converts the fs format to the xrd format.
- Parameters:
f (array) – array of values for which the failure/death was observed
s (array) – array of right censored observation values
- Returns:
x (array) – sorted array of values of variable for which observations were made.
r (array) – array of count of units/people at risk at time x (including if it had an event at ‘x’).
d (array) – array of the count of failures/deaths at each time x.
Examples
>>> from surpyval import fs_to_xrd >>> f = [1, 4, 5] >>> s = [2, 3] >>> x, r, d = fs_to_xrd(f, s) >>> x array([1., 2., 3., 4., 5.]) >>> r array([5, 4, 3, 2, 1]) >>> d array([1, 0, 0, 1, 1])
- surpyval.utils.xcnt_to_xrd(x: ArrayLike, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, **kwargs: Any) tuple[NDArray, NDArray, NDArray]
Converts the xcnt format to the xrd format: the distinct times, the number at risk at each and the number of deaths at each. The data is validated with
xcnt_handler()first. Only observed and right censored rows without right truncation can be converted; left truncation is allowed and sets when each item enters the risk set, under the (entry, exit] convention.- Parameters:
x (array) – array of values of variable for which observations were made.
c (array, optional (default: None)) – array of censoring values (0 or 1) corresponding to x. If None, an array of 0s is created corresponding to each x.
n (array, optional (default: None)) – array of count of observations at each x and with censoring c. If None, an array of ones is created.
t (array, optional (default: None)) – array of shape (?, 2) of truncation bounds; the right bounds must be infinite.
kwargs (keywords for truncation,
tlandtr, used in place of) –tas inxcnt_handler()
- Returns:
x (array) – sorted array of values of variable for which observations were made.
r (array) – array of count of units/people at risk at time x (including if it had an event at ‘x’).
d (array) – array of the count of failures/deaths at each time x.
- Raises:
ValueError – If any row is left (-1) or interval (2) censored, or right truncated.
Examples
>>> x = np.array([1, 2, 3, 4, 5]) >>> c = np.array([0, 1, 1, 0, 0]) >>> n = np.array([1, 1, 1, 1, 1]) >>> x, r, d = xcnt_to_xrd(x, c, n) >>> x array([1., 2., 3., 4., 5.]) >>> r array([5, 4, 3, 2, 1]) >>> d array([1, 0, 0, 1, 1]) >>> # Using left truncated data: under the (entry, exit] convention a >>> # subject entering exactly at an event time is not at risk for it. >>> x = np.array([1, 2, 3, 4, 5]) >>> tl = np.array([0, 1, 2, 3, 4]) >>> x, r, d = xcnt_to_xrd(x, tl=tl) >>> x array([1., 2., 3., 4., 5.]) >>> r array([1, 1, 1, 1, 1]) >>> d array([1, 1, 1, 1, 1])
- surpyval.utils.xrd_to_xcnt(x: ArrayLike, r: ArrayLike, d: ArrayLike) tuple[NDArray, NDArray, NDArray, NDArray]
Converts the xrd format to the xcnt format. Each death becomes an observed row, and the items that leave the risk set without dying between
x[j]andx[j + 1],r[j] - d[j] - r[j + 1], become right censored rows atx[j](after the last time,r - d). The input is validated withxrd_handler(), so the times must be distinct; rows out of order are sorted first. The result has no truncation and no left or interval censoring.Note: left truncation cannot be recovered from the xrd format because the at-risk count r collapses per-subject truncation times into a single scalar. Use xcnt format directly when left truncation is present.
- Parameters:
x (array) – array of values of variable for which observations were made.
r (array) – array of at risk items at each value of x
d (array) – array of failures / deaths at each value of x
- Returns:
x (array) – array of values of variable for which observations were made.
c (array) – array of censoring values (0 or 1) corresponding to x
n (array) – array of count of observations at each x and with censoring c
t (array) – array of values with shape (?, 2) with the left and right value of truncation (all
[-inf, inf])
- Raises:
ValueError – If
xrd_handler()refuses the data, or if the risk set grows from one time to the next (late entry), which the xcnt output cannot represent.
Examples
>>> x = np.array([1, 2, 3, 4, 5]) >>> r = np.array([5, 4, 3, 2, 1]) >>> d = np.array([1, 0, 0, 1, 1]) >>> x, c, n, t = xrd_to_xcnt(x, r, d) >>> x array([1., 2., 3., 4., 5.]) >>> c array([0, 1, 1, 0, 0]) >>> n array([1, 1, 1, 1, 1]) >>> t array([[-inf, inf], [-inf, inf], [-inf, inf], [-inf, inf], [-inf, inf]])
- surpyval.utils.round_sig(points: NDArray, sig: int = 2) list
Round each value to
sigsignificant figures (used for the tick labels of probability plots).- Parameters:
points (array or scalar) – The values to round. 0 (which has no leading digit) and the non-finite values are returned as they are.
sig (int, optional) – The number of significant figures. Defaults to 2.
- Returns:
The rounded values: a list for an array, a scalar for a scalar.
- Return type:
list or scalar
Examples
>>> import numpy as np >>> from surpyval import round_sig >>> round_sig(np.array([1234.5, 0.012345, -0.5678, 0.0]), 2) [np.float64(1200.0), np.float64(0.012), np.float64(-0.57), np.float64(0.0)] >>> round_sig(0) np.int64(0)
Random draws follow one seeding rule (see Conventions); SurPyval
turns every random_state or seed argument into a generator with:
- surpyval.utils.rng.as_generator(random_state: Any = None) Generator
A numpy
Generatorforrandom_state.- Parameters:
random_state (None, int, array_like of ints, numpy.random.SeedSequence,) – numpy.random.BitGenerator or numpy.random.Generator
Noneseeds the generator from numpy’s global RNG, sonp.random.seedmakes the draw reproducible (and advances the global stream by one draw). Anything else is passed tonumpy.random.default_rng: a seed gives a reproducible stream independent of the global state, and aGeneratoris used as is.- Return type:
numpy.random.Generator
Examples
>>> import numpy as np >>> from surpyval.utils.rng import as_generator >>> np.random.seed(0); a = as_generator().uniform(size=3) >>> np.random.seed(0); b = as_generator().uniform(size=3) >>> bool(np.array_equal(a, b)) True >>> bool(np.array_equal(as_generator(1).uniform(size=3), ... np.random.default_rng(1).uniform(size=3))) True
Lower-level helpers in surpyval.utils, used by the handlers above:
- surpyval.utils.xcnt_sort(x: NDArray, c: NDArray, n: NDArray, t: NDArray) tuple[NDArray, NDArray, NDArray, NDArray]
Sort
xcntarrays byx(the interval midpoint for 2-Dx), breaking ties by the lower truncation bound and then by the censoring flag (so at a tied time left-censored rows come first, then observed, then right-censored, then interval-censored).- Returns:
x, c, n, t – The same arrays, reordered together.
- Return type:
arrays
- surpyval.utils.group_xcnt(x: NDArray, c: NDArray, n: NDArray, t: NDArray) tuple[NDArray, NDArray, NDArray, NDArray]
Collapse identical
(x, c, t)rows, summing their counts.Takes and returns the four
xcntarrays (tof shape(k, 2)); each group is represented by its first row, in order of first appearance. Rows containing NaN are never merged.It sorts and counts with
np.bincount: a walk over every observation in Python (a triple-nesteddefaultdict) is O(N) but with a very large constant – roughly 13 microseconds per observation, 94% of a 50,000-point Normal fit and two seconds at 100,000 points.The group order is x-major, which is subtler than it looks: outer by first appearance of
x, then of(x, c)within it, then of the full key.xcnt_sortruns immediately after this and re-sorts onc,t.min(axis=1)andx– but it is a stable sort, so rows tying on all three of those keep whatever order arrived. Rows sharing anxandcwith differenttrbut an equalt.min()are exactly such a tie, so a plain sortednp.uniquewould silently reorder them; the lexsort below keeps the nesting.When nothing needs grouping the inputs are returned as they are, rather than copied. Callers inside the package sort immediately afterwards, which copies, so nothing aliases in practice.
- surpyval.utils.coerce_xcnt_x(x: ArrayLike) NDArray
Coerce the
xvariable of xcnt-format data into a float numpy array.Accepts a scalar (a single observation), a 1D array of event values, or a 2D array / list-of-pairs of
[left, right]interval bounds. In a list (or tuple), any element that is itself a list, tuple or array is an interval row, and the scalar elements are rows with equal ends. Validates dimensionality, the absence of NaNs and the interval ordering (left <= right). Shared by the univariate (xcnt_handler) and recurrent (handle_xicn) handlers. Durations and dates are refused (seerefuse_time_values).
- surpyval.utils.format_truncation(t: ArrayLike | None, tl: ArrayLike | Number | None, tr: ArrayLike | Number | None, n_rows: int) NDArray
Build the
(n_rows, 2)truncation array from either atmatrix or separatetl/trbounds (scalars, including 0-d arrays, broadcast to all rows). The default window is the whole real line[-inf, inf]. NaN bounds are refused. Shared byxcnt_handlerandhandle_xicn.
- surpyval.utils.is_missing_event(value: Any) bool
Whether a competing-risks event value marks no attributed cause – a censored observation. That is Python
Noneor any missing value (NaN, pandasNA).
- surpyval.utils.missing_events(values: NDArray) NDArray
is_missing_event()of each element of a 1-D object array.A per-element loop over
is_missing_eventis a noticeable part of a competing-risks fit at 1e5 rows (#515).pandas.isnaof an object array checks each element asisnachecks a scalar, and so agrees withis_missing_eventfor every label type – except an object that is array-like without being a numpy scalar, which the scalarisnaconverts to an array first. Such labels take the loop.
- surpyval.utils.resolve_cr_censoring(e: ArrayLike, c: ArrayLike | None) tuple[NDArray, NDArray]
Canonicalise a competing-risks event vector and its censoring flag.
A missing event value (
None,NaNor pandasNA) marks a right-censored observation with no attributed cause; such values are canonicalised toNone. Whencis not supplied it is derived from the events – a missing event is censored (c = 1), an event present is observed (c = 0) – so competing-risks data may be given as(x, e)alone, without a separate censoring array.Returns the canonicalised event array (object dtype,
Nonefor censored) and the censoring flag (unchanged if supplied, otherwise derived).lifelines, scikit-survival and R’s
cmprskcode competing-risks data as one integer column with 0 for a censored row. Such data read here as a cause called 0 and no censoring at all, and every incidence is wrong, silently (#486). So wherecis not given, no label is missing and the labels are numbers including 0, this warns, saying how to convert the data; passingcsays which rows are censored and silences it.
The functions above documented as surpyval.utils.<name> are what that
namespace promises (its __all__). It also holds the input validators the fitters call
(check_* and validate_* functions such as validate_coxph and
validate_cr_inputs, and optional_column, validate_1d and
validate_float_array). They are internal: their
behaviour is covered by the fitters’ own documentation, and they may
change without notice.
Abstract Base Classes
Every model derives from one of these, exported from surpyval; they
define the minimal interface a model provides and are useful for
isinstance checks and for writing new model types.
- class surpyval.distribution.Distribution
Bases:
ABCRoot abstract base class that every surpyval model inherits from.
The contract shared by all models – parametric, nonparametric, mixtures, compositions and degenerate models – is the survival interface:
sf()the survival (reliability) functionff()the cumulative distribution (failure) function
Hf()is provided as a default derived fromsf()and may be overridden by models with a closed form. Statistical summaries such asmoment,entropy,randomandto_dictare not part of this minimal contract; theParametricDistributionandNonParametricDistributionsubclasses add the ones appropriate to their model family.Examples
It is abstract; every fitted model is one, whatever its family:
>>> import numpy as np >>> from surpyval import KaplanMeier, Weibull >>> from surpyval.distribution import Distribution >>> x = np.array([1, 2, 3, 4, 5]) >>> models = [Weibull.fit(x), KaplanMeier.fit(x)] >>> all(isinstance(m, Distribution) for m in models) True >>> [m.sf([2.5]).round(4) for m in models] [array([0.609]), array([0.6])]
- Hf(x: ArrayLike, *args: Any, **kwargs: Any) ArrayLike
The cumulative hazard,
-log sf(x)unless overridden.
- abstractmethod ff(x: ArrayLike, *args: Any, **kwargs: Any) ArrayLike
The failure (cumulative distribution) function at
x.
- abstractmethod sf(x: ArrayLike, *args: Any, **kwargs: Any) ArrayLike
The survival (reliability) function at
x.
- class surpyval.distribution.ParametricDistribution
Bases:
DistributionA fully specified parametric model. In addition to the survival interface it supports random sampling and the standard statistical summaries (moments, entropy) and can be serialised with
to_dict.Examples
It is abstract; a model from a parametric fitter is one:
>>> from surpyval import Weibull >>> from surpyval.distribution import ParametricDistribution >>> model = Weibull.from_params([10, 2]) >>> isinstance(model, ParametricDistribution) True >>> model.sf([5, 10]).round(4) array([0.7788, 0.3679]) >>> round(float(model.moment(1)), 4) 8.8623
- Hf(x: ArrayLike, *args: Any, **kwargs: Any) ArrayLike
The cumulative hazard,
-log sf(x)unless overridden.
- abstractmethod entropy(*args: Any, **kwargs: Any) ArrayLike
The differential entropy.
- abstractmethod ff(x: ArrayLike, *args: Any, **kwargs: Any) ArrayLike
The failure (cumulative distribution) function at
x.
- abstractmethod moment(n: int, *args: Any, **kwargs: Any) ArrayLike
The
n-th raw moment.
- abstractmethod random(size: int | tuple[int, ...], *args: Any, **kwargs: Any) ArrayLike
Draw random samples from the model.
random_state(an int or anumpy.random.Generator) gives a draw of its own;Nonedraws from numpy’s global stream, sonp.random.seedreproduces it (Conventions, “Random draws and seeds”).
- abstractmethod sf(x: ArrayLike, *args: Any, **kwargs: Any) ArrayLike
The survival (reliability) function at
x.
- abstractmethod to_dict() dict
Serialise the model to a plain dictionary.
- class surpyval.distribution.NonParametricDistribution
Bases:
DistributionAn empirical model produced by a nonparametric estimator (Kaplan-Meier, Nelson-Aalen, Fleming-Harrington or Turnbull). Adds random sampling from the fitted estimate to the survival interface.
Examples
It is abstract; a model from a non-parametric fitter is one:
>>> import numpy as np >>> from surpyval import KaplanMeier >>> from surpyval.distribution import NonParametricDistribution >>> model = KaplanMeier.fit(np.array([1, 2, 3, 4, 5])) >>> isinstance(model, NonParametricDistribution) True >>> model.sf([2.5, 4]).round(4) array([0.6, 0.2])
- Hf(x: ArrayLike, *args: Any, **kwargs: Any) ArrayLike
The cumulative hazard,
-log sf(x)unless overridden.
- abstractmethod ff(x: ArrayLike, *args: Any, **kwargs: Any) ArrayLike
The failure (cumulative distribution) function at
x.
- abstractmethod random(size: int, *args: Any, **kwargs: Any) ArrayLike
Draw random samples from the fitted estimate, with
random_stateas forParametricDistribution.random().
- abstractmethod sf(x: ArrayLike, *args: Any, **kwargs: Any) ArrayLike
The survival (reliability) function at
x.
- class surpyval.distribution.MultivariateDistribution
Bases:
ABCA jointly-specified model of several correlated event-time series.
Unlike
Distribution, whose functions take a single random variable, the multivariate interface is evaluated at a point in several dimensions –xis array-like with one column per series. Concrete implementations (e.g. copula models) glue together ordinary univariate surpyval margins with a dependence structure, so the contract is the joint survival interface plus sampling:cdf()the joint cumulative distributionP(X_1<=x_1, ...)sf()the joint survivalP(X_1>x_1, ...)pdf()the joint densityrandom()draw correlated samples (one row per realisation)
Examples
It is abstract; a copula model is one:
>>> from surpyval import Weibull >>> from surpyval.distribution import MultivariateDistribution >>> from surpyval.multivariate import Clayton >>> margins = [ ... Weibull.from_params([10, 2]), ... Weibull.from_params([20, 3]), ... ] >>> model = Clayton.from_params([2.0], margins) >>> isinstance(model, MultivariateDistribution) True >>> model.sf([[5, 15], [10, 20]]).round(4) array([0.624 , 0.2354])
- abstractmethod cdf(x: ArrayLike, *args: Any, **kwargs: Any) ArrayLike
The joint CDF at each row of
x.
- abstractmethod pdf(x: ArrayLike, *args: Any, **kwargs: Any) ArrayLike
The joint density at each row of
x.
- abstractmethod random(size: int | tuple[int, ...], *args: Any, **kwargs: Any) ArrayLike
Draw correlated samples, one row per realisation, with
random_stateas forParametricDistribution.random().
- abstractmethod sf(x: ArrayLike, *args: Any, **kwargs: Any) ArrayLike
The joint survival function at each row of
x.