Multivariate Modelling

Joint models for several correlated event-time series. A copula separates the problem into two independent choices: the marginal distribution of each dimension, which is any surpyval univariate distribution, and the dependence structure between them, which is the copula itself.

For a narrative introduction with worked examples, see Multivariate Modelling with SurPyval.

Note

surpyval.multivariate.Gumbel is the Gumbel copula and is a different object from surpyval.Gumbel, the univariate Gumbel distribution. The two are unrelated; import the copulas from surpyval.multivariate to keep them apart.

Copulas

Each copula below is exported as a ready-to-use instance – Clayton, Gumbel, Frank, Gaussian, Independence, Joe, AMH and StudentT – in the same way the univariate distributions are. The classes are documented here; the instances carry the same methods.

class surpyval.multivariate.parametric.copula.copula.Copula

Bases: object

Bivariate copula family.

Subclasses define cdf() (and, for speed/stability, may override the partial derivatives, dependence measures and sampler). The fitting, likelihood and default autograd-based derivatives live here.

cdf(u: Any, v: Any, *params: Any) → Any

The copula \(C(u, v)\); defined by each family.

closed_bounds: tuple = ()

Names of parameters whose finite bounds are themselves valid values (Gumbel’s theta = 1 is the independence copula). The fitter never reaches a bound, but from_params may be given one.

dependence_limits: dict = {}

The Frechet bounds the family reaches only as its parameter runs to a limit: +1 the comonotone copula (perfect positive dependence), -1 the countermonotone one, each mapped to that limit as the fit’s warning names it (see _warn_if_perfectly_dependent()). Empty for a family that is not known to reach either.

du(u: Any, v: Any, *params: Any) → Any

dC/du – the h-function \(P(V \le v \mid U = u)\).

The default differentiates cdf() with autograd; override it when a closed form is known. u and v broadcast against each other, and the result has their common shape.

dv(u: Any, v: Any, *params: Any) → Any

dC/dv, the h-function \(P(U \le u \mid V = v)\); u and v broadcast as for du().

fit(x: ArrayLike, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, margins: Any = None, how: str = 'IFM', xl: ArrayLike | None = None, xr: ArrayLike | None = None, init: ArrayLike | None = None, rotation: int = 0) → Any

Fit the copula and its margins to multivariate survival data.

Parameters:
  • x – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • c – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • n – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • t – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xl – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xr – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • margins (sequence of length D) – Either surpyval distribution classes (e.g. surpyval.Weibull) to be fitted, or already-fitted models exposing ff/df. Under "IFM" a fitted model is used as it is. Under "MLE" a fitted parametric model supplies the starting values and its configuration: it is re-estimated jointly with the copula with the same offset, limited-failure or zero-inflated option, and any parameters it was fitted with fixed stay at their values. A non-parametric margin (a class such as surpyval.KaplanMeier, or a fitted non-parametric model) can only be used with "IFM": it gives the semi-parametric pseudo-likelihood estimator, and the likelihood and criteria then compare copula families with the same margins only.

  • how ({"IFM", "MLE"}) – "IFM" (default) fits each margin independently then the single copula parameter (robust two-stage estimation). "MLE" jointly optimises copula parameter + margin parameters.

  • init (array like, optional) – Starting value of the copula parameter(s) for the search, one per entry of parameter_names, each strictly inside the family’s bounds. By default each family starts from its own guess (the built-in families match the empirical Kendall’s tau).

  • rotation ({0, 90, 180, 270}, optional) – Rotate the copula by this many degrees, in the convention of R’s VineCopula (families 13, 23, 33 for the Clayton, …): 180 is the survival copula, C(u, v) = u + v - 1 + C_0(1 - u, 1 - v), with the tail dependence moved to the other tail (a Clayton’s to the upper tail, a Gumbel’s or Joe’s to the lower); 90, C(u, v) = v - C_0(1 - u, v), and 270, C(u, v) = u - C_0(u, 1 - v), give negative dependence with the family’s shape, its tail in a corner where one series is short and the other long. The parameter keeps its own range (pyvinecopulib’s convention; VineCopula writes it negated for 90 and 270). Only the Clayton, Gumbel and Joe copulas are rotated; for a radially symmetric family (Frank, Gaussian, Student-t) 180 is the family itself and 90 its negative parameter. Default 0, the family as it is.

Returns:

The fitted model: the copula parameter(s) params and the fitted margins, with the joint sf/cdf/pdf, sampling, dependence measures and the likelihood-based log_likelihood/neg_ll/aic/bic.

Return type:

CopulaModel

Warns:

UserWarning – “No finite maximum” when the rows observed in both dimensions are perfectly dependent (Kendall’s tau of +-1) and the family reaches that dependence only as its parameter runs to a limit: the returned parameter is then meaningless.

Examples

Simulate from a Clayton copula with Weibull margins, then recover it:

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> model = Clayton.fit(X, margins=[Weibull, Weibull])
>>> model.params.round(3)
array([2.293])
>>> round(float(model.kendall_tau()), 3)
0.534
fit_from_df(df: Any, x_cols: list[str], c_cols: list[str] | None = None, n_col: str | None = None, xl_cols: list[str] | None = None, xr_cols: list[str] | None = None, tl_cols: list[str] | None = None, tr_cols: list[str] | None = None, **fit_options: Any) → Any

Fit the copula and its margins to the columns of a pandas.DataFrame.

Each argument names, per dimension, the columns fit() takes as arrays: x_cols=["a", "b"] reads the two series from columns a and b. The names are those of the univariate fit_from_df (Weibull.fit_from_df(df, x_col=..., c_col=...)) with _cols for a list of columns, one per dimension (principle 21); every other fit() option (margins, how, init) is passed to it unchanged.

Parameters:
  • df (pandas.DataFrame) – The data, one row per unit.

  • x_cols (list of str) – The column of each dimension’s values.

  • c_cols (list of str, optional) – The column of each dimension’s censoring flags. Defaults to every value observed.

  • n_col (str, optional) – The column of row counts.

  • xl_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • xr_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • tl_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • tr_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • **fit_options – Every other option of fit().

Returns:

The model fit() returns for the same arrays.

Return type:

CopulaModel

Examples

>>> import pandas as pd
>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> df = pd.DataFrame(X, columns=["pump", "motor"])
>>> model = Clayton.fit_from_df(
...     df, x_cols=["pump", "motor"], margins=[Weibull, Weibull]
... )
>>> model.params.round(3)
array([2.293])
from_params(params: Any, margins: Any, rotation: int = 0) → Any

Build a CopulaModel from a known parameter and margins, without fitting.

Parameters:
  • params (array like) – The copula parameter(s), e.g. [theta] (empty for the independence copula): one per entry of parameter_names, each strictly inside the family’s bounds (Gumbel’s theta = 1, the independence copula, is allowed too). Anything else raises ValueError.

  • margins (sequence of length 2) – Fitted (or from_params) univariate models, one per dimension, each exposing ff and df.

  • rotation ({0, 90, 180, 270}, optional) – The rotation of the copula, as for fit().

Returns:

The model, for evaluation and simulation.

Return type:

CopulaModel

Examples

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> model = Clayton.from_params([2.0], margins)
>>> round(float(model.kendall_tau()), 3)
0.5
kendall_tau(*params: float) → float

Kendall’s tau.

The default integrates \(\tau = 1 - 4 \int_0^1 \int_0^1 \frac{\partial C}{\partial u} \frac{\partial C}{\partial v} \, du \, dv\) by Gauss-Legendre quadrature (400 nodes per margin: accurate to about 1e-10 at a tau of 0.5 and 1e-8 at 0.8, the integrand sharpening along the diagonal as the dependence grows); families with a closed form override it. It used to estimate tau from 50 000 simulated pairs, with an error near 1e-3.

neg_ll(params: Any, dims: list, weights: NDArray) → float

The copula-stage negative log-likelihood (used by the fit).

pdf(u: Any, v: Any, *params: Any) → Any

d2C/du dv – the copula density. The default differentiates du() with autograd; u and v broadcast as for du().

rotatable: bool = False

the families whose rotations are new copulas (the asymmetric Archimedean ones).

Type:

Whether rotated() applies

rotated(rotation: int) → Copula

The family rotated by rotation degrees (0, 90, 180 or 270), in R’s VineCopula convention; see the rotation option of fit(), which uses it. rotated(0) is the family itself.

Examples

>>> from surpyval.multivariate import Clayton
>>> survival = Clayton.rotated(180)
>>> survival.tail_dependence(2.0)
(0.0, 0.7071067811865476)
>>> Clayton.rotated(90).kendall_tau(2.0)
-0.5
rotation: int = 0

0 for the family as it is.

Type:

The rotation of the copula in degrees (see rotated())

sample_uv(size: int, params: Any, random_state: int | None = None) → tuple[NDArray, NDArray]

Draw (u, v) pairs by conditional inversion of the h-function.

u is uniform; given u and a uniform w, v solves dC/du(u, v) = w (a CDF in v, hence monotone) by bisection. Override for families with a direct sampler (e.g. Gaussian).

spearman_rho(*params: float) → float

Spearman’s rho.

The default integrates \(\rho_S = 12 \int_0^1 \int_0^1 C(u, v) \, du \, dv - 3\) by Gauss-Legendre quadrature (400 nodes per margin, accurate to about 1e-11 for the built-in families); families with a closed form override it. It used to estimate rho from 50 000 simulated pairs, which was up to 5e-3 off (Clayton, Gumbel).

tail_dependence(*params: float) → tuple

Lower/upper tail-dependence coefficients (lambda_L, lambda_U).

Default (0.0, 0.0) (no tail dependence); families override.

class surpyval.multivariate.parametric.copula.archimedean.IndependenceCopula

Bases: Copula

The independence copula C(u, v) = u v (no parameter).

Fitting it fits the two margins separately; the joint survival is then the product of theirs. It is the baseline a dependent copula is compared against.

Examples

>>> import numpy as np
>>> from surpyval import Weibull
>>> from surpyval.multivariate import Independence
>>> rng = np.random.default_rng(0)
>>> x1 = 10 * rng.weibull(2, 100)
>>> x2 = 20 * rng.weibull(3, 100)
>>> model = Independence.fit([x1, x2], margins=[Weibull, Weibull])
>>> [m.params.round(3) for m in model.margins]
[array([10.763,  1.91 ]), array([20.847,  3.463])]
>>> model.kendall_tau()
0.0
>>> model.sf([[5, 15]]).round(4)
array([0.5762])
cdf(u: Any, v: Any, *params: Any) → Any

The copula \(C(u, v)\); defined by each family.

closed_bounds: tuple = ()

Names of parameters whose finite bounds are themselves valid values (Gumbel’s theta = 1 is the independence copula). The fitter never reaches a bound, but from_params may be given one.

dependence_limits: dict = {}

The Frechet bounds the family reaches only as its parameter runs to a limit: +1 the comonotone copula (perfect positive dependence), -1 the countermonotone one, each mapped to that limit as the fit’s warning names it (see _warn_if_perfectly_dependent()). Empty for a family that is not known to reach either.

du(u: Any, v: Any, *params: Any) → Any

dC/du – the h-function \(P(V \le v \mid U = u)\).

The default differentiates cdf() with autograd; override it when a closed form is known. u and v broadcast against each other, and the result has their common shape.

dv(u: Any, v: Any, *params: Any) → Any

dC/dv, the h-function \(P(U \le u \mid V = v)\); u and v broadcast as for du().

fit(x: ArrayLike, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, margins: Any = None, how: str = 'IFM', xl: ArrayLike | None = None, xr: ArrayLike | None = None, init: ArrayLike | None = None, rotation: int = 0) → Any

Fit the margins only (the independence copula has no parameter); arguments as for Copula.fit(), with how ignored (init, if given, must be empty; the copula is not rotated).

fit_from_df(df: Any, x_cols: list[str], c_cols: list[str] | None = None, n_col: str | None = None, xl_cols: list[str] | None = None, xr_cols: list[str] | None = None, tl_cols: list[str] | None = None, tr_cols: list[str] | None = None, **fit_options: Any) → Any

Fit the copula and its margins to the columns of a pandas.DataFrame.

Each argument names, per dimension, the columns fit() takes as arrays: x_cols=["a", "b"] reads the two series from columns a and b. The names are those of the univariate fit_from_df (Weibull.fit_from_df(df, x_col=..., c_col=...)) with _cols for a list of columns, one per dimension (principle 21); every other fit() option (margins, how, init) is passed to it unchanged.

Parameters:
  • df (pandas.DataFrame) – The data, one row per unit.

  • x_cols (list of str) – The column of each dimension’s values.

  • c_cols (list of str, optional) – The column of each dimension’s censoring flags. Defaults to every value observed.

  • n_col (str, optional) – The column of row counts.

  • xl_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • xr_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • tl_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • tr_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • **fit_options – Every other option of fit().

Returns:

The model fit() returns for the same arrays.

Return type:

CopulaModel

Examples

>>> import pandas as pd
>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> df = pd.DataFrame(X, columns=["pump", "motor"])
>>> model = Clayton.fit_from_df(
...     df, x_cols=["pump", "motor"], margins=[Weibull, Weibull]
... )
>>> model.params.round(3)
array([2.293])
from_params(params: Any, margins: Any, rotation: int = 0) → Any

Build a CopulaModel from a known parameter and margins, without fitting.

Parameters:
  • params (array like) – The copula parameter(s), e.g. [theta] (empty for the independence copula): one per entry of parameter_names, each strictly inside the family’s bounds (Gumbel’s theta = 1, the independence copula, is allowed too). Anything else raises ValueError.

  • margins (sequence of length 2) – Fitted (or from_params) univariate models, one per dimension, each exposing ff and df.

  • rotation ({0, 90, 180, 270}, optional) – The rotation of the copula, as for fit().

Returns:

The model, for evaluation and simulation.

Return type:

CopulaModel

Examples

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> model = Clayton.from_params([2.0], margins)
>>> round(float(model.kendall_tau()), 3)
0.5
kendall_tau(*params: float) → float

Kendall’s tau.

The default integrates \(\tau = 1 - 4 \int_0^1 \int_0^1 \frac{\partial C}{\partial u} \frac{\partial C}{\partial v} \, du \, dv\) by Gauss-Legendre quadrature (400 nodes per margin: accurate to about 1e-10 at a tau of 0.5 and 1e-8 at 0.8, the integrand sharpening along the diagonal as the dependence grows); families with a closed form override it. It used to estimate tau from 50 000 simulated pairs, with an error near 1e-3.

neg_ll(params: Any, dims: list, weights: NDArray) → float

The copula-stage negative log-likelihood (used by the fit).

pdf(u: Any, v: Any, *params: Any) → Any

d2C/du dv – the copula density. The default differentiates du() with autograd; u and v broadcast as for du().

rotatable: bool = False

the families whose rotations are new copulas (the asymmetric Archimedean ones).

Type:

Whether rotated() applies

rotated(rotation: int) → Copula

The family rotated by rotation degrees (0, 90, 180 or 270), in R’s VineCopula convention; see the rotation option of fit(), which uses it. rotated(0) is the family itself.

Examples

>>> from surpyval.multivariate import Clayton
>>> survival = Clayton.rotated(180)
>>> survival.tail_dependence(2.0)
(0.0, 0.7071067811865476)
>>> Clayton.rotated(90).kendall_tau(2.0)
-0.5
rotation: int = 0

0 for the family as it is.

Type:

The rotation of the copula in degrees (see rotated())

sample_uv(size: int, params: Any, random_state: int | None = None) → tuple[NDArray, NDArray]

Draw (u, v) pairs by conditional inversion of the h-function.

u is uniform; given u and a uniform w, v solves dC/du(u, v) = w (a CDF in v, hence monotone) by bisection. Override for families with a direct sampler (e.g. Gaussian).

spearman_rho(*params: float) → float

Spearman’s rho.

The default integrates \(\rho_S = 12 \int_0^1 \int_0^1 C(u, v) \, du \, dv - 3\) by Gauss-Legendre quadrature (400 nodes per margin, accurate to about 1e-11 for the built-in families); families with a closed form override it. It used to estimate rho from 50 000 simulated pairs, which was up to 5e-3 off (Clayton, Gumbel).

tail_dependence(*params: float) → tuple

Lower/upper tail-dependence coefficients (lambda_L, lambda_U).

Default (0.0, 0.0) (no tail dependence); families override.

class surpyval.multivariate.parametric.copula.archimedean.ClaytonCopula

Bases: Copula

Clayton copula (lower-tail dependence), theta > 0.

cdf(u: Any, v: Any, theta: Any) → Any

The copula \(C(u, v)\); defined by each family.

closed_bounds: tuple = ()

Names of parameters whose finite bounds are themselves valid values (Gumbel’s theta = 1 is the independence copula). The fitter never reaches a bound, but from_params may be given one.

dependence_limits: dict = {1: 'theta grows without bound'}

The Frechet bounds the family reaches only as its parameter runs to a limit: +1 the comonotone copula (perfect positive dependence), -1 the countermonotone one, each mapped to that limit as the fit’s warning names it (see _warn_if_perfectly_dependent()). Empty for a family that is not known to reach either.

du(u: Any, v: Any, theta: Any) → Any

dC/du – the h-function \(P(V \le v \mid U = u)\).

The default differentiates cdf() with autograd; override it when a closed form is known. u and v broadcast against each other, and the result has their common shape.

dv(u: Any, v: Any, theta: Any) → Any

dC/dv, the h-function \(P(U \le u \mid V = v)\); u and v broadcast as for du().

fit(x: ArrayLike, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, margins: Any = None, how: str = 'IFM', xl: ArrayLike | None = None, xr: ArrayLike | None = None, init: ArrayLike | None = None, rotation: int = 0) → Any

Fit the copula and its margins to multivariate survival data.

Parameters:
  • x – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • c – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • n – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • t – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xl – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xr – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • margins (sequence of length D) – Either surpyval distribution classes (e.g. surpyval.Weibull) to be fitted, or already-fitted models exposing ff/df. Under "IFM" a fitted model is used as it is. Under "MLE" a fitted parametric model supplies the starting values and its configuration: it is re-estimated jointly with the copula with the same offset, limited-failure or zero-inflated option, and any parameters it was fitted with fixed stay at their values. A non-parametric margin (a class such as surpyval.KaplanMeier, or a fitted non-parametric model) can only be used with "IFM": it gives the semi-parametric pseudo-likelihood estimator, and the likelihood and criteria then compare copula families with the same margins only.

  • how ({"IFM", "MLE"}) – "IFM" (default) fits each margin independently then the single copula parameter (robust two-stage estimation). "MLE" jointly optimises copula parameter + margin parameters.

  • init (array like, optional) – Starting value of the copula parameter(s) for the search, one per entry of parameter_names, each strictly inside the family’s bounds. By default each family starts from its own guess (the built-in families match the empirical Kendall’s tau).

  • rotation ({0, 90, 180, 270}, optional) – Rotate the copula by this many degrees, in the convention of R’s VineCopula (families 13, 23, 33 for the Clayton, …): 180 is the survival copula, C(u, v) = u + v - 1 + C_0(1 - u, 1 - v), with the tail dependence moved to the other tail (a Clayton’s to the upper tail, a Gumbel’s or Joe’s to the lower); 90, C(u, v) = v - C_0(1 - u, v), and 270, C(u, v) = u - C_0(u, 1 - v), give negative dependence with the family’s shape, its tail in a corner where one series is short and the other long. The parameter keeps its own range (pyvinecopulib’s convention; VineCopula writes it negated for 90 and 270). Only the Clayton, Gumbel and Joe copulas are rotated; for a radially symmetric family (Frank, Gaussian, Student-t) 180 is the family itself and 90 its negative parameter. Default 0, the family as it is.

Returns:

The fitted model: the copula parameter(s) params and the fitted margins, with the joint sf/cdf/pdf, sampling, dependence measures and the likelihood-based log_likelihood/neg_ll/aic/bic.

Return type:

CopulaModel

Warns:

UserWarning – “No finite maximum” when the rows observed in both dimensions are perfectly dependent (Kendall’s tau of +-1) and the family reaches that dependence only as its parameter runs to a limit: the returned parameter is then meaningless.

Examples

Simulate from a Clayton copula with Weibull margins, then recover it:

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> model = Clayton.fit(X, margins=[Weibull, Weibull])
>>> model.params.round(3)
array([2.293])
>>> round(float(model.kendall_tau()), 3)
0.534
fit_from_df(df: Any, x_cols: list[str], c_cols: list[str] | None = None, n_col: str | None = None, xl_cols: list[str] | None = None, xr_cols: list[str] | None = None, tl_cols: list[str] | None = None, tr_cols: list[str] | None = None, **fit_options: Any) → Any

Fit the copula and its margins to the columns of a pandas.DataFrame.

Each argument names, per dimension, the columns fit() takes as arrays: x_cols=["a", "b"] reads the two series from columns a and b. The names are those of the univariate fit_from_df (Weibull.fit_from_df(df, x_col=..., c_col=...)) with _cols for a list of columns, one per dimension (principle 21); every other fit() option (margins, how, init) is passed to it unchanged.

Parameters:
  • df (pandas.DataFrame) – The data, one row per unit.

  • x_cols (list of str) – The column of each dimension’s values.

  • c_cols (list of str, optional) – The column of each dimension’s censoring flags. Defaults to every value observed.

  • n_col (str, optional) – The column of row counts.

  • xl_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • xr_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • tl_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • tr_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • **fit_options – Every other option of fit().

Returns:

The model fit() returns for the same arrays.

Return type:

CopulaModel

Examples

>>> import pandas as pd
>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> df = pd.DataFrame(X, columns=["pump", "motor"])
>>> model = Clayton.fit_from_df(
...     df, x_cols=["pump", "motor"], margins=[Weibull, Weibull]
... )
>>> model.params.round(3)
array([2.293])
from_params(params: Any, margins: Any, rotation: int = 0) → Any

Build a CopulaModel from a known parameter and margins, without fitting.

Parameters:
  • params (array like) – The copula parameter(s), e.g. [theta] (empty for the independence copula): one per entry of parameter_names, each strictly inside the family’s bounds (Gumbel’s theta = 1, the independence copula, is allowed too). Anything else raises ValueError.

  • margins (sequence of length 2) – Fitted (or from_params) univariate models, one per dimension, each exposing ff and df.

  • rotation ({0, 90, 180, 270}, optional) – The rotation of the copula, as for fit().

Returns:

The model, for evaluation and simulation.

Return type:

CopulaModel

Examples

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> model = Clayton.from_params([2.0], margins)
>>> round(float(model.kendall_tau()), 3)
0.5
kendall_tau(theta: float) → float

Kendall’s tau.

The default integrates \(\tau = 1 - 4 \int_0^1 \int_0^1 \frac{\partial C}{\partial u} \frac{\partial C}{\partial v} \, du \, dv\) by Gauss-Legendre quadrature (400 nodes per margin: accurate to about 1e-10 at a tau of 0.5 and 1e-8 at 0.8, the integrand sharpening along the diagonal as the dependence grows); families with a closed form override it. It used to estimate tau from 50 000 simulated pairs, with an error near 1e-3.

neg_ll(params: Any, dims: list, weights: NDArray) → float

The copula-stage negative log-likelihood (used by the fit).

pdf(u: Any, v: Any, theta: Any) → Any

d2C/du dv – the copula density. The default differentiates du() with autograd; u and v broadcast as for du().

rotatable: bool = True

the families whose rotations are new copulas (the asymmetric Archimedean ones).

Type:

Whether rotated() applies

rotated(rotation: int) → Copula

The family rotated by rotation degrees (0, 90, 180 or 270), in R’s VineCopula convention; see the rotation option of fit(), which uses it. rotated(0) is the family itself.

Examples

>>> from surpyval.multivariate import Clayton
>>> survival = Clayton.rotated(180)
>>> survival.tail_dependence(2.0)
(0.0, 0.7071067811865476)
>>> Clayton.rotated(90).kendall_tau(2.0)
-0.5
rotation: int = 0

0 for the family as it is.

Type:

The rotation of the copula in degrees (see rotated())

sample_uv(size: int, params: Any, random_state: int | None = None) → tuple[NDArray, NDArray]

Draw (u, v) pairs by conditional inversion of the h-function.

u is uniform; given u and a uniform w, v solves dC/du(u, v) = w (a CDF in v, hence monotone) by bisection. Override for families with a direct sampler (e.g. Gaussian).

spearman_rho(*params: float) → float

Spearman’s rho.

The default integrates \(\rho_S = 12 \int_0^1 \int_0^1 C(u, v) \, du \, dv - 3\) by Gauss-Legendre quadrature (400 nodes per margin, accurate to about 1e-11 for the built-in families); families with a closed form override it. It used to estimate rho from 50 000 simulated pairs, which was up to 5e-3 off (Clayton, Gumbel).

tail_dependence(theta: float) → tuple

Lower/upper tail-dependence coefficients (lambda_L, lambda_U).

Default (0.0, 0.0) (no tail dependence); families override.

class surpyval.multivariate.parametric.copula.archimedean.GumbelCopula

Bases: Copula

Gumbel-Hougaard copula (upper-tail dependence), theta >= 1.

cdf(u: Any, v: Any, theta: Any) → Any

The copula \(C(u, v)\); defined by each family.

closed_bounds: tuple = ('theta',)

Names of parameters whose finite bounds are themselves valid values (Gumbel’s theta = 1 is the independence copula). The fitter never reaches a bound, but from_params may be given one.

dependence_limits: dict = {1: 'theta grows without bound'}

The Frechet bounds the family reaches only as its parameter runs to a limit: +1 the comonotone copula (perfect positive dependence), -1 the countermonotone one, each mapped to that limit as the fit’s warning names it (see _warn_if_perfectly_dependent()). Empty for a family that is not known to reach either.

du(u: Any, v: Any, theta: Any) → Any

dC/du – the h-function \(P(V \le v \mid U = u)\).

The default differentiates cdf() with autograd; override it when a closed form is known. u and v broadcast against each other, and the result has their common shape.

dv(u: Any, v: Any, theta: Any) → Any

dC/dv, the h-function \(P(U \le u \mid V = v)\); u and v broadcast as for du().

fit(x: ArrayLike, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, margins: Any = None, how: str = 'IFM', xl: ArrayLike | None = None, xr: ArrayLike | None = None, init: ArrayLike | None = None, rotation: int = 0) → Any

Fit the copula and its margins to multivariate survival data.

Parameters:
  • x – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • c – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • n – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • t – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xl – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xr – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • margins (sequence of length D) – Either surpyval distribution classes (e.g. surpyval.Weibull) to be fitted, or already-fitted models exposing ff/df. Under "IFM" a fitted model is used as it is. Under "MLE" a fitted parametric model supplies the starting values and its configuration: it is re-estimated jointly with the copula with the same offset, limited-failure or zero-inflated option, and any parameters it was fitted with fixed stay at their values. A non-parametric margin (a class such as surpyval.KaplanMeier, or a fitted non-parametric model) can only be used with "IFM": it gives the semi-parametric pseudo-likelihood estimator, and the likelihood and criteria then compare copula families with the same margins only.

  • how ({"IFM", "MLE"}) – "IFM" (default) fits each margin independently then the single copula parameter (robust two-stage estimation). "MLE" jointly optimises copula parameter + margin parameters.

  • init (array like, optional) – Starting value of the copula parameter(s) for the search, one per entry of parameter_names, each strictly inside the family’s bounds. By default each family starts from its own guess (the built-in families match the empirical Kendall’s tau).

  • rotation ({0, 90, 180, 270}, optional) – Rotate the copula by this many degrees, in the convention of R’s VineCopula (families 13, 23, 33 for the Clayton, …): 180 is the survival copula, C(u, v) = u + v - 1 + C_0(1 - u, 1 - v), with the tail dependence moved to the other tail (a Clayton’s to the upper tail, a Gumbel’s or Joe’s to the lower); 90, C(u, v) = v - C_0(1 - u, v), and 270, C(u, v) = u - C_0(u, 1 - v), give negative dependence with the family’s shape, its tail in a corner where one series is short and the other long. The parameter keeps its own range (pyvinecopulib’s convention; VineCopula writes it negated for 90 and 270). Only the Clayton, Gumbel and Joe copulas are rotated; for a radially symmetric family (Frank, Gaussian, Student-t) 180 is the family itself and 90 its negative parameter. Default 0, the family as it is.

Returns:

The fitted model: the copula parameter(s) params and the fitted margins, with the joint sf/cdf/pdf, sampling, dependence measures and the likelihood-based log_likelihood/neg_ll/aic/bic.

Return type:

CopulaModel

Warns:

UserWarning – “No finite maximum” when the rows observed in both dimensions are perfectly dependent (Kendall’s tau of +-1) and the family reaches that dependence only as its parameter runs to a limit: the returned parameter is then meaningless.

Examples

Simulate from a Clayton copula with Weibull margins, then recover it:

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> model = Clayton.fit(X, margins=[Weibull, Weibull])
>>> model.params.round(3)
array([2.293])
>>> round(float(model.kendall_tau()), 3)
0.534
fit_from_df(df: Any, x_cols: list[str], c_cols: list[str] | None = None, n_col: str | None = None, xl_cols: list[str] | None = None, xr_cols: list[str] | None = None, tl_cols: list[str] | None = None, tr_cols: list[str] | None = None, **fit_options: Any) → Any

Fit the copula and its margins to the columns of a pandas.DataFrame.

Each argument names, per dimension, the columns fit() takes as arrays: x_cols=["a", "b"] reads the two series from columns a and b. The names are those of the univariate fit_from_df (Weibull.fit_from_df(df, x_col=..., c_col=...)) with _cols for a list of columns, one per dimension (principle 21); every other fit() option (margins, how, init) is passed to it unchanged.

Parameters:
  • df (pandas.DataFrame) – The data, one row per unit.

  • x_cols (list of str) – The column of each dimension’s values.

  • c_cols (list of str, optional) – The column of each dimension’s censoring flags. Defaults to every value observed.

  • n_col (str, optional) – The column of row counts.

  • xl_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • xr_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • tl_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • tr_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • **fit_options – Every other option of fit().

Returns:

The model fit() returns for the same arrays.

Return type:

CopulaModel

Examples

>>> import pandas as pd
>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> df = pd.DataFrame(X, columns=["pump", "motor"])
>>> model = Clayton.fit_from_df(
...     df, x_cols=["pump", "motor"], margins=[Weibull, Weibull]
... )
>>> model.params.round(3)
array([2.293])
from_params(params: Any, margins: Any, rotation: int = 0) → Any

Build a CopulaModel from a known parameter and margins, without fitting.

Parameters:
  • params (array like) – The copula parameter(s), e.g. [theta] (empty for the independence copula): one per entry of parameter_names, each strictly inside the family’s bounds (Gumbel’s theta = 1, the independence copula, is allowed too). Anything else raises ValueError.

  • margins (sequence of length 2) – Fitted (or from_params) univariate models, one per dimension, each exposing ff and df.

  • rotation ({0, 90, 180, 270}, optional) – The rotation of the copula, as for fit().

Returns:

The model, for evaluation and simulation.

Return type:

CopulaModel

Examples

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> model = Clayton.from_params([2.0], margins)
>>> round(float(model.kendall_tau()), 3)
0.5
kendall_tau(theta: float) → float

Kendall’s tau.

The default integrates \(\tau = 1 - 4 \int_0^1 \int_0^1 \frac{\partial C}{\partial u} \frac{\partial C}{\partial v} \, du \, dv\) by Gauss-Legendre quadrature (400 nodes per margin: accurate to about 1e-10 at a tau of 0.5 and 1e-8 at 0.8, the integrand sharpening along the diagonal as the dependence grows); families with a closed form override it. It used to estimate tau from 50 000 simulated pairs, with an error near 1e-3.

neg_ll(params: Any, dims: list, weights: NDArray) → float

The copula-stage negative log-likelihood (used by the fit).

pdf(u: Any, v: Any, theta: Any) → Any

d2C/du dv – the copula density. The default differentiates du() with autograd; u and v broadcast as for du().

rotatable: bool = True

the families whose rotations are new copulas (the asymmetric Archimedean ones).

Type:

Whether rotated() applies

rotated(rotation: int) → Copula

The family rotated by rotation degrees (0, 90, 180 or 270), in R’s VineCopula convention; see the rotation option of fit(), which uses it. rotated(0) is the family itself.

Examples

>>> from surpyval.multivariate import Clayton
>>> survival = Clayton.rotated(180)
>>> survival.tail_dependence(2.0)
(0.0, 0.7071067811865476)
>>> Clayton.rotated(90).kendall_tau(2.0)
-0.5
rotation: int = 0

0 for the family as it is.

Type:

The rotation of the copula in degrees (see rotated())

sample_uv(size: int, params: Any, random_state: int | None = None) → tuple[NDArray, NDArray]

Draw (u, v) pairs by conditional inversion of the h-function.

u is uniform; given u and a uniform w, v solves dC/du(u, v) = w (a CDF in v, hence monotone) by bisection. Override for families with a direct sampler (e.g. Gaussian).

spearman_rho(*params: float) → float

Spearman’s rho.

The default integrates \(\rho_S = 12 \int_0^1 \int_0^1 C(u, v) \, du \, dv - 3\) by Gauss-Legendre quadrature (400 nodes per margin, accurate to about 1e-11 for the built-in families); families with a closed form override it. It used to estimate rho from 50 000 simulated pairs, which was up to 5e-3 off (Clayton, Gumbel).

tail_dependence(theta: float) → tuple

Lower/upper tail-dependence coefficients (lambda_L, lambda_U).

Default (0.0, 0.0) (no tail dependence); families override.

class surpyval.multivariate.parametric.copula.archimedean.FrankCopula

Bases: Copula

Frank copula (symmetric, no tail dependence), theta != 0.

Every primitive is evaluated in log space from terms of \(C(u, v) = -\frac{1}{\theta} \log\left(1 + \frac{(e^{-\theta u} - 1)(e^{-\theta v} - 1)}{e^{-\theta} - 1} \right)\) that are sums of non-negative numbers, so they stay accurate for any \(\theta\). The textbook form overflowed once \(e^{-\theta}\) rounded against 1 (\(\theta \gtrsim 37\), Kendall’s tau 0.9): the CDF and density became infinite near the upper corner, the likelihood +inf and the sampler’s second margin far from uniform. theta = 0 is the independence copula.

cdf(u: Any, v: Any, theta: Any) → Any

The copula \(C(u, v)\); defined by each family.

closed_bounds: tuple = ()

Names of parameters whose finite bounds are themselves valid values (Gumbel’s theta = 1 is the independence copula). The fitter never reaches a bound, but from_params may be given one.

dependence_limits: dict = {-1: 'theta falls without bound', 1: 'theta grows without bound'}

The Frechet bounds the family reaches only as its parameter runs to a limit: +1 the comonotone copula (perfect positive dependence), -1 the countermonotone one, each mapped to that limit as the fit’s warning names it (see _warn_if_perfectly_dependent()). Empty for a family that is not known to reach either.

du(u: Any, v: Any, theta: Any) → Any

dC/du – the h-function \(P(V \le v \mid U = u)\).

The default differentiates cdf() with autograd; override it when a closed form is known. u and v broadcast against each other, and the result has their common shape.

dv(u: Any, v: Any, theta: Any) → Any

dC/dv, the h-function \(P(U \le u \mid V = v)\); u and v broadcast as for du().

fit(x: ArrayLike, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, margins: Any = None, how: str = 'IFM', xl: ArrayLike | None = None, xr: ArrayLike | None = None, init: ArrayLike | None = None, rotation: int = 0) → Any

Fit the copula and its margins to multivariate survival data.

Parameters:
  • x – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • c – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • n – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • t – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xl – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xr – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • margins (sequence of length D) – Either surpyval distribution classes (e.g. surpyval.Weibull) to be fitted, or already-fitted models exposing ff/df. Under "IFM" a fitted model is used as it is. Under "MLE" a fitted parametric model supplies the starting values and its configuration: it is re-estimated jointly with the copula with the same offset, limited-failure or zero-inflated option, and any parameters it was fitted with fixed stay at their values. A non-parametric margin (a class such as surpyval.KaplanMeier, or a fitted non-parametric model) can only be used with "IFM": it gives the semi-parametric pseudo-likelihood estimator, and the likelihood and criteria then compare copula families with the same margins only.

  • how ({"IFM", "MLE"}) – "IFM" (default) fits each margin independently then the single copula parameter (robust two-stage estimation). "MLE" jointly optimises copula parameter + margin parameters.

  • init (array like, optional) – Starting value of the copula parameter(s) for the search, one per entry of parameter_names, each strictly inside the family’s bounds. By default each family starts from its own guess (the built-in families match the empirical Kendall’s tau).

  • rotation ({0, 90, 180, 270}, optional) – Rotate the copula by this many degrees, in the convention of R’s VineCopula (families 13, 23, 33 for the Clayton, …): 180 is the survival copula, C(u, v) = u + v - 1 + C_0(1 - u, 1 - v), with the tail dependence moved to the other tail (a Clayton’s to the upper tail, a Gumbel’s or Joe’s to the lower); 90, C(u, v) = v - C_0(1 - u, v), and 270, C(u, v) = u - C_0(u, 1 - v), give negative dependence with the family’s shape, its tail in a corner where one series is short and the other long. The parameter keeps its own range (pyvinecopulib’s convention; VineCopula writes it negated for 90 and 270). Only the Clayton, Gumbel and Joe copulas are rotated; for a radially symmetric family (Frank, Gaussian, Student-t) 180 is the family itself and 90 its negative parameter. Default 0, the family as it is.

Returns:

The fitted model: the copula parameter(s) params and the fitted margins, with the joint sf/cdf/pdf, sampling, dependence measures and the likelihood-based log_likelihood/neg_ll/aic/bic.

Return type:

CopulaModel

Warns:

UserWarning – “No finite maximum” when the rows observed in both dimensions are perfectly dependent (Kendall’s tau of +-1) and the family reaches that dependence only as its parameter runs to a limit: the returned parameter is then meaningless.

Examples

Simulate from a Clayton copula with Weibull margins, then recover it:

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> model = Clayton.fit(X, margins=[Weibull, Weibull])
>>> model.params.round(3)
array([2.293])
>>> round(float(model.kendall_tau()), 3)
0.534
fit_from_df(df: Any, x_cols: list[str], c_cols: list[str] | None = None, n_col: str | None = None, xl_cols: list[str] | None = None, xr_cols: list[str] | None = None, tl_cols: list[str] | None = None, tr_cols: list[str] | None = None, **fit_options: Any) → Any

Fit the copula and its margins to the columns of a pandas.DataFrame.

Each argument names, per dimension, the columns fit() takes as arrays: x_cols=["a", "b"] reads the two series from columns a and b. The names are those of the univariate fit_from_df (Weibull.fit_from_df(df, x_col=..., c_col=...)) with _cols for a list of columns, one per dimension (principle 21); every other fit() option (margins, how, init) is passed to it unchanged.

Parameters:
  • df (pandas.DataFrame) – The data, one row per unit.

  • x_cols (list of str) – The column of each dimension’s values.

  • c_cols (list of str, optional) – The column of each dimension’s censoring flags. Defaults to every value observed.

  • n_col (str, optional) – The column of row counts.

  • xl_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • xr_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • tl_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • tr_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • **fit_options – Every other option of fit().

Returns:

The model fit() returns for the same arrays.

Return type:

CopulaModel

Examples

>>> import pandas as pd
>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> df = pd.DataFrame(X, columns=["pump", "motor"])
>>> model = Clayton.fit_from_df(
...     df, x_cols=["pump", "motor"], margins=[Weibull, Weibull]
... )
>>> model.params.round(3)
array([2.293])
from_params(params: Any, margins: Any, rotation: int = 0) → Any

Build a CopulaModel from a known parameter and margins, without fitting.

Parameters:
  • params (array like) – The copula parameter(s), e.g. [theta] (empty for the independence copula): one per entry of parameter_names, each strictly inside the family’s bounds (Gumbel’s theta = 1, the independence copula, is allowed too). Anything else raises ValueError.

  • margins (sequence of length 2) – Fitted (or from_params) univariate models, one per dimension, each exposing ff and df.

  • rotation ({0, 90, 180, 270}, optional) – The rotation of the copula, as for fit().

Returns:

The model, for evaluation and simulation.

Return type:

CopulaModel

Examples

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> model = Clayton.from_params([2.0], margins)
>>> round(float(model.kendall_tau()), 3)
0.5
kendall_tau(theta: float) → float

Kendall’s tau.

The default integrates \(\tau = 1 - 4 \int_0^1 \int_0^1 \frac{\partial C}{\partial u} \frac{\partial C}{\partial v} \, du \, dv\) by Gauss-Legendre quadrature (400 nodes per margin: accurate to about 1e-10 at a tau of 0.5 and 1e-8 at 0.8, the integrand sharpening along the diagonal as the dependence grows); families with a closed form override it. It used to estimate tau from 50 000 simulated pairs, with an error near 1e-3.

neg_ll(params: Any, dims: list, weights: NDArray) → float

The copula-stage negative log-likelihood (used by the fit).

pdf(u: Any, v: Any, theta: Any) → Any

d2C/du dv – the copula density. The default differentiates du() with autograd; u and v broadcast as for du().

rotatable: bool = False

the families whose rotations are new copulas (the asymmetric Archimedean ones).

Type:

Whether rotated() applies

rotated(rotation: int) → Copula

The family rotated by rotation degrees (0, 90, 180 or 270), in R’s VineCopula convention; see the rotation option of fit(), which uses it. rotated(0) is the family itself.

Examples

>>> from surpyval.multivariate import Clayton
>>> survival = Clayton.rotated(180)
>>> survival.tail_dependence(2.0)
(0.0, 0.7071067811865476)
>>> Clayton.rotated(90).kendall_tau(2.0)
-0.5
rotation: int = 0

0 for the family as it is.

Type:

The rotation of the copula in degrees (see rotated())

sample_uv(size: Any, params: Any, random_state: int | None = None) → tuple[NDArray, NDArray]

Draw (u, v) pairs, inverting the h-function in closed form: given u and a uniform w,

\[v = -\frac{1}{\theta} \log \frac{w e^{-\theta} + (1 - w) e^{-\theta u}} {w + (1 - w) e^{-\theta u}},\]

with both sums formed in log space (exact for any theta).

spearman_rho(theta: float) → float

Spearman’s rho in closed form, \(1 - \frac{12}{\theta}(D_1(\theta) - D_2(\theta))\) with \(D_k\) the Debye functions.

tail_dependence(*params: float) → tuple

Lower/upper tail-dependence coefficients (lambda_L, lambda_U).

Default (0.0, 0.0) (no tail dependence); families override.

class surpyval.multivariate.parametric.copula.archimedean.JoeCopula

Bases: Copula

Joe copula (upper-tail dependence), theta >= 1.

\[C(u, v) = 1 - \left(\bar u^\theta + \bar v^\theta - \bar u^\theta \bar v^\theta\right)^{1/\theta}, \qquad \bar u = 1 - u,\]

the parameterisation of R’s copula::joeCopula and of VineCopula (family 6). theta = 1 is the independence copula, and the dependence grows with theta towards the comonotone copula. Like the Gumbel it has upper-tail dependence only, \(\lambda_U = 2 - 2^{1/\theta}\), but for a given Kendall’s tau a stronger one: at tau = 0.5 the Joe has \(\lambda_U = 0.71\) (theta = 2.86), the Gumbel 0.59.

With \(a = 1 - \bar u^\theta\) and \(b = 1 - \bar v^\theta\), the bracket is \(A = 1 - a b\), and every primitive is formed from log A: as log1p(-a b) near the lower corner (where a b is small and C is tiny), as a log-sum-exp of \(\bar u^\theta\) and \(\bar v^\theta a\) near the upper one.

Examples

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Joe
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> model = Joe.from_params([2.0], margins)
>>> round(model.kendall_tau(), 4)
0.3551
>>> [round(x, 4) for x in model.tail_dependence()]
[0.0, 0.5858]
cdf(u: Any, v: Any, theta: Any) → Any

The copula \(C(u, v)\); defined by each family.

closed_bounds: tuple = ('theta',)

Names of parameters whose finite bounds are themselves valid values (Gumbel’s theta = 1 is the independence copula). The fitter never reaches a bound, but from_params may be given one.

dependence_limits: dict = {1: 'theta grows without bound'}

The Frechet bounds the family reaches only as its parameter runs to a limit: +1 the comonotone copula (perfect positive dependence), -1 the countermonotone one, each mapped to that limit as the fit’s warning names it (see _warn_if_perfectly_dependent()). Empty for a family that is not known to reach either.

du(u: Any, v: Any, theta: Any) → Any

dC/du – the h-function \(P(V \le v \mid U = u)\).

The default differentiates cdf() with autograd; override it when a closed form is known. u and v broadcast against each other, and the result has their common shape.

dv(u: Any, v: Any, theta: Any) → Any

dC/dv, the h-function \(P(U \le u \mid V = v)\); u and v broadcast as for du().

fit(x: ArrayLike, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, margins: Any = None, how: str = 'IFM', xl: ArrayLike | None = None, xr: ArrayLike | None = None, init: ArrayLike | None = None, rotation: int = 0) → Any

Fit the copula and its margins to multivariate survival data.

Parameters:
  • x – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • c – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • n – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • t – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xl – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xr – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • margins (sequence of length D) – Either surpyval distribution classes (e.g. surpyval.Weibull) to be fitted, or already-fitted models exposing ff/df. Under "IFM" a fitted model is used as it is. Under "MLE" a fitted parametric model supplies the starting values and its configuration: it is re-estimated jointly with the copula with the same offset, limited-failure or zero-inflated option, and any parameters it was fitted with fixed stay at their values. A non-parametric margin (a class such as surpyval.KaplanMeier, or a fitted non-parametric model) can only be used with "IFM": it gives the semi-parametric pseudo-likelihood estimator, and the likelihood and criteria then compare copula families with the same margins only.

  • how ({"IFM", "MLE"}) – "IFM" (default) fits each margin independently then the single copula parameter (robust two-stage estimation). "MLE" jointly optimises copula parameter + margin parameters.

  • init (array like, optional) – Starting value of the copula parameter(s) for the search, one per entry of parameter_names, each strictly inside the family’s bounds. By default each family starts from its own guess (the built-in families match the empirical Kendall’s tau).

  • rotation ({0, 90, 180, 270}, optional) – Rotate the copula by this many degrees, in the convention of R’s VineCopula (families 13, 23, 33 for the Clayton, …): 180 is the survival copula, C(u, v) = u + v - 1 + C_0(1 - u, 1 - v), with the tail dependence moved to the other tail (a Clayton’s to the upper tail, a Gumbel’s or Joe’s to the lower); 90, C(u, v) = v - C_0(1 - u, v), and 270, C(u, v) = u - C_0(u, 1 - v), give negative dependence with the family’s shape, its tail in a corner where one series is short and the other long. The parameter keeps its own range (pyvinecopulib’s convention; VineCopula writes it negated for 90 and 270). Only the Clayton, Gumbel and Joe copulas are rotated; for a radially symmetric family (Frank, Gaussian, Student-t) 180 is the family itself and 90 its negative parameter. Default 0, the family as it is.

Returns:

The fitted model: the copula parameter(s) params and the fitted margins, with the joint sf/cdf/pdf, sampling, dependence measures and the likelihood-based log_likelihood/neg_ll/aic/bic.

Return type:

CopulaModel

Warns:

UserWarning – “No finite maximum” when the rows observed in both dimensions are perfectly dependent (Kendall’s tau of +-1) and the family reaches that dependence only as its parameter runs to a limit: the returned parameter is then meaningless.

Examples

Simulate from a Clayton copula with Weibull margins, then recover it:

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> model = Clayton.fit(X, margins=[Weibull, Weibull])
>>> model.params.round(3)
array([2.293])
>>> round(float(model.kendall_tau()), 3)
0.534
fit_from_df(df: Any, x_cols: list[str], c_cols: list[str] | None = None, n_col: str | None = None, xl_cols: list[str] | None = None, xr_cols: list[str] | None = None, tl_cols: list[str] | None = None, tr_cols: list[str] | None = None, **fit_options: Any) → Any

Fit the copula and its margins to the columns of a pandas.DataFrame.

Each argument names, per dimension, the columns fit() takes as arrays: x_cols=["a", "b"] reads the two series from columns a and b. The names are those of the univariate fit_from_df (Weibull.fit_from_df(df, x_col=..., c_col=...)) with _cols for a list of columns, one per dimension (principle 21); every other fit() option (margins, how, init) is passed to it unchanged.

Parameters:
  • df (pandas.DataFrame) – The data, one row per unit.

  • x_cols (list of str) – The column of each dimension’s values.

  • c_cols (list of str, optional) – The column of each dimension’s censoring flags. Defaults to every value observed.

  • n_col (str, optional) – The column of row counts.

  • xl_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • xr_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • tl_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • tr_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • **fit_options – Every other option of fit().

Returns:

The model fit() returns for the same arrays.

Return type:

CopulaModel

Examples

>>> import pandas as pd
>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> df = pd.DataFrame(X, columns=["pump", "motor"])
>>> model = Clayton.fit_from_df(
...     df, x_cols=["pump", "motor"], margins=[Weibull, Weibull]
... )
>>> model.params.round(3)
array([2.293])
from_params(params: Any, margins: Any, rotation: int = 0) → Any

Build a CopulaModel from a known parameter and margins, without fitting.

Parameters:
  • params (array like) – The copula parameter(s), e.g. [theta] (empty for the independence copula): one per entry of parameter_names, each strictly inside the family’s bounds (Gumbel’s theta = 1, the independence copula, is allowed too). Anything else raises ValueError.

  • margins (sequence of length 2) – Fitted (or from_params) univariate models, one per dimension, each exposing ff and df.

  • rotation ({0, 90, 180, 270}, optional) – The rotation of the copula, as for fit().

Returns:

The model, for evaluation and simulation.

Return type:

CopulaModel

Examples

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> model = Clayton.from_params([2.0], margins)
>>> round(float(model.kendall_tau()), 3)
0.5
kendall_tau(theta: float) → float

Kendall’s tau, \(1 - \frac{2}{\theta}\, \frac{\psi(2 + \delta) - \psi(2)}{\delta}\) with \(\delta = 2/\theta - 1\) and \(\psi\) the digamma function (the closed form of R’s copula::tau for the Joe copula, written so that theta = 2 is not 0/0).

neg_ll(params: Any, dims: list, weights: NDArray) → float

The copula-stage negative log-likelihood (used by the fit).

pdf(u: Any, v: Any, theta: Any) → Any

d2C/du dv – the copula density. The default differentiates du() with autograd; u and v broadcast as for du().

rotatable: bool = True

the families whose rotations are new copulas (the asymmetric Archimedean ones).

Type:

Whether rotated() applies

rotated(rotation: int) → Copula

The family rotated by rotation degrees (0, 90, 180 or 270), in R’s VineCopula convention; see the rotation option of fit(), which uses it. rotated(0) is the family itself.

Examples

>>> from surpyval.multivariate import Clayton
>>> survival = Clayton.rotated(180)
>>> survival.tail_dependence(2.0)
(0.0, 0.7071067811865476)
>>> Clayton.rotated(90).kendall_tau(2.0)
-0.5
rotation: int = 0

0 for the family as it is.

Type:

The rotation of the copula in degrees (see rotated())

sample_uv(size: int, params: Any, random_state: int | None = None) → tuple[NDArray, NDArray]

Draw (u, v) pairs by conditional inversion of the h-function.

u is uniform; given u and a uniform w, v solves dC/du(u, v) = w (a CDF in v, hence monotone) by bisection. Override for families with a direct sampler (e.g. Gaussian).

spearman_rho(*params: float) → float

Spearman’s rho.

The default integrates \(\rho_S = 12 \int_0^1 \int_0^1 C(u, v) \, du \, dv - 3\) by Gauss-Legendre quadrature (400 nodes per margin, accurate to about 1e-11 for the built-in families); families with a closed form override it. It used to estimate rho from 50 000 simulated pairs, which was up to 5e-3 off (Clayton, Gumbel).

tail_dependence(theta: float) → tuple

Lower/upper tail-dependence coefficients (lambda_L, lambda_U).

Default (0.0, 0.0) (no tail dependence); families override.

class surpyval.multivariate.parametric.copula.archimedean.AMHCopula

Bases: Copula

Ali-Mikhail-Haq copula (weak dependence), -1 <= theta <= 1.

\[C(u, v) = \frac{u v}{1 - \theta (1 - u)(1 - v)},\]

the parameterisation of R’s copula::amhCopula. theta = 0 is the independence copula. The family only reaches weak dependence: Kendall’s tau lies in \([-0.1817, 1/3]\) and Spearman’s rho in \([-0.2711, 0.4784]\), both bounds attained at theta = -1 and 1. A fit to data more strongly dependent than that runs to the bound (a valid copula) and returns it, without a warning, as a Clayton fit to negatively dependent data runs to independence; use the Clayton, Frank or Gaussian copula there. It has no tail dependence, except \(\lambda_L = 1/2\) at theta = 1 (where it is the Clayton copula with theta = 1).

Examples

>>> from surpyval import Weibull
>>> from surpyval.multivariate import AMH
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> model = AMH.from_params([0.5], margins)
>>> round(model.kendall_tau(), 4), round(model.spearman_rho(), 4)
(0.1288, 0.1924)
cdf(u: Any, v: Any, theta: Any) → Any

The copula \(C(u, v)\); defined by each family.

closed_bounds: tuple = ('theta',)

Names of parameters whose finite bounds are themselves valid values (Gumbel’s theta = 1 is the independence copula). The fitter never reaches a bound, but from_params may be given one.

dependence_limits: dict = {}

The Frechet bounds the family reaches only as its parameter runs to a limit: +1 the comonotone copula (perfect positive dependence), -1 the countermonotone one, each mapped to that limit as the fit’s warning names it (see _warn_if_perfectly_dependent()). Empty for a family that is not known to reach either.

du(u: Any, v: Any, theta: Any) → Any

dC/du – the h-function \(P(V \le v \mid U = u)\).

The default differentiates cdf() with autograd; override it when a closed form is known. u and v broadcast against each other, and the result has their common shape.

dv(u: Any, v: Any, theta: Any) → Any

dC/dv, the h-function \(P(U \le u \mid V = v)\); u and v broadcast as for du().

fit(x: ArrayLike, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, margins: Any = None, how: str = 'IFM', xl: ArrayLike | None = None, xr: ArrayLike | None = None, init: ArrayLike | None = None, rotation: int = 0) → Any

Fit the copula and its margins to multivariate survival data.

Parameters:
  • x – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • c – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • n – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • t – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xl – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xr – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • margins (sequence of length D) – Either surpyval distribution classes (e.g. surpyval.Weibull) to be fitted, or already-fitted models exposing ff/df. Under "IFM" a fitted model is used as it is. Under "MLE" a fitted parametric model supplies the starting values and its configuration: it is re-estimated jointly with the copula with the same offset, limited-failure or zero-inflated option, and any parameters it was fitted with fixed stay at their values. A non-parametric margin (a class such as surpyval.KaplanMeier, or a fitted non-parametric model) can only be used with "IFM": it gives the semi-parametric pseudo-likelihood estimator, and the likelihood and criteria then compare copula families with the same margins only.

  • how ({"IFM", "MLE"}) – "IFM" (default) fits each margin independently then the single copula parameter (robust two-stage estimation). "MLE" jointly optimises copula parameter + margin parameters.

  • init (array like, optional) – Starting value of the copula parameter(s) for the search, one per entry of parameter_names, each strictly inside the family’s bounds. By default each family starts from its own guess (the built-in families match the empirical Kendall’s tau).

  • rotation ({0, 90, 180, 270}, optional) – Rotate the copula by this many degrees, in the convention of R’s VineCopula (families 13, 23, 33 for the Clayton, …): 180 is the survival copula, C(u, v) = u + v - 1 + C_0(1 - u, 1 - v), with the tail dependence moved to the other tail (a Clayton’s to the upper tail, a Gumbel’s or Joe’s to the lower); 90, C(u, v) = v - C_0(1 - u, v), and 270, C(u, v) = u - C_0(u, 1 - v), give negative dependence with the family’s shape, its tail in a corner where one series is short and the other long. The parameter keeps its own range (pyvinecopulib’s convention; VineCopula writes it negated for 90 and 270). Only the Clayton, Gumbel and Joe copulas are rotated; for a radially symmetric family (Frank, Gaussian, Student-t) 180 is the family itself and 90 its negative parameter. Default 0, the family as it is.

Returns:

The fitted model: the copula parameter(s) params and the fitted margins, with the joint sf/cdf/pdf, sampling, dependence measures and the likelihood-based log_likelihood/neg_ll/aic/bic.

Return type:

CopulaModel

Warns:

UserWarning – “No finite maximum” when the rows observed in both dimensions are perfectly dependent (Kendall’s tau of +-1) and the family reaches that dependence only as its parameter runs to a limit: the returned parameter is then meaningless.

Examples

Simulate from a Clayton copula with Weibull margins, then recover it:

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> model = Clayton.fit(X, margins=[Weibull, Weibull])
>>> model.params.round(3)
array([2.293])
>>> round(float(model.kendall_tau()), 3)
0.534
fit_from_df(df: Any, x_cols: list[str], c_cols: list[str] | None = None, n_col: str | None = None, xl_cols: list[str] | None = None, xr_cols: list[str] | None = None, tl_cols: list[str] | None = None, tr_cols: list[str] | None = None, **fit_options: Any) → Any

Fit the copula and its margins to the columns of a pandas.DataFrame.

Each argument names, per dimension, the columns fit() takes as arrays: x_cols=["a", "b"] reads the two series from columns a and b. The names are those of the univariate fit_from_df (Weibull.fit_from_df(df, x_col=..., c_col=...)) with _cols for a list of columns, one per dimension (principle 21); every other fit() option (margins, how, init) is passed to it unchanged.

Parameters:
  • df (pandas.DataFrame) – The data, one row per unit.

  • x_cols (list of str) – The column of each dimension’s values.

  • c_cols (list of str, optional) – The column of each dimension’s censoring flags. Defaults to every value observed.

  • n_col (str, optional) – The column of row counts.

  • xl_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • xr_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • tl_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • tr_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • **fit_options – Every other option of fit().

Returns:

The model fit() returns for the same arrays.

Return type:

CopulaModel

Examples

>>> import pandas as pd
>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> df = pd.DataFrame(X, columns=["pump", "motor"])
>>> model = Clayton.fit_from_df(
...     df, x_cols=["pump", "motor"], margins=[Weibull, Weibull]
... )
>>> model.params.round(3)
array([2.293])
from_params(params: Any, margins: Any, rotation: int = 0) → Any

Build a CopulaModel from a known parameter and margins, without fitting.

Parameters:
  • params (array like) – The copula parameter(s), e.g. [theta] (empty for the independence copula): one per entry of parameter_names, each strictly inside the family’s bounds (Gumbel’s theta = 1, the independence copula, is allowed too). Anything else raises ValueError.

  • margins (sequence of length 2) – Fitted (or from_params) univariate models, one per dimension, each exposing ff and df.

  • rotation ({0, 90, 180, 270}, optional) – The rotation of the copula, as for fit().

Returns:

The model, for evaluation and simulation.

Return type:

CopulaModel

Examples

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> model = Clayton.from_params([2.0], margins)
>>> round(float(model.kendall_tau()), 3)
0.5
kendall_tau(theta: float) → float

Kendall’s tau, \(1 - \frac{2}{3 \theta^2}(\theta + (1 - \theta)^2 \log(1 - \theta))\) (Nelsen 2006, example 5.4), from -0.1817 at theta = -1 to 1/3 at theta = 1.

neg_ll(params: Any, dims: list, weights: NDArray) → float

The copula-stage negative log-likelihood (used by the fit).

pdf(u: Any, v: Any, theta: Any) → Any

d2C/du dv – the copula density. The default differentiates du() with autograd; u and v broadcast as for du().

rotatable: bool = False

the families whose rotations are new copulas (the asymmetric Archimedean ones).

Type:

Whether rotated() applies

rotated(rotation: int) → Copula

The family rotated by rotation degrees (0, 90, 180 or 270), in R’s VineCopula convention; see the rotation option of fit(), which uses it. rotated(0) is the family itself.

Examples

>>> from surpyval.multivariate import Clayton
>>> survival = Clayton.rotated(180)
>>> survival.tail_dependence(2.0)
(0.0, 0.7071067811865476)
>>> Clayton.rotated(90).kendall_tau(2.0)
-0.5
rotation: int = 0

0 for the family as it is.

Type:

The rotation of the copula in degrees (see rotated())

sample_uv(size: int, params: Any, random_state: int | None = None) → tuple[NDArray, NDArray]

Draw (u, v) pairs by conditional inversion of the h-function.

u is uniform; given u and a uniform w, v solves dC/du(u, v) = w (a CDF in v, hence monotone) by bisection. Override for families with a direct sampler (e.g. Gaussian).

spearman_rho(theta: float) → float

Spearman’s rho, \(\frac{12 (1 + \theta)}{\theta^2} \mathrm{Li}_2(\theta) - \frac{24 (1 - \theta)}{\theta^2} \log(1 - \theta) - \frac{3 (\theta + 12)}{\theta}\) (Nelsen 2006, exercise 5.10, with the dilogarithm \(\mathrm{Li}_2\)), from -0.2711 at theta = -1 to 0.4784 at theta = 1.

tail_dependence(theta: float) → tuple

Lower/upper tail-dependence coefficients (lambda_L, lambda_U).

Default (0.0, 0.0) (no tail dependence); families override.

class surpyval.multivariate.parametric.copula.elliptical.GaussianCopula

Bases: Copula

Gaussian copula, rho in (-1, 1) (no tail dependence).

Every function is accurate for any |rho| < 1: the formulas are written in 1 - |rho| and in the difference of the two normal quantiles, and the CDF is scipy’s bivariate normal CDF (Genz’s algorithm), within 1e-14 of a 40-digit integration up to rho = 1 - 2**-52.

cdf(u: Any, v: Any, rho: Any) → Any

The copula \(C(u, v)\); defined by each family.

closed_bounds: tuple = ()

Names of parameters whose finite bounds are themselves valid values (Gumbel’s theta = 1 is the independence copula). The fitter never reaches a bound, but from_params may be given one.

dependence_limits: dict = {-1: 'rho tends to -1', 1: 'rho tends to 1'}

The Frechet bounds the family reaches only as its parameter runs to a limit: +1 the comonotone copula (perfect positive dependence), -1 the countermonotone one, each mapped to that limit as the fit’s warning names it (see _warn_if_perfectly_dependent()). Empty for a family that is not known to reach either.

du(u: Any, v: Any, rho: Any) → Any

dC/du – the h-function \(P(V \le v \mid U = u)\).

The default differentiates cdf() with autograd; override it when a closed form is known. u and v broadcast against each other, and the result has their common shape.

dv(u: Any, v: Any, rho: Any) → Any

dC/dv, the h-function \(P(U \le u \mid V = v)\); u and v broadcast as for du().

fit(x: ArrayLike, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, margins: Any = None, how: str = 'IFM', xl: ArrayLike | None = None, xr: ArrayLike | None = None, init: ArrayLike | None = None, rotation: int = 0) → Any

Fit the copula and its margins to multivariate survival data.

Parameters:
  • x – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • c – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • n – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • t – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xl – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xr – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • margins (sequence of length D) – Either surpyval distribution classes (e.g. surpyval.Weibull) to be fitted, or already-fitted models exposing ff/df. Under "IFM" a fitted model is used as it is. Under "MLE" a fitted parametric model supplies the starting values and its configuration: it is re-estimated jointly with the copula with the same offset, limited-failure or zero-inflated option, and any parameters it was fitted with fixed stay at their values. A non-parametric margin (a class such as surpyval.KaplanMeier, or a fitted non-parametric model) can only be used with "IFM": it gives the semi-parametric pseudo-likelihood estimator, and the likelihood and criteria then compare copula families with the same margins only.

  • how ({"IFM", "MLE"}) – "IFM" (default) fits each margin independently then the single copula parameter (robust two-stage estimation). "MLE" jointly optimises copula parameter + margin parameters.

  • init (array like, optional) – Starting value of the copula parameter(s) for the search, one per entry of parameter_names, each strictly inside the family’s bounds. By default each family starts from its own guess (the built-in families match the empirical Kendall’s tau).

  • rotation ({0, 90, 180, 270}, optional) – Rotate the copula by this many degrees, in the convention of R’s VineCopula (families 13, 23, 33 for the Clayton, …): 180 is the survival copula, C(u, v) = u + v - 1 + C_0(1 - u, 1 - v), with the tail dependence moved to the other tail (a Clayton’s to the upper tail, a Gumbel’s or Joe’s to the lower); 90, C(u, v) = v - C_0(1 - u, v), and 270, C(u, v) = u - C_0(u, 1 - v), give negative dependence with the family’s shape, its tail in a corner where one series is short and the other long. The parameter keeps its own range (pyvinecopulib’s convention; VineCopula writes it negated for 90 and 270). Only the Clayton, Gumbel and Joe copulas are rotated; for a radially symmetric family (Frank, Gaussian, Student-t) 180 is the family itself and 90 its negative parameter. Default 0, the family as it is.

Returns:

The fitted model: the copula parameter(s) params and the fitted margins, with the joint sf/cdf/pdf, sampling, dependence measures and the likelihood-based log_likelihood/neg_ll/aic/bic.

Return type:

CopulaModel

Warns:

UserWarning – “No finite maximum” when the rows observed in both dimensions are perfectly dependent (Kendall’s tau of +-1) and the family reaches that dependence only as its parameter runs to a limit: the returned parameter is then meaningless.

Examples

Simulate from a Clayton copula with Weibull margins, then recover it:

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> model = Clayton.fit(X, margins=[Weibull, Weibull])
>>> model.params.round(3)
array([2.293])
>>> round(float(model.kendall_tau()), 3)
0.534
fit_from_df(df: Any, x_cols: list[str], c_cols: list[str] | None = None, n_col: str | None = None, xl_cols: list[str] | None = None, xr_cols: list[str] | None = None, tl_cols: list[str] | None = None, tr_cols: list[str] | None = None, **fit_options: Any) → Any

Fit the copula and its margins to the columns of a pandas.DataFrame.

Each argument names, per dimension, the columns fit() takes as arrays: x_cols=["a", "b"] reads the two series from columns a and b. The names are those of the univariate fit_from_df (Weibull.fit_from_df(df, x_col=..., c_col=...)) with _cols for a list of columns, one per dimension (principle 21); every other fit() option (margins, how, init) is passed to it unchanged.

Parameters:
  • df (pandas.DataFrame) – The data, one row per unit.

  • x_cols (list of str) – The column of each dimension’s values.

  • c_cols (list of str, optional) – The column of each dimension’s censoring flags. Defaults to every value observed.

  • n_col (str, optional) – The column of row counts.

  • xl_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • xr_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • tl_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • tr_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • **fit_options – Every other option of fit().

Returns:

The model fit() returns for the same arrays.

Return type:

CopulaModel

Examples

>>> import pandas as pd
>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> df = pd.DataFrame(X, columns=["pump", "motor"])
>>> model = Clayton.fit_from_df(
...     df, x_cols=["pump", "motor"], margins=[Weibull, Weibull]
... )
>>> model.params.round(3)
array([2.293])
from_params(params: Any, margins: Any, rotation: int = 0) → Any

Build a CopulaModel from a known parameter and margins, without fitting.

Parameters:
  • params (array like) – The copula parameter(s), e.g. [theta] (empty for the independence copula): one per entry of parameter_names, each strictly inside the family’s bounds (Gumbel’s theta = 1, the independence copula, is allowed too). Anything else raises ValueError.

  • margins (sequence of length 2) – Fitted (or from_params) univariate models, one per dimension, each exposing ff and df.

  • rotation ({0, 90, 180, 270}, optional) – The rotation of the copula, as for fit().

Returns:

The model, for evaluation and simulation.

Return type:

CopulaModel

Examples

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> model = Clayton.from_params([2.0], margins)
>>> round(float(model.kendall_tau()), 3)
0.5
kendall_tau(rho: float) → float

Kendall’s tau.

The default integrates \(\tau = 1 - 4 \int_0^1 \int_0^1 \frac{\partial C}{\partial u} \frac{\partial C}{\partial v} \, du \, dv\) by Gauss-Legendre quadrature (400 nodes per margin: accurate to about 1e-10 at a tau of 0.5 and 1e-8 at 0.8, the integrand sharpening along the diagonal as the dependence grows); families with a closed form override it. It used to estimate tau from 50 000 simulated pairs, with an error near 1e-3.

neg_ll(params: Any, dims: list, weights: NDArray) → float

The copula-stage negative log-likelihood (used by the fit).

pdf(u: Any, v: Any, rho: Any) → Any

d2C/du dv – the copula density. The default differentiates du() with autograd; u and v broadcast as for du().

rotatable: bool = False

the families whose rotations are new copulas (the asymmetric Archimedean ones).

Type:

Whether rotated() applies

rotated(rotation: int) → Copula

The family rotated by rotation degrees (0, 90, 180 or 270), in R’s VineCopula convention; see the rotation option of fit(), which uses it. rotated(0) is the family itself.

Examples

>>> from surpyval.multivariate import Clayton
>>> survival = Clayton.rotated(180)
>>> survival.tail_dependence(2.0)
(0.0, 0.7071067811865476)
>>> Clayton.rotated(90).kendall_tau(2.0)
-0.5
rotation: int = 0

0 for the family as it is.

Type:

The rotation of the copula in degrees (see rotated())

sample_uv(size: int, params: Any, random_state: int | None = None) → tuple[NDArray, NDArray]

Draw (u, v) pairs by conditional inversion of the h-function.

u is uniform; given u and a uniform w, v solves dC/du(u, v) = w (a CDF in v, hence monotone) by bisection. Override for families with a direct sampler (e.g. Gaussian).

spearman_rho(rho: float) → float

Spearman’s rho.

The default integrates \(\rho_S = 12 \int_0^1 \int_0^1 C(u, v) \, du \, dv - 3\) by Gauss-Legendre quadrature (400 nodes per margin, accurate to about 1e-11 for the built-in families); families with a closed form override it. It used to estimate rho from 50 000 simulated pairs, which was up to 5e-3 off (Clayton, Gumbel).

tail_dependence(*params: float) → tuple

Lower/upper tail-dependence coefficients (lambda_L, lambda_U).

Default (0.0, 0.0) (no tail dependence); families override.

class surpyval.multivariate.parametric.copula.elliptical.StudentTCopula

Bases: Copula

Student-t copula: correlation rho in (-1, 1) and degrees of freedom nu > 0 (symmetric tail dependence).

The copula of the bivariate t distribution with correlation rho and nu degrees of freedom, the parameterisation of R’s copula::tCopula (param = rho, df = nu) and of VineCopula family 2 (par, par2). Unlike the Gaussian copula (its limit as nu grows) it has tail dependence, equal in both tails,

\[\lambda_L = \lambda_U = 2\, T_{\nu + 1}\left(-\sqrt{ \frac{(\nu + 1)(1 - \rho)}{1 + \rho}}\right),\]

so it is the elliptical copula for joint extremes: failures that come together early, or survivals that come together late. Its Kendall’s tau is the Gaussian’s, \(\frac{2}{\pi}\arcsin\rho\), whatever nu. The fit starts rho from the data’s Kendall’s tau and nu from the best of 1, 2, 4, …, 64.

When the data show no more tail dependence than a Gaussian copula, the likelihood keeps increasing with nu (towards that limit) and has no finite maximum; the fit then warns and recommends the Gaussian copula. The criterion is that the Gaussian copula, with its own best rho and the same margins, is at least as likely (to the optimiser’s tolerance, 1e-4) as the t copula the search reached.

Examples

>>> from surpyval import Weibull
>>> from surpyval.multivariate import StudentT
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> model = StudentT.from_params([0.7, 4.0], margins)
>>> model
Copula SurPyval Model
=====================
Copula    : StudentT
Parameters: rho=0.7, nu=4
Margins   : Weibull, Weibull
Fitted by : given
>>> round(model.kendall_tau(), 4)
0.4936
>>> [round(x, 4) for x in model.tail_dependence()]
[0.3907, 0.3907]
cdf(u: Any, v: Any, rho: Any, nu: Any) → Any

The copula, the bivariate t CDF at the margins’ t quantiles.

It is the integral of the h-function over the first margin,

\[C(u, v) = \int_0^u P(V \le v \mid U = s) \, ds = v - \int_u^1 P(V \le v \mid U = s) \, ds\]

(the shorter of the two is used), evaluated by tanh-sinh quadrature (105 nodes per piece), split where the integrand changes fastest – where the conditional median of Y passes y, at \(s^* = T_\nu(y / \rho)\) – so that a sharp step at strong dependence sits at the end of a piece, where the rule’s nodes crowd. The rule also absorbs the algebraic singularities of the integrand at s = 0 and 1 (it approaches the tail-dependence limits as a power of s). It agrees with the exact bivariate t CDF of Genz (2004; R’s mvtnorm::pmvt(algorithm = TVPACK()), integer nu only) to about 1e-11 (7e-11 at rho = 0.999), and with mpmath’s 30-digit integration at non-integer nu to 1e-15. scipy’s multivariate_t.cdf is a randomised quasi-Monte Carlo integration: about 1e-4 off at its default tolerances and different on every call, which an optimiser cannot use; vinecopulib interpolates linearly between the integers either side of a non-integer nu (1.4e-4 off at nu = 2.5).

closed_bounds: tuple = ()

Names of parameters whose finite bounds are themselves valid values (Gumbel’s theta = 1 is the independence copula). The fitter never reaches a bound, but from_params may be given one.

dependence_limits: dict = {-1: 'rho tends to -1', 1: 'rho tends to 1'}

The Frechet bounds the family reaches only as its parameter runs to a limit: +1 the comonotone copula (perfect positive dependence), -1 the countermonotone one, each mapped to that limit as the fit’s warning names it (see _warn_if_perfectly_dependent()). Empty for a family that is not known to reach either.

du(u: Any, v: Any, rho: Any, nu: Any) → Any

dC/du – the h-function \(P(V \le v \mid U = u)\).

The default differentiates cdf() with autograd; override it when a closed form is known. u and v broadcast against each other, and the result has their common shape.

dv(u: Any, v: Any, rho: Any, nu: Any) → Any

dC/dv, the h-function \(P(U \le u \mid V = v)\); u and v broadcast as for du().

fit(x: ArrayLike, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, margins: Any = None, how: str = 'IFM', xl: ArrayLike | None = None, xr: ArrayLike | None = None, init: ArrayLike | None = None, rotation: int = 0) → Any

Fit the copula and its margins to multivariate survival data.

Parameters:
  • x – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • c – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • n – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • t – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xl – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • xr – Multivariate survival data; see MultivariateSurpyvalData for the accepted shapes.

  • margins (sequence of length D) – Either surpyval distribution classes (e.g. surpyval.Weibull) to be fitted, or already-fitted models exposing ff/df. Under "IFM" a fitted model is used as it is. Under "MLE" a fitted parametric model supplies the starting values and its configuration: it is re-estimated jointly with the copula with the same offset, limited-failure or zero-inflated option, and any parameters it was fitted with fixed stay at their values. A non-parametric margin (a class such as surpyval.KaplanMeier, or a fitted non-parametric model) can only be used with "IFM": it gives the semi-parametric pseudo-likelihood estimator, and the likelihood and criteria then compare copula families with the same margins only.

  • how ({"IFM", "MLE"}) – "IFM" (default) fits each margin independently then the single copula parameter (robust two-stage estimation). "MLE" jointly optimises copula parameter + margin parameters.

  • init (array like, optional) – Starting value of the copula parameter(s) for the search, one per entry of parameter_names, each strictly inside the family’s bounds. By default each family starts from its own guess (the built-in families match the empirical Kendall’s tau).

  • rotation ({0, 90, 180, 270}, optional) – Rotate the copula by this many degrees, in the convention of R’s VineCopula (families 13, 23, 33 for the Clayton, …): 180 is the survival copula, C(u, v) = u + v - 1 + C_0(1 - u, 1 - v), with the tail dependence moved to the other tail (a Clayton’s to the upper tail, a Gumbel’s or Joe’s to the lower); 90, C(u, v) = v - C_0(1 - u, v), and 270, C(u, v) = u - C_0(u, 1 - v), give negative dependence with the family’s shape, its tail in a corner where one series is short and the other long. The parameter keeps its own range (pyvinecopulib’s convention; VineCopula writes it negated for 90 and 270). Only the Clayton, Gumbel and Joe copulas are rotated; for a radially symmetric family (Frank, Gaussian, Student-t) 180 is the family itself and 90 its negative parameter. Default 0, the family as it is.

Returns:

The fitted model: the copula parameter(s) params and the fitted margins, with the joint sf/cdf/pdf, sampling, dependence measures and the likelihood-based log_likelihood/neg_ll/aic/bic.

Return type:

CopulaModel

Warns:

UserWarning – “No finite maximum” when the rows observed in both dimensions are perfectly dependent (Kendall’s tau of +-1) and the family reaches that dependence only as its parameter runs to a limit: the returned parameter is then meaningless.

Examples

Simulate from a Clayton copula with Weibull margins, then recover it:

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> model = Clayton.fit(X, margins=[Weibull, Weibull])
>>> model.params.round(3)
array([2.293])
>>> round(float(model.kendall_tau()), 3)
0.534
fit_from_df(df: Any, x_cols: list[str], c_cols: list[str] | None = None, n_col: str | None = None, xl_cols: list[str] | None = None, xr_cols: list[str] | None = None, tl_cols: list[str] | None = None, tr_cols: list[str] | None = None, **fit_options: Any) → Any

Fit the copula and its margins to the columns of a pandas.DataFrame.

Each argument names, per dimension, the columns fit() takes as arrays: x_cols=["a", "b"] reads the two series from columns a and b. The names are those of the univariate fit_from_df (Weibull.fit_from_df(df, x_col=..., c_col=...)) with _cols for a list of columns, one per dimension (principle 21); every other fit() option (margins, how, init) is passed to it unchanged.

Parameters:
  • df (pandas.DataFrame) – The data, one row per unit.

  • x_cols (list of str) – The column of each dimension’s values.

  • c_cols (list of str, optional) – The column of each dimension’s censoring flags. Defaults to every value observed.

  • n_col (str, optional) – The column of row counts.

  • xl_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • xr_cols (list of str, optional) – The columns of each dimension’s interval ends, where the censoring flag is 2.

  • tl_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • tr_cols (list of str, optional) – The columns of each dimension’s left / right truncation.

  • **fit_options – Every other option of fit().

Returns:

The model fit() returns for the same arrays.

Return type:

CopulaModel

Examples

>>> import pandas as pd
>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> df = pd.DataFrame(X, columns=["pump", "motor"])
>>> model = Clayton.fit_from_df(
...     df, x_cols=["pump", "motor"], margins=[Weibull, Weibull]
... )
>>> model.params.round(3)
array([2.293])
from_params(params: Any, margins: Any, rotation: int = 0) → Any

Build a CopulaModel from a known parameter and margins, without fitting.

Parameters:
  • params (array like) – The copula parameter(s), e.g. [theta] (empty for the independence copula): one per entry of parameter_names, each strictly inside the family’s bounds (Gumbel’s theta = 1, the independence copula, is allowed too). Anything else raises ValueError.

  • margins (sequence of length 2) – Fitted (or from_params) univariate models, one per dimension, each exposing ff and df.

  • rotation ({0, 90, 180, 270}, optional) – The rotation of the copula, as for fit().

Returns:

The model, for evaluation and simulation.

Return type:

CopulaModel

Examples

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> model = Clayton.from_params([2.0], margins)
>>> round(float(model.kendall_tau()), 3)
0.5
kendall_tau(rho: float, nu: float) → float

Kendall’s tau, \(\frac{2}{\pi}\arcsin\rho\) (the same for every elliptical copula, so independent of nu).

neg_ll(params: Any, dims: list, weights: NDArray) → float

The copula-stage negative log-likelihood (used by the fit).

pdf(u: Any, v: Any, rho: Any, nu: Any) → Any

The density, the bivariate t density over the product of its margins’ densities, in log space.

rotatable: bool = False

the families whose rotations are new copulas (the asymmetric Archimedean ones).

Type:

Whether rotated() applies

rotated(rotation: int) → Copula

The family rotated by rotation degrees (0, 90, 180 or 270), in R’s VineCopula convention; see the rotation option of fit(), which uses it. rotated(0) is the family itself.

Examples

>>> from surpyval.multivariate import Clayton
>>> survival = Clayton.rotated(180)
>>> survival.tail_dependence(2.0)
(0.0, 0.7071067811865476)
>>> Clayton.rotated(90).kendall_tau(2.0)
-0.5
rotation: int = 0

0 for the family as it is.

Type:

The rotation of the copula in degrees (see rotated())

sample_uv(size: int, params: Any, random_state: int | None = None) → tuple[NDArray, NDArray]

Draw (u, v) pairs by conditional inversion in closed form: given u (x) and a uniform w, \(v = T_\nu(\rho x + \sigma(x) T_{\nu+1}^{-1}(w))\).

spearman_rho(rho: float, nu: float) → float

Spearman’s rho, \(12\,E[UV] - 3\), which has no closed form for the t copula.

\(E[UV] = \int_0^1 u\, E[V \mid U = u] \, du\), and given U = u (X = x) the second coordinate is Y = rho x + sigma(x) Z with Z a t variate with nu + 1 degrees of freedom, so \(E[V \mid U = u] = \int_0^1 T_\nu(\rho x + \sigma(x) T_{\nu+1}^{-1}(q)) \, dq\). Both integrals are taken by tanh-sinh quadrature; the result agrees with the integral of the copula’s CDF to about 1e-10.

tail_dependence(rho: float, nu: float) → tuple

Lower/upper tail-dependence coefficients (lambda_L, lambda_U).

Default (0.0, 0.0) (no tail dependence); families override.

Fitted Model

class surpyval.multivariate.parametric.copula.copula_model.CopulaModel(copula: Any, params: ArrayLike, margins: Any, data: Any = None, how: str = 'given', k: int | None = None)

Bases: SerialisableMixin, MultivariateDistribution

A fitted bivariate copula glued to two univariate margins.

copula

The copula family.

Type:

Copula

params

The fitted copula parameter(s) (empty for the independence copula).

Type:

numpy.ndarray

margins

The fitted margin models (each exposes ff/df/qf).

Type:

list

k

The number of parameters the fit estimated: the copula’s plus those of every margin the fit estimated (all of them for how="MLE"; under how="IFM" those passed as distributions, not a margin passed already fitted). None for from_params.

Type:

int or None

Examples

Copula.fit and Copula.from_params return one:

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> model = Clayton.from_params([2.0], margins)
>>> model
Copula SurPyval Model
=====================
Copula    : Clayton
Parameters: theta=2
Margins   : Weibull, Weibull
Fitted by : given
>>> model.sf([[5, 15], [10, 20]]).round(4)
array([0.624 , 0.2354])
>>> round(float(model.kendall_tau()), 3)
0.5
aic() → float

Akaike’s information criterion, \(2k - 2\ln L\), with k the number of estimated parameters (see the class docstring). Lower is better.

bic() → float

The Bayesian information criterion, \(k \ln N - 2\ln L\), with N the number of joint observations (rows, weighted by n) in which at least one series failed – was observed exactly, or left- or interval-censored – or, when no row has a failure, the number of rows. It is the sample size of every SurPyval BIC: a unit right-censored in every series adds nothing. Lower is better.

cdf(x: ArrayLike) → NDArray

Joint CDF P(X_1 <= x_1, X_2 <= x_2).

conditional_cdf(x: ArrayLike, given_dim: int = 0) → NDArray

P(X_other <= x_other | X_d = x_d) – the copula h-function.

given_dim=0 conditions on the first series, giving \(P(X_2 \le x_2 \mid X_1 = x_1)\); given_dim=1 conditions on the second. Any other value raises ValueError (it used to be read silently as 1).

ff(x: ArrayLike) → NDArray

Alias of cdf() for consistency with surpyval naming.

classmethod from_dict(model_dict: dict) → CopulaModel

Rebuild a copula model from a to_dict() dictionary.

classmethod from_json(fp: str | PathLike) → Any

Load a model from a JSON file written by to_json(), or from the JSON text it returned (a string starting with {).

kendall_tau() → float

Kendall’s rank correlation implied by the fitted copula.

property log_likelihood: float

The maximised log-likelihood, -neg_ll().

neg_ll() → float

The negative log-likelihood of the fitted model: the full joint likelihood of the data it was fitted to, with each row’s censoring (right, left, interval, per series), its truncation and its count n, and the margins’ densities for the observed entries. Raises ValueError for a from_params model.

Examples

>>> from surpyval import Weibull
>>> from surpyval.multivariate import Clayton
>>> margins = [
...     Weibull.from_params([10, 2]),
...     Weibull.from_params([20, 3]),
... ]
>>> X = Clayton.from_params([2.0], margins).random(300, random_state=0)
>>> model = Clayton.fit(X, margins=[Weibull, Weibull])
>>> round(model.neg_ll(), 3)
1729.371
>>> model.k, round(model.aic(), 3)
(5, 3468.741)
property parameter_names: list[str]

the copula family’s parameter_names (["theta"], ["rho"], or [] for the independence copula). The margins’ parameters are on the margins.

Type:

The names of params, entry by entry

pdf(x: ArrayLike) → NDArray

Joint density c(F_1, F_2) f_1 f_2.

random(size: int | tuple[int, ...], random_state: int | None = None) → NDArray

Draw correlated samples: an array of shape (size, 2) for an integer size, one row per draw, or (*size, 2) for a tuple (the two series on the last axis).

sf(x: ArrayLike) → NDArray

Joint survival P(X_1 > x_1, X_2 > x_2).

spearman_rho() → float

Spearman’s rank correlation implied by the fitted copula.

tail_dependence() → tuple

The lower and upper tail-dependence coefficients (lambda_L, lambda_U) of the fitted copula.

to_dict() → dict

Serialise to a plain dictionary: the copula family, its parameter(s), the fit method and each margin’s own to_dict. The data is not stored, but for a fitted model the negative log-likelihood, parameter count and BIC’s sample size are, so the restored model still reports neg_ll/aic/bic. Restore with from_dict() or surpyval.from_dict.

to_json(fp: str | PathLike | None = None, with_data: bool = False) → str | None

Write to_dict() to fp as strict JSON, or return it.

Parameters:
  • fp (str or os.PathLike, optional) – The file to write. Without it the JSON is returned as a string (as pandas.DataFrame.to_json does), which from_json also reads.

  • with_data (bool, optional) – Write to_dict(with_data=True), which also stores the fitted data, for the models whose to_dict takes with_data (the univariate Parametric and NonParametric); a TypeError for any other model. Defaults to False.

Data

class surpyval.multivariate.parametric.data.MultivariateSurpyvalData(x: ArrayLike, c: ArrayLike | None = None, n: ArrayLike | None = None, t: ArrayLike | None = None, xl: ArrayLike | None = None, xr: ArrayLike | None = None)

Bases: object

Normalise and hold row-aligned multivariate survival data.

Parameters:
  • x (array-like, shape (N, D) or sequence of D length-N arrays) – Point values per dimension. For an interval-censored entry (c == 2) the point value is ignored and xl/xr are used.

  • c (array-like, shape (N, D) or (D,), optional) – Per-dimension censoring codes in {0, 1, -1, 2}. A single row of D codes applies to every row. Defaults to all observed.

  • n (array-like, shape (N,), optional) – Integer weight (count) of each row. Defaults to all ones.

  • t (array-like, shape (N, D, 2), optional) – Per-dimension truncation window [tl, tr]. Defaults to (-inf, inf) (no truncation).

  • xl (array-like, shape (N, D), optional) – Interval-censoring bounds, required where c == 2.

  • xr (array-like, shape (N, D), optional) – Interval-censoring bounds, required where c == 2.

Raises:

ValueError – If an array has the wrong shape, a censoring code is not one of the four, or a series (with the counts and its own truncation window) is not valid univariate data: a NaN value, a count that is not a positive whole number, an interval with xl >= xr, or a value outside its truncation window. The message names the series ("Series 0: ...").

Examples

A list is read as one sequence per series. Here three rows of two series, the second series right censored in every row:

>>> from surpyval.multivariate import MultivariateSurpyvalData
>>> data = MultivariateSurpyvalData(
...     [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], c=[0, 1]
... )
>>> data.N, data.D
(3, 2)
>>> data.c
array([[0, 1],
       [0, 1],
       [0, 1]])
>>> data.dimension(1)[:2]
(array([4., 5., 6.]), array([1, 1, 1]))
dimension(d: int) → tuple

Return (x, c, xl, xr, tl, tr) arrays for series d.