Design Principles
These are the rules every part of SurPyval keeps, whichever model you use. They exist so that what you learn about one model holds for the others, and so that a contributor or reviewer can check a change against a fixed list rather than against memory. Conventions gives the details of several of them (the data format, missing values, seeds, saving and loading).
A principle is only as good as its check, so each one names the tests that
enforce it. Most are properties: general statements checked for every
registered model by the conformance suite (surpyval/tests/conformance,
see Contributing), or on generated data by the property-based tests
(surpyval/tests/properties). Where a model is known to break a
principle, the test marks it as a strict expected failure naming the issue
that tracks the fix, so the suite stays green and turns red the day the fix
lands.
Each principle is marked checked (every model it applies to is tested), partly checked (tested for some models or some cases; the gap is named), or judgement (a matter of review that no test can fully decide).
Inputs
One data model. Every fitter takes
x,c,n,t,tlandtrwith the same meanings: censoring flag 0 observed, 1 right, -1 left, 2 interval; intervals are \((l, r]\) and truncation windows \((t_l, t_r]\).Partly checked. The property tests generate data in every form (mixed censoring, ties, counts, truncation, tiny samples) for the univariate, regression and recurrent fitters, and the reference tests check the conventions against R; the other families are checked through their conformance fixtures only.
Invalid input raises a
ValueErrorthat names the argument and says how to fix it – never aTypeError, anIndexErroror a silently wrong answer.Partly checked by
properties/test_validation.py, for the univariate fitters.Missing values.
nanin,nanout at prediction. At fit, a row with a missing value is dropped with one warning where rows are independent, and refused where a row is part of one unit (a recurrent item, a degradation path).Checked by
conformance/test_missing.py.Order doesn’t matter. The order of the data rows never changes a fit, and permuting a query permutes the result.
Checked by
conformance/test_metamorphic.pyandconformance/test_vectorisation.py, and on generated data by the property tests.Counts equal repetition.
n = kgives the same answer askidentical rows.Checked by
conformance/test_metamorphic.py.Units don’t matter. Rescaling time rescales the answer and nothing else. Nor does a covariate’s origin: with
center=Trueevery regression gives the same model when a constant is added to a covariate, and so do the defaults of Cox, Fine-Gray and the families whose baseline maps exactly between origins.Checked by
conformance/test_metamorphic.py(test_covariate_origin*for covariates). One family is excepted: the Beta4’s likelihood is unbounded, so its maximum-likelihood fit can depend on the units, and warns when it does; its MPS fit does not (#385).
Outputs
Shape in, shape out. A scalar query gives a numpy scalar, a 1-D or 2-D query a result of its shape, and an empty query an empty result of its shape; a two-sided confidence bound adds a last
[lower, upper]axis. With covariates the shape is that of the times, rows and times paired: one row for every time, or one time for every row; other counts raise.grid=Trueon the Cox and parametric regression functions gives the row-by-time grid,(n_rows,) + x.shape, which the survival tree and forest return by default.surpyval.utils.shapesapplies the rule at every model’s public methods.Checked by
conformance/test_vectorisation.pyandcb_shapeinconformance/test_options.py, for every registered model, and for the time-varyingsf_tvcbyconformance/test_tvc.py.The functions of a model agree with each other. \(S + F = 1\), \(H = -\log S\), \(f = h S\),
qfinvertsff, and the causes’ cumulative incidences sum to \(1 - S\).Checked by
conformance/test_identities.py.Valid and accurate values. Survival stays in \([0, 1]\) and never increases; cumulative quantities never decrease. The documented exception is the parametric additive hazards models, whose survival can exceed 1 where \(h_0 + \beta'Z < 0\) (#376); the Lin-Ying model predicts with the running maximum of its estimate. A distribution’s functions are accurate to double precision wherever the value is representable, in the tails and at extreme parameters too.
Checked by
conformance/test_bounds.py, and for accuracy byreference/test_tails.pyagainst 50-digit mpmath values, for every distribution with a closed form.Covariate rows are independent. Evaluating rows together gives the same as evaluating them one at a time.
Checked by
conformance/test_vectorisation.py.Behaviour outside the data is defined and documented. A parametric model is defined everywhere by its formula. An estimate with no formula for its shape – a step estimate, a semi-parametric baseline – starts at its initial value before the first time (survival 1, everything cumulative 0; at time 0 for the additive hazards model, whose covariate effect acts from time 0), and after the last time either holds its last value (the single-event estimates) or is
nan(the recurrent mean cumulative functions), the same for all of a model’s functions.set_support(lower, upper)gives a non-parametric estimate an explicit support: its start value fromlowerto the first time, its last value held up toupper, andnanoutside, for every function,interpand confidence bound.Checked by
conformance/test_outside_data.py, which also sets bounds on every non-parametric estimate, requires each to haveset_support, and requires every bound to benanoutside the data when none is set; and for time-varying covariates byconformance/test_tvc.py, for a path that starts before 0, at 0 or later, and for queries at 0 and below.
Estimation
A fit returns what it claims: the optimum of its stated estimator. If the estimator has no optimum on the data, the fit refuses or warns; it never returns a silent degenerate answer. Where the data do not determine a coefficient – a covariate column that is constant where the model has an intercept, or a combination of the others – the fit says so: the coefficient is
nanand listed inaliased, with one warning naming the column, rather than an arbitrary value.Partly checked. The property tests check that parametric fits, on generated data with every kind of censoring and truncation, are local optima of a likelihood the test computes itself from the fitted model’s functions (
properties/test_parametric.py), and the reference tests compare fits with R, lifelines and scikit-survival;calibration/test_refit_registry.pyrefits every registered model to data drawn from itself (nightly). Where the likelihood has no finite maximum, univariate MLE refuses and the regression, frailty, Fine-Gray, copula, mixture and degradation fits warn “No finite maximum” (#392); known gap: abutting intervals such as (1, 3] and (3, 5], whose likelihood has a flat ridge.conformance/test_aliasing.pyrefits every registered model that has coefficients with a repeated covariate column, and with a constant one where it has an intercept, and requires the aliasing and otherwise the fit without the column; the time-varying fits are checked inunivariate/regression/test_aliasing.py. The derivatives a fit takes – the gradient and Hessian with which it searches, verifies its maximum and computes its covariance – agree with finite differences at the fit, for every registered model that takes them (conformance/test_derivatives.py, #562).Failure is never silent. An optimiser that does not converge warns, and a fit never quietly returns its starting values.
Checked by
conformance/test_convergence.py: every iterative fit is starved (an iteration limit of 1, a start a million times the answer, or data with no maximum) and must warn, raise, or still reach the maximum; a closed-form or exact estimator is excluded, with the reason. A fit accepts an optimiser’s answer only when it is a verified maximum (zero gradient, positive-definite Hessian), and a fit giveninitis also started from the default start. A likelihood with no finite maximum warns so (#392), whatever the model. And byconformance/test_maximum.py: every maximum-likelihood fit in the registry – the univariate distributions, mixtures, the parametric and semi-parametric regressions, frailty, competing-risks, recurrence and copula models – records what it reached as its model’smaximum("verified","unverified"or"no finite maximum"), warns exactly when that is not a verified maximum, its fixture’s fit, its starved fit and its time-varying-covariate fit alike; and a verified maximum passes an independent check at the reported parameters (the gradient of the model’s own likelihood ~0 and its Hessian positive definite, a parameter on a boundary of its space held out where the likelihood does not rise off it). Known gap: the degradation process and destructive fits (#564).Entry points agree.
fit,fit_from_df, a formula,from_paramsandfit_tvcgive the same model for the same data.Checked by
conformance/test_fit_paths.py, and for a regression model’s attributes (every builder,from_dictincluded, gives the same declared attributes) byconformance/test_attributes.py.Defaults are the statistically best standard choice, and the same everywhere. For example, every Cox fit defaults to Efron’s tie handling, which is far less biased than Breslow’s.
The first half is judgement, informed by the calibration studies and the literature. The second half is checked by
conformance/test_defaults.py: every entry point of a fitter gives a shared argument the same default, and the tie-handling default is the same across the Cox-based models.Conventions follow the established references – R’s
survival,cmprskandpec, lifelines, scikit-survival – unless there is a documented reason to differ.Checked by
surpyval/tests/reference, which compares results with values those packages computed (regenerated byscripts/reference/regenerate.sh).
Uncertainty
Intervals achieve their stated coverage, and tests their stated size.
Partly checked by the calibration studies (
surpyval/tests/calibration, run nightly), which cover the main parametric, non-parametric, Cox, regression, degradation and recurrent bounds and the hypothesis tests, not every model.Intervals behave consistently. Bounds contain the estimate and stay in the valid range; a one-sided bound is the matching end of the two-sided bound; a higher confidence level gives a wider interval.
Checked by
conformance/test_options.py, over every uncertainty method, confidence level and option of every registered model.
Behaviour and API
One seed rule. Every method that draws takes
random_state.Nonedraws from numpy’s global generator, sonp.random.seedreproduces it; an explicit seed gets its own stream and leaves the global one alone.Checked by
conformance/test_seeds.py, for every registered model that draws.Saving and loading. Every model round-trips through strict JSON with identical predictions, stamped with the oldest schema version that can read it.
Checked by
conformance/test_serialisation.pyandproperties/test_serialisation.py, and for tuple and mixedstr/intcause labels byconformance/test_labels.py.Consistent names. The same option has the same name, meaning and default everywhere (
alpha_ci,bound,on,interp,Z,random_state,n_boot,tie_method,event, andxfor the times andpfor a quantile’s probability), and so does the same attribute: every model’s fitted values areparams, named entry by entry by the attributeparameter_names. Every DataFrame entry point (fit_from_df,fit_tvc_from_df,fit_tvc_timeline_from_df) names a column argument after thefitargument it fills with a_colsuffix,_colsfor a list of columns:x_col,c_col,n_col,xl_col,xr_col,tl_col,tr_col,i_col,e_col,y_col,Z_cols. When a name changes, the old one keeps working for one release with aDeprecationWarningnaming the new one.Checked by
conformance/test_options.py,conformance/test_params.pyand, for the column names,conformance/test_fit_paths.py.Warnings and errors. One warning per problem, with counts, saying what happened and what to do about it. No raw numpy warning escapes from package code.
Checked by
conformance/test_warnings.py, which fails for any raw numerical warning raised inside SurPyval while fitting or predicting with a registered model.Documentation. Every public item is documented with a runnable example, and every number quoted in the prose is checked.
Checked for what exists: the documentation build runs every example and the hidden checks of the quoted numbers (see Contributing), and the docstring examples run as tests. Checked for completeness by
conformance/test_documentation.py: every public item has a docstring with an example, and a new one without fails.Simple by default; more as an option. When a method reaches its limit – data it cannot handle, an approximation that breaks down, a question that needs a heavier computation – the new approach is added as an option beside it, not put in its place. The default stays the simple, standard method that serves the usual case, so a plain call stays fast and easy to explain, and its results do not move. A default changes only when it is wrong for the usual case (principle 15; for example a band that did not hold its level, #390), not because a better method exists for a harder one. For example, trees split greedily by default and take conditional inference with
selection="ctree";cbgives Wald bounds andbootstrap_cbresamples.Judgement, applied in review: a change to a default says which principle the old default broke.
Adding to the list
When a bug is fixed, ask which principle it broke. If the principle’s check did not catch it, extend the check – a new property in the conformance suite covers every model at once. If no principle covers it, it may be a new one: add it here with its check.