Changelog
v0.22 (3 October 2026)
Upgrading from 0.21. A breaking release. Most code needs no change; the items most likely to need one are listed here, and every entry below says what changed.
Arguments 0.21 deprecated now raise
TypeError(see Removed). Run your code on 0.21 withpython -W error::DeprecationWarningfirst.The life models moved to
surpyval.life_models(life_models.Power,life_models.Exponential, …); the old top-level names warn until v0.23.Error and warning messages have one wording per condition: code that matches message text (an unknown option,
alpha_ci, a cause, a covariance, column lengths, “Monotone partial likelihood”, which is now “No finite maximum: the partial likelihood keeps increasing …”) must update its patterns.model.formulais thestryou gave, for every model (Cox, Buckley-James and competing-risks PH kept a parsedFormula).Some results move, each towards a better answer: fits that stopped short now reach a verified maximum or warn (every likelihood fit has
maximum); accelerated-life and AFT time-varying standard errors use the exact information (an example’sse(a)was 317, correctly 574); Gaussian and Student-t copulas no longer clip rho at 0.9999.The bundled lung and Rossi data’s event columns are 1 for an event, as in R and lifelines.
Importing surpyval corrects autograd’s derivative of
np.wherefor the whole process (#562).
Versioning. From this release, versions have two parts,
MAJOR.MINOR (0.22, tagged v0.22); every release takes the
next minor number. pip compares 0.22 and 0.22.0 as equal.
Removed. The names 0.21 deprecated (#422) are gone. An old argument
name (seed, confidence, B, t, q, u, CoxPH’s
method, id_col, time_col, the competing-risks how and
cause, CompetingRisks’ method) is now an unknown argument and
raises TypeError. The degradation calls in the old positional order
(Z last) are no longer recognised and raise, except that a process
model fitted with stress now reads random(size, a, b) as Z=a,
random_state=b. Also removed: the fitted CompetingRisks.method and
CompetingRisksProportionalHazards.how aliases (use .how and
.model), the surpyval.experimental alias (use surpyval.beta.ml)
and band’s unused n_sims and random_state. Saved models still
load. On 0.21, run your code or tests with python -W
error::DeprecationWarning first to find the calls to update.
Behaviour changes. Fits accept an optimiser’s answer only when it is a
verified maximum, and a fit given init is also started from the default
start, so a few fits that stopped short silently now reach a better
maximum or warn; data with no maximum raise ValueError. A
non-parametric df is the probability of each step. Kaplan-Meier and
Nelson-Aalen keep the estimate over a step with no one at risk, as R’s
survfit does. Gray’s test and the competing-risks Cox incidences now
match R. Unknown option values raise ValueError everywhere. The
Uniform’s MLE refuses censored data again. Bernoulli’s sf is
P(X > x), as for every other discrete distribution. Fits whose data
have no finite maximum warn “No finite maximum”. A covariate column
the others already account for gets a nan coefficient and a warning,
as in R. FrailtyModel.summary() returns a DataFrame. Covariate rows
that cannot be paired with the times raise ValueError. The bundled
Rossi data’s arrest is 1 for an arrest. Trend tests report a trend
only when it is significant. qf outside [0, 1] is nan.
param_names is deprecated in favour of parameter_names. The
bundled lung data’s status is 1 for a death. CoxPH.check_ph()
returns a DataFrame. Durations
and dates are refused. Probability plots draw failures only. fit_best
no longer considers the Uniform and Beta4 by default. Small-sample Wald
bands change (#477).
Changed: load_lung()’s status means a death (#509). It was stored as SurPyval’s censoring flag (0 = death), the opposite of lifelines, so
c = 1 - statusfitted the complement (a Kaplan-Meier median of 588 days instead of 310).statusis now 1 for a death, as in lifelines and R; passc = 1 - status.A fitted model’s printout shows its data (#508). For example
Data : 60 units: 9 events at 9 unique times, 51 right censored, with left, interval and truncated counts when present, counted in units (weighted byn), and the number of distinct event times. Parametric, mixture, non-parametric, parametric regression, Cox and Buckley-James models print it, and keep it throughto_dict. A “1 = failed” column passed ascis now visible at a glance.Changed: probability plots take label= and color= (#510).
Parametric.plotandMixtureModel.plotdraw the points, fitted line and bounds in one colour (by default the axes’ next colour), withlabel=on the fitted line, so fits overlaid on one plot can be told apart. The fitted line is solid and the bounds dashed; they were a black dashed line and red bounds. Overlaid plots keep both ranges in view.Changed: CoxPH.check_ph() returns a table (#514). A
DataFrameas R’scox.zphprints it: a row per covariate and aGLOBALrow, withstatistic,dfandp. The old dictionary isproportional_hazards.diagnostics.check_ph(model).Deprecated: cs(x, X) is cs(x, given) (#514). The time already survived is named
given, as in the regression models’sf_tvc(..., given=);X=works until v0.23 with aDeprecationWarning. Called by position, nothing changes.Plots label their axes (#514). The time axis is “Time” unless it is already labelled; non-parametric plots say “Survival probability” and “Kaplan-Meier estimate” (etc.), not “R” and “Model Survival Plot”. A failure at exactly 0 now points to
zi=True.Imperfect-repair fits are 10-270 times faster (#515). The ARA, Kijima-II and G1 renewal likelihoods took a Python step per item and per event on every evaluation; they now step through event positions across all items at once, with the same virtual ages bit for bit (G1’s likelihood to the last digit).
ARA.fiton 100 items went from 10-16 s to 0.45-0.7 s and on 1000 items from 73 s to 2.4 s;GeneralizedRenewal(kijima="ii")on 1000 items from 16.5 s to 1.8 s;GeneralizedOneRenewalon 100 items from 44 s to 0.16 s.Faster saving and loading (#515).
to_dictandfrom_dictvisited every number of a model’s arrays one at a time; number arrays now go through in one pass. The saved documents are byte-identical. A Kaplan-Meier model with 100,000 rows of data loads in 0.82 s (was 1.22 s).Changed: renewal models test the repair against perfect and minimal repair (#513).
repair_test(alpha_ci=0.05)gives two likelihood-ratio tests: against perfect repair (Kijima q = 0, ARA rho = 1; ARI’s rho = 1 is maximal repair) and against minimal repair (q = 1, rho = 0; G1 has none). Each refits the model with the restoration parameter fixed and reports the statistic and p-value, halved where the tested value is on the edge of the parameter’s range (Self and Liang 1987). The printed model shows the conclusion, for example “consistent with minimal repair; perfect repair rejected” on the issue’s eight haul trucks (against q = 1: LR 0.40, p 0.53; against q = 0: LR 34.9, p 2e-9), rather than a fixed rule on the width of the interval. The refits run the first time the model is printed or tested, and are cached.REML warns when its between-unit covariance is on the boundary.
DegradationAnalysis(population_method="reml")could return a singular covariance (an intercept-slope correlation of 0.999998 on six units) without comment, while the default method warned on the same data. It now warns when the covariance with its smallest eigenvalue removed fits the data as well (tosqrt(eps)): the estimate is on the boundary of the valid covariances, where standard errors and intervals are unreliable. On 144 simulated datasets it flagged all 18 boundary fits and none of the others. The default method’s warning now says that REML may land on the boundary too.Conditional sf_tvc is 1 at and before given (#523).
sf_tvc(x, schedule, given=g)returned \(S(x)/S(g)\) for x < g as well, above 1 (1.75 for WeibullPO and 1.64 for CoxPH on the conformance fixtures). It is now 1 for x <= g, for the parametric regressions and Cox, along step schedules and covariate paths; a conformance property checks it.Changed / deprecated: DataFrame columns are named with _col (principle 21). Every
fit_from_df(andfit_tvc_from_df,fit_tvc_timeline_from_df) names a column argument after thefitargument it fills, with_col(_colsfor a list).Weibull.fit_from_df(df, x=, c=, n=, xl=, xr=, tl=, tr=)is nowx_col=, c_col=, n_col=, xl_col=, xr_col=, tl_col=, tr_col=(tl_colandtr_colalso take a number shared by every row), andDegradationAnalysis,WienerProcessandGammaProcesstakex_col=, y_col=, i_col=, as the regression, recurrent and competing-risks fitters already did. The 0.21 names work until v0.23 with aDeprecationWarning. A conformance test checks the rule on every DataFrame entry point.Changed: the concordance index leaves out tied event times by default.
sp.metrics.concordance_index, the regression models’concordance()andRandomSurvivalForest.scoretaketies="therneau"(the default: two events at the same time are not a usable pair, as in R’ssurvival::concordanceand lifelines) orties="harrell"(Harrell’s original definition, which counted them). On the lung Cox model (age, sex, ph.ecog), with 28 pairs of tied deaths, C is 0.637135, as R and lifelines give, instead of 0.636942. The deprecatedsurpyval.utils.score.scorekeeps Harrell’s convention.fit_from_df on every fitter (#511). Kaplan-Meier, Nelson-Aalen, Fleming-Harrington, Turnbull, RoystonParmar, MixtureModel, the closed-form distributions, the copulas, FineGray, DestructiveDegradation, the survival trees and forest, and the recurrent fitters had no
fit_from_df. They now take aDataFrame, naming the columns as their family already did (x=,c=,xl=,tl=for one lifetime per row;x_col=,i_col=,c_col=,Z_cols=for recurrent and regression data), and give the modelfitgives on the same arrays, which the conformance suite checks for every registered model.Weibull.fit_from_dfno longer castscto an integer, which turned a missing flag into -9.2e18.The concordance index is fast, and a metric (#512). Harrell’s C was a pairwise Python loop: 2.9 s at 5,000 subjects and about 5 minutes at 50,000. A merge sort over the ranked scores gives the same value, with the same tie rules, in 0.01 s and 0.15 s. It is
sp.metrics.concordance_index(x, c, risk), and every regression model hasconcordance(), scoring its training data by default.surpyval.utils.score.scoreis deprecated until v0.23. With tied event times the value differs slightly from R and lifelines (0.6369 vs 0.6371 on the lung Cox model), which do not count two events at the same time as a usable pair.Changed: renewal models print the restoration factor’s uncertainty (#513). A generalized renewal fit to minimal-repair data printed q = 2.63 with no sign that its 95% interval was [0.094, 73.4]. The printout now gives each parameter’s standard error and Wald interval (also
summary()) and says whenqorrhois not determined by the data or sits at the edge of its range. The docstrings say whatqandrhomean.repair_test()tests the fit against minimal repair (on that data LR = 0.40, p = 0.53).Bounds, mean life and accelerated life along a covariate path (#172, phase 2).
cb_tvc(x, Z, xl=None, given=None, on="sf", ...)boundssf,ffandHfalong a step schedule or aCovariatePath, by the delta method on the same scale ascb, with the quadrature mesh held at the fitted parameters; a constant path givescb, and in 1,000 simulated fits the 95% bounds covered the truth 94.6-96.2% of the time.mean_tvc(Z, xl=None, given=None)gives the mean (or withgiventhe mean residual life) in one cumulative pass, accurate to about 1e-15, andinfwith a warning where survival levels off.AcceleratedLifemodels, which refused every path, now follow Nelson’s cumulative exposure along steps and paths for the Weibull, Exponential, Gamma and LogNormal (location families still refuse, saying why). AFT and accelerated life integrate one period of a periodic path, so 10 million cycles take about 1 ms.Fixed: likelihood-ratio bounds widen with the confidence level (#535). An ExpoWeibull 99% likelihood-ratio interval did not contain the 95% one, because the search stopped on a local extreme of the long, curved likelihood region. The 99%
hf(13)lower bound was 0.1046, above the 95% one of 0.1017, and theqf(0.95)upper bound was 36.44, below the 95% one of 40.37. The search now follows each bound outward through the regions at 1/4, 1/2, 3/4 and all of the critical value, and gives 0.0714 and 80.6; the region approaches 0.0708 and 87.0 as alpha goes to 0. Every other registered model’s likelihood-ratio bounds are unchanged to the bit.Breaking: ARI takes its baseline intensity as baseline= (#507). In
ARA,GeneralizedRenewalandGeneralizedOneRenewal,distis a lifetime distribution, but inARIit was the baseline intensity model (CrowAMSAA,Duane,CoxLewis), sodist=sp.Weibullwas an easy mistake.fit,fit_from_recurrent_data,fit_from_dfandfit_from_parametersnow takebaseline=(andfit_from_parameters’sdist_paramsisbaseline_params); the old names work until v0.23 with aDeprecationWarning. Saved models are unchanged, and old files load as before.New: every likelihood fit says whether it reached a verified maximum. The regression fits (PH, AFT, PO, AH, accelerated life and their time-varying forms), Cox, parametric and Cox frailty, proportional odds, Fine-Gray and competing-risks PH, mixture, Royston-Parmar, HPP, NHPP, proportional-intensity, renewal and copula models have
maximum, asParametricdoes:"verified","unverified"or"no finite maximum", agreeing with the fit’s warnings, and saved byto_dict(an older dict reads"unknown"). Lin-Ying and Buckley-James, which solve estimating equations, are"not applicable". A parameter on its bound where the likelihood is highest (a frailty variance of 0, an AMH copula attheta = 1) is a verified boundary maximum. A new conformance property,maximum, checks every likelihood fit in the registry and verifies each answer independently; the degradation process and destructive fits are its known gap (#564).Fixed: fits that kept an unverified answer now polish it or say so. The Nelder-Mead fits (NHPP, proportional intensity, renewal, copulas), Fine-Gray’s BFGS, Cox’s fallback root-finder, the HPP and Royston-Parmar took their optimiser’s answer unchecked. Where results move, they move towards the maximum: Cox-Lewis by 1.7e-5 relative (log-likelihood up 7e-9), a Gaussian copula’s
rhoby 1.9e-6, a G1 renewal’sqby 1e-4.Fixed: the truncated mixture fit kept L-BFGS-B’s answer unverified (#560). It is now polished and verified, as the EM fit is; two equivalent forms of the same data agree to 5e-7 (they differed by 1.8e-5).
Changed: one wording for no finite maximum. Cox’s, Fine-Gray’s, competing-risks PH’s and Cox frailty’s “Monotone partial likelihood: …” now reads “No finite maximum: the partial likelihood keeps increasing …”; update code that matches the old text. Cox frailty’s EM warning says it “did not reach a verified maximum”, and a mixture with a point-mass component warns once, not twice.
Fixed: rows censored with a finite truncation bound are read as the intervals they are in the fit checks (#559). Data whose every row is right censored with a finite
tr(or left censored with a finitetl) was refused as having no failure, from the raw censoring codes. The checks and the start guess now read each row as the likelihood does. Such data still bounds no failure from one side, so its likelihood has no finite maximum unless a parameter is fixed: a free two-parameter fit is refused, as the same rows written as intervals are; a fit with a parameter fixed equals the interval fit; and the Exponential, Rayleigh, Poisson, Geometric, NegativeBinomial and DiscreteWeibull fits, which stopped atfailure_rate = 3.6e-7orsigma = 313.5and reported it verified, warn “No finite maximum”. MPS still refuses data with no exact value.Fixed: distribution functions are quiet in the far tail and right at infinity (#561).
Weibull.sffar in the tail warned “overflow encountered in power” though its 0 was right; an overflow or division by zero inside a distribution function is no longer warned about (an invalid operation still is). A sweep of every registered distribution at extremexand parameters also found 218 wrongnanvalues, each now the right value: a Gamma’s or Poisson’ssf(inf), a Weibull’sdf(1e300), Gumbel’sdf(inf), a NegativeBinomial past 1.3e154 trials, a discretised Weibull’s hazard from k = 1e6, a zero-inflated model’s hazard at largex. Hazards take their limits at infinity (a Gamma’s rate, a NegativeBinomial’s p). Fitted results are bit-identical.Fixed: derivatives through ``np.where`` with a broadcast argument (#562). autograd’s rule for
np.where(c, x, y)did not reduce the gradient to the shape of a broadcastxory: the gradient came back the wrong shape or raised, and where a later step summed it away the second derivative was silently wrong (7.52 instead of 4.46 in the old accelerated-life substitution). SurPyval registers a broadcast-aware rule when imported (surpyval/utils/autograd_where_compat.py), so every model and every custom distribution written withsurpyval.npgets the right derivatives; no fitted value changes. Behaviour change: importing surpyval changesautograd.numpy.where’s derivative for the whole process. A test notices when autograd fixes this itself, so the patch can go.Fixed: ``Discretize`` has an exact Hessian at the first bin (#562). The probability of
k = 1tookR(0)through the continuous formula, whose Hessian is not finite there (a Weibull’s(0/α)^β), so the covariance silently fell back to a numerical Hessian.Ris now exactly 1 at the start of the support; values are unchanged.A ``derivatives`` conformance property (#562). For every registered model whose fit or inference differentiates its likelihood, the gradient and Hessian it takes at the fit, in its search space, agree with Richardson finite differences to 1e-6 of the standard-error scale (1e-5 through the numerical incomplete gamma and beta shape derivatives); so do a parametric
cb’s delta-method gradients, the copulas’ h-functions and density, and the degradation paths’ Jacobians.Fixed: accelerated life and AFT time-varying fits warn of no finite maximum and use the exact information (#555). Stress levels without failures, or a covariate level with no events, let parameters run off silently; these fits now warn “No finite maximum” once, as the other regressions do. The covariance is the exact observed information instead of a numerical Hessian, which was up to 1e-3 of the standard errors off and much worse on ill-conditioned fits: the Arrhenius example’s
se(a)was 317 where the correct value is 574, so AL standard errors and bounds change. The AL likelihood’s second derivatives were also wrong (autograd’snp.wherewith a broadcast argument: a LogNormalHfwas 13% off; #562 audits the rest). AFT time-varying fits now search with their exact gradient and reach a slightly better maximum (parameters move by up to 6e-5 relative), in half the time.Fixed: HPP proportional intensity restarts and warns (#554). From a poor
initthe single BFGS search returned a rate of 0 withnancoefficients, or its starting point, without a word. A user’sinitis now followed by the default start, the better answer kept, and an answer that is not a verified maximum warns. Default fits are unchanged.AFT models fit a covariate timeline from a DataFrame (#553):
WeibullAFT.fit_tvc_timeline_from_df(every AFT model), as PH, AH, PO and Cox have.Fixed (breaking): the Gaussian and Student-t copulas no longer clip rho to ±0.9999 (#541). Every formula silently evaluated
rho = 0.9999for any larger value:from_params([0.99995])gave a density of 81.11 at (0.3, 0.3001) where the true value is 114.68. They are now accurate for any|rho| < 1(within 1e-14 of a 40-digit integration);rho = ±1is still refused. A fit whose rho runs to ±1, which stopped silently near 0.99999, gives the standard no-finite-maximum warning recommending the comonotone or countermonotone model. Ordinary fits change only in the last digits.Changed: ``RandomSurvivalForest.fit`` is quiet and takes ``n_jobs`` (#546). It printed joblib’s progress on every fit.
n_jobs(default 1, the old sequential behaviour; -1 for every core) grows the trees in worker processes: 20 trees on 400 rows in 3.9 s with 2 jobs against 6.6 s with 1. A seeded forest is identical whatevern_jobsis.Fixed: a non-parametric survival tree or forest crashed on data mixing exact, right- and interval-censored rows (#543). A node holding only exact and right-censored rows kept its parent’s two-column times; it is now split and leafed as its rows would be on their own.
Fixed: mixture fits with a censored row and finite truncation (#544). A right-censored row with a finite
tr(or left-censored with a finitetl) is the interval [x, tr] (or [tl, x]) since #310, but the mixture likelihood grouped rows by censoring code and raisedIndexError. It now uses the same observation masks as every other fitter, and equals a fit to the rows written as intervals.Fixed: zero-inflated fits with a left truncation below 0 (#548). A
tlbelow 0 counted the mass at 0 as already excluded, so a no-optl=-1sentf0to 1 with an unbounded likelihood. The mass now enters a window only from 0 on, as in the model’sff, sotl < 0gives exactly the untruncated fit (withlfpandoffsettoo);tl = 0still excludes it, as for the discrete distributions.Performance sweep. Results are unchanged to the bit except where noted, checked by the new equivalence harness:
Cox residuals,
check_phand robust standard errors are linear time. They looped over the event times with a mask of every row: dfbeta on 10,000 rows took 4.3 s and now takes 0.01 s;check_phwith four residual kinds on 100,000 rows 6.5 s, now 0.3 s.A truncation time at or below the support’s edge truncates nothing and is no longer evaluated. A
tlof 0 sent the covariance to the numerical Hessian: a left-truncated Weibull at 100,000 rows fits in 0.37 s instead of 0.80 s, and the covariance is now the analytic one (the registry’s mixed-censoring Weibull covariance was 0.09% off; the Gamma’s 4e-6). With an offset, rows truncated below the threshold gave a nan gradient, so the fit fell back to Nelder-Mead and stopped 1.23 log-likelihood units short of the maximum with a “not verified” warning; it now converges, in 0.15 s instead of 4.6 s.Likelihood-ratio bounds keep each likelihood their searches evaluate (18% were repeats): 10-60% faster.
Censored Gamma, Beta and Negative Binomial fits compute each incomplete function’s complement only where it is used: a censored Gamma at 100,000 rows takes 0.77 s instead of 1.25 s.
Kaplan-Meier, Nelson-Aalen and Fleming-Harrington sort and validate their data once: 1,000,000 rows in 0.49 s instead of 0.72 s.
metrics.auc_tdcounts pairs by binary search: 100,000 rows and 20 horizons in 0.48 s instead of 15.5 s.Saving and reading models: the schema check walks the document once, and Cox models are read into plain arrays. A Cox JSON round trip at 100,000 rows takes 0.41 s instead of 2.35 s.
Removed: the empty surpyval.alpha package, which has held no models since v0.17.0, and the unused
surpyval.utils.validate_tv_coxph_df_inputs.Development: large modules split (maintainability sweep, phase 1). Code was moved only, checked bit-exact with the equivalence harness. The likelihood-ratio bounds are in
univariate/parametric/_likelihood_ratio.py, the optimised fits inoptimised_fit.pyand their input checks and starts in_fit_inputs.py; the non-parametric support helpers in_support.pyand its bands in_bands.py;surpyval/utils/__init__.pyis split intodata_formats,validation,covariates,numericandwarnings; and the remaining-useful-life classes are indegradation/rul.py. Old import paths keep working.surpyval.utilsnow has an__all__of its 18 documented handlers, converters and helpers; everything else it exports is internal. A test stops new imports of another package’s private names. In the same way, the parametric regression model’s time-varying evaluation is inunivariate/regression/_tvc_evaluation.pyand its covariance and Wald bounds in_inference.py; the Cox partial likelihood (tie terms, likelihood generators and the Newton-Raphson solver) is inproportional_hazards/cox_likelihood.py; the fittedDegradationModelis indegradation/degradation_model.py; andbootstrap_cbis innonparametric/_bands.pybesideband.A model’s formula is the str you gave. Cox, Buckley-James and the competing-risks PH model kept a parsed
formulaic.Formulainmodel.formula, every other model the str; now all of them keep the str, and a saved Cox model’s formula no longer changes on a round trip (a time-varying Cox model saved"1 + dose"for"dose"). Code that read.formulaas aFormulashould parse the str withformulaic.Formula(model.formula).The semi-parametric fits share one input check. Proportional odds, Lin-Ying additive hazards, Buckley-James, Fine-Gray and the competing-risks PH model now drop a row with a missing covariate (with the usual warning) before checking that the observed times are finite, as Cox did; they raised “must be finite” on such a row. The degradation models’ input errors share one wording and name the offending input (“y must contain only finite values”; “x, y, and i must have the same length; got 55, 54, and 55”).
Error and warning wording unified (messages only; no numbers change). Code that matches the old texts must update its patterns.
An unknown option value (
bound,on,how,method,tie_method,kindand the other enumerated arguments) raises'<name>' must be one of 'a', 'b' or 'c'; got <value>.alpha_cioutside (0, 1):'alpha_ci' must be strictly between 0 and 1; got <value>.An unknown cause:
Unknown cause 'x'; the causes are [...](CauseSpecificMCF.mcfandmcf_cbraisedKeyError; they now raiseValueError). A missing cause:<what> is of one cause at a time; pass `event`.No covariance, a singular information matrix, unknown parameter names in
param_cbandfixed(the univariate fit now lists every unknown name), andc,n,tlortrof the wrong length ('c' must be the same length as 'x') each have one wording.fit_from_dfgiven neitherZ_colsnorformula:One of 'Z_cols' or 'formula' must be provided.Every fit that stops short of a verified maximum gives one warning, “… did not reach a verified maximum of the likelihood (<reason>)”, at the caller’s line. The univariate “Precision was lost” and the regression and NHPP “did not converge” warnings are this warning now, and
quiet_maximum_warningsholds it back for every fit, not only the univariate MLE.RandomSurvivalForest.sfrefuses an unknownensemble_method; it used'sf'for anything other than'Hf'.
Development: duplicated code merged (consolidation sweep, phases 1-3). Fitted numbers are bit-identical, checked with the equivalence harness. The PH, AFT and PO fits share
fit_log_linearandsplit_log_linear; the accelerated-life and AFT time-varying fits are assembled byassemble_regression_model; the time-varying DataFrame fits sharefit_tvc_df; the Efron and Breslow Cox generators share one risk-set setup; the semi-parametric models share one covariate centring and oneLinearPredictorMixin; the plain and proportional-intensity NHPP fitters, and HPP, share one log-likelihood (#350); the bootstrap tails sharepercentile_bounds(#351); and the degradation inputs are checked by onevalidate_xy(#352); option checks go throughutils.validation.check_optionand unverified-maximum warnings throughutils.no_maximum.warn_unverified; and thefit_from_dfdesign matrices are built bydesign_matrix_from_dfalone (wrangle_and_check_form_and_Z_colsis removed). The accelerated-life fitter now has the deprecatedparam_namesalias the other fitters have.Development: fitted regression models declare their attributes (maintainability sweep, phase 2).
ParametricRegressionModeland the Cox, ProportionalOdds, Lin-Ying, Buckley-James and Fine-Gray model classes declare every attribute their builders set, with its type and meaning. The covariate links (Phi, the additive link andfrom_dict’s namespace) are oneCovariateLink(name,phi_param_map,phi;Phiremains as an alias so old pickles load).fitnow also setsdist, andfrom_dictsetsdistribution_param_mapandphi_param_map, which only the fits set before. A new conformance property,attributes, checks thatfit,fit_from_df, a formula,fit_tvcandfrom_dictgive a model the same declared attributes (from_dictless the data and what was computed from it) and nothing undeclared; before this change it failed 33 cases. The regression family names are constants inregression/_kinds.py, and the model branches on_is_accelerated_life()and_is_additive()instead of comparingkindwith string literals;kindand its values are unchanged.Development: long functions split into named steps; flake8 ``max-complexity`` lowered from 70 to 25 (maintainability sweep, phase 2).
handle_xicn, the tvc-schedule expression evaluator, the MLE fitter,xcnt_handler,turnbull,DegradationAnalysis.fit, the fit-input validation and the harness’sdiffare split into named steps, and the likelihood-ratio search behindcb(method="lr")is a_PsiBoundSearchclass rather than eleven closures. No result changes: the equivalence harness is bit-identical, and old-against-new runs over thousands of inputs per function give identical values, warnings and error messages. Comments in these modules state current behaviour rather than its history, and the copula and renewal tolerances are named for what they are (_U_CLIP,_LOG_FLOOR,_MACHINE_EPS).Development: tests organised by feature (maintainability sweep, phase 3). The 59 test modules named after fix rounds (
*_fixesN.py,*_roundN.py,test_tvc_phase2.py, …) are renamed as, or merged into, feature modules; the 17,023 collected tests are unchanged apart from their paths. Helpers copied between modules live once insurpyval/tests/_helpers.py, the conformance registry is split into fixtures, family helpers, cases and known failures (still imported throughregistry.py), and Contributing says where a fix’s regression test goes.Development: refactors are proven bit-identical.
scripts/refactor/snapshot.pyrecords what every registered model computes and says (fits, predictions, every bound,to_dict, printouts, warnings, and errors on invalid input), plus time-varying, bootstrap and recurrent fits, the public API, the import set and the test IDs, socompareshows a change altered nothing (see Contributing). CI lints with isort as well, and coversconftest.pyandscripts/; flake8 caps function complexity (now at 25); mypy reports unusedtype: ignorecomments, redundant casts and impossible comparisons; a test fails once the version reachesREMOVED_IN(0.23) while deprecated names are still accepted; and the nightly refit study covers the seven models added in 0.22 (#545).Added: Joe, Ali-Mikhail-Haq and Student-t copulas, and rotations (#157).
surpyval.multivariate.Joe,AMHandStudentThave censoring- and truncation-aware likelihoods, Kendall’s tau, Spearman’s rho and tail dependence, conditional-inversion sampling and serialisation. The Student-t CDF is a deterministic integral within 1e-11 of Genz’s exact algorithm (scipy’smultivariate_t.cdfis randomised and about 1e-4 off). A Student-t fit to data with no tail dependence warns thatnuhas no finite maximum and recommends the Gaussian copula.rotation=(90, 180, 270) for the Clayton, Gumbel and Joe copulas gives dependence in the other tail; 180 is the survival copula. Parameter recovery was checked over 200 replications with censoring for every family.Fixed: copula accuracy (#291). A numerical review of every copula against reference values (pyvinecopulib, R’s mvtnorm, 50-digit mpmath):
Spearman’s rho for Clayton and Gumbel was estimated from 50,000 simulated pairs, up to 5e-3 off (Clayton at theta 0.5: 0.2901 for 0.2949). The default Kendall’s tau and Spearman’s rho of any copula are now integrals, accurate to about 1e-11.
Gumbel’s h-function and density underflowed near the upper corner: the density was inf at (0.98, 0.98) for theta 100. They are now closed forms in log space, within 1e-9 of 50-digit references.
Breaking: with margins passed already fitted, the data were not checked: a NaN returned the starting value with a NaN likelihood, and negative counts,
xl > xrand values outside the truncation window were used. Each series is now checked like univariate data, and the error names the series.
Added: log-normal shared frailty (#343).
Frailty(dist, family="lognormal")fits \(u = e^w\), \(w \sim N(0, \theta)\), as frailtypack, coxme andsurvival::frailty(dist="gaussian")define it, sothetais the variance of \(\log u\) and the median frailty is 1. Each group’s likelihood is integrated by 30-node adaptive Gauss-Hermite quadrature, centred on the group’s mode: within 3e-10 of scipy’s adaptive quadrature for \(\theta \le 1\). On the kidney data it matches R’s lme4 (nAGQ = 25) to 6e-13 in the log-likelihood. Gamma stays the default.frailty_variance(Var(u)/E(u)^2) and the newkendall_taucompare the two families on one scale. Breaking: an unknownfamilyraisesValueError(it wasNotImplementedError).Added: shared frailty with a Cox baseline (#342).
CoxFrailtyfits a gamma frailty with an unspecified baseline, by EM over the frailties withCoxPH’s partial likelihood as the M-step (accelerated by SQUAREM), andthetamaximising the profile likelihood. It reproduces R’scoxph(... + frailty(id, dist="gamma"))on the kidney data: theta 0.40777, thefemalecoefficient -1.5832 with standard error 0.4484, and I-likelihood -181.6386, with Efron ties (Breslow too). It predicts marginally or for an observed group, and saves and loads. In 200 simulated fits its coefficient and theta intervals covered 94% and 97%.load_kidney()adds the kidney catheter data (McGilchrist and Aisbett 1991; R’ssurvival::kidney).Added: a semi-parametric proportional odds model (#341).
surpyval.ProportionalOddsis the proportional-odds counterpart ofCoxPH: the covariates multiply the survival odds of a baseline left to the data, \(S(x|Z)/F(x|Z) = e^{\beta'Z} S_0(x)/F_0(x)\). It is fitted by nonparametric maximum likelihood (Murphy, Rossini and van der Vaart 1997), with an exact solve for the baseline at each step, and its standard errors come from the profile likelihood. A positive coefficient means a longer life, as inLogisticPO; R’stimereg::prop.oddsreports the negatives. It takes observed and right-censored data with left truncation, and refuses other censoring with aValueError. The coefficients agree with R’ssurvival::coxphwith a unit gamma frailty per subject (the same model) to 1e-6 on the lung data and 1e-7 on the Rossi data, and the intervals covered 94-96% in 1,000-replication studies.Changed: the life models are in surpyval.life_models.
Power,InversePower,Eyring,InverseEyring,Linear,InverseExponential,DualPower,DualExponential,PowerExponential,GeneralLogLinearand theLifeModelbase class are used asAcceleratedLife(Weibull, life_models.Power). The exponential (Arrhenius) life model islife_models.Exponential: at the top level that name is the distribution, so it wasExponentialLifeModel. The old top-level names still work until v0.23, with aDeprecationWarningnaming the new one. Every life model’s docstring now gives its formula, what each parameter means (aof the Arrhenius model is \(E_a / k_B\), for example), its constraints and an example.Breaking: GeneralLogLinear fixed, with a constant term, and exported (#530, #345).
AcceleratedLife(dist, GeneralLogLinear).fitraised an autograd broadcastValueErroron any data. The life model is now L(Z) = c exp(sum of beta_j Z_j), with one coefficient per column ofZ; with no constant, L(0) was 1 in whatever unit the times were in (principle 6). With the Weibull it reaches the Weibull AFT maximum (log-likelihood -422.414831 in both). It is importable assp.life_models.GeneralLogLinear, round-trips through JSON (a model saved now cannot be read by 0.21), and is in the conformance registry.LifeModel.resolve(n_stresses)builds the model for a number of columns, so its parameter map and bounds are the typesLifeModeldeclares.Forest importance no longer silently NaN (#533). When an out-of-bag row had zero ensemble probability,
feature_importancesreturned NaN for every feature without a warning (3 trees, 1 row of 46). Each drop is now over the rows that are finite before and after the shuffle, the same as before when every row is, and one warning gives the counts and recommends more trees orkind="exponential".oob_log_likelihoodwarns the same way when it returns -inf.Normal and LogNormal 2-4 times faster (#469).
sf,ff,df,qfand their logs usescipy.specialdirectly instead ofscipy.stats.norm, with bit-identical values:Normal.sfon 1,024 values takes 41 µs instead of 96 µs,qf33 µs instead of 110 µs, and a censored fit 30 ms instead of 42 ms. The density no longer emits numpy’s overflow warning at |x| near 1e300.import surpyval is 0.3 s faster (#470). The regression, competing-risks, recurrent and degradation models and the metrics load on first use, and pandas and formulaic are imported only where used (1.16 s to 0.86 s). Every name,
dir(surpyval)and attribute access such assurpyval.recurrent.laplacework as before.Faster no-maximum check (#501). A coefficient at or near 0, which Newton’s step leaves unchecked, had its profile read with a third derivative and two full Hessians: 65% of a 100,000-row LogNormal AFT fit. It now uses Hessian-vector products. The criterion and its verdicts are unchanged, and its curvature is exact where the old one lost digits. The fit takes 2.7 s instead of 4.8 s.
Faster likelihood-ratio bounds (#519). The searches evaluate the distribution’s formulas without the guards the data cannot trip (bit-identical values), and for two-parameter models the region’s boundary is traced once and each search starts from it. A Weibull
cb(method="lr")at 20 points on 1,000 units takes 2.4 s instead of 9.4 s, and at one point 0.56 s instead of 1.1 s. Bounds are unchanged to 2.3e-7. A steep LogNormal hazard bound that ended 5e-8 outside the region now sits on its boundary.Incomplete beta tails (#520). Where a small tail’s
xrounds towards 1, as with a NegativeBinomialrnear 1e-172, the continued fraction ran to 100,000 terms (4.6 s per call) and was 0.5 nats off. That tail now comes from the other side’s power series, to 1e-15 of mpmath, in milliseconds, and an unconverged continued fraction gives nan instead of a wrong value.Faster degradation bootstrap bounds (#522). Each refit reuses the units’ path fits and warm-starts the life fit from the full-data estimate, and every refit still reaches a verified maximum. 200 units with 200 resamples take 5.4 s instead of 9.7 s, and the bounds change by at most 4e-7.
MCF variance in linear time (#521). The Lawless-Nadeau variance and each item’s observation window took a pass over all times or rows per item. At 5,000 items the fit takes 0.27 s instead of 1.8 s, with the same variance to 2e-13.
Design principle 24: simple by default, more as an option. When a method reaches its limit, the new approach is added as an option beside it. The default stays the simple, standard method, and changes only when it is wrong for the usual case.
Fixed: survival along a continuous path at far-out times.
Hazard shut off by the path.
sf_tvcalong aCovariatePathstarted with one quadrature panel from the last earlier edge to the query. When the path shuts the hazard off, all of the hazard sits near the model’s time scale, and that panel never sampled it. For example, with a Weibull(10, 2) PH model andZ = -0.2t,sf_tvc(1e21)gave 1.0, or sf(1) when queried with t = 1, against the exact exp(-0.5) = 0.6065. The starting mesh now has edges a factor of 2 apart from below the model’s time scale up to the query, about 110 panels for 1e21.Hazard that overflows. A hazard that grows without bound along the path overflowed to inf - inf and gave NaN with a raw numpy warning. It now gives survival 0.
Fixed: an additive hazards model’s cb below 0. It gave a band where
sfis 1 (WeibullAHcb(-1, z)was [0.81, 0.97]); it is now [1, 1].Survival forests grow 8-90 times faster (#190, #518). The Weibull split ran a Nelder-Mead fit for every candidate child (2 trees at n = 300: 21 s). On observed and right-censored data each child’s maximum is now found directly, every candidate of a feature at once: the exponential rate and Weibull scale in closed form, the Weibull shape from its profile likelihood (0.23 s). The log-rank split sorts each feature once and scores every threshold from cumulative counts (10 trees at n = 1000: 2.2 s, was 16.5 s). The chosen splits are unchanged, so seeded forests predict exactly as before. Data with left or interval censoring or truncation still runs an optimiser for each candidate.
A forest’s first prediction is about 20 times faster. Each Weibull or exponential leaf was fitted by a full
Weibull.fitthe first time a prediction reached it: 93 s for a 20-tree forest on 1,000 rows, five times as long as growing it. On observed and right-censored data a leaf is now the maximum found as the split search finds a child’s (the exponential rate in closed form, the Weibull from its profile likelihood, with the shape searched as widely asWeibull.fitdoes), built from its parameters: 4.3 s. The parameters agree withWeibull.fitto its tolerance (about 1e-5), so predictions move in their last digits. Such a leaf has nocb()of its own.Changed: trees and forests know their covariate names (#192).
fit_from_dftakes aformulaas well asZ_cols, andfittakes aDataFrameZ; the fitted model keepsfeature_namesand serialises them,print(tree)shows the splits by name (temp <= 42) and each leaf’s model, and predictions read aDataFrameby name.RandomSurvivalForest.feature_importancesis apandas.Serieskeyed by feature name (it was an array).min_split_gain for the likelihood trees (#189). A
"weibull"or"exponential"node splits only if its best cut raises the maximised log-likelihood by more thanmin_split_gain: a number,"aic"(the kind’s parameters, 1 or 2) or"bic". The default, 0, keeps the old behaviour, which suits a forest;"aic"is the setting for a single tree (on no-effect data, 3.4 leaves on average instead of 32).Split housekeeping (#193).
log_rank_splitcalled directly on left- or interval-censored or right-truncated data raisedIndexErroror returned a wrong split; it now raisesValueError.min_leaf_failurescounts failures weighted bynin every split, so a row with count n and n identical rows give the same tree.Non-parametric trees on truncated data (#188).
SurvivalTreeandRandomSurvivalForestwithkind="non-parametric"refused right-truncated data, and truncated data with left or interval censoring. They now take the Turnbull-score split with Turnbull leaves: a truncated row’s log-rank score is that of its likelihood given its truncation window (the event’s score less the window’s, under the pooled Turnbull estimate fitted with the truncation). On left-truncated right-censored data these are the delayed-entry martingale residuals. Without the window term, a covariate that changed only the truncation was found significant in 39-49% of data sets at the 5% level; with it, 0.5-3%. Every tree kind now accepts the full data model.Breaking: Kaplan-Meier bands hold their level (#390). The equal-precision band covered 0.87-0.89 for a nominal 0.95 (0.83 untransformed), most misses at the first events, where its boundary is unbounded and the estimate rests on a few failures.
band()now forms its bands on the arcsine-square-root scale by default (bound_type="arcsine"; Borgan & Liestøl 1990), and the equal-precision band covers 0.1 <= a <= 0.9 by default, NaN outside;x_range=(t_L, t_U)sets any range. Coverage is now 0.94-0.96 at n = 40-400; Hall-Wellner covers about 0.95. Passbound_type="exp"for the old scale.cb()is unchanged.Breaking: parametric regression Wald bands on the baseline family’s scale (#504). As for the univariate (#477) and degradation models, the
sf/ff/Hfband is now formed on ln H (Weibull, Exponential, Rayleigh, Gumbel), the normal quantile of F (Normal, LogNormal) or the logit (the rest), from the cumulative hazard, andcb_tvcfollows. A model with its coefficient fixed at 0 now gives the univariate band (it was up to 120% away). On ten-point samples the Weibull PH and AFT bands turned back in a tail in 200 of 200 fits and now never do. Small-sample bands change; at n = 2000 they move by at most 1.4% of their width. The three families of model share one helper insurpyval.utils.linalg.Proportional-intensity regressions alias (#502). A repeated column was split (-0.231 as -0.116 / -0.116) and a constant column took part of the baseline (rate 0.0807 became 0.0764), silently. Both now give a NaN coefficient,
model.aliased, one warning, and the fit without the column. A constant is aliased where the baseline has a scale (has_scale: HPP, Duane, Crow-AMSAA, Cox-Lewis).Dual-stress life models alias an undetermined stress effect (#503). With equal (or collinear) stresses, DualPower and DualExponential split one effect between two parameters (-1.174 as -0.568 / -0.605), silently. The later stress’s parameter is now NaN, with one warning naming the column, and the fit is the single-stress one. PowerExponential is unaffected.
Renewal models: intervals at a boundary (#461). A
qdriven to 0, or arhoto 1 or 0, gave NaNparam_cbfor it and foralpha(whose variance was -10.6). The restoration parameter now gets its one-sided profile-likelihood interval (e.g. q in [0, 0.069]); the others get Wald intervals of the model held at the edge, and the printed table says so.MixtureModel: no false “max iterations” warning, and faster (#506). EM on a censored mixture crawled for 1000 iterations and warned at the maximum. After at most 20 EM iterations the fit is polished by direct maximum likelihood with autograd gradients, and a verified maximum is accepted. The issue’s case takes about 1 s (7.4 s before on the same machine), with no warning and a slightly better maximum (738.046941 against 738.047053).
Proportional-odds cumulative hazard is accurate where it is small (#528).
Hf,log_sfandHf_tvcof every PO model computed H0 - ln(phi) + ln(F0 + phi S0), whose terms cancel to about 1e-16 in absolute terms: 20% wrong at H = 4e-16, andlog_sfwas -inf where S0 underflows. They now use ln(1 + F0/(phi S0)). Checked against 50-digit values for all seven baselines; fitted parameters are unchanged.ExpoWeibull likelihood is accurate at extreme shapes (#472). The log-density and hazard added terms of size beta ln(x/alpha) that cancel, an error of about 1e4 per point at beta = 7e19, so the likelihood came out up to 3.7e6 in deviance above the fitted maximum and a likelihood-ratio search met -2e139. They are now computed without the cancellation, and the likelihood-ratio band’s walk starts from the points its direct searches reached.
Every parametric fit searches in the units of its start (#366). Only offset fits did; the others searched a scale as a log below 1 and linearly above it, a different search in every set of units. Fits to data in millionths and in millions now agree to rounding (1e-16 to 1e-11; up to 2e-6 before, and 3% for an ExpoWeibull MSE fit). Fits move only in their last digits. A Beta4 fit on data where its likelihood is unbounded can now warn “No finite maximum” instead of “did not reach a verified maximum”.
Fine-Gray fits in linear time (#517).
FineGray.fit(andCompetingRisksProportionalHazards(model="Fine-Gray")) built a dense events-by-rows matrix of censoring weights and used it in every likelihood, gradient and Hessian evaluation: 1.3 s and 476 MiB at 10,000 rows, and impossible at 100,000 (22 GB for the matrix alone). The risk sets are now cumulative sums over the rows in time order, and the censoring Kaplan-Meier (which the prediction metrics use too) is a suffix sum, bit-identical. A fit takes 0.036 s at 10,000 rows and 0.63 s at 100,000, with memory linear in the rows (34 MiB). Results agree to about 1e-14, except in a few percent of large fits where the optimiser stops one iteration apart (coefficients within its tolerance, up to 3.4e-6 relative); it now stops on “precision loss” less often.CoxPH fits are 4-10 times faster at 100,000 rows (#516). The information matrix was built as one p x p matrix per event time, with truncation terms computed even with no truncation, and the solver built it about 14 times per fit. It is now one product over the rows, the truncation terms are skipped when nothing is truncated, and the coefficients come from Newton-Raphson with step-halving (as in R’s
coxph), about 5 steps, falling back to the previous solver where it fails, as where the likelihood has no finite maximum (that warning is unchanged). At 100,000 rows and 5 covariates, Efron takes 0.28 s with ties (was 1.6 s) and 0.66 s without (was 5.6 s); time-varying fits on 40,000 intervals take 0.24 s (was 0.85 s). The score is now solved to rounding, so results change only in the last digits (at most 4e-13 relative).tolis the largest step, in standard errors, at which the fit stops.Faster Efron Cox fits with tied times (#515). The Efron score was computed on a masked (times x largest tie x covariates) array: one 51-way tie among 30,000 rows took 10.3 s instead of 1.1 s. The sum over tied deaths is now factored and stored per death: 1.4 s. Time-varying Cox fits on 40,000 rows take 0.53 s (was 8.3 s). Untied and Breslow fits are bit-identical; tied Efron fits agree to the last digits.
Frailty and parametric additive hazards fits use the gradient (#515). They began with thousands of Nelder-Mead evaluations. They now run a gradient search first and keep its answer when it is a verified maximum, falling back to the old search otherwise (the #376 and #392 warnings are unchanged). At 10,000 rows: WeibullFrailty 2.2 s to 0.13 s, GammaFrailty 23 s to 1.4 s, WeibullAH 1.0 s to 0.21 s. On some data the old frailty search stopped with the variance at 0, up to 0.1 log-likelihood units short of an interior maximum the new one finds.
Faster log-rank test, Turnbull, Fleming-Harrington and competing-risks fits (#515). Results are bit-identical.
logrankbuilt an at-risk array of rows by event times: 13 s and 2 GB for 30,000 rows in three groups, and out of memory at 100,000; with running totals it takes 0.02 s and 0.06 s. The Fleming-Harrington estimator,Turnbull’s default, summed each step’s tied events in a Python loop on every EM iteration: a Turnbull fit to 1,000 random intervals takes 0.22 s (was 7.9 s) andFlemingHarrington.fiton 100,000 rows 0.05 s (was 0.74 s).CompetingRisks.fitsearched the distinct times for every row: 0.07 s at 100,000 rows (was 1.65 s); the newsurpyval.utils.missing_eventsfinds the censored rows in one pass. Grouping tied rows sorts the data once instead of three times, halving a tied 100,000-row Weibull fit.Faster Kaplan-Meier, Nelson-Aalen, AFT and proportional-odds fits (#498, #499). Greenwood’s and the Nelson-Aalen variance snapped each
d / rto a whole number in a Python loop, most of a large fit; vectorised, a 100,000-row Kaplan-Meier fit takes 50 ms instead of 207 ms (Nelson-Aalen 53 ms, was 228), with bit-identical variances and bounds. AFT and PO fits started with Nelder-Mead, hundreds of derivative-free evaluations; with a differentiable likelihood they now take the gradient ladder the PH models use first, and fall back to the old ladder only when that cannot verify its optimum. At 100,000 rows and 5 covariates a WeibullAFT fit takes 0.48 s (was 2.0 s) and WeibullPO 0.64 s (3.9 s); at 1,000 rows AFT and PO fits are 4-6 times faster. They reach the same maximum to 1e-6 in the log-likelihood, or a slightly higher one. The time-varying AFT likelihood is not differentiable and keeps Nelder-Mead.Likelihood-ratio bounds at the edge, and along valleys (#421). Profiles are now searched on the log or logit scale of each parameter, each point starting from the ones already solved. The old search on raw parameters, restarted from the fit every time, overstated the profiles: ExpoWeibull 42.4 for 0.29 at mu = 1e-4, NegativeBinomial 3.05 for 2.35 at p = 0.999999. Where a profile levels off below the critical value, the bound is the edge of the parameter’s space. A NegativeBinomial
rtends to a shifted Poisson with deviance 2.345, so its 95% upper bound isinf, not 1.7e16, and its 80% bound (36.2) is finite. ExpoWeibullbetais now [0.053, inf] at 95% and [0.486, inf] at 80%; the old [0.171, 367.5] and [0.486, 372.0] were not nested. A band is where the function’s own profile reaches the critical value, andsf,ffandHfshare one band. NegativeBinomialhf(2)at 95% is now 0.305 (was 0.207), and the ExpoWeibull bands that werenanare found. The NegativeBinomial sweep no longer leaks over 14,000 raw numpy warnings. Likelihood-ratio bounds are slower for the simple families (a Weibullcbwent from about 0.06 to 0.2-0.5 s).Continuously varying covariates: CovariatePath (#172, phase 1).
sf_tvcandHf_tvcaccepted only step schedules, so a ramp-stress profile or a thermal cycle had to be cut into steps. That was slow (150 ms for 1000 steps), and only accurate to the square of the step width, with no error reported.CovariatePath.from_points(times, values, period=None)(straight lines; a repeated time is a jump) andCovariatePath.from_callable(func, p=1, breakpoints=None, period=None)now describe the path. PH, AH, PO and AFT integrate the hazard (AFT the accelerated age, Nelson’s cumulative exposure) by adaptive Gauss-Kronrod quadrature to 1e-10 relative error on H, with oneRuntimeWarningif that is missed. Cox sums its baseline jumps along the path exactly. The error against closed forms is 5e-16 to 9e-14, in about 2 ms for 200 times.given=integrates from the conditioning age. A step schedule is still summed exactly, and a flat path gives the same result to rounding. A path evaluates a fitted model only: fitting still uses steps (fit_tvc).Out-of-bag log-likelihood and permutation importance for the random survival forest (#186). The forest could only be scored by concordance, which needs right-censored data.
RandomSurvivalForest.oob_log_likelihood()now scores every training row by its full likelihood (density, S, F or interval probability, over the truncation probability) under the trees that did not see it, for every censoring type and truncation.feature_importances(n_repeats=5, random_state=None)reports how much that score drops when a feature is shuffled among each tree’s out-of-bag rows. For this score a non-parametric leaf is read as a continuous distribution (linear between its drops, exponential after the last). On the docs example the score rises from -2.644 without splits to -2.484 with them.Non-parametric survival trees on left- and interval-censored data (#188, stage 1).
kind="non-parametric"raised on such data. It now splits on the log-rank scores of the node’s pooled Turnbull estimate (the standardised left-child sum, with its permutation variance), with Turnbull leaves. On right-censored data the scores are exactly the classic log-rank scores, so the split chooses as the log-rank split does. Truncation combined with left or interval censoring, and right truncation, still raise (stage 2).Added: random_state for SurvivalTree and RandomSurvivalForest (#471). The bootstrap samples and each split’s candidate features were drawn from numpy’s global stream, so a forest could only be reproduced by seeding numpy globally, and fitting one disturbed the global stream.
random_state=Nonestill draws from the global stream exactly as before, so forests undernp.random.seedare identical. An int orGeneratorgives the forest its own stream, with a child stream per tree, and leaves the global one alone.Added: conditional-inference trees, selection=”ctree” (#188). Greedy search prefers covariates with many values and always splits: in 60 simulated data sets it chose a noise covariate 47% of the time over a two-valued covariate with a real effect. With
selection="ctree"each node chooses its covariate by the Bonferroni-adjusted p-value of its maximally selected score statistic, and splits only if that p is belowalpha_split(0.05). The scores are log-rank for"non-parametric", and the working model’s score contributions for"exponential"and"weibull". The p-value is exact for the statistic’s asymptotic chain. Noise is then chosen 10% of the time, and on null data 96% of trees stay a single leaf, where greedy search always splits. It works for every kind and every censoring type; the default is unchanged.Beta4: a fit with no maximum says so, and MPS is recommended (#385). The four-parameter Beta’s likelihood is unbounded: a shape below 1 makes the density infinite at a support end. So a maximum-likelihood fit could run an end onto the smallest or largest observation and stop wherever its search gave up. The answer depended on the data’s units: shapes 1.00, 1.19 on a fixture, and 0.18, 0.18 on the same data times 7.3. Such a fit now warns “No finite maximum” and recommends
how="MPS". Maximum product of spacings scores an end gap of zero as minus infinity, so its estimates are finite and the same in any units. MLE stays the default; the class docstring explains when to prefer MPS.Changed: Bernoulli’s survival function is P(X > x) (#344). It was
P(X >= x), sosfwas [1, p] at the outcomes 0 and 1 andffwasP(X < x), which never reaches 1, soqfcould not invert it.sfis now [p, 0] andff[1 - p, 1], as forBinomialwithn = 1,scipy.stats.bernoulliand every other discrete distribution in the package.Hfis [-log p, inf];df,hf,qf,mean,randomand the fittedpare unchanged. Code that read the probability of the1outcome (a one-shot device working on demand) assf(1)wantssf(0), orpitself.The Uniform’s MLE refuses censored data again (#460). 0.21.0 fitted right- and left-censored data by maximum likelihood. The estimates were right, but they sit on a wall of the likelihood (the smallest or largest observation), where its curvature says nothing about their uncertainty. The covariance it reported was not positive definite, so Wald bounds were NaN or silently several times too wide: an
sfbound of [0.25, 0.98] where the likelihood-ratio bound is [0.72, 0.85].Uniform.fitnow raisesValueErroron any censored value, as it always did for interval-censored ones, and names the methods that take censored data (how="MPS","MPP","MSE"). Exactly observed data, truncated or not, fit as before, still with no covariance.Fits with no finite maximum warn instead of returning silently (#392). Some data leave the likelihood with no maximum: a covariate level with no events, perfectly dependent pairs, a mixture component on a point mass, noise-free degradation readings. These fits returned wherever their optimiser stopped, silently:
a WeibullPH coefficient of -14.7 (+32 to +36 for the PO models, about -32 for most frailty models);
Fine-Gray -12.9, with BFGS reporting success;
copula dependence of Clayton theta 3.2e6, Frank 1.2e7, Gumbel 105.5 (log-likelihood inf) and Gaussian rho at its 0.9999 cap;
a mixture component with beta 9100;
GammaProcess alpha at the end of its search range (1e6), and DestructiveDegradation sigma 9.9e-16;
BetaGeometric in its Geometric limit (a, b about 1e5, 3.5e5).
Each now gives one
UserWarningstarting “No finite maximum” at the caller’s line. It names the parameter that runs away and what to do instead (for example “use Geometric”), and the fit still returns the model it reached. Fine-Gray andCompetingRisksProportionalHazards(Fine-Gray) use CoxPH’s “Monotone partial likelihood” warning.The regression test is Newton’s. Along each coefficient’s profile, Kantorovich’s
h = |f'''| |f'| / f''^2is 1 on the way to a supremum, however far the optimiser went. At the maximum of an ordinary fit it is at most 2e-4, over the test registry and 1360 calibration refits. It costs one Hessian at the fitted values, and a coefficient’s profile is read only when its Newton step there exceeds 1/709.8 of its value (the log of the largest double): every coefficient running off to infinity exceeds it and a converged one does not. That Hessian is kept:covariance(),standard_errors(),cb()andparam_cb()of the PH, AFT, PO, AH, frailty and Fine-Gray fits invert it instead of computing a numerical Hessian on every call, so a PH fit followed by its standard errors and a band is 8-42% faster than before, and the standard errors move by at most 5e-4 relative (rounding in the numerical Hessian). The numerical Hessian is still used for reloaded models, accelerated-life and AFTfit_tvcfits, fits with no finite maximum, and Hessians that are not positive definite. The additive-hazards models, whose likelihood rises without bound on such data, now say so instead of reporting a positivity boundary (all except GammaAH). The Gumbel copula no longer leaks about 230 raw numpy overflow warnings. Univariate MLE now also refuses a point mass at the edge of a truncation window (Weibull beta 455.6, Normal sigma 0.037, silently).Parametric additive hazards warn where their hazard is negative (#376).
h_0(x) + beta'Zhas nothing keeping it positive away from the observed failures, so for a protective covariate row the cumulative hazard falls:sfexceeded 1 (1.03 for aWeibullAH) andffanddfwent negative, silently. The values are still the model’s, butsf,ff,df,hf,Hf,cb,sf_tvcandHf_tvcnow give oneRuntimeWarningper call naming how many queried points have a negative hazard and how farsfexceeds 1. (The semi-parametricAdditiveHazardsalready predicts with the running maximum of its cumulative hazard, since #462.)Additive hazards bounds no longer leak a numpy overflow warning (#465). Where the fitted cumulative hazard is negative, the logit-scale
sfbound computed1 / (1 + exp(-t))withthugely negative, and numpy warned “overflow encountered in exp” on the way to the right answer, 0. It now usesscipy.special.expit; the bounds are unchanged.Cox models no longer break on a covariate far from zero (#459).
CoxPHfitted on the raw covariates, so a column such as a year or a date overflowedexp(beta'Z): on 200 rows, adding 2000 to a N(0, 1) covariate moved beta from 0.860 to 0.768, made the survival and p-values NaN, gave a false “monotone partial likelihood” warning and leaked numpy overflow warnings. The fit now always runs on the covariates centred on their (n-weighted) means, as R’scoxph, lifelines and scikit-survival do; beta and every prediction are unchanged by any shift of a column, for plain, stratified and time-varying fits, residuals,check_ph, robust errors and the cause-specific Cox model. The reported baseline (h0,H0) stays atZ = 0(R’sbasehaz(fit, centered = FALSE)), as before, andphi(Z)isexp(beta'Z); predictions combine the two on the log scale, so a tiny baseline and a huge multiplier lose nothing. Where the baseline atZ = 0cannot be represented (it over- or underflows), the fit raisesValueErrorand says to pass the newcenter=True(onfit,fit_from_dfand thefit_tvcvariants), which keeps the baseline at the means,model.center(R’sbasehaz(fit)), withphi(Z) = exp(beta'(Z - center)), and saves"center"(schema 2). A default fit’s file has the layout of before (schema 1), and files saved by 0.21.0 and earlier load as fitted.Parametric regressions and Fine-Gray no longer break on a covariate far from zero (#463).
exp(beta'Z)overflowed in these fits too: adding 2000 to a N(0, 1) covariate turned aWeibullPHcoefficient of 0.707 into 0.0247 with a scale of 2.4e18, silently, andFineGray.fitandCompetingRisksProportionalHazards(model="Fine-Gray")raisedLinAlgError: SVD did not converge.FineGrayand the competing- risks model now centre as Cox does, and takecenter=with the same meaning. The parametric families (PH,AFT,PO,AHand theirfit_from_df/fit_tvcvariants) takecenter=False: by default the baseline is reported atZ = 0, as before. Where the family and link have an exact map between the covariate means andZ = 0(a Weibull, Rayleigh, Exponential or Gumbel PH; every AFT baseline SurPyval has; a LogLogistic or Logistic PO) the fit runs on centred covariates and maps back, so beta, the predictions and the bounds are the same whatever the covariates’ origin, and it raisesValueError, pointing tocenter=True, when the baseline at 0 over- or underflows. For the other pairs the default fit is atZ = 0, unchanged and unchecked, so on covariates far from zero it can still fail or stop at a poor answer (a LogNormal PH coefficient of 0.006 against 0.64 at a shift of 300).center=Truefits any family with its baseline at the means (model.center, shown in the summary and saved, schema 2), where a shift of a column changes nothing. For additive hazards,center=Trueis a different model,h0 + beta'(Z - center). On ordinary data the reported parameters are those of before to optimiser tolerance (log-likelihoods within 2e-8 on 48 of 49 test fits; one Logistic PO fit on 34 tires stops 1.3e-5 nats short of the old point), and the fits are no slower.Fits no longer stop far from the maximum in silence (#427, #428, #429). The maximum-likelihood ladder took the first optimiser that reported success, and from a poor start BFGS, TNC and Newton-CG report it where the likelihood first looks flat: a Weibull started at alpha 1e7 returned beta 0.099 with a log-likelihood 40 below the maximum. A rung now counts only when its answer has a zero gradient and a positive-definite Hessian; a fit given
initis also started from the default start and keeps the best likelihood; an answer still not a verified maximum warns. Accelerated-life fits (InversePower was 14.7 below the maximum), additive-hazards fits, and NHPP, proportional intensity and renewal fits (ARI 227 below; Duane NaN) check their answer and use the default start the same way. Fits the first rung solves cost the same as before.Data with no maximum are refused (#392). A parametric MLE whose data one failure time explains completely returned a spike: an exact 0.5 with a left-censored 1 gave a Weibull beta of 395.7 and a Normal sigma of 5e-324; the intervals (1, 3] and (2, 4] gave beta 57.9. These raise
ValueErrornow, as tied values already did.The truncated likelihood is exact in the upper tail (#412, #393). A truncation or interval window where F rounds to 1 gave +inf or NaN (a LogNormal left-truncated at 1 with mu = -5 gave +inf, not -23.73), which made truncated fits depend on the data’s units (Normal 8.536, 2.461 on the data but 48.87, 23.04 on the data x 7.3). Such windows are computed from the log survival function now; truncated fits are unit-free, reach the maximum (8.745, 2.459), and run about 40 times faster.
Continuous distributions are accurate in their tails (#410, #442, #443, #444, #447). Functions no longer take the log or complement of a probability already rounded to 1, 0 or inf.
Logistic.sfwas NaN more than 709 scales below the location;Hfandlog_sfwere -0.0 deep in the left tail (true 1e-300) andlog_ff0 deep in the right (true -1e-30); log forms were +-inf where the probability underflows (Weibulllog_ffat 1e-400 is -921); hazards were NaN where density and survival both underflow (Normal.hfuses the asymptotic series above z = 100); LogNormal, Weibull, LogLogistic and Gamma were NaN at x = 0; GammadfraisedOverflowErrorat shape 1000;Rayleigh.qf(1e-30)was 0 andHypoexponential.qf(1e-30, 1, 2)4.3e-61 against 1.0e-15. Hypoexponential uses a power series near the origin and gainslog_sf,log_ffandlog_df. Every function is within 1e-8 of 50-digit mpmath values on the tail grid, and fitted parameters are unchanged.Discrete and Beta-family distributions are accurate in their tails (#442-#447, #449, #458). Geometric used
log(1 - p)(8 digits lost at p = 1e-9); BetaGeometric tookln B(a, b + k)as a difference ofgammalnvalues of size 2.6e13, losing up to 22% offfat k = 1e12, and itsqf(1)was finite; Beta4 raisedOverflowErrorat extreme shapes. The references generated for Beta, Binomial, DiscreteWeibull and NegativeBinomial (#448) exposed 102 more failing groups – a Betasfof 0 against 1e-30, a NaN density at alpha = beta = 1000, a Binomial hazard of 0 against 0.99999, a NegativeBinomialqf2.5% short at p = 1e-6 – and Poisson and Beta4 lost their tails the same way. The incomplete beta gains log-scale tails (betainclnand the newbetainccln, a continued fraction where a tail is below 1e-3), with accurateln Band gamma ratios and exact discrete quantiles. Every one of the 1,404 tail groups now agrees with 50-digit mpmath values; fitted parameters are unchanged.Discretize: ``qf(ff(k))`` is k (#383). It returned k + 1 where
ceilrounded k + 1e-15 up.Distribution functions accept lists and tuples (#424).
Gamma.sf([5, 10], 8, 3)returned six values (Python list repetition) andWeibull.sf([5, 10], 8, 3)raisedTypeError; every distribution’s functions take a list or tuple as an array now.A missing query gives NaN everywhere (#382). The constant-hazard models (Exponential and its regressions, Geometric, the HPP
iif), Uniform, Beta4, Binomial, Bernoulli, FixedEventProbability, ExactEventTime, NeverOccurs, InstantlyOccurs, BetaGeometric and RoystonParmarqf, the Gaussian copula’scdfand the non-parametric and renewalmcfreturned a number, the last value or raised at a NaN; they give NaN there now.Fixed: Turnbull keeps its last piece when every row is right truncated (#391). The ladder assumed the last bound was +inf and dropped the piece ending at the largest
tr: one failure at 1 observable up to 1 gave sf(1) = 1 (0 withturnbull_estimator='Kaplan-Meier'), and a left-censored row at its truncation time raisedIndexError.Changed: a non-parametric ``df`` is the step probability (#408). It is the drop in
sfover the stephfdifferences, nothf * exp(-Hf), which was inf x 0 = NaN (with a raw warning) where a Kaplan-Meier reaches zero:KaplanMeier.fit([1, 2, 3]).df([2.5, 3.5, 4.5])was[inf, nan, nan]and is[1/3, 1/3, 1/3]. Values move slightly elsewhere (the Nelson-Aalen example: 0.2047 to 0.1811).Fixed: cubic confidence bounds close onto the estimate (#417). A non-linear
interp’s bounds are the interpolated estimate plus the interpolated distance over the times with a variance, so they close ontosfand always contain it (atalpha_ci-> 1, 0.1873 against sf 0.1948 before); linear bounds are unchanged.Fixed: ``band`` at a large ``alpha_ci`` (#420). The critical-value search climbed from far below the root:
alpha_ci = 0.9took 13 s and1 - 1e-6tried to allocate 158 TiB or hung. Now 0.1-6 s; critical values at the usual levels are unchanged.Changed: a step with no one at risk keeps the estimate (#425).
kaplan_meierandnelson_aalentook a step with no one at risk and no events to zero whilefleming_harringtonkept its value. All three, and the variances, now carry the estimate there, as R’ssurvfitdoes; a step with events and no one at risk raisesValueError.Semi-parametric fitters refuse an infinite event time (#394).
CoxPH(stratified too),AdditiveHazards,BuckleyJames,CompetingRisksProportionalHazards,FineGrayandCompetingRiskstook an observedx = infas an event: Cox returned beta 19.4 on two rows and Lin-Ying died in aLinAlgError. They now raise the univariate fitters’ValueError; an infinite censoring time is still accepted.Buckley-James takes one covariate row per time (#426).
sf,ffandHftook one vector only and raised numpy’s bare matmul error for paired rows; they now pair row i withx[i]like every other regression model, and refuse a mismatchedZwith a message.Frailty models drop a row with a missing group (#388). A
NaNlabel was a group of its own (9 groups instead of 8) andNoneraisedTypeError; such rows are dropped with one “Dropped k of n rows” warning, andgroup=nanpredictsnan.Coefficients the data cannot separate are NaN, as in R (#476, #409). A covariate column that adds nothing to the others – a constant column where the baseline already has a scale, a duplicated column, dummy columns that add up to another column – has no estimate of its own: shifting weight between it and the columns it repeats fits the data equally well. The fit used to return whichever split its search stopped at. A constant column moved WeibullAFT’s
alphafrom 54.06 to 51.02, a repeated column split a Cox effect into 0.148 and -0.264, and a constant column on separated data got 3.1e14 with an all-NaN baseline. Now the fit leaves such columns out, as R’scoxphandsurvregdo: their coefficients arenan, withnanstandard error and p-value,model.aliasedlists them, and one warning names them. The other coefficients and the predictions are those of the fit without the column. This coversCoxPH(with strata and time-varying covariates), the parametric proportional hazards, AFT, proportional odds and additive hazards models and theirfit_tvc(the AFT one split a repeated column’s 0.338 into 1.685 and -1.346), Fine-Gray, the competing-risks Cox model, the frailty models, and the Lin-YingAdditiveHazardsandBuckleyJamesmodels. Changed: those two raisedValueErrorfor a constant column, a single observation or collinear columns; they now fit the other columns and warn. Fits without such a column are unchanged. The proportional-intensity recurrent regressions (#502) and the dual-stress life models (#503) do not alias yet.conformance/test_aliasing.pychecks every registered model (principle 12).Regression models print a coefficient table: summary() (#484). The Cox and parametric regression models’
summary()returns aDataFramein the layout of R’ssummary(coxph)and lifelines:coef,exp(coef),se(coef), Wald intervals for both,zandp, one row per covariate, named from theDataFramecolumns forfit_from_df. On the Rossi data it matches R (fin: coef -0.37942, se 0.19138, p 0.04742). The parametric models add the baseline’s parameters, with the intervals ofparam_cb, so the Weibull shapebetais no longer printed among the coefficientsbeta_0,beta_1, …. The models’reprprints the table.Changed: FrailtyModel.summary() returns the table (#484). It returned the text its
reprprints; it now returns the sameDataFrameas the parametric models, with afrailtyrow fortheta.print(model)gives the text.A cause label 0 with no censoring warns (#486). lifelines, scikit-survival and R’s
cmprskcode a censored row as cause 0; SurPyval codes it as a missing cause (NoneorNaN), or withc. Data coded the other way fitted without complaint as a model with an extra cause “0” and no censoring. Numeric cause labels that include 0, with none missing and noc, now warn and give the one-line conversion; passingcsays 0 really is a cause and silences it.Covariate rows and times that cannot be paired raise (#488). A regression model pairs row
iofZwith timex[i](one row for every time, or one time for every row). Any other count was a raw numpy broadcast error; it is now aValueErrorsaying so. For a survival curve per subject, the Cox and parametric regression functions takegrid=True, which gives every row at every time, shape(len(Z),) + x.shape, as the survival tree and forest do.Changed: load_rossi_static()’s arrest means an arrest (#479). It stored the censoring flag (1 = not arrested) as a float under the name
arrest, soc = 1 - df["arrest"], the natural call coming from R or lifelines, fitted the complement without complaint (a Cox coefficient of -0.001 for age instead of -0.057).arrestis now an integer, 1 for an arrest, as in R’scarData::Rossiand lifelines. Code that passedc=df["arrest"]must passc=1 - df["arrest"]. The bundled rossi, heart and lung data also lose their saved row-index columns (Unnamed: 0).Changed: trend tests name a trend only when it is significant (#481).
laplaceandmil_hdbk_189cprinted “Suggested trend: increasing” at p = 0.25.trendis now the conclusion at a keyword-onlyalpha_ci=0.05("none"unless p <alpha_ci) and the newdirectionis the sign of the statistic. The models’trend_test()takealpha_ci, and the tests takec=by keyword only.MixtureModel.fit returns the model (#482). It returned
None.MixtureModel.fit(x, dist=Weibull, m=2)also builds and fits in one call. An unfitted model’sreprnames it, and a truncated fit reports “Fitted by: MLE”, which it is, not “EM”.params and param_names on FrailtyModel and RenewalModel (#483). Frailty: the baseline, the coefficients, then
theta. Renewal:qorrho, then the distribution’s parameters, the order ofstandard_errors.AcceleratedLife reports its life parameter as the life model’s (#489). It printed “alpha: 1.0”, a placeholder, as if fitted; it now prints
alpha: L(Z) of the Power life model,param_cbrefuses that parameter, and the model haslife_parameter. Every parametric regression model hasparam_names. Data at one stress level raise “needs at least two distinct stress levels”, and too few levels withinitwarn that the life-stress relationship cannot be identified.The repair models say which kind of dist they take (#495). A lifetime distribution passed to
ARI(anAttributeError) and an intensity model passed toARA,GeneralizedRenewalorGeneralizedOneRenewal(“more than one right censored time”) raise aValueErrornaming the right fitter.API papercuts (#485).
to_json()without a path returns the JSON text, andfrom_jsonreads it. A model with no data (fromfrom_params, or loaded) plots its CDF.howis case-insensitive.x,candngiven as (n, 1) columns are read as one value per row. Changed:qfoutside [0, 1] isnan(it wasinfor 0), a continuous distribution’sqf(0)is the start of its support (a Normal’s-inf, not 0), and a bounded one’sqf(1)its end (Uniform(2, 5): 5, notinf). Kaplan-Meier, Nelson-Aalen and Fleming-Harrington given left- or interval-censored or right-truncated data point to Turnbull.logrankwarns when most groups have one member.fit_bestpasses over candidates whose support excludes the data quietly and reports other failures once. The MPP heuristic error lists the valid names, andsp.CrowAMSAAand the like say which subpackage to import from.CoxPH.fit_tvc_from_dfandfit_tvc_timeline_from_dftakeformula=, and non-numericZ_colssuggest it.CompetingRisks.plot()is new. Signatures print readably:Weibull.fit’s is 593 characters, was 3,218.Changed / deprecated: one name for parameter names, parameter_names (principle 21). A model’s parameter names were spelt three ways:
param_names, the regression models’parameter_names()method and the recurrent models’parameter_namesproperty (which a model built withfrom_paramsrefused). Every distribution and every model withparamsnow hasparameter_names, a list namingparamsentry by entry (Weibull.fit(x).parameter_namesis['alpha', 'beta'],WeibullPH’s['alpha', 'beta', 'beta_0']), including models that had none (CoxPH,AdditiveHazards,BuckleyJames,RoystonParmar,MixtureModel,CopulaModel). AProportionalIntensityModelnamesparamsthencoeffs, the order of itscovariance. Until v0.23 the old spellings work with aDeprecationWarning: theparam_namesattribute, callingparameter_names(), theparam_names=keyword ofCustomDistribution, and aparam_namesclass attribute on your ownPathModel,CopulaorCountingProcesssubclass.paramsis unchanged, and saved files keep the key"param_names", so they move both ways between 0.21 and 0.22.Fits record whether they reached a maximum. A parametric model has
maximum:"verified"(zero gradient and a positive-definite Hessian, or an exact estimator),"unverified","no finite maximum","not applicable"(not a maximum-likelihood fit, orfrom_params) or"unknown"(loaded from an older save). It matches the fit’s warnings and is saved byto_dict.fit_bestsets candidates aside by it rather than by the text of their warnings (#492).Changed: two-parameter fits no longer hide an unverified maximum. A maximum-likelihood fit that did not reach a verified maximum warned only for families with more than two parameters: the exemption meant for the Uniform, whose support ends are parameters, matched every two-parameter family (Weibull, Gamma, LogNormal, …). Such fits now warn.
formula= for the parametric time-varying fits.
fit_tvc_from_df(PH, AH, PO, AFT) andfit_tvc_timeline_from_df(PH, AH, PO) take a formula instead ofZ_cols, asCoxPH’s do (#485): categorical columns are coded, the model predicts from aDataFramewith the same coding, and an aliased column is named.Z_colsnow defaults toNone.The recurrent-event, competing-risks and degradation models are importable from surpyval.
sp.CrowAMSAA,sp.ARA,sp.FineGray,sp.CompetingRisks,sp.DegradationAnalysisand the other model classes are now at the top level, as the regression models already were, and stay in their packages too. Helper functions and result types (laplace,mil_hdbk_189c,TrendTestResult) and the generically named copulas (Gaussian,Frank, …) stay in their packages; asking for one at the top level says where it is.The bundled Claude Code skill matches 0.22, and its code is tested (#491). Every code block in it runs as a test.
Offset fits with no maximum warn (#487). With a shape below 1 the density is infinite at the offset, so the likelihood grows without bound as
gammaruns onto the first failure. Such a fit returned a degenerate model silently or behind “MLE Failed”: the 3-parameter Weibull on [55, …, 140] reached a shape of 0.09. Every offset family now warns “No finite maximum” and recommendshow="MPS", and returns the point its search reached. The Exponential’s genuine maximum at the first failure is unchanged.Changed: durations and dates are refused (#480).
timedelta64anddatetime64input was fitted in its storage ticks (a scale of 5.0e5 for durations of days held in seconds, 6.9e14 in nanoseconds), and predictions raised numpy’sTypeError. SurPyval has no time unit, so such values now raise aValueErrorwherever a time is accepted, with the conversion to use (x / pd.Timedelta(days=1)).Changed: suspensions are not drawn as points (#478). A probability plot drew each suspension at the
Fof the failure before it, where it looked like one more failure.plot()now draws the failures only, as in Abernethy’s New Weibull Handbook and Weibull++, andplot(show_censored=True)(also onMixtureModel.plot) marks the suspension times with ticks on the time axis, which still spans every time.get_plot_data()keeps its meaning:x_andFhold every row as before, and a new booleanfailedmask selects the failures that are drawn;x_censoredholds the suspension times. The non-parametricget_plot_datareturnsfailedtoo.weibayes (#493).
surpyval.weibayes(x, c, n, beta)gives the Weibayes lower confidence bound on a Weibull’s scale of known shape from few or no failures (Nelson 1985; Abernethy), as a Weibull model: ten units run 500 hours with no failure and a shape of 2 give a 95% lower bound on the scale of 913.5 hours.fitstill refuses data with no failures, and now points to it.Changed: fit_best ranks regular maxima only (#492). AIC and BIC assume a regular maximum. The Uniform and Beta4, whose support ends are parameters, are no longer default candidates (
include=still tries them), and a fit with no finite maximum or an unverified one is ranked only when no regular candidate fits, with one warning naming it. The Beta4 no longer “wins” on [1, …, 7], nor the Uniform on 50 Weibull draws.Changed: Wald bands are monotone (#477). Wald bounds on
sf,ffandHfwere formed on the logit ofsffor every family, and on small samples turned back: the issue’s lower bound onFfell from 0.39 to 0.00004 as time went on. They are now formed on each family’s probability-plot scale (log(-log S)for the Weibull, the normal quantile for the Normal and LogNormal), where they are monotone whenever the shape’s own interval excludes 0; that bound is now 0.83. Large samples are unchanged to 2e-4. The degradation models’ two-stage band is formed on the same scale, so it still contains the life model’s own.plotandget_plot_datatakemethod=to draw the likelihood-ratio band.quantile_cb and mean_cb (#494). Parametric models give confidence bounds on a quantile (a B-life) and on the mean, by Wald (matching R’s
survregto 1e-7) or likelihood ratio (method="lr", better on small samples), withbound=andalpha_cias oncb.Competing-risks Cox incidences add up to 1 - sf (#384). They were built on the product-limit survival while
sfisexp(-H)(summing to 1.0 againstff= 0.975 at t = 30 on the conformance fixture). Each step now uses the matrix-exponential transition probabilities of R’s multi-statecoxph, which it matches to 7 digits.Gray’s test matches cmprsk (#380). The variance was SurPyval’s own linearisation: 6.741 against
cuminc’s 7.015 on tied data. The score, variance and rho-weight incidence now follow cmprsk’scrstroutine and agree withcumincto about 1e-14, with ties and any rho.Lin-Ying survival stays in [0, 1] (#376).
AdditiveHazards.sfrose above 1 (1.21 inside the data, 57.8 at a row with a negative hazard) because the estimate falls between event times. It now predicts with the running maximum of its cumulative-hazard estimate from time 0; the fittedH0is unchanged. The parametric additive hazards models are not yet changed.Wald bounds that do not exist say so (#411).
param_cbandcbwere a silent[nan, nan]where a variance was negative (GeneralizedRenewal’sq= 2.7e-16 had variance -0.031), and ARI’srho = 1.0raisedZeroDivisionError. They give NaN with one warning naming the parameter and the reason now.Rate bounds at zero, and discrete hazards (#413, #414).
cb(on='hf'/'df')was[nan, nan]where the rate is 0; it is[0, 0]. The discrete hazard bound was centred ondf(k)/sf(k)(Poissonhf(10)= 0.719 had bounds [1.73, 3.78]); it uses the model’sdf(k)/sf(k-1)on the logit scale now ([0.634, 0.791]), andhfof a discrete limited-failure or zero-inflated model isdf(k)/sf(k-1)too (it gave 0.271 for 0.213).Royston-Parmar one-sided bounds (#415).
bound='lower'onfforHfreturned the upper end, andbound='both'was taken as'upper'; both are right now, and'both'raises.Regression cumulative-hazard bounds have no ceiling (#418).
sfwas clipped at 1e-15, soHfbounds stopped at 34.54 (GumbelPHHf= 110.6 had [34.54, 34.54]); they are formed fromHfnow ([4.4e-18, 261]).Likelihood-ratio bounds (#421, partly). A search that stopped on the wrong side of the estimate is retried (Rayleigh’s
dflower bound was 0.0502 against an estimate of 0.0359), one-parameter bands are the exact extreme over the profile interval, and the Uniform’s search respects the data’s extremes (30 s to 1.3 s).Count-terminated simulation of a falling intensity is refused with the reason (#386). A CoxLewis with beta < 0 expects only
cif(inf)events ever (6.04 on the conformance fixture), so a sequence can stop short of the count: one seed failed with “Event times ‘x’ must be finite” and others returned a sample silently.count_terminated_simulation(and_data) raise aValueErrorfor every seed now, givingcif(inf), the chance of falling short and the time-terminated alternative.CoxLewis least-squares fit (#419). The search started at alpha = beta = 1, where
cif(60)is about 1e26, and BFGS stopped far off: on a sample of the fitted modelcif(55)was 113.4 against 4.40 by MLE. It starts from the constant rate through the MCF now, and a BFGS stop that did not converge is finished by Nelder-Mead: 4.32, matching a direct minimisation. Fits in hours rather than days agree too (both gave acifof inf).One rule for unknown option values (#416).
mcf_cb(bound='both')raisedUnboundLocalErrorand an unknowninterpthere returned the event-time bounds; an unknowninterpon the non-parametric estimators raised scipy’sNotImplementedError; the cause-specific Cox model accepted anyinterpand ignored it. All raiseValueError: '<arg>' must be one of (...); got ...now; the cause-specific Cox model takesinterp="step"only, andDestructiveDegradation.cbacceptson='R'/'F'like every othercb.
v0.21.0 (28 September 2026)
Upgrading to 0.21. Most code runs unchanged, but these changes can alter results or break code without a warning:
Shapes. A scalar query now returns a numpy scalar, not a
(1,)array, and a two-sided bound ends in a[lower, upper]axis: code that indexed a scalar result (km.sf(5)[0]) raisesIndexErrorand uses the result directly instead.Cox ties.
CoxPH.fitdefaults to Efron’s tie handling (it was Breslow’s), so fits to tied data change; passtie_method="breslow"for the old model.Draws.
random()of a non-parametric estimate draws from the estimate, withinffor the probability beyond the last time; of a limited-failure or zero-inflated model it returns lifetimes (inffor a unit that never fails), and the old survival-data draw israndom_data(). Seeded draws of those models give different numbers.Limited-failure summaries.
mean(),moment()andvar()of a model withp < 1areinf;defective=Truegives the old values.Criteria and extrapolation.
ParametricCompetingRisks.bic()is the joint criterion (larger than before), and the additive hazards model holds its estimate after the last observed time instead of extending it.Missing values. A missing time, covariate or probability gives NaN at prediction, and fitting drops rows with a missing covariate with one warning (or refuses them where a row is part of one unit).
Recurrent interval levels are
alpha_ci=0.05, by keyword only: an old positionalconfidencelevel inmcf_cbor the recurrentplotmethods raisesTypeError.
Renamed arguments keep working, with a DeprecationWarning naming the new
name, until v0.22.0, which removes them together with the
surpyval.experimental alias (use surpyval.beta.ml) and band’s
unused n_sims and random_state. To find the calls to update, run your
code or tests with python -W error::DeprecationWarning. Saved models from
earlier versions still load; a model saved by 0.21 with a feature older
versions cannot read (a set_support support, a truncated band’s sample
size, some formula terms) is stamped schema 2 and refused by them with a
request to upgrade.
Changed: one name per option (#422, principle 21). The same option had different names in different parts of the package; each now has one, and the old name keeps working until v0.22.0 with a
DeprecationWarningnaming the new one (surpyval.utils.deprecationdoes this for every rename):Interval level:
alpha_ci=0.05everywhere (the recurrentmcf_cband plots tookconfidence=0.95). Seeds:random_stateeverywhere (some recurrent, Buckley-James and degradation methods tookseed). Bootstrap size:n_boot(BinNonParametric.bootstrap_cb).Times are
xand a quantile’s probabilitypeverywhere (Parametric.cb, Royston-Parmar and the degradation models tookt;qftookuorqin a few models).Regression and competing risks:
CoxPH.fit,fit_from_dfandfit_tvc*taketie_method(wasmethod;CoxPH.baselineand the competing-risks Cox already did); thefit_tvc*_from_dfmethods takei_col(wasid_col: the data argument isi) andfit_tvc_timeline_from_dftakesx_col(wastime_col);BuckleyJamesModel.bootstrap_citakesrandom_state(wasseed);CompetingRisksProportionalHazards.fit/fit_from_dftakemodel="Cox"or"Fine-Gray"(washow, the estimation method everywhere else; the fitted.howis.model); andFineGray.fitandgray_testtakeevent(wascause). The non-parametricCompetingRisks.fit/fit_from_dfchoose their survival estimator withhow(wasmethod; the fitted.methodis.how).Recurrent events:
NonParametricCounting.mcf_cband.plot,CauseSpecificMCF.mcf_cband.plot, and the parametric and proportional-intensityplottakealpha_ci=0.05(wasconfidence=0.95;confidence=0.9is read asalpha_ci=0.1), keyword-only, so an old level passed by position raises rather than silently meaning its complement. The simulations, the simulatedmcfandplotand everycramer_von_misestakerandom_state(wasseed), andCauseSpecificMCF/CauseSpecificNHPPtakeevent(wascause). Plot labels read “95%”, not “95.0%”.Degradation: times are
x, nott, in the Wiener and Gamma process models (sf,ff,df,hf,Hf) andDestructiveDegradationModel(sf,ff,df,Hf,cb,median_degradation,degradation_quantile, whose probability isp, notq).Zcomes straight after the query:DegradationModel.cb(x, Z, on, ...)and the process models’random(size, Z, random_state), asDegradationModel.random;DegradationModel.induced_life(n_samples, *, Z, random_state),DegradationModel.predict_rul(x, y, *, Z, Z_future, alpha_ci, n_samples, random_state)and the process models’predict_rul(current_degradation, *, Z, alpha_ci)takeZfirst and the rest by keyword, so any argument they are given by position after the query is read in the old order, with a warning. A call in the old positional order (a string second argument tocb, or two positional arguments aftersizeinrandom, read as(random_state, Z)) still works with a warning;random(size, v)on a model fitted with stress, which raised for the missingZ, now draws at stressv.
Every ``random`` takes ``random_state`` (#389). The univariate distributions and models (
random,random_data), mixture models (which accepted and ignored it), Royston-Parmar and the PH, AH and accelerated-life regressions now take a keyword-onlyrandom_state:Nonedraws from numpy’s global stream exactly as before, and a seed gives a stream of its own (numpy.random.default_rng(seed)) that leaves the global one alone.conformance/test_seeds.pychecks this for every registered model that draws.``Binomial.random`` accepts a fitted model’s parameters. A fitted
nis a float (5.0), andrandomraisedTypeError: Cannot cast scalar from dtype('float64') to dtype('int64'); a whole-number float is accepted now, and a fractionalnraises aValueError.Time-varying covariate paths start at 0 (#433). A schedule starting before 0 was counted as age in the AFT
sf_tvcand the degradation stress clock: a constantWeibullAFTpath from -10 gavesf_tvc(20) = 0.8626againstsf(20, Z) = 0.9362, and the clock gave F(100) = 0.662 against 0.489. A schedule starting after a query time raised. Every schedule is now clipped to start at 0 – the part before 0 is ignored, a later start holds its first value back to 0 – for every family andStressClock(whosetau(0)is now 0).``StepSchedule.from_expression`` means what Python means (#434).
and/orreturned a bool, so"(t > 50) and 2.0 or 1.0"was 1.0 everywhere; they return an operand now (2.0 after t = 50). Keyword arguments were dropped (round(t/10, ndigits=1)gave 0 at t = 1, 2) andround(t/10, 1)raised aTypeError; both give 0.1, 0.2 now, and a keyword a function cannot take raises aValueErrornaming it.``sf_tvc`` / ``Hf_tvc`` accept any time (#435). Time 0 raised “x must contain a positive time”; it gives sf 1 (or the baseline’s value at 0) now, and negative times match
sf(NormalAFTgave 0.9725 at -5, not 0.9801).givenat or below 0 now conditions as documented (a Logistic baseline: sf(20 | 0) = 0.9378, not the unconditional 0.8831). The newconformance/test_tvc.pychecks every model withsf_tvcagainstsffor constant paths.Kaplan-Meier no longer fails when the estimate underflows (#450). Once the product fell below the smallest float – 1100 staggered entries with two at risk at each failure, R = 0.5^k –
KaplanMeier.fitraisedFloatingPointErrorfrom its log-space fallback and leaked “divide by zero in log”. It gives 0 there now, quietly, and matches scikit-survival 0.28 exactly at all 1100 times. A step with no one at risk no longer leaks “invalid value” either.``bootstrap_cb`` is NaN outside the data, like ``cb`` (#452). Without a support it carried its step convention past the data: Kaplan-Meier of 1..10 (last censored) gave
bootstrap_cb([0.5, 11]) = [[1, 1], [0, 0.548]]wherecbis NaN, and a missing time got the last bounds.cb,R_cbandbootstrap_cbnow share one rule: NaN outside the data and at a missing time, or the support’s values underset_support.A left-truncated estimate’s band survives saving without its data (#451). A restored model took
band’s N from the largest risk set: 37 instead of 60 in one example, moving the band at the 20% time from [0.4900, 0.8917] to [0.4535, 0.9018] (a truncated Turnbull fit went the other way, 8 instead of 4).to_dictstores it as"band_n"where it differs, stamped schema 2; older dictionaries keep the old fallback and untruncated models are still schema 1.Recurrent cause labels (#440).
CauseSpecificNHPPandCauseSpecificMCFhandle cause labels with the same code as the univariate competing-risks models. A tuple label did not load back (from_dictraised “unhashable type: ‘list’”), mixed labels such as's'and2raised a bareTypeErrorfrom sorting, and a tuple mark on every row was split into a column so the fit raised. All three fit, predict per cause and round-trip now. The new conformance propertytest_labels.pychecks tuple and mixed labels for every model fitted withe=.Additive-hazards draws below 0 (#441).
random()of an additive hazards model on a Normal, Gumbel or Logistic baseline searches the whole support, so the shareff(0)of its mass below 0 is drawn there (GumbelAH: 0.0347 of the draws againstff(0)= 0.0361); it returned 2.7e-20 for all of them.Changed: ``ParametricCompetingRisks.bic()`` is the joint criterion,
2 neg_ll + K ln(n)with K the parameters of all causes and n the failures of any cause, as every other SurPyval BIC counts n. It was the sum of the causes’ BICs, which charged each cause onlylnof its own failures: the competing-risks guide’s example goes from 2220.8 to 2223.1 (Weibull + Exponential).aic()was already the joint AIC.ExpoWeibull is accurate in both tails (#436).
1 - exp(-t)rounded to 0 below t = 1e-16 andx / alphaoverflowed:log_df(1e-4, 10, 4, 0.5)was +inf (true -13.12),ff(1e-3, 10, 4, 2)23% high, and at mu = 1hf,Hfandlog_sfat x = 100 were NaN, inf and -inf (the Weibull’s 30, 1000 and -1000). Every function is now computed on the log scale with exact branches for each regime and exact values at 0,momentandmeantake array parameters, and a fit that failed on data with a value at 1e-4 (returning its start, neg_ll 124.97) now reaches 90.97.``CustomDistribution`` checks its inputs (#437). A parameter named after a model attribute (
k,dist,data,method, …) overwrote it silently –kmoved the AIC from 289.50 to 290.08 – and is now refused with aValueErrorlisting the reserved names.qfoutside [0, 1] is NaN (it was the support’s lower bound), any(x, *args)signature is accepted, and reusing a name warns that it replaces the registry entry used to restore saved models.``fit_from_non_parametric`` matches ``fit(how=’MPP’)`` (#438). It plotted the censored times too (alpha 10.711 instead of 10.597 on censored data); it plots the failure times only now.
fit_from_ecdfraises aValueErrorfor an F outside [0, 1] or NaN (dropped silently before) and for unequal lengths (anIndexErrorbefore).Probability plots with a tick at 0 (#439).
round_sigtooklog10(0), soNormal.fit([-1, 0.5, 2, 3, 5]).plot()raisedOverflowError. Zero, negative and non-finite ticks work, andround_sig(0)is 0.Mutation testing pilot (#396).
scripts/mutation/run.shruns mutmut on a module in a copy of the repository, andrecheck.pychecks new tests against its survivors. On the non-parametric estimators 593 of 2,702 mutants survived the test suite (score 78.1%); 273 were real gaps, now covered bysurpyval/tests/mutation, which raises the score to 88.6% (93.2% without the equivalent and dead-code mutants). Among the gaps no test checked: the Hall-Wellner band’s width off by a factor of N,dfignoringinterp, andrmst_diff’s interval and ratio. It found #450-#452, pinned as strict expected failures, and that a warning raised inside a shape-wrapped method pointed at the wrapper instead of the caller (fixed).Changed: shape in, shape out, for every model (#381, #435). A function evaluated at query points –
sf,ff,Hf,hf,df,qf, the per-cause and recurrentcif,iif,mcf,sf_tvc,Hf_tvc,smoothed_hfand every confidence bound – returns the query’s shape: a scalar gives a numpy scalar, 1-D and 2-D queries keep their shape, an empty query gives an empty array, and a two-sided bound adds a trailing[lower, upper]axis. The non-parametric estimates, Royston-Parmar, the AFT, PO and AL regressions, Cox, the competing-risks and recurrent models, the degradation models and everycbreturned(1,)for a scalar ((1, 2)for a bound); several raised on a 2-D or empty query; and some gave a right-looking shape with wrong values – a Kaplan-Meiercbof a (2, 2) query had its axes transposed (lower 0.724 above upper 0.063), and a copula’s (2, 2, 2) query mixed its coordinates. Of 18,342 surveyed calls, 5,506 changed shape and no 1-D value changed. Survival trees and forests keep their row-by-time grid, now(n_rows,) + x.shape. Code that indexed a scalar query’s result (km.sf(5)[0]) now uses the result directly. Parametricsf_tvc(..., given=nan)is now NaN.Tail accuracy is checked against 50-digit references (#398).
reference/test_tails.pycomparessf,ff,df,hf,Hf,qfand the log forms of 17 distributions with mpmath values (stored intails_mpmath.json, written byscripts/reference/tails_mpmath.py, so CI needs no mpmath) on a grid of extreme parameters and times, from survival 1e-300 to 1e-300 of failure. It needs relative accuracy 1e-8 where the value is a normal double, or 64 ulps of the inputs’ own sensitivity where the function is ill-conditioned. 212 groups of values fail, pinned by cause: cancellation near probability 1 (#442), log-scale functions that under- or overflow (#443), NaN at valid arguments (#444), overflow errors at extreme shapes (#445), Geometric at small p (#446),qfat tiny probabilities (#447), BetaGeometric (#449), and the ExpoWeibull (#436) and Logistic (#410) forms. Beta, NegativeBinomial, DiscreteWeibull and Binomial await their references (#448).Changed: ``random()`` of a non-parametric estimate draws from the estimate itself. It drew each observed value with the estimate’s probability there, but where the estimate does not reach zero it spread the remaining probability over the observed values, so the draws disagreed with the model’s own
sf(by 0.12 for one Turnbull fit) and an all-censored fit raised. Each draw is nowqf(u)for one uniformu, and the probability left beyond the last time is drawn asinf, as for a parametric model’s never-failing units (#403).Every model is refitted to data drawn from itself (#397). A new nightly study,
calibration/test_refit_registry.py, takes each model in the conformance registry that can simulate (120 of 128; the rest are excluded with a reason), draws a few hundred units from its fitted fixture 20-100 times, refits, and requires the mean estimate within3/sqrt(reps) + 0.2standard deviations of the truth and the mean curve within 3 Monte Carlo standard errors + 0.02. A likelihood that ignored delayed entry shows as a 1.44 sd bias against a tolerance of 0.5. It found thatrandom()of an additive-hazards model on a Normal, Gumbel or Logistic baseline never draws below 0, putting that mass (3.6% for one fixture) at 2.7e-20 instead (#441).Added: ``set_support`` for the non-parametric estimates. Outside the data a non-parametric estimate only had a convention: the step curves started at 1 and held their last value however far away, while the interpolated forms, the confidence bounds and the mean cumulative functions were NaN.
KaplanMeier,NelsonAalen,FlemingHarringtonandTurnbullmodels,CompetingRisks(both methods),NonParametricCountingandCauseSpecificMCFnow takemodel.set_support(lower, upper): every function, everyinterpand the pointwise bounds are then at their start value (sf1, the rest 0) fromlowerto the first observed value, hold the last value up toupper, and are NaN outside. Negative and infinite bounds are allowed (the variable need not be time). The bounds are the model’ssupport, as for the parametric models, and are saved byto_dict(schema 2). Without the call nothing changes.Fixed: a failing non-parametric call no longer silences numpy for the whole process.
cb,R_cb,bandand the Turnbull fit turned numpy’s floating-point warnings off withnp.seterrand back on afterwards; a call that raised in between –cb(bound_type="bogus"), an unknowninterp,band(alpha_ci=2.0)– left them off for every later computation in the session. They now usenp.errstate, which restores the state however the call ends.Fixed: cubic non-parametric curves no longer dip below 0 (#417). At the last time of a Kaplan-Meier estimate that falls to 0,
sf(x, interp="cubic")was -2.3e-17, soHfthere was NaN with a raw warning instead of inf. The PCHIP curve is now clipped to the range of its knots.Silent non-convergence is checked for every model (#401). A new conformance property,
test_convergence.py, forces each iterative fit to fail – an iteration limit of 1, a start a million times the answer, or data whose likelihood has no maximum – and requires a warning, aValueError, or the true maximum. It found 62 fits that return a wrong model without a word, pinned as known failures: for exampleWeibull.fitfrominit=[1.03e7, 2.32]returns alpha 1.03e7, beta 0.099 (log-likelihood -78.2 against -37.9; #427), and every parametric PH/AFT/PO model gives a group with no events a finite coefficient (-16.3 for WeibullPH) whereCoxPHwarns (#392); also #428 and #429.Fixed: ``init`` with an offset is checked in the right order. The check read
[gamma, *params]as[*params, gamma], so it refused valid starts (Exponential.fit(..., offset=True, init=[6, 10]): “gamma = 10.0 lies outside its bounds”) and let an offset beyond the first observation through to fail later as “MLE Failed”.Changed: ``random()`` of a limited-failure or zero-inflated model draws lifetimes (#403). It returned
(x, c, n, t)survival data whenp < 1and an array otherwise, and drew zero-inflated samples by a binomial count and a shuffle, so a seed did not giveqf(u). It now returns an array for every model,qf(u)from one uniform per draw:inffor a unit that never fails, 0 for one dead on arrival, and the same numbers asqf(np.random.random_sample(size))after the same seed. The survival-data draw is the newrandom_data(), which censors the never-failing units after the last failure, ready to refit.Changed: ``mean()``, ``moment(n)`` and ``var()`` of a limited-failure model are infinite (#404). They returned the defective values (79.76 for
Weibull.from_params([100, 2], p=0.9)), which code readingmean()as the mean life took at face value. A fraction1 - pnever fails, so they are nowinf;defective=Truegives the old values. Models withp = 1are unchanged.Added: ``df(x, continuous=True)`` (#405). A zero-inflated model’s
df(0)is the point massf0, so integratingdfon a grid from 0 counted a spuriousf0 * dx / 2(0.95 instead of 0.90 forf0 = 0.1).continuous=Truereturns the continuous part alone,(p - f0)times the base density; the docstring says whatdf(0)is. A zero-inflateddfat a scalar now returns a scalar.Added: ``model.with_params(params)`` and ``model.extras`` (#406).
from_params(model.params)silently dropped the offset,pandf0(sf(50)0.7788 instead of 0.7533).extrasis the dict of those the model has ({"gamma": 5.0, "p": 0.9, "f0": 0.1}, empty for a plain model), andwith_paramsrebuilds the same model with other parameters, validated asfrom_paramsvalidates them.Fixed: loose ends of the zero-inflation mass at 0 (#407).
Hfbefore time 0 was-0.0; it is0.0. The entropy docstring and error message still put the mass at the offset; they now say 0.Every public item has a runnable example (#402). 49 public classes, functions and fitters had a docstring but no example, and the low-level
kaplan_meier,nelson_aalenandfleming_harringtonhad no docstring. Each now has a short, seeded example that runs as a doctest in CI, and the conformance check allows no public item without one.Changed: the additive hazards model holds its estimate past the last observed time (#400).
AdditiveHazardsModel.Hfkept changing after the last observed time, at the last interval’s rate \(\beta'(Z - \bar Z)\), where there is no risk set to estimate anything from: on one fit,Hfwas 3.52 at the last time and 23.0 at 100 times it. It now holds its value there, as every other semi-parametric estimate does, sosfandffhold andhfanddfare 0.hfat a NaN time is now NaN rather than \(\beta' Z\).Design principles (#379). A new page, Design Principles, lists the rules every model keeps – one data format,
nanin and out, order, units and counts not mattering, consistent shapes and identities, behaviour outside the data, entry points agreeing, the same defaults and names everywhere, calibrated and consistent intervals, one seed rule, useful warnings, documented examples – each with the tests that enforce it and the issues where a model does not yet comply. The README summarises them. Three new conformance checks fill the gaps: behaviour outside the data (test_outside_data.py: a step or semi-parametric estimate starts at its initial value and, past the last time, holds or isnanfor every function alike; found #400, additive hazards extrapolating), the same defaults across a fitter’s entry points (test_defaults.py), documentation (test_documentation.py: every public item has a docstring with an example; 49 are listed against #402 and the list can only shrink).Option sweeps in the conformance suite (#379). The other checks call each model with its default options, where many past bugs lived in the others.
test_options.pysweeps every confidence-bound method of every registered model overon=,bound=,alpha_ciand their variants, everyinterp=value and every estimation option. It checks that the bounds contain the estimate and stay in range, that one-sided and two-sided bounds agree, that intervals nest and close onto the estimate asalpha_ciapproaches 1, thesf/ff/Hftransforms and the output shapes, and that shared options have one name and default across models. It found bound failures now tracked as #411 and #413-#419 (for example discretecb(on="hf")centred ondf/sf(k)instead ofhf, Royston-Parmar one-sidedff/Hfbounds on the wrong side, andHfbounds capped at 34.54) and ten naming inconsistencies (#422), each a strict expected failure.Raw numerical warnings no longer leak from Kaplan-Meier, Binomial, Weibull and LogNormal functions.
KaplanMeier.Hf/hfpast the time the estimate reaches zero,Binomial.Hffromx = non,Weibull.df/hfat 0 with a shape below 1, andLogNormal.sfat 0 gave numpy “divide by zero” or “invalid value” warnings; their values (inf, or 1 forsf(0)) were already right and are now returned without a warning. The conformance and property suites now fail on any raw numpy, scipy or autograd warning that escapes the package, and check that a fit or prediction gives each deliberate warning at most once (#379). The check found five wrong results hidden behind warnings, tracked as #408-#412.Changed: ``CoxPH.fit`` defaults to Efron ties, with the matching Efron baseline (#387).
fitdefaulted to Breslow whilefit_from_df, the time-varying-covariate fits and the competing-risks Cox model defaulted to Efron, so the same tied data gave different models by different routes. Every route is now Efron. Efron is chosen on merit: with ties from rounding a continuous time, Breslow biases the coefficients towards zero (in a simulation with true \(\beta = 0.7\), by -0.06 to -0.21 as the ties grow, against -0.006 to -0.06 for Efron), at no saving worth having. An Efron fit’s baseline hazard now takes the same tie correction as its likelihood: the \(d\) deaths tied at a time leave the risk set a fraction at a time, and the step is \(\sum_{l<d} 1 / (R - \tfrac{l}{d} R_D)\) instead of Breslow’s \(d / R\) – the covariate-weighted Fleming-Harrington estimator, as Breslow’s is the covariate-weighted Nelson-Aalen. It matches R’ssurvfit.coxphafter an Efron fit (checked against it in the reference tests) and the Efron residuals, which already used it. Results change on tied data: passtie_method="breslow"for the old fit. Without ties every method gives the same model as before.Property-based tests (#379). Hypothesis generates data with mixed censoring, ties, counts, truncation and tiny samples, and checks the non-parametric estimators, parametric fits, regression, competing-risks, recurrent-event and serialisation paths against general properties (valid curves, local optimality against an independent likelihood, invariance to row order, units and counts,
ValueErroron invalid input), shrinking any failure to a minimal case. The default run takes under a minute;SURPYVAL_HYPOTHESIS_PROFILE=nightlysearches thoroughly in the nightly workflow.hypothesisis a new test-only dependency. It found four bugs, pinned as strict expected failures: Turnbull dropping its last piece under right truncation (#391), silent degenerate fits where the likelihood has no maximum (#392), unit-dependent fits to truncated data (#393), andCoxPHaccepting an infinite event time (#394).Statistical calibration suite and nightly run (#379). Simulation studies in
surpyval/tests/calibration(opt in with--run-calibration) check that results are statistically right, not only consistent: confidence-interval coverage for parametric, non-parametric, Cox and parametric regression, degradation and recurrent bounds; parameter recovery with truncation, interval censoring, limited failure, frailty and renewal models; size and power of the log-rank, stratified log-rank, Gray, Laplace, MIL-HDBK-189C and Cramer-von Mises tests; and Brier/AUC bias with tied times. Each passes within 3 Monte Carlo standard errors plus a stated slack, with fixed seeds, and would have caught the old Gray’s test (size 0.20 against 0.05) and the #365 Brier bias. The schedulednightly.ymlruns the full suite on three Pythons, the docs build and the calibration suite againstdevelopdaily, once it is onmaster. Found: the equal-precision (nair) Kaplan-Meier band covers about 0.89 for a nominal 0.95 (#390).A conformance suite checks every model against the same properties (#379). Bugs kept reappearing as old kinds of failure in new models (unsorted input, units, row routing, missing values, serialisation), because each fix tested only its own case.
surpyval/tests/conformanceregisters every public model (128 cases) and runs each through the identities between its functions; scalar, 2-D and empty queries; query and row order; units, data-row order and counts; valid values; the missing-value rule; seeds; the strict-JSON round trip; and agreement of its fit paths. A test fails when a public model is left unregistered. The fast set runs on every pull request (about 40 s). The 58 failures it found are strict expected failures, each naming its issue (#381-#388).Numbers quoted in the documentation are checked (#379). The prose around executed examples quoted outputs (“a shape of about 2.1”, “the lower AIC”) that nothing verified, so they went stale when outputs changed. Hidden cells now assert 263 such claims across 16 pages, and the documentation build fails when one no longer holds. The first pass found two stale statements in the offset section of Parametric SurPyval Modelling: the starting offset is
min(x)minus the data’s mean spacing, notmin(x) - 1, and the example’s quoted moment-based shape was from a different sample. See “Checking the numbers quoted in the text” in Contributing.Stored results from R and Python survival software (#379).
surpyval/tests/referencecompares SurPyval with 82 results computed once on shared fixtures (lung, heart, aml, ovarian, PBC, and small sets with ties, left truncation, interval censoring and competing risks) by R survival 3.5-8, cmprsk 2.2-11, timereg 2.0.5, pec, riskRegression, npsurv and fitdistrplus, lifelines 0.30.3 and scikit-survival 0.28, so CI needs neither R nor lifelines;scripts/reference/regenerate.shrebuilds them. Kaplan-Meier, Nelson-Aalen, restricted mean, log-rank, MCF, Aalen-Johansen, Lin-Ying, Brier score and AUC agree to rounding; Cox (Breslow, Efron, strata, left truncation, start-stop), survreg AFT fits, Fine-Gray and Turnbull to between 1e-6 and 5e-4. Deliberate differences are asserted and recorded with their reason. Gray’s test disagrees with cmprsk in its variance (#380).Degradation: missing values give NaN, and predictions read a DataFrame by name (#375, #374). Gamma- and Wiener-process models gave sf = 1 and
ff = Hf = hf = df = 0at a NaN time andqf(nan) = inf, raised on a NaN stress, andpredict_rul(current_degradation=nan)never returned.DegradationModel.qfreturned inf for a NaN covariate or p and used only the first row ofZ;InducedFailureDistributiongaveff(nan) = 0; the bootstrapcbraised on a NaN stress. A missing time, stress or probability now gives NaN for that element only, andqfpairs each p with its row ofZ.predict_rul,predict_failure_timeandinduced_lifestill raise for a missing value (they describe one unit), and the process quantile search can no longer loop forever.DegradationAnalysis.fit_from_dfand the newWienerProcess.fit_from_df/GammaProcess.fit_from_dfrecord the stress columns asZ_cols(kept throughto_dict), so every method that takesZaccepts a DataFrame and selects those columns by name; a model fitted from arrays refuses one with an accurate message (it used to say “fit the model withfit_from_df” to a model fitted that way).One rule for missing values (#375). Prediction: a missing covariate, time or probability gives NaN for exactly the outputs that depend on it, and a method whose input is one unit’s history (
predict_rul,induced_life,sf_tvc,mcf) raises instead. Fitting: rows with a missing covariate are dropped with one warning where rows are independent observations, and refused where a row is only part of one; a missing time or response always raises. See Conventions. Fixed to follow it:Survival trees and forests sent a missing covariate right at every split, so it predicted like +inf (tree sf 0.6974 for both), and kept such rows in the fit without a warning. They are now dropped with a warning, predict NaN, and
RandomSurvivalForest.scoreis NaN when a score is missing.Proportional-intensity
mcfwith a missing covariate ran every sequence tomax_eventsand then reported a missing time; it and the simulation entry points now refuse a missing or mis-shapedZby name before simulating.survival_probabilitycastZto float, so a formula fit with string levels could not be scored; a DataFrame is now passed tomodel.sfas it is.A missing time at prediction returned the value at t = inf in
CoxPH(every method, plain and stratified: sf 0.0102,hfanddf0), competing-risks Cox (cif 0.385) and Fine-Gray (0.354), and sf = 1 in Buckley-James; it now gives NaN. Coxpredict_tvcrefuses a covariate path with a missing value by name.Fine-Gray and the competing-risks Cox array path dropped rows with a missing covariate silently; they now warn like every other fitter (and drop infinite covariates too).
Stratified
CoxPHraised aTypeErroron a missing stratum label; such observations are now dropped with one warning, and the array path warns once in total rather than once per stratum.Kaplan-Meier, Nelson-Aalen, Fleming-Harrington and Turnbull (every function,
cbandband) and non-parametricCompetingRisksreturned the value at t = inf for a missing time (a NaN sorts past the last step), and the arrayhf/dfcopied a neighbour’s increment into it; parametricsf_tvc/Hf_tvcraised anIndexError. They now give NaN for that time only.Covariates given as a list or object array holding
Noneraised aTypeErrorin the parametric PH and AH families, the accelerated-life fit andAdditiveHazardsprediction; they are now read as floats, soNoneis a missing value.
Competing-risks Cox pairs each time with its own covariate row. With one row per time and unsorted times,
hf/Hf/sf/ff/dfread the baseline at the sorted times but used the rows in the given order (times [10, 1], rows [2], [-2]:Hfgave [0.139, 0.590] instead of [4.961, 0.017]).cifwas not affected.A declared category level with no fitted rows is refused at prediction (#377). A level listed in
C(g, levels=[...]), or an unused category of apd.Categoricalcolumn, got a coefficient with nothing to estimate it, so its predictions were made up (the reference level’s forWeibullPHandCoxPH; a drifted coefficient forWeibullAFT, S = 0.863 against 0.803). The fit now warns once, naming the column and the empty levels, and keeps the column so coding stays the same across data splits; predicting for such a level raises the same “not fitted with”ValueErroras an unseen level, fitted or restored.AdditiveHazardsand Buckley-James, which reject an all-zero covariate column, still refuse such a fit after the warning. A saved model with an empty level needs schema 2.Turnbull reaches the maximum-likelihood estimate with interval censoring and right truncation (#368). Two index searches were one Turnbull piece off: a right-censored observation could not fail in the piece just after its censoring time, and a right-truncated window
(tl, tr]took in the piece just aftertr. Exact and right-censored data were unaffected; otherwise the EM converged to a curve that was not the NPMLE (one failure in (1, 2] and one unit censored at 1.5 were fitted at a likelihood of 0.375 instead of 1). On 600 random small data sets the old fits fell up to 1.4 log-likelihood units short without right truncation and 25 to 68 with it; some right-truncated fits had likelihood zero, and a few doubly truncated ones raisedIndexError. Every fit whose NPMLE exists now matches an independent maximisation to 2e-9, and thenpmleverdict, now built on the corrected supports, agreed with the EM’s behaviour on all 399 of those data sets where it gave a firm verdict. Delayed-entry data in which a unit is censored before a later unit enters is now reported"not unique"(with a warning), since the mass between them is not determined; the Kaplan-Meier option still returns the delayed-entry Kaplan-Meier. The Nair interval example in the docs rises from -59.52 to -58.06 in log-likelihood.Formula models refuse a category level they were not fitted with (#371). Predicting for a level absent from the fitted data coded it silently as the reference level (
WeibullPHgave S(5) = 0.5283 for bothg="a"and an unknowng="d"), with only formulaic’sDataMismatchWarning. Every family that takes aformula(parametric PH/AFT/PO/AH,AcceleratedLife,CoxPH,AdditiveHazards, Buckley-James, frailty, competing-risks Cox and Fine-Gray) now raises aValueErrornaming the column and the unknown levels, fitted or restored. A fit whose data has a level outside itsC(g, levels=[...])list raises too; declared levels count as known. A missing categorical value still predicts NaN in place, as a missing numeric one does. Buckley-James returned survival 0 for any missing covariate and now returns NaN.Competing-risks Cox predicts from a DataFrame (#370).
CompetingRisksProportionalHazardsread a DataFrame by column position: withZ_cols=["z", "w"], passing the columns as[w, z]changed S(5) from 0.655 to 0.914, and aformulafit could not expand raw covariates at all.sf,ff,Hf,hf,df,cif,phiandphi_enow select and encode the columns recorded byfit_from_df, asCoxPHdoes, fitted or restored. Arrays work as before.Proportional odds fits time-varying covariates (#372).
PO(dist)models could be evaluated along a step covariate path but not fitted to one; the docs said PO lacked the structure. It does not: the PO hazard \(h_0 / (F_0 + \phi S_0)\) depends only on the time and the current covariate, so splitting a subject into delayed-entry intervals is exact.fit_tvc,fit_tvc_timelineand their_from_dfforms now work forWeibullPO/PO(dist). On simulated step-path data (8 x 2,000 subjects, truth [10, 2, 1, -0.5]) the mean estimate is [10.03, 1.98, 0.98, -0.50], and the fitted negative log-likelihood equals the path likelihood fromsf_tvc/hfto about 1e-12.``fit_tvc`` no longer truncates at time 0. For PH, AH and PO models with a baseline defined below zero (Normal, Gumbel, Logistic), each subject’s first interval was treated as left-truncated at 0, conditioning the fit on surviving to 0, so a constant covariate split into intervals did not reproduce
fit(LogisticPO scale 5.22 against 9.33, NormalPH 5.56 against 9.65). A first interval starting at 0 is now untruncated, matchingfitandsf_tvc. Baselines on the positive axis are unchanged.Survival tree predictions for several subjects (#369).
SurvivalTree.sf(x, Z)(andff,df,hf,Hf) routed a covariate matrix by a row,Z[split_index], instead of a column. With one covariate every subject silently got the first subject’s curve (S(5) = 0.8811 for all rows, where row by row gives 0.2955 for half of them); with two or more it raised.survival_probability, and so the Brier score and AUC, were wrong for a single tree. Each row now goes to its own leaf, and a 2-DZreturns an(n_rows, n_times)grid equal to stacking the per-row results, asRandomSurvivalForestdoes; a 1-DZ(one subject) is unchanged. The forest, already correct, now evaluates the whole matrix in one call per tree: identical results, about 3x faster insurvival_probabilityandscore.Turnbull decides from the data whether its estimate exists (#327). A fit warned “not identifiable” when more than 90% of its mass sat on pieces some observation gains from and none pays for, or when the EM did not converge: a cut-off tuned on simulated samples. On samples whose estimate does not exist that share ranged from 0.11 to 0.99 depending on how far the EM had got, so half were caught only because they had not converged; other non-existent estimates (a delayed-entry Kaplan-Meier that drops to zero before a later entry, Lynden-Bell and doubly truncated exact data) were reported only as not converged, and flat likelihoods not at all. The new
model.npmleis"exists","not unique","does not exist"or"undetermined", from a structural criterion checked before the EM runs: Vardi and Wang’s graph condition for exact data, and a hazard-scale gap argument for one-sided truncation with any censoring. It takes a few milliseconds on thousands of rows, and the warnings name the case and the time involved. On 240 simulated left-truncated samples, all 63 “does not exist” fits drifted to the boundary and none of the 176 “exists” fits did. With censoring and truncation on both sides existence can depend on the counts, and such fits are reported as"undetermined". The fitted estimate is unchanged, andexploitable_massis still reported as a diagnostic.Every regression formula round-trips through serialisation (#244). A model fitted with
fit_from_df(..., formula=...)refusedto_dictfor wrapped categoricals (C(g),C(g, levels=...),C(g, contr.sum)) and fitted transforms (scale,center,poly,bs,cs). It restored integer-level categoricals with string levels, so every row was coded as the reference level (sf off by up to 0.09), and Cox and competing-risks Cox models lost a0 +from the formula, so a restored model had 3 design columns for 4 coefficients and could not predict.to_dictnow stores each factor’s levels, in order and with their types, and each transform’s fitted state as strict JSON;from_dictrebuilds the same design-matrix transformer, and restored models predict identically (rtol 1e-12) across the PH/AFT/PO/AH, accelerated-life, Cox, Lin-Ying, Buckley-James, frailty and competing-risks families. A formula is checked when saving, so anything that cannot be restored raises into_dict. A formula that SurPyval 0.20 cannot rebuild is stamped schema 2, so 0.20 asks for an upgrade instead of failing with a formula error; plain columns and string categoricals stay schema 1. Old files still load, and a Cox file missing its0 +is repaired. Also fixed: predicting with plain integers for a column fitted as an integerCategoricaltreated it as numeric (sf 0.372 instead of 0.083).Brier score and time-dependent AUC with tied event and censoring times (#365, #290). The censoring survival \(\hat G\) behind the inverse-probability-of-censoring weights counted an event as still at risk of being censored at its own time, and weighted it by \(1/\hat G(x_i)\). The metrics now use the events-first reverse Kaplan-Meier (as
prodlimand scikit-survival) and weight an event by \(1/\hat G(x_i-)\) (aspec; Gerds and Schumacher 2006). On a data set whose true values are known exactly, the Brier score at t = 2 was 0.2330 against a true 0.2250 (now exact) and the AUC 0.6703 against 2/3; in simulation with discrete times the old Brier score was biased by -0.029 and is now unbiased (scikit-survival’s \(1/\hat G(x_i)\) weighting gives +0.011). Without such ties the results are unchanged and equal scikit-survival’s.censoring_survivalgainsties=; Fine-Gray keeps itscmprskconvention. Also:integrated_brier_scoresorts an unsorted grid (0.1908 became 0.1949 on one example);cmust be 0 or 1 and matchxin length (a left-censored row was scored as a survivor);x_trainneedsc_train; and a horizon that needs the training \(\hat G\) where it has fallen to 0 scores NaN rather than being biased towards 0 (0.179 against a true 0.25).Proportional odds along a time-varying covariate path (#236).
sf_tvc/Hf_tvcraisedNotImplementedErrorforPO(dist)models. The PO hazard \(h_0 / (F_0 + e^{\beta'z} S_0)\) depends only on the time and the current covariate, so the cumulative hazard along a step path is exactly the sum of the constant-covariate increments, as for PH; PO now takes that path. It matches a numerical integral of the hazard to 1e-9, and a constant path givessf(x, Z)to 1e-13. PO’sHfis now computed as \(H_0 - \ln\phi + \ln(F_0 + \phi S_0)\) rather than-log(sf): before, a change-point where the baseline survival underflows made everysf_tvcvalue NaN (WeibullPOwith a change at t = 1500: S(5) was NaN, now 0.8387). Time-varying fitting is still not available for PO.``sf_tvc`` for PH and AH with a baseline defined below zero. For a Normal, Gumbel or Logistic baseline a constant covariate path gave the survival conditional on surviving to time 0, not
sf(x, Z)(at x = 5: PH(Normal) 0.91842 against 0.91551, PH(Gumbel) 0.90613 against 0.87800). The first segment now starts at the bottom of the support, so a constant path reproducessfexactly.Recurrent-event simulations are much faster (#362); seeded results change.
mcf,plot,time_terminated_simulation,count_terminated_simulation(and their..._dataversions) and the renewal models’cramer_von_misesbootstrap used to simulate one item and one event at a time, with two model calls per event. Every item is now advanced together, one event per round, with one array operation per round, and the renewal models’ root finding is vectorised too. Anmcfover 1000 items is 23-46x faster for the Kijima and ARA models (GeneralizedRenewalwith a Weibull lifetime and Kijima II: 2.0 s to 0.06 s), 9-18x for ARI and the intensity models, and 4x for the already cheap G1 model. The simulated MCF also no longer computes the Lawless-Nadeau variance it then discarded, which was most of the time for the intensity models with many items. The draws follow the same processes (checked against one-sequence-at-a-time references to round-off), but the uniforms are assigned to events in a different order, so a givenseednow gives different simulated values and bootstrap p-values than in 0.20. The unused uniform-pool helpersinitialize_simulation,get_uniform_random_numberandclear_simulationare removed. Three examples in Recurrent Event Modelling with SurPyval use new seeds so that they still illustrate what the text describes.One seeding rule for every random draw (#361). With the default
random_state=None(orseed=None), the non-parametricrandom()andbootstrap_cb(),ParametricCompetingRisks.random(), the copulas’sample_uv()(and sorandom()), the degradation models’random(),induced_life(),predict_rul()and bootstrap bounds, the Buckley-James bootstrap and the recurrent-event goodness-of-fit p-values used a fresh OS-seeded generator on every call, sonp.random.seedhad no effect on them while it did controlParametric.randomand the recurrent simulations.Nonenow draws from numpy’s global RNG throughout (surpyval.utils.rng.as_generator). An explicit seed orGeneratorgives the same stream as before. See Conventions.``import surpyval`` no longer imports matplotlib (#363). pyplot is imported inside the plotting methods, which saves about 0.3 s on every cold start of a program that never plots. Plotting is unchanged.
Design changes approved after the third documentation review.
One sample size for BIC and AIC_c. Every model that reports a BIC or AIC_c uses the number of observed failures (exact, left- and interval-censored, weighted by their counts), or the number of observations when there is none. Recurrent-event models count observed events; copulas count rows in which at least one series failed. Previously univariate models counted failures for BIC but all units for AIC_c, regression counted exact failures only (
-infwithout one), recurrent models counted exact events (NaN without one), and copula and Royston-Parmar models counted every row. Univariate BIC on exact and right-censored data is unchanged; AIC_c on censored data now uses the failures. Restored models store the sample size ("ic_n"), sobic()andaic_c()work without the data.fit_best(metric="aic_c")raises a clear error when no candidate has a finite AIC_c instead of returningNone.Fine-Gray follows cmprsk on tied times. The censoring weights are \(\hat{G}(t-)/\hat{G}(x_i-)\), with the censoring Kaplan-Meier read just before each time, as in R’s
cmprsk::crr. Results are unchanged when no censoring time equals an event time.Readings from a coarse gauge.
GammaProcess.fit(gauge=...)(withrounding,exact_startandgauge_method) maximises the probability that each unit’s path passes through its recorded gauge bins. With a gauge step near the mean increment, taking rounded increments at face value more than doubledalphaand shrank stress coefficients; the gauge likelihood recovers the unrounded estimates. The default fit is unchanged.Scale-equivariant parametric fits. Every continuous distribution and method, with or without an offset, now gives the same answer in any units from 1e-4 to 1e5: the search is scaled per coordinate, MLE normalises its objective per observation, MOM uses the scaled search, MPS no longer evaluates the CDF at the support edge, and offsets start one data spacing (not one unit) below the smallest value, with MPP searching the offset in the data’s spacing. At a data scale of 1e-3, Rayleigh MOM had been 1.2% off, Uniform MPS 0.15% and Beta4 MLE 0.1%, and many offset fits never left their start. Fits in their own units move by about 1e-6 relative, towards the optimum.
Strict-JSON serialisation.
to_dict/to_jsonno longer emitNaN/Infinity: non-finite values are written asnulland listed under"non_finite"(JSON Pointers by kind), and every reader restores them. Each file is stamped with the oldest schema version that reads it: 2 when it records non-finite values this way, and 1 (the layout SurPyval 0.20 reads, which loads it identically) otherwise; older dictionaries and files still load.to_json(path, with_data=True)works forParametricandNonParametric, and every class-levelfrom_dictapplies the package reader’s checks.Non-parametric copula margins. Under
how="IFM"a margin can beKaplanMeier(or any fitted non-parametric model), giving the semi-parametric estimator of Genest, Ghoudi and Rivest (1995); it used to crash.Datasets.
load_framingham,load_pbc2andload_support2expose bundled data that had no loader; the undocumentedsynthetic_dataset.csvis removed. The G1 example data are credited to Kaminskiy and Krivtsov (2010).
Bug fixes found in the third documentation review. This review probed the documented behaviour adversarially (identities, round trips, cross-method agreement, edge cases). Each fix has a regression test that fails on the old code.
Wrong results that are now correct.
Degradation: the default two-sided analytic
cbband was a 90% band (each side used the fullalpha_ci). Units already past the threshold at their first measurement were treated as survivors; they are left-censored there. Wienersf/ffreturned NaN for low noise;GammaProcessModel.meancould be negative; zero Gamma increments are censored below aresolutioninstead of a 1e-12 nudge.Gray’s test is Gray’s (1988) statistic with group-specific censoring; the pooled version rejected a true null up to 89% of the time when groups were censored differently.
Parametric competing risks:
cifandprobability_of_causeare integrated per query instead of on a fixed grid (they could sum to 0.81, or 0.0005).Recurrent events: GRP/ARA simulation was wrong at long horizons and NHPP simulation failed past ~745 expected events; renewal fits could keep a worse optimum than one they found (boundary optima such as ARA
rho -> 1); left-censored counts now cover(tl, x].Regression: a missing or infinite covariate made PH/AH/frailty return their starting values (rows are now dropped with a warning in every fitter); stratified Cox
predict_tvcused the first stratum’s baseline; Lin-Ying predictions depended on covariate centring; counts were treated as clusters in robust standard errors,check_phranks and the Buckley-James bootstrap; some AFT/PO fits stopped short of the maximum (PO(Weibull)by 5.5 nats) and are finished by a gradient-based optimiser.Parametric:
csignoredp,f0and the offset; likelihood-ratio bands collapsed onto the estimate when the inner search failed (Geometric coverage 0.65); the mixture EM stalled on alog(0); ExpoWeibull andCustomDistributionmoments were wrong away from unit scale; MPS and MSE were not scale invariant; raw distribution functions were evaluated outside their support;bic()was-infwithout exact failures; LogNormal and Gamma hazards overflowed in the far tail.Non-parametric:
qf/medianhad no round-off tolerance (the median of 1..30 was 16);band()critical values were ~1.5% low; log-rank counted groups never at risk in its degrees of freedom; the Turnbull identifiability warning fired on correct fits.Copulas: Frank overflowed for \(\theta \gtrsim 37\) and Clayton collapsed at extreme \(\theta\); joint MLE dropped pre-fitted margins’ options; automatic derivatives summed over broadcast axes.
Data layer: NaN truncation bounds were read differently by each fitter; unsorted xrd input gave a wrong estimate.
Crashes and unclear errors. Two-column
xwithout intervals now works in every fitter; truncation rules are identical for one- and two-columnx; badfixed/init/bound/howarguments, wrong covariate row counts, degenerate data, out-of-range parameters infrom_params/fit_from_parameters, non-integer data for discrete distributions and corrupt serialised dictionaries raise clear errors. Models restored without their data explain what needs it.Serialisation. Discretize,
CustomDistribution(after re-construction),NeverOccurs/InstantlyOccurs, destructive degradation models with any distribution, and competing-risks models with mixed or tuple labels round-trip; Cox dictionaries are strict JSON; likelihood-ratio bounds work after awith_datarestore.Behaviour changes to note. BIC’s sample size for univariate models counts every non-right-censored failure;
aic_cis NaN when \(N \le k + 1\); recurrent covariates must be constant within an item; Coxmodel.phiis a method; stratifiedsf_tvcrequiresstratum=;band()’sn_sims/random_stateare deprecated; two-columnxwithout intervals is stored as one column.Bug fixes found in the second documentation review. Each was reproduced first and has a regression test that fails on the old code; documentation describing the old behaviour was updated.
Non-parametric.
Turnbull confidence bounds stepped a piece too early. The variance at each value included the next piece’s expected failures, so
cb()on interval-censored data gave[0, 1]where the estimate was 1. Ther,dand variance ladders now line up withR; the Kaplan-Meier option without truncation starts the EM on Turnbull’s innermost intervals; MPP fits withheuristic='Turnbull'pair each failure with the CDF after its drop (their estimates move slightly).smoothed_hf()works on Turnbull models; Turnbullbootstrap_cb()refits with the fit’stolandmax_iter; Greenwood’s variance no longer blows up (about 1e14) from round-off at the last value.'Benard'plotting positions use Benard’s (i - 0.3)/(N + 0.4); an unknownturnbull_estimatorraises aValueErrorup front.
Parametric.
Distribution parameters named
pclashed with the limited-failure proportion.param_cb('p')on a Geometric or NegativeBinomial now bounds the distribution’s ownp, andlfp=Trueworks for them (the proportion is namedlfp_p).Uniform MLE with censoring was not the maximum.
(min, max)is used only for exact data; censored data gets a bounded search (b = 24.75, not 10, in the reported example).Method of moments could return its start as the fit. Beta-Geometric MOM has a closed form, non-finite moments at the start raise, and the search backs away from regions without moments. With
fixed, MOM matches one moment per free parameter.CustomDistributiongains finite moments andvar(), a fast MOM, and the standard MPP refusal;offset=Trueis refused for discrete distributions;param_cb('gamma')and unknown names give clear errors;var()of LFP and zero-inflated models followsmean()’s convention; the MSE/MPS fallback ends on Nelder-Mead as documented.The gradients of the incomplete gamma/beta helpers had the wrong shape when a scalar argument met an array partner (NegativeBinomial fits that differentiate the CDF crashed).
Information criteria.
AIC, AICc and BIC count only estimated parameters. Fixed parameters (and the accelerated-life placeholder) no longer add to
k, in parametric and regression models; restored models agree. A Weibull with its shape fixed now scores the same as the equivalent Rayleigh, andAcceleratedLife(Weibull, Power)AICs drop by 2.Recurrent models use one BIC sample size, the number of exactly observed events (it was every row, or only failures for ARI).
Regression.
The exact Cox tie methods are fast.
'kp'uses the Gail-Lubin-Rubinstein recursion and'exact'the DeLong-Guirguis-So integral, with analytic score and information: a 107-way tie that took over nine minutes fits in 0.06 s, and the twelve-tie cap is gone.Cox silently fitted left-censored rows as right-censored and failed on interval rows; both are refused with a clear error.
CoxPH.fit_from_dfgainstl_col, and strata labels follow the missing-covariate row mask.A parametric regression fit could return its starting values when the start’s log-likelihood was infinite; it now falls back to the default start with a warning, or raises.
AH(...).randomworks with several covariate rows; the unreachable accelerated-life(low, high)stress option is removed and a scalar stress works; additive-hazards fits held at the positivity boundary warn;FrailtyModelgainsneg_ll/aic/bic/aic_cand is warning-free near \(\theta = 0\).
Competing risks and copulas.
gray_testcounted a NaN cause as a competing failure (onlyNonewas censored); it uses the shared missing-cause rule.CompetingRisksProportionalHazards(Cox and Fine-Gray) can be saved and restored, reproducing every prediction.Copula fits start strictly inside the family’s bounds and take a public
init;CopulaModelreportslog_likelihood,neg_ll(),aic()andbic()from the full censored and truncated likelihood.
Recurrent events.
The cause-specific MCF’s bounds were too narrow: it used the per-step variance; each cause now gets the Lawless-Nadeau robust variance. The per-step variance itself was wrong with ties.
The non-parametric MCF accepts right truncation (
tr); models without data raise an informative error; G1, ARI and NHPP fits are warning-free (NHPP searches run on an unconstrained scale, moving fitted values only within the optimiser tolerance); the renewal goodness-of-fit bootstrap resimulates each item as it was observed.
Degradation.
Offset-exponential models could not be reloaded; every built-in path now round-trips.
DestructiveDegradationModelgainsto_json/from_jsonand keeps its data, so a reloaded model’scbmatches the original.
Bug fixes found while rewriting the documentation. Each was reproduced first, is covered by a new test, and the documentation that described the old behaviour (or a workaround for it) has been updated.
Regression.
Cox predictions before the first event were wrong. The baseline lookup index was -1 there, which wrapped to the last baseline value, so
Hf(0.1)returned the end-of-data cumulative hazard andsfwas near 0 where it should be 1. It is now 0 before the first event.Cox predictions paired unsorted times with the wrong covariate rows. The query times were sorted for the baseline lookup but
phi(Z)was not, soHf([3, 1], Z)gave[9.69, 0.39]instead of[1.87, 2.03]. Times and rows now stay paired, in the order given.The gamma-frailty fit broke when there was little frailty. Its group log-likelihood subtracted terms of size \((1/\theta)\log(1/\theta)\) to get a difference of order \(\theta\); as \(\theta \to 0\) round-off swamped it and the fit chased the noise (
neg_llof -1985 against the PH fit’s 755) or divided by an underflowed \(\theta\). It is now computed in a form that tends to the PH contribution, and the fit coincides with the ordinary PH fit when there is no frailty.A 1-D ``Z`` raised ``IndexError`` in the PH, AFT, PO and AH fitters; it is read as one covariate, as the other fitters already did.
Accelerated life with a Gamma baseline had the life model backwards. Gamma’s
betais a rate, so the life now enters as1 / life(as for the Exponential); a longer modelled life had meant a faster rate.PH random draws returned ``inf`` for a tiny hazard multiplier; they now invert through the cumulative hazard.
Docstring: Schoenfeld residuals follow the input order of the event rows, not event-time order.
Competing risks.
``CompetingRisksProportionalHazards`` (``how=”Cox”``) incidence could exceed 1. Its CIF weighted each hazard increment by
exp(-H)(the #278 defect, fixed for the non-parametric CIF but not here): total incidence reached 1.07-1.18 in small samples, and 118 in one. The weight is the product-limit survival, and the CIFs now sum to exactly1 - S.``CompetingRisks(method=”Kaplan-Meier”)`` only changed ``S``;
sf,ffandHfstill usedexp(-H). They now follow the method (the default is unchanged; the method is serialised).The causes of a Cox competing-risks model are sorted, so the row order of
betasno longer depends on the hash seed. Docstrings forhow/tie_method, FineGray’sc, and API entries forFineGrayModelandGrayTestResult.
Parametric.
Two-sided parameter bounds other than (0, 1) were ignored (a spline knot bounded to (0, 50) was fitted at 89.7). Every finite interval is now enforced.
Poor default starts reported success at poor optima. A
CustomDistributionstarted each positive parameter at 1, where for data on another scale the likelihood is flat to machine precision; it now starts from the best of a grid of magnitudes, with the old start as a second try. A limited-failure-population fit could settle on the worse of two optima (Meeker’s data:p = 0.116,neg_ll302.9, instead ofp = 0.0067, 293.0); it is also started from the failures alone. An MLE fit withoutinitkeeps the best of these starts.Restored models:
neg_ll()andaic()work from the stored likelihood;bic(),aic_c()andplot()explain that the data were not saved instead of raisingTypeError.Docstrings for
from_params’p, zero-inflation inqfandrandom,tl/tr,MixtureModel.loglikeand theCustomDistributionexample.
Non-parametric.
``Turnbull.fit`` modified its input. Infinite interval endpoints were rewritten in the caller’s array, so refitting the same array gave a different answer.
Fleming-Harrington could return ``nan`` when the Turnbull EM’s risk and death sets carried float noise (
1 + 2e-16); near-integer sets are treated as integers.Docstrings for
turnbull_estimator(Fleming-Harrington, the default, was missing),hf(it returns increments, not a rate), and the Turnbull EM (the estimator option acts inside the EM too).
Multivariate (copulas).
IFM ignored counts and truncation in the margins. The first stage now fits each margin with the row counts and its own series’ truncation window.
Clayton had a spurious likelihood maximum near \(\theta = 0\). The closed form rounded to
C = 1and density1/(uv)below \(\theta \approx 10^{-16}\), which negatively dependent data drove the fit into; computed throughlog1p/expm1it tends to the independence copula.MultivariateSurpyvalData: interval rows withoutxl/xrare rejected (the check could never fire), a single row ofDcensoring codes broadcasts, and noRuntimeWarningcomes from the default infinite truncation bounds.
Recurrent events.
The MCF variance ignored within-item covariance. It is now the Lawless-Nadeau robust variance; the per-step variance was about eight times too small when items differ in their rates.
``mcf_cb(bound_type=”normal”)`` scaled the standard error by the estimate a second time, giving far too wide, negative bounds; it is now
M +- z SE.``ProportionalIntensityNHPP`` stopped short of the optimum (Duane baseline: 15-30 AIC worse than the same model as Crow-AMSAA). It starts from the covariate-free baseline fit and searches on an unconstrained scale; on one-event-per-item data it now reproduces Weibull PH exactly. Crow-AMSAA’s
betais bounded below by 0.CauseSpecificNHPP(dist=HPP)works (it raisedTypeError) andHPPhasfrom_params; model summaries say how the model was obtained instead of always “MLE”; a covariate dict of scalars and a 1-DZare accepted;CauseSpecificMCF.plotdraws confidence bounds and honoursconfidence/plot_bounds.Docstrings: residuals are not exactly i.i.d. Exp(1) when an item’s window closes before an event; the Rossi example passes
iandcby keyword.
Degradation.
Bootstrap bounds failed on a reloaded model.
from_dictdid not restore the fitter the refits need, so every refit failed; it is now recovered from the restored life model (the distribution, or the regression fitter of an accelerated model).
REML population fits are 20-60x faster.
population_method="reml"– the plain and stress-dependent (links) populations, linear and nonlinear paths – now evaluates the same REML objective through the Woodbury identity (ap x pcomputation per unit instead of ann_i x n_ifactorisation) and searches it by BFGS with a Nelder-Mead fallback. The Nelder-Mead search it replaces used an absolute function tolerance that round-off could prevent it from meeting, and then ran to its 20,000-evaluation cap in every Lindstrom-Bates iteration: a nonlinear fit whose units’ time scales differ several-fold ran for more than ten minutes, and now takes 0.05 s. The estimates move only within the optimiser tolerance (at most ~1e-5 relative across the fingerprinted REML fits, with the REML objective at the new optimum equal to the old to 3e-10); moments fits are bit-identical. The degradation test suite runs in half the time.Accelerated degradation, Stage 2: stress-conditional predictions (<#155>, second half). A model fitted with
linksnow uses its stress-conditional path population,eta ~ N(D(z) gamma, Sigma)on the link scale, for prediction:predict_rul(x, y, Z=...)updates a new unit’s trajectory against the population of units at its stress rather than the pooled population that mixes every tested stress. The posterior is taken on the link scale, so a log-linked rate stays positive, andposterior_mean/posterior_covare reported there.induced_life(Z=...)gives the Lu-Meeker induced failure-time distribution at any stress – until now it was refused for every accelerated model. Outside the tested range the mechanism, not a curve through the pseudo failure times, carries the extrapolation. The induced distribution records the stress it was taken at (stress, serialised and shown in itsrepr).path_param_link_mean(Z)andpath_param_median(Z)expose the population at a stress, the latter on the natural scale (the median of each parameter, exactly, since each link is monotone).
A linked model requires
Zin these calls and a model withoutlinksrefuses it, each with a message naming the fix. Everything withoutlinksis bit-identical to before (66 fingerprinted outputs acrosspredict_rulandinduced_lifeon plain, nonlinear and Stage-1 accelerated models). The degradation theory page gains a section on the stress-dependent model and the how-to a worked example.Accelerated degradation, Stage 3: step-stress process models (<#155>).
WienerProcess.fitandGammaProcess.fittakeZ(one stress row per measurement, the stress applied over the interval ending there) andstress_ref, so constant-stress and step-stress accelerated tests can be fitted. Stress accelerates the process clock,AF(z) = exp(gamma'(z - z_ref)): the process runs on the operational timetau(t) = integral of AF(z(s)) ds(Whitmore & Schenkelberg, 1997), the fitted process parameters are the reference-stress values and the newgammacoefficients how strongly stress speeds degradation up (z = 1/Tgives Arrhenius).ff,sf,df,hf,Hf,qf,mean,randomandpredict_rultakeZ: a single stress row, or aStepSchedulefor a stress that changes over time. The life under any profile is closed form,F(t) = F0(tau(t)), for both processes; forpredict_rulthe schedule starts now.acceleration_factor(Z)andis_accelerated;gammaandstress_refare serialised and shown in therepr.For the Wiener process stress scales the diffusion along with the drift, the assumption that makes the model identifiable and the life closed form.
A stressed model requires
Zfor predictions and a stress-free one refuses it. Fits and predictions withoutZare bit-identical to before (84 fingerprinted outputs). The degradation theory page gains a section on time-varying stress and the how-to a step-stress worked example.Accelerated degradation, Stage 3: step-stress general-path models (<#155>).
DegradationAnalysis.fit(..., Z=Z, acceleration="clock", stress_ref=...)lets a unit’s stress change during its test. Stress speeds up the clock of every unit’s path,AF(z) = exp(gamma'(z - z_ref)): the path is the ordinary path model on the reference-stress time the unit has aged (the cumulative-exposure model, Nelson 1980, which for these paths is also the rate-based model – damage carries over at a step).Zhas one row per measurement, the stress over the interval ending there.population_method="moments"estimatesgammaby profile least squares from the units whose stress steps (and refuses data with no steps);"reml"fits the mixed model – the FOCE profile likelihood ofgamma– which also identifies it from units at different constant stresses.The path parameters, their population and the pseudo failure times are on the reference-stress clock, the life distribution is fitted to those reference-stress lifetimes, and
sf,ff,df,hf,Hf,qf,meanandrandomtakeZas a stress row or aStepSchedule:F(t) = F0(tau(t)).gamma,stress_ref,acceleration_factor(Z);path(t, unit)andplotfollow each unit’s own stress history. Serialised and shown in therepr.Prediction for a monitored unit on its own clock:
predict_rul,predict_failure_timeandpredict_remaining_lifetake the unit’s stress history asZ(one row per measurement, or one row) and the planned stress from its last measurement asZ_future(a row or aStepSchedulestarting now; by default the last stress is held). The posterior is taken against the reference-stress population and each draw’s failure time is mapped back to calendar time along the history and the plan.induced_life(Z=...)under a stress row or a profile, and two-stage bootstrap boundscb(..., Z=..., method="bootstrap")that resample units with their stress histories and re-estimate the clock on every refit. The analytic correction andlife_parameter_covarianceare not derived for a clock model (its pseudo failure times also depend on the estimated clock) and say so, pointing to the bootstrap.linksandpath="best"cannot be combined with the clock.
The clock leaves every existing fit bit-identical (118 fingerprinted outputs of plain, REML, nonlinear,
path="best", Stage-1 andlinksmodels, and the 84 process-model outputs); the only change is that the error for aZthat varies within a unit now points toacceleration="clock". The process models’ clock moved to a shared module unchanged, and the REML step gained a Woodbury-identity variant used by the clock fit. The theory page gains a section on the accelerated clock and the how-to a step-stress worked example, including remaining life under two stress plans.
v0.20.0 (23 September 2026)
Offset moments are exact.
ParametricFitter._momentwith an offset – what the method-of-moments fit andParametric.varuse – computed \(E[(\gamma + X)^n]\) by integrating the shifted density to infinity withquad, even for distributions whose moments have closed forms. It now takes the binomial expansion of the un-offset raw moments, asParametric.momentalready did: no quadrature for closed-form distributions, and noIntegrationWarningon machines where the shifted integral hitquad’s roundoff limit (which failed the warnings-as-errors documentation build for this release).New distribution: Hypoexponential. The sum of independent Exponential stages with distinct rates (the generalised Erlang), which is the lifetime of a load-sharing group or a warm/hot standby system – anything that passes through several memoryless stages in series.
Hypoexponential.from_params([r1, r2, ...])takes any number of stage rates and returns an ordinaryParametricmodel with that many parameters (lambda_1 ... lambda_m), sosf,ff,df,hf,Hf,qf(bisection between exact exponential brackets),mean,var,moment,entropy,random(one exponential draw per stage), offsets, limited-failure and zero-inflated variants andto_dict/from_dictall come with it. The distribution functions also take the rates directly,Hypoexponential.sf(x, r1, r2, ...). Rates must be strictly positive and distinct: the partial-fraction coefficients blow up with alternating signs as two rates approach, so near-equal rates raise a clear error pointing atGamma(equal rates are the Erlang). There is nofit; construct it from known stage rates.To make a variable-parameter-count distribution deserialisable,
ParametricFittergained a_for_paramshook (returnsselffor every fixed-arity distribution) thatParametric.from_dictconsults, so the restored model reports the rightk.Accelerated degradation, Stage 2: stress-dependent path parameters (<#155>, first half).
DegradationAnalysis.fittakeslinksalongsideZto model the degradation mechanism against stress rather than only the pseudo failure times: the path parameters named inlinksdepend on the unit’s stress on an"identity"or"log"link (a log-linked rate withZ = 1/Tis the Arrhenius relationship), the others are common, and a per-unit random effect sits on top –eta_i = D(z_i) gamma + u_iwithu_i ~ MVN(0, Sigma).gammaandSigmaare estimated by the same two-stage (Lu-Meeker) or REML route as the plain population and stored aspath_param_fixed(labelled bypath_param_fixed_names) andpath_param_link_cov; the fitted model round-trips throughto_dict/from_dictand shows the fixed effects in itsrepr. The life model is still the Stage-1 regression on the pseudo failure times, so every existing prediction method is unchanged; the stress-conditional prior forpredict_rulandinduced_lifeis the second half.Under the hood a
LinkedPathModelpresents any path model on its link scale, so the per-unit fits and the FOCE linearisation apply unchanged, and the REML routines take an optional fixed-effects design (a_mat_list/d_mat_list). Withoutlinksthe pipeline is bit-identical to before (verified by fingerprinting 55 numeric outputs across the moments, REML, nonlinear-REML, best-path, Stage-1 ADT and bootstrap surfaces).API reference completed for the remaining public surfaces (<#141>). New autodoc pages for every distribution that had none: the discrete lifetimes (Geometric, Poisson, Binomial, Negative Binomial, Beta-Geometric, discrete Weibull and
Discretize), the per-demand and degenerate models (Bernoulli, FixedEventProbability, ExactEventTime, InstantlyOccurs/NeverOccurs), the continuous stragglers (Beta, Rayleigh, GumbelLEV) and the Royston-Parmar flexible parametric model. The competing-risks subpackage – absent from the API tree entirely – has a page (Aalen-JohansenCompetingRisks,ParametricCompetingRisks,FineGray,CompetingRisksProportionalHazards), as do model persistence (from_dict/from_json) and the recurrent trend tests (laplace,mil_hdbk_189c).fit_bestgained its first docstring and joined the comparison-and-validation page, and the regression pages now document the fitted-model classes (ParametricRegressionModel,AdditiveHazardsModel).Fixed along the way: every recurrent-events API page still targeted the
ARA_-style shadow classes that thesingleton_fitterrefactor removed, so their method documentation had silently dropped out of clean builds.Third duplicate-code consolidation. Another sweep for repeated definitions, this time at the small end (exact duplicates the earlier sweeps’ size thresholds skipped, plus inline fragments):
Every
from_dictopened with the same three-line “wrong dict” guard, written out 21 times with hand-composed messages. They now callrequire_model_tag(surpyval.serialisation), which raises the same “Must create … from a <Tag> dict”ValueErrorwith the model tag always present. One message changed wording:FrailtyModel.from_dictsaid “from its own dict” and now names the tag like every other model.The five models that persist covariate metadata (feature names, formula, formula terms – <#244>) carried the same
to_dict/from_dictblocks; they now callserialise_covariate_meta/restore_covariate_metainregression_data.The
exp(beta'Z)covariate link was still restated in six places after <#295> introducedLogLinearPhi: the AFT fitter’s private copy, PO’s methods, the AFT time-varying-covariate fit’s local class, PH’s lambdas and the deserialiser’s lambda. All now useLogLinearPhi, whose two historical serialisation names (PH’se^vs AFT/PO’sexp) are class constants. The PH constructor’s phi signature check now compares parameter names and kinds instead of the signature’s string form, so the annotated shared function passes.The optimiser objective every regression
fitbuilt inline is nowmake_objectivein the fit skeleton, and the sharedx - gammaprobability-plot transform lives onHazardIdentitiesMixin.Small orphans:
_check_has_datamoved toLikelihoodInferenceMixin; the NHPP baselines’ identical all-onesparameter_initialiserbecame theIntensityModeldefault; the two competing-risksfit_from_dfhelpers shareoptional_columninutils.
Verified by fingerprinting 101 numeric outputs across the touched surfaces on both sides of the change: all identical except the one reworded frailty message. The heavier parameterised rewrites found in the same sweep are filed as <#350>, <#351> and <#352>.
Type-hint ratchet: finished. Every function in the package is annotated – 1750 of 1750 defs – and the per-module ratchet list is gone:
disallow_untyped_defsnow holds package-wide, with only the test suite and the alpha tree exempt (they are exercised, not typed, but are still checked against the package’s annotations). <#143> can close.The final pass covered the remaining seven areas in sequence – the regression fitters and
cox_ph, theutilswrangling surface and theautograd_gamma_compatshim, competing risks, the whole recurrent package, degradation, the copulas and thebeta.mlforest. The conventions are the ones the earlier passes established:Numeric/Boxablewherever autograd differentiates,npt.ArrayLikeat user entry points and never in arithmetic, declaration blocks for fit-populated model attributes, andTYPE_CHECKINGcontract stubs where a mixin calls methods its host supplies.Typing the bodies kept finding things, as it has all along:
SemiParametricRegressionModel.neg_ll/jacwere declared as a float and an array; they hold the fit’s closures.Parametric.randomis documented and annotated to return an array, but returns xcnt-format(x, c, n, t)arrays for limited-failure-population and zero-inflated models – now annotated and documented honestly.singleton_fitteris typed(cls: type[T]) -> T, so the checker finally knows every fitter singleton is an instance; onetype(...)-indirection call site simplified away.NonParametricCounting.varis honestly Optional (simulated MCFs carry no variance), and three error paths that crashed with unpacking/attribute errors now raise namedValueErrors:BuckleyJamesModel.bootstrap_ciand destructive-degradation bootstrap bounds on models carrying no fit data, andmcf_cbon a simulated MCF.The recurrent intensity contract moved off the documented
ArrayLiketrap onto the honestBoxableunion – those functions are differentiated by autograd in the NHPP likelihoods.A handful of genuine signature divergences (the covariate-extended simulation methods, the named-single-parameter intensity and copula families against their variadic base contracts) are documented at the definition with targeted ignores rather than silently widened.
Type-hint ratchet: the regression package’s shared plumbing and its two model classes. Coverage moves to 1041/1747 (60%), tracked in <#143>. Typed in dependency order – the base layer before the fitters that build on it, the same sequencing the parametric package used – so the coming fitter pass starts from typed call boundaries instead of
Anyflowing in.Eight modules join the ratchet:
_fit_skeleton(the fitting spine every family imports –LogLinearPhi,HazardIdentitiesMixin, the prepare/assemble pair, both optimiser ladders andmirror_distribution),_likelihood(the shared censoring- and truncation-aware negative log-likelihood),_bounds,tvc_schedule(theStepSchedulemachinery), theparametric_regression_modelandsemi_parametric_regression_modelclasses, and thePH/AHfactory__init__s.The conventions carry over from the parametric package:
Numeric/Boxablefor the distribution-function surface (HazardIdentitiesMixinandLogLinearPhi.phiare differentiated under autograd),npt.ArrayLikeat the user entry points, and aTYPE_CHECKINGcontract block onHazardIdentitiesMixindeclaring theHf/hfits identities call – the host class supplies them, and one that forgets still gets theAttributeErrorthat names it.Three annotations followed the code rather than the reverse:
prepare_regression_fit’sphi_bounds/phi_param_mapreally are callables or static values (both branches are live);_safe_evalin the schedule expression interpreter returnsfloat | boolbecause comparisons are values in that grammar; and_ic_countsmatches thetuple[int, int]its mixin supertype declares. No behaviour changed – annotations are erased at runtime, and the only body edits are local renames where a variable was reused with a second type (thesegmentsaccumulation list, the expression interpreter’s comparison operator).The structural duplicates: shared bases extracted where whole class bodies were copied. The body-level sweep’s deeper findings, where the fix is a base class or driver rather than a moved function.
WienerProcessModelandGammaProcessModelnow shareFirstPassageProcessModel: both reduce their failure-time distribution to one hook – the probability the process has crossed a distance by timet– and everything expressible in terms of that CDF (ff/sf, the hazard identities, the bracket-and-brentqquantile,predict_rul, serialisation) had been written out twice, verbatim. The density, mean, sampling and repr stay per class: those genuinely differ.ARA,ARIandGeneralizedRenewalshared their entire fitting spine – multi-start Nelder-Mead over[restoration, *dist params]in the unconstrained transform space – as three copies that the near-match pass scored at 0.84-0.94 similarity: already drifting. It is nowRenewalFitMixin._fit_restoration_ml; each family supplies its restoration parameter’s name, bounds and start grid.GeneralizedOneRenewalkeeps its own optimiser call deliberately: its likelihood needs onlyq > -1, so it runs under box bounds rather than a transform.Two five-way wrapper stacks collapsed to dispatchers:
ParametricRegressionModel’ssf/ff/df/hf/Hfcarried the same coerce-resolve-evaluate body five times (now_eval), andDegradationModel’s five carried the same accelerated-or-plain dispatch (now_life_fn). The named methods and their docstrings remain. The four regression fitters’__init__blocks mirrored the same six distribution attributes verbatim; that is now_fit_skeleton.mirror_distribution.Investigated and left where they are:
hpp.fitand the NHPP fitter’sfit(identical one-call wrappers over genuinely different fitting routines), the forest and tree prediction methods (already two-line delegations to each class’s dispatcher – the end state, not duplication), and the renewalfit/_refitwrappers (two-line delegations whose docstrings carry the per-family defaults).Behaviour was checked rather than assumed: 82 fingerprints – the 47 from the previous sweeps plus both process models’ full surface (fit, all distribution functions, quantiles, RUL, seeded sampling and a serialisation round trip), all four renewal fits, and the regression and degradation models’ five prediction functions – are bit-identical before and after, with the baseline verified to import the pre-change code.
A second duplication sweep, this time by function body. The first sweep matched helper names; this one normalised every function and method in the package at the AST level – identifiers abstracted, docstrings stripped – and compared the 1,483 non-trivial bodies for exact and near matches. Three findings were acted on; the rest are either deliberate parallels (the distribution API restates
hfandHfper class, each with its own closed-form docstring) or structural refactors queued with their areas (the renewal family’s triplicated fitting loop, the two process-model classes sharing verbatimpredict_rul/qf/ff).The legacy AFT fitter was dead code, and two of its methods lived on as orphans.
accelerated_failure_time/accelerated_failure_time.py– the pre-skeletonAcceleratedFailureTimeFitter, 220 lines – was imported by nothing: the package__init__re-exports fromaft_fitter, and no test touches it. It carried the only live copy of_parameter_initialiser_dist; the verbatim copies on the proportional-hazards fitter and the accelerated-life parameter-substitution fitter had no callers at all. All three are deleted along with the module.The Bernoulli / FixedEventProbability split had copied its estimation machinery wholesale. The 0.20.0 split gave each class its own verbatim
fit,from_params,entropyandrandom– the largest exact duplicate in the package. They now shareSingleProbabilityMixin(distributions/_single_probability.py, ratcheted from birth): one probability in(0, 1)fitted from 0/1 data by a weighted mean is the same estimation problem for both models, while everything distributional –sf,ff, supports and each model’s own convention docstrings – stays on the classes. Consolidating also fixed a copy-paste artifact:FixedEventProbability.from_params’s docstring said “Create a Bernoulli model”.The support-respecting Wald transform existed twice. The four-branch core of
param_cb– generalised logit for an interval-bounded parameter, log distance for one-sided, natural scale otherwise – was verbatim between the recurrent-event inference mixin and the parametric regression model, each wrapped in its own parameter lookup. It is nowutils.linalg.wald_bound_on_support; bothparam_cbs keep their lookup and delegate.Two hash matches were investigated and deliberately left:
sf_tvcand_prepare_Zare five-line and one-line wrappers over machinery that is already shared (Hf_tvcgenuinely differs per family;prepare_Zis common), and their docstrings carry per-family content worth keeping.Behaviour was checked rather than assumed: 47 fingerprints – the 37 from the previous consolidation plus
param_cbon both a bounded distribution parameter and an unbounded coefficient, and the Bernoulli / FixedEventProbabilityfit/from_params/entropyand seededrandom– are bit-identical againstdevelop, with the baseline run verified to import the pre-change code.The duplicated numeric helpers are consolidated into two new utils modules. A sweep of every module-level helper in the package found the same functions written repeatedly, three of them verbatim.
surpyval.utils.linalgnow holds the single copy of each:numerical_hessian,delta_method_se,bound_signsandlog_transformed_cbwere duplicated wholesale betweenrecurrent.inferenceandunivariate.regression._bounds– the drift-prone verbatim-copy pattern that produced <#288> – with two further hand-rolled Hessians (a different step rule,1e-5against cube-root-of-epsilon) onroyston_parmarand the frailty fitter, which now pass their step explicitly. Theinv-then-pinvfallback, written out at seven call sites, issafe_invandsafe_quadform(the latter theu'V^{-1}utest-statistic shape shared by the log-rank and Gray’s tests). The eigenvalue-surgery family from the degradation package – symmetrise,eigh, repair the spectrum, reconstruct, five sites in three flavours – ispsd_project,psd_floor,psd_precisionandpsd_root, with each call site’s own floor convention preserved as arguments.surpyval.utils.ipcwholds the censoring-distribution Kaplan-Meier (censoring_survival) and the right-continuous step lookup (step_at) that Gray’s test, Fine-Gray and the prediction metrics each carried privately – three copies of each, under three names (_G_at/_step/_g_atfor the same four lines). The copies had begun to drift: the metrics copy silently ignored count weights, consistent with its callers today but a trap for the next reuse. The shared implementation is weighted, with nonas the unweighted case.utils.validate_1djoins the wrangling helpers for the 1-D float coercion the metrics module had as_as_1d.The bodies are transplants, not rewrites, and behaviour was checked rather than assumed: 37 fingerprints across every touched path – Gray’s test (both
rho), Fine-Gray coefficients/covariance/CIF, the competing-risks PH wrapper, Brier/IBS/AUC, a three-group log-rank, Coxcheck_phand dfbeta residuals, additive-hazards fit, Royston-Parmar and frailty covariances, Crow-AMSAA/HPP standard errors and bounds, WeibullPHsf/hfbounds, the degradation fit with its corrected life covariance, REML, and the seeded induced-life sample – are bit-identical before and after. One caller-facing rename:delta_method_std_errors(therecurrent.inferencespelling) is nowdelta_method_seeverywhere, matching the regression package’s name for the identical function.Both new modules are fully annotated and under the mypy ratchet (<#143>) from birth, so the coming
utilstyping pass types each of these once instead of three times.A behavioural consistency sweep across the base distributions. The previous sweep compared annotations; this one compares what the distributions actually compute. Every identity that should hold for all of them –
sf + ff == 1,Hf == -ln sf,log_df == ln df,hf == df/R(k-1),qf(ff(x)) == x,mean == moment(1)– was evaluated across all twenty-three, and the disagreements chased down.Six discrete distributions returned nonsense below their support. Geometric, DiscreteWeibull, BetaGeometric and NegativeBinomial live on \(\{1, 2, 3, \dots\}\); Poisson and Binomial on \(\{0, 1, 2, \dots\}\). Their closed forms are algebraic and did not know where the support started, so evaluating one step below it gave
Geometric.df(0) == 0.43– a positive probability outside the distribution, growing without bound askdecreases –BetaGeometric.sf(-1) == 2.0, a survival above one thathfdivided by,DiscreteWeibull.df(0) == 0.0355+0.5468j, a complex number from a negative base to a fractional power, and NaN from the incomplete gamma and beta forms in Poisson and NegativeBinomial. The pmf now sums to one whether or not the sum starts below the support; it did not for three of them before.The fitter’s interior check kept these values out of a likelihood, which is why nothing failed, but
dfandsfare public: anyone plotting a pmf from zero got them. Each is now guarded at the first mass point. The guards clamp the input, not just the result, so the discarded branch of thenp.wherenever evaluates the invalid expression – otherwise it still computes the NaN and warns before throwing it away.Three quantile functions did not invert their own CDF.
Geometric,DiscreteWeibullandBetaGeometricansweredk + 1for authat came straight out of their ownff. \(F(k) = 1 - R(k)\) is formed by cancellation, so recoveringkfrom it lands a few ulp above the integer andceilrounds away from it. The first two snap a near-integer before the ceiling; the third compares with a relative slack in its bisection.``BetaGeometric.moment`` reported finite values for moments that do not exist. The survival decays as \(k^{-a}\), so \(E[T^m]\) converges only for \(a > m\) – the condition
meanalready applied at \(m = 1\). A truncated sum cannot see divergence; ata = 2, b = 3it returned about 25 for a second moment that is infinite. It now returnsinf, andmoment(1)uses the closed form, so it agrees withmeanexactly rather than to three decimal places.Two distributions were missing methods that are well defined.
FixedEventProbabilityhad noHf, solog_sfandlog_ff– which the base class writes in terms of it – raisedAttributeErrorinstead of returning constants. Itsdf,hf,qfandmeanremain absent deliberately:Fis flat, so the mass is an atom rather than a density.Hfis the exception, exactly as forExactEventTime, whoseHfexists while itshfdoes not.ExactEventTimeitself gainedqf,meanandmoment: a point mass has no density, but its quantile isTfor everyu, its mean isTand its m-th moment isT**m.Binomial’s support excluded two of its own outcomes.
supportis a pair of exclusive bounds –_validate_fit_inputsrejectsx <= support[0]andx >= support[1]– so a distribution declares them one step outside its first and last mass points, which is whyPoissondeclares-1andGeometricdeclares0.BinomialhadGeometric’s lower bound withPoisson’s first mass point:0, saying that zero events in n trials lies outside the distribution when its probability is 0.168 at n = 5, p = 0.3.fitandfrom_paramsset[0, n], excluding n events as well. The bounds are now(-1, n + 1).Nothing had observed this: the check lives on
OptimisedFitMixin, whichBinomialdoes not inherit – it is one of the three closed-form distributions that validate their own inputs – so the field was inert metadata that would have become live the moment anything else read it. All of its values are unchanged, which was checked: 18 fingerprints across both constructors are bit-identical.Behaviour on the support is unchanged and was checked rather than assumed: 58 fingerprints – every function over its support for all six discrete distributions, plus each one’s fitted parameters and
neg_llfitted plain and right-censored – are bit-identical before and after. The only intended change isBetaGeometric.moment. Nine new tests – 37 cases once parametrised across the distributions – cover the below-support behaviour, the pmf total, the quantile round trip, the divergence rule and the support bounds.A consistency sweep across the base distributions. With every distribution now annotated, the annotations themselves could be read as data and compared. Ten argument slots and thirteen returns disagreed across the twenty-two modules – drift from having typed them a batch at a time rather than a deliberate difference.
Most of it was cosmetic and is now uniform. The three
mpp_*transforms take annpt.NDArray: every call site in the package passes one, eight of the fifteen implementations index their argument, and probability plotting is a least-squares regression on plotting positions that is never differentiated, so the input is never an autograd box and never a scalar. Their returns stayBoxable, because the bodies delegate toqf; narrowing them would mean changing code to suit a type hint, which is the wrong way round.randomreturns annpt.NDArrayeverywhere –GeometricandDiscreteWeibullreturnedself.qf(...)straight through, and now wrap it, which is honest for the same reason in reverse:qfisBoxablebecause a fit differentiates it, and sampling never does._momistuple[float, float]throughout.One difference was a real error rather than an inconsistency.
NumericandBoxableboth excludelist, andfitandfrom_paramswere typed with them on four distributions – yet every one of those accepts a list, as their own docstring examples show (Binomial.from_params([5, 0.3])). These are the entry points a user reaches for with whatever data they have. They are nownpt.ArrayLike, which is the correct type here precisely because the value is converted withnp.asarrayon the first line rather than used in arithmetic.Binomial.from_paramsalready had it right;Bernoulli,FixedEventProbabilityandExactEventTimedid not.Eight differences remain and each is deliberate:
ExactEventTime’ssf,ff,df,hfandHfreturn the narrowernpt.NDArray, which is a stronger promise rather than a broken one – they are step functions built withnp.atleast_1dand provably return a real array – andExpoWeibull.unpack_rrreturns three values where the two-parameter distributions return two.Five tests were added to the shared-signature guard, so a future distribution cannot reintroduce any of this: the distribution functions take a
Numericand return aBoxable, parameters areBoxable, thempp_*family takes arrays,randomreturns one, and the user entry points accept array-likes. Twenty-two tests in that file now. No behaviour changed – annotations are erased at runtime, and the twonp.asarraywraps were checked to produce identical samples.Type-hint ratchet: ``univariate.parametric`` is finished. Coverage moves from 869/1760 (49%) to 995/1771 (56%), tracked in <#143>. Every module in the package – the fitters, the model, the base class and the mixture – is now under
disallow_untyped_defs.Two structural additions came out of it, both of the same kind. A
TYPE_CHECKINGblock onParametricFitternow declares the distribution functions its own methods call –csdivides twosfs,log_sfnegatesHf,randominvertsqf, and the fourll_*methods are written in terms ofhf,Hfand the log densities. The class docstring already stated that contract in prose (“a distribution needs onlyhfandHf, orsf,ffanddf”); this is the same statement in a form the checker reads, and it mirrors the blockOptimisedFitMixinalready carried for the estimation machinery. Declared rather than defined, so a distribution that forgets one still gets theAttributeErrorthat names it instead of a silently wrong inherited implementation.MixtureModel’s fitted state –data,params,w,pandloglike– is annotated where it is initialised toNone.Three annotations had to follow the code rather than the reverse, each a small fact:
probability_plot_data’sffis the failure function, not an array of values;bounds_convertreturns five things, not three; andfallback_minimize’sjacandhessare declared optional but are supplied by every caller.Where a value comes back from scipy or autograd and genuinely has no narrower type – the confidence-bound closures, the mixture’s prediction inputs – it is
Anyrather thannpt.ArrayLike. That is the same trap theNumeric/Boxablecomment inparametric_fitteralready documents:ArrayLikeadmitsstrandbytes, so arithmetic on it does not type check, and thenp.asarraythat clears the error destroys an autograd box.Behaviour is unchanged and was checked rather than assumed: four distributions fitted plain, right- and left-censored, interval censored, truncated, with a limited-failure population and with zero inflation, plus
neg_ll,aicand a two-component mixture fit – bit-identical before and after.Type-hint ratchet: the remaining eleven distributions. Coverage moves from 665/1760 (38%) to 869/1760 (49%), tracked in <#143>. Every distribution module is now under
disallow_untyped_defsexceptgeneral_log_linear’s counterpart concerns (<#345>).rayleigh,beta,beta4,gamma,gumbel,gumbel_lev,loglogistic,exponential,uniform,degenerateandexpo_weibull– 202 signatures. The bulk was mechanical, generated from each distribution’s ownparam_namesso thatxis aNumeric, a parameter is aBoxableand the return follows the method. What was not mechanical were the places the generated guess was wrong, and each of those is a small fact about the code:Rayleigh.mppandExponential.mpptreat the output ofmpp_y_transformas an array – indexing it, and passing it tonp.polyfitandnp.linalg.lstsq– while the transform is declared to return aBoxable. Wrapped at the call site rather than widening the transform, which is shared.Gamma._moment_estimateand the two_momhelpers return 2-tuples, not arrays.Exponential._closed_form_mleandUniform._closed_form_mlereturnNonewhen the closed form does not apply to the data, so they arenpt.NDArray | None.ExpoWeibull.unpack_rrreturns three values where every other distribution’s returns two.degenerate’s classes inheritDistribution, notParametricFitter, and its signatures have to match that supertype rather than the distribution convention.ExpoWeibull._gumbel_seedreadsgumb.res, which aParametriconly carries after an MLE fit – the branch that reads it is the one that asked for MLE, so it is annotated as deliberate rather than made unconditional.
Behaviour is unchanged, and checked rather than assumed: every one of the eleven distributions was fitted by MLE, MPP, MSE and MOM, and its
entropyand second moment evaluated, before and after. All 66 results are bit-identical.``Logistic`` ratcheted, and ``mgf`` made private.
Logisticwas the only distribution with a publicmgf, which read as a method the other twenty-two were missing.It is not an orphan and is not removed:
Logistic.momentdifferentiates itmtimes with autograd to get the m-th raw moment, and the results are exact –Logistic(mu=3, sigma=2) moment(1) = 3.0000000000 exact mu = 3 moment(2) = 22.1594725348 exact mu^2 + s^2 pi^2/3 moment(3) = 145.4352528131 exact mu^3 + 3 mu s^2 pi^2/3
The general closed form for a logistic raw moment needs Bernoulli numbers, so differentiating the MGF is both shorter and exact. What was wrong was its visibility: it is machinery for
moment, not part of the distribution surface. It is_mgfnow, alongside the other private helpers on distributions (_closed_form_mle,_moment_estimate,_gumbel_seed). Nothing outside the class ever referenced it.The module is now fully annotated and added to the ratchet (#143). Two annotations had to follow the code rather than the other way round:
mpp_y_transformindexesy, so it takes annpt.NDArrayrather than aNumericthat includesfloat, andunpack_rrreturns a tuple of two values, not an array – both matching howWeibullalready declares them.New tests pin the three low-order Logistic moments against the algebra rather than against another numerical method, check the variance comes out as \(\sigma^2\pi^2/3\), and assert that no distribution exposes a public
mgf.Type-hint ratchet: the accelerated-life package, plus nine modules that were already complete. Coverage across the package moves from 611/1755 (35%) to 646/1760 (37%), tracked in <#143>.
Nine modules were fully annotated but not listed under
disallow_untyped_defs, so nothing stopped them slipping back. They are listed now:fit_best,utils.recurrent_utils,utils.score,recurrent.tests,recurrent.parametric.counting_process,univariate.regression.regression_data,univariate.regression.tvc_fit,univariate.regression.frailtyanddistributions.fixed_event_probability. Onlycounting_processneeded work – four*paramsthat an AST scan counts as annotated and mypy does not.Eleven of the twelve accelerated-life modules follow, and locking them in turned up four real problems that annotations made visible:
``GeneralLogLinear``’s constructor arguments were swapped. The bounds lambda sat in the
phi_param_mapslot and the param-map lambda in thephi_boundsslot. Nothing consumed either, so it had no observable effect, but it would have bitten whoever finished the model. That module stays out of the ratchet: itsphi_param_mapandphi_boundsare callables of the covariate dimension rather than thedictandtupleLifeModeldeclares, which is why it is already excluded fromLIFE_MODELS(<#345>).``LifeModel.phi_bounds`` was annotated as a one-element tuple while every caller passes two or three. Now variadic.
Two dead branches around ``phi_init``. The fitter chose between three shapes – a
"(Z)"-only signature selected by comparingstr(inspect.signature(...)), the two-argument form, and a non-callablephi_init. All ten life models are callable with(life, Z), so only one branch could ever run.``AcceleratedLife`` deserialisation accepted a distribution it cannot fit. The guard established a
ParametricFitter, which admitsBernoulli,BinomialandExactEventTime– none of them fittable. Since the dict is untrusted input, such a name got through and failed deep inside the fitter on a missing attribute; it now raises where the mistake is.
hfis also declared inOptimisedFitMixin’sTYPE_CHECKINGblock, wheresf,ff,df,Hfandqfalready were. Its absence was invisible until a typed caller reached for it.Behaviour is unchanged throughout: the accelerated-life fit, prediction,
randomand serialisation round-trip all produce bit-identical results before and after.``Bernoulli.qf``. The quantile function, added after the rest of the distribution:
Bernoulli.qf([0.1, 0.7, 0.75, 0.99], 0.3) -> array([0., 0., 1., 1.])
It inverts \(P(X \leq x)\) – the ordinary CDF – stepping from 0 to 1 at
u = 1 - p. On the open interval it matchesBinomial.qf(u, 1, p)andscipy.stats.binom.ppfexactly. Atu = 0those answer-1, one below the support; this answers 0, the smallest outcome there is.It is deliberately not the inverse of this class’s
ff, and that follows from the survival convention rather than being an oversight.R(x) = P(X \geq x)forcesF(x) = P(X < x)if the two are to sum to one, andP(X < x)never exceeds1 - panywhere on{0, 1}– so onceupasses1 - pnoxin the support satisfiesF(x) >= u. The other discrete distributions, whoseR(k)isP(X > k), do not have this split, and the usualff(qf(u)) >= ucheck still holds for them. A test pins the difference in both directions so it stays a known consequence rather than becoming a surprise.What the definition does buy is the property worth having:
qf(U)for uniformUreproduces the distribution, which is howParametricFitter.randomsamples. Tested at 200,000 draws, and at the degenerate endsp = 0andp = 1.BREAKING: ``Bernoulli`` is now a Bernoulli distribution. It was not one.
F(x)returnedpat everyx– includingx = -100– which is a flat curve with no time axis, not a coin flip. Meanwhilemoment,entropy,randomandfitall described a genuine{0, 1}variable:E[X^m] = p, the binary entropy, draws of 0 and 1, and a fit that rejects anything else. The class was two models at once, anddf,hf,Hfandmeanwere missing because they are the four places the contradiction cannot be papered over.Bernoulliis now the coin flip the name promises.xis the outcome, so 0 and 1 are the only values accepted and anything else raises:Bernoulli.sf([0, 1], 0.3) -> array([1. , 0.3]) Bernoulli.df([0, 1], 0.3) -> array([0.7, 0.3]) Bernoulli.hf([0, 1], 0.3) -> array([0.7, 1. ]) Bernoulli.sf(37.5, 0.3) -> ValueError
The survival function is \(R(x) = P(X \geq x)\), so
R(0) = 1andR(1) = p: read as a one-shot device,pis the probability it works when demanded.df,hf,Hfandmeanare added and every internal identity now holds – the mass sums to one,h = f/R,H = -ln R, andE[X]from the mass equals bothmeanandmoment(1).moment,entropy,randomandfitare unchanged, because they already described this model.``p`` has changed direction. It was documented as the probability of failure; it is now the probability of the
1outcome, which under the survival reading is the probability of surviving. Code that coded failures as 1 now fits the survival probability and wants1 - p.log_dfis defined on the class rather than inherited. Neither base relation fits:DiscreteParametricFitterusesf(k) = h(k) R(k - 1), which assumesR(k) = P(X > k), and here the at-risk set atxisR(x)itself.The flat model is not gone. It survives unchanged as
FixedEventProbability, which until now was a second instance of the same class and is now its own. It is the two-point mixture ofInstantlyOccurs(weightp) andNeverOccurs(weight1 - p) – which is whydegenerate.pyalready described those two as its limits atp = 1andp = 0. Itsdf,hf,qfandmeanremain absent, correctly: a constantFhas no density, no invertible quantile and no time to average.Both names serialise and round-trip under their own identities, so stored models keep pointing at the model they were fitted with – but a stored
Bernoullifitted before 0.20.0 will now be read with the new semantics, and itspreinterpreted as above.binomial.pyclaimed Bernoulli was “the special casen = 1”. That was false of the old model and is now true of the mass function:Bernoulli.dfandBinomial.df(..., 1, p)agree exactly. The survival functions remain offset by one by convention, and the docstring now says so.``ExpoWeibull.moment``. It was the only continuous distribution without a public
moment, while already havingmeanandentropy.The exponentiated Weibull has a closed form – an infinite series in \(\binom{\mu-1}{i}(-1)^{i}(i+1)^{-(1+m/\beta)}\) – but it only terminates when \(\mu\) is a positive integer. For other \(\mu\) it is alternating and slow to converge, losing significance to cancellation as \(\mu\) grows. The integral is quadrature either way, so
momenttakes it directly, asentropyalready does and asmeanalready did.meannow delegates tomoment(1)rather than repeating the integral.Checked against two references with no integration in them: at \(\mu = 1\) the distribution collapses to the Weibull, whose m-th moment is \(\alpha^{m}\Gamma(1 + m/\beta)\) exactly; and for integer \(\mu\) the series terminates and can be summed. Both agree to about 1e-14. The exponentiated-exponential case (\(\alpha = \beta = 1\)) is also pinned against the harmonic number \(H_{\mu}\), which is its mean.
ExpoWeibulljoins themomentcomparison against quantile-bounded numerical integration intest_distributions_math.py, which had excluded it by name. That check is not circular despite both sides integrating: the reference integrates between quantiles with breakpoints,momentintegrates from zero to infinity.This does not change fitting.
ParametricFitter._momentalready had a quadrature fallback for distributions without amoment, sohow="MOM"worked forExpoWeibullbefore this and still does. What was missing was the public method.Fixed: ``Binomial.log_df`` returned the wrong mass, and ``ExactEventTime`` answered ``df`` and ``hf`` with ``inf``.
Two consequences of continuous-distribution assumptions reaching distributions that are not continuous.
ParametricFitter.log_dfislog(hf) - Hf, which encodes the continuous identity \(f = h R(x)\). On the integers the mass atkis \(P(T = k) = h(k) R(k - 1)\) – the hazard there times the survival to just before it. The two differ by a factor \(R(k)/R(k-1)\), which is not a rounding difference:Binomial.log_df(3, 10, 0.3) -> -1.887 (was) log(Binomial.df(3, 10, 0.3)) -> -1.321
Across
k = 1..7the returned mass ran from 0.88 of the truth down to 0.15. Five of the six discrete distributions overridelog_dfwith a closed-form log-pmf and were unaffected; Binomial did not, and reached the continuous identity.DiscreteParametricFitternow supplies the discrete relation, so Binomial is correct and any future discrete distribution inherits the right one. The class already documented the convention –hfisP(T = k) / R(k - 1)– it simply had nolog_dfto match it.The bug was latent rather than live:
Binomial.fitis analytic (pis the observed mean over the trial count) and never evaluates a log-density, so no fit was affected.Binomial.log_dfis public, though, and generic code that calls it got the wrong numbers.Separately,
ExactEventTimeis a point mass, so its density is a Dirac delta: zero everywhere, infinite at one point, integrating to one. There is no function ofxthat represents it.dfreturnedinfatTand 0 elsewhere, which integrates toinfrather than 1;hfreturnedinfatTand at every x after it; and the inheritedlog_dfcomputedlog(inf) - infand returnednan. All three now raiseNotImplementedErrorexplaining why and pointing at the functions that are defined. Aninfpropagates into a plot, a likelihood or a mixture weight and surfaces far from its cause; a raise stops at the call site.Bernoullialready omitteddf,hfandHffor the same reason.ExactEventTime.Hfis kept and is unchanged in value – it is \(-\log R(x)\), stepping from 0 to infinity atT, which is well defined. It had been written as an alias forhf, which happened to take the same two values; it is now written as itself.sf,ff,qfand fitting are untouched.New tests cover the discrete mass identity for all six distributions, Binomial’s log-pmf against scipy, that the discrete hazard is a probability (a continuous-convention hazard can exceed one, which is how the mix-up shows itself), that
Hfaccumulates as \(-\sum \log(1 - h)\) rather than \(\sum h\), and the degenerate refusals alongside proof that fitting and serialisation still work.``Beta.mpp`` and ``Beta4.mpp`` removed as unreachable. Both bodies were a single
raise NotImplementedError, and neither could ever run. Refusing probability plotting is declarative –supports_mpp = False, checked infitbefore the fitter is dispatched – and both distributions already set it, so the guard raised aValueErrornaming the distribution and the alternatives three frames before the method was reachable.Deleting them changes no behaviour.
Beta,Beta4,GammaandExpoWeibullall still refusehow="MPP"from the same guard, with the same message.mppis now defined only byExponentialandRayleigh, which is where the hook means something: absence ofmppsends a distribution to the generic plotting path, so the method is an override for a closed form, never a way to decline.Two invariants in
test_shared_signatures.pykeep the two mechanisms from drifting back together: no distribution may declaresupports_mpp = Falseand definemppas well, and every distribution that refuses must refuse through the shared guard rather than an exception of its own. The second covers nine distributions and is scoped to those whosefittakes ahowat all –Bernoulli,BinomialandExactEventTimeoverridefitwith a narrow signature that has none, so asking them for MPP is aTypeErrorfrom argument binding. That is the separatefitdivergence, still open.``cs`` is inherited rather than restated on every distribution, and Gamma’s ``cs`` documentation no longer describes the exponential. Twelve distributions defined a conditional survival function. Eleven of the twelve had the same body as
ParametricFitter.cs, differing only in spelling the parameters out instead of taking*params:return self.sf(x + X, alpha, beta) / self.sf(X, alpha, beta)
The duplication had already rotted.
Gamma.cscarried\[R(x) = e^{-\lambda x}\]which is the exponential survival function – copy-pasted from
exponential.py, where both methods sat at line 136. The body computed the ratio correctly, so the code was right and the documentation above it described a different distribution. Gamma is not memoryless and its conditional survival is not its survival.The eleven pass-through overrides are removed (395 lines), and
ParametricFitter.cs– which had no docstring at all, socswas undocumented anywhere the override was absent – now carries the definition, the parameter descriptions and a worked example. The wrong Gamma formula goes with the override it lived on, and Gamma inherits the correct generic statement.Exponential.csis kept. The exponential is memoryless, so \(R(x, X) = R(x)\), which is oneexprather than two and a division, and avoids the cancellation the ratio suffers far into the tail.Ten of the removed docstrings carried doctested examples, and those were the only per-distribution numerical check on
cs. Their values are preserved insurpyval/tests/univariate/parametric/test_conditional_survival.py, alongside tests that each distribution’scsequals the survival ratio (which is what checks Exponential’s shortcut against the long way), thatcs(0, X) == 1, that the exponential is memoryless for any conditioning time, and that the discrete distributions reach a working inheritedcs.BREAKING: shared methods now have one signature across every distribution. A distribution is reached through a
ParametricFitterreference all over the package –fit_bestiterates a list of them,DiscretizeandMixtureModelwrap one, the regression fitters hold one asself.dist– so code written against that reference has to work for every member. Three shared methods disagreed about what their leading argument was called, which made a keyword call correct for a subset and aTypeErrorfor the rest:Weibull.qf(p=0.5, alpha=10, beta=2) worked Poisson.qf(p=0.5, mu=3) TypeError Poisson.qf(u=0.5, mu=3) worked Weibull.moment(n=2, alpha=10, beta=2) worked Poisson.moment(n=2, mu=3) TypeError
This is the defect that made the narrow
from_paramsoverrides onBernoulli,BinomialandExactEventTimeworth fixing earlier in this release, applied to the rest of the surface.qf’s first argument isuin all 22 implementations. It waspin 14,uin 7 andqinBinomial.pcannot be the shared name because it is an actual parameter ofBernoulli,Binomial,GeometricandNegativeBinomial, andqis one ofDiscreteWeibull’s – which is why the two obvious choices had been avoided piecemeal in the first place.moment’s first argument ismin all 21. It wasnin 13, andnisBinomial’s trial count.mpp_x_transformtakesxalone in all 15. Eleven of them also took agammathey subtracted, and the other four did not. No caller ever passed it: the MPP fitter subtracts the offset fromxbefore calling (fitters/mpp.py), so a caller that did pass it would have subtracted the offset twice. Removed rather than added to the other four.
Positional calls – which is what every docstring example, every call inside the package, and every notebook uses – are unaffected. No keyword call to any of the three exists in the package, its tests or its documentation. There is no deprecation shim: keeping the old name as an alias would preserve exactly the ambiguity the change removes.
momentis also now typedm: intuniformly, and nine docstrings that promised “integer or numpy array of integers” are narrowed to “integer”. Only six of the twenty implementations actually accepted an array of orders; the rest raised, because they delegate toscipy.stats:LogNormal.moment(np.array([1, 2]), 3., 4.) -> [5.99e+04, 3.19e+16] Normal.moment(np.array([1, 2]), 3., 4.) -> ValueError
surpyval/tests/univariate/parametric/test_shared_signatures.pyreads the signatures rather than asserting a list of names, so a distribution added later is covered without touching the test, and an open-ended guard fails on any method implemented by five or more distributions whose leading data argument disagrees. Parameter names are excluded from that guard:Weibull.mean(alpha, beta)againstPoisson.mean(mu)is not a divergence, it is what the distributions are.API reference pages for the surfaces that only had narrative docs. The multivariate copulas and the beta survival tree and forest had no autodoc coverage at all, and the degradation page stopped at the path models. Closes the second half of #141.
New:
surpyval.multivariate(theCopulabase, the five copula classes,CopulaModelandMultivariateSurpyvalData) andsurpyval.beta(RandomSurvivalForest,SurvivalTreeand the node classes a serialised tree is built from). The degradation page gains the Wiener and gamma stochastic-process models,ProcessRULand destructive degradation._boundsandpopulationare deliberately left out: neither is exported from the package’s__init__, so both are internal rather than API.Two surfaces the issue did not name but which fit its description also had no autodoc, and now have pages: shared frailty models and Buckley-James. And
surpyval.regressionhad two headings – “Accelerated Time Models” and “Accelerated Life Models” – with nothing underneath them, rendering as empty sections while the content sat inregression/parametric; that page is reorganised into semi-parametric, parametric and correlated-observations.Three
automethoddirectives on the NHPP regression page pointed atcif,iifandinv_cifon the fitter, where they do not exist – they are on the model the fit returns. Those were three of the 18 warnings standing between the build and-W. All 138 autodoc targets across the documentation now resolve.The estimation machinery moved off the distribution base class.
ParametricFitter.fittakes 18 named arguments –how,offset,zi,lfp,fixed, the truncation and interval bounds – and three distributions cannot honour any of them.Bernoulli,BinomialandExactEventTimeestimate their parameters in closed form and accept onlyxand at mostc,nandt. They overrodefitwith a narrower signature, which is a real divergence and not a typing nicety:Bernoulli.fit([0, 1], c=[0, 0]) TypeError Bernoulli.fit([0, 1], how="MLE") TypeError Binomial.fit([1, 2], n_trials=3, how="MLE") TypeError
Code written against a
ParametricFittertherefore broke on exactly those three, and nothing said so until it ran.fitand the twelve methods it needs now live on a newOptimisedFitMixin, which the 21 distributions that have them inherit alongsideParametricFitter. Nothing was removed and no behaviour changed: every distribution is still aParametricFitter, which is what theisinstancechecks in the model, mixture, regression, frailty and renewal code test, and what carries the distribution functions, the likelihood andfrom_params.The point of the split is that the wrong thing is now unwriteable rather than merely undocumented.
fit_best’s candidate list is typedlist[OptimisedFitMixin], so adding one of the three to it is a type error instead of a runtime one.Annotate a parameter
OptimisedFitMixinwhen it must be fittable by a chosen method, andParametricFitterwhen only the distribution functions are needed.BREAKING:
Bernoulli.from_paramsandExactEventTime.from_paramsrenamed their first argument toparams. It waspandTrespectively, while the base calls itparams, so positional calls worked and keyword calls raised:Bernoulli.from_params(0.5) OK Bernoulli.from_params(params=0.5) TypeError
That is the shape of bug a test suite never catches, because every internal call and every docstring example passes positionally.
Bernoulli’s was worse than a rename. The base’s
pis the proportion that never fails, sop=0.5meant the never-fails fraction on twenty-four distributions and the event probability on Bernoulli – the same keyword, sibling classes, unrelated meanings, and no error either way.Bernoulli.from_params(p=...)andExactEventTime.from_params(T=...)now raiseTypeError. Positional calls are unaffected, and no call changes meaning silently:paramshas no default, so the old keyword forms fail loudly rather than being reinterpreted.All three also accept
gamma,pandf0now, and reject them with aValueErrornaming the distribution. Accepting-and-rejecting rather than omitting is what makes the signatures match the base, so these can be called through aParametricFitterreference at all – and it removes the last two# type: ignore[override]markers.Every distribution now exports its own type. The 17 that read
Weibull: ParametricFitter = Weibull_("Weibull")erased the concrete class, and since the base declares none ofsf,ff,df,hf,Hf,qformean, the example in each distribution’s own docstring did not type check for anyone whose checker honourspy.typed:Weibull.sf(x, 3, 4) error: "ParametricFitter" has no attribute "sf"
The annotation cannot simply be dropped: the regression subpackages and
fit_bestimport these names, and without an explicit type mypy cannot resolve them through that cycle. They name the concrete class instead.``Normal`` and ``Gumbel`` ignored their own documented default.
ParametricFitterdocuments the initialiser signature as(self, x, c=None, n=None, t=None, offset=False), butNormaltested2 in cand indexedx[c != -1], andGumbeltested(2 in c) or (-1 in c), before either defaultedc. Calling either as documented raisedTypeError: argument of type 'NoneType' is not iterable. Every caller inside the package passescandn, which is why it went unnoticed;GumbelLEVis unaffected because it forwardsctofitwithout inspecting it. A sweep of all nineteen distributions found these two and no others.BREAKING: ``_parameter_initialiser`` takes a ``SurpyvalData``. The signature was
(self, x, c=None, n=None, t=None, offset=False), and every one of the 21 implementations spent its opening lines re-establishing conventions that had already been established – inconsistently, and in some cases wrongly.Normaldefaultedcandn;Gumbelguardedcwithis not None;Betatested(c is not None) and (c == 0).all();Beta4tested bothcandn;LogLogisticran a wholexcnt_handlerround trip in its offset branch. Two of those checks were absent until this release and raisedTypeErrorfor the documented call.None of it was ever needed. The one production caller,
_initial_guess, is reached fromfit_from_surpyval_data, which is handed aSurpyvalData– an object whose entire purpose is to guarantee thatx,c,nandtare present, validated and in xcnt form – and destructured it into loose arrays on the first line of its body. The convention was rebuilt three layers below the object that had already established it.The signature is now
(self, data: SurpyvalData, offset: bool = False).offsetstays a separate argument because it describes the model being requested, not the data._initial_guessand_fit_numericallytake the object rather than loose arrays for the same reason. Seven defaulting checks are gone, along with 63 optional data parameters (27 of them explicitly annotated| None), and the initialisers that used to round-trip their arrays back throughfit(re-runningxcnt_handlerand rebuilding the object the caller already held) now callfit_from_surpyval_datadirectly.tis not passed to the initialisers, and never was: no caller has ever supplied it._initial_guessimputes interval- and left-censored points to midpoints before seeding, which can put an observation at or before its own left-truncation bound – dataxcnt_handlerrejects outright (#260) – so the working copy it builds is deliberately untruncated. That is what every initialiser has always received; it is now explicit rather than accidental.This is a breaking change for anyone who has written their own distribution class. There is no shim: a bare array now fails at the first attribute access rather than being half-accepted. Every one of the 38 seeds – each distribution, plain and offset – is identical before and after.
Every ``_parameter_initialiser`` now returns the same thing. The initial-guess seed a distribution hands the optimiser came back in four different containers across the 21 implementations: a tuple in nine, a numpy array in six, a Python list in one, a fitted model’s
.paramsin five – and a bare scalar inRayleigh. Two files disagreed with themselves:exponentialreturned a tuple in its offset branch and an array in the other,rayleigha tuple and a scalar.It worked because the one caller,
_initial_guess, doesnp.array(init), which flattens tuple, list and array alike. It stopped working at the scalar, becausenp.arrayof a scalar is 0-dimensional rather than length-1, and thelfpandzipaths concatenate onto the seed.All 28 return statements now construct a 1-D float array explicitly, so the shape is decided where the values are known rather than inferred downstream, and a 0-dimensional seed is no longer expressible. No seed changed: all 38 – every distribution, plain and offset – were compared before and after and are identical.
The seed itself is unchanged in layout, and it is flat rather than nested:
[gamma]when an offset is requested, then thekdistribution parameters, then[p]for a limited failure population and[f0]for zero inflation, appended by the caller. The arity therefore depends on bothkand the structural flags.Limited-failure and zero-inflated Rayleigh models could not be fit.
Rayleigh.fit(x, lfp=True)andRayleigh.fit(x, zi=True)both raisedValueError: zero-dimensional arrays cannot be concatenated.Rayleigh is the only single-parameter distribution here, and its
_parameter_initialiserreturned the sigma seed as a bare scalar rather than a sequence.np.array(init)in_initial_guessthen produced a 0-dimensional array instead of a length-1 one, and thelfpandzipaths append theirpandf0seeds withnp.concatenate, which a 0-d array cannot take. The seed is now a one-tuple. Plain and offset fits are unchanged.Found by surveying every
_parameter_initialiserin the library after the type-hint work turned up three different return shapes; a sweep of all fourteen continuous distributions across both paths confirmed Rayleigh was the only one affected.Type-hint coverage is now enforced, for twenty-one modules (#143).
surpyval.distribution,surpyval.serialisation,surpyval.metrics,surpyval.univariate.information_criteria,surpyval.datasets, the Weibull, the Normal, the LogNormal, the eight discrete distributions,CustomDistributionandExactEventTime, and all ofsurpyval.univariate.nonparametric,surpyval.recurrent.nonparametricandsurpyval.univariate.regression.frailtyhavedisallow_untyped_defsset inpyproject.toml, so an unannotated function in any of them is a mypy error. That covers the abstract base classes every model inherits from, the Kaplan-Meier, Nelson-Aalen, Fleming-Harrington and Turnbull estimators, the log-rank test, the plotting positions, the non-parametric MCF, the shared-frailty fitter, the bundled datasets and thirteen of the 25 parametric distributions.LogNormal.momentis annotatedn: NumericwhereNormal’s isn: int, and the difference is real rather than an oversight. Both docstrings promise “integer or numpy array of integers”. LogNormal’s closed form is vectorised and delivers that;Normal,GumbelLEVandLogLogisticdelegate toscipy.stats, which raisesValueError: The truth value of an array ... is ambiguouson an array of orders. The annotations now say which is which; the three docstrings that overpromise are not yet corrected.CustomDistributionneeded restructuring rather than only annotating. It assigned its distribution functions onto the instance –self.Hf = fun, then lambdas forhf,sf,ffanddf– which stopped being possible onceOptimisedFitMixindeclared those names for its own use, because a subclass inherits the declarations and assigning to an inherited method is an error. The function is stored as_funand the five are real methods delegating to it. Equivalent by construction: the oldself.Hf = funwas an unbound instance attribute, soself.Hf(x, *params)calledfun(x, *params)either way. The autograd-derivedhfanddfwere checked numerically against the previous implementation, gradients included.Its
_parameter_initialiserreturns a list, where Weibull returns a tuple and the discrete distributions return an array – three shapes for one contract the base never pinned down. Noted in the signatures rather than unified, since every caller coerces.handle_xicngained@overloaddeclarations as part of this. Its return shape is decided byas_recurrent_data, but its signature only said “one or the other”, so all nine callers taking the default were handed a union to narrow themselves. The overloads say which argument decides, once, for all seventeen call sites.The package ships
py.typed, which tells a user’s type checker that the annotations are there to be trusted, and mypy already ran in CI – but nothing required an annotation to exist, so mypy checked only the ones that happened to be written. That madepy.typeda promise the package kept unevenly.This is deliberately a ratchet rather than a target. A module is added to the list once it is clean, and from then on it cannot regress; the remaining ~1350 unannotated functions do not have to be finished first for the enforced part to start holding.
Turning it on immediately found something.
SerialisableMixin.to_jsonandfrom_jsoncallself.to_dict()andcls.from_dict(), which the mixin never declares – every class using it supplies them, but that contract existed only in the docstring, and mypy skips the bodies of unannotated functions, so the calls had never been checked. They are now declared underTYPE_CHECKING: a real stub raisingNotImplementedErrorwould read better, but it would be inherited, andcopula_modeldecides whether a margin is serialisable withhasattr(m, "to_dict")– which an inherited stub would answer True for every time.The non-parametric package turned up a second one. Its
ESTIMATOR_FUNCStable was built fromnonp.nelson_aalenand its two siblings, each of which shares its name with the submodule that defines it – so the attribute is the function only after the package__init__has bound it over the submodule, and that table is built while the__init__is still running. It worked because the__init__happens to import the estimators first, which is a load order rather than a guarantee. mypy resolved the names to the modules and reported the table as not callable. The three are now imported from the modules that define them; the other uses of the package namespace in that file are inside functions, so they resolve after initialisation and are left alone.Annotating a function makes mypy check its body, and that found a real inconsistency the ratchet had been hiding behind an over-wide annotation.
turnbull,rank_adjustandNonParametricCounting.from_xrddeclare array-like parameters and then index, slice and divide them directly – which array-like does not support, because it also coversstr,bytesand scalars. Each now takes its arguments as arrays before using them as arrays, so the signature and the body agree._logrank_z_vlikewise declaredcandnas arrays while handlingNonefor both internally, andNonParametricCounting.fitdeclaredwindowsas array-like when it is the{item: [(start, end), ...]}dictionary its own docstring describes.success_runwas the same shape of problem in its argument handling: it testedconfidenceandalphafor truthiness, so passing both with either set to zero skipped the “only one of” raise, andconfidence=0fell through every branch and leftalphaunset. Both are now tested againstNone.The distributions needed a vocabulary before any of them could be annotated, and
parametric_fitternow defines it. A distribution function deals in two kinds of value, and only one of them can be an autograd box:Numericis what the function is evaluated at (times, or probabilities forqf), always real data;Boxableis a parameter, or anything computed from one. Maximum likelihood differentiates these functions, so during a fit autograd substitutes anArrayBoxfor each parameter to carry the derivative. A box is neither a float nor an ndarray, which is whyBoxableis not narrowed to a numpy type – and why the “array-like in, array out” convention used elsewhere in the package must not be applied here.np.asarrayon a box does not reject it. It wraps the box in a 0-d object-dtype array, which still computes the right value, because object arrays dispatch arithmetic back to the box. Only the derivative is damaged, and how depends on the arithmetic: an operation whose backward pass needs a ufunc the box does not implement raises aTypeError, but a plain product silently returns a zero gradient for that parameter. A zero gradient is not an error to an optimiser – it means “this parameter does not affect the likelihood” – so the fit leaves the parameter at its initial guess and reports success.BoxablenamesArrayBoxrather than beingAny, which is what makes the parameter positions checkable rather than merely annotated: underAny,Weibull.Hf(1.0, "not a number", 4.0)was accepted in silence. autograd ships no type information, sostubs/autograd/numpy/numpy_boxes.pyidescribes the one type of its that appears in surpyval’s own signatures. Supplying any stub for a package makes mypy consider the whole package described, so the__getattr__stubs beside it keep the rest of autograd as untyped as it was.Weibulland the eight discrete distributions –Bernoulli,BetaGeometric,Binomial,DiscreteWeibull,Geometric,NegativeBinomial,Poissonand theDiscretizewrapper – are annotated against that vocabulary and are on the enforced list.Checking their bodies turned up three things about the base class.
Weibullwas exported asWeibull: ParametricFitter, which erases the concrete type; since the base declares none ofsf,ff,df,hf,Hf,qformean, and the package shipspy.typed, the example insf’s own docstring did not type check for a user:Weibull.sf(x, 3, 4) error: "ParametricFitter" has no attribute "sf"
It is now exported as
Weibull_. Sixteen other distributions carry the same erasure and are corrected as each is annotated.Discretizehits the same gap from the other side: it must hold the distribution it wraps asAny, because every delegation would otherwise be an error against the declared type.Bernoulli.fitandBinomial.fitdo not honourParametricFitter.fit, and the divergence is real rather than a typing artefact –Bernoulli.fit(x, c=...)andBinomial.fit(x, n_trials=k, how=...)raiseTypeError. Both have closed-form maximum likelihood estimates and support neither censoring nor an alternative estimation method. Generic code written againstParametricFitterwill fail on them. This is recorded at each site rather than changed here._parameter_initialiseris also inconsistent across the base’s implementations:Weibullreturns a tuple, the discrete distributions return an array. Callers coerce either, so nothing is broken, but the signatures now say which is which.The lint job had been failing on a single over-long line. The
:class:cross-reference added toRandomSurvivalForest’s docstring while clearing the documentation build warnings pushed the line to 80 characters. flake8 runs before mypy in the lint job, so from that commit on mypy was skipped rather than run, and the type errors described above reached CI unchecked. The line is rewrapped and the errors are fixed.The survival tree’s log-rank split statistic was wrong (#287).
kind="non-parametric"trees, and anyRandomSurvivalForestbuilt from them, selected splits on a statistic off by factors of several – and not merely inflated, but reordered: of the two cases in the issue, the weaker separation scored 1.856 against a true 0.276 while the stronger scored 0.276 against a true 1.225. Trees were choosing the wrong split.The log-rank statistic sums over the pooled event times of both children, so the left child’s at-risk count is needed at times where the left child itself has no observation. Those were filled in by carrying its own risk ladder forward, which was wrong twice: the carried value did not subtract the deaths and censorings that occurred at the time it was carried from, and the tail past the last observation subtracted only deaths, so a child ending in a censored observation kept someone at risk for ever. Both inflate the count, which biases the numerator and the variance.
The at-risk count is now computed directly – at each pooled time \(t\), the observations with \(t_l < t \leq x\), which is the
(entry, exit]conventionxcnt_to_xrdalready uses, so the left child’s \(Y_L\) and the pooled \(Y\) agree about what “at risk” means. There is nothing to extrapolate, so both the leading and trailing special cases disappear along with the forward fill.Tests check the statistic against a deliberately naive implementation of its definition: the two cases from the issue, a censored tail, left truncation, ties shared across children, a child starting after the other’s first event, and 200 random partitions. One test pins
at_risk_on_gridagainstxcnt_to_xrddirectly, since a disagreement between them is what makes \(Y_L / Y\) stop being a proportion.The documentation build fails on any warning.
-W --keep-goingin the CI docs job, andfail_on_warningin.readthedocs.yaml. They are set together deliberately: if only one has it, that one goes green while the other publishes a broken page.A Sphinx warning is rarely cosmetic. A broken cross-reference renders as plain text, a mistyped
autoclasspath drops the class from the page altogether, a page missing from every toctree is published and unreachable. In each case the build reports success and ships something wrong, and the only evidence is a line in a log nobody reads. That is precisely how the three brokenProportionalIntensityNHPPautodoc targets survived – three methods absent from the rendered documentation, behind a “build succeeded, 18 warnings”.Warnings-as-errors only works from zero, so the sixteen were cleared first:
Twelve duplicate labels, from
autosectionlabelminting a cross-reference target for every heading while the changelog necessarily repeats “Serialisation”, “Degradation” and so on once per release.autosectionlabel_maxdepth = 1keeps the labels a:ref:between pages actually wants and stops minting the rest.The
sphinx_rtd_themeget_html_theme_pathdeprecation, whose own message said the call was safe to remove.A title-level inconsistency in the non-parametric page, reported by docutils as
CRITICALrather than a warning.RandomSurvivalForest’s docstring, which was not valid reStructuredText – and which nothing surfaced until the new API page began rendering it.The interval-censored Turnbull example, whose EM ran out of iterations. It converges at
max_iter=10000; the example now passes it and the prose explains why, rather than marking the warning expected and hiding a usable answer.
``scripts/check_all_pythons.py`` runs the CI checks locally on every supported interpreter. With the suite no longer running on pull requests into
develop(below), this is the other half of the trade: one command runs the test suite and both doctest passes on 3.11, 3.12 and 3.13, and refuses to say “passed” unless all of them did.It keeps its environments in a git-ignored
.venvs/and reuses them, so only the first run pays for the installs; it usesuvwhen available and falls back tovenvandpipwhen not, and reports an interpreter that is not installed rather than failing on it. The command list is deliberately a copy of the workflow’s, so what it runs is what CI would have run.The test suite runs on the release pull request, not on every one. Pull requests into
developnow run lint only, about a minute against the nine the suite takes across three interpreters. The suite still runs in full on the release pull request intomasterand on pushes tomaster.The reason is the edit-review loop: with a single maintainer running the suite locally before pushing, the pull-request run was mostly confirming what was already known, and it was the slowest part of working on the package.
What this gives up is stated rather than glossed: a failure that appears on only one interpreter is now found when the release is prepared, with a release’s worth of commits to search rather than one. That is not hypothetical – the doctest numeric comparison two entries below landed green on 3.11 and failed on 3.12 and 3.13, and it was the pull-request run that caught it.
Contributing.rstnow says which jobs run on which event, and what to run locally to compensate.The documentation build runs in CI on the release pull request. The docs execute every
.. jupyter-execute::cell as they build, so they are a second test suite that exercises the public API for real – and one that a change touching no documentation file at all can break, as the Gamma entry below did. Read the Docs builds onlymasterand tags, so until now that break would have surfaced as a failed hosted build after a release.The new job is conditioned on
github.base_ref == 'master', which is set only for pull requests, so it runs on thedevelop->masterrelease pull request and nowhere else. It is not run on pushes tomastereither: Read the Docs rebuilds there anyway, and by then the gate has nothing left to gate. It matches.readthedocs.yamlrather than the test jobs – Python 3.12, the package installed via its owndocsextra – because its purpose is to reproduce the hosted build, and it uploads the rendered HTML as an artifact.It does not build with
-W; the build currently emits 18 warnings, mostly duplicate labels fromautosectionlabelmeeting the changelog’s repeated section headings. Clearing those and then failing on warning here and in.readthedocs.yamltogether is worth doing separately.The residual gap is deliberate: a documentation break introduced on a pull request into
developis caught when the release is prepared, not when it lands. Building on every pull request would cost minutes on each, and a path filter would not have helped here – the change that broke the build was ingamma.py, not underdocs/.Fixed a documentation build broken by the Gamma MPP removal. The offset-threshold section of Parametric SurPyval Modelling ran a
jupyter-executecell looping over['MPP', 'MOM', 'MSE', 'MPS', 'MLE']for a shifted Gamma. SinceGamma.fit(how="MPP")now raises, that cell raised, and because documentation cells are executed during the build the whole build failed.Nothing caught it: continuous integration does not build the documentation, and Read the Docs builds only
masterand tags, so it would have surfaced as a failed hosted build at the next release rather than on the change that caused it. It was found by running a build to validate thedocsextra below.The prose around the cell had gone stale the same way – it described the multi-start probability-plotting search that the removal deleted, and quoted an
MPPtolerance fromtest_offset_divergence.pythat no longer exists. It now explains why the Gamma has no probability plot at all: the shape sits inside the regularised incomplete gamma rather than outside as an exponent, so the only straight-line axis is the inverse incomplete gamma, which needs the very shape being estimated.Gamma.plot()is unaffected, since by then the parameters are known.The documentation toolchain is a ``docs`` extra.
pip install -e ".[docs]"now installs everything needed to build the documentation, alongside thetestsextra that was already there.docs/requirements.txtis gone: its pins moved intopyproject.tomlverbatim, and Read the Docs installs the extra directly viaextra_requirements. Keeping both would have meant two copies of the same pinned toolchain, which is the arrangement that drifts.The pins are unchanged, including the
ipykernel==6.31.0cap and the reason for it – jupyter-sphinx notebook execution dies against the ipykernel 7 line.matplotlibis not repeated in the extra; it is a runtime dependency of the package, which is installed alongside.Part of #141.
CI now runs the docstring examples.
pytest --doctest-modulesover the package is a new step in the deployment workflow, and every one of the 229 docstring examples passes. It was 59 failing tests when the flag was first turned on.A docstring example is a promise about what the library prints, and it is the one users and coding agents reach for first –
help(Weibull.fit)is faster than opening the docs. Nothing was checking it, so it drifted: examples recorded the output of an optimiser two rewrites ago, of numpy 1’s scalar repr, of a module that has since moved.What the run found, beyond the cosmetic drift:
Twelve examples could not run at all. Six regression docstrings (
PH,AH,PO,AFT,AcceleratedLife,Frailty) were sketches –model = PH(Weibull).fit(x, Z=covariates, c=c)withx,covariatesandcnever defined. Four more used>>>on the continuation lines of a multi-line call, so pasting them raisedSyntaxError.plotting_positionsimported fromsurpyval.nonparametric, which moved tosurpyval.univariate.nonparametricseveral releases ago.ParametricFitter.fitdemonstratedhow='MPP'on interval-censored input, which now (correctly) requires the Turnbull heuristic and raises without it. All are now runnable, with data.The five
ParametricRegressionModelprediction examples (sf,ff,df,hf,Hf) had been copied from the univariateParametricclass and never adapted: they built aWeibull.from_params([10, 3])and called it with no covariates at all, documenting a signature the method does not have. They now fit aWeibullPHand passZ.Parametric.var()claimed 11.229 for a Weibull(10, 3). The variance is 10.533 (100 Gamma(5/3) - (10 Gamma(4/3))^2); the code was right.Several examples fitted unseeded random data and then recorded specific digits, which cannot be reproducible. They now seed.
Parametric.hfandParametric.Hfreturned a 0-d array (array(0.012)) for scalar input wheresf,ff,dfandqfall returned a numpy scalar, andcsdid the same; their ownReturnssections promised “the scalar value … if a scalar was passed”. That is now true. The 0-d array came fromnp.where, which does not collapse.The numbers in the examples are compared as numbers. doctest compares printed output as text, which is the wrong test for a library whose examples end in an optimiser: the same
Duanefit lands onb = 4.1995e-05under Python 3.11 and4.2032e-05under 3.12, and numpy prints eight significant digits either way. Sixteen of the 229 examples disagree between those two Pythons somewhere in their digits.The obvious workaround – trimming each documented number back to the digits that agree everywhere – makes the docstring show something the reader’s own session will not produce, which is precisely what these examples exist to avoid. So the examples record the real output, in full, and
conftest.pyinstalls a fallback comparison that runs only after the ordinary text comparison has failed. It fires when the two outputs are identical apart from their numeric literals – same words, same brackets, same integer-versus-float shape, so1never matches1.and a dtype change is still a failure – and then compares the numbers withrel_tol=1e-3, set by the loosest genuine disagreement between supported Pythons with no margin beyond it, andabs_tol=1e-12for a restoration factor whose true value is zero and which surfaces as1e-16with whatever mantissa the optimiser stopped on.What that forgives is a value drifting inside the tolerance. What it still catches is every defect listed above: a stale value from another parameterisation, the wrong function being called, the wrong shape, an exception, a missing import.
surpyval/tests/test_doctest_checker.pypins both halves of that, using the real output pairs observed on different Pythons.The fallback only runs when an example has actually drifted, which on any one machine is a handful of them – and a different handful on each. A gap in it is therefore invisible locally and surfaces in CI, on whichever Python computed a different last digit. So the doctest step runs twice: once normally, and once under
--doctest-force-numeric, which routes every example whose output contains a number through the numeric comparison. Fifteen seconds, and the fallback is exercised against all 229 examples rather than today’s accidental few.NORMALIZE_WHITESPACEis set inpyproject.tomlfor the same reason: numpy picks its own line breaks and column padding for an array and both move with the width of the widest element.This closes #158.
The distribution docstring examples now show what you actually see.
pytest --doctest-modulesoverdistributions/is green: 139 examples, no failures. Previously 39 failed.Most were the numpy 2 scalar repr.
Weibull.mean(3, 4)printsnp.float64(2.7192074311664314), where the docstring recorded the bare2.7192074311664314that numpy 1 used to print. The examples now record the wrapper, because that is what appears at a prompt – the alternative wasnp.set_printoptions(legacy="1.25")in a test fixture, which would have kept the docstrings prettier by showing readers something their own session will not produce.Four
qfexamples printed wider than the 79-column limit once the real output was recorded, numpy having rewrapped the arrays. Rather than hand-wrap them into something numpy would not emit, those examples take fewer probabilities: what is shown is exactly what that input produces.Two scalar examples had drifted in the last digit, and are re-recorded.
With the examples now true,
--doctest-modulesis worth running in CI, which is what stops this recurring; it is turned on above.``Gamma`` no longer offers probability plotting as a fit method.
Gamma.fit(x, how="MPP")now raises, joiningBetaandExpoWeibull, which already declined for the same reason.A probability plot works by rearranging the survival function so some transform of the data falls on a straight line. For a Weibull,
log(-log S) = beta log x - beta log alpha— the axes do not depend on the answer, so you can draw them before knowing anything. The Gamma has no such rearrangement: its CDF is the regularised incomplete gamma function and the shape sits inside that special function rather than outside as an exponent. The only straight-line y-axis is the inverse incomplete gamma, which needs the shape. To draw the axis you need the answer; to get the answer you need the axis.The code broke the circle by guessing the shape from moments, drawing the plot on that guess, and regressing. When the guess was off, the axis was the wrong axis, the points were no longer straight on it, and the regression fitted a line through a curve — returning a confident wrong estimate rather than an error. An offset made it worse: the shift distorts the low-
xend hardest, which is exactly where the shape information lives.plot()is unaffected. It transforms with the fitted parameters, so by the time the plot is drawn the axis is the right one — the probability plot of an MLE-fitted Gamma remains a valid diagnostic. Fitting is unchanged for MLE (the default), MSE and MOM.The 118-line
Gamma.mppoverride is deleted with it, which removes therr="x"mis-inversion and the censored-dataLinAlgErrorfrom #257 by making both paths unreachable. The MPP sweep intest_fit.pynow gates on each distribution’ssupports_mppflag instead of a hardcoded exclusion list, so it stays correct without editing.Nine wrong examples in the distribution docstrings. Running
pytest --doctest-modulesoverdistributions/gives 39 failures. Thirty are the numpy-2 scalar repr (np.float64(2.719...)against a recorded2.719...) and are cosmetic. Nine were not.Five documented outputs were simply wrong.
Uniform.ff’s example calledUniform.sf, andExpoWeibull.cs’s calledExpoWeibull.sf– in both the printed values were right for the function being documented and wrong for the one being called, so the example read as if the two were the same.LogLogistic.sfcarried values from some other parameterisation entirely (0.622 where the answer is 0.988),LogLogistic.mean(3, 4)claimed3against3.3322(the closed form isalpha (pi/beta) / sin(pi/beta)), andExponential.qfhad stale digits.The other four were the
CustomDistributionexample – the Gompertz walkthrough – whose multi-linedefused>>>where doctest needs..., so pasting it raisedIndentationError.In every case the code was right and the documentation was wrong, which is the reassuring direction, but a reader checking their understanding against these would have been misled. They accumulated precisely because the doctests were not run, which is addressed above.
Documented why a Turnbull fit does not equal a Kaplan-Meier fit.
Turnbull.fitdefaults toturnbull_estimator="Fleming-Harrington"whileKaplanMeier.fitis, unsurprisingly, KM. The Turnbull EM recovers the sameranddeither way; the three estimator options then differ in how those become a survival curve. Comparing the default againstKaplanMeierand reading the gap as a defect is an easy mistake — it is the one #260 was filed on, and the one made again while checking whether #260 was still open.On
x=[2,3,3,4,5,6], tl=[0,0,1,1,2,2]the survival at 2 is 0.750 under KM, 0.765 under FH and 0.779 under NA. With the estimator matched, Turnbull agrees withKaplanMeierto around 1e-9 on bothsfandcb, across right-censored and left-truncated data.Only the KM option is the non-parametric MLE. Maximising the truncated likelihood directly over the mass vector gives 0.750; FH’s 0.765 scores worse on that same likelihood, as an
exp(-H)construction should. FH is the default because it behaves better in the far tails and on zero-inflated data (v0.8.0), not because it maximises anything. The docstring now says all of this, and a test pins the three figures and the NPMLE identity against a brute-force maximisation.No behaviour change.
v0.19.0 (4 August 2026)
Confidence bounds no longer turn silently to nan on data measured in large units, and those fits are around 17x faster. A Weibull fit to the same lifetimes expressed in hours had standard errors; in seconds it returned
nanfor every one of them, with no warning and a perfectly good set of parameters alongside.The cause is the
np.wheretrap again, this time in the parameter transform rather than a likelihood. Every parameter bounded on(0, inf)is mapped to the unbounded space the optimiser searches byadj_relu, which chose betweenx + 1andexp(x)withnp.where. Autograd evaluates both branches, soexp(x)was taped even wherex + 1was selected, and abovex = 709.78it overflows to inf. The inf then poisoned the derivative of the branch that was chosen, so the jacobian of the transform came back nan – and with itcov_matrix, which is that jacobian either side of the inverse hessian.The threshold is a property of the fitted parameter, not of the sample size or the conditioning, which is why it looked so arbitrary: a fit died as soon as any
(0, inf)parameter exceeded about 710. A Weibull withalpha = 10lost its bounds once the data was scaled past about 70x, while a Gumbel, whose location is unbounded and so untransformed, survived to 350x on the same data. The Gamma failed at the small end instead, its rate parameter growing as the data shrinks. The Normal, LogNormal and Exponential were immune throughout because they have closed-form estimators and never touch the transform; the Uniform reports no covariance at any scale by design, its MLE being an order statistic rather than a stationary point.Clamping the dead branch’s argument fixes it: the branch is responsible only for
x < 0, so restricting what it may be handed leaves its value and derivative untouched where it is used, and bounded where it is not.The hessian was never the problem – the numerical fallback (#270) produced a finite, well conditioned matrix at every scale – which is why this presented as nan bounds rather than as a warning or a failure. It also explains the speed: the same nan reached the objective’s gradient, so BFGS, TNC and Newton-CG each gave up and Nelder-Mead finished the fit derivative free. Across twelve distributions at seven scales the sweep goes from 19.59s to 1.18s, and every fit that used to end on Nelder-Mead now ends on BFGS.
Results that already worked are unchanged: 65 of 96 reference fits are bit-identical, and the other 31 are the restorations, where the objective agrees to fifteen significant figures and the parameters to nine.
Restoring the gradient exposed a second, smaller scale problem underneath, now fixed with it. scipy stops BFGS when
max|grad| < gtol, and its default of 1e-5 is an absolute threshold on a quantity that is not scale free: a log-likelihood’s gradient shrinks like1/theta, so on data measured in tens of thousands the test is met well short of the optimum and BFGS reports success on its first check. This had been invisible because those fits used to end on Nelder-Mead, which is derivative free and so kept going. Three reference fits on real data of that magnitude landed 1e-2 away in relative terms, at a likelihood 2e-3 below the answer they had been recorded from.There is a second dimension to the same problem, in the opposite direction. A log-likelihood is a sum over observations, so its gradient grows like
nas well as shrinking like1/theta. At n = 1e5 it is five orders of magnitude larger, the same absolute threshold is correspondingly unreachable, and BFGS gives up on censored samples it should handle easily – found by benchmarking against lifelines with censoring, where it was the one configuration in 64 where surpyval was slower.So the problem is rescaled in both of its dimensions, and a single constant then means the same thing for every fit:
s = max(|u0|, 1) f0 = max(|f(u0)|, 1) v = u / s g(v) = f(s v) / f0
The starting point is order 1 in every component and so is the objective, whatever the units and whatever the sample size. Since
dg/dv = s * df/du / f0, withsgrowing exactly asdf/dushrinks andf0growing liken, so is the gradient. Thegtolof 1e-6 applied there is a genuine relative tolerance rather than the dimensioned constant scipy’s default is.Tuning the threshold was tried first and does not work. Three criteria were measured against the same reference set: an absolute
gtol, agtolscaled by the gradient at the initial guess, and BFGS’s step-size testxrtol. Swapping the whole method for L-BFGS-B to reach its relativeftolwas measured too. None is scale free in practice.Scaling
gtolby the initial gradient in particular looks scale free and is not: the initialiser scales with the data too, so that gradient is itself roughly scale invariant – a Weibull at scale 1, 1e4 and 1e6 all came out withgtol = 1.86e-6. Nor is there a constant that serves every case: tight enough for a Weibull at 1e6 is unreachable for an n=8 sample, which then drops out of BFGS into TNC.xrtolandftolfail differently – both stop on how the optimiser is behaving rather than on the quantity that is zero at the answer, so they quit early along flat directions, which is precisely where the standard error is largest and most needs to be right. The ExpoWeibull, three parameters and a flat surface, was 17% out under both. The full measurements are in #323.Rescaling beats every one of them, and is faster than the tolerance it replaces:
scale-equivariance of
se(8 distributions)relative
gtolrescaled problem
worst deviation
1.6e-3
2.0e-5
cases above 1e-5
1 of 16
1 of 16
Weibull n=1e5, 30% censored, scale 1
0.163s
0.153s
Weibull n=1e5, 60% censored, scale 1e6
0.250s
0.119s
Rescaling the parameters alone reaches 2.8e-8 on the first row, better than the 2.0e-5 above, but is the version that leaves BFGS failing at large
n: those two censored fits take 0.370s and 0.590s under it. Normalising the objective as well trades a little of that accuracy for a criterion that holds across sample sizes too, which is the point.Both mappings are fixed before the search begins and neither can move the optimum: a diagonal linear change of variable relocates a minimum no more than dividing the objective by a positive constant does. They change the route taken and the units of the convergence test, nothing else.
res.xandres.funare both mapped back inside the helper, so no scaled quantity exists anywhere else in the package, not even transiently: the covariance step,cband serialisation all receive exactly what they received before. Nothing needs to know which parameter is a scale and which a shape, which is what made the internal-rescaling proposal in #323 risky; preconditioning needs none of it.Two consequences worth noting. Fits now converge further than before wherever BFGS wins, so a handful of pinned numbers moved in their last few digits – always towards a better likelihood. The two Monte Carlo simulation tests in
test_counting.pyalso had their tolerance loosened fromallclose’s 1e-5 to 1e-3: they drive a 5000-run simulation from an optimiser’s output, where a change in the seventh significant figure of a parameter moves the simulated MCF in the fourth. That is convergence noise, and what those tests exist to catch would break far more loudly.The second is that it flushed out a separate defect, which is fixed alongside it and described next.
#323 is now rescoped. It had proposed rescaling the model – fit on transformed data, then map the parameters and the covariance back – to fix both the bounds and the speed. Neither needs it. The bounds were the
np.whereoverflow above, and re-measured with the gradient working, rescaling the data is between 4.2x faster and 6x slower depending on the distribution. What was left, the convergence criterion, is what preconditioning the search addresses, so the twelve per-distribution back-transforms are not needed for any of it.Cox fits are 3x to 10x faster, with the answers unchanged to every digit. None of this is a change to the maths; the coefficients still match
lifelinesto between 1e-06 and 1e-12, exactly as before.Four things were costing the time (#329).
_GroupBy.sumusednp.add.atfor its multi-dimensional case, an unbuffered scatter with no fast path, called about ten times perjac_hesson arrays of shape(n, p, p). It is now a sortednp.add.reduceat, with two shortcuts: nothing to permute when the keys already arrive grouped, and nothing to add when every key is distinct — one-element groups in order are the input array, which is what continuous event times give you.The hessian was a Python double loop over event times and tied deaths, with an
np.outerinside it. The sum over tied deaths turns out to factor out of thep x ppart entirely: onlyc = j / ddepends onj, so five scalar sums per event time carry the whole ragged axis and the covariate blocks are formed once. That drops the cost fromO(times x ties x p^2)toO(times x ties + times x p^2), and it removes the separate no-ties case rather than special-casing it — an untied time is a singlej = 0term withc = 0.np.einsum("ij,ik->ijk", Z, Z)was rebuilt insidejac_hesson every root-finding iteration, thoughZis fixed for the life of the fit. It is hoisted, and the two remaining weighted-outer einsums are plain broadcasts.Finally the rows are put in event-time order once when the closures are built, so
_GroupBynever has to permute an(n, p, p)array again. Nothing downstream depends on row order — every quantity is aggregated to unique event times first — and the model still stores the caller’s unsorted arrays for the residual and diagnostic code.Wall clock, Efron ties, 40% censored, against
lifelines:n
p
before
after
lifelines
500
2
0.038s
0.012s
0.060s
2 000
2
0.146s
0.024s
0.146s
10 000
2
0.890s
0.105s
0.619s
50 000
2
5.71s
0.663s
3.13s
10 000
5
1.379s
0.339s
0.656s
50 000
5
7.57s
2.01s
3.18s
2 000
10
0.378s
0.107s
0.164s
50 000
10
15.39s
8.31s
4.03s
Heavily tied event times — dates, rounded durations — gain the most: n=20 000 with five covariates goes from 1.91s to 0.33s. Breslow, which shares
_GroupByand the einsum hoist, improves alongside it. The delayed-entry and time-varying-covariate paths, already well ahead, roughly halve again: 43 384 rows with delayed entry from 6.86s to 2.62s, and 20 000 start-stop rows from 1.07s to 0.44s.Two cases remain slower than
lifelines: fifty thousand rows with ten covariates (8.3s against 4.0s), and heavy ties (0.33s against 0.08s). Both are now dominated by materialising(n, p, p)arrays — 40MB apiece at that size, several per iteration. Getting past that means accumulating thep x pinformation incrementally per event time instead of building per-observation outer products, which is a change of algorithm rather than of implementation, and is tracked separately as #332.Timings are single runs and drift by 10–20% between them; the
lifelinescolumn moves about as much as the surpyval one does.Parametric proportional hazards fits are up to 6x faster and no longer degrade on data measured in large units. A
WeibullPHfit at data scale 1e6 settled 1.5 nats of log-likelihood short of the optimum, and 1e-2 away in the covariate coefficients – a different fitted model, not a tolerance artefact. The same data in unit scale fitted correctly, so nothing about it looked wrong.The PH ladder was
minimize(fun, init_t)followed by TNC, and three things were the matter with it. The objective closes overregression_neg_ll, which is written inautograd.numpyand is therefore differentiable, but nojacwas passed – so scipy fell back to a two-point finite difference and paidp + 1extra objective evaluations per gradient, which is what made the fit slow down as the covariate count rose. The search was not preconditioned, so PH inherited the scale sensitivity fixed for univariate MLE elsewhere in this release: scipy stops BFGS on an absolute threshold applied to a gradient that shrinks with the data magnitude and grows with the sample size. And TNC’s answer was returned whether or not it had converged, so a rung that can only ever be an improvement was free to be a regression – the AFT and PO ladder, defined immediately below it, already guarded against exactly that.The ladder is now preconditioned BFGS on the analytic gradient, then TNC, then Nelder-Mead, stopping at the first rung that converges and never returning a worse point than it started from. The derivative-free rung stays for fits where the gradient is unusable.
Measured against
lifelines, scoring both packages’ answers on an independently written Weibull PH log-likelihood: the scale-1e6 shortfall goes from 1.468 nats to 1.1e-08, and a 50 000 x 10 fit drops from 1.53s to 0.40s againstlifelines’ 2.38s.ExponentialPHat the same size goes from 1.21s to 0.14s. Coefficients continue to agree withlifelinesto around 1e-06.preconditioned_bfgsmoved fromfitters/mle.pyup tofitters/__init__.py, alongsidebounds_convertandfallback_minimize, so the univariate and regression ladders share one copy rather than two. Behaviour of the univariate ladder is unchanged.optimise_nm_tnc, which serves AFT and PO, has the same missing gradient and missing preconditioning. Its first rung is Nelder-Mead, which is derivative free and so cannot fail the way BFGS did here, so it is being measured before it is changed rather than assumed to need the same fix — #331.Turnbull no longer rejects left-censored observations under left truncation. An entry time below every observation excludes nobody, so it should leave a fit untouched. With any left-censored row present it raised instead:
ValueError: An observation's censoring interval does not intersect its own truncation window ...
A support index
jstands for the half-open interval(bounds[j], bounds[j+1]], so an event placed there is already strictly afterbounds[j]. The first index a row entering attlmay use is therefore the last bound equal totl– that interval is(tl, next]. The window construction took one index further on, discarding it.It mattered most for left censoring because such an event lies in
(-inf, xr], which under an entry attlis the single interval(tl, xr]– frequently the only one the row has. Dropping it left the row with an empty support, hence the rejection.Neither endpoint of the search alone is correct, which is what made this awkward.
side="left"keeps the zero-width(tl, tl]interval that a duplicated exact event time creates, readmitting an event at exactly the entry time and breaking the strict(entry, exit]convention (#260);side="right"discards(tl, next]as well, one too many.side="right" - 1lands between them, and does so exactly, because every finite truncation time is itself inbounds.A Turnbull fit that is not identifiable now says so instead of returning a collapsed curve quietly. Left-censored observations combined with two or more distinct entry times admit a flat direction in the likelihood, and where the data leans on it the estimate is worthless while looking ordinary.
An interval that one observation could have failed in, but that precedes another observation’s entry, is worth mass to the first and costs the second nothing. The second’s contribution is conditional on its own entry, so mass it never had the chance to see divides out of both its numerator and its denominator exactly. On a six-point example the estimator drives 99.995% of the mass into a single such interval, reaching a log-likelihood of -6.14 against -9.36 for the sensible answer.
So the estimator is not misbehaving. It is maximising correctly, and the likelihood has no interior maximum to find – it climbs towards the boundary, which is why raising
max_iternever helps. Both ingredients are needed: left censoring, the only kind whose support reaches back into the entry region, and two distinct entry times, so that such an interval exists at all. Six distinct entry times with no left censoring fit flawlessly; one common entry time with left censoring round-trips exactly.The fit is still returned, because meeting the condition does not mean the data is spoilt. Across 240 simulated samples that all met it, the proportion actually degenerating ran from 8% to 72%, rising with the share left censored – rejecting on structure would refuse far more good data than bad. What separates the two is how much mass ends up on the flat direction: healthy fits reached at most 0.836 of it, spoilt ones a median of 0.994. Warning above 0.9 caught them without a single false alarm across those samples; 0.7 would have cost 9% and 0.5 40%.
The share is reported as
model.exploitable_massso a borderline fit can be judged rather than guessed at. Note that the structural condition is not used as a trigger on its own: ordinary staggered-entry data meets it routinely and estimates perfectly well, and pairing it with non-convergence would have mis-advised the #203 case, which is structurally exploitable but converges given the iterations.This is the second half of #308, which closes with it; the first half, an off-by-one that made these same inputs raise, is above. The threshold is a measured cut-off standing in for a property that is actually decidable: Vardi (1985) and Wang (1991) give a graphical condition on the data that settles whether the NPMLE exists, exists but is not unique, or does not exist at all, with nothing to tune. Adopting it, and the question of what a non-identifiable fit should return rather than merely report, are #327. Worth noting alongside that left truncation with interval censoring is documented as yielding an inconsistent NPMLE, so this is a known limit of the estimator for this data shape rather than something particular to surpyval.
Truncated parametric regression fits could report a log-likelihood tens of thousands higher than their parameters earn, and be optimised towards it.
truncation_correctioncomputed the mass in each observation’s truncation window as a difference of CDFs, floored at the smallest positive float:np.log(np.maximum(right - left, _TINY))
Under left truncation that difference is
1 - F(tl), the survival probability at the truncation bound, which underflows to exactly zero as the fitted scale shrinks. The floor then capped the correction atlog(tiny) = -708rather than letting it grow without bound – and since the correction is subtracted, every truncated row appeared to contribute +708 to the log-likelihood. A region the data rules out entirely became the best fit on offer, and the optimiser walked straight into it.A
WeibullPHfit to left-truncated data reportedneg_ll-21311.40 at parameters whose true value is +17118.30, against 927.83 at the correct answer: wrong by 38,000, and pointing the wrong way. Recomputing the likelihood by hand from the model definition is what settled it – at the correct parameters surpyval agrees to the digit, so the objective is right everywhere except where the floor engages.One-sided windows are now evaluated in log space through
log_sfandlog_ff, which stay finite where the difference cannot, so there is nothing to floor. Only a genuinely two-sided window still takes a difference, and there both bounds are finite and the mass is not driven to zero by the scale alone. As elsewhere, thenp.wherebranches are evaluated at substituted-finite arguments so that an infinity in an unselected branch cannot poison the gradient of the selected one.This is in
_likelihood.py, which serves proportional hazards, proportional odds, accelerated failure time and accelerated life alike, so any left- or right-truncated parametric regression fit was exposed – it needed only the optimiser to wander far enough for the underflow to bite. Nothing warned when it did.Found by the rescaling change above, which perturbed an initial guess by six parts in ten million and was enough to tip one fit over. The first diagnosis was wrong: it looked like a genuinely unbounded truncated likelihood being followed legitimately, and the arithmetic disproved that. #326 records both. The regression test asserts the reported objective equals an independently computed one and that shrinking the scale below the truncation bounds always scores worse.
A truncated fit is around 60x faster, and the truncation term is evaluated once per distinct window rather than once per row. Any fit with a truncation bound on one side only had no usable gradient. The window probability chose between the CDF and an analytic limit with
np.where, which picks the right value but evaluates both branches – soff(inf)was still recorded by autograd, and its nan derivative propagated through the selection whichever side won.Nothing warned. The objective was correct throughout; only the gradient was nan. So BFGS and Newton-CG each gave up after a single evaluation, TNC spent its whole 1000-evaluation budget discovering the same thing, and Nelder-Mead finished the job derivative free. A Weibull that fits in 0.014s took 1.36s, and a
tlof 0 – a no-op, sinceF(0) = 0– cost exactly as much as a real truncation, which is what gives the cause away. Windows with both bounds finite were always fast, because no infinity ever reached the tape.The infinity is now substituted out of the argument before the CDF sees it, so a single vectorised call covers every row whatever its pattern of bounds, and the surviving
np.whereonly ever chooses between two values that are already finite. The stand-in cannot be an arbitrary constant: zero looks natural and is wrong, because a Weibull withbeta < 1has an unbounded density derivative at the origin, which would swap one nan gradient for another. Reusing a bound that is genuinely present keeps it inside the support and at the data’s own magnitude; its value never reaches the result, only its derivative has to be finite.Separately, the truncation correction depends only on the observation window, not on where in it the observation fell, so it is now evaluated once per distinct window. Truncation is nearly always common to a whole sample – one burn-in time, one study entry date – which collapsed 360 CDF evaluations per likelihood call to one in the test case, and the likelihood is called hundreds of times per fit.
fit
before
after
plain
0.014s
0.015s
left truncated
1.398s
0.022s
tl = 0(a no-op)1.362s
0.041s
right truncated
1.344s
0.024s
both bounds finite
0.019s
0.025s
Fitted results are unchanged: all 330 reference fits across thirteen distributions and five methods are bit-identical, and BFGS now wins every truncated fit where Nelder-Mead used to.
Confidence bounds were never affected. The covariance step already recomputes a numerical hessian whenever the autograd one comes back nan or asymmetric (#270), so it caught this on every truncated fit and produced correct bounds by the slow route – checked against
906f0cb~1, where the standard errors are identical to eight figures. That fallback was part of what made these fits slow.The slow parts of the test suite are opt in, and there is a new invariant sweep behind the same mechanism.
pytestalone now runs in about two minutes rather than three: the beta survival tree and forest tests were 97 of the suite’s 180 seconds for 85 of its 2000-odd tests. They run with--run-ml, and continuous integration passes the flag, so coverage is unchanged. Theconftest.pythat defines the flags lives at the repository root:pytest_addoptionis only honoured in initial conftest files, and the CI invocation selects with--ignoreand names no path, so one undersurpyval/testswould be loaded too late to register them.--run-invariantsaddstest_fit_invariants.py, a wide net over the fitting API. Every defect found in this release cycle slipped past the whole suite, and each lived at an intersection of dimensions the suite tests one at a time – a censored observation that was also truncated, an offset combined with a particular method, an offset combined with a large shift magnitude, a sample below the 1000 floor ofFIT_SIZES. The full cross of distributions, methods, censoring, truncation, structural flags, sizes and scales is around 600,000 cells, so the sweep does not attempt it. It asserts cheap invariants instead – finite parameters, finiteneg_ll, a survival function that stays in [0, 1] and never increases, and maximum likelihood attaining the lowest negative log-likelihood of the five methods – over a seeded sample of that space. Four of the five defects would have failed the first two assertions.Data scale is included as an axis because it was previously untested anywhere, despite the maximum likelihood failure warning itself advising users to rescale towards 1. 270 cases, three and a half minutes.
Maximum likelihood fits are about 2.2x faster. Every MLE fit ran five optimisers – Nelder-Mead, Powell, BFGS, TNC and Newton-CG – and kept the best result. Over 102 fits across eleven distributions, five data shapes and two sample sizes, all five agreed on the objective to 1e-10. The last four were confirming what an earlier one had already found.
That confirmation was not cheap. Nelder-Mead and Powell are derivative free, so they pay for robustness in function evaluations – 50 and 22 against BFGS’s 21 – and every evaluation costs O(n). On a million observations those two alone were 42% of the fit.
The gradient methods now run first and the search stops at the first that converges, with the derivative-free pair kept as the fallback. Order and early exit had to change together: stopping early without reordering halts at Nelder-Mead, which is both the most expensive rung and the one with the worst objective, while reordering without stopping early saves nothing. Cold-start BFGS now wins 83 of 102 fits, TNC takes 10 and Newton-CG one; the eight that Nelder-Mead or Powell used to win now land on a gradient method at the same objective, so they were winning ties on ordering rather than finding better optima. The derivative-free methods still start from the cold initial guess when they are reached, so the multi-start behaviour survives for the fits that need it.
Fitted parameters can move in about the seventh significant digit. All 102 objectives are identical to 1e-10 and one improved, so this is optimiser tolerance rather than a change of answer, but it is not bit-identical: the median shift is 3e-8 and the 90th percentile 8e-7. The documented
GeneralizedOneRenewalexample and the two tests that pin it have been regenerated. Those numbers were always a snapshot of the library’s own output rather than an external reference, and their tolerance has deliberately been left tight, so that any future change to the optimiser surfaces as a decision rather than passing unnoticed.Degenerate data is rejected with an explanation instead of an ``IndexError`` from inside numdifftools.
Weibull.fiton three tied observations died four steps from the cause: a probability plot has no slope through a single distinct abscissa, sopolyfitreturned a nan; the nan seeded the maximum likelihood fit, which started at nan and produced a nan hessian; the numerical fallback then asked numdifftools for one, and its list of finite-difference steps came back empty. Neither truncation nor censoring was involved, despite where the symptom was first seen.GammaandBetafailed the same way but in silence. Their moment-based initialisers divide by a variance that is exactly zero for tied data, giving(inf, inf), and since a failed optimiser reports its initial guess (#261) those infinities were returned as a fitted model.Three changes. The probability-plot regression falls back to a unit slope through the centroid when it is rank deficient – zero slope would be the more literal reading, but every
unpack_rrdivides by the slope to recover a scale, so it only moves the nan one step later.GammaandBetaseed the exponential and uniform cases rather than dividing by zero. And a fit now refuses to return a non-finite parameter whatever produced it.The fit is then rejected when the data cannot pin down the free parameters: fewer distinct non-right-censored values than free parameters means a flat – for a Weibull on tied data, unbounded – direction in the likelihood, and the answer would be wherever the optimiser stopped. Three tied observations returned
beta = 512withsuccess=Trueand no warning once the nan was fixed.The count is of free parameters, so fixing one buys back a degree of freedom:
Weibull.fit([10.], fixed={'beta': 2})is well posed and now returnsalpha = 10, where before it raised. One-parameter distributions are unaffected –ExponentialandRayleighfit tied data exactly as they should. Probability plotting is exempt, being a regression rather than a likelihood maximisation, and is how several distributions seed themselves.All 330 reference fits across thirteen distributions, five methods and plain, right-censored and offset data are bit-identical.
v0.18.0 (2 August 2026)
Documentation caught up with the estimator changes below. The most consequential correction is in Parametric SurPyval Modelling, whose section on offset unidentifiability demonstrated the hazard of threshold parameters by fitting
Gamma(3, 2) + 10and reporting thatMPPreturned a negativegammawith a shape parameter inflated by two orders of magnitude. That has not been true since #257 and #313: every fit method now recovers the offset to three decimal places, at offsets from 10 to 1000 and at both small and large samples. The page had also drifted from the test file it cites –test_offset_divergence.pywas tightened by #257 and #275 to assert parameter recovery, the opposite of what the prose claimed.The underlying theory is kept, since it is exactly what made a poor starting point so damaging:
gammatrades off against the shape and scale, and the likelihood is flat along that ridge. It now reads as a caution for your own data rather than a demonstration of broken output, and names both causes of the old behaviour separately – the probability-plotting search stranded by a single starting shape, and the moment-based initialisers taking their moments before the shift.The maximum likelihood notes in Parametric Estimation now cover the observation that is both censored and truncated – the contribution it makes, why the numerator has to be capped by the truncation bound, and the fact that surpyval recasts such a point as interval censored in its internal representation rather than special-casing the likelihood. The user’s own
x,c,nandtare untouched, which is now said explicitly.The method of moments section explains that the optimisation matches scaled central moments rather than raw ones, and why that matters for offset data, where every raw moment is dominated by the offset.
Parametric SurPyval Modelling gains a worked comparison of
neg_ll,aicandbicacross all five fit methods, showing both that the criteria are available whatever the method and that MLE attains the lowest negative log-likelihood – a check that could not be run before.An offset ``ExpoWeibull`` fit now seeds itself from the shifted data.
ExpoWeibullstarts from a Gumbel fit tolog(x), since a Weibull’s logs are Gumbel distributed. Withoffset=Trueit took those logs before removing the shift, so it readlog(x)where the model wantslog(x - gamma). A large offset compresses those logs into a narrow band, the Gumbelsigmacollapses, andbeta = 1 / sigmaexplodes: on 500 points fromExpoWeibull(10, 2, 1) + 100the seed came back asalpha = 111, beta = 23.5against a true 10 and 2. The maximum likelihood fit then failed outright, returningnanand warning its way back to the MPP estimate.This was not an edge case. Over 120 offset fits – four offsets, five parameter sets, six replicates each – 54 returned nan, which is every configuration at an offset of 100 or 1000. All 54 now converge, none of the 66 that already worked changed for the worse, and the whole sweep takes 25.5 seconds against 188.3, since a hopeless starting point is expensive to fail from.
The offset is now estimated first and the shape parameters read off
x - gamma. It is estimated asmin(x) - 1, which is what the fitter installs regardless of what the initialiser returns – seeding against a different shift than the one being optimised under defeats the point.The nested Gumbel MLE that refines the offset seed is kept. Removing it was tried, on the reasoning that shifting the data correctly makes the probability plot alone good enough; it is not, and five of 48 offset fits landed on a worse optimum without it.
``aic``, ``bic``, ``aic_c`` and ``neg_ll`` now work for every fit method. They were available only after a maximum-likelihood or closed-form fit, because only those compute a log-likelihood on the way to the answer. A model fitted with
how='MPS','MSE','MOM'or'MPP'raisedAttributeErrorfrom all four – which meant the usual way of choosing between distributions was unavailable for four of the five methods, and failed with a message that did not say why.The log-likelihood is a property of the parameters and the data, not of the search that found them, so it is now evaluated after any fit. Methods that already reported one keep it exactly: maximum likelihood’s is the optimiser’s own final objective, which on its fallback path is deliberately taken at the initial guess rather than at the failed result (#261).
A worked consequence, on 500 Weibull(10, 2) points – the maximum likelihood estimator attaining the maximum likelihood, which was not checkable before:
method
neg_llMLE
1466.3254
MPS
1466.3865
MOM
1466.3979
MSE
1466.6759
MPP
1470.7119
MSE and MPS fits try BFGS before Newton-CG, and are several times faster for it. Both go through a shared fallback that reached for Newton-CG first, which needs a hessian. Building one is disproportionately expensive for the distributions whose derivatives autograd cannot take analytically – the incomplete gamma is central-differenced, so every second-order entry costs a difference of differences. An offset Gamma MSE fit at n=5000 spent 8.2 of its 8.3 seconds there, and BFGS reached a marginally better optimum in half a second.
Reversing the order was checked over 132 fits: MSE and MPS, nine distributions, at two sample sizes, on plain, right-censored, left-censored and offset data. 129 objectives came back identical, three improved, none got worse, for 3.9x less time overall. Newton-CG is still there, escalated to when BFGS fails, and Nelder-Mead behind it; the zero-hessian guard is kept, since a hessian of zeros makes Newton-CG stop at the initial guess while reporting success, so there is nothing to escalate to and Nelder-Mead should take over.
scipy.optimize.least_squareswas tried first, on the reasoning that the MSE objective is a sum of squares and Gauss-Newton should exploit it. It is far worse: it needs the full residual jacobian, so n=5000 means 5000 rows each paying that central-differenced derivative, and the same fit took 239 seconds against the scalar gradient’s three numbers.ExpoWeibull no longer runs a nested optimiser ladder to build its initial guess. It seeds itself from a Gumbel fit to
log(x), and that inner fit was a full maximum likelihood run – an optimiser ladder, to produce a starting point for another optimiser. It cost 15-30% of the fit (20 ms at n=200, 41 ms at n=5000) and the probability plot alone turned out to be just as good a seed: across 54 parameter combinations, plus right-censored, left-censored and heavily tied data, every fit reached the same optimum to the optimiser’s own tolerance. Ordinary fits are about 20% faster.The offset path keeps the refinement. There the seed reads
log(x)of the unshifted data, so a large shift compresses the logs into a narrow band and the probability plot is a poor starting point – dropping it moved one fit in twelve to a worse optimum (861.898 to 861.962) and made that fit seven times slower.The ARI likelihood is evaluated for the whole sample at once, making imperfect-repair fits 30-160x faster. It used to walk every event in Python, calling the baseline
cifandiifand rebuilding the intensity reduction from the failure history at each step – around ten milliseconds per event, repeated for every one of the optimiser’s several hundred objective evaluations. Fitting 250 items took 19 seconds and 1000 items was impractical.The apparent obstacle is that the reduction depends on the failure history, which looks inherently sequential. It is not: summing over the reduction’s window offset rather than over the failures turns it into a handful of whole-array passes, and the offset only ever runs to
min(m, longest item)– exactly one pass for ARI1. Each failure’s ordinal within its own item bounds the window, which is what stops one item’s history leaking into the next.fit
before
after
35 items x 6 events
2.11 s
0.07 s
250 items x 8 events
19.02 s
0.12 s
Cramer-von Mises, 10 boot
24.16 s
1.71 s
1000 items x 10 events
~80 s
0.53 s
Results are unchanged to floating point: over 312 captured values – the reduction helper across memory regimes, the objective on a fixed parameter grid, fitted parameters and the rescaled-increment residuals – the largest relative difference is 1.1e-14, from summation order. The original per-event implementation is kept in the test suite as an oracle so the two cannot drift.
One behavioural nicety: a non-positive reduced intensity is outside the model’s support, and the scalar loop returned
infearly on reaching one. The vectorised form tests the whole array before taking logs, so it returnsinfrather than warning its way to anan.Method of moments now matches central moments rather than raw ones. The two describe the same estimator – the binomial transform between them is exact and bijective, so matching the first
kcentral moments is matching the firstkraw moments – but raw moments hide the answer from the optimiser once a distribution is offset.E[X^k]is then dominated bygamma^kand the shape contributes only a fractional correction: 0.5% ofE[X^3]for a Gamma(3, 4) shifted by 10. Fitting three parameters off the third decimal place of a large number degenerates, and offset fits settled on parameters that matched the sample moments better than the true parameters did while being nowhere near them – a shape of 17.7 against a true 3.0, unchanged at any sample size.Central moments remove the offset by construction, so the shape is the whole of the third moment rather than a rounding error in it. An offset Gamma at n=5000 goes from
gamma=49.01, alpha=15.74, beta=9.07togamma=50.01, alpha=2.74, beta=3.75against a true(50, 3, 4), and the fit drops from up to 25 s to under a second. Unshifted fits are unaffected – Weibull, Gamma, Normal, LogNormal, Logistic, Gumbel and Exponential all agree with the previous results to at least four decimal places, because there the conditioning was never the problem.The terms are scaled by the sample’s own
sigma^k, so each is dimensionless: the mean in units of sigma, the relative variance error, then the skewness difference. The mismatch warning’s threshold moves from 1e-4 to 1e-2 to suit those units. Healthy fits land near 1e-12 when the moment equations have an exact solution and near 1e-3 when sampling noise means none exists and the optimiser returns the closest match; a fit that has actually failed sits near 0.5.Offset Gamma fits no longer start from a corrupted initial guess. Every offset-capable distribution returns the shift first in its parameter vector, because
_initial_guessoverwrites that slot with its own estimate of the shift.Gammareturned it last, so the overwrite landed on the shape parameter and destroyed it, while the initialiser’s own copy of the shift stayed behind in the scale slot: the seed came back as(offset, shape-ish, offset).Compounding it, the shape approximation was computed on the raw
x. On offset data the constant squashess = log(mean x) - mean(log x)towards zero, and since the shape grows like1 / 12sthe estimate exploded – 649 for a true shape of 3. The moments are now taken after the shift is removed.The consequences were silent wrong answers, not just slow ones. A 600-point sample from
Gamma(3, 4)shifted up by 10 fitted by MSE returned a negative shift of -1.35 with a shape of 63.8; another sample stopped after 0.03 s at the seed itself, reporting a scale equal to the offset. Both now recover the shift, and the fit is also 4x faster (6.2 s to 1.6 s) because the optimiser no longer has to travel back from a nonsense starting point. MLE was unaffected – it found its way regardless – and non-offset fits are untouched.Method of moments still disagrees on offset Gamma, but that is not a defect: its solution matches the sample moments better than the true parameters do (first three moments 10.76 / 116 / 1253 against the truth’s 10.75 / 115.7 / 1248). The three-parameter moment system with a threshold is close to non-identifiable, which is why
MOMis not among the offset methods exercised in the test suite.Fitted parameters change for censored *and* truncated data: the likelihood was unbounded there (#310). A censored observation is only ever known to lie inside its own truncation window – it could not have been observed otherwise – so its likelihood numerator has to be the probability of that intersection. Every likelihood in the package instead used the unconditional form,
F(x)for left censoring andS(x)for right, and divided by a separately accumulated window probability. That counts territory the truncation has already ruled out, so the contribution exceeds one, and the excess grows without limit as the fitted distribution’s mass slides out of the window.The consequence was a silent wrong answer. On 200 left-censored LogNormal points with a true
muof 0 and mild left truncation, the fit returnedmu = -7.81withneg_llof-infandres.successset toTrue. Across eight distributions, 21 of 48 censoring/truncation combinations returned a non-finite likelihood, reporting parameters such as a Weibullalphaof 1.3e81 or a Normalmuof -1.77e4. Regression was affected too, and failed more quietly: adding left truncation to a workingWeibullAFTfit returned an entirely plausible-looking parameter vector whose covariate coefficient had collapsed from a true 0.5 to 0.0018.Rather than teach each likelihood about truncation, such rows are now handed over as intervals: a left-censored row truncated at
tlbecomes[tl, x], a right-censored row truncated attrbecomes[x, tr]. The interval term is already a difference of CDFs, so it computes the correct numerator with no change to any likelihood function – which is also why interval-censored data never had the bug. One change inSurpyvalDatatherefore fixes the parametric, regression, Royston-Parmar and mixture likelihoods together.Only rows with a finite bound on the relevant side are recast, which is exactly where the defect lived. Untruncated fits are bit-identical: verified over 292 cases spanning ten distributions, seven censoring regimes, two sample sizes, weighted and unweighted, and three regression fitters, compared as exact float bit patterns. The restriction also keeps right censoring on the exact
log_sfpath, since expressing it as1 - F(x)loses all precision onceF(x)rounds to one – alog_sfof -49 comes back as-inf, and the optimiser does evaluate the likelihood that far from the data.The LogNormal case above now fits
mu = -0.56, matching the maximum of the correctly conditioned likelihood, and all 48 combinations return finite results. Right censoring combined with finite right truncation remains contradictory data and keeps its existing warning; what changes is that it now yields the coherent conditionalP(x < X <= tr)rather than an unbounded direction.``group_xcnt`` no longer walks every observation in Python. The step that collapses duplicate
(x, c, t)rows accumulated into a triple-nesteddefaultdict, one iteration per observation. That is linear but with a very large constant – around 13 microseconds an observation – which made it the single dominant cost of fitting once samples grew: 94% of a 50,000-point Normal fit, and two seconds at 100,000 points. It is now a sort plusnp.bincount. Fitted values are bit-identical.Group ordering is preserved exactly, which matters more than it appears:
xcnt_sortruns immediately afterwards and is a stable sort keyed onc,t.min(axis=1)andx, so any rows tying on all three keep whatever order grouping produced. Rows sharing anxandcwith differenttrbut equalt.min()are exactly such a tie, and a plain sortednp.uniquewould silently reorder them, so the original x-major nesting is reproduced instead. Integer counts also stay integer (np.bincountreturns float64), andnanentries keep their own groups as they did under dictionary keying.Measured end to end: a Normal fit at n=100,000 goes from 1137 ms to 125 ms (9.1x), Weibull at n=100,000 from 3890 ms to 926 ms (4.2x), and small fits improve too – Exponential at n=1000 from 5.9 ms to 2.4 ms. Kaplan-Meier at n=100,000 now takes 140 ms.
On top of that, grouping is now skipped entirely when there is nothing to group. Continuous measurements have distinct values, so every row is already its own group in input order and the operation is the identity – but the data handler runs it several times per fit regardless. Distinct values in the leading column of
xare enough to establish this (they make whole rows distinct whatevercandthold), which costs one sort of one column against three sorts of the full key. Only tied data – rounded, discrete, or heavily weighted – takes the grouping path now. Grouping 100,000 distinct points falls from 74 ms to 1.2 ms, taking the Normal fit above to 64 ms, Weibull to 505 ms, and Kaplan-Meier to 63 ms. Repeatednanvalues deliberately fail the check and fall through to grouping, sincenan != nanmeans they must stay separate.Exact closed-form maximum likelihood for the Exponential, Normal and LogNormal, where one exists. These have analytic MLEs – the Exponential’s events-over-exposure ratio, the Normal’s mean and standard deviation – so the fit no longer builds an initial guess or runs the five-optimiser ladder. Because the closed form is exact, the result is not merely faster but at least as good: verified that its log-likelihood is never worse than the optimiser’s, and its parameters agree to within the optimiser’s own convergence tolerance (~1e-8), as do its confidence bounds. At n=1000 a LogNormal fit goes from 135 ms to 10 ms, a Normal from 42 ms to 5 ms, and an Exponential from 11 ms to 5 ms; Weibull and the rest are untouched.
The applicability conditions are exact. The Exponential admits right censoring and left truncation (which only moves each unit’s exposure from
xtox - tl), but falls back to the optimiser for left or interval censoring and for right truncation, each of which makes the score transcendental. The Normal and LogNormal need complete, untruncated data: any censoring makes them the Tobit model and any truncation introduces a normal-CDF normaliser.Fixed: closed-form fits silently ignored ``lfp``, ``zi``, ``offset`` and fixed parameters. The hook that short-circuited to a distribution’s analytic MLE fired before any of these were checked, so
Uniform.fit(x, lfp=True)returnedp = 1.0andUniform.fit(x, fixed={"a": 0.0})ignored the held value – in both cases without warning. Requests carrying that structure now go to the optimiser, which estimates them.Fixed: ``Uniform`` fits had no usable log-likelihood.
Uniform.fit(x).aic()raisedAttributeErrorandcb()raised for want of a covariance. The log density is now defined directly rather than through the genericlog(hf) - Hfidentity, which isnanat the upper support edge (wheresfis 0) – exactly where the MLE putsb.neg_ll,aic,bicandaic_care now correct. No parameter covariance is offered, deliberately: the Uniform MLE is an order statistic sitting on the support edge rather than an interior stationary point, so the observed information is not positive definite and its inverse carries negative variances;cbrefuses rather than returning silentnanbounds.Tests: a breadth sweep over Turnbull’s supported inputs. Every combination of censoring type (observed, left, right, interval, and all four mixed), truncation form (none, left, right, both) and hazard estimator (Nelson-Aalen, Kaplan-Meier, Fleming-Harrington) is now fitted and checked for a converged, valid, monotone survival curve with a coherent risk-set ladder, complementing the existing tests that each pin one regime. The sweep also pins that the estimator choice is honoured, and that Fleming-Harrington coincides with Nelson-Aalen exactly when event times are distinct (its tied-event correction being the only difference). Building it surfaced #308.
Known issue: left censoring combined with left truncation (#308). Turnbull either raises “censoring interval does not intersect its own truncation window” on data where the intersection is plainly non-empty, or fails to converge and returns a degenerate estimate. A left-censored row lives on the
(-inf, x]bound, and the support-window intersection added for #273 drops that bound whenever an entry time sits above its lower edge, even though the event interval(tl, x]is non-empty. Only fits with both left-censored observations and left truncation are affected. The sweep marks this regimexfail(strict), so it will report as soon as it is fixed.Fixed: ``xcnt_to_xrd`` was quadratic in time and memory, and raised ``MemoryError`` past roughly 50,000 observations (#306). The at-risk entry count was built as an
N x Kcomparison matrix (observations × distinct times): 20,000 observations needed a 3.2 GB intermediate and ~15 s, and 50,000 attempted an 18.6 GiB allocation and failed. Because this conversion feeds every nonparametric estimator — and the MLE initial guess, which comes from probability plotting — the ceiling applied to most of the package:Weibull.fiton 50,000 points raisedMemoryErroreven though the likelihood itself was fine. The entry count is now computed in two linear branches: a constant when nothing is left truncated (the common case, where the matrix was entirelyTrueand merely recomputedn.sum()), and a sortedsearchsortedlookup otherwise.side="left"counts strictly-less-than exactly as the previous<did, so the(entry, exit]convention from #260 is unchanged, and integer counts make the cumulative sum exact — values are bit-identical. AWeibull.fitat n=10,000 goes from 1,836 ms to 182 ms; 200,000 points now fit in 5.6 s and a 500,000-point Kaplan-Meier in 8.1 s, where both previously failed.
v0.17.0 (1 August 2026)
Discrete distributions are now structurally separated from the continuous catalogue. A new
DiscreteParametricFitterbase class (Geometric, Poisson, DiscreteWeibull, NegativeBinomial, Binomial, Bernoulli/FixedEventProbability, BetaGeometric, andDiscretizewrappers) is the single home for what discreteness means when fitting: thediscretetrait, the centralsupports_mpp = False(each class previously set its own flag), and a new clear rejection ofhow="MPS"— spacings are increments of a continuous CDF and repeated integers make them degenerate, so this now raises instead of fitting nonsense. MLE, MSE and MOM behaviour is unchanged.``InstantlyOccurs`` and ``NeverOccurs`` are now first-class degenerate distributions in
univariate/parametric/distributions/degenerate.py(previously partial-API classes tucked intoparametric/__init__.py). They gain the missingdf/hf/qf/meanmethods and serialisation:to_dictstamps the schema andsurpyval.from_dictrestores the class itself (identity preserved, as the survival-tree leaves require). Historical import paths keep working.Simplification: one shared ``fit()`` skeleton for the parametric regression families (#295). PH, AFT, PO and parametric AH carried five copy-pasted versions of the same fit pipeline — data prep, the #251 param-map offset merge,
bounds_convert, optimisation, model assembly — which had already drifted (different optimiser ladders; only some families settingdist_params/phi_params). The skeleton now lives once in_fit_skeleton.py(each family supplies its optimiser strategy and covariate-link object), along with a singleLogLinearPhifor theexp(beta'Z)link previously defined inline in seven places. Fitted values are bit-identical; each family’s historical optimiser ladder and serialisation name tags are preserved exactly.Simplification: shared hazard identities and information criteria (#297, #298). The six
sf/ff/df/log_*identities derived fromHf/hfwere repeated in four regression fitters; they now live in oneHazardIdentitiesMixin. PH and parametric AHffnow use-expm1(-H)(matching AFT/AL), which is more accurate in the deep left tail whereHis tiny; all other values are bit-identical.neg_ll/aic/bic/aic_cwere duplicated betweenParametricandParametricRegressionModel; oneInformationCriteriaMixinnow serves both, preserving each class’s historicalaic_cparameter-count convention exactly.Simplification: NHPP likelihood data split hoisted (#296). The five-way censoring/interval split of recurrent-event data (the code that drifted into #288) was duplicated between the NHPP fitter and the proportional-intensity NHPP fitter; it now lives once as
RecurrentEventData.split_for_nhpp_likelihood. Fitted values are bit-identical.Simplification: numerical dedup batch (#299). The Aalen-Johansen
S(t-)incidence weighting (the pattern behind #253/#278) was implemented three times — nonparametricCompetingRisks, the competing-risks PHcifand Gray’s pooled CIF — and now lives once inaalen_johansen_iif(Gray’s pooled CIF is vectorised in the process). The Cox at-risk rule (entry-stricttl < tau, exit-inclusivex >= tau) is now documented in onecox_at_risk_maskhelper used by the exact-tie preparation and the Schoenfeld risk-set means, andCoxPH.baselinereplaces its O(K·N) Python loop with the same suffix-sum subtraction the Efron generator uses (values agree to ~1e-15 relative; pinned by the R/lifelines comparison tests). The degradationbootstrap_cbandbootstrap_cb_acceleratedmerged into one function (Z=Noneselects the plain path; the plain path now also drops non-finite refit curves instead of letting them poison the quantiles).CopulaModelserialisation is now round-trippable:to_dictstamps the schema version and a newfrom_dict(registered withsurpyval.from_dictunder the"copula"parameterization) rebuilds the model, where previously the dictionary was written in a form nothing could read. The CB transform sharing and Cox TVC wrapper collapse from #299 are deferred.Simplification: low-risk cleanup batch from the code-simplification review (#295-#299 track the medium-risk remainder). Dead code removed: the unused
surv_sksurv_transformationsmodule,init_from_bounds,_scaleandxcn_to_fslin utils,ParametricFitter. parameter_transform(would have crashed if called), unusedmpp_inv_x_transformmethods, commented-out blocks, the always-truehessflag in the Cox generator contract (generators now return(neg_ll, jac_hess)pairs), write-onlyfitting_infokeys, and the dead constant-rate methods on the PI-NHPP fitter. Duplication collapsed: a sharedSerialisableMixinnow providesto_json/from_jsonfor ~22 model classes; the degradation package imports the delta-method helpers fromrecurrent.inferenceinstead of carrying verbatim copies;fitand_fit_stratifiedshare one solve/p-value helper incox_ph.py;NonParametricCounting.from_xrdis the single home of the MCF estimator;predict_tvcreuses_tvc_cumhaz;mean_cbdelegates tormst. Plotting-position heuristics moved to dispatch tables.ParametricFittergains a default conditional-survivalcs(fixingAttributeErrorfor the discrete distributions); three docstring-free identitycscopies and Weibull’s less-stablelog_ffoverride were deleted. Convention fixes:validate_tv_coxphno longer double-validates (and now masks the truncation bounds alongside the data when covariate rows are dropped);RecurrentEventDataiterates statelessly and ordersitemsdeterministically; stalesurpyval.alphapointers updated.Removed: the alpha-stage ``SeriesModel``/``ParallelModel`` reliability-block composition (#284). Nested composition produced incorrect survival functions (
ParallelModel | ParallelModelreturned a parallel model; mixed-type composition flattened blocks instead of nesting them), and reliability block diagrams are covered by the Repyability package. Thesurpyval.experimentalshim now re-exports only the tree/forest models.Fixed: ``NonParametricCounting.mcf_cb`` corrupted bounds for off-grid queries (#285). The out-of-range masks were applied to the grid-length bound array before indexing by query position — zeroing the whole upper-bound column, wrapping out-of-range queries to the last grid value, and raising
IndexErrorwhen queries outnumbered the two bound rows. Bounds are now selected per query then masked (below-min → 0, above-max/negative → NaN, mirroringmcf), and two-sided output is now ordered[lower, upper], consistent with the parametriccif_cb(previously[upper, lower]).Fixed: ``CoxLewis`` constrained the log-intensity intercept to be non-negative (#286), silently pinning fits at
alpha = 0for any process with a baseline rate below one event per time unit. The intercept is now unbounded; a simulated(alpha, beta) = (-1, 0.05)process is recovered to(-1.006, 0.050).Fixed: recurrent-fitter batch (#288). A typo (
x[:, 0]forx_prev[:, 0]) cancelled the observed-event exposure term for 2-D event input without interval rows — degenerate[t, t]pairs now fit identically to 1-D input. The dead (and would-be-wrong) Cox-Lewis MCF correction in the simulator was deleted. The proportional- intensity HPP/NHPP fitters now honour a user-suppliedinit(previously silently overwritten) and validate its length.Fixed: round-2 follow-ups (#289). MPS tie densities are evaluated only at genuinely tied points (untied points contributed
0 * log(0) = NaNwhere a clean infinite penalty was intended); the additive-hazards kernel bandwidth falls back to the time scale when event times are (nearly) coincident instead of returning Dirac spikes; andBeta4.hfis 0 below andinfat/above the support instead of NaN above it.Fixed: Cox residuals, ``check_ph`` and robust standard errors now apply the Efron tie correction for Efron fits (#279). All residuals used plain Breslow risk-set means and increments regardless of the tie method, so heavily tied Efron fits (the
fit_from_dfdefault) disagreed with R/lifelines —check_phkm statistic 0.36 vs 0.72, robust SEs ~20% small. Schoenfeld residuals andcheck_ph(km, identity, log transforms) now match lifelines to 6+ figures under heavy ties; martingale and score residual sums vanish at the MLE for both tie methods; dfbeta correlates 0.999 with exact leave-one-out influence. The"rank"transform now uses average ranks for ties (R’scox.zphconvention; lifelines’ cumulative-count variant is nonstandard).Fixed: the concordance index credited 0.5 to a discordant event/censored pair tied in time (#276). Harrell’s C treats the censored subject as having outlived the tied event, so the pair is fully comparable: 1/0.5/0 by score order. Tie-heavy data was biased toward 0.5; results now match lifelines up to the (documented) both-events-tied-time convention difference.
Fixed: Lin-Ying additive-hazards ``hf``/``df`` added the baseline *jump* to a hazard *rate* (#277) — dimensionally incoherent, and as n grows the baseline vanished entirely (
hf -> beta'Z). The baseline rate is now a kernel-smoothed (Ramlau-Hansen, Epanechnikov) estimate from the corrected cumulative-baseline increments, with abandwidthargument.phi()on additive-hazards models now raises a clearNotImplementedError(the covariate effect is additive, not a multiplier) instead of anAttributeError.Fixed: competing-risks CIFs could exceed 1 with the default Nelson-Aalen method (#278). The Aalen-Johansen increment paired the discrete hazard
d/rwith the exponential survivalexp(-H); only the product-limit survival satisfies the telescoping identity, so total incidence reached 1.22-1.31 in small samples. Increments now always use the product-limitS(t-); the reportedsfkeeps the requested estimator.Fixed: distribution edge cases (#280). LogLogistic:
sf/ffare defined atx = 0(previouslyZeroDivisionError/NaN) andlog_sf/log_ffuse alogaddexpform that no longer overflows to-inffor largealpha**beta. Beta4:df/hfare 0 outside the support instead of arbitrary/negative/NaN values. Rayleigh, Gamma and Exponential custom probability-plotting paths now forward truncation bounds (previously silently dropped), and Rayleigh masks plotting positions at F = 1 (the ECDF heuristic returned NaN parameters). Uniform’s closed-form MLE rejects interval-censored data with a clear error instead of a crypticIndexError.Fixed: ``xrd_to_xcnt`` silently corrupted late-entry data (#281). A risk set that grows between observation times (left truncation) cannot be represented in xcnt output; the
np.absof the risk-set differences masked the increase and returned a different study. It now raises an informativeValueError.Fixed: container and robustness batch (#282).
SurpyvalData: scalar indexing on interval-censored data no longer flattens the interval row (IndexError), slicing carries covariatesZthrough, andto_xrdcaches per estimator instead of returning the first call’s result for every later estimator. Nonparametric models: scalarhf/dfreturn the step’s hazard increment instead of always NaN, and confidence bounds fall back to the point estimate when no point on the curve has a finite variance (single-observation fits returned NaN bounds).check_phno longer emits a spurious “ignoring left truncated values” warning for models fit with a constant entry column (tl = 0), and the stale pre-#260xcnt_to_xrddocstring example was updated.Fixed: the MPS estimator returned wrong parameters for censored, tied, truncated, and offset-truncated data (#268). Four defects: the censored/ties block was divided by a different count than the spacings, making the estimator inconsistent even without truncation (integer-tied Weibull data fit as
(13.7, 1.52)vs the true(10, 2)); censored survivor/CDF terms were not conditioned on the truncation window (truncated + censored fits biased to(11.7, 4.36)); offset fits passed unshifted truncation bounds to the shifted distribution (objective infinite at the true parameters); and interval-censored input crashed deep innp.hstackinstead of a clear validation error. The objective is now the Cheng-Amin sum form (spacings + tie densities + conditional censored terms in one sum), bounds are shifted with the data (clamped at the support), and interval data raises an informativeValueError. All four scenarios now track MLE to within ~2%.Fixed: censored/truncated Gamma and Beta fits had a silently corrupted Wald covariance (#270). The autograd shims for the incomplete gamma/beta functions stripped the derivative trace in their shape-parameter VJPs, zeroing every second-derivative contribution through a shape parameter: the stored covariance was wrong (12x the true sampling variance in one repro) and not even symmetric, corrupting
param_cb,cb, plot bands and the serialised covariance while the point estimates were fine. The shape derivatives are now traced primitives with numerical second-derivative VJPs, so autograd Hessians match the true observed information (verified against numerical differentiation to ~1e-6 for censored Gamma and Beta, including offset fits);mleadditionally validates Hessian symmetry and falls back to a numerical Hessian if a corrupted one ever reappears.Fixed: Turnbull excluded the right endpoint of interval- and left-censored observations from their support (#272). An interval
(l, r]whose right endpoint coincided with an exactly observed event time was forbidden from having failed atr(and a left-censored observation from having failed at its own bound), pushing its mass onto earlier atoms —sfbetween the atoms was 0.45 where the (l, r] NPMLE (Turnbull 1976, lifelines, icenReg) gives 0.83. Supports now include the atom at the right endpoint, matching the (entry, exit] convention adopted in #260.Fixed: Turnbull under truncation — variance ladder, support windows, and degenerate-interval inputs (#273). (1) The truncated variance ladder redistributed right-censored mass as fractional later events and kept censored items at risk via conditional tail probabilities — the anti-conservative mechanism #260 removed for untruncated data — and at the last event produced huge negative Greenwood increments that passed the finiteness guard. It now uses observed counts (events at exact atoms, censored items leave at censoring), reducing exactly to the delayed-entry Kaplan-Meier Greenwood ladder for exact + right-censored data. (2) Each observation’s support is now intersected with its own truncation window: mass can no longer be redistributed to times where the observed event provably cannot be, which previously drove the EM to a degenerate all-zero fixed point on valid left-censored + delayed-entry data — including the original #203 reproduction, which now converges to a healthy estimate. An empty intersection raises an informative
ValueError. (3) The KM-reducible variance branch now recognises exact + right-censored data expressed as degenerate intervals (xl == xr/xr = inf), which previously fell back to the anti-conservative expected-count ladder.Fixed: LFP fits with left truncation maximised an unbounded likelihood and returned degenerate parameters with optimiser success (#269). The truncation normaliser used
(p - f0) * (1 - F0(tl))— dropping the never-failing mass from the survival at entry — instead of the mixture survival1 - f0 - (p - f0) * F0(tl).Weibull.fit(x, c, tl=..., lfp=True)returnedalpha ~ 1e-42on healthy data; it now recovers the true parameters. Finite-bound windows (interval censoring, double truncation) are algebraically unchanged.Fixed: parametric PH ``random()`` sampled the wrong distribution and crashed with two or more covariates (#271). The sampler inverted
qf(U ** phi)where PH requiresqf(1 - U ** (1 / phi)), so every draw with a non-zero covariate effect came from the wrong distribution (empirical SF 0.67 vs model 0.89 in the repro); the covariate broadcast also raisedValueErrorfor multi-covariate models. Draws now reproduce the model’s ownsfand the returned covariates have shape(size, p).Fixed: Royston-Parmar silently returned NaN models (#274). The BFGS polish replaced the finite Nelder-Mead result even when it diverged (e.g. on doubly-truncated data); it is now kept only when finite and better, and a non-finite final likelihood raises. Quantile knot placement over too-few or tied event times produced coincident knots and an all-NaN model with no warning;
fitnow validates that the data contain at leastdf + 1distinct event times and that knots are distinct, raising an informativeValueError.Fixed: numeric MOM fits stopped far from the moment-matching solution (#275). The optimiser ran with
tol=1e-1and no convergence check, sohow="MOM"withoffset=Trueorfixedreturned e.g.beta ~ 3-4for truebeta = 2silently. The path now optimises tightly, polishes with Nelder-Mead when needed, and warns if the sample moments remain unmatched.Fixed: Cox delayed-entry / start-stop (TVC) fits had corrupted scores and Hessians whenever any covariate value was negative (#250). The left-truncation risk-set adjustment was forward-filled with
np.minimum.accumulate— valid for the scalar (positive, non-increasing) sum but wrong for the signed Z-weighted score and information sums, which it clamped to a stale running minimum. The optimiser could “converge” to a spurious zero of the corrupted score (wrong coefficients with no warning), and even rescued fits carried garbage standard errors, p-values,check_phand cluster-robust covariance. The adjustment is now an exact suffix-sum gather (not_yet_entered), valid for signed quantities; the analytic score and information now match numerical differentiation of the partial log-likelihood under delayed entry. All-positive covariates were unaffected.Fixed: parametric PH ``fixed={“beta_0”: …}`` silently pinned the first distribution parameter instead of the covariate coefficient (#251). The covariate parameter map was merged without the distribution-parameter offset (AFT/PO/AH were unaffected), so
WeibullPH.fit(x, Z, fixed={"beta_0": v})fixedalphatovand leftbeta_0free, corrupting the fit and its covariance. The map is now offset like the other regression families.Fixed: nonparametric competing-risks CIFs were systematically underestimated (#253). The Aalen-Johansen incidence increment weighted each cause-specific hazard by the survival after the jump,
S(t), instead ofS(t-)— with one cause and no censoring the CIF topped out at ~0.72 instead of 1. Cause-specific CIFs now sum exactly to1 - S(Kaplan-Meier weighting). The same correction applies to the Cox-path cause-specificcif. Also fixed: query times before the first observed event wrapped to the last step value (sf(0.1)on data starting at 1 returned the final survival instead of 1) in both the nonparametric and Cox-path predictors, andCompetingRisks.fit_from_dfstored the source DataFrame asmodel.df, shadowing the density method — it is nowmodel.source_df.Fixed: likelihood-ratio confidence bounds ignored user-fixed parameters (#255). Profiling silently re-freed a parameter fixed at fit time, letting the profile drop below the fitted negative log-likelihood and inflating the interval several-fold (
Weibull.fit(x, fixed={"beta": 5})gave analphaLR interval ~5x the Wald width). Fixed parameters now stay pinned during both the parameter profile and the function-band constrained search, and requesting an LR bound on a fixed parameter raises a clearValueError.Fixed: MixtureModel likelihood and EM corrections (#254). Counts from grouped/tied data were applied as per-component likelihood powers before mixing (
sum w_i f_i^n != (sum w_i f_i)^n), so any tied data (e.g. rounded measurements) silently skewed the mixing weights — a true 50/50 Weibull mixture fit as 14/86. Counts now multiply the mixture log-likelihood; the mixing-weight update is count-weighted; the M-step now minimises the proper EM Q-function (responsibilities times component log-likelihoods) instead of an ad-hoc responsibilities-as-weights objective; and the interval-censored contribution wasF(l) - F(r)(negative) — nowF(r) - F(l). Truncation (tl/tr/t) was accepted but silently ignored; truncated data is now fitted by direct maximum likelihood on the truncation-corrected observed likelihood (the window couples the components, so label-based EM does not apply). Also fixed:xl/xr-only input crashed onlen(None), anddfcrashed on integer input.Fixed: LFP / zero-inflated / offset parametric model conventions made mutually consistent (#256).
df/hffor combined LFP+ZI models used(1 - f0) * pwheresf/ffand the likelihood use(p - f0)— the density did not integrate to the failure probability.mean()andmoment()ignoredf0entirely.qf/randomplaced the zero-inflation mass at the offsetgammawhiledf/ffplace it at 0, soqfdid not invertff. Offset models returnedff < 0/sf > 1/ NaNs belowgamma— now clamped to the boundary values.cbreturned NaN where the point estimate sits on the boundary (sf == 1, e.g.t <= gamma) — now the boundary.random()crashed for LFP models when the binomial draw produced zero failures.aic_cpenalised a different parameter count thanaic. The numerically stable left-censored likelihood branch was unreachable (invertedf0check).Fixed: distribution-level defects (#257).
LogNormal.fitcrashed for any data with geometric mean < 1 (the locationmuwas wrongly bounded positive).Bernoulli.fitwas broken for essentially every input (it broadcastxagainst the literal[0, 1]and mishandledn=None). Offset MPP fits withrr="x"mis-inverted the regression for Exponential and Gamma (silently wronglambda/gamma); the Gamma offset MPP also seeded its shape search from unshifted-data moments and now multi-starts it. Gamma’s censored non-offsetrr="x"crashed on a length mismatch.ExpoWeibull.sf/Hf/log_sfunderflowed to 0/inf/-inf in the (reachable) right tail — rewritten in a cancellation-freeexpm1/log1pform. Probability-plot y-axis inverse transforms were not inverses for Exponential, GumbelLEV and Beta (silently mislabelled plot axes).Logisticlog-functions overflowed in the deep tail (nowlogaddexp).ExactEventTime.fitwithout both censoring sides now raises an informative error.Fixed: formula fits with a categorical covariate were non-identified — categoricals are now reference-level coded (#252).
fit_from_df(..., formula="age + sex")used to expandsexinto a full one-hot (sex[F],sex[M]) whose columns sum to a constant — exactly collinear with the baseline distribution’s scale (or the Cox baseline), so the likelihood was flat along a ridge and the reported coefficients and standard errors were optimizer-path noise (predictions were unaffected, which is why it went unseen). Formulas are now materialised with their implicit intercept, giving categoricals standard treatment coding, and the intercept column is dropped (the baseline provides it).Migration note: feature names and coefficient meanings change for formula fits with categoricals —
['sex[F]', 'sex[M]']becomes['sex[T.M]'], and the coefficient is the log-hazard-ratio (or equivalent) of that level versus the reference (first) level, matching R, lifelines and statsmodels. Predictions from refitted models are unchanged. An explicit"0 + ..."formula opts back into full-rank coding. This also fixes theLinAlgErrorcrash in Buckley-James formula fits with categoricals.Fixed: regression serialisation and robustness batch (#261). Buckley-James and Lin-Ying additive-hazards models now persist their formula encoder state (the #244 treatment), so restored models predict from DataFrames with transforms/categoricals; repeated save/load cycles of a parametric regression model no longer silently drop the stored covariance;
ParametricRegressionModel.random(broken on every path — it ignoredZ) now dispatches to the fitter’s covariate-aware sampler; theAcceleratedLifefitter is no longer stateful across fits, keeps user-fixed parameters inmodel.fixed(SEs were reported for constrained parameters), and accepts 1-D stress vectors;WeibullPH.fitaccepts plain-list covariates;fit(init=<ndarray>)no longer crashes; deserialised univariate models supportbic/aic_c/re-serialisation and carry their support interval; interval/left-censored observations below the distribution’s support are rejected at validation instead of producing a NaN likelihood and a silent initial-guess “fit” (whose reported likelihood now matches its returned parameters); and invalidcb/param_cbarguments raiseValueErrorinstead ofUnboundLocalError.Fixed: TVC prediction and alignment (#259). Predicting along a covariate schedule treated intervals as
[xl, xr), so a baseline-hazard jump exactly at a covariate-change time was weighted by the new covariate while the fitted likelihood uses(xl, xr]— predictions now match the fit, and a query time returns the same value regardless of the other query points. Cluster-robust standard errors on start-stop (TVC) fits now permute user-supplied per-row cluster labels into the internal row order (previously silently misassigned unless the input was already sorted), and default to clustering by subject. An exactly singular information matrix now degrades to the pseudo-inverse/NaN path instead of crashing.Fixed: AFT time-varying-covariate fits now refuse delayed entry and observation gaps instead of silently dropping the missing exposure (#258). The accumulated accelerated age
psi(T)integrates the covariate path from time 0; a subject entering observation late (or with gaps) has unobserved covariates over the uncovered window, and the likelihood previously treated that time as contributing zero ageing — shifting every subject’s window by +5 returned bit-identical parameters. Correct conditioning would require the unobserved pre-entry covariate path, so rather than guess it the fit raises an informative error pointing to Cox TVC (CoxPH.fit_tvc), which handles delayed entry and gaps exactly.Fixed: frailty models handle the ``theta -> 0`` (no-frailty) limit (#262). A frailty variance that underflows to zero — frailty-free data, or a restored model — gave NaN marginal predictions (division by
theta) and a NaN/crashing Wald interval; the marginal now takes the well-defined proportional-hazards limiteta * H0, and a boundary estimate returns a zero-width interval instead of dividing by zero.Fixed: Turnbull confidence intervals and the delayed-entry risk-set convention (#260). On plain right-censored data (where Turnbull reduces exactly to Kaplan-Meier) the variance was computed from the EM’s expected-count ladder, which redistributes censored mass as fractional later events and silently understated it — confidence intervals were anti-conservative (e.g. Var(H) 0.47 vs the correct Greenwood 0.63). The variance now uses the observed-count ladder in that regime and matches Kaplan-Meier’s Greenwood intervals exactly; genuinely interval-censored data keeps the expected-count approximation (use
bootstrap_cbfor calibrated intervals there).Convention change: delayed-entry risk sets now follow the standard
(entry, exit]convention (Rsurvival/ lifelines): a subject entering observation exactly at an event time is not at risk for that event. Kaplan-Meier/Nelson-Aalen previously counted it, disagreeing with Turnbull’s NPMLE on identical data; the two now agree. Fits only change where an entry time exactly ties an event time. Consistently, a value at exactly its own left-truncation time (a zero-length observation window) is now rejected at validation instead of silently distorting the estimate, andTurnbull.fit(..., max_iter=0)raises instead of crashing. The truncated-fit degeneracy detector now inspects only the identifiable region, so partial collapses are reported as degenerate rather than as generic non-convergence.Changed: the proportional-hazards test now uses the standard Grambsch-Therneau forms (#262). The per-covariate statistic is
d (Vu)_j^2 / (Sgc2 V_jj)withVthe inverse information — the form used by R’scox.zphand lifelines — replacing the previous information-diagonal variant (both are valid chi-square screens, but they weight cross-covariate information differently, so surpyval could flag a different covariate than R/lifelines on the same data). The"km"time transform is now the true1 - KM(t)fit on the full data (censoring included) rather than the censoring-blind ECDF of event times.check_phnow matches lifelines to numerical precision (verified against lifelines 0.30.3); reported per-covariate statistics change for multi-covariate models. The global test was already the standard form and is unchanged.Royston-Parmar flexible parametric models.
RoystonParmar.fit(x, c=..., df=..., scale=...)fits a flexible parametric survival model that replaces the straight log-cumulative-hazard-vs-log-time line of a Weibull with a restricted cubic spline, giving a smooth, fully parametric baseline of arbitrary shape – flexible like a Cox baseline but extrapolable like a parametric one. Three link scales:"hazard"(proportional hazards;df= 1 is a Weibull),"odds"(proportional odds), and"normal"(probit;df= 1 is a log-normal). Knots are placed at quantiles of the event log-times by default (or supplied explicitly), and beyond the boundary knots the spline is linear, so the model extrapolates with a Weibull-like tail – which pairs naturally with the restricted-mean survival time added in 0.16. The fittedRoystonParmarModelexposessf/ff/hf/Hf/df/qf/random/mean, a linear-predictor confidence band (cb),aic/bicfor choosingdf, andto_dict/from_dict. The likelihood supports the full arbitrary censoring/truncation surface – observed, right-, left- and interval-censored observations (passxl/xror 2-elementxrows), with left- and/or right-truncation (tl/tr/t) and observation weights (n).Shared-frailty proportional-hazards models (Gamma frailty). A new
Frailty(distribution)factory (with pre-builtWeibullFrailty,ExponentialFrailty,LogNormalFrailty,GammaFrailtyinstances) fits a proportional-hazards model with a random hazard multiplier shared within a group –h(t | Z, u) = u h0(t) exp(beta'Z),udrawn once per group from a Gamma of mean 1 and variancetheta..fit(x, Z, c, groups=...)and.fit_from_df(..., group_col=...)maximise the closed-form marginal likelihood (the Gamma frailty integrates out per group), so it captures unobserved between-group heterogeneity and the within-group correlation it induces – the conditional/random-effects complement to the cluster-robust standard errors added in 0.16. The fittedFrailtyModelreports the frailty variancetheta(with a Wald CI), the per-group posterior (empirical-Bayes) frailties, and predicts either marginally (population-averaged, the default –S = (1 + theta e^{beta'Z} H0)^{-1/theta}) or conditionally on an observed group or a supplied frailty value viasf(x, Z, group=...)/sf(x, Z, frailty=...). OmittingZgives a pure random-effects survival model. Serialises withto_dict/from_dict. Gamma frailty only for now; log-normal, Cox, and nested/hierarchical frailty are planned.Fixed: formula-fit regression models now round-trip through serialisation (#244). A regression model fit with
fit_from_df(..., formula=...)using a categorical term dropped its design-matrix transformer onto_dict/from_dict, so a restored model failed to evaluate from raw covariates (['sex[F]', 'sex[M]'] not in dataframe columns).to_dictnow persists the categorical factor levels and numeric column names, andfrom_dictrebuilds an equivalentformulaicmodel spec, so a restored model expands raw covariates identically to the original – for the parametric families (PH/AFT/PO/AH) and Cox. Data-dependent transforms (scale()/center()) keep fitted statistics that cannot be restored from levels, so serialising such a formula now raises early atto_dictrather than round-tripping to a silently wrong encoding.Likelihood-ratio confidence bounds on model functions.
cbgains the samemethodargument:method="lr"returns a profile-likelihood band onsf/ff/Hf/hf/df. At each time the bound is the extreme value of the function over the parameter confidence region \(\{\theta : 2[\text{nll}(\theta) - \text{nll}_{\hat{}}] \le \chi^2_1\}\), found by constrained optimisation with a warm-started sweep over the time grid. Like the parameter version it is transformation-invariant and better behaved in small samples than the Wald/delta band, needs the original data, and does not yet cover offset / LFP / ZI models.Likelihood-ratio confidence bounds on parameters. A fitted parametric model’s
param_cbgains amethodargument:method="wald"(the existing default) ormethod="lr"for a profile-likelihood (likelihood-ratio) bound. The interval is the set of parameter values whose profile deviance stays below the \(\chi^2_1\) critical value, with the remaining parameters re-optimised at each candidate. Unlike the Wald bound it is transformation-invariant, respects the parameter’s support boundary, and need not be symmetric about the estimate – usually better small-sample coverage, and the reliability-engineering default. It needs the original data (a deserialised model raises, directing you tomethod="wald"); offset / LFP / ZI models are not yet supported.
v0.16.0 (22 Jul 2026)
Diagnostics & validation
Cox model diagnostics (#211). A fitted
CoxPHmodel now exposescompute_residuals(kind=...)– Schoenfeld, scaled Schoenfeld, martingale, deviance, score and dfbeta residuals – andcheck_ph(), the Grambsch-Therneau proportional-hazards test (a per-covariate and a joint global test against a transform of time; a smallp-value is evidence against proportional hazards). All residuals respect delayed entry (tl) and count weights. The residual identities are exact at the MLE (Schoenfeld, score and martingale residuals sum to zero) and the PH test is validated for both power (it detects a genuine time-varying coefficient) and calibration (its p-values are ~Uniform under true proportional hazards).Restricted mean survival time (#213). A fitted non-parametric model (e.g.
KaplanMeier) gainsrmst(tau)– the area under the survival curve to a horizon with its standard error and confidence interval – and the package-levelsurpyval.rmst_diff(model_a, model_b, tau)compares two groups’ RMST (difference, ratio, CI and a two-sided p-value). The RMST-difference is the assumption-light alternative to the hazard ratio when proportional hazards fails; the estimate matches its analytic value and the two-group test is calibrated under the null.Cluster-robust standard errors (#215).
CoxPHmodels gainrobust_covariance(cluster=...)androbust_summary(cluster=...)– the Lin-Wei sandwich variance for clustered / correlated data (repeated events per subject, grouped sampling), built from the dfbeta residuals. On independent data it agrees with the model-based errors; on exactly replicated clusters it inflates by the theoretically exactsqrt(cluster size).Gray’s test (#216). The package-level
surpyval.gray_testcompares cumulative incidence functions across groups for a specified cause in the presence of competing risks – the subdistribution analogue of the log-rank test. Unlike a cause-specific log-rank, it keeps competing-cause failures in the risk set with an inverse-probability-of-censoring weight, so it tests the CIFs directly. Returns a chi-squared statistic, degrees of freedom and p-value. Validated for calibration under the null (including under heavy censoring, which exercises the IPCW weighting) and for power against genuine CIF differences.Stratified Cox and stratified log-rank (#214).
CoxPH.fit/fit_from_dfacceptstrata(orstrata_col) to fit a stratified proportional-hazards model: a separate baseline hazard per stratum with shared coefficients, the partial likelihood summed within strata. Prediction (sf/Hf/…) then takes astratumargument to select that stratum’s baseline.surpyval.logrankgains astrataargument for the stratified log-rank test (per-stratum observed-minus-expected and variance summed before forming the statistic). Both are the standard remedy when proportional hazards fails for a nuisance covariate. Validated by simulation: the stratified estimators recover the truth (and stay calibrated) in a confounded design where the pooled versions are badly biased / over-reject, reduce exactly to their unstratified counterparts with a single stratum, and the stratified Cox partial likelihood factorises into the per-stratum contributions.Prediction-validation metrics (#212). A new
surpyval.metricsmodule scores a predicted survival function against right-censored outcomes with inverse-probability-of-censoring weighting:brier_score/integrated_brier_score(the time-dependent Brier score of Graf et al. 1999 and its integral – calibration and discrimination together, lower is better) andauc_td(Uno’s 2007 cumulative/dynamic time-dependent AUC – discrimination as a function of the horizon). All are model-agnostic; thesurvival_probabilityhelper builds the required survival matrix from any fitted model exposingsf(x, Z)(the parametric regression families,CoxPHand thebeta.mlforest), giving the ML-flavoured workflow its first proper validation-and-comparison story. Validated against known answers: without censoring the Brier score is exactly the mean squared error; a well-specified model beats the marginal Kaplan-Meier reference (and a constant predictor is worse); and the AUC is ~1 for a near-perfect ordering and ~0.5 for a random one.
Correctness
Turnbull EM under truncation (#203). Three statistical defects in the truncated Turnbull NPMLE are fixed. (1) The EM now iterates with the Kaplan-Meier self-consistency update (
pproportional to the expected countsd), the canonical M-step; theFleming-Harrington/Nelson-Aaleninner estimators setR = exp(-H), which violates that fixed point and left even healthy truncated fits reporting tol-level non-convergence – they now converge, and the requested hazard-form estimator is applied to the converged ladder. (2) The expected counts are confined to the identifiable support each iteration, stopping the ghost step from migrating mass below every entry window. (3) The convergence check is no longer NaN-blind: a non-finite update or a total mass collapse is detected as a degenerate, non-identifiable fixed point and reported with an explicit warning and adegenerateflag on the model, instead of a silent all-zero survival curve. Untruncated fits are unchanged. Validated: the issue’s degenerate reproduction is now flagged and warned; a left-truncated sample recoversS(median)to within 0.04 with all three inner estimators; and the documented untruncated example is byte-for-byte identical.
Degradation
Destructive degradation modelling (#153). New
surpyval.degradation.DestructiveDegradationfor tests whose measurement destroys the specimen, so each unit yields a single(time, degradation)point (material/adhesive strength, breakdown voltage, …). With no per-unit paths to fit, the population degradation distribution is modelled directly as a location-scale regression on a time transform,Y | t ~ dist(loc = β₀ + β₁·φ(t), σ)(LogNormalorNormal;φ= linear / log / sqrt / reciprocal, ortransform="best"by AICc), and the lifetime distribution is induced by crossing the failure threshold (sf/ff/Hf/df), with the increasing (wear) vs decreasing (strength-loss) direction inferred automatically. Censored measurements (a strength below the test floor, a specimen that did not break) are handled through the ordinarycconvention;cbgives bootstrap bounds and the model round-trips throughto_dict/from_dict. This completes the degradation half of #153 alongside the stochastic-process models.
Regression
Time-varying-covariate fitting for accelerated failure time (#150).
WeibullAFT(and everyAFT(dist)) gainsfit_tvc/fit_tvc_timelineand the DataFrame variants, taking the same start-stop / timeline input (i/xl/xr/c) as the other families. Because AFT rescales the time axis, a subject’s likelihood depends on its accumulated accelerated ageψ = Σ exp(β'z)(b − a)across intervals and does not factorise into independent left-truncated rows the way the proportional/additive-hazards families do, so it is fit with a dedicated accumulated-age likelihood (a within-subject scan each optimiser step) rather than the reshape-and-refit used for PH/AH. The shared MLE code is untouched: the fit binds the custom likelihood onto its own result object, so confidence bounds (a numerical Hessian of that likelihood) are correct, and information criteria are reported on the subject count rather than the episode rows. This closes the last open part of #150; with #170’s evaluation side, AFT now has full time-varying-covariate support.Evaluate a fitted regression model along a time-varying covariate path (#170). A fitted
WeibullPH(anyPH(dist)),WeibullAH(anyAH(dist)) orWeibullAFT(anyAFT(dist)) gainssf_tvc(andHf_tvc): given a piecewise-constant covariate scheduleZ(t)it returns the resulting survivalS(t), with an optionalgiven=age for conditional survival. For proportional and additive hazards the cumulative hazard is additive over disjoint intervals, so the survival along a step path is the exact sum of the per-segment increments; for accelerated failure time the path instead accumulates an accelerated ageψ(x) = Σ exp(β'z)·(b − a)fed once through the baseline. Either way it reduces to ordinarysffor a constant covariate. The covariate path is described by a newStepSchedule, built structurally (from_changepoints/from_intervals/cyclicfor duty cycles) or from a step-valued expression string int(from_expression, e.g."0.9 if t % 24 < 8 else 0.3"or"0.3 * 2 ** floor(t / 1000)"). Expressions are proved piecewise-constant from their syntax tree before evaluation –tmay reach the value only through a quantizer (floor/ceil///) or a comparison – so a continuously-varying covariate (0.3 + 1e-4 * t,sin(t)) is rejected withStepValuedErrorrather than silently returning a wrong answer.sf_tvcmay be given(xl, Z)arrays directly or aStepSchedule. The semi-parametricCoxPHgains the samesf_tvc/Hf_tvcandStepScheduleconvention (summing the fitted baseline-hazard jumps along the path); the existing interval-orientedpredict_tvcis unchanged andsf_tvcagrees with it exactly. Only proportional odds does not yet expose a time-varying-covariate evaluation and raises.Time-varying covariates for the parametric PH and additive-hazards families (#150).
WeibullPH(and everyPH(dist)) andWeibullAH(everyAH(dist)) gainfit_tvc/fit_tvc_timelineand the DataFrame variants, taking the same start-stop / timeline input asCoxPH.fit_tvc(i/xl/xr/c, surpyval’s censoring convention). For these families the cumulative hazard is additive over time intervals, so a time-varying-covariate subject factorises exactly into one left-truncated observation per constant-covariate interval; the fitter simply reshapes the data and reuses the ordinary parametric MLE, giving the same fit as the equivalent non-time-varying data. Accelerated failure time and proportional odds do not compose this way (they need an accumulated accelerated age / have no additive structure), so they do not exposefit_tvc.Timeline (xicnt-style) input for time-varying-covariate Cox.
CoxPH.fit_tvc_timeline/fit_tvc_timeline_from_dfaccept a covariate timeline – one row per covariate change per subject (i,x,Z,c) with the terminal event / censoring on the subject’s last row – as an alternative to writing explicit(xl, xr]intervals forfit_tvc. Each covariate value holds from its time until the subject’s next row, the first time is the (delayed-)entry time and the last is the exit; the timeline is expanded to start-stop intervals and fitted identically, so it gives the same fit as the equivalentfit_tvcdata.Time-varying-covariate Cox input harmonised to the surpyval convention. The start-stop interface (
CoxPH.fit_tvc/fit_tvc_from_df/predict_tvcandhandle_tvc) is renamed to match surpyval’s vocabulary: the subject id isi(wasident), the interval bounds arexl/xr(werestart/stop), and the status isc(wasevent).cnow follows the standard surpyval censoring convention –0= event atxr,1= right-censored – which is the inverse of the oldeventflag (event=1->c=0). The DataFrame entry point’s columns are namedxl_col/xr_col/c_colaccordingly. Positional calls are unaffected; keyword calls and theeventvalues need updating.Accelerated Life with an Exponential distribution now fits.
AcceleratedLife(Exponential, life_model).fit(...)raisedKeyError: 'lambda'because the life-parameter map named the Exponential’s parameter"lambda"while the distribution actually calls it"failure_rate". The name is corrected (thelife <-> ratetransforms were already right), so Exponential accelerated-life models fit, predict and serialise; a guard test now checks every distribution’s declared life parameter is a real parameter of that distribution.Exact and Kalbfleisch-Prentice tie handling for Cox (#142).
CoxPH.fitgains two furthermethodchoices beyond'breslow'and'efron':'exact'(the average-over-orderings exact partial likelihood, for ties that arise from coarse rounding of an underlying continuous time) and'kalbfleisch-prentice'(alias'kp'– the exact discrete / conditional-logistic likelihood, for genuinely discrete time). Both honour delayed entry (tl), stratification and count weights, and reduce to Breslow/Efron when there are no ties. The KP denominator is the elementary symmetric polynomial of the risk-set scores, computed by the standard polynomial recursion; the exact term is summed over tied-death orderings by anO(2^d)subset recursion, which is guarded against oversized tie sets. Validated by matching a brute-force per-tie likelihood exactly, and by score/Hessian agreement with finite differences. These methods are niche – Breslow and Efron already match what R’ssurvivaland lifelines use by default – and correspondingly more expensive under heavy ties.
Serialisation
Survival tree & forest serialisation (#191).
SurvivalTreeandRandomSurvivalForestnow implementto_dict/from_dict(andto_json/from_json), completing the serialisation campaign that had deferred them while the forest was crash-prone. A tree serialises as its recursive node structure with each leaf stored as its own fitted model (Parametric/NonParametric, or a sentinel for the emptyNeverOccursleaf), so a restored tree predicts identically without re-fitting; a forest is the ensemble settings plus its trees. Both carry a"model"class tag and dispatch through the package-levelsurpyval.from_dict/surpyval.from_json, are schema-stamped, and are BSON-native for MongoDB. In the course of this, a latent leak was fixed inParametric.to_dict:_neg_ll(always) andgamma/p/f0(for offset / LFP / zero-inflated models) were emitted as NumPy scalars, which MongoDB’s BSON encoder rejects; they are now native floats.Accelerated Life model serialisation. Fitted Accelerated Life parameter-substitution models (
AcceleratedLife(dist, life_model)) now round-trip throughto_dict/from_dict/to_json/from_jsonand the package-levelsurpyval.from_dict. Previously only the fixed-form covariate families (AFT, PH, PO, AH) serialised and any Accelerated Life model raisedNotImplementedError. The model is rebuilt from the stored distribution and built-in life-model names (Power,Eyring,Linear, the Arrhenius-styleExponential, the dual-stressDualPower/DualExponential/PowerExponential, and their inverses), so the restored model predicts identically and, when a covariance was stored, reproduces the same confidence bounds. A genuinely custom life model (whose parameterisation is not a fixed name map) is still refused with a clear error.
v0.15.2 (20 Jul 2026)
Data handling
xcnt_handlernow warns when right-censored observations carry a finite right-truncation time (#195). The combination is contradictory – right truncation means the unit was only observable because its event occurred beforetr, while right censoring says the event is after the censoring time – and such rows can make truncation-adjusted likelihoods unbounded.
Serialisation
RenewalModel.from_dictnow validates that the stored distribution name resolves to a genuine distribution fitter (#206), matching the guard used by every other reader, so an untrusted document cannot resolve arbitrary package attributes.
Misc
The bundled dataset loaders use pandas’ default (C) CSV engine instead of
engine="python"(#207) – identical parses, faster, and one less thing for security scanners to worry about; the loaders are now covered by tests.Modernised the documentation build toolchain (
docs/requirements.txt): the 2022-era pins (sphinx 5.3,jupyter-sphinx 0.4) leftipykernelunpinned, and against current ipykernel 7 the notebook execution hangs or crashes – one of the reasons hosted docs builds kept failing. The new set (sphinx 8.2, sphinx-rtd-theme 3.1, jupyter-sphinx 0.5.3, ipykernel capped below 7) is fully pinned and validated by a complete docs build in a clean virtualenv.
v0.15.1 (20 Jul 2026)
Non-parametric
Fixed Turnbull fits with truncation hanging indefinitely (this also hung the documentation builds, which is why the hosted docs went stale). The Fleming-Harrington tie ladder (
fh_h/fh_var_h) was a per-event Python loop; the Turnbull EM feeds it fractional expected event counts which, under heavy truncation, can grow without bound between iterations – the loop then effectively (or with an infinite count, literally) never returned. The ladder is now evaluated in closed form (digamma/trigamma harmonic sums) beyond a small exact loop, so its cost is O(1) in the event count: identical results for ordinary tie counts, and pathological counts now yield a diverging hazard (inf) instead of a hang. Note that the truncated NPMLE itself remains delicate on small or heavily truncated samples (it can be non-identifiable and the EM converges to a degenerate estimate); such fits now terminate and are flagged, and the docs note the caveat.
v0.15.0 (20 Jul 2026)
Serialisation
Every serialised model dictionary now carries a schema version (
"schema": 1), stamped by everyto_dict. The version is bumped only when a dictionary’s shape changes incompatibly, so documents stored today (in files or MongoDB) stay recognisable to future SurPyval versions: the package-levelsurpyval.from_dictrefuses documents written by a newer schema with a clear error, and treats documents with no"schema"key (written before versioning) as schema 0, which remains loadable.MongoDB compatibility, verified for every serialisable model: BSON is stricter than JSON (numpy integer scalars and arrays are rejected, and dictionary keys must be strings), so every model’s
to_dictoutput is now tested through the full MongoDB path –bson.encode(whatinsert_onedoes), decode, add the_idfieldfind_onereturns, and restore viasurpyval.from_dictwith predictions reproduced. The cause-label fields of the competing-risks containers are now normalised to native Python types with a newsurpyval.serialisation.to_nativehelper (numpy labels passed by the caller no longer leak into the document), andpymongowas added to the test dependencies for the BSON round-trip tests.Added package-level readers for serialised models:
surpyval.from_dict(model_dict)andsurpyval.from_json(fp)restore a model of the right class from any model’sto_dictdictionary /to_jsonfile, so the caller no longer needs to know which class wrote it. Dispatch reads the serialised dictionary itself: the"model"class tag written by most models, or the"parameterization"marker ("parametric","non-parametric","parametric-regression") of the core univariate families. The class-level readers are unchanged.
Package structure
Pre-stable models are now tiered by maturity:
surpyval.alpha(exploratory; the interfaces may change or disappear – currently theParallelModel/SeriesModelsystem models, previously insurpyval.experimental) andsurpyval.beta(functionally complete and tested, interface not yet part of the release contract – the survival tree and random survival forest insurpyval.beta.ml).surpyval.experimentalremains as a deprecated re-export of both and warns on import.
Machine learning
The survival tree and random survival forest graduated from
surpyval.experimentalto the beta tier:from surpyval.beta.ml import SurvivalTree, RandomSurvivalForest. The oldsurpyval.experimentalimports still work as re-exports. Their test suite now runs in CI, expanded with behavioural and structural tests: prediction coherence (ff = 1 - sf,Hf = -log(sf), monotone boundedsf),max_depth/min_leaf_samples/min_leaf_failuresguarantees, seeded determinism, degenerate inputs (all-censored, constant covariates, tiny samples, tied times, count weights), forest ensemble maths (the forestsfis exactly the tree average; the"Hf"method averages cumulative hazards), prediction shapes, mortality ordering and a concordance sanity check.Fixed the concordance index (
surpyval.utils.score.score, used byRandomSurvivalForest.score): pairs were ordered by censoring flag instead of by time before comparison, which pushed the c-index of even a strongly informative forest towards 0.5. Pairs are now ordered by time (event first on exact ties), soscorereturns Harrell’s c-index for mortality-like scores (1 = perfectly concordant).forest.scorealso now respects itstie_tolargument.
Competing risks & mixtures
Added serialisation to the competing-risks and mixture models:
MixtureModel(EM mixture of a base family),FineGrayModel(subdistribution-hazard regression),ParametricCompetingRisks(one distribution per cause) and the nonparametricCompetingRisksnow haveto_dict/from_dictandto_json/from_json. The mixture stores its base-family name, component parameters and weights; Fine-Gray stores its coefficients, covariance and subdistribution-baseline step arrays; and the competing-risks models store their per-cause sub-models (via each cause’s ownto_dict) or per-event step arrays. Every reloaded model reproduces its predictions exactly.
Degradation
Added serialisation to the fitted degradation models:
DegradationModel, the stochastic-process modelsWienerProcessModelandGammaProcessModel, and the Monte-CarloInducedFailureDistributionnow haveto_dict/from_dictandto_json/from_json. The process models store their few parameters; the induced distribution stores its samples (theinfnever-fails draws are written asnullso the result is valid JSON); andDegradationModelstores its raw data, the path model (by name) and per-unit fits, the population summaries, and the fitted life model (via its ownto_dict– plain or accelerated), so the reloaded model reproduces its predictions and per-unit paths and (because the data is kept) its bootstrap confidence bounds too.
Recurrent events
Added serialisation to the renewal / imperfect-repair models (
RenewalModel): the generalized-renewal (Kijima-I/II), G1 renewal, ARA and ARI families now haveto_dict/from_dictandto_json/from_json. These processes have no closed-form intensity (their MCF comes from a sampler closure that cannot be pickled), so the dict stores the family, the underlying distribution (by name) and its parameters, the restoration parameter and the family option (kijima_typeor memorym); on load the family’s fitter rebuilds the sampler from those, so the simulated MCF reproduces exactly. This completes serialisation coverage of every non-experimental fitted model in the package.Added serialisation to the fitted recurrent-event models:
ParametricRecurrenceModel(NHPP/HPP intensity fits),NonParametricCounting(the MCF estimate),ProportionalIntensityModel(proportional-intensity regression), and the competing-risks containersCauseSpecificMCFandCauseSpecificNHPPnow haveto_dict/from_dictandto_json/from_json. The intensity model is stateless, so each stores its name plus the fitted parameters (or, for the MCF, thex/mcf_hat/varstep arrays), and the reloaded model reproducescif/iif/mcfexactly. Intensity models are resolved by name from a restricted set. The likelihood/data state is not stored, so a reloaded model behaves like afrom_paramsone for confidence bounds and diagnostics.
Regression
Added serialisation to the semi-parametric regression models, each on its own result class: Cox proportional hazards (
SemiParametricRegressionModel), the Lin-Ying additive-hazards model (AdditiveHazardsModel), and the Buckley-James AFT (BuckleyJamesModel) now haveto_dict/from_dictandto_json/from_json. Because the baseline is nonparametric, the coefficients plus the fitted baseline step arrays (or, for Buckley-James, the residual survival) are stored, so the reloaded model predicts identically – including Cox’spredict_tvcfor a time-varying-covariate fit, the additive model’s covariance / standard errors, and Buckley-James’sbootstrap_ci(its fit data is kept).SemiParametricRegressionModelis now exported fromsurpyval.univariate.regression.Added serialisation to the parametric regression models:
ParametricRegressionModelnow hasto_dict/from_dictandto_json/from_json, so a fitted Accelerated Failure Time, Proportional Hazards, Proportional Odds or (parametric) Additive Hazards model can be saved and rebuilt without the training data. The restored model predicts identically (sf/ff/df/hf/Hf/phi/random); if the fit’s parameter covariance was computable it is stored too, so the reloaded model also produces confidence bounds (cb/param_cb/standard_errors). Distribution and family are resolved by name from a restricted set, so an untrusted dict cannot load arbitrary objects. Models with a bespoke covariate link (an Accelerated Life parameter-substitution model) are refused with a clear error.ParametricRegressionModelis now exported fromsurpyval.univariate.regression.
Experimental
Breaking (experimental API): the survival tree/forest now take a single
kindparameter that couples the split criterion with its matching leaf model, replacing the independentsplit_rule/leaf_type/parametricknobs (whose free combination invited mismatched trees and whose defaults disagreed between entry points).kind="weibull"(the new default) adds the Weibull deviance split – a 2-d.f. likelihood-ratio gain computed with the full likelihood, with power against scale and shape differences (e.g. crossing-hazards populations that the exponential rule and the log-rank statistic largely miss) – paired with Weibull MLE leaves.kind="exponential"is the Davis-Anderson rule with Exponential leaves, andkind="non-parametric"is the risk-set log-rank with Nelson-Aalen leaves (observed/right-censored data, optionally left-truncated; raises otherwise). Parametric kinds now stay parametric all the way down: the degenerate-leaf rescue ladder is Weibull -> Exponential -> crude rate, never a nonparametric leaf. Split-search child fits warm-start from the parent’s optimum, which also guarantees a non-negative split gain in the 2-parameter case. The internal Weibull MLE is cross-validated againstWeibull.fiton every data configuration. now supports the full SurPyval data model: observed, left-, right- and interval-censored observations with optional left and/or right truncation. The risk-set log-rank split only exists for observed / right-censored (optionally left-truncated) data, so the tree gains a second split criterion – the full-likelihood exponential deviance split of Davis & Anderson (1989) – in which every candidate split is scored by the joint maximised exponential log-likelihood of its children, with each observation type contributing its exact likelihood term (including theS(t_l) - S(t_r)truncation correction). A newsplit_ruleparameter ("auto"default) keeps the log-rank split for data it is defined on – existing behaviour is unchanged – and switches to the deviance split otherwise; forcing"log-rank"on incompatible data raises a clear error. All candidate children within a node are scored over a common parameter window so the criterion is monotone (a split can never score below its parent), and splits with no likelihood gain stop the branch. Nonparametric leaves now use the Turnbull NPMLE when the data has left or interval censoring or right truncation (Nelson-Aalen otherwise, as before); parametric (Weibull) leaves already supported the full data model.fitalso accepts thexl/xrandtl/trconveniences.Fixed a crash in the experimental
RandomSurvivalForest: a degenerate bootstrap sample (e.g. heavily tied event times) could make a terminal node’s Weibull covariance step raise, taking down the whole forest fit. A terminal node now falls back to progressively simpler, more robust fits (Exponential, then Nelson-Aalen). The experimental modules are also excluded from the CI test run, since they are not part of the release contract.
Degradation
Added two-stage confidence bounds for the accelerated-degradation (covariate) life fit:
DegradationModel.cbnow accepts a stress vectorZand, withmethod="bootstrap", resamples units (each carrying its stress) and reruns the whole ADT pipeline to fold the first-stage path/extrapolation uncertainty into the reliability atZ. PreviouslycbraisedNotImplementedErrorfor covariate models; the analytic (generated-regressor) correction remains underived for the regression fit, so bootstrap is required there. The bootstrap holds the selected path model fixed, so it composes cleanly withpath="best"(no per-resample path re-selection).cbalso now validatesZ(required for covariate models, rejected for plain ones).Extended
population_method="reml"to nonlinear path models (exponential, power, Gompertz, …). Previously REML population estimation was restricted to paths linear in their parameters; nonlinear paths are now fitted with the Lindstrom-Bates (1990) FOCE alternating algorithm – each unit’s parameters are estimated at their conditional (penalised-least- squares) mode, the path is linearised about that mode into a working linear mixed model, and the linear REML step is iterated to convergence. This gives a positive-definitepath_param_covby construction (no PSD clipping) for nonlinear paths too, which is the more robust population estimate when the unit count is small. On a linear-in-parameters path the routine reduces exactly to the previous linear REML fit in a single pass.Added the Lu-Meeker induced failure-time distribution:
DegradationModel.induced_lifederives the population failure-time distribution directly from the fitted path-parameter distribution – drawing path parameterstheta ~ N(path_param_mean, path_param_cov)and pushing each through the path model’sinv_path(threshold)by Monte Carlo – rather than via each unit’s noisy pseudo failure time. It returns anInducedFailureDistributionexposingsf/ff/qf/mean/median/random(with aninf“never fails” mass reported asprob_never_fails), a diagnostic complement to the pseudo-failure-time life fit that the two can be overlaid to check.Added stochastic-process degradation models that model the degradation increments directly, deriving the failure-time distribution from the process’s first passage to the threshold (rather than via pseudo failure times), and handling irregular measurement spacing naturally. Two complementary processes are provided in
surpyval.degradation:WienerProcess(Brownian motion with drift, for non-monotone / noisy signals; its first passage is a closed-form Inverse-Gaussian law) andGammaProcess(monotone increasing increments, for irreversible damage such as wear, corrosion or crack growth; its first-passage distribution comes from the incomplete gamma function). Both fit by maximum likelihood from(x, y, i)measurement data and expose the induced failure-time distribution (sf/ff/df/hf/Hf/qf/mean/random) plus apredict_rulremaining-useful-life summary. The degradation documentation gained an expansive section explaining both processes, what each parameter means, the first-passage failure-time derivation, worked runnable examples, and guidance on choosing between them.
v0.14.0 (19 Jul 2026)
Documentation
Substantially expanded the recurrent-event documentation for the release. The theory pages now cover the arithmetic-reduction (ARA/ARI) models, the geometric-process view of the G1 renewal process, the time-rescaling residual / trend-test / Cramer-von Mises diagnostics, marked (competing-risks) recurrent events, gapped multi-window observation, and truncation, each with a short References section. The worked-example pages gained runnable demonstrations of ARA/ARI, renewal-model checking, gapped observation, the cause-specific MCF and intensity models, and a full build-out of the proportional-intensity regression examples.
Fixed and completed the recurrent-event API reference. Every model’s autodoc page (HPP, Duane, Cox-Lewis, Crow-AMSAA, the renewal and proportional-intensity models) previously rendered as an empty “alias of object” because the fitters are exposed as singletons; the pages now document each model’s methods. Added missing API pages for
ARA,ARI,NonParametricCounting,CauseSpecificMCF,CauseSpecificNHPPand the fittedRenewalModelobject.
Recurrent events
Added residual (
residuals:cumulative_hazard/pit/martingale), trend-test (trend_test) and Cramer-von Mises goodness-of-fit (cramer_von_mises) diagnostics to the renewal / virtual-age imperfect-repair models (GeneralizedRenewal,GeneralizedOneRenewal,ARA,ARI), completing the diagnostic coverage of the recurrent module. These processes have no marginal cumulative intensity, so the time-rescaling residuals come from each one’s conditional intensity – the cumulative hazard accumulated over each interarrival given the model’s virtual age (Kijima / ARA), time scaling (G1R) or intensity reduction (ARI) – and are iid Exp(1) under the fitted model. The Cramer-von Mises transforms use the compensator built from those increments (there being no closed-form intensity), and its p-value comes from a parametric bootstrap that resimulates each item and refits the full imperfect-repair model per replicate.Added support for gapped (multi-window) observation: an item can be observed over several disjoint time windows with unobserved gaps in between (events may occur during a gap but are not recorded). Pass
windows={item: [(start, end), ...]}to the intensity fitters (HPP,CrowAMSAA,Duane,CoxLewis) and the nonparametricNonParametricCountingMCF; every row ofxis then an observed event and the windows supply the end-of-window censoring. Because event counts over disjoint windows are independent for an NHPP, each window is fitted as its own observation period, so the intensity likelihood and the MCF at-risk set (an item is absent from the risk set during its gaps) both handle the gaps exactly. The virtual-age / renewal models (GeneralizedRenewal,GeneralizedOneRenewal,ARA,ARI) reject gapped data, since the virtual age at the start of a later window depends on the unobserved events during the gap.Recurrent event marks (competing-risks recurrent events) are now first class.
handle_xicntakes an event-type markeper row (withNone/NaNmarks normalised to a single “no cause” sentinel), so marked data gets the same validation, sorting and truncation handling as every other recurrent fit.CauseSpecificMCFnow routes through that handler and gains afit_from_df. NewCauseSpecificNHPPfits a parametric cause-specific intensity model – one NHPP (CrowAMSAAby default, or any counting-process fitter) per event type. Because a marked Poisson process decomposes into independent thinned Poisson processes, each cause is fitted to its own events over the full observation window of every item (other-cause events are ignored, exactly as a censored period would be), so each per-cause model is an ordinary fitted recurrence model with its fullcif/iif, inference and diagnostics;total_cifsums them for the overall event intensity.
v0.13.0 (18 Jul 2026)
Distributions
Added three Tier-2 discrete distributions:
Poisson(the count distribution on{0, 1, 2, ...}, distinct from the recurrent Poisson processes),BetaGeometric(a discrete-time frailty model — Geometric with a Beta-mixed failure probability, whose marginal hazard decreases with time), andDiscretize(distribution), a factory that turns any non-negative continuous distribution into its integer-binned counterpart (K = ceil(T), soP(K=k) = F(k) - F(k-1)and the discrete survival equals the continuous survival), fit by MLE on the underlying parameters.Beta.fit(how="MPP")now raises a clearValueError(the Beta has no linearising probability plot) instead of a rawNotImplementedError, and points toMLE/MSE/MOM.Parametric.momentnow works for limited-failure, zero-inflated and offset models (it previously raisedNotImplementedErrorunder a cure fraction, and silently dropped the offset). It returns the defective moment of the failure-time density, consistent withmean(moment(1) == mean()): the offset shifts the failure times and the cured fraction contributes nothing.Parametric.entropylikewise handles the offset (differential entropy is translation-invariant) and now raises a clearValueErrorfor models with a probability atom (a limited-failure mass at infinity or a zero-inflation mass at the offset), where a single differential entropy does not exist – it previously returned a wrong value for zero-inflated models.Parametric.qfnow works for limited-failure, zero-inflated and offset models (it previously raisedNotImplementedErrorwhenever a cure fraction was present). It inverts the full mixtureF(x) = f0 + (p - f0) F0(x - gamma): quantiles at or below the zero-inflation massf0return the offset, and quantiles at or above the attainable proportionpare infinite (that cured fraction never fails, so e.g. the median of a majority-cured population isinf). This also fixes the quantile of a zero-inflated (p == 1,f0 > 0) model, which previously ignoredf0and returned the wrong value.
Competing risks
Added
ParametricCompetingRisks, a fully parametric competing-risks model: a parametric distribution is fitted to each cause’s cause-specific hazard (the joint likelihood factorises, so each cause is fitted with the other causes’ events treated as right-censored) and smooth, extrapolatable cumulative-incidence functions are assembled from them. Providesfit/fit_from_df(with a per-cause distribution mapping), all-cause and cause-specifichf/Hf/sf/ff, the subdistribution densityiif, the cumulative incidencecif,probability_of_cause, sampling viarandom, andaic/bic/neg_ll. Complements the existing nonparametricCompetingRisksestimator and the semi-parametric cause-specific Cox / Fine-Gray regression models.ParametricCompetingRisks.from_fittedassembles a competing-risks model from already-fitted per-cause models, each of any family and configuration (e.g. a limited-failure Weibull for one cause, a LogNormal for another): pass a{cause: model}mapping or a sequence of models. Sampling handles cure fractions – when every cause carries one, some units never fail and are returned with causeNone.Every competing-risks model (parametric, nonparametric, and the Fine-Gray / cause-specific Cox regression) now treats a missing event value (
None,NaNor pandasNA) as a censored observation with no attributed cause, and derives the censoring flagcfrom the events when it is not supplied – so competing-risks data can be given as(x, e)alone, and a pandas cause column withNaNfor censored rows works directly.
Recurrent events
Added residual (
residuals:cumulative_hazard/pit/martingale), trend-test (trend_test) and Cramer-von Mises goodness-of-fit (cramer_von_mises) diagnostics to the proportional-intensity regression models (ProportionalIntensityHPP/ProportionalIntensityNHPP), matching those already on the parametric recurrence models. Each item’s time-rescaling residuals and conditionally- uniform transforms use its own covariate-scaled cumulative intensityLambda_0(t) exp(Z'beta), and the Cramer-von Mises p-value comes from a parametric bootstrap that refits the full regression model per replicate.
Regression — Cox proportional hazards
Added time-varying-covariate support in counting-process (start-stop) format:
CoxPH.fit_tvc/fit_tvc_from_dftake one row per interval(ident, start, stop, event, Z), validated byhandle_tvc, andSemiParametricRegressionModel.predict_tvcgives a subject’s survival along a supplied covariate path.Fixed the Breslow baseline hazard to respect left-truncation / delayed entry (
tl) and case weights (n);H0was previously wrong for any delayed-entry fit even though the coefficients were correct.CoxPH.fitgained a minimisation fallback so staggered delayed-entry data (e.g. the start-stop representation) converges where the root-finder stalled.Right / interval truncation is now rejected with a clear, Cox-specific error (a 2-D
tl), since the forward partial likelihood cannot express it.
Truncation
Verified and tested that the parametric AFT / PO / PH truncation correction uses each row’s own covariates: a covariate-recovery test confirms the coefficient and scale are recovered from left-, right-, interval- and partially-truncated data.
Documentation
Added worked, executed examples for regression confidence bounds, Buckley-James AFT, competing-risks regression (Fine-Gray + cause-specific Cox), degradation ADT covariates and two-stage bounds, the copula module, and the combined data-input flexibility; wrote the Maximum Product of Spacings (MPS) estimation theory section.
v0.12.0 (15 Jul 2026)
A large release consolidating the regression, recurrent-event, competing-risks,
degradation, and multivariate work accumulated since v0.10.1. Requires
Python 3.11+ and NumPy 2.
Regression
Standardised every univariate regression fitter (accelerated failure time, proportional hazards, proportional odds, additive hazards, accelerated life) on a common instance-based
fit()/fit_from_df()API with pandas and formulaic formula support.CoxPHgained the Efron tie handling in addition to Breslow, and its analytic (Efron) information matrix is now correct, so standard errors and p-values are produced for tied data.Added delta-method confidence bounds to the parametric regression models:
cb()on a predicted function at a covariate vector,param_cb()on a single coefficient, andcovariance()/standard_errors()/parameter_names()on the fitted parameters.Added
BuckleyJames, a semi-parametric accelerated-failure-time model with an unspecified error distribution (the accelerated-time counterpart of Cox), fitted by the Buckley-James imputation iteration with percentile-bootstrap coefficient intervals.Added a parametric
AdditiveHazardsregression fitter.
Competing risks
Added a competing-risks regression module with a cause-specific Cox model and a Fine-Gray subdistribution-hazard model (
CompetingRisksProportionalHazards), each withfit()/fit_from_df()and cumulative-incidence prediction.
Recurrent events
Standardised the recurrent-model API on the same instance-based fitters the univariate distributions use:
HPP,CrowAMSAA,Duane,CoxLewis,NonParametricCounting, the renewal fitters (GeneralizedRenewal/GeneralizedOneRenewal/ARA/ARI) and the proportional-intensity fitters are now configured singleton instances with an instance-methodfit(). PublicModel.fit(...)calls are unchanged; internally provided by thesurpyval.utils.fitter.singleton_fitterdecorator. Removed the unusedParametricRecurrenceRegressionModelstub.Added parameter-uncertainty and diagnostic support to the recurrent models, and removed the
dist='t'heuristic from the recurrentmcf_cb.
Degradation
Added the
surpyval.degradationpseudo-failure-time analysis module: per-unit path fits over a library of path models, extrapolation to a failure threshold, and a fitted life distribution, with population path-parameter estimation (Lu-Meeker two-stage and REML) and Bayesian remaining-useful-life prediction (predict_rul).Added two-stage (delta-method and bootstrap) confidence bounds on the fitted life model that fold in the first-stage path/extrapolation uncertainty (
DegradationModel.cb/life_parameter_covariance).Added Stage-1 accelerated degradation testing (ADT) covariates: passing
ZtoDegradationAnalysis.fitfits a regression life model on the pseudo failure times so life can be predicted at any stress condition.
Multivariate
Added a
surpyval.multivariatemodule with copula models over the univariate distributions.
Distributions and core
Added discrete lifetime distributions.
Hardened input validation in the
handle_xicn/xcnt_handlerdata handlers, and fixed a reserved-attribute clash.Simulation and
dist='t'cleanups.
v0.10.1.0 (25 Mar 2022)
Changed plot methods to now take ‘Axis’ object. This allows a user to pass in an existing axis.
plot functions now return an Axis object instead of the Lines2D object. Allows for easy user update after plotting.
Added fs_to_xcn as it was dropped in 10.0.1.
Changed all imports for numpy to be done from the surpyval module. This will allow for easy maintenance in future in the event of deprecated autograd.
v0.10.0.1 (22 Nov 2021)
Removed fsl_to_xcn function and replaced with fsli_to_xcn function that is able to take any combination of fsli.
v0.10.0 (9 Aug 2021)
Version snapshot for JOSS review
v0.9.0 (5 Aug 2021)
Better initial estimates in the
_parameter_initialiserfor the lfp data (use max F from nonp estimate…)issue #13 - Better failures when insufficient data provided.
issue #12 - Created
fsli_to_xcnhelper function.Fixed bug in confidence bounds implementation for offset distributions. CBs were not using the offset and were therefore way out. Now fixed.
Created a
NonParametric.cb()method to matchParametricAPI for confidence bounds.Cleaned up NonParametric code (removed some technical debt and duplicated code).
Changed the
__repr__function inNonParametricto be aligned toParametricUpdated the docstring for
fit()forNonParametricFixed bug in
NonParametricthat required thexinput to be in order for the functions (e.g.dfetc.).CoxPHreleased.General AL fitter in beta
General PH fitter in beta
Created
Linear,Power,InversePower,Exponential,InverseExponential,Eyring,InverseEyring,DualPower,PowerExponential,DualExponentiallife models.Created
GeneralLogLinearlife model for variable stress count input.For each combination of a SurPyval distribution and life model, there is an instance to use
fit(). For example there areWeibullDualExponential,LogNormalPower,ExponentialExponentialetc.- Docs Updates:
- Add application examples to docs:
Reliability Engineering
Actuary / Demography
Boston Housing
Medical science
Biology - Ware, J.H., Demets, D.L.: Reanalysis of some baboon descent data. Biometrics 459–463 (1976).
v0.8.0 (27 July 2021)
Made backwards incompatible changes to
LFPmodels, these are now created with thelfp=Truekeyword in thefit()methodCreated ability to fit zero-inflated models. Simply pass the
zi=Trueoption to thefit()method.Chanages to
utils.xcnt_handlerto ensurex,xl, andxrare handled consistently.changed the way
__repr__displays a Parametric object.Changed the default for plotting to be
Fleming-Harrington. This was a result of seeing how poorly theNelson-Aalenmethod fits zero inflated models. FH therefore offers the best performance of a Non-Parametric estimate at the low values of the survival function (as KM reaches 0 for fully observed data) and at high values (KM is good but NA is poor).Added a Fleming-Harrington method to the Turnbull class.
Improved stability with dedicated
log_sf,log_ff, andlog_dffunctions. Less chance of overflows and therefore better convergence.Changed interpolation method of
NonParametric. Allows for use of cubic interpolationChanged
from_paramsto accept lfp and zi (or any combo)Changed
random()inParametricso that lfp or zi models can be simulated!Improved the way surpyval fails
Substantial docs updates.
v0.7.0 (19 July 2021)
Major changes to the confidence bounds for
Parametricmodels. Now use thecb()method for every bound.Removed the
OffsetParametricclass and madeParametricclass now work with (or without) an offset.Minor doc updates.