Saving and Loading Models
A fitted model can be saved and restored later, or in another process,
without refitting. Almost every fitted SurPyval model can be serialised
to a plain dictionary with to_dict() and written to JSON with
to_json(path). The dictionaries hold only plain Python types, so they
can also go straight into a document store (the round trip through
BSON/MongoDB is tested). What a restored model keeps, and what needs the
original data, is described in Conventions.
Restoring takes one call whichever class wrote the file: the package-level readers dispatch on the serialised dictionary itself.
import surpyval
model = surpyval.Weibull.fit(x)
model.to_json("weibull.json")
restored = surpyval.from_json("weibull.json") # any model's file
restored = surpyval.from_dict(model_dict) # any model's dict
The class-level readers on the fitted-model classes
(Parametric.from_dict, SemiParametricRegressionModel.from_dict,
… and the matching from_json) remain available when the model class
is known up front; each rejects a dictionary written by a different class
with a ValueError. The readers belong to the model classes, not to
the fitters: Weibull.from_dict and CoxPH.from_dict do not exist.
CompetingRisksProportionalHazards serialises the same way, for both
model="Cox" and model="Fine-Gray"; its per-cause optimiser results
(results) are not stored.
Some details differ between families:
Only the univariate
ParametricandNonParametricmodels taketo_dict(with_data=True), which stores the fitted data as well so thatplotand likelihood-ratio bounds (parametric) orbootstrap_cb(non-parametric) work after restoring;to_json(path, with_data=True)writes the same to a file (any other model’sto_jsonrefuseswith_data=Truewith aTypeError). The information criteria need no data: a fitted univariate or regression model stores their sample size ("ic_n"), sobicandaic_cwork on the restored model either way.The degenerate
NeverOccursandInstantlyOccursdistributions have no fitted state: the class itself is the model, soto_dict,to_json,from_dictandfrom_jsonare called on the class (NeverOccurs.to_json(path)), andsurpyval.from_jsonreturns the class.Every reader – the package-level ones and each class’s own
from_dict/from_jsonalike – checks the"schema"(an integer no newer than this SurPyval), names the entry a truncated or hand-edited dictionary is missing, and refuses the parameters of a univariate parametric model that fall outside the distribution’s bounds, each with aValueError.These cannot be saved and raise an error from
to_dict: a stratified Cox model, an accelerated-life model with a user-defined life model, and a copula of a custom family. A regression fitted with a formula is saved with its levels and fitted transform statistics (C(...),scale(),poly(), splines); a formula whose state cannot be stored as JSON raises atto_dictinstead.A model of a
Discretize(...)distribution is read back like any other. A model of aCustomDistributionstores the distribution’s name only (its cumulative hazard is a Python function):from_dictreads it back in a session that has constructed the sameCustomDistributionagain, and otherwise raises aValueErrorsaying so.
Infinite and NaN values
A fitted model holds values that are not finite numbers: an untruncated
bound (-inf, inf), a cumulative hazard after the last death
(inf), an undefined variance (nan). JSON cannot write them –
Python’s json writes the non-standard Infinity and NaN, which
strict parsers (JavaScript’s JSON.parse, many databases) refuse – so
the dictionaries use a convention of their own, and are strict JSON:
each non-finite value is written as
null;the dictionary holding it records what every such
nullstood for under"non_finite": for each kind present ("inf","-inf","nan") a list of JSON Pointers to the values, relative to that dictionary. A model dictionary nested in another (a copula’s margins, a tree’s leaves) carries its own record. Anullthat no record names is an ordinary missing value.
>>> surpyval.KaplanMeier.fit([3, 4, 5]).to_dict()
{..., 'H': [0.405..., 1.098..., None],
'greenwood': [0.166..., 0.666..., None],
'non_finite': {'inf': ['/H/2'], 'nan': ['/greenwood/2']}, 'schema': 2}
Every reader restores the original values, so a round trip is exact. A
consumer in another language sees null where no number applies, and
can use the record to recover the values. Dictionaries and files written
before this convention ("schema" 0 or 1, holding Infinity or
NaN, which Python’s json reads) still load.
Each file is stamped with the oldest schema version that reads it
correctly (required_schema()). A file with no
non-finite value – most fitted parametric, regression, recurrent, copula
and degradation models – has exactly the schema-1 layout and is stamped
1, so SurPyval 0.20, which reads schema 1, still loads it (the test suite
checks this against that release). A file that records non-finite values as
null, as most non-parametric models and anything saved with
with_data=True do, is stamped 2; an older SurPyval refuses it with an
error asking for an upgrade rather than misreading its null values. So is
a regression model fitted with a formula that SurPyval 0.20 cannot rebuild
(a wrapped categorical such as C(g), integer category levels, or a
fitted transform such as scale(z)); a formula of plain columns and
string categoricals is stamped 1. So is a regression model fitted with
center=True (CoxPH, FineGray,
CompetingRisksProportionalHazards or a parametric regression) whose
covariate means center are not all zero: its baseline is that of a unit
at center, which SurPyval 0.20 would ignore, reading the baseline as at
Z = 0. A model with its baseline at Z = 0, the default, has no
center and the layout of before. SurpyvalData.to_json uses the same
convention. encode_non_finite() and
decode_non_finite() apply and undo it on
any dictionary.
- surpyval.serialisation.from_dict(model_dict: dict) Any
Restore any serialised SurPyval model from its dictionary.
Reads the dictionary written by any fitted model’s
to_dictand dispatches to the right class’sfrom_dict, so the caller does not need to know which class wrote it.- Parameters:
model_dict (dict) – A dictionary produced by a SurPyval model’s
to_dict.- Return type:
The restored model, of whichever class serialised the dictionary.
- Raises:
ValueError – If the dictionary is not recognisable as a serialised SurPyval model, lacks an entry its reader needs (the message names it), has a
"schema"that is not a non-negative integer or was written by a newer SurPyval (a higher"schema"version), holds parameters outside the distribution’s bounds (for a univariate parametric model), or names a distribution the reader does not know – which includes aCustomDistributionthat has not been constructed again in this session (a dictionary stores only its name, since its cumulative hazard is a Python function). Also if its"non_finite"record (seeencode_non_finite()) is corrupt. Each class’s ownfrom_dictraises the same errors.
Notes
What a restored model keeps differs by family: in general the parameters and whatever predictions need, but not the fitted data, so methods that need the data (
plot, bootstrap and likelihood-ratio bounds, residuals) raise on the restored model – except a univariate parametric model’splot, which draws the model’s CDF without data points. A fitted univariate parametric, regression or copula model keeps the likelihood and sample size of its information criteria, soaicandbic(andaic_c, where the model has one) work on its restored copy. See “Saving and Loading Models” in the Conventions page.Examples
>>> import surpyval >>> from surpyval import Weibull >>> model = Weibull.fit([3.0, 4.0, 5.0, 6.0, 7.0]) >>> restored = surpyval.from_dict(model.to_dict()) >>> restored.dist.name 'Weibull'
- surpyval.serialisation.from_json(fp: str | Path) Any
Restore any serialised SurPyval model from a JSON file or string.
Reads a file written by any fitted model’s
to_json(or the JSON textto_json()returns without a path) and dispatches to the right class’s reader; seefrom_dict().- Parameters:
fp (str | Path) – Path to a JSON file written by a SurPyval model’s
to_json, or the JSON text itself (a string starting with{).- Return type:
The restored model, of whichever class serialised the file.
Examples
>>> import os, tempfile >>> import surpyval >>> from surpyval import Weibull >>> model = Weibull.fit([3.0, 4.0, 5.0, 6.0, 7.0]) >>> path = os.path.join(tempfile.mkdtemp(), "weibull.json") >>> model.to_json(path) >>> restored = surpyval.from_json(path) >>> restored.params array([5.53092634, 4.04187535]) >>> surpyval.from_json(model.to_json()).params array([5.53092634, 4.04187535])
- surpyval.serialisation.encode_non_finite(model_dict: dict) dict
Make a serialised dictionary strict JSON, in place.
json.dumpswritesinf,-infandnanas the literalsInfinity,-InfinityandNaN, which are not JSON: strict parsers (JavaScript’sJSON.parse, many databases) refuse the whole document. Yet the values are meaningful in a fitted model – an untruncated bound, the cumulative hazard after the last death, an undefined variance – so they cannot simply be dropped.The convention: each non-finite float is written as
null, and the dictionary records what every suchnullstood for under"non_finite", as lists of RFC 6901 JSON Pointers relative to the dictionary, grouped by kind:{"H": [0.1, 0.4, null], "greenwood": [0.01, 0.05, null], "non_finite": {"inf": ["/H/2"], "nan": ["/greenwood/2"]}}
A
nullthat no pointer names is an ordinaryNone. Readers put the floats back withdecode_non_finite(), which everyfrom_dictdoes; a consumer in another language seesnullwhere no number applies and can use the record to recover the exact values. The kinds with no values are omitted, as is the record itself when the dictionary holds no non-finite float.Numpy arrays and scalars are converted to native Python types on the way, so the dictionary is also BSON-native. A record the dictionary already carries (
to_dictoutput re-encoded) is extended, and the records of nested model dictionaries are left in place.Examples
>>> import numpy as np >>> from surpyval.serialisation import encode_non_finite >>> encode_non_finite({"H": [0.5, np.inf], "var": np.nan}) {'H': [0.5, None], 'var': None, 'non_finite': {'inf': ['/H/1'], 'nan': ['/var']}}
- surpyval.serialisation.decode_non_finite(model_dict: dict) dict
Undo
encode_non_finite(): the dictionary with everynullits"non_finite"records name put back toinf,-infornan, and the records removed.The records of nested model dictionaries are applied too. The input is not modified; a dictionary without records (including those written before schema 2, whose non-finite values were stored as floats) is returned unchanged. A record naming a missing entry or a value that is not
nullraises aValueError.Examples
>>> from surpyval.serialisation import decode_non_finite >>> decode_non_finite( ... {"H": [0.5, None], "non_finite": {"inf": ["/H/1"]}} ... ) {'H': [0.5, inf]}
- surpyval.serialisation.required_schema(model_dict: dict) int
The oldest schema version that reads
model_dictcorrectly.2 if the dictionary, or a model dictionary nested in it, records non-finite values as
null(a"non_finite"record, seeencode_non_finite()), which a schema-1 reader would take for missing entries, holds a regression formula that only a schema-2 reader can rebuild (wrapped categoricals such asC(g), integer levels, or fitted transforms such asscale(z)), or holds the"support"of a non-parametric estimate’sset_supportor the"band_n"of itsband, or the nonzero covariate"center"of a regression model fitted withcenter=True, which a schema-1 reader would silently ignore; 1 otherwise, the layout SurPyval v0.20 reads. This is the versionstamp_schema()writes.Examples
>>> from surpyval.serialisation import required_schema >>> required_schema({"params": [10.0, 2.0]}) 1 >>> required_schema({"H": [0.1, None], "non_finite": {"inf": ["/H/1"]}}) 2 >>> required_schema({"x": [1.0, 2.0], "support": [0.0, 5.0]}) 2 >>> required_schema({"beta": [0.5], "center": [0.0]}) 1 >>> required_schema({"beta": [0.5], "center": [2000.0]}) 2