Example data sets used throughout the documentation and the docstrings,
each returned as a pandas.DataFrame by a load_* function in
surpyval.datasets. Check each loader’s description for how its
censoring column is coded: several are already coded as SurPyval’s
censoring flag c (0 observed, 1 right-censored) rather than as an
event indicator.
Returns a Pandas DataFrame containing the data of
the tensile strength of Bofors Steel from [2].
Grouped data in ten rows: x is the strength and n the number
of specimens with that strength, so pass the columns as x and
n. Every value is an exact observation.
This is a well-known data set used in machine learning. It can be
analysed with survival analysis methods by considering the
fact that the highest prices appear to be right censored.
One row per census tract (506), with the thirteen covariates of the
original data and medv, the median house value in thousands of
dollars. medv is capped at 50: the 16 tracts at 50.0 are right
censored there. There is no censoring column; build one from the cap.
The Framingham Heart Study teaching dataset, from [13].
A longitudinal extract: 4434 participants (RANDID) examined at up
to three visits (PERIOD 1-3, 11627 rows), with the risk factors
measured at each visit (SEX 1 male / 2 female, AGE,
TOTCHOL, SYSBP, DIABP, CURSMOKE, CIGPDAY, BMI,
DIABETES, BPMEDS, HEARTRTE, GLUCOSE, educ,
HDLC, LDLC; TIME is the visit’s day since baseline) and
PREV* flags for conditions present at the visit.
Follow-up is 24 years (8766 days). Each outcome has an event flag and
a time in days: DEATH/TIMEDTH, ANGINA/TIMEAP,
HOSPMI/TIMEMI, MI_FCHD/TIMEMIFC, ANYCHD/TIMECHD,
STROKE/TIMESTRK, CVD/TIMECVD and
HYPERTEN/TIMEHYP. A flag of 1 means the event happened at that
time, so c=1-flag; outcome columns repeat on every visit row,
so take one row per participant (e.g. PERIOD==1) for a
time-to-event fit. 1550 participants died during follow-up.
The teaching data are anonymised and perturbed by the NHLBI; they are
meant for learning methods, not for publishing epidemiological
results.
Data on the survival of a repairable system from [4].
One system with twelve failures. x is the cumulative time of
each failure (the times between failures are 3, 6, 11, 5, 16, 9, 19,
22, 37, 23, 31 and 45); every row is an observed failure and there is
no item or censoring column.
Data on the survival of patients who may or may not have received a
heart transplant, from [5].
Start-stop form: each patient (id) has one row before a
transplant and, if transplanted, one after (transplant 1), over
(start,stop]; event is 1 for a death at stop (so
c=1-event). There are 172 rows for 103 patients, 75 of whom
died. age is the age at acceptance minus 48 years, year the
date of acceptance in years after 1 November 1967 and surgery 1
for prior bypass surgery.
Times to infection at the catheter insertion point of 38 kidney
patients on portable dialysis, from McGilchrist and Aisbett [16], as
R’s survival::kidney: two times per patient, the classic example
of a shared frailty (the patient).
id is the patient (the frailty group), time the time to
infection or censoring in days and status 1 for an infection and 0
for a censored time, so pass c=1-status. The covariates are
age, sex (1 male, 2 female) and disease (Other,
GN, AN or PKD); frail is the frailty estimate of the
original paper. 76 rows, 58 infections.
Data on the survival of patients with advanced lung cancer from [6].
time is the survival time in days. status is 1 for a death
at time and 0 for a patient censored (alive at last follow-up),
as in lifelines’ load_lung (R’s survival::lung codes it 2 and
1). SurPyval’s censoring flag is its complement, so pass
c=1-status. (Before v0.22 this copy stored status already
inverted to the censoring flag.) There are 228 patients, 165 of whom
died. The other columns are the covariates
of the original data (inst, age, sex with 1 male and 2
female, ph.ecog, ph.karno, pat.karno, meal.cal and
wt.loss, several with missing values).
Data on failures of integrated circuits from [12].
Very difficult for LFP calculations since the data is heavily
right censored: 4,156 circuits were tested for 1,370 hours and 28
failed. The data is in xcnt form: x the time in hours, c
the censoring flag and n the count, with the 4,128 survivors in
one row right censored at 1,370.
Data on the survival of a repairable system from [7].
Recurrent event data for six systems: x is the cumulative time of
each event on system i, and c is 1 on each system’s last
row, the end of its observation (0 for a failure).
The Mayo Clinic primary biliary cirrhosis (PBC) trial, longitudinal
version, from [14].
312 patients (id) randomised to D-penicillamine or placebo
(drug), with repeated laboratory measurements: one row per visit
(1945 rows), year being the visit time in years since
enrolment. Per-patient columns (repeated on every row) are years,
the follow-up time in years, and status ("alive",
"transplanted" or "dead") at that time; status2 is 1 for
death and 0 otherwise (alive or transplanted), so for survival with
transplantation treated as censoring c=1-status2. 140
patients died and 29 were transplanted. The per-visit covariates are
ascites, hepatomegaly, spiders, edema, serBilir,
serChol, albumin, alkaline, SGOT, platelets,
prothrombin and histologic; age and sex are at
baseline and sno. is a row index.
Transplantation is a competing event for death, so the data also
suit competing risks
(causes from status).
Data on the recidivism of released prisoners from [8]. Uses only
static covariates.
One row per prisoner (432): week is the week of first arrest after
release, or 52 for those not arrested in the year of follow-up, and
the covariates are fin (financial aid), age, race,
wexp (work experience), mar (married), paro (released on
parole) and prio (number of prior convictions).
arrest has the original coding, as in R’s carData::Rossi and
lifelines’ load_rossi: 1 for a prisoner arrested during follow-up
(an observed event at week), 0 for one not arrested (right
censored at week 52). SurPyval’s censoring flag is its complement, so
pass c=1-arrest. (Before v0.22 this copy stored arrest
already inverted to the censoring flag.) There are 114 arrests.
Data on the recidivism of released prisoners from [9]. Includes time
varying covariates.
Start-stop (counting-process) form: one row per prisoner-week, with
id the prisoner, start and stop the week’s interval,
event 1 if the prisoner was arrested at the end of that week (so
c=1-event), and employed the time-varying covariate. Here
arrest has the original coding (1 = arrested during follow-up), as
in load_rossi_static(). fin, race, wexp, mar,
paro and employed are stored as strings ("yes" / "no"
and so on), as in R; encode them as numbers, or use a formula, before
fitting.
The SUPPORT study of seriously ill hospitalised adults, from [15].
9105 patients, one row each (sno is a row index). d.time is
the follow-up time in days and death 1 for a death during
follow-up (6201 patients), so c=1-death; hospdead flags a
death in hospital and slos is the days from study entry to
discharge. dzgroup/dzclass give the disease group (acute
respiratory failure or multiple organ system failure, CHF, COPD,
cirrhosis, colon or lung cancer, coma). The remaining columns are
demographics (age, sex, race, edu, income),
comorbidity and severity scores (num.co, scoma, sps,
aps, avtisst), physiology on day 3 (meanbp, wblc,
hrt, resp, temp, pafi, alb, bili, crea,
sod, ph, glucose, bun, urine), costs, activities
of daily living and the study’s own model and physician survival
estimates (surv2m, surv6m, prg2m, prg6m). Many
physiology columns have missing values.
One row per tire (34). Survival is the (normalised) time and
Censoring is coded as SurPyval’s censoring flag – 1 for a tire
that had not failed, 0 for a failure (11 of them) – so use it
directly as c. The other seven columns are the covariates.