| Title: | Surveillance Data Cleaning and Preparation for Public Health |
| Version: | 0.7.8 |
| Description: | Clean, prepare, and aggregate surveillance data for public health analysis. Provides structural data cleaning and standardisation (clean_the_nest()), age categorisation against ~50 published schemes with publication-ready labelling (preening()), time-unit aggregation with zero-filling and seasonal awareness (roost()), joint aggregation of several linked event dates (e.g. onset, admission, ICU, complication, fatality) into one table of comparable rate columns (flyway()), under-ascertainment correction via a stratified, time-varying multiplier factor supplied directly, derived by the ratio (multiplier) method, or derived by inverting an externally sourced severity rate (e.g. an infection-fatality-rate anchor) against an observed severity ratio (corncrake()), comorbidity detection from ICD-10-AM clinical coding (plumage()), vaccine coverage data construction (brood()), hash-based de-identification (molting()), and relinking of previously de-identified data (homing()). brood() produces a brood_df object supporting two population models: pre-aggregated denominators (population_model = "pre_aggregated") and record-level cohort designs (population_model = "cohort"). The cohort model handles single time-point coverage snapshots, interrupted time series analysis via a built-in sweep returning monthly coverage rates (time_series = TRUE), and birth cohort designs with person-time computation. This cohort/time-series coverage model was applied in Roughan et al. (2026) <doi:10.33321/cdi.2026.50.031> to estimate infant immunisation coverage against respiratory syncytial virus over an 18-month period. Both wide format (one row per person with dose columns, from 'starling'::murmuration()) and long format (one row per dose) are accepted. corncrake() returns both a point-corrected count and uncertainty bounds wherever they can be derived, including the inverse relationship between a severity-anchored factor and the bounds of its own reference rate. Built for Australian public health surveillance practice but not specific to it – see individual function documentation for notes on non-Australian use (e.g. Northern Hemisphere season boundaries). |
| License: | MIT + file LICENSE |
| Depends: | R (≥ 4.1) |
| Imports: | dplyr (≥ 1.1.0), tidyr (≥ 1.3.0), lubridate (≥ 1.9.0), stringr (≥ 1.5.0), rlang (≥ 1.1.0), tibble (≥ 3.2.0), digest (≥ 0.6.30), janitor (≥ 2.2.0), utils, stats |
| Suggests: | testthat (≥ 3.0.0), knitr (≥ 1.42), rmarkdown (≥ 2.20), usethis (≥ 2.1.0), gtsummary, ggplot2 |
| Config/testthat/edition: | 3 |
| Encoding: | UTF-8 |
| Language: | en-GB |
| LazyData: | true |
| RoxygenNote: | 8.0.0 |
| VignetteBuilder: | knitr |
| URL: | https://github.com/nrsmoll/mudnester |
| BugReports: | https://github.com/nrsmoll/mudnester/issues |
| NeedsCompilation: | no |
| Author: | Nicolas Smoll |
| Maintainer: | Nicolas Smoll <nicolas.smoll@health.qld.gov.au> |
| Packaged: | 2026-09-23 03:20:56 UTC; SmollN |
| Repository: | CRAN |
| Date/Publication: | 2026-10-02 11:50:02 UTC |
mudnester: Surveillance Data Cleaning and Preparation for Public Health
Description
Bird note: The White-winged Chough (Corcorax melanorhamphos)
belongs to the Corcoracidae family – Australia's mud-nesters.
These cooperative breeders build their nests from mud in careful,
methodical layers, returning again and again to reinforce weak spots and
add fresh material before anything else is placed on top. A group of
Choughs will work together on the same nest for weeks, each layer
dependent on the one beneath it. No other bird in Australia builds with
such deliberate, cumulative care. mudnester is named for this
behaviour: every function in the package is a layer of the nest, and
nothing downstream is trustworthy until the layers beneath it are
structurally sound.
The individual functions continue the bird theme across the full aviary ecosystem:
-
clean_the_nest()– the foundational mud layer. The Chough reinforces the base before adding anything else. -
preening()– how a bird sorts every feather into a functional position; here, the same raw age column sorted into one of ~50 standard groupings on demand. -
plumage()– a bird's plumage is the outward sign of its underlying condition; here, ICD-10-AM codes turned into a visible comorbidity profile. -
roost()– a roost is where many individual birds settle into one countable gathering at the end of the day; here, many individual surveillance records resolving into counts by time unit – the ecosystem's epicurve builder, for case incidence, hospitalisations, ICU admissions, deaths, or vaccine doses alike. -
flyway()– a flyway is the migratory corridor linking a bird population's successive stopover sites, the same birds traced through each waypoint of one journey; here, the same cohort traced through several linked milestone dates (onset, admission, ICU, death) into one aligned table.roost()'s own wrapper for that shape of problem. -
corncrake()– corncrakes are detected far more often by ear than by eye, and field surveyors have long applied a calibrated call-based correction to convert a count of detections into an estimate of the true, largely-unheard population behind it; here, a notified surveillance count corrected for under-ascertainment by the same logic. -
brood()– a brood is the full clutch under a parent bird's care: every egg counted, every hatchling tracked, none overlooked. The gap between the clutch and the hatchlings is not an absence; it is information. Here, the full eligible population (the clutch) and the vaccinated subset (those that have hatched) are structured into a validatedbrood_dfready forbowerbird::brood_plot(). Accepts wide format (one row per person withvax_date_1tovax_date_Ncolumns – the natural output ofstarling::murmuration()withlie_nest_flat = TRUE) and long format (one row per dose). Supports pre-aggregated and cohort population models, the latter with person-time and age-based eligibility windows, single time-point and interrupted-time-series (time_series = TRUE) assessment. Thevalidity_daysargument applies a dose expiry window so only currently protective doses are counted, andvax_targetfilters to specific vaccine types by partial string match across all dose columns. -
molting()– moulting is how a bird sheds its identifying plumage and replaces it with something uniform and anonymous; here, PII stripped and replaced with a cryptographic hash. -
homing()– homing pigeons return unerringly to their loft; here, de-identified data navigated back to original identifiers using themolting()lookup table.
Worldwide use, Australian home turf
Nothing in mudnester's cleaning or aggregation logic is
Australia-specific – it works on surveillance data from any
jurisdiction. Its defaults, its ~50 age-banding schemes in
preening, and its worked examples are tuned for Australian
public health practice (Southern Hemisphere seasons by default in
roost, ATAGI-aligned age bands, jurisdictional datasets like
the Australian Immunisation Register and NNDSS in the surrounding
ecosystem) because that is the practice this package was built inside of,
at the Sunshine Coast Public Health Unit. Individual function
documentation notes where a setting needs to change for non-Australian
use (e.g. northern_hemisphere = TRUE in roost).
Funding
Development of this package was supported by the Moderna Global Research Fellowship.
Author(s)
Maintainer: Nicolas Smoll nicolas.smoll@health.qld.gov.au (ORCID)
Authors:
Nicolas Smoll nicolas.smoll@health.qld.gov.au (ORCID)
Other contributors:
Moderna (Support for this package's development was provided via the Moderna Global Research Fellowship) [funder]
See Also
Useful links:
Age Banding Scheme Definitions
Description
A tibble of ~50 named age-banding schemes used by preening()
and list_age_schemes(). Each row defines one scheme: its
breaks, labels, family, focus tags, source citation, and audit metadata.
Usage
age_schemes
Format
A tibble with 50 rows and 11 variables:
- scheme
Character. Unique
snake_casescheme identifier.- family
Factor. One of
"national_stats","international_stats","vaccination","surveillance","clinical_developmental","disease_specific".- focus_tags
List-column of character vectors. One or more focus tags from the fixed vocabulary:
paediatric,aged_care,who_standard,vaccination,surveillance,broad,fine_grained,national_au,research.- n_bands
Integer. Number of age bands (length of
labels).- breaks
List-column of numeric vectors. Break points; always starts at 0, ends at
Inf, strictly increasing.- labels
List-column of character vectors. Band labels; length equals
length(breaks) - 1.- age_unit
Factor.
"years"or"days".- source
Character. Short citation.
- source_url
Character. URL to the source document (NA if none).
- last_checked
Date. Date the scheme was last verified against its source.
- is_illustrative
Logical.
TRUEif the scheme follows a documented convention but lacks a single primary citation.
Source
See vignette("age-schemes", package = "mudnester") for the full
catalogue with source citations. Constructed in
data-raw/age_schemes.R.
See Also
Construct a Vaccine Coverage Data Object
Description
Bird note: A brood is the full clutch under a parent bird's care –
every egg counted, every hatchling tracked, none overlooked. The parent
knows precisely how many eggs were laid (the eligible population), how many
have hatched (the vaccinated), and which ones are still waiting. The gap
between the clutch and the hatchlings is not an absence; it is information.
brood() does the same for a vaccination cohort: it takes the full
eligible population and the vaccinated subset, structures them into a
validated, self-documenting brood_df object, and makes the coverage
gap as visible as the coverage itself. The resulting brood_df is
ready for bowerbird::brood_plot().
Constructs and validates a brood_df object for vaccine coverage
analysis. Accepts two input formats and two population models:
Input formats:
-
Wide format (
data_format = "wide"): one row per person with vaccination dates in columnsvax_date_1throughvax_date_Nand optionally paired type columns. This is the natural output ofstarling::murmuration(). -
Long format (
data_format = "long"): one row per dose, withid_col,vax_date_col, and optionallyvaccine_type_col.
Population models:
-
Pre-aggregated (
population_model = "pre_aggregated", default): supply a summary data frame withn_vaccinatedandn_eligiblealready computed per stratum (and optionally per time period viadate_col). Does not perform per-person dose logic. Acceptsperson_time_colfor person-time denominators. -
Cohort (
population_model = "cohort"): record-level data, one row per person. This is the backbone model. Handles all individual-level designs:-
Single time point: set
windowandreference_date(orintervention_date) for a snapshot coverage figure. -
Time series (
time_series = TRUE): sweepswindow = "current_at_date"across monthly (or other) intervals betweents_startandts_end. Returns one row per stratum per time point, with areference_datecolumn and anintervention_periodcolumn auto-populated whenintervention_dateis supplied. This is the primary design for interrupted time series analysis of vaccination coverage. -
Birth cohort: a special case of the cohort model. Set
entry_date_col = "dob"and supplyeligibility_daysandcensor_dateto get birth-cohort behaviour with person-time computation.
-
Assessment windows (population_model = "cohort"):
-
"current_at_date"(default): currently covered atreference_date. Used automatically for every time point whentime_series = TRUE. -
"at_exit": vaccinated at or beforeexit_date_col. -
"pre_entry": vaccinated before cohort entry (baseline). -
"post_entry_days": vaccinated withinwindow_daysdays of cohort entry. -
"post_intervention": vaccinated on or afterintervention_date.
Usage
brood(
data,
data_format = "wide",
population_model = "pre_aggregated",
stratum_col = "stratum",
vaccine_type_col = NULL,
dose_col = NULL,
date_col = NULL,
n_vaccinated_col = "n_vaccinated",
n_eligible_col = "n_eligible",
person_time_col = NULL,
vax_date_cols = NULL,
vax_type_cols = NULL,
id_col = "id_var",
vax_date_col = "vax_date",
vax_target = NULL,
vax_target_exact = FALSE,
validity_days = Inf,
entry_date_col = "entry_date",
exit_date_col = "exit_date",
window = "current_at_date",
window_days = 365L,
intervention_date = NULL,
post_window_days = NULL,
reference_date = NULL,
eligibility_days = NULL,
censor_date = NULL,
time_series = FALSE,
ts_start = NULL,
ts_end = NULL,
ts_by = "month",
denominator_notes = NULL
)
Arguments
data |
A data frame. For |
data_format |
Character. |
population_model |
Character. |
stratum_col |
Character or NULL. Stratification column name. Default
|
vaccine_type_col |
Character or NULL. Vaccine type column for long
format or pre-aggregated data. For wide format, use |
dose_col |
Character or NULL. Dose number/label column. Default
|
date_col |
Character or NULL. Date column for pre-aggregated
time-varying input. Default |
n_vaccinated_col |
Character. Column containing vaccinated count.
Default |
n_eligible_col |
Character or NULL. Column containing eligible
population count. Default |
person_time_col |
Character or NULL. Column containing person-time
denominator. Default |
vax_date_cols |
Character vector. Names of the dose date columns
(e.g. |
vax_type_cols |
Character vector or NULL. Names of paired vaccine type
columns. Required if |
id_col |
Character. Person identifier column. Default |
vax_date_col |
Character. Single dose date column. Default
|
vax_target |
Character vector or NULL. Vaccine type(s) to count.
|
vax_target_exact |
Logical. Use exact string matching if |
validity_days |
Numeric. Days after administration after which a dose
is considered expired. Default |
entry_date_col |
Character. Cohort entry date column. For birth cohort
studies, pass |
exit_date_col |
Character or NULL. Cohort exit date column. Required
for |
window |
Character. Coverage assessment anchor for single-timepoint
designs. One of |
window_days |
Integer. Days after cohort entry for
|
intervention_date |
Date or NULL. For |
post_window_days |
Integer or NULL. Days forward from
|
reference_date |
Date or NULL. Anchor date for single-timepoint
|
eligibility_days |
Integer or NULL. For birth cohort studies: days
from |
censor_date |
Date or NULL. Administrative censoring date for birth
cohort designs. Persons still within their eligibility window at
|
time_series |
Logical. If |
ts_start |
Date or NULL. Start of the time-series sweep. Required when
|
ts_end |
Date or NULL. End of the time-series sweep. Required when
|
ts_by |
Character. Time interval between successive reference dates.
One of |
denominator_notes |
Character or NULL. Free-text denominator
description. Stored as a metadata attribute and displayed by
|
Value
A brood_df object: a classed data frame with columns:
stratumStratum label (character).
reference_dateReference date for cohort time-series output (Date).
NAfor pre-aggregated output.intervention_periodFor time-series with
intervention_datesupplied:"Pre-intervention"or"Post-intervention"(character).NAotherwise.vaccine_typeVaccine product label (character or NA).
dose_numberDose label (character or NA).
n_vaccinatedVaccinated count, numerator (integer).
n_eligibleEligible population count, denominator (integer or NA).
person_timePerson-time denominator (numeric or NA – only populated for birth cohort designs).
coverageProportion vaccinated, 0-1 scale (numeric).
window_descriptionAuto-generated description of the coverage window applied (character).
Metadata attributes:
-
attr(x, "population_model"):"pre_aggregated"or"cohort" -
attr(x, "data_format"):"wide"or"long" -
attr(x, "time_series"):TRUEorFALSE -
attr(x, "time_series_by"): interval used (time-series only) -
attr(x, "n_time_points"): number of time points (time-series) -
attr(x, "intervention_date_ts"): intervention date supplied (time-series only, forbowerbird::brood_plot()to draw a vertical reference line) -
attr(x, "window"): single-timepoint window used -
attr(x, "validity_days"): validity window applied -
attr(x, "vax_target"): vaccine type filter applied -
attr(x, "denominator_notes"): free-text denominator notes -
attr(x, "n_strata"): number of unique strata -
attr(x, "mudnester_version"): package version -
attr(x, "created_at"): POSIXct creation timestamp
Pre-aggregated denominator arguments (population_model = "pre_aggregated")
Used when data is already a stratum-level summary.
Wide format dose arguments
Used when data_format = "wide" with record-level data.
Long format dose arguments
Used when data_format = "long" with record-level data.
Vaccine targeting
Controls which vaccine type(s) count as valid doses.
Vaccine validity
Controls how long a dose is considered currently protective.
Cohort arguments
Apply when population_model = "cohort".
Time-series arguments
Apply when population_model = "cohort" and time_series = TRUE.
This is the primary mode for interrupted time series analysis of vaccination
coverage.
See Also
bowerbird::brood_plot() for visualising brood_df objects.
roost for aggregating cases and hospitalisations.
clean_the_nest for upstream data cleaning.
starling::murmuration() for the linkage step that produces the
wide-format linked dataset used as input for cohort coverage designs.
Examples
# Pre-aggregated coverage (no dose-level data needed)
brood(
data.frame(stratum = c("0-17","18-49","50-64","65+"),
n_vaccinated = c(480L,3550L,2370L,1760L),
n_eligible = c(1000L,5000L,3000L,2000L)),
denominator_notes = "ABS ERP 2024, SC LGA."
)
# -- Time-series: monthly Shingrix coverage in a JAK-inhibitor cohort -----
# A small stand-in for what starling::murmuration() would return in a real
# pipeline: one row per person, entry_date_nominal plus up to 2 doses.
linked_cohort <- data.frame(
entry_date_nominal = as.Date(c("2024-01-15", "2024-03-20", "2024-06-10",
"2024-09-05", "2025-01-12", "2025-04-18")),
vax_date_1 = as.Date(c("2024-11-02", NA, "2025-02-14",
"2025-01-20", NA, "2025-06-01")),
vax_type_1 = c("Shingrix", NA, "Shingrix", "Fluvax", NA, "Shingrix"),
vax_date_2 = as.Date(c(NA, "2025-08-15", NA, NA, NA, NA)),
vax_type_2 = c(NA, "Shingrix", NA, NA, NA, NA)
)
vax_date_cols <- grep("^vax_date_", names(linked_cohort), value = TRUE)
vax_type_cols <- grep("^vax_type_", names(linked_cohort), value = TRUE)
cov_ts <- brood(
linked_cohort,
data_format = "wide",
population_model = "cohort",
vax_date_cols = vax_date_cols,
vax_type_cols = vax_type_cols,
vax_target = "Shingrix",
validity_days = Inf,
entry_date_col = "entry_date_nominal",
time_series = TRUE,
ts_start = as.Date("2024-12-01"),
ts_end = as.Date("2026-06-01"),
ts_by = "month",
intervention_date = as.Date("2025-12-01"),
denominator_notes = "JAK-inhibitor cohort, SC HHS. Shingrix only, no expiry."
)
print(cov_ts)
# -- Single time point: pre-intervention snapshot -------------------------
cov_pre <- brood(
linked_cohort,
data_format = "wide",
population_model = "cohort",
vax_date_cols = vax_date_cols,
vax_type_cols = vax_type_cols,
vax_target = "Shingrix",
validity_days = Inf,
entry_date_col = "entry_date_nominal",
window = "current_at_date",
reference_date = as.Date("2025-11-30"),
denominator_notes = "JAK-inhibitor cohort. Pre-intervention Shingrix coverage."
)
# -- Birth cohort: nirsevimab coverage (6-month eligibility window) --------
birth_records <- data.frame(
dob = as.Date(c("2024-05-03", "2024-06-14", "2024-07-22", "2024-08-30")),
nirsevimab_date = as.Date(c("2024-05-10", NA, "2024-07-25", "2025-04-01")),
vax_type_1 = c("nirsevimab", NA, "nirsevimab", "nirsevimab"),
gestation_category = c("Term", "Preterm", "Term", "Preterm")
)
cov_birth <- brood(
birth_records,
data_format = "wide",
population_model = "cohort",
vax_date_cols = "nirsevimab_date",
vax_type_cols = "vax_type_1",
vax_target = "nirsevimab",
entry_date_col = "dob",
eligibility_days = 180L,
censor_date = as.Date("2024-12-31"),
stratum_col = "gestation_category",
denominator_notes = "SCPHU birth cohort 2024. Eligible = 0 to 180 days from birth."
)
Clean and Standardise Surveillance Data
Description
Bird note: The White-winged Chough (Corcorax melanorhamphos) — after
which mudnester is named — builds its mud nest in meticulous layers,
returning again and again to reinforce weak spots before adding anything new
on top. clean_the_nest() is that foundational layer: nothing downstream is
trustworthy until the structure here is sound. Of note, Starlings are cavity
nesters, meaning they prefer to build their homes inside existing holes and
crevices — the cleaned nest clean_the_nest() prepares is what
starling::murmuration() moves into.
Cleans surveillance dataset types and prepares them for record linkage.
This is the first step in the mudnester pipeline. Works with case notification
datasets (linelists such as Notifiable Conditions registers), hospitalisation
datasets (administrative datasets like iPM), vaccination datasets (such as the
Australian Immunisation Register), and generic person-level cohorts
(data_type = "cohort") — any linelist of people (birth cohort, study cohort,
outbreak linelist, flight manifest) being prepared for linkage but without
disease-specific semantics of its own. All datasets should include information
suitable for data linkage — first name, last name, date of birth, Medicare
number, gender, and/or postcode.
Classic pipeline:
-
clean_the_nest()— clean and standardise. Pay close attention to linkage variables (letternames, dob, Medicare number, gender, postcode). -
starling::murmuration()— link cases to vaccination data (c2v). -
starling::murmuration()— linkc2vto hospitalisation data (c2v2h). Vaccination linkage is optional. -
preening()— age categorisation and publication-ready labelling. Works well withgtsummary::tbl_summary(). -
roost()— aggregate to a chosen time unit for visualisation. -
molting()— de-identify before sharing or archiving.
Usage
clean_the_nest(
data,
id_var = NULL,
event_id_var = NULL,
drop_eggs = FALSE,
data_type = NULL,
lie_nest_flat = FALSE,
drop_the_na_vax = TRUE,
keep_vars = NULL,
diagnosis = NULL,
lettername1 = NULL,
lettername2 = NULL,
dob = NULL,
age = NULL,
medicare = NULL,
postcode = NULL,
gender = NULL,
fn = NULL,
latitude = NULL,
longitude = NULL,
onset_date = NULL,
cohort_entry_date = NULL,
cohort_exit_date = NULL,
vax_type = NULL,
vax_date = NULL,
lag = 0,
admission_date = NULL,
discharge_date = NULL,
hospital = NULL,
icd_code = NULL,
diagnosis_description = NULL,
drg = NULL,
icu_date = NULL,
icu_hours = NULL,
dialysis = NULL,
genomics = NULL,
dod = NULL,
died = NULL
)
Arguments
data |
A data frame. Can be a case notifications dataset (infections), hospital admissions, or vaccination dataset. Dates must be in Date format before calling this function. |
id_var |
Character. Name of the column uniquely identifying each
individual. Critical for linkage — must be non-missing. For |
event_id_var |
Character. Name of the column uniquely identifying each event (hospitalisation or vaccination). Must be unique across the whole dataset. Some datasets have duplicate event IDs — this is checked and raises an error. |
drop_eggs |
Logical. If |
data_type |
Character. One of |
lie_nest_flat |
Logical. If |
drop_the_na_vax |
Logical. If |
keep_vars |
Character vector of additional column names to retain when
|
diagnosis |
Character. Name of the column containing the infectious
disease diagnosis (e.g. |
lettername1 |
Character. Name of the first name column. Non-alphanumeric characters are stripped and values lowercased. Middle names are removed (only the first token is kept). |
lettername2 |
Character. Name of the last name column. Non-alphanumeric characters are stripped and values lowercased. Two-part surnames are kept. |
dob |
Character. Name of the date-of-birth column (must be Date class).
Used for record linkage and age calculation. In birth-cohort studies
(e.g. nirsevimab or Abrysvo effectiveness), |
age |
Character. Name of a pre-calculated age column (numeric). Supply
only when you do not want age recalculated from |
medicare |
Character. Name of the Medicare number column. Australian Medicare number structure (11 digits total when the IRN is included):
Derived output columns:
|
postcode |
Character. Name of the postcode column. |
gender |
Character. Name of the gender column. Values are left as-is
(clean to a consistent format before calling, e.g. all |
fn |
Character. Name of the First Nations status column. |
latitude |
Character. Name of the latitude column. |
longitude |
Character. Name of the longitude column. |
onset_date |
Character. Name of the onset/diagnosis date column
(must be Date class). Age is calculated from |
cohort_entry_date |
Character. Name of the cohort entry date column
(must be Date class). The date an individual formally begins accumulating
time-at-risk. Conceptually distinct from |
cohort_exit_date |
Character. Name of the cohort exit date column (must be Date class). The date an individual leaves the analytic cohort (end of follow-up or administrative censoring). |
vax_type |
Character. Name of the vaccine type/brand/antigen column. |
vax_date |
Character. Name of the vaccination event date column (Date
class). Arrange rows in chronological order before calling if you want
|
lag |
Numeric. Days to add to |
admission_date |
Character. Name of the hospital admission date column (must be Date class). |
discharge_date |
Character. Name of the discharge date column (must be
Date class). Length of stay ( |
hospital |
Character. Name of the hospital identifier column. |
icd_code |
Character. Name of the ICD-10-AM code column. |
diagnosis_description |
Character. Name of the ICD code description column (for human readability; not critical for analysis). |
drg |
Character. Name of the AR-DRG column. |
icu_date |
Character. Name of the ICU admission date column (Date class). |
icu_hours |
Character. Name of the ICU hours column (numeric). ICU days are derived automatically. |
dialysis |
Character. Name of the dialysis indicator column (0/1). |
genomics |
Character. Name of the genomics/variant column (e.g. SARS-CoV-2 variant, Hepatitis A genotype). |
dod |
Character. Name of the date-of-death column (Date class). If
both |
died |
Character. Name of a pre-existing death indicator column (0/1).
When |
Value
A data frame with standardised column names, derived linkage
variables (blocking variables, composite name strings), and domain-specific
derived columns (length of stay, age, ICU outcome, death outcome, Medicare
variants). If drop_eggs = TRUE, retains only linkage and analysis
columns.
See Also
preening for age categorisation after cleaning,
roost for time-unit aggregation,
molting for de-identification before sharing,
homing for relinking de-identified data.
Examples
# Case notifications
case_data <- data.frame(
identity = 1:2,
first_name = c("Alex", "Sam"),
surname = c("Nguyen", "Patel"),
dob = as.Date(c("1990-02-14", "1985-11-03"))
)
df_diag <- clean_the_nest(
case_data, data_type = "cases", drop_eggs = TRUE,
id_var = "identity",
lettername1 = "first_name",
lettername2 = "surname",
dob = "dob"
)
# Hospital admissions
hosp_data <- data.frame(
patient_id = 1:2,
firstname = c("Alex", "Sam"),
last_name = c("Nguyen", "Patel"),
birth_date = as.Date(c("1990-02-14", "1985-11-03")),
admission_date = as.Date(c("2025-01-10", "2025-02-20"))
)
df_hosp <- clean_the_nest(
hosp_data, data_type = "hospital", drop_eggs = TRUE,
id_var = "patient_id",
lettername1 = "firstname",
lettername2 = "last_name",
dob = "birth_date",
admission_date = "admission_date"
)
# Vaccination records -- pivot one-row-per-dose data to one row per person
vax_data <- data.frame(
patient_id = c(1, 1, 2),
firstname = c("Alex", "Alex", "Sam"),
last_name = c("Nguyen", "Nguyen", "Patel"),
vax_type = c("Influenza", "COVID-19", "Influenza"),
vax_date = as.Date(c("2025-04-01", "2025-04-01", "2025-04-15"))
)
df_vax <- clean_the_nest(
vax_data, data_type = "vaccination", lie_nest_flat = TRUE,
id_var = "patient_id",
lettername1 = "firstname",
lettername2 = "last_name",
vax_type = "vax_type",
vax_date = "vax_date"
)
# Birth-cohort study (nirsevimab / Abrysvo effectiveness):
# dob is the cohort entry date -- pass the same column name to both.
birth_cohort_data <- data.frame(
baby_id = 1:2,
first_name = c("Ivy", "Leo"),
last_name = c("Chen", "Brown"),
babys_date_of_birth = as.Date(c("2024-03-01", "2024-04-15")),
end_of_followup_date = as.Date(c("2024-09-01", "2024-10-15"))
)
df_cohort <- clean_the_nest(
birth_cohort_data, data_type = "cohort",
id_var = "baby_id",
lettername1 = "first_name",
lettername2 = "last_name",
dob = "babys_date_of_birth",
cohort_entry_date = "babys_date_of_birth",
cohort_exit_date = "end_of_followup_date"
)
Correct Surveillance Counts for Under-Ascertainment
Description
Bird note: Corncrakes are famously detected far more often by ear than
by eye – the overwhelming majority of records are calls, not sightings.
Field surveyors have long applied call-based correction factors to convert
a count of detections into an estimate of the true, largely-unheard
population behind it. corncrake() does the same to a surveillance count:
it takes an observed count, together with a factor that may vary by
stratum and by time, and returns an estimate of the true total sitting
behind the ascertained one.
Applies a stratified, time-varying ascertainment (multiplier) factor to a
count column, appending the resolved factor and a corrected count to the
data – with uncertainty bounds wherever they can be derived. Designed to
receive output from roost (default count_col = "n"
matches directly, and time_col auto-detects from a roost_tbl),
but works on any tidy data frame with a count column.
Factor sources (method):
-
User-supplied (
method = "user_supplied", default): supplyfactor_table, a lookup table of ascertainment factors that may vary bygroup_bystratum and by adate_start/date_endvalidity window. This is the appropriate choice when factors come from a published estimate, an external evaluation study, or expert judgement. -
Ratio estimate (
method = "ratio_estimate"): the factor is derived internally assecondary_count / countat each stratum/time point, using a second, more-complete data stream supplied viasecondary_dataandsecondary_count_col(passed through...). This is the classic surveillance "multiplier method" – e.g. dividing total positive laboratory tests by notified cases to estimate under-notification. -
Severity anchor (
method = "severity_anchor"): the factor is derived by comparing an observed severity ratio indata– e.g. a case-fatality or case-hospitalisation rate, computed asseverity_count_col / count_col– against areference_raterepresenting the believed-true rate from a well-ascertained source (e.g. an externally published infection- fatality rate). This is the case-fatality/infection-fatality-rate anchor inversion described in Smoll et al.'s Queensland COVID-19 under-ascertainment analysis (UAF = CFR_{obs} / IFR_{ref}): under-ascertainment inflates the observed severity ratio above the true one, and their ratio directly recovers the ascertainment factor. It requires no seroprevalence survey, which is its main advantage early in a novel outbreak (WHO's "Disease X" scenario) before any other anchor is available.reference_rate(passed through...) can be a single scalar (with optionalreference_rate_lower/reference_rate_upperfor uncertainty) or afactor_table-shaped lookup with aratecolumn (optionallyrate_lower/rate_upper) for a stratified or time-varying reference rate. The relationship is inverse: because the factor is divided by the reference rate, a higher reference rate produces a lower factor. Bounds therefore cross over – the factor's lower bound is computed from the reference rate's upper bound, and vice versa – whichcorncrake()handles internally so it doesn't need to be worked out by hand for every use. -
Known UAF (
method = "known_uaf"): the reverse direction of"severity_anchor". Rather than deriving the under-ascertainment factor, you supply it – borrowed from a published estimate, another jurisdiction, or a prior"severity_anchor"run – andcorncrake()applies it to recover the true infection burden and corrected severity rates:N_{true} = N_{obs} \times UAFcCHR = hospitalisations / N_{true}cCFR = deaths / N_{true}Supply
uaf(through...) as a single positive number (with optionaluaf_lower/uaf_upper) or afactor_table-shaped data frame with afactorcolumn for a stratified or time-varying UAF (e.g. the quarterly UAFs from the Queensland analysis). Supplyseverity_count_coland/ordeath_count_col(through...) to also computecCHR/cCFR; omit both to getn_trueonly. This is the natural companion to"severity_anchor": one derives a UAF from a reference rate, the other applies a known UAF to get corrected rates – together they cover the "Disease X" preparedness workflow of borrowing another jurisdiction's cCFR/cCHR to estimate your own UAF and true burden, or vice versa. As with"severity_anchor", thecCHR/cCFRbounds invert (a largerN_{true}gives a smaller corrected rate), whichcorncrake()handles internally.
Usage
corncrake(
data,
count_col = "n",
method = "user_supplied",
factor_table = NULL,
group_by = NULL,
time_col = NULL,
on_missing = "warn_na",
ci_method = "table",
denominator_col = NULL,
rate_multiplier = 1e+05,
...
)
Arguments
data |
A data frame, typically the output of |
count_col |
Character. Name of the observed count column. Default
|
method |
Character. |
factor_table |
A data frame of ascertainment factors. Required for
|
group_by |
Character vector of stratifying column names present in
both |
time_col |
Character or |
on_missing |
Character. Behaviour when a row has no matching
ascertainment factor for its stratum/time: |
ci_method |
Character. How |
denominator_col |
Character or |
rate_multiplier |
Numeric. Multiplier used for |
... |
For |
Value
data with the following columns appended, and the same
class as the input (so a roost_tbl stays a roost_tbl –
see roost):
ascertainment_factorThe resolved point factor for that row.
ascertainment_factor_lower,ascertainment_factor_upper-
From
factor_table, if supplied; from the reference rate's own bounds formethod = "severity_anchor"(inverted – see Description);NAformethod = "ratio_estimate", which derives a point estimate only. corrected_countcount_col * ascertainment_factor.corrected_count_lower,corrected_count_upperPer
ci_method;NAifci_method = "none"or unavailable.corrected_rate*Only if
denominator_colsupplied.n_true,n_true_lower,n_true_upperOnly for
method = "known_uaf": the corrected true infection count (N_{obs} \times UAF) and its bounds – identical values tocorrected_count, exposed under the manuscript's name.cCHR,cCHR_lower,cCHR_upperOnly for
method = "known_uaf"whenseverity_count_colis supplied: corrected case-hospitalisation rate (hospitalisations / N_{true}), with inverted bounds.cCFR,cCFR_lower,cCFR_upperOnly for
method = "known_uaf"whendeath_count_colis supplied: corrected case-fatality rate (deaths / N_{true}), with inverted bounds.
Metadata attributes attr(x, "ascertainment_method"),
attr(x, "ascertainment_source"), and attr(x, "corncrake_meta")
(a list: method, count_col, time_col, group_by, ci_method, on_missing,
n_rows, n_missing, mudnester_version, created_at) are attached so
bowerbird::roost_plot() can auto-label a corrected line without
extra arguments.
References
Smoll N, Carroll H, Andrews R, Lambert SB, Khandaker G. Queensland's COVID-19
surveillance: a Bayesian Markov chain Monte Carlo analysis. Manuscript in
submission, Eurosurveillance (2026). The "severity_anchor" and
"known_uaf" methods implement the two directions of the
under-ascertainment-factor relationship used in this analysis
(UAF = CFR_{obs} / IFR_{ref}; N_{true} = N_{obs} \times UAF;
corrected cCHR/cCFR = severity / N_{true}). Note the manuscript's
core estimation is a full seroprevalence-anchored Bayesian MCMC model
(implemented in Stan), which corncrake() deliberately does not
reimplement – these methods are the lightweight, closed-form anchor and
its inverse, suitable for "Disease X" preparedness and for reusing published
UAF/severity estimates from other jurisdictions, not a substitute for the
full Bayesian model.
See Also
roost to build the aggregated counts corncrake() corrects.
flyway to build a single table with several linked event
counts (e.g. cases and deaths together), directly usable as
severity_count_col for method = "severity_anchor".
preening for the age_group (or other) stratification labels
typically used in group_by.
Examples
# -- Stratified + time-varying, user-supplied factors -------------------
cases <- data.frame(
month = as.Date(c("2024-01-01", "2024-02-01", "2024-01-01", "2024-02-01")),
age_group = c("0-17", "0-17", "18+", "18+"),
n = c(10L, 14L, 40L, 55L)
)
factors <- data.frame(
age_group = c("0-17", "0-17", "18+", "18+"),
date_start = as.Date(c("2024-01-01", "2024-02-01", "2024-01-01", "2024-02-01")),
date_end = as.Date(c("2024-01-31", "2024-02-29", "2024-01-31", "2024-02-29")),
factor = c(2.5, 2.2, 1.4, 1.3),
factor_lower = c(1.9, 1.7, 1.1, 1.0),
factor_upper = c(3.3, 2.9, 1.8, 1.7),
source = "Illustrative multiplier, SCPHU surveillance evaluation 2025"
)
corncrake(
cases,
count_col = "n",
factor_table = factors,
group_by = "age_group",
time_col = "month"
)
# Chained straight after roost() -- time_col auto-detected as "month"
df_diag <- data.frame(
onset_date = as.Date(c("2024-01-05", "2024-01-20", "2024-02-02",
"2024-02-14", "2024-01-10", "2024-02-25")),
age_group = c("0-17", "0-17", "18+", "18+", "0-17", "18+")
)
cases_monthly <- roost(df_diag, date_col = "onset_date", time_unit = "month",
group_cols = "age_group")
cases_corrected <- corncrake(
cases_monthly,
factor_table = factors,
group_by = "age_group"
)
# Ratio (multiplier) method: derive the factor from a more-complete
# secondary stream (e.g. all positive laboratory tests) instead of
# supplying one directly
lab_positive_monthly <- data.frame(
age_group = c("0-17", "0-17", "18+", "18+"),
month = as.Date(c("2024-01-01", "2024-02-01", "2024-01-01", "2024-02-01")),
n = c(25L, 30L, 90L, 110L)
)
cases_corrected2 <- corncrake(
cases_monthly,
method = "ratio_estimate",
group_by = "age_group",
secondary_data = lab_positive_monthly,
secondary_count_col = "n"
)
# Severity anchor: early in a novel ("Disease X") outbreak, before any
# seroprevalence survey exists, invert an externally published
# infection-fatality rate against locally observed deaths and cases --
# requires flyway(), not roost(), since both a case count and a death
# count are needed together
df_linked <- data.frame(
onset_date = as.Date(c("2025-01-05", "2025-01-20", "2025-02-02",
"2025-02-14", "2025-03-01", "2025-03-10")),
fatality_date = as.Date(c(NA, NA, "2025-02-10", NA, NA, NA)),
age_group = c("18-64", "65+", "65+", "18-64", "65+", "18-64")
)
linked_monthly <- flyway(
df_linked,
events = c(cases = "onset_date", deaths = "fatality_date"),
group_cols = "age_group"
)
cases_corrected3 <- corncrake(
linked_monthly,
count_col = "n_cases",
method = "severity_anchor",
group_by = "age_group",
severity_count_col = "n_deaths",
reference_rate = 0.01, # reference IFR, e.g. from WHO or a
# high-ascertainment reference jurisdiction
reference_rate_lower = 0.005,
reference_rate_upper = 0.020,
reference_source = "WHO Disease X planning scenario, IFR 1.0% (0.5-2.0%)"
)
# Known UAF (reverse direction): a UAF has already been estimated (here, the
# cumulative Queensland 2022 value of 2.98) -- apply it to recover the true
# infection count and the corrected severity rates from observed counts.
obs <- data.frame(
quarter = as.Date(c("2022-01-01", "2022-04-01",
"2022-07-01", "2022-10-01")),
n = c(670083L, 432457L, 246464L, 80652L), # notified cases
hospitalisations = c(6902L, 5478L, 5093L, 4278L),
deaths = c(734L, 578L, 699L, 318L)
)
qld_corrected <- corncrake(
obs,
count_col = "n",
method = "known_uaf",
time_col = "quarter",
uaf = 2.98, # cumulative UAF (95% CrI 2.78-3.29)
uaf_lower = 2.78,
uaf_upper = 3.29,
uaf_source = "Smoll et al. 2026, Queensland COVID-19 cumulative UAF",
severity_count_col = "hospitalisations",
death_count_col = "deaths"
)
# -> adds n_true, cCHR (+bounds), cCFR (+bounds)
# A time-varying UAF is supplied the same way, as a factor_table-shaped
# `uaf` with a `factor` column (e.g. the manuscript's quarterly UAFs).
quarterly_uaf <- data.frame(
date_start = as.Date(c("2022-01-01", "2022-04-01",
"2022-07-01", "2022-10-01")),
date_end = as.Date(c("2022-03-31", "2022-06-30",
"2022-09-30", "2022-12-31")),
factor = c(2.35, 2.60, 3.10, 4.08),
factor_lower = c(2.08, 2.30, 2.60, 2.41),
factor_upper = c(2.68, 3.00, 3.90, 6.38)
)
qld_corrected_tv <- corncrake(
obs,
count_col = "n",
method = "known_uaf",
time_col = "quarter",
uaf = quarterly_uaf,
severity_count_col = "hospitalisations",
death_count_col = "deaths"
)
Aggregate Multiple Sequential Event Dates in One Table
Description
Bird note: A flyway is the established migratory corridor linking a
bird population's successive stopover and staging sites – the same birds,
traced through each waypoint of a single journey. flyway() does the same
for a linked case cohort: it traces the same population through however
many clinical milestones matter (onset, admission, ICU, complication,
death), aggregating each waypoint separately by its own date, and folding
the results into one table aligned by time unit and stratum.
Calls roost once per date column named in events, then
joins the results on the shared time-unit and group_cols columns,
producing one row per stratum/time point with one count column per event,
named n_ followed by that event's name (e.g. n_cases). This is the
natural next step after a
case-to-hospital (or similar) linkage via starling::murmuration(),
where a single linked record carries several milestone dates for the same
person – e.g. to compute an observed case-hospitalisation rate as
n_hospitalisations / n_cases directly from one table, or to feed
n_hospitalisations/n_deaths into corncrake's
method = "severity_anchor".
Implemented as a thin wrapper, not a change to roost itself:
each event's own calendar-grid zero-filling, NA-date handling, and
validation is exactly roost's own, unchanged. A stratum/time
point present for one event but absent for another (e.g. no deaths yet
recorded in a stratum where cases already exist) is zero-filled after the
join, consistent with roost's own zero-filling philosophy.
Usage
flyway(
data,
events,
time_unit = "month",
group_cols = NULL,
week_start = 1,
hemisphere = c("southern", "northern"),
warn_na = TRUE,
anchor_event = NULL
)
Arguments
data |
A data frame of linked surveillance records, with one date column per milestone event. |
events |
A named character vector (or list) mapping short event
names to date column names in |
time_unit |
Passed through to every underlying |
group_cols |
Passed through to every underlying |
week_start |
|
hemisphere |
|
warn_na |
|
anchor_event |
|
Value
A flyway_tbl object (which also inherits roost_tbl, so it
works directly with corncrake's time_col auto-detection
and with the existing roost_tbl print/subset methods): a tibble with
any group_cols, optionally epiyear, the named time-unit column, and
one n_-prefixed count column per entry in events. Metadata is
attached as attr(x, "roost_meta") (matching roost's
own shape) and attr(x, "flyway_meta") (per-event detail,
including each event's own roost_meta, and – when set –
which event anchor_event named).
Two aggregation modes – independent vs. case-linked (anchor_event)
By default (anchor_event = NULL), every event is bucketed by ITS OWN
date, independently – n_cases in a given period counts onsets falling
in that period; n_hospitalisations counts admissions falling in that
same period, a separate population, not "of the cases whose onset was in
this period, how many were hospitalised". Since a downstream milestone
(admission, ICU, death) can fall some days or weeks after the case's own
onset, this is a good approximation summed over a long enough period –
but it can produce an impossible rate (a hospitalised count exceeding the
case count) over a narrow window, whenever admission volume in a nearby
period doesn't track onset volume in the window being looked at. This is
a real, confirmed failure mode, not a hypothetical one – see the
anchor_event example below.
Setting anchor_event (naming one entry in events, typically the case/
onset event) switches to case-linked aggregation instead: every OTHER
event is bucketed by the ANCHOR's own date, not its own – "of the rows
whose anchor fell in this period, how many also have a non-NA value in
this other event's date column". A case and its own downstream
hospitalisation/ICU/death are then always attributed to the same period,
whichever period the milestone itself actually fell in, which is what
makes the resulting counts a valid ratio (e.g. n_hospitalisations / n_cases) even over a single narrow period, not just when summed over a
long one. The anchor event's own row-level accounting (calendar-grid
building, zero-filling, NA-dropping) is still exactly roost's
own, applied to the anchor's date column alone.
See Also
roost, which flyway() calls once per event.
corncrake, which can consume flyway() output directly –
e.g. count_col = "n_cases", method = "severity_anchor",
severity_count_col = "n_deaths".
Examples
# After a case-to-hospital linkage via starling::murmuration(), a single
# linked record carries several milestone dates for the same person
linked_cohort <- data.frame(
onset_date = as.Date(c("2025-01-05", "2025-01-20", "2025-02-02",
"2025-02-14", "2025-03-01", "2025-03-10")),
admission_date = as.Date(c("2025-01-08", NA, "2025-02-05",
NA, "2025-03-03", NA)),
icu_date = as.Date(c(NA, NA, "2025-02-07", NA, NA, NA)),
complication_date = as.Date(c(NA, NA, NA, NA, "2025-03-06", NA)),
fatality_date = as.Date(c(NA, NA, "2025-02-10", NA, NA, NA)),
age_group = c("18-64", "65+", "65+", "18-64", "65+", "18-64")
)
linked_monthly <- flyway(
linked_cohort,
events = c(
cases = "onset_date",
hospitalisations = "admission_date",
icu = "icu_date",
complications = "complication_date",
deaths = "fatality_date"
),
time_unit = "month",
group_cols = "age_group"
)
# Observed (uncorrected) case-hospitalisation rate, directly from one table
linked_monthly$chr_obs <- linked_monthly$n_hospitalisations / linked_monthly$n_cases
# anchor_event: the same call, but computing a rate over a narrow window
# (the most recent 12 months, say) needs case-linked aggregation, not the
# default independent bucketing above -- without anchor_event, a burst of
# admissions linked to cases whose own onset fell just before the window
# can push n_hospitalisations above n_cases for that same window, an
# impossible rate.
linked_monthly_linked <- flyway(
linked_cohort,
events = c(
cases = "onset_date",
hospitalisations = "admission_date",
icu = "icu_date",
complications = "complication_date",
deaths = "fatality_date"
),
time_unit = "month",
group_cols = "age_group",
anchor_event = "cases"
)
Relink De-identified Data Using the molting() Lookup Table
Description
Bird note: Homing pigeons are renowned for their extraordinary ability
to return to their loft from any location — navigating hundreds of kilometres
to find their way back to the place they started. homing() does the same
thing for a de-identified dataset: it takes records that have been stripped
of their identifiers by molting and navigates them back to the
original identities, using the lookup table as its internal compass. The
path home is only available to those who hold the lookup.
Relinks de-identified data with original identifiers using the lookup table
created by molting. The hash column present in both the
de-identified data and the lookup table is the join key. Typically called
only when authorised re-identification is required (e.g. clinical follow-up,
audit).
Usage
homing(
deidentified_data,
lookup_table,
hash_col_name = "row_hash",
keep_hash = TRUE
)
Arguments
deidentified_data |
A de-identified data frame containing a hash column.
Typically |
lookup_table |
The lookup table data frame mapping hashes back to
original identifiers. Typically |
hash_col_name |
Character. Name of the hash column used as the join key.
Must exist in both |
keep_hash |
Logical. If |
Value
A data frame with the original identifiers merged back into the
de-identified data. The class and column order of deidentified_data are
preserved, with identifier columns appended on the right (or left of the
hash column if keep_hash = FALSE).
See Also
molting to create the de-identified data and lookup table,
clean_the_nest for the upstream data preparation step.
Examples
# Create sample data and de-identify it (same data molting()'s own
# example uses, so the two paired examples stay consistent)
patient_data <- data.frame(
patient_name = c("John Doe", "Jane Smith"),
dob = as.Date(c("1980-01-01", "1975-05-15")),
mrn = c("12345", "67890"),
age5cat = factor(c("18-64", "18-64")),
diagnosis = c("Condition A", "Condition B"),
lab_value = c(120, 95)
)
result <- suppressMessages(molting(patient_data))
# Relink when authorised follow-up is needed
relinked <- homing(
deidentified_data = result$deidentified,
lookup_table = result$lookup
)
# Drop the hash column from the final relinked dataset
relinked_clean <- homing(
result$deidentified,
result$lookup,
keep_hash = FALSE
)
List Available Age Banding Schemes
Description
Bird note: see preening() — this is preening's companion lookup function,
a quick way to see every "arrangement" of feathers (age schemes) on offer before
picking one, narrowed down to just the arrangements relevant to the task at hand.
Prints (and invisibly returns) a data frame of age schemes available to
preening(): scheme name, family, focus tags, age range covered, number of
bands, and source citation. By default lists all ~50 schemes; supply family,
focus, and/or max_bands to return only a relevant subset — e.g. just the
paediatric-focused schemes, or just the WHO-standard ones — instead of scrolling
through all 50 every time.
Usage
list_age_schemes(family = NULL, focus = NULL, max_bands = NULL)
Arguments
family |
Optional character vector filter, one or more of
|
focus |
Optional character vector filter on cross-cutting focus tags, e.g.
|
max_bands |
Optional integer. Only include schemes with at most this many age
bands. |
Value
Invisibly, a tibble with columns scheme, family, focus
(list-column of tags), n_bands, age_range, source. Printed to
console by default, sorted by family then scheme name.
Examples
list_age_schemes() # all ~50
list_age_schemes(focus = "paediatric") # paediatric-focused only
list_age_schemes(focus = "who_standard") # WHO standard schemes only
list_age_schemes(family = "surveillance", max_bands = 6) # compact surveillance schemes
De-identify a Dataset with Hash-Based Relinking
Description
Bird note: Moulting is how a bird sheds its identifying plumage —
distinctive breeding colours, worn feathers, recognisable patterns — and
replaces it with something uniform and anonymous. molting() does the same
to a dataset: it strips the identifying columns and replaces them with a
cryptographic hash for each row, so the bird (patient) is still
distinguishable by its new uniform feathers, but cannot be named from them.
The old plumage is kept safely in a lookup table, which homing
uses to re-dress the bird when authorised relinking is needed.
Removes personally identifiable information (PII) from a data frame using
column-name pattern matching, creates a cryptographic row hash from the
removed identifiers, and returns the de-identified data alongside a secure
lookup table for later relinking via homing. Age category
variables (age2cat, age5cat, etc.) are automatically preserved as they
are not directly identifying.
Security note: The lookup table contains sensitive information and must be stored securely with appropriate access controls, separate from the de-identified dataset. Consider encrypting the lookup file before archiving.
Usage
molting(
data,
id_cols = NULL,
pii_patterns = NULL,
additional_pii_cols = NULL,
hash_method = "sha256",
hash_col_name = "row_hash",
return_lookup = TRUE,
seed = NULL
)
Arguments
data |
A data frame to de-identify. |
id_cols |
Optional character vector of column names to use for hashing.
If |
pii_patterns |
Optional character vector of regular expression patterns used to detect PII columns for removal. Defaults to a standard list covering common identifiers (name, address, phone, email, dob, mrn, etc.). |
additional_pii_cols |
Optional character vector of specific column names to remove as PII in addition to those detected by pattern matching. Useful for dataset-specific identifiers without modifying patterns. |
hash_method |
Character. Hashing algorithm. One of |
hash_col_name |
Character. Name for the new hash column. Default
|
return_lookup |
Logical. If |
seed |
Optional integer seed for reproducible hashing with seeded
algorithms. Default |
Value
If return_lookup = TRUE (default): a named list with elements
$deidentified (de-identified data frame with hash column) and $lookup
(the lookup table mapping hash to removed identifier columns). If
return_lookup = FALSE: only the de-identified data frame.
See Also
homing to relink de-identified data using the lookup table,
clean_the_nest to prepare and standardise data before
de-identification, roost to aggregate data (aggregated counts
can often be shared without de-identification).
Examples
# Create sample data
patient_data <- data.frame(
patient_name = c("John Doe", "Jane Smith"),
dob = as.Date(c("1980-01-01", "1975-05-15")),
mrn = c("12345", "67890"),
age5cat = factor(c("18-64", "18-64")),
diagnosis = c("Condition A", "Condition B"),
lab_value = c(120, 95)
)
# Basic de-identification — age categories automatically retained
result <- suppressMessages(molting(patient_data))
names(result$deidentified)
# Use MD5 for a shorter hash
result_md5 <- suppressMessages(molting(patient_data, hash_method = "md5"))
# Return only de-identified data (no lookup — irreversible)
deidentified_only <- suppressMessages(molting(patient_data, return_lookup = FALSE))
# Specify exactly which columns to hash
result_ids <- suppressMessages(molting(patient_data, id_cols = c("mrn", "dob")))
Identify Chronic Comorbidities Using ICD-10-AM U-Codes and Acute Admission Codes
Description
Bird note: A bird's plumage is the outward, visible sign of what's
underneath – condition, age, health, even diet all show up in feather
quality and colour to a practised observer. plumage() does the same for
a hospitalisation record: the ICD-10-AM codes on file are the visible
markings, and this function reads them to surface the chronic conditions
underneath, turning a code list into a visible comorbidity profile.
Analyses a hospitalisation dataset to identify chronic conditions based on both ICD-10-AM U-codes (supplementary chronic condition codes) and principal/secondary acute admission ICD-10-AM codes.
The dual-code strategy is necessary because patients are frequently admitted under
an acute diagnosis code (e.g., J44.1 for COPD exacerbation) without a supplementary
U-code being recorded. Using only U-codes will undercount comorbidity burden, and
will systematically miss exactly the admissions where the condition is the ACTIVE
reason for the encounter rather than a background comorbidity noted in passing –
see active_inactive below.
Usage
plumage(
df,
icd_column,
prefix = NULL,
decimal = TRUE,
drop_eggs = FALSE,
include_drg = FALSE,
drg_column = NULL,
active_inactive = "both"
)
Arguments
df |
A data frame containing hospitalisation records. |
icd_column |
Character string specifying the name of the column containing ICD-10-AM codes. Multiple codes per row should be separated by any non-alphanumeric character (comma, space, semicolon, pipe, etc.). |
prefix |
Optional character string to prefix all output column names.
Default is |
decimal |
Logical. Should codes be matched with decimal points ( |
drop_eggs |
Logical. If |
include_drg |
Logical. If |
drg_column |
Optional character string specifying a separate column containing
AR-DRG codes. Used only when |
active_inactive |
Character string: one of
|
Details
Code Strategy
Each chronic condition is detected by searching for any of the following that appear in the specified column(s):
-
U-code — ICD-10-AM supplementary chronic condition code (e.g.,
U83.2for COPD). This is the inactive signal: it marks the condition as a known comorbidity, without indicating it was the reason for this particular admission. -
Acute ICD-10-AM code — Principal or secondary admission code that unambiguously indicates the chronic condition (e.g.,
J44codes for COPD,E84for cystic fibrosis). This is the active signal: it marks the condition as clinically active in this admission. -
AR-DRG code (optional) — Australian Refined Diagnosis Related Group code (e.g.,
E65A/E65Bfor COPD), searched only wheninclude_drg = TRUE. Also counted as an active signal.
A patient is flagged for a condition (the <condition> column) according to
active_inactive: "both" (default, and the only behaviour previous versions of
this function had) if any matching code is detected; "active" for the acute
ICD-10-AM/AR-DRG signal alone; "inactive" for the U-code signal alone. See that
argument's description above.
icd_stems (the acute-code definitions) were originally populated only for the
Respiratory conditions (plus cystic fibrosis). This revision adds icd_stems for
the remaining 23 conditions, using standard ICD-10-AM chapter-level codes, so that
active_inactive = "active" is meaningful for all conditions listed below, not
only the respiratory ones – confirmed accurate. No AR-DRG codes were added for
these 23 – drg_codes remains empty for all of them, so include_drg still only
adds signal for the original six respiratory/CF conditions.
Conditions Detected
Metabolic / Endocrine:
Obesity — U78.1; E66
Cystic fibrosis — U78.2; E84, E84.0–E84.9 (also counted in Respiratory)
Mental Health:
Dementia — U79.1; F00–F03, G30
Schizophrenia — U79.2; F20
Depression — U79.3; F32, F33
Intellectual/developmental disability — U79.4; F70–F73, F78, F79
Neurological:
Parkinson's disease — U80.1; G20
Multiple sclerosis — U80.2; G35
Epilepsy — U80.3; G40
Cerebral palsy — U80.4; G80
Paralysis — U80.5; G81, G82, G83
Cardiovascular:
Ischaemic heart disease — U82.1; I20–I25
Heart failure — U82.2; I50
Hypertension — U82.3; I10, I11, I12, I13, I15
Respiratory:
Emphysema — U83.1; J43, J43.0–J43.9, J98.2, J98.3
COPD — U83.2; J44, J44.0–J44.9; DRG E65 (root; any complexity split)
Asthma/Chronic bronchitis — U83.3; J41, J41.0–J41.8, J42, J45, J45.0–J45.9, J46; DRG E69 (root; any complexity split)
Bronchiectasis — U83.4; J47; DRG E77 (root; any complexity split)
Chronic respiratory failure — U83.5; J96.1, J96.10–J96.19
Cystic fibrosis — U78.2; E84, E84.0–E84.9; DRG E60 (root; any complexity split)
Gastrointestinal:
Crohn's disease — U84.1; K50
Ulcerative colitis — U84.2; K51
Liver failure — U84.3; K72
Musculoskeletal:
Rheumatoid arthritis — U86.1; M05, M06
Osteoarthritis — U86.2; M15–M19
Systemic lupus erythematosus — U86.3; M32
Osteoporosis — U86.4; M80, M81
Renal:
Chronic kidney disease — U87.1; N18
Congenital:
Spina bifida — U88.1; Q05
Down syndrome — U88.2; Q90
ICD-10-AM Code Matching
Acute ICD-10-AM codes are matched at the category level using anchored regex. For
example, specifying J44 matches J44, J44.0, J44.1, J44.8, and J44.9 but
does NOT match J440 or J449 when decimal = TRUE (the default). When
decimal = FALSE, the same block-level match applies without a decimal separator.
Value
The input data frame with additional columns appended:
Individual condition indicators (unless drop_eggs = TRUE): one binary (0/1)
column per condition, optionally prefixed. active_inactive controls what "1"
means for each of these columns (see that argument's description above) – it
does not add any extra columns.
Summary columns:
-
total_conditions— sum of all individual condition flags -
total_metabolic_conditions -
total_mental_health_conditions -
total_neurological_conditions -
total_cardiovascular_conditions -
total_respiratory_conditions -
total_gastrointestinal_conditions -
total_musculoskeletal_conditions -
total_renal_conditions -
total_congenital_conditions -
conditions_category— ordered factor with levels"0","1","2","3+"
All summary column names are prefixed when prefix is specified. Summary totals
are computed from whichever active_inactive mode was selected – e.g. with
active_inactive = "active", total_conditions counts only active conditions.
Examples
# --- Example 1: mixed U-code and acute ICD-10-AM codes ---
hospital_data <- data.frame(
patient_id = 1:5,
icd_codes = c(
"K29.70",
"U78.1, U83.2, U82.3", # U-codes: obesity, COPD, hypertension
"J44.1, U79.3", # Acute COPD admission + depression U-code
"J43.2, J47", # Acute emphysema + bronchiectasis (no U-codes)
"E84.0, U80.3" # Acute cystic fibrosis + epilepsy U-code
)
)
results <- plumage(hospital_data, "icd_codes")
results[, c("patient_id", "copd", "emphysema", "bronchiectasis",
"cystic_fibrosis", "total_respiratory_conditions")]
# --- Example 2: no-decimal format (codes stored without dots) ---
hospital_data2 <- data.frame(
patient_id = 1:2,
icd_codes = c("U832 U823", "J441 J431")
)
results2 <- plumage(hospital_data2, "icd_codes", decimal = FALSE)
# --- Example 3: include AR-DRG codes from a separate column ---
hospital_data3 <- data.frame(
patient_id = 1:3,
icd_codes = c("K29.70", "J44.1", "J45.0"),
drg_codes = c("G07B", "E65A", "E69B")
)
results3 <- plumage(hospital_data3, "icd_codes",
include_drg = TRUE, drg_column = "drg_codes")
# --- Example 4: prefix + drop individual columns ---
results4 <- plumage(hospital_data, "icd_codes",
prefix = "chr_", drop_eggs = TRUE)
# --- Example 5: active vs inactive comorbidity status ---
hospital_data5 <- data.frame(
patient_id = 1:3,
icd_codes = c(
"N18.5", # CKD as the acute reason for admission -> active only
"U87.1", # CKD noted as a background comorbidity -> inactive only
"N18.5, U87.1" # both present
)
)
results5_both <- plumage(hospital_data5, "icd_codes") # active_inactive = "both" (default)
results5_active <- plumage(hospital_data5, "icd_codes", active_inactive = "active")
results5_inactive <- plumage(hospital_data5, "icd_codes", active_inactive = "inactive")
data.frame(
patient_id = 1:3,
both = results5_both$kidney_disease,
active = results5_active$kidney_disease,
inactive = results5_inactive$kidney_disease
)
Age Categorisation and Variable Labelling
Description
Bird note: Preening is how a bird sorts every feather into one of a small set
of functional positions, re-arranging the same plumage to suit the moment without
changing the bird underneath. preening() does the same to a population: the
same raw age column is re-sorted into whichever of the ~50 standard schemes
an analysis calls for.
Optional ad hoc utility for age categorisation and publication-ready variable
labelling. Can be applied at any point in the pipeline — before or after
roost(). Draws on a library of ~50 named age-banding schemes spanning
Australian and international statistical standards, vaccination/clinical
guidance, surveillance-system conventions, and disease-specific research bands
(see vignette("age-schemes") for the full referenced list).
roost() does not require preening() to have been called first.
Users are not required to know all ~50 scheme names. scheme can be a single
exact name, OR left NULL and narrowed instead with family, focus,
and/or max_bands — the same filters exposed by list_age_schemes() — so a
paediatric-only analysis, say, only ever sees paediatric-relevant schemes rather than
all 50.
Usage
preening(
data,
age_col = NULL,
scheme = NULL,
family = NULL,
focus = NULL,
max_bands = NULL,
age_unit = "years",
age_breaks = NULL,
age_labels = NULL,
label_vars = NULL,
label_style = "figure"
)
Arguments
data |
A data frame at any stage of processing. |
age_col |
Name of numeric age-in-years column (or age-in-days for neonatal/infant
schemes — see |
scheme |
Name of a single age banding scheme to apply, e.g.
|
family |
Optional character vector to filter candidate schemes by family before
selection — one or more of |
focus |
Optional character vector to filter candidate schemes by cross-cutting
focus tag, e.g. |
max_bands |
Optional integer. Restrict candidates to schemes with at most this many
age bands. Ignored (with a warning) if |
age_unit |
Unit of the values in |
age_breaks |
Numeric vector of custom break points (only used if
|
age_labels |
Character vector of custom labels (length must equal
|
label_vars |
Character vector of column names to apply publication-ready
labels to. Independent of age banding — can be used with or without |
label_style |
|
Details
Selecting a scheme — three ways:
-
Exact:
preening(data, age_col = "age", scheme = "atagi_covid19_2025")— applies that one scheme directly, no filtering involved. -
Filtered, single match:
preening(data, age_col = "age", family = "vaccination", focus = "paediatric")— if exactly one scheme matches the filter, it is applied automatically (with amessage()naming which scheme was selected, so the choice is never silent). -
Filtered, multiple matches: if more than one scheme matches the filter,
preening()does not guess — itstop()s and prints the matching scheme names so the caller can pick one explicitly viascheme.
Call list_age_schemes() for the full catalogue, or with the same family /
focus / max_bands filters to preview what a given filter combination would
match before calling preening().
Value
Input data frame with new or relabelled columns appended.
Original columns are always preserved. If age banding was performed, a new
age_group column is added as an ordered factor, and the scheme's source
citation is attached as attr(x, "age_scheme_source").
Examples
df <- data.frame(age = c(2, 8, 15, 35, 62, 80))
# Exact scheme
preening(df, age_col = "age", scheme = "atagi_covid19_2025")
# All paediatric-focused schemes, regardless of family -- errors with a
# printed list of matches, since several paediatric schemes exist. Wrapped
# in try() here so the error (the point of this example) doesn't fail the
# example check itself.
try(preening(df, age_col = "age", focus = "paediatric"))
# Narrow further: paediatric AND vaccination family -- confirmed against a
# real check run to still match 2 schemes (pneumococcal_program,
# rsv_maternal_infant), not 1 as filters alone might suggest -- also
# wrapped in try() for the same reason, then resolved by supplying scheme
# explicitly, using one of the two names the error itself lists
try(preening(df, age_col = "age", family = "vaccination", focus = "paediatric"))
preening(df, age_col = "age", scheme = "rsv_maternal_infant")
# WHO-standard schemes only -- wrapped defensively as well, since whether
# this resolves to one match or several depends on the real age_schemes
# data and hasn't been independently confirmed the way the case above was
try(preening(df, age_col = "age", focus = "who_standard"))
# Custom breaks, bypassing the scheme library entirely -- no dependency on
# age_schemes data at all, so no ambiguity is possible here
preening(df, age_col = "age", scheme = "custom",
age_breaks = c(0, 18, 65, Inf),
age_labels = c("0-17", "18-64", "65+"))
Print Method for brood_df Objects
Description
Print Method for brood_df Objects
Usage
## S3 method for class 'brood_df'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments passed to |
Value
Invisibly returns x.
Build an Epicurve: Aggregate Surveillance Records by Time Unit
Description
Bird note: A roost is where a flock settles at the end of the day —
many individual birds resolving into one countable gathering at a shared
site and time. roost() does exactly the same thing to records: many
individual rows resolve into counts, grouped by whichever time unit
matters for the analysis — day, ISO week, fortnight, month, quarter, year,
epidemiological week, or season.
roost() is the epicurve builder underneath almost everything on the
surveillance side of the aviary ecosystem. Anywhere you'd draw an epicurve
— case incidence, hospitalisations, ICU admissions, deaths, vaccine doses
delivered — roost() is the function turning line-list rows into the
counts that curve is built from. It fills zero-count periods explicitly
(so epi curves never skip a quiet week) and applies hemisphere-aware
season and season-year labels. Designed to receive output from
clean_the_nest and optionally preening (though
fully standalone — neither is required), producing a roost_tbl ready for
bowerbird::roost_plot().
For a single cohort that needs tracing through several downstream
milestone dates at once (onset, admission, ICU, death) into one aligned
table, see flyway, roost()'s own wrapper for exactly that
shape of problem (Example 2 below) — the natural next step after a
starling::murmuration() linkage, and the way to get case,
hospitalisation, ICU, and death counts into one data frame for a severity
cascade or feeding corncrake.
Usage
roost(
data,
date_col,
time_unit = "month",
event_cols = NULL,
group_cols = NULL,
week_start = 1,
hemisphere = c("southern", "northern"),
warn_na = TRUE
)
Arguments
data |
A data frame of surveillance records, typically the output of
|
date_col |
Character. The name of the date column (quoted, e.g.
|
time_unit |
Character. Aggregation unit. One of |
event_cols |
Character vector of binary (0/1) column names to sum
alongside the row count. Use for outcomes such as ICU admission or death.
|
group_cols |
Character vector of stratification column names
(e.g. |
week_start |
Integer 1–7 (1 = Monday, 7 = Sunday). Sets the first day
of the ISO week when |
hemisphere |
One of |
warn_na |
Logical. Warn when rows with |
Value
A roost_tbl object: a tibble with columns for any group_cols,
optionally epiyear (when time_unit = "epiweek"), the named time-unit
column, n (row count), and any event_cols. Zero-count periods are
filled with 0. Metadata is attached as attr(x, "roost_meta").
Worldwide use, Australian home turf
roost() and the rest of the aviary ecosystem work on surveillance data
from anywhere — there's nothing Australia-specific in the aggregation
logic itself. The defaults, age schemes, and worked examples are tuned for
Australian public health practice (Southern Hemisphere seasons by default,
ATAGI-aligned age bands in preening, jurisdictional datasets
like AIR and NNDSS in the surrounding tooling) because that's the practice
this ecosystem was built inside of. Set hemisphere = "northern" for
time_unit = "season" or "season_year" elsewhere, and everything else
applies unchanged.
See Also
clean_the_nest to prepare data before aggregation,
preening to add age groups before stratified aggregation,
flyway to aggregate several milestone dates for the same
cohort into one table, corncrake to correct a roost_tbl's
counts for under-ascertainment, molting to de-identify the
aggregated output before sharing.
Examples
# ---------------------------------------------------------------------------
# Example 1: Yearly RSV case counts
# A straightforward epicurve -- confirmed RSV notifications aggregated to
# yearly counts for a multi-year burden-of-disease summary.
# ---------------------------------------------------------------------------
rsv_cases <- data.frame(
notification_date = as.Date(c("2022-04-12", "2022-07-03", "2023-05-20",
"2023-08-14", "2024-04-30", "2024-09-02"))
)
rsv_yearly <- roost(rsv_cases, date_col = "notification_date", time_unit = "year")
rsv_yearly
# ---------------------------------------------------------------------------
# Example 2: Influenza, 3 seasons, severity cascade via flyway()
# Confirmed influenza cases across three seasons, linked (via
# starling::murmuration()) to hospitalisation, ICU, and death dates.
# flyway() runs roost() once per milestone and joins the results into one
# table -- the shape a severity cascade or n_hospitalisations/n_cases
# observed-CHR calculation needs.
# ---------------------------------------------------------------------------
influenza_linked_cohort <- data.frame(
onset_date = as.Date(c("2023-06-10", "2023-07-22", "2024-06-05",
"2024-08-14", "2025-06-20", "2025-07-30")),
admission_date = as.Date(c("2023-06-13", NA, "2024-06-08",
NA, "2025-06-23", NA)),
icu_date = as.Date(c(NA, NA, "2024-06-10", NA, NA, NA)),
fatality_date = as.Date(c(NA, NA, "2024-06-15", NA, NA, NA))
)
influenza_by_season <- flyway(
influenza_linked_cohort,
events = c(
cases = "onset_date",
hospitalisations = "admission_date",
icu = "icu_date",
deaths = "fatality_date"
),
time_unit = "season_year"
)
influenza_by_season$chr_obs <-
influenza_by_season$n_hospitalisations / influenza_by_season$n_cases
# ---------------------------------------------------------------------------
# Example 5: Disease X, multiple time units and age bands via preening()
# The same case data viewed two ways: a fine-grained weekly curve for
# outbreak monitoring, and a coarser monthly curve stratified by a
# preening()-derived age band for a routine surveillance report.
# ---------------------------------------------------------------------------
disease_x_cases <- data.frame(
onset_date = as.Date(c("2025-01-05", "2025-01-12", "2025-01-20",
"2025-02-02", "2025-02-14", "2025-02-25")),
age_years = c(5, 34, 68, 12, 45, 80)
)
disease_x_cases <- preening(disease_x_cases, age_col = "age_years",
scheme = "atagi_covid19_2025")
disease_x_weekly <- roost(disease_x_cases, date_col = "onset_date",
time_unit = "isoweek")
disease_x_monthly_by_age <- roost(disease_x_cases, date_col = "onset_date",
time_unit = "month",
group_cols = "age_group")
# ---------------------------------------------------------------------------
# Example 3: RSV hospitalisations only, for an interrupted time series
# Monthly RSV-associated hospitalisation counts in infants, stratified by
# era (pre/post a nirsevimab programme start date) -- straight into
# bowerbird::roost_plot(facet_by = "era") for a pre/post comparison panel.
# ---------------------------------------------------------------------------
rsv_infant_admissions <- data.frame(
admission_date = as.Date(c("2023-05-02", "2023-06-18", "2023-07-30",
"2024-05-14", "2024-06-22", "2024-08-05")),
era = c("Pre-nirsevimab", "Pre-nirsevimab", "Pre-nirsevimab",
"Post-nirsevimab", "Post-nirsevimab", "Post-nirsevimab")
)
rsv_hosp_monthly <- roost(
rsv_infant_admissions,
date_col = "admission_date",
time_unit = "month",
group_cols = "era"
)
# ---------------------------------------------------------------------------
# Example 4: Measles vaccine doses delivered during a rollout
# Total measles vaccine doses administered per week during an outbreak-
# response rollout -- the delivery-tempo curve, not a coverage/eligibility
# question (that's brood()'s job -- see brood()'s own examples).
# ---------------------------------------------------------------------------
measles_vax_records <- data.frame(
vax_date = as.Date(c("2025-03-03", "2025-03-05", "2025-03-10",
"2025-03-12", "2025-03-17", "2025-03-19"))
)
measles_doses_weekly <- roost(
measles_vax_records,
date_col = "vax_date",
time_unit = "isoweek"
)