In Queensland hospital admission data, a chronic condition is not
always coded consistently. A patient with COPD admitted for an
exacerbation may be coded under the acute admission code
J44.1 without the supplementary U-code U83.2
being recorded alongside it. A strategy that searches U-codes alone will
systematically undercount comorbidity burden — sometimes
substantially.
plumage() solves this by flagging a condition as present
if any of up to three independent code types match: an
ICD-10-AM supplementary U-code, a matching acute principal or secondary
ICD-10-AM admission code, or (optionally) an AR-DRG code. These are
combined with OR logic — one match is enough.
The name fits: just as experienced birders read a bird’s plumage as a
proxy for its underlying physiological condition, plumage()
reads a patient record’s clinical coding as a proxy for chronic
condition burden.
plumage() detects 29 chronic conditions
across 9 body-system categories.
| Category | Condition | Column name |
|---|---|---|
| Metabolic/Endocrine | Obesity | obesity |
| Metabolic/Endocrine | Cystic fibrosis† | cystic_fibrosis |
| Mental Health | Dementia | dementia |
| Mental Health | Schizophrenia | schizophrenia |
| Mental Health | Depression | depression |
| Mental Health | Intellectual/developmental disability | intellectual_dev |
| Neurological | Parkinson’s disease | parkinsons |
| Neurological | Multiple sclerosis | multiple_sclerosis |
| Neurological | Epilepsy | epilepsy |
| Neurological | Cerebral palsy | cerebral_palsy |
| Neurological | Paralysis | paralysis |
| Cardiovascular | Ischaemic heart disease | ihd |
| Cardiovascular | Heart failure | heart_failure |
| Cardiovascular | Hypertension | hypertension |
| Respiratory | Emphysema | emphysema |
| Respiratory | COPD | copd |
| Respiratory | Asthma/chronic bronchitis | asthma |
| Respiratory | Bronchiectasis | bronchiectasis |
| Respiratory | Chronic respiratory failure | respiratory_failure |
| Respiratory | Cystic fibrosis† | cystic_fibrosis |
| Gastrointestinal | Crohn’s disease | crohns |
| Gastrointestinal | Ulcerative colitis | ulcerative_colitis |
| Gastrointestinal | Liver failure | liver_failure |
| Musculoskeletal | Rheumatoid arthritis | rheumatoid_arthritis |
| Musculoskeletal | Osteoarthritis | osteoarthritis |
| Musculoskeletal | SLE | lupus |
| Musculoskeletal | Osteoporosis | osteoporosis |
| Renal | Chronic kidney disease | kidney_disease |
| Congenital | Spina bifida | spina_bifida |
| Congenital | Down syndrome | downs |
hospital_data <- data.frame(
patient_id = 1:5,
icd_codes = c(
"K29.70", # gastritis only — no chronic comorbidities
"U78.1, U83.2, U82.3", # obesity + COPD + hypertension (all U-codes)
"J44.1, U79.3", # COPD via acute ICD + depression via U-code
"J43.2, J47", # emphysema + bronchiectasis (acute ICD, no U-codes)
"E84.0, U80.3" # cystic fibrosis (acute ICD) + epilepsy (U-code)
)
)
results <- plumage(hospital_data, "icd_codes")
# View key columns
results[, c("patient_id", "copd", "emphysema", "bronchiectasis",
"cystic_fibrosis", "total_conditions", "conditions_category")]
#> patient_id copd emphysema bronchiectasis cystic_fibrosis total_conditions
#> 1 1 0 0 0 0 0
#> 2 2 1 0 0 0 3
#> 3 3 1 0 0 0 2
#> 4 4 0 1 1 0 2
#> 5 5 0 0 0 1 2
#> conditions_category
#> 1 0
#> 2 3+
#> 3 2
#> 4 2
#> 5 2
Notice that row 3 has copd = 1 detected purely from the
acute code J44.1 (no U83.2 present), and row 4
has both emphysema and bronchiectasis detected from acute codes alone.
This is exactly the dual-code advantage.
conditions_category summaryEvery run of plumage() produces a
conditions_category ordered factor — a coarse summary of
comorbidity burden useful for stratified analyses and tables.
table(results$conditions_category)
#>
#> 0 1 2 3+
#> 1 0 3 1
The ordering (0 < 1 < 2 < 3+) is preserved in
gtsummary::tbl_summary() and ggplot2 without
any extra setup.
Some Queensland datasets store ICD codes without decimal points
(e.g. U832 instead of U83.2). Set
decimal = FALSE to match this format.
df_nodot <- data.frame(
icd = c("U832 U823", "J441 J431"),
stringsAsFactors = FALSE
)
plumage(df_nodot, "icd", decimal = FALSE)[, c("copd", "hypertension", "emphysema")]
#> copd hypertension emphysema
#> 1 1 1 0
#> 2 1 0 1
When DRG codes are available in a separate column,
include_drg = TRUE adds them as a third detection pathway —
particularly useful for COPD, asthma, and bronchiectasis where DRGs are
well-specified.
df_drg <- data.frame(
patient_id = 1:3,
icd_codes = c("K29.70", "J44.1", "K29.70"), # row 3: no respiratory ICD
drg_codes = c("G07B", "E65A", "E65A") # row 3: COPD DRG only
)
# Without DRG: row 3 missed entirely
plumage(df_drg, "icd_codes", include_drg = FALSE)[, c("patient_id", "copd")]
#> patient_id copd
#> 1 1 0
#> 2 2 1
#> 3 3 0
# With DRG: row 3 caught via E65A
plumage(df_drg, "icd_codes", include_drg = TRUE,
drg_column = "drg_codes")[, c("patient_id", "copd")]
#> patient_id copd
#> 1 1 0
#> 2 2 1
#> 3 3 1
When combining plumage() output with other flag columns,
use prefix to avoid name collisions.
res_prefixed <- plumage(hospital_data, "icd_codes", prefix = "chr_")
names(res_prefixed)[grepl("^chr_", names(res_prefixed))] |> head(8)
#> [1] "chr_obesity" "chr_cystic_fibrosis" "chr_dementia"
#> [4] "chr_schizophrenia" "chr_depression" "chr_intellectual_dev"
#> [7] "chr_parkinsons" "chr_multiple_sclerosis"
drop_eggsFor downstream modelling where you only need summary counts,
drop_eggs = TRUE removes the 29 individual binary columns
and retains only the 11 summary columns — a substantial reduction in
width for large datasets.
res_lean <- plumage(hospital_data, "icd_codes", drop_eggs = TRUE)
names(res_lean)
#> [1] "patient_id" "icd_codes"
#> [3] "total_conditions" "total_metabolic_conditions"
#> [5] "total_mental_health_conditions" "total_neurological_conditions"
#> [7] "total_cardiovascular_conditions" "total_respiratory_conditions"
#> [9] "total_gastrointestinal_conditions" "total_musculoskeletal_conditions"
#> [11] "total_renal_conditions" "total_congenital_conditions"
#> [13] "conditions_category"
plumage(): what comes nextThe output of plumage() integrates naturally with the
rest of the mudnester pipeline:
# Typical hospitalisation workflow
df_hosp <- clean_the_nest(hosp_raw, data_type = "hospital", ...)
df_hosp <- plumage(df_hosp, icd_column = "icd_code")
df_hosp <- preening(df_hosp, age_col = "age", scheme = "geriatric_fine")
# Stratify comorbidity burden by age group before aggregation
roost(df_hosp, date_col = "admission_date", time_unit = "month",
group_cols = c("age_group", "conditions_category"))
Before sharing or archiving the enriched dataset, pass it through
molting() (see vignette("molting")). The
conditions_category column is retained by default — it
matches the age\d+cat preservation pattern — but check that
total_conditions and the individual binary columns are
appropriately handled for your sharing context.
U78–U88) does not exist in ICD-10.plumage() operates row-wise. If your dataset has multiple
rows per patient (one per admission), a patient will be flagged for a
condition in any row where the relevant code appears. Aggregate across
admissions first
(e.g. group_by(patient_id) |> summarise(copd = max(copd)))
if you want one row per patient.See vignette("mudnester-getting-started") for the full
pipeline context.