| Title: | Publication-Ready Tables for Descriptive Statistics and Regression Models |
| Version: | 0.13.0 |
| Description: | Provides publication-ready tables for descriptive statistics and regression models: frequency tables and cross-tabulations with association measures (Cramer's V, Kendall's Tau-b, and others), categorical and continuous summary tables, by group or from a complex survey design, and regression tables for one or more models side by side, across more than thirty model classes from mixed-effects to survival and Bayesian, with robust standard errors, average marginal effects, and univariable screening. Tables follow APA conventions by default, can switch to named journal styles such as JAMA, NEJM, or The Lancet, and render identically in the console and in 'gt', 'tinytable', 'flextable', 'Word', 'Excel', or the clipboard. Declared missing values in labelled data are honored and disclosed throughout the descriptive tables. Helpers cover codebooks, variable inspection, and row-wise summaries. |
| License: | MIT + file LICENSE |
| URL: | https://amaltawfik.github.io/spicy/, https://github.com/amaltawfik/spicy |
| BugReports: | https://github.com/amaltawfik/spicy/issues |
| Encoding: | UTF-8 |
| Language: | en-US |
| Imports: | crayon, dplyr, labelled, rlang (≥ 1.1.0), sandwich (≥ 3.1-2), stats, stringr, tibble, tidyselect, utils |
| Suggests: | bit64, boot, broom, car, clipr, clubSandwich, DT, effectsize, emmeans, estimatr (≥ 1.0.0), fixest (≥ 0.11), flexsurv (≥ 2.2), flextable, geepack (≥ 1.3.9), gt, haven, htmltools, insight, glmmTMB (≥ 1.1.7), knitr, lme4 (≥ 1.1-35), lmerTest (≥ 3.1-3), lmtest, merDeriv (≥ 0.2.6), marginaleffects, modelsummary, collapse, MASS, mgcv (≥ 1.9), nlme (≥ 3.1-160), nnet, numDeriv, ordinal (≥ 2022.11), AER (≥ 1.2), betareg (≥ 3.1), mlogit (≥ 1.1), pscl (≥ 1.5), quantreg (≥ 5.86), RLRsim, reformulas, rms (≥ 6.7), sampleSelection (≥ 1.2), posterior (≥ 1.5.0), rstanarm (≥ 2.21), brms (≥ 2.20), loo (≥ 2.6), rstan, survey (≥ 4.5), survival (≥ 3.5), officer, openxlsx2, parameters, performance, rmarkdown, spelling, testthat (≥ 3.0.0), tinytable, withr |
| VignetteBuilder: | knitr |
| Depends: | R (≥ 4.1.0) |
| Config/testthat/edition: | 3 |
| LazyData: | true |
| Config/roxygen2/version: | 8.1.0 |
| NeedsCompilation: | no |
| Packaged: | 2026-09-30 23:11:30 UTC; at |
| Author: | Amal Tawfik |
| Maintainer: | Amal Tawfik <amal.tawfik@hesav.ch> |
| Repository: | CRAN |
| Date/Publication: | 2026-10-01 08:10:02 UTC |
spicy: publication-ready tables for descriptive statistics and regression models
Description
spicy provides publication-ready tables for descriptive statistics and regression models: frequency tables and cross-tabulations with association measures, categorical and continuous summary tables (by group or from a complex survey design), and regression tables for one or more models side by side, across more than thirty model classes, with robust standard errors, average marginal effects, and univariable screening. Tables follow APA conventions by default, can switch to named journal styles, and render identically in the console and the rich outputs. Declared missing values in labelled data are honored and disclosed throughout the descriptive tables; helpers cover codebooks, variable inspection, and row-wise summaries.
API stability
spicy is in active pre-1.0 development. Breaking changes are
made deliberately at minor-version bumps and are always
announced in NEWS.md. The API surface is partitioned as
follows; users planning to embed spicy in production pipelines
or downstream packages should rely on the stable surface.
Stable (signature and behaviour preserved across 0.y.z and into 1.0.0; documented changes only):
Frequency / cross-tabs:
freq(),cross_tab()Variable inspection:
varlist()/vl(),code_book(),label_from_names()Clipboard export:
copy_clipboard()Association measures (point estimates and documented CIs):
cramer_v(),phi(),contingency_coef(),yule_q(),gamma_gk(),kendall_tau_b(),kendall_tau_c(),somers_d(),lambda_gk(),goodman_kruskal_tau(),uncertainty_coef()
Stabilising (still maturing; argument names may be tightened
before 1.0 with a NEWS.md entry, but no silent behavioural
changes):
Summary table builders:
table_categorical(),table_continuous(),table_continuous_lm(),table_outcome()Survey-design summary tables:
table_continuous_svy(),table_categorical_svy()Regression tables:
table_regression(),table_regression_uv(),table_regression_models(),as_structured()Inline citation:
inline()Table styles:
spicy_style(),spicy_style_names()Omnibus association overview:
assoc_measures()
Experimental (new in this cycle; the shape of the output and
the argument names may still move, with a NEWS.md entry, on their
OWN clock rather than the parent family's):
Display language and label overrides:
spicy_labels()(withoptions(spicy.language)/options(spicy.labels))
Internal API (not part of the public surface; can change without notice – avoid calling directly from downstream code):
ASCII rendering primitive:
spicy_print_table()(build_ascii_table()is no longer exported)
broom output shape
The broom::tidy() and broom::glance() methods on
spicy_categorical_table, spicy_continuous_table,
spicy_continuous_lm_table, spicy_continuous_svy_table,
spicy_categorical_svy_table, and spicy_regression_table follow
the standard broom column conventions (outcome, term,
estimate, std.error, conf.low, conf.high, statistic,
p.value, df, df.residual, r.squared, adj.r.squared,
nobs, ...). The set of columns produced by each method is
considered stabilising: existing columns will not be silently
renamed or have their semantics changed within 0.y.z, and any
breaking change is announced in NEWS.md. Adding optional new
columns (e.g. covariate-adjustment metadata) is not a breaking
change. Numeric columns keep the types downstream
broom-consumers expect: test degrees of freedom that are integer
by construction (chi-squared, factor-comparison F tests) stay
integer, while every degrees-of-freedom column that can be
fractional – df.residual, the regression method's
per-coefficient df, and Welch-corrected test df – is
numeric double, so Satterthwaite-corrected degrees of freedom
from cluster-robust variance modes are preserved verbatim
(matching lmerTest::glance() and the afex output
convention).
Classed conditions
All errors and warnings emitted by the stable / stabilising
surfaces carry classed conditions so downstream code can
dispatch on class via tryCatch() / withCallingHandlers()
instead of matching message strings. Each condition has a
package-wide parent class plus a leaf class describing the
specific cause:
spicy_errorCatch-all parent for every error raised by spicy. Leaves:
-
spicy_invalid_input– bad argument value or type. -
spicy_invalid_data– bad data shape or content (not a data.frame, length mismatch,bit64::integer64columns, degenerate grouping). -
spicy_missing_pkg– a Suggests dependency is required by the requested operation but not installed. -
spicy_missing_column– a referenced column is not indata. -
spicy_unsupported– the operation is not applicable to this input (e.g., Phi requested on a non-2x2 table). -
spicy_ame_satt_unsupported_formula– signaled together withspicy_unsupportedwhen AME Satterthwaite degrees of freedom are unavailable for the model's formula structure; normally caught internally and surfaced as aspicy_fallbackwarning. -
spicy_unsupported_class– the model class has no regression-frame method, sotable_regression()cannot render it. -
spicy_unsupported_vcov– the requestedvcovmode is not available for this model class. -
spicy_unsupported_standardized– the requestedstandardizedmode is not available for this model class. -
spicy_invalid_frame– an object failed the structural validation of the internal regression-frame contract behindtable_regression(). -
spicy_resampling_failed– bootstrap / jackknife resampling produced too few valid replicates to estimate the requested statistic. -
spicy_defunct– an argument removed in a pre-1.0 hard break; the message names the replacement. Signaled together withspicy_invalid_inputso generic input handlers still catch it. -
spicy_internal– an internal precondition failed; this is a bug in spicy, please report it. -
spicy_internal_invariant– an internal consistency check on a spicy-built object failed and the result cannot be trusted (see the warning leaf of the same name for the renderable case).
-
spicy_warningCatch-all parent for every warning. Leaves:
-
spicy_undefined_stat– the requested statistic is undefined for this input; result isNA(e.g., Tau-b on a table with all-zero marginals). -
spicy_negative_weights_no_test– signaled together withspicy_undefined_statwhentable_continuous_svy()ortable_categorical_svy()withholds a group comparison because the analytic sample carries negatively weighted rows; the estimates are still reported and the table note says so. -
spicy_dropped_na–NAobservations were silently excluded from the computation (e.g.,NAweights). -
spicy_ignored_arg– an argument was ignored due to context (e.g.,correct = TRUEon a non-2x2 table). -
spicy_no_selection– a column selector produced an empty set; an empty result is returned rather than erroring. -
spicy_fallback– the requested computation failed; a simpler estimator was used instead. -
spicy_caveat– the computation succeeded but its interpretation carries a non-trivial methodological caveat (e.g., standardized coefficients on non-additive terms). -
spicy_bayes_diagnostics– signaled together withspicy_caveatwhen a Bayesian fit's sampler or predictive-accuracy diagnostics miss their targets (R-hat, ESS, divergences, E-BFMI, Pareto k, p_waic). -
spicy_nonconvergence– signaled together withspicy_caveatwhen the fitting engine reports that its optimizer did not converge. The table still shows the numbers the model object holds, and a footer note says what they are worth. -
spicy_model_choice– a defaulted modeling choice was made on the user's behalf and is disclosed (e.g., the linear probability modeltable_regression_uv()fits to a binary outcome under the defaultmethod). -
spicy_passthrough– a third-party warning captured during an operation (e.g., the clipboard copy) and re-emitted under the spicy taxonomy. -
spicy_summary_failed–varlist()could not summarise one column; the rest of the table is fine. -
spicy_renamed_column– a user data column or factor level collided with a spicy-internal name and was auto-renamed to preserve the data (emitted bycross_tab()). -
spicy_internal_invariant– an internal consistency check on a spicy-built object failed but the output still renders, so the user sees both the table and the diagnostic.
-
spicy_infoParent for informational messages (emitted via
rlang::inform(); muffle withwithCallingHandlers(spicy_info = ...)). Leaf:spicy_silent_reference– reference levels are displayed nowhere underreference_style = "none"withfactor_layout = "flat". The once-per-session hint on ordered-factor polynomial contrasts carries its own class,spicy_polynomial_contrasts_info.
Author(s)
Maintainer: Amal Tawfik amal.tawfik@hesav.ch (ORCID) (ROR) [copyright holder]
Authors:
Amal Tawfik amal.tawfik@hesav.ch (ORCID) (ROR) [copyright holder]
See Also
Useful links:
Report bugs at https://github.com/amaltawfik/spicy/issues
Coerce a spicy_categorical_svy_table to a data frame or tibble
Description
These S3 methods strip the "spicy_categorical_svy_table" class and
the rendering-only attributes, keeping the wide compute frame and the
three provenance markers (group_var, design_meta, note).
Usage
## S3 method for class 'spicy_categorical_svy_table'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)
## S3 method for class 'spicy_categorical_svy_table'
as_tibble(x, ...)
Arguments
x |
A |
row.names, optional, ... |
Passed on for method compatibility; ignored. |
Value
A plain data.frame (or a tbl_df for as_tibble()).
See Also
Coerce a spicy_categorical_table to a plain data frame or tibble
Description
These S3 methods strip the "spicy_categorical_table" /
"spicy_table" classes and the rendering-only attributes
(display_df, indent_text, align, decimal_mark,
long_data, ...) from an object returned by table_categorical()
so the underlying wide-format data can be manipulated with
downstream tools (dplyr, tidyr, etc.) under the standard
data.frame / tbl_df contract. The single attribute
"group_var" is preserved as a lightweight provenance marker; all
other spicy attributes are dropped. The original x is unaffected,
and print(x) continues to render the formatted ASCII table.
Usage
## S3 method for class 'spicy_categorical_table'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)
## S3 method for class 'spicy_categorical_table'
as_tibble(x, ...)
Arguments
x |
A |
row.names, optional |
Standard |
... |
Further arguments passed to |
Details
The returned data is the wide raw representation (one row per
(variable x level), group columns side by side). For the
tidy long format – one row per (variable x level x group) –
use tidy.spicy_categorical_table() or call table_categorical()
directly with output = "long".
Value
A plain data.frame (or tbl_df) with the same rows and
columns as the wide raw output of table_categorical().
See Also
tidy.spicy_categorical_table(),
glance.spicy_categorical_table().
Coerce a spicy_continuous_lm_table to a plain data frame or tibble
Description
These S3 methods strip the "spicy_continuous_lm_table" /
"spicy_table" classes and the rendering-only attributes
(digits, decimal_mark, ci_level, ...) from an object returned
by table_continuous_lm() so the underlying long-format data can
be manipulated with downstream tools (dplyr, tidyr, etc.) under
the standard data.frame / tbl_df contract. The single attribute
"by_var" is preserved as a lightweight provenance marker; all
other spicy attributes are dropped. The original x is unaffected,
and print(x) continues to render the formatted ASCII table.
Usage
## S3 method for class 'spicy_continuous_lm_table'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)
## S3 method for class 'spicy_continuous_lm_table'
as_tibble(x, ...)
Arguments
x |
A |
row.names, optional |
Standard |
... |
Further arguments passed to |
Value
A plain data.frame (or tbl_df) with the same rows and
columns as the long output of table_continuous_lm().
See Also
tidy.spicy_continuous_lm_table(),
glance.spicy_continuous_lm_table() for cleaner broom-style
pivots tailored to downstream pipelines.
Coerce a spicy_continuous_svy_table to a plain data frame or tibble
Description
These S3 methods strip the "spicy_continuous_svy_table" class and
the rendering-only attributes, keeping the long compute frame and the
three provenance markers (group_var, design_meta, note).
Usage
## S3 method for class 'spicy_continuous_svy_table'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)
## S3 method for class 'spicy_continuous_svy_table'
as_tibble(x, ...)
Arguments
x |
A |
row.names, optional, ... |
Passed on for method compatibility; ignored. |
Value
A plain data.frame (or a tbl_df for as_tibble()).
See Also
Coerce a spicy_continuous_table to a plain data frame or tibble
Description
These S3 methods strip the "spicy_continuous_table" /
"spicy_table" classes and the rendering-only attributes
(digits, decimal_mark, ci_level, align, p_digits, ...)
from an object returned by table_continuous() so the underlying
long-format data can be manipulated with downstream tools (dplyr,
tidyr, etc.) under the standard data.frame / tbl_df contract.
The single attribute "group_var" is preserved as a lightweight
provenance marker; all other spicy attributes are dropped. The
original x is unaffected, and print(x) continues to render the
formatted ASCII table.
Usage
## S3 method for class 'spicy_continuous_table'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)
## S3 method for class 'spicy_continuous_table'
as_tibble(x, ...)
Arguments
x |
A |
row.names, optional |
Standard |
... |
Further arguments passed to |
Details
The returned data is identical to what output = "long" (or
output = "data.frame") returns directly from table_continuous();
use whichever entry point reads better in your pipeline.
Value
A plain data.frame (or tbl_df) with one row per
(variable x group) (or one row per variable when by is not
used).
See Also
tidy.spicy_continuous_table(),
glance.spicy_continuous_table() for cleaner broom-style pivots
tailored to downstream pipelines.
Coerce a spicy_outcome_table to a plain data frame or tibble
Description
These S3 methods strip the "spicy_outcome_table" class and the
rendering-only attributes from an object returned by
table_outcome(), so the underlying long-format data can be
manipulated with downstream tools under the standard data.frame /
tbl_df contract. The "outcome" and "select" attributes are kept as
lightweight provenance markers. The original x is unaffected, and
print(x) continues to render the formatted table.
Usage
## S3 method for class 'spicy_outcome_table'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)
## S3 method for class 'spicy_outcome_table'
as_tibble(x, ...)
Arguments
x |
A |
row.names, optional |
Standard |
... |
Further arguments passed to |
Details
The returned data is identical to what output = "long" (or
output = "data.frame") returns directly from table_outcome().
Value
A plain data.frame (or tbl_df), one row per displayed
row of the table.
See Also
tidy.spicy_outcome_table(), glance.spicy_outcome_table().
Convert a spicy_regression_table to a plain data.frame / tibble
Description
Strips the spicy_regression_table / spicy_table classes and
the internal analytic attributes (spicy_long, spicy_fit_stats),
returning the wide character display as a plain data.frame (or
tbl_df via as_tibble()). The title, note, provenance
(model_ids, outcome), and rendering (col_spec) attributes are
preserved.
Usage
## S3 method for class 'spicy_regression_table'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)
## S3 method for class 'spicy_regression_table'
as_tibble(x, ...)
Arguments
x |
A |
row.names, optional |
Standard |
... |
Currently ignored. |
Details
as.data.frame() is equivalent to passing output = "data.frame"
to table_regression(): the two paths return identical objects
(same cells, classes, and attributes).
Value
A plain data.frame (for as.data.frame()) or a
tbl_df (for as_tibble()).
See Also
tidy.spicy_regression_table(),
glance.spicy_regression_table() for broom-canonical long
views.
Convert a spicy flextable output to a plain flextable
Description
table_regression() and table_continuous_lm() return their
output = "flextable" tables with a lightweight spicy_flextable
class tag whose only job is HTML note styling. Every flextable
verb already works on the tagged object; this method returns the
clean underlying flextable::flextable() – the note is already
part of its footer – for workflows that want the untagged object
(e.g. custom knit hooks, flextable::save_as_docx() pipelines, or
composition with other flextable tooling), mirroring
gtsummary::as_flex_table().
Usage
as_flextable.spicy_flextable(x, ...)
Arguments
x |
A |
... |
Unused; for generic consistency. |
Details
Rendering in Quarto / R Markdown does NOT require this conversion:
the knit_print method auto-detects the output format and
delegates to flextable's native rendering for non-HTML targets
(Word, PowerPoint, PDF).
Value
A flextable object.
Extract the typed (structured) view of a spicy table
Description
spicy's tables return a display representation by default – a
character data.frame with stars suffixes, en-dash for reference
rows, bracketed "[L, U]" confidence intervals, and APA padding on
p-values. This accessor returns the typed view that the output
engines (Excel, gt, tinytable, flextable, clipboard) consume
internally: a fully numeric body with CI pre-split into LL / UL
columns, NAs for non-applicable / reference cells, plus per-cell
markers and a format specification.
Usage
as_structured(x)
Arguments
x |
A spicy table built with |
Details
This is the right entry point for users who want to:
-
Filter coefficients programmatically, e.g.
as_structured(tbl)$body[which(as_structured(tbl)$body$p < 0.05), ](which()drops the structurally empty rows – factor headers, reference levels – whosepisNA; a bare logical index would keep them as all-NArows). -
Aggregate raw values across rows, e.g.
mean(as_structured(tbl)$body[["B"]], na.rm = TRUE). -
Build a custom downstream renderer that consumes the same structured contract as spicy's built-in engines.
Value
A list with the structured view (see Details for the schema).
Schema
-
version– integer contract version (see Versioning below). -
body–data.framewith aVariablecharacter column, one or more numeric value columns, and four dot-prefixed identity columns at the end. Confidence intervals are split intoLL/ULcolumns named like"95% CI: LL"/"95% CI: UL"(or prefixed with the model label in multi-model output). A cell with no value (reference level, term absent from a model, factor header) isNA. -
body$.variable– source variable of the row: the factor name on a factor header / level / reference row, the term otherwise, and the fit-statistic token ("nobs","r2","fixed_effects", ...) on a fit-statistic row. -
body$.level– the factor level (or the absorbed factor of a fixed-effects block, or the grouping factor of an"N (...)"row);NAoutside a factor. -
body$.row_role– what the row is:"coef","factor_header","level","reference","fit_stat","outcome","vc"(variance component) in a regression table, plus"summary"(a row summarising one variable),"group"(a row keyed by one level ofby) and"missing"(a row keyed by the missing value) in the descriptive ones. The role is the key a consumer matches on:"(Missing)"is a display label – auto-renamed on collision, translatable – and the role survives both. -
body$.indent– display indent depth of the label (0or1). Renderers indent.indent > 0rows; the label text itself is already indented in the character body only. -
cell_status– per-cell semantics, keyed by column name, one character vector as long asbodyper column that needs one:"reference"(reference level, in this estimate block and this model),"undefined"(the statistic applies to the row but no number expresses it – an unavailable variance-component standard error, a fit statistic undefined for one model's class),""otherwise. Both marked statuses display as an en-dash. A cell whose value isNAwith no status is absent and displays blank. Columns with no marked cell are omitted. -
outcome_labels_by_col– for the outcome row (explicitoutcome_labelswith two or more models), the display label keyed by each model's first structured column name. -
col_meta– per-column metadata keyed by structured column name (token, model_id, precision, p-style, below-threshold, CI pair / role / label). A column whose cells cannot be reconstructed from one number – the"events/N"counts ofshow_columns = "n_events"– also carriesdisplay_cells, a character vector as long asbodyholding the display string of each cell (NAwhere the number formats normally). A renderer must prefer it over the numeric value. -
stars–NULLunlessstarswas requested, otherwise a list withthresholds(symbol to p cutoff) andmarkers(per-cell marker strings, keyed by column name,""where a cell takes none). -
spanners– named list mapping a grouping label to its column indices inbody: the model labels of a multi-model regression table, thebylevels of a categorical one (each spanning itsn/%pair). -
ci_pairs– list of(label, cols)entries describing each CI pair inbody. -
format_spec– global format defaults (decimal mark, digits, p-style, CI level, etc.).
Descriptive tables
The three descriptive families return the same schema; only the
col_meta tokens and the row roles they emit differ.
-
table_categorical()– one"factor_header"row per variable then one"level"(or"missing") row per category, exactly as the console lays them out. Tokens:"n","pct","p","assoc"(the association measure),"assoc_ci"(its bounds) and"smd"(the standardized mean difference, whensmd = TRUE). In abytable each group owns ann/%pair with its own spanner, and the margin column is flaggedcol_meta$<col>$total, never matched on the label"Total". -
table_continuous()– one"summary"row per variable, or one"group"row per level ofby("missing"for the missing-bygroup), with the level in.level. Tokens are theshow_columnsvocabulary itself ("m","sd","med","iqr","med_iqr","q1","q3","min","max","ci","med_ci","n","weighted_n") plus"statistic","p","es"and"smd"for the group comparison – the last four arecol_metatokens without beingshow_columnsvalues, because they belong to a comparison BETWEEN groups rather than to any one group. A statistic another variable displays is absent (NA, no status); one that applies but has no value is"undefined". -
table_continuous_lm()– one"summary"row per outcome. Tokens:"emmean"(a marginal mean; the column says which level incol_meta$<col>$level),"delta"(the contrast),"b"(a numeric predictor's slope),"ci","statistic","p","r2","es","n","weighted_n".
Cells the console builds from more than one number – a
"Med [Q1, Q3]", a test gloss, an effect size with its interval –
keep their numeric anchor in body and their exact printed string
in col_meta$<col>$display_cells. stars is always NULL:
descriptive tables carry no significance markers.
Versioning
version says which contract an object carries. Version 3 moved
row identity out of index vectors and into the body itself:
-
reference_rowsandreference_models_by_rowbecomecell_status, which marks the reference cell instead of the whole row – the row-scoped flag was blanking estimate blocks that have no per-level reference at all. -
factor_header_rowsbecomes.row_role == "factor_header". -
fit_stat_rowsbecomes.row_role == "fit_stat". -
level_rowsbecomes.indent > 0. -
outcome_rowbecomes.row_role == "outcome".
Index vectors are the structure that corrupts as soon as two bodies are stacked or merged, so they were removed rather than kept alongside. An object carrying an older contract (or none) is refused with the correspondence above; rebuild it with the function that produced it. An object carrying a newer contract than the spicy reading it is refused as well.
Version 3 also opened the accessor to the descriptive families,
which had no typed view before: .row_role gained "summary",
"group" and "missing". The vocabulary is extended by addition
– an existing role never changes meaning.
See Also
table_regression(), table_categorical(),
table_continuous(), table_continuous_lm() for the
user-facing entry points.
Examples
fit <- lm(mpg ~ wt + factor(cyl), data = mtcars)
tbl <- table_regression(fit)
s <- as_structured(tbl)
s$body # raw numeric body
s$body[which(s$body$p < 0.05), ] # filter significant rows
# which() drops the structural NA rows (headers, reference levels)
s$body$.row_role # what each row is
s$body[s$body$.variable == "wt", ] # address a row by its variable
s$col_meta$B # column metadata for B
# The same schema on a descriptive table.
ct <- table_categorical(mtcars, c(cyl, gear), by = am)
sc <- as_structured(ct)
sc$body[, c("Variable", ".variable", ".level", ".row_role")]
sc$spanners # one per `by` group
Association measures summary table
Description
assoc_measures() computes a range of association measures for a
two-way contingency table and returns them in a tidy data frame.
Usage
assoc_measures(
x,
type = c("all", "nominal", "ordinal"),
conf_level = 0.95,
digits = 3L
)
Arguments
x |
A contingency table (of class |
type |
Which family of measures to compute:
|
conf_level |
A single number strictly between 0 and 1 giving
the confidence level (default |
digits |
Number of decimal places used when printing the
result (default |
Details
type = "all" (the default) returns all nominal and ordinal
measures. Use type = "nominal" or type = "ordinal" to
restrict the output to a single family.
The nominal family includes cramer_v(), contingency_coef(),
lambda_gk(), goodman_kruskal_tau(), uncertainty_coef(),
and (for 2x2 tables) phi() and yule_q().
The ordinal family includes gamma_gk(), kendall_tau_b(),
kendall_tau_c(), and somers_d().
Measures that are undefined on the given table appear as NA
rows (printed as an en dash). The classed warnings the individual
functions raise (e.g. spicy_undefined_stat) are re-emitted
once per distinct message after the table is assembled, so
condition handlers and suppressWarnings() behave as they do
for the individual functions.
Standard error formulas follow the DescTools implementations
(Signorell et al., 2024), except for Kendall's Tau-b, whose
ASE follows Brown and Benedetti (1977) as printed by SPSS /
PSPP CROSSTABS; see kendall_tau_b().
Value
A data frame with columns measure, estimate, se,
ci_lower, ci_upper, and p_value. The p_value comes
from two test families:
-
Pearson chi-squared test of independence for Cramer's V, Phi, and the Contingency Coefficient (the three chi-squared-derived nominal measures). All three carry the same chi-squared p-value on a given table.
-
Wald z-test of H0: measure = 0 for every other measure: Yule's Q, Lambda, Goodman-Kruskal's Tau, the Uncertainty Coefficient, and all ordinal measures (Gamma, Tau-b, Tau-c, Somers' D).
Direction-dependent measures (lambda_gk(),
goodman_kruskal_tau(), uncertainty_coef(), somers_d())
contribute one row per direction (symmetric / R|C / C|R
where applicable), so the output has more rows than the
number of helper functions.
References
Agresti, A. (2002). Categorical Data Analysis (2nd ed.). Wiley.
Brown, M. B., & Benedetti, J. K. (1977). Sampling behavior of tests for correlation in two-way contingency tables. Journal of the American Statistical Association, 72(358), 309-315. doi:10.1080/01621459.1977.10480995
Liebetrau, A. M. (1983). Measures of Association. Sage.
Signorell, A. et al. (2024). DescTools: Tools for Descriptive Statistics. R package.
See Also
cramer_v(), gamma_gk(), kendall_tau_b()
Other association measures:
contingency_coef(),
cramer_v(),
gamma_gk(),
goodman_kruskal_tau(),
kendall_tau_b(),
kendall_tau_c(),
lambda_gk(),
phi(),
somers_d(),
uncertainty_coef(),
yule_q()
Examples
tab <- table(sochealth$smoking, sochealth$education)
assoc_measures(tab)
assoc_measures(tab, type = "nominal")
assoc_measures(tab, type = "ordinal")
Build a formatted ASCII table
Description
Low-level rendering engine that constructs a visually aligned
ASCII table from a data.frame. Supports Unicode line-drawing
characters, ANSI colors via crayon, automatic
colored-text-aware width detection, configurable padding, and
per-column alignment.
Usage
build_ascii_table(
x,
padding = 2L,
first_column_line = TRUE,
row_total_line = TRUE,
bottom_line = FALSE,
lines_color = "darkgrey",
align_left_cols = c(1L, 2L),
align_center_cols = integer(0),
center_headers = FALSE,
spanners = NULL,
group_sep_rows = integer(0),
total_row_idx = NULL,
display_labels = NULL,
...
)
Arguments
x |
A |
padding |
Non-negative integer giving the number of extra
characters added to each column's auto-computed width (the
maximum of the cell-content width and the header width).
Defaults to The string choices |
first_column_line |
Logical. If |
row_total_line |
Logical. Controls the horizontal rule drawn
before a total row. Defaults to |
bottom_line |
Logical. If |
lines_color |
Character. Color used for table separators. Defaults to |
align_left_cols |
Integer vector of column indices to
left-align. Defaults to |
align_center_cols |
Integer vector of column indices to
center-align. Defaults to |
center_headers |
Logical. When |
spanners |
Optional named list defining a column group row
drawn above the column headers (the "spanner" / "supra-header"
convention; APA Manual 7 §7.13).
Names are spanner labels; values are integer vectors of 1-based
column indices the label spans (must be contiguous). A thin
underline rule is drawn below each spanner across its span.
Used by |
group_sep_rows |
Integer vector of row indices before which a
light dashed separator line is drawn. Defaults to |
total_row_idx |
Optional integer vector of 1-based row indices
identifying the totals rows; a horizontal rule is drawn just
before each. When |
display_labels |
Optional character vector of length |
... |
Additional arguments (currently ignored). |
Details
Most users should not call this directly: it is wrapped by
spicy_print_table() and the internal print.spicy_* methods,
which add titles, notes, and table-type-aware alignment defaults.
Reach for build_ascii_table() only when you need to render an
arbitrary data.frame to a string with the same look as
spicy's tables.
Value
A single character string containing the full ASCII-formatted table,
suitable for direct printing with cat().
See Also
spicy_print_table() for a user-facing wrapper that adds titles and notes.
Generate an interactive variable codebook
Description
code_book() creates an interactive and exportable codebook summarizing
selected variables of a data frame. It builds upon varlist() to provide
an overview of variable names, labels, classes, and representative values in
a sortable, searchable table.
The output is displayed as an interactive DT::datatable() in the Viewer pane
(for example in RStudio or Positron), allowing searching, sorting, and export
(copy, print, CSV, Excel, PDF) directly.
Usage
code_book(
x,
...,
values = FALSE,
include_na = FALSE,
title = "Codebook",
filename = NULL,
factor_levels = c("all", "observed"),
user_na = TRUE
)
Arguments
x |
A data frame or tibble. |
... |
Optional tidyselect-style column selectors (e.g.
|
values |
Logical. If |
include_na |
Logical. If |
title |
Optional character string displayed as the table caption.
Defaults to |
filename |
Optional character string used as the base for exported CSV,
Excel, and PDF filenames. If |
factor_levels |
Character. Controls how factor values are displayed
in |
user_na |
Logical. If |
Details
The interactive
datatablesupports column sorting, global searching, and client-side export (copy, print, CSV, Excel, PDF) directly from the Viewer.Variable selection uses the same tidyselect interface as
varlist(); the underlying summary tibble is built byvarlist()withtbl = TRUE.
Value
A DT::datatable object.
Dependencies
Requires the following package:
-
DT
Declared missing values
Survey files imported with haven often carry declared missing
values: codes such as 8 = Don't know or 9 = Refused that the
source file marks as missing while keeping them distinct from a
plain NA. Two kinds of declaration exist: na_values / na_range
metadata on haven::labelled_spss() vectors, and tagged missing
values created by haven::tagged_na() (the Stata .a, .b, ...
convention).
spicy honors the declaration by default (user_na = TRUE):
declared missing values are excluded from every statistic exactly
like NA – valid percentages, means, chi-squared tests,
association measures, row-wise summaries, and group definitions –
but they are not erased from display. freq() lists each observed
declared value as its own row of the Missing block, with its value
label; cross_tab(), table_categorical(), and
table_continuous() disclose the exclusion in the table note
(Declared missing values removed: x (2).); varlist() and
code_book() count them as missing in N_valid / NAs /
N_distinct while still listing the declared codes in Values.
Every function involved offers the same escape hatch: set
user_na = FALSE to ignore the declaration and treat the declared
codes as valid values (the behavior of spicy before 0.13.0).
Tagged missing values are genuine NAs either way; for them,
user_na = FALSE only collapses the per-tag breakdown back into
the regular NA count.
See Also
varlist() for generating the underlying variable summaries.
Other variable inspection:
label_from_names(),
varlist()
Examples
## Not run:
if (requireNamespace("DT", quietly = TRUE)) {
code_book(sochealth)
code_book(sochealth, starts_with("bmi"))
code_book(sochealth, starts_with("bmi"), values = TRUE, include_na = TRUE)
factors <- data.frame(
group = factor(c("A", "B", NA), levels = c("A", "B", "C"))
)
code_book(
factors,
values = TRUE,
include_na = TRUE,
factor_levels = "observed"
)
code_book(
sochealth,
starts_with("bmi"),
title = "BMI codebook",
filename = "bmi_codebook"
)
}
## End(Not run)
Pearson's contingency coefficient
Description
contingency_coef() computes Pearson's contingency coefficient C
for a two-way contingency table.
Usage
contingency_coef(x, detail = FALSE, conf_level = 0.95, digits = 3L)
Arguments
x |
A contingency table (of class |
detail |
Logical. If |
conf_level |
A single number strictly between 0 and 1 giving
the confidence level (default |
digits |
Number of decimal places used when printing the
result (default |
Details
The contingency coefficient is
C = \sqrt{\chi^2 / (\chi^2 + n)}.
It ranges from 0 (independence) to a maximum that depends on
the table dimensions. No standard asymptotic standard error exists,
so the confidence interval is not computed.
Value
Same structure as cramer_v(): a scalar when
detail = FALSE, a named vector when detail = TRUE.
The p-value tests the null hypothesis of no association
(Pearson chi-squared test). CI values are NA because no
standard asymptotic SE exists for C.
References
Liebetrau, A. M. (1983). Measures of Association. Sage.
See Also
Other association measures:
assoc_measures(),
cramer_v(),
gamma_gk(),
goodman_kruskal_tau(),
kendall_tau_b(),
kendall_tau_c(),
lambda_gk(),
phi(),
somers_d(),
uncertainty_coef(),
yule_q()
Examples
tab <- table(sochealth$smoking, sochealth$education)
contingency_coef(tab)
Copy data to the clipboard
Description
Copies a data.frame, matrix, 2D or higher array, table, or
atomic vector to the system clipboard, ready to paste into a
text editor, spreadsheet, or word processor. Wraps
clipr::write_clip() (a Suggests dependency); requires clipr
to be installed and a clipboard backend to be available on the
platform.
Usage
copy_clipboard(
x,
row_names_as_col = FALSE,
row_names = TRUE,
col_names = TRUE,
show_message = TRUE,
quiet = FALSE,
...
)
Arguments
x |
A |
row_names_as_col |
Logical or character. If |
row_names |
Logical. If |
col_names |
Logical. If |
show_message |
Logical. If |
quiet |
Logical. If |
... |
Additional arguments passed to |
Details
Objects that are not data.frames or 2D matrices (atomic
vectors, arrays, tables) are automatically coerced to character
on the way to the clipboard, as required by
clipr::write_clip(). The caller's object is never modified in
place; transformations happen on a local copy (see the return
value).
Multidimensional arrays (3D and higher) are flattened to a 1D
character vector with one element per line. To preserve a
tabular layout, extract a 2D slice first, e.g.
copy_clipboard(my_array[, , 1]).
Messages and warnings raised by the clipboard backend are
re-emitted as regular R conditions, so suppressMessages() /
suppressWarnings() work as usual; quiet = TRUE silences them
all at once.
Value
Invisibly returns the object as it was sent to the
clipboard: identical to x by default, but reflecting the
row_names_as_col transformation when one was requested (e.g.
a matrix comes back as a data.frame with the promoted
row-name column). The function is called for its clipboard
side effect.
Examples
# Writes to the system clipboard, so never run by checks:
## Not run:
if (clipr::clipr_available()) {
# Data frame
copy_clipboard(sochealth)
# Data frame with row names as column
copy_clipboard(head(sochealth), row_names_as_col = "id")
# Matrix
mat <- matrix(1:6, nrow = 2)
copy_clipboard(mat)
# Table
tbl <- table(sochealth$education)
copy_clipboard(tbl)
# Array (3D) -- flattened to character
arr <- array(1:8, dim = c(2, 2, 2))
copy_clipboard(arr)
# Recommended: copy 2D slice for tabular layout
copy_clipboard(arr[, , 1])
# Numeric vector
copy_clipboard(c(3.14, 2.71, 1.618))
# Character vector
copy_clipboard(c("apple", "banana", "cherry"))
# Quiet mode (no messages shown)
copy_clipboard(sochealth, quiet = TRUE)
}
## End(Not run)
Row-wise count of specific or special values
Description
Counts, for each row of a data.frame or matrix, how many
times one or more values appear across selected columns. Supports
type-safe comparison (allow_coercion = FALSE), case-insensitive
string matching (ignore_case = TRUE), and detection of special
values (NA, NaN, Inf, -Inf) via special. Designed to
flow inside dplyr::mutate() pipelines.
Usage
count_n(
data = NULL,
select = tidyselect::everything(),
exclude = NULL,
count = NULL,
special = NULL,
allow_coercion = TRUE,
ignore_case = FALSE,
regex = FALSE,
verbose = FALSE,
user_na = TRUE
)
Arguments
data |
A |
select |
Columns to include. Defaults to |
exclude |
Columns to exclude after selection (names or positions,
as accepted by |
count |
Value(s) to count. Defaults to |
special |
Character vector of special values to count: |
allow_coercion |
Logical. If |
ignore_case |
Logical. If |
regex |
Logical. If |
verbose |
Logical. If |
user_na |
Logical. If |
Value
A numeric vector of row-wise counts (unnamed), of length
nrow(data). Missing values never match a regular count value,
so an all-NA row counts 0 unless special targets missing
values. If the selection resolves to zero usable columns, a
classed warning (spicy_no_selection) is emitted and NA is
returned for all rows, as in mean_n() and sum_n().
Strict matching (allow_coercion = FALSE)
Comparison falls back to identical() when types differ, which
also inspects factor levels. Two consequences:
-
count = "b"does not match a factor"b"value: pass a factor, e.g.count = factor("b", levels = levels(df$x)). Even with a factor
count, comparisons against columns whose level set differs will return0. To guarantee a perfect match (label and levels), reuse a value taken from the data itself (e.g.df$x[2]).
Case-insensitive matching (ignore_case = TRUE)
All values are converted to lowercase via tolower() before
matching; factor columns are first coerced to character. This
mode takes precedence over allow_coercion: equality becomes
lowercase string equality, so "b" and "B" match even when
allow_coercion = FALSE.
Coercion of count itself
R coerces mixed-type vectors at construction time: count = c(2, "2") becomes c("2", "2") before the function ever sees it.
To get type-sensitive matching, keep count homogeneous.
Declared missing values
Survey files imported with haven often carry declared missing
values: codes such as 8 = Don't know or 9 = Refused that the
source file marks as missing while keeping them distinct from a
plain NA. Two kinds of declaration exist: na_values / na_range
metadata on haven::labelled_spss() vectors, and tagged missing
values created by haven::tagged_na() (the Stata .a, .b, ...
convention).
spicy honors the declaration by default (user_na = TRUE):
declared missing values are excluded from every statistic exactly
like NA – valid percentages, means, chi-squared tests,
association measures, row-wise summaries, and group definitions –
but they are not erased from display. freq() lists each observed
declared value as its own row of the Missing block, with its value
label; cross_tab(), table_categorical(), and
table_continuous() disclose the exclusion in the table note
(Declared missing values removed: x (2).); varlist() and
code_book() count them as missing in N_valid / NAs /
N_distinct while still listing the declared codes in Values.
Every function involved offers the same escape hatch: set
user_na = FALSE to ignore the declaration and treat the declared
codes as valid values (the behavior of spicy before 0.13.0).
Tagged missing values are genuine NAs either way; for them,
user_na = FALSE only collapses the per-tag breakdown back into
the regular NA count.
See Also
datawizard::row_count() for a closely related row-wise
counter; count_n() adds element-wise type-safe matching,
multi-value count, and special-value detection.
Other row-wise summaries:
mean_n(),
sum_n()
Examples
library(dplyr)
library(tibble)
library(labelled)
# Basic usage
df <- tibble(
x = c(1, 2, 2, 3, NA),
y = c(2, 2, NA, 3, 2),
z = c("2", "2", "2", "3", "2")
)
count_n(df, count = 2)
count_n(df, count = 2, allow_coercion = FALSE)
df |> mutate(num_twos = count_n(count = 2))
# Mixed types and special values
df <- tibble(
num = c(1, 2, NA, -Inf, NaN),
char = c("a", "B", "b", "a", NA),
fact = factor(c("a", "b", "b", "a", "c")),
date = as.Date(c("2023-01-01", "2023-01-01", NA, "2023-01-02", "2023-01-01")),
lab = labelled(c(1, 2, 1, 2, NA), labels = c(No = 1, Yes = 2)),
logic = c(TRUE, FALSE, NA, TRUE, FALSE)
)
count_n(df, count = 2)
count_n(df, count = "b", ignore_case = TRUE)
count_n(df, count = "a", select = fact)
count_n(df, count = as.Date("2023-01-01"), select = date)
# Count special values
count_n(df, special = "NA")
# Column selection strategies
df <- tibble(
score_math = c(1, 2, 2, 3, NA),
score_science = c(2, 2, NA, 3, 2),
score_lang = c("2", "2", "2", "3", "2"),
name = c("Jean", "Marie", "Ali", "Zoe", "Nina")
)
count_n(df, select = c(score_math, score_science), count = 2)
count_n(df, select = starts_with("score_"), exclude = "score_lang", count = 2)
count_n(df, select = "^score_", regex = TRUE, count = 2)
df |> mutate(nb_two = count_n(count = 2))
# Strict type-safe matching with factor columns
df <- tibble(
x = factor(c("a", "b", "c")),
y = factor(c("b", "B", "a"))
)
# Coercion: character "b" matches both x and y
count_n(df, count = "b")
# Strict match: fails because "b" is character, not factor (returns only 0s)
count_n(df, count = "b", allow_coercion = FALSE)
# Strict match with factor value: works only where levels match
count_n(df, count = factor("b", levels = levels(df$x)), allow_coercion = FALSE)
Cramer's V
Description
cramer_v() computes Cramer's V for a two-way contingency table,
measuring the strength of association between two categorical variables.
Usage
cramer_v(x, detail = FALSE, conf_level = 0.95, digits = 3L)
Arguments
x |
A contingency table (of class |
detail |
Logical. If |
conf_level |
A single number strictly between 0 and 1 giving
the confidence level (default |
digits |
Number of decimal places used when printing the
result (default |
Details
Cramer's V is computed as
V = \sqrt{\chi^2 / (n \cdot (k - 1))}, where \chi^2
is the Pearson chi-squared statistic, n is the total count,
and k = \min(r, c). The point estimate matches the
DescTools (Signorell et al., 2024) and SPSS implementations.
The confidence interval uses the Fisher z-transformation
on V (\tanh(\mathrm{atanh}(V) \pm z_{\alpha/2} /
\sqrt{n - 3})), which differs from the noncentral chi-squared
or bootstrap CIs reported by DescTools::CramerV().
Value
When detail = FALSE: a single numeric value (the
estimate).
When detail = TRUE and conf_level is non-NULL:
c(estimate, se, ci_lower, ci_upper, p_value).
When detail = TRUE and conf_level = NULL:
c(estimate, se, p_value).
The p-value tests the null hypothesis of no association
(Pearson chi-squared test).
References
Agresti, A. (2002). Categorical Data Analysis (2nd ed.). Wiley.
Brown, M. B., & Benedetti, J. K. (1977). Sampling behavior of tests for correlation in two-way contingency tables. Journal of the American Statistical Association, 72(358), 309-315. doi:10.1080/01621459.1977.10480995
Liebetrau, A. M. (1983). Measures of Association. Sage.
Signorell, A. et al. (2024). DescTools: Tools for Descriptive Statistics. R package.
See Also
phi(), contingency_coef(), assoc_measures()
Other association measures:
assoc_measures(),
contingency_coef(),
gamma_gk(),
goodman_kruskal_tau(),
kendall_tau_b(),
kendall_tau_c(),
lambda_gk(),
phi(),
somers_d(),
uncertainty_coef(),
yule_q()
Examples
tab <- table(sochealth$smoking, sochealth$education)
cramer_v(tab)
cramer_v(tab, detail = TRUE)
cramer_v(tab, detail = TRUE, conf_level = NULL)
Cross-tabulation
Description
Computes a two-way cross-tabulation with optional weights, grouping
(including combinations of multiple variables via interaction()),
row / column percentages, and inferential statistics (Chi-squared
test with an APA-style association measure).
Both x and y are required; for one-way frequency tables, use
freq().
Usage
cross_tab(
data,
x,
y = NULL,
by = NULL,
weights = NULL,
rescale = FALSE,
percent = c("none", "column", "row"),
include_stats = TRUE,
assoc_measure = c("auto", "cramer_v", "phi", "gamma", "tau_b", "tau_c", "somers_d",
"lambda", "none"),
assoc_ci = FALSE,
correct = FALSE,
simulate_p = FALSE,
simulate_B = 2000,
digits = NULL,
output = c("default", "data.frame"),
show_n = TRUE,
decimal_mark = ".",
p_digits = 3L,
user_na = TRUE,
styled
)
Arguments
data |
A data frame. Alternatively, a vector when using the vector-based interface. |
x |
Row variable (unquoted). |
y |
Column variable (unquoted). Required; the |
by |
Optional grouping variable or expression. Can be a single variable
or a combination of multiple variables (e.g. |
weights |
Optional numeric weights. A logical vector is also accepted and coerced to 1/0 (include / exclude). |
rescale |
Logical. If |
percent |
One of |
include_stats |
Logical. If |
assoc_measure |
Character. Which association measure to report.
|
assoc_ci |
Logical. If |
correct |
Logical. If |
simulate_p |
Logical. If |
simulate_B |
Integer. Number of replicates for Monte Carlo simulation.
Defaults to |
digits |
Number of decimals for cell values: a single
non-negative integer. Defaults to |
output |
Output format. |
show_n |
Logical. If |
decimal_mark |
Character used as the decimal mark in printed
numeric values (cells, chi-squared, association estimate, CI
bounds, p-value, table note). Either |
p_digits |
Integer number of decimals used to format the
p-value (and to determine the small- |
user_na |
Logical. If |
styled |
Defunct. |
Value
Depends on output and by:
-
output = "default", noby: aspicy_cross_tableobject (adata.framecarrying rendering metadata as attributes:title,digits,decimal_mark,n_row_idx,n_col_name, and the inferential block wheninclude_stats = TRUE). Printing dispatches toprint.spicy_cross_table(). -
output = "default",bysupplied: aspicy_cross_table_list, i.e. a named list ofspicy_cross_tableobjects (one element per group level, named by that level). Printing dispatches toprint.spicy_cross_table_list()which renders each table in turn separated by a blank line. -
output = "data.frame": the same payload returned as a plaindata.frame(or named list ofdata.frames withby), stripped of thespicy_*classes and of every metadata attribute (title,note,n_total,chi2,p_value,assoc_*, ...). For programmatic access to the statistics, read the attributes of the default object, e.g.attr(cross_tab(...), "p_value").
Cell columns are the levels of y; rows are the levels of x.
When percent != "none", the N column (or N row) is added
according to show_n. When include_stats = TRUE, the result
carries a Chi-squared row (statistic, df, p) and an
association-measure row (estimate, optional CI via assoc_ci).
Global Options
The function recognizes the following global options that modify its default behavior:
-
options(spicy.percent = "column")Sets the default percentage mode for all calls tocross_tab(). Valid values are"none","row", and"column". Equivalent to settingpercent = "column"(or another choice) in each call. -
options(spicy.simulate_p = TRUE)Enables Monte Carlo simulation for all Chi-squared tests by default. Equivalent to settingsimulate_p = TRUEin every call. -
options(spicy.rescale = TRUE)Automatically rescales weights so that total weighted N equals the raw N. Equivalent to settingrescale = TRUEin each call. Also read byfreq(), so one option governs both tabulators.
These options are convenient for users who wish to enforce consistent
behavior across multiple calls: spicy.percent and spicy.simulate_p
apply to cross_tab(), and spicy.rescale applies to both
cross_tab() and freq().
They can be disabled or reset by setting them to NULL:
options(spicy.percent = NULL, spicy.simulate_p = NULL, spicy.rescale = NULL).
Example:
options(spicy.simulate_p = TRUE, spicy.rescale = TRUE) cross_tab(sochealth, smoking, education, weights = weight)
Declared missing values
Survey files imported with haven often carry declared missing
values: codes such as 8 = Don't know or 9 = Refused that the
source file marks as missing while keeping them distinct from a
plain NA. Two kinds of declaration exist: na_values / na_range
metadata on haven::labelled_spss() vectors, and tagged missing
values created by haven::tagged_na() (the Stata .a, .b, ...
convention).
spicy honors the declaration by default (user_na = TRUE):
declared missing values are excluded from every statistic exactly
like NA – valid percentages, means, chi-squared tests,
association measures, row-wise summaries, and group definitions –
but they are not erased from display. freq() lists each observed
declared value as its own row of the Missing block, with its value
label; cross_tab(), table_categorical(), and
table_continuous() disclose the exclusion in the table note
(Declared missing values removed: x (2).); varlist() and
code_book() count them as missing in N_valid / NAs /
N_distinct while still listing the declared codes in Values.
Every function involved offers the same escape hatch: set
user_na = FALSE to ignore the declaration and treat the declared
codes as valid values (the behavior of spicy before 0.13.0).
Tagged missing values are genuine NAs either way; for them,
user_na = FALSE only collapses the per-tag breakdown back into
the regular NA count.
Examples
# Basic crosstab
cross_tab(sochealth, smoking, education)
# Column percentages
cross_tab(sochealth, smoking, education, percent = "column")
# Weighted (rescaled)
cross_tab(sochealth, smoking, education, weights = weight, rescale = TRUE)
# Grouped by sex
cross_tab(sochealth, smoking, education, by = sex)
# Grouped by combination of variables
cross_tab(sochealth, smoking, education, by = interaction(sex, age_group))
# Ordinal variables: auto-selects Kendall's Tau-b
cross_tab(sochealth, education, self_rated_health)
# 2x2 table with Yates correction
cross_tab(sochealth, smoking, physical_activity, correct = TRUE)
# APA-style p-value precision and European decimal mark
cross_tab(sochealth, smoking, education, decimal_mark = ",", p_digits = 4)
Frequency Table
Description
Creates a frequency table for a vector or variable from a data frame, with options for weighting, sorting, handling labelled data, defining custom missing values, and displaying cumulative percentages.
With output = "default", the function returns a
spicy_freq_table object that auto-prints as a spicy-formatted ASCII
table via print.spicy_freq_table() and spicy_print_table(); with
output = "data.frame", it returns a plain data.frame containing
frequencies and proportions.
Usage
freq(
data,
x = NULL,
weights = NULL,
digits = 1L,
valid = TRUE,
cum = FALSE,
sort = "",
na_val = NULL,
labelled_levels = c("prefixed", "labels", "values"),
factor_levels = c("observed", "all"),
rescale = FALSE,
decimal_mark = ".",
output = c("default", "data.frame"),
user_na = TRUE,
styled
)
Arguments
data |
A |
x |
A variable from |
weights |
Optional numeric vector of weights (same length as |
digits |
Number of decimal digits to display for percentages
(default: |
valid |
Logical. If |
cum |
Logical. If |
sort |
Sorting method for values:
For labelled variables displayed with their codes
( |
na_val |
Atomic vector of numeric or character values to be treated as missing ( For labelled variables (from haven or labelled), this argument must refer to the underlying coded values, not the visible labels. Example: x <- labelled(c(1, 2, 3, 1, 2, 3), c("Low" = 1, "Medium" = 2, "High" = 3))
freq(x, na_val = 1) # Treat all "Low" as missing
|
labelled_levels |
For
|
factor_levels |
Character. Controls how factor and labelled values
are displayed in the frequency table. |
rescale |
Logical. If |
decimal_mark |
Character used as the decimal mark in printed
percentages. Either |
output |
Output format. |
user_na |
Logical. If |
styled |
Defunct. |
Details
Designed to mimic common frequency procedures from SPSS or Stata
while integrating the flexibility of R's data structures. The
input type (vector, factor, labelled) is auto-detected; see
@param labelled_levels and @param factor_levels for the
schema-vs-observed level controls, and @param na_val for
optional sentinel-value recoding.
Weighting (weights): frequencies and percentages are computed
proportionally to the weights. Missing values in weights cause
those observations to be dropped from the table entirely (with a
warning), matching the behaviour of cross_tab() in spicy
0.11.0+. With rescale = TRUE, the remaining (non-NA-weighted)
weights are normalised so the total weighted N equals the count
of non-NA-weighted rows. With rescale = FALSE, the total
weighted N is the actual sum of non-NA weights.
For schema-level inspection without computing frequencies, use
varlist() or code_book().
Value
With output = "data.frame", a plain data.frame with no extra
attributes and columns:
-
value- unique values or factor levels -
n- frequency count (weighted if applicable) -
prop- proportion of total -
valid_prop- proportion of valid responses (ifvalid = TRUE) -
cum_prop,cum_valid_prop- cumulative percentages (ifcum = TRUE)
With output = "default" (the default), a spicy_freq_table object: the same
data.frame carrying rendering metadata as attributes (digits,
data_name, var_name, var_label, class_name, n_total,
n_valid, weighted, rescaled, weight_var) used by
print.spicy_freq_table(). The object is returned visibly, so a
bare freq(...) call auto-prints at the console while
f <- freq(...) stays silent (print f to display the table).
Declared missing values
Survey files imported with haven often carry declared missing
values: codes such as 8 = Don't know or 9 = Refused that the
source file marks as missing while keeping them distinct from a
plain NA. Two kinds of declaration exist: na_values / na_range
metadata on haven::labelled_spss() vectors, and tagged missing
values created by haven::tagged_na() (the Stata .a, .b, ...
convention).
spicy honors the declaration by default (user_na = TRUE):
declared missing values are excluded from every statistic exactly
like NA – valid percentages, means, chi-squared tests,
association measures, row-wise summaries, and group definitions –
but they are not erased from display. freq() lists each observed
declared value as its own row of the Missing block, with its value
label; cross_tab(), table_categorical(), and
table_continuous() disclose the exclusion in the table note
(Declared missing values removed: x (2).); varlist() and
code_book() count them as missing in N_valid / NAs /
N_distinct while still listing the declared codes in Values.
Every function involved offers the same escape hatch: set
user_na = FALSE to ignore the declaration and treat the declared
codes as valid values (the behavior of spicy before 0.13.0).
Tagged missing values are genuine NAs either way; for them,
user_na = FALSE only collapses the per-tag breakdown back into
the regular NA count.
See Also
cross_tab() for two-way cross-tabulations;
table_categorical() for multi-variable categorical summary
tables; varlist() / code_book() for variable inspection;
print.spicy_freq_table() for formatted printing;
spicy_print_table() for the underlying ASCII rendering engine.
Examples
# Frequency table with labelled ordered factor
freq(sochealth, education)
freq(sochealth, self_rated_health, sort = "-")
library(labelled)
# Simple numeric vector
x <- c(1, 2, 2, 3, 3, 3, NA)
freq(x)
# Plain vector with a sentinel value recoded as missing
freq(c(1, 2, 3, 99, 99), na_val = 99)
# Labelled variable (haven-style)
x_lbl <- labelled(
c(1, 2, 3, 1, 2, 3, 1, 2, NA),
labels = c("Low" = 1, "Medium" = 2, "High" = 3)
)
var_label(x_lbl) <- "Satisfaction level"
# Treat value 1 ("Low") as missing
freq(x_lbl, na_val = 1)
# Display only labels, add cumulative %
freq(x_lbl, labelled_levels = "labels", cum = TRUE)
# Display values only, sorted descending
freq(x_lbl, labelled_levels = "values", sort = "-")
# Show all declared factor levels, including unused ones (n = 0).
# The default "observed" mirrors Stata's `tab` and SPSS FREQUENCIES,
# which both drop unused levels.
f <- factor(c("Yes", "No", "Yes"), levels = c("Yes", "No", "Maybe"))
freq(f, factor_levels = "all")
# With weighting
df <- data.frame(
sex = factor(c("Male", "Female", "Female", "Male", NA, "Female")),
weight = c(12, 8, 10, 15, 7, 9)
)
# Weighted frequencies (raw weighted counts, the default)
freq(df, sex, weights = weight)
# Weighted frequencies rescaled so the total matches the sample size
freq(df, sex, weights = weight, rescale = TRUE)
# Base R style, with weights and cumulative percentages
freq(df$sex, weights = df$weight, cum = TRUE)
# Piped version (tidy syntax) and sort alphabetically descending ("name-")
df |> freq(sex, sort = "name-")
# European decimal mark (matches `cross_tab()` and the `table_*()` family)
freq(sochealth, education, decimal_mark = ",")
# Plain data.frame return (for programmatic use)
f <- freq(df, sex, output = "data.frame")
head(f)
Goodman-Kruskal Gamma
Description
gamma_gk() computes the Goodman-Kruskal Gamma statistic for a
two-way contingency table of ordinal variables.
Usage
gamma_gk(x, detail = FALSE, conf_level = 0.95, digits = 3L)
Arguments
x |
A contingency table (of class |
detail |
Logical. If |
conf_level |
A single number strictly between 0 and 1 giving
the confidence level (default |
digits |
Number of decimal places used when printing the
result (default |
Details
Gamma is computed as \gamma = (C - D) / (C + D), where
C and D are the numbers of concordant and
discordant pairs. It ignores tied pairs, making it appropriate
for ordinal variables with many ties.
When the asymptotic standard error is zero (e.g. a perfect
association), the Wald z-test is undefined and the p-value is
NA, matching the other measures in the family.
Standard error formulas follow the DescTools implementations
(Signorell et al., 2024); see cramer_v() for full references.
Value
Same structure as cramer_v(): a scalar when
detail = FALSE, a named vector when detail = TRUE.
The p-value tests H0: gamma = 0 (Wald z-test).
References
Goodman, L. A., & Kruskal, W. H. (1954). Measures of association for cross classifications. Journal of the American Statistical Association, 49(268), 732-764. doi:10.2307/2281536
Brown, M. B., & Benedetti, J. K. (1977). Sampling behavior of tests for correlation in two-way contingency tables. Journal of the American Statistical Association, 72(358), 309-315. doi:10.1080/01621459.1977.10480995
See Also
kendall_tau_b(), kendall_tau_c(), somers_d(),
assoc_measures()
Other association measures:
assoc_measures(),
contingency_coef(),
cramer_v(),
goodman_kruskal_tau(),
kendall_tau_b(),
kendall_tau_c(),
lambda_gk(),
phi(),
somers_d(),
uncertainty_coef(),
yule_q()
Examples
tab <- table(sochealth$education, sochealth$self_rated_health)
gamma_gk(tab)
gamma_gk(tab, detail = TRUE)
Goodman-Kruskal's Tau
Description
goodman_kruskal_tau() computes Goodman-Kruskal's Tau, a
proportional reduction in error (PRE) measure for nominal
variables.
Usage
goodman_kruskal_tau(
x,
direction = c("row", "column"),
detail = FALSE,
conf_level = 0.95,
digits = 3L
)
Arguments
x |
A contingency table (of class |
direction |
Direction of prediction:
|
detail |
Logical. If |
conf_level |
A single number strictly between 0 and 1 giving
the confidence level (default |
digits |
Number of decimal places used when printing the
result (default |
Details
Unlike lambda_gk(), Goodman-Kruskal's Tau uses all cell
frequencies rather than only the modal categories, making it
more sensitive to association patterns where lambda may be
zero. Goodman-Kruskal's Tau is intrinsically directional and
has no canonical symmetric form (unlike lambda_gk() or
uncertainty_coef()); only "row" and "column" are
supported.
Standard error formulas follow the DescTools implementations
(Signorell et al., 2024); see cramer_v() for full references.
Value
Same structure as cramer_v(): a scalar when
detail = FALSE, a named vector when detail = TRUE.
The p-value tests H0: tau = 0 (Wald z-test).
References
Goodman, L. A., & Kruskal, W. H. (1954). Measures of association for cross classifications. Journal of the American Statistical Association, 49(268), 732-764. doi:10.2307/2281536
See Also
lambda_gk(), uncertainty_coef(), assoc_measures()
Other association measures:
assoc_measures(),
contingency_coef(),
cramer_v(),
gamma_gk(),
kendall_tau_b(),
kendall_tau_c(),
lambda_gk(),
phi(),
somers_d(),
uncertainty_coef(),
yule_q()
Examples
tab <- table(sochealth$smoking, sochealth$education)
goodman_kruskal_tau(tab)
goodman_kruskal_tau(tab, direction = "column", detail = TRUE)
Cite a table cell in inline text
Description
Returns one cell of a spicy table as a character scalar, formatted exactly as the table displays it – same decimals, same p style, same interval punctuation, same journal style. Designed for inline R chunks in Quarto / R Markdown:
Smokers had higher odds (`r inline(tbl, smoking, "Yes", "or")`).
Usage
inline(x, variable, level = NULL, column = NULL, model = NULL)
Arguments
x |
A table returned by |
variable |
The source variable, unquoted or as a string; or a
fit-statistic token ( |
level |
For a factor variable, the level, as a string.
|
column |
A column token, or a |
model |
In a multi-model table, the model: its label (as displayed in the column spanners) or its position. |
Value
A character scalar.
Addressing
The row is found by identity, not by display text: variable
names the source column (.variable in the typed body), level
the level (.level). Custom labels, a style, or a translated
display never change the call. As a convenience, a variable
that matches no source column is looked up among the displayed
labels before erroring. The missing-value category is addressed by
level = "(Missing)" whatever its displayed (possibly deduplicated)
label, through its row role. Fit statistics are addressed by their
token as variable (inline(tbl, "n"), inline(tbl, "r2")).
A statistic that belongs to a whole variable rather than to one of
its levels sits on the variable's own row: the p of a
table_categorical() block, its association measure, its SMD. Leave
level out to cite it (inline(tbl, smoking, column = "p")).
table_continuous_lm() lays its groups out sideways – one row per
outcome, the by levels as columns (M (Female), M (Male)) –
so there level names the group whose column you want:
inline(tbl, bmi, "Female", "emmean"). The columns that belong to
no group (the contrast, its interval, p, n) are cited without
one, as before.
The column is a token of the typed contract ("b", "se",
"p", "ci", "or", "ame", "n", "pct", "m", ... – see
as_structured()'s col_meta), never a display header. "ci"
composes the interval with the style's brackets and separator, and
so does every other interval token the table carries ("med_ci",
"ame_ci", "assoc_ci"): each addresses its own bounds. In
a multi-model table, model selects the model by its spanner
label or position; in a by table, the spanners are the groups,
so model selects the group the same way.
Patterns
A column containing { is a pattern: each {token} is replaced
by the corresponding cell, so one call quotes a full sentence
fragment:
inline(tbl, smoking, "Yes", "{or} ({ci_label} {ci}; p = {p})")
{ci_label} inserts the interval label of the interval the pattern
cites (95% CI, or Med 95% CI in a pattern quoting {med_ci}) –
the first one when it cites several, the table's first when it cites
none. Note that {p} carries the floor operator when the table does
(<.001), so write p {p} rather than p = {p} in patterns that
may hit the floor.
Errors
Every misaddressing is a classed error that lists the available choices: unknown variables list the variables, missing levels list the levels, unknown tokens list the table's tokens, ambiguous models list the spanner labels. A cell the table itself displays as undefined (an aliased coefficient's en-dash) refuses with the reason rather than pasting a dash into a sentence.
See Also
as_structured() for the typed contract behind the
addressing.
Examples
fit <- lm(wellbeing_score ~ age + sex, data = sochealth)
tbl <- table_regression(fit)
inline(tbl, age, column = "b")
inline(tbl, sex, "Male", "{b} ({ci_label} {ci}; p {p})")
inline(tbl, "n")
Kendall's Tau-b
Description
kendall_tau_b() computes Kendall's Tau-b for a two-way
contingency table of ordinal variables.
Usage
kendall_tau_b(x, detail = FALSE, conf_level = 0.95, digits = 3L)
Arguments
x |
A contingency table (of class |
detail |
Logical. If |
conf_level |
A single number strictly between 0 and 1 giving
the confidence level (default |
digits |
Number of decimal places used when printing the
result (default |
Details
Kendall's Tau-b is computed as
\tau_b = (C - D) / \sqrt{(n_0 - n_1)(n_0 - n_2)},
where n_0 = n(n-1)/2, n_1 is the number of
pairs tied on the row variable, and n_2 is the number
tied on the column variable. Tau-b corrects for ties and is
appropriate for square tables.
When the asymptotic standard error is zero (e.g. a perfect
association), the Wald z-test is undefined and the p-value is
NA, matching the other measures in the family.
The asymptotic standard error is the Brown and Benedetti (1977)
ASE1, as printed by SPSS / PSPP CROSSTABS. It deliberately
diverges from DescTools::KendallTauB(), whose implementation
mis-scales one margin term of the gradient; see cramer_v()
for full references.
Value
Same structure as cramer_v(): a scalar when
detail = FALSE, a named vector when detail = TRUE.
The p-value tests H0: tau-b = 0 (Wald z-test).
References
Kendall, M. G. (1938). A new measure of rank correlation. Biometrika, 30(1-2), 81-93. doi:10.2307/2332226
Brown, M. B., & Benedetti, J. K. (1977). Sampling behavior of tests for correlation in two-way contingency tables. Journal of the American Statistical Association, 72(358), 309-315. doi:10.1080/01621459.1977.10480995
See Also
kendall_tau_c(), gamma_gk(), somers_d(),
assoc_measures()
Other association measures:
assoc_measures(),
contingency_coef(),
cramer_v(),
gamma_gk(),
goodman_kruskal_tau(),
kendall_tau_c(),
lambda_gk(),
phi(),
somers_d(),
uncertainty_coef(),
yule_q()
Examples
tab <- table(sochealth$education, sochealth$self_rated_health)
kendall_tau_b(tab)
Kendall's Tau-c (Stuart's Tau-c)
Description
kendall_tau_c() computes Stuart's Tau-c (also known as
Kendall's Tau-c) for a two-way contingency table of ordinal
variables.
Usage
kendall_tau_c(x, detail = FALSE, conf_level = 0.95, digits = 3L)
Arguments
x |
A contingency table (of class |
detail |
Logical. If |
conf_level |
A single number strictly between 0 and 1 giving
the confidence level (default |
digits |
Number of decimal places used when printing the
result (default |
Details
Stuart's Tau-c is computed as
\tau_c = 2m(C - D) / (n^2(m - 1)), where
m = \min(r, c). It is designed for rectangular tables;
the estimate is bounded by [-1, 1] only when the table is
square, and may fall outside that range otherwise.
When the asymptotic standard error is zero (e.g. a perfect
association), the Wald z-test is undefined and the p-value is
NA, matching the other measures in the family.
When one variable is constant (all observations in a single row
or column), there are no untied pairs and the statistic
degenerates to a meaningless 0: the function returns NA with a
spicy_undefined_stat warning, like its siblings, matching the
SPSS / PSPP behavior of reporting no value.
Standard error formulas follow the DescTools implementations
(Signorell et al., 2024); see cramer_v() for full references.
Value
Same structure as cramer_v(): a scalar when
detail = FALSE, a named vector when detail = TRUE.
The p-value tests H0: tau-c = 0 (Wald z-test).
References
Stuart, A. (1953). The estimation and comparison of strengths of association in contingency tables. Biometrika, 40(1-2), 105-110. doi:10.2307/2333101
Brown, M. B., & Benedetti, J. K. (1977). Sampling behavior of tests for correlation in two-way contingency tables. Journal of the American Statistical Association, 72(358), 309-315. doi:10.1080/01621459.1977.10480995
See Also
kendall_tau_b(), gamma_gk(), somers_d(),
assoc_measures()
Other association measures:
assoc_measures(),
contingency_coef(),
cramer_v(),
gamma_gk(),
goodman_kruskal_tau(),
kendall_tau_b(),
lambda_gk(),
phi(),
somers_d(),
uncertainty_coef(),
yule_q()
Examples
tab <- table(sochealth$education, sochealth$self_rated_health)
kendall_tau_c(tab)
Derive variable labels from column names name<sep>label
Description
Splits each column name at the first occurrence of sep,
renames the column to the part before sep (the name, trimmed
of surrounding whitespace), and assigns the part after sep as a
"label" attribute on the column. The label attribute follows
the haven convention also used
by labelled::var_label(), so labelled-aware tooling
(labelled, haven, varlist(), code_book(), ...) reads it
transparently. Splitting at the first sep means the label
itself may contain the separator.
Usage
label_from_names(df, sep = ". ")
Arguments
df |
A |
sep |
Character string used as separator between name and
label. Default |
Details
Designed primarily for LimeSurvey CSV exports with Headings:
Question code & question text, which produce column names like
"code. question text". The default separator ". " matches
that export.
LimeSurvey question codes (the part before sep) are
restricted to alphanumerics, must start with a letter, and
contain no spaces – so the column name has to carry both the
code and the question text. If your export uses Headings:
Question code (codes only), re-export with Question code &
question text before calling this function; there is no way to
recover a label from a code alone.
Whitespace handling: the name (left of sep) is trimmed
of surrounding whitespace, because R column names are intended
to be referenced bare (without backticks) and leading / trailing
whitespace would force quoting throughout the user's downstream
code. The label (right of sep) is preserved verbatim,
following the Stata / SPSS convention that variable labels are
faithful user content – spicy does not silently mutate label
strings. To trim labels yourself, post-process with
labelled::var_label(df) <- lapply(labelled::var_label(df), trimws).
Value
An object of the same class as df – a base
data.frame if df was a base data.frame, a tbl_df if df
was a tibble. The output has column names equal to the trimmed
names (before sep) and, for every column whose original name
contained sep, a "label" attribute equal to the label (after
sep). Columns whose name does not contain sep are passed
through unchanged with no label attached.
Errors
The function raises an actionable error – rather than letting the downstream constructor raise a cryptic one – when the split produces:
duplicate column names that the renaming itself creates (two original names share the same prefix before
sep, or a new name collides with an existing one). Names that were already duplicated in the input (check.names = FALSEdata) are passed through untouched, not blamed on the split; oran empty column name (the original name starts with
sepand has nothing before it).
See Also
labelled::var_label() reads the "label" attribute set
by this function; varlist() and code_book() surface it in
their inspection outputs.
Other variable inspection:
code_book(),
varlist()
Examples
# LimeSurvey-style column names (default sep = ". ").
df <- data.frame(
"age. Age of respondent" = c(25, 30),
"score. Total score. Manually computed." = c(12, 14),
check.names = FALSE
)
out <- label_from_names(df)
attr(out$age, "label")
attr(out$score, "label")
# Custom separator.
df2 <- data.frame(
"id|Identifier" = 1:3,
"score|Total score" = c(10, 20, 30),
check.names = FALSE
)
out2 <- label_from_names(df2, sep = "|")
Goodman-Kruskal's Lambda
Description
lambda_gk() computes Goodman-Kruskal's Lambda, a proportional
reduction in error (PRE) measure for nominal variables.
Usage
lambda_gk(
x,
direction = c("symmetric", "row", "column"),
detail = FALSE,
conf_level = 0.95,
digits = 3L
)
Arguments
x |
A contingency table (of class |
direction |
Direction of prediction:
|
detail |
Logical. If |
conf_level |
A single number strictly between 0 and 1 giving
the confidence level (default |
digits |
Number of decimal places used when printing the
result (default |
Details
Lambda measures how much prediction error is reduced when the independent variable is used to predict the dependent variable. It ranges from 0 (no reduction) to 1 (perfect prediction). Lambda can equal zero even when variables are associated if the modal category dominates in every column (or row).
The default direction = "symmetric" follows the SPSS and
DescTools convention: symmetric lambda is a standard,
well-defined variant with its own asymptotic standard error.
somers_d() deliberately differs (its default is "row")
because its symmetric form is a derived quantity without an
analytic SE; see its documentation.
Standard error formulas follow the DescTools implementations
(Signorell et al., 2024); see cramer_v() for full references.
Value
Same structure as cramer_v(): a scalar when
detail = FALSE, a named vector when detail = TRUE.
The p-value tests H0: lambda = 0 (Wald z-test).
References
Goodman, L. A., & Kruskal, W. H. (1954). Measures of association for cross classifications. Journal of the American Statistical Association, 49(268), 732-764. doi:10.2307/2281536
See Also
goodman_kruskal_tau(), uncertainty_coef(),
assoc_measures()
Other association measures:
assoc_measures(),
contingency_coef(),
cramer_v(),
gamma_gk(),
goodman_kruskal_tau(),
kendall_tau_b(),
kendall_tau_c(),
phi(),
somers_d(),
uncertainty_coef(),
yule_q()
Examples
tab <- table(sochealth$smoking, sochealth$education)
lambda_gk(tab)
lambda_gk(tab, direction = "row")
lambda_gk(tab, direction = "column", detail = TRUE)
Row means with an optional minimum-valid-values rule
Description
Computes row-wise means across selected numeric columns of a
data.frame or matrix. Missing values are handled per row via
min_valid (an integer count or proportion of non-NA values
required); rows that fail the rule return NA, and rows with no
valid values at all return NA even when min_valid = 0.
Non-numeric columns are dropped silently (set verbose = TRUE to
see which).
Designed to flow inside dplyr::mutate(): when called without
an explicit data argument, the current data context is used.
Usage
mean_n(
data = NULL,
select = tidyselect::everything(),
exclude = NULL,
min_valid = NULL,
digits = NULL,
regex = FALSE,
verbose = FALSE,
user_na = TRUE
)
Arguments
data |
A |
select |
Columns to include. If |
exclude |
Columns to exclude (default: |
min_valid |
Minimum number of valid (non-
Non-integer values Rows with zero valid values always return |
digits |
Optional non-negative integer giving the number of
decimal places to round the result to. Defaults to |
regex |
Logical. If |
verbose |
Logical. If |
user_na |
Logical. If |
Value
A numeric vector of row-wise means.
Declared missing values
Survey files imported with haven often carry declared missing
values: codes such as 8 = Don't know or 9 = Refused that the
source file marks as missing while keeping them distinct from a
plain NA. Two kinds of declaration exist: na_values / na_range
metadata on haven::labelled_spss() vectors, and tagged missing
values created by haven::tagged_na() (the Stata .a, .b, ...
convention).
spicy honors the declaration by default (user_na = TRUE):
declared missing values are excluded from every statistic exactly
like NA – valid percentages, means, chi-squared tests,
association measures, row-wise summaries, and group definitions –
but they are not erased from display. freq() lists each observed
declared value as its own row of the Missing block, with its value
label; cross_tab(), table_categorical(), and
table_continuous() disclose the exclusion in the table note
(Declared missing values removed: x (2).); varlist() and
code_book() count them as missing in N_valid / NAs /
N_distinct while still listing the declared codes in Values.
Every function involved offers the same escape hatch: set
user_na = FALSE to ignore the declaration and treat the declared
codes as valid values (the behavior of spicy before 0.13.0).
Tagged missing values are genuine NAs either way; for them,
user_na = FALSE only collapses the per-tag breakdown back into
the regular NA count.
See Also
Other row-wise summaries:
count_n(),
sum_n()
Examples
library(dplyr)
# Create a simple numeric data frame
df <- tibble(
var1 = c(10, NA, 30, 40, 50),
var2 = c(5, NA, 15, NA, 25),
var3 = c(NA, 30, 20, 50, 10)
)
# Compute row-wise mean (all values must be valid by default)
mean_n(df)
# Require at least 2 valid (non-NA) values per row
mean_n(df, min_valid = 2)
# Require at least 50% valid (non-NA) values per row
mean_n(df, min_valid = 0.5)
# Round the result to 1 decimal
mean_n(df, digits = 1)
# Select specific columns
mean_n(df, select = c(var1, var2))
# Select specific columns using a pipe
df |>
select(var1, var2) |>
mean_n()
# Exclude a column
mean_n(df, exclude = "var3")
# Select columns ending with "1"
mean_n(df, select = ends_with("1"))
# Use with native pipe
df |> mean_n(select = starts_with("var"))
# Use inside dplyr::mutate()
df |> mutate(mean_score = mean_n(min_valid = 2))
# Select columns directly inside mutate()
df |> mutate(mean_score = mean_n(select = c(var1, var2), min_valid = 1))
# Select columns before mutate
df |>
select(var1, var2) |>
mutate(mean_score = mean_n(min_valid = 1))
# Show verbose processing info
df |> mutate(mean_score = mean_n(min_valid = 2, digits = 1, verbose = TRUE))
# Add character and grouping columns
df_mixed <- mutate(df,
name = letters[1:5],
group = c("A", "A", "B", "B", "A")
)
df_mixed
# Non-numeric columns are ignored
mean_n(df_mixed)
# Use within mutate() on mixed data
df_mixed |> mutate(mean_score = mean_n(select = starts_with("var")))
# Use everything() but exclude non-numeric columns manually
mean_n(df_mixed, select = everything(), exclude = "group")
# Select columns using regex
mean_n(df_mixed, select = "^var", regex = TRUE)
mean_n(df_mixed, select = "ar", regex = TRUE)
# Apply to a subset of rows (first 3)
df_mixed[1:3, ] |> mean_n(select = starts_with("var"))
# Store the result in a new column
df_mixed$mean_score <- mean_n(df_mixed, select = starts_with("var"))
df_mixed
# With a numeric matrix
mat <- matrix(c(1, 2, NA, 4, 5, NA, 7, 8, 9), nrow = 3, byrow = TRUE)
mat
mat |> mean_n(min_valid = 2)
Phi coefficient
Description
phi() computes the phi coefficient for a 2x2 contingency table.
Usage
phi(x, detail = FALSE, conf_level = 0.95, digits = 3L)
Arguments
x |
A contingency table (of class |
detail |
Logical. If |
conf_level |
A single number strictly between 0 and 1 giving
the confidence level (default |
digits |
Number of decimal places used when printing the
result (default |
Details
The phi coefficient is \phi = \sqrt{\chi^2 / n}. It is
equivalent to Cramer's V for 2x2 tables and equals the absolute
value of the Pearson correlation between the two binary
variables – spicy returns only the magnitude (always
non-negative), matching the DescTools (Signorell et al., 2024)
and PSPP conventions. SPSS itself SIGNS phi on 2x2 tables (its
CROSSTABS algorithm sets the sign to that of the Pearson
correlation), so SPSS output can show a negative value of the
same magnitude. To recover the signed direction of the
2x2 association, compute the Pearson correlation directly
(e.g. cor(x, y) after coding both variables 0/1).
The confidence interval uses the Fisher z-transformation on
\phi; see cramer_v() for the formula and full references.
Value
Same structure as cramer_v(): a scalar when
detail = FALSE, a named vector when detail = TRUE.
The p-value tests the null hypothesis of no association
(Pearson chi-squared test).
References
Liebetrau, A. M. (1983). Measures of Association. Sage.
See Also
cramer_v(), yule_q(), assoc_measures()
Other association measures:
assoc_measures(),
contingency_coef(),
cramer_v(),
gamma_gk(),
goodman_kruskal_tau(),
kendall_tau_b(),
kendall_tau_c(),
lambda_gk(),
somers_d(),
uncertainty_coef(),
yule_q()
Examples
tab <- table(sochealth$smoking, sochealth$sex)
phi(tab)
phi(tab, detail = TRUE)
Print a detailed association measure result
Description
Formats a spicy_assoc_detail vector (returned by association
functions with detail = TRUE) with fixed decimal places and
< 0.001 notation for small p-values.
Usage
## S3 method for class 'spicy_assoc_detail'
print(x, digits = attr(x, "digits") %||% 3L, ...)
Arguments
x |
A |
digits |
Number of decimal places for the estimate, SE, and
confidence interval. Defaults to 3. The p-value is always
formatted separately using APA notation ( |
... |
Ignored. |
Value
x, invisibly.
See Also
Print an association measures summary table
Description
Formats a spicy_assoc_table data frame (returned by
assoc_measures()) with fixed decimal places, aligned columns,
and APA-style <.001 notation for small p-values (same helper as
cross_tab() and the table_*() family).
Usage
## S3 method for class 'spicy_assoc_table'
print(x, digits = attr(x, "digits") %||% 3L, ...)
Arguments
x |
A |
digits |
Number of decimal places for estimates, SE, and
confidence intervals. Defaults to 3. The p-value is always
formatted separately using APA notation ( |
... |
Ignored. |
Value
x, invisibly.
See Also
Print method for categorical survey-design tables
Description
Formats and prints a spicy_categorical_svy_table object as a
styled ASCII table using spicy_print_table().
Usage
## S3 method for class 'spicy_categorical_svy_table'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (currently ignored). |
Value
Invisibly returns x.
See Also
table_categorical_svy(), spicy_print_table()
Print method for categorical summary tables
Description
Formats and prints a spicy_categorical_table object as a styled ASCII table using
spicy_print_table().
Usage
## S3 method for class 'spicy_categorical_table'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (currently ignored). |
Value
Invisibly returns x.
See Also
table_categorical(), spicy_print_table()
Print method for bivariate linear-model tables
Description
Formats and prints a spicy_continuous_lm_table object as a styled
ASCII table using spicy_print_table().
Usage
## S3 method for class 'spicy_continuous_lm_table'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (currently ignored). |
Value
Invisibly returns x.
See Also
table_continuous_lm(), spicy_print_table()
Print method for continuous survey-design tables
Description
Formats and prints a spicy_continuous_svy_table object as a styled
ASCII table using spicy_print_table().
Usage
## S3 method for class 'spicy_continuous_svy_table'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (currently ignored). |
Value
Invisibly returns x.
See Also
table_continuous_svy(), spicy_print_table()
Print method for continuous summary tables
Description
Formats and prints a spicy_continuous_table object as a styled ASCII
table using spicy_print_table().
Usage
## S3 method for class 'spicy_continuous_table'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (currently ignored). |
Value
Invisibly returns x.
See Also
table_continuous(), spicy_print_table()
Print method for spicy_cross_table objects
Description
Prints a formatted SPSS-like crosstable created by cross_tab().
Usage
## S3 method for class 'spicy_cross_table'
print(x, digits = NULL, decimal_mark = NULL, ...)
Arguments
x |
A |
digits |
Optional integer; number of decimal places to display for cell values. Defaults to the value stored in the object. |
decimal_mark |
Optional character ( |
... |
Additional arguments passed to internal formatting functions. |
Value
Invisibly returns x.
Internal print method for lists of cross-tab tables
Description
Prints each element of a spicy_cross_table_list object on its own,
inserting a blank line between tables.
Usage
## S3 method for class 'spicy_cross_table_list'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments passed to individual print methods. |
Value
Invisibly returns x.
Print method for spicy-tagged flextables
Description
Prints a spicy_flextable object – the flextable::flextable()
returned by output = "flextable", tagged so the table note keeps
its styling in interactive HTML display. Every flextable verb works
on the tagged object directly; printing delegates to flextable's
own rendering (as_flextable.spicy_flextable() returns the
untagged object).
Usage
## S3 method for class 'spicy_flextable'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (currently ignored). |
Value
Invisibly returns NULL (HTML display path) or the result
of flextable's own print method.
See Also
table_regression(), as_flextable.spicy_flextable()
Print method for freq() tables
Description
Formats and prints a spicy_freq_table object as a styled ASCII
table using spicy_print_table().
Usage
## S3 method for class 'spicy_freq_table'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (currently ignored). |
Value
Invisibly returns x.
See Also
Print method for spicy-tagged gt tables
Description
Prints a spicy_gt object – the gt::gt() table returned by
output = "gt", tagged so the table note keeps its styling in
interactive HTML display. Every gt verb works on the tagged object
directly; printing delegates to gt's own rendering.
Usage
## S3 method for class 'spicy_gt'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (currently ignored). |
Value
Invisibly returns NULL (HTML display path) or the result
of gt's own print method.
See Also
Print method for outcome tables
Description
Formats and prints a spicy_outcome_table object as a styled ASCII
table using spicy_print_table().
Usage
## S3 method for class 'spicy_outcome_table'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (currently ignored). |
Value
Invisibly returns x.
See Also
table_outcome(), spicy_print_table()
Print method for internal regression frames
Description
Prints a compact one-glance summary of a spicy_regression_frame
(the internal intermediate representation behind
table_regression()): model class, sample size, coefficient
dimensions, family / link, CI method, and the capability flags the
frame advertises.
Usage
## S3 method for class 'spicy_regression_frame'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (currently ignored). |
Value
Invisibly returns x.
See Also
Print method for regression tables
Description
Formats and prints a spicy_regression_table object as a styled
ASCII table: title banner, spanner row for multi-model tables,
decimal-aligned body with factor grouping and reference rows,
fit-statistics block, and footer note.
Usage
## S3 method for class 'spicy_regression_table'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (currently ignored). |
Value
Invisibly returns x.
See Also
table_regression(), spicy_print_table()
Print method for table styles
Description
Prints what a spicy_style() encodes: for a named theme, the
journal, the rules it applies, and the official document they come
from; then the levers themselves.
Usage
## S3 method for class 'spicy_style'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (currently ignored). |
Value
Invisibly returns x.
See Also
Simulated social-health survey
Description
A simulated dataset of 1200 respondents from a fictional social-health survey, designed to illustrate the main features of the spicy package: variable labels, ordered factors, survey weights, association measures, and publication-ready reporting.
Usage
sochealth
Format
A tibble with 1200 rows and 24 variables:
- sex
Factor. Sex of the respondent.
- age
Numeric. Age in years (25–75).
- age_group
Ordered factor. Age group (25–34, 35–49, 50–64, 65–75).
- education
Ordered factor. Highest education level (Lower secondary, Upper secondary, Tertiary).
- social_class
Ordered factor. Subjective social class (Lower, Working, Lower middle, Middle, Upper middle).
- region
Factor. Region of residence (6 regions).
- employment_status
Factor. Employment status (Employed, Student, Unemployed, Inactive).
- income_group
Ordered factor. Household income group (Low, Lower middle, Upper middle, High). Contains missing values.
- income
Numeric. Monthly household income in CHF (1000–7400).
- smoking
Factor. Current smoker (No, Yes). Contains missing values.
- physical_activity
Factor. Regular physical activity (No, Yes).
- dentist_12m
Factor. Dentist visit in the last 12 months (No, Yes).
- self_rated_health
Ordered factor. Self-rated health (Poor, Fair, Good, Very good). Contains missing values.
- wellbeing_score
Numeric. WHO-5 wellbeing index (0–100).
- bmi
Numeric. Body mass index in kg/m
^2(16–39). Contains missing values.- bmi_category
Ordered factor. BMI category (Normal weight, Overweight, Obesity). Contains missing values.
- institutional_trust
Ordered factor. Trust in institutions (Very low, Low, High, Very high).
- political_position
Numeric. Political position on a 0 (left) to 10 (right) scale. Contains missing values.
- life_sat_health
Integer. Satisfaction with own health (1–5 Likert scale). Contains missing values.
- life_sat_work
Integer. Satisfaction with work or main activity (1–5 Likert scale). Contains missing values.
- life_sat_relationships
Integer. Satisfaction with personal relationships (1–5 Likert scale). Contains missing values.
- life_sat_standard
Integer. Satisfaction with standard of living (1–5 Likert scale). Contains missing values.
- response_date
POSIXct. Date and time of survey response (September–November 2024).
- weight
Numeric. Survey design weight (range 0.29–3.45); calibrated so that
sum(weight)matches the unweighted N andmean(weight)is approximately 1. SeeDetails.
Details
Every variable carries a "label" attribute (read by
labelled::var_label() and surfaced by varlist() /
code_book()). The mix of factor types is deliberate: nominal
factors (sex, region, ...) and ordered factors (education,
self_rated_health, ...) live side by side so that
cross_tab() and table_categorical() can demonstrate the
automatic ordinal-vs-nominal dispatch (Cramer's V, Phi, Kendall's
Tau-b, Goodman-Kruskal Gamma) on the same dataset.
Survey weights (weight) are calibrated: sum(weight) matches
the unweighted N to within rounding (\approx 1200) and
mean(weight) is \approx 1. Weighted means therefore agree
with unweighted means up to sampling noise without further
rescaling.
Source
Simulated data for illustration purposes; reproducible by
sourcing data-raw/sochealth.R. The script seeds the main
generation block with set.seed(2025), and the two
missing-value injection blocks with set.seed(2027) (the
four life_sat_* items) and set.seed(2026) (smoking,
self_rated_health, income_group, political_position,
bmi).
Examples
data(sochealth)
varlist(sochealth)
freq(sochealth, education)
cross_tab(sochealth, education, self_rated_health)
Somers' D
Description
somers_d() computes Somers' D for a two-way contingency
table of ordinal variables.
Usage
somers_d(
x,
direction = c("row", "column", "symmetric"),
detail = FALSE,
conf_level = 0.95,
digits = 3L
)
Arguments
x |
A contingency table (of class |
direction |
Direction of prediction:
|
detail |
Logical. If |
conf_level |
A single number strictly between 0 and 1 giving
the confidence level (default |
digits |
Number of decimal places used when printing the
result (default |
Details
Somers' D is an asymmetric ordinal measure defined as
d = (C - D) / (C + D + T), where T is the
number of pairs tied on the independent variable. The
symmetric version (direction = "symmetric") is the
harmonic mean of the two asymmetric values, matching the
SPSS / PSPP convention; this is not identical to
Kendall's Tau-b (which is the geometric mean of the same
two quantities), although the two often agree to two
decimals. It is computed via the equivalent closed form
2(C - D) divided by the sum of the two asymmetric
denominators, so a table with exactly as many concordant as
discordant pairs (e.g. an independence pattern) yields the
well-defined value 0 – the harmonic-mean form is 0/0 there –
as printed by SPSS / PSPP. The symmetric estimate
is NA only when one of the asymmetric directions is itself
undefined (with the same spicy_undefined_stat warning). No
analytic SE / CI is reported for the symmetric form: its se
is always NA, matching PSPP, which prints no ASE for it
(DescTools offers no symmetric form at all).
The default direction = "row" differs deliberately from
lambda_gk() and uncertainty_coef() (which default to
"symmetric"): Somers' D is intrinsically asymmetric, its
symmetric form is a derived convention without inference, and
the reference implementation (DescTools::SomersDelta())
offers only the two asymmetric directions, defaulting to
"row".
Standard error formulas for the asymmetric directions follow
the DescTools implementations (Signorell et al., 2024); see
cramer_v() for full references.
Value
Same structure as cramer_v(): a scalar when
detail = FALSE, a named vector when detail = TRUE.
The p-value tests H0: D = 0 (Wald z-test).
References
Somers, R. H. (1962). A new asymmetric measure of association for ordinal variables. American Sociological Review, 27(6), 799-811. doi:10.2307/2090408
Brown, M. B., & Benedetti, J. K. (1977). Sampling behavior of tests for correlation in two-way contingency tables. Journal of the American Statistical Association, 72(358), 309-315. doi:10.1080/01621459.1977.10480995
See Also
kendall_tau_b(), gamma_gk(), assoc_measures()
Other association measures:
assoc_measures(),
contingency_coef(),
cramer_v(),
gamma_gk(),
goodman_kruskal_tau(),
kendall_tau_b(),
kendall_tau_c(),
lambda_gk(),
phi(),
uncertainty_coef(),
yule_q()
Examples
tab <- table(sochealth$education, sochealth$self_rated_health)
somers_d(tab, direction = "row")
somers_d(tab, direction = "column", detail = TRUE)
Table labels and their language
Description
Every string a reader of a spicy table sees – column headers, row
labels, titles, table footnotes – is held under a stable key, and
spicy_labels() returns those keys with the label each one currently
resolves to. It is the companion of the two options that move them.
Usage
spicy_labels(language = NULL)
Arguments
language |
A language name ( |
Value
A named character vector, one element per key, in registry order.
Global options
-
options(spicy.language = "fr")The language of the table, for the whole document – its labels and the typography its numbers are written in. Two sets ship:"en"(the default) and"fr". The language of a report is a property of the report, so this is set once in a setup chunk rather than passed to each table. -
options(spicy.labels = list(<key> = "<label>"))A per-label override, for the case where one word has to change and a language does not: a named list (or named character vector) whose names are keys ofspicy_labels(). An unknown key is an error.
A label resolves through spicy.labels, then the spicy.language
set, then English. A set carries only the keys it translates, so
anything it does not name falls back to English rather than erroring
or coming out blank. Both options are cleared with NULL.
What a language does not change
Only DISPLAY strings translate. The column names of the exported
frames (as.data.frame(), tidy(), as_structured()) are a
documented contract that user code indexes into, so out[["Yes %"]]
resolves under every language, and so do the block identities, the
encoded cell tokens and the mathematical glyphs. Errors, warnings and
messages stay English: they are read by developers and quoted in bug
reports.
One exception, and it follows from the same rule. A column named
after a LEVEL of by takes that level's own spelling – which is why
"Yes %" stays "Yes %": your data is never translated. spicy's own
missing category is such a level, so it is the one name that does
move: table_categorical(by = ) on a variable with missing values
gives "(Missing) n" in English and "(Manquant) n" under "fr"
(or whatever row_missing_level is overridden to). Address that
column through row_missing_level rather than by typing it.
Numbers follow the language too
A language brings its typographic locale with it, so
options(spicy.language = "fr") is one gesture for a coherent
French report table: French words, comma decimal mark, and the
leading zero French typography keeps on a p-value (0,003). The
sources are the BIPM's SI brochure and the European Union's Code
de redaction interinstitutionnel; spicy_style() quotes them.
"en" brings no locale, and nothing changes for anyone who sets no
language.
The locale rides the style layer, so it reaches the reporting
families – table_regression(), table_categorical(),
table_continuous(), table_continuous_lm(), table_outcome()
and the survey twins. The language reaches every table, the
exploration pair included: freq() and cross_tab() have no
style layer, so it sets the DEFAULT of their decimal_mark – the
one typographic lever they carry – and an argument you type wins
over it. Under a comma their p-value keeps its leading zero
(p = 0,659), the form French typography requires.
The locale sits at the BOTTOM of the formatting resolution. A
journal style outranks it – style = "jama" under "fr" gives
JAMA's p-values and the French comma, style = "lancet" keeps the
journal's midline decimal point – and an argument you type outranks
both, so decimal_mark = "." gives French words with a decimal
point (the p-value keeps its locale zero; spicy_style() documents
that lever).
Two timing notes. Figures are frozen when a table is BUILT, like every formatting argument, while words resolve when it is printed – so set the language once, at the top of the document. And a comma mark makes an Excel export write its body as text rather than live numbers, since Excel would otherwise re-punctuate a numeric cell with the viewer's own locale.
See Also
spicy_style() for the journal styles, which outrank the
language's own typography.
Examples
head(spicy_labels())
# The same keys in French.
fr <- spicy_labels("fr")
fr[["header_mean"]]
fr[["row_missing_level"]]
# One label, not a language: the missing CATEGORY of the grouping
# variable is a refusal to answer here, not an absent value.
old <- options(spicy.labels = list(row_missing_level = "(No answer)"))
table_categorical(sochealth, select = sex, by = smoking)
options(old)
Print a spicy-formatted ASCII table
Description
User-facing helper that prints a spicy-styled ASCII table to the
console with optional title and note, table-type-aware alignment
defaults, and automatic horizontal panelling when the table is
wider than the console. Wraps the internal renderer
build_ascii_table().
Usage
spicy_print_table(
x,
title = attr(x, "title"),
note = attr(x, "note"),
padding = 2L,
first_column_line = TRUE,
row_total_line = TRUE,
bottom_line = FALSE,
lines_color = "darkgrey",
align_left_cols = NULL,
align_center_cols = integer(0),
center_headers = FALSE,
spanners = NULL,
group_sep_rows = integer(0),
total_row_idx = attr(x, "total_row_idx"),
display_labels = NULL,
fit_stats_start = NULL,
qualify_companions = FALSE,
...
)
Arguments
x |
A |
title |
Optional title displayed above the table. Defaults to the
|
note |
Optional note displayed below the table. Defaults to the |
padding |
Non-negative integer giving the number of extra
characters added to each column's auto-computed width
(max of cell-content width and header width). Defaults to
|
first_column_line |
Logical. If |
row_total_line, bottom_line |
Logical flags controlling the
horizontal line before a total row and the closing line at the
bottom of the table. |
lines_color |
Character. Color for table separators. Defaults to |
align_left_cols |
Integer vector of column indices to left-align.
If
|
align_center_cols |
Integer vector of column indices to
center-align. Defaults to |
center_headers |
Logical. When |
spanners |
Optional named list of column-group labels
(label -> integer column indices). Passed through to
|
group_sep_rows |
Integer vector of row indices before which a
light dashed separator line is drawn. Defaults to |
total_row_idx |
Optional integer vector of 1-based row indices
identifying the totals rows; defaults to the |
display_labels |
Optional character vector of length |
fit_stats_start |
Optional 1-based index of the first
model-level statistics row (the block below the dashed rule in
regression tables). When the table splits into stacked panels,
continuation panels drop the rows of that block whose every
visible data cell is blank – model-level statistics print once,
under the panel that carries their values, instead of leaving
empty |
qualify_companions |
Logical. When the table splits into stacked
panels, should a companion column ( |
... |
Additional arguments passed to |
Details
Table type is auto-detected from x and drives the default
alignment when align_left_cols = NULL:
-
frequency table (a
Categorycolumn is present): the first two columns (Category,Values) are left-aligned. -
cross table (otherwise): only the first column (row variable) is left-aligned.
If the table is wider than the console, it is split into stacked
horizontal panels with the left-most identifier columns repeated
on each panel. Unicode line-drawing characters are used by
default; coloured separators are drawn when the terminal supports
ANSI colour (crayon::has_color()) and fall back to monochrome
otherwise.
The layout arguments spanners, display_labels,
fit_stats_start, total_row_idx and group_sep_rows are
plumbing consumed by spicy's own print methods; they are
documented for completeness and are rarely useful when calling
this function directly.
Value
Invisibly returns x, after printing the formatted ASCII table to the console.
See Also
build_ascii_table() for the underlying text rendering engine.
print.spicy_freq_table() for the specialized printing method used by freq().
Examples
# Simple demonstration
df <- data.frame(
Category = c("Valid", "", "Missing", "Total"),
Values = c("Yes", "No", "NA", ""),
Freq. = c(12, 8, 1, 21),
Percent = c(57.1, 38.1, 4.8, 100.0)
)
spicy_print_table(df,
title = "Frequency table: Example",
note = "Class: data.frame\nData: demo"
)
Build or select a table style
Description
A style is a small set of number-formatting rules – how many
decimals a p-value gets, where it bottoms out, whether it keeps its
leading zero, what the decimal mark is, how a confidence interval is
written. spicy_style() composes one by hand; the named themes
("jama", "nejm", "lancet", "annals", "apa", "aer" –
spicy_style_names() returns the list) are pre-composed ones, each
encoding rules taken verbatim from an official document of the
institution.
Usage
spicy_style(
base = NULL,
...,
p_style = NULL,
p_digits = NULL,
p_floor = NULL,
p_bands = NULL,
p_sigfig = NULL,
decimal_mark = NULL,
ci_sep = NULL,
ci_brackets = NULL,
stars = NULL,
digits = NULL,
effect_size_digits = NULL,
fit_digits = NULL,
ic_digits = NULL,
percent_digits = NULL,
v_digits = NULL
)
spicy_style_names()
Arguments
base |
Optional theme to start from – a theme name or another
|
... |
Must be empty. Any argument landing here is a misspelt lever and raises an error rather than being ignored. |
p_style |
How p-values carry their leading zero: |
p_digits |
Decimal places for p-values (a positive integer).
Under |
p_floor |
The value below which a p-value prints as |
p_bands |
Decimal places that vary with the size of the
p-value: a list of two-element numeric vectors
|
p_sigfig |
Significant figures for p-values (a positive
integer), capped at |
decimal_mark |
A single character; the mark between the integer and the decimal part of every number in the table. |
ci_sep |
The string between the two bounds of a confidence
interval. |
ci_brackets |
A character vector of length two, the opening and
closing delimiters of a confidence interval, e.g. |
stars |
Significance stars: |
digits, effect_size_digits, fit_digits, ic_digits |
Decimal places for estimates, effect sizes, fit statistics and information criteria. Each is applied only by the table functions that have the matching argument. |
percent_digits, v_digits |
Decimal places for percentages and
for the association measure of |
Details
A style is accepted by the style argument of table_regression(),
table_categorical(), table_continuous() and
table_continuous_lm(), either as a theme name or as the object
returned here, and by options(spicy.style = ) for document-wide
scope.
Value
spicy_style() returns an object of class spicy_style: a
named list holding only the levers that were set. Themes returned
by name carry a provenance attribute with the journal, its
source document and the exact list of encoded rules; print()
shows it. spicy_style_names() returns the character vector of
available theme names.
What a theme claims, and what it does not
A theme covers numeric formatting conformity only – not full editorial conformity. It does not check reporting guidelines, table structure, footnote symbols, units, abbreviation policy, or anything else a journal's instructions ask of a manuscript. Each theme below lists the exact rules it encodes, with the sentence it encodes them from; anything not listed is spicy's own default, not the journal's rule.
Themes only move defaults. An argument you type wins over the
theme, always: table_regression(fit, style = "lancet", digits = 3)
is a Lancet table with three decimals. One consequence worth
stating: typing p_digits is a request for that many decimals on
every p-value, so it also switches off the theme's own ways of
choosing p-precision – p_bands, p_sigfig, and the p_floor
derived from them. The leading-zero rule is orthogonal and stays.
Resolution order
An argument you type, then the style argument, then
getOption("spicy.style"), then the typography of
getOption("spicy.language"), then spicy's defaults. Within a
style, an explicit function argument beats the style's value for
the same lever.
The language's locale
A language brings its typography with it, at the bottom of that
order. options(spicy.language = "fr") is therefore one gesture
for a coherent French report table – French words, and numbers
written the French way: comma decimal mark, and the leading zero
French typography keeps on a p-value (0,003). "La virgule est
utilisee pour separer les unites des decimales" (European Union,
Code de redaction interinstitutionnel, French edition 2022,
point 6.5, https://style-guide.europa.eu/fr); "Si le nombre se
situe entre +1 et -1, le separateur decimal est toujours precede
d'un zero" (BIPM, Le Systeme international d'unites, 9th
edition, 2019, section 5.4.4,
https://www.bipm.org/documents/20126/41483022/SI-Brochure-9.pdf).
"en" brings no locale, and nothing changes for anyone who sets
no language. The locale reaches the exploration pair as well:
freq() and cross_tab() have no style layer, so the language
sets the DEFAULT of their decimal_mark and nothing else. An
argument you type still wins. The leading zero of a p-value works
the same way in both worlds: its DEFAULT follows the mark – a
comma keeps it (0,003), a point drops it (.003) – and
p_style is the explicit lever that overrides that default
wherever there is one, which is why a theme's rule survives any
mark in the table_*() families while the pair, having no such
lever, always follows its mark.
A theme composes with a locale rather than fighting it, because a
theme encodes only what its own source states: "jama" fixes no
decimal mark, so JAMA under a French language gives JAMA's p-value
rules and the French comma. Where the two do meet, the theme wins –
you asked for it by name. "lancet" keeps its midline decimal
point; "apa" keeps its missing leading zero, which under a comma
prints ,003, a form the SI brochure forbids. That is the price of
an explicit gesture; to keep APA's other rules and restore the
zero, compose the way out yourself:
style = spicy_style("apa", p_style = "standard"). One lever bends
the other way: a theme's ", " interval separator was sourced
under a dot mark, and under a comma mark it would BE the mark – so
it yields to the derived "; ", exactly as the French adaptations
of APA style themselves write ([3,45; 6,78]). Any other
separator (" to ", an en dash) is unambiguous and stays.
An argument beats both, which is the escape hatch for a bilingual
table: decimal_mark = "." under a French language gives French
words and a decimal point. It moves only the mark: the p-value
keeps the locale's leading zero (0.003), because p_style is a
style lever with no argument of its own –
style = spicy_style(p_style = "apa") drops it again.
See spicy_labels() for the language option itself.
Themes
"jama" – JAMA / JAMA Network
Source: Instructions for Table Creation, JAMA Network author
document, 23 February 2016
(https://jamanetwork.com/DocumentLibrary/InstructionsForAuthors/InstructionsForTableCreation.pdf),
consulted 2026-08-14.
Encoded:
p-values on two decimals, three below
.01, floored at<.001, with no leading zero – "All P values should be reported to exact numbers to 2 digits past the decimal point, regardless of statistical significance. For values lower than .01, present the P value to 3 digits. Express any values lower than .001 as P<.001."
Not encoded: everything else in that document (percentages carrying their numerator and denominator, one datum per cell, footnote letters, SI conversion factors) is table construction, not number formatting. The document states no rule for the confidence-interval separator or the decimal mark, so spicy's defaults apply.
"nejm" – The New England Journal of Medicine
Source: NEJM Author Center, New Manuscripts, section Statistical
Reporting Guidelines
(https://www.nejm.org/author-center/new-manuscripts), a living
page; its full text was read in a browser by the package maintainer
on 2026-08-14 – the site refuses automated retrieval, which is why
earlier surveys reported these rules as unavailable.
Encoded:
tiered p-values, leading zero kept – "In general, P values larger than 0.01 should be reported to two decimal places, and those between 0.01 and 0.001 to three decimal places; P values smaller than 0.001 should be reported as P<0.001." The guideline's own text writes
0.01/0.001: unlike JAMA, the leading zero stays.measures of association on two decimals – "measures of association, such as odds ratios, should ordinarily be reported to two decimal places" (a pin of spicy's default).
Not encoded: the guideline's stated exceptions to the p rule
(stopping-rule tests, genomewide studies) are analysis contexts the
style layer cannot see; its inference policy (no p-values without a
prespecified multiplicity plan, estimates + 95% CI instead, no
p-values in the Table 1 of a randomised trial) is about what to
report, not how to format it – request those layouts through
show_columns and p_value = FALSE where you need them.
"lancet" – The Lancet
Sources: Information for Authors, April 2026; Randomised trials
in The Lancet: formatting guidelines and Observational studies in
The Lancet: formatting guidelines, both last updated July 2025
(https://www.thelancet.com/pb-assets/Lancet/authors/tl-info-for-authors-1690986041530.pdf),
consulted 2026-08-14.
Below, the journal's midline decimal point (Unicode U+00B7, MIDDLE
DOT) is written [.] so it cannot be confused with an ordinary
full stop.
Encoded:
midline decimal mark, U+00B7, on every number in the table – "Type decimal points midline (ie, 23
[.]4, not 23.4)."p-values on two significant figures, capped at four decimals, floored at
p<0[.]0001, leading zero kept – "Supply p values to two significant figures (capped at four decimal places), or p<0[.]0001."en dash between the bounds of a confidence interval. This one is not a written rule: the journal states none. It matches the intervals printed throughout the journal's own model tables and figures (
0[.]78 (0[.]60-1[.]00)) and is encoded as conformity to the journal's published examples, nothing stronger.
Not encoded: the empty-cell filler, the ban on p-values in a randomised trial's baseline table, and the absolute-rather-than- relative effect rule are content decisions, not number formats.
One caveat on the en dash. The journal's examples are ratio
measures, whose bounds are positive. On the identity scale a
negative lower bound puts a minus sign next to the dash
([-5.17--2.58]), which reads badly. The journal states no rule
for that case and none is invented here; override the separator
when it arises:
spicy_style("lancet", ci_sep = " to ").
"annals" – Annals of Internal Medicine
Source: Information for Authors, American College of Physicians,
document publication date 08/04/2026
(https://www.acpjournals.org/pb-assets/pdf/AnnalsAuthorInfo-1755188286957.pdf),
consulted 2026-08-14.
Encoded:
p-values on three decimals up to 0.20 and two decimals above it, floored at
P<0.001, leading zero kept – "For P values between 0.001 and 0.20, please report the value to the nearest thousandth. For P values greater than 0.20, please report the value to the nearest hundredth. For P values less than 0.001, report as 'P<0.001.'"
Not encoded: the percentage rule ("Report percentages to one decimal
place ... when sample size is \ge 200", no decimals below
200) is conditional on a per-group sample size that the style layer
does not see; set percent_digits yourself. The mean (SD)
notation the journal asks for, and its refusal of \eqn{\pm}{+/-},
are already how spicy renders dispersion.
"apa" – APA Style, 7th edition
Sources: APA Style numbers and statistics guide, last updated
11 September 2024
(https://apastyle.apa.org/instructional-aids/numbers-statistics-guide.pdf);
Sample tables, last updated June 2024
(https://apastyle.apa.org/style-grammar-guidelines/tables-figures/sample-tables),
both consulted 2026-08-14.
Encoded – and this theme pins rather than changes: spicy's
defaults already follow the guide, so style = "apa" is a promise
that a future change of those defaults will not move an
APA-formatted table.
p-values on three decimals, floored at
<.001, no leading zero – "Report exact p values to two or three decimals (e.g., p = .006, p = .03)", "report p values less than .001 as 'p < .001.'", and "Do not use a zero before a decimal when the statistic cannot be greater than 1 (proportion, correlation, level of statistical significance)."estimates and dispersion on two decimals – "Report other means and standard deviations and correlations, proportions, and inferential statistics (t, F, chi-square) to two decimals."
confidence intervals in square brackets, bounds separated by a comma – the Sample tables page gives
95% CI [LL, UL]as one of the two official layouts.
Not encoded: the one-decimal rule for means of integer scales
(depends on what the variable measures, which spicy cannot know),
and the thousands separator (see Known gaps below). Nor the star
thresholds. The official correlation and ANOVA sample tables mark
.05 / .01 / .001, which is what stars = TRUE gives you –
but as spicy's own default, not as anything this theme pins: one
stars lever carries both the thresholds and the decision to show
them, and the APA regression sample table shows none. So the theme
sets no stars, and style = "apa" never switches them on.
"aer" – American Economic Review / AEA journals
Source: AER Style Guide, American Economic Association
(https://www.aeaweb.org/journals/aer/style-guide), consulted
2026-08-14. The Tables section is shared with AEJ: Applied, AEJ:
Macro, AEJ: Policy, JEL and AEA P&P.
Encoded:
leading zero on every decimal fraction, p-values included – "Place a zero in front of the decimal point in all decimal fractions (e.g., 0.357, not .357)."
no significance stars, pinned – "Do not use asterisks to denote significance of estimation results. Report the standard errors in parentheses." This is the only outright written ban on stars in the whole surveyed corpus, so it is encoded even though spicy already defaults to
stars = FALSE. Ask forshow_columns = c("b", "se")for the standard-error layout.
Not encoded: the guide fixes no number of decimals, no p-value floor, no interval format and no decimal mark, so spicy's defaults apply. Horizontal-rules-only, no shading, a nine-column maximum and "Panel A / Panel B" blocks are layout, not number formatting.
Known gaps
Two rules the sources state but this release does not encode:
-
thousands separator (APA's comma, the EU code's thin no-break space). spicy formats numbers through more than one path, and a separator applied to some of them only would print
1,234in one column and1234in the next – the exact inconsistency the SI brochure forbids ("the format used should not vary within one column"). -
significant-figure rounding of estimates (Epidemiology's
nn/n.n/0.nn, Science's "report only significant digits"). spicy'sdigitsis a decimals contract applied per column family; a blanket significant-figure mode would also round fit statistics, which those rules do not ask for.
Journals deliberately absent: BMJ (its house-style page could not be read from any official source, and the widely repeated "95% CI 1.2 to 3.4" rule traces to no BMJ document), Econometrica (verified negative: the official guidelines contain no numeric table rule), QJE (its instructions page could not be read), Epidemiology (see Known gaps). Naming any of them would mean inventing their rules.
Examples
fit <- lm(mpg ~ wt + hp, data = mtcars)
# A named theme.
table_regression(fit, style = "jama")
# An argument you type beats the theme.
table_regression(fit, style = "jama", p_digits = 4)
# A style composed by hand.
table_regression(fit, style = spicy_style(decimal_mark = ",",
p_style = "standard"))
# A theme with one rule changed.
table_regression(fit, style = spicy_style("lancet", ci_sep = " to "))
# Document-wide scope.
old <- options(spicy.style = "apa")
table_continuous(mtcars, c(mpg, wt))
options(old)
# What a theme encodes, and where it comes from.
spicy_style("lancet")
Spicy table engine
Description
Index page for the ASCII rendering engine that produces every
spicy console table (freq(), cross_tab(), the table_*()
family, and the association-measure printers). The engine
supports Unicode line drawing, ANSI colours via crayon
(with monochrome fallback), automatic colour-aware width
detection, configurable integer padding (0L / 2L / 4L),
per-column alignment, and horizontal panelling for tables
wider than the console.
User-facing entry points
-
freq()– one-way frequency tables -
cross_tab()– two-way cross-tabulations -
table_categorical(),table_continuous(),table_continuous_lm()– multi-variable summary tables
Rendering primitives (internal API)
-
spicy_print_table()– user-facing wrapper that adds title, note, table-type-aware alignment defaults, and panelling. -
build_ascii_table()– the underlying string renderer.
See Also
spicy for the full package overview, including the API stability tiers and the classed-condition taxonomy.
Row sums with an optional minimum-valid-values rule
Description
Computes row-wise sums across selected numeric columns of a
data.frame or matrix. Missing values are handled per row via
min_valid (an integer count or proportion of non-NA values
required); rows that fail the rule return NA, and rows with no
valid values at all return NA even when min_valid = 0.
Non-numeric columns are dropped silently (set verbose = TRUE to
see which).
Designed to flow inside dplyr::mutate(): when called without
an explicit data argument, the current data context is used.
Usage
sum_n(
data = NULL,
select = tidyselect::everything(),
exclude = NULL,
min_valid = NULL,
digits = NULL,
regex = FALSE,
verbose = FALSE,
user_na = TRUE
)
Arguments
data |
A |
select |
Columns to include. If |
exclude |
Columns to exclude (default: |
min_valid |
Minimum number of valid (non-
Non-integer values Rows with zero valid values always return |
digits |
Optional non-negative integer giving the number of
decimal places to round the result to. Defaults to |
regex |
Logical. If |
verbose |
Logical. If |
user_na |
Logical. If |
Value
A numeric vector of row-wise sums.
Declared missing values
Survey files imported with haven often carry declared missing
values: codes such as 8 = Don't know or 9 = Refused that the
source file marks as missing while keeping them distinct from a
plain NA. Two kinds of declaration exist: na_values / na_range
metadata on haven::labelled_spss() vectors, and tagged missing
values created by haven::tagged_na() (the Stata .a, .b, ...
convention).
spicy honors the declaration by default (user_na = TRUE):
declared missing values are excluded from every statistic exactly
like NA – valid percentages, means, chi-squared tests,
association measures, row-wise summaries, and group definitions –
but they are not erased from display. freq() lists each observed
declared value as its own row of the Missing block, with its value
label; cross_tab(), table_categorical(), and
table_continuous() disclose the exclusion in the table note
(Declared missing values removed: x (2).); varlist() and
code_book() count them as missing in N_valid / NAs /
N_distinct while still listing the declared codes in Values.
Every function involved offers the same escape hatch: set
user_na = FALSE to ignore the declaration and treat the declared
codes as valid values (the behavior of spicy before 0.13.0).
Tagged missing values are genuine NAs either way; for them,
user_na = FALSE only collapses the per-tag breakdown back into
the regular NA count.
See Also
Other row-wise summaries:
count_n(),
mean_n()
Examples
library(dplyr)
# Create a simple numeric data frame
df <- tibble(
var1 = c(10, NA, 30, 40, 50),
var2 = c(5, NA, 15, NA, 25),
var3 = c(NA, 30, 20, 50, 10)
)
# Compute row-wise sums (all values must be valid by default)
sum_n(df)
# Require at least 2 valid (non-NA) values per row
sum_n(df, min_valid = 2)
# Require at least 50% valid (non-NA) values per row
sum_n(df, min_valid = 0.5)
# Round the results to 1 decimal
sum_n(df, digits = 1)
# Select specific columns
sum_n(df, select = c(var1, var2))
# Select specific columns using a pipe
df |>
select(var1, var2) |>
sum_n()
# Exclude a column
sum_n(df, exclude = "var3")
# Select columns ending with "1"
sum_n(df, select = ends_with("1"))
# Use with native pipe
df |> sum_n(select = starts_with("var"))
# Use inside dplyr::mutate()
df |> mutate(sum_score = sum_n(min_valid = 2))
# Select columns directly inside mutate()
df |> mutate(sum_score = sum_n(select = c(var1, var2), min_valid = 1))
# Select columns before mutate
df |>
select(var1, var2) |>
mutate(sum_score = sum_n(min_valid = 1))
# Show verbose message
df |> mutate(sum_score = sum_n(min_valid = 2, digits = 1, verbose = TRUE))
# Add character and grouping columns
df_mixed <- mutate(df,
name = letters[1:5],
group = c("A", "A", "B", "B", "A")
)
df_mixed
# Non-numeric columns are ignored
sum_n(df_mixed)
# Use inside mutate with mixed data
df_mixed |> mutate(sum_score = sum_n(select = starts_with("var")))
# Use everything(), but exclude known non-numeric
sum_n(df_mixed, select = everything(), exclude = "group")
# Select columns using regex
sum_n(df_mixed, select = "^var", regex = TRUE)
sum_n(df_mixed, select = "ar", regex = TRUE)
# Apply to a subset of rows
df_mixed[1:3, ] |> sum_n(select = starts_with("var"))
# Store the result in a new column
df_mixed$sum_score <- sum_n(df_mixed, select = starts_with("var"))
df_mixed
# With a numeric matrix
mat <- matrix(c(1, 2, NA, 4, 5, NA, 7, 8, 9), nrow = 3, byrow = TRUE)
mat
mat |> sum_n(min_valid = 2)
Categorical summary table
Description
Builds a publication-ready frequency or cross-tabulation table for one or many categorical variables selected with tidyselect syntax.
With by, produces grouped cross-tabulation summaries (using
cross_tab() internally) with Chi-squared p-values and optional
association measures.
Without by, produces one-way frequency-style summaries.
Multiple output formats are available via output: a printed ASCII
table ("default"), a wide or long numeric data.frame
("data.frame", "long"), or publication-ready tables
("tinytable", "gt", "flextable", "excel", "clipboard",
"word").
Usage
table_categorical(
data,
select = tidyselect::everything(),
by = NULL,
labels = NULL,
levels_keep = NULL,
include_total = TRUE,
drop_na = FALSE,
weights = NULL,
rescale = FALSE,
correct = FALSE,
simulate_p = FALSE,
simulate_B = 2000,
percent_digits = 1,
p_digits = 3,
v_digits = 2,
assoc_measure = "auto",
assoc_ci = FALSE,
smd = FALSE,
decimal_mark = ".",
align = c("decimal", "center", "right"),
output = c("default", "data.frame", "long", "tinytable", "gt", "flextable", "excel",
"clipboard", "word"),
indent_text = " ",
indent_text_excel_clipboard = strrep("Â ", 6),
add_multilevel_header = TRUE,
blank_na_wide = FALSE,
excel_path = NULL,
excel_sheet = NULL,
clipboard_delim = "\t",
word_path = NULL,
user_na = TRUE,
style = NULL
)
Arguments
data |
A data frame. |
select |
Columns to include as row variables. Supports tidyselect
syntax and character vectors of column names. When omitted,
defaults to every eligible categorical column in |
by |
Optional grouping column used for columns/groups. Accepts an
unquoted column name or a single character column name. Factor
levels keep their declared order; any other |
labels |
An optional named character vector of variable
labels whose names match column names in |
levels_keep |
Optional character vector of levels to keep/order for row
modalities. If |
include_total |
Logical. If |
drop_na |
Logical. If |
weights |
Optional weights. Either |
rescale |
Logical. If |
correct |
Logical. If |
simulate_p |
Logical. If |
simulate_B |
Integer. Number of Monte Carlo replicates when
|
percent_digits |
Number of digits for percentages in report outputs.
Defaults to |
p_digits |
Integer >= 1. Number of decimal places used to
render p-values in the |
v_digits |
Number of digits for the association measure. Defaults
to |
assoc_measure |
Which association measure to report alongside the chi-squared p-value. Accepts four input shapes:
When a single measure is used for every row, the column header is
that measure's name (e.g.
|
assoc_ci |
Passed to |
smd |
Logical. If |
decimal_mark |
Decimal separator ( |
align |
Horizontal alignment of numeric columns in the
printed ASCII table and in the
In the |
output |
Output format. One of:
|
indent_text |
Prefix used for modality labels in report table building.
Defaults to |
indent_text_excel_clipboard |
Stronger indentation used in Excel and clipboard exports. Defaults to six non-breaking spaces. |
add_multilevel_header |
Logical. If |
blank_na_wide |
Logical. If |
excel_path |
Path for |
excel_sheet |
Sheet name for Excel export. |
clipboard_delim |
Delimiter for clipboard text export. Defaults
to |
word_path |
File path for |
user_na |
Logical. If |
style |
A journal style: a theme name ( |
Value
Depends on output:
-
"default": the underlyingdata.framecarrying the rendering metadata as attributes (S3 class"spicy_categorical_table"). The object is returned visibly, so a baretable_categorical(...)call auto-prints the styled ASCII table at the console whilet <- table_categorical(...)stays silent (printtto display the table). -
"data.frame": a widedata.framewith one row per variable–level combination. Whenbyis used, the columns areVariable,Level, and one pair ofn/\%columns per group level (plusTotalwheninclude_total = TRUE), followed byChi2,df,p, and the association measure column. Whenby = NULL, the columns areVariable,Level,n,\%. -
"long": a longdata.framewith columnsvariable,level,n,pct(plusgroup,chi2,df,pwhenbyis used). The association measure is always calledeffect_size, whichever measure it is, andeffect_size_typenames that measure per row ("cramer_v","phi", ...), or isNAon the rows of a variable givenassoc_measure = "none". The wide outputs instead name the column after the measure, orEffect sizewhen the row variables do not share one. Withsmd = TRUEthis output also carriessmdandsmd_type("binary"or"multinomial", the kernel the value came from); the wide outputs name that columnSMD. Like the association columns, both are ABSENT when the statistic is not requested. -
"tinytable": atinytableobject. -
"gt": agt_tblobject. -
"flextable": aflextableobject. -
"excel"/"word": writes to disk and returns the file path invisibly. -
"clipboard": copies the table and returns the displaydata.frameinvisibly.
The drop_na = TRUE disclosure travels with the table on every
route, not just the console: "default" prints it under the ASCII
table, "tinytable" / "gt" / "flextable" / "word" carry it as
a table note, "excel" writes it below the body, and
"data.frame" keeps the sentence verbatim in the
missing_note attribute (attr(x, "missing_note"), NULL when
nothing was removed) so a pipeline that renders the numbers itself
can still state what left the table. On the "tinytable" route the
note is set one size down; options(spicy.note_style) governs that
(see table_regression()).
The Excel sheet carries the same title the console prints on its first row; the table itself starts on row 3.
Tests
When by is used, each selected variable is cross-tabulated
against the grouping variable with cross_tab() and the omnibus
chi-squared p-value is reported in the p column. See
@param correct / simulate_p to switch on Yates' continuity
correction or Monte Carlo p-values, and @param assoc_measure
for the per-row dispatch table used by "auto" (2x2 -> Phi,
both ordered -> Kendall's Tau-b, otherwise Cramer's V). Without
by, the table reports the marginal frequency distribution of
each variable with no inferential statistics.
For model-based comparisons (cluster-robust SE, weighted contrasts,
fitted means) on continuous outcomes, see table_continuous_lm().
For descriptive (empirical) comparisons on continuous outcomes, see
table_continuous().
Standardized mean difference
smd = TRUE adds an SMD column with the balance diagnostic of
the Table 1 literature, on the variable row beside p. For a
two-category variable it is the Bernoulli form,
\mathrm{SMD} = \frac{p_1 - p_2}{\sqrt{(p_1(1-p_1) +
p_2(1-p_2)) / 2}}
with p the proportion of the SECOND category, signed,
group 1 minus group 2 in the order the table displays them. Note
the denominator: the Bernoulli variance p(1-p) at n, not
var() at n-1, which would be 19% off on a small table.
"Second category" is the order the table shows, which is worth
knowing for a logical variable: spicy displays TRUE then
FALSE, so the sign is taken on FALSE, where tableone and
cobalt coerce with factor() (FALSE, TRUE) and take it on
TRUE – the same magnitude with the opposite sign. Convert to a
factor with the level order you want if the direction matters.
For three or more categories it is the multivariate form of Yang and Dalton (2012, SAS Global Forum 335-2012),
\mathrm{SMD} = \sqrt{T' S^{-} T}
with T the difference of the two profiles of proportions
(first category dropped) and S the mean of their multinomial
covariance matrices. This is a Mahalanobis distance: it is
unsigned, it is not bounded by 1, and S^{-} is a
pseudo-inverse, because a declared-but-unobserved category makes
S singular and solve() would abort where the
pseudo-inverse returns exactly the value that category's absence
implies. Which kernel a row took is published as smd_type in the
"long" output, and the unsigned reading is stated in the table
note whenever a variable has more than two categories. The
MASS package is needed for this arm only.
Two profiles with no category in common have an infinite
standardized distance. The pseudo-inverse would quietly publish a
finite number there, so the cell is an en-dash and a classed
warning says why. The same applies when each group is constant on
a different category, where the naive route publishes 0 –
"perfectly balanced" for the most imbalanced variable possible.
Conventions shared with table_continuous(): exactly two groups
(three or more are refused, not averaged over pairs); complete
cases on the observed groups, so a drop_na = FALSE "(Missing)"
level is displayed and never enters the diagnostic; no confidence
interval and no p-value, by design. Under weights the profiles
are the weighted proportions, which makes this column agree with
both the frequency and the survey-design readings – a profile of
proportions is invariant to a global rescaling of the weights, so
rescale cannot move it. (Only the continuous arm has a
convention to choose there.)
The SMD cell keeps its leading zero where the association cell
drops it: the APA strip belongs to a bounded measure, and this one
is not bounded. The two columns therefore print 0.45 and .45
side by side, on purpose.
Two limits of the current grammar. This function has no p_value
argument, so the p column cannot be switched off here as it can in
table_continuous(); a complete balance table mixing continuous
and categorical variables will show a categorical p beside a
continuous column you removed. And inline() cannot quote this
SMD cell: like p and the association measure, it lives on the
variable row, which inline() cannot address on a variable that
has levels. The continuous SMD cell is quotable
(inline(tbl, x, "A", column = "smd")); for the categorical one,
read output = "long".
Display conventions
Decimal alignment, p-value formatting, and required suggested
packages per output engine are documented under @param align,
@param p_digits, and @param output respectively.
Counts are displayed as integers: weighted counts are rounded
(ties half to even, the R convention) at display time only, in
cells and margins alike – the SPSS Crosstabs convention. Cells
and margins are rounded independently, so small display
discrepancies are possible (e.g. two cells of exactly 0.5 each
display as 0 while their Total of 1.0 displays as 1). The
machine outputs ("data.frame", "long") carry the exact
weighted counts and full-precision percentages.
Declared missing values
Survey files imported with haven often carry declared missing
values: codes such as 8 = Don't know or 9 = Refused that the
source file marks as missing while keeping them distinct from a
plain NA. Two kinds of declaration exist: na_values / na_range
metadata on haven::labelled_spss() vectors, and tagged missing
values created by haven::tagged_na() (the Stata .a, .b, ...
convention).
spicy honors the declaration by default (user_na = TRUE):
declared missing values are excluded from every statistic exactly
like NA – valid percentages, means, chi-squared tests,
association measures, row-wise summaries, and group definitions –
but they are not erased from display. freq() lists each observed
declared value as its own row of the Missing block, with its value
label; cross_tab(), table_categorical(), and
table_continuous() disclose the exclusion in the table note
(Declared missing values removed: x (2).); varlist() and
code_book() count them as missing in N_valid / NAs /
N_distinct while still listing the declared codes in Values.
Every function involved offers the same escape hatch: set
user_na = FALSE to ignore the declaration and treat the declared
codes as valid values (the behavior of spicy before 0.13.0).
Tagged missing values are genuine NAs either way; for them,
user_na = FALSE only collapses the per-tag breakdown back into
the regular NA count.
See Also
table_continuous() for empirical comparisons on
continuous outcomes; table_continuous_lm() for the model-based
companion (heteroskedasticity-consistent / cluster-robust /
bootstrap / jackknife SE, fitted means, weighted contrasts);
cross_tab() for two-way cross-tabulations; freq() for
one-way frequency tables.
Other spicy tables:
table_continuous(),
table_continuous_lm()
Examples
# --- Basic usage ---------------------------------------------------------
# Default: ASCII console table grouped by sex.
table_categorical(
sochealth,
select = c(smoking, physical_activity),
by = sex
)
# One-way frequency-style table (no `by`).
table_categorical(
sochealth,
select = c(smoking, physical_activity)
)
# Pretty labels keyed by column name.
table_categorical(
sochealth,
select = c(smoking, physical_activity),
by = education,
labels = c(
smoking = "Current smoker",
physical_activity = "Physical activity"
)
)
# Survey weights with rescaling.
table_categorical(
sochealth,
select = c(smoking, physical_activity),
by = education,
weights = "weight",
rescale = TRUE
)
# Confidence interval for the association measure.
table_categorical(
sochealth,
select = smoking,
by = education,
assoc_ci = TRUE
)
# --- Per-variable association measure ----------------------------------
# Default (`assoc_measure = "auto"`): one measure per row variable based on
# the variable type (2x2 -> Phi, both ordered factors -> Kendall's Tau-b,
# otherwise Cramer's V). When the chosen measures differ across rows, the
# column header collapses to `"Effect size"` and an APA-style `Note.` line
# documents which measure was used for which variable.
table_categorical(
sochealth,
select = c(smoking, education),
by = sex
)
# Force a uniform measure across all row variables.
table_categorical(
sochealth,
select = c(smoking, education),
by = sex,
assoc_measure = "cramer_v"
)
# Per-variable override (recommended named form).
table_categorical(
sochealth,
select = c(smoking, education, self_rated_health),
by = sex,
assoc_measure = c(
smoking = "phi", # binary x binary
education = "cramer_v", # multi-category nominal
self_rated_health = "tau_b" # ordinal x binary, Tau-b
)
)
# --- Output formats -----------------------------------------------------
# The rendered outputs below all wrap the same call:
# table_categorical(sochealth,
# select = c(smoking, physical_activity),
# by = sex)
# only `output` changes. Assign each result to a variable -- some
# engines auto-print as a console-friendly text fallback inside
# the `?` help viewer.
# Wide data.frame (one row per modality).
table_categorical(
sochealth,
select = c(smoking, physical_activity),
by = sex,
output = "data.frame"
)
# Long data.frame (one row per (modality x group)).
table_categorical(
sochealth,
select = c(smoking, physical_activity),
by = sex,
output = "long"
)
# Rendered HTML / docx objects -- best viewed inside a
# Quarto / R Markdown document or a pkgdown article.
if (requireNamespace("tinytable", quietly = TRUE)) {
tt <- table_categorical(
sochealth, select = c(smoking, physical_activity), by = sex,
output = "tinytable"
)
}
if (requireNamespace("gt", quietly = TRUE)) {
tbl <- table_categorical(
sochealth, select = c(smoking, physical_activity), by = sex,
output = "gt"
)
}
if (requireNamespace("flextable", quietly = TRUE)) {
ft <- table_categorical(
sochealth, select = c(smoking, physical_activity), by = sex,
output = "flextable"
)
}
# Excel and Word: write to a temporary file.
if (requireNamespace("openxlsx2", quietly = TRUE)) {
tmp <- tempfile(fileext = ".xlsx")
table_categorical(
sochealth, select = c(smoking, physical_activity), by = sex,
output = "excel", excel_path = tmp
)
unlink(tmp)
}
if (
requireNamespace("flextable", quietly = TRUE) &&
requireNamespace("officer", quietly = TRUE)
) {
tmp <- tempfile(fileext = ".docx")
table_categorical(
sochealth, select = c(smoking, physical_activity), by = sex,
output = "word", word_path = tmp
)
unlink(tmp)
}
## Not run:
# Clipboard: writes to the system clipboard.
table_categorical(
sochealth, select = c(smoking, physical_activity), by = sex,
output = "clipboard"
)
## End(Not run)
Categorical summary table from a survey design
Description
The design twin of table_categorical(): counts and estimated
percentages of categorical variables computed from a
survey::svydesign() or survey::as.svrepdesign() object instead
of a data frame.
Every statistic is survey's. survey::svymean() estimates the
percentages and their design effects, survey::svyciprop() their
confidence intervals, and survey::svychisq() tests the
association, Rao-Scott corrected and referred to the design degrees
of freedom.
Usage
table_categorical_svy(
design,
select = tidyselect::everything(),
by = NULL,
labels = NULL,
levels_keep = NULL,
include_total = TRUE,
drop_na = FALSE,
proportion_ci = FALSE,
ci_method = c("logit", "likelihood", "asin", "beta", "mean", "xlogit", "wilson"),
ci_level = 0.95,
chisq_statistic = c("F", "Chisq", "Wald", "adjWald", "saddlepoint"),
deff = FALSE,
df = NULL,
p_value = NULL,
percent_digits = 1,
p_digits = 3,
decimal_mark = ".",
align = c("decimal", "center", "right"),
output = c("default", "data.frame", "long", "tinytable", "gt", "flextable", "excel",
"clipboard", "word"),
indent_text = " ",
indent_text_excel_clipboard = strrep("Â ", 6),
excel_path = NULL,
excel_sheet = NULL,
clipboard_delim = "\t",
word_path = NULL,
user_na = TRUE,
style = NULL
)
Arguments
design |
A survey design: |
select |
Columns to tabulate, as a tidyselect expression on the design's variables. |
by |
A single grouping column: one column block per level. |
labels |
Named character vector of display labels. |
levels_keep |
Levels to keep, as a character vector (all variables) or a named list (per variable). |
include_total |
Add a |
drop_na |
Drop missing values (default |
proportion_ci |
Add the confidence interval of each percentage. |
ci_method |
Interval method passed to |
ci_level |
Coverage of the interval. |
chisq_statistic |
Statistic for |
deff |
Show the design effect of each percentage: |
df |
Degrees of freedom for the intervals. |
p_value |
Show the p-value column (defaults to |
percent_digits, p_digits, decimal_mark |
Number formatting. |
align |
Numeric-cell alignment: |
output |
One of |
indent_text, indent_text_excel_clipboard |
Level-row indentation, for the console and for the plain-text engines. |
excel_path, excel_sheet, clipboard_delim, word_path |
Output
destinations, as in |
user_na |
Honour declared missing values (see |
style |
A journal style; see |
Value
A spicy_categorical_svy_table: the wide compute frame,
with the display frame and the typed view attached.
output = "data.frame" / "long" returns the compute frame
unclassed – the two tokens are synonyms and return identical
objects.
What the columns are
n is the observed count – the number of rows in the sample, not
an estimated population size. % is the estimated percentage
within its column: without by it is the distribution of the
variable in the population, with by the distribution inside that
domain. The table note gives the sample size and the estimated
population together, because neither alone tells the reader what
they are looking at.
proportion_ci = TRUE adds the interval of each percentage.
ci_method chooses among the seven survey::svyciprop() offers;
the default "logit" is bounded inside 0 to 100, which the Wald
interval ("mean") is not. The percentage itself always comes from
survey::svymean(), so it does not move when ci_method does.
The test
svychisq() with chisq_statistic = "F" (the default) is the
Pearson chi-square with the Rao-Scott second-order correction,
referred to F(ndf, survey::degf(design)). It is survey's own
default and the one Stata's svy: tabulate reports.
It runs on the complete cases of the two variables, and on their
observed levels: a (Missing) row and a declared-but-unobserved
level are descriptive, and neither belongs to the null hypothesis.
The p-value is therefore the same whether drop_na shows those rows
or removes them, and the intervals beside it describe the same
domain – the two families test the same table.
"Chisq" shows the p-value only: survey adjusts the statistic in
the "F" branch and only the p-value in the "Chisq" one, so the
statistic there is not the one the p-value came from. "Wald",
"adjWald" and "saddlepoint" are available; "lincom" and
"wls-score" are refused, the first because its integration is
documented as failing in the far tail (?pchisqsum), the second
because it has no reporting convention here.
Stability
This function is stabilising in the sense ?spicy defines: the
names of its design-specific arguments may still be tightened
before 1.0 – with a NEWS.md entry – but the behaviour does not
change silently. It was experimental through 0.13; the shape of the
table has held across the cycle, so it now moves on the parent
family's clock rather than its own. The numbers themselves are
survey's and do not move with it.
What is absent, and why
weights and rescale (the weighting is the design). correct
(Yates), simulate_p and simulate_B, which have no meaning once
the reference distribution is Rao-Scott's. And the association
measures: Cramer's V, phi, tau-b/c, gamma, Somers' D and lambda
have no established design-based variance, and the intervals
table_categorical() gives them assume simple random sampling. The
design-based measure of association here is the Rao-Scott test in
the p column; for an effect size, model it with
table_regression(survey::svyglm(...)).
See Also
table_categorical() for the data-frame sibling,
table_continuous_svy() for continuous variables.
Examples
data(api, package = "survey")
dclus1 <- survey::svydesign(
id = ~dnum, weights = ~pw, data = apiclus1, fpc = ~fpc
)
table_categorical_svy(dclus1, select = c(stype, awards))
table_categorical_svy(dclus1, select = stype, by = sch.wide)
table_categorical_svy(
dclus1,
select = stype,
proportion_ci = TRUE,
deff = TRUE
)
Continuous summary table
Description
Computes descriptive statistics (mean, SD, min, max, confidence interval of the mean, n) for one or many continuous variables selected with tidyselect syntax.
With by, produces grouped summaries and reports a group-comparison
p-value by default (Welch test; change via test). Additional
inferential output is opt-in: test statistics (statistic) and
effect sizes (effect_size / effect_size_ci). Set p_value = FALSE
to suppress the p-value column. Without by, produces one-way
descriptive summaries.
Multiple output formats are available via output: a printed ASCII
table ("default"), a plain data.frame ("data.frame" or
"long" – synonyms for the underlying long-format data, see
Details), or publication-ready tables ("tinytable", "gt",
"flextable", "excel", "clipboard", "word").
This is the descriptive companion to table_continuous_lm(). The
two functions share their layout, alignment, and reporting precision
so descriptive and model-based analyses of the same data look
uniform side by side – with one exception, documented under
align: only table_continuous() carries align into the excel
output. Use table_continuous_lm() when you need
robust SE, weighted contrasts, fitted means, or covariate
adjustment.
Usage
table_continuous(
data,
select = tidyselect::everything(),
by = NULL,
exclude = NULL,
regex = FALSE,
drop_na = TRUE,
weights = NULL,
rescale = FALSE,
test = c("welch", "student", "nonparametric"),
p_value = NULL,
statistic = FALSE,
show_n = TRUE,
show_columns = NULL,
effect_size = c("none", "auto", "hedges_g", "eta_sq", "r_rb", "epsilon_sq"),
effect_size_ci = FALSE,
smd = FALSE,
ci = TRUE,
labels = NULL,
ci_level = 0.95,
digits = 2,
effect_size_digits = 2,
p_digits = 3,
decimal_mark = ".",
align = c("decimal", "center", "right"),
output = c("default", "data.frame", "long", "tinytable", "gt", "flextable", "excel",
"clipboard", "word"),
excel_path = NULL,
excel_sheet = NULL,
clipboard_delim = "\t",
word_path = NULL,
verbose = FALSE,
user_na = TRUE,
style = NULL
)
Arguments
data |
A |
select |
Columns to include. If |
by |
Optional grouping column. Accepts an unquoted column name
or a single character column name. Coerced to factor for
grouping; non-numeric grouping columns (factor, character,
logical) are supported as-is. Factor levels keep their declared
order; any other |
exclude |
Columns to exclude. Supports tidyselect syntax and character vectors of column names. |
regex |
Logical. If |
drop_na |
Logical. Controls how missing values in the |
weights |
Optional case weights: an unquoted column name, a
character column name, or a numeric vector of length |
rescale |
Logical. If |
test |
Character. Statistical test to use when comparing groups.
One of
Used whenever |
p_value |
Logical or |
statistic |
Logical. If |
show_n |
Logical. If |
show_columns |
Statistics to display, as a character vector of
tokens or a named list of such vectors (one per variable). |
effect_size |
Effect-size measure to include in the rendered outputs. One of:
For backward compatibility, |
effect_size_ci |
Logical. If |
smd |
Logical. If |
ci |
Logical. If |
labels |
An optional named character vector of variable labels.
Names must match column names in |
ci_level |
Confidence level for the mean confidence interval
(default: |
digits |
Number of decimal places for descriptive values and test
statistics (default: |
effect_size_digits |
Number of decimal places for effect-size values
in formatted displays (default: |
p_digits |
Integer >= 1. Number of decimal places used to
render p-values in the |
decimal_mark |
Character used as decimal separator.
Either |
align |
Horizontal alignment of numeric columns in the printed
ASCII table and in the
|
output |
Output format. One of:
|
excel_path |
File path for |
excel_sheet |
Sheet name for |
clipboard_delim |
Delimiter for |
word_path |
File path for |
verbose |
Logical. If |
user_na |
Logical. If |
style |
A journal style: a theme name ( |
Value
Depends on output:
-
"default": the underlyingdata.framecarrying the rendering metadata as attributes (S3 class"spicy_continuous_table"/"spicy_table"). The object is returned visibly, so a baretable_continuous(...)call auto-prints the styled ASCII table at the console whilet <- table_continuous(...)stays silent (printtto display the table). The object can be re-coerced viaas.data.frame.spicy_continuous_table()or piped intobroom::tidy()/broom::glance(). -
"data.frame"/"long": a plaindata.framewith columnsvariable,label,group(whenbyis used),mean,sd,min,max,ci_lower,ci_upper,median,q1,q3,iqr,med_ci_lower,med_ci_upper,n. Every statistic is computed whatevershow_columnsdisplays. Whenbyis used together withp_value = TRUE,statistic = TRUE, oreffect_size != "none", additional columns are appended (populated on the first row of each variable block only):-
test_type– test identifier (e.g.,"welch_t","welch_anova","student_t","anova","wilcoxon","kruskal"). -
statistic,df1,df2,p.value– test results. -
es_type– effect-size identifier ("hedges_g","eta_sq","r_rb", or"epsilon_sq"), wheneffect_size != "none". -
es_value,es_ci_lower,es_ci_upper– effect-size estimate and confidence interval bounds.
A
byframe ALSO carriessmd_typeandsmd_valueunconditionally –NAthroughout whensmd = FALSE– so the schema a pipeline indexes into does not move with an argument (theweighted_nrule).smd_typenames the kernel the value came from:"continuous"here. The two names"data.frame"and"long"are synonyms (the descriptive output is naturally already long). Pick whichever reads better in your code. -
-
"tinytable": atinytableobject. -
"gt": agt_tblobject. -
"flextable": aflextableobject. -
"excel"/"word": writes to disk and returns the file path invisibly. -
"clipboard": copies the table and returns the displaydata.frameinvisibly.
The missing-value disclosure (values excluded from the summaries,
and rows removed for a missing by value under drop_na = TRUE)
travels with the table on every route, not just the console:
"default" prints it under the ASCII table, "tinytable" / "gt" /
"flextable" / "word" carry it as a table note, "excel" writes
it below the body, and
"data.frame" / "long" keep the sentence verbatim in the
missing_note attribute (attr(x, "missing_note"), NULL when
nothing was removed) so a pipeline that renders the numbers itself
can still state what left the table. On the "tinytable" route the
note is set one size down; options(spicy.note_style) governs that
(see table_regression()).
The Excel sheet carries the same title the console prints on its first row; the table itself starts on row 3.
Choosing the statistics
show_columns selects which statistics the table displays. The
tokens, and the column each one produces:
| Token | Column | Statistic |
"m" | M | mean |
"sd" | SD | standard deviation |
"med" | Med | median (stats::median()) |
"iqr" | IQR | interquartile width, Q3 - Q1 |
"med_iqr" | Med [Q1, Q3] | median and the interquartile interval, in one compact column |
"q1" / "q3" | Q1 / Q3 | first / third quartile |
"min" / "max" | Min / Max | extremes |
"ci" | <level>% CI LL / UL | t confidence interval of the mean |
"med_ci" | Med <level>% CI LL / UL | exact confidence interval of the median |
"n" | n | valid observations |
"weighted_n" | Weighted n | sum of weights (requires weights)
|
Quartiles use stats::quantile()'s default type 7. "iqr" is the
width (one number, the rank mirror of SD); "med_iqr" shows the
interval with its bounds. Columns appear in the canonical order of
the table above, whatever order they were written in.
Weights
With weights, every displayed statistic uses the
frequency-expansion convention: for integer weights each
statistic equals its unweighted version computed on the data with
every row repeated w times (rep(x, w)), exactly; with all
weights equal to 1 every statistic equals its unweighted sibling.
The formulas, with W = \sum w_i:
mean:
\sum w_i x_i / W;SD:
\sqrt{\sum w_i (x_i - \bar{x}_w)^2 / (W - 1)};quantiles: type-7 positions on the cumulative-weight scale (the
Hmisc::wtd.quantile()default algorithm);CI of the mean:
\bar{x}_w \pm t_{W-1} \, s_w / \sqrt{W};-
ncounts the rows used;"weighted_n"reportsW.
These are the conventions of Hmisc::wtd.mean() / wtd.var() /
wtd.quantile() (defaults), matrixStats::weightedSd(), and
DescTools::Quantile(), and – for integer weights – of Stata's
[fweight] and SPSS's WEIGHT BY. With rescale = TRUE the
weights are normalised to sum to the number of observations first,
which makes every result invariant to the scale of the weights and
makes the SD equal Stata's [aweight] / survey::svyvar() value
– the reading appropriate for sampling weights. Weighted-quantile
conventions genuinely differ across software (Stata interpolates
nowhere, SAS refuses analytic-weighted quantiles, the survey
package offers twelve rules); spicy states its rule here rather
than leaving it implicit.
Two deliberate refusals: the "med_ci" token (an order-statistic
interval with no weighted version) and, under by, the group tests
and effect sizes – a t-test printed next to weighted descriptives
would silently be unweighted. Set p_value = FALSE for weighted
descriptives by group, or use table_continuous_lm() with
weights for weighted comparisons. Note that
table_continuous_lm()'s residual SD answers a different question
(model-based, precision-weight convention) and is not expected to
match the descriptive SD here.
A named list applies a different selection to each variable, with
.default covering the variables it does not name – the case of a
table where a skewed variable must be reported as a median while the
others keep the mean:
show_columns = list(
mvpa = c("med_iqr", "n"),
sitting = c("med_iqr", "n"),
.default = c("m", "sd", "n")
)
The table's columns are the union of the requested tokens; a cell of a column the variable did not ask for is left blank (structurally empty, not an en dash, which is reserved for an undefined statistic).
The table tests what it shows. A variable displaying a median
without a mean takes the rank-based test – Wilcoxon rank-sum for
two groups, Kruskal-Wallis beyond – and the rank effect size
(rank-biserial r, \varepsilon^2) when effect_size is
"auto". The switch is per variable, so a mixed table carries a
rank test on its median rows and Welch on its mean rows, and the
table note names which test each variable carries. An explicit
test is sovereign: it applies to every variable, with a warning
naming the ones displayed as medians.
"med_ci" is the exact order-statistic (sign-test) confidence
interval: the tightest interval [x_{(k)}, x_{(n-k+1)}] whose
binomial coverage still reaches ci_level. It is distribution-free
and deterministic – no bootstrap, no seed – and its coverage is at
least nominal, the same convention as SAS PROC UNIVARIATE
(CIPCTLDF) and DescTools::MedianCI(method = "exact"). Below
about six observations no interval reaches the requested level; the
cells then show an en dash rather than a false interval.
"ci" is the confidence interval of the mean: requested without
"m" it is dropped with a warning pointing at "med_ci", and
"med_ci" without a displayed median is dropped likewise. When
show_columns is supplied it decides the n and CI columns on its
own, and a contradictory show_n / ci is reported.
Tests
The omnibus test is computed only when by is supplied and at
least two groups remain after dropping NAs, with every group
contributing at least two observations. Choice of test family is
driven by test (see the @param entry for the full dispatch
and the underlying stats:: functions called).
For model-based contrasts (heteroskedasticity-consistent SE,
cluster-robust SE, weighted contrasts, fitted means, covariate
adjustment), use table_continuous_lm().
Effect sizes
See @param effect_size for the dispatch table (canonical
measure for each (test, n_groups) combination) and the
validation rules applied to explicit requests.
Confidence intervals (enabled with effect_size_ci = TRUE) use
noncentral F inversion for \eta^2, the Hedges-Olkin
normal approximation for g, the Fisher z-transform for r,
and percentile bootstrap (2,000 replicates) for
\varepsilon^2. The bootstrap bounds depend on the random
number generator state: call set.seed() before the table for
reproducible \varepsilon^2 intervals (the other three CIs
are closed-form and deterministic).
For Cohen's d, Hays' \omega^2, and Cohen's f^2
(derived from a fitted, possibly weighted lm()), use the
model-based companion table_continuous_lm().
Standardized mean difference
smd = TRUE adds an SMD column with the balance diagnostic of
the Table 1 literature, in Austin's form (Austin 2009, Stat Med
28:3083-3107; Austin 2011, Multivar Behav Res 46:399-424):
\mathrm{SMD} = \frac{\bar{x}_1 - \bar{x}_2}{\sqrt{(s_1^2 +
s_2^2) / 2}}
The denominator is the root mean of the two group variances,
each at n - 1, not the degrees-of-freedom pooled SD. At
equal group sizes those two denominators are the same, so the
SMD is exactly Cohen's d; at unequal sizes they part company (on
a 4-versus-3 split the SMD is -0.51 and d is -0.54).
effect_size = "hedges_g", two columns to the left, is a third
number: g applies the small-sample correction J on top of d,
so it never equals the SMD. At equal group sizes the ratio
g / \mathrm{SMD} is exactly J – 0.80 at n = 3 per group,
0.96 at n = 10 – approaching 1 only as the sample grows. Read
each for what it is; do not recompute one from the other. (The
divergence is nameable upstream:
cobalt::col_w_smd(s.d.denom = "pooled") reproduces this column,
s.d.denom = "hedges" reproduces hedges_g.)
Conventions, all deliberate:
-
Signed, group 1 minus group 2 in the order the table displays the groups – the two groups sit side by side, so a bare magnitude would make the reader re-derive a direction the row already gives. (
tableonepublishes the magnitude;cobaltandarsenalsign it the other way, guessing the second level as "treated".) The threshold in the table note is read on|\mathrm{SMD}|; the column keeps the sign. No conditional formatting: spicy never highlights a threshold. -
No confidence interval and no p-value, ever. The SMD is a descriptive diagnostic; attaching an interval to it reintroduces the test reasoning the balance literature asks the reader to drop. This is not a missing feature.
-
Exactly two groups. A
bywith three or more is refused rather than averaged over pairs: an average has no published reading, and it can sit under the usual threshold while one pair sits well over it. -
Complete cases on the observed groups. A
drop_na = FALSE"(Missing)" group is displayed and never enters the diagnostic, exactly as the test and the effect size behave. -
Independent of
p_value. Turning the SMD on turns nothing else off. The balance-table idiom issmd = TRUE, p_value = FALSE; you have to write both.
Under weights, the means and variances are the weighted ones the
M and SD columns already display – the frequency convention
of the Weights section, from the same producer, so the column
cannot contradict its neighbours. One consequence follows and is
intended: a frequency weight is a number of copies, so the
weighted SMD is not invariant to the scale of the weights
(multiplying every weight by ten moves it, as it moves the SD
column). rescale = TRUE normalises the weights to sum to n,
restores scale invariance, and is the form to use for sampling
weights until the dedicated survey-design functions land.
A cell is an en-dash when the diagnostic applies but cannot be
estimated: both groups constant at different values (an infinite
standardized distance, disclosed by a warning), or a group with
too little data to have a variance (silent – the SD cell beside
it already says so). Two groups constant at the same value are
perfectly balanced and print 0.00.
Display conventions
Decimal alignment, p-value formatting, and required suggested
packages per output engine are documented under @param align,
@param p_digits, and @param output respectively.
Non-numeric columns are silently dropped (set verbose = TRUE to
see which columns were excluded). When a constant column is
passed, its statistics are reported exactly: SD is 0.00 and the
CI degenerates to [m, m]. An en-dash cell appears only when a
statistic is undefined (fewer than two valid observations).
Declared missing values
Survey files imported with haven often carry declared missing
values: codes such as 8 = Don't know or 9 = Refused that the
source file marks as missing while keeping them distinct from a
plain NA. Two kinds of declaration exist: na_values / na_range
metadata on haven::labelled_spss() vectors, and tagged missing
values created by haven::tagged_na() (the Stata .a, .b, ...
convention).
spicy honors the declaration by default (user_na = TRUE):
declared missing values are excluded from every statistic exactly
like NA – valid percentages, means, chi-squared tests,
association measures, row-wise summaries, and group definitions –
but they are not erased from display. freq() lists each observed
declared value as its own row of the Missing block, with its value
label; cross_tab(), table_categorical(), and
table_continuous() disclose the exclusion in the table note
(Declared missing values removed: x (2).); varlist() and
code_book() count them as missing in N_valid / NAs /
N_distinct while still listing the declared codes in Values.
Every function involved offers the same escape hatch: set
user_na = FALSE to ignore the declaration and treat the declared
codes as valid values (the behavior of spicy before 0.13.0).
Tagged missing values are genuine NAs either way; for them,
user_na = FALSE only collapses the per-tag breakdown back into
the regular NA count.
See Also
table_outcome() for the transposed shape – ONE
continuous outcome across the levels of SEVERAL groupings, one
block of rows per grouping. Several outcomes across one grouping
is this function; one outcome across one or more groupings is
that one;
table_continuous_lm() for the model-based companion
(heteroskedasticity-consistent SE, cluster-robust SE, weighted
contrasts, fitted means);
table_categorical() for categorical variables;
freq() for one-way frequency tables;
cross_tab() for two-way cross-tabulations.
Other spicy tables:
table_categorical(),
table_continuous_lm()
Examples
# --- Basic usage ---------------------------------------------------------
# Default: ASCII console table.
table_continuous(
sochealth,
select = c(bmi, wellbeing_score)
)
# Grouped by education (Welch p-value added by default).
table_continuous(
sochealth,
select = c(bmi, wellbeing_score),
by = education
)
# Test statistic alongside the p-value.
table_continuous(
sochealth,
select = c(bmi, wellbeing_score),
by = education,
statistic = TRUE
)
# --- Choosing the statistics --------------------------------------------
# Median and interquartile range instead of mean and SD.
table_continuous(
sochealth,
select = c(bmi, wellbeing_score),
show_columns = c("med_iqr", "n")
)
# Median with its exact (order-statistic) confidence interval.
table_continuous(
sochealth,
select = bmi,
show_columns = c("med", "iqr", "med_ci", "n")
)
# One selection per variable. A skewed variable that a scoring
# protocol requires in median and IQR (the IPAQ case) sits next to
# variables kept in mean and SD; each row is tested the way it is
# displayed, and the note says so.
table_continuous(
sochealth,
select = c(bmi, life_sat_health, wellbeing_score),
by = sex,
show_columns = list(
life_sat_health = c("med_iqr", "n"),
.default = c("m", "sd", "n")
)
)
# --- Effect sizes -------------------------------------------------------
# Auto-selected effect size with confidence interval (Hedges' g for
# binary `by`, eta-squared for k > 2).
table_continuous(
sochealth,
select = wellbeing_score,
by = sex,
effect_size = "auto",
effect_size_ci = TRUE
)
# Explicit effect-size measure.
table_continuous(
sochealth,
select = wellbeing_score,
by = education,
effect_size = "eta_sq",
effect_size_ci = TRUE,
effect_size_digits = 3
)
# --- Selection helpers --------------------------------------------------
# Regex selection.
table_continuous(
sochealth,
select = "^life_sat",
regex = TRUE
)
# Pretty labels keyed by column name.
table_continuous(
sochealth,
select = c(bmi, life_sat_health),
labels = c(
bmi = "Body mass index",
life_sat_health = "Satisfaction with health"
)
)
# --- Output formats -----------------------------------------------------
# The rendered outputs below all wrap the same call:
# table_continuous(sochealth,
# select = c(bmi, wellbeing_score),
# by = sex)
# only `output` changes. Assign each result to a variable -- some
# engines auto-print as a console-friendly text fallback inside
# the `?` help viewer.
# Wide / long data.frame (synonyms): one row per (variable x group).
table_continuous(
sochealth,
select = c(bmi, wellbeing_score),
by = sex,
output = "data.frame"
)
# Rendered HTML / docx objects -- best viewed inside a
# Quarto / R Markdown document or a pkgdown article.
if (requireNamespace("tinytable", quietly = TRUE)) {
tt <- table_continuous(
sochealth, select = c(bmi, wellbeing_score), by = sex,
output = "tinytable"
)
}
if (requireNamespace("gt", quietly = TRUE)) {
tbl <- table_continuous(
sochealth, select = c(bmi, wellbeing_score), by = sex,
output = "gt"
)
}
if (requireNamespace("flextable", quietly = TRUE)) {
ft <- table_continuous(
sochealth, select = c(bmi, wellbeing_score), by = sex,
output = "flextable"
)
}
# Excel and Word: write to a temporary file.
if (requireNamespace("openxlsx2", quietly = TRUE)) {
tmp <- tempfile(fileext = ".xlsx")
table_continuous(
sochealth, select = c(bmi, wellbeing_score), by = sex,
output = "excel", excel_path = tmp
)
unlink(tmp)
}
if (
requireNamespace("flextable", quietly = TRUE) &&
requireNamespace("officer", quietly = TRUE)
) {
tmp <- tempfile(fileext = ".docx")
table_continuous(
sochealth, select = c(bmi, wellbeing_score), by = sex,
output = "word", word_path = tmp
)
unlink(tmp)
}
## Not run:
# Clipboard: writes to the system clipboard.
table_continuous(
sochealth, select = c(bmi, wellbeing_score), by = sex,
output = "clipboard"
)
## End(Not run)
Continuous-outcome linear-model table
Description
Builds publication-ready summary tables from a series of linear models for one or many continuous outcomes selected with tidyselect syntax.
A single focal predictor is supplied with by; each selected numeric
outcome is fit as lm(outcome ~ by, ...), optionally extended with
additive covariates via covariates and case weights via weights.
Categorical by produces model-based estimated marginal means by
level (covariate-adjusted via adjustment when covariates are
present), plus an optional single difference for dichotomous
predictors. Numeric by produces the slope and its confidence
interval.
Inference adapts via vcov: classical OLS, "HC0"-"HC5"
(heteroscedasticity-consistent), "CR0"-"CR3" (cluster-robust,
requires cluster), or "bootstrap" / "jackknife" resampling.
Effect sizes (Cohen's "d", Hedges' "g", Hays' "omega2",
Cohen's "f2") are reported with optional noncentral t / F
confidence intervals via effect_size_ci, and adapt under
covariate adjustment (see effect_size).
Multiple output formats are available via output: a printed ASCII table
("default"), a plain wide data.frame ("data.frame"), a raw long
data.frame ("long"), or rendered outputs ("tinytable", "gt",
"flextable", "excel", "clipboard", "word").
Usage
table_continuous_lm(
data,
select = tidyselect::everything(),
by,
covariates = NULL,
adjustment = c("proportional", "balanced"),
exclude = NULL,
regex = FALSE,
weights = NULL,
vcov = c("classical", "HC0", "HC1", "HC2", "HC3", "HC4", "HC4m", "HC5", "CR0", "CR1",
"CR2", "CR3", "bootstrap", "jackknife"),
cluster = NULL,
boot_n = 1000,
contrast = c("auto", "none"),
statistic = FALSE,
p_value = TRUE,
show_n = TRUE,
show_weighted_n = FALSE,
effect_size = c("none", "f2", "d", "g", "omega2"),
effect_size_ci = FALSE,
r2 = c("r2", "adj_r2", "none"),
ci = TRUE,
labels = NULL,
ci_level = 0.95,
digits = 2,
fit_digits = 2,
effect_size_digits = 2,
p_digits = 3,
decimal_mark = ".",
align = c("decimal", "center", "right"),
output = c("default", "data.frame", "long", "tinytable", "gt", "flextable", "excel",
"clipboard", "word"),
excel_path = NULL,
excel_sheet = NULL,
clipboard_delim = "\t",
word_path = NULL,
verbose = FALSE,
user_na = TRUE,
style = NULL
)
Arguments
data |
A |
select |
Outcome columns to include. If |
by |
A single predictor column. Accepts an unquoted column name or a single character column name. The predictor can be:
Rows with |
covariates |
Optional additive covariates to adjust each per-outcome
linear model for. Accepts a tidyselect expression (e.g.
When non-empty, each model is fitted as v1 supports additive covariates only. Formula syntax with
interactions or transforms ( Rows with |
adjustment |
How the covariate-adjusted estimated marginal
means (the
Both methods reduce to the same linear-contrast formula
|
exclude |
Columns to exclude from |
regex |
Logical. If |
weights |
Optional case weights. Accepts:
Validation: weights must be finite, non-negative, and contain at least
one positive value (otherwise the function errors). Rows with |
vcov |
Variance estimator used for standard errors, confidence intervals, and Wald test statistics. One of:
The |
cluster |
Cluster identifier for cluster-aware variance
estimators. Required when
Rows with |
boot_n |
Integer. Number of bootstrap replicates used when
|
contrast |
Contrast display for categorical predictors. One of:
|
statistic |
Logical. If |
p_value |
Logical. If |
show_n |
Logical. If |
show_weighted_n |
Logical. If |
effect_size |
Character. Effect-size column to include in the wide and rendered outputs. One of:
When Under covariate adjustment (
|
effect_size_ci |
Logical. If |
r2 |
Character. Fit statistic to include in the wide and rendered outputs. One of:
When |
ci |
Logical. If |
labels |
An optional named character vector of outcome labels. Names
must match column names in |
ci_level |
Confidence level for coefficient and model-based mean
intervals (default: |
digits |
Number of decimal places for descriptive values, regression
coefficients, and test statistics (default: |
fit_digits |
Number of decimal places for model-fit columns ( |
effect_size_digits |
Number of decimal places for the effect-size
column ( |
p_digits |
Integer >= 1. Number of decimal places used to render
p-values in the |
decimal_mark |
Character used as decimal separator. Either |
align |
Horizontal alignment of numeric columns in the printed
ASCII table and in the
The |
output |
Output format. One of:
|
excel_path |
File path for |
excel_sheet |
Sheet name for |
clipboard_delim |
Delimiter for |
word_path |
File path for |
verbose |
Logical. If |
user_na |
Logical. If |
style |
A journal style: a theme name ( |
Value
Depends on output:
-
"default": the underlying longdata.framewith class"spicy_continuous_lm_table"/"spicy_table". The object is returned visibly, so a baretable_continuous_lm(...)call auto-prints the styled ASCII table at the console whilet <- table_continuous_lm(...)stays silent (printtto display the table). -
"data.frame": a plain widedata.framewith one row per outcome and numeric columns for means (categoricalby) or slope (numericby), optional contrast and CI, optional test statistic,p, fit statistic (\eqn{R^2}{R^2}or adjusted\eqn{R^2}{R^2}), effect size, optionales_ci_lower/es_ci_upper(wheneffect_size_ci = TRUE),n, andWeighted n. -
"long": a rawdata.framewith one block per outcome and 28 columns covering identification (variable,label,predictor_type,predictor_label,level,reference), fitted means and their CI (emmean,emmean_se,emmean_ci_lower,emmean_ci_upper), contrast or slope estimates and CI (estimate_type,estimate,estimate_se,estimate_ci_lower,estimate_ci_upper), inferential output (test_type,statistic,df1,df2,p.value), effect size with its CI (es_type,es_value,es_ci_lower,es_ci_upper), fit (r2,adj_r2), and sample size (n,weighted_n). -
"tinytable": atinytableobject. -
"gt": agt_tblobject. -
"flextable": aflextableobject. -
"excel"/"word": writes to disk and returns the file path. -
"clipboard": copies the wide table and returns it invisibly.
The Excel sheet carries the same title the console prints on its first row; the table itself starts on row 3, and the note lines sit below the body.
If no numeric outcome columns remain after applying select, exclude,
and regex, the function emits a warning and returns an empty
data.frame() regardless of output.
Model and outputs
table_continuous_lm() is designed for article-style reporting around
a single focal predictor: one model per selected continuous outcome,
fitted as lm(outcome ~ by, ...) and optionally extended with case
weights and additive covariates (lm(outcome ~ by + cov1 + ...)).
For categorical by, the reported means are model-based fitted means
(or covariate-adjusted estimated marginal means; see adjustment) for
each level, and contrasts come from the same fitted linear model. For
an unweighted lm(y ~ factor) with classical variance and no
covariates, the fitted means coincide numerically with empirical
subgroup means; the model-based qualifier matters because (a) under
weights the means become weighted least-squares estimates, (b) their
CIs derive from the model vcov (classical, HC*, CR*,
bootstrap or jackknife), (c) under covariates they become
adjusted marginal means, and (d) tests, p-values and effect sizes
all come from the same fitted model, keeping the table internally
consistent.
Compared with table_continuous(), this function is the model-based
companion: choose it when you want heteroskedasticity-consistent standard
errors (vcov = "HC*"), model fit statistics, or case weights via
lm(..., weights = ...). Because the function exists to report a fitted
model, its inferential output is on by default: p_value = TRUE and
r2 = "r2" are the defaults; set p_value = FALSE or r2 = "none" to
suppress them.
Effect sizes
Effect size is selected explicitly via effect_size (defaults to
"none"). All variants are derived from the same fitted model as the
displayed coefficients, \eqn{R^2}{R^2}, and CIs, so the effect size stays
internally consistent with the rest of the table.
-
"f2": Cohen's\eqn{f^2}{f^2} = \eqn{R^2}{R^2} / (1 - \eqn{R^2}{R^2})(Cohen 1988). Defined for any predictor type. For a single-predictor model,\eqn{f^2}{f^2}is a monotone transform of\eqn{R^2}{R^2}and adds no information beyond it; its primary use is in a priori power analysis (e.g. G*Power). -
"d","g": standardized mean difference (Cohen's d or Hedges' g), defined only whenbyhas exactly two non-empty levels.d = beta_hat / sigma_hatwithsigma_hat = summary(fit)$sigma(the pooled within-group SD for the unweighted two-group case);g = J * dwithJ = 1 - 3 / (4 * df_resid - 1)(Hedges and Olkin 1985). The sign matches the displayedDelta (level2 - level1). For published reports of two-group comparisons, g is the convention recommended by Hedges and Olkin (1985). -
"omega2": Hays'\omega^2, computed from weighted sums of squares as(SS_effect - df_effect * MSE) / (SS_total + MSE)and truncated at 0 for small or null effects (Hays 1963; Olejnik and Algina 2003). Less biased than\eta^2(which equalsR^2in this single-predictor design) and recommended for reporting variance explained in ANOVA-style designs (Olejnik and Algina 2003).
All four effect sizes are point estimates derived from the OLS/WLS fit
and are invariant to vcov: choosing HC* changes the SE, CI, and
test statistic of the contrast but not the standardized magnitude
itself.
Under covariate adjustment (covariates non-empty), "f2" and
"omega2" become the partial f^2 / partial \omega^2 of by, derived
from the partial F restricted to the focal term (the Type-II
test of by after all covariates, equal to stats::drop1() in
this additive setting). "d" and "g" raise a spicy_unsupported error: the pooled
standard deviation has no canonical extension under adjustment, so
Cohen's d and Hedges' g are undefined for adjusted models. See
effect_size for the full dispatch.
Confidence intervals for the effect size are available via
effect_size_ci = TRUE and use the modern noncentral-distribution
inversion approach, the consensus standard in commercial statistical
software (Stata esize / estat esize, SAS PROC TTEST and
PROC GLM EFFECTSIZE 14.2+) and in mainstream R packages
(effectsize, MOTE, TOSTER, effsize):
-
"d","g": noncentral t inversion (Steiger and Fouladi 1997; Goulet-Pelletier and Cousineau 2018). Empirical coverage is nominal across sample sizes (Fitts 2021), unlike the older Hedges-Olkin normal approximation which is biased for small samples. For Hedges' g the bounds inherit the J small-sample correction. -
"omega2","f2": noncentral F inversion (Steiger 2004). Without covariates (model-level effect sizes), bounds are converted from the noncentrality parameter usingomega^2 = ncp / (ncp + N)and\eqn{f^2}{f^2} = ncp / Nrespectively, withN = df1 + df2 + 1(total sample size). Under covariate adjustment (partial effect sizes), the bounds use the partial transforms instead: partial\eqn{f^2}{f^2} = ncp / df2, and for partialomega^2the inversion runs at the F-value equivalent of the omega-squared point estimate with boundsncp / (ncp + df2)– theeffectsize::omega_squared()partial = TRUEconvention, which is narrower than the partial eta-squared CI because omega-squared shrinks the point estimate.
For the weighted case, the CI uses raw (unweighted) group counts and
df.residual(fit) = n - p, consistent with the WLS reporting convention
(DuMouchel and Duncan 1983). For propensity-score balance assessment or
complex-survey designs, dedicated packages (cobalt::bal.tab() for the
Austin and Stuart 2015 formulation; survey for design-based effect
sizes) are more appropriate.
Robust standard errors
When vcov is one of the HC* variants, the standard errors, CIs, and
Wald test statistics use a heteroskedasticity-consistent sandwich
estimator computed via sandwich::vcovHC() (Zeileis 2004), the
canonical R implementation. For a brief guide:
-
"HC0"is the original White (1980) form;"HC1"adds then / (n - p)correction (MacKinnon and White 1985), Stata's, robustdefault. -
"HC2"and"HC3"use leverage-based residual rescalings (MacKinnon and White 1985);"HC3"is thesandwich::vcovHC()default for small to moderate samples (Long and Ervin 2000). -
"HC4"adapts the leverage exponent for influential observations (Cribari-Neto 2004);"HC4m"is a modified-exponent refinement (Cribari-Neto and da Silva 2011);"HC5"is an alternative leverage-adaptive variant (Cribari-Neto, Souza and Vasconcellos 2007).
When observations are not independent (repeated measurements per
individual, students nested in classes, patients in hospitals,
country-year panels), classical and HC* standard errors are biased
downward. Use the CR* variants together with cluster = id_var to
get cluster-robust inference (Liang and Zeger 1986). The
implementation dispatches to clubSandwich::vcovCR() for the
variance and to clubSandwich::coef_test() (single-coefficient,
Satterthwaite t) and clubSandwich::Wald_test() (multi-coefficient
Hotelling-T-squared with Satterthwaite df, "HTZ") for inference.
"CR2" (Bell and McCaffrey 2002; Pustejovsky and Tipton 2018) is the
modern recommended default; it generally produces fractional
Satterthwaite degrees of freedom in df2, which the displayed
t(df) / F(df1, df2) header renders to one decimal. "CR1"
applies the G/(G-1) correction only – Stata's , vce(cluster id)
uses the larger G(N-1)/((G-1)(N-p)) factor with t(G-1) inference
(clubSandwich's "CR1S"; exposed for lm fits in
table_regression(), not here), so "CR1" does not
reproduce Stata. Effect sizes remain invariant
to vcov (including CR*); only the SE, CI, test statistic, and
df2 of the contrast change.
Two resampling-based estimators are also available without adding
any dependency: vcov = "bootstrap" (nonparametric resampling-cases
bootstrap; Davison and Hinkley 1997) and vcov = "jackknife"
(leave-one-out delete-1; Quenouille 1956; MacKinnon and White 1985).
Supplying cluster switches both to their cluster-aware variants
(cluster bootstrap, Cameron, Gelbach and Miller 2008;
leave-one-cluster-out jackknife). The number of bootstrap replicates
is controlled by boot_n (default 1000); replicates that fail to
fit on rank-deficient resamples are dropped, with an explicit warning
if more than half fail. Fewer than 10 valid bootstrap replicates (or
fewer than 2 jackknife leave-outs) raises spicy_resampling_failed
rather than silently reporting a different variance estimator.
Inference for both estimators is
asymptotic (z for single-coefficient contrasts, chi^2(q) for the
multi-coefficient global Wald test on k > 2 categorical
predictors), reflected in the displayed test header. Use the
bootstrap when the residual distribution is non-standard or the
sample is small; use the jackknife as a closed-form, deterministic
alternative.
\eqn{R^2}{R^2}, adjusted \eqn{R^2}{R^2}, and the effect sizes remain ordinary
least-squares (or weighted least-squares) statistics regardless of
vcov.
Weights
When weights is supplied, table_continuous_lm() fits weighted
linear models via lm(..., weights = ...). Means become weighted
least-squares estimates and contrasts and slopes are weighted. The
fit statistics \eqn{R^2}{R^2} and adjusted \eqn{R^2}{R^2}, as well as Hays' omega^2
and Cohen's \eqn{f^2}{f^2}, use the corresponding weighted sums of squares
from the WLS fit. Cohen's d and Hedges' g use the WLS
coefficient and the model's weighted residual standard deviation
(summary(fit)$sigma), which is the standard convention for
case-weighted regression-style reporting (DuMouchel and Duncan
1983); the noncentral t CI for d / g uses the raw (unweighted)
group counts and the residual degrees of freedom of the WLS fit
(n - p). This case-weighted workflow is appropriate for weighted
article tables, but is not a substitute for a full complex-survey
design (see e.g. the survey package), nor for propensity-score
balance assessment under the Austin and Stuart (2015) convention
(see e.g. cobalt::bal.tab()).
The n column always reports the unweighted analytic sample size for
each outcome. When show_weighted_n = TRUE, an additional
Weighted n column reports the sum of case weights in the same
analytic sample.
Display conventions
For dichotomous categorical predictors, the wide outputs report fitted
means in reference-level order and label the contrast column
explicitly as Delta (level2 - level1). For categorical predictors
with more than two levels, no single contrast or contrast CI is shown
in the wide outputs; instead, the table reports level-specific means
plus the overall F test when statistic = TRUE (or F(df1, df2)
when the degrees of freedom are constant across outcomes).
When covariates is non-empty, the printed ASCII table appends an
APA-style footer naming the covariates and the chosen estimand, e.g.
Note. Adjusted for age, education (proportional).
The rendering engines carry that same footer as a table note. On the
"tinytable" route the note is set one size down;
options(spicy.note_style) governs that (see table_regression()).
Optional output engines require the corresponding suggested packages:
-
tinytable for
output = "tinytable" -
gt for
output = "gt" -
flextable for
output = "flextable" -
flextable + officer for
output = "word" -
openxlsx2 for
output = "excel" -
clipr for
output = "clipboard"
Declared missing values
Survey files imported with haven often carry declared missing
values: codes such as 8 = Don't know or 9 = Refused that the
source file marks as missing while keeping them distinct from a
plain NA. Two kinds of declaration exist: na_values / na_range
metadata on haven::labelled_spss() vectors, and tagged missing
values created by haven::tagged_na() (the Stata .a, .b, ...
convention).
spicy honors the declaration by default (user_na = TRUE):
declared missing values are excluded from every statistic exactly
like NA – valid percentages, means, chi-squared tests,
association measures, row-wise summaries, and group definitions –
but they are not erased from display. freq() lists each observed
declared value as its own row of the Missing block, with its value
label; cross_tab(), table_categorical(), and
table_continuous() disclose the exclusion in the table note
(Declared missing values removed: x (2).); varlist() and
code_book() count them as missing in N_valid / NAs /
N_distinct while still listing the declared codes in Values.
Every function involved offers the same escape hatch: set
user_na = FALSE to ignore the declaration and treat the declared
codes as valid values (the behavior of spicy before 0.13.0).
Tagged missing values are genuine NAs either way; for them,
user_na = FALSE only collapses the per-tag breakdown back into
the regular NA count.
References
Austin, P. C., & Stuart, E. A. (2015). Moving towards best practice when using inverse probability of treatment weighting (IPTW) using the propensity score to estimate causal treatment effects in observational studies. Statistics in Medicine, 34(28), 3661–3679. doi:10.1002/sim.6607
Bell, R. M., & McCaffrey, D. F. (2002). Bias reduction in standard errors for linear regression with multi-stage samples. Survey Methodology, 28(2), 169–181.
Cameron, A. C., Gelbach, J. B., & Miller, D. L. (2008). Bootstrap-based improvements for inference with clustered errors. Review of Economics and Statistics, 90(3), 414–427. doi:10.1162/rest.90.3.414
Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Hillsdale, NJ: Lawrence Erlbaum.
Cribari-Neto, F. (2004). Asymptotic inference under heteroskedasticity of unknown form. Computational Statistics & Data Analysis, 45(2), 215–233. doi:10.1016/S0167-9473(02)00366-3
Cribari-Neto, F., Souza, T. C., & Vasconcellos, K. L. P. (2007). Inference under heteroskedasticity and leveraged data. Communications in Statistics – Theory and Methods, 36(10), 1877–1888. doi:10.1080/03610920601126589
Cribari-Neto, F., & da Silva, W. B. (2011). A new heteroskedasticity-consistent covariance matrix estimator for the linear regression model. AStA Advances in Statistical Analysis, 95(2), 129–146. doi:10.1007/s10182-010-0141-2
Davison, A. C., & Hinkley, D. V. (1997). Bootstrap Methods and Their Application. Cambridge: Cambridge University Press. doi:10.1017/CBO9780511802843
DuMouchel, W. H., & Duncan, G. J. (1983). Using sample survey weights in multiple regression analyses of stratified samples. Journal of the American Statistical Association, 78(383), 535–543. doi:10.1080/01621459.1983.10478006
Fitts, D. A. (2021). Expected and empirical coverages of different methods for generating noncentral t confidence intervals for a standardized mean difference. Behavior Research Methods, 53(6), 2412–2429. doi:10.3758/s13428-021-01550-4
Goulet-Pelletier, J.-C., & Cousineau, D. (2018). A review of effect sizes and their confidence intervals, Part I: The Cohen's d family. The Quantitative Methods for Psychology, 14(4), 242–265. doi:10.20982/tqmp.14.4.p242
Hays, W. L. (1963). Statistics for Psychologists. New York: Holt, Rinehart and Winston.
Hedges, L. V., & Olkin, I. (1985). Statistical Methods for Meta-Analysis. Orlando, FL: Academic Press.
Long, J. S., & Ervin, L. H. (2000). Using heteroscedasticity consistent standard errors in the linear regression model. The American Statistician, 54(3), 217–224. doi:10.1080/00031305.2000.10474549
Liang, K.-Y., & Zeger, S. L. (1986). Longitudinal data analysis using generalized linear models. Biometrika, 73(1), 13–22. doi:10.1093/biomet/73.1.13
MacKinnon, J. G., & White, H. (1985). Some heteroskedasticity-consistent covariance matrix estimators with improved finite sample properties. Journal of Econometrics, 29(3), 305–325. doi:10.1016/0304-4076(85)90158-7
Olejnik, S., & Algina, J. (2003). Generalized eta and omega squared statistics: Measures of effect size for some common research designs. Psychological Methods, 8(4), 434–447. doi:10.1037/1082-989X.8.4.434
Pustejovsky, J. E., & Tipton, E. (2018). Small-sample methods for cluster-robust variance estimation and hypothesis testing in fixed effects models. Journal of Business & Economic Statistics, 36(4), 672–683. doi:10.1080/07350015.2016.1247004
Quenouille, M. H. (1956). Notes on bias in estimation. Biometrika, 43(3/4), 353–360. doi:10.1093/biomet/43.3-4.353
Steiger, J. H. (2004). Beyond the F test: Effect size confidence intervals and tests of close fit in the analysis of variance and contrast analysis. Psychological Methods, 9(2), 164–182. doi:10.1037/1082-989X.9.2.164
Steiger, J. H., & Fouladi, R. T. (1997). Noncentrality interval estimation and the evaluation of statistical models. In L. L. Harlow, S. A. Mulaik, & J. H. Steiger (Eds.), What if there were no significance tests? (pp. 221–257). Mahwah, NJ: Lawrence Erlbaum.
White, H. (1980). A heteroskedasticity-consistent covariance matrix estimator and a direct test for heteroskedasticity. Econometrica, 48(4), 817–838. doi:10.2307/1912934
Zeileis, A. (2004). Econometric computing with HC and HAC covariance matrix estimators. Journal of Statistical Software, 11(10), 1–17. doi:10.18637/jss.v011.i10
See Also
table_continuous(), table_categorical().
For broader workflows on the same statistical building blocks:
sandwich::vcovHC() (the canonical R implementation of the HC*
sandwich estimators, used internally for vcov = "HC*");
clubSandwich::vcovCR(), clubSandwich::coef_test() and
clubSandwich::Wald_test() (the canonical R implementation of
cluster-robust variance and Satterthwaite-style inference, used
internally for vcov = "CR*"); effectsize::cohens_d(),
effectsize::hedges_g(), and effectsize::omega_squared()
(alternative effect-size computations and CIs); cobalt::bal.tab()
for propensity-score covariate balance with weighted standardized
mean differences (Austin and Stuart 2015); the
survey package for
design-based inference on complex-survey samples.
Other spicy tables:
table_categorical(),
table_continuous()
Examples
# --- Basic usage ---------------------------------------------------------
# Default: ASCII table with model-based means, p, and \eqn{R^2}{R^2}.
table_continuous_lm(
sochealth,
select = c(wellbeing_score, bmi),
by = sex
)
# --- Effect sizes -------------------------------------------------------
# Cohen's d (binary by required).
table_continuous_lm(
sochealth,
select = c(wellbeing_score, bmi),
by = sex,
effect_size = "d"
)
# Hedges' g with weighted analysis and weighted n column.
table_continuous_lm(
sochealth,
select = c(wellbeing_score, bmi),
by = sex,
weights = weight,
statistic = TRUE,
effect_size = "g",
show_weighted_n = TRUE
)
# Hedges' g with noncentral t confidence interval (bracket notation).
table_continuous_lm(
sochealth,
select = c(wellbeing_score, bmi),
by = sex,
effect_size = "g",
effect_size_ci = TRUE
)
# Cohen's \eqn{f^2}{f^2} alongside \eqn{R^2}{R^2} (familiar power-analysis effect size).
table_continuous_lm(
sochealth,
select = c(wellbeing_score, bmi),
by = sex,
effect_size = "f2"
)
# Hays' omega-squared for a 3-level predictor (d / g would error here).
table_continuous_lm(
sochealth,
select = c(wellbeing_score, bmi),
by = education,
effect_size = "omega2"
)
# --- Robust SE for a numeric predictor ----------------------------------
# HC3 standard errors for the slope of a continuous predictor.
table_continuous_lm(
sochealth,
select = c(wellbeing_score, bmi),
by = age,
vcov = "HC3",
ci = FALSE
)
# Cluster-robust SE for repeated-measures data: the `sleep` dataset
# has 10 subjects measured twice (one observation per group).
if (requireNamespace("clubSandwich", quietly = TRUE)) {
table_continuous_lm(
sleep,
select = extra,
by = group,
cluster = ID,
vcov = "CR2"
)
}
# --- Covariate adjustment ----------------------------------------------
# Adjust the comparison of `wellbeing_score` and `bmi` by `sex` for `age`
# and `education`. The footer surfaces the adjustment estimand
# ("proportional" by default = G-computation, matching Stata `margins`).
table_continuous_lm(
sochealth,
select = c(wellbeing_score, bmi),
by = sex,
covariates = c(age, education),
vcov = "HC3"
)
# Same model with the emmeans / SPSS UNIANOVA convention (equal-weight
# marginal means on a synthetic covariate grid).
table_continuous_lm(
sochealth,
select = c(wellbeing_score, bmi),
by = sex,
covariates = c(age, education),
adjustment = "balanced",
vcov = "HC3"
)
# Effect sizes adjust automatically: f2 / omega2 become partial
# effect sizes via the partial F restricted to the focal `by`.
# d / g are undefined under adjustment and raise spicy_unsupported.
table_continuous_lm(
sochealth,
select = c(wellbeing_score, bmi),
by = sex,
covariates = c(age, education),
effect_size = "f2",
effect_size_ci = TRUE
)
# --- Article-style polish -----------------------------------------------
# Pretty outcome labels and adjusted \eqn{R^2}{R^2}.
table_continuous_lm(
sochealth,
select = c(wellbeing_score, bmi),
by = sex,
labels = c(
wellbeing_score = "WHO-5 wellbeing (0-100)",
bmi = "Body-mass index (kg/m^2)"
),
r2 = "adj_r2"
)
# European decimal comma.
table_continuous_lm(
sochealth,
select = c(wellbeing_score, bmi),
by = sex,
decimal_mark = ","
)
# Regex selection of all columns starting with "life_sat".
table_continuous_lm(
sochealth,
select = "^life_sat",
by = sex,
regex = TRUE
)
# --- Output formats -----------------------------------------------------
# The rendered outputs below all wrap the same call:
# table_continuous_lm(sochealth,
# select = c(wellbeing_score, bmi),
# by = sex)
# only `output` changes. Assign to a variable to avoid the
# console-friendly text fallback that some engines fall back to
# when printed directly in `?` help.
# Wide data.frame (one row per outcome).
table_continuous_lm(
sochealth,
select = c(wellbeing_score, bmi),
by = sex,
output = "data.frame"
)
# Raw long data.frame (one block per outcome).
table_continuous_lm(
sochealth,
select = c(wellbeing_score, bmi),
by = sex,
output = "long"
)
# Rendered HTML / docx objects -- best viewed inside a
# Quarto / R Markdown document or a pkgdown article.
if (requireNamespace("tinytable", quietly = TRUE)) {
tt <- table_continuous_lm(
sochealth, select = c(wellbeing_score, bmi), by = sex,
output = "tinytable"
)
}
if (requireNamespace("gt", quietly = TRUE)) {
tbl <- table_continuous_lm(
sochealth, select = c(wellbeing_score, bmi), by = sex,
output = "gt"
)
}
if (requireNamespace("flextable", quietly = TRUE)) {
ft <- table_continuous_lm(
sochealth, select = c(wellbeing_score, bmi), by = sex,
output = "flextable"
)
}
# Excel and Word: write to a temporary file.
if (requireNamespace("openxlsx2", quietly = TRUE)) {
tmp <- tempfile(fileext = ".xlsx")
table_continuous_lm(
sochealth, select = c(wellbeing_score, bmi), by = sex,
output = "excel", excel_path = tmp
)
unlink(tmp)
}
if (
requireNamespace("flextable", quietly = TRUE) &&
requireNamespace("officer", quietly = TRUE)
) {
tmp <- tempfile(fileext = ".docx")
table_continuous_lm(
sochealth, select = c(wellbeing_score, bmi), by = sex,
output = "word", word_path = tmp
)
unlink(tmp)
}
## Not run:
# Clipboard: writes to the system clipboard.
table_continuous_lm(
sochealth, select = c(wellbeing_score, bmi), by = sex,
output = "clipboard"
)
## End(Not run)
Descriptive statistics from a survey design
Description
The design twin of table_continuous(): the same table of means,
standard deviations, intervals and counts, computed from a
survey::svydesign() or survey::as.svrepdesign() object instead
of a data frame.
Not one statistic is computed here. survey::svymean() gives the
mean, its standard error and its design effect, survey::svyvar()
the standard deviation, survey::svyquantile() the quantiles, and
survey::svyttest() / survey::regTermTest() /
survey::svyranktest() the group comparison. Every interval and
every test is referred to survey::degf(design).
Usage
table_continuous_svy(
design,
select = tidyselect::everything(),
by = NULL,
exclude = NULL,
regex = FALSE,
drop_na = TRUE,
deff = FALSE,
qrule = "math",
df = NULL,
test = c("welch", "student", "nonparametric"),
p_value = NULL,
statistic = FALSE,
show_n = TRUE,
show_columns = NULL,
ci = TRUE,
labels = NULL,
ci_level = 0.95,
digits = 2,
p_digits = 3,
decimal_mark = ".",
align = c("decimal", "center", "right"),
output = c("default", "data.frame", "long", "tinytable", "gt", "flextable", "excel",
"clipboard", "word"),
excel_path = NULL,
excel_sheet = NULL,
clipboard_delim = "\t",
word_path = NULL,
verbose = FALSE,
user_na = TRUE,
style = NULL
)
Arguments
design |
A survey design: |
select |
Columns to summarize, as a tidyselect expression on the design's variables. |
by |
A single grouping column. One domain per level. |
exclude |
Columns to drop from |
regex |
Treat |
drop_na |
Drop observations with a missing |
deff |
Show the design effect: |
qrule |
Quantile rule: |
df |
Degrees of freedom for the intervals. |
test |
Group comparison: |
p_value |
Show the p-value column (defaults to |
statistic |
Show the test-statistic column. |
show_n |
Show the count column. |
show_columns |
Character vector of statistic tokens; |
ci, ci_level |
The mean's confidence interval and its level. |
labels |
Named character vector of display labels. |
digits, p_digits, decimal_mark |
Number formatting. |
align |
Numeric-cell alignment: |
output |
One of |
excel_path, excel_sheet, clipboard_delim, word_path |
Output
destinations, as in |
verbose |
Report the columns skipped as non-numeric. |
user_na |
Honour declared missing values (see |
style |
A journal style; see |
Value
A spicy_continuous_svy_table: the compute frame, with the
display frame and the typed view attached. output = "data.frame"
/ "long" returns the compute frame unclassed – the two tokens
are synonyms and return identical objects.
Which function do I need?
A data frame with a column of weights is table_continuous()
(weights = ). A survey design object – strata, clusters, finite
population correction, calibration, replicate weights – is this
function. Passing one to the other errors with the name of the
right one; there is no silent coercion, because the design-based
standard errors, degrees of freedom and tests cannot be recovered
from the weights alone.
Two conventions, one bridge
table_continuous(weights = ) implements the frequency-expansion
convention: a weight is a number of copies, and SD has denominator
sum(w) - 1. This function implements the sampling-weight
convention: a weight is a number of units represented, and SD is
sqrt(survey::svyvar()), whose denominator is n - 1 on weights
normalised to sum to n. These are two estimands, not two
approximations of one.
rescale = TRUE is the bridge, and it is an identity rather than a
coincidence. Writing w' = w * n / sum(w), the rescaled weighted
variance is
sum(w' (x - xbar)^2) / (sum(w') - 1) = n / (n - 1) * sum(w (x - xbar)^2) / sum(w),
which is what survey::svyvar() computes. So on a design that
declares nothing but weights, table_continuous(weights = w, rescale = TRUE) and this function return the same mean and the same
standard deviation. The default rescale = FALSE does not, and that
is the estimand boundary, not a bug.
The mean is continuous across both regimes: sum(w x) / sum(w) does
not move when the weights are rescaled.
Choosing the statistics
show_columns takes the tokens of table_continuous() with two
additions and one removal:
-
"se"– the design-based standard error of the mean; -
"deff"– the design effect (requiresdeff = TRUE); -
"med_ci"is refused. The exact interval of the sibling inverts a binomial sign test on independent observations, which a clustered or stratified sample is not.
Quantiles
qrule = "math" is the default and estimates inf{x : F(x) >= p},
the quantile of the population. qrule = "spicy" switches to the
type-7 interpolation table_continuous() uses, for a reader who
needs the two tables to agree cell for cell; any other value –
including a function – is handed to survey::svyquantile()
untouched. The note always says which rule produced the numbers.
Groups and degrees of freedom
by = cuts one domain per group with [ on the design. survey
recomputes the degrees of freedom on the primary sampling units and
strata each domain retains, so a grouped table generally carries a
different df per row; the note gives the span when they differ.
A group with a missing value is a domain like any other:
drop_na = FALSE gives it a (Missing) row, with its own degrees
of freedom. A domain reduced to one primary sampling unit has none,
and its interval shows the undefined dash rather than an interval
built on qt(p, df = 0).
The comparison is a single test on the whole design, not a set of
pairwise ones: survey::svyttest() with two observed groups,
survey::regTermTest() on survey::svyglm() with three or more,
or survey::svyranktest() under test = "nonparametric". Under a
design the Welch / Student distinction does not exist – the
variance is the design's – so test = "student" warns and behaves
like "welch".
Stability
This function is stabilising in the sense ?spicy defines: the
names of its design-specific arguments may still be tightened
before 1.0 – with a NEWS.md entry – but the behaviour does not
change silently. It was experimental through 0.13; the shape of the
table has held across the cycle, so it now moves on the parent
family's clock rather than its own. The numbers themselves are
survey's and do not move with it.
What is absent, and why
weights and rescale (the weighting is the design), effect_size
and smd (no established design-based variance), and data.
See Also
table_continuous() for the data-frame sibling,
table_categorical_svy() for categorical variables,
table_regression() on a survey::svyglm() fit for a model.
Examples
data(api, package = "survey")
dclus1 <- survey::svydesign(
id = ~dnum, weights = ~pw, data = apiclus1, fpc = ~fpc
)
table_continuous_svy(dclus1, select = c(api00, api99))
table_continuous_svy(dclus1, select = api00, by = stype)
table_continuous_svy(
dclus1,
select = api00,
show_columns = c("m", "se", "ci", "deff", "n"),
deff = TRUE
)
Describe one continuous outcome across several groupings
Description
Summarises one continuous outcome across the levels of several
categorical variables, one block of rows per variable. It is the
inverse layout of table_continuous(), which puts several outcomes
in rows and one grouping in columns.
Each block reports the outcome's statistics level by level, plus its
own group comparison on the block's header row, and an Overall row
gives the marginal summary of the whole analytic sample.
Usage
table_outcome(
data,
outcome,
select,
labels = NULL,
overall = TRUE,
drop_na = FALSE,
weights = NULL,
rescale = FALSE,
test = c("welch", "student", "nonparametric"),
p_value = NULL,
statistic = FALSE,
show_n = TRUE,
show_columns = NULL,
effect_size = c("none", "auto", "hedges_g", "eta_sq", "r_rb", "epsilon_sq"),
effect_size_ci = FALSE,
ci = TRUE,
ci_level = 0.95,
digits = 2,
effect_size_digits = 2,
p_digits = 3,
decimal_mark = ".",
align = c("decimal", "center", "right"),
output = c("default", "data.frame", "long", "tinytable", "gt", "flextable", "excel",
"clipboard", "word"),
indent_text = " ",
indent_text_excel_clipboard = strrep("Â ", 6),
excel_path = NULL,
excel_sheet = NULL,
clipboard_delim = "\t",
word_path = NULL,
user_na = TRUE,
style = NULL,
by
)
Arguments
data |
A data frame. |
outcome |
The continuous outcome, unquoted or as a string. Exactly one column. |
select |
The grouping characteristics to describe the outcome
across, as a tidyselect expression or a character vector of column
names. One block of rows per variable, in the order given. As
everywhere in the family, |
labels |
Named character vector of display labels, for the
outcome and for the |
overall |
Show the marginal |
drop_na |
Drop rows with a missing |
weights, rescale |
Frequency weights and whether to rescale them
to sum to the sample size, as in |
test |
Group comparison for every block: |
p_value |
Show the p-value column (default |
statistic |
Show the test statistic column. |
show_n |
Show the count column. |
show_columns |
Character vector of statistic tokens; |
effect_size, effect_size_ci |
Effect size per block and its
confidence interval, as in |
ci, ci_level |
The mean's confidence interval and its level. |
digits, effect_size_digits, p_digits, decimal_mark |
Number formatting. |
align |
Numeric-cell alignment: |
output |
One of |
indent_text, indent_text_excel_clipboard |
Level-row indentation, for the console and for the plain-text engines. |
excel_path, excel_sheet, clipboard_delim, word_path |
Output
destinations, as in |
user_na |
Honour declared missing values (see |
style |
A journal style; see |
by |
Defunct. The grouping characteristics are selected with
|
Value
A spicy_outcome_table: the compute frame, with the display
frame and the typed view attached. output = "data.frame" /
"long" returns the compute frame unclassed – the two tokens are
synonyms and return identical objects.
Which shape do I need?
Several continuous variables across one grouping is
table_continuous() (select = , by = ). One continuous variable
across one or several groupings is this function. A single select is
legitimate here – it is the natural way in when you know more
groupings are coming – but with several outcomes and one grouping,
the sibling is the table you want.
Choosing the statistics
show_columns takes the same tokens as table_continuous(), with
the same meanings; see the show_columns section of
?table_continuous for the vocabulary. Only the character-vector
form is accepted here: there is one outcome, so a per-variable list
would name nothing.
Weights
weights applies the frequency-expansion convention of the family:
all weights 1 reproduces the unweighted table, and integer weights
reproduce the STATISTICS of the data duplicated that many times –
n stays the raw count of rows that carried the weights. Rows
with a missing or zero weight leave the analytic sample; the note
counts the missing ones.
rescale is the switch between the two readings of a weight: the
frequency reading above, and the sampling-weight reading, where the
weights are normalised to sum to the sample size. See the Weights
section of table_continuous() for the choice in full.
rescale = TRUE normalises the weights over the outcome's whole
surviving sample, once, never per level – a per-level rescale
would destroy the relative weights across levels, which is the
entire information a sampling weight carries into this table. The
means are unchanged by it; the standard deviations move, because
their denominator is sum(w) - 1.
A weighted table refuses the group comparison. The estimates and
their interval have no weighted version here, and a p-value or an
effect size silently computed unweighted beside weighted
descriptives is the one thing that must not happen: set p_value = FALSE (and statistic = FALSE, effect_size = "none"), or use
table_continuous_lm() for a weighted comparison. The
order-statistic median interval is refused for the same reason.
Blocks and the group comparison
Every block is a separate one-way comparison of the outcome across the levels of that variable. Nothing in this table adjusts one block for another, and the table note says so. Read the blocks as a set of bivariate descriptions, not as a model.
Each block chooses its test independently: with two observed levels
test = "welch" is the Welch t-test, with three or more it is the
Welch one-way ANOVA, and test = "nonparametric" is the
Wilcoxon rank-sum or the Kruskal-Wallis test on the same rule. A
block with fewer than two observed levels, or with a level holding a
single observation, is not tested; its statistics stay empty and the
other blocks are unaffected.
The Overall row
overall = TRUE puts the marginal summary of the whole analytic
sample on the first row. Under the default drop_na = FALSE the
levels of every block partition that sample – the (Missing)
display level included – so each block's counts add up to the
Overall count exactly, which is what makes it a usable
denominator.
The row reads Overall, not Total, and the distinction is
deliberate. Total is the word of a COUNT margin: the column of
table_categorical() where frequencies add up. This row is the
whole analytic sample, where a mean is recomputed over every
observation and nothing is added. A mean is not a total.
Choosing the select columns
The canonical form is select = where(is.factor), or an explicit
enumeration. Negation (select = -c(x, y)) is not recommended: it
sweeps in every remaining column, and a numeric one becomes a block
with one LEVEL per distinct value, in order of first appearance. A
variable producing more than 20 levels raises a warning for that
reason – an arbitrary threshold, but a sixty-row block where a
reader expects a handful of categories is not a table.
A haven_labelled column used as select shows its numeric CODES, not
its value labels, as it does in table_continuous(). Convert it
first (haven::as_factor()) to get the labels in the stub.
The table note
One note sits under the table and states what left the analytic
sample, which group comparison ran in each block, what the displayed
columns mean, and how the blocks and the Overall row are to be
read. The rendering engines carry the same sentence as a table note.
On the "tinytable" route it is set one size down;
options(spicy.note_style) governs that (see table_regression()).
See Also
table_continuous() for the transposed shape,
table_categorical() for categorical outcomes.
Examples
table_outcome(sochealth, bmi, select = c(sex, smoking))
table_outcome(sochealth, wellbeing_score, select = where(is.factor))
Regression coefficient summary table
Description
Publication-ready coefficient table from one or more fitted
lm / glm models. Supports standardised coefficients
(\beta), average marginal effects (AME), partial
effect sizes (f^2 / \eta^2 /
\omega^2 for lm; partial
\chi^2 for glm), pseudo-R^2
(glm), and a full vocabulary of variance estimators
(classical / HC* / cluster-robust with Satterthwaite-corrected
df / bootstrap / jackknife). glm covers binomial / poisson /
Gamma / inverse.gaussian / quasi families with any link.
Usage
table_regression(
models,
vcov = "classical",
cluster = NULL,
ci_level = 0.95,
ci_method = c("wald", "profile", "boot_percentile", "hdi"),
boot_n = 1000L,
tau = NULL,
at_time = NULL,
standardized = c("none", "refit", "posthoc", "basic", "smart", "pseudo"),
exponentiate = FALSE,
p_adjust = "none",
show_columns = NULL,
keep = NULL,
drop = NULL,
show_intercept = TRUE,
show_thresholds = TRUE,
show_components = TRUE,
intercept_position = c("first", "last"),
factor_layout = c("grouped", "flat"),
reference_style = c("row", "annotation", "footer", "none"),
reference_label = "(ref.)",
show_fit_stats = NULL,
fit_stats_layout = c("first_col", "merged"),
show_re = TRUE,
re_scale = c("sd", "variance"),
re_columns = c("est", "se", "ci"),
re_test = c("none", "lrt", "rlrt"),
re_ci = c("wald", "profile"),
model_labels = NULL,
outcome_labels = NULL,
stars = FALSE,
nested = FALSE,
digits = 2L,
p_digits = 3L,
effect_size_digits = 2L,
fit_digits = 2L,
ic_digits = 1L,
decimal_mark = ".",
align = c("decimal", "center", "right"),
padding = 0L,
labels = NULL,
title = NULL,
note = NULL,
output = c("default", "data.frame", "long", "gt", "flextable", "tinytable", "excel",
"clipboard", "word"),
excel_path = NULL,
excel_sheet = NULL,
clipboard_delim = "\t",
word_path = NULL,
word_template = NULL,
style = NULL
)
Arguments
models |
A fitted model object, or a list of such fits
(named or unnamed; classes may be mixed). Single fits are
auto-promoted to a 1-element list. A broad set of model classes
is supported – linear / generalised linear ( |
vcov |
Variance-covariance estimator: |
cluster |
Cluster identifier for cluster-robust variance
(used when
For multi-model use, pass a list of one form per model
(mix-and-match allowed). A bare unquoted name
( |
ci_level |
Confidence level for all reported CIs (B, |
ci_method |
CI construction. |
boot_n |
Number of bootstrap replicates when
|
tau |
RMST horizon: the |
at_time |
Landmark time for the |
standardized |
Standardisation method for the |
exponentiate |
Logical. When |
p_adjust |
Multiple-comparison adjustment method applied
to the family of estimated coefficient p-values within each
model (intercept and reference rows excluded). One of
|
show_columns |
Character vector of tokens selecting the
per-coefficient columns and their display order. Accepts
atomic tokens ( |
keep |
Character vector of regexes. Only coefficient rows
whose term name (as in |
drop |
Character vector of regexes. Coefficient rows
matching any pattern are removed. Mutually exclusive with
|
show_intercept |
Whether to display the intercept row.
Default |
show_thresholds |
For ordinal cumulative-link models
( |
show_components |
For models with secondary components,
whether to display them as labelled subordinate blocks of rows
below the primary (count / conditional / location) coefficients.
Default
Component rows carry full Wald inference (B / SE / z / p / CI),
join the |
intercept_position |
Where to place the intercept when
shown. |
factor_layout |
Layout of factor predictors. Applies to
any categorical predictor –
|
reference_style |
Rendering of factor reference levels. Four modes, distinguishing WHERE the reference information is exposed (in a row, inline, in the footer, or nowhere):
Ordered factors with AME: under R's default |
reference_label |
Suffix shown after the reference level in
|
show_fit_stats |
Character vector of tokens for the
model-level rows below the coefficients; row order follows
token order.
Under |
fit_stats_layout |
Layout of the fit-stat values (
Cell merging is supported by
Decimal alignment of every numeric column is preserved in
both modes: the |
show_re |
Logical. |
re_scale |
One of
Correlation rows ( |
re_columns |
Character vector. Subset of
Note. Under the default For |
re_test |
One of
The test statistic and df stay out of the displayed t/z column
(they are chi-square-scale, not t/z) but are carried in
|
re_ci |
One of
|
model_labels |
Per-model labels used as the column-group
spanner above each model's sub-columns (console + gt /
flextable / tinytable / Excel / Word renderers). |
outcome_labels |
Optional Outcome body row override.
|
stars |
Significance asterisks. |
nested |
Whether to inject pairwise change-statistic rows
for adjacent models (M2 vs M1, M3 vs M2, ...). |
digits |
Decimal places for general numeric tokens
( |
p_digits |
Decimal places for p-values ( |
effect_size_digits |
Decimals for per-coefficient effect
sizes ( |
fit_digits |
Decimals for variance-explained / model-level
effect-size fit stats ( |
ic_digits |
Decimals for information criteria ( |
decimal_mark |
Decimal mark used in numeric display.
|
align |
Numeric column alignment.
|
padding |
Non-negative integer giving the extra characters
added to each data column's auto-computed width when the
default |
labels |
Named character vector overriding per-coefficient
row labels. Names are coefficient term names (from
The names are model identifiers, not display text: they are the
term and coefficient names R itself produces, and they keep that
form regardless of the language the table is rendered in. A name
that is neither a term label nor a coefficient name is rejected,
so the header of a subordinate block – With several models, a name is checked against the union of their terms: a key naming a term only one model has is legal, and the label lands on the rows where that term exists. |
title, note |
Override or suppress the auto-built caption / methodological footer. Three modes per argument:
Validation messages, the spanner row, and the in-body change- stat rows are not affected – they belong to the table structure, not to the banner. |
output |
Output type. |
excel_path |
File path for |
excel_sheet |
Sheet name when writing to Excel. |
clipboard_delim |
Field delimiter for
Paste behaviour by target:
|
word_path |
File path for R Markdown / Quarto: for embedded use, prefer |
word_template |
Optional path to a custom .docx file used as the
template for Customising the caption appearance: the table caption is
tagged with the Word named style |
style |
A journal style: a theme name ( |
Value
A spicy_regression_table object (a data.frame
subclass with classes c("spicy_regression_table", "spicy_table", "data.frame")) when output = "default".
The result carries rendering attributes (title, note,
align, padding) and provenance attributes (outcome,
model_ids) consumed by the print method and the broom
methods. For other output values, returns the
format-specific object (gt_tbl, flextable, tinytable,
data.frame, tbl_df, or invisible(x) for side-effect
outputs).
Vocabulary tokens
Two vector arguments – show_columns and show_fit_stats –
accept named tokens that select what to display and in
what order. All tokens are lowercase
(snake_case for compound tokens). Group tokens
("all_b", "all_ame", ...) expand to a fixed vector of
atomic tokens; see show_columns below.
show_columns – per-coefficient columns
Each token = one displayed column.
Coefficient family:
"b","beta"(standardised),"se","ci","t","p".Marginal effects:
"ame","ame_se","ame_ci","ame_p"."p"always refers to the B-coefficient p-value; for the AME-specific p-value use"ame_p".Partial effect sizes –
lmonly:"partial_f2","partial_eta2","partial_omega2", each with a paired_cicompanion ("partial_f2_ci", ...).Partial effect size –
glmand mixed-effects models:"partial_chi2", the term-level Type-II partial chi-square. Each term is tested under the marginality-respecting hypothesis that excludes any higher-order interaction containing it (thecar::Anova(type = 2)convention), so the value does not depend on the factor coding. Forglmthe statistic is the likelihood-ratio chi-square (Long & Freese 2014 Sections 3.2.2, 3.2.4); forlmer/glmer/glmmTMB/nlme::lmeit is the Wald chi-square built from the same Type-II hypothesis – a deliberate departure from the Type-III default of SAS PROC MIXED andlmerTest(the two conventions coincide for models without interactions). Rendered asvalue (df)to disambiguate factor terms (k-1 df) from numeric terms (1 df).Variance explained per fit –
lmonly:"r2","adj_r2". Like"n", these are per-fit columns: they are populated bytable_regression_uv()screens, where each predictor block is its own model and theR^2answers "how much of the outcome does this predictor explain on its own". A model whoseR^2is a single model-level number drops the column and reports it in the fit-statistics rows instead (the"r2"/"adj_r2"tokens ofshow_fit_stats, on by default forlm). Refused for classes with no least-squares variance partition (glm,coxph, ...), where only competing pseudo-R^2measures exist: name the one you want inshow_fit_stats.Sample columns:
"n"– per-row N, populated bytable_regression_uv()screens (each predictor block is its own fit); models without per-row N data drop the column."n_events"– outcome event counts asevents/N, per factor level (reference row included) and model totals on continuous rows, computed on each model's own estimation sample (STROBE item 16's "data behind the association"; NEJM-styleno. of events/total no.). Binary (binomial) outcomes and right-censoredcoxphfits – other models raisespicy_invalid_input.Survival estimands –
coxphonly:"rmst","rmst_se","rmst_ci","rmst_p"(restricted-mean- survival-time difference over[0, tau]) and"risk_diff","risk_diff_se","risk_diff_ci","risk_diff_p"(cumulative-incidence difference atat_time), by g-computation from the fitted model with bootstrap inference. The horizon arguments are required; seetau/at_time.
Group tokens (presets) expand to a fixed atomic vector before validation:
-
"all_b"->c("b", "se", "ci", "p") -
"all_b_compact"->c("b", "se", "p") -
"all_b_full"->c("b", "se", "ci", "t", "p") -
"all_beta"->c("b", "beta", "se", "ci", "p") -
"all_ame"->c("ame", "ame_se", "ame_ci", "ame_p") -
"all_ame_compact"->c("ame", "ame_p") -
"all_f2"/"all_eta2"/"all_omega2"->partial_*+ its_cicompanion.
Mix groups and atomic tokens:
show_columns = c("all_b", "ame", "ame_p"). Duplicates after
expansion are deduplicated; the order of tokens controls the
order of the displayed columns. If standardized != "none" and
"beta" is not already requested, it is auto-injected after
"b". Asking for "beta" while standardized = "none" raises
spicy_invalid_input.
Default (show_columns = NULL) is context-aware:
"all_b" for a single model (APA-7 Section 6.46 publication layout),
"all_b_compact" for two or more models (CI dropped to fit the
side-by-side layout; restore it explicitly when needed).
show_fit_stats – model-level rows below the coefficients
Counts:
"nobs","weighted_nobs","n_events"(Cox models: number of events, reported alongsidenper the field convention; blank for other classes).Variance explained (
lmonly):"r2","adj_r2","omega2".Pseudo-
R^2(glmand ordinalpolr/clm):"pseudo_r2_mcfadden"(McFadden 1974),"pseudo_r2_nagelkerke"(Nagelkerke 1991),"pseudo_r2_tjur"(Tjur 2009; binomial only).Residual scale:
"sigma"(lm\hat{\sigma}/ glm dispersion),"rmse".Survey designs (
survey::svyglmonly):"eff_p", the effective number of parameters of the design (the first element of survey'sextractAIC, the sum of the Rao-Scott eigenvalues). Opt-in and distinct from"aic", which reports the design-based AIC of Lumley & Scott (2015). Blank for other classes.Bayesian fits only:
"r2_bayes"(posterior-median BayesianR^2, Gelman et al. 2019 – in the all-Bayesian default),"elpd_loo"and"looic"(PSIS-LOO expected log predictive density and its deviance-scale twin, Vehtari et al. 2017; opt-in, a few seconds per model; the footer discloses the elpd standard error), and"waic"(Watanabe-Akaike; PSIS-LOO is generally preferred).Negative-binomial dispersion (
MASS::glm.nbonly):"theta"(V = \mu + \mu^2/\theta) and"alpha"(= 1/\theta, the Statanbregconvention). Refused for other families.Beta-regression precision (
betareg::betaregonly):"phi"(Var(y) = \mu(1-\mu)/(1+\phi), Ferrari & Cribari-Neto 2004; higher\phi= less dispersion). A constant precision is back-transformed from its link, soy ~ x | 1reports the same\phiasy ~ x. Refused for other families and when the precision has covariates (y ~ x | z), so it is not a single number.GEE (
geepack::geeglmonly):"qic"and"qicu"– the quasi-likelihood information criteria (Pan 2001; QIC for working-correlation choice, QICu for covariate choice) viageepack::QIC(), computed only when requested (QIC silently refits the independence working model);"scale"– the estimated scale (dispersion) parameter, shown only when the fit estimated it (blank underscale.fix = TRUE, matching geepack's own summary);"max_cluster_size"– the largest cluster in the estimation sample (on by default with"n_groups", which renders theN (<id>)cluster count). Cluster statistics count the clusters geepack itself used (geese$clusz, consecutive runs ofid =values), so they always describe the displayed sandwich inference. Refused when no model in the table is a geeglm fit; the likelihood-based tokens ("aic", pseudo-R^2, ...) are refused in return for all-GEE tables – quasi-likelihood has no likelihood.Absorbed fixed effects (
fixestfits, andestimatrfits built withfixed_effects =):"fixed_effects"renders aFixed effects:block at the top of the fit statistics – one Yes / No row per absorbed factor (theetable/esttabconvention), blank cells for models without the concept in mixed tables – and is on by default for both. Varying-slope-only factors (Origin[[x]]) absorb no intercept: they read No when another model absorbs that factor, and contribute no row otherwise."within_r2"is the FE-partialled within R-squared (a default forfeols, opt-in forestimatr; GLM-family fixest fits report fixest's McFaddenpr2instead). Both tokens are refused when no model in the table absorbs fixed effects.Effect size:
"f2".Information criteria:
"aic","aicc","bic","deviance"(lowercase like every other token; the rendered row labels stayAIC/AICc/BIC).Change-stats for hierarchical comparison (active under
nested = TRUE; see Hierarchical comparison below):"r2_change","adj_r2_change","f_change","f2_change","lrt_change","aic_change","aicc_change","bic_change","deviance_change","p_change".
Default (resolved when NULL) is class-aware: lm and
lm_robust fits get
c("nobs", "r2", "adj_r2"); glm and ordinal polr / clm fits get
c("nobs", "pseudo_r2_mcfadden", "pseudo_r2_nagelkerke", "aic");
mixed lm + glm sets union both groups (the renderer per-row
en-dashes the inappropriate cell); Cox fits get
c("nobs", "n_events", "aic"); GEE fits get
c("nobs", "n_groups", "max_cluster_size") (the cluster
structure, Stata xtgee-header style). When nested = TRUE, the
class-aware default is extended with change tokens
(c("r2_change", "f_change", "p_change") for lm,
c("lrt_change", "p_change") for glm). The order of tokens in
show_fit_stats controls the order of the rows.
Multi-model semantics
Pass a single fit or a list() of fits. Multi-model layout
draws a centred spanner label above each model's
sub-columns:
-
list("Naive" = m1, "Adjusted" = m2)-> spanner labels"Naive"/"Adjusted". Partial naming (list("Naive" = m1, m2)) auto-fills missing slots as"Model <position>". -
list(m1, m2)(unnamed) -> if all response variables differ, the bare DV name (fromformula(fit)[[2]]) becomes the spanner label and the redundant Outcome body row is suppressed. If DVs match, the labels default to"Model 1, 2, ...". -
model_labels = c("A", "B")overrides everything.
Duplicate explicit names in the list are rejected
(spicy_invalid_input) – they would silently collide in the
internal model_id key. So is a name that collides with the
auto-filled label of another slot (list("Model 2" = m1, m2)):
two models under one spanner cannot be told apart in the table,
nor addressed by inline(). Rename the model, or pass
model_labels.
Inference and standard errors
vcov selects the variance-covariance estimator:
-
"classical"– OLS (lm) / Fisher information (glm). -
"HC0"to"HC5"– heteroskedasticity-consistent (viasandwich::vcovHC()). -
"CR0"to"CR3"– cluster-robust with Satterthwaite-corrected df (viaclubSandwich::vcovCR()). Requirescluster. -
"bootstrap"– nonparametric or cluster bootstrap (boot_nreplicates). -
"jackknife"– leave-one-out / leave-one-cluster-out.
For multi-model use, both vcov and cluster accept a single
value (recycled to all models) or a list (one per model). The
same fit can appear several times with different estimators to
compare standard errors side-by-side.
Inferential regimes for lm / glm (B and AME share the same
regime):
-
classical,HC*-> t withdf.residual(lm) / z-asymptotic (glm). -
bootstrap,jackknife-> z asymptotic. -
CR0-CR3-> t with Satterthwaite-corrected df (B viaclubSandwich::coef_test(); AME viaclubSandwich::linear_contrast(); Pustejovsky & Tipton 2018). Under non-linear terms (poly(),I(),log(),splines::ns()), AME falls back to z-asymptotic with aspicy_fallbackwarning.
For other classes the t-vs-z axis follows the estimator's native
reference distribution (e.g. glm, glmmTMB, survival, ordinal,
betareg, mlogit use z; lm, lme, lmer, rms::ols use t).
Multinomial models: outcome categories as columns
A single nnet::multinom model renders the publication layout:
predictors as rows, one column group per non-reference outcome
category (spanner = category name), so a predictor's effect can
be compared across equations along its row. The mandatory footer
note Reference outcome: <level>. names the base category (Stata's
"base outcome" line); outcome_labels relabels the spanners (e.g.
"Student vs Employed"). Fit statistics print once, under the
first group. When per-category AMEs are requested the reference
category appears as a last, AME-only group (its coefficient cells
are empty – the reference has no equation; its AMEs complete the
zero-sum across categories). The default show_columns compacts
to B / SE / p exactly like multi-model tables; request atomic
tokens to restore CIs. Multi-model and nested = TRUE multinomial
tables keep the one-row-per-(category, predictor) layout –
categories-within-models would need two spanner levels.
tidy() and output = "long" always return the long form
("<category>: <term>" rows), whatever the display;
as_structured() mirrors the displayed table (one column set per
category), as it does for every other layout.
Robust SE availability by model class
Not every estimator is defined for every class. A robust vcov
the class cannot honour fails fast with spicy_unsupported_vcov
– never a silent model-based result under a robust label:
lm,glm,MASS::glm.nball of
classical,HC*,CR*,bootstrap,jackknife.lmadditionally takes"CR1S"– the full Stataregress, vce(cluster)convention: the CR1S small-sample scaling (equal tosandwich::vcovCL()withtype = "HC1") witht(G - 1)inference,Gthe number of clusters. Use it to reproduce Stata tables;"CR2"(Bell-McCaffrey, Satterthwaite df) remains the recommended modern choice."CR1S"is refused forglmwith an explanation: Stata's ML commands use a different convention (G/(G-1)scaling, z), so the label would not match Stata output there.mlogitclassical+CR*only (cluster at the choice-situation level) –sandwich::vcovHC()mis-scales the sandwich for mlogit's per-choice-situation scores, soHC*is refused.lmer,lme,coxph,survreg,mgcv::gam/bam,polr,clm,betareg,nnet::multinom,pscl::zeroinfl/hurdle,rms(ols/lrm/cph/Glm)classical+CR*only –HC*and the resamplers (which refitlm/glm) are not defined for these.clmwith a scale / nominal (partial-PO) component isclassicalonly.multinomneeds sandwich >= 3.1-2 (which added itsestfun()method); itsclusteris one entry per observation. For the two-partpsclfits the cluster sandwich covers both components (count and zero) at once.quantreg::rqits own estimator family, not the sandwich vocabulary:
classical(="nid", quantreg's large-sample default),"iid","ker","rank"(intervals only) and a native"bootstrap", clustered via the wild gradient bootstrap.HC*,CR*andjackknifeare refused, each with its own reason – see thevcovargument.geepack::geeglmno spicy-side estimator at all: GEE inference is robust by construction, so the fit's own sandwich (or jackknife) standard errors – chosen by geeglm's
std.err =option, clustered on itsid =– are the displayed inference.HC*/CR*andclusterare refused with a pointer to those fit options.survey::svyglmclassicalonly, on principle: the design-based Taylor / replicate variance is already the robust variance for the declared design, and clustering belongs in the design itself (survey::svydesign(ids = )), not in the table call.estimatr(lm_robust/iv_robust) andfixestthese fits carry the variance they were computed with – estimatr's
se_type =, and whatever fixest's ownvcovinterface produced – and spicy never overwrites it. Avcovrequest is refused; set the estimator through the fitting package instead (se_type =for estimatr;vcov =at estimation or insummary()for fixest).- Other classes (
glmer,glmmTMB,rstanarm/brms, ...) classical(model-based) only (clubSandwich has no working backend forglmer/glmmTMB).
Cluster-robust backends differ by class but are each cross-validated
to the field-standard oracle: lm/glm/lmer/lme use
clubSandwich (CR2 = Bell-McCaffrey, with Satterthwaite df for
lm/lme/lmer); coxph/cph use the Lin-Wei grouped-dfbeta
sandwich (identical to coxph(..., cluster=));
survreg/gam/polr/clm/betareg/mlogit/multinom and the
pscl two-part fits use
sandwich::vcovCL(); rms fits use rms::robcov() (which
needs the fit's x = TRUE, y = TRUE). These single cluster
sandwiches have no CR0-CR3 bias-reduction variants, so the requested
CR* maps to the one available estimator. cluster length is one
entry per observation, except mlogit (one per choice situation)
and censored coxph (one per subject).
How to specify cluster
Three accepted forms, in order of preference:
-
Formula –
cluster = ~region(orcluster = ~region:yearfor the interaction of two variables). The variables are looked up inmodel.frame(fit)first, then in the originaldataargument captured by the fit. Recommended: independent of the dataset's name, composable for multi-way clustering, consistent withsandwich::vcovCL()/clubSandwich::vcovCR(). -
String –
cluster = "region". A single column name resolved the same way as the formula. Convenient but cannot express interactions. -
Vector –
cluster = df$region. An atomic vector of lengthnobs(fit). Use this when the cluster key is derived on the fly (cluster = interaction(df$region, df$year),cluster = as.integer(format(df$date, "%Y"))), comes from a different dataset with matching row order, or is otherwise not a column of the model'sdata.
A bare unquoted name (cluster = region) is not column
selection: it is evaluated as an ordinary R variable. If an object
region exists in the calling environment, its value is used as
the vector form above; if it does not (the typical "unquoted
column name" intent), the call fails with a migration error
pointing at ~region / "region". Bare-name column selection is
deliberately unsupported – it would require non-standard
evaluation magic that breaks under programmatic use (function
wrapping, dynamic column choice, loops).
For multi-model use, mix forms freely:
cluster = list(~region, "region", df$region).
Bayesian fits (rstanarm / brms)
stanreg and brmsfit models are summarized from their posterior
draws: B is the posterior median, SE the posterior MAD SD (the
scaled median absolute deviation – the pairing of Regression and
Other Stories and rstanarm's own print; equal to the posterior SD
for a normal posterior), and the interval an equal-tailed credible
interval (header 95% CrI; ci_method = "hdi" opts into the
highest-density interval, header 95% HDI). There is no p-value
or t-statistic; group presets ("all_b", ...) expand without
them, and the opt-in "pd" column reports the probability of
direction with a footer definition (Makowski et al. 2019). Under
exponentiate = TRUE all quantities are computed from the
exponentiated draws directly (the SE is the posterior MAD SD on
the ratio scale, not a delta-method approximation). The AME
columns are equally draws-native: avg_slopes() computes the
marginal effects per posterior draw, and the table reports their
posterior median, MAD SD and credible interval (equal-tailed or
HDI per ci_method); "ame_p" has nothing to fill and follows
the p-column policy (dropped from presets, refused as an atomic
token, dashed in mixed tables). Standardized betas follow the same
draws logic: "posthoc" / "basic" / "smart" are exact affine
rescales of the link-scale summaries (the Gaussian family divides
by SD(y); every other family standardizes predictors only, the
frequentist glm convention); "refit" and "pseudo" are refused.
Fit statistics: "r2_bayes" (in the Bayesian default) plus the
opt-in "elpd_loo" / "looic" / "waic", whose standard errors
are disclosed in the footer; unreliable estimates (PSIS-LOO Pareto
k above the sample-size-specific threshold
\min(1 - 1/\log_{10} S, 0.7) of
Vehtari et al. 2024 – the same bound loo::loo() prints – and
WAIC p_waic > 0.4) add a footer caveat instead of being
silenced. Every table is backed by an automatic sampler-diagnostics
guard (R-hat >= 1.01, ESS below 100 per chain – floored at 400 so
fewer chains never weaken the bar –, divergent
transitions, E-BFMI < 0.2, per Vehtari et al. 2021): problems add
a footer line and raise a warning classed
spicy_bayes_diagnostics (nested under spicy_caveat, so
withCallingHandlers(spicy_bayes_diagnostics = ...) mutes the
guard selectively); clean fits print nothing. Per-coefficient
"rhat" / "ess_bulk" / "ess_tail" columns are available in
all-Bayesian tables, as is "mcse" – the Monte Carlo standard
error of the displayed posterior median, the criterion for how
many digits a table can honestly show (a displayed digit is
Monte-Carlo stable when twice the MCSE stays below it; Gelman,
Vehtari, McElreath et al. 2026, sec. 11.6). Under
exponentiate = TRUE the MCSE is recomputed on the exponentiated
draws.
Refused on principle (classed error, never a silent fallback):
likelihood-based fit statistics ("aic", pseudo-R², ...),
p_adjust, robust / cluster vcov (model the clustering with
group-level terms instead), ci_method = "profile" /
"boot_percentile", and non-MCMC fits
(algorithm = "meanfield" / "optimizing" – refit with
algorithm = "sampling"). See the
Bayesian regression tables
article.
Hierarchical (nested) model comparison
nested = TRUE adds per-pair change statistics as in-table
rows (APA Table 7.13 / Stata esttab / SPSS Model Summary
convention). Each adjacent pair (M2 vs M1, M3 vs M2, ...)
contributes one column of change stats; the FIRST model column
gets en-dashes (no previous model to compare to). Validation
requires identical nobs and identical response variable
across all models.
Structural refusals. Four hierarchies carry no valid change
statistic whatever their nobs and response, and are refused with
spicy_invalid_input rather than rendered with empty or
meaningless cells:
-
A REML pair whose fixed effects differ (
nlme::lme(method = "REML"),glmmTMB(REML = TRUE)). The restricted likelihood depends on the fixed-effects design matrix, so the two criteria are not on a common scale and the likelihood-ratio test is not valid (Pinheiro & Bates 2000 Section 2.4.2). Refit both by maximum likelihood (method = "ML",REML = FALSE). A REML pair that differs only in its random structure stays valid and is accepted. -
Two
fixestfits that absorbed different fixed effects. fixest counts every absorbed level as an estimated parameter, so a test across two| festructures measures the change in the absorbed effects together with the change in the coefficients, and the two cannot be separated. Hold the| fepart fixed across the hierarchy. -
Mixed-effects fits from different engines (
lme4,glmmTMB,nlme::lme). The change statistics come from the engine's own two-modelanova(), and no engine's method accepts a fit produced by another. Refit one of the two with the other's engine. -
A
MASS::rlmhierarchy. M-estimation minimises a bounded loss on the scaled residuals: it partitions no sums of squares and it is not a likelihood, so there is neither a likelihood-ratio test nor a partial F to report. Every change token is refused – including on a single fit or anested = FALSEpair, where the token is the explicit request. Compare the least-squares fits (lm()) if the hierarchy itself is the question.
In the first three, nested = FALSE renders the models side by
side instead.
Default change tokens auto-injected when show_fit_stats is
NULL:
All-lm (and
nls):c("r2_change", "f_change", "p_change")– APA hierarchical regression standard.All-glm:
c("lrt_change", "p_change")– Hosmer & Lemeshow Section 3.5; Long & Freese 2014 Section 3.2.4.Any other hierarchy that carries a likelihood (
survreg,polr,clm,coxph,gls,betareg,multinom,fixest, ...): the samec("lrt_change", "p_change").Mixed-effects (
lmer/glmer/glmmTMB/nlme::lme):c("aic_change", "bic_change", "lrt_change", "p_change")– Pinheiro & Bates 2000 Section 2.4.1.Quantile (
rq):c("f_change", "p_change")– the Wald-type testquantreg::anova.rq()reports.
To customise, pass the change tokens directly to
show_fit_stats. The variance-explained change tokens
("r2_change", "adj_r2_change", "f_change",
"f2_change") raise spicy_invalid_input on any hierarchy
whose nested comparison is a likelihood-ratio test – glm,
and equally every other likelihood class and the
mixed-effects families: that comparison reports a chi-square,
not a variance-explained change, and the message points at
"lrt_change" + "p_change". A quantile hierarchy refuses
them too, and "lrt_change" with them, pointing at
"f_change" + "p_change" instead. lm and nls keep the
least-squares tokens.
Standardised coefficients
standardized controls the method when "beta" is in
show_columns:
-
"refit"– refit on z-scored data. Forlmboth X and Y are z-scored (Cohen et al. 2003 gold standard); forglmonly numeric X (Long & Freese 2014 Section 4.7.2 "x-standardization"). -
"posthoc"– post-hoc scaling. lm:\beta = B \times SD(X) / SD(Y); glm: X-only\beta = B \times SD(X)(Y is undefined on the link scale). -
"basic"– like"posthoc"but factor dummies are scaled by their column SD. -
"smart"– Gelman (2008): continuous numeric inputs are scaled by2 * SD(X)(a +/-1 SD swing then spans the same range as a binary's 0 to 1 step); binary inputs – numeric 0/1 and factor dummies – are left unscaled. (Before 0.13.0 the rule was applied inverted; see NEWS.) -
"pseudo"– glm only. Menard (2004, 2011) fully-standardised\beta = B \times SD(X) / SD(Y^*), withY^*the latent variable on the link scale andSD(Y^*) = \sqrt{Var(\hat{\eta}) + Var_{link}}(\pi^2/3logit,1probit,\pi^2/6cloglog). Binomial families only; non-binomial returns NA with aspicy_caveat. -
"none"(default) – no\betacomputed.
Interactions and transforms. Under "refit", an interaction's
\beta is the coefficient of the product of the z-scored
components – the recommended treatment for such models (Cohen et
al. 2003 Section 7.7; Aiken & West 1991; Friedrich 1982). Under
"posthoc", "basic", and "smart", the product / transformed
design column is treated as a single numeric column and scaled by
its own SD (under "smart", by 2 * SD when the product column is
continuous; a binary product column – e.g. a binary-by-binary
interaction – stays unscaled like any binary input) – the
convention of SPSS beta, Stata regress, beta, SAS PROC REG STB, and
lm.beta::lm.beta(), identical to
effectsize::standardize_parameters(method = "basic") on those
columns. The two conventions differ whenever the components are
correlated; a spicy_caveat warns and the footer names the
convention actually used. "refit" declines formulas with inline
transforms (log(x), poly(), factor() written in the formula):
the model frame's evaluated columns cannot be re-evaluated on
z-scored data, so spicy falls back to "posthoc" with a warning
(effectsize's refit instead standardizes before the transform
– a different estimand). Pre-build transformed columns in data
for an exact refit.
Multiple-comparison adjustment
Adjusting the p-values of all coefficients of a single
regression model is not the standard convention. Each
coefficient tests a distinct hypothesis on a distinct
predictor – not the situation multiple-testing procedures
were designed for (Rothman 1990; Greenland 2017; APA Manual 7
Section 6.46; Harrell Regression Modeling Strategies Section 5.4; Gelman,
Hill & Yajima 2012). Hence the default p_adjust = "none".
Adjustment is appropriate for: mass screening with no prior
hypothesis (typically "BH" / FDR), pre-registered
multi-endpoint confirmatory designs (typically "holm"), or
when a journal / SAP explicitly requests it.
The adjustment runs before any keep / drop filtering,
so the family is the model's full coefficient set (intercept
and reference rows excluded), not the displayed subset –
filtering is a display choice and must not change the
inferential family.
Tied event times and the survival estimands
A Cox model has to decide what to do when several events share
the same time, and survival::coxph() decides with
ties = "efron" by default. The absolute estimands do not make
a second decision:
The baseline hazard behind the RMST-difference and
risk-difference columns follows the tie-handling convention of
the fit, exactly as survival::survfit() and
survival::basehaz() do. A ties = "breslow" fit gives a
Breslow baseline; the default Efron fit gives an Efron baseline.
One rule, so the hazard-ratio column and the dRMST column always
come from the same likelihood. To change the convention, change
the fit.
When the choice matters is not simply "when there are many ties". What governs it is the fraction of tied events within the risk set, together with the effect size and the spread of the covariates: a large risk set dilutes a tied block, a small one – fine strata, matched or nested case-control designs – does not. A file with hundreds of tied days can be insensitive to the convention while a small stratified one with a handful of ties is not.
And Efron is the better default, not the correct answer: the approximations are all biased at finite sample size and agree asymptotically, and Efron's virtue is a much smaller small-sample bias (Hertz-Picciotto & Rockhill 1997). Efron himself, quoting Peto, put it as "it probably doesn't make much difference" (Efron 1977).
Output formats and broom integration
output selects the return type:
-
"default"– aspicy_regression_table(data.framesubclass) printed viaspicy_print_table(). -
"data.frame"/"long"– raw data.frame / long-format tibble. -
"gt"/"flextable"/"tinytable"– rich-format HTML / Word / PDF tables (require the corresponding Suggests package). -
"excel"– writes toexcel_pathviaopenxlsx2::write_xlsx(). -
"word"– writes toword_pathviaflextable::save_as_docx(). -
"clipboard"– copies to the system clipboard viaclipr::write_clip().
broom::tidy() returns a long tibble with one row per
(model_id, term, estimate_type) and broom-canonical column
names (estimate, std.error, conf.low, conf.high,
statistic, p.value). broom::glance() returns one row per
model with the model-level statistics; df.residual is kept
numeric so cluster-robust Satterthwaite df is preserved.
Global options
-
options(spicy.style = )– the journal style applied by default to all four table families (table_regression(),table_categorical(),table_continuous(),table_continuous_lm()). Takes a theme name ("jama","nejm","lancet","annals","apa","aer") or aspicy_style()object. The style of a report is a property of the document, so it is set once in the setup chunk; thestyleargument overrides it per call, and any formatting argument you type overrides both.options(spicy.style = NULL)restores spicy's own defaults. Seespicy_style()for the exact rules each theme encodes and the official document they come from. -
options(spicy.note_style = )– how the table note is rendered by the"tinytable"engine (all four table families). A note is subordinate to the table it documents, so it is set one size down, in black:0.9em, matching the"gt"engine and the 10pt-on-11pt of the Word engine. Two opt-ins:-
"none"– no intervention; the note is rendered exactly as the receiving document template styles it. any other string – extra arguments for the Typst
text()call around the note, appended to the size (e.g.options(spicy.note_style = "fill: luma(89)")for a grey note). Typst only; the HTML note keeps the plain0.9em.
options(spicy.note_style = NULL)restores the default. -
Weights
No weights argument: weights are a property of the fit
(extracted via stats::weights()). Pass them when fitting:
lm(y ~ x, data = df, weights = w). All downstream
computations (vcov, AME, standardisation, weighted_nobs)
extract them automatically.
Internationalisation
Output is in English. Override user-facing strings via
reference_label, model_labels, outcome_labels, and
labels. The title and footer are post-processable via
attr(result, "title") and attr(result, "note").
Classed conditions
Every error and warning emitted by table_regression() carries
a classed condition for programmatic dispatch via tryCatch()
or withCallingHandlers(). Errors inherit from spicy_error
(root); warnings from spicy_warning. Specific leaves used by
this function include spicy_invalid_input,
spicy_invalid_data, spicy_unsupported,
spicy_unsupported_vcov, spicy_unsupported_standardized,
spicy_missing_pkg, spicy_missing_column,
spicy_ignored_arg, spicy_caveat, spicy_fallback. See
spicy for the full taxonomy.
References
APA Manual 7 (American Psychological Association, 2020), Tables 7.13-7.15.
Aiken, L.S. & West, S.G. (1991). Multiple regression: Testing and interpreting interactions.
Cohen, J., Cohen, P., West, S.G., & Aiken, L.S. (2003). Applied multiple regression / correlation analysis for the behavioral sciences (3rd ed.). Lawrence Erlbaum.
Davison, A.C. & Hinkley, D.V. (1997). Bootstrap methods and their application. Cambridge University Press.
Efron, B. (1977). The efficiency of Cox's likelihood function for censored data. Journal of the American Statistical Association, 72(359), 557-565.
Friedrich, R.J. (1982). In defense of multiplicative terms in multiple regression equations. American Journal of Political Science, 26(4), 797-833.
Hertz-Picciotto, I. & Rockhill, B. (1997). Validity and efficiency of approximation methods for tied survival times in Cox regression. Biometrics, 53(3), 1151-1156.
Pustejovsky, J.E. & Tipton, E. (2018). Small-sample methods for cluster-robust variance estimation and hypothesis testing in fixed effects models. Journal of Business & Economic Statistics, 36(4), 672-683.
Wasserstein, R.L., Schirm, A.L., & Lazar, N.A. (2019). Moving to a world beyond "p < 0.05". The American Statistician, 73(sup1), 1-19.
See Also
table_regression_models() for the registry of supported model
classes and the per-family behaviour reference (also reachable as
?table_regression_mixed, ?table_regression_ordinal, ...). If a
class is not listed there, try table_regression(fit) anyway –
unsupported classes error with a clear message.
Other regression-table functions:
table_continuous_lm() for one-predictor-by-many-outcomes
descriptive tables.
Other spicy table functions:
freq(), cross_tab(), table_categorical(),
table_continuous().
Underlying machinery:
spicy_print_table() for ASCII rendering;
build_ascii_table() for the low-level renderer.
Inferential infrastructure (internal):
compute_model_vcov(), compute_coef_inference(),
compute_wald_test().
broom integration:
broom::tidy(), broom::glance().
Examples
# ---- Single-model usage ------------------------------------------
fit <- lm(wellbeing_score ~ age + sex + smoking, data = sochealth)
# Default APA layout: B / SE / 95% CI / p plus the n / R^2 /
# Adj. R^2 fit-stats footer. Factor reference level is annotated
# with `(ref.)` and shows an en dash in the statistic columns.
table_regression(fit)
# Standardised coefficients (beta) injected next to B. "refit"
# is the Cohen et al. (2003) refit-on-z-scores convention;
# "basic" reproduces the SPSS / Stata regress, beta definition.
table_regression(fit, standardized = "refit")
# Custom column set: B + AME + AME-specific p-value. Note that
# the `p` token always belongs to B, never to AME -- use the
# explicit `ame_p` token for AME inference.
if (requireNamespace("marginaleffects", quietly = TRUE)) {
table_regression(
fit,
show_columns = c("b", "p", "ame", "ame_ci", "ame_p")
)
}
# Group-token shortcut: "all_b" + "all_ame" expands to the full
# B / AME column families side by side.
if (requireNamespace("marginaleffects", quietly = TRUE)) {
table_regression(fit, show_columns = c("all_b", "all_ame"))
}
# ---- Cluster-robust variance -------------------------------------
# CR2 (Bell-McCaffrey) with Satterthwaite-corrected df is the
# recommended default under few clusters. Three forms are accepted
# for `cluster`; the formula is preferred for composability with
# multi-way clustering and for programmatic robustness.
if (requireNamespace("clubSandwich", quietly = TRUE)) {
table_regression(fit, vcov = "CR2", cluster = ~region)
}
if (requireNamespace("clubSandwich", quietly = TRUE)) {
table_regression(fit, vcov = "CR2", cluster = "region")
}
if (requireNamespace("clubSandwich", quietly = TRUE)) {
table_regression(fit, vcov = "CR2", cluster = ~region:age_group)
}
# ---- Hierarchical (nested) regression ----------------------------
# Adds in-table change-statistic rows (Delta R^2 / F-change /
# p-change for lm; LRT / p-change for glm) below the fit-stats.
# Note: hierarchical comparison requires identical observations
# across all models -- prepare a complete-case subset first so
# R's listwise deletion does not produce different `nobs` per
# model (which the function rejects).
sochealth_cc <- na.omit(
sochealth[, c("wellbeing_score", "age", "sex", "smoking")]
)
m1 <- lm(wellbeing_score ~ age, data = sochealth_cc)
m2 <- lm(wellbeing_score ~ age + sex, data = sochealth_cc)
m3 <- lm(wellbeing_score ~ age + sex + smoking, data = sochealth_cc)
table_regression(
list("Step 1" = m1, "Step 2" = m2, "Step 3" = m3),
nested = TRUE
)
# ---- Side-by-side variance comparison ----------------------------
# Same fit, three vcovs in one wide table. Useful for showing the
# sensitivity of inference to the variance assumption.
if (requireNamespace("clubSandwich", quietly = TRUE)) {
table_regression(
list("Classical" = fit, "HC3" = fit, "CR2" = fit),
vcov = list("classical", "HC3", "CR2"),
cluster = list(NULL, NULL, ~region)
)
}
# ---- Tidy long format for downstream pipelines -------------------
if (requireNamespace("broom", quietly = TRUE)) {
broom::tidy(table_regression(fit))
}
# ---- Mixed-effects models ----------------------------------------
# Linear mixed-effects (lme4). The footer adds a random-effects
# panel with sigma + Wald SE / CI from `merDeriv`, the Nakagawa
# marginal / conditional R^2 fit-stats, and a per-class p-value
# annotation line.
if (requireNamespace("lme4", quietly = TRUE)) {
fit <- lme4::lmer(Reaction ~ Days + (Days | Subject),
data = lme4::sleepstudy)
table_regression(fit)
# Switch to the variance scale (sigma^2 instead of sigma).
table_regression(fit, re_scale = "variance")
# Minimal random-effects display: estimates only, no SE / CI.
table_regression(fit, re_columns = "est")
# Suppress the random-effects panel entirely.
table_regression(fit, show_re = FALSE)
}
# Hierarchical mixed-effects comparison (nested LRT).
if (requireNamespace("lme4", quietly = TRUE)) {
m1 <- lme4::lmer(Reaction ~ 1 + (1 | Subject),
data = lme4::sleepstudy, REML = FALSE)
m2 <- lme4::lmer(Reaction ~ Days + (1 | Subject),
data = lme4::sleepstudy, REML = FALSE)
m3 <- lme4::lmer(Reaction ~ Days + (Days | Subject),
data = lme4::sleepstudy, REML = FALSE)
table_regression(list(m1, m2, m3), nested = TRUE)
}
## Not run:
# ---- Rich-format outputs (require optional Suggests packages) ----
table_regression(fit, output = "gt")
table_regression(fit, output = "flextable")
table_regression(fit, output = "tinytable")
# ---- File outputs ------------------------------------------------
table_regression(fit, output = "excel",
excel_path = tempfile(fileext = ".xlsx"))
table_regression(fit, output = "word",
word_path = tempfile(fileext = ".docx"))
# ---- System clipboard (interactive use) --------------------------
table_regression(fit, output = "clipboard")
## End(Not run)
Supported models and per-family behaviour of table_regression()
Description
table_regression_models() returns the registry of model classes
supported by table_regression(), one row per engine, with each class's
family, average-marginal-effects estimand, exponentiate semantics, and
labelled table blocks. The same registry drives this page's table, so the
published list cannot drift from the code.
This page is also the reference for per-family behaviour (the
sections below). It is reachable as ?table_regression_models,
?table_regression_mixed, ?table_regression_ordinal,
?table_regression_counts, ?table_regression_categorical,
?table_regression_survival, ?table_regression_robust, or
?table_regression_bayesian.
If a class is not listed: fit the model and call table_regression(fit)
anyway – unsupported classes error with a clear message naming the
supported set. Feature requests are welcome on the issue tracker.
Usage
table_regression_models()
Value
A data frame with one row per supported engine and columns
family, class, engine, ame, exponentiate, blocks.
Supported classes
| Family | Class | Engine | AME | Exponentiate | Blocks |
| Linear and generalized linear | lm | stats::lm() | yes | - | - |
| Linear and generalized linear | glm | stats::glm() | yes | OR / IRR / RR / MR / HR (link) | - |
| Linear and generalized linear | negbin | MASS::glm.nb() | yes | IRR | - |
| Linear and generalized linear | rlm | MASS::rlm() | yes | - | - |
| Linear and generalized linear | nls | stats::nls() | no | - | - |
| Robust, IV, quantile, panel | lm_robust | estimatr::lm_robust() | yes | - | - |
| Robust, IV, quantile, panel | iv_robust | estimatr::iv_robust() | yes | - | - |
| Robust, IV, quantile, panel | ivreg | AER::ivreg() | yes | - | - |
| Robust, IV, quantile, panel | tobit | AER::tobit() | yes | - | - |
| Robust, IV, quantile, panel | rq | quantreg::rq() | yes | - | - |
| Robust, IV, quantile, panel | fixest | fixest::feols(), fixest::feglm(), fixest::fepois(), fixest::fenegbin() | yes | feglm: OR / IRR | - |
| Mixed effects | lmerMod | lme4::lmer() | yes | - | Random effects |
| Mixed effects | glmerMod | lme4::glmer() | yes | OR / IRR (link) | Random effects |
| Mixed effects | glmmTMB | glmmTMB::glmmTMB() | yes | link-dependent (IRR for count families) | Random effects; Zero-inflation; Dispersion |
| Mixed effects | lme | nlme::lme() | yes | - | Random effects |
| Mixed effects | gls | nlme::gls() | yes | - | - |
| Population-averaged (GEE) | geeglm | geepack::geeglm() | yes | OR / IRR / RR / MR / HR (link) | - |
| Ordinal | polr | MASS::polr() | per category | OR (logit) | Thresholds |
| Ordinal | clm | ordinal::clm() | per category | OR (logit) | Thresholds; Non-proportional effects |
| Categorical | multinom | nnet::multinom() | per outcome | OR | per-outcome blocks |
| Categorical | mlogit | mlogit::mlogit() | no | OR | per-alternative rows |
| Counts, two-part | zeroinfl | pscl::zeroinfl() | yes (combined response) | IRR (count) + OR (logit zero part) | Zero-inflation |
| Counts, two-part | hurdle | pscl::hurdle() | yes (combined response) | IRR (count) + OR (logit zero part) | Zero hurdle |
| Survival | coxph | survival::coxph() | RMST / risk diff | HR | - |
| Survival | survreg | survival::survreg() | yes + RMST / risk diff | TR (log-scale distributions) | - |
| Survival | cph | rms::cph() | no | HR | - |
| Survival | flexsurvreg | flexsurv::flexsurvreg() | no | TR / HR (dist) | distribution parameters |
| Survey-weighted | svyglm | survey::svyglm() | yes (design-based) | OR / IRR | - |
| Survey-weighted | svyolr | survey::svyolr() | per category (design-based) | OR (logit) | Thresholds |
| Survey-weighted | svycoxph | survey::svycoxph() | no | HR | - |
| Additive, proportions, selection | gam | mgcv::gam(), mgcv::bam() | yes | OR / IRR (link) | - |
| Additive, proportions, selection | betareg | betareg::betareg() | yes | OR (mean link) | - |
| Additive, proportions, selection | selection | sampleSelection::selection() | no | - | selection component |
| rms | ols | rms::ols() | yes | - | - |
| rms | lrm | rms::lrm() | yes | OR | - |
| rms | Glm | rms::Glm() | yes | link-dependent | - |
| Bayesian | stanreg | rstanarm::stan_glm(), rstanarm::stan_glmer() | yes (draws) | link-dependent | Random effects (if multilevel) |
| Bayesian | brmsfit | brms::brm() | yes (draws) | link-dependent | Random effects (if multilevel) |
Shared semantics (all classes)
A robust
vcovrequest is honoured through the class's field-standard backend, or refused with a clear error naming the supported set; the footer always names the estimator actually applied.-
exponentiate = TRUEis link-gated: it produces a labelled ratio (OR / IRR / HR / RR / MR / TR) only where the link warrants one. Identity-link fits warn and are left untouched; non-ratio links (probit, cauchit, inverse, ...) are refused with a clear error. Class-specific structure renders as labelled subordinate blocks of rows in the same table, each explained by a footer line.
Fit statistics default to the family's field standard (
show_fit_statsoverrides; class-inappropriate tokens are rejected with a pointer to the right ones).Everything is available programmatically:
broom::tidy(),glance(),as_structured(),as.data.frame().
Mixed effects
Fixed effects: Satterthwaite t (lmer + lmerTest), Wald z
(glmer, glmmTMB), containment-df t (lme). Random effects render as
a Random effects block of rows (SD / correlation / residual with SE and
CI; re_scale, re_columns), deliberately with no per-row p-value
(boundary-invalid Wald; Self & Liang 1987) – the footer carries the
chi-bar-squared LR test of the whole random part, and
re_test = "lrt" / "rlrt" adds an opt-in boundary-correct per-term
test. N (groups) and ICC are fit-stat rows; Nakagawa marginal /
conditional R-squared are the default R-squared family. CR*
cluster-robust standard errors are available for lmer and lme fits,
via clubSandwich with Satterthwaite degrees of freedom. glmer
and glmmTMB fits keep their model-based standard errors: a CR*
request on them is refused with a clear error.
Population-averaged (GEE) models
geepack::geeglm() fits are read on their own terms: the sandwich
standard errors the fit computed (its std.err = option, clustered
on its id =) are the displayed inference – GEE is robust by
construction, so spicy's vcov / cluster arguments are refused
with a pointer to the fit options. Coefficients are
population-averaged (marginal) effects; the footer discloses the
working correlation structure with its estimated alpha. Wald z
inference; exponentiate follows the usual link gates (OR / IRR /
RR). Default fit statistics report the cluster structure (n,
N (<id>), largest cluster); the quasi-likelihood information
criteria "qic" / "qicu" (Pan 2001) and the "scale"
(dispersion) parameter are opt-in – there is no likelihood, so
AIC, pseudo-R-squared, nested = TRUE, and standardized are
refused. See the population-averaged section of the
Mixed-effects regression tables
article for the contrast with subject-specific mixed models.
Ordinal models
Cut-points render as a Thresholds block (log-odds scale, never
exponentiated; show_thresholds). Partial-proportional-odds clm terms
render as a Non-proportional effects block, one coefficient per
cut-point. exponentiate yields proportional odds ratios under logit;
ci_method = "profile" profiles the predictor coefficients. AME is
per-category (the marginal effect on each P(Y = k)). Defaults include
McFadden and Nagelkerke pseudo-R-squared. See the
Ordinal regression tables
article.
Counts and two-part models
Two-part models show their full model: the zero component renders as a
Zero-inflation block (zeroinfl, glmmTMB ziformula: probability of a
structural zero) or a Zero hurdle block (hurdle: probability of a
nonzero count – the opposite direction, hence the distinct label), and a
Dispersion block when dispformula has covariates. Component
coefficients join the p_adjust family and take stars; a zero component
is exponentiated only under a logit link (odds ratio). AME is the
combined-response effect on E(Y). CR* for pscl fits covers both components
via sandwich::vcovCL(). Opt out with show_components = FALSE.
Categorical outcomes
multinom renders per non-reference outcome; exponentiate yields
odds ratios of each outcome against the reference outcome – the
baseline-category logits are log-odds (Agresti; SAS prints
"Odds Ratio Estimates" under its generalized-logit link; Stata's mlogit, rrr labels
the same quantity a relative-risk ratio). AME is per-outcome.
nested = TRUE compares nested multinom fits by likelihood-ratio test
(the anova.multinom() convention). Cluster-robust CR* is
available (one cluster value per observation; sandwich >= 3.1-2)
and the AME columns honour it; HC* is refused – a multi-equation
model has no working residuals.
mlogit renders
per-alternative rows; AME is refused (no slopes() method exists for
its data format). CR* is available with one cluster value per choice
situation, and n counts choice situations; HC* is refused
(sandwich::vcovHC() mis-scales the meat for mlogit's per-chooser
score structure).
Survival models
Cox models exponentiate to hazard ratios; survreg log-scale
distributions to time ratios (identity-scale distributions are left
untouched). AME is refused for Cox fits (no marginal-probability effect
on the hazard scale); their absolute-effect columns are the
"rmst" and "risk_diff" families instead – covariate-adjusted
RMST and cumulative-incidence differences by g-computation, with
the mandatory tau / at_time horizons. For coxph:
right-censored single-record fits, strata() supported
(within-stratum baselines), tt() refused. For survreg: the
closed-form AFT curves are standardized directly (stratified
survreg refused).
CR* uses the Lin-Wei grouped-dfbeta sandwich
(coxph) or rms::robcov() (cph, needs x = TRUE, y = TRUE).
nested = TRUE compares nested Cox fits by likelihood-ratio test.
Survey-design models
Fits from a survey::svydesign() or survey::as.svrepdesign()
design – svyglm (and its replicate sibling svrepglm), svyolr,
svycoxph (and svrepcoxph) – are read as design-based
throughout: the coefficients, the variance and the reference
distribution all come from survey.
Inference is Wald t at the degrees of freedom survey writes on
the FIT – df.residual for svyglm / svyolr, degf.resid or
degf.residual for the two Cox engines – which is what
survey::regTermTest() takes as its denominator. It is not
survey::degf(design), and it is not re-derived here: the six engines
of survey do not share one expression and are not harmonised (a Cox
fit carries degf(design) - p + 1 although it has no intercept for
the + 1 to cancel, so it ends one above the two other classes). The
value is read off the object. The footer names the design and prints
the number, and the average marginal effects answer to the same
distribution as the coefficient rows.
The average marginal effect is the Horvitz-Thompson estimator: the mean unit-level effect weighted by the sampling weights of the analytic sample, with its variance from the delta method on the design vcov.
Counts are both reported: the observed n and the Weighted n the
estimates describe. svycoxph adds the number of events, and its
concordance goes to the footer.
What is refused, and why: every model-derived variance (HC*,
CR*, bootstrap, jackknife) – the design is the variance
authority, and the way to change the estimator is to change the
design; every likelihood statistic (AIC, BIC, logLik, deviance,
pseudo-R-squared) for svyolr and svycoxph – there is no
likelihood, and survey's own deviance() returns a sign-flipped
likelihood-ratio statistic on one Cox engine and a bare zero on the
other; the AME for svycoxph, on the same ground as for a plain Cox
fit; and the "rmst" / "risk_diff" columns for svycoxph, whose
uncertainty comes from resampling subjects and so ignores the strata
and clusters the design declares (use survey::svykm() for a
marginal curve).
nested = TRUE is refused for a design-based table: there is no
likelihood to compare, so every change statistic would be empty, and
a block of empty rows reads like an answer. Put the models side by
side with nested = FALSE, and test a term under the design with
survey::regTermTest().
svyglm keeps an AIC row: survey's extractAIC.svyglm computes
the design-based AIC of Lumley & Scott (2015), the one information
criterion published for this class, and
show_fit_stats = "eff_p" reports the effective number of design
parameters beside it. BIC.svyglm requires a maximal model and has
no default, so it stays blank. See the
Summary tables from a survey design
article.
Robust, IV, quantile and panel models
estimatr fits keep their own robust SEs (never overwritten);
quantreg::rq() defaults to the heteroskedasticity-robust "nid"
sandwich (quantreg's own large-sample default), with "iid",
"ker", "rank" (CIs only) and a native "bootstrap" –
clustered via the wild gradient bootstrap – as vcov options
(the footer names the estimator);
fixest fits – and estimatr fits built with fixed_effects = –
disclose their absorbed fixed effects as a default-on
Fixed effects: block (one Yes / No row per factor;
varying-slope-only factors are not absorbed intercepts and read
No), with per-factor N (<factor>) counts via the opt-in
n_groups token and the within R-squared in the default fit
statistics for fixest (opt-in via within_r2 for estimatr).
Bayesian models
Posterior median, posterior MAD SD, and equal-tailed credible
intervals (ci_method = "hdi" opts into the highest-density
interval); deliberately no p-value column and no stars – the
probability of direction ("pd") is the opt-in posterior summary.
A sampler-diagnostics guard checks every fit (R-hat, ESS,
divergences, E-BFMI) and per-coefficient "rhat" / "ess_bulk" /
"ess_tail" / "mcse" columns are available. The AME columns are
draws-native (posterior median, MAD SD and credible interval of
the per-draw avg_slopes(); no "ame_p"), and so are the
standardized betas ("posthoc" / "basic" / "smart", exact
affine rescales of the draws) on fixed-effects fits:
stan_glm-style models and standard-formula brm() models,
whose design matrix is recovered through insight.
Multilevel fits, stan_polr / stan_betareg, brms formulas with
distributional or special terms, and "refit" / "pseudo" are
refused with a pre-standardization hint.
Multilevel fits
(stan_glmer, brm with grouping terms) report their random
effects as a block – posterior median SD and credible interval per
component, from the draws – with no likelihood-ratio line.
p_adjust and likelihood-based fit-statistic tokens are refused
(no p-values, no likelihood-based information criteria in a
posterior); "r2_bayes" is in the default fit statistics and
"elpd_loo" / "looic" / "waic" are opt-in, with standard
errors and reliability caveats in the footer; compare models with
loo::loo_compare() outside the table.
See Also
table_regression(); the
Publication-ready regression tables
and
Ordinal regression tables
articles.
Examples
table_regression_models()
# All engines of one family:
subset(table_regression_models(), family == "Mixed effects")
Univariable screening table (with optional multivariable merge)
Description
Fits one model per candidate predictor (the univariable screen)
and renders them as a single table with one row block per
predictor. With multivariable = TRUE (default), the full model
containing all predictors is merged side by side under
"Univariable" / "Multivariable" column groups – the standard
presentation of applied epidemiology (the
gtsummary::tbl_uvregression() + tbl_merge() workflow).
Usage
table_regression_uv(
data,
outcome,
predictors,
method = c("lm", "glm", "coxph"),
family = stats::binomial(),
multivariable = TRUE,
complete_cases = FALSE,
show_columns = c("n", "b", "ci", "p"),
show_intercept = FALSE,
title = NULL,
...
)
Arguments
data |
A data frame. |
outcome |
The outcome column (unquoted name, tidyselect). For
|
predictors |
Candidate predictor columns (tidyselect, e.g.
|
method |
|
family |
A stats::family for |
multivariable |
Logical, default |
complete_cases |
Logical, default |
show_columns |
Passed to |
show_intercept |
Display the |
title |
Table title; |
... |
Passed to |
Value
See table_regression() (same output contract).
Sample sizes
By default each univariable model is fit on its own complete
cases, so N varies across predictors – that is what the N
column discloses (shown on the first row of each block), and a
table note states it whenever the Ns differ. The multivariable
model is fit on the complete cases of all its variables (its
n appears in the fit-statistics rows). Pass
complete_cases = TRUE to restrict every model – univariable and
multivariable – to the common complete-case sample.
Variance explained
show_columns = c("n", "b", "ci", "p", "r2") adds an
R^2 column to the screen (method = "lm"): each
predictor block reports its own model's R^2, on the
first row of the block like N. It answers what a coefficient and
its interval cannot – how much of the outcome the predictor
accounts for – and often shows that a firmly established
association still explains a small share of the variance. Add
"adj_r2" for the adjusted form. On the multivariable side the
R^2 is one number for the whole model, so it stays in
the fit-statistics rows (where it is shown by default) instead of
being repeated down a column. Not available for method = "glm"
or "coxph": outside least squares only competing
pseudo-R^2 measures exist, and spicy asks you to name
the one you want (show_fit_stats = "pseudo_r2_mcfadden").
Multiplicity
p_adjust (passed through to table_regression()) treats the
whole univariable screen as ONE family (all screened coefficients
together); the multivariable model is its own family, as in any
multi-model table.
Why the default screen is linear
The default method = "lm" fits the linear screen: R's canonical
model, and – when the outcome is continuous – the estimand with
the most direct reading. If the outcome looks binary under this
default, the screen proceeds as a linear probability model and
says so in a classed warning: LPM coefficients are probability
differences, comparable across models and samples in a way that
odds ratios are not (Mood, 2010), but the model's built-in
heteroskedasticity calls for vcov = "HC3". A two-level factor
(or logical) outcome is coded 0/1 on its second level – the glm
convention – and the warning names the modeled probability; an
outcome with more observed levels is refused (a multinomial
outcome has no linear screen). The classical
epidemiological screen is one argument away: method = "glm"
(with the default family = binomial()) gives the logistic
screen, and supplying any family selects the glm screen
directly.
Intercepts
Hidden by default on both sides (each univariable fit has its own
nuisance intercept), matching gtsummary::tbl_regression()'s
intercept = FALSE default. Pass show_intercept = TRUE to
display them: each univariable block then opens with its own fit's
(Intercept) row, and the multivariable model shows its intercept
as in any table_regression() table.
References
Batra, N. et al. (Eds.) (2021). The Epidemiologist R Handbook, Univariate and multivariable regression. https://epirhandbook.com/en/new_pages/regression.html
Mood, C. (2010). Logistic regression: Why we cannot do what we think we can do, and what we can do about it. European Sociological Review, 26(1), 67-82. doi:10.1093/esr/jcp006
Sjoberg, D.D., Whiting, K., Curry, M., Lavery, J.A., & Larmarange, J. (2021). Reproducible summary tables with the gtsummary package. The R Journal, 13(1), 570-580.
Examples
table_regression_uv(
sochealth,
outcome = smoking,
predictors = c(age, sex, education),
family = binomial(),
exponentiate = TRUE
)
# Linear screen with the share of variance each predictor
# explains on its own (the multivariable model reports its own
# R-squared in the fit-statistics rows).
table_regression_uv(
sochealth,
outcome = wellbeing_score,
predictors = c(age, sex, bmi),
show_columns = c("n", "b", "ci", "p", "r2")
)
Terms method for univariable screen bundles
Description
Returns the stats::terms() object of the formula
outcome ~ predictor_1 + ... + predictor_k spanning every
predictor screened by table_regression_uv(). Non-syntactic
column names are backtick-quoted, so the terms are valid whatever
the input names. Used internally by the label validator, which
reads term labels off every model in a table.
Usage
## S3 method for class 'spicy_uv_screen'
terms(x, ...)
Arguments
x |
A |
... |
Additional arguments (currently ignored). |
Value
A terms object for outcome ~ all screened predictors.
See Also
Tidying methods for a spicy_categorical_svy_table
Description
Standard broom::tidy() and broom::glance() interfaces for an
object returned by table_categorical_svy().
Usage
## S3 method for class 'spicy_categorical_svy_table'
tidy(x, ...)
## S3 method for class 'spicy_categorical_svy_table'
glance(x, ...)
Arguments
x |
A |
... |
Ignored, for S3 compatibility. |
Details
tidy() returns one row per (variable x level x column block), which
is the LONG reading of a table whose blocks are columns. Columns:
variable, label, level, group (the by level, or the margin,
NA without by), total (whether that block is the margin),
n (observed), estimate (the estimated percentage), conf.low,
conf.high, deff. The header rows carry no level statistic and do
not appear; their p-value is what glance() is for.
glance() returns one row per variable: variable, label,
n_levels, p.value, statistic_type (the svychisq() statistic
asked for), degf (the design's own), nobs, weighted.nobs.
n_levels counts the levels the TABLE displays, so a (Missing)
display level counts; the test behind p.value runs on the complete
cases and the observed levels only, as it does in
table_categorical().
Value
A tbl_df (or data.frame when tibble is not installed).
Tidying methods for a spicy_categorical_table
Description
Standard broom::tidy() and broom::glance() interfaces for an
object returned by table_categorical(). They re-shape the
underlying long-format data (stored on the object as the
"long_data" attribute) into the two canonical broom views so the
table can be consumed by any downstream tidyverse-stats pipeline.
Usage
## S3 method for class 'spicy_categorical_table'
tidy(x, ...)
## S3 method for class 'spicy_categorical_table'
glance(x, ...)
Arguments
x |
A |
... |
Currently ignored. Present for compatibility with the
|
Details
tidy() returns one row per (variable x level) – or per
(variable x level x group) when by is used – with
broom-conventional columns: outcome, level, group (when
applicable), n, proportion (the percentage divided by 100).
glance() returns one row per outcome with the omnibus
chi-squared test (when by is used) and the requested association
measure: outcome, test_type ("chi_squared"), statistic
(chi-squared), df, p.value, assoc_type, assoc_value,
assoc_ci_lower, assoc_ci_upper, n_total. Without by, only
outcome and n_total are populated; the other columns are NA.
Value
A tbl_df.
See Also
as.data.frame.spicy_categorical_table() for the raw
wide-format access; tidy.spicy_continuous_table() for the
continuous-descriptive companion.
Tidying methods for a spicy_continuous_lm_table
Description
Standard broom::tidy() and broom::glance() interfaces for an
object returned by table_continuous_lm(). They re-shape the
underlying long-format data into the two canonical broom views so
the table can be consumed by any downstream tidyverse-stats
pipeline.
Usage
## S3 method for class 'spicy_continuous_lm_table'
tidy(x, ...)
## S3 method for class 'spicy_continuous_lm_table'
glance(x, ...)
Arguments
x |
A |
... |
Currently ignored. Present for compatibility with the
|
Details
tidy() returns one row per estimated parameter across all
outcomes:
One row per fitted level mean (
estimate_type = "emmean") for categorical predictors, with the level name interm.One row per contrast (
estimate_type = "difference") when a binary contrast is shown, withterm = "<level2> - <level1>".One row per slope (
estimate_type = "slope") for numeric predictors, withterm = predictor_label.
Standard broom columns: outcome, label, term,
estimate_type, estimate, std.error, conf.low, conf.high,
statistic, p.value. The outcome column carries the original
variable name; label carries the human-readable label.
glance() returns one row per outcome with model-level
statistics. Columns: outcome, label, predictor_type
("categorical" or "continuous"), test_type ("F" for
categorical predictors, "t" for continuous ones),
statistic, df, df.residual, p.value, r.squared,
adj.r.squared, es_type, es_value, es_ci_lower,
es_ci_upper, nobs, weighted_n.
Value
A tbl_df.
See Also
as.data.frame.spicy_continuous_lm_table() for the raw
long-format access.
Tidying methods for a spicy_continuous_svy_table
Description
Standard broom::tidy() and broom::glance() interfaces for an
object returned by table_continuous_svy().
Usage
## S3 method for class 'spicy_continuous_svy_table'
tidy(x, ...)
## S3 method for class 'spicy_continuous_svy_table'
glance(x, ...)
Arguments
x |
A |
... |
Ignored, for S3 compatibility. |
Details
tidy() returns one row per displayed row: one per variable, or one
per (variable x group) with by. Columns: variable, label,
group (NA without by), estimate (the mean), std.error (the
design-based standard error, from survey::svymean() and never
recomputed from sd / sqrt(n) – under a design those are different
quantities), conf.low, conf.high, df (the degrees of freedom
the interval used), n (observed), weighted.n (the sum of the
sampling weights), median, q1, q3, min, max, sd, deff.
glance() returns one row per variable with its group comparison:
variable, label, n_groups, test_type, statistic, df,
df.residual, p.value, degf (the design's own), nobs,
weighted.nobs. One row per variable even without by, where the
comparison columns are NA – a fixed schema a pipeline can index
into by NAME.
Value
A tbl_df (or data.frame when tibble is not installed).
Tidying methods for a spicy_continuous_table
Description
Standard broom::tidy() and broom::glance() interfaces for an
object returned by table_continuous(). They re-shape the
underlying long-format data into the two canonical broom views so
the descriptive table can be consumed by any downstream
tidyverse-stats pipeline.
Usage
## S3 method for class 'spicy_continuous_table'
tidy(x, ...)
## S3 method for class 'spicy_continuous_table'
glance(x, ...)
Arguments
x |
A |
... |
Currently ignored. Present for compatibility with the
|
Details
tidy() returns one row per (variable x group) (or per
variable when by is not used) with broom-conventional columns:
outcome, label, group (when applicable), estimate (the
empirical mean), std.error (sd / sqrt(n)), conf.low,
conf.high (the mean confidence interval at ci_level), n,
min, max, sd. The outcome column carries the variable name
and label the human-readable label.
glance() returns one row per outcome with the omnibus group
comparison (when by is used), the requested effect size and the
balance diagnostic. Columns: outcome, label, test_type,
statistic, df, df.residual, p.value, es_type, es_value,
es_ci_lower, es_ci_upper, smd_type, smd_value, n_total.
Without by, only outcome, label, and n_total are populated;
the other columns are NA. The schema is fixed: smd_type /
smd_value are present whether or not smd = TRUE, like every
other comparison column here, and are NA when the table carries
no standardized mean difference. They sit before n_total, so
index this frame by NAME rather than by position.
glance.spicy_categorical_table() does not carry the SMD: each
family documents its own broom contract, and the categorical one
publishes the association measure alone. On a mixed balance table,
read the categorical SMD from output = "long".
Value
A tbl_df.
See Also
as.data.frame.spicy_continuous_table() for the raw
long-format access; tidy.spicy_continuous_lm_table() for the
model-based companion.
Tidying methods for a spicy_outcome_table
Description
Standard broom::tidy() and broom::glance() interfaces for an
object returned by table_outcome().
Usage
## S3 method for class 'spicy_outcome_table'
tidy(x, ...)
## S3 method for class 'spicy_outcome_table'
glance(x, ...)
Arguments
x |
A |
... |
Ignored, for S3 compatibility. |
Details
tidy() returns the DESCRIBED rows: the marginal Overall row and
one row per (grouping x level). Columns: outcome (the outcome
name, constant down the frame), variable (the grouping, or the
outcome itself on the marginal row), label, level (NA on the
marginal row), estimate (the mean), std.error (sd / sqrt(n)),
conf.low, conf.high, n, min, max, sd.
Two identity columns where the sibling has one, and deliberately:
here the outcome is fixed and the variable changes, so a single
outcome column would have to mean two different things down the
frame. The schema reads without knowing which function produced it.
glance() returns one row per grouping – one BLOCK – with that
block's own comparison. Columns: outcome, variable, label,
n_levels, test_type, statistic, df, df.residual,
p.value, es_type, es_value, es_ci_lower, es_ci_upper,
smd_type, smd_value, n_total.
n_levels counts the levels the TABLE displays, so a missing-value
display level counts; the comparison behind test_type runs on the
observed levels only, as it does everywhere in the family.
The schema is FIXED: smd_type / smd_value are present and NA
from the first version, so the day a standardized mean difference
enters this table it cannot break a pipeline that indexes the
frame. Index by NAME rather than by position.
Value
A tbl_df (or data.frame when tibble is not installed).
Tidy / glance methods for spicy_regression_table
Description
Standard broom::tidy() / broom::glance() interfaces for an
object returned by table_regression(). They re-shape the
underlying long-format data into the two canonical broom views
so the table can be consumed by any downstream tidyverse-stats
pipeline.
Usage
## S3 method for class 'spicy_regression_table'
tidy(x, ...)
## S3 method for class 'spicy_regression_table'
glance(x, ...)
Arguments
x |
A |
... |
Currently ignored. Present for compatibility with the
|
Details
tidy() returns one row per (model_id, term, estimate_type, outcome_level) combination, with estimate_type in
c("B", "beta", "ame", "partial_f2", "partial_eta2", "partial_omega2").
outcome_level names the response category of per-category rows
(ordinal / multinomial average marginal effects) and is NA for
single-outcome models. Reference-row placeholders (factor reference
levels) and singular coefficients (NA estimates) are dropped. Columns:
model_id, outcome, outcome_level, term, estimate_type, estimate, std.error, conf.low, conf.high, statistic, df, p.value, test_type, is_intercept, factor_term, factor_level.
glance() returns one row per (model_id, outcome) with
model-level statistics. Columns: model_id, outcome, nobs, weighted_nobs, r.squared, adj.r.squared, omega2, sigma, rmse, f2, AIC, AICc, BIC, deviance, df.residual (numeric –
Satterthwaite-safe).
Value
A tbl_df.
See Also
as.data.frame.spicy_regression_table() for the wide
raw view.
Uncertainty Coefficient
Description
uncertainty_coef() computes the Uncertainty Coefficient
(Theil's U) for a two-way contingency table, based on
information entropy.
Usage
uncertainty_coef(
x,
direction = c("symmetric", "row", "column"),
detail = FALSE,
conf_level = 0.95,
digits = 3L
)
Arguments
x |
A contingency table (of class |
direction |
Direction of prediction:
|
detail |
Logical. If |
conf_level |
A single number strictly between 0 and 1 giving
the confidence level (default |
digits |
Number of decimal places used when printing the
result (default |
Details
The uncertainty coefficient measures association using Shannon
entropy. Let H_X and H_Y be the marginal entropies
of the row and column variables respectively, and
H_{XY} the joint entropy.
-
direction = "row"(column predicts row):U = (H_X + H_Y - H_{XY}) / H_X. -
direction = "column"(row predicts column):U = (H_X + H_Y - H_{XY}) / H_Y. -
direction = "symmetric":U = 2 (H_X + H_Y - H_{XY}) / (H_X + H_Y).
The default direction = "symmetric" follows the SPSS and
DescTools convention: the symmetric coefficient is a standard,
well-defined variant with its own asymptotic standard error.
somers_d() deliberately differs (its default is "row")
because its symmetric form is a derived quantity without an
analytic SE; see its documentation.
When the marginal entropy in the denominator is zero (the
predicted variable is constant, e.g. an unused factor level),
the coefficient is the undefined form 0/0: the function
returns NA with a spicy_undefined_stat warning, like the
other measures in the family. For direction = "symmetric"
this happens only when both variables are constant; with a
single constant variable the symmetric coefficient is a
well-defined 0.
The entropy terms use the standard mathematical convention
0 \log 0 = 0, matching SPSS / PSPP CROSSTABS and the
definition in Cover & Thomas (2006). Note that
DescTools::UncertCoef() applies an additional Laplace
correction (replacing zero cells with 1/n^2) before the
entropy computation, which produces slightly different point
estimates on tables with empty cells; that correction is
uncommon in the information-theory literature and is not used
here. The asymptotic standard errors follow the DescTools delta
method; see cramer_v() for full references.
Value
Same structure as cramer_v(): a scalar when
detail = FALSE, a named vector when detail = TRUE.
The p-value tests H0: U = 0 (Wald z-test).
References
Theil, H. (1970). On the estimation of relationships involving qualitative variables. American Journal of Sociology, 76(1), 103-154. doi:10.1086/224909
See Also
lambda_gk(), goodman_kruskal_tau(), assoc_measures()
Other association measures:
assoc_measures(),
contingency_coef(),
cramer_v(),
gamma_gk(),
goodman_kruskal_tau(),
kendall_tau_b(),
kendall_tau_c(),
lambda_gk(),
phi(),
somers_d(),
yule_q()
Examples
tab <- table(sochealth$smoking, sochealth$education)
uncertainty_coef(tab)
uncertainty_coef(tab, direction = "row", detail = TRUE)
Generate a comprehensive summary of the variables
Description
varlist() lists the variables of a data frame and extracts
essential metadata: variable names, labels, summary values,
classes, number of distinct values, number of valid (non-missing)
observations, and number of missing values. Tidyselect-style
selectors can be supplied to pick or reorder columns dynamically.
vl() is a convenient shorthand for varlist() that offers identical
functionality with a shorter name.
Usage
varlist(
x,
...,
values = FALSE,
tbl = FALSE,
include_na = FALSE,
factor_levels = c("observed", "all"),
user_na = TRUE
)
vl(
x,
...,
values = FALSE,
tbl = FALSE,
include_na = FALSE,
factor_levels = c("observed", "all"),
user_na = TRUE
)
Arguments
x |
A data frame, or a transformation of one. |
... |
Optional tidyselect-style column selectors (e.g.
|
values |
Logical. If |
tbl |
Logical. If |
include_na |
Logical. If |
factor_levels |
Character. Controls how factor values are displayed
in |
user_na |
Logical. If |
Details
In an interactive session (RStudio, Positron, ...), the summary
opens in the Viewer pane with a contextual title like
vl: sochealth. If the data frame has been transformed or
subsetted, the title is suffixed with * (e.g. vl: sochealth*);
anonymous or ambiguous calls fall back to vl: <data>. Pass
tbl = TRUE to return a tibble instead.
The default factor_levels = "observed" mirrors what is actually
in the data; code_book() defaults to "all" to document the
declared schema. See @param factor_levels to override either
default.
Value
A tibble with one row per selected variable, containing the following columns:
-
Variable: variable names -
Label: variable labels (if available via thelabelattribute) -
Values: a summary of the variable's values, depending on thevaluesandinclude_naarguments. Ifvalues = FALSE, a compact summary is shown: all unique values when there are at most four, otherwise 3 + ... + last. Ifvalues = TRUE, all unique non-missing values are displayed. For labelled variables, prefixed labels are displayed usinglabelled::to_factor(levels = "prefixed"). For factors, levels are displayed according tofactor_levels. Matrix and array columns are summarized by their dimensions.difftimevalues are annotated with their units, e.g.1.5, 2.5 (hours). Missing value markers (<NA>,<NaN>) are optionally appended at the end (controlled viainclude_na). Literal strings"NA","NaN", and""are quoted to distinguish them from missing markers. -
Class: the class of each variable (possibly multiple, e.g."labelled", "numeric") -
N_distinct: number of distinct non-missing values -
N_valid: number of non-missing observations -
NAs: number of missing observations
For matrix and array columns, observations are counted per row:
a row is treated as missing if any of its cells is NA. N_valid
/ NAs therefore count complete vs. incomplete rows, not
individual cells.
With tbl = FALSE (the default) the tibble is sent to the
Viewer (interactive) or surfaced via a message (non-interactive)
and the function returns invisibly NULL. Set tbl = TRUE to
return the tibble directly for downstream use.
Declared missing values
Survey files imported with haven often carry declared missing
values: codes such as 8 = Don't know or 9 = Refused that the
source file marks as missing while keeping them distinct from a
plain NA. Two kinds of declaration exist: na_values / na_range
metadata on haven::labelled_spss() vectors, and tagged missing
values created by haven::tagged_na() (the Stata .a, .b, ...
convention).
spicy honors the declaration by default (user_na = TRUE):
declared missing values are excluded from every statistic exactly
like NA – valid percentages, means, chi-squared tests,
association measures, row-wise summaries, and group definitions –
but they are not erased from display. freq() lists each observed
declared value as its own row of the Missing block, with its value
label; cross_tab(), table_categorical(), and
table_continuous() disclose the exclusion in the table note
(Declared missing values removed: x (2).); varlist() and
code_book() count them as missing in N_valid / NAs /
N_distinct while still listing the declared codes in Values.
Every function involved offers the same escape hatch: set
user_na = FALSE to ignore the declaration and treat the declared
codes as valid values (the behavior of spicy before 0.13.0).
Tagged missing values are genuine NAs either way; for them,
user_na = FALSE only collapses the per-tag breakdown back into
the regular NA count.
See Also
Other variable inspection:
code_book(),
label_from_names()
Examples
varlist(sochealth, tbl = TRUE)
sochealth |> varlist(tbl = TRUE)
varlist(sochealth, where(is.numeric), values = TRUE, tbl = TRUE)
varlist(
sochealth,
starts_with("bmi"),
values = TRUE,
include_na = TRUE,
tbl = TRUE
)
df <- data.frame(
group = factor(c("A", "B", NA), levels = c("A", "B", "C"))
)
varlist(
df,
values = TRUE,
include_na = TRUE,
factor_levels = "all",
tbl = TRUE
)
vl(sochealth, tbl = TRUE)
sochealth |> vl(tbl = TRUE)
vl(sochealth, starts_with("bmi"), tbl = TRUE)
vl(sochealth, where(is.numeric), values = TRUE, tbl = TRUE)
Yule's Q
Description
yule_q() computes Yule's Q coefficient of association for a 2x2
contingency table.
Usage
yule_q(x, detail = FALSE, conf_level = 0.95, digits = 3L)
Arguments
x |
A contingency table (of class |
detail |
Logical. If |
conf_level |
A single number strictly between 0 and 1 giving
the confidence level (default |
digits |
Number of decimal places used when printing the
result (default |
Details
For a 2x2 table with cells a, b, c, d, Yule's Q is
Q = (ad - bc) / (ad + bc). It is equivalent to the
Goodman-Kruskal Gamma for 2x2 tables. The asymptotic standard
error is
SE = 0.5 (1 - Q^2) \sqrt{1/a + 1/b + 1/c + 1/d}.
Edge cases: when ad + bc = 0, Q itself is undefined and the
function returns NA with a spicy_undefined_stat warning.
When any cell is zero (and ad + bc > 0), Q is well-defined
but the SE formula divides by zero – the point estimate is
returned, and se, ci_lower, ci_upper, and p_value are
all NA.
Standard error formulas follow the DescTools implementations
(Signorell et al., 2024); see cramer_v() for full references.
Value
Same structure as cramer_v(): a scalar when
detail = FALSE, a named vector when detail = TRUE.
The p-value tests H0: Q = 0 (Wald z-test).
References
Yule, G. U. (1900). On the association of attributes in statistics. Philosophical Transactions of the Royal Society of London, Series A, 194, 257-319. doi:10.1098/rsta.1900.0019
See Also
phi(), gamma_gk(), assoc_measures()
Other association measures:
assoc_measures(),
contingency_coef(),
cramer_v(),
gamma_gk(),
goodman_kruskal_tau(),
kendall_tau_b(),
kendall_tau_c(),
lambda_gk(),
phi(),
somers_d(),
uncertainty_coef()
Examples
tab <- table(sochealth$smoking, sochealth$sex)
yule_q(tab)