Package {spicy}


Title: Publication-Ready Tables for Descriptive Statistics and Regression Models
Version: 0.13.0
Description: Provides publication-ready tables for descriptive statistics and regression models: frequency tables and cross-tabulations with association measures (Cramer's V, Kendall's Tau-b, and others), categorical and continuous summary tables, by group or from a complex survey design, and regression tables for one or more models side by side, across more than thirty model classes from mixed-effects to survival and Bayesian, with robust standard errors, average marginal effects, and univariable screening. Tables follow APA conventions by default, can switch to named journal styles such as JAMA, NEJM, or The Lancet, and render identically in the console and in 'gt', 'tinytable', 'flextable', 'Word', 'Excel', or the clipboard. Declared missing values in labelled data are honored and disclosed throughout the descriptive tables. Helpers cover codebooks, variable inspection, and row-wise summaries.
License: MIT + file LICENSE
URL: https://amaltawfik.github.io/spicy/, https://github.com/amaltawfik/spicy
BugReports: https://github.com/amaltawfik/spicy/issues
Encoding: UTF-8
Language: en-US
Imports: crayon, dplyr, labelled, rlang (≥ 1.1.0), sandwich (≥ 3.1-2), stats, stringr, tibble, tidyselect, utils
Suggests: bit64, boot, broom, car, clipr, clubSandwich, DT, effectsize, emmeans, estimatr (≥ 1.0.0), fixest (≥ 0.11), flexsurv (≥ 2.2), flextable, geepack (≥ 1.3.9), gt, haven, htmltools, insight, glmmTMB (≥ 1.1.7), knitr, lme4 (≥ 1.1-35), lmerTest (≥ 3.1-3), lmtest, merDeriv (≥ 0.2.6), marginaleffects, modelsummary, collapse, MASS, mgcv (≥ 1.9), nlme (≥ 3.1-160), nnet, numDeriv, ordinal (≥ 2022.11), AER (≥ 1.2), betareg (≥ 3.1), mlogit (≥ 1.1), pscl (≥ 1.5), quantreg (≥ 5.86), RLRsim, reformulas, rms (≥ 6.7), sampleSelection (≥ 1.2), posterior (≥ 1.5.0), rstanarm (≥ 2.21), brms (≥ 2.20), loo (≥ 2.6), rstan, survey (≥ 4.5), survival (≥ 3.5), officer, openxlsx2, parameters, performance, rmarkdown, spelling, testthat (≥ 3.0.0), tinytable, withr
VignetteBuilder: knitr
Depends: R (≥ 4.1.0)
Config/testthat/edition: 3
LazyData: true
Config/roxygen2/version: 8.1.0
NeedsCompilation: no
Packaged: 2026-09-30 23:11:30 UTC; at
Author: Amal Tawfik ORCID iD ROR ID [aut, cre, cph]
Maintainer: Amal Tawfik <amal.tawfik@hesav.ch>
Repository: CRAN
Date/Publication: 2026-10-01 08:10:02 UTC

spicy: publication-ready tables for descriptive statistics and regression models

Description

spicy provides publication-ready tables for descriptive statistics and regression models: frequency tables and cross-tabulations with association measures, categorical and continuous summary tables (by group or from a complex survey design), and regression tables for one or more models side by side, across more than thirty model classes, with robust standard errors, average marginal effects, and univariable screening. Tables follow APA conventions by default, can switch to named journal styles, and render identically in the console and the rich outputs. Declared missing values in labelled data are honored and disclosed throughout the descriptive tables; helpers cover codebooks, variable inspection, and row-wise summaries.

API stability

spicy is in active pre-1.0 development. Breaking changes are made deliberately at minor-version bumps and are always announced in NEWS.md. The API surface is partitioned as follows; users planning to embed spicy in production pipelines or downstream packages should rely on the stable surface.

Stable (signature and behaviour preserved across 0.y.z and into 1.0.0; documented changes only):

Stabilising (still maturing; argument names may be tightened before 1.0 with a NEWS.md entry, but no silent behavioural changes):

Experimental (new in this cycle; the shape of the output and the argument names may still move, with a NEWS.md entry, on their OWN clock rather than the parent family's):

Internal API (not part of the public surface; can change without notice – avoid calling directly from downstream code):

broom output shape

The broom::tidy() and broom::glance() methods on spicy_categorical_table, spicy_continuous_table, spicy_continuous_lm_table, spicy_continuous_svy_table, spicy_categorical_svy_table, and spicy_regression_table follow the standard broom column conventions (outcome, term, estimate, std.error, conf.low, conf.high, statistic, p.value, df, df.residual, r.squared, adj.r.squared, nobs, ...). The set of columns produced by each method is considered stabilising: existing columns will not be silently renamed or have their semantics changed within ⁠0.y.z⁠, and any breaking change is announced in NEWS.md. Adding optional new columns (e.g. covariate-adjustment metadata) is not a breaking change. Numeric columns keep the types downstream broom-consumers expect: test degrees of freedom that are integer by construction (chi-squared, factor-comparison F tests) stay integer, while every degrees-of-freedom column that can be fractional – df.residual, the regression method's per-coefficient df, and Welch-corrected test df – is numeric double, so Satterthwaite-corrected degrees of freedom from cluster-robust variance modes are preserved verbatim (matching lmerTest::glance() and the afex output convention).

Classed conditions

All errors and warnings emitted by the stable / stabilising surfaces carry classed conditions so downstream code can dispatch on class via tryCatch() / withCallingHandlers() instead of matching message strings. Each condition has a package-wide parent class plus a leaf class describing the specific cause:

spicy_error

Catch-all parent for every error raised by spicy. Leaves:

  • spicy_invalid_input – bad argument value or type.

  • spicy_invalid_data – bad data shape or content (not a data.frame, length mismatch, bit64::integer64 columns, degenerate grouping).

  • spicy_missing_pkg – a Suggests dependency is required by the requested operation but not installed.

  • spicy_missing_column – a referenced column is not in data.

  • spicy_unsupported – the operation is not applicable to this input (e.g., Phi requested on a non-2x2 table).

  • spicy_ame_satt_unsupported_formula – signaled together with spicy_unsupported when AME Satterthwaite degrees of freedom are unavailable for the model's formula structure; normally caught internally and surfaced as a spicy_fallback warning.

  • spicy_unsupported_class – the model class has no regression-frame method, so table_regression() cannot render it.

  • spicy_unsupported_vcov – the requested vcov mode is not available for this model class.

  • spicy_unsupported_standardized – the requested standardized mode is not available for this model class.

  • spicy_invalid_frame – an object failed the structural validation of the internal regression-frame contract behind table_regression().

  • spicy_resampling_failed – bootstrap / jackknife resampling produced too few valid replicates to estimate the requested statistic.

  • spicy_defunct – an argument removed in a pre-1.0 hard break; the message names the replacement. Signaled together with spicy_invalid_input so generic input handlers still catch it.

  • spicy_internal – an internal precondition failed; this is a bug in spicy, please report it.

  • spicy_internal_invariant – an internal consistency check on a spicy-built object failed and the result cannot be trusted (see the warning leaf of the same name for the renderable case).

spicy_warning

Catch-all parent for every warning. Leaves:

  • spicy_undefined_stat – the requested statistic is undefined for this input; result is NA (e.g., Tau-b on a table with all-zero marginals).

  • spicy_negative_weights_no_test – signaled together with spicy_undefined_stat when table_continuous_svy() or table_categorical_svy() withholds a group comparison because the analytic sample carries negatively weighted rows; the estimates are still reported and the table note says so.

  • spicy_dropped_na – NA observations were silently excluded from the computation (e.g., NA weights).

  • spicy_ignored_arg – an argument was ignored due to context (e.g., correct = TRUE on a non-2x2 table).

  • spicy_no_selection – a column selector produced an empty set; an empty result is returned rather than erroring.

  • spicy_fallback – the requested computation failed; a simpler estimator was used instead.

  • spicy_caveat – the computation succeeded but its interpretation carries a non-trivial methodological caveat (e.g., standardized coefficients on non-additive terms).

  • spicy_bayes_diagnostics – signaled together with spicy_caveat when a Bayesian fit's sampler or predictive-accuracy diagnostics miss their targets (R-hat, ESS, divergences, E-BFMI, Pareto k, p_waic).

  • spicy_nonconvergence – signaled together with spicy_caveat when the fitting engine reports that its optimizer did not converge. The table still shows the numbers the model object holds, and a footer note says what they are worth.

  • spicy_model_choice – a defaulted modeling choice was made on the user's behalf and is disclosed (e.g., the linear probability model table_regression_uv() fits to a binary outcome under the default method).

  • spicy_passthrough – a third-party warning captured during an operation (e.g., the clipboard copy) and re-emitted under the spicy taxonomy.

  • spicy_summary_failed – varlist() could not summarise one column; the rest of the table is fine.

  • spicy_renamed_column – a user data column or factor level collided with a spicy-internal name and was auto-renamed to preserve the data (emitted by cross_tab()).

  • spicy_internal_invariant – an internal consistency check on a spicy-built object failed but the output still renders, so the user sees both the table and the diagnostic.

spicy_info

Parent for informational messages (emitted via rlang::inform(); muffle with withCallingHandlers(spicy_info = ...)). Leaf: spicy_silent_reference – reference levels are displayed nowhere under reference_style = "none" with factor_layout = "flat". The once-per-session hint on ordered-factor polynomial contrasts carries its own class, spicy_polynomial_contrasts_info.

Author(s)

Maintainer: Amal Tawfik amal.tawfik@hesav.ch (ORCID) (ROR) [copyright holder]

Authors:

See Also

Useful links:


Coerce a spicy_categorical_svy_table to a data frame or tibble

Description

These S3 methods strip the "spicy_categorical_svy_table" class and the rendering-only attributes, keeping the wide compute frame and the three provenance markers (group_var, design_meta, note).

Usage

## S3 method for class 'spicy_categorical_svy_table'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)

## S3 method for class 'spicy_categorical_svy_table'
as_tibble(x, ...)

Arguments

x

A spicy_categorical_svy_table returned by table_categorical_svy().

row.names, optional, ...

Passed on for method compatibility; ignored.

Value

A plain data.frame (or a tbl_df for as_tibble()).

See Also

table_categorical_svy().


Coerce a spicy_categorical_table to a plain data frame or tibble

Description

These S3 methods strip the "spicy_categorical_table" / "spicy_table" classes and the rendering-only attributes (display_df, indent_text, align, decimal_mark, long_data, ...) from an object returned by table_categorical() so the underlying wide-format data can be manipulated with downstream tools (dplyr, tidyr, etc.) under the standard data.frame / tbl_df contract. The single attribute "group_var" is preserved as a lightweight provenance marker; all other spicy attributes are dropped. The original x is unaffected, and print(x) continues to render the formatted ASCII table.

Usage

## S3 method for class 'spicy_categorical_table'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)

## S3 method for class 'spicy_categorical_table'
as_tibble(x, ...)

Arguments

x

A spicy_categorical_table returned by table_categorical().

row.names, optional

Standard base::as.data.frame() arguments. Currently ignored.

...

Further arguments passed to tibble::as_tibble() (for the tibble method) or ignored (for the as.data.frame() method).

Details

The returned data is the wide raw representation (one row per ⁠(variable x level)⁠, group columns side by side). For the tidy long format – one row per ⁠(variable x level x group)⁠ – use tidy.spicy_categorical_table() or call table_categorical() directly with output = "long".

Value

A plain data.frame (or tbl_df) with the same rows and columns as the wide raw output of table_categorical().

See Also

tidy.spicy_categorical_table(), glance.spicy_categorical_table().


Coerce a spicy_continuous_lm_table to a plain data frame or tibble

Description

These S3 methods strip the "spicy_continuous_lm_table" / "spicy_table" classes and the rendering-only attributes (digits, decimal_mark, ci_level, ...) from an object returned by table_continuous_lm() so the underlying long-format data can be manipulated with downstream tools (dplyr, tidyr, etc.) under the standard data.frame / tbl_df contract. The single attribute "by_var" is preserved as a lightweight provenance marker; all other spicy attributes are dropped. The original x is unaffected, and print(x) continues to render the formatted ASCII table.

Usage

## S3 method for class 'spicy_continuous_lm_table'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)

## S3 method for class 'spicy_continuous_lm_table'
as_tibble(x, ...)

Arguments

x

A spicy_continuous_lm_table returned by table_continuous_lm().

row.names, optional

Standard base::as.data.frame() arguments. Currently ignored: the long format already carries integer row names and explicit columns.

...

Further arguments passed to tibble::as_tibble() (for the tibble method) or ignored (for the as.data.frame() method).

Value

A plain data.frame (or tbl_df) with the same rows and columns as the long output of table_continuous_lm().

See Also

tidy.spicy_continuous_lm_table(), glance.spicy_continuous_lm_table() for cleaner broom-style pivots tailored to downstream pipelines.


Coerce a spicy_continuous_svy_table to a plain data frame or tibble

Description

These S3 methods strip the "spicy_continuous_svy_table" class and the rendering-only attributes, keeping the long compute frame and the three provenance markers (group_var, design_meta, note).

Usage

## S3 method for class 'spicy_continuous_svy_table'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)

## S3 method for class 'spicy_continuous_svy_table'
as_tibble(x, ...)

Arguments

x

A spicy_continuous_svy_table returned by table_continuous_svy().

row.names, optional, ...

Passed on for method compatibility; ignored.

Value

A plain data.frame (or a tbl_df for as_tibble()).

See Also

table_continuous_svy().


Coerce a spicy_continuous_table to a plain data frame or tibble

Description

These S3 methods strip the "spicy_continuous_table" / "spicy_table" classes and the rendering-only attributes (digits, decimal_mark, ci_level, align, p_digits, ...) from an object returned by table_continuous() so the underlying long-format data can be manipulated with downstream tools (dplyr, tidyr, etc.) under the standard data.frame / tbl_df contract. The single attribute "group_var" is preserved as a lightweight provenance marker; all other spicy attributes are dropped. The original x is unaffected, and print(x) continues to render the formatted ASCII table.

Usage

## S3 method for class 'spicy_continuous_table'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)

## S3 method for class 'spicy_continuous_table'
as_tibble(x, ...)

Arguments

x

A spicy_continuous_table returned by table_continuous().

row.names, optional

Standard base::as.data.frame() arguments. Currently ignored: the long format already carries integer row names and explicit columns.

...

Further arguments passed to tibble::as_tibble() (for the tibble method) or ignored (for the as.data.frame() method).

Details

The returned data is identical to what output = "long" (or output = "data.frame") returns directly from table_continuous(); use whichever entry point reads better in your pipeline.

Value

A plain data.frame (or tbl_df) with one row per ⁠(variable x group)⁠ (or one row per variable when by is not used).

See Also

tidy.spicy_continuous_table(), glance.spicy_continuous_table() for cleaner broom-style pivots tailored to downstream pipelines.


Coerce a spicy_outcome_table to a plain data frame or tibble

Description

These S3 methods strip the "spicy_outcome_table" class and the rendering-only attributes from an object returned by table_outcome(), so the underlying long-format data can be manipulated with downstream tools under the standard data.frame / tbl_df contract. The "outcome" and "select" attributes are kept as lightweight provenance markers. The original x is unaffected, and print(x) continues to render the formatted table.

Usage

## S3 method for class 'spicy_outcome_table'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)

## S3 method for class 'spicy_outcome_table'
as_tibble(x, ...)

Arguments

x

A spicy_outcome_table returned by table_outcome().

row.names, optional

Standard base::as.data.frame() arguments, currently ignored.

...

Further arguments passed to tibble::as_tibble() (for the tibble method) or ignored.

Details

The returned data is identical to what output = "long" (or output = "data.frame") returns directly from table_outcome().

Value

A plain data.frame (or tbl_df), one row per displayed row of the table.

See Also

tidy.spicy_outcome_table(), glance.spicy_outcome_table().


Convert a spicy_regression_table to a plain data.frame / tibble

Description

Strips the spicy_regression_table / spicy_table classes and the internal analytic attributes (spicy_long, spicy_fit_stats), returning the wide character display as a plain data.frame (or tbl_df via as_tibble()). The title, note, provenance (model_ids, outcome), and rendering (col_spec) attributes are preserved.

Usage

## S3 method for class 'spicy_regression_table'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)

## S3 method for class 'spicy_regression_table'
as_tibble(x, ...)

Arguments

x

A spicy_regression_table returned by table_regression().

row.names, optional

Standard as.data.frame() arguments (currently ignored – the table's row layout is preserved).

...

Currently ignored.

Details

as.data.frame() is equivalent to passing output = "data.frame" to table_regression(): the two paths return identical objects (same cells, classes, and attributes).

Value

A plain data.frame (for as.data.frame()) or a tbl_df (for as_tibble()).

See Also

tidy.spicy_regression_table(), glance.spicy_regression_table() for broom-canonical long views.


Convert a spicy flextable output to a plain flextable

Description

table_regression() and table_continuous_lm() return their output = "flextable" tables with a lightweight spicy_flextable class tag whose only job is HTML note styling. Every flextable verb already works on the tagged object; this method returns the clean underlying flextable::flextable() – the note is already part of its footer – for workflows that want the untagged object (e.g. custom knit hooks, flextable::save_as_docx() pipelines, or composition with other flextable tooling), mirroring gtsummary::as_flex_table().

Usage

as_flextable.spicy_flextable(x, ...)

Arguments

x

A spicy_flextable object.

...

Unused; for generic consistency.

Details

Rendering in Quarto / R Markdown does NOT require this conversion: the knit_print method auto-detects the output format and delegates to flextable's native rendering for non-HTML targets (Word, PowerPoint, PDF).

Value

A flextable object.


Extract the typed (structured) view of a spicy table

Description

spicy's tables return a display representation by default – a character data.frame with stars suffixes, en-dash for reference rows, bracketed "[L, U]" confidence intervals, and APA padding on p-values. This accessor returns the typed view that the output engines (Excel, gt, tinytable, flextable, clipboard) consume internally: a fully numeric body with CI pre-split into LL / UL columns, NAs for non-applicable / reference cells, plus per-cell markers and a format specification.

Usage

as_structured(x)

Arguments

x

A spicy table built with output = "default": table_regression(), table_categorical(), table_continuous() or table_continuous_lm(). All four return the same schema.

Details

This is the right entry point for users who want to:

Value

A list with the structured view (see Details for the schema).

Schema

Descriptive tables

The three descriptive families return the same schema; only the col_meta tokens and the row roles they emit differ.

Cells the console builds from more than one number – a "Med [Q1, Q3]", a test gloss, an effect size with its interval – keep their numeric anchor in body and their exact printed string in ⁠col_meta$<col>$display_cells⁠. stars is always NULL: descriptive tables carry no significance markers.

Versioning

version says which contract an object carries. Version 3 moved row identity out of index vectors and into the body itself:

Index vectors are the structure that corrupts as soon as two bodies are stacked or merged, so they were removed rather than kept alongside. An object carrying an older contract (or none) is refused with the correspondence above; rebuild it with the function that produced it. An object carrying a newer contract than the spicy reading it is refused as well.

Version 3 also opened the accessor to the descriptive families, which had no typed view before: .row_role gained "summary", "group" and "missing". The vocabulary is extended by addition – an existing role never changes meaning.

See Also

table_regression(), table_categorical(), table_continuous(), table_continuous_lm() for the user-facing entry points.

Examples

fit <- lm(mpg ~ wt + factor(cyl), data = mtcars)
tbl <- table_regression(fit)
s <- as_structured(tbl)
s$body                               # raw numeric body
s$body[which(s$body$p < 0.05), ]     # filter significant rows
# which() drops the structural NA rows (headers, reference levels)
s$body$.row_role                     # what each row is
s$body[s$body$.variable == "wt", ]   # address a row by its variable
s$col_meta$B                         # column metadata for B

# The same schema on a descriptive table.
ct <- table_categorical(mtcars, c(cyl, gear), by = am)
sc <- as_structured(ct)
sc$body[, c("Variable", ".variable", ".level", ".row_role")]
sc$spanners                          # one per `by` group


Association measures summary table

Description

assoc_measures() computes a range of association measures for a two-way contingency table and returns them in a tidy data frame.

Usage

assoc_measures(
  x,
  type = c("all", "nominal", "ordinal"),
  conf_level = 0.95,
  digits = 3L
)

Arguments

x

A contingency table (of class table).

type

Which family of measures to compute: "all" (default), "nominal", or "ordinal".

conf_level

A single number strictly between 0 and 1 giving the confidence level (default 0.95). Set to NULL to omit the confidence interval. Any other value – including percentages such as 95 – raises a classed error (spicy_invalid_input).

digits

Number of decimal places used when printing the result (default 3).

Details

type = "all" (the default) returns all nominal and ordinal measures. Use type = "nominal" or type = "ordinal" to restrict the output to a single family.

The nominal family includes cramer_v(), contingency_coef(), lambda_gk(), goodman_kruskal_tau(), uncertainty_coef(), and (for 2x2 tables) phi() and yule_q().

The ordinal family includes gamma_gk(), kendall_tau_b(), kendall_tau_c(), and somers_d().

Measures that are undefined on the given table appear as NA rows (printed as an en dash). The classed warnings the individual functions raise (e.g. spicy_undefined_stat) are re-emitted once per distinct message after the table is assembled, so condition handlers and suppressWarnings() behave as they do for the individual functions.

Standard error formulas follow the DescTools implementations (Signorell et al., 2024), except for Kendall's Tau-b, whose ASE follows Brown and Benedetti (1977) as printed by SPSS / PSPP CROSSTABS; see kendall_tau_b().

Value

A data frame with columns measure, estimate, se, ci_lower, ci_upper, and p_value. The p_value comes from two test families:

Direction-dependent measures (lambda_gk(), goodman_kruskal_tau(), uncertainty_coef(), somers_d()) contribute one row per direction (symmetric / R|C / C|R where applicable), so the output has more rows than the number of helper functions.

References

Agresti, A. (2002). Categorical Data Analysis (2nd ed.). Wiley.

Brown, M. B., & Benedetti, J. K. (1977). Sampling behavior of tests for correlation in two-way contingency tables. Journal of the American Statistical Association, 72(358), 309-315. doi:10.1080/01621459.1977.10480995

Liebetrau, A. M. (1983). Measures of Association. Sage.

Signorell, A. et al. (2024). DescTools: Tools for Descriptive Statistics. R package.

See Also

cramer_v(), gamma_gk(), kendall_tau_b()

Other association measures: contingency_coef(), cramer_v(), gamma_gk(), goodman_kruskal_tau(), kendall_tau_b(), kendall_tau_c(), lambda_gk(), phi(), somers_d(), uncertainty_coef(), yule_q()

Examples

tab <- table(sochealth$smoking, sochealth$education)
assoc_measures(tab)
assoc_measures(tab, type = "nominal")
assoc_measures(tab, type = "ordinal")


Build a formatted ASCII table

Description

Low-level rendering engine that constructs a visually aligned ASCII table from a data.frame. Supports Unicode line-drawing characters, ANSI colors via crayon, automatic colored-text-aware width detection, configurable padding, and per-column alignment.

Usage

build_ascii_table(
  x,
  padding = 2L,
  first_column_line = TRUE,
  row_total_line = TRUE,
  bottom_line = FALSE,
  lines_color = "darkgrey",
  align_left_cols = c(1L, 2L),
  align_center_cols = integer(0),
  center_headers = FALSE,
  spanners = NULL,
  group_sep_rows = integer(0),
  total_row_idx = NULL,
  display_labels = NULL,
  ...
)

Arguments

x

A data.frame or spicy_table object containing the table to format. Typically, this includes columns such as Category, Values, Freq., Percent, etc.

padding

Non-negative integer giving the number of extra characters added to each column's auto-computed width (the maximum of the cell-content width and the header width). Defaults to 2L, which gives a Stata- / cli-like compact look. Each cell additionally receives a one-space gutter on each side, so a padding = 2L column whose content is at most 5 characters wide occupies 9 characters in total (1 + 5 + 2 + 1).

The string choices "compact", "normal" and "wide" from spicy ⁠< 0.11.0⁠ were removed; pass 0L, 2L or 4L instead. Passing a string raises an actionable error.

first_column_line

Logical. If TRUE (the default), a vertical separator is drawn after the first column (useful for separating categories from data).

row_total_line

Logical. Controls the horizontal rule drawn before a total row. Defaults to TRUE.

bottom_line

Logical. If FALSE (the default), no closing line is drawn. If TRUE, draws a closing line at the bottom of the table.

lines_color

Character. Color used for table separators. Defaults to "darkgrey". The color is applied only when ANSI color support is available (see crayon::has_color()).

align_left_cols

Integer vector of column indices to left-align. Defaults to c(1L, 2L) (the layout used by freq()-style tables); pass an explicit vector for other layouts.

align_center_cols

Integer vector of column indices to center-align. Defaults to integer(0) (no centered columns). Columns not in align_left_cols or align_center_cols are right-aligned.

center_headers

Logical. When TRUE, column headers are centered above their column content even when the data itself is right-aligned (the publication convention for coefficient / summary tables; matches Stata regress and SPSS REGRESSION output). Left-aligned columns (per align_left_cols) keep their header on the left. Defaults to FALSE for backward compatibility; the print.spicy_regression_table method enables it.

spanners

Optional named list defining a column group row drawn above the column headers (the "spanner" / "supra-header" convention; APA Manual 7 §7.13). Names are spanner labels; values are integer vectors of 1-based column indices the label spans (must be contiguous). A thin underline rule is drawn below each spanner across its span. Used by print.spicy_regression_table() to display the model name above each model's block of sub-columns. Defaults to NULL (no spanner row).

group_sep_rows

Integer vector of row indices before which a light dashed separator line is drawn. Defaults to integer(0).

total_row_idx

Optional integer vector of 1-based row indices identifying the totals rows; a horizontal rule is drawn just before each. When NULL (the default), falls back to a regex match on "Total" / "Column_Total" in the formatted row text, which can mis-fire if a user category is literally named "Total" or "Sub-Total". Cross-tabs and frequency tables built by cross_tab() and freq() set this attribute on their result so the print methods are immune to that false positive.

display_labels

Optional character vector of length ncol(x) used to override colnames(x) for the rendered header row only. The data.frame's actual names are kept for indexing; only the visual header is swapped. Used by print.spicy_regression_table() so a B + AME table shows the bare label (⁠95% CI⁠, p) in both blocks rather than R's deduplicated ⁠95% CI.2⁠ / p.2. Defaults to NULL (use colnames(x) verbatim).

...

Additional arguments (currently ignored).

Details

Most users should not call this directly: it is wrapped by spicy_print_table() and the internal ⁠print.spicy_*⁠ methods, which add titles, notes, and table-type-aware alignment defaults. Reach for build_ascii_table() only when you need to render an arbitrary data.frame to a string with the same look as spicy's tables.

Value

A single character string containing the full ASCII-formatted table, suitable for direct printing with cat().

See Also

spicy_print_table() for a user-facing wrapper that adds titles and notes.


Generate an interactive variable codebook

Description

code_book() creates an interactive and exportable codebook summarizing selected variables of a data frame. It builds upon varlist() to provide an overview of variable names, labels, classes, and representative values in a sortable, searchable table.

The output is displayed as an interactive DT::datatable() in the Viewer pane (for example in RStudio or Positron), allowing searching, sorting, and export (copy, print, CSV, Excel, PDF) directly.

Usage

code_book(
  x,
  ...,
  values = FALSE,
  include_na = FALSE,
  title = "Codebook",
  filename = NULL,
  factor_levels = c("all", "observed"),
  user_na = TRUE
)

Arguments

x

A data frame or tibble.

...

Optional tidyselect-style column selectors (e.g. starts_with("var"), where(is.numeric), etc.). Columns can be selected or reordered, but renaming selections is not supported.

values

Logical. If FALSE (the default), displays a compact summary of the variable's values. For numeric, character, date/time, labelled, and factor variables, all unique non-missing values are shown when there are at most four; otherwise the first three values, an ellipsis (...), and the last value are shown. Values are sorted when appropriate (e.g., numeric, character, date). For factors, factor_levels controls whether observed or all declared levels are shown; level order is preserved. For labelled variables, prefixed labels are displayed via labelled::to_factor(levels = "prefixed"). If TRUE, all unique non-missing values are displayed.

include_na

Logical. If TRUE, unique missing value markers (⁠<NA>⁠, ⁠<NaN>⁠) are explicitly appended at the end of the Values summary when present in the variable. This applies to all variable types. Literal strings "NA", "NaN", and "" are quoted to distinguish them from missing markers. If FALSE (the default), missing values are omitted from Values but still counted in the NAs column.

title

Optional character string displayed as the table caption. Defaults to "Codebook". Set to NULL to remove the title completely. When filename = NULL, the title is also used as the base for export filenames after conversion to a portable ASCII name.

filename

Optional character string used as the base for exported CSV, Excel, and PDF filenames. If NULL (the default), a portable filename is derived from title, falling back to "Codebook" when needed. File extensions are added by the browser/export engine.

factor_levels

Character. Controls how factor values are displayed in Values. "all" (the default; varlist() uses "observed") shows all declared levels, including unused levels. "observed" shows only levels present in the data, preserving factor level order.

user_na

Logical. If TRUE (the default), declared missing values count as missing in N_valid, NAs, and N_distinct (all three columns share one missing definition). If FALSE, they count as valid. Either way, the declared codes remain listed in Values (with their value labels when declared) – a codebook documents the full coding scheme. See the "Declared missing values" section of freq().

Details

Value

A DT::datatable object.

Dependencies

Requires the following package:

Declared missing values

Survey files imported with haven often carry declared missing values: codes such as ⁠8 = Don't know⁠ or 9 = Refused that the source file marks as missing while keeping them distinct from a plain NA. Two kinds of declaration exist: na_values / na_range metadata on haven::labelled_spss() vectors, and tagged missing values created by haven::tagged_na() (the Stata .a, .b, ... convention).

spicy honors the declaration by default (user_na = TRUE): declared missing values are excluded from every statistic exactly like NA – valid percentages, means, chi-squared tests, association measures, row-wise summaries, and group definitions – but they are not erased from display. freq() lists each observed declared value as its own row of the Missing block, with its value label; cross_tab(), table_categorical(), and table_continuous() disclose the exclusion in the table note (⁠Declared missing values removed: x (2).⁠); varlist() and code_book() count them as missing in N_valid / NAs / N_distinct while still listing the declared codes in Values.

Every function involved offers the same escape hatch: set user_na = FALSE to ignore the declaration and treat the declared codes as valid values (the behavior of spicy before 0.13.0). Tagged missing values are genuine NAs either way; for them, user_na = FALSE only collapses the per-tag breakdown back into the regular NA count.

See Also

varlist() for generating the underlying variable summaries.

Other variable inspection: label_from_names(), varlist()

Examples

## Not run: 
if (requireNamespace("DT", quietly = TRUE)) {
  code_book(sochealth)
  code_book(sochealth, starts_with("bmi"))
  code_book(sochealth, starts_with("bmi"), values = TRUE, include_na = TRUE)

  factors <- data.frame(
    group = factor(c("A", "B", NA), levels = c("A", "B", "C"))
  )
  code_book(
    factors,
    values = TRUE,
    include_na = TRUE,
    factor_levels = "observed"
  )

  code_book(
    sochealth,
    starts_with("bmi"),
    title = "BMI codebook",
    filename = "bmi_codebook"
  )
}

## End(Not run)


Pearson's contingency coefficient

Description

contingency_coef() computes Pearson's contingency coefficient C for a two-way contingency table.

Usage

contingency_coef(x, detail = FALSE, conf_level = 0.95, digits = 3L)

Arguments

x

A contingency table (of class table).

detail

Logical. If FALSE (default), return the estimate as a numeric scalar. If TRUE, return a named numeric vector including confidence interval and p-value.

conf_level

A single number strictly between 0 and 1 giving the confidence level (default 0.95). Only used when detail = TRUE. Set to NULL to omit the confidence interval. Any other value – including percentages such as 95 – raises a classed error (spicy_invalid_input).

digits

Number of decimal places used when printing the result (default 3). Only affects the detail = TRUE output.

Details

The contingency coefficient is C = \sqrt{\chi^2 / (\chi^2 + n)}. It ranges from 0 (independence) to a maximum that depends on the table dimensions. No standard asymptotic standard error exists, so the confidence interval is not computed.

Value

Same structure as cramer_v(): a scalar when detail = FALSE, a named vector when detail = TRUE. The p-value tests the null hypothesis of no association (Pearson chi-squared test). CI values are NA because no standard asymptotic SE exists for C.

References

Liebetrau, A. M. (1983). Measures of Association. Sage.

See Also

cramer_v(), assoc_measures()

Other association measures: assoc_measures(), cramer_v(), gamma_gk(), goodman_kruskal_tau(), kendall_tau_b(), kendall_tau_c(), lambda_gk(), phi(), somers_d(), uncertainty_coef(), yule_q()

Examples

tab <- table(sochealth$smoking, sochealth$education)
contingency_coef(tab)


Copy data to the clipboard

Description

Copies a data.frame, matrix, 2D or higher array, table, or atomic vector to the system clipboard, ready to paste into a text editor, spreadsheet, or word processor. Wraps clipr::write_clip() (a Suggests dependency); requires clipr to be installed and a clipboard backend to be available on the platform.

Usage

copy_clipboard(
  x,
  row_names_as_col = FALSE,
  row_names = TRUE,
  col_names = TRUE,
  show_message = TRUE,
  quiet = FALSE,
  ...
)

Arguments

x

A data.frame, matrix, 2D array, 3D array, table, or atomic vector to be copied.

row_names_as_col

Logical or character. If FALSE (the default), row names are not added as a column. If TRUE, a column named "rownames" is prepended. If a character string, it is used as the column name for the promoted row names. Ignored (with a warning) when x is neither a data.frame nor a strict matrix.

row_names

Logical. If TRUE (the default), row names are included in the clipboard output; FALSE omits them.

col_names

Logical. If TRUE (the default), column names are included in the clipboard output; FALSE omits them.

show_message

Logical. If TRUE (the default), prints a success message after copying.

quiet

Logical. If FALSE (the default), messages are shown. If TRUE, suppresses all messages, including the success message, coercion notices, and warnings.

...

Additional arguments passed to clipr::write_clip(). The pre-0.13.0 dot.case argument names (row.names.as.col, row.names, col.names) are trapped here and raise an error naming their snake_case replacements.

Details

Objects that are not data.frames or 2D matrices (atomic vectors, arrays, tables) are automatically coerced to character on the way to the clipboard, as required by clipr::write_clip(). The caller's object is never modified in place; transformations happen on a local copy (see the return value).

Multidimensional arrays (3D and higher) are flattened to a 1D character vector with one element per line. To preserve a tabular layout, extract a 2D slice first, e.g. copy_clipboard(my_array[, , 1]).

Messages and warnings raised by the clipboard backend are re-emitted as regular R conditions, so suppressMessages() / suppressWarnings() work as usual; quiet = TRUE silences them all at once.

Value

Invisibly returns the object as it was sent to the clipboard: identical to x by default, but reflecting the row_names_as_col transformation when one was requested (e.g. a matrix comes back as a data.frame with the promoted row-name column). The function is called for its clipboard side effect.

Examples

# Writes to the system clipboard, so never run by checks:
## Not run: 
if (clipr::clipr_available()) {
  # Data frame
  copy_clipboard(sochealth)

  # Data frame with row names as column
  copy_clipboard(head(sochealth), row_names_as_col = "id")

  # Matrix
  mat <- matrix(1:6, nrow = 2)
  copy_clipboard(mat)

  # Table
  tbl <- table(sochealth$education)
  copy_clipboard(tbl)

  # Array (3D) -- flattened to character
  arr <- array(1:8, dim = c(2, 2, 2))
  copy_clipboard(arr)

  # Recommended: copy 2D slice for tabular layout
  copy_clipboard(arr[, , 1])

  # Numeric vector
  copy_clipboard(c(3.14, 2.71, 1.618))

  # Character vector
  copy_clipboard(c("apple", "banana", "cherry"))

  # Quiet mode (no messages shown)
  copy_clipboard(sochealth, quiet = TRUE)
}

## End(Not run)

Row-wise count of specific or special values

Description

Counts, for each row of a data.frame or matrix, how many times one or more values appear across selected columns. Supports type-safe comparison (allow_coercion = FALSE), case-insensitive string matching (ignore_case = TRUE), and detection of special values (NA, NaN, Inf, -Inf) via special. Designed to flow inside dplyr::mutate() pipelines.

Usage

count_n(
  data = NULL,
  select = tidyselect::everything(),
  exclude = NULL,
  count = NULL,
  special = NULL,
  allow_coercion = TRUE,
  ignore_case = FALSE,
  regex = FALSE,
  verbose = FALSE,
  user_na = TRUE
)

Arguments

data

A data.frame or matrix. Optional inside dplyr::mutate(), where the current data context is used automatically.

select

Columns to include. Defaults to tidyselect::everything(). Uses tidyselect helpers like tidyselect::starts_with(), etc.; a character vector of names is validated with tidyselect::all_of(), so unknown names raise an error (as in mean_n() and sum_n()). If regex = TRUE, select is treated as a regex string.

exclude

Columns to exclude after selection (names or positions, as accepted by tidyselect::any_of()). Defaults to NULL (no exclusion).

count

Value(s) to count. Defaults to NULL. Ignored if special is used. Multiple values are allowed (e.g., count = c(1, 2, 3) or count = c("yes", "no")). R automatically coerces all values in count to a common type (e.g., c(2, "2") becomes c("2", "2")), so all values are expected to be of the same final type. If allow_coercion = FALSE, matching is type-safe using identical(), and the type of count must match that of the values in the data. A zero-length count (e.g. an upstream filter that emptied it) raises a classed error rather than silently counting nothing, and so does a count made only of missing values (use special = "NA" / "NaN").

special

Character vector of special values to count: "NA", "NaN", "Inf", "-Inf", or "all". Defaults to NULL. Every entry is validated, including alongside "all": any other value raises an error, and so does an empty vector (special = character(0) selects nothing to count). "NA" uses is.na(), and therefore includes both NA and NaN values. "NaN" uses is.nan() to match only actual NaN values.

allow_coercion

Logical. If TRUE (the default), values are compared after coercion. If FALSE, uses strict matching via identical().

ignore_case

Logical. If FALSE (the default), comparisons are case-sensitive. If TRUE, performs case-insensitive string comparisons.

regex

Logical. If FALSE (the default), uses tidyselect helpers. If TRUE, interprets select as a regular expression pattern.

verbose

Logical. If FALSE (the default), messages are suppressed. If TRUE, prints processing messages.

user_na

Logical. If TRUE (the default), special = "NA" counts declared missing values together with regular NA (the SPSS MISSING / NMISS convention). If FALSE, only genuine NA / NaN count as missing. The ⁠count =⁠ path matches the underlying codes in both modes, so explicitly listed declared codes are always counted. See the "Declared missing values" section of freq().

Value

A numeric vector of row-wise counts (unnamed), of length nrow(data). Missing values never match a regular count value, so an all-NA row counts 0 unless special targets missing values. If the selection resolves to zero usable columns, a classed warning (spicy_no_selection) is emitted and NA is returned for all rows, as in mean_n() and sum_n().

Strict matching (allow_coercion = FALSE)

Comparison falls back to identical() when types differ, which also inspects factor levels. Two consequences:

Case-insensitive matching (ignore_case = TRUE)

All values are converted to lowercase via tolower() before matching; factor columns are first coerced to character. This mode takes precedence over allow_coercion: equality becomes lowercase string equality, so "b" and "B" match even when allow_coercion = FALSE.

Coercion of count itself

R coerces mixed-type vectors at construction time: count = c(2, "2") becomes c("2", "2") before the function ever sees it. To get type-sensitive matching, keep count homogeneous.

Declared missing values

Survey files imported with haven often carry declared missing values: codes such as ⁠8 = Don't know⁠ or 9 = Refused that the source file marks as missing while keeping them distinct from a plain NA. Two kinds of declaration exist: na_values / na_range metadata on haven::labelled_spss() vectors, and tagged missing values created by haven::tagged_na() (the Stata .a, .b, ... convention).

spicy honors the declaration by default (user_na = TRUE): declared missing values are excluded from every statistic exactly like NA – valid percentages, means, chi-squared tests, association measures, row-wise summaries, and group definitions – but they are not erased from display. freq() lists each observed declared value as its own row of the Missing block, with its value label; cross_tab(), table_categorical(), and table_continuous() disclose the exclusion in the table note (⁠Declared missing values removed: x (2).⁠); varlist() and code_book() count them as missing in N_valid / NAs / N_distinct while still listing the declared codes in Values.

Every function involved offers the same escape hatch: set user_na = FALSE to ignore the declaration and treat the declared codes as valid values (the behavior of spicy before 0.13.0). Tagged missing values are genuine NAs either way; for them, user_na = FALSE only collapses the per-tag breakdown back into the regular NA count.

See Also

datawizard::row_count() for a closely related row-wise counter; count_n() adds element-wise type-safe matching, multi-value count, and special-value detection.

Other row-wise summaries: mean_n(), sum_n()

Examples

library(dplyr)
library(tibble)
library(labelled)

# Basic usage
df <- tibble(
  x = c(1, 2, 2, 3, NA),
  y = c(2, 2, NA, 3, 2),
  z = c("2", "2", "2", "3", "2")
)
count_n(df, count = 2)
count_n(df, count = 2, allow_coercion = FALSE)
df |> mutate(num_twos = count_n(count = 2))

# Mixed types and special values
df <- tibble(
  num   = c(1, 2, NA, -Inf, NaN),
  char  = c("a", "B", "b", "a", NA),
  fact  = factor(c("a", "b", "b", "a", "c")),
  date  = as.Date(c("2023-01-01", "2023-01-01", NA, "2023-01-02", "2023-01-01")),
  lab   = labelled(c(1, 2, 1, 2, NA), labels = c(No = 1, Yes = 2)),
  logic = c(TRUE, FALSE, NA, TRUE, FALSE)
)
count_n(df, count = 2)
count_n(df, count = "b", ignore_case = TRUE)
count_n(df, count = "a", select = fact)
count_n(df, count = as.Date("2023-01-01"), select = date)

# Count special values
count_n(df, special = "NA")

# Column selection strategies
df <- tibble(
  score_math    = c(1, 2, 2, 3, NA),
  score_science = c(2, 2, NA, 3, 2),
  score_lang    = c("2", "2", "2", "3", "2"),
  name          = c("Jean", "Marie", "Ali", "Zoe", "Nina")
)
count_n(df, select = c(score_math, score_science), count = 2)
count_n(df, select = starts_with("score_"), exclude = "score_lang", count = 2)
count_n(df, select = "^score_", regex = TRUE, count = 2)
df |> mutate(nb_two = count_n(count = 2))

# Strict type-safe matching with factor columns
df <- tibble(
  x = factor(c("a", "b", "c")),
  y = factor(c("b", "B", "a"))
)

# Coercion: character "b" matches both x and y
count_n(df, count = "b")

# Strict match: fails because "b" is character, not factor (returns only 0s)
count_n(df, count = "b", allow_coercion = FALSE)

# Strict match with factor value: works only where levels match
count_n(df, count = factor("b", levels = levels(df$x)), allow_coercion = FALSE)


Cramer's V

Description

cramer_v() computes Cramer's V for a two-way contingency table, measuring the strength of association between two categorical variables.

Usage

cramer_v(x, detail = FALSE, conf_level = 0.95, digits = 3L)

Arguments

x

A contingency table (of class table).

detail

Logical. If FALSE (default), return the estimate as a numeric scalar. If TRUE, return a named numeric vector including confidence interval and p-value.

conf_level

A single number strictly between 0 and 1 giving the confidence level (default 0.95). Only used when detail = TRUE. Set to NULL to omit the confidence interval. Any other value – including percentages such as 95 – raises a classed error (spicy_invalid_input).

digits

Number of decimal places used when printing the result (default 3). Only affects the detail = TRUE output.

Details

Cramer's V is computed as V = \sqrt{\chi^2 / (n \cdot (k - 1))}, where \chi^2 is the Pearson chi-squared statistic, n is the total count, and k = \min(r, c). The point estimate matches the DescTools (Signorell et al., 2024) and SPSS implementations. The confidence interval uses the Fisher z-transformation on V (\tanh(\mathrm{atanh}(V) \pm z_{\alpha/2} / \sqrt{n - 3})), which differs from the noncentral chi-squared or bootstrap CIs reported by DescTools::CramerV().

Value

When detail = FALSE: a single numeric value (the estimate). When detail = TRUE and conf_level is non-NULL: c(estimate, se, ci_lower, ci_upper, p_value). When detail = TRUE and conf_level = NULL: c(estimate, se, p_value). The p-value tests the null hypothesis of no association (Pearson chi-squared test).

References

Agresti, A. (2002). Categorical Data Analysis (2nd ed.). Wiley.

Brown, M. B., & Benedetti, J. K. (1977). Sampling behavior of tests for correlation in two-way contingency tables. Journal of the American Statistical Association, 72(358), 309-315. doi:10.1080/01621459.1977.10480995

Liebetrau, A. M. (1983). Measures of Association. Sage.

Signorell, A. et al. (2024). DescTools: Tools for Descriptive Statistics. R package.

See Also

phi(), contingency_coef(), assoc_measures()

Other association measures: assoc_measures(), contingency_coef(), gamma_gk(), goodman_kruskal_tau(), kendall_tau_b(), kendall_tau_c(), lambda_gk(), phi(), somers_d(), uncertainty_coef(), yule_q()

Examples

tab <- table(sochealth$smoking, sochealth$education)
cramer_v(tab)
cramer_v(tab, detail = TRUE)
cramer_v(tab, detail = TRUE, conf_level = NULL)


Cross-tabulation

Description

Computes a two-way cross-tabulation with optional weights, grouping (including combinations of multiple variables via interaction()), row / column percentages, and inferential statistics (Chi-squared test with an APA-style association measure).

Both x and y are required; for one-way frequency tables, use freq().

Usage

cross_tab(
  data,
  x,
  y = NULL,
  by = NULL,
  weights = NULL,
  rescale = FALSE,
  percent = c("none", "column", "row"),
  include_stats = TRUE,
  assoc_measure = c("auto", "cramer_v", "phi", "gamma", "tau_b", "tau_c", "somers_d",
    "lambda", "none"),
  assoc_ci = FALSE,
  correct = FALSE,
  simulate_p = FALSE,
  simulate_B = 2000,
  digits = NULL,
  output = c("default", "data.frame"),
  show_n = TRUE,
  decimal_mark = ".",
  p_digits = 3L,
  user_na = TRUE,
  styled
)

Arguments

data

A data frame. Alternatively, a vector when using the vector-based interface.

x

Row variable (unquoted).

y

Column variable (unquoted). Required; the NULL default in the signature is a placeholder and triggers an error if left unset (use freq() for one-way tables).

by

Optional grouping variable or expression. Can be a single variable or a combination of multiple variables (e.g. interaction(vs, am)).

weights

Optional numeric weights. A logical vector is also accepted and coerced to 1/0 (include / exclude).

rescale

Logical. If FALSE (the default), weights are used as-is. If TRUE, rescales weights so total weighted N matches raw N.

percent

One of "none" (the default), "column", or "row". Unique abbreviations are accepted (e.g. "n", "c", "r").

include_stats

Logical. If TRUE (the default), computes Chi-squared and an association measure (see assoc_measure).

assoc_measure

Character. Which association measure to report. "auto" (default) selects Kendall's Tau-b when both variables are ordered factors and Cramer's V otherwise. Other choices: "cramer_v", "phi", "gamma", "tau_b", "tau_c", "somers_d", "lambda", "none".

assoc_ci

Logical. If TRUE, includes the 95 percent confidence interval of the association measure in the note. Defaults to FALSE.

correct

Logical. If FALSE (the default), no continuity correction is applied. If TRUE, applies Yates correction (only for 2x2 tables).

simulate_p

Logical. If FALSE (the default), uses asymptotic p-values. If TRUE, uses Monte Carlo simulation.

simulate_B

Integer. Number of replicates for Monte Carlo simulation. Defaults to 2000.

digits

Number of decimals for cell values: a single non-negative integer. Defaults to NULL, which is resolved to 1 when percent != "none" and 0 when percent = "none" (counts are integers unless fractional weights are used; raise digits to display fractional weighted counts exactly). Same role as digits in freq(), which formats percentages only and therefore uses a fixed default of 1. Displayed values round ties half to even (the R / IEC 60559 convention, shared with Stata), so an exact tie like 6.25 prints as 6.2 where SPSS would print 6.3.

output

Output format. "default" (the default) returns a spicy_cross_table object (for formatted printing); "data.frame" returns a plain data.frame. The values match the output argument of the ⁠table_*()⁠ family; the rendered engines that family also accepts ("tinytable", "gt", "flextable", ...) are not available in cross_tab().

show_n

Logical. If TRUE (the default), adds marginal N totals when percent != "none".

decimal_mark

Character used as the decimal mark in printed numeric values (cells, chi-squared, association estimate, CI bounds, p-value, table note). Either "." or ",". The default follows the language: options(spicy.language = "fr") gives the comma here and in freq(), exactly as it does in the reporting ⁠table_*()⁠ family. An argument you type wins, so decimal_mark = "." under a French language gives French words and a decimal point. Under a comma the p-value keeps its leading zero (⁠p = 0,659⁠), the form French typography requires. The resolved mark is frozen on the object when it is built, like every other formatting argument: a table built under a French language still prints its commas when it is printed later with the language option cleared.

p_digits

Integer number of decimals used to format the p-value (and to determine the small-p threshold below which ⁠< .001⁠ notation is used). Defaults to 3 (the APA standard); matches the p_digits argument of the ⁠table_*()⁠ family.

user_na

Logical. If TRUE (the default), declared missing values in x, y, or by are treated as missing: they are excluded from the table and its statistics like NA, and the exclusion is disclosed in the table note (⁠Declared missing values removed: ...⁠). If FALSE, the declared codes tabulate as categories (and, in by, define groups). See the "Declared missing values" section of freq().

styled

Defunct. styled = TRUE is now output = "default" (the default) and styled = FALSE is now output = "data.frame"; supplying styled is an error.

Value

Depends on output and by:

Cell columns are the levels of y; rows are the levels of x. When percent != "none", the N column (or N row) is added according to show_n. When include_stats = TRUE, the result carries a Chi-squared row (statistic, df, p) and an association-measure row (estimate, optional CI via assoc_ci).

Global Options

The function recognizes the following global options that modify its default behavior:

These options are convenient for users who wish to enforce consistent behavior across multiple calls: spicy.percent and spicy.simulate_p apply to cross_tab(), and spicy.rescale applies to both cross_tab() and freq(). They can be disabled or reset by setting them to NULL: options(spicy.percent = NULL, spicy.simulate_p = NULL, spicy.rescale = NULL).

Example:

options(spicy.simulate_p = TRUE, spicy.rescale = TRUE)
cross_tab(sochealth, smoking, education, weights = weight)

Declared missing values

Survey files imported with haven often carry declared missing values: codes such as ⁠8 = Don't know⁠ or 9 = Refused that the source file marks as missing while keeping them distinct from a plain NA. Two kinds of declaration exist: na_values / na_range metadata on haven::labelled_spss() vectors, and tagged missing values created by haven::tagged_na() (the Stata .a, .b, ... convention).

spicy honors the declaration by default (user_na = TRUE): declared missing values are excluded from every statistic exactly like NA – valid percentages, means, chi-squared tests, association measures, row-wise summaries, and group definitions – but they are not erased from display. freq() lists each observed declared value as its own row of the Missing block, with its value label; cross_tab(), table_categorical(), and table_continuous() disclose the exclusion in the table note (⁠Declared missing values removed: x (2).⁠); varlist() and code_book() count them as missing in N_valid / NAs / N_distinct while still listing the declared codes in Values.

Every function involved offers the same escape hatch: set user_na = FALSE to ignore the declaration and treat the declared codes as valid values (the behavior of spicy before 0.13.0). Tagged missing values are genuine NAs either way; for them, user_na = FALSE only collapses the per-tag breakdown back into the regular NA count.

Examples

# Basic crosstab
cross_tab(sochealth, smoking, education)

# Column percentages
cross_tab(sochealth, smoking, education, percent = "column")

# Weighted (rescaled)
cross_tab(sochealth, smoking, education, weights = weight, rescale = TRUE)

# Grouped by sex
cross_tab(sochealth, smoking, education, by = sex)

# Grouped by combination of variables
cross_tab(sochealth, smoking, education, by = interaction(sex, age_group))

# Ordinal variables: auto-selects Kendall's Tau-b
cross_tab(sochealth, education, self_rated_health)

# 2x2 table with Yates correction
cross_tab(sochealth, smoking, physical_activity, correct = TRUE)

# APA-style p-value precision and European decimal mark
cross_tab(sochealth, smoking, education, decimal_mark = ",", p_digits = 4)


Frequency Table

Description

Creates a frequency table for a vector or variable from a data frame, with options for weighting, sorting, handling labelled data, defining custom missing values, and displaying cumulative percentages.

With output = "default", the function returns a spicy_freq_table object that auto-prints as a spicy-formatted ASCII table via print.spicy_freq_table() and spicy_print_table(); with output = "data.frame", it returns a plain data.frame containing frequencies and proportions.

Usage

freq(
  data,
  x = NULL,
  weights = NULL,
  digits = 1L,
  valid = TRUE,
  cum = FALSE,
  sort = "",
  na_val = NULL,
  labelled_levels = c("prefixed", "labels", "values"),
  factor_levels = c("observed", "all"),
  rescale = FALSE,
  decimal_mark = ".",
  output = c("default", "data.frame"),
  user_na = TRUE,
  styled
)

Arguments

data

A data.frame, vector, or factor. If a data frame is provided, specify the target variable x. If both data and x are supplied as vectors, data is ignored with a warning.

x

A variable from data (unquoted).

weights

Optional numeric vector of weights (same length as x). A logical vector is also accepted and coerced to 1/0 (include / exclude). The variable may be referenced as a bare name when it belongs to data, or as a qualified expression like other$w (evaluated in the calling environment), which always takes precedence over data lookup. Observations with NA weights are dropped from the table with a warning; see Details.

digits

Number of decimal digits to display for percentages (default: 1). Same role as digits in cross_tab(), where the NULL default resolves to the same 1 decimal whenever percentages are shown. Displayed values round ties half to even (the R / IEC 60559 convention, shared with Stata), so an exact tie like 6.25 prints as 6.2 where SPSS would print 6.3.

valid

Logical. If TRUE (default), display valid percentages (excluding missing values).

cum

Logical. If FALSE (the default), cumulative percentages are omitted. If TRUE, adds cumulative percentages.

sort

Sorting method for values:

  • "" - no sorting (default)

  • "+" - increasing frequency

  • "-" - decreasing frequency

  • "name+" - alphabetical A-Z

  • "name-" - alphabetical Z-A

For labelled variables displayed with their codes (labelled_levels "prefixed" or "values"), "name+" / "name-" sort by the underlying code (so ⁠[10]⁠ follows ⁠[2]⁠, as in SPSS), not by the display string. With labelled_levels = "labels", labels sort alphabetically.

na_val

Atomic vector of numeric or character values to be treated as missing (NA).

For labelled variables (from haven or labelled), this argument must refer to the underlying coded values, not the visible labels.

Example:

x <- labelled(c(1, 2, 3, 1, 2, 3), c("Low" = 1, "Medium" = 2, "High" = 3))
freq(x, na_val = 1) # Treat all "Low" as missing
labelled_levels

For labelled variables, defines how labels and values are displayed:

  • "prefixed" or "p" - show labels as ⁠[value] label⁠ (default)

  • "labels" or "l" - show only labels

  • "values" or "v" - show only numeric codes

factor_levels

Character. Controls how factor and labelled values are displayed in the frequency table. "observed" (the default; matches Stata's tab and SPSS FREQUENCIES, which both list only values present in the data) shows only levels present in the data. "all" (code_book()'s default) keeps every declared level, including unused ones, which appear with n = 0.

rescale

Logical. If FALSE (the default), weights are used as-is. If TRUE, rescale weights so that their total equals the unweighted sample size (length(weights)). When the argument is not supplied, the default can be set globally with options(spicy.rescale = TRUE), which is read by both freq() and cross_tab(). See Details for the interaction with NA weights.

decimal_mark

Character used as the decimal mark in printed percentages. Either "." or ",". The default follows the language: options(spicy.language = "fr") gives the comma here and in cross_tab(), exactly as it does in the reporting ⁠table_*()⁠ families. An argument you type wins, so decimal_mark = "." under a French language gives French words and a decimal point. The resolved mark is frozen on the object when it is built: a table built under a French language still prints its commas when it is printed later with the language option cleared. Its words are not – freq() resolves its labels at print time – so that table prints English headings over French numbers.

output

Output format. "default" (the default) returns a spicy_freq_table object that auto-prints as a formatted spicy table; "data.frame" returns a plain data.frame with frequency values. The values match the output argument of the ⁠table_*()⁠ family; the rendered engines that family also accepts ("tinytable", "gt", "flextable", ...) are not available in freq().

user_na

Logical. If TRUE (the default), declared missing values are treated as missing: each observed declared value becomes its own row of the Missing block (with its value label), and valid percentages exclude those observations. If FALSE, the declaration is ignored and the declared codes tabulate as valid categories. See the "Declared missing values" section.

styled

Defunct. styled = TRUE is now output = "default" (the default) and styled = FALSE is now output = "data.frame"; supplying styled is an error.

Details

Designed to mimic common frequency procedures from SPSS or Stata while integrating the flexibility of R's data structures. The input type (vector, factor, labelled) is auto-detected; see ⁠@param labelled_levels⁠ and ⁠@param factor_levels⁠ for the schema-vs-observed level controls, and ⁠@param na_val⁠ for optional sentinel-value recoding.

Weighting (weights): frequencies and percentages are computed proportionally to the weights. Missing values in weights cause those observations to be dropped from the table entirely (with a warning), matching the behaviour of cross_tab() in spicy 0.11.0+. With rescale = TRUE, the remaining (non-NA-weighted) weights are normalised so the total weighted N equals the count of non-NA-weighted rows. With rescale = FALSE, the total weighted N is the actual sum of non-NA weights.

For schema-level inspection without computing frequencies, use varlist() or code_book().

Value

With output = "data.frame", a plain data.frame with no extra attributes and columns:

With output = "default" (the default), a spicy_freq_table object: the same data.frame carrying rendering metadata as attributes (digits, data_name, var_name, var_label, class_name, n_total, n_valid, weighted, rescaled, weight_var) used by print.spicy_freq_table(). The object is returned visibly, so a bare freq(...) call auto-prints at the console while f <- freq(...) stays silent (print f to display the table).

Declared missing values

Survey files imported with haven often carry declared missing values: codes such as ⁠8 = Don't know⁠ or 9 = Refused that the source file marks as missing while keeping them distinct from a plain NA. Two kinds of declaration exist: na_values / na_range metadata on haven::labelled_spss() vectors, and tagged missing values created by haven::tagged_na() (the Stata .a, .b, ... convention).

spicy honors the declaration by default (user_na = TRUE): declared missing values are excluded from every statistic exactly like NA – valid percentages, means, chi-squared tests, association measures, row-wise summaries, and group definitions – but they are not erased from display. freq() lists each observed declared value as its own row of the Missing block, with its value label; cross_tab(), table_categorical(), and table_continuous() disclose the exclusion in the table note (⁠Declared missing values removed: x (2).⁠); varlist() and code_book() count them as missing in N_valid / NAs / N_distinct while still listing the declared codes in Values.

Every function involved offers the same escape hatch: set user_na = FALSE to ignore the declaration and treat the declared codes as valid values (the behavior of spicy before 0.13.0). Tagged missing values are genuine NAs either way; for them, user_na = FALSE only collapses the per-tag breakdown back into the regular NA count.

See Also

cross_tab() for two-way cross-tabulations; table_categorical() for multi-variable categorical summary tables; varlist() / code_book() for variable inspection; print.spicy_freq_table() for formatted printing; spicy_print_table() for the underlying ASCII rendering engine.

Examples

# Frequency table with labelled ordered factor
freq(sochealth, education)
freq(sochealth, self_rated_health, sort = "-")

library(labelled)

# Simple numeric vector
x <- c(1, 2, 2, 3, 3, 3, NA)
freq(x)

# Plain vector with a sentinel value recoded as missing
freq(c(1, 2, 3, 99, 99), na_val = 99)

# Labelled variable (haven-style)
x_lbl <- labelled(
  c(1, 2, 3, 1, 2, 3, 1, 2, NA),
  labels = c("Low" = 1, "Medium" = 2, "High" = 3)
)
var_label(x_lbl) <- "Satisfaction level"

# Treat value 1 ("Low") as missing
freq(x_lbl, na_val = 1)

# Display only labels, add cumulative %
freq(x_lbl, labelled_levels = "labels", cum = TRUE)

# Display values only, sorted descending
freq(x_lbl, labelled_levels = "values", sort = "-")

# Show all declared factor levels, including unused ones (n = 0).
# The default "observed" mirrors Stata's `tab` and SPSS FREQUENCIES,
# which both drop unused levels.
f <- factor(c("Yes", "No", "Yes"), levels = c("Yes", "No", "Maybe"))
freq(f, factor_levels = "all")

# With weighting
df <- data.frame(
  sex = factor(c("Male", "Female", "Female", "Male", NA, "Female")),
  weight = c(12, 8, 10, 15, 7, 9)
)

# Weighted frequencies (raw weighted counts, the default)
freq(df, sex, weights = weight)

# Weighted frequencies rescaled so the total matches the sample size
freq(df, sex, weights = weight, rescale = TRUE)

# Base R style, with weights and cumulative percentages
freq(df$sex, weights = df$weight, cum = TRUE)

# Piped version (tidy syntax) and sort alphabetically descending ("name-")
df |> freq(sex, sort = "name-")

# European decimal mark (matches `cross_tab()` and the `table_*()` family)
freq(sochealth, education, decimal_mark = ",")

# Plain data.frame return (for programmatic use)
f <- freq(df, sex, output = "data.frame")
head(f)


Goodman-Kruskal Gamma

Description

gamma_gk() computes the Goodman-Kruskal Gamma statistic for a two-way contingency table of ordinal variables.

Usage

gamma_gk(x, detail = FALSE, conf_level = 0.95, digits = 3L)

Arguments

x

A contingency table (of class table).

detail

Logical. If FALSE (default), return the estimate as a numeric scalar. If TRUE, return a named numeric vector including confidence interval and p-value.

conf_level

A single number strictly between 0 and 1 giving the confidence level (default 0.95). Only used when detail = TRUE. Set to NULL to omit the confidence interval. Any other value – including percentages such as 95 – raises a classed error (spicy_invalid_input).

digits

Number of decimal places used when printing the result (default 3). Only affects the detail = TRUE output.

Details

Gamma is computed as \gamma = (C - D) / (C + D), where C and D are the numbers of concordant and discordant pairs. It ignores tied pairs, making it appropriate for ordinal variables with many ties. When the asymptotic standard error is zero (e.g. a perfect association), the Wald z-test is undefined and the p-value is NA, matching the other measures in the family. Standard error formulas follow the DescTools implementations (Signorell et al., 2024); see cramer_v() for full references.

Value

Same structure as cramer_v(): a scalar when detail = FALSE, a named vector when detail = TRUE. The p-value tests H0: gamma = 0 (Wald z-test).

References

Goodman, L. A., & Kruskal, W. H. (1954). Measures of association for cross classifications. Journal of the American Statistical Association, 49(268), 732-764. doi:10.2307/2281536

Brown, M. B., & Benedetti, J. K. (1977). Sampling behavior of tests for correlation in two-way contingency tables. Journal of the American Statistical Association, 72(358), 309-315. doi:10.1080/01621459.1977.10480995

See Also

kendall_tau_b(), kendall_tau_c(), somers_d(), assoc_measures()

Other association measures: assoc_measures(), contingency_coef(), cramer_v(), goodman_kruskal_tau(), kendall_tau_b(), kendall_tau_c(), lambda_gk(), phi(), somers_d(), uncertainty_coef(), yule_q()

Examples

tab <- table(sochealth$education, sochealth$self_rated_health)
gamma_gk(tab)
gamma_gk(tab, detail = TRUE)


Goodman-Kruskal's Tau

Description

goodman_kruskal_tau() computes Goodman-Kruskal's Tau, a proportional reduction in error (PRE) measure for nominal variables.

Usage

goodman_kruskal_tau(
  x,
  direction = c("row", "column"),
  detail = FALSE,
  conf_level = 0.95,
  digits = 3L
)

Arguments

x

A contingency table (of class table).

direction

Direction of prediction: "row" (default, column predicts row) or "column" (row predicts column).

detail

Logical. If FALSE (default), return the estimate as a numeric scalar. If TRUE, return a named numeric vector including confidence interval and p-value.

conf_level

A single number strictly between 0 and 1 giving the confidence level (default 0.95). Only used when detail = TRUE. Set to NULL to omit the confidence interval. Any other value – including percentages such as 95 – raises a classed error (spicy_invalid_input).

digits

Number of decimal places used when printing the result (default 3). Only affects the detail = TRUE output.

Details

Unlike lambda_gk(), Goodman-Kruskal's Tau uses all cell frequencies rather than only the modal categories, making it more sensitive to association patterns where lambda may be zero. Goodman-Kruskal's Tau is intrinsically directional and has no canonical symmetric form (unlike lambda_gk() or uncertainty_coef()); only "row" and "column" are supported.

Standard error formulas follow the DescTools implementations (Signorell et al., 2024); see cramer_v() for full references.

Value

Same structure as cramer_v(): a scalar when detail = FALSE, a named vector when detail = TRUE. The p-value tests H0: tau = 0 (Wald z-test).

References

Goodman, L. A., & Kruskal, W. H. (1954). Measures of association for cross classifications. Journal of the American Statistical Association, 49(268), 732-764. doi:10.2307/2281536

See Also

lambda_gk(), uncertainty_coef(), assoc_measures()

Other association measures: assoc_measures(), contingency_coef(), cramer_v(), gamma_gk(), kendall_tau_b(), kendall_tau_c(), lambda_gk(), phi(), somers_d(), uncertainty_coef(), yule_q()

Examples

tab <- table(sochealth$smoking, sochealth$education)
goodman_kruskal_tau(tab)
goodman_kruskal_tau(tab, direction = "column", detail = TRUE)


Cite a table cell in inline text

Description

Returns one cell of a spicy table as a character scalar, formatted exactly as the table displays it – same decimals, same p style, same interval punctuation, same journal style. Designed for inline R chunks in Quarto / R Markdown:

Smokers had higher odds (`r inline(tbl, smoking, "Yes", "or")`).

Usage

inline(x, variable, level = NULL, column = NULL, model = NULL)

Arguments

x

A table returned by table_regression(), table_categorical(), table_continuous(), table_continuous_lm(), table_outcome(), table_categorical_svy() or table_continuous_svy() (default output) – every family as_structured() accepts.

variable

The source variable, unquoted or as a string; or a fit-statistic token ("n", "r2", ...).

level

For a factor variable, the level, as a string. "(Missing)" addresses the missing-value category by role. On table_continuous_lm(), whose groups are columns rather than rows, it names the group.

column

A column token, or a {token} pattern. NULL (the default) returns the estimate-like column of the row when it is unambiguous: the family's primary estimate. That is the coefficient for table_regression() – always token "b": an exponentiated table changes its header to OR, IRR or HR, never its token – the contrast ("delta") or, across a numeric by, the slope ("b") for table_continuous_lm(), the mean ("m") or, on a median-only table, the median ("med", "med_iqr") for table_continuous(), and the count ("n") for table_categorical(). A row carrying none of them refuses and lists its tokens.

model

In a multi-model table, the model: its label (as displayed in the column spanners) or its position.

Value

A character scalar.

Addressing

The row is found by identity, not by display text: variable names the source column (.variable in the typed body), level the level (.level). Custom labels, a style, or a translated display never change the call. As a convenience, a variable that matches no source column is looked up among the displayed labels before erroring. The missing-value category is addressed by level = "(Missing)" whatever its displayed (possibly deduplicated) label, through its row role. Fit statistics are addressed by their token as variable (inline(tbl, "n"), inline(tbl, "r2")).

A statistic that belongs to a whole variable rather than to one of its levels sits on the variable's own row: the p of a table_categorical() block, its association measure, its SMD. Leave level out to cite it (inline(tbl, smoking, column = "p")).

table_continuous_lm() lays its groups out sideways – one row per outcome, the by levels as columns (M (Female), M (Male)) – so there level names the group whose column you want: inline(tbl, bmi, "Female", "emmean"). The columns that belong to no group (the contrast, its interval, p, n) are cited without one, as before.

The column is a token of the typed contract ("b", "se", "p", "ci", "or", "ame", "n", "pct", "m", ... – see as_structured()'s col_meta), never a display header. "ci" composes the interval with the style's brackets and separator, and so does every other interval token the table carries ("med_ci", "ame_ci", "assoc_ci"): each addresses its own bounds. In a multi-model table, model selects the model by its spanner label or position; in a by table, the spanners are the groups, so model selects the group the same way.

Patterns

A column containing ⁠{⁠ is a pattern: each {token} is replaced by the corresponding cell, so one call quotes a full sentence fragment:

inline(tbl, smoking, "Yes", "{or} ({ci_label} {ci}; p = {p})")

{ci_label} inserts the interval label of the interval the pattern cites (⁠95% CI⁠, or ⁠Med 95% CI⁠ in a pattern quoting {med_ci}) – the first one when it cites several, the table's first when it cites none. Note that {p} carries the floor operator when the table does (⁠<.001⁠), so write ⁠p {p}⁠ rather than p = {p} in patterns that may hit the floor.

Errors

Every misaddressing is a classed error that lists the available choices: unknown variables list the variables, missing levels list the levels, unknown tokens list the table's tokens, ambiguous models list the spanner labels. A cell the table itself displays as undefined (an aliased coefficient's en-dash) refuses with the reason rather than pasting a dash into a sentence.

See Also

as_structured() for the typed contract behind the addressing.

Examples

fit <- lm(wellbeing_score ~ age + sex, data = sochealth)
tbl <- table_regression(fit)
inline(tbl, age, column = "b")
inline(tbl, sex, "Male", "{b} ({ci_label} {ci}; p {p})")
inline(tbl, "n")

Kendall's Tau-b

Description

kendall_tau_b() computes Kendall's Tau-b for a two-way contingency table of ordinal variables.

Usage

kendall_tau_b(x, detail = FALSE, conf_level = 0.95, digits = 3L)

Arguments

x

A contingency table (of class table).

detail

Logical. If FALSE (default), return the estimate as a numeric scalar. If TRUE, return a named numeric vector including confidence interval and p-value.

conf_level

A single number strictly between 0 and 1 giving the confidence level (default 0.95). Only used when detail = TRUE. Set to NULL to omit the confidence interval. Any other value – including percentages such as 95 – raises a classed error (spicy_invalid_input).

digits

Number of decimal places used when printing the result (default 3). Only affects the detail = TRUE output.

Details

Kendall's Tau-b is computed as \tau_b = (C - D) / \sqrt{(n_0 - n_1)(n_0 - n_2)}, where n_0 = n(n-1)/2, n_1 is the number of pairs tied on the row variable, and n_2 is the number tied on the column variable. Tau-b corrects for ties and is appropriate for square tables. When the asymptotic standard error is zero (e.g. a perfect association), the Wald z-test is undefined and the p-value is NA, matching the other measures in the family.

The asymptotic standard error is the Brown and Benedetti (1977) ASE1, as printed by SPSS / PSPP CROSSTABS. It deliberately diverges from DescTools::KendallTauB(), whose implementation mis-scales one margin term of the gradient; see cramer_v() for full references.

Value

Same structure as cramer_v(): a scalar when detail = FALSE, a named vector when detail = TRUE. The p-value tests H0: tau-b = 0 (Wald z-test).

References

Kendall, M. G. (1938). A new measure of rank correlation. Biometrika, 30(1-2), 81-93. doi:10.2307/2332226

Brown, M. B., & Benedetti, J. K. (1977). Sampling behavior of tests for correlation in two-way contingency tables. Journal of the American Statistical Association, 72(358), 309-315. doi:10.1080/01621459.1977.10480995

See Also

kendall_tau_c(), gamma_gk(), somers_d(), assoc_measures()

Other association measures: assoc_measures(), contingency_coef(), cramer_v(), gamma_gk(), goodman_kruskal_tau(), kendall_tau_c(), lambda_gk(), phi(), somers_d(), uncertainty_coef(), yule_q()

Examples

tab <- table(sochealth$education, sochealth$self_rated_health)
kendall_tau_b(tab)


Kendall's Tau-c (Stuart's Tau-c)

Description

kendall_tau_c() computes Stuart's Tau-c (also known as Kendall's Tau-c) for a two-way contingency table of ordinal variables.

Usage

kendall_tau_c(x, detail = FALSE, conf_level = 0.95, digits = 3L)

Arguments

x

A contingency table (of class table).

detail

Logical. If FALSE (default), return the estimate as a numeric scalar. If TRUE, return a named numeric vector including confidence interval and p-value.

conf_level

A single number strictly between 0 and 1 giving the confidence level (default 0.95). Only used when detail = TRUE. Set to NULL to omit the confidence interval. Any other value – including percentages such as 95 – raises a classed error (spicy_invalid_input).

digits

Number of decimal places used when printing the result (default 3). Only affects the detail = TRUE output.

Details

Stuart's Tau-c is computed as \tau_c = 2m(C - D) / (n^2(m - 1)), where m = \min(r, c). It is designed for rectangular tables; the estimate is bounded by [-1, 1] only when the table is square, and may fall outside that range otherwise. When the asymptotic standard error is zero (e.g. a perfect association), the Wald z-test is undefined and the p-value is NA, matching the other measures in the family. When one variable is constant (all observations in a single row or column), there are no untied pairs and the statistic degenerates to a meaningless 0: the function returns NA with a spicy_undefined_stat warning, like its siblings, matching the SPSS / PSPP behavior of reporting no value. Standard error formulas follow the DescTools implementations (Signorell et al., 2024); see cramer_v() for full references.

Value

Same structure as cramer_v(): a scalar when detail = FALSE, a named vector when detail = TRUE. The p-value tests H0: tau-c = 0 (Wald z-test).

References

Stuart, A. (1953). The estimation and comparison of strengths of association in contingency tables. Biometrika, 40(1-2), 105-110. doi:10.2307/2333101

Brown, M. B., & Benedetti, J. K. (1977). Sampling behavior of tests for correlation in two-way contingency tables. Journal of the American Statistical Association, 72(358), 309-315. doi:10.1080/01621459.1977.10480995

See Also

kendall_tau_b(), gamma_gk(), somers_d(), assoc_measures()

Other association measures: assoc_measures(), contingency_coef(), cramer_v(), gamma_gk(), goodman_kruskal_tau(), kendall_tau_b(), lambda_gk(), phi(), somers_d(), uncertainty_coef(), yule_q()

Examples

tab <- table(sochealth$education, sochealth$self_rated_health)
kendall_tau_c(tab)


Derive variable labels from column names ⁠name<sep>label⁠

Description

Splits each column name at the first occurrence of sep, renames the column to the part before sep (the name, trimmed of surrounding whitespace), and assigns the part after sep as a "label" attribute on the column. The label attribute follows the haven convention also used by labelled::var_label(), so labelled-aware tooling (labelled, haven, varlist(), code_book(), ...) reads it transparently. Splitting at the first sep means the label itself may contain the separator.

Usage

label_from_names(df, sep = ". ")

Arguments

df

A data.frame or tibble with column names of the form ⁠"name<sep>label"⁠ (e.g. "code. question text").

sep

Character string used as separator between name and label. Default ". " (LimeSurvey's default); any literal string can be used. Matched as a fixed string, so regex metacharacters such as . or | carry no special meaning.

Details

Designed primarily for LimeSurvey CSV exports with Headings: Question code & question text, which produce column names like "code. question text". The default separator ". " matches that export.

LimeSurvey question codes (the part before sep) are restricted to alphanumerics, must start with a letter, and contain no spaces – so the column name has to carry both the code and the question text. If your export uses Headings: Question code (codes only), re-export with Question code & question text before calling this function; there is no way to recover a label from a code alone.

Whitespace handling: the name (left of sep) is trimmed of surrounding whitespace, because R column names are intended to be referenced bare (without backticks) and leading / trailing whitespace would force quoting throughout the user's downstream code. The label (right of sep) is preserved verbatim, following the Stata / SPSS convention that variable labels are faithful user content – spicy does not silently mutate label strings. To trim labels yourself, post-process with labelled::var_label(df) <- lapply(labelled::var_label(df), trimws).

Value

An object of the same class as df – a base data.frame if df was a base data.frame, a tbl_df if df was a tibble. The output has column names equal to the trimmed names (before sep) and, for every column whose original name contained sep, a "label" attribute equal to the label (after sep). Columns whose name does not contain sep are passed through unchanged with no label attached.

Errors

The function raises an actionable error – rather than letting the downstream constructor raise a cryptic one – when the split produces:

See Also

labelled::var_label() reads the "label" attribute set by this function; varlist() and code_book() surface it in their inspection outputs.

Other variable inspection: code_book(), varlist()

Examples

# LimeSurvey-style column names (default sep = ". ").
df <- data.frame(
  "age. Age of respondent" = c(25, 30),
  "score. Total score. Manually computed." = c(12, 14),
  check.names = FALSE
)
out <- label_from_names(df)
attr(out$age, "label")
attr(out$score, "label")

# Custom separator.
df2 <- data.frame(
  "id|Identifier" = 1:3,
  "score|Total score" = c(10, 20, 30),
  check.names = FALSE
)
out2 <- label_from_names(df2, sep = "|")


Goodman-Kruskal's Lambda

Description

lambda_gk() computes Goodman-Kruskal's Lambda, a proportional reduction in error (PRE) measure for nominal variables.

Usage

lambda_gk(
  x,
  direction = c("symmetric", "row", "column"),
  detail = FALSE,
  conf_level = 0.95,
  digits = 3L
)

Arguments

x

A contingency table (of class table).

direction

Direction of prediction: "symmetric" (default), "row" (column predicts row), or "column" (row predicts column).

detail

Logical. If FALSE (default), return the estimate as a numeric scalar. If TRUE, return a named numeric vector including confidence interval and p-value.

conf_level

A single number strictly between 0 and 1 giving the confidence level (default 0.95). Only used when detail = TRUE. Set to NULL to omit the confidence interval. Any other value – including percentages such as 95 – raises a classed error (spicy_invalid_input).

digits

Number of decimal places used when printing the result (default 3). Only affects the detail = TRUE output.

Details

Lambda measures how much prediction error is reduced when the independent variable is used to predict the dependent variable. It ranges from 0 (no reduction) to 1 (perfect prediction). Lambda can equal zero even when variables are associated if the modal category dominates in every column (or row).

The default direction = "symmetric" follows the SPSS and DescTools convention: symmetric lambda is a standard, well-defined variant with its own asymptotic standard error. somers_d() deliberately differs (its default is "row") because its symmetric form is a derived quantity without an analytic SE; see its documentation.

Standard error formulas follow the DescTools implementations (Signorell et al., 2024); see cramer_v() for full references.

Value

Same structure as cramer_v(): a scalar when detail = FALSE, a named vector when detail = TRUE. The p-value tests H0: lambda = 0 (Wald z-test).

References

Goodman, L. A., & Kruskal, W. H. (1954). Measures of association for cross classifications. Journal of the American Statistical Association, 49(268), 732-764. doi:10.2307/2281536

See Also

goodman_kruskal_tau(), uncertainty_coef(), assoc_measures()

Other association measures: assoc_measures(), contingency_coef(), cramer_v(), gamma_gk(), goodman_kruskal_tau(), kendall_tau_b(), kendall_tau_c(), phi(), somers_d(), uncertainty_coef(), yule_q()

Examples

tab <- table(sochealth$smoking, sochealth$education)
lambda_gk(tab)
lambda_gk(tab, direction = "row")
lambda_gk(tab, direction = "column", detail = TRUE)


Row means with an optional minimum-valid-values rule

Description

Computes row-wise means across selected numeric columns of a data.frame or matrix. Missing values are handled per row via min_valid (an integer count or proportion of non-NA values required); rows that fail the rule return NA, and rows with no valid values at all return NA even when min_valid = 0. Non-numeric columns are dropped silently (set verbose = TRUE to see which). Designed to flow inside dplyr::mutate(): when called without an explicit data argument, the current data context is used.

Usage

mean_n(
  data = NULL,
  select = tidyselect::everything(),
  exclude = NULL,
  min_valid = NULL,
  digits = NULL,
  regex = FALSE,
  verbose = FALSE,
  user_na = TRUE
)

Arguments

data

A data.frame or matrix. Optional inside dplyr::mutate(), where the current grouping/data context is used automatically.

select

Columns to include. If regex = FALSE, use tidyselect syntax (default: tidyselect::everything()). If regex = TRUE, provide a regular expression pattern (character string).

exclude

Columns to exclude (default: NULL).

min_valid

Minimum number of valid (non-NA) values required per row. Accepts:

  • NULL (the default) – every selected column must be valid.

  • a proportion in ⁠(0, 1)⁠ – round(ncol(x) * min_valid) valid columns required (e.g. min_valid = 0.5 requires at least half of the selected columns to be non-NA).

  • a non-negative integer count up to the number of selected numeric columns.

Non-integer values ⁠>= 1⁠ (e.g. 1.5) and counts greater than ncol(x) raise an actionable error.

Rows with zero valid values always return NA, even when min_valid = 0: an empty row-wise summary is undefined, so the raw rowMeans() / rowSums() identities (NaN / 0) are never returned.

digits

Optional non-negative integer giving the number of decimal places to round the result to. Defaults to NULL (no rounding).

regex

Logical. If FALSE (the default), uses tidyselect helpers. If TRUE, the select argument is treated as a regular expression.

verbose

Logical. If FALSE (the default), messages are suppressed. If TRUE, prints a message about non-numeric columns excluded.

user_na

Logical. If TRUE (the default), declared missing values count as missing – both in the computed summary and in the min_valid valid-count gate. If FALSE, the declared codes are treated as ordinary numbers. See the "Declared missing values" section of freq().

Value

A numeric vector of row-wise means.

Declared missing values

Survey files imported with haven often carry declared missing values: codes such as ⁠8 = Don't know⁠ or 9 = Refused that the source file marks as missing while keeping them distinct from a plain NA. Two kinds of declaration exist: na_values / na_range metadata on haven::labelled_spss() vectors, and tagged missing values created by haven::tagged_na() (the Stata .a, .b, ... convention).

spicy honors the declaration by default (user_na = TRUE): declared missing values are excluded from every statistic exactly like NA – valid percentages, means, chi-squared tests, association measures, row-wise summaries, and group definitions – but they are not erased from display. freq() lists each observed declared value as its own row of the Missing block, with its value label; cross_tab(), table_categorical(), and table_continuous() disclose the exclusion in the table note (⁠Declared missing values removed: x (2).⁠); varlist() and code_book() count them as missing in N_valid / NAs / N_distinct while still listing the declared codes in Values.

Every function involved offers the same escape hatch: set user_na = FALSE to ignore the declaration and treat the declared codes as valid values (the behavior of spicy before 0.13.0). Tagged missing values are genuine NAs either way; for them, user_na = FALSE only collapses the per-tag breakdown back into the regular NA count.

See Also

Other row-wise summaries: count_n(), sum_n()

Examples

library(dplyr)

# Create a simple numeric data frame
df <- tibble(
  var1 = c(10, NA, 30, 40, 50),
  var2 = c(5, NA, 15, NA, 25),
  var3 = c(NA, 30, 20, 50, 10)
)

# Compute row-wise mean (all values must be valid by default)
mean_n(df)

# Require at least 2 valid (non-NA) values per row
mean_n(df, min_valid = 2)

# Require at least 50% valid (non-NA) values per row
mean_n(df, min_valid = 0.5)

# Round the result to 1 decimal
mean_n(df, digits = 1)

# Select specific columns
mean_n(df, select = c(var1, var2))

# Select specific columns using a pipe
df |>
  select(var1, var2) |>
  mean_n()

# Exclude a column
mean_n(df, exclude = "var3")

# Select columns ending with "1"
mean_n(df, select = ends_with("1"))

# Use with native pipe
df |> mean_n(select = starts_with("var"))

# Use inside dplyr::mutate()
df |> mutate(mean_score = mean_n(min_valid = 2))

# Select columns directly inside mutate()
df |> mutate(mean_score = mean_n(select = c(var1, var2), min_valid = 1))

# Select columns before mutate
df |>
  select(var1, var2) |>
  mutate(mean_score = mean_n(min_valid = 1))

# Show verbose processing info
df |> mutate(mean_score = mean_n(min_valid = 2, digits = 1, verbose = TRUE))

# Add character and grouping columns
df_mixed <- mutate(df,
  name = letters[1:5],
  group = c("A", "A", "B", "B", "A")
)
df_mixed

# Non-numeric columns are ignored
mean_n(df_mixed)

# Use within mutate() on mixed data
df_mixed |> mutate(mean_score = mean_n(select = starts_with("var")))

# Use everything() but exclude non-numeric columns manually
mean_n(df_mixed, select = everything(), exclude = "group")

# Select columns using regex
mean_n(df_mixed, select = "^var", regex = TRUE)
mean_n(df_mixed, select = "ar", regex = TRUE)

# Apply to a subset of rows (first 3)
df_mixed[1:3, ] |> mean_n(select = starts_with("var"))

# Store the result in a new column
df_mixed$mean_score <- mean_n(df_mixed, select = starts_with("var"))
df_mixed

# With a numeric matrix
mat <- matrix(c(1, 2, NA, 4, 5, NA, 7, 8, 9), nrow = 3, byrow = TRUE)
mat
mat |> mean_n(min_valid = 2)


Phi coefficient

Description

phi() computes the phi coefficient for a 2x2 contingency table.

Usage

phi(x, detail = FALSE, conf_level = 0.95, digits = 3L)

Arguments

x

A contingency table (of class table).

detail

Logical. If FALSE (default), return the estimate as a numeric scalar. If TRUE, return a named numeric vector including confidence interval and p-value.

conf_level

A single number strictly between 0 and 1 giving the confidence level (default 0.95). Only used when detail = TRUE. Set to NULL to omit the confidence interval. Any other value – including percentages such as 95 – raises a classed error (spicy_invalid_input).

digits

Number of decimal places used when printing the result (default 3). Only affects the detail = TRUE output.

Details

The phi coefficient is \phi = \sqrt{\chi^2 / n}. It is equivalent to Cramer's V for 2x2 tables and equals the absolute value of the Pearson correlation between the two binary variables – spicy returns only the magnitude (always non-negative), matching the DescTools (Signorell et al., 2024) and PSPP conventions. SPSS itself SIGNS phi on 2x2 tables (its CROSSTABS algorithm sets the sign to that of the Pearson correlation), so SPSS output can show a negative value of the same magnitude. To recover the signed direction of the 2x2 association, compute the Pearson correlation directly (e.g. cor(x, y) after coding both variables 0/1).

The confidence interval uses the Fisher z-transformation on \phi; see cramer_v() for the formula and full references.

Value

Same structure as cramer_v(): a scalar when detail = FALSE, a named vector when detail = TRUE. The p-value tests the null hypothesis of no association (Pearson chi-squared test).

References

Liebetrau, A. M. (1983). Measures of Association. Sage.

See Also

cramer_v(), yule_q(), assoc_measures()

Other association measures: assoc_measures(), contingency_coef(), cramer_v(), gamma_gk(), goodman_kruskal_tau(), kendall_tau_b(), kendall_tau_c(), lambda_gk(), somers_d(), uncertainty_coef(), yule_q()

Examples

tab <- table(sochealth$smoking, sochealth$sex)
phi(tab)
phi(tab, detail = TRUE)


Print a detailed association measure result

Description

Formats a spicy_assoc_detail vector (returned by association functions with detail = TRUE) with fixed decimal places and ⁠< 0.001⁠ notation for small p-values.

Usage

## S3 method for class 'spicy_assoc_detail'
print(x, digits = attr(x, "digits") %||% 3L, ...)

Arguments

x

A spicy_assoc_detail object.

digits

Number of decimal places for the estimate, SE, and confidence interval. Defaults to 3. The p-value is always formatted separately using APA notation (⁠<.001⁠ or three decimal places, no leading zero), via the shared format_p_value() helper used by cross_tab() and the ⁠table_*()⁠ family.

...

Ignored.

Value

x, invisibly.

See Also

cramer_v(), assoc_measures()


Print an association measures summary table

Description

Formats a spicy_assoc_table data frame (returned by assoc_measures()) with fixed decimal places, aligned columns, and APA-style ⁠<.001⁠ notation for small p-values (same helper as cross_tab() and the ⁠table_*()⁠ family).

Usage

## S3 method for class 'spicy_assoc_table'
print(x, digits = attr(x, "digits") %||% 3L, ...)

Arguments

x

A spicy_assoc_table object.

digits

Number of decimal places for estimates, SE, and confidence intervals. Defaults to 3. The p-value is always formatted separately using APA notation (⁠<.001⁠ or three decimal places, no leading zero), via the shared format_p_value() helper used by cross_tab() and the ⁠table_*()⁠ family.

...

Ignored.

Value

x, invisibly.

See Also

assoc_measures()


Print method for categorical survey-design tables

Description

Formats and prints a spicy_categorical_svy_table object as a styled ASCII table using spicy_print_table().

Usage

## S3 method for class 'spicy_categorical_svy_table'
print(x, ...)

Arguments

x

A data.frame of class "spicy_categorical_svy_table" as returned by table_categorical_svy().

...

Additional arguments (currently ignored).

Value

Invisibly returns x.

See Also

table_categorical_svy(), spicy_print_table()


Print method for categorical summary tables

Description

Formats and prints a spicy_categorical_table object as a styled ASCII table using spicy_print_table().

Usage

## S3 method for class 'spicy_categorical_table'
print(x, ...)

Arguments

x

A data.frame of class "spicy_categorical_table" as returned by table_categorical() with output = "default".

...

Additional arguments (currently ignored).

Value

Invisibly returns x.

See Also

table_categorical(), spicy_print_table()


Print method for bivariate linear-model tables

Description

Formats and prints a spicy_continuous_lm_table object as a styled ASCII table using spicy_print_table().

Usage

## S3 method for class 'spicy_continuous_lm_table'
print(x, ...)

Arguments

x

A data.frame of class "spicy_continuous_lm_table" as returned by table_continuous_lm().

...

Additional arguments (currently ignored).

Value

Invisibly returns x.

See Also

table_continuous_lm(), spicy_print_table()


Print method for continuous survey-design tables

Description

Formats and prints a spicy_continuous_svy_table object as a styled ASCII table using spicy_print_table().

Usage

## S3 method for class 'spicy_continuous_svy_table'
print(x, ...)

Arguments

x

A data.frame of class "spicy_continuous_svy_table" as returned by table_continuous_svy().

...

Additional arguments (currently ignored).

Value

Invisibly returns x.

See Also

table_continuous_svy(), spicy_print_table()


Print method for continuous summary tables

Description

Formats and prints a spicy_continuous_table object as a styled ASCII table using spicy_print_table().

Usage

## S3 method for class 'spicy_continuous_table'
print(x, ...)

Arguments

x

A data.frame of class "spicy_continuous_table" as returned by table_continuous().

...

Additional arguments (currently ignored).

Value

Invisibly returns x.

See Also

table_continuous(), spicy_print_table()


Print method for spicy_cross_table objects

Description

Prints a formatted SPSS-like crosstable created by cross_tab().

Usage

## S3 method for class 'spicy_cross_table'
print(x, digits = NULL, decimal_mark = NULL, ...)

Arguments

x

A spicy_cross_table object.

digits

Optional integer; number of decimal places to display for cell values. Defaults to the value stored in the object.

decimal_mark

Optional character ("." or ",") used as the decimal mark. Defaults to the value stored in the object.

...

Additional arguments passed to internal formatting functions.

Value

Invisibly returns x.


Internal print method for lists of cross-tab tables

Description

Prints each element of a spicy_cross_table_list object on its own, inserting a blank line between tables.

Usage

## S3 method for class 'spicy_cross_table_list'
print(x, ...)

Arguments

x

A spicy_cross_table_list object.

...

Additional arguments passed to individual print methods.

Value

Invisibly returns x.


Print method for spicy-tagged flextables

Description

Prints a spicy_flextable object – the flextable::flextable() returned by output = "flextable", tagged so the table note keeps its styling in interactive HTML display. Every flextable verb works on the tagged object directly; printing delegates to flextable's own rendering (as_flextable.spicy_flextable() returns the untagged object).

Usage

## S3 method for class 'spicy_flextable'
print(x, ...)

Arguments

x

A spicy_flextable object.

...

Additional arguments (currently ignored).

Value

Invisibly returns NULL (HTML display path) or the result of flextable's own print method.

See Also

table_regression(), as_flextable.spicy_flextable()


Print method for freq() tables

Description

Formats and prints a spicy_freq_table object as a styled ASCII table using spicy_print_table().

Usage

## S3 method for class 'spicy_freq_table'
print(x, ...)

Arguments

x

A data.frame of class "spicy_freq_table" as returned by freq() (with the default output = "default"). Rendering metadata is read from attributes set by freq().

...

Additional arguments (currently ignored).

Value

Invisibly returns x.

See Also

freq(), spicy_print_table()


Print method for spicy-tagged gt tables

Description

Prints a spicy_gt object – the gt::gt() table returned by output = "gt", tagged so the table note keeps its styling in interactive HTML display. Every gt verb works on the tagged object directly; printing delegates to gt's own rendering.

Usage

## S3 method for class 'spicy_gt'
print(x, ...)

Arguments

x

A spicy_gt object.

...

Additional arguments (currently ignored).

Value

Invisibly returns NULL (HTML display path) or the result of gt's own print method.

See Also

table_regression()


Print method for outcome tables

Description

Formats and prints a spicy_outcome_table object as a styled ASCII table using spicy_print_table().

Usage

## S3 method for class 'spicy_outcome_table'
print(x, ...)

Arguments

x

A data.frame of class "spicy_outcome_table" as returned by table_outcome().

...

Additional arguments (currently ignored).

Value

Invisibly returns x.

See Also

table_outcome(), spicy_print_table()


Print method for internal regression frames

Description

Prints a compact one-glance summary of a spicy_regression_frame (the internal intermediate representation behind table_regression()): model class, sample size, coefficient dimensions, family / link, CI method, and the capability flags the frame advertises.

Usage

## S3 method for class 'spicy_regression_frame'
print(x, ...)

Arguments

x

A spicy_regression_frame object.

...

Additional arguments (currently ignored).

Value

Invisibly returns x.

See Also

table_regression()


Print method for regression tables

Description

Formats and prints a spicy_regression_table object as a styled ASCII table: title banner, spanner row for multi-model tables, decimal-aligned body with factor grouping and reference rows, fit-statistics block, and footer note.

Usage

## S3 method for class 'spicy_regression_table'
print(x, ...)

Arguments

x

A data.frame of class "spicy_regression_table" as returned by table_regression() with output = "default".

...

Additional arguments (currently ignored).

Value

Invisibly returns x.

See Also

table_regression(), spicy_print_table()


Print method for table styles

Description

Prints what a spicy_style() encodes: for a named theme, the journal, the rules it applies, and the official document they come from; then the levers themselves.

Usage

## S3 method for class 'spicy_style'
print(x, ...)

Arguments

x

A spicy_style object.

...

Additional arguments (currently ignored).

Value

Invisibly returns x.

See Also

spicy_style()


Simulated social-health survey

Description

A simulated dataset of 1200 respondents from a fictional social-health survey, designed to illustrate the main features of the spicy package: variable labels, ordered factors, survey weights, association measures, and publication-ready reporting.

Usage

sochealth

Format

A tibble with 1200 rows and 24 variables:

sex

Factor. Sex of the respondent.

age

Numeric. Age in years (25–75).

age_group

Ordered factor. Age group (25–34, 35–49, 50–64, 65–75).

education

Ordered factor. Highest education level (Lower secondary, Upper secondary, Tertiary).

social_class

Ordered factor. Subjective social class (Lower, Working, Lower middle, Middle, Upper middle).

region

Factor. Region of residence (6 regions).

employment_status

Factor. Employment status (Employed, Student, Unemployed, Inactive).

income_group

Ordered factor. Household income group (Low, Lower middle, Upper middle, High). Contains missing values.

income

Numeric. Monthly household income in CHF (1000–7400).

smoking

Factor. Current smoker (No, Yes). Contains missing values.

physical_activity

Factor. Regular physical activity (No, Yes).

dentist_12m

Factor. Dentist visit in the last 12 months (No, Yes).

self_rated_health

Ordered factor. Self-rated health (Poor, Fair, Good, Very good). Contains missing values.

wellbeing_score

Numeric. WHO-5 wellbeing index (0–100).

bmi

Numeric. Body mass index in kg/m^2 (16–39). Contains missing values.

bmi_category

Ordered factor. BMI category (Normal weight, Overweight, Obesity). Contains missing values.

institutional_trust

Ordered factor. Trust in institutions (Very low, Low, High, Very high).

political_position

Numeric. Political position on a 0 (left) to 10 (right) scale. Contains missing values.

life_sat_health

Integer. Satisfaction with own health (1–5 Likert scale). Contains missing values.

life_sat_work

Integer. Satisfaction with work or main activity (1–5 Likert scale). Contains missing values.

life_sat_relationships

Integer. Satisfaction with personal relationships (1–5 Likert scale). Contains missing values.

life_sat_standard

Integer. Satisfaction with standard of living (1–5 Likert scale). Contains missing values.

response_date

POSIXct. Date and time of survey response (September–November 2024).

weight

Numeric. Survey design weight (range 0.29–3.45); calibrated so that sum(weight) matches the unweighted N and mean(weight) is approximately 1. See Details.

Details

Every variable carries a "label" attribute (read by labelled::var_label() and surfaced by varlist() / code_book()). The mix of factor types is deliberate: nominal factors (sex, region, ...) and ordered factors (education, self_rated_health, ...) live side by side so that cross_tab() and table_categorical() can demonstrate the automatic ordinal-vs-nominal dispatch (Cramer's V, Phi, Kendall's Tau-b, Goodman-Kruskal Gamma) on the same dataset.

Survey weights (weight) are calibrated: sum(weight) matches the unweighted N to within rounding (\approx 1200) and mean(weight) is \approx 1. Weighted means therefore agree with unweighted means up to sampling noise without further rescaling.

Source

Simulated data for illustration purposes; reproducible by sourcing data-raw/sochealth.R. The script seeds the main generation block with set.seed(2025), and the two missing-value injection blocks with set.seed(2027) (the four ⁠life_sat_*⁠ items) and set.seed(2026) (smoking, self_rated_health, income_group, political_position, bmi).

Examples

data(sochealth)
varlist(sochealth)
freq(sochealth, education)
cross_tab(sochealth, education, self_rated_health)

Somers' D

Description

somers_d() computes Somers' D for a two-way contingency table of ordinal variables.

Usage

somers_d(
  x,
  direction = c("row", "column", "symmetric"),
  detail = FALSE,
  conf_level = 0.95,
  digits = 3L
)

Arguments

x

A contingency table (of class table).

direction

Direction of prediction: "row" (default, column predicts row), "column" (row predicts column), or "symmetric" (harmonic mean of both directions).

detail

Logical. If FALSE (default), return the estimate as a numeric scalar. If TRUE, return a named numeric vector including confidence interval and p-value.

conf_level

A single number strictly between 0 and 1 giving the confidence level (default 0.95). Only used when detail = TRUE. Set to NULL to omit the confidence interval. Any other value – including percentages such as 95 – raises a classed error (spicy_invalid_input).

digits

Number of decimal places used when printing the result (default 3). Only affects the detail = TRUE output.

Details

Somers' D is an asymmetric ordinal measure defined as d = (C - D) / (C + D + T), where T is the number of pairs tied on the independent variable. The symmetric version (direction = "symmetric") is the harmonic mean of the two asymmetric values, matching the SPSS / PSPP convention; this is not identical to Kendall's Tau-b (which is the geometric mean of the same two quantities), although the two often agree to two decimals. It is computed via the equivalent closed form 2(C - D) divided by the sum of the two asymmetric denominators, so a table with exactly as many concordant as discordant pairs (e.g. an independence pattern) yields the well-defined value 0 – the harmonic-mean form is 0/0 there – as printed by SPSS / PSPP. The symmetric estimate is NA only when one of the asymmetric directions is itself undefined (with the same spicy_undefined_stat warning). No analytic SE / CI is reported for the symmetric form: its se is always NA, matching PSPP, which prints no ASE for it (DescTools offers no symmetric form at all).

The default direction = "row" differs deliberately from lambda_gk() and uncertainty_coef() (which default to "symmetric"): Somers' D is intrinsically asymmetric, its symmetric form is a derived convention without inference, and the reference implementation (DescTools::SomersDelta()) offers only the two asymmetric directions, defaulting to "row".

Standard error formulas for the asymmetric directions follow the DescTools implementations (Signorell et al., 2024); see cramer_v() for full references.

Value

Same structure as cramer_v(): a scalar when detail = FALSE, a named vector when detail = TRUE. The p-value tests H0: D = 0 (Wald z-test).

References

Somers, R. H. (1962). A new asymmetric measure of association for ordinal variables. American Sociological Review, 27(6), 799-811. doi:10.2307/2090408

Brown, M. B., & Benedetti, J. K. (1977). Sampling behavior of tests for correlation in two-way contingency tables. Journal of the American Statistical Association, 72(358), 309-315. doi:10.1080/01621459.1977.10480995

See Also

kendall_tau_b(), gamma_gk(), assoc_measures()

Other association measures: assoc_measures(), contingency_coef(), cramer_v(), gamma_gk(), goodman_kruskal_tau(), kendall_tau_b(), kendall_tau_c(), lambda_gk(), phi(), uncertainty_coef(), yule_q()

Examples

tab <- table(sochealth$education, sochealth$self_rated_health)
somers_d(tab, direction = "row")
somers_d(tab, direction = "column", detail = TRUE)


Table labels and their language

Description

Every string a reader of a spicy table sees – column headers, row labels, titles, table footnotes – is held under a stable key, and spicy_labels() returns those keys with the label each one currently resolves to. It is the companion of the two options that move them.

Usage

spicy_labels(language = NULL)

Arguments

language

A language name ("en", "fr"), or NULL (the default) to report the labels in force. Any options(spicy.labels) override applies either way.

Value

A named character vector, one element per key, in registry order.

Global options

A label resolves through spicy.labels, then the spicy.language set, then English. A set carries only the keys it translates, so anything it does not name falls back to English rather than erroring or coming out blank. Both options are cleared with NULL.

What a language does not change

Only DISPLAY strings translate. The column names of the exported frames (as.data.frame(), tidy(), as_structured()) are a documented contract that user code indexes into, so out[["Yes %"]] resolves under every language, and so do the block identities, the encoded cell tokens and the mathematical glyphs. Errors, warnings and messages stay English: they are read by developers and quoted in bug reports.

One exception, and it follows from the same rule. A column named after a LEVEL of by takes that level's own spelling – which is why "Yes %" stays "Yes %": your data is never translated. spicy's own missing category is such a level, so it is the one name that does move: table_categorical(by = ) on a variable with missing values gives "(Missing) n" in English and "(Manquant) n" under "fr" (or whatever row_missing_level is overridden to). Address that column through row_missing_level rather than by typing it.

Numbers follow the language too

A language brings its typographic locale with it, so options(spicy.language = "fr") is one gesture for a coherent French report table: French words, comma decimal mark, and the leading zero French typography keeps on a p-value (⁠0,003⁠). The sources are the BIPM's SI brochure and the European Union's Code de redaction interinstitutionnel; spicy_style() quotes them. "en" brings no locale, and nothing changes for anyone who sets no language.

The locale rides the style layer, so it reaches the reporting families – table_regression(), table_categorical(), table_continuous(), table_continuous_lm(), table_outcome() and the survey twins. The language reaches every table, the exploration pair included: freq() and cross_tab() have no style layer, so it sets the DEFAULT of their decimal_mark – the one typographic lever they carry – and an argument you type wins over it. Under a comma their p-value keeps its leading zero (⁠p = 0,659⁠), the form French typography requires.

The locale sits at the BOTTOM of the formatting resolution. A journal style outranks it – style = "jama" under "fr" gives JAMA's p-values and the French comma, style = "lancet" keeps the journal's midline decimal point – and an argument you type outranks both, so decimal_mark = "." gives French words with a decimal point (the p-value keeps its locale zero; spicy_style() documents that lever).

Two timing notes. Figures are frozen when a table is BUILT, like every formatting argument, while words resolve when it is printed – so set the language once, at the top of the document. And a comma mark makes an Excel export write its body as text rather than live numbers, since Excel would otherwise re-punctuate a numeric cell with the viewer's own locale.

See Also

spicy_style() for the journal styles, which outrank the language's own typography.

Examples

head(spicy_labels())

# The same keys in French.
fr <- spicy_labels("fr")
fr[["header_mean"]]
fr[["row_missing_level"]]

# One label, not a language: the missing CATEGORY of the grouping
# variable is a refusal to answer here, not an absent value.
old <- options(spicy.labels = list(row_missing_level = "(No answer)"))
table_categorical(sochealth, select = sex, by = smoking)
options(old)


Print a spicy-formatted ASCII table

Description

User-facing helper that prints a spicy-styled ASCII table to the console with optional title and note, table-type-aware alignment defaults, and automatic horizontal panelling when the table is wider than the console. Wraps the internal renderer build_ascii_table().

Usage

spicy_print_table(
  x,
  title = attr(x, "title"),
  note = attr(x, "note"),
  padding = 2L,
  first_column_line = TRUE,
  row_total_line = TRUE,
  bottom_line = FALSE,
  lines_color = "darkgrey",
  align_left_cols = NULL,
  align_center_cols = integer(0),
  center_headers = FALSE,
  spanners = NULL,
  group_sep_rows = integer(0),
  total_row_idx = attr(x, "total_row_idx"),
  display_labels = NULL,
  fit_stats_start = NULL,
  qualify_companions = FALSE,
  ...
)

Arguments

x

A spicy_table or data.frame to be printed.

title

Optional title displayed above the table. Defaults to the "title" attribute of x if present.

note

Optional note displayed below the table. Defaults to the "note" attribute of x if present.

padding

Non-negative integer giving the number of extra characters added to each column's auto-computed width (max of cell-content width and header width). Defaults to 2L. See build_ascii_table() for the precise formula and the migration note from the pre-0.11.0 string enum.

first_column_line

Logical. If TRUE (the default), adds a vertical separator after the first column.

row_total_line, bottom_line

Logical flags controlling the horizontal line before a total row and the closing line at the bottom of the table. row_total_line defaults to TRUE; bottom_line defaults to FALSE.

lines_color

Character. Color for table separators. Defaults to "darkgrey". Only applied if the output supports ANSI colors (see crayon::has_color()).

align_left_cols

Integer vector of column indices to left-align. If NULL (the default), alignment is auto-detected based on x:

  • For freq tables -> c(1, 2)

  • For cross tables -> 1

align_center_cols

Integer vector of column indices to center-align. Defaults to integer(0).

center_headers

Logical. When TRUE, column headers are centered above their column content even when the data itself is right-aligned. Passed through to build_ascii_table(). Defaults to FALSE.

spanners

Optional named list of column-group labels (label -> integer column indices). Passed through to build_ascii_table(); when the table is split into horizontal panels each panel keeps only the spanners whose columns are fully contained in it. Defaults to NULL (no spanner row).

group_sep_rows

Integer vector of row indices before which a light dashed separator line is drawn. Defaults to integer(0).

total_row_idx

Optional integer vector of 1-based row indices identifying the totals rows; defaults to the "total_row_idx" attribute of x (set by cross_tab()). See build_ascii_table().

display_labels

Optional character vector of length ncol(x) used to override colnames(x) for the printed header text only. Sliced per panel when the table is split across stacked panels. Forwarded to build_ascii_table(); see that function for full semantics. Defaults to NULL.

fit_stats_start

Optional 1-based index of the first model-level statistics row (the block below the dashed rule in regression tables). When the table splits into stacked panels, continuation panels drop the rows of that block whose every visible data cell is blank – model-level statistics print once, under the panel that carries their values, instead of leaving empty n / AIC stub rows on every continuation panel. NULL (default) keeps all rows on all panels.

qualify_companions

Logical. When the table splits into stacked panels, should a companion column (SE, p, an interval) that a width split separated from its carrier name that carrier in its header – ⁠95% CI⁠ becoming ⁠95% CI (B)⁠? Set it TRUE only for a layout where those headers really are companions of the estimate column on their left, as in a coefficient table. In a layout whose p is the omnibus test of a whole block it has no carrier, and qualifying it would attribute it to whichever column happened to sit before it. Defaults to FALSE.

...

Additional arguments passed to build_ascii_table().

Details

Table type is auto-detected from x and drives the default alignment when align_left_cols = NULL:

If the table is wider than the console, it is split into stacked horizontal panels with the left-most identifier columns repeated on each panel. Unicode line-drawing characters are used by default; coloured separators are drawn when the terminal supports ANSI colour (crayon::has_color()) and fall back to monochrome otherwise.

The layout arguments spanners, display_labels, fit_stats_start, total_row_idx and group_sep_rows are plumbing consumed by spicy's own print methods; they are documented for completeness and are rarely useful when calling this function directly.

Value

Invisibly returns x, after printing the formatted ASCII table to the console.

See Also

build_ascii_table() for the underlying text rendering engine. print.spicy_freq_table() for the specialized printing method used by freq().

Examples

# Simple demonstration
df <- data.frame(
  Category = c("Valid", "", "Missing", "Total"),
  Values = c("Yes", "No", "NA", ""),
  Freq. = c(12, 8, 1, 21),
  Percent = c(57.1, 38.1, 4.8, 100.0)
)

spicy_print_table(df,
  title = "Frequency table: Example",
  note = "Class: data.frame\nData: demo"
)


Build or select a table style

Description

A style is a small set of number-formatting rules – how many decimals a p-value gets, where it bottoms out, whether it keeps its leading zero, what the decimal mark is, how a confidence interval is written. spicy_style() composes one by hand; the named themes ("jama", "nejm", "lancet", "annals", "apa", "aer" – spicy_style_names() returns the list) are pre-composed ones, each encoding rules taken verbatim from an official document of the institution.

Usage

spicy_style(
  base = NULL,
  ...,
  p_style = NULL,
  p_digits = NULL,
  p_floor = NULL,
  p_bands = NULL,
  p_sigfig = NULL,
  decimal_mark = NULL,
  ci_sep = NULL,
  ci_brackets = NULL,
  stars = NULL,
  digits = NULL,
  effect_size_digits = NULL,
  fit_digits = NULL,
  ic_digits = NULL,
  percent_digits = NULL,
  v_digits = NULL
)

spicy_style_names()

Arguments

base

Optional theme to start from – a theme name or another spicy_style – whose levers the arguments below override. This is the composition escape hatch: spicy_style("lancet", ci_sep = " to ") is The Lancet's rules with one of them changed. The theme's provenance travels with the result and names the levers you overrode, so a modified theme never passes for the theme itself.

...

Must be empty. Any argument landing here is a misspelt lever and raises an error rather than being ignored.

p_style

How p-values carry their leading zero: "apa" drops it (.003), "standard" keeps it (0.003). NULL leaves spicy's default, which follows the mark: dropped under a point, kept under a comma.

p_digits

Decimal places for p-values (a positive integer). Under p_sigfig it acts as the decimal cap instead.

p_floor

The value below which a p-value prints as ⁠<floor⁠ rather than exactly, e.g. 0.001 for ⁠<.001⁠. A number strictly between 0 and 1. NULL uses 10^-p_digits.

p_bands

Decimal places that vary with the size of the p-value: a list of two-element numeric vectors c(cutoff, digits), cutoffs strictly increasing, the last one Inf. list(c(0.01, 3), c(Inf, 2)) reads as "three decimals below .01, two otherwise". The comparison is p < cutoff. Mutually exclusive with p_sigfig.

p_sigfig

Significant figures for p-values (a positive integer), capped at p_digits decimal places. Mutually exclusive with p_bands.

decimal_mark

A single character; the mark between the integer and the decimal part of every number in the table.

ci_sep

The string between the two bounds of a confidence interval. NULL keeps spicy's default, which is ", " and turns into "; " when the decimal mark is a comma.

ci_brackets

A character vector of length two, the opening and closing delimiters of a confidence interval, e.g. c("[", "]") or c("(", ")").

stars

Significance stars: FALSE to suppress them, or a named numeric vector of symbol-to-threshold pairs such as c("***" = .001, "**" = .01, "*" = .05).

digits, effect_size_digits, fit_digits, ic_digits

Decimal places for estimates, effect sizes, fit statistics and information criteria. Each is applied only by the table functions that have the matching argument.

percent_digits, v_digits

Decimal places for percentages and for the association measure of table_categorical().

Details

A style is accepted by the style argument of table_regression(), table_categorical(), table_continuous() and table_continuous_lm(), either as a theme name or as the object returned here, and by options(spicy.style = ) for document-wide scope.

Value

spicy_style() returns an object of class spicy_style: a named list holding only the levers that were set. Themes returned by name carry a provenance attribute with the journal, its source document and the exact list of encoded rules; print() shows it. spicy_style_names() returns the character vector of available theme names.

What a theme claims, and what it does not

A theme covers numeric formatting conformity only – not full editorial conformity. It does not check reporting guidelines, table structure, footnote symbols, units, abbreviation policy, or anything else a journal's instructions ask of a manuscript. Each theme below lists the exact rules it encodes, with the sentence it encodes them from; anything not listed is spicy's own default, not the journal's rule.

Themes only move defaults. An argument you type wins over the theme, always: table_regression(fit, style = "lancet", digits = 3) is a Lancet table with three decimals. One consequence worth stating: typing p_digits is a request for that many decimals on every p-value, so it also switches off the theme's own ways of choosing p-precision – p_bands, p_sigfig, and the p_floor derived from them. The leading-zero rule is orthogonal and stays.

Resolution order

An argument you type, then the style argument, then getOption("spicy.style"), then the typography of getOption("spicy.language"), then spicy's defaults. Within a style, an explicit function argument beats the style's value for the same lever.

The language's locale

A language brings its typography with it, at the bottom of that order. options(spicy.language = "fr") is therefore one gesture for a coherent French report table – French words, and numbers written the French way: comma decimal mark, and the leading zero French typography keeps on a p-value (⁠0,003⁠). "La virgule est utilisee pour separer les unites des decimales" (European Union, Code de redaction interinstitutionnel, French edition 2022, point 6.5, ⁠https://style-guide.europa.eu/fr⁠); "Si le nombre se situe entre +1 et -1, le separateur decimal est toujours precede d'un zero" (BIPM, Le Systeme international d'unites, 9th edition, 2019, section 5.4.4, ⁠https://www.bipm.org/documents/20126/41483022/SI-Brochure-9.pdf⁠). "en" brings no locale, and nothing changes for anyone who sets no language. The locale reaches the exploration pair as well: freq() and cross_tab() have no style layer, so the language sets the DEFAULT of their decimal_mark and nothing else. An argument you type still wins. The leading zero of a p-value works the same way in both worlds: its DEFAULT follows the mark – a comma keeps it (⁠0,003⁠), a point drops it (.003) – and p_style is the explicit lever that overrides that default wherever there is one, which is why a theme's rule survives any mark in the ⁠table_*()⁠ families while the pair, having no such lever, always follows its mark.

A theme composes with a locale rather than fighting it, because a theme encodes only what its own source states: "jama" fixes no decimal mark, so JAMA under a French language gives JAMA's p-value rules and the French comma. Where the two do meet, the theme wins – you asked for it by name. "lancet" keeps its midline decimal point; "apa" keeps its missing leading zero, which under a comma prints ⁠,003⁠, a form the SI brochure forbids. That is the price of an explicit gesture; to keep APA's other rules and restore the zero, compose the way out yourself: style = spicy_style("apa", p_style = "standard"). One lever bends the other way: a theme's ", " interval separator was sourced under a dot mark, and under a comma mark it would BE the mark – so it yields to the derived "; ", exactly as the French adaptations of APA style themselves write (⁠[3,45; 6,78]⁠). Any other separator (" to ", an en dash) is unambiguous and stays.

An argument beats both, which is the escape hatch for a bilingual table: decimal_mark = "." under a French language gives French words and a decimal point. It moves only the mark: the p-value keeps the locale's leading zero (0.003), because p_style is a style lever with no argument of its own – style = spicy_style(p_style = "apa") drops it again.

See spicy_labels() for the language option itself.

Themes

"jama" – JAMA / JAMA Network

Source: Instructions for Table Creation, JAMA Network author document, 23 February 2016 (⁠https://jamanetwork.com/DocumentLibrary/InstructionsForAuthors/InstructionsForTableCreation.pdf⁠), consulted 2026-08-14.

Encoded:

Not encoded: everything else in that document (percentages carrying their numerator and denominator, one datum per cell, footnote letters, SI conversion factors) is table construction, not number formatting. The document states no rule for the confidence-interval separator or the decimal mark, so spicy's defaults apply.

"nejm" – The New England Journal of Medicine

Source: NEJM Author Center, New Manuscripts, section Statistical Reporting Guidelines (⁠https://www.nejm.org/author-center/new-manuscripts⁠), a living page; its full text was read in a browser by the package maintainer on 2026-08-14 – the site refuses automated retrieval, which is why earlier surveys reported these rules as unavailable.

Encoded:

Not encoded: the guideline's stated exceptions to the p rule (stopping-rule tests, genomewide studies) are analysis contexts the style layer cannot see; its inference policy (no p-values without a prespecified multiplicity plan, estimates + 95% CI instead, no p-values in the Table 1 of a randomised trial) is about what to report, not how to format it – request those layouts through show_columns and p_value = FALSE where you need them.

"lancet" – The Lancet

Sources: Information for Authors, April 2026; Randomised trials in The Lancet: formatting guidelines and Observational studies in The Lancet: formatting guidelines, both last updated July 2025 (⁠https://www.thelancet.com/pb-assets/Lancet/authors/tl-info-for-authors-1690986041530.pdf⁠), consulted 2026-08-14.

Below, the journal's midline decimal point (Unicode U+00B7, MIDDLE DOT) is written ⁠[.]⁠ so it cannot be confused with an ordinary full stop.

Encoded:

Not encoded: the empty-cell filler, the ban on p-values in a randomised trial's baseline table, and the absolute-rather-than- relative effect rule are content decisions, not number formats.

One caveat on the en dash. The journal's examples are ratio measures, whose bounds are positive. On the identity scale a negative lower bound puts a minus sign next to the dash (⁠[-5.17--2.58]⁠), which reads badly. The journal states no rule for that case and none is invented here; override the separator when it arises: spicy_style("lancet", ci_sep = " to ").

"annals" – Annals of Internal Medicine

Source: Information for Authors, American College of Physicians, document publication date 08/04/2026 (⁠https://www.acpjournals.org/pb-assets/pdf/AnnalsAuthorInfo-1755188286957.pdf⁠), consulted 2026-08-14.

Encoded:

Not encoded: the percentage rule ("Report percentages to one decimal place ... when sample size is \ge 200", no decimals below 200) is conditional on a per-group sample size that the style layer does not see; set percent_digits yourself. The mean (SD) notation the journal asks for, and its refusal of ⁠\eqn{\pm}{+/-}⁠, are already how spicy renders dispersion.

"apa" – APA Style, 7th edition

Sources: APA Style numbers and statistics guide, last updated 11 September 2024 (⁠https://apastyle.apa.org/instructional-aids/numbers-statistics-guide.pdf⁠); Sample tables, last updated June 2024 (⁠https://apastyle.apa.org/style-grammar-guidelines/tables-figures/sample-tables⁠), both consulted 2026-08-14.

Encoded – and this theme pins rather than changes: spicy's defaults already follow the guide, so style = "apa" is a promise that a future change of those defaults will not move an APA-formatted table.

Not encoded: the one-decimal rule for means of integer scales (depends on what the variable measures, which spicy cannot know), and the thousands separator (see Known gaps below). Nor the star thresholds. The official correlation and ANOVA sample tables mark .05 / .01 / .001, which is what stars = TRUE gives you – but as spicy's own default, not as anything this theme pins: one stars lever carries both the thresholds and the decision to show them, and the APA regression sample table shows none. So the theme sets no stars, and style = "apa" never switches them on.

"aer" – American Economic Review / AEA journals

Source: AER Style Guide, American Economic Association (⁠https://www.aeaweb.org/journals/aer/style-guide⁠), consulted 2026-08-14. The Tables section is shared with AEJ: Applied, AEJ: Macro, AEJ: Policy, JEL and AEA P&P.

Encoded:

Not encoded: the guide fixes no number of decimals, no p-value floor, no interval format and no decimal mark, so spicy's defaults apply. Horizontal-rules-only, no shading, a nine-column maximum and "Panel A / Panel B" blocks are layout, not number formatting.

Known gaps

Two rules the sources state but this release does not encode:

Journals deliberately absent: BMJ (its house-style page could not be read from any official source, and the widely repeated "95% CI 1.2 to 3.4" rule traces to no BMJ document), Econometrica (verified negative: the official guidelines contain no numeric table rule), QJE (its instructions page could not be read), Epidemiology (see Known gaps). Naming any of them would mean inventing their rules.

Examples

fit <- lm(mpg ~ wt + hp, data = mtcars)

# A named theme.
table_regression(fit, style = "jama")

# An argument you type beats the theme.
table_regression(fit, style = "jama", p_digits = 4)

# A style composed by hand.
table_regression(fit, style = spicy_style(decimal_mark = ",",
                                          p_style = "standard"))

# A theme with one rule changed.
table_regression(fit, style = spicy_style("lancet", ci_sep = " to "))

# Document-wide scope.
old <- options(spicy.style = "apa")
table_continuous(mtcars, c(mpg, wt))
options(old)

# What a theme encodes, and where it comes from.
spicy_style("lancet")

Spicy table engine

Description

Index page for the ASCII rendering engine that produces every spicy console table (freq(), cross_tab(), the ⁠table_*()⁠ family, and the association-measure printers). The engine supports Unicode line drawing, ANSI colours via crayon (with monochrome fallback), automatic colour-aware width detection, configurable integer padding (0L / 2L / 4L), per-column alignment, and horizontal panelling for tables wider than the console.

User-facing entry points

Rendering primitives (internal API)

See Also

spicy for the full package overview, including the API stability tiers and the classed-condition taxonomy.


Row sums with an optional minimum-valid-values rule

Description

Computes row-wise sums across selected numeric columns of a data.frame or matrix. Missing values are handled per row via min_valid (an integer count or proportion of non-NA values required); rows that fail the rule return NA, and rows with no valid values at all return NA even when min_valid = 0. Non-numeric columns are dropped silently (set verbose = TRUE to see which). Designed to flow inside dplyr::mutate(): when called without an explicit data argument, the current data context is used.

Usage

sum_n(
  data = NULL,
  select = tidyselect::everything(),
  exclude = NULL,
  min_valid = NULL,
  digits = NULL,
  regex = FALSE,
  verbose = FALSE,
  user_na = TRUE
)

Arguments

data

A data.frame or matrix. Optional inside dplyr::mutate(), where the current grouping/data context is used automatically.

select

Columns to include. If regex = FALSE, use tidyselect syntax (default: tidyselect::everything()). If regex = TRUE, provide a regular expression pattern (character string).

exclude

Columns to exclude (default: NULL).

min_valid

Minimum number of valid (non-NA) values required per row. Accepts:

  • NULL (the default) – every selected column must be valid.

  • a proportion in ⁠(0, 1)⁠ – round(ncol(x) * min_valid) valid columns required (e.g. min_valid = 0.5 requires at least half of the selected columns to be non-NA).

  • a non-negative integer count up to the number of selected numeric columns.

Non-integer values ⁠>= 1⁠ (e.g. 1.5) and counts greater than ncol(x) raise an actionable error.

Rows with zero valid values always return NA, even when min_valid = 0: an empty row-wise summary is undefined, so the raw rowMeans() / rowSums() identities (NaN / 0) are never returned.

digits

Optional non-negative integer giving the number of decimal places to round the result to. Defaults to NULL (no rounding).

regex

Logical. If FALSE (the default), uses tidyselect helpers. If TRUE, the select argument is treated as a regular expression.

verbose

Logical. If FALSE (the default), messages are suppressed. If TRUE, prints a message about non-numeric columns excluded.

user_na

Logical. If TRUE (the default), declared missing values count as missing – both in the computed summary and in the min_valid valid-count gate. If FALSE, the declared codes are treated as ordinary numbers. See the "Declared missing values" section of freq().

Value

A numeric vector of row-wise sums.

Declared missing values

Survey files imported with haven often carry declared missing values: codes such as ⁠8 = Don't know⁠ or 9 = Refused that the source file marks as missing while keeping them distinct from a plain NA. Two kinds of declaration exist: na_values / na_range metadata on haven::labelled_spss() vectors, and tagged missing values created by haven::tagged_na() (the Stata .a, .b, ... convention).

spicy honors the declaration by default (user_na = TRUE): declared missing values are excluded from every statistic exactly like NA – valid percentages, means, chi-squared tests, association measures, row-wise summaries, and group definitions – but they are not erased from display. freq() lists each observed declared value as its own row of the Missing block, with its value label; cross_tab(), table_categorical(), and table_continuous() disclose the exclusion in the table note (⁠Declared missing values removed: x (2).⁠); varlist() and code_book() count them as missing in N_valid / NAs / N_distinct while still listing the declared codes in Values.

Every function involved offers the same escape hatch: set user_na = FALSE to ignore the declaration and treat the declared codes as valid values (the behavior of spicy before 0.13.0). Tagged missing values are genuine NAs either way; for them, user_na = FALSE only collapses the per-tag breakdown back into the regular NA count.

See Also

Other row-wise summaries: count_n(), mean_n()

Examples

library(dplyr)

# Create a simple numeric data frame
df <- tibble(
  var1 = c(10, NA, 30, 40, 50),
  var2 = c(5, NA, 15, NA, 25),
  var3 = c(NA, 30, 20, 50, 10)
)

# Compute row-wise sums (all values must be valid by default)
sum_n(df)

# Require at least 2 valid (non-NA) values per row
sum_n(df, min_valid = 2)

# Require at least 50% valid (non-NA) values per row
sum_n(df, min_valid = 0.5)

# Round the results to 1 decimal
sum_n(df, digits = 1)

# Select specific columns
sum_n(df, select = c(var1, var2))

# Select specific columns using a pipe
df |>
  select(var1, var2) |>
  sum_n()

# Exclude a column
sum_n(df, exclude = "var3")

# Select columns ending with "1"
sum_n(df, select = ends_with("1"))

# Use with native pipe
df |> sum_n(select = starts_with("var"))

# Use inside dplyr::mutate()
df |> mutate(sum_score = sum_n(min_valid = 2))

# Select columns directly inside mutate()
df |> mutate(sum_score = sum_n(select = c(var1, var2), min_valid = 1))

# Select columns before mutate
df |>
  select(var1, var2) |>
  mutate(sum_score = sum_n(min_valid = 1))

# Show verbose message
df |> mutate(sum_score = sum_n(min_valid = 2, digits = 1, verbose = TRUE))

# Add character and grouping columns
df_mixed <- mutate(df,
  name = letters[1:5],
  group = c("A", "A", "B", "B", "A")
)
df_mixed

# Non-numeric columns are ignored
sum_n(df_mixed)

# Use inside mutate with mixed data
df_mixed |> mutate(sum_score = sum_n(select = starts_with("var")))

# Use everything(), but exclude known non-numeric
sum_n(df_mixed, select = everything(), exclude = "group")

# Select columns using regex
sum_n(df_mixed, select = "^var", regex = TRUE)
sum_n(df_mixed, select = "ar", regex = TRUE)

# Apply to a subset of rows
df_mixed[1:3, ] |> sum_n(select = starts_with("var"))

# Store the result in a new column
df_mixed$sum_score <- sum_n(df_mixed, select = starts_with("var"))
df_mixed

# With a numeric matrix
mat <- matrix(c(1, 2, NA, 4, 5, NA, 7, 8, 9), nrow = 3, byrow = TRUE)
mat
mat |> sum_n(min_valid = 2)


Categorical summary table

Description

Builds a publication-ready frequency or cross-tabulation table for one or many categorical variables selected with tidyselect syntax.

With by, produces grouped cross-tabulation summaries (using cross_tab() internally) with Chi-squared p-values and optional association measures. Without by, produces one-way frequency-style summaries.

Multiple output formats are available via output: a printed ASCII table ("default"), a wide or long numeric data.frame ("data.frame", "long"), or publication-ready tables ("tinytable", "gt", "flextable", "excel", "clipboard", "word").

Usage

table_categorical(
  data,
  select = tidyselect::everything(),
  by = NULL,
  labels = NULL,
  levels_keep = NULL,
  include_total = TRUE,
  drop_na = FALSE,
  weights = NULL,
  rescale = FALSE,
  correct = FALSE,
  simulate_p = FALSE,
  simulate_B = 2000,
  percent_digits = 1,
  p_digits = 3,
  v_digits = 2,
  assoc_measure = "auto",
  assoc_ci = FALSE,
  smd = FALSE,
  decimal_mark = ".",
  align = c("decimal", "center", "right"),
  output = c("default", "data.frame", "long", "tinytable", "gt", "flextable", "excel",
    "clipboard", "word"),
  indent_text = "  ",
  indent_text_excel_clipboard = strrep(" ", 6),
  add_multilevel_header = TRUE,
  blank_na_wide = FALSE,
  excel_path = NULL,
  excel_sheet = NULL,
  clipboard_delim = "\t",
  word_path = NULL,
  user_na = TRUE,
  style = NULL
)

Arguments

data

A data frame.

select

Columns to include as row variables. Supports tidyselect syntax and character vectors of column names. When omitted, defaults to every eligible categorical column in data: factor, character, logical, and labelled (haven / labelled) columns, excluding the by column – matching the select-less defaults of table_continuous() and table_continuous_lm(). An explicit select is taken verbatim (numeric columns included), so numeric-coded categorical variables can still be tabulated by naming them.

by

Optional grouping column used for columns/groups. Accepts an unquoted column name or a single character column name. Factor levels keep their declared order; any other by (character, numeric, haven labelled) forms group columns in order of first appearance in the data – the same convention as table_continuous(). For a haven labelled by, the group headers are the raw codes (value labels are not used for group headers – the family convention shared with table_continuous() and table_continuous_lm()); declared missing values follow user_na as usual.

labels

An optional named character vector of variable labels whose names match column names in data (e.g. c(smoking = "Current smoker")) – the same contract as table_continuous() and table_continuous_lm(). Only listed columns are relabelled. For the remaining columns (and when labels = NULL, the default), labels are auto-detected from the variable's label attribute (e.g. from haven); if none is found, the column name is used. Unnamed (positional) label vectors, accepted before 0.13.0, now raise an error.

levels_keep

Optional character vector of levels to keep/order for row modalities. If NULL, all observed levels are kept. Entries must match the level strings the table displays (for labelled columns these are the "[code] label" strings, not the bare label text). When nothing matches for a selected variable, that variable is dropped from the table with a classed warning (spicy_no_selection) listing the available level strings.

include_total

Logical. If TRUE (the default), includes a Total group when available.

drop_na

Logical. If FALSE (the default), missing values are displayed as a dedicated "(Missing)" level (and, under by, a "(Missing)" group column) – the field convention for descriptive tables (gtsummary's "Unknown" row, janitor's NA row; see the Epidemiologist R Handbook, Descriptive tables). If TRUE, rows with NA in the tabulated variable (and in by, when supplied) are removed BEFORE each cross-tabulation, and the removal is disclosed in a table note ("Missing values removed: ...") rather than silent. Before 0.13.0 the default was TRUE with no disclosure.

weights

Optional weights. Either NULL (the default), a numeric vector of length nrow(data), or a single column in data supplied as an unquoted name or a character string.

rescale

Logical. If FALSE (the default), weights are used as-is. If TRUE, rescales weights so total weighted N matches raw N. Passed to spicy::cross_tab(). When the argument is not supplied, the default is read from options(spicy.rescale) (falling back to FALSE), matching cross_tab().

correct

Logical. If FALSE (the default), no continuity correction is applied. If TRUE, applies Yates correction in 2x2 chi-squared contexts. Passed to spicy::cross_tab().

simulate_p

Logical. If FALSE (the default), uses asymptotic p-values. If TRUE, uses Monte Carlo simulation. Passed to spicy::cross_tab().

simulate_B

Integer. Number of Monte Carlo replicates when simulate_p = TRUE. Defaults to 2000.

percent_digits

Number of digits for percentages in report outputs. Defaults to 1.

p_digits

Integer >= 1. Number of decimal places used to render p-values in the p column (default: 3, the APA Publication Manual standard). Both the displayed precision and the small-p threshold derive from this argument: p_digits = 3 prints .045 and ⁠<.001⁠; p_digits = 4 prints .0451 and ⁠<.0001⁠. Leading zeros are always stripped, following APA convention.

v_digits

Number of digits for the association measure. Defaults to 2.

assoc_measure

Which association measure to report alongside the chi-squared p-value. Accepts four input shapes:

  • "none" – drop the column entirely.

  • "auto" (the default) – pick a measure per row variable based on the variable type: a 2x2 table (binary row variable vs. binary by) uses phi, a pair of ordered factors uses tau_b, every other case uses cramer_v.

  • a single string from c("cramer_v", "phi", "gamma", "tau_b", "tau_c", "somers_d", "lambda") – applied uniformly to every row variable.

  • a character vector with one entry per row variable. Both named (c(smoking = "phi", health = "tau_b"), recommended; unnamed variables fall back to "auto") and unnamed positional (c("phi", "tau_b", "auto"), paired up with select) are accepted. Named is more robust to reordering of select.

When a single measure is used for every row, the column header is that measure's name (e.g. "Cramer's V"). When multiple measures are used (typically with "auto" on a heterogeneous select), the header collapses to "Effect size" and an APA-style Note. line is appended documenting which measure was used for which variable.

phi requires a 2x2 table; if explicitly requested for a non-2x2 variable, an error is raised so the user can choose another measure or fall back to "auto".

assoc_ci

Passed to cross_tab(). If TRUE, includes the confidence interval of the association measure. In wide raw outputs ("data.frame", "excel", "clipboard"), two extra columns ⁠CI lower⁠ / ⁠CI upper⁠ are added; in the long raw output ("long") the bounds appear as ci_lower / ci_upper. In rendered formats ("gt", "tinytable", "flextable", "word"), the CI is shown inline (e.g., .14 [.08, .19]). Defaults to FALSE.

smd

Logical. If TRUE, adds an SMD column holding the standardized mean difference between the two groups of by, the balance diagnostic of the Table 1 literature. Requires exactly two groups; the value sits on the variable row beside p, never on a level row. Signed for a two-category variable (group 1 minus group 2 on the second category), unsigned for three or more, where it is a distance. No confidence interval and no p-value, by design. Rounded with v_digits. See the "Standardized mean difference" section below. Defaults to FALSE.

decimal_mark

Decimal separator ("." or ","). Defaults to ".".

align

Horizontal alignment of numeric columns in the printed ASCII table and in the tinytable, gt, flextable, word, and clipboard outputs. The first column (Variable) is always left-aligned. One of:

  • "decimal" (default): align numeric columns on the decimal mark, the standard scientific-publication convention used by SPSS, SAS, and LaTeX siunitx. Numeric cells are pre-padded with figure-spaces (U+2007, digit-width) so every string in a column has the same width with the decimal mark at the same internal position; centring those uniform-width strings then stacks the decimal points vertically. The same pad-then-centre strategy is applied on every rendering engine (gt, tinytable, flextable, word, ASCII print) for a homogeneous rendering, matching table_regression() and table_continuous_lm(). The clipboard output is delimited text meant to be parsed rather than read at a fixed width, so its cells travel unpadded (a padded number pastes as text next to an unpadded number).

  • "center": center-align all numeric columns.

  • "right": right-align all numeric columns.

In the excel output, "center" centres the numeric columns and "right" is the same rendering as the default: cell-string padding does not align decimals under a proportional font, so "decimal" right-aligns instead, which combined with the per-column numfmt already produces dot-aligned columns. Same default and same three values as table_continuous() / table_continuous_lm(), whose workbooks resolve "decimal" differently: table_continuous() right-aligns only the counts and the p-value there, and table_continuous_lm() applies that convention at every align.

output

Output format. One of:

  • "default" (an ASCII table object, printed when the call is bare)

  • "data.frame" (a wide numeric data.frame)

  • "long" (a long numeric data.frame)

  • "tinytable" (requires tinytable)

  • "gt" (requires gt)

  • "flextable" (requires flextable)

  • "excel" (requires openxlsx2)

  • "clipboard" (requires clipr)

  • "word" (requires flextable and officer)

indent_text

Prefix used for modality labels in report table building. Defaults to " " (two spaces).

indent_text_excel_clipboard

Stronger indentation used in Excel and clipboard exports. Defaults to six non-breaking spaces.

add_multilevel_header

Logical. If TRUE (the default), merges top headers in Excel export. Only consulted for output = "excel" on a grouped table (by supplied); like the other output-scoped presentation arguments (excel_sheet, clipboard_delim, ...), it is silently unused in every other output and in one-way tables, which have a single header row to begin with.

blank_na_wide

Logical. If FALSE (the default), NA values are kept as-is in wide raw output. If TRUE, replaces them with empty strings.

excel_path

Path for output = "excel". Defaults to NULL.

excel_sheet

Sheet name for Excel export. NULL (the default) uses "Categorical".

clipboard_delim

Delimiter for clipboard text export. Defaults to "\t". A cell holding the delimiter itself, a double quote or a line break is quoted RFC 4180-style, so the grid survives whatever delimiter you choose.

word_path

File path for output = "word". Defaults to NULL. Before 0.13.0, supplying it with output = "flextable" also wrote a .docx as a side effect; it is now used exclusively by output = "word" (the contract shared with the rest of the table family) and is ignored, with a warning, under output = "flextable".

user_na

Logical. If TRUE (the default), declared missing values in the row variables and in by are treated as missing: they join the "(Missing)" level under drop_na = FALSE, and under drop_na = TRUE they are removed with a dedicated disclosure line (⁠Declared missing values removed: ...⁠). If FALSE, the declared codes stay valid categories. See the "Declared missing values" section of freq().

style

A journal style: a theme name ("jama", "nejm", "lancet", "annals", "apa", "aer"), a spicy_style() object, or NULL (the default). A style only changes DEFAULTS – any argument you pass explicitly wins over it. Set options(spicy.style = ) for document-wide scope. A theme covers numeric formatting conformity only, not full editorial conformity; ?spicy_style lists the exact rules each one encodes and the official document they come from. An unknown name is an error.

Value

Depends on output:

The drop_na = TRUE disclosure travels with the table on every route, not just the console: "default" prints it under the ASCII table, "tinytable" / "gt" / "flextable" / "word" carry it as a table note, "excel" writes it below the body, and "data.frame" keeps the sentence verbatim in the missing_note attribute (attr(x, "missing_note"), NULL when nothing was removed) so a pipeline that renders the numbers itself can still state what left the table. On the "tinytable" route the note is set one size down; options(spicy.note_style) governs that (see table_regression()).

The Excel sheet carries the same title the console prints on its first row; the table itself starts on row 3.

Tests

When by is used, each selected variable is cross-tabulated against the grouping variable with cross_tab() and the omnibus chi-squared p-value is reported in the p column. See ⁠@param correct⁠ / simulate_p to switch on Yates' continuity correction or Monte Carlo p-values, and ⁠@param assoc_measure⁠ for the per-row dispatch table used by "auto" (2x2 -> Phi, both ordered -> Kendall's Tau-b, otherwise Cramer's V). Without by, the table reports the marginal frequency distribution of each variable with no inferential statistics.

For model-based comparisons (cluster-robust SE, weighted contrasts, fitted means) on continuous outcomes, see table_continuous_lm(). For descriptive (empirical) comparisons on continuous outcomes, see table_continuous().

Standardized mean difference

smd = TRUE adds an SMD column with the balance diagnostic of the Table 1 literature, on the variable row beside p. For a two-category variable it is the Bernoulli form,

\mathrm{SMD} = \frac{p_1 - p_2}{\sqrt{(p_1(1-p_1) + p_2(1-p_2)) / 2}}

with p the proportion of the SECOND category, signed, group 1 minus group 2 in the order the table displays them. Note the denominator: the Bernoulli variance p(1-p) at n, not var() at n-1, which would be 19% off on a small table.

"Second category" is the order the table shows, which is worth knowing for a logical variable: spicy displays TRUE then FALSE, so the sign is taken on FALSE, where tableone and cobalt coerce with factor() (FALSE, TRUE) and take it on TRUE – the same magnitude with the opposite sign. Convert to a factor with the level order you want if the direction matters.

For three or more categories it is the multivariate form of Yang and Dalton (2012, SAS Global Forum 335-2012),

\mathrm{SMD} = \sqrt{T' S^{-} T}

with T the difference of the two profiles of proportions (first category dropped) and S the mean of their multinomial covariance matrices. This is a Mahalanobis distance: it is unsigned, it is not bounded by 1, and S^{-} is a pseudo-inverse, because a declared-but-unobserved category makes S singular and solve() would abort where the pseudo-inverse returns exactly the value that category's absence implies. Which kernel a row took is published as smd_type in the "long" output, and the unsigned reading is stated in the table note whenever a variable has more than two categories. The MASS package is needed for this arm only.

Two profiles with no category in common have an infinite standardized distance. The pseudo-inverse would quietly publish a finite number there, so the cell is an en-dash and a classed warning says why. The same applies when each group is constant on a different category, where the naive route publishes 0 – "perfectly balanced" for the most imbalanced variable possible.

Conventions shared with table_continuous(): exactly two groups (three or more are refused, not averaged over pairs); complete cases on the observed groups, so a drop_na = FALSE "(Missing)" level is displayed and never enters the diagnostic; no confidence interval and no p-value, by design. Under weights the profiles are the weighted proportions, which makes this column agree with both the frequency and the survey-design readings – a profile of proportions is invariant to a global rescaling of the weights, so rescale cannot move it. (Only the continuous arm has a convention to choose there.)

The SMD cell keeps its leading zero where the association cell drops it: the APA strip belongs to a bounded measure, and this one is not bounded. The two columns therefore print 0.45 and .45 side by side, on purpose.

Two limits of the current grammar. This function has no p_value argument, so the p column cannot be switched off here as it can in table_continuous(); a complete balance table mixing continuous and categorical variables will show a categorical p beside a continuous column you removed. And inline() cannot quote this SMD cell: like p and the association measure, it lives on the variable row, which inline() cannot address on a variable that has levels. The continuous SMD cell is quotable (inline(tbl, x, "A", column = "smd")); for the categorical one, read output = "long".

Display conventions

Decimal alignment, p-value formatting, and required suggested packages per output engine are documented under ⁠@param align⁠, ⁠@param p_digits⁠, and ⁠@param output⁠ respectively.

Counts are displayed as integers: weighted counts are rounded (ties half to even, the R convention) at display time only, in cells and margins alike – the SPSS Crosstabs convention. Cells and margins are rounded independently, so small display discrepancies are possible (e.g. two cells of exactly 0.5 each display as 0 while their Total of 1.0 displays as 1). The machine outputs ("data.frame", "long") carry the exact weighted counts and full-precision percentages.

Declared missing values

Survey files imported with haven often carry declared missing values: codes such as ⁠8 = Don't know⁠ or 9 = Refused that the source file marks as missing while keeping them distinct from a plain NA. Two kinds of declaration exist: na_values / na_range metadata on haven::labelled_spss() vectors, and tagged missing values created by haven::tagged_na() (the Stata .a, .b, ... convention).

spicy honors the declaration by default (user_na = TRUE): declared missing values are excluded from every statistic exactly like NA – valid percentages, means, chi-squared tests, association measures, row-wise summaries, and group definitions – but they are not erased from display. freq() lists each observed declared value as its own row of the Missing block, with its value label; cross_tab(), table_categorical(), and table_continuous() disclose the exclusion in the table note (⁠Declared missing values removed: x (2).⁠); varlist() and code_book() count them as missing in N_valid / NAs / N_distinct while still listing the declared codes in Values.

Every function involved offers the same escape hatch: set user_na = FALSE to ignore the declaration and treat the declared codes as valid values (the behavior of spicy before 0.13.0). Tagged missing values are genuine NAs either way; for them, user_na = FALSE only collapses the per-tag breakdown back into the regular NA count.

See Also

table_continuous() for empirical comparisons on continuous outcomes; table_continuous_lm() for the model-based companion (heteroskedasticity-consistent / cluster-robust / bootstrap / jackknife SE, fitted means, weighted contrasts); cross_tab() for two-way cross-tabulations; freq() for one-way frequency tables.

Other spicy tables: table_continuous(), table_continuous_lm()

Examples

# --- Basic usage ---------------------------------------------------------

# Default: ASCII console table grouped by sex.
table_categorical(
  sochealth,
  select = c(smoking, physical_activity),
  by = sex
)

# One-way frequency-style table (no `by`).
table_categorical(
  sochealth,
  select = c(smoking, physical_activity)
)

# Pretty labels keyed by column name.
table_categorical(
  sochealth,
  select = c(smoking, physical_activity),
  by = education,
  labels = c(
    smoking           = "Current smoker",
    physical_activity = "Physical activity"
  )
)

# Survey weights with rescaling.
table_categorical(
  sochealth,
  select = c(smoking, physical_activity),
  by = education,
  weights = "weight",
  rescale = TRUE
)

# Confidence interval for the association measure.
table_categorical(
  sochealth,
  select = smoking,
  by = education,
  assoc_ci = TRUE
)

# --- Per-variable association measure ----------------------------------

# Default (`assoc_measure = "auto"`): one measure per row variable based on
# the variable type (2x2 -> Phi, both ordered factors -> Kendall's Tau-b,
# otherwise Cramer's V). When the chosen measures differ across rows, the
# column header collapses to `"Effect size"` and an APA-style `Note.` line
# documents which measure was used for which variable.
table_categorical(
  sochealth,
  select = c(smoking, education),
  by = sex
)

# Force a uniform measure across all row variables.
table_categorical(
  sochealth,
  select = c(smoking, education),
  by = sex,
  assoc_measure = "cramer_v"
)

# Per-variable override (recommended named form).
table_categorical(
  sochealth,
  select = c(smoking, education, self_rated_health),
  by = sex,
  assoc_measure = c(
    smoking           = "phi",        # binary x binary
    education         = "cramer_v",   # multi-category nominal
    self_rated_health = "tau_b"       # ordinal x binary, Tau-b
  )
)

# --- Output formats -----------------------------------------------------

# The rendered outputs below all wrap the same call:
#   table_categorical(sochealth,
#                     select = c(smoking, physical_activity),
#                     by = sex)
# only `output` changes. Assign each result to a variable -- some
# engines auto-print as a console-friendly text fallback inside
# the `?` help viewer.

# Wide data.frame (one row per modality).
table_categorical(
  sochealth,
  select = c(smoking, physical_activity),
  by = sex,
  output = "data.frame"
)

# Long data.frame (one row per (modality x group)).
table_categorical(
  sochealth,
  select = c(smoking, physical_activity),
  by = sex,
  output = "long"
)


# Rendered HTML / docx objects -- best viewed inside a
# Quarto / R Markdown document or a pkgdown article.
if (requireNamespace("tinytable", quietly = TRUE)) {
  tt <- table_categorical(
    sochealth, select = c(smoking, physical_activity), by = sex,
    output = "tinytable"
  )
}
if (requireNamespace("gt", quietly = TRUE)) {
  tbl <- table_categorical(
    sochealth, select = c(smoking, physical_activity), by = sex,
    output = "gt"
  )
}
if (requireNamespace("flextable", quietly = TRUE)) {
  ft <- table_categorical(
    sochealth, select = c(smoking, physical_activity), by = sex,
    output = "flextable"
  )
}

# Excel and Word: write to a temporary file.
if (requireNamespace("openxlsx2", quietly = TRUE)) {
  tmp <- tempfile(fileext = ".xlsx")
  table_categorical(
    sochealth, select = c(smoking, physical_activity), by = sex,
    output = "excel", excel_path = tmp
  )
  unlink(tmp)
}
if (
  requireNamespace("flextable", quietly = TRUE) &&
    requireNamespace("officer", quietly = TRUE)
) {
  tmp <- tempfile(fileext = ".docx")
  table_categorical(
    sochealth, select = c(smoking, physical_activity), by = sex,
    output = "word", word_path = tmp
  )
  unlink(tmp)
}


## Not run: 
# Clipboard: writes to the system clipboard.
table_categorical(
  sochealth, select = c(smoking, physical_activity), by = sex,
  output = "clipboard"
)

## End(Not run)


Categorical summary table from a survey design

Description

The design twin of table_categorical(): counts and estimated percentages of categorical variables computed from a survey::svydesign() or survey::as.svrepdesign() object instead of a data frame.

Every statistic is survey's. survey::svymean() estimates the percentages and their design effects, survey::svyciprop() their confidence intervals, and survey::svychisq() tests the association, Rao-Scott corrected and referred to the design degrees of freedom.

Usage

table_categorical_svy(
  design,
  select = tidyselect::everything(),
  by = NULL,
  labels = NULL,
  levels_keep = NULL,
  include_total = TRUE,
  drop_na = FALSE,
  proportion_ci = FALSE,
  ci_method = c("logit", "likelihood", "asin", "beta", "mean", "xlogit", "wilson"),
  ci_level = 0.95,
  chisq_statistic = c("F", "Chisq", "Wald", "adjWald", "saddlepoint"),
  deff = FALSE,
  df = NULL,
  p_value = NULL,
  percent_digits = 1,
  p_digits = 3,
  decimal_mark = ".",
  align = c("decimal", "center", "right"),
  output = c("default", "data.frame", "long", "tinytable", "gt", "flextable", "excel",
    "clipboard", "word"),
  indent_text = "  ",
  indent_text_excel_clipboard = strrep(" ", 6),
  excel_path = NULL,
  excel_sheet = NULL,
  clipboard_delim = "\t",
  word_path = NULL,
  user_na = TRUE,
  style = NULL
)

Arguments

design

A survey design: survey::svydesign() or survey::as.svrepdesign().

select

Columns to tabulate, as a tidyselect expression on the design's variables.

by

A single grouping column: one column block per level.

labels

Named character vector of display labels.

levels_keep

Levels to keep, as a character vector (all variables) or a named list (per variable).

include_total

Add a Total column block with the whole design's percentages (default TRUE, only with by).

drop_na

Drop missing values (default FALSE: they show as a (Missing) level). Shown or dropped, they never enter the test: the p-value is computed on the complete cases either way, which is the convention table_categorical() applies.

proportion_ci

Add the confidence interval of each percentage.

ci_method

Interval method passed to survey::svyciprop(): "logit" (default), "likelihood", "asin", "beta", "mean", "xlogit" or "wilson".

ci_level

Coverage of the interval.

chisq_statistic

Statistic for survey::svychisq(): "F" (default), "Chisq", "Wald", "adjWald" or "saddlepoint". "saddlepoint" is refused on a replicate-weights design: survey computes its p-value without the denominator degrees of freedom there, so it comes out too small.

deff

Show the design effect of each percentage: FALSE (default), TRUE or "replace".

df

Degrees of freedom for the intervals. NULL (default) uses survey::degf() on each domain. It does not reach the test: survey::svychisq() has no df argument, so the Rao-Scott reference distribution keeps the design's own degrees of freedom and the note says so.

p_value

Show the p-value column (defaults to TRUE with by). A variable whose complete cases include negatively weighted rows is not tested: the Rao-Scott correction is a function of the design variance, which is not defined when the weights change sign. Its percentages are still reported, the note says which tests were withheld, and the call warns (spicy_negative_weights_no_test).

percent_digits, p_digits, decimal_mark

Number formatting.

align

Numeric-cell alignment: "decimal", "center" or "right".

output

One of "default", "data.frame", "long", or a rendering engine: "tinytable", "gt", "flextable", "excel", "clipboard", "word". "data.frame" and "long" are synonyms here and return the same object: the wide compute frame, one row per level with a pair of columns per group. Note the difference from table_categorical(), where the two tokens return genuinely different shapes – this design carries a single compute frame, and both names reach it.

indent_text, indent_text_excel_clipboard

Level-row indentation, for the console and for the plain-text engines.

excel_path, excel_sheet, clipboard_delim, word_path

Output destinations, as in table_categorical().

user_na

Honour declared missing values (see ?freq).

style

A journal style; see spicy_style().

Value

A spicy_categorical_svy_table: the wide compute frame, with the display frame and the typed view attached. output = "data.frame" / "long" returns the compute frame unclassed – the two tokens are synonyms and return identical objects.

What the columns are

n is the observed count – the number of rows in the sample, not an estimated population size. ⁠%⁠ is the estimated percentage within its column: without by it is the distribution of the variable in the population, with by the distribution inside that domain. The table note gives the sample size and the estimated population together, because neither alone tells the reader what they are looking at.

proportion_ci = TRUE adds the interval of each percentage. ci_method chooses among the seven survey::svyciprop() offers; the default "logit" is bounded inside 0 to 100, which the Wald interval ("mean") is not. The percentage itself always comes from survey::svymean(), so it does not move when ci_method does.

The test

svychisq() with chisq_statistic = "F" (the default) is the Pearson chi-square with the Rao-Scott second-order correction, referred to F(ndf, survey::degf(design)). It is survey's own default and the one Stata's svy: tabulate reports.

It runs on the complete cases of the two variables, and on their observed levels: a (Missing) row and a declared-but-unobserved level are descriptive, and neither belongs to the null hypothesis. The p-value is therefore the same whether drop_na shows those rows or removes them, and the intervals beside it describe the same domain – the two families test the same table.

"Chisq" shows the p-value only: survey adjusts the statistic in the "F" branch and only the p-value in the "Chisq" one, so the statistic there is not the one the p-value came from. "Wald", "adjWald" and "saddlepoint" are available; "lincom" and "wls-score" are refused, the first because its integration is documented as failing in the far tail (?pchisqsum), the second because it has no reporting convention here.

Stability

This function is stabilising in the sense ?spicy defines: the names of its design-specific arguments may still be tightened before 1.0 – with a NEWS.md entry – but the behaviour does not change silently. It was experimental through 0.13; the shape of the table has held across the cycle, so it now moves on the parent family's clock rather than its own. The numbers themselves are survey's and do not move with it.

What is absent, and why

weights and rescale (the weighting is the design). correct (Yates), simulate_p and simulate_B, which have no meaning once the reference distribution is Rao-Scott's. And the association measures: Cramer's V, phi, tau-b/c, gamma, Somers' D and lambda have no established design-based variance, and the intervals table_categorical() gives them assume simple random sampling. The design-based measure of association here is the Rao-Scott test in the p column; for an effect size, model it with table_regression(survey::svyglm(...)).

See Also

table_categorical() for the data-frame sibling, table_continuous_svy() for continuous variables.

Examples


data(api, package = "survey")
dclus1 <- survey::svydesign(
  id = ~dnum, weights = ~pw, data = apiclus1, fpc = ~fpc
)
table_categorical_svy(dclus1, select = c(stype, awards))
table_categorical_svy(dclus1, select = stype, by = sch.wide)
table_categorical_svy(
  dclus1,
  select = stype,
  proportion_ci = TRUE,
  deff = TRUE
)


Continuous summary table

Description

Computes descriptive statistics (mean, SD, min, max, confidence interval of the mean, n) for one or many continuous variables selected with tidyselect syntax.

With by, produces grouped summaries and reports a group-comparison p-value by default (Welch test; change via test). Additional inferential output is opt-in: test statistics (statistic) and effect sizes (effect_size / effect_size_ci). Set p_value = FALSE to suppress the p-value column. Without by, produces one-way descriptive summaries.

Multiple output formats are available via output: a printed ASCII table ("default"), a plain data.frame ("data.frame" or "long" – synonyms for the underlying long-format data, see Details), or publication-ready tables ("tinytable", "gt", "flextable", "excel", "clipboard", "word").

This is the descriptive companion to table_continuous_lm(). The two functions share their layout, alignment, and reporting precision so descriptive and model-based analyses of the same data look uniform side by side – with one exception, documented under align: only table_continuous() carries align into the excel output. Use table_continuous_lm() when you need robust SE, weighted contrasts, fitted means, or covariate adjustment.

Usage

table_continuous(
  data,
  select = tidyselect::everything(),
  by = NULL,
  exclude = NULL,
  regex = FALSE,
  drop_na = TRUE,
  weights = NULL,
  rescale = FALSE,
  test = c("welch", "student", "nonparametric"),
  p_value = NULL,
  statistic = FALSE,
  show_n = TRUE,
  show_columns = NULL,
  effect_size = c("none", "auto", "hedges_g", "eta_sq", "r_rb", "epsilon_sq"),
  effect_size_ci = FALSE,
  smd = FALSE,
  ci = TRUE,
  labels = NULL,
  ci_level = 0.95,
  digits = 2,
  effect_size_digits = 2,
  p_digits = 3,
  decimal_mark = ".",
  align = c("decimal", "center", "right"),
  output = c("default", "data.frame", "long", "tinytable", "gt", "flextable", "excel",
    "clipboard", "word"),
  excel_path = NULL,
  excel_sheet = NULL,
  clipboard_delim = "\t",
  word_path = NULL,
  verbose = FALSE,
  user_na = TRUE,
  style = NULL
)

Arguments

data

A data.frame.

select

Columns to include. If regex = FALSE, use tidyselect syntax or a character vector of column names (default: tidyselect::everything()). If regex = TRUE, provide a regular expression pattern (character string).

by

Optional grouping column. Accepts an unquoted column name or a single character column name. Coerced to factor for grouping; non-numeric grouping columns (factor, character, logical) are supported as-is. Factor levels keep their declared order; any other by (character, numeric, haven labelled) forms groups in order of first appearance in the data – the same convention as table_categorical(). For a haven labelled by, the group headers are the raw codes (value labels are not used for grouping headers – the family convention shared with table_categorical() and table_continuous_lm()); declared missing values follow user_na as usual.

exclude

Columns to exclude. Supports tidyselect syntax and character vectors of column names.

regex

Logical. If FALSE (the default), uses tidyselect helpers. If TRUE, the select argument is treated as a regular expression.

drop_na

Logical. Controls how missing values in the by column are handled – the same argument as table_categorical(), with one structural difference: a continuous summary has no "(Missing)" row for the summarized variable itself (a mean cannot include NA), so NAs in each summarized variable are always excluded from that variable's statistics and the exclusion is disclosed in a table note ("Missing values removed: ...") rather than silent. If TRUE (the default, preserving this function's historical behavior; table_categorical() defaults to FALSE), rows with NA in by are removed from the grouped summaries, with a warning and a dedicated note line ("Rows with missing ... removed"). If FALSE, rows with NA in by form a dedicated "(Missing)" group – the field convention for descriptive tables (gtsummary's "Unknown" row; see the Epidemiologist R Handbook, Descriptive tables) – while the group-comparison test and effect size are still computed on the observed groups only (show the missing, test the observed, matching table_categorical()). Ignored (with a warning) when by is not used.

weights

Optional case weights: an unquoted column name, a character column name, or a numeric vector of length nrow(data). Weights must be non-negative and finite; rows with NA or zero weight leave every statistic (including Min / Max) and NA weights are disclosed in the table note. The weighted formulas follow the frequency-expansion convention – see the Weights section. The weighted table names its weights in the note ("Statistics weighted by ..."). Group tests and effect sizes are not computed under weights: use table_continuous_lm() for weighted comparisons.

rescale

Logical. If TRUE, weights are first normalised so that they sum to the number of observations used for each variable – the same rescale grammar as table_categorical(), read from options(spicy.rescale) when not supplied. This is the sampling-weights reading: results become invariant to the scale of the weights, and the weighted SD then equals Stata's ⁠[aweight]⁠ / survey::svyvar() value exactly. The default FALSE uses the weights as given (frequency reading).

test

Character. Statistical test to use when comparing groups. One of "welch" (default), "student", or "nonparametric".

  • "welch": Welch t-test (2 groups) or Welch one-way ANOVA (3+ groups). Does not assume equal variances.

  • "student": Student t-test (2 groups) or classic one-way ANOVA (3+ groups). Assumes equal variances.

  • "nonparametric": Wilcoxon rank-sum / Mann–Whitney U (2 groups) or Kruskal–Wallis H (3+ groups).

Used whenever by is supplied (since p_value defaults to TRUE in that case) or when statistic = TRUE / effect_size = TRUE. Ignored when by is not used, or when all three display toggles are turned off.

p_value

Logical or NULL. If TRUE and by is used, adds a p-value column from the test specified by test. When NULL (the default), the p-value is shown automatically whenever by is supplied, and hidden otherwise. Pass p_value = FALSE to suppress the column explicitly. Ignored when by is not used.

statistic

Logical. If TRUE and by is used, the test statistic is shown in an additional column (e.g., t(df) = ..., F(df1, df2) = ..., W = ..., or H(df) = ...). Both p_value and statistic are independent; either or both can be enabled. Defaults to FALSE. Ignored when by is not used.

show_n

Logical. If TRUE, includes an unweighted n column in the printed ASCII table and in every rendered output (tinytable, gt, flextable, word, excel, clipboard). Set to FALSE to drop the n column structurally from those outputs (no empty placeholder, no spanner). The n column is always present in the raw output = "data.frame" / "long" for downstream programmatic access. Defaults to TRUE. Ignored (with a warning) when show_columns is supplied.

show_columns

Statistics to display, as a character vector of tokens or a named list of such vectors (one per variable). NULL (the default) keeps the historical display: mean, SD, min, max, the mean CI (see ci) and n (see show_n). See the "Choosing the statistics" section for the token vocabulary, the per-variable form, and the test that follows a median.

effect_size

Effect-size measure to include in the rendered outputs. One of:

  • "none" (default): no effect-size column.

  • "auto": auto-select the canonical measure for the chosen test and group count – Hedges' g (parametric, 2 groups), eta-squared (parametric, 3+ groups), rank-biserial r (nonparametric, 2 groups), epsilon-squared (nonparametric, 3+ groups).

  • "hedges_g": Hedges' g (bias-corrected standardised mean difference, 2 groups, parametric). CI via the Hedges & Olkin normal approximation.

  • "eta_sq": Eta-squared (\eta^2, parametric ANOVA-style SS_between / SS_total). CI via inversion of the noncentral F distribution.

  • "r_rb": Rank-biserial r from the Wilcoxon / Mann-Whitney statistic (2 groups, nonparametric). CI via Fisher z-transform.

  • "epsilon_sq": Epsilon-squared (\varepsilon^2) from the Kruskal-Wallis statistic (3+ groups, nonparametric). CI via percentile bootstrap (2 000 replicates).

For backward compatibility, effect_size = TRUE is silently coerced to "auto" and effect_size = FALSE to "none". Explicit choices are validated against the active test and the number of groups; an incompatible request (e.g. "eta_sq" with two groups, or "hedges_g" with test = "nonparametric") triggers an actionable error. Ignored when by is not used.

effect_size_ci

Logical. If TRUE, appends the confidence interval of the effect size in brackets (e.g., g = 0.45 [0.22, 0.68]). Implies a non-"none" effect size: if left at the default effect_size = "none", the function warns and promotes effect_size to "auto" so the requested CI can be shown. Defaults to FALSE.

smd

Logical. If TRUE, adds an SMD column holding the standardized mean difference between the two groups of by, the balance diagnostic of the Table 1 literature. Requires exactly two groups; signed, group 1 minus group 2 in the order the table displays them; no confidence interval and no p-value, by design. It is independent of p_value and of effect_size: turning it on turns nothing else off. Rounded with effect_size_digits. See the "Standardized mean difference" section below. Defaults to FALSE.

ci

Logical. If TRUE, includes the mean confidence interval columns (⁠<level>% CI LL⁠ / ⁠<level>% CI UL⁠) and their spanner in the printed ASCII table and in every rendered output (tinytable, gt, flextable, word, excel, clipboard). Set to FALSE to drop both columns and the CI spanner structurally from those outputs (no empty placeholders, no border lines under an empty header). The CI bounds are always present as ci_lower / ci_upper in the raw output = "data.frame" / "long" for downstream programmatic access. Defaults to TRUE. The CI level is taken from ci_level. Ignored (with a warning) when show_columns is supplied.

labels

An optional named character vector of variable labels. Names must match column names in data. When NULL (the default), labels are auto-detected from variable attributes (e.g., haven labels); if none are found, the column name is used.

ci_level

Confidence level for the mean confidence interval (default: 0.95). Must be between 0 and 1 exclusive.

digits

Number of decimal places for descriptive values and test statistics (default: 2).

effect_size_digits

Number of decimal places for effect-size values in formatted displays (default: 2).

p_digits

Integer >= 1. Number of decimal places used to render p-values in the p column (default: 3, the APA Publication Manual standard). Both the displayed precision and the small-p threshold derive from this argument: p_digits = 3 prints .045 and ⁠<.001⁠; p_digits = 4 prints .0451 and ⁠<.0001⁠; p_digits = 2 prints .05 and ⁠<.01⁠. Useful for genomics / GWAS contexts with very small p-values, or for journals using a coarser convention. Leading zeros are always stripped, following APA convention.

decimal_mark

Character used as decimal separator. Either "." (default) or ",".

align

Horizontal alignment of numeric columns in the printed ASCII table and in the tinytable, gt, flextable, word, excel, and clipboard outputs. The first column (Variable) and Group (when present) are always left-aligned. One of:

  • "decimal" (default): align numeric columns on the decimal mark, the standard scientific-publication convention used by SPSS, SAS, and LaTeX siunitx. Numeric cells are pre-padded with figure-spaces (U+2007, digit-width) so every string in a column has the same width with the decimal mark at the same internal position; centring those uniform-width strings then stacks the decimal points vertically. The same pad-then-centre strategy is applied on every rendering engine (gt, tinytable, flextable, word, ASCII print) for a homogeneous rendering, matching table_regression() and table_continuous_lm(). The clipboard output is delimited text meant to be parsed rather than read at a fixed width, so its cells travel unpadded (a padded number pastes as text next to an unpadded number).

  • "center": center-align all numeric columns.

  • "right": right-align all numeric columns.

"center" and "right" reach the excel output too. "decimal" does not: Excel cells are written unpadded, because cell-string padding does not align decimals under a proportional font, so the workbook keeps the engine's own convention instead – counts and the p-value right-aligned, the other numeric columns centred. Same default and same three values as table_continuous_lm(), whose excel output still uses that convention at every align.

output

Output format. One of:

  • "default": an ASCII table object, printed when the call is bare.

  • "data.frame" / "long": a plain data.frame with one row per ⁠(variable x group)⁠ (or one row per variable when by is not used). The two names are synonyms; pick whichever reads better in your pipeline ("long" matches table_continuous_lm()'s naming).

  • "tinytable" (requires tinytable)

  • "gt" (requires gt)

  • "flextable" (requires flextable)

  • "excel" (requires openxlsx2)

  • "clipboard" (requires clipr)

  • "word" (requires flextable and officer)

excel_path

File path for output = "excel".

excel_sheet

Sheet name for output = "excel". NULL (the default) uses "Descriptives".

clipboard_delim

Delimiter for output = "clipboard" (default: "\t"). A cell holding the delimiter itself, a double quote or a line break is quoted RFC 4180-style, so the grid survives whatever delimiter you choose.

word_path

File path for output = "word".

verbose

Logical. If TRUE, prints messages about excluded non-numeric columns (default: FALSE).

user_na

Logical. If TRUE (the default), declared missing values never reach the numeric summaries: they are excluded like NA and disclosed in the table note (⁠Declared missing values removed: ...⁠); declared-missing by values form no group. If FALSE, the declared codes are summarized as ordinary numbers. See the "Declared missing values" section of freq().

style

A journal style: a theme name ("jama", "nejm", "lancet", "annals", "apa", "aer"), a spicy_style() object, or NULL (the default). A style only changes DEFAULTS – any argument you pass explicitly wins over it. Set options(spicy.style = ) for document-wide scope. A theme covers numeric formatting conformity only, not full editorial conformity; ?spicy_style lists the exact rules each one encodes and the official document they come from. An unknown name is an error.

Value

Depends on output:

The missing-value disclosure (values excluded from the summaries, and rows removed for a missing by value under drop_na = TRUE) travels with the table on every route, not just the console: "default" prints it under the ASCII table, "tinytable" / "gt" / "flextable" / "word" carry it as a table note, "excel" writes it below the body, and "data.frame" / "long" keep the sentence verbatim in the missing_note attribute (attr(x, "missing_note"), NULL when nothing was removed) so a pipeline that renders the numbers itself can still state what left the table. On the "tinytable" route the note is set one size down; options(spicy.note_style) governs that (see table_regression()).

The Excel sheet carries the same title the console prints on its first row; the table itself starts on row 3.

Choosing the statistics

show_columns selects which statistics the table displays. The tokens, and the column each one produces:

Token Column Statistic
"m" M mean
"sd" SD standard deviation
"med" Med median (stats::median())
"iqr" IQR interquartile width, Q3 - Q1
"med_iqr" Med [Q1, Q3] median and the interquartile interval, in one compact column
"q1" / "q3" Q1 / Q3 first / third quartile
"min" / "max" Min / Max extremes
"ci" ⁠<level>% CI LL⁠ / UL t confidence interval of the mean
"med_ci" ⁠Med <level>% CI LL⁠ / UL exact confidence interval of the median
"n" n valid observations
"weighted_n" ⁠Weighted n⁠ sum of weights (requires weights)

Quartiles use stats::quantile()'s default type 7. "iqr" is the width (one number, the rank mirror of SD); "med_iqr" shows the interval with its bounds. Columns appear in the canonical order of the table above, whatever order they were written in.

Weights

With weights, every displayed statistic uses the frequency-expansion convention: for integer weights each statistic equals its unweighted version computed on the data with every row repeated w times (rep(x, w)), exactly; with all weights equal to 1 every statistic equals its unweighted sibling. The formulas, with W = \sum w_i:

These are the conventions of Hmisc::wtd.mean() / wtd.var() / wtd.quantile() (defaults), matrixStats::weightedSd(), and DescTools::Quantile(), and – for integer weights – of Stata's ⁠[fweight]⁠ and SPSS's ⁠WEIGHT BY⁠. With rescale = TRUE the weights are normalised to sum to the number of observations first, which makes every result invariant to the scale of the weights and makes the SD equal Stata's ⁠[aweight]⁠ / survey::svyvar() value – the reading appropriate for sampling weights. Weighted-quantile conventions genuinely differ across software (Stata interpolates nowhere, SAS refuses analytic-weighted quantiles, the survey package offers twelve rules); spicy states its rule here rather than leaving it implicit.

Two deliberate refusals: the "med_ci" token (an order-statistic interval with no weighted version) and, under by, the group tests and effect sizes – a t-test printed next to weighted descriptives would silently be unweighted. Set p_value = FALSE for weighted descriptives by group, or use table_continuous_lm() with weights for weighted comparisons. Note that table_continuous_lm()'s residual SD answers a different question (model-based, precision-weight convention) and is not expected to match the descriptive SD here.

A named list applies a different selection to each variable, with .default covering the variables it does not name – the case of a table where a skewed variable must be reported as a median while the others keep the mean:

show_columns = list(
  mvpa    = c("med_iqr", "n"),
  sitting = c("med_iqr", "n"),
  .default = c("m", "sd", "n")
)

The table's columns are the union of the requested tokens; a cell of a column the variable did not ask for is left blank (structurally empty, not an en dash, which is reserved for an undefined statistic).

The table tests what it shows. A variable displaying a median without a mean takes the rank-based test – Wilcoxon rank-sum for two groups, Kruskal-Wallis beyond – and the rank effect size (rank-biserial r, \varepsilon^2) when effect_size is "auto". The switch is per variable, so a mixed table carries a rank test on its median rows and Welch on its mean rows, and the table note names which test each variable carries. An explicit test is sovereign: it applies to every variable, with a warning naming the ones displayed as medians.

"med_ci" is the exact order-statistic (sign-test) confidence interval: the tightest interval [x_{(k)}, x_{(n-k+1)}] whose binomial coverage still reaches ci_level. It is distribution-free and deterministic – no bootstrap, no seed – and its coverage is at least nominal, the same convention as SAS ⁠PROC UNIVARIATE⁠ (CIPCTLDF) and DescTools::MedianCI(method = "exact"). Below about six observations no interval reaches the requested level; the cells then show an en dash rather than a false interval.

"ci" is the confidence interval of the mean: requested without "m" it is dropped with a warning pointing at "med_ci", and "med_ci" without a displayed median is dropped likewise. When show_columns is supplied it decides the n and CI columns on its own, and a contradictory show_n / ci is reported.

Tests

The omnibus test is computed only when by is supplied and at least two groups remain after dropping NAs, with every group contributing at least two observations. Choice of test family is driven by test (see the ⁠@param⁠ entry for the full dispatch and the underlying ⁠stats::⁠ functions called).

For model-based contrasts (heteroskedasticity-consistent SE, cluster-robust SE, weighted contrasts, fitted means, covariate adjustment), use table_continuous_lm().

Effect sizes

See ⁠@param effect_size⁠ for the dispatch table (canonical measure for each (test, n_groups) combination) and the validation rules applied to explicit requests.

Confidence intervals (enabled with effect_size_ci = TRUE) use noncentral F inversion for \eta^2, the Hedges-Olkin normal approximation for g, the Fisher z-transform for r, and percentile bootstrap (2,000 replicates) for \varepsilon^2. The bootstrap bounds depend on the random number generator state: call set.seed() before the table for reproducible \varepsilon^2 intervals (the other three CIs are closed-form and deterministic).

For Cohen's d, Hays' \omega^2, and Cohen's f^2 (derived from a fitted, possibly weighted lm()), use the model-based companion table_continuous_lm().

Standardized mean difference

smd = TRUE adds an SMD column with the balance diagnostic of the Table 1 literature, in Austin's form (Austin 2009, Stat Med 28:3083-3107; Austin 2011, Multivar Behav Res 46:399-424):

\mathrm{SMD} = \frac{\bar{x}_1 - \bar{x}_2}{\sqrt{(s_1^2 + s_2^2) / 2}}

The denominator is the root mean of the two group variances, each at n - 1, not the degrees-of-freedom pooled SD. At equal group sizes those two denominators are the same, so the SMD is exactly Cohen's d; at unequal sizes they part company (on a 4-versus-3 split the SMD is -0.51 and d is -0.54).

effect_size = "hedges_g", two columns to the left, is a third number: g applies the small-sample correction J on top of d, so it never equals the SMD. At equal group sizes the ratio g / \mathrm{SMD} is exactly J – 0.80 at n = 3 per group, 0.96 at n = 10 – approaching 1 only as the sample grows. Read each for what it is; do not recompute one from the other. (The divergence is nameable upstream: cobalt::col_w_smd(s.d.denom = "pooled") reproduces this column, s.d.denom = "hedges" reproduces hedges_g.)

Conventions, all deliberate:

Under weights, the means and variances are the weighted ones the M and SD columns already display – the frequency convention of the Weights section, from the same producer, so the column cannot contradict its neighbours. One consequence follows and is intended: a frequency weight is a number of copies, so the weighted SMD is not invariant to the scale of the weights (multiplying every weight by ten moves it, as it moves the SD column). rescale = TRUE normalises the weights to sum to n, restores scale invariance, and is the form to use for sampling weights until the dedicated survey-design functions land.

A cell is an en-dash when the diagnostic applies but cannot be estimated: both groups constant at different values (an infinite standardized distance, disclosed by a warning), or a group with too little data to have a variance (silent – the SD cell beside it already says so). Two groups constant at the same value are perfectly balanced and print 0.00.

Display conventions

Decimal alignment, p-value formatting, and required suggested packages per output engine are documented under ⁠@param align⁠, ⁠@param p_digits⁠, and ⁠@param output⁠ respectively.

Non-numeric columns are silently dropped (set verbose = TRUE to see which columns were excluded). When a constant column is passed, its statistics are reported exactly: SD is 0.00 and the CI degenerates to ⁠[m, m]⁠. An en-dash cell appears only when a statistic is undefined (fewer than two valid observations).

Declared missing values

Survey files imported with haven often carry declared missing values: codes such as ⁠8 = Don't know⁠ or 9 = Refused that the source file marks as missing while keeping them distinct from a plain NA. Two kinds of declaration exist: na_values / na_range metadata on haven::labelled_spss() vectors, and tagged missing values created by haven::tagged_na() (the Stata .a, .b, ... convention).

spicy honors the declaration by default (user_na = TRUE): declared missing values are excluded from every statistic exactly like NA – valid percentages, means, chi-squared tests, association measures, row-wise summaries, and group definitions – but they are not erased from display. freq() lists each observed declared value as its own row of the Missing block, with its value label; cross_tab(), table_categorical(), and table_continuous() disclose the exclusion in the table note (⁠Declared missing values removed: x (2).⁠); varlist() and code_book() count them as missing in N_valid / NAs / N_distinct while still listing the declared codes in Values.

Every function involved offers the same escape hatch: set user_na = FALSE to ignore the declaration and treat the declared codes as valid values (the behavior of spicy before 0.13.0). Tagged missing values are genuine NAs either way; for them, user_na = FALSE only collapses the per-tag breakdown back into the regular NA count.

See Also

table_outcome() for the transposed shape – ONE continuous outcome across the levels of SEVERAL groupings, one block of rows per grouping. Several outcomes across one grouping is this function; one outcome across one or more groupings is that one; table_continuous_lm() for the model-based companion (heteroskedasticity-consistent SE, cluster-robust SE, weighted contrasts, fitted means); table_categorical() for categorical variables; freq() for one-way frequency tables; cross_tab() for two-way cross-tabulations.

Other spicy tables: table_categorical(), table_continuous_lm()

Examples

# --- Basic usage ---------------------------------------------------------

# Default: ASCII console table.
table_continuous(
  sochealth,
  select = c(bmi, wellbeing_score)
)

# Grouped by education (Welch p-value added by default).
table_continuous(
  sochealth,
  select = c(bmi, wellbeing_score),
  by = education
)

# Test statistic alongside the p-value.
table_continuous(
  sochealth,
  select = c(bmi, wellbeing_score),
  by = education,
  statistic = TRUE
)

# --- Choosing the statistics --------------------------------------------

# Median and interquartile range instead of mean and SD.
table_continuous(
  sochealth,
  select = c(bmi, wellbeing_score),
  show_columns = c("med_iqr", "n")
)

# Median with its exact (order-statistic) confidence interval.
table_continuous(
  sochealth,
  select = bmi,
  show_columns = c("med", "iqr", "med_ci", "n")
)

# One selection per variable. A skewed variable that a scoring
# protocol requires in median and IQR (the IPAQ case) sits next to
# variables kept in mean and SD; each row is tested the way it is
# displayed, and the note says so.
table_continuous(
  sochealth,
  select = c(bmi, life_sat_health, wellbeing_score),
  by = sex,
  show_columns = list(
    life_sat_health = c("med_iqr", "n"),
    .default = c("m", "sd", "n")
  )
)

# --- Effect sizes -------------------------------------------------------

# Auto-selected effect size with confidence interval (Hedges' g for
# binary `by`, eta-squared for k > 2).
table_continuous(
  sochealth,
  select = wellbeing_score,
  by = sex,
  effect_size = "auto",
  effect_size_ci = TRUE
)

# Explicit effect-size measure.
table_continuous(
  sochealth,
  select = wellbeing_score,
  by = education,
  effect_size = "eta_sq",
  effect_size_ci = TRUE,
  effect_size_digits = 3
)

# --- Selection helpers --------------------------------------------------

# Regex selection.
table_continuous(
  sochealth,
  select = "^life_sat",
  regex = TRUE
)

# Pretty labels keyed by column name.
table_continuous(
  sochealth,
  select = c(bmi, life_sat_health),
  labels = c(
    bmi = "Body mass index",
    life_sat_health = "Satisfaction with health"
  )
)

# --- Output formats -----------------------------------------------------

# The rendered outputs below all wrap the same call:
#   table_continuous(sochealth,
#                    select = c(bmi, wellbeing_score),
#                    by = sex)
# only `output` changes. Assign each result to a variable -- some
# engines auto-print as a console-friendly text fallback inside
# the `?` help viewer.

# Wide / long data.frame (synonyms): one row per (variable x group).
table_continuous(
  sochealth,
  select = c(bmi, wellbeing_score),
  by = sex,
  output = "data.frame"
)


# Rendered HTML / docx objects -- best viewed inside a
# Quarto / R Markdown document or a pkgdown article.
if (requireNamespace("tinytable", quietly = TRUE)) {
  tt <- table_continuous(
    sochealth, select = c(bmi, wellbeing_score), by = sex,
    output = "tinytable"
  )
}
if (requireNamespace("gt", quietly = TRUE)) {
  tbl <- table_continuous(
    sochealth, select = c(bmi, wellbeing_score), by = sex,
    output = "gt"
  )
}
if (requireNamespace("flextable", quietly = TRUE)) {
  ft <- table_continuous(
    sochealth, select = c(bmi, wellbeing_score), by = sex,
    output = "flextable"
  )
}

# Excel and Word: write to a temporary file.
if (requireNamespace("openxlsx2", quietly = TRUE)) {
  tmp <- tempfile(fileext = ".xlsx")
  table_continuous(
    sochealth, select = c(bmi, wellbeing_score), by = sex,
    output = "excel", excel_path = tmp
  )
  unlink(tmp)
}
if (
  requireNamespace("flextable", quietly = TRUE) &&
    requireNamespace("officer", quietly = TRUE)
) {
  tmp <- tempfile(fileext = ".docx")
  table_continuous(
    sochealth, select = c(bmi, wellbeing_score), by = sex,
    output = "word", word_path = tmp
  )
  unlink(tmp)
}


## Not run: 
# Clipboard: writes to the system clipboard.
table_continuous(
  sochealth, select = c(bmi, wellbeing_score), by = sex,
  output = "clipboard"
)

## End(Not run)


Continuous-outcome linear-model table

Description

Builds publication-ready summary tables from a series of linear models for one or many continuous outcomes selected with tidyselect syntax.

A single focal predictor is supplied with by; each selected numeric outcome is fit as lm(outcome ~ by, ...), optionally extended with additive covariates via covariates and case weights via weights. Categorical by produces model-based estimated marginal means by level (covariate-adjusted via adjustment when covariates are present), plus an optional single difference for dichotomous predictors. Numeric by produces the slope and its confidence interval.

Inference adapts via vcov: classical OLS, "HC0"-"HC5" (heteroscedasticity-consistent), "CR0"-"CR3" (cluster-robust, requires cluster), or "bootstrap" / "jackknife" resampling. Effect sizes (Cohen's "d", Hedges' "g", Hays' "omega2", Cohen's "f2") are reported with optional noncentral t / F confidence intervals via effect_size_ci, and adapt under covariate adjustment (see effect_size).

Multiple output formats are available via output: a printed ASCII table ("default"), a plain wide data.frame ("data.frame"), a raw long data.frame ("long"), or rendered outputs ("tinytable", "gt", "flextable", "excel", "clipboard", "word").

Usage

table_continuous_lm(
  data,
  select = tidyselect::everything(),
  by,
  covariates = NULL,
  adjustment = c("proportional", "balanced"),
  exclude = NULL,
  regex = FALSE,
  weights = NULL,
  vcov = c("classical", "HC0", "HC1", "HC2", "HC3", "HC4", "HC4m", "HC5", "CR0", "CR1",
    "CR2", "CR3", "bootstrap", "jackknife"),
  cluster = NULL,
  boot_n = 1000,
  contrast = c("auto", "none"),
  statistic = FALSE,
  p_value = TRUE,
  show_n = TRUE,
  show_weighted_n = FALSE,
  effect_size = c("none", "f2", "d", "g", "omega2"),
  effect_size_ci = FALSE,
  r2 = c("r2", "adj_r2", "none"),
  ci = TRUE,
  labels = NULL,
  ci_level = 0.95,
  digits = 2,
  fit_digits = 2,
  effect_size_digits = 2,
  p_digits = 3,
  decimal_mark = ".",
  align = c("decimal", "center", "right"),
  output = c("default", "data.frame", "long", "tinytable", "gt", "flextable", "excel",
    "clipboard", "word"),
  excel_path = NULL,
  excel_sheet = NULL,
  clipboard_delim = "\t",
  word_path = NULL,
  verbose = FALSE,
  user_na = TRUE,
  style = NULL
)

Arguments

data

A data.frame.

select

Outcome columns to include. If regex = FALSE, use tidyselect syntax or a character vector of column names (default: tidyselect::everything()). If regex = TRUE, provide a regular expression pattern (character string).

by

A single predictor column. Accepts an unquoted column name or a single character column name. The predictor can be:

  • numeric (continuous): treated as a continuous regressor. The table reports the slope of by and its CI from lm(y ~ by, ...).

  • factor or ordered factor: treated as categorical. Level order is preserved as declared; the first level is the reference for the displayed contrast (R's default treatment-contrast convention). An ordered factor is refit as an unordered factor with the same level order, so the model uses treatment contrasts – not the polynomial contrasts (contr.poly) lm() would apply by default. The ordering only determines the display order and the reference level; per-level means, the displayed difference, and its CI are identical to those of an unordered factor with the same levels.

  • character: coerced to factor with factor(by), which orders the levels alphabetically. To control the reference level, supply by as an explicit factor with the desired level ordering (e.g. via forcats::fct_relevel() or factor(..., levels = ...)).

  • logical: coerced to factor with levels "FALSE", "TRUE" (in that order, since FALSE < TRUE). The reference level is "FALSE", so a binary contrast displays as Delta (TRUE - FALSE).

  • haven labelled with value labels (an SPSS / Stata import): treated as categorical over the raw codes, in ascending code order (the reference level is the lowest code; the model needs a fixed level order). table_continuous() and table_categorical() form the same groups from such a column but display them in order of first appearance in the data, so the group ORDER can differ between the sibling tables. A labelled vector without value labels is treated as a continuous regressor. Declared missing values follow user_na as usual.

Rows with NA in by are excluded from the analytic sample for each outcome (NAs in y and weights are also excluded; see Details).

covariates

Optional additive covariates to adjust each per-outcome linear model for. Accepts a tidyselect expression (e.g. covariates = c(age, sex), covariates = tidyselect::all_of(cov_vec), covariates = tidyselect::starts_with("control_")) or a literal character vector of column names. Each covariate must be numeric, integer, logical, factor, or character; covariates that also appear in select are silently auto-excluded from the outcome list (a variable cannot be both outcome and adjustment), and a covariate that equals by raises an error (a variable cannot be both predictor and adjustment).

When non-empty, each model is fitted as lm(y ~ by + cov1 + cov2 + ...) and the reported estimate / SE / p-value / CI on by are covariate-adjusted via the focal coefficient. For categorical by, the displayed emmean is the covariate-adjusted estimated marginal mean – see adjustment for the choice of estimand (G-computation by default vs. equal-weight averaging). The omnibus test of by is the Wald F restricted to the focal coefficients (computed via sandwich / clubSandwich for HC* / CR* mode), so adding covariates does not contaminate the omnibus statistic with covariate contributions. Effect sizes adapt automatically – see effect_size.

v1 supports additive covariates only. Formula syntax with interactions or transforms (covariates = ~ age * sex, covariates = ~ I(age^2)) is reserved for a future release; passing a formula raises a spicy_unsupported error with a migration hint.

Rows with NA in any covariate are dropped from the analytic sample for each outcome (complete-cases per outcome, matching the existing by / weights NA handling).

adjustment

How the covariate-adjusted estimated marginal means (the emmean / emmean_se / ⁠emmean_ci_*⁠ columns) are computed when covariates is non-empty. One of:

  • "proportional" (the default; matches Stata margins and marginaleffects::avg_predictions()): G-computation on the observed sample. For each focal level of by, the model predicts at every observation with by set to that level (covariates kept at their observed values), and the predictions are averaged. Population-weighted by construction – the empirical joint distribution of covariates is the reference. Under weights, the averaging uses the case weights (the Stata margins convention after a weighted regression; equivalent to marginaleffects::avg_predictions(wts = )), so the reference distribution is the weighted one. Best when the goal is "what is the predicted mean in this population if everyone had by = lvl".

  • "balanced" (matches emmeans::emmeans() default and the SPSS UNIANOVA EMMEANS / SAS LSMEANS conventions): synthetic grid of factor-covariate level combinations x numeric covariates fixed at their sample mean, with each grid cell weighted equally. Treats the design as if covariates were balanced – the "marginal mean assuming a balanced design" estimand. Best when the goal is to report a covariate-purified comparison independent of the empirical covariate distribution.

Both methods reduce to the same linear-contrast formula emmean = avg_row %*% beta and inherit the spicy variance pipeline (HC* / CR* / bootstrap / jackknife). They give the same answer when there are no covariates, and also when all covariates are numeric (fixed at their means either way). The two estimands diverge when a factor, character, or logical covariate has non-uniform observed proportions: those are levels to balance, and "balanced" weights them equally (a logical is a two-level factor to lm(), balanced FALSE/TRUE – the emmeans convention).

exclude

Columns to exclude from select. Supports tidyselect syntax and character vectors of column names.

regex

Logical. If FALSE (the default), uses tidyselect helpers. If TRUE, the select argument is treated as a regular expression.

weights

Optional case weights. Accepts:

  • NULL (default): an ordinary unweighted lm() is fit.

  • an unquoted numeric column name present in data.

  • a single character column name present in data.

  • a numeric vector of length nrow(data) evaluated in the calling environment.

Validation: weights must be finite, non-negative, and contain at least one positive value (otherwise the function errors). Rows with NA in weights are excluded from the analytic sample for each outcome, alongside rows with NA in y or by. When supplied, weights are passed to lm(..., weights = ...), so coefficients become weighted least-squares estimates and ⁠\eqn{R^2}{R^2}⁠, adjusted ⁠\eqn{R^2}{R^2}⁠, and the four effect sizes are computed from the corresponding weighted sums of squares (see the Weights section in Details).

vcov

Variance estimator used for standard errors, confidence intervals, and Wald test statistics. One of:

  • "classical" (default): the ordinary OLS/WLS variance from vcov(lm), which assumes homoscedastic errors.

  • "HC0": the original Eicker–White heteroskedasticity-consistent sandwich estimator (White 1980), with no finite-sample correction.

  • "HC1": HC0 multiplied by n / (n - p) (MacKinnon and White 1985). Matches Stata's ⁠, robust⁠ default.

  • "HC2": residuals divided by sqrt(1 - h_ii) (MacKinnon and White 1985).

  • "HC3": residuals divided by (1 - h_ii) (MacKinnon and White 1985). A common default for small to moderate samples (Long and Ervin 2000).

  • "HC4": leverage-adaptive variant designed for influential observations (Cribari-Neto 2004).

  • "HC4m": refinement of HC4 with a modified leverage exponent (Cribari-Neto and da Silva 2011).

  • "HC5": alternative leverage-adaptive variant designed for leveraged data (Cribari-Neto, Souza and Vasconcellos 2007).

  • "CR0", "CR1", "CR2", "CR3": cluster-robust sandwich estimators for non-independent observations (Liang and Zeger 1986); requires cluster. "CR2" is the modern default (Bell and McCaffrey 2002; Pustejovsky and Tipton 2018), with Satterthwaite degrees of freedom for inference; the fractional df is reported in the df2 column and in the t(df) / F(df1, df2) test header. "CR1" applies the G/(G-1) correction only; Stata's ⁠, vce(cluster id)⁠ uses the larger G(N-1)/((G-1)(N-p)) factor with t(G-1) inference (clubSandwich's "CR1S"; exposed for lm fits in table_regression(), not here), so "CR1" does not reproduce Stata. Cluster-robust variants are dispatched to clubSandwich::vcovCR() and inference uses clubSandwich::coef_test() / clubSandwich::Wald_test(); install clubSandwich to use them.

  • "bootstrap": nonparametric (resampling cases) or cluster bootstrap variance, depending on whether cluster is supplied (Davison and Hinkley 1997; Cameron, Gelbach and Miller 2008). The number of replicates is set by boot_n. Inference is asymptotic (z for single contrasts, chi^2(q) for the global Wald test); CIs are Wald-type around the point estimate.

  • "jackknife": leave-one-out variance, or leave-one-cluster-out when cluster is supplied (Quenouille 1956; MacKinnon and White 1985). Inference is asymptotic (z / chi^2(q)).

The ⁠HC*⁠ variants are computed via sandwich::vcovHC(). Coefficients (means, contrasts, slopes), ⁠\eqn{R^2}{R^2}⁠, and the standardized effect sizes (f2, d, g, omega2) are point estimates from the OLS/WLS fit and are not affected by vcov; only their standard errors, CIs, and the test statistic of the contrast change.

cluster

Cluster identifier for cluster-aware variance estimators. Required when vcov is one of the ⁠CR*⁠ variants; optional and triggers a cluster bootstrap or leave-one-cluster-out jackknife when vcov is "bootstrap" / "jackknife"; forbidden for the other (independent-observation) variants. Accepts:

  • NULL (default): no cluster structure.

  • an unquoted column name in data.

  • a single character column name in data.

  • a one-sided formula naming a column in data (~ region), the sandwich / fixest convention shared with table_regression().

  • an atomic vector of length nrow(data) evaluated in the calling environment (factor, character, integer, etc.).

Rows with NA in cluster are excluded from the analytic sample for each outcome (alongside rows with NA in y, by, or weights). At least two distinct non-missing cluster values are required. Multi-way clustering (a list / data.frame of multiple cluster vectors) is not supported; use sandwich::vcovCL() or clubSandwich::vcovCR() directly on the fitted model for that case.

boot_n

Integer. Number of bootstrap replicates used when vcov = "bootstrap". Defaults to 1000. Ignored otherwise. Must be at least 50 (values below that floor raise an error: fewer replicates make the variance estimate too noisy to use); non-integer values are truncated (e.g. 500.9 becomes 500). Larger values reduce Monte-Carlo error in the bootstrap variance; typical values for inference are 500-2000.

contrast

Contrast display for categorical predictors. One of:

  • "auto" (default): show a single reference contrast Delta (level2 - level1) only when by has exactly two non-empty levels. The reference level is the first level of the factor (R's default treatment-contrast convention, getOption("contrasts")[1]). To change which level acts as the reference, re-level by upstream (for example with forcats::fct_relevel() or stats::relevel()).

  • "none": suppress the contrast column for categorical predictors. Level-specific means are still displayed.

statistic

Logical. If TRUE, includes a test-statistic column in the wide and rendered outputs. Defaults to FALSE.

p_value

Logical. If TRUE, includes a p column in the wide and rendered outputs. Defaults to TRUE.

show_n

Logical. If TRUE, includes an unweighted n column in the wide and rendered outputs. Defaults to TRUE.

show_weighted_n

Logical. If TRUE and weights is supplied, includes a ⁠Weighted n⁠ column equal to the sum of case weights in the analytic sample. Defaults to FALSE.

effect_size

Character. Effect-size column to include in the wide and rendered outputs. One of:

  • "none" (the default): no effect-size column.

  • "f2": Cohen's ⁠\eqn{f^2}{f^2} = \eqn{R^2}{R^2} / (1 - \eqn{R^2}{R^2})⁠. Defined for any predictor type. Familiar from Cohen (1988); standard input for a-priori power analysis. Note that for a single-predictor model, ⁠\eqn{f^2}{f^2}⁠ is a monotone transform of ⁠\eqn{R^2}{R^2}⁠ and adds no information beyond it.

  • "d": Cohen's d = beta_hat / sigma_hat, where beta_hat is the model coefficient (the displayed difference) and sigma_hat is the residual standard deviation from the fitted model. Defined only when by has exactly two non-empty levels; otherwise the function errors. The sign matches the displayed Delta (level2 - level1).

  • "g": Hedges' g = J * d with the small-sample correction J = 1 - 3 / (4 * df_resid - 1). Same domain as "d".

  • "omega2": Hays' omega-squared, a bias-corrected estimator of the population variance explained, less optimistic than ⁠\eqn{R^2}{R^2}⁠ for small samples. Defined for any predictor type and truncated at 0.

When weights is supplied, "d", "g", and "omega2" are derived from the weighted least-squares fit (using weighted sums of squares and the model's weighted residual standard deviation), keeping them consistent with the weighted contrast and its CI shown in the table. All effect sizes are point estimates derived from the OLS/WLS fit and are not affected by vcov.

Under covariate adjustment (covariates non-empty):

  • "f2" and "omega2" become the partial f^2 / partial \omega^2, derived from the partial F of by (the Type-II test of by after all covariates, equal to stats::drop1() in this additive setting) – the correctly-defined effect size when the model is adjusted. For numeric by, partial f^2 equals the squared partial correlation of by with the outcome, divided by ⁠(1 - r^2_partial)⁠.

  • "d" and "g" raise a spicy_unsupported error: Cohen's d and Hedges' g have no canonical extension to adjusted models (the pooled SD is undefined under adjustment). Use "f2" or "omega2" instead – both generalise via partial F.

effect_size_ci

Logical. If TRUE and effect_size != "none", adds a confidence interval for the effect size derived from inversion of the appropriate noncentral distribution (noncentral t for "d" / "g"; noncentral F for "omega2" / "f2"). The CI level is taken from ci_level. In the long output (output = "long"), the bounds are always present in es_ci_lower / es_ci_upper (numeric). In the wide raw output (output = "data.frame"), the bounds appear under the same names, es_ci_lower / es_ci_upper (numeric). In the printed ASCII table and rendered outputs ("tinytable", "gt", "flextable", "word", "excel", "clipboard"), the effect-size column shows the value followed by the CI in brackets (e.g. 0.18 [0.07, 0.30]). Defaults to FALSE. When effect_size = "none", this argument is ignored with a warning.

r2

Character. Fit statistic to include in the wide and rendered outputs. One of:

  • "r2" (default): the model ⁠\eqn{R^2}{R^2}⁠ (summary(lm)$r.squared).

  • "adj_r2": adjusted ⁠\eqn{R^2}{R^2}⁠, penalising for df_effect relative to the residual degrees of freedom.

  • "none": omit the fit-statistic column.

When weights is supplied, ⁠\eqn{R^2}{R^2}⁠ and adjusted ⁠\eqn{R^2}{R^2}⁠ are the weighted least-squares versions reported by summary(lm(..., weights = ...)).

ci

Logical. If TRUE, includes contrast confidence-interval columns in the wide and rendered outputs when a single contrast is shown. Defaults to TRUE.

labels

An optional named character vector of outcome labels. Names must match column names in data. When NULL (the default), labels are auto-detected from variable attributes; if none are found, the column name is used.

ci_level

Confidence level for coefficient and model-based mean intervals (default: 0.95). Must be between 0 and 1 exclusive.

digits

Number of decimal places for descriptive values, regression coefficients, and test statistics (default: 2). Must be a single non-negative number; non-integer values are truncated (e.g. 2.9 becomes 2). Same constraint for fit_digits and effect_size_digits.

fit_digits

Number of decimal places for model-fit columns (⁠\eqn{R^2}{R^2}⁠ or adjusted ⁠\eqn{R^2}{R^2}⁠) in wide and rendered outputs (default: 2).

effect_size_digits

Number of decimal places for the effect-size column (f2, d, g, or omega2) in wide and rendered outputs (default: 2).

p_digits

Integer >= 1. Number of decimal places used to render p-values in the p column (default: 3, the APA Publication Manual standard). Both the displayed precision and the small-p threshold derive from this argument: p_digits = 3 prints .045 and ⁠<.001⁠; p_digits = 4 prints .0451 and ⁠<.0001⁠; p_digits = 2 prints .05 and ⁠<.01⁠. Useful for genomics / GWAS contexts where adjusted p-values can be very small, or for journals using a coarser convention. Leading zeros are always stripped, following APA convention. Values below 1 raise an error; non-integer values are truncated (e.g. 3.7 becomes 3).

decimal_mark

Character used as decimal separator. Either "." (default) or ",".

align

Horizontal alignment of numeric columns in the printed ASCII table and in the tinytable, gt, flextable, word, and clipboard outputs. The first column (Variable) is always left-aligned. One of:

  • "decimal" (default): align numeric columns on the decimal mark, the standard scientific-publication convention used by SPSS, SAS, and LaTeX siunitx. Numeric cells are pre-padded with figure-spaces (U+2007, digit-width) so every string in a column has the same width with the decimal mark at the same internal position; centring those uniform-width strings then stacks the decimal points vertically. The same pad-then-centre strategy is applied on every rendering engine (gt, tinytable, flextable, word, ASCII print) for a homogeneous rendering, matching table_regression(). The clipboard output is delimited text meant to be parsed rather than read at a fixed width, so its cells travel unpadded (a padded number pastes as text next to an unpadded number).

  • "center": center-align all numeric columns.

  • "right": right-align all numeric columns.

The excel output uses the engine's default alignment in any case: cell-string padding does not align decimals under proportional fonts, and writing raw numbers with a numeric format would require a separate refactor.

output

Output format. One of:

  • "default": an ASCII table object, printed when the call is bare

  • "data.frame": a plain wide data.frame

  • "long": a raw long data.frame

  • "tinytable" (requires tinytable)

  • "gt" (requires gt)

  • "flextable" (requires flextable)

  • "excel" (requires openxlsx2)

  • "clipboard" (requires clipr)

  • "word" (requires flextable and officer)

excel_path

File path for output = "excel".

excel_sheet

Sheet name for output = "excel". NULL (the default) uses "Linear models".

clipboard_delim

Delimiter for output = "clipboard" (default: "\t"). A cell holding the delimiter itself, a double quote or a line break is quoted RFC 4180-style, so the grid survives whatever delimiter you choose.

word_path

File path for output = "word".

verbose

Logical. If TRUE, prints messages about ignored non-numeric selected outcomes (default: FALSE).

user_na

Logical. If TRUE (the default), declared missing values in the outcomes, in by, and in covariates are treated as missing and excluded from the fitted models (reflected in the per-group n). If FALSE, the declared codes enter the fits as ordinary values. See the "Declared missing values" section of freq().

style

A journal style: a theme name ("jama", "nejm", "lancet", "annals", "apa", "aer"), a spicy_style() object, or NULL (the default). A style only changes DEFAULTS – any argument you pass explicitly wins over it. Set options(spicy.style = ) for document-wide scope. A theme covers numeric formatting conformity only, not full editorial conformity; ?spicy_style lists the exact rules each one encodes and the official document they come from. An unknown name is an error.

Value

Depends on output:

The Excel sheet carries the same title the console prints on its first row; the table itself starts on row 3, and the note lines sit below the body.

If no numeric outcome columns remain after applying select, exclude, and regex, the function emits a warning and returns an empty data.frame() regardless of output.

Model and outputs

table_continuous_lm() is designed for article-style reporting around a single focal predictor: one model per selected continuous outcome, fitted as lm(outcome ~ by, ...) and optionally extended with case weights and additive covariates (lm(outcome ~ by + cov1 + ...)). For categorical by, the reported means are model-based fitted means (or covariate-adjusted estimated marginal means; see adjustment) for each level, and contrasts come from the same fitted linear model. For an unweighted lm(y ~ factor) with classical variance and no covariates, the fitted means coincide numerically with empirical subgroup means; the model-based qualifier matters because (a) under weights the means become weighted least-squares estimates, (b) their CIs derive from the model vcov (classical, ⁠HC*⁠, ⁠CR*⁠, bootstrap or jackknife), (c) under covariates they become adjusted marginal means, and (d) tests, p-values and effect sizes all come from the same fitted model, keeping the table internally consistent.

Compared with table_continuous(), this function is the model-based companion: choose it when you want heteroskedasticity-consistent standard errors (vcov = "HC*"), model fit statistics, or case weights via lm(..., weights = ...). Because the function exists to report a fitted model, its inferential output is on by default: p_value = TRUE and r2 = "r2" are the defaults; set p_value = FALSE or r2 = "none" to suppress them.

Effect sizes

Effect size is selected explicitly via effect_size (defaults to "none"). All variants are derived from the same fitted model as the displayed coefficients, ⁠\eqn{R^2}{R^2}⁠, and CIs, so the effect size stays internally consistent with the rest of the table.

All four effect sizes are point estimates derived from the OLS/WLS fit and are invariant to vcov: choosing ⁠HC*⁠ changes the SE, CI, and test statistic of the contrast but not the standardized magnitude itself.

Under covariate adjustment (covariates non-empty), "f2" and "omega2" become the partial f^2 / partial \omega^2 of by, derived from the partial F restricted to the focal term (the Type-II test of by after all covariates, equal to stats::drop1() in this additive setting). "d" and "g" raise a spicy_unsupported error: the pooled standard deviation has no canonical extension under adjustment, so Cohen's d and Hedges' g are undefined for adjusted models. See effect_size for the full dispatch.

Confidence intervals for the effect size are available via effect_size_ci = TRUE and use the modern noncentral-distribution inversion approach, the consensus standard in commercial statistical software (Stata esize / ⁠estat esize⁠, SAS ⁠PROC TTEST⁠ and ⁠PROC GLM EFFECTSIZE⁠ 14.2+) and in mainstream R packages (effectsize, MOTE, TOSTER, effsize):

For the weighted case, the CI uses raw (unweighted) group counts and df.residual(fit) = n - p, consistent with the WLS reporting convention (DuMouchel and Duncan 1983). For propensity-score balance assessment or complex-survey designs, dedicated packages (cobalt::bal.tab() for the Austin and Stuart 2015 formulation; survey for design-based effect sizes) are more appropriate.

Robust standard errors

When vcov is one of the ⁠HC*⁠ variants, the standard errors, CIs, and Wald test statistics use a heteroskedasticity-consistent sandwich estimator computed via sandwich::vcovHC() (Zeileis 2004), the canonical R implementation. For a brief guide:

When observations are not independent (repeated measurements per individual, students nested in classes, patients in hospitals, country-year panels), classical and ⁠HC*⁠ standard errors are biased downward. Use the ⁠CR*⁠ variants together with cluster = id_var to get cluster-robust inference (Liang and Zeger 1986). The implementation dispatches to clubSandwich::vcovCR() for the variance and to clubSandwich::coef_test() (single-coefficient, Satterthwaite t) and clubSandwich::Wald_test() (multi-coefficient Hotelling-T-squared with Satterthwaite df, "HTZ") for inference. "CR2" (Bell and McCaffrey 2002; Pustejovsky and Tipton 2018) is the modern recommended default; it generally produces fractional Satterthwaite degrees of freedom in df2, which the displayed t(df) / F(df1, df2) header renders to one decimal. "CR1" applies the G/(G-1) correction only – Stata's ⁠, vce(cluster id)⁠ uses the larger G(N-1)/((G-1)(N-p)) factor with t(G-1) inference (clubSandwich's "CR1S"; exposed for lm fits in table_regression(), not here), so "CR1" does not reproduce Stata. Effect sizes remain invariant to vcov (including ⁠CR*⁠); only the SE, CI, test statistic, and df2 of the contrast change.

Two resampling-based estimators are also available without adding any dependency: vcov = "bootstrap" (nonparametric resampling-cases bootstrap; Davison and Hinkley 1997) and vcov = "jackknife" (leave-one-out delete-1; Quenouille 1956; MacKinnon and White 1985). Supplying cluster switches both to their cluster-aware variants (cluster bootstrap, Cameron, Gelbach and Miller 2008; leave-one-cluster-out jackknife). The number of bootstrap replicates is controlled by boot_n (default 1000); replicates that fail to fit on rank-deficient resamples are dropped, with an explicit warning if more than half fail. Fewer than 10 valid bootstrap replicates (or fewer than 2 jackknife leave-outs) raises spicy_resampling_failed rather than silently reporting a different variance estimator. Inference for both estimators is asymptotic (z for single-coefficient contrasts, chi^2(q) for the multi-coefficient global Wald test on k > 2 categorical predictors), reflected in the displayed test header. Use the bootstrap when the residual distribution is non-standard or the sample is small; use the jackknife as a closed-form, deterministic alternative.

⁠\eqn{R^2}{R^2}⁠, adjusted ⁠\eqn{R^2}{R^2}⁠, and the effect sizes remain ordinary least-squares (or weighted least-squares) statistics regardless of vcov.

Weights

When weights is supplied, table_continuous_lm() fits weighted linear models via lm(..., weights = ...). Means become weighted least-squares estimates and contrasts and slopes are weighted. The fit statistics ⁠\eqn{R^2}{R^2}⁠ and adjusted ⁠\eqn{R^2}{R^2}⁠, as well as Hays' omega^2 and Cohen's ⁠\eqn{f^2}{f^2}⁠, use the corresponding weighted sums of squares from the WLS fit. Cohen's d and Hedges' g use the WLS coefficient and the model's weighted residual standard deviation (summary(fit)$sigma), which is the standard convention for case-weighted regression-style reporting (DuMouchel and Duncan 1983); the noncentral t CI for d / g uses the raw (unweighted) group counts and the residual degrees of freedom of the WLS fit (n - p). This case-weighted workflow is appropriate for weighted article tables, but is not a substitute for a full complex-survey design (see e.g. the survey package), nor for propensity-score balance assessment under the Austin and Stuart (2015) convention (see e.g. cobalt::bal.tab()).

The n column always reports the unweighted analytic sample size for each outcome. When show_weighted_n = TRUE, an additional ⁠Weighted n⁠ column reports the sum of case weights in the same analytic sample.

Display conventions

For dichotomous categorical predictors, the wide outputs report fitted means in reference-level order and label the contrast column explicitly as Delta (level2 - level1). For categorical predictors with more than two levels, no single contrast or contrast CI is shown in the wide outputs; instead, the table reports level-specific means plus the overall F test when statistic = TRUE (or F(df1, df2) when the degrees of freedom are constant across outcomes).

When covariates is non-empty, the printed ASCII table appends an APA-style footer naming the covariates and the chosen estimand, e.g. ⁠Note. Adjusted for age, education (proportional).⁠

The rendering engines carry that same footer as a table note. On the "tinytable" route the note is set one size down; options(spicy.note_style) governs that (see table_regression()).

Optional output engines require the corresponding suggested packages:

Declared missing values

Survey files imported with haven often carry declared missing values: codes such as ⁠8 = Don't know⁠ or 9 = Refused that the source file marks as missing while keeping them distinct from a plain NA. Two kinds of declaration exist: na_values / na_range metadata on haven::labelled_spss() vectors, and tagged missing values created by haven::tagged_na() (the Stata .a, .b, ... convention).

spicy honors the declaration by default (user_na = TRUE): declared missing values are excluded from every statistic exactly like NA – valid percentages, means, chi-squared tests, association measures, row-wise summaries, and group definitions – but they are not erased from display. freq() lists each observed declared value as its own row of the Missing block, with its value label; cross_tab(), table_categorical(), and table_continuous() disclose the exclusion in the table note (⁠Declared missing values removed: x (2).⁠); varlist() and code_book() count them as missing in N_valid / NAs / N_distinct while still listing the declared codes in Values.

Every function involved offers the same escape hatch: set user_na = FALSE to ignore the declaration and treat the declared codes as valid values (the behavior of spicy before 0.13.0). Tagged missing values are genuine NAs either way; for them, user_na = FALSE only collapses the per-tag breakdown back into the regular NA count.

References

Austin, P. C., & Stuart, E. A. (2015). Moving towards best practice when using inverse probability of treatment weighting (IPTW) using the propensity score to estimate causal treatment effects in observational studies. Statistics in Medicine, 34(28), 3661–3679. doi:10.1002/sim.6607

Bell, R. M., & McCaffrey, D. F. (2002). Bias reduction in standard errors for linear regression with multi-stage samples. Survey Methodology, 28(2), 169–181.

Cameron, A. C., Gelbach, J. B., & Miller, D. L. (2008). Bootstrap-based improvements for inference with clustered errors. Review of Economics and Statistics, 90(3), 414–427. doi:10.1162/rest.90.3.414

Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Hillsdale, NJ: Lawrence Erlbaum.

Cribari-Neto, F. (2004). Asymptotic inference under heteroskedasticity of unknown form. Computational Statistics & Data Analysis, 45(2), 215–233. doi:10.1016/S0167-9473(02)00366-3

Cribari-Neto, F., Souza, T. C., & Vasconcellos, K. L. P. (2007). Inference under heteroskedasticity and leveraged data. Communications in Statistics – Theory and Methods, 36(10), 1877–1888. doi:10.1080/03610920601126589

Cribari-Neto, F., & da Silva, W. B. (2011). A new heteroskedasticity-consistent covariance matrix estimator for the linear regression model. AStA Advances in Statistical Analysis, 95(2), 129–146. doi:10.1007/s10182-010-0141-2

Davison, A. C., & Hinkley, D. V. (1997). Bootstrap Methods and Their Application. Cambridge: Cambridge University Press. doi:10.1017/CBO9780511802843

DuMouchel, W. H., & Duncan, G. J. (1983). Using sample survey weights in multiple regression analyses of stratified samples. Journal of the American Statistical Association, 78(383), 535–543. doi:10.1080/01621459.1983.10478006

Fitts, D. A. (2021). Expected and empirical coverages of different methods for generating noncentral t confidence intervals for a standardized mean difference. Behavior Research Methods, 53(6), 2412–2429. doi:10.3758/s13428-021-01550-4

Goulet-Pelletier, J.-C., & Cousineau, D. (2018). A review of effect sizes and their confidence intervals, Part I: The Cohen's d family. The Quantitative Methods for Psychology, 14(4), 242–265. doi:10.20982/tqmp.14.4.p242

Hays, W. L. (1963). Statistics for Psychologists. New York: Holt, Rinehart and Winston.

Hedges, L. V., & Olkin, I. (1985). Statistical Methods for Meta-Analysis. Orlando, FL: Academic Press.

Long, J. S., & Ervin, L. H. (2000). Using heteroscedasticity consistent standard errors in the linear regression model. The American Statistician, 54(3), 217–224. doi:10.1080/00031305.2000.10474549

Liang, K.-Y., & Zeger, S. L. (1986). Longitudinal data analysis using generalized linear models. Biometrika, 73(1), 13–22. doi:10.1093/biomet/73.1.13

MacKinnon, J. G., & White, H. (1985). Some heteroskedasticity-consistent covariance matrix estimators with improved finite sample properties. Journal of Econometrics, 29(3), 305–325. doi:10.1016/0304-4076(85)90158-7

Olejnik, S., & Algina, J. (2003). Generalized eta and omega squared statistics: Measures of effect size for some common research designs. Psychological Methods, 8(4), 434–447. doi:10.1037/1082-989X.8.4.434

Pustejovsky, J. E., & Tipton, E. (2018). Small-sample methods for cluster-robust variance estimation and hypothesis testing in fixed effects models. Journal of Business & Economic Statistics, 36(4), 672–683. doi:10.1080/07350015.2016.1247004

Quenouille, M. H. (1956). Notes on bias in estimation. Biometrika, 43(3/4), 353–360. doi:10.1093/biomet/43.3-4.353

Steiger, J. H. (2004). Beyond the F test: Effect size confidence intervals and tests of close fit in the analysis of variance and contrast analysis. Psychological Methods, 9(2), 164–182. doi:10.1037/1082-989X.9.2.164

Steiger, J. H., & Fouladi, R. T. (1997). Noncentrality interval estimation and the evaluation of statistical models. In L. L. Harlow, S. A. Mulaik, & J. H. Steiger (Eds.), What if there were no significance tests? (pp. 221–257). Mahwah, NJ: Lawrence Erlbaum.

White, H. (1980). A heteroskedasticity-consistent covariance matrix estimator and a direct test for heteroskedasticity. Econometrica, 48(4), 817–838. doi:10.2307/1912934

Zeileis, A. (2004). Econometric computing with HC and HAC covariance matrix estimators. Journal of Statistical Software, 11(10), 1–17. doi:10.18637/jss.v011.i10

See Also

table_continuous(), table_categorical(). For broader workflows on the same statistical building blocks: sandwich::vcovHC() (the canonical R implementation of the ⁠HC*⁠ sandwich estimators, used internally for vcov = "HC*"); clubSandwich::vcovCR(), clubSandwich::coef_test() and clubSandwich::Wald_test() (the canonical R implementation of cluster-robust variance and Satterthwaite-style inference, used internally for vcov = "CR*"); effectsize::cohens_d(), effectsize::hedges_g(), and effectsize::omega_squared() (alternative effect-size computations and CIs); cobalt::bal.tab() for propensity-score covariate balance with weighted standardized mean differences (Austin and Stuart 2015); the survey package for design-based inference on complex-survey samples.

Other spicy tables: table_categorical(), table_continuous()

Examples

# --- Basic usage ---------------------------------------------------------

# Default: ASCII table with model-based means, p, and \eqn{R^2}{R^2}.
table_continuous_lm(
  sochealth,
  select = c(wellbeing_score, bmi),
  by = sex
)

# --- Effect sizes -------------------------------------------------------

# Cohen's d (binary by required).
table_continuous_lm(
  sochealth,
  select = c(wellbeing_score, bmi),
  by = sex,
  effect_size = "d"
)

# Hedges' g with weighted analysis and weighted n column.
table_continuous_lm(
  sochealth,
  select = c(wellbeing_score, bmi),
  by = sex,
  weights = weight,
  statistic = TRUE,
  effect_size = "g",
  show_weighted_n = TRUE
)

# Hedges' g with noncentral t confidence interval (bracket notation).
table_continuous_lm(
  sochealth,
  select = c(wellbeing_score, bmi),
  by = sex,
  effect_size = "g",
  effect_size_ci = TRUE
)

# Cohen's \eqn{f^2}{f^2} alongside \eqn{R^2}{R^2} (familiar power-analysis effect size).
table_continuous_lm(
  sochealth,
  select = c(wellbeing_score, bmi),
  by = sex,
  effect_size = "f2"
)

# Hays' omega-squared for a 3-level predictor (d / g would error here).
table_continuous_lm(
  sochealth,
  select = c(wellbeing_score, bmi),
  by = education,
  effect_size = "omega2"
)

# --- Robust SE for a numeric predictor ----------------------------------

# HC3 standard errors for the slope of a continuous predictor.
table_continuous_lm(
  sochealth,
  select = c(wellbeing_score, bmi),
  by = age,
  vcov = "HC3",
  ci = FALSE
)

# Cluster-robust SE for repeated-measures data: the `sleep` dataset
# has 10 subjects measured twice (one observation per group).
if (requireNamespace("clubSandwich", quietly = TRUE)) {
  table_continuous_lm(
    sleep,
    select = extra,
    by = group,
    cluster = ID,
    vcov = "CR2"
  )
}

# --- Covariate adjustment ----------------------------------------------

# Adjust the comparison of `wellbeing_score` and `bmi` by `sex` for `age`
# and `education`. The footer surfaces the adjustment estimand
# ("proportional" by default = G-computation, matching Stata `margins`).
table_continuous_lm(
  sochealth,
  select = c(wellbeing_score, bmi),
  by = sex,
  covariates = c(age, education),
  vcov = "HC3"
)

# Same model with the emmeans / SPSS UNIANOVA convention (equal-weight
# marginal means on a synthetic covariate grid).
table_continuous_lm(
  sochealth,
  select = c(wellbeing_score, bmi),
  by = sex,
  covariates = c(age, education),
  adjustment = "balanced",
  vcov = "HC3"
)

# Effect sizes adjust automatically: f2 / omega2 become partial
# effect sizes via the partial F restricted to the focal `by`.
# d / g are undefined under adjustment and raise spicy_unsupported.
table_continuous_lm(
  sochealth,
  select = c(wellbeing_score, bmi),
  by = sex,
  covariates = c(age, education),
  effect_size = "f2",
  effect_size_ci = TRUE
)

# --- Article-style polish -----------------------------------------------

# Pretty outcome labels and adjusted \eqn{R^2}{R^2}.
table_continuous_lm(
  sochealth,
  select = c(wellbeing_score, bmi),
  by = sex,
  labels = c(
    wellbeing_score = "WHO-5 wellbeing (0-100)",
    bmi = "Body-mass index (kg/m^2)"
  ),
  r2 = "adj_r2"
)

# European decimal comma.
table_continuous_lm(
  sochealth,
  select = c(wellbeing_score, bmi),
  by = sex,
  decimal_mark = ","
)

# Regex selection of all columns starting with "life_sat".
table_continuous_lm(
  sochealth,
  select = "^life_sat",
  by = sex,
  regex = TRUE
)

# --- Output formats -----------------------------------------------------

# The rendered outputs below all wrap the same call:
#   table_continuous_lm(sochealth,
#                       select = c(wellbeing_score, bmi),
#                       by = sex)
# only `output` changes. Assign to a variable to avoid the
# console-friendly text fallback that some engines fall back to
# when printed directly in `?` help.

# Wide data.frame (one row per outcome).
table_continuous_lm(
  sochealth,
  select = c(wellbeing_score, bmi),
  by = sex,
  output = "data.frame"
)

# Raw long data.frame (one block per outcome).
table_continuous_lm(
  sochealth,
  select = c(wellbeing_score, bmi),
  by = sex,
  output = "long"
)


# Rendered HTML / docx objects -- best viewed inside a
# Quarto / R Markdown document or a pkgdown article.
if (requireNamespace("tinytable", quietly = TRUE)) {
  tt <- table_continuous_lm(
    sochealth, select = c(wellbeing_score, bmi), by = sex,
    output = "tinytable"
  )
}
if (requireNamespace("gt", quietly = TRUE)) {
  tbl <- table_continuous_lm(
    sochealth, select = c(wellbeing_score, bmi), by = sex,
    output = "gt"
  )
}
if (requireNamespace("flextable", quietly = TRUE)) {
  ft <- table_continuous_lm(
    sochealth, select = c(wellbeing_score, bmi), by = sex,
    output = "flextable"
  )
}

# Excel and Word: write to a temporary file.
if (requireNamespace("openxlsx2", quietly = TRUE)) {
  tmp <- tempfile(fileext = ".xlsx")
  table_continuous_lm(
    sochealth, select = c(wellbeing_score, bmi), by = sex,
    output = "excel", excel_path = tmp
  )
  unlink(tmp)
}
if (
  requireNamespace("flextable", quietly = TRUE) &&
    requireNamespace("officer", quietly = TRUE)
) {
  tmp <- tempfile(fileext = ".docx")
  table_continuous_lm(
    sochealth, select = c(wellbeing_score, bmi), by = sex,
    output = "word", word_path = tmp
  )
  unlink(tmp)
}


## Not run: 
# Clipboard: writes to the system clipboard.
table_continuous_lm(
  sochealth, select = c(wellbeing_score, bmi), by = sex,
  output = "clipboard"
)

## End(Not run)


Descriptive statistics from a survey design

Description

The design twin of table_continuous(): the same table of means, standard deviations, intervals and counts, computed from a survey::svydesign() or survey::as.svrepdesign() object instead of a data frame.

Not one statistic is computed here. survey::svymean() gives the mean, its standard error and its design effect, survey::svyvar() the standard deviation, survey::svyquantile() the quantiles, and survey::svyttest() / survey::regTermTest() / survey::svyranktest() the group comparison. Every interval and every test is referred to survey::degf(design).

Usage

table_continuous_svy(
  design,
  select = tidyselect::everything(),
  by = NULL,
  exclude = NULL,
  regex = FALSE,
  drop_na = TRUE,
  deff = FALSE,
  qrule = "math",
  df = NULL,
  test = c("welch", "student", "nonparametric"),
  p_value = NULL,
  statistic = FALSE,
  show_n = TRUE,
  show_columns = NULL,
  ci = TRUE,
  labels = NULL,
  ci_level = 0.95,
  digits = 2,
  p_digits = 3,
  decimal_mark = ".",
  align = c("decimal", "center", "right"),
  output = c("default", "data.frame", "long", "tinytable", "gt", "flextable", "excel",
    "clipboard", "word"),
  excel_path = NULL,
  excel_sheet = NULL,
  clipboard_delim = "\t",
  word_path = NULL,
  verbose = FALSE,
  user_na = TRUE,
  style = NULL
)

Arguments

design

A survey design: survey::svydesign() or survey::as.svrepdesign(). Two-phase, pps, database-backed and multiframe designs are refused with a classed error.

select

Columns to summarize, as a tidyselect expression on the design's variables.

by

A single grouping column. One domain per level.

exclude

Columns to drop from select.

regex

Treat select as a regular expression.

drop_na

Drop observations with a missing by value (default TRUE). With FALSE they form a (Missing) domain of their own – an ordinary subpopulation, with its own degrees of freedom – which is excluded from the group comparison.

deff

Show the design effect: FALSE (default), TRUE (against sampling without replacement) or "replace" (against sampling with replacement, ignoring the finite population correction).

qrule

Quantile rule: "math" (default), "spicy", or anything survey::svyquantile() accepts, including a function.

df

Degrees of freedom for the intervals. NULL (default) uses survey::degf() on each domain. It does not reach the group comparison: survey::svyttest() and survey::svyranktest() have no df argument, so the test keeps the design's own degrees of freedom and the note says so.

test

Group comparison: "welch" (default), "student" (warns; identical under a design) or "nonparametric". A variable whose complete cases include negatively weighted rows is not compared: a design-based test is not defined when the weights change sign. Its estimates are still reported, the note says which comparisons were withheld, and the call warns (spicy_negative_weights_no_test).

p_value

Show the p-value column (defaults to TRUE with by).

statistic

Show the test-statistic column.

show_n

Show the count column.

show_columns

Character vector of statistic tokens; NULL keeps the default display.

ci, ci_level

The mean's confidence interval and its level.

labels

Named character vector of display labels.

digits, p_digits, decimal_mark

Number formatting.

align

Numeric-cell alignment: "decimal", "center" or "right".

output

One of "default", "data.frame", "long", or a rendering engine: "tinytable", "gt", "flextable", "excel", "clipboard", "word". "data.frame" and "long" are synonyms and return the same object: the compute frame is already long, one row per ⁠(variable x group)⁠. Same pair, same reason, as in table_continuous().

excel_path, excel_sheet, clipboard_delim, word_path

Output destinations, as in table_continuous().

verbose

Report the columns skipped as non-numeric.

user_na

Honour declared missing values (see ?freq).

style

A journal style; see spicy_style().

Value

A spicy_continuous_svy_table: the compute frame, with the display frame and the typed view attached. output = "data.frame" / "long" returns the compute frame unclassed – the two tokens are synonyms and return identical objects.

Which function do I need?

A data frame with a column of weights is table_continuous() (⁠weights = ⁠). A survey design object – strata, clusters, finite population correction, calibration, replicate weights – is this function. Passing one to the other errors with the name of the right one; there is no silent coercion, because the design-based standard errors, degrees of freedom and tests cannot be recovered from the weights alone.

Two conventions, one bridge

table_continuous(weights = ) implements the frequency-expansion convention: a weight is a number of copies, and SD has denominator sum(w) - 1. This function implements the sampling-weight convention: a weight is a number of units represented, and SD is sqrt(survey::svyvar()), whose denominator is n - 1 on weights normalised to sum to n. These are two estimands, not two approximations of one.

rescale = TRUE is the bridge, and it is an identity rather than a coincidence. Writing ⁠w' = w * n / sum(w)⁠, the rescaled weighted variance is ⁠sum(w' (x - xbar)^2) / (sum(w') - 1) = n / (n - 1) * sum(w (x - xbar)^2) / sum(w)⁠, which is what survey::svyvar() computes. So on a design that declares nothing but weights, table_continuous(weights = w, rescale = TRUE) and this function return the same mean and the same standard deviation. The default rescale = FALSE does not, and that is the estimand boundary, not a bug.

The mean is continuous across both regimes: ⁠sum(w x) / sum(w)⁠ does not move when the weights are rescaled.

Choosing the statistics

show_columns takes the tokens of table_continuous() with two additions and one removal:

Quantiles

qrule = "math" is the default and estimates ⁠inf{x : F(x) >= p}⁠, the quantile of the population. qrule = "spicy" switches to the type-7 interpolation table_continuous() uses, for a reader who needs the two tables to agree cell for cell; any other value – including a function – is handed to survey::svyquantile() untouched. The note always says which rule produced the numbers.

Groups and degrees of freedom

⁠by = ⁠ cuts one domain per group with [ on the design. survey recomputes the degrees of freedom on the primary sampling units and strata each domain retains, so a grouped table generally carries a different df per row; the note gives the span when they differ.

A group with a missing value is a domain like any other: drop_na = FALSE gives it a (Missing) row, with its own degrees of freedom. A domain reduced to one primary sampling unit has none, and its interval shows the undefined dash rather than an interval built on qt(p, df = 0).

The comparison is a single test on the whole design, not a set of pairwise ones: survey::svyttest() with two observed groups, survey::regTermTest() on survey::svyglm() with three or more, or survey::svyranktest() under test = "nonparametric". Under a design the Welch / Student distinction does not exist – the variance is the design's – so test = "student" warns and behaves like "welch".

Stability

This function is stabilising in the sense ?spicy defines: the names of its design-specific arguments may still be tightened before 1.0 – with a NEWS.md entry – but the behaviour does not change silently. It was experimental through 0.13; the shape of the table has held across the cycle, so it now moves on the parent family's clock rather than its own. The numbers themselves are survey's and do not move with it.

What is absent, and why

weights and rescale (the weighting is the design), effect_size and smd (no established design-based variance), and data.

See Also

table_continuous() for the data-frame sibling, table_categorical_svy() for categorical variables, table_regression() on a survey::svyglm() fit for a model.

Examples


data(api, package = "survey")
dclus1 <- survey::svydesign(
  id = ~dnum, weights = ~pw, data = apiclus1, fpc = ~fpc
)
table_continuous_svy(dclus1, select = c(api00, api99))
table_continuous_svy(dclus1, select = api00, by = stype)
table_continuous_svy(
  dclus1,
  select = api00,
  show_columns = c("m", "se", "ci", "deff", "n"),
  deff = TRUE
)


Describe one continuous outcome across several groupings

Description

Summarises one continuous outcome across the levels of several categorical variables, one block of rows per variable. It is the inverse layout of table_continuous(), which puts several outcomes in rows and one grouping in columns.

Each block reports the outcome's statistics level by level, plus its own group comparison on the block's header row, and an Overall row gives the marginal summary of the whole analytic sample.

Usage

table_outcome(
  data,
  outcome,
  select,
  labels = NULL,
  overall = TRUE,
  drop_na = FALSE,
  weights = NULL,
  rescale = FALSE,
  test = c("welch", "student", "nonparametric"),
  p_value = NULL,
  statistic = FALSE,
  show_n = TRUE,
  show_columns = NULL,
  effect_size = c("none", "auto", "hedges_g", "eta_sq", "r_rb", "epsilon_sq"),
  effect_size_ci = FALSE,
  ci = TRUE,
  ci_level = 0.95,
  digits = 2,
  effect_size_digits = 2,
  p_digits = 3,
  decimal_mark = ".",
  align = c("decimal", "center", "right"),
  output = c("default", "data.frame", "long", "tinytable", "gt", "flextable", "excel",
    "clipboard", "word"),
  indent_text = "  ",
  indent_text_excel_clipboard = strrep(" ", 6),
  excel_path = NULL,
  excel_sheet = NULL,
  clipboard_delim = "\t",
  word_path = NULL,
  user_na = TRUE,
  style = NULL,
  by
)

Arguments

data

A data frame.

outcome

The continuous outcome, unquoted or as a string. Exactly one column.

select

The grouping characteristics to describe the outcome across, as a tidyselect expression or a character vector of column names. One block of rows per variable, in the order given. As everywhere in the family, select is what structures the rows.

labels

Named character vector of display labels, for the outcome and for the select variables alike.

overall

Show the marginal Overall row (default TRUE).

drop_na

Drop rows with a missing select value from that block (default FALSE: they are shown as a (Missing) level and excluded from the comparison).

weights, rescale

Frequency weights and whether to rescale them to sum to the sample size, as in table_continuous().

test

Group comparison for every block: "welch" (default), "student" or "nonparametric".

p_value

Show the p-value column (default TRUE).

statistic

Show the test statistic column.

show_n

Show the count column.

show_columns

Character vector of statistic tokens; NULL keeps the historical display.

effect_size, effect_size_ci

Effect size per block and its confidence interval, as in table_continuous().

ci, ci_level

The mean's confidence interval and its level.

digits, effect_size_digits, p_digits, decimal_mark

Number formatting.

align

Numeric-cell alignment: "decimal", "center" or "right".

output

One of "default", "data.frame", "long", or a rendering engine: "tinytable", "gt", "flextable", "excel", "clipboard", "word". "data.frame" and "long" are synonyms and return the same object: the compute frame is already long (one row per level, plus the block header and the overall row). Both names are kept so the argument reads the same here as in table_continuous(), where the pair is also synonymous.

indent_text, indent_text_excel_clipboard

Level-row indentation, for the console and for the plain-text engines.

excel_path, excel_sheet, clipboard_delim, word_path

Output destinations, as in table_continuous().

user_na

Honour declared missing values (see ?freq).

style

A journal style; see spicy_style().

by

Defunct. The grouping characteristics are selected with select since spicy 0.13.0; supplying by is an error.

Value

A spicy_outcome_table: the compute frame, with the display frame and the typed view attached. output = "data.frame" / "long" returns the compute frame unclassed – the two tokens are synonyms and return identical objects.

Which shape do I need?

Several continuous variables across one grouping is table_continuous() (⁠select = ⁠, ⁠by = ⁠). One continuous variable across one or several groupings is this function. A single select is legitimate here – it is the natural way in when you know more groupings are coming – but with several outcomes and one grouping, the sibling is the table you want.

Choosing the statistics

show_columns takes the same tokens as table_continuous(), with the same meanings; see the show_columns section of ?table_continuous for the vocabulary. Only the character-vector form is accepted here: there is one outcome, so a per-variable list would name nothing.

Weights

weights applies the frequency-expansion convention of the family: all weights 1 reproduces the unweighted table, and integer weights reproduce the STATISTICS of the data duplicated that many times – n stays the raw count of rows that carried the weights. Rows with a missing or zero weight leave the analytic sample; the note counts the missing ones.

rescale is the switch between the two readings of a weight: the frequency reading above, and the sampling-weight reading, where the weights are normalised to sum to the sample size. See the Weights section of table_continuous() for the choice in full.

rescale = TRUE normalises the weights over the outcome's whole surviving sample, once, never per level – a per-level rescale would destroy the relative weights across levels, which is the entire information a sampling weight carries into this table. The means are unchanged by it; the standard deviations move, because their denominator is sum(w) - 1.

A weighted table refuses the group comparison. The estimates and their interval have no weighted version here, and a p-value or an effect size silently computed unweighted beside weighted descriptives is the one thing that must not happen: set p_value = FALSE (and statistic = FALSE, effect_size = "none"), or use table_continuous_lm() for a weighted comparison. The order-statistic median interval is refused for the same reason.

Blocks and the group comparison

Every block is a separate one-way comparison of the outcome across the levels of that variable. Nothing in this table adjusts one block for another, and the table note says so. Read the blocks as a set of bivariate descriptions, not as a model.

Each block chooses its test independently: with two observed levels test = "welch" is the Welch t-test, with three or more it is the Welch one-way ANOVA, and test = "nonparametric" is the Wilcoxon rank-sum or the Kruskal-Wallis test on the same rule. A block with fewer than two observed levels, or with a level holding a single observation, is not tested; its statistics stay empty and the other blocks are unaffected.

The Overall row

overall = TRUE puts the marginal summary of the whole analytic sample on the first row. Under the default drop_na = FALSE the levels of every block partition that sample – the (Missing) display level included – so each block's counts add up to the Overall count exactly, which is what makes it a usable denominator.

The row reads Overall, not Total, and the distinction is deliberate. Total is the word of a COUNT margin: the column of table_categorical() where frequencies add up. This row is the whole analytic sample, where a mean is recomputed over every observation and nothing is added. A mean is not a total.

Choosing the select columns

The canonical form is select = where(is.factor), or an explicit enumeration. Negation (select = -c(x, y)) is not recommended: it sweeps in every remaining column, and a numeric one becomes a block with one LEVEL per distinct value, in order of first appearance. A variable producing more than 20 levels raises a warning for that reason – an arbitrary threshold, but a sixty-row block where a reader expects a handful of categories is not a table.

A haven_labelled column used as select shows its numeric CODES, not its value labels, as it does in table_continuous(). Convert it first (haven::as_factor()) to get the labels in the stub.

The table note

One note sits under the table and states what left the analytic sample, which group comparison ran in each block, what the displayed columns mean, and how the blocks and the Overall row are to be read. The rendering engines carry the same sentence as a table note. On the "tinytable" route it is set one size down; options(spicy.note_style) governs that (see table_regression()).

See Also

table_continuous() for the transposed shape, table_categorical() for categorical outcomes.

Examples

table_outcome(sochealth, bmi, select = c(sex, smoking))
table_outcome(sochealth, wellbeing_score, select = where(is.factor))

Regression coefficient summary table

Description

Publication-ready coefficient table from one or more fitted lm / glm models. Supports standardised coefficients (\beta), average marginal effects (AME), partial effect sizes (f^2 / \eta^2 / \omega^2 for lm; partial \chi^2 for glm), pseudo-R^2 (glm), and a full vocabulary of variance estimators (classical / HC* / cluster-robust with Satterthwaite-corrected df / bootstrap / jackknife). glm covers binomial / poisson / Gamma / inverse.gaussian / quasi families with any link.

Usage

table_regression(
  models,
  vcov = "classical",
  cluster = NULL,
  ci_level = 0.95,
  ci_method = c("wald", "profile", "boot_percentile", "hdi"),
  boot_n = 1000L,
  tau = NULL,
  at_time = NULL,
  standardized = c("none", "refit", "posthoc", "basic", "smart", "pseudo"),
  exponentiate = FALSE,
  p_adjust = "none",
  show_columns = NULL,
  keep = NULL,
  drop = NULL,
  show_intercept = TRUE,
  show_thresholds = TRUE,
  show_components = TRUE,
  intercept_position = c("first", "last"),
  factor_layout = c("grouped", "flat"),
  reference_style = c("row", "annotation", "footer", "none"),
  reference_label = "(ref.)",
  show_fit_stats = NULL,
  fit_stats_layout = c("first_col", "merged"),
  show_re = TRUE,
  re_scale = c("sd", "variance"),
  re_columns = c("est", "se", "ci"),
  re_test = c("none", "lrt", "rlrt"),
  re_ci = c("wald", "profile"),
  model_labels = NULL,
  outcome_labels = NULL,
  stars = FALSE,
  nested = FALSE,
  digits = 2L,
  p_digits = 3L,
  effect_size_digits = 2L,
  fit_digits = 2L,
  ic_digits = 1L,
  decimal_mark = ".",
  align = c("decimal", "center", "right"),
  padding = 0L,
  labels = NULL,
  title = NULL,
  note = NULL,
  output = c("default", "data.frame", "long", "gt", "flextable", "tinytable", "excel",
    "clipboard", "word"),
  excel_path = NULL,
  excel_sheet = NULL,
  clipboard_delim = "\t",
  word_path = NULL,
  word_template = NULL,
  style = NULL
)

Arguments

models

A fitted model object, or a list of such fits (named or unnamed; classes may be mixed). Single fits are auto-promoted to a 1-element list. A broad set of model classes is supported – linear / generalised linear (lm, glm, MASS::glm.nb), mixed-effects (lmer, lme, glmmTMB), survival (coxph, survreg), ordinal (polr, clm), mgcv::gam/bam, betareg, mlogit, survey::svyglm, rms (ols/lrm/cph/Glm), and Bayesian (rstanarm/brms), among others. An unsupported class raises spicy_unsupported. Raw data + formula is not accepted – fit-only API.

vcov

Variance-covariance estimator: "classical", "HC0"-"HC5", "CR0"-"CR3", "bootstrap", or "jackknife". A scalar is recycled to all models; a list (one string per model) allows mixed estimators. Default "classical". The resampling estimators (lm / glm, including MASS::glm.nb) refit each replicate on resampled rows of the fixed evaluated design; for glm.nb the dispersion parameter theta is held at its full-sample estimate (Stata's ⁠nbreg, vce(bootstrap)⁠ re-estimates it per replicate – the two conventions differ slightly). Resampling that fails on nearly every replicate raises spicy_resampling_failed rather than silently reporting a different estimator. Quantile regression (quantreg::rq) uses its own estimator family instead: "classical" resolves to the heteroskedasticity-robust "nid" sandwich (Hendricks-Koenker, Hall-Sheather bandwidth – quantreg's own large-sample default and the Stata vce(robust) analogue), with "iid" (Koenker-Bassett homoskedastic, Stata vce(iid) parity), "ker" (Powell kernel) and "rank" (rank-score inversion: genuine CIs, no SE / t / p – those cells render as dashes) as opt-ins, and "bootstrap" running quantreg's native boot.rq (bsmethod = "xy"; with ⁠cluster =⁠, the Hagemann 2017 wild gradient cluster bootstrap – the one cluster-robust route for rq; ⁠CR*⁠ and ⁠HC*⁠ are refused for rq, as is "jackknife", whose leave-one-out form is inconsistent for quantiles). See Inference and standard errors.

cluster

Cluster identifier for cluster-robust variance (used when vcov is "CR0"-"CR3" or a cluster-bootstrap / cluster-jackknife). Three accepted forms (see How to specify cluster in the details):

  • Formula: ~region, ~region:year (recommended).

  • String column name: "region".

  • Atomic vector of length nobs(fit): df$region, interaction(df$region, df$year), ... (for keys derived on the fly).

For multi-model use, pass a list of one form per model (mix-and-match allowed). A bare unquoted name (cluster = region) is not column selection: it is evaluated as an ordinary variable (the vector form when region exists in the calling environment, a migration error pointing at ~region / "region" otherwise). Default NULL (no clustering).

ci_level

Confidence level for all reported CIs (B, \beta, AME, partial effect sizes). Default 0.95.

ci_method

CI construction. "wald" (default) uses ⁠estimate +/- z x SE⁠ (⁠t x SE⁠ for lm). "profile" (glm and ordinal MASS::polr / ordinal::clm) uses the profile-likelihood CI from stats::confint() – MASS::confint.glm() for glm, confint.polr() / confint.clm() for ordinal fits – asymmetric, exact for likelihood-based inference (Venables & Ripley MASS Section 7.2). Only the CI bounds change; estimate, SE, statistic and p-value remain Wald. For ordinal fits the profile covers the predictor coefficients (the cut-point thresholds stay Wald), and a robust vcov takes precedence (its Wald-robust CIs are used instead, with a consolidated warning). "profile" with lm raises spicy_invalid_input. "boot_percentile" (requires vcov = "bootstrap" for every model) replaces the coefficient CI bounds with equal-tailed percentile intervals of the bootstrap replicates (the boot::boot.ci() type = "perc" convention; Davison & Hinkley 1997, ch. 5), reusing the same resamples as the bootstrap SEs. Only the CI bounds change; estimate, SE, statistic and p-value remain Wald from the bootstrap covariance – the Stata convention (normal-based table CIs by default, percentile on request via ⁠estat bootstrap⁠). AME and standardized-beta CIs are not covered. With exponentiate = TRUE the percentile bounds are exponentiated (percentile intervals are transformation-respecting). "hdi" (Bayesian stanreg / brmsfit fits only) replaces the default equal-tailed credible interval with the highest-density interval – the shortest interval containing ci_level of the posterior draws (Kruschke 2015), relabelling the column header ⁠95% HDI⁠. Unlike the equal-tailed interval the HDI is not transformation-invariant, so under exponentiate = TRUE it is recomputed on the exponentiated draws rather than transformed. Requesting "hdi" for a frequentist fit, or "profile" / "boot_percentile" for a Bayesian fit, raises spicy_invalid_input.

boot_n

Number of bootstrap replicates when vcov = "bootstrap". Single positive integer. Default 1000L.

tau

RMST horizon: the "rmst" column family reports the restricted-mean-survival-time difference over ⁠[0, tau]⁠ (coxph and survreg fits). A positive time on the outcome's scale, or "minmax" for the smallest per-group maximum follow-up (the resolved value is disclosed in the table note). No default: the horizon defines the estimand. Stratified fits (strata()) are supported: standardization keeps each subject's own stratum baseline, disclosed in the table note.

at_time

Landmark time for the "risk_diff" column family: the difference in cumulative incidence at at_time (coxph and survreg fits). A positive time on the outcome's scale; no default.

standardized

Standardisation method for the "beta" column. One of "none" (default), "refit", "posthoc", "basic", "smart", "pseudo". "pseudo" is glm only (Menard 2011 fully-standardised); using it with lm() raises spicy_invalid_input. Supported classes: lm, glm (incl. MASS::glm.nb), the mixed engines (lmer / glmer / glmmTMB / nlme::lme), and fixed-effects stan_glm-style stanreg fits for the algebraic flavors "posthoc" / "basic" / "smart" – exact affine rescales of the posterior draws (median, MAD SD and credible bounds all scale by the same positive SD ratio; the beta stays on the link scale under exponentiate = TRUE). Standard-formula fixed-effects brmsfit models are supported too (their design matrix is recovered through insight); the scale factors are engine-invariant, so the same model fit with rstanarm or brms gets identical betas up to sampling noise. Bayesian "refit" (would re-run the sampler) and "pseudo" (per-draw latent variance; planned) are refused. Also refused: multilevel fits (stan_glmer / stan_lmer and brm() with group-level terms; a Bayesian beta would need an explicit sd(Y) decomposition that the frequentist mixed engines resolve by refitting on z-scored data), non-GLM stanreg subclasses (stan_polr, stan_betareg), and brms formulas with distributional, multivariate or special terms (mo(), s(), ...). In all refused Bayesian cases, standardize predictors before fitting instead. Any other class raises spicy_unsupported_standardized rather than rendering an empty beta column. See the Standardised coefficients section.

exponentiate

Logical. When TRUE, B, the CI bounds, and the SE (delta method: ⁠SE_OR = OR x SE_log-odds⁠) are exp()-transformed for links where the result is a ratio: logit (OR), log (IRR / RR / MR per family, generic exp(B) for other log-link families – a genuine ratio of means), and binomial / ordinal cloglog (HR; grouped-time proportional hazards, Prentice & Gloeckler 1978). For the ordinal (cumulative) cloglog model the displayed HR is exp(-B), not exp(B): the cumulative parametrisation ⁠cloglog P(Y <= j) = zeta_j - xB⁠ places the hazard of the grouped event on -B, so exp(B) would be the reciprocal of the hazard ratio (the CI endpoints are negated and swapped accordingly, and the table note discloses the convention; a cloglog clm with ⁠nominal =⁠ terms refuses exponentiate because threshold-side coefficients have no covariate-HR reading). The statistic and p-value stay on the link scale (invariant under monotone transformation). Identity-link fits are left untouched (a spicy_ignored_arg warning fires when no model in the table exponentiates), so mixed lm + logit tables keep working. Any other link (probit, cauchit, inverse – the Gamma() default –, 1/mu^2, sqrt, ordinal loglog, ...) raises spicy_invalid_input: the exponential of such a coefficient has no ratio interpretation, and reporting it would mislabel the estimate. Report response-scale effects for those models via the AME column instead. Note that the displayed CI is the link-scale CI with exponentiated endpoints (Wald or profile per ci_method), so it is asymmetric around the ratio and cannot be reconstructed as ⁠estimate ± z × SE⁠; the delta-method SE follows the Stata convention (⁠[R] logistic⁠, Methods and formulas: ⁠se(OR) = OR × se(b)⁠) and the log-scale CI endpoints carry the uncertainty more faithfully than this SE. The footer states the SE scale and the CI asymmetry whenever an SE column is displayed. Default FALSE.

p_adjust

Multiple-comparison adjustment method applied to the family of estimated coefficient p-values within each model (intercept and reference rows excluded). One of "none" (default), "holm", "hochberg", "hommel", "bonferroni", "BH" / "fdr", or "BY". Delegated to stats::p.adjust(); applied per-model and per estimate_type (B and AME p-values are adjusted independently within their own families). Active adjustments are documented in the footer (method + family size). See the Multiple-comparison adjustment section for when this is and is not appropriate.

show_columns

Character vector of tokens selecting the per-coefficient columns and their display order. Accepts atomic tokens ("b", "se", "ci", "t", "p", "beta", "n", "n_events", "pd" (probability of direction, Bayesian fits only), "rhat" / "ess_bulk" / "ess_tail" / "mcse" (per-coefficient sampler diagnostics and the Monte Carlo standard error of the displayed posterior median, all-Bayesian tables only), "ame", "ame_se", "ame_ci", "ame_p", "rmst" + "rmst_se" / "rmst_ci" / "rmst_p", "risk_diff" + its ⁠_se⁠ / ⁠_ci⁠ / ⁠_p⁠ companions, "partial_f2" + "partial_f2_ci", "partial_eta2" + "partial_eta2_ci", "partial_omega2" + "partial_omega2_ci", "partial_chi2", "r2" / "adj_r2" (per-fit variance explained, lm screens)) and group tokens ("all_b", "all_b_compact", "all_b_full", "all_beta", "all_ame", "all_ame_compact", "all_f2", "all_eta2", "all_omega2"). See Vocabulary tokens in the details for the full enumeration. Default NULL selects a context-aware layout: "all_b" (single model) or "all_b_compact" (multi-model, and the single-multinomial outcome-as-columns layout, which has the same width pressure – restore CIs with atomic tokens, e.g. c("b", "se", "ci", "p")). The "p" token is always the B / beta p-value; for the AME-specific p-value use "ame_p".

keep

Character vector of regexes. Only coefficient rows whose term name (as in stats::coef() – e.g. "wt", "cyl6", "factor(cyl)8") matches at least one pattern are kept. Mutually exclusive with drop. Filtering is a display choice; p_adjust runs against the full coefficient family before filtering. Default NULL (no filter).

drop

Character vector of regexes. Coefficient rows matching any pattern are removed. Mutually exclusive with keep. Default NULL.

show_intercept

Whether to display the intercept row. Default TRUE (APA convention). Hide via FALSE.

show_thresholds

For ordinal cumulative-link models (MASS::polr, ordinal::clm), whether to display the estimated category thresholds (cut-points) as a subordinate "Thresholds" block of rows below the predictors, carrying B / SE / CI / p like the predictor rows. Default TRUE. FALSE collapses them to a compact one-line footer note instead. Thresholds are reported on the log-odds (B) scale and are never exponentiated (under exponentiate = TRUE their rows stay on the log-odds scale). Has no effect on non-ordinal models, and the rows are shown only when a coefficient column ("b"/"beta") is in show_columns.

show_components

For models with secondary components, whether to display them as labelled subordinate blocks of rows below the primary (count / conditional / location) coefficients. Default TRUE:

  • pscl::zeroinfl and glmmTMB(ziformula = ): a Zero-inflation block – the model for the probability of a structural (excess) zero.

  • pscl::hurdle: a ⁠Zero hurdle⁠ block – the model for the probability of a nonzero count (note the opposite direction vs zero-inflation; the footer names each block's meaning).

  • glmmTMB(dispformula = ): a Dispersion block (only when dispersion was actually modelled; log scale, never exponentiated).

  • ⁠ordinal::clm(scale = ~)⁠: a ⁠Scale effects⁠ block – the covariate effects on the log standard deviation of the latent response. Never exponentiated: their exponential is a ratio of latent standard deviations, not an odds ratio.

Component rows carry full Wald inference (B / SE / z / p / CI), join the p_adjust family, and take significance stars. Under exponentiate = TRUE a component is exponentiated only when its link makes the result an odds ratio (the logit zero components); probit / cauchit / cloglog zero links and count-type hurdle zero parts stay on the link scale, disclosed in the footer. FALSE omits the blocks (the title still names the model type).

intercept_position

Where to place the intercept when shown. "first" (default, APA) or "last" (Stata-style, intercept just above the fit-stats footer). Ignored when show_intercept = FALSE (with spicy_ignored_arg warning).

factor_layout

Layout of factor predictors. Applies to any categorical predictor – factor, ordered, character, or logical (R coerces the latter two to factors at fit time). Two options:

  • "grouped" (default): the variable name on its own header row ending with : (e.g., ⁠education:⁠); each level follows as an indented sub-row with the bare level name. APA convention.

  • "flat": each non-reference dummy is one row with the ⁠<variable><level>⁠ form (e.g., educationUpper); no header, no indent. Econometrics convention.

reference_style

Rendering of factor reference levels. Four modes, distinguishing WHERE the reference information is exposed (in a row, inline, in the footer, or nowhere):

  • "row" (default): explicit row Female (ref.) with en-dashes in all stat columns (NEJM / BMJ clinical convention). reference_label controls the suffix.

  • "annotation": the row is dropped and the reference is shown inline. Under factor_layout = "grouped" the factor header reads ⁠education: [ref: Lower]⁠; under factor_layout = "flat" the marker ⁠[vs Lower]⁠ is attached to the first non-reference dummy of each factor (subsequent dummies inherit the same reference).

  • "footer": the row is dropped and a single line ⁠Reference categories: education = Lower; sex = Female.⁠ is added to the footer note. SAS ⁠PROC LOGISTIC⁠ / SPSS "Categorical Variables Codings" convention. Best for publication-grade dense multi-factor tables.

  • "none": the row is dropped and no reference information is displayed anywhere. The user is responsible for stating the reference convention elsewhere (article text, table caption). Under factor_layout = "flat", an informational message is emitted to flag the silent omission.

Ordered factors with AME: under R's default contr.poly, ordered factors have B coefficients named .L / .Q / .C (orthogonal polynomial trends) which have no per-level reference semantics. When "ame" is in show_columns, however, the AME block is per-level contrasts against levels()[1]. A synthetic reference row anchored on levels()[1] is therefore emitted so the reader sees the AME baseline explicitly, with the same reference_style handling as plain treatment-coded factors. The ⁠[vs <ref>]⁠ annotation in "annotation" mode is attached to the first AME row, not to the polynomial-trend rows.

reference_label

Suffix shown after the reference level in reference_style = "row" mode. Default "(ref.)". Ignored by the other three modes (which use structural English wording – "ref:", "vs", "Reference categories:").

show_fit_stats

Character vector of tokens for the model-level rows below the coefficients; row order follows token order. NULL (default) resolves class-aware:

  • lm, estimatr::lm_robust(): c("nobs", "r2", "adj_r2"). estimatr::iv_robust() gets "nobs" alone – the 2SLS R-squared is not the classical one.

  • glm, ordinal polr / clm: c("nobs", "pseudo_r2_mcfadden", "pseudo_r2_nagelkerke", "aic") (McFadden = Stata ologit default, Nagelkerke = SPSS PLUM).

  • mixed lm + glm: the union of the two (the renderer en-dashes per cell the stat not defined for a given model class).

Under nested = TRUE the default is extended with the class-appropriate change-stat tokens (e.g. "r2_change", "f_change" for lm). See Vocabulary tokens (show_fit_stats subsection) and Hierarchical (nested) model comparison in the details for the full vocabulary.

fit_stats_layout

Layout of the fit-stat values (n, R^2, AIC, ...) within each model's column group. Two options:

  • "first_col" (default): the value is placed in the FIRST numeric sub-column of each model (typically B); the model's remaining sub-columns (SE, LL, UL, p, ...) are left empty for that row. The APA Manual 7 Table 7.13 layout.

  • "merged": the model's numeric sub-columns are merged into a single wide cell containing the fit-stat value, centred under the model spanner. Stata esttab layout / Econometrica and AER journal convention. Resolves the mixed-precision look of "first_col" (an integer n row sharing the B column with two-decimal coefficients).

Cell merging is supported by excel, flextable, and word (via flextable). gt, tinytable, clipboard, and default (console) always render in "first_col" mode regardless of this setting:

  • gt lacks a native row-spanning cell-merge API (tab_spanner covers columns, not row-cell ranges).

  • tinytable's style_tt(colspan = N) emits HTML colspan only on header rows, not on body cells.

  • clipboard ships TSV plaintext.

  • default ships fixed-width ASCII.

Decimal alignment of every numeric column is preserved in both modes: the B column decimal-aligns its coefficient values plus any fit-stat value(s) in "first_col" mode (native primitives handle the mixed-precision case), and trivially decimal-aligns in "merged" mode (the fit-stat values move out of the B column into the merged cell).

show_re

Logical. TRUE (default) renders the random-effects variance components of a mixed-effects fit (lmer, glmer, glmmTMB, lme) as a subordinate "Random effects" block of table rows below the fixed effects: one row per standard deviation / correlation per grouping factor, plus the residual, each with its estimate, SE, and CI in the shared coefficient columns. The group sizes (N (groups)) and the ICC render as fit-statistic rows; the footer reports the estimation method (REML / ML) and the likelihood-ratio test of the whole random part against the no-random-effects model, with the boundary-corrected chi-bar-squared p-value (Self & Liang 1987; Stram & Lee 1994). Variance-component rows deliberately carry no per-row p-value: a Wald test of a variance is invalid at the boundary of the parameter space, and no reporting guideline requests one (see the Mixed-effects models section of the Publication-ready regression tables article). FALSE suppresses the block. No effect on fits without random effects (lm, glm, coxph, ...).

re_scale

One of "sd" (default) or "variance". Controls the display scale of the random-effects rows:

  • "sd": report the random-effect standard deviation \sigma (Gelman 2005, Technometrics: "directly interpretable as the size of the variation across groups"). Standard error and CI converted via the Delta method: SE(\sigma) = SE(\sigma^2) / (2\sigma); CI(\sigma) = \sqrt{CI(\sigma^2)}.

  • "variance": report \sigma^2 (the canonical internal scale; SE and CI come straight from the Hessian / nlme::intervals() / glmmTMB::confint() without rescaling).

Correlation rows (\rho) are unitless and pass through either way.

re_columns

Character vector. Subset of c("est", "se", "ci") controlling which cells of the random-effects rows are displayed ("est" is mandatory); deselected SE / CI cells render as an en-dash on those rows only. Useful for slimming output (re_columns = "est") or for journals that want only standard errors (re_columns = c("est", "se")). Display-only: the underlying data (broom::tidy(), as_structured()) always carries the full SE + CI.

Note. Under the default re_ci = "wald" the SE and CI are Wald (⁠est ± z * SE⁠, clamped at 0 for variances). Wald can be optimistic near the variance boundary (Self & Liang 1987 chi-bar-squared); request boundary-respecting profile-likelihood intervals with re_ci = "profile" when robustness is critical. See the Mixed-effects models section of the Publication-ready regression tables article.

For lmer / glmer fits these SEs come from merDeriv, whose cost grows superlinearly with the number of observations (about a minute at n of roughly 2,700). Above options("spicy.re_se_max_n") (default 1000) they are skipped: the rows keep their estimates, the SE / CI cells render as en-dashes, a table note states the omission, and a spicy_caveat warning points here. Raise the cap (e.g. options(spicy.re_se_max_n = Inf)) to force the computation, or test the random terms with re_test = "lrt".

re_test

One of "none" (default), "lrt", or "rlrt". Opt-in per-term significance test for the random-effect variance components, filling the otherwise-empty p column of the Random effects rows. Never a Wald test (invalid at the boundary sigma = 0):

  • "lrt": likelihood-ratio test of each random term vs the model refitted without it (the term's variance plus its covariances with the other terms of its bar), referred to the boundary-corrected chi-bar-squared mixture ⁠0.5 chi2(q-1) + 0.5 chi2(q)⁠ (Self & Liang 1987; Stram & Lee 1994). The reduction scheme matches lmerTest::ranova(); the mixture reference makes the p-value exact-asymptotic rather than conservative. A bar's intercept is tested only when it is the bar's single term. Supported: lmer, glmer, glmmTMB, and lme with a simple random = ~ terms | group structure.

  • "rlrt": exact restricted likelihood-ratio test with a simulated finite-sample null (RLRsim::exactRLRT(); Crainiceanu & Ruppert 2004). Only defined for a Gaussian lmer / lme fit with a single variance component.

The test statistic and df stay out of the displayed t/z column (they are chi-square-scale, not t/z) but are carried in broom::tidy() (test_type "chibar2" / "rlrt"). The whole-block LR test in the footer is unaffected. Correlation and residual rows are never tested (a correlation is tested jointly with its slope; the residual has no zero-variance null). Refits happen once per random term: expect a noticeable cost on large models. For structures outside these two routes (e.g. multiple variance components needing a finite-sample null), pbkrtest::PBmodcomp() offers a parametric-bootstrap LR test on the same nested pair of fits (Halekoh & Hojsgaard 2014) – run it directly and report its p-value alongside the table.

re_ci

One of "wald" (default) or "profile". Uncertainty route for the random-effect variance-component rows of lmer / glmer fits:

  • "wald": SE and symmetric CI from the observed information (merDeriv), subject to the options("spicy.re_se_max_n") size cap (see re_columns).

  • "profile": profile-likelihood CIs via confint(fit, method = "profile") – the route lme4 itself documents and defaults to. The intervals are asymmetric and respect the boundary at 0; no SE is shown (lme4's position: a symmetric SE misdescribes the skewed sampling distribution of a variance). Sidesteps the size cap entirely (roughly two seconds per variance parameter even at n ~ 7,000; glmer profiles cost more). The footer discloses the method, and likelihood intervals transform exactly, so re_scale = "variance" shows the profile CI of the variance itself.

glmmTMB and nlme::lme fits keep their engine-native CIs (TMB's sdreport; nlme's apVar) and refuse "profile".

model_labels

Per-model labels used as the column-group spanner above each model's sub-columns (console + gt / flextable / tinytable / Excel / Word renderers). NULL (default) resolves automatically; see Multi-model semantics for the full rule. A character vector of length length(models) overrides. Refused (error) for a single multinomial model: there the column groups are outcome categories, relabelled via outcome_labels.

outcome_labels

Optional Outcome body row override. NULL (default) hides the row entirely – under the multi-model spanner the DV is already visible above the data. A character vector of length length(models) forces an explicit Outcome row with those values (the spanner stays as "Model 1, ..." unless model_labels is also supplied). FALSE also suppresses the row. Single multinomial model (outcome-as-columns layout): repurposed as the category spanner override – a character vector with one label per non-reference outcome category (unique, in model order), e.g. c("Student vs Employed", ...); the reference category's AME-only group (when displayed) keeps its own name, and FALSE is a no-op (there is no Outcome row to suppress).

stars

Significance asterisks. FALSE (default, APA 7 Section 6.46) – no stars. TRUE – APA cutoffs c("*" = 0.05, "**" = 0.01, "***" = 0.001). A named numeric vector specifies custom thresholds, e.g. c("+" = 0.10, "*" = 0.05, "**" = 0.01, "***" = 0.001). With output = "excel" a marked estimate is written as text ("64.63***" is not a number); the cells with no marker stay numeric.

nested

Whether to inject pairwise change-statistic rows for adjacent models (M2 vs M1, M3 vs M2, ...). FALSE (default) – pure side-by-side display. TRUE – requires identical nobs and identical response variable across all models. See Hierarchical (nested) model comparison.

digits

Decimal places for general numeric tokens (b, beta, se, ci, t, f_change, lrt_change, deviance, deviance_change, ame, ame_se, weighted_nobs). Default 2L.

p_digits

Decimal places for p-values (p, ame_p, p_change). APA-strict: leading zero stripped, ⁠<.001⁠ (or ⁠<.0001⁠ etc. depending on p_digits) for small values. Default 3L.

effect_size_digits

Decimals for per-coefficient effect sizes (partial_f2, partial_eta2, partial_omega2). Default 2L.

fit_digits

Decimals for variance-explained / model-level effect-size fit stats (r2, adj_r2, r2_change, adj_r2_change, omega2, f2, f2_change, sigma, rmse). Default 2L.

ic_digits

Decimals for information criteria (AIC, AICc, BIC, and their ⁠_change⁠ form). Default 1L.

decimal_mark

Decimal mark used in numeric display. "." (default) or "," (European convention). When "," is used, the CI bracket separator switches to "; " automatically to avoid "0,18 [0,07, 0,30]" ambiguity. With output = "excel" a numeric cell is displayed by Excel with the viewer's locale separator, which a file cannot set, so a non-default mark makes the body go out as pre-formatted text (the default "." keeps every cell a real number).

align

Numeric column alignment. "decimal" (default) – pre-pad cells so decimal marks line up vertically (publication-style). For CI cells (⁠[LL, UL]⁠) the left bracket, the LL decimal point, the comma separator, the UL decimal point, and the right bracket are independently aligned across rows. "center" and "right" apply the same alignment to every numeric column.

padding

Non-negative integer giving the extra characters added to each data column's auto-computed width when the default print method renders the table. Default 0L (compact – fits more models in the same console / page width). Use 2L (Stata-like) or 4L for a more spacious look. Headers stay centered above the data region regardless of padding.

labels

Named character vector overriding per-coefficient row labels. Names are coefficient term names (from stats::terms()); values are the displayed labels. E.g. c("age" = "Age (years)", "sexM" = "Male (vs Female)"). Default NULL (use raw term names).

The names are model identifiers, not display text: they are the term and coefficient names R itself produces, and they keep that form regardless of the language the table is rendered in. A name that is neither a term label nor a coefficient name is rejected, so the header of a subordinate block – ⁠Random effects⁠, Thresholds, Zero-inflation, ... – cannot be relabelled here; those headers are set by the package. Classes that carry no terms component at all (nls(), sampleSelection::selection()) are keyed on their coefficient (or parameter) names alone.

With several models, a name is checked against the union of their terms: a key naming a term only one model has is legal, and the label lands on the rows where that term exists.

title, note

Override or suppress the auto-built caption / methodological footer. Three modes per argument:

  • NULL (default): the package builds the standard caption ("Linear regression on ⁠<DV>⁠" / "Hierarchical linear regression on ⁠<DV>⁠" / ...) and a methodological note (VCV type, p-adjust method, reference categories, ...).

  • FALSE: the corresponding banner row is omitted from every output engine. Use when the surrounding manuscript provides its own caption / note.

  • character string (length 1): replaces the auto-built text verbatim. The renderer applies no APA formatting on top – supply the exact string you want displayed (multi- line notes accepted via embedded "\n").

Validation messages, the spanner row, and the in-body change- stat rows are not affected – they belong to the table structure, not to the banner.

output

Output type. "default" (a printable spicy_regression_table); "data.frame" / "long" (raw data); "gt" / "flextable" / "tinytable" (rich-format tables); "excel" (writes to excel_path); "clipboard" (copies to system clipboard); "word" (writes flextable to word_path).

excel_path

File path for output = "excel". Default NULL (required when output = "excel").

excel_sheet

Sheet name when writing to Excel. NULL (the default) uses "Regression".

clipboard_delim

Field delimiter for output = "clipboard". Default "\t" (tab-separated, pastes cleanly into Excel / Google Sheets / Word). The clipboard payload mirrors the Excel layout (title row, spanner row, header, body, footer note) but is plain text – horizontal rules, cell merging, decimal alignment and the monospace font cannot be encoded in delimited text and are therefore absent from the paste. A cell holding the delimiter itself (a label with a comma under clipboard_delim = ",", or any number under decimal_mark = ","), a double quote or a line break is quoted RFC 4180-style, so the grid survives whatever delimiter you choose.

Paste behaviour by target:

  • Excel / Google Sheets: numerics are auto-detected and right-aligned; text cells stay left-aligned. (P-values such as .005 get re-parsed as 0.005 by Excel's auto-format – to preserve the APA leading-zero-dropped display, prefer output = "excel".)

  • Word: the paste is converted to a Word table; all cells start left-aligned. Apply a Table Style (Insert > Table > Design) for APA-style borders, and set right-alignment on numeric columns (Layout > Align Right). For a self-contained Word file with borders and alignment pre-applied, use output = "word" instead.

word_path

File path for output = "word". Default NULL (required when output = "word"). The Word table inherits the flextable styling (Calibri font, APA borders, decimal-aligned numerics) and adds Word-specific features: an auto-numbered caption ("Table 1: ...", "Table 2: ...") via Word's SEQ field so multiple table_regression() calls in one document number consecutively; a re-printed header row on each page break; row split prevention so a single coefficient row never wraps across two pages; and an APA-styled note line (⁠*Note.*⁠ italic prefix per APA Manual 7 §7.14).

R Markdown / Quarto: for embedded use, prefer output = "flextable" (returns the flextable object that knits to docx/HTML/PDF natively). output = "word" writes a standalone .docx file, suited to scripted exports rather than chunk-level rendering.

word_template

Optional path to a custom .docx file used as the template for output = "word". The template's header, footer, page size, margins, and named styles ("Table Caption" in particular) are honoured; the table is appended to the template body. Useful for institutional templates with pre-set headers ("APA Style", "Manuscript Submission Template", etc.). Default NULL (uses flextable's stock template).

Customising the caption appearance: the table caption is tagged with the Word named style "Table Caption". The visual rendering (italic / bold / colour / font) follows whatever that style is set to in the docx template. The stock Word template renders "Table Caption" in italic — the APA Manual 7 §7.10 condensed convention. For a different appearance (Nature-style bold non-italic, APA-strict 2-line bold-number / italic-title, etc.), edit the "Table Caption" style in a docx template and pass it via word_template = "your_template.docx". Style-based delegation keeps the rendered caption consistent with the surrounding document and lets editorial conventions (Nature, APA-strict, journal-specific) be applied without modifying the call site.

style

A journal style: a theme name ("jama", "nejm", "lancet", "annals", "apa", "aer"), a spicy_style() object, or NULL (the default). A style only changes DEFAULTS – any argument you pass explicitly wins over it. Set options(spicy.style = ) for document-wide scope. A theme covers numeric formatting conformity only, not full editorial conformity; ?spicy_style lists the exact rules each one encodes and the official document they come from. An unknown name is an error.

Value

A spicy_regression_table object (a data.frame subclass with classes c("spicy_regression_table", "spicy_table", "data.frame")) when output = "default". The result carries rendering attributes (title, note, align, padding) and provenance attributes (outcome, model_ids) consumed by the print method and the broom methods. For other output values, returns the format-specific object (gt_tbl, flextable, tinytable, data.frame, tbl_df, or invisible(x) for side-effect outputs).

Vocabulary tokens

Two vector arguments – show_columns and show_fit_stats – accept named tokens that select what to display and in what order. All tokens are lowercase (snake_case for compound tokens). Group tokens ("all_b", "all_ame", ...) expand to a fixed vector of atomic tokens; see show_columns below.

show_columns – per-coefficient columns

Each token = one displayed column.

Group tokens (presets) expand to a fixed atomic vector before validation:

Mix groups and atomic tokens: show_columns = c("all_b", "ame", "ame_p"). Duplicates after expansion are deduplicated; the order of tokens controls the order of the displayed columns. If standardized != "none" and "beta" is not already requested, it is auto-injected after "b". Asking for "beta" while standardized = "none" raises spicy_invalid_input.

Default (show_columns = NULL) is context-aware: "all_b" for a single model (APA-7 Section 6.46 publication layout), "all_b_compact" for two or more models (CI dropped to fit the side-by-side layout; restore it explicitly when needed).

show_fit_stats – model-level rows below the coefficients

Default (resolved when NULL) is class-aware: lm and lm_robust fits get c("nobs", "r2", "adj_r2"); glm and ordinal polr / clm fits get c("nobs", "pseudo_r2_mcfadden", "pseudo_r2_nagelkerke", "aic"); mixed lm + glm sets union both groups (the renderer per-row en-dashes the inappropriate cell); Cox fits get c("nobs", "n_events", "aic"); GEE fits get c("nobs", "n_groups", "max_cluster_size") (the cluster structure, Stata xtgee-header style). When nested = TRUE, the class-aware default is extended with change tokens (c("r2_change", "f_change", "p_change") for lm, c("lrt_change", "p_change") for glm). The order of tokens in show_fit_stats controls the order of the rows.

Multi-model semantics

Pass a single fit or a list() of fits. Multi-model layout draws a centred spanner label above each model's sub-columns:

Duplicate explicit names in the list are rejected (spicy_invalid_input) – they would silently collide in the internal model_id key. So is a name that collides with the auto-filled label of another slot (list("Model 2" = m1, m2)): two models under one spanner cannot be told apart in the table, nor addressed by inline(). Rename the model, or pass model_labels.

Inference and standard errors

vcov selects the variance-covariance estimator:

For multi-model use, both vcov and cluster accept a single value (recycled to all models) or a list (one per model). The same fit can appear several times with different estimators to compare standard errors side-by-side.

Inferential regimes for lm / glm (B and AME share the same regime):

For other classes the t-vs-z axis follows the estimator's native reference distribution (e.g. glm, glmmTMB, survival, ordinal, betareg, mlogit use z; lm, lme, lmer, rms::ols use t).

Multinomial models: outcome categories as columns

A single nnet::multinom model renders the publication layout: predictors as rows, one column group per non-reference outcome category (spanner = category name), so a predictor's effect can be compared across equations along its row. The mandatory footer note ⁠Reference outcome: <level>.⁠ names the base category (Stata's "base outcome" line); outcome_labels relabels the spanners (e.g. "Student vs Employed"). Fit statistics print once, under the first group. When per-category AMEs are requested the reference category appears as a last, AME-only group (its coefficient cells are empty – the reference has no equation; its AMEs complete the zero-sum across categories). The default show_columns compacts to B / SE / p exactly like multi-model tables; request atomic tokens to restore CIs. Multi-model and nested = TRUE multinomial tables keep the one-row-per-(category, predictor) layout – categories-within-models would need two spanner levels. tidy() and output = "long" always return the long form ("<category>: <term>" rows), whatever the display; as_structured() mirrors the displayed table (one column set per category), as it does for every other layout.

Robust SE availability by model class

Not every estimator is defined for every class. A robust vcov the class cannot honour fails fast with spicy_unsupported_vcov – never a silent model-based result under a robust label:

lm, glm, MASS::glm.nb

all of classical, ⁠HC*⁠, ⁠CR*⁠, bootstrap, jackknife. lm additionally takes "CR1S" – the full Stata ⁠regress, vce(cluster)⁠ convention: the CR1S small-sample scaling (equal to sandwich::vcovCL() with type = "HC1") with t(G - 1) inference, G the number of clusters. Use it to reproduce Stata tables; "CR2" (Bell-McCaffrey, Satterthwaite df) remains the recommended modern choice. "CR1S" is refused for glm with an explanation: Stata's ML commands use a different convention (G/(G-1) scaling, z), so the label would not match Stata output there.

mlogit

classical + ⁠CR*⁠ only (cluster at the choice-situation level) – sandwich::vcovHC() mis-scales the sandwich for mlogit's per-choice-situation scores, so ⁠HC*⁠ is refused.

lmer, lme, coxph, survreg, mgcv::gam/bam, polr, clm, betareg, nnet::multinom, pscl::zeroinfl/hurdle, rms (ols/lrm/cph/Glm)

classical + ⁠CR*⁠ only – ⁠HC*⁠ and the resamplers (which refit lm/glm) are not defined for these. clm with a scale / nominal (partial-PO) component is classical only. multinom needs sandwich >= 3.1-2 (which added its estfun() method); its cluster is one entry per observation. For the two-part pscl fits the cluster sandwich covers both components (count and zero) at once.

quantreg::rq

its own estimator family, not the sandwich vocabulary: classical (= "nid", quantreg's large-sample default), "iid", "ker", "rank" (intervals only) and a native "bootstrap", clustered via the wild gradient bootstrap. ⁠HC*⁠, ⁠CR*⁠ and jackknife are refused, each with its own reason – see the vcov argument.

geepack::geeglm

no spicy-side estimator at all: GEE inference is robust by construction, so the fit's own sandwich (or jackknife) standard errors – chosen by geeglm's ⁠std.err =⁠ option, clustered on its ⁠id =⁠ – are the displayed inference. ⁠HC*⁠ / ⁠CR*⁠ and cluster are refused with a pointer to those fit options.

survey::svyglm

classical only, on principle: the design-based Taylor / replicate variance is already the robust variance for the declared design, and clustering belongs in the design itself (survey::svydesign(ids = )), not in the table call.

estimatr (lm_robust/iv_robust) and fixest

these fits carry the variance they were computed with – estimatr's ⁠se_type =⁠, and whatever fixest's own vcov interface produced – and spicy never overwrites it. A vcov request is refused; set the estimator through the fitting package instead (⁠se_type =⁠ for estimatr; ⁠vcov =⁠ at estimation or in summary() for fixest).

Other classes (glmer, glmmTMB, rstanarm/brms, ...)

classical (model-based) only (clubSandwich has no working backend for glmer / glmmTMB).

Cluster-robust backends differ by class but are each cross-validated to the field-standard oracle: lm/glm/lmer/lme use clubSandwich (CR2 = Bell-McCaffrey, with Satterthwaite df for lm/lme/lmer); coxph/cph use the Lin-Wei grouped-dfbeta sandwich (identical to coxph(..., cluster=)); survreg/gam/polr/clm/betareg/mlogit/multinom and the pscl two-part fits use sandwich::vcovCL(); rms fits use rms::robcov() (which needs the fit's ⁠x = TRUE, y = TRUE⁠). These single cluster sandwiches have no CR0-CR3 bias-reduction variants, so the requested ⁠CR*⁠ maps to the one available estimator. cluster length is one entry per observation, except mlogit (one per choice situation) and censored coxph (one per subject).

How to specify cluster

Three accepted forms, in order of preference:

  1. Formula – cluster = ~region (or cluster = ~region:year for the interaction of two variables). The variables are looked up in model.frame(fit) first, then in the original data argument captured by the fit. Recommended: independent of the dataset's name, composable for multi-way clustering, consistent with sandwich::vcovCL() / clubSandwich::vcovCR().

  2. String – cluster = "region". A single column name resolved the same way as the formula. Convenient but cannot express interactions.

  3. Vector – cluster = df$region. An atomic vector of length nobs(fit). Use this when the cluster key is derived on the fly (cluster = interaction(df$region, df$year), cluster = as.integer(format(df$date, "%Y"))), comes from a different dataset with matching row order, or is otherwise not a column of the model's data.

A bare unquoted name (cluster = region) is not column selection: it is evaluated as an ordinary R variable. If an object region exists in the calling environment, its value is used as the vector form above; if it does not (the typical "unquoted column name" intent), the call fails with a migration error pointing at ~region / "region". Bare-name column selection is deliberately unsupported – it would require non-standard evaluation magic that breaks under programmatic use (function wrapping, dynamic column choice, loops).

For multi-model use, mix forms freely: cluster = list(~region, "region", df$region).

Bayesian fits (rstanarm / brms)

stanreg and brmsfit models are summarized from their posterior draws: B is the posterior median, SE the posterior MAD SD (the scaled median absolute deviation – the pairing of Regression and Other Stories and rstanarm's own print; equal to the posterior SD for a normal posterior), and the interval an equal-tailed credible interval (header ⁠95% CrI⁠; ci_method = "hdi" opts into the highest-density interval, header ⁠95% HDI⁠). There is no p-value or t-statistic; group presets ("all_b", ...) expand without them, and the opt-in "pd" column reports the probability of direction with a footer definition (Makowski et al. 2019). Under exponentiate = TRUE all quantities are computed from the exponentiated draws directly (the SE is the posterior MAD SD on the ratio scale, not a delta-method approximation). The AME columns are equally draws-native: avg_slopes() computes the marginal effects per posterior draw, and the table reports their posterior median, MAD SD and credible interval (equal-tailed or HDI per ci_method); "ame_p" has nothing to fill and follows the p-column policy (dropped from presets, refused as an atomic token, dashed in mixed tables). Standardized betas follow the same draws logic: "posthoc" / "basic" / "smart" are exact affine rescales of the link-scale summaries (the Gaussian family divides by SD(y); every other family standardizes predictors only, the frequentist glm convention); "refit" and "pseudo" are refused.

Fit statistics: "r2_bayes" (in the Bayesian default) plus the opt-in "elpd_loo" / "looic" / "waic", whose standard errors are disclosed in the footer; unreliable estimates (PSIS-LOO Pareto k above the sample-size-specific threshold \min(1 - 1/\log_{10} S, 0.7) of Vehtari et al. 2024 – the same bound loo::loo() prints – and WAIC p_waic > 0.4) add a footer caveat instead of being silenced. Every table is backed by an automatic sampler-diagnostics guard (R-hat >= 1.01, ESS below 100 per chain – floored at 400 so fewer chains never weaken the bar –, divergent transitions, E-BFMI < 0.2, per Vehtari et al. 2021): problems add a footer line and raise a warning classed spicy_bayes_diagnostics (nested under spicy_caveat, so withCallingHandlers(spicy_bayes_diagnostics = ...) mutes the guard selectively); clean fits print nothing. Per-coefficient "rhat" / "ess_bulk" / "ess_tail" columns are available in all-Bayesian tables, as is "mcse" – the Monte Carlo standard error of the displayed posterior median, the criterion for how many digits a table can honestly show (a displayed digit is Monte-Carlo stable when twice the MCSE stays below it; Gelman, Vehtari, McElreath et al. 2026, sec. 11.6). Under exponentiate = TRUE the MCSE is recomputed on the exponentiated draws.

Refused on principle (classed error, never a silent fallback): likelihood-based fit statistics ("aic", pseudo-R², ...), p_adjust, robust / cluster vcov (model the clustering with group-level terms instead), ci_method = "profile" / "boot_percentile", and non-MCMC fits (algorithm = "meanfield" / "optimizing" – refit with algorithm = "sampling"). See the Bayesian regression tables article.

Hierarchical (nested) model comparison

nested = TRUE adds per-pair change statistics as in-table rows (APA Table 7.13 / Stata esttab / SPSS Model Summary convention). Each adjacent pair (M2 vs M1, M3 vs M2, ...) contributes one column of change stats; the FIRST model column gets en-dashes (no previous model to compare to). Validation requires identical nobs and identical response variable across all models.

Structural refusals. Four hierarchies carry no valid change statistic whatever their nobs and response, and are refused with spicy_invalid_input rather than rendered with empty or meaningless cells:

In the first three, nested = FALSE renders the models side by side instead.

Default change tokens auto-injected when show_fit_stats is NULL:

To customise, pass the change tokens directly to show_fit_stats. The variance-explained change tokens ("r2_change", "adj_r2_change", "f_change", "f2_change") raise spicy_invalid_input on any hierarchy whose nested comparison is a likelihood-ratio test – glm, and equally every other likelihood class and the mixed-effects families: that comparison reports a chi-square, not a variance-explained change, and the message points at "lrt_change" + "p_change". A quantile hierarchy refuses them too, and "lrt_change" with them, pointing at "f_change" + "p_change" instead. lm and nls keep the least-squares tokens.

Standardised coefficients

standardized controls the method when "beta" is in show_columns:

Interactions and transforms. Under "refit", an interaction's \beta is the coefficient of the product of the z-scored components – the recommended treatment for such models (Cohen et al. 2003 Section 7.7; Aiken & West 1991; Friedrich 1982). Under "posthoc", "basic", and "smart", the product / transformed design column is treated as a single numeric column and scaled by its own SD (under "smart", by 2 * SD when the product column is continuous; a binary product column – e.g. a binary-by-binary interaction – stays unscaled like any binary input) – the convention of SPSS beta, Stata ⁠regress, beta⁠, SAS PROC REG STB, and lm.beta::lm.beta(), identical to effectsize::standardize_parameters(method = "basic") on those columns. The two conventions differ whenever the components are correlated; a spicy_caveat warns and the footer names the convention actually used. "refit" declines formulas with inline transforms (log(x), poly(), factor() written in the formula): the model frame's evaluated columns cannot be re-evaluated on z-scored data, so spicy falls back to "posthoc" with a warning (effectsize's refit instead standardizes before the transform – a different estimand). Pre-build transformed columns in data for an exact refit.

Multiple-comparison adjustment

Adjusting the p-values of all coefficients of a single regression model is not the standard convention. Each coefficient tests a distinct hypothesis on a distinct predictor – not the situation multiple-testing procedures were designed for (Rothman 1990; Greenland 2017; APA Manual 7 Section 6.46; Harrell Regression Modeling Strategies Section 5.4; Gelman, Hill & Yajima 2012). Hence the default p_adjust = "none".

Adjustment is appropriate for: mass screening with no prior hypothesis (typically "BH" / FDR), pre-registered multi-endpoint confirmatory designs (typically "holm"), or when a journal / SAP explicitly requests it.

The adjustment runs before any keep / drop filtering, so the family is the model's full coefficient set (intercept and reference rows excluded), not the displayed subset – filtering is a display choice and must not change the inferential family.

Tied event times and the survival estimands

A Cox model has to decide what to do when several events share the same time, and survival::coxph() decides with ties = "efron" by default. The absolute estimands do not make a second decision:

The baseline hazard behind the RMST-difference and risk-difference columns follows the tie-handling convention of the fit, exactly as survival::survfit() and survival::basehaz() do. A ties = "breslow" fit gives a Breslow baseline; the default Efron fit gives an Efron baseline. One rule, so the hazard-ratio column and the dRMST column always come from the same likelihood. To change the convention, change the fit.

When the choice matters is not simply "when there are many ties". What governs it is the fraction of tied events within the risk set, together with the effect size and the spread of the covariates: a large risk set dilutes a tied block, a small one – fine strata, matched or nested case-control designs – does not. A file with hundreds of tied days can be insensitive to the convention while a small stratified one with a handful of ties is not.

And Efron is the better default, not the correct answer: the approximations are all biased at finite sample size and agree asymptotically, and Efron's virtue is a much smaller small-sample bias (Hertz-Picciotto & Rockhill 1997). Efron himself, quoting Peto, put it as "it probably doesn't make much difference" (Efron 1977).

Output formats and broom integration

output selects the return type:

broom::tidy() returns a long tibble with one row per ⁠(model_id, term, estimate_type)⁠ and broom-canonical column names (estimate, std.error, conf.low, conf.high, statistic, p.value). broom::glance() returns one row per model with the model-level statistics; df.residual is kept numeric so cluster-robust Satterthwaite df is preserved.

Global options

Weights

No weights argument: weights are a property of the fit (extracted via stats::weights()). Pass them when fitting: lm(y ~ x, data = df, weights = w). All downstream computations (vcov, AME, standardisation, weighted_nobs) extract them automatically.

Internationalisation

Output is in English. Override user-facing strings via reference_label, model_labels, outcome_labels, and labels. The title and footer are post-processable via attr(result, "title") and attr(result, "note").

Classed conditions

Every error and warning emitted by table_regression() carries a classed condition for programmatic dispatch via tryCatch() or withCallingHandlers(). Errors inherit from spicy_error (root); warnings from spicy_warning. Specific leaves used by this function include spicy_invalid_input, spicy_invalid_data, spicy_unsupported, spicy_unsupported_vcov, spicy_unsupported_standardized, spicy_missing_pkg, spicy_missing_column, spicy_ignored_arg, spicy_caveat, spicy_fallback. See spicy for the full taxonomy.

References

APA Manual 7 (American Psychological Association, 2020), Tables 7.13-7.15.

Aiken, L.S. & West, S.G. (1991). Multiple regression: Testing and interpreting interactions.

Cohen, J., Cohen, P., West, S.G., & Aiken, L.S. (2003). Applied multiple regression / correlation analysis for the behavioral sciences (3rd ed.). Lawrence Erlbaum.

Davison, A.C. & Hinkley, D.V. (1997). Bootstrap methods and their application. Cambridge University Press.

Efron, B. (1977). The efficiency of Cox's likelihood function for censored data. Journal of the American Statistical Association, 72(359), 557-565.

Friedrich, R.J. (1982). In defense of multiplicative terms in multiple regression equations. American Journal of Political Science, 26(4), 797-833.

Hertz-Picciotto, I. & Rockhill, B. (1997). Validity and efficiency of approximation methods for tied survival times in Cox regression. Biometrics, 53(3), 1151-1156.

Pustejovsky, J.E. & Tipton, E. (2018). Small-sample methods for cluster-robust variance estimation and hypothesis testing in fixed effects models. Journal of Business & Economic Statistics, 36(4), 672-683.

Wasserstein, R.L., Schirm, A.L., & Lazar, N.A. (2019). Moving to a world beyond "p < 0.05". The American Statistician, 73(sup1), 1-19.

See Also

table_regression_models() for the registry of supported model classes and the per-family behaviour reference (also reachable as ?table_regression_mixed, ?table_regression_ordinal, ...). If a class is not listed there, try table_regression(fit) anyway – unsupported classes error with a clear message. Other regression-table functions: table_continuous_lm() for one-predictor-by-many-outcomes descriptive tables. Other spicy table functions: freq(), cross_tab(), table_categorical(), table_continuous(). Underlying machinery: spicy_print_table() for ASCII rendering; build_ascii_table() for the low-level renderer. Inferential infrastructure (internal): compute_model_vcov(), compute_coef_inference(), compute_wald_test(). broom integration: broom::tidy(), broom::glance().

Examples

# ---- Single-model usage ------------------------------------------
fit <- lm(wellbeing_score ~ age + sex + smoking, data = sochealth)

# Default APA layout: B / SE / 95% CI / p plus the n / R^2 /
# Adj. R^2 fit-stats footer. Factor reference level is annotated
# with `(ref.)` and shows an en dash in the statistic columns.
table_regression(fit)


# Standardised coefficients (beta) injected next to B. "refit"
# is the Cohen et al. (2003) refit-on-z-scores convention;
# "basic" reproduces the SPSS / Stata regress, beta definition.
table_regression(fit, standardized = "refit")



# Custom column set: B + AME + AME-specific p-value. Note that
# the `p` token always belongs to B, never to AME -- use the
# explicit `ame_p` token for AME inference.
if (requireNamespace("marginaleffects", quietly = TRUE)) {
  table_regression(
    fit,
    show_columns = c("b", "p", "ame", "ame_ci", "ame_p")
  )
}



# Group-token shortcut: "all_b" + "all_ame" expands to the full
# B / AME column families side by side.
if (requireNamespace("marginaleffects", quietly = TRUE)) {
  table_regression(fit, show_columns = c("all_b", "all_ame"))
}


# ---- Cluster-robust variance -------------------------------------
# CR2 (Bell-McCaffrey) with Satterthwaite-corrected df is the
# recommended default under few clusters. Three forms are accepted
# for `cluster`; the formula is preferred for composability with
# multi-way clustering and for programmatic robustness.

if (requireNamespace("clubSandwich", quietly = TRUE)) {
  table_regression(fit, vcov = "CR2", cluster = ~region)
}
if (requireNamespace("clubSandwich", quietly = TRUE)) {
  table_regression(fit, vcov = "CR2", cluster = "region")
}
if (requireNamespace("clubSandwich", quietly = TRUE)) {
  table_regression(fit, vcov = "CR2", cluster = ~region:age_group)
}



# ---- Hierarchical (nested) regression ----------------------------
# Adds in-table change-statistic rows (Delta R^2 / F-change /
# p-change for lm; LRT / p-change for glm) below the fit-stats.
# Note: hierarchical comparison requires identical observations
# across all models -- prepare a complete-case subset first so
# R's listwise deletion does not produce different `nobs` per
# model (which the function rejects).
sochealth_cc <- na.omit(
  sochealth[, c("wellbeing_score", "age", "sex", "smoking")]
)
m1 <- lm(wellbeing_score ~ age,                  data = sochealth_cc)
m2 <- lm(wellbeing_score ~ age + sex,            data = sochealth_cc)
m3 <- lm(wellbeing_score ~ age + sex + smoking,  data = sochealth_cc)
table_regression(
  list("Step 1" = m1, "Step 2" = m2, "Step 3" = m3),
  nested = TRUE
)



# ---- Side-by-side variance comparison ----------------------------
# Same fit, three vcovs in one wide table. Useful for showing the
# sensitivity of inference to the variance assumption.
if (requireNamespace("clubSandwich", quietly = TRUE)) {
  table_regression(
    list("Classical" = fit, "HC3" = fit, "CR2" = fit),
    vcov    = list("classical", "HC3", "CR2"),
    cluster = list(NULL, NULL, ~region)
  )
}


# ---- Tidy long format for downstream pipelines -------------------
if (requireNamespace("broom", quietly = TRUE)) {
  broom::tidy(table_regression(fit))
}


# ---- Mixed-effects models ----------------------------------------
# Linear mixed-effects (lme4). The footer adds a random-effects
# panel with sigma + Wald SE / CI from `merDeriv`, the Nakagawa
# marginal / conditional R^2 fit-stats, and a per-class p-value
# annotation line.
if (requireNamespace("lme4", quietly = TRUE)) {
  fit <- lme4::lmer(Reaction ~ Days + (Days | Subject),
                     data = lme4::sleepstudy)
  table_regression(fit)

  # Switch to the variance scale (sigma^2 instead of sigma).
  table_regression(fit, re_scale = "variance")

  # Minimal random-effects display: estimates only, no SE / CI.
  table_regression(fit, re_columns = "est")

  # Suppress the random-effects panel entirely.
  table_regression(fit, show_re = FALSE)
}

# Hierarchical mixed-effects comparison (nested LRT).
if (requireNamespace("lme4", quietly = TRUE)) {
  m1 <- lme4::lmer(Reaction ~ 1     + (1 | Subject),
                    data = lme4::sleepstudy, REML = FALSE)
  m2 <- lme4::lmer(Reaction ~ Days  + (1 | Subject),
                    data = lme4::sleepstudy, REML = FALSE)
  m3 <- lme4::lmer(Reaction ~ Days  + (Days | Subject),
                    data = lme4::sleepstudy, REML = FALSE)
  table_regression(list(m1, m2, m3), nested = TRUE)
}


## Not run: 
# ---- Rich-format outputs (require optional Suggests packages) ----
table_regression(fit, output = "gt")
table_regression(fit, output = "flextable")
table_regression(fit, output = "tinytable")

# ---- File outputs ------------------------------------------------
table_regression(fit, output = "excel",
                 excel_path = tempfile(fileext = ".xlsx"))
table_regression(fit, output = "word",
                 word_path  = tempfile(fileext = ".docx"))

# ---- System clipboard (interactive use) --------------------------
table_regression(fit, output = "clipboard")

## End(Not run)


Supported models and per-family behaviour of table_regression()

Description

table_regression_models() returns the registry of model classes supported by table_regression(), one row per engine, with each class's family, average-marginal-effects estimand, exponentiate semantics, and labelled table blocks. The same registry drives this page's table, so the published list cannot drift from the code.

This page is also the reference for per-family behaviour (the sections below). It is reachable as ?table_regression_models, ?table_regression_mixed, ?table_regression_ordinal, ?table_regression_counts, ?table_regression_categorical, ?table_regression_survival, ?table_regression_robust, or ?table_regression_bayesian.

If a class is not listed: fit the model and call table_regression(fit) anyway – unsupported classes error with a clear message naming the supported set. Feature requests are welcome on the issue tracker.

Usage

table_regression_models()

Value

A data frame with one row per supported engine and columns family, class, engine, ame, exponentiate, blocks.

Supported classes

Family Class Engine AME Exponentiate Blocks
Linear and generalized linear lm stats::lm() yes - -
Linear and generalized linear glm stats::glm() yes OR / IRR / RR / MR / HR (link) -
Linear and generalized linear negbin MASS::glm.nb() yes IRR -
Linear and generalized linear rlm MASS::rlm() yes - -
Linear and generalized linear nls stats::nls() no - -
Robust, IV, quantile, panel lm_robust estimatr::lm_robust() yes - -
Robust, IV, quantile, panel iv_robust estimatr::iv_robust() yes - -
Robust, IV, quantile, panel ivreg AER::ivreg() yes - -
Robust, IV, quantile, panel tobit AER::tobit() yes - -
Robust, IV, quantile, panel rq quantreg::rq() yes - -
Robust, IV, quantile, panel fixest fixest::feols(), fixest::feglm(), fixest::fepois(), fixest::fenegbin() yes feglm: OR / IRR -
Mixed effects lmerMod lme4::lmer() yes - Random effects
Mixed effects glmerMod lme4::glmer() yes OR / IRR (link) Random effects
Mixed effects glmmTMB glmmTMB::glmmTMB() yes link-dependent (IRR for count families) Random effects; Zero-inflation; Dispersion
Mixed effects lme nlme::lme() yes - Random effects
Mixed effects gls nlme::gls() yes - -
Population-averaged (GEE) geeglm geepack::geeglm() yes OR / IRR / RR / MR / HR (link) -
Ordinal polr MASS::polr() per category OR (logit) Thresholds
Ordinal clm ordinal::clm() per category OR (logit) Thresholds; Non-proportional effects
Categorical multinom nnet::multinom() per outcome OR per-outcome blocks
Categorical mlogit mlogit::mlogit() no OR per-alternative rows
Counts, two-part zeroinfl pscl::zeroinfl() yes (combined response) IRR (count) + OR (logit zero part) Zero-inflation
Counts, two-part hurdle pscl::hurdle() yes (combined response) IRR (count) + OR (logit zero part) Zero hurdle
Survival coxph survival::coxph() RMST / risk diff HR -
Survival survreg survival::survreg() yes + RMST / risk diff TR (log-scale distributions) -
Survival cph rms::cph() no HR -
Survival flexsurvreg flexsurv::flexsurvreg() no TR / HR (dist) distribution parameters
Survey-weighted svyglm survey::svyglm() yes (design-based) OR / IRR -
Survey-weighted svyolr survey::svyolr() per category (design-based) OR (logit) Thresholds
Survey-weighted svycoxph survey::svycoxph() no HR -
Additive, proportions, selection gam mgcv::gam(), mgcv::bam() yes OR / IRR (link) -
Additive, proportions, selection betareg betareg::betareg() yes OR (mean link) -
Additive, proportions, selection selection sampleSelection::selection() no - selection component
rms ols rms::ols() yes - -
rms lrm rms::lrm() yes OR -
rms Glm rms::Glm() yes link-dependent -
Bayesian stanreg rstanarm::stan_glm(), rstanarm::stan_glmer() yes (draws) link-dependent Random effects (if multilevel)
Bayesian brmsfit brms::brm() yes (draws) link-dependent Random effects (if multilevel)

Shared semantics (all classes)

Mixed effects

Fixed effects: Satterthwaite t (lmer + lmerTest), Wald z (glmer, glmmTMB), containment-df t (lme). Random effects render as a ⁠Random effects⁠ block of rows (SD / correlation / residual with SE and CI; re_scale, re_columns), deliberately with no per-row p-value (boundary-invalid Wald; Self & Liang 1987) – the footer carries the chi-bar-squared LR test of the whole random part, and re_test = "lrt" / "rlrt" adds an opt-in boundary-correct per-term test. N (groups) and ICC are fit-stat rows; Nakagawa marginal / conditional R-squared are the default R-squared family. ⁠CR*⁠ cluster-robust standard errors are available for lmer and lme fits, via clubSandwich with Satterthwaite degrees of freedom. glmer and glmmTMB fits keep their model-based standard errors: a ⁠CR*⁠ request on them is refused with a clear error.

Population-averaged (GEE) models

geepack::geeglm() fits are read on their own terms: the sandwich standard errors the fit computed (its ⁠std.err =⁠ option, clustered on its ⁠id =⁠) are the displayed inference – GEE is robust by construction, so spicy's vcov / cluster arguments are refused with a pointer to the fit options. Coefficients are population-averaged (marginal) effects; the footer discloses the working correlation structure with its estimated alpha. Wald z inference; exponentiate follows the usual link gates (OR / IRR / RR). Default fit statistics report the cluster structure (n, ⁠N (<id>)⁠, largest cluster); the quasi-likelihood information criteria "qic" / "qicu" (Pan 2001) and the "scale" (dispersion) parameter are opt-in – there is no likelihood, so AIC, pseudo-R-squared, nested = TRUE, and standardized are refused. See the population-averaged section of the Mixed-effects regression tables article for the contrast with subject-specific mixed models.

Ordinal models

Cut-points render as a Thresholds block (log-odds scale, never exponentiated; show_thresholds). Partial-proportional-odds clm terms render as a ⁠Non-proportional effects⁠ block, one coefficient per cut-point. exponentiate yields proportional odds ratios under logit; ci_method = "profile" profiles the predictor coefficients. AME is per-category (the marginal effect on each P(Y = k)). Defaults include McFadden and Nagelkerke pseudo-R-squared. See the Ordinal regression tables article.

Counts and two-part models

Two-part models show their full model: the zero component renders as a Zero-inflation block (zeroinfl, glmmTMB ziformula: probability of a structural zero) or a ⁠Zero hurdle⁠ block (hurdle: probability of a nonzero count – the opposite direction, hence the distinct label), and a Dispersion block when dispformula has covariates. Component coefficients join the p_adjust family and take stars; a zero component is exponentiated only under a logit link (odds ratio). AME is the combined-response effect on E(Y). ⁠CR*⁠ for pscl fits covers both components via sandwich::vcovCL(). Opt out with show_components = FALSE.

Categorical outcomes

multinom renders per non-reference outcome; exponentiate yields odds ratios of each outcome against the reference outcome – the baseline-category logits are log-odds (Agresti; SAS prints "Odds Ratio Estimates" under its generalized-logit link; Stata's ⁠mlogit, rrr⁠ labels the same quantity a relative-risk ratio). AME is per-outcome. nested = TRUE compares nested multinom fits by likelihood-ratio test (the anova.multinom() convention). Cluster-robust ⁠CR*⁠ is available (one cluster value per observation; sandwich >= 3.1-2) and the AME columns honour it; ⁠HC*⁠ is refused – a multi-equation model has no working residuals. mlogit renders per-alternative rows; AME is refused (no slopes() method exists for its data format). ⁠CR*⁠ is available with one cluster value per choice situation, and n counts choice situations; ⁠HC*⁠ is refused (sandwich::vcovHC() mis-scales the meat for mlogit's per-chooser score structure).

Survival models

Cox models exponentiate to hazard ratios; survreg log-scale distributions to time ratios (identity-scale distributions are left untouched). AME is refused for Cox fits (no marginal-probability effect on the hazard scale); their absolute-effect columns are the "rmst" and "risk_diff" families instead – covariate-adjusted RMST and cumulative-incidence differences by g-computation, with the mandatory tau / at_time horizons. For coxph: right-censored single-record fits, strata() supported (within-stratum baselines), tt() refused. For survreg: the closed-form AFT curves are standardized directly (stratified survreg refused). ⁠CR*⁠ uses the Lin-Wei grouped-dfbeta sandwich (coxph) or rms::robcov() (cph, needs ⁠x = TRUE, y = TRUE⁠). nested = TRUE compares nested Cox fits by likelihood-ratio test.

Survey-design models

Fits from a survey::svydesign() or survey::as.svrepdesign() design – svyglm (and its replicate sibling svrepglm), svyolr, svycoxph (and svrepcoxph) – are read as design-based throughout: the coefficients, the variance and the reference distribution all come from survey.

Inference is Wald t at the degrees of freedom survey writes on the FIT – df.residual for svyglm / svyolr, degf.resid or degf.residual for the two Cox engines – which is what survey::regTermTest() takes as its denominator. It is not survey::degf(design), and it is not re-derived here: the six engines of survey do not share one expression and are not harmonised (a Cox fit carries degf(design) - p + 1 although it has no intercept for the + 1 to cancel, so it ends one above the two other classes). The value is read off the object. The footer names the design and prints the number, and the average marginal effects answer to the same distribution as the coefficient rows.

The average marginal effect is the Horvitz-Thompson estimator: the mean unit-level effect weighted by the sampling weights of the analytic sample, with its variance from the delta method on the design vcov.

Counts are both reported: the observed n and the ⁠Weighted n⁠ the estimates describe. svycoxph adds the number of events, and its concordance goes to the footer.

What is refused, and why: every model-derived variance (⁠HC*⁠, ⁠CR*⁠, bootstrap, jackknife) – the design is the variance authority, and the way to change the estimator is to change the design; every likelihood statistic (AIC, BIC, logLik, deviance, pseudo-R-squared) for svyolr and svycoxph – there is no likelihood, and survey's own deviance() returns a sign-flipped likelihood-ratio statistic on one Cox engine and a bare zero on the other; the AME for svycoxph, on the same ground as for a plain Cox fit; and the "rmst" / "risk_diff" columns for svycoxph, whose uncertainty comes from resampling subjects and so ignores the strata and clusters the design declares (use survey::svykm() for a marginal curve).

nested = TRUE is refused for a design-based table: there is no likelihood to compare, so every change statistic would be empty, and a block of empty rows reads like an answer. Put the models side by side with nested = FALSE, and test a term under the design with survey::regTermTest().

svyglm keeps an AIC row: survey's extractAIC.svyglm computes the design-based AIC of Lumley & Scott (2015), the one information criterion published for this class, and show_fit_stats = "eff_p" reports the effective number of design parameters beside it. BIC.svyglm requires a maximal model and has no default, so it stays blank. See the Summary tables from a survey design article.

Robust, IV, quantile and panel models

estimatr fits keep their own robust SEs (never overwritten); quantreg::rq() defaults to the heteroskedasticity-robust "nid" sandwich (quantreg's own large-sample default), with "iid", "ker", "rank" (CIs only) and a native "bootstrap" – clustered via the wild gradient bootstrap – as vcov options (the footer names the estimator); fixest fits – and estimatr fits built with ⁠fixed_effects =⁠ – disclose their absorbed fixed effects as a default-on ⁠Fixed effects:⁠ block (one Yes / No row per factor; varying-slope-only factors are not absorbed intercepts and read No), with per-factor ⁠N (<factor>)⁠ counts via the opt-in n_groups token and the within R-squared in the default fit statistics for fixest (opt-in via within_r2 for estimatr).

Bayesian models

Posterior median, posterior MAD SD, and equal-tailed credible intervals (ci_method = "hdi" opts into the highest-density interval); deliberately no p-value column and no stars – the probability of direction ("pd") is the opt-in posterior summary. A sampler-diagnostics guard checks every fit (R-hat, ESS, divergences, E-BFMI) and per-coefficient "rhat" / "ess_bulk" / "ess_tail" / "mcse" columns are available. The AME columns are draws-native (posterior median, MAD SD and credible interval of the per-draw avg_slopes(); no "ame_p"), and so are the standardized betas ("posthoc" / "basic" / "smart", exact affine rescales of the draws) on fixed-effects fits: stan_glm-style models and standard-formula brm() models, whose design matrix is recovered through insight. Multilevel fits, stan_polr / stan_betareg, brms formulas with distributional or special terms, and "refit" / "pseudo" are refused with a pre-standardization hint. Multilevel fits (stan_glmer, brm with grouping terms) report their random effects as a block – posterior median SD and credible interval per component, from the draws – with no likelihood-ratio line. p_adjust and likelihood-based fit-statistic tokens are refused (no p-values, no likelihood-based information criteria in a posterior); "r2_bayes" is in the default fit statistics and "elpd_loo" / "looic" / "waic" are opt-in, with standard errors and reliability caveats in the footer; compare models with loo::loo_compare() outside the table.

See Also

table_regression(); the Publication-ready regression tables and Ordinal regression tables articles.

Examples

table_regression_models()

# All engines of one family:
subset(table_regression_models(), family == "Mixed effects")

Univariable screening table (with optional multivariable merge)

Description

Fits one model per candidate predictor (the univariable screen) and renders them as a single table with one row block per predictor. With multivariable = TRUE (default), the full model containing all predictors is merged side by side under "Univariable" / "Multivariable" column groups – the standard presentation of applied epidemiology (the gtsummary::tbl_uvregression() + tbl_merge() workflow).

Usage

table_regression_uv(
  data,
  outcome,
  predictors,
  method = c("lm", "glm", "coxph"),
  family = stats::binomial(),
  multivariable = TRUE,
  complete_cases = FALSE,
  show_columns = c("n", "b", "ci", "p"),
  show_intercept = FALSE,
  title = NULL,
  ...
)

Arguments

data

A data frame.

outcome

The outcome column (unquoted name, tidyselect). For method = "coxph", a Surv(time, status) expression evaluated in data (the tbl_uvregression convention).

predictors

Candidate predictor columns (tidyselect, e.g. c(age, sex, education) or where(is.numeric)). The outcome column(s) are dropped from the selection automatically.

method

"lm" (default), "glm", or "coxph" (requires the survival package; estimates render as HRs with exponentiate = TRUE).

family

A stats::family for method = "glm", in any of the three forms stats::glm() accepts: a family object (binomial()), its name ("binomial"), or the bare constructor (binomial). Default binomial(), so method = "glm" alone is the logistic screen; supplying family without method selects the glm screen directly (a family can only mean that). Refused for method = "coxph", and gaussian() with the identity link is refused too: use method = "lm" for the linear screen. With method = "lm", any non-gaussian family is refused the same way (use method = "glm"), and a supplied gaussian() is ignored with a classed warning – the linear screen already fits it.

multivariable

Logical, default TRUE: merge the full model (all predictors together) as a second column group.

complete_cases

Logical, default FALSE. TRUE restricts ALL models to the rows complete on outcome + every predictor (common-sample comparison); the reduction is disclosed in the table note.

show_columns

Passed to table_regression(). Default c("n", "b", "ci", "p") – the tbl_uvregression column set; "n" is the per-predictor sample size. The multivariable group carries no N column (its single n is a fit-statistics row, as in the reference layouts). For binary outcomes, add "n_events" for outcome event counts as events/N per factor level (each column group counts on its own estimation sample). For method = "lm", add "r2" (and/or "adj_r2") for the share of outcome variance each predictor explains on its own – see Variance explained below. For method = "coxph", the RMST / risk-difference families ("rmst", "risk_diff", ...) work with an explicit numeric tau / at_time shared by every column: each univariable fit runs its own boot_n-replicate bootstrap, and the multivariable group reports the covariate-adjusted estimand from the full fit. tau = "minmax" is refused (per-fit horizons would make the column incomparable across predictors).

show_intercept

Display the (Intercept) rows? Default FALSE – the opposite of table_regression()'s default, because each univariable fit carries its own nuisance intercept. See Intercepts.

title

Table title; NULL (default) builds "Univariable and multivariable <type> regression: <outcome>".

...

Passed to table_regression() (exponentiate, vcov, cluster, p_adjust, digits, labels, output, ...). nested is not meaningful for a screen and is refused. cluster must be a single vector with one value per row of data; it is aligned to each fit's own estimation sample automatically.

Value

See table_regression() (same output contract).

Sample sizes

By default each univariable model is fit on its own complete cases, so N varies across predictors – that is what the N column discloses (shown on the first row of each block), and a table note states it whenever the Ns differ. The multivariable model is fit on the complete cases of all its variables (its n appears in the fit-statistics rows). Pass complete_cases = TRUE to restrict every model – univariable and multivariable – to the common complete-case sample.

Variance explained

show_columns = c("n", "b", "ci", "p", "r2") adds an R^2 column to the screen (method = "lm"): each predictor block reports its own model's R^2, on the first row of the block like N. It answers what a coefficient and its interval cannot – how much of the outcome the predictor accounts for – and often shows that a firmly established association still explains a small share of the variance. Add "adj_r2" for the adjusted form. On the multivariable side the R^2 is one number for the whole model, so it stays in the fit-statistics rows (where it is shown by default) instead of being repeated down a column. Not available for method = "glm" or "coxph": outside least squares only competing pseudo-R^2 measures exist, and spicy asks you to name the one you want (show_fit_stats = "pseudo_r2_mcfadden").

Multiplicity

p_adjust (passed through to table_regression()) treats the whole univariable screen as ONE family (all screened coefficients together); the multivariable model is its own family, as in any multi-model table.

Why the default screen is linear

The default method = "lm" fits the linear screen: R's canonical model, and – when the outcome is continuous – the estimand with the most direct reading. If the outcome looks binary under this default, the screen proceeds as a linear probability model and says so in a classed warning: LPM coefficients are probability differences, comparable across models and samples in a way that odds ratios are not (Mood, 2010), but the model's built-in heteroskedasticity calls for vcov = "HC3". A two-level factor (or logical) outcome is coded 0/1 on its second level – the glm convention – and the warning names the modeled probability; an outcome with more observed levels is refused (a multinomial outcome has no linear screen). The classical epidemiological screen is one argument away: method = "glm" (with the default family = binomial()) gives the logistic screen, and supplying any family selects the glm screen directly.

Intercepts

Hidden by default on both sides (each univariable fit has its own nuisance intercept), matching gtsummary::tbl_regression()'s intercept = FALSE default. Pass show_intercept = TRUE to display them: each univariable block then opens with its own fit's (Intercept) row, and the multivariable model shows its intercept as in any table_regression() table.

References

Batra, N. et al. (Eds.) (2021). The Epidemiologist R Handbook, Univariate and multivariable regression. https://epirhandbook.com/en/new_pages/regression.html

Mood, C. (2010). Logistic regression: Why we cannot do what we think we can do, and what we can do about it. European Sociological Review, 26(1), 67-82. doi:10.1093/esr/jcp006

Sjoberg, D.D., Whiting, K., Curry, M., Lavery, J.A., & Larmarange, J. (2021). Reproducible summary tables with the gtsummary package. The R Journal, 13(1), 570-580.

Examples


table_regression_uv(
  sochealth,
  outcome    = smoking,
  predictors = c(age, sex, education),
  family     = binomial(),
  exponentiate = TRUE
)

# Linear screen with the share of variance each predictor
# explains on its own (the multivariable model reports its own
# R-squared in the fit-statistics rows).
table_regression_uv(
  sochealth,
  outcome      = wellbeing_score,
  predictors   = c(age, sex, bmi),
  show_columns = c("n", "b", "ci", "p", "r2")
)


Terms method for univariable screen bundles

Description

Returns the stats::terms() object of the formula outcome ~ predictor_1 + ... + predictor_k spanning every predictor screened by table_regression_uv(). Non-syntactic column names are backtick-quoted, so the terms are valid whatever the input names. Used internally by the label validator, which reads term labels off every model in a table.

Usage

## S3 method for class 'spicy_uv_screen'
terms(x, ...)

Arguments

x

A spicy_uv_screen bundle (the internal object wrapping the univariable fits of table_regression_uv()).

...

Additional arguments (currently ignored).

Value

A terms object for ⁠outcome ~ all screened predictors⁠.

See Also

table_regression_uv()


Tidying methods for a spicy_categorical_svy_table

Description

Standard broom::tidy() and broom::glance() interfaces for an object returned by table_categorical_svy().

Usage

## S3 method for class 'spicy_categorical_svy_table'
tidy(x, ...)

## S3 method for class 'spicy_categorical_svy_table'
glance(x, ...)

Arguments

x

A spicy_categorical_svy_table returned by table_categorical_svy().

...

Ignored, for S3 compatibility.

Details

tidy() returns one row per (variable x level x column block), which is the LONG reading of a table whose blocks are columns. Columns: variable, label, level, group (the by level, or the margin, NA without by), total (whether that block is the margin), n (observed), estimate (the estimated percentage), conf.low, conf.high, deff. The header rows carry no level statistic and do not appear; their p-value is what glance() is for.

glance() returns one row per variable: variable, label, n_levels, p.value, statistic_type (the svychisq() statistic asked for), degf (the design's own), nobs, weighted.nobs.

n_levels counts the levels the TABLE displays, so a (Missing) display level counts; the test behind p.value runs on the complete cases and the observed levels only, as it does in table_categorical().

Value

A tbl_df (or data.frame when tibble is not installed).


Tidying methods for a spicy_categorical_table

Description

Standard broom::tidy() and broom::glance() interfaces for an object returned by table_categorical(). They re-shape the underlying long-format data (stored on the object as the "long_data" attribute) into the two canonical broom views so the table can be consumed by any downstream tidyverse-stats pipeline.

Usage

## S3 method for class 'spicy_categorical_table'
tidy(x, ...)

## S3 method for class 'spicy_categorical_table'
glance(x, ...)

Arguments

x

A spicy_categorical_table returned by table_categorical().

...

Currently ignored. Present for compatibility with the broom::tidy() / broom::glance() generics.

Details

tidy() returns one row per ⁠(variable x level)⁠ – or per ⁠(variable x level x group)⁠ when by is used – with broom-conventional columns: outcome, level, group (when applicable), n, proportion (the percentage divided by 100).

glance() returns one row per outcome with the omnibus chi-squared test (when by is used) and the requested association measure: outcome, test_type ("chi_squared"), statistic (chi-squared), df, p.value, assoc_type, assoc_value, assoc_ci_lower, assoc_ci_upper, n_total. Without by, only outcome and n_total are populated; the other columns are NA.

Value

A tbl_df.

See Also

as.data.frame.spicy_categorical_table() for the raw wide-format access; tidy.spicy_continuous_table() for the continuous-descriptive companion.


Tidying methods for a spicy_continuous_lm_table

Description

Standard broom::tidy() and broom::glance() interfaces for an object returned by table_continuous_lm(). They re-shape the underlying long-format data into the two canonical broom views so the table can be consumed by any downstream tidyverse-stats pipeline.

Usage

## S3 method for class 'spicy_continuous_lm_table'
tidy(x, ...)

## S3 method for class 'spicy_continuous_lm_table'
glance(x, ...)

Arguments

x

A spicy_continuous_lm_table returned by table_continuous_lm().

...

Currently ignored. Present for compatibility with the broom::tidy() / broom::glance() generics.

Details

tidy() returns one row per estimated parameter across all outcomes:

Standard broom columns: outcome, label, term, estimate_type, estimate, std.error, conf.low, conf.high, statistic, p.value. The outcome column carries the original variable name; label carries the human-readable label.

glance() returns one row per outcome with model-level statistics. Columns: outcome, label, predictor_type ("categorical" or "continuous"), test_type ("F" for categorical predictors, "t" for continuous ones), statistic, df, df.residual, p.value, r.squared, adj.r.squared, es_type, es_value, es_ci_lower, es_ci_upper, nobs, weighted_n.

Value

A tbl_df.

See Also

as.data.frame.spicy_continuous_lm_table() for the raw long-format access.


Tidying methods for a spicy_continuous_svy_table

Description

Standard broom::tidy() and broom::glance() interfaces for an object returned by table_continuous_svy().

Usage

## S3 method for class 'spicy_continuous_svy_table'
tidy(x, ...)

## S3 method for class 'spicy_continuous_svy_table'
glance(x, ...)

Arguments

x

A spicy_continuous_svy_table returned by table_continuous_svy().

...

Ignored, for S3 compatibility.

Details

tidy() returns one row per displayed row: one per variable, or one per (variable x group) with by. Columns: variable, label, group (NA without by), estimate (the mean), std.error (the design-based standard error, from survey::svymean() and never recomputed from sd / sqrt(n) – under a design those are different quantities), conf.low, conf.high, df (the degrees of freedom the interval used), n (observed), weighted.n (the sum of the sampling weights), median, q1, q3, min, max, sd, deff.

glance() returns one row per variable with its group comparison: variable, label, n_groups, test_type, statistic, df, df.residual, p.value, degf (the design's own), nobs, weighted.nobs. One row per variable even without by, where the comparison columns are NA – a fixed schema a pipeline can index into by NAME.

Value

A tbl_df (or data.frame when tibble is not installed).


Tidying methods for a spicy_continuous_table

Description

Standard broom::tidy() and broom::glance() interfaces for an object returned by table_continuous(). They re-shape the underlying long-format data into the two canonical broom views so the descriptive table can be consumed by any downstream tidyverse-stats pipeline.

Usage

## S3 method for class 'spicy_continuous_table'
tidy(x, ...)

## S3 method for class 'spicy_continuous_table'
glance(x, ...)

Arguments

x

A spicy_continuous_table returned by table_continuous().

...

Currently ignored. Present for compatibility with the broom::tidy() / broom::glance() generics.

Details

tidy() returns one row per ⁠(variable x group)⁠ (or per variable when by is not used) with broom-conventional columns: outcome, label, group (when applicable), estimate (the empirical mean), std.error (sd / sqrt(n)), conf.low, conf.high (the mean confidence interval at ci_level), n, min, max, sd. The outcome column carries the variable name and label the human-readable label.

glance() returns one row per outcome with the omnibus group comparison (when by is used), the requested effect size and the balance diagnostic. Columns: outcome, label, test_type, statistic, df, df.residual, p.value, es_type, es_value, es_ci_lower, es_ci_upper, smd_type, smd_value, n_total. Without by, only outcome, label, and n_total are populated; the other columns are NA. The schema is fixed: smd_type / smd_value are present whether or not smd = TRUE, like every other comparison column here, and are NA when the table carries no standardized mean difference. They sit before n_total, so index this frame by NAME rather than by position.

glance.spicy_categorical_table() does not carry the SMD: each family documents its own broom contract, and the categorical one publishes the association measure alone. On a mixed balance table, read the categorical SMD from output = "long".

Value

A tbl_df.

See Also

as.data.frame.spicy_continuous_table() for the raw long-format access; tidy.spicy_continuous_lm_table() for the model-based companion.


Tidying methods for a spicy_outcome_table

Description

Standard broom::tidy() and broom::glance() interfaces for an object returned by table_outcome().

Usage

## S3 method for class 'spicy_outcome_table'
tidy(x, ...)

## S3 method for class 'spicy_outcome_table'
glance(x, ...)

Arguments

x

A spicy_outcome_table returned by table_outcome().

...

Ignored, for S3 compatibility.

Details

tidy() returns the DESCRIBED rows: the marginal Overall row and one row per (grouping x level). Columns: outcome (the outcome name, constant down the frame), variable (the grouping, or the outcome itself on the marginal row), label, level (NA on the marginal row), estimate (the mean), std.error (sd / sqrt(n)), conf.low, conf.high, n, min, max, sd.

Two identity columns where the sibling has one, and deliberately: here the outcome is fixed and the variable changes, so a single outcome column would have to mean two different things down the frame. The schema reads without knowing which function produced it.

glance() returns one row per grouping – one BLOCK – with that block's own comparison. Columns: outcome, variable, label, n_levels, test_type, statistic, df, df.residual, p.value, es_type, es_value, es_ci_lower, es_ci_upper, smd_type, smd_value, n_total.

n_levels counts the levels the TABLE displays, so a missing-value display level counts; the comparison behind test_type runs on the observed levels only, as it does everywhere in the family.

The schema is FIXED: smd_type / smd_value are present and NA from the first version, so the day a standardized mean difference enters this table it cannot break a pipeline that indexes the frame. Index by NAME rather than by position.

Value

A tbl_df (or data.frame when tibble is not installed).


Tidy / glance methods for spicy_regression_table

Description

Standard broom::tidy() / broom::glance() interfaces for an object returned by table_regression(). They re-shape the underlying long-format data into the two canonical broom views so the table can be consumed by any downstream tidyverse-stats pipeline.

Usage

## S3 method for class 'spicy_regression_table'
tidy(x, ...)

## S3 method for class 'spicy_regression_table'
glance(x, ...)

Arguments

x

A spicy_regression_table returned by table_regression().

...

Currently ignored. Present for compatibility with the broom::tidy() / broom::glance() generics.

Details

tidy() returns one row per ⁠(model_id, term, estimate_type, outcome_level)⁠ combination, with estimate_type in c("B", "beta", "ame", "partial_f2", "partial_eta2", "partial_omega2"). outcome_level names the response category of per-category rows (ordinal / multinomial average marginal effects) and is NA for single-outcome models. Reference-row placeholders (factor reference levels) and singular coefficients (NA estimates) are dropped. Columns: ⁠model_id, outcome, outcome_level, term, estimate_type, estimate, std.error, conf.low, conf.high, statistic, df, p.value, test_type, is_intercept, factor_term, factor_level⁠.

glance() returns one row per ⁠(model_id, outcome)⁠ with model-level statistics. Columns: ⁠model_id, outcome, nobs, weighted_nobs, r.squared, adj.r.squared, omega2, sigma, rmse, f2, AIC, AICc, BIC, deviance, df.residual⁠ (numeric – Satterthwaite-safe).

Value

A tbl_df.

See Also

as.data.frame.spicy_regression_table() for the wide raw view.


Uncertainty Coefficient

Description

uncertainty_coef() computes the Uncertainty Coefficient (Theil's U) for a two-way contingency table, based on information entropy.

Usage

uncertainty_coef(
  x,
  direction = c("symmetric", "row", "column"),
  detail = FALSE,
  conf_level = 0.95,
  digits = 3L
)

Arguments

x

A contingency table (of class table).

direction

Direction of prediction: "symmetric" (default), "row" (column predicts row), or "column" (row predicts column).

detail

Logical. If FALSE (default), return the estimate as a numeric scalar. If TRUE, return a named numeric vector including confidence interval and p-value.

conf_level

A single number strictly between 0 and 1 giving the confidence level (default 0.95). Only used when detail = TRUE. Set to NULL to omit the confidence interval. Any other value – including percentages such as 95 – raises a classed error (spicy_invalid_input).

digits

Number of decimal places used when printing the result (default 3). Only affects the detail = TRUE output.

Details

The uncertainty coefficient measures association using Shannon entropy. Let H_X and H_Y be the marginal entropies of the row and column variables respectively, and H_{XY} the joint entropy.

The default direction = "symmetric" follows the SPSS and DescTools convention: the symmetric coefficient is a standard, well-defined variant with its own asymptotic standard error. somers_d() deliberately differs (its default is "row") because its symmetric form is a derived quantity without an analytic SE; see its documentation.

When the marginal entropy in the denominator is zero (the predicted variable is constant, e.g. an unused factor level), the coefficient is the undefined form 0/0: the function returns NA with a spicy_undefined_stat warning, like the other measures in the family. For direction = "symmetric" this happens only when both variables are constant; with a single constant variable the symmetric coefficient is a well-defined 0.

The entropy terms use the standard mathematical convention 0 \log 0 = 0, matching SPSS / PSPP CROSSTABS and the definition in Cover & Thomas (2006). Note that DescTools::UncertCoef() applies an additional Laplace correction (replacing zero cells with 1/n^2) before the entropy computation, which produces slightly different point estimates on tables with empty cells; that correction is uncommon in the information-theory literature and is not used here. The asymptotic standard errors follow the DescTools delta method; see cramer_v() for full references.

Value

Same structure as cramer_v(): a scalar when detail = FALSE, a named vector when detail = TRUE. The p-value tests H0: U = 0 (Wald z-test).

References

Theil, H. (1970). On the estimation of relationships involving qualitative variables. American Journal of Sociology, 76(1), 103-154. doi:10.1086/224909

See Also

lambda_gk(), goodman_kruskal_tau(), assoc_measures()

Other association measures: assoc_measures(), contingency_coef(), cramer_v(), gamma_gk(), goodman_kruskal_tau(), kendall_tau_b(), kendall_tau_c(), lambda_gk(), phi(), somers_d(), yule_q()

Examples

tab <- table(sochealth$smoking, sochealth$education)
uncertainty_coef(tab)
uncertainty_coef(tab, direction = "row", detail = TRUE)


Generate a comprehensive summary of the variables

Description

varlist() lists the variables of a data frame and extracts essential metadata: variable names, labels, summary values, classes, number of distinct values, number of valid (non-missing) observations, and number of missing values. Tidyselect-style selectors can be supplied to pick or reorder columns dynamically.

vl() is a convenient shorthand for varlist() that offers identical functionality with a shorter name.

Usage

varlist(
  x,
  ...,
  values = FALSE,
  tbl = FALSE,
  include_na = FALSE,
  factor_levels = c("observed", "all"),
  user_na = TRUE
)

vl(
  x,
  ...,
  values = FALSE,
  tbl = FALSE,
  include_na = FALSE,
  factor_levels = c("observed", "all"),
  user_na = TRUE
)

Arguments

x

A data frame, or a transformation of one.

...

Optional tidyselect-style column selectors (e.g. starts_with("var"), where(is.numeric), etc.). Columns can be selected or reordered, but renaming selections is not supported.

values

Logical. If FALSE (the default), displays a compact summary of the variable's values. For numeric, character, date/time, labelled, and factor variables, all unique non-missing values are shown when there are at most four; otherwise the first three values, an ellipsis (...), and the last value are shown. Values are sorted when appropriate (e.g., numeric, character, date). For factors, factor_levels controls whether observed or all declared levels are shown; level order is preserved. For labelled variables, prefixed labels are displayed via labelled::to_factor(levels = "prefixed"). If TRUE, all unique non-missing values are displayed.

tbl

Logical. If FALSE (the default), opens the summary in the Viewer if the session is interactive. If TRUE, returns a tibble.

include_na

Logical. If TRUE, unique missing value markers (⁠<NA>⁠, ⁠<NaN>⁠) are explicitly appended at the end of the Values summary when present in the variable. This applies to all variable types. Literal strings "NA", "NaN", and "" are quoted to distinguish them from missing markers. If FALSE (the default), missing values are omitted from Values but still counted in the NAs column.

factor_levels

Character. Controls how factor values are displayed in Values. "observed" (the default; code_book() uses "all") shows only levels present in the data, preserving factor level order. "all" shows all declared levels, including unused levels. An explicit NA level (e.g. from addNA()) is displayed as ⁠<NA>⁠ among the declared levels.

user_na

Logical. If TRUE (the default), declared missing values count as missing in N_valid, NAs, and N_distinct (all three columns share one missing definition). If FALSE, they count as valid. Either way, the declared codes remain listed in Values (with their value labels when declared) – a codebook documents the full coding scheme. See the "Declared missing values" section of freq().

Details

In an interactive session (RStudio, Positron, ...), the summary opens in the Viewer pane with a contextual title like vl: sochealth. If the data frame has been transformed or subsetted, the title is suffixed with * (e.g. ⁠vl: sochealth*⁠); anonymous or ambiguous calls fall back to ⁠vl: <data>⁠. Pass tbl = TRUE to return a tibble instead.

The default factor_levels = "observed" mirrors what is actually in the data; code_book() defaults to "all" to document the declared schema. See ⁠@param factor_levels⁠ to override either default.

Value

A tibble with one row per selected variable, containing the following columns:

For matrix and array columns, observations are counted per row: a row is treated as missing if any of its cells is NA. N_valid / NAs therefore count complete vs. incomplete rows, not individual cells.

With tbl = FALSE (the default) the tibble is sent to the Viewer (interactive) or surfaced via a message (non-interactive) and the function returns invisibly NULL. Set tbl = TRUE to return the tibble directly for downstream use.

Declared missing values

Survey files imported with haven often carry declared missing values: codes such as ⁠8 = Don't know⁠ or 9 = Refused that the source file marks as missing while keeping them distinct from a plain NA. Two kinds of declaration exist: na_values / na_range metadata on haven::labelled_spss() vectors, and tagged missing values created by haven::tagged_na() (the Stata .a, .b, ... convention).

spicy honors the declaration by default (user_na = TRUE): declared missing values are excluded from every statistic exactly like NA – valid percentages, means, chi-squared tests, association measures, row-wise summaries, and group definitions – but they are not erased from display. freq() lists each observed declared value as its own row of the Missing block, with its value label; cross_tab(), table_categorical(), and table_continuous() disclose the exclusion in the table note (⁠Declared missing values removed: x (2).⁠); varlist() and code_book() count them as missing in N_valid / NAs / N_distinct while still listing the declared codes in Values.

Every function involved offers the same escape hatch: set user_na = FALSE to ignore the declaration and treat the declared codes as valid values (the behavior of spicy before 0.13.0). Tagged missing values are genuine NAs either way; for them, user_na = FALSE only collapses the per-tag breakdown back into the regular NA count.

See Also

Other variable inspection: code_book(), label_from_names()

Examples

varlist(sochealth, tbl = TRUE)
sochealth |> varlist(tbl = TRUE)
varlist(sochealth, where(is.numeric), values = TRUE, tbl = TRUE)
varlist(
  sochealth,
  starts_with("bmi"),
  values = TRUE,
  include_na = TRUE,
  tbl = TRUE
)

df <- data.frame(
  group = factor(c("A", "B", NA), levels = c("A", "B", "C"))
)
varlist(
  df,
  values = TRUE,
  include_na = TRUE,
  factor_levels = "all",
  tbl = TRUE
)

vl(sochealth, tbl = TRUE)
sochealth |> vl(tbl = TRUE)
vl(sochealth, starts_with("bmi"), tbl = TRUE)
vl(sochealth, where(is.numeric), values = TRUE, tbl = TRUE)

Yule's Q

Description

yule_q() computes Yule's Q coefficient of association for a 2x2 contingency table.

Usage

yule_q(x, detail = FALSE, conf_level = 0.95, digits = 3L)

Arguments

x

A contingency table (of class table).

detail

Logical. If FALSE (default), return the estimate as a numeric scalar. If TRUE, return a named numeric vector including confidence interval and p-value.

conf_level

A single number strictly between 0 and 1 giving the confidence level (default 0.95). Only used when detail = TRUE. Set to NULL to omit the confidence interval. Any other value – including percentages such as 95 – raises a classed error (spicy_invalid_input).

digits

Number of decimal places used when printing the result (default 3). Only affects the detail = TRUE output.

Details

For a 2x2 table with cells a, b, c, d, Yule's Q is Q = (ad - bc) / (ad + bc). It is equivalent to the Goodman-Kruskal Gamma for 2x2 tables. The asymptotic standard error is SE = 0.5 (1 - Q^2) \sqrt{1/a + 1/b + 1/c + 1/d}.

Edge cases: when ad + bc = 0, Q itself is undefined and the function returns NA with a spicy_undefined_stat warning. When any cell is zero (and ad + bc > 0), Q is well-defined but the SE formula divides by zero – the point estimate is returned, and se, ci_lower, ci_upper, and p_value are all NA.

Standard error formulas follow the DescTools implementations (Signorell et al., 2024); see cramer_v() for full references.

Value

Same structure as cramer_v(): a scalar when detail = FALSE, a named vector when detail = TRUE. The p-value tests H0: Q = 0 (Wald z-test).

References

Yule, G. U. (1900). On the association of attributes in statistics. Philosophical Transactions of the Royal Society of London, Series A, 194, 257-319. doi:10.1098/rsta.1900.0019

See Also

phi(), gamma_gk(), assoc_measures()

Other association measures: assoc_measures(), contingency_coef(), cramer_v(), gamma_gk(), goodman_kruskal_tau(), kendall_tau_b(), kendall_tau_c(), lambda_gk(), phi(), somers_d(), uncertainty_coef()

Examples

tab <- table(sochealth$smoking, sochealth$sex)
yule_q(tab)