title: “Using pubformat” output: rmarkdown::html_vignette vignette: > % % % ————————-

Overview

pubformat provides tools for converting numeric statistical results into consistent, publication-ready character output.

The package is designed around a simple principle:

Statistical analysis and statistical formatting should remain separate.

R or another statistical procedure produces the numeric result. pubformat handles the reporting step without changing the underlying statistical result or decision.

library(pubformat)

Formatting p-values

format_p() converts numeric p-values into publication-ready text.

format_p(0.048)
#> [1] "p = .048"
format_p(0.0002)
#> [1] "p < .001"

The reporting threshold is strict. A value equal to .001 is retained, while a value below .001 is reported using less-than notation.

format_p(0.001)
#> [1] "p = .001"
format_p(0.0009)
#> [1] "p < .001"

Optional significance codes can be added when useful for tables or other compact displays.

format_p(0.048, sig = TRUE)
#> [1] "p = .048 *"
format_p(0.006, sig = TRUE)
#> [1] "p = .006 **"
format_p(0.0002, sig = TRUE)
#> [1] "p < .001 ***"

Importantly, significance codes are determined from the original numeric p-value rather than its rounded display value.

format_p(0.0499, sig = TRUE)
#> [1] "p = .050 *"

Formatting confidence intervals

format_ci() combines numeric lower and upper confidence limits into publication-ready interval notation.

format_ci(1.08, 1.87)
#> [1] "95% CI [1.08, 1.87]"

The confidence level can be changed.

format_ci(1.08, 1.87, level = 0.99)
#> [1] "99% CI [1.08, 1.87]"

Leading zeros are retained by default because the appropriate convention depends on the statistic being reported.

format_ci(0.21, 0.48)
#> [1] "95% CI [0.21, 0.48]"

For statistics that cannot exceed 1 in absolute value, such as correlations, leading zeros can be omitted.

format_ci(0.21, 0.48, leading_zero = FALSE)
#> [1] "95% CI [.21, .48]"

Formatting correlations

For Pearson correlations, format_r() provides a concise formatter.

format_r(0.42)
#> [1] "r = .42"
format_r(-0.31)
#> [1] "r = -.31"

For Pearson, Spearman, or Kendall correlations, format_cor() selects the appropriate reporting symbol.

format_cor(0.42, method = "pearson")
#> [1] "r = .42"
format_cor(0.42, method = "spearman")
#> [1] "ρ = .42"
format_cor(0.42, method = "kendall")
#> [1] "τ = .42"

format_cor() formats an existing coefficient. It does not calculate the correlation itself.

For example:

rho <- cor(
  mtcars$mpg,
  mtcars$wt,
  method = "spearman"
)

format_cor(rho, method = "spearman")
#> [1] "ρ = -.89"

Formatting percentages

format_percent() interprets numeric values as proportions by default.

format_percent(0.423)
#> [1] "42.3%"

Values that are already expressed as percentages can be identified explicitly.

format_percent(42.3, input = "percent")
#> [1] "42.3%"

This distinction prevents accidental rescaling of values.

format_percent(0.42)
#> [1] "42.0%"
format_percent(42, input = "percent")
#> [1] "42.0%"

The function is vectorized.

format_percent(c(0.25, 0.50, 0.75))
#> [1] "25.0%" "50.0%" "75.0%"

A typical analysis-to-publication workflow

Consider a simple linear model using the built-in mtcars data.

model <- lm(mpg ~ wt, data = mtcars)

summary(model)
#> 
#> Call:
#> lm(formula = mpg ~ wt, data = mtcars)
#> 
#> Residuals:
#>     Min      1Q  Median      3Q     Max 
#> -4.5432 -2.3647 -0.1252  1.4096  6.8727 
#> 
#> Coefficients:
#>             Estimate Std. Error t value Pr(>|t|)    
#> (Intercept)  37.2851     1.8776  19.858  < 2e-16 ***
#> wt           -5.3445     0.5591  -9.559 1.29e-10 ***
#> ---
#> Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
#> 
#> Residual standard error: 3.046 on 30 degrees of freedom
#> Multiple R-squared:  0.7528, Adjusted R-squared:  0.7446 
#> F-statistic: 91.38 on 1 and 30 DF,  p-value: 1.294e-10

The analysis produces numeric statistical results. Those values can then be extracted and formatted separately.

P-value

p_value <- summary(model)$coefficients["wt", "Pr(>|t|)"]

p_value
#> [1] 1.293959e-10
format_p(p_value)
#> [1] "p < .001"

Confidence interval

ci <- confint(model, "wt")

ci
#>        2.5 %    97.5 %
#> wt -6.486308 -4.202635
format_ci(ci[1], ci[2])
#> [1] "95% CI [-6.49, -4.20]"

Correlation

r_value <- cor(mtcars$mpg, mtcars$wt)

r_value
#> [1] -0.8676594
format_r(r_value)
#> [1] "r = -.87"

The workflow is therefore:

statistical analysis
        |
        v
numeric statistical results
        |
        v
pubformat
        |
        v
publication-ready output

pubformat does not replace statistical analysis functions. Its purpose is to make the final reporting step more consistent, reproducible, and convenient.

Vectorized workflows

The formatting functions are designed to accept vectors, which makes them useful when preparing tables or formatting several results at once.

p_values <- c(0.032, 0.006, 0.0004, 0.213)

format_p(p_values)
#> [1] "p = .032" "p = .006" "p < .001" "p = .213"
correlations <- c(0.42, -0.31, 0.18)

format_r(correlations)
#> [1] "r = .42"  "r = -.31" "r = .18"
proportions <- c(0.24, 0.51, 0.83)

format_percent(proportions)
#> [1] "24.0%" "51.0%" "83.0%"

Missing values are preserved rather than silently converted into text.

format_p(c(0.032, NA, 0.213))
#> [1] "p = .032" NA         "p = .213"

Design philosophy

The package separates three concepts that can otherwise become conflated:

  1. Statistical calculation — performed by functions such as lm(), cor(), confint(), or other statistical packages.
  2. Statistical decisions — made using the original numeric results and the researcher’s chosen inferential criteria.
  3. Publication formatting — performed by pubformat after the analysis is complete.

This separation allows results to be formatted consistently without changing the values on which statistical conclusions are based.