title: “Using pubformat” output: rmarkdown::html_vignette vignette: > % % % ————————-
pubformat provides tools for converting numeric
statistical results into consistent, publication-ready character
output.
The package is designed around a simple principle:
Statistical analysis and statistical formatting should remain separate.
R or another statistical procedure produces the numeric result.
pubformat handles the reporting step without changing the
underlying statistical result or decision.
library(pubformat)
format_p() converts numeric p-values into
publication-ready text.
format_p(0.048)
#> [1] "p = .048"
format_p(0.0002)
#> [1] "p < .001"
The reporting threshold is strict. A value equal to .001
is retained, while a value below .001 is reported using
less-than notation.
format_p(0.001)
#> [1] "p = .001"
format_p(0.0009)
#> [1] "p < .001"
Optional significance codes can be added when useful for tables or other compact displays.
format_p(0.048, sig = TRUE)
#> [1] "p = .048 *"
format_p(0.006, sig = TRUE)
#> [1] "p = .006 **"
format_p(0.0002, sig = TRUE)
#> [1] "p < .001 ***"
Importantly, significance codes are determined from the original numeric p-value rather than its rounded display value.
format_p(0.0499, sig = TRUE)
#> [1] "p = .050 *"
format_ci() combines numeric lower and upper confidence
limits into publication-ready interval notation.
format_ci(1.08, 1.87)
#> [1] "95% CI [1.08, 1.87]"
The confidence level can be changed.
format_ci(1.08, 1.87, level = 0.99)
#> [1] "99% CI [1.08, 1.87]"
Leading zeros are retained by default because the appropriate convention depends on the statistic being reported.
format_ci(0.21, 0.48)
#> [1] "95% CI [0.21, 0.48]"
For statistics that cannot exceed 1 in absolute value, such as correlations, leading zeros can be omitted.
format_ci(0.21, 0.48, leading_zero = FALSE)
#> [1] "95% CI [.21, .48]"
For Pearson correlations, format_r() provides a concise
formatter.
format_r(0.42)
#> [1] "r = .42"
format_r(-0.31)
#> [1] "r = -.31"
For Pearson, Spearman, or Kendall correlations,
format_cor() selects the appropriate reporting symbol.
format_cor(0.42, method = "pearson")
#> [1] "r = .42"
format_cor(0.42, method = "spearman")
#> [1] "ρ = .42"
format_cor(0.42, method = "kendall")
#> [1] "τ = .42"
format_cor() formats an existing coefficient. It does
not calculate the correlation itself.
For example:
rho <- cor(
mtcars$mpg,
mtcars$wt,
method = "spearman"
)
format_cor(rho, method = "spearman")
#> [1] "ρ = -.89"
format_percent() interprets numeric values as
proportions by default.
format_percent(0.423)
#> [1] "42.3%"
Values that are already expressed as percentages can be identified explicitly.
format_percent(42.3, input = "percent")
#> [1] "42.3%"
This distinction prevents accidental rescaling of values.
format_percent(0.42)
#> [1] "42.0%"
format_percent(42, input = "percent")
#> [1] "42.0%"
The function is vectorized.
format_percent(c(0.25, 0.50, 0.75))
#> [1] "25.0%" "50.0%" "75.0%"
Consider a simple linear model using the built-in mtcars
data.
model <- lm(mpg ~ wt, data = mtcars)
summary(model)
#>
#> Call:
#> lm(formula = mpg ~ wt, data = mtcars)
#>
#> Residuals:
#> Min 1Q Median 3Q Max
#> -4.5432 -2.3647 -0.1252 1.4096 6.8727
#>
#> Coefficients:
#> Estimate Std. Error t value Pr(>|t|)
#> (Intercept) 37.2851 1.8776 19.858 < 2e-16 ***
#> wt -5.3445 0.5591 -9.559 1.29e-10 ***
#> ---
#> Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
#>
#> Residual standard error: 3.046 on 30 degrees of freedom
#> Multiple R-squared: 0.7528, Adjusted R-squared: 0.7446
#> F-statistic: 91.38 on 1 and 30 DF, p-value: 1.294e-10
The analysis produces numeric statistical results. Those values can then be extracted and formatted separately.
p_value <- summary(model)$coefficients["wt", "Pr(>|t|)"]
p_value
#> [1] 1.293959e-10
format_p(p_value)
#> [1] "p < .001"
ci <- confint(model, "wt")
ci
#> 2.5 % 97.5 %
#> wt -6.486308 -4.202635
format_ci(ci[1], ci[2])
#> [1] "95% CI [-6.49, -4.20]"
r_value <- cor(mtcars$mpg, mtcars$wt)
r_value
#> [1] -0.8676594
format_r(r_value)
#> [1] "r = -.87"
The workflow is therefore:
statistical analysis
|
v
numeric statistical results
|
v
pubformat
|
v
publication-ready output
pubformat does not replace statistical analysis
functions. Its purpose is to make the final reporting step more
consistent, reproducible, and convenient.
The formatting functions are designed to accept vectors, which makes them useful when preparing tables or formatting several results at once.
p_values <- c(0.032, 0.006, 0.0004, 0.213)
format_p(p_values)
#> [1] "p = .032" "p = .006" "p < .001" "p = .213"
correlations <- c(0.42, -0.31, 0.18)
format_r(correlations)
#> [1] "r = .42" "r = -.31" "r = .18"
proportions <- c(0.24, 0.51, 0.83)
format_percent(proportions)
#> [1] "24.0%" "51.0%" "83.0%"
Missing values are preserved rather than silently converted into text.
format_p(c(0.032, NA, 0.213))
#> [1] "p = .032" NA "p = .213"
The package separates three concepts that can otherwise become conflated:
lm(), cor(), confint(),
or other statistical packages.pubformat after the analysis is complete.This separation allows results to be formatted consistently without changing the values on which statistical conclusions are based.