Getting started with causalfrag

Overview

causalfrag runs sensitivity analyses for unmeasured confounding across several frameworks, classifies and interprets the results, and writes structured reports. Statistical results always come from transparent R functions; an optional language model only rephrases template text, and the package works fully offline.

The OLS crosswalk

Several widely reported sensitivity statistics are computed from the same two numbers. Under ordinary least squares with model-based standard errors, the point and inference Robustness Values (Cinelli and Hazlett, 2020), the conditional ITCV (Frank, 2000) and the RIR percentage (Frank et al., 2013) are all functions of the focal |t| and the residual degrees of freedom. Reporting them together is useful, but their agreement is largely fixed by construction and should not be described as several methods independently corroborating a conclusion. crosswalk() makes that explicit.

library(causalfrag)
fit <- lm(mpg ~ am + wt + hp, data = mtcars)
crosswalk(fit, treatment = "am")
#> 
#> ── OLS crosswalk: statistics determined by |t| and df ──
#> 
#>   t for 'am' = 1.5139, residual df = 28, alpha = 0.05
#> 
#>   Statistic                           Value   Target
#>   Partial correlation r              0.2751   effect size
#>   Robustness Value (point)           0.2481   point estimate to zero
#>   Robustness Value (inference)       0.0000   loss of significance
#>   Conditional ITCV                  -0.1345   loss of significance
#>   RIR (percent)                      -35.3%   loss of significance
#> ℹ All five values are functions of the same |t| and df. Agreement among them is largely fixed by construction and is not independent corroboration.
#> ! The estimate is not significant at this alpha; ITCV and RIR then describe the change needed to reach significance.
#> ℹ E-values, Oster's delta and benchmark bounds use additional information and are outside this crosswalk.

crosswalk() also accepts a t statistic and degrees of freedom directly:

crosswalk(4.2228, df = 5668)
#> 
#> ── OLS crosswalk: statistics determined by |t| and df ──
#> 
#>   t = 4.2228, residual df = 5668, alpha = 0.05
#> 
#>   Statistic                           Value   Target
#>   Partial correlation r              0.0560   effect size
#>   Robustness Value (point)           0.0545   point estimate to zero
#>   Robustness Value (inference)       0.0296   loss of significance
#>   Conditional ITCV                   0.0308   loss of significance
#>   RIR (percent)                       53.6%   loss of significance
#> 
#>   RIR as cases: about 3038 of 5670 observations would need to be replaced.
#> ℹ All five values are functions of the same |t| and df. Agreement among them is largely fixed by construction and is not independent corroboration.
#> ℹ E-values, Oster's delta and benchmark bounds use additional information and are outside this crosswalk.

Just above the conventional significance threshold there is a narrow band in which ITCV and RIR are already positive while the inference Robustness Value is still zero, because the two conventions use slightly different degrees of freedom in their critical values. crosswalk() flags it:

crosswalk(1.972, df = 244)
#> 
#> ── OLS crosswalk: statistics determined by |t| and df ──
#> 
#>   t = 1.972, residual df = 244, alpha = 0.05
#> 
#>   Statistic                           Value   Target
#>   Partial correlation r              0.1253   effect size
#>   Robustness Value (point)           0.1185   point estimate to zero
#>   Robustness Value (inference)       0.0000   loss of significance
#>   Conditional ITCV                   0.0002   loss of significance
#>   RIR (percent)                        0.1%   loss of significance
#> 
#>   RIR as cases: about 0 of 246 observations would need to be replaced.
#> ℹ All five values are functions of the same |t| and df. Agreement among them is largely fixed by construction and is not independent corroboration.
#> ! Boundary band: |t| lies between 1.9697 and 1.9738, so ITCV and RIR are positive while the inference Robustness Value is still zero. The two conventions use different degrees of freedom in their critical values.
#> ℹ E-values, Oster's delta and benchmark bounds use additional information and are outside this crosswalk.

E-values, Oster’s delta and benchmark-based bounds use information beyond |t| and the degrees of freedom and are outside the crosswalk.

Running the frameworks

result <- fragility_report(model = fit, treatment = "am", data = mtcars)
print(result)        # includes the crosswalk for lm fits
cat(result$narrative)

Or step by step:

res <- run_sensitivity(fit, treatment = "am", data = mtcars)
res <- flag_fragility(res)          # per-framework classification
res <- interpret_sensitivity(res)   # template narrative (or LLM if configured)
cat(generate_report(res))

Deprecated: the Causal Fragility Index

Earlier versions combined these statistics into a 0–100 Causal Fragility Index. Because most of its components share the same test information, the composite counted that information several times. The CFI functions are deprecated as of version 0.2.0 and will be removed; use crosswalk() and the per-framework classifications instead.

sens_report() was renamed fragility_report() so that it no longer masks confoundvis::sens_report().

References

Cinelli, C., & Hazlett, C. (2020). Making sense of sensitivity: Extending omitted variable bias. Journal of the Royal Statistical Society: Series B, 82(1), 39–67.

Frank, K. A. (2000). Impact of a confounding variable on a regression coefficient. Sociological Methods & Research, 29(2), 147–194.

Frank, K. A., Maroulis, S. J., Duong, M. Q., & Kelcey, B. M. (2013). What would it take to change an inference? Using Rubin’s causal model to interpret the robustness of causal inferences. Educational Evaluation and Policy Analysis, 35(4), 437–460.

VanderWeele, T. J., & Ding, P. (2017). Sensitivity analysis in observational research: Introducing the E-value. Annals of Internal Medicine, 167(4), 268–274.