This vignette runs the package on the empirical application of
Frandsen, Leslie and McIntyre (2025): weekend bail hearings in
Miami-Dade County, where quasi-randomly assigned bail judges differ in
leniency, judge identities instrument for pretrial release, and errors
are clustered by courtroom shift. It is a large-scale worked example:
91,421 defendants, 146 judges, and several thousand shift clusters. The
companion vignette vignette("queens-workflow") introduces
the same workflow on a smaller design.
This vignette is precomputed: the data file cannot be redistributed inside an R package, so the code below was executed by the maintainer against a local copy and the outputs are stored. Every displayed result comes from the displayed code.
The analysis file clean_Miami_data.dta is the cleaned
weekend-arraignment file from the published replication materials of
Frandsen, Leslie and McIntyre (2025, Review of Economics and
Statistics); the underlying case records are those of Dobbie,
Goldin and Yang (2018). Obtain it from the journal’s replication archive
and set data_path to your local copy:
The file is one row per defendant. Following the paper, the sample is restricted to judges with at least 200 cases, and a cluster is a courtroom shift (court by hearing date):
library(clusterIV)
df <- as.data.frame(haven::read_dta(data_path))
df <- df[df$n_cases >= 200, , drop = FALSE] # judges with >= 200 cases
shift <- as.integer(factor(paste(df$court, df$clean_baildate)))
y <- as.numeric(df$any_guilty) # outcome
x <- as.numeric(df$bail_met) # endogenous: met bail
Z <- as.matrix(df[, startsWith(names(df), "jfe_")]) # judge dummies
W <- as.matrix(df[, startsWith(names(df), "fe_")]) # time-by-place dummies
storage.mode(Z) <- "double"; storage.mode(W) <- "double"
c(n = nrow(df), judges = length(unique(df$judgeid)),
shifts = length(unique(shift)))
#> n judges shifts
#> 91421 146 1831Judge-dummy instrument sets need one hygiene step in any software: after the exogenous controls are partialled out, dummies for judges who never appear in the estimation sample (or are absorbed by the controls) carry no variation, and if every in-sample judge keeps a dummy their sum is collinear with the intercept. The package deliberately reports rank-deficient instrument matrices as errors instead of silently dropping columns, so the drop is made explicit:
qC <- qr(cbind(1, W), tol = 1e-10, LAPACK = FALSE)
Zres <- qr.resid(qC, Z)
keep <- sqrt(colSums(Zres^2)) > 1e-8 * max(sqrt(colSums(Zres^2)))
Z <- Z[, keep, drop = FALSE]
Zres <- Zres[, keep, drop = FALSE]
qZ <- qr(Zres, tol = 1e-10, LAPACK = FALSE)
if (qZ$rank < ncol(Zres)) {
keep_idx <- sort(qZ$pivot[seq_len(qZ$rank)])
Z <- Z[, keep_idx, drop = FALSE]
}
ncol(Z) # instrument columns entering the analysis
#> [1] 145iv_compare() reports OLS, 2SLS, the improved jackknife
IV (IJIVE row, historically labelled JIVE), and CJIVE on the identical
design. The time-by-place dummies enter as dense controls, exactly as in
the paper:
cmp <- iv_compare(y, x, Z,
cluster = shift, # courtroom-shift clusters
controls = W) # time-by-place fixed effects
print(cmp, digits = 3)
#> estimator coefficient se statistic p.value conf.low conf.high
#> 1 OLS -0.233 0.00658 -35.33 2.25e-273 -0.245 -0.2196
#> 2 2SLS -0.274 0.06567 -4.18 2.96e-05 -0.403 -0.1456
#> 3 JIVE -0.298 0.10671 -2.79 5.22e-03 -0.507 -0.0889
#> 4 CJIVE -0.438 0.20206 -2.17 3.01e-02 -0.834 -0.0422The published Table 1 values are OLS −0.232 (0.007), 2SLS −0.275
(0.066), IJIVE −0.299 (0.107), and CJIVE −0.444 (0.206). The OLS, 2SLS
and IJIVE rows above reproduce the published values to three decimals.
The CJIVE estimate is validated differently: run on this deposited
sample, the authors’ own released Stata implementation
(cjive.ado) produces the same constructed leave-cluster-out
instrument pointwise and the same coefficient to the ado’s
single-precision limit (about 1e−7), which is the strongest available
check that the package computes the estimator the paper defines.
fit <- cjive(y, x, Z, cluster = shift, controls = W)
summary(fit)
#> Cluster-jackknife IV (CJIVE)
#> Call: cjive.default(y = y, x = x, z = Z, cluster = shift, controls = W)
#>
#> coefficient = -0.4383 cluster-robust SE = 0.2021
#> z = -2.169 p = 0.03008 95% CI = [-0.8343, -0.04225]
#> n = 91421 G = 1831 clusters k = 145 instruments path = dense
#> max within-cluster leverage = 0.341
#>
#> Coefficients:
#> Estimate Std. Error z value Pr(>|z|)
#> x -0.4383 0.2021 -2.169 0.0301 *
#> ---
#> Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
#>
#> Instrument strength:
#> F_CJ requires a cjar() fit on the same design; iv_infer() reports the
#> estimate, both strength statistics and the robust confidence set in one call.
#> F_eff = 1.69 vs critical value 12.41 (Montiel Olea-Pflueger effective F, simplified TSLS, tau = 10%, alpha = 5%, K_eff = 59.61)
#> (Confidence-set topology is read from the endpoint matrix printed
#> above; no reporting branch infers it from F_eff.)Two diagnostics matter in a design of this size. maxlev
is the maximum within-cluster leverage of the first stage; values near 1
mean some shift nearly spans the instrument space and the
leave-cluster-out fit is at the conditioning frontier.
F_eff is the clustered Montiel Olea–Pflueger effective
first-stage F with its simplified-TSLS critical value.
Frandsen, Leslie and McIntyre publish no weak-instrument-robust
interval for this application. With 100+ instruments, the CJIVE
t-interval’s validity rests on instrument strength; the
cluster-jackknife Anderson–Rubin and score tests of Ligtenberg (2025)
stay valid when instruments are weak or many. iv_infer()
returns the estimate and both tests from one shared preprocessing
pass:
panel <- iv_infer(y, x, Z, cluster = shift, controls = W)
panel
#> Cluster IV inference panel (CJIVE + CJAR + CJS)
#> Call: iv_infer.default(y = y, x = x, z = Z, cluster = shift, controls = W)
#>
#> CJIVE/Wald (H0: beta = 0): coefficient = -0.4383 cluster-robust SE = 0.2021 z = -2.169 p = 0.03008
#> 95% Wald interval = [-0.8343, -0.04225]
#> CJAR (H0: beta = 0): T = 3.681 one-sided p = 0.0004972
#> 95% confidence set = {} <- empty: no beta is accepted at this level
#> CJS (H0: beta = 0): LM = 4.759 p = 0.02914
#> 95% confidence set = [-0.9866, -0.05018]
#> F_CJ = 4.237 critical value = 1.709
#> F_CJS^2 = 13.9 critical value = 3.841
#> effective F (Montiel Olea-Pflueger) = 1.69 critical value (tau = 10%, alpha = 5%) = 12.41 K_eff = 59.61
#> n = 91421 G = 1831 clusters k = 145 instruments
#> max within-cluster leverage = 0.341This panel is a lesson in reading weak-instrument-robust output rather than a clean confirmation. The effective F (1.69) is far below its critical value: with 145 instrument columns the first stage is weak by the Montiel Olea–Pflueger standard, so the Wald interval’s nominal coverage is not guaranteed — precisely the regime the robust tests are for. The two robust results then differ in kind. The CJS test gives a bounded 95% set that excludes zero and comfortably contains the CJIVE estimate. The CJAR set is empty: no coefficient value is accepted at the 5% level. An empty AR set is a valid outcome, not a numerical failure — the AR statistic aggregates all 145 moment conditions, so it also has power against violations of the exclusion restrictions themselves, and with this much overidentification it can reject every candidate coefficient. It should be read as evidence of tension in the full instrument set rather than as an interval estimate (see the FAQ in the README). The sets are computed by analytic polynomial inversion, not a parameter grid, so empty, disjoint, and unbounded outcomes are exact statements. The p-value curves make all of this visible:
At this size (n = 91,421, k > 100 instruments, dense controls) the
whole panel above runs in minutes on a laptop. The package keeps
Imports: limited to stats; the loading step
above uses haven only to read the Stata file, which is not
a package dependency. For designs with many more fixed-effect levels,
pass them as fixed_effects = factors instead of dense dummy
columns: the matrix-free absorption never forms the dummy matrix. Note
the caveat in ?cjive: with many absorbed controls relative
to the sample, the plug-in CJIVE standard error can over-reject
(Kolesár, Min, Wang and Zhang 2026); the time-by-place controls here are
few relative to n.
Dobbie, W., Goldin, J. and Yang, C. S. (2018). The effects of pretrial detention on conviction, future crime, and employment. American Economic Review, 108(2), 201–240. The underlying Miami-Dade case records.
Frandsen, B., Leslie, E. and McIntyre, S. (2025). Cluster jackknife instrumental variables estimation. Review of Economics and Statistics. doi:10.1162/rest.a.263. The estimator, the application, and the replication materials used here.
Ligtenberg, J. W. (2025). Inference in clustered IV models with many
and weak instruments. arXiv:2306.08559v3. The CJAR and CJS tests
reported by iv_infer().