Case study: judge leniency and pretrial detention (Miami-Dade)

This vignette runs the package on the empirical application of Frandsen, Leslie and McIntyre (2025): weekend bail hearings in Miami-Dade County, where quasi-randomly assigned bail judges differ in leniency, judge identities instrument for pretrial release, and errors are clustered by courtroom shift. It is a large-scale worked example: 91,421 defendants, 146 judges, and several thousand shift clusters. The companion vignette vignette("queens-workflow") introduces the same workflow on a smaller design.

This vignette is precomputed: the data file cannot be redistributed inside an R package, so the code below was executed by the maintainer against a local copy and the outputs are stored. Every displayed result comes from the displayed code.

Data

The analysis file clean_Miami_data.dta is the cleaned weekend-arraignment file from the published replication materials of Frandsen, Leslie and McIntyre (2025, Review of Economics and Statistics); the underlying case records are those of Dobbie, Goldin and Yang (2018). Obtain it from the journal’s replication archive and set data_path to your local copy:

data_path <- "clean_Miami_data.dta"

The file is one row per defendant. Following the paper, the sample is restricted to judges with at least 200 cases, and a cluster is a courtroom shift (court by hearing date):

library(clusterIV)

df <- as.data.frame(haven::read_dta(data_path))
df <- df[df$n_cases >= 200, , drop = FALSE]        # judges with >= 200 cases
shift <- as.integer(factor(paste(df$court, df$clean_baildate)))

y <- as.numeric(df$any_guilty)                     # outcome
x <- as.numeric(df$bail_met)                       # endogenous: met bail
Z <- as.matrix(df[, startsWith(names(df), "jfe_")])  # judge dummies
W <- as.matrix(df[, startsWith(names(df), "fe_")])   # time-by-place dummies
storage.mode(Z) <- "double"; storage.mode(W) <- "double"

c(n = nrow(df), judges = length(unique(df$judgeid)),
  shifts = length(unique(shift)))
#>      n judges shifts 
#>  91421    146   1831

Judge-dummy instrument sets need one hygiene step in any software: after the exogenous controls are partialled out, dummies for judges who never appear in the estimation sample (or are absorbed by the controls) carry no variation, and if every in-sample judge keeps a dummy their sum is collinear with the intercept. The package deliberately reports rank-deficient instrument matrices as errors instead of silently dropping columns, so the drop is made explicit:

qC   <- qr(cbind(1, W), tol = 1e-10, LAPACK = FALSE)
Zres <- qr.resid(qC, Z)
keep <- sqrt(colSums(Zres^2)) > 1e-8 * max(sqrt(colSums(Zres^2)))
Z    <- Z[, keep, drop = FALSE]
Zres <- Zres[, keep, drop = FALSE]
qZ   <- qr(Zres, tol = 1e-10, LAPACK = FALSE)
if (qZ$rank < ncol(Zres)) {
  keep_idx <- sort(qZ$pivot[seq_len(qZ$rank)])
  Z <- Z[, keep_idx, drop = FALSE]
}
ncol(Z)   # instrument columns entering the analysis
#> [1] 145

Estimator comparison

iv_compare() reports OLS, 2SLS, the improved jackknife IV (IJIVE row, historically labelled JIVE), and CJIVE on the identical design. The time-by-place dummies enter as dense controls, exactly as in the paper:

cmp <- iv_compare(y, x, Z,
                  cluster  = shift,   # courtroom-shift clusters
                  controls = W)       # time-by-place fixed effects
print(cmp, digits = 3)
#>   estimator coefficient      se statistic   p.value conf.low conf.high
#> 1       OLS      -0.233 0.00658    -35.33 2.25e-273   -0.245   -0.2196
#> 2      2SLS      -0.274 0.06567     -4.18  2.96e-05   -0.403   -0.1456
#> 3      JIVE      -0.298 0.10671     -2.79  5.22e-03   -0.507   -0.0889
#> 4     CJIVE      -0.438 0.20206     -2.17  3.01e-02   -0.834   -0.0422

The published Table 1 values are OLS −0.232 (0.007), 2SLS −0.275 (0.066), IJIVE −0.299 (0.107), and CJIVE −0.444 (0.206). The OLS, 2SLS and IJIVE rows above reproduce the published values to three decimals. The CJIVE estimate is validated differently: run on this deposited sample, the authors’ own released Stata implementation (cjive.ado) produces the same constructed leave-cluster-out instrument pointwise and the same coefficient to the ado’s single-precision limit (about 1e−7), which is the strongest available check that the package computes the estimator the paper defines.

The CJIVE fit and its diagnostics

fit <- cjive(y, x, Z, cluster = shift, controls = W)
summary(fit)
#> Cluster-jackknife IV (CJIVE)
#> Call: cjive.default(y = y, x = x, z = Z, cluster = shift, controls = W)
#> 
#>   coefficient = -0.4383   cluster-robust SE = 0.2021
#>   z = -2.169   p = 0.03008   95% CI = [-0.8343, -0.04225]
#>   n = 91421   G = 1831 clusters   k = 145 instruments   path = dense
#>   max within-cluster leverage = 0.341
#> 
#> Coefficients:
#>   Estimate Std. Error z value Pr(>|z|)  
#> x  -0.4383     0.2021  -2.169   0.0301 *
#> ---
#> Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
#> 
#> Instrument strength:
#>   F_CJ requires a cjar() fit on the same design; iv_infer() reports the
#>   estimate, both strength statistics and the robust confidence set in one call.
#>   F_eff = 1.69   vs critical value 12.41   (Montiel Olea-Pflueger effective F, simplified TSLS, tau = 10%, alpha = 5%, K_eff = 59.61)
#>   (Confidence-set topology is read from the endpoint matrix printed
#>    above; no reporting branch infers it from F_eff.)

Two diagnostics matter in a design of this size. maxlev is the maximum within-cluster leverage of the first stage; values near 1 mean some shift nearly spans the instrument space and the leave-cluster-out fit is at the conditioning frontier. F_eff is the clustered Montiel Olea–Pflueger effective first-stage F with its simplified-TSLS critical value.

Weak-instrument-robust inference

Frandsen, Leslie and McIntyre publish no weak-instrument-robust interval for this application. With 100+ instruments, the CJIVE t-interval’s validity rests on instrument strength; the cluster-jackknife Anderson–Rubin and score tests of Ligtenberg (2025) stay valid when instruments are weak or many. iv_infer() returns the estimate and both tests from one shared preprocessing pass:

panel <- iv_infer(y, x, Z, cluster = shift, controls = W)
panel
#> Cluster IV inference panel (CJIVE + CJAR + CJS)
#> Call: iv_infer.default(y = y, x = x, z = Z, cluster = shift, controls = W)
#> 
#>   CJIVE/Wald (H0: beta = 0): coefficient = -0.4383   cluster-robust SE = 0.2021   z = -2.169   p = 0.03008
#>     95% Wald interval = [-0.8343, -0.04225]
#>   CJAR (H0: beta = 0): T = 3.681   one-sided p = 0.0004972
#>     95% confidence set = {}  <- empty: no beta is accepted at this level
#>   CJS (H0: beta = 0): LM = 4.759   p = 0.02914
#>     95% confidence set = [-0.9866, -0.05018]
#>   F_CJ = 4.237   critical value = 1.709
#>   F_CJS^2 = 13.9   critical value = 3.841
#>   effective F (Montiel Olea-Pflueger) = 1.69   critical value (tau = 10%, alpha = 5%) = 12.41   K_eff = 59.61
#>   n = 91421   G = 1831 clusters   k = 145 instruments
#>   max within-cluster leverage = 0.341

This panel is a lesson in reading weak-instrument-robust output rather than a clean confirmation. The effective F (1.69) is far below its critical value: with 145 instrument columns the first stage is weak by the Montiel Olea–Pflueger standard, so the Wald interval’s nominal coverage is not guaranteed — precisely the regime the robust tests are for. The two robust results then differ in kind. The CJS test gives a bounded 95% set that excludes zero and comfortably contains the CJIVE estimate. The CJAR set is empty: no coefficient value is accepted at the 5% level. An empty AR set is a valid outcome, not a numerical failure — the AR statistic aggregates all 145 moment conditions, so it also has power against violations of the exclusion restrictions themselves, and with this much overidentification it can reject every candidate coefficient. It should be read as evidence of tension in the full instrument set rather than as an interval estimate (see the FAQ in the README). The sets are computed by analytic polynomial inversion, not a parameter grid, so empty, disjoint, and unbounded outcomes are exact statements. The p-value curves make all of this visible:

plot(panel)
plot of chunk pcurve
plot of chunk pcurve

Notes on scale

At this size (n = 91,421, k > 100 instruments, dense controls) the whole panel above runs in minutes on a laptop. The package keeps Imports: limited to stats; the loading step above uses haven only to read the Stata file, which is not a package dependency. For designs with many more fixed-effect levels, pass them as fixed_effects = factors instead of dense dummy columns: the matrix-free absorption never forms the dummy matrix. Note the caveat in ?cjive: with many absorbed controls relative to the sample, the plug-in CJIVE standard error can over-reject (Kolesár, Min, Wang and Zhang 2026); the time-by-place controls here are few relative to n.

References

Dobbie, W., Goldin, J. and Yang, C. S. (2018). The effects of pretrial detention on conviction, future crime, and employment. American Economic Review, 108(2), 201–240. The underlying Miami-Dade case records.

Frandsen, B., Leslie, E. and McIntyre, S. (2025). Cluster jackknife instrumental variables estimation. Review of Economics and Statistics. doi:10.1162/rest.a.263. The estimator, the application, and the replication materials used here.

Ligtenberg, J. W. (2025). Inference in clustered IV models with many and weak instruments. arXiv:2306.08559v3. The CJAR and CJS tests reported by iv_infer().