Calls to Azure show output recorded from a live run, and setup code is shown but not run.
This article codes course evaluation comments for one primary theme and sentiment, then turns those labels into an estimate. The same pattern applies to interview transcripts, support tickets and open-ended survey answers. The sections below add, in order, a versioned codebook, readable labels, a check against hand codes, a stability check, an interval and a provenance record. A reviewer can ask about each one.
The codebook is the measurement instrument. It records the instructions, schema, examples and version that define what a label means.
comments <- tibble::tibble(
comment_id = sprintf("c%02d", 1:10),
comment = c(
"The lectures were clear and the examples made regression feel concrete.",
"The weekly quizzes felt rushed and did not match the homework.",
"Office hours helped me catch up after I missed the first lab.",
"The slides were hard to follow because notation changed between weeks.",
"The final project connected the material to real policy questions.",
"I needed more feedback before the midterm.",
"The instructor explained difficult topics patiently.",
"The reading packet was useful, but several links were broken.",
"Group work helped, although the grading rubric came too late.",
"More examples before the final exam would have helped."
)
)course_schema <- foundry_schema(
theme = schema_enum(
c("instruction", "assessment", "support", "materials"),
description = paste(
"Primary theme: instruction, assessment, support, or materials."
)
),
sentiment = schema_enum(
c("positive", "negative", "mixed"),
description = "Overall sentiment toward the course element."
)
)
course_instructions <- paste(
"Code one course evaluation comment.",
"Choose exactly one primary theme.",
"Use instruction for teaching clarity or examples.",
"Use assessment for quizzes, exams, projects, grading or feedback.",
"Use support for office hours or help outside class.",
"Use materials for slides, readings, links or course files.",
"Choose positive, negative or mixed sentiment from the student's wording."
)
course_codebook <- foundry_codebook(
name = "course-evaluation-codes",
version = "1.0.0",
instructions = course_instructions,
schema = course_schema,
examples = list(
list(
text = "The lectures were clear.",
theme = "instruction",
sentiment = "positive"
),
list(
text = "The rubric came too late.",
theme = "assessment",
sentiment = "negative"
)
)
)
course_codebook
#> foundry codebook: course-evaluation-codes
#> version: 1.0.0
#> hash: 44dc8b370cb4
#> variables:
#> - theme: string [instruction, assessment, support, materials] (Primary theme: instruction, assessment, support, or materials.)
#> - sentiment: string [positive, negative, mixed] (Overall sentiment toward the course element.)
#> examples: 2Versioning matters because label changes change the estimand. If “workload” becomes a valid theme, yesterday’s “assessment” labels are no longer directly comparable with tomorrow’s labels.
course_schema_v2 <- foundry_schema(
theme = schema_enum(
c("instruction", "assessment", "support", "materials", "workload"),
description = paste(
"Primary theme: instruction, assessment, support, materials, or workload."
)
),
sentiment = schema_enum(
c("positive", "negative", "mixed"),
description = "Overall sentiment toward the course element."
)
)
course_codebook_v2 <- foundry_codebook(
name = "course-evaluation-codes",
version = "1.1.0",
instructions = paste(
course_instructions,
"Use workload for comments about pacing or volume that are not mainly assessment."
),
schema = course_schema_v2,
examples = course_codebook$examples
)
codebook_diff(course_codebook, course_codebook_v2)
#> Codebook diff
#> old: course-evaluation-codes 1.0.0 44dc8b370cb45b83fccb26264225d07b7af182a050fd67a6e222a6776e4dad89
#> new: course-evaluation-codes 1.1.0 25a223fd7de1c485c3112ba0bf59597ad13322c9dffa6f4b0612f14f16f27282
#>
#> Instructions:
#> --- old instructions
#> +++ new instructions
#> @@
#> -Code one course evaluation comment. Choose exactly one primary theme. Use instruction for teaching clarity or examples. Use assessment for quizzes, exams, projects, grading or feedback. Use support for office hours or help outside class. Use materials for slides, readings, links or course files. Choose positive, negative or mixed sentiment from the student's wording.
#> +Code one course evaluation comment. Choose exactly one primary theme. Use instruction for teaching clarity or examples. Use assessment for quizzes, exams, projects, grading or feedback. Use support for office hours or help outside class. Use materials for slides, readings, links or course files. Choose positive, negative or mixed sentiment from the student's wording. Use workload for comments about pacing or volume that are not mainly assessment.
#>
#> Schema:
#> ~ theme.description: "Primary theme: instruction, assessment, support, or materials." -> "Primary theme: instruction, assessment, support, materials, or workload."
#> ~ theme.enum: +workload
#> sentiment: no change
#>
#> Examples:
#> (no changes)The diff is short enough for a methods appendix or code review. It shows the schema and instruction changes behind any shift in label rates.
foundry_extract() returns input columns, extracted
fields, dot-prefixed metadata and the raw response. Print the fields
needed for the next decision, not the whole object.
coding_model <- "gpt-5-nano" # your deployment name
model_labels <- foundry_extract(
comments,
text_col = "comment",
schema = course_codebook$schema,
instructions = course_codebook$instructions,
model = coding_model
)
model_labels |>
select(comment_id, theme, sentiment)
#> # A tibble: 10 × 3
#> comment_id theme sentiment
#> <chr> <chr> <chr>
#> 1 c01 instruction positive
#> 2 c02 assessment negative
#> 3 c03 support positive
#> 4 c04 materials negative
#> 5 c05 assessment positive
#> 6 c06 assessment negative
#> 7 c07 instruction positive
#> 8 c08 materials mixed
#> 9 c09 assessment mixed
#> 10 c10 instruction negativeReadable output is a quality control step. It lets a reviewer notice a systematic problem, such as materials comments being coded as instruction, before the labels become an estimate.
The next check is agreement with a reference sample. The labels below are hand codes written for this example. In a study, draw this sample before looking at model errors, and keep a second labeled sample if you tune the prompt or codebook.
hand_codes <- tibble::tibble(
comment_id = c("c01", "c02", "c03", "c04", "c05", "c06"),
human_theme = c(
"instruction",
"assessment",
"support",
"materials",
"assessment",
"assessment"
),
human_sentiment = c(
"positive",
"negative",
"positive",
"negative",
"positive",
"negative"
)
)
hand_codes
#> # A tibble: 6 × 3
#> comment_id human_theme human_sentiment
#> <chr> <chr> <chr>
#> 1 c01 instruction positive
#> 2 c02 assessment negative
#> 3 c03 support positive
#> 4 c04 materials negative
#> 5 c05 assessment positive
#> 6 c06 assessment negativevalidation_sample <- model_labels |>
select(comment_id, theme, sentiment) |>
inner_join(hand_codes, by = "comment_id")
theme_agreement <- foundry_agreement(
validation_sample,
estimate = "theme",
truth = "human_theme"
) |>
mutate(variable = "theme")
sentiment_agreement <- foundry_agreement(
validation_sample,
estimate = "sentiment",
truth = "human_sentiment"
) |>
mutate(variable = "sentiment")
bind_rows(theme_agreement, sentiment_agreement) |>
select(variable, metric, value, n)
#> # A tibble: 12 × 4
#> variable metric value n
#> <chr> <chr> <dbl> <int>
#> 1 theme accuracy 1 6
#> 2 theme precision_macro 1 6
#> 3 theme recall_macro 1 6
#> 4 theme f1_macro 1 6
#> 5 theme cohen_kappa 1 6
#> 6 theme krippendorff_alpha 1 6
#> 7 sentiment accuracy 1 6
#> 8 sentiment precision_macro 1 6
#> 9 sentiment recall_macro 1 6
#> 10 sentiment f1_macro 1 6
#> 11 sentiment cohen_kappa 1 6
#> 12 sentiment krippendorff_alpha 1 6All 6 theme pairs agree in this recording. That says little on its own, because 6 comments written to be unambiguous are an easy test. A real validation sample is drawn at random from the data being coded and is large enough for an interval on accuracy to be informative.
Accuracy is the share of complete pairs where the model and human label match. Macro precision, recall and F1 compute a per-class value and average across classes, so a rare class can matter as much as a common class. foundryR follows the yardstick convention: when a class has an undefined denominator for a macro metric, that class is dropped from that macro average with a warning.
Cohen’s kappa discounts the agreement you would expect by chance, given how often each coder uses each label. foundryR computes Krippendorff’s alpha for nominal labels from two coders. Both are useful when a high raw accuracy could come from a dominant class. None of these metrics proves the hand codes are true, unbiased or complete. They measure agreement with the reference labels you supplied.
Agreement checks accuracy against people. Stability checks whether the same instrument gives the same label when it is run again. A model can be stable and wrong.
stability <- foundry_consistency(
comments$comment[1:4],
schema = course_codebook$schema,
n = 3,
instructions = course_codebook$instructions,
model = coding_model
)
stability |>
select(.input_idx, successful_runs, failed_runs, modal_share, entropy)
#> # A tibble: 4 × 5
#> .input_idx successful_runs failed_runs modal_share entropy
#> <int> <int> <int> <dbl> <dbl>
#> 1 1 3 0 1 0
#> 2 2 3 0 1 0
#> 3 3 3 0 1 0
#> 4 4 3 0 1 0In this recording 4 of 4 comments received the same record in all 3 runs. Short, clear comments are the easy case; run the same check on the ambiguous comments your hand coders disagreed about.
modal_share is the largest repeated-label pattern’s
share of successful runs. With n = 3 and all three runs
successful, it can only be one third, two thirds or one.
entropy is in bits and increases when repeated runs split
across several distinct records. These values summarize stability only.
They do not say whether the modal label matches a human code.
The final estimate should carry the instrument and model that produced it. That record lets another analyst connect a table of estimates back to the exact schema.
run_provenance <- foundry_provenance(
model = coding_model,
schema = course_codebook$schema,
metadata = list(
codebook = course_codebook$name,
codebook_version = course_codebook$version,
codebook_hash = course_codebook$hash
)
)
run_provenance |>
select(model, schema_hash, package_version, captured_at)
#> # A tibble: 1 × 4
#> model schema_hash package_version captured_at
#> <chr> <chr> <chr> <dttm>
#> 1 gpt-5-nano c5ee5dc66f90546f2c3f053bd61847… 1.0.0 2026-10-01 15:58:32The codebook hash and schema hash are not substitutes for archiving the codebook. They are compact checks that the labels and report refer to the same instrument.
For a defensible text-to-estimate workflow, report the choices that affect the estimate. A short methods note should cover:
n, successful runs,
failed runs, modal share and entropy;