Calls to Azure show output recorded from a live run, and setup code is shown but not run.
Batch is worth the extra setup when you have thousands of independent texts, can wait for the 24-hour processing window, and want the lower batch price instead of immediate responses. It is a production path for stable annotation jobs, not the first place to debug a new prompt.
Batch jobs need a deployment made for batch inference. Microsoft
Learn states that the Foundry portal shows batch deployment types as
Global-Batch and Data Zone Batch, and that the
model field in each request line must match the Global
Batch deployment name. This article uses
gpt-4.1-nano-batch, a Global Batch deployment of
gpt-4.1-nano. See Microsoft’s Azure
OpenAI batch deployments guide.
foundry_batch_requests() writes the JSONL file that the
service reads. The call is local, so it can run before credentials are
configured.
comments <- tibble::tibble(
comment_id = sprintf("c%02d", 1:6),
comment = c(
"The lectures were clear and the examples made regression feel concrete.",
"The weekly quizzes felt rushed and did not match the homework.",
"Office hours helped me catch up after I missed the first lab.",
"The slides were hard to follow because notation changed between weeks.",
"The final project connected the material to real policy questions.",
"I needed more feedback before the midterm."
)
)
course_schema <- foundry_schema(
theme = schema_enum(
c("instruction", "assessment", "support", "materials"),
description = "Primary course evaluation theme."
),
sentiment = schema_enum(
c("positive", "negative", "mixed"),
description = "Overall sentiment toward the course element."
)
)
course_instructions <- paste(
"Code one course evaluation comment.",
"Choose exactly one primary theme and one sentiment.",
"Return only fields that conform to the schema."
)jsonl <- tempfile(fileext = ".jsonl")
request_info <- foundry_batch_requests(
comments,
input = "comment",
path = jsonl,
model = "gpt-4.1-nano-batch",
custom_id = "comment_id",
schema = course_schema,
schema_name = "CourseEvaluationCodes",
instructions = course_instructions,
overwrite = TRUE
)
request_info |>
select(requests, endpoint)
#> # A tibble: 1 × 2
#> requests endpoint
#> <int> <chr>
#> 1 6 /v1/responsesPrint one request line, formatted for inspection, rather than printing the temporary path. The request line is the object you need to review.
jsonlite::prettify(readLines(jsonl, n = 1))
#> {
#> "custom_id": "c01",
#> "method": "POST",
#> "url": "/v1/responses",
#> "body": {
#> "model": "gpt-4.1-nano-batch",
#> "input": "The lectures were clear and the examples made regression feel concrete.",
#> "instructions": "Code one course evaluation comment. Choose exactly one primary theme and one sentiment. Return only fields that conform to the schema.",
#> "text": {
#> "format": {
#> "type": "json_schema",
#> "name": "CourseEvaluationCodes",
#> "schema": {
#> "type": "object",
#> "properties": {
#> "theme": {
#> "type": "string",
#> "description": "Primary course evaluation theme.",
#> "enum": [
#> "instruction",
#> "assessment",
#> "support",
#> "materials"
#> ]
#> },
#> "sentiment": {
#> "type": "string",
#> "description": "Overall sentiment toward the course element.",
#> "enum": [
#> "positive",
#> "negative",
#> "mixed"
#> ]
#> }
#> },
#> "required": [
#> "theme",
#> "sentiment"
#> ],
#> "additionalProperties": false
#> },
#> "strict": true
#> }
#> }
#> }
#> }
#> The custom_id is the join key. Use a stable ID from your
data instead of relying on file order.
The low-level path makes each service object visible: uploaded file, batch job, final status, parsed results and usage.
uploaded_file <- foundry_file_upload(jsonl, purpose = "batch")
batch <- foundry_batch_create(
input_file_id = uploaded_file$file_id,
endpoint = request_info$endpoint
)
batch |>
select(status, completion_window, batch_id)
#> # A tibble: 1 × 3
#> status completion_window batch_id
#> <chr> <chr> <chr>
#> 1 validating 24h batch_4c084697-726e-4b46-b252-54c5df527796foundry_batch_wait() polls until the batch reaches a
terminal state. Batch jobs target completion within 24 hours, so use a
longer polling interval in a large live workflow.
completed_batch <- foundry_batch_wait(
batch$batch_id,
interval = 60
)
completed_batch |>
select(
status,
request_counts_total,
request_counts_completed,
request_counts_failed
)
#> # A tibble: 1 × 4
#> status request_counts_total request_counts_completed request_counts_failed
#> <chr> <int> <int> <int>
#> 1 completed 6 6 0foundry_batch_results() downloads and parses the output
file and, when present, the error file. Failed request rows have
.error = TRUE and an .error_msg.
batch_results <- foundry_batch_results(completed_batch$batch_id)
batch_results |>
select(custom_id, output_text, .error, input_tokens, output_tokens)
#> # A tibble: 6 × 5
#> custom_id output_text .error input_tokens output_tokens
#> <chr> <chr> <lgl> <int> <int>
#> 1 c02 "{\"theme\":\"assessment\",\"sent… FALSE 114 11
#> 2 c04 "{\"theme\":\"materials\",\"senti… FALSE 114 11
#> 3 c03 "{\"theme\":\"support\",\"sentime… FALSE 115 11
#> 4 c06 "{\"theme\":\"assessment\",\"sent… FALSE 111 11
#> 5 c05 "{\"theme\":\"assessment\",\"sent… FALSE 113 11
#> 6 c01 "{\"theme\":\"instruction\",\"sen… FALSE 114 11The results came back in the order c02, c04, c03, c06, c05, c01, not
in the order the requests were written. Join on custom_id,
never on row position.
The coded fields arrive as JSON text in output_text and,
parsed, in the structured list-column. The extraction path
below turns them into ordinary columns joined to your rows.
When the service writes an error file,
completed_batch$error_file_id names it. Those rows are
included in foundry_batch_results() with
.error columns so you can separate successful labels from
failed requests before joining.
Usage is reported in tokens. Pass your own per-token rates because Azure prices change and differ by deployment. The rates below are illustrative, not current prices.
Most annotation projects do not need the low-level objects in
analysis code. foundry_extract_batch(wait = FALSE) writes
the request file, uploads it and creates the batch. It returns the batch
immediately.
extract_batch <- foundry_extract_batch(
comments,
text_col = "comment",
schema = course_schema,
model = "gpt-4.1-nano-batch",
wait = FALSE,
instructions = course_instructions
)
extract_batch |>
select(status, completion_window, batch_id)
#> # A tibble: 1 × 3
#> status completion_window batch_id
#> <chr> <chr> <chr>
#> 1 validating 24h batch_5dfcd819-e17a-4b4c-951a-fbbb6c6d0dc1Save extract_batch$batch_id with the data and schema.
Later, even in a new R session, check that the batch has finished, then
call foundry_extract_batch_results() with the batch ID,
original rows and the same schema. foundryR joins results back to the
original rows using the row-N IDs written by
foundry_extract_batch(), and warns about any row that has
no result.
foundry_batch_wait(extract_batch$batch_id, interval = 60) |>
select(status, request_counts_completed, request_counts_failed)
#> # A tibble: 1 × 3
#> status request_counts_completed request_counts_failed
#> <chr> <int> <int>
#> 1 completed 6 0
extract_results <- foundry_extract_batch_results(
extract_batch$batch_id,
data = comments,
schema = course_schema,
text_col = "comment"
)
extract_results |>
select(comment_id, theme, sentiment, .error)
#> # A tibble: 6 × 4
#> comment_id theme sentiment .error
#> <chr> <chr> <chr> <lgl>
#> 1 c01 instruction positive FALSE
#> 2 c02 assessment negative FALSE
#> 3 c03 support positive FALSE
#> 4 c04 materials negative FALSE
#> 5 c05 assessment positive FALSE
#> 6 c06 assessment mixed FALSEThis path is the one to use in a reproducible analysis. The original tibble stays in your project, and the service result is collected into the same row structure after the batch finishes.
Delete service files that are no longer needed. Keep the batch IDs, codebook version and schema with your analysis so the results can be traced.