---
title: "Transcribe and translate audio"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Transcribe and translate audio}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
fixture_dir <- "audio"
recording <- nzchar(Sys.getenv("FOUNDRY_RECORD_DOCS"))
have_fixtures <- dir.exists(fixture_dir) && length(list.files(fixture_dir)) > 0
run_api <- requireNamespace("httptest2", quietly = TRUE) &&
  (recording || have_fixtures)
library(foundryR)
if (run_api) {
  httptest2::start_vignette(fixture_dir)
}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>", eval = run_api,
  fig.width = 7, fig.height = 4.5, out.width = "100%")

embed_audio <- function(path) {
  if (!requireNamespace("base64enc", quietly = TRUE)) {
    return(invisible(NULL))
  }
  uri <- paste0("data:audio/mpeg;base64,", base64enc::base64encode(path))
  knitr::asis_output(
    sprintf(
      '<audio controls src="%s">Your browser does not support audio playback.</audio>',
      uri
    )
  )
}
```

```{r library, eval = TRUE, message = FALSE}
library(foundryR)
```

Calls to Azure show output recorded from a live run; setup code is shown but not run.

Audio workflows usually start with transcription and end with a coding or analysis table. The examples use the bundled JFK sample and a short Spanish clip that the article generates, and they keep each service result in a tibble.

```{r sample, eval = TRUE}
sample_audio <- system.file("extdata", "samples", "jfk.wav", package = "foundryR")
basename(sample_audio)
```

## Fast transcription

`foundry_transcribe()` defaults to Speech fast transcription with `service = "speech"`. This path does not need a model or deployment name, so the returned `model` column is `NA`. The `language` column holds the locale the service detected.

```{r transcribe}
transcript <- foundry_transcribe(sample_audio)

transcript$text
transcript[, c("model", "duration_ms", "language")]
```

The `phrases` list-column holds phrase-level timing when the service returns it. Use those offsets to spot-check long recordings before coding them.

```{r transcribe-phrases}
head(transcript$phrases[[1]][, c("text", "offset_ms", "duration_ms")])
```

## MAI-Transcribe enhanced mode

MAI-Transcribe is opt-in. Supplying a MAI model selects Speech enhanced mode, which is preview and region-limited. In live testing, MAI-Transcribe worked on a Sweden Central resource, while East US 2 rejected enhanced mode. foundryR adds region guidance to that error.

```{r mai-transcribe, eval = FALSE}
foundry_set_speech_endpoint(Sys.getenv("AZURE_FOUNDRY_SPEECH_ENDPOINT"))
foundry_set_speech_key("your-speech-key")

foundry_transcribe(
  sample_audio,
  model = "mai-transcribe-2"
)
```

## Transcribe non-English audio

Interviews are often recorded in more than one language. This example makes a Spanish clip with `foundry_speak()`, so the article does not depend on real participant audio.

```{r translate-synth}
spanish_path <- tempfile(fileext = ".mp3")
spanish_clip <- foundry_speak(
  paste(
    "Buenos dias a todos. En la reunion de hoy revisamos la encuesta a los docentes.",
    "La mayoria pidio mas tiempo para preparar las clases y menos tareas administrativas.",
    "El equipo presentara un plan el proximo mes."
  ),
  model = "gpt-4o-mini-tts",
  voice = "verse",
  path = spanish_path
)
```

```{r translate-play, echo = FALSE}
embed_audio(spanish_clip$path)
```

Pass the locale to fast transcription, then translate the text with a model if readers need English.

```{r spanish-transcribe}
spanish <- foundry_transcribe(spanish_clip$path, locales = "es-ES")
spanish$text

english <- foundry_response(
  spanish$text,
  instructions = "Translate the text into English. Return only the translation."
)
english$output_text
```

Keep the Spanish transcript as the record and treat the English text as a reading aid. The translation is another model output, so have a bilingual reader spot-check it before you quote it.

## Audio translation endpoints

`foundry_translate_audio()` sends the audio to a translation endpoint instead. Speech translation needs LLM Speech enhanced mode, and it failed in both regions tested, East US 2 and Sweden Central. A classic Azure `whisper` deployment accepts translation requests through the legacy deployment route. Azure `whisper` version 001 retires on 2026-12-15.

```{r translate}
whisper <- foundry_translate_audio(
  spanish_clip$path,
  service = "openai",
  model = "whisper",
  api = "deployment"
)

whisper$text
```

```{r translate-summary, echo = FALSE, results = "asis"}
words <- function(x) {
  x <- tolower(iconv(x, "UTF-8", "ASCII//TRANSLIT"))
  unique(strsplit(gsub("[^a-z ]", " ", x), "\\s+")[[1]])
}
source_words <- setdiff(words(spanish$text), "")
shared <- mean(source_words %in% words(whisper$text))
if (shared > 0.6) {
  cat(sprintf(
    "The Whisper translation route returned the clip in Spanish: %.0f%% of the words in the Spanish transcript appear in its output. Whisper translation should return English, so check the language of every output before you use this route in a pipeline.\n",
    100 * shared
  ))
} else {
  cat("Whisper returned English text for this clip. Check the language of every output before you use this route in a pipeline, because Whisper can return untranslated text.\n")
}
```

## Code the transcript

The next step is to turn transcript text into variables you can inspect, join, or model. Closed label sets work best here: a fixed list of topics and tones gives you codes you can count. The model can code the Spanish transcript directly with English labels, so the codes do not depend on the translation.

```{r code-transcript}
transcript_schema <- foundry_schema(
  topic = schema_enum(
    c("workload", "curriculum", "facilities", "other"),
    "Main topic of the passage."
  ),
  tone = schema_enum(
    c("formal", "informal", "urgent", "reflective"),
    "Overall tone of the speaker."
  )
)

meeting_code <- foundry_extract(
  spanish$text,
  schema = transcript_schema,
  schema_name = "TranscriptCode"
)

meeting_code[, c("topic", "tone", ".status", ".error")]
```

Check `.error` before you use the codes. When a response stops early, for example because the deployment's content filter blocked the output or the token limit ran out, `.status` is `"incomplete"`, `.error` is `TRUE`, and `.error_msg` names the reason. Rerun or drop those rows; their missing fields are not answers. The same call on the JFK transcript shows what a failed row looks like.

```{r code-speech}
speech_code <- foundry_extract(
  transcript$text,
  schema = transcript_schema,
  schema_name = "TranscriptCode"
)

speech_code[, c(".status", ".error")]
speech_code$.error_msg
```

```{r code-speech-summary, echo = FALSE, results = "asis"}
if (isTRUE(speech_code$.error) && grepl("content_filter", speech_code$.error_msg)) {
  cat("In this recording the content filter stopped the response, although the request asks only for two labels. Widely quoted text, such as a famous speech, can trip the filter. If you code texts like these, ask your Azure administrator about the content filter configuration on the deployment, and report how many rows failed next to any coded shares.\n")
} else {
  cat("In this recording the JFK transcript was coded without an error. Deployments with stricter content filters can still stop responses about widely quoted text, so count failed rows before you report coded shares.\n")
}
```

## Synthesize speech

`foundry_speak()` writes generated audio to disk and returns metadata about the file. Use an explicit path for audio you want to keep.

```{r speak}
speech_path <- tempfile(fileext = ".mp3")
speech <- foundry_speak(
  "Please read each survey question before choosing an answer.",
  model = "gpt-4o-mini-tts",
  voice = "verse",
  path = speech_path
)

speech[, c("bytes", "model", "voice", "format")]
```

```{r speak-play, echo = FALSE}
embed_audio(speech$path)
```

## Notes for analysis

Inspect `transcript$phrases[[1]]` before processing long recordings, because segment structure affects coding and quality checks. Keep raw audio out of your repository; store transcript text, stable IDs, model metadata, and links to controlled storage.

```{r cleanup, include = FALSE, eval = TRUE}
if (run_api) {
  unlink(c(speech_path, spanish_path))
  httptest2::end_vignette()
}
```
