UlavalSSD provides fixed snapshots for practising data cleaning and exploratory analysis. These examples run offline and use base R. The same objects can be used with the tidyverse when those packages are installed.
MeteoQuebec retains the raw column names and types
expected by the course. Month and day are character strings. The first
column, ...1, is an export identifier, not a measurement.
Explicitly converting to a base data frame makes subsetting behaviour
independent of whether tibble is installed.
weather <- as.data.frame(MeteoQuebec)
weather$date <- as.Date(with(weather, paste(year, month, day, sep = "-")))
range(weather$date)
#> [1] "1970-01-01" "2025-01-22"
head(weather[c("date", "min_temp", "max_temp")])
#> date min_temp max_temp
#> 1 1970-01-01 -19.4 -12.8
#> 2 1970-01-02 -19.4 -12.8
#> 3 1970-01-03 -18.9 -13.3
#> 4 1970-01-04 -20.6 -13.3
#> 5 1970-01-05 -21.1 -12.8
#> 6 1970-01-06 -18.3 -10.0A complete sequence of dates does not mean every measurement is observed. Count the missing values before choosing a summary or a model.
measurements <- c(
"max_temp", "mean_temp", "min_temp", "total_precip",
"total_rain", "total_snow", "snow_grnd"
)
data.frame(
variable = measurements,
missing = vapply(weather[measurements], function(x) sum(is.na(x)), integer(1)),
missing_proportion = vapply(
weather[measurements], function(x) mean(is.na(x)),
numeric(1)
),
row.names = NULL
)
#> variable missing missing_proportion
#> 1 max_temp 53 0.002635374
#> 2 mean_temp 54 0.002685098
#> 3 min_temp 40 0.001988961
#> 4 total_precip 97 0.004823231
#> 5 total_rain 10490 0.521605092
#> 6 total_snow 10587 0.526428323
#> 7 snow_grnd 8912 0.443140570For a descriptive plot, retain dates alongside observations. Missing values remain visible as gaps rather than being replaced by zero.
one_year <- weather[weather$year == 2024, , drop = FALSE]
plot(one_year$date, one_year$min_temp,
type = "l",
xlab = "Date", ylab = "Daily minimum temperature (degrees C)"
)An independent reconstruction from the official ECCC service matched every stored column on 2026-09-22. Station 5251 is used through 1995 and station 26892 from 1996. These data support cleaning exercises; a climate-trend study would require checking station history, measurement flags and consistency first.
listecondamnation contains historical records. It is not
a representative sample of food establishments. One establishment may
appear more than once. Counts of these records cannot estimate the
probability of an offence or an establishment’s present-day operating
conditions.
records <- as.data.frame(listecondamnation)
range(as.Date(records$Date_publication))
#> [1] "2023-02-13" "2025-02-10"
head(sort(table(records$Type_etablissement), decreasing = TRUE))
#>
#> RESTAURANT REST. SERVICE RAPIDE
#> 1353 187
#> RESTAURANT SERVICE RAPIDE RESTAURANT METS A EMPORTER
#> 137 35Amende is text. Examine its formats before parsing
currency. The following example validates a deliberately small set of
known formats, rather than silently discarding every non-numeric
character in an arbitrary string.
amount_text <- c("1 500 $", "250,50 $", NA_character_)
amount_clean <- gsub("[ $]", "", amount_text)
amount_clean <- sub(",", ".", amount_clean, fixed = TRUE)
valid <- is.na(amount_clean) | grepl("^[0-9]+([.][0-9]{1,2})?$", amount_clean)
stopifnot(all(valid))
amount <- as.numeric(amount_clean)
data.frame(amount_text, amount)
#> amount_text amount
#> 1 1 500 $ 1500.0
#> 2 250,50 $ 250.5
#> 3 <NA> NAThis example uses synthetic amounts. Applying a conversion to the full dataset requires checking all observed formats and preserving the original text for verification.
French output remains the default. The row feedback is a fixed answer key for the original penguin exercise file distributed on the course website. It is valid only before rows are filtered or reordered.
consulter_taches("histogramme", lang = "en")
#> [1] "Show the distribution of penguin flipper lengths. There seems to be an error in the data. Can you find it?"
verifier_valeur_aberrante(11, lang = "en")
#> [1] "Exercise correction: body mass should be 3300 g (3.3 kg) and bill length should be 37.8 mm."Install the optional package lintr to use
eval_tidyverse_style(). It examines code without executing
it. In a Quarto document it inspects fenced R chunks and reports line
numbers in the original file.
if (requireNamespace("lintr", quietly = TRUE)) {
path <- tempfile(fileext = ".R")
writeLines(c("daily_mean <- mean(c(2, 4, 6))", "print(daily_mean)"), path)
feedback <- eval_tidyverse_style(path)
feedback[c("total", "status", "diagnostics")]
unlink(path)
}The score is a transparent formative indicator based on seven
observable criteria. It is not a validated assessment of programming
competence. Empty or invalid code has no score, and the clarity of
reasoning and usefulness of comments remain for human review. See
?eval_tidyverse_style for the criteria and Quarto fence
support.
Weather observations are attributed to Environment and Climate Change
Canada; administrative records are attributed to MAPAQ via Donnees
Quebec. See ?MeteoQuebec, ?listecondamnation,
and the installed attribution file:
system.file("COPYRIGHTS", package = "UlavalSSD")
#> [1] "/tmp/Rtmp3kXTFK/Rinst16ba4426980a/UlavalSSD/COPYRIGHTS"The datasets do not refresh during installation, loading, examples or tests.