---
title: "dataprep: descriptive statistics and diagnostic plots"
output: rmarkdown::html_vignette
vignette: >
  %\documentclass{article}
  %\VignetteIndexEntry{dataprep: descriptive statistics and diagnostic plots}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse  = TRUE,
  comment   = "#>",
  fig.align = "center",
  fig.width = 6,
  fig.height = 5.5,
  out.width = "80%",
  fig.retina = 2
)
```

```{r}
library(dataprep)
library(ggplot2)
```

## Overview

`dataprep` provides two families of plotting helpers, both built on
top of the C++ descriptive backends:

* **`descplot()`** — descriptive statistics (`n`, `na`, `mean`, `sd`,
  `median`, `trimmed`, `min`, `max`, `IQR`) computed in C++ and
  displayed as line or bar charts.

* **`percplot()`** — top and bottom percentile summaries, useful for
  detecting heavy tails and percentile-based outlier cutoffs. The
  geometry depends on the column names: lines when they are
  numeric, grouped bars otherwise.

Both share the same interface style: a data frame, a numeric range,
and optional grouping. Use `data1` (7,640 rows × 7 columns) for
quick demos, and `data` (7,640 × 65) for full-size examples.

Under the hood, `descplot()` calls `descdata()` (which calls
`desc_stats_cpp()`), and `percplot()` calls `percdata()` (which
calls `quantile()` from base R on each column). Both return a
`ggplot` object, so all usual `ggplot2` layers apply.

## Descriptive statistics

### Line plot, numeric variable names

When variable names are essentially numeric (e.g. particle
diameters such as 3.16, 3.55, ...), `descplot()` draws a line
plot with a log-scaled x axis.

```{r, fig.height = 5}
descplot(data1, cols = 3:7) +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
```

### Selected statistics

Pass a subset of statistics by index or by name to focus the plot.

```{r, fig.height = 4}
descplot(data1, cols = 3:7,
         stats = c("na", "min", "max", "IQR")) +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
```

The `stats` argument accepts both forms — numeric indices
(`1:9`) and character names (`"na"`, `"min"`, `"max"`, `"IQR"`)
— and can mix them.

### Bar chart, character variable names

When variable names are character (e.g. aerosol mode names
`Nucleation`, `Aitken`, `Accumulation`), `descplot()` falls back
to a bar chart.

```{r, fig.height = 5}
descplot(data1, cols = 3:7) +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
```

### Control facet layout

```{r, fig.height = 3}
descplot(data1, cols = 3:7, stats = c("min", "max", "IQR")) +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
```

### Full-size data

```{r, fig.height = 4.5}
descplot(data, cols = 5:65)
```

### The underlying table

If you only need the numbers (not the plot), call `descdata()`
directly. It returns a data frame with one row per variable and
one column per statistic.

```{r, fig.height = 4}
descdata(data1, cols = 3:7, stats = c(2, 3, 4, 7:9))
```

## Percentile plots

### Full percentile range

`percplot()` computes the extreme percentiles (0 to 0.5 and 99.5 to
100 by default) and draws them against the variable axis. This is
the visual companion to `condextr()` and `percoutl()`.

```{r, fig.height = 6, fig.width = 7}
percplot(data1, cols = 3:7, group = 2) +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
```

### Top percentiles only

```{r, fig.height = 3, fig.width = 7}
percplot(data1, cols = 3:7, group = 2, part = "top") +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
```

### Bottom percentiles only

```{r, fig.height = 3, fig.width = 7}
percplot(data1, cols = 3:7, group = 2, part = "bottom") +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
```

### Percentile curves of the raw data

When the column names are numeric, `percplot()` draws curves
instead of bars. The 61 size-bin columns of `data` (`cols = 5:65`)
have numeric names, so the default `num_xaxis = "auto"` selects a
log-scaled x axis:

```{r, fig.height = 4, fig.width = 7}
percplot(data, cols = 5:65, group = 4) +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
```

### Numeric axis control

For numeric variable names, the x axis can be forced to linear
scale with `num_xaxis = "numeric"`.

```{r, fig.height = 4, fig.width = 7}
percplot(data, cols = 5:65, group = 4, num_xaxis = "numeric") +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
```

The `num_xaxis` argument controls how the x axis is treated when
column names are numeric:

* `"auto"` (default) — uses a log scale if the column names are
  evenly spaced on a log scale with a range ratio of at least
  1000; otherwise keeps them as factor levels.
* `"log"` / `"numeric"` — force the corresponding scale.
* `"character"` / `"factor"` / `"keep"` / `FALSE` — keep the
  column names as factor levels.

### The underlying table

`percdata()` returns the same table that `percplot()` draws.

```{r}
percdata(data1, cols = 3:7, group = 2, part = "top")
```

## Combining with `ggplot2`

Both `descplot()` and `percplot()` return `ggplot` objects, so all
usual `ggplot2` layers apply.

```{r, fig.height = 4, fig.width = 7}
percplot(data1, cols = 3:7, group = 2) +
  ggplot2::theme_bw(base_size = 11) +
  ggplot2::labs(title = "Percentile plots by month",
                x = "Variable", y = "Value") +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
```

## A diagnostic workflow

A typical diagnostic workflow combines `data_report()`,
`na_diagnose()`, and the two plot families:

```{r, eval = FALSE}
# 1. Overview of the whole table
data_report(data, cols = 5:65)

# 2. Per-column NA run statistics
na_diagnose(data, cols = 5:65)

# 3. Descriptive statistics of the raw data
descplot(data, cols = 5:65)

# 4. Percentile curves of the raw data
percplot(data, cols = 5:65, group = 4)
```

After running `dataprep()` you can compare the raw and cleaned
versions in the same plot by stacking them with a `g` column:

```{r, eval = FALSE}
cleaned <- dataprep(data, cols = 5:65, group = 4)

percplot(
  rbind(
    transform(data[names(cleaned)], g = "original"),
    transform(cleaned,              g = "preprocessed")
  ),
  cols  = 5:ncol(cleaned),
  group = ncol(cleaned) + 1
)
```

## Where to go next

* **Design philosophy** — `vignette("dataprep-philosophy")`.
  Why the cleaning pipeline has the shape it does.

* **Cleaning pipeline walkthrough** — `vignette("dataprep-cleaning")`.
  Step-by-step execution of the four cleaning steps on `data`.

* **Performance and cross-engine consistency** —
  `vignette("dataprep-performance")`. Benchmark tables and
  8-engine consistency checks.

* **Upgrading from 0.1.5 to 0.1.8** —
  `vignette("dataprep-migration")`. Behaviour changes and the
  migration checklist.

* **Fast reshaping** — `vignette("dataprep-melt-dcast")`.

## Session info

```{r}
sessionInfo()
```
