---
title: "Introductory rtables - Basic Table Layout Instructions"
author: "Gabriel Becker"
date: "`r Sys.Date()`"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Introductory rtables - Basic Table Layout Instructions}
  %\VignetteEncoding{UTF-8}
  %\VignetteEngine{knitr::rmarkdown}
editor_options:
  chunk_output_type: console
---

# Introduction

To create a table using `rtables`, we first declare its structure
using the layout engine. We do this by defining four aspects of the
table we intend to create:

1. Individual Row and Cell Contents,
2. Column Structure,
3. Row Faceting Structure, and
4. Marginal (/Summary) Cell Contents

Each of these aspects corresponds with a set of layout functions which we will go through now.

# Declaring Rows And Cell Contents With `analyze()`

## Core Idea

We declare rows and cell values by `analyze()`ing variables with *analysis functions*.

Analysis functions: 
- are applied to data by `rtables` to calculate cell values at table
  construction time
- can generate multiple rows
- *do not* handle subsetting data into relevant subsets for each cell

## Basic Usage

In `rtables` we declare individual rows and the cell contents for
those rows by setting an *analysis function* via an `analyze()`
call. By default, `analyze()` declares the `simple_analysis()`
analysis, which will create a single row with the man value(s) for a
numeric variable and a row per level with counts for a factor (or
character) variable. For this introductory portion of the tour we will
simply use this default analysis function, as it suffices to
illustrate the core behaviors we are discussing.

Consider a trivial table, with a single column representing all of our
data. We can analyze `AGE` (a numeric variable) to declare a single
row displaying the mean of patient ages in our data:

```{r}
library(rtables)
lyt <- basic_table() |>
  analyze("AGE")

build_table(lyt, ex_adsl)
```

Alternatively, if we `analyze` `BMRKR2`, a simulated categorical
biomarker, we get rows with counts for each level:

```{r}
lyt2 <- basic_table() |>
  analyze("BMRKR2")

build_table(lyt2, ex_adsl)
```

Note: in production tables we will typically not use the default
analysis function; we are using it here to separate the discussion of
what `analyze()` does from discussion of any particular analysis
function. See the intermediate guided tour, specifically [Identifying
Required Analysis Behavior](./guided_intermediate_afun_reqs.html) for
more details on how we select analysis functions in more realistic
scenarios.

# Declaring Columns With `split_cols_by()`

## Core Idea

Individual columns are typically defined via `split_cols_by()`, which defines *column faceting*. This can be nested which is discussed in detail in the [next portion](./guided_intro_nesting.html) of this guide. 

Column faceting:
- defines individual columns
- defines subsetting for each column which applies to all rows
  - typically a partition on categorical variable
  - *can* be overlapping or non-exhaustive as needed

## Basic Usage

By default, faceting (which we also call *splitting*) partitions the data; in the case of `split_cols_by`, declaring columns.

For example, we can define a column for each trial arm in our simulated data by splitting on `ARM`:

```{r}
lyt3 <- basic_table() |>
  split_cols_by("ARM") |>
  analyze("BMRKR2")

build_table(lyt3, ex_adsl)
```

Here we combine our column splitting with the analysis of `BMRKR2` to illustrate how splitting interacts with analysis functions that generate more than one row.

We can see that we now have three columns - one for each arm - and each column has three cells - one for each `BMRKR2` level.

Note: our analysis function is only called here three times - once for
each column - as there is no additional row faceting and each call
generates three values.


# Declaring Row-Grouping With `split_rows_by()`

## Core Idea

We declare *structural groups* of rows via row faceting. This defines
what data our analysis function is passed *from the row structure
perspective* to (repeatedly) create our cells and individual rows
during tabulation.

Row faceting defines groups that:

- will contain one or more individual rows, and
- are eligible for marginal summaries (see next section)

## Basic Usage

Row facets represent subsets of our data which should be analyzed. For
example, we might want to analyze patient `BMRKR2` status for each
gender:

```{r}
lyt4 <- basic_table() |>
  split_rows_by("SEX") |>
  analyze("BMRKR2")

build_table(lyt4, ex_adsl)
```

# Adding Marginal Summaries With `summarize_row_groups`

## Core Idea

Marginal row-group summaries generate cells that provide context for cells underneath them in the row structure; they *typically* replace label a facet's row with more informative row(s). 

We declare marginal summaries by declaring a *content function* via `summarize_row_groups()`, which modifies the the currently active (most recent) row split such that:

- the content function will be called for each facet generated by the split, and
- the label row for each facet will be replaced with the row(s) generated by the content function

## Basic Usage

We might want an overall count for each gender in addition to those of each `BMRKR2` level within those genders:

```{r}
lyt5 <- basic_table() |>
  split_rows_by("SEX") |>
  summarize_row_groups("SEX") |>
  analyze("BMRKR2")

build_table(lyt5, ex_adsl)
```

## Some Relevant Details

- By default label rows are hidden when a marginal summary is present
  - this can be disabled via `child_labels = 'visible'` in the
    `split_rows_by` call
  
# All Together - A Basic Rectangular Table

We combine our column splitting, row splitting, group summary, and
analysis instructions to create a full table layout, like so:

```{r}
lyt_basic <- basic_table() |>
  split_cols_by("ARM") |>
  split_rows_by("SEX") |>
  summarize_row_groups("SEX") |>
  analyze("BMRKR2")

build_table(lyt_basic, ex_adsl)
```

Thus we have created a basic table. In practice, our table structures
are generally significantly more complex than this; rtables supports
these myriad structures by allowing us to control the *nesting*
behavior of both splitting and analysis instructions. We cover this in
detail in the [next section](./guided_intro_nesting.html) of this
guide.
