---
title: "Introductory `rtables` - Facet And Analysis Nesting"
author: "Gabriel Becker"
date: "`r Sys.Date()`"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Introductory rtables - Facet And Analysis Nesting}
  %\VignetteEncoding{UTF-8}
  %\VignetteEngine{knitr::rmarkdown}
editor_options:
  chunk_output_type: console
---

# Introduction


`rtables` models data-summarizing tables as *faceted data
visualizations*, analogous to a `ggplot2` plot using `facet_grid` or a
`lattice` plot conditioned on multiple factors.

We saw in the previous section that we use:

- `split_cols_by` to declare *columns*,
- `split_rows_by` to declare *groups of individual rows*,
- `summarize_row_groups` to declare *marginal summary rows* for groups
  of individual rows, and
- `analyze` to declare (sets of) *individual rows*.

Combining a single call each to `split_cols_by`, `split_rows_by` and
`analyze` creates a rectangular table, while adding
`summarize_row_groups` after the `split_rows_by` adds marginal summary
rows for each group.

Often we need tables with more complex structure, whether it is
multiple top-level sections of the table; tables which analyze
multiple variables simultaneously; nested faceting in row structure,
column structure, or both; or combinations of all three of these.

We achieve all of these by leveraging *nesting* of layout instructions.

Throughout this vignette we will use a custom split function (`keep_2_levels`)
for table brevity, defined as follows:

```{r}
keep_2_levels <- function(varnm, dat = ex_adsl) {
  keep_split_levels(levels(dat[[varnm]])[1:2])
}
```

# Nesting

*Nesting* is how we talk about *where* a layout instruction fits with
respect to the existing state of the layout. We say an instruction is
*nested within* a preceding faceting instruction (`split_rows_by` or
`split_cols_by`) if the new instruction *should be applied separately
within each facet generated during tabulation from the previous
instruction. This is analogous to what we see with `facet_*` in
`ggplot2` when we give multiple variables for a single faceting
dimension.

By default, each layout instruction is nested within the directly
preceding layout instruction - if any - in its dimension (row or
column), with a couple caveats we discuss later. We see this default
behavior below:

```{r}
library(rtables)
lyt <- basic_table() |>
  split_cols_by("ARM") |>
  split_cols_by("STRATA1") |>
  split_rows_by("SEX", split_fun = keep_2_levels("SEX")) |>
  split_rows_by("BMRKR2", split_fun = keep_2_levels("BMRKR2")) |>
  analyze("AGE")
table_structure(build_table(lyt, ex_adsl))
```

When `analyze` instructions are 'nested within' another `analyze`, the
analyses are bundled into a 'multi-analysis' parent structure. This
parent structure as a whole, then, has the nesting behavior that a
single `analyze` call would have in its place.

```{r}
lyt2 <- basic_table() |>
  split_cols_by("ARM") |>
  split_cols_by("STRATA1") |>
  split_rows_by("SEX", split_fun = keep_2_levels("SEX")) |>
  split_rows_by("BMRKR2", split_fun = keep_2_levels("BMRKR2")) |>
  analyze("AGE") |>
  analyze("BMRKR1")
table_structure(build_table(lyt2, ex_adsl))
```

By default:

- `analyze` calls nest within the most recently preceding
  `split_rows_by` or instruction
  - multiple `analyze` calls that nest within the 

## Creating 'Multi-Section' Tables With `nested = FALSE`

We often want to create tables with rows grouped into two or more
logical or analytical sections. For example we might want to analyze
`AGE` overall, then separately by `SEX` and by `RACE`. We can do this by:

1. `analyze`ing `AGE`, then
2. splitting by `SEX` and `analyze`ing `AGE`, and finally
3. splitting by `RACE` and `analyze`ing `AGE` again.

We will start each section after the first with `nested = FALSE` to
delineate it from the previous portion of the layout.

NOTE: while we will do it explicitly for illustration purposes, any
`split_rows_by` layout instruction that follows an `analyze` 
defaults to `nested = FALSE`.

Thus we can create our table with the code below:

Note: we set a top level section divider to make our different
sections concrete; section dividers will be covered in a later part of
this guide and can be taken as is for now.

```{r}
trim_adsl <- subset(ex_adsl, RACE %in% levels(ex_adsl$RACE)[1:3] & SEX %in% c("F", "M"))
trim_adsl$RACE <- factor(trim_adsl$RACE)
trim_adsl$SEX <- factor(trim_adsl$SEX)

nice_mean <- function(x) {
  in_rows("Average Age" = mean(x), .formats = list("Average Age" = "xx.x"))
}

lyt3 <- basic_table(top_level_section_div = "-") |>
  split_cols_by("ARM") |>
  analyze("AGE", afun = nice_mean) |>
  split_rows_by("SEX", nested = FALSE) |>
  analyze("AGE", afun = nice_mean) |>
  split_rows_by("RACE", nested = FALSE) |>
  analyze("AGE", afun = nice_mean)

tbl3 <- build_table(lyt3, trim_adsl)
tbl3
```

We see three clear top-level 'sections' of our table in row-space, as
desired.  Contrast this with our result without `nested = FALSE` (and
with the first two `analyze` calls replaced with
`summarize_row_groups`:


```{r}
nice_mean_cfun <- function(x, labelstr) {
  lbl <- paste0(labelstr, " (Ave. Age)")
  in_rows(mean(x), .labels = lbl, .formats = "xx.x")
}

lyt3b <- basic_table(top_level_section_div = "-") |>
  split_cols_by("ARM") |>
  summarize_row_groups("AGE", cfun = nice_mean_cfun) |>
  split_rows_by("SEX", split_fun = keep_2_levels("SEX")) |>
  summarize_row_groups("AGE", cfun = nice_mean_cfun) |>
  split_rows_by("RACE", split_fun = keep_2_levels("RACE")) |>
  analyze("AGE", afun = nice_mean)

tbl3b <- build_table(lyt3b, trim_adsl)
head(tbl3b)
```

Here the faceting on `RACE` occurs *nested within* the faceting on
`SEX`, whereas above it occurs *alongside* it.

### A Brief Note On Removing Sections

Sometimes we will receive a template script that does more than we
want, but is close to meeting our needs. For example, imagine we
wanted the above table (non-nested version) but only wanted the
overall and race portions, removing the age within gender analysis.

We do this by simply identifying all of the layout instructions
corresponding to that portion of the table and removing them. Our top
level section dividers can help us reason about this, and can be added
to the template if they were not there originally.

In our case, the instructions for our section to remove are the
`split_rows_by("SEX", nested = FALSE)`, and directly following
`analyze("AGE")` calls. By starting with our code above and removing
those, we would get our desired table:


```{r}
lyt3c <- basic_table(top_level_section_div = "-") |>
  split_cols_by("ARM") |>
  analyze("AGE", afun = nice_mean) |>
  ## split_rows_by("SEX", nested = FALSE) |>
  ## analyze("AGE", afun = nice_mean) |>
  split_rows_by("RACE", nested = FALSE) |>
  analyze("AGE", afun = nice_mean)

tbl3c <- build_table(lyt3c, trim_adsl)
tbl3c
```

When performing this kind of layout pruning in the wild, it is
important to remember that `split_rows_by` calls that follow `analyze`
calls default to `nested = FALSE`, even if that is not made explicit
in the template script you are starting from.

It is also important to not remove all `analyze` call(s) nested within
any series of row faceting (`split_rows_by*` calls), as this will
result in an degenerate (invalidly structured) table which could have
undefined behavior when passed to some other aspects of the `rtables`
and `formatters` APIs.

## Intermediate Nesting

As of `rtables` `0.7.0`, we can declare *intermediate* nesting, rather
than simply full -- the previous and now default behavior when `nested
= TRUE` -- and no -- the `nested = FALSE` behavior -- nesting.

We do this via the new `at_sibling` parameter the `split_rows_by*` and
`analyze*` families of layout functions now accept. `at_sibling` allows
us to specify the *nesting anchor* for a row split or analyze directive;
when we do so, the table resulting from our new directive will appear
*as a direct sibling* to that resulting from our anchor in the
created table.


Consider where our `BMRKR2` analysis is placed in when using the
following layouts to build tables:

The default behavior:
```{r}
lyt4 <- basic_table() |>
  split_cols_by("ARM") |>
  split_rows_by("STRATA1", split_fun = keep_2_levels("STRATA1")) |>
  split_rows_by("SEX", split_fun = keep_2_levels("SEX")) |>
  analyze("AGE") |>
  analyze("BMRKR2")

build_table(lyt4, ex_adsl)
```

Anchoring the analysis on `"SEX"`:
```{r}
lyt4a <- basic_table() |>
  split_cols_by("ARM") |>
  split_rows_by("STRATA1", split_fun = keep_2_levels("STRATA1")) |>
  split_rows_by("SEX", split_fun = keep_2_levels("SEX")) |>
  analyze("AGE") |>
  analyze("BMRKR2", at_sibling = "SEX", show_labels = "visible")

build_table(lyt4a, ex_adsl)
```

Anchoring the analysis on `"STRATA1"`
```{r}
lyt4b <- basic_table() |>
  split_cols_by("ARM") |>
  split_rows_by("STRATA1", split_fun = keep_2_levels("STRATA1")) |>
  split_rows_by("SEX", split_fun = keep_2_levels("SEX")) |>
  analyze("AGE") |>
  analyze("BMRKR2", at_sibling = "STRATA1", show_labels = "visible")

build_table(lyt4b, ex_adsl)
```

Analysis is fully non-nested:
```{r}
lyt4c <- basic_table() |>
  split_cols_by("ARM") |>
  split_rows_by("STRATA1", split_fun = keep_2_levels("STRATA1")) |>
  split_rows_by("SEX", split_fun = keep_2_levels("SEX")) |>
  analyze("AGE") |>
  analyze("BMRKR2", nested = FALSE, show_labels = "visible")

build_table(lyt4c, ex_adsl)
```

Note that because our `STRATA` split is a top-level split, anchoring
our analysis to it is equivalent to simply using `nested =
FALSE`. While these result in identically-rendering tables, they will
not if our current layout is placed under a new split, such as when we
want the same table structure both globally and split by subgroups or
parameters:



Anchoring the analysis on `"STRATA1"`
```{r}
lyt4d <- basic_table() |>
  split_cols_by("ARM") |>
  split_rows_by("RACE", split_fun = keep_2_levels("RACE")) |>
  split_rows_by("STRATA1", split_fun = keep_2_levels("STRATA1")) |>
  split_rows_by("SEX", split_fun = keep_2_levels("SEX")) |>
  analyze("AGE") |>
  analyze("BMRKR2", at_sibling = "STRATA1", show_labels = "visible")

build_table(lyt4d, ex_adsl)
```

Analysis is fully non-nested:
```{r}
lyt4c <- basic_table() |>
  split_cols_by("ARM") |>
  split_rows_by("RACE", split_fun = keep_2_levels("RACE")) |>
  split_rows_by("STRATA1", split_fun = keep_2_levels("STRATA1")) |>
  split_rows_by("SEX", split_fun = keep_2_levels("SEX")) |>
  analyze("AGE") |>
  analyze("BMRKR2", nested = FALSE, show_labels = "visible")

build_table(lyt4c, ex_adsl)
```

Thus whether to use explicit anchoring to generate top-level sections
is a trade-off between explicit clarity (`nested = FALSE`) and
robustness to this sort of slotting the described structure into a
larger table (`at_sibling = `).


### Anchor Resolution

Given a pre-existing layout, only certain elements are eligible to act
as nesting anchors.  For clarity, we will use *element* to refer to
any individual layout instruction that effect the resulting table
row-structure (i.e., `split_rows_by*` and `analyze`); furthermore we
will refer to an element named by `at_sibling` as the *anchor point*
and an element placed via `at_sibling` as the *anchored element*. For
convenience we will refer to elements which do not act as anchor
points nor anchored elements as *standard elements*.

Using this terminology, the general rules are as follows:


1. Elements nested within previous top-level elements are *not eligible*,
2. Elements nested within previous anchor points are *not eligible*,
3. For each previous anchor point, elements nested within anchored
   elements other than the most recently placed one are *not
   eligible*.


We can re-frame this into an algorithm to determine the list of
eligible elements like so:

1. All previous top-level elements,
2. the current top-level element and all standard elements nested
   directly within it until the first anchor point,
3. the anchor point and all of its anchored elements,
4. all standard elements nested within the most recent of this anchor
   point's anchored elements, until the next anchor point,
5. repeat (3)-(4) until no more anchor points along the path exist.

Viewed a certain way, this algorithm defines a horizon along the edge
of the branching structure defined by a layout.

To illustrate these rules, and this concept of a horizon, consider the
following illustrative - if analytically nonsensical - complex row
layout:


```{r}
complex_lyt <- basic_table() |>
  split_rows_by("STRATA1", split_fun = keep_2_levels("RACE")) |>
  split_rows_by("STRATA2", split_fun = keep_2_levels("STRATA2")) |>
  analyze("ARM") |>
  split_rows_by("SEX", split_fun = keep_2_levels("SEX")) |>
  split_rows_by("RACE", split_fun = keep_2_levels("RACE")) |>
  split_rows_by("STRATA1", split_fun = keep_2_levels("STRATA1")) |>
  analyze("BMRKR1") |>
  split_rows_by("BMRKR2", split_fun = keep_2_levels("BMRKR2"), at_sibling = "RACE") |>
  split_rows_by("COUNTRY", split_fun = keep_2_levels("COUNTRY")) |>
  analyze("AGE") |>
  split_rows_by("SITEID", split_fun = drop_split_levels, at_sibling = "RACE") |>
  split_rows_by("BEP01FL", split_fun = keep_2_levels("BEP01FL")) |>
  analyze("AGE")
```

We can get the list of eligible anchor points via `get_anchor_list`:


```{r}
get_row_anchor_list(complex_lyt)
```

We can see that our first `STRATA1` split is eligible, but the
`STRATA2` split nested within it and the `ARM` analysis nested within
that are not. Then, for the current top-level structure, `SEX`(std element)
`RACE`(anchor pt), `BMRKR2` (anchored element), `SITEID` (anchored element),
`BEP01FL` (std element), and `AGE` (std element) are eligible.


### Order Of Intermediate Nesting Placement

The rules above imply a particular order required to place
intermediately nested elements anchored to different points within the
same top-level structure:

*When you intend to anchor multiple points to different elements in a
sequence of splits, place them in order from most deeply nested anchor
point to least deeply nested anchor point.*


We can see this in practice, consider the following sequence of
splitting layout instructions (ending, as always, with an `analyze`):

```{r}
lyt_stack <- basic_table() |>
  split_rows_by("STRATA1", split_fun = keep_2_levels("STRATA1")) |>
  split_rows_by("STRATA2", split_fun = keep_2_levels("STRATA2")) |>
  split_rows_by("RACE", split_fun = keep_2_levels("RACE")) |>
  split_rows_by("SEX", split_fun = keep_2_levels("SEX")) |>
  analyze("AGE")
```

Now suppose we want to place additional `analyze`'s as siblings to the
`STRATA2` and `RACE` splits.

If we do so in that order, our first placement will work:

```{r}
lyt_stack2 <- lyt_stack |>
  analyze("BMRKR1", at_sibling = "STRATA2")
```

But we will get an error when attempting to anchor another `analyze` onto `RACE`:

```{r, error = TRUE}
lyt_stack2 |>
  analyze("BMRKR2", at_sibling = "RACE")
```

If, however, we anchor our `BMRKR2` analyze to `RACE` *first*, and
then place our `BMRKR1` analyze to `STRATA2`, we can achieve both
placements:


```{r}
lyt_stack3 <- lyt_stack |>
  analyze("BMRKR2", at_sibling = "RACE", show_labels = "visible") |>
  analyze("BMRKR1", at_sibling = "STRATA2", show_labels = "visible")
```


Thus we can build the (somewhat lengthy) desired table:

```{r}
build_table(lyt_stack3, ex_adsl)
```

Phrased a different way, placing an anchored element at an anchor
point diverts the stream of eligible nested elements after that point
from those nested within the anchor point, or previously placed
anchored elements, to those nested within the newly placed anchored
element.

We can consider the current state to help us visualize the eligible
anchor points by printing our current layout:

```{r}
lyt_stack
```


# Designing Multi-Section Row-Layouts To Support Subgroup Variants

In setting with standardized table outputs, we commonly want both
all-patient and split-by-subgroups variants of a given core table
structure. Intermediate nesting allows us to develop layouts with this
in mind as we will see in this section.

Consider a table layout with multiple sections in row space, e.g., an
overall analysis, an analysis split by `RACE` and the same analysis
split by `SEX`:

```{r}
lyt <- basic_table() |>
  split_cols_by("ARM") |>
  analyze("AGE") |>
  split_rows_by("RACE", split_fun = keep_2_levels("RACE")) |>
  analyze("AGE") |>
  split_rows_by("SEX", split_fun = keep_2_levels("SEX")) |>
  analyze("AGE")

build_table(lyt, ex_adsl)
```

Now supposing we want the same structure for each strata in our
sample, if we apply `split_rows_by("STRATA1")` as the first row
instruction, we do not get the desired table, because each
`split_rows_by` that follows an `analyze` is `nested = FALSE` by
default, bringing it all the way to the top level, ie.e, outside of
our new strata splitting:


```{r}
lyt2 <- basic_table() |>
  split_cols_by("ARM") |>
  split_rows_by("STRATA1", split_fun = keep_2_levels("STRATA1")) |>
  analyze("AGE") |>
  split_rows_by("RACE", split_fun = keep_2_levels("RACE")) |>
  analyze("AGE") |>
  split_rows_by("SEX", split_fun = keep_2_levels("SEX")) |>
  analyze("AGE")

build_table(lyt2, ex_adsl)
```

This is not a subgroup variant of our table.

If we anchor those `split_rows_by` (each of which essentially starts a
new section of the layout in row space) to our overall `AGE` analyze
call, we will get the same table for the full variant:

```{r}
lyt_good <- basic_table() |>
  split_cols_by("ARM") |>
  analyze("AGE") |>
  split_rows_by("RACE",
    split_fun = keep_2_levels("RACE"),
    at_sibling = "AGE"
  ) |>
  analyze("AGE") |>
  split_rows_by("SEX",
    split_fun = keep_2_levels("SEX"),
    at_sibling = "AGE"
  ) |>
  analyze("AGE")

build_table(lyt_good, ex_adsl)
```

Crucially, however, when we prepend a new row splitting instruction to
the sequence of row layout instructions, we immediately get our
desired subgroup variant with no extra effort required:


```{r}
lyt_good_subgrp <- basic_table() |>
  split_cols_by("ARM") |>
  split_rows_by("STRATA1", split_fun = keep_2_levels("STRATA1")) |>
  analyze("AGE") |>
  split_rows_by("RACE",
    split_fun = keep_2_levels("RACE"),
    at_sibling = "AGE"
  ) |>
  analyze("AGE") |>
  split_rows_by("SEX",
    split_fun = keep_2_levels("SEX"),
    at_sibling = "AGE"
  ) |>
  analyze("AGE")

build_table(lyt_good, ex_adsl)
```

Note can use either standard splitting or splitting with `page_by =
TRUE` when injecting our subgroups, depending on the desired behavior,
with no other changes.

Thus, it is good practice to anchor all top level (seemingly
non-nested) row instructions after the first to that first instruction
to make our layouts easily support the creation of subgroup variants.
