---
title: "Getting started with arulesCBA"
author: "Michael Hahsler"
output:
  rmarkdown::html_vignette:
    toc: true
vignette: >
  %\VignetteIndexEntry{Getting started with arulesCBA}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
library(arulesCBA)
set.seed(1234)
```

`arulesCBA` builds classifiers from class association rules (CARs). A CAR has
predictor conditions on its left-hand side and a class label on its right-hand
side. This makes the resulting model inspectable: a prediction can be traced
back to one or more rules.

This guide shows the typical workflow:

1. split a data frame into training and test data;
2. learn a CBA classifier;
3. inspect its rule base; and
4. predict classes and evaluate accuracy.

## Installation

Install the released version from CRAN:

```{r install, eval=FALSE}
install.packages("arulesCBA")
```

Then load the package:

```{r load-package}
library(arulesCBA)
```

Loading `arulesCBA` also makes the transaction and rule infrastructure from
`arules` available.

## Train a classifier

We use the numeric measurements in `iris` to predict `Species`. Keep a test set
aside so that evaluation uses observations that were not used to build the
classifier.

```{r split-data}
train_id <- sample(seq_len(nrow(iris)), 100)
iris_train <- iris[train_id, ]
iris_test <- iris[-train_id, ]

table(iris_train$Species)
table(iris_test$Species)
```

`CBA()` accepts the same formula-and-data interface as many R modeling
functions. Numeric predictors are discretized automatically because
association rules operate on items rather than continuous values.

```{r train-model}
classifier <- CBA(Species ~ ., data = iris_train)
classifier
```

The printed model reports the number of retained rules, the default class, and
the classification strategy. CBA ranks candidate rules, prunes the rule base,
and uses the first matching rule for prediction. If no rule matches, it uses
the default class.

## Inspect the learned rules

The rules are stored in `classifier$rules` as an `arules` `rules` object. The
first few rules can be displayed with `inspect()`:

```{r inspect-rules}
inspect(head(classifier$rules, 5))
```

A rule such as
`{Petal.Length=[-Inf,1.9)} => {Species=setosa}` says that observations whose
petal length falls in that interval are predicted as `setosa`. The exact
intervals depend on the training sample.

The most important rule quality measures are:

- **support:** the proportion of training observations covered by the rule's
  conditions;
- **confidence:** the proportion of covered observations that have the class
  shown on the right-hand side; and
- **lift:** the confidence relative to the prevalence of that class.

Quality measures are available as a data frame, so they can be summarized or
used to select rules.

```{r rule-quality}
summary(quality(classifier$rules)[, c("support", "confidence", "lift")])
```

## Predict and evaluate

Pass the original, undiscretized data frame to `predict()`. The classifier
stores the cut points learned from the training data and applies the same cut
points to new observations.

```{r predict}
prediction <- predict(classifier, iris_test)
head(prediction)
```

A confusion matrix shows which classes were confused. `accuracy()` calculates
the overall fraction classified correctly.

```{r evaluate}
table(predicted = prediction, observed = iris_test$Species)
accuracy(prediction, iris_test$Species)
```

For a reliable estimate of predictive performance, use repeated train/test
splits or cross-validation rather than reporting accuracy on the training
data.

## Control rule mining

The most commonly adjusted mining parameters are minimum support, minimum
confidence, and maximum rule length. They can be supplied directly to `CBA()`:

```{r tune-model}
classifier_tuned <- CBA(
  Species ~ .,
  data = iris_train,
  support = 0.05,
  confidence = 0.9,
  maxlen = 4
)
classifier_tuned
```

Lower support typically produces more candidate rules, while higher confidence
requires rules to be more reliable on the training data. Increasing `maxlen`
allows more conditions in a rule. These settings affect both runtime and model
complexity; very low support or a large `maxlen` can produce an extremely large
candidate rule set.

The same settings can instead be passed in a named list:

```{r parameter-list, eval=FALSE}
classifier <- CBA(
  Species ~ .,
  data = iris_train,
  parameter = list(support = 0.05, confidence = 0.9, maxlen = 4)
)
```

For imbalanced data, `balanceSupport = TRUE` lowers the minimum support for
minority classes relative to the majority class:

```{r balanced-support, eval=FALSE}
classifier_balanced <- CBA(
  class ~ .,
  data = training_data,
  support = 0.1,
  confidence = 0.8,
  balanceSupport = TRUE
)
```

## Prepare and mine rules separately

`CBA()` handles data preparation, rule mining, and pruning in one call. The
individual steps are also available when more control is needed.

`prepareTransactions()` discretizes numeric predictors and converts every row
to a transaction. Class values become items that can appear on the right-hand
side of a CAR.

```{r prepare-transactions}
iris_transactions <- prepareTransactions(Species ~ ., iris_train)
iris_transactions
inspect(head(iris_transactions, 3))
```

`mineCARs()` restricts the right-hand side of every mined rule to a value of
the response variable.

```{r mine-cars}
cars <- mineCARs(
  Species ~ .,
  iris_transactions,
  support = 0.1,
  confidence = 0.8,
  maxlen = 4,
  verbose = FALSE
)
cars
inspect(head(cars, 5))
```

These lower-level functions are useful for studying the candidate rules or for
constructing a custom classifier with `CBA_ruleset()`. For a first analysis,
the higher-level `CBA()` interface is usually sufficient.

## Other classifiers

The package includes several other associative classification algorithms,
including `FOIL()`, `RCAR()`, `CMAR()`, `CPAR()`, and `PRM()`. It also provides
wrappers for rule learners from `RWeka`. Some of these methods require Java and
additional suggested packages; see their help pages for requirements and
algorithm-specific options.

```{r help, eval=FALSE}
help(package = "arulesCBA")
?CBA
?mineCARs
?CBA_ruleset
```
