SSLfmm: A Practical Workflow

Geoffrey J. McLachlan and Jinran Wu

Overview

SSLfmm fits semi-supervised Gaussian finite-mixture classifiers when class labels are observed for only part of a sample. The package provides a common interface for complete-case (cc), missing-completely-at-random (mcar), entropy-dependent missing-at-random (mar), and mixed MCAR/MAR analyses. This vignette illustrates the main simulation, fitting, prediction, and assessment workflow using only the public package API.

Simulate partially labelled data

library(SSLfmm)

mu <- matrix(c(-1.5, 1.5), nrow = 1, ncol = 2)
sim <- simulate_mixed_missingness(
  n = 120,
  pi = c(0.5, 0.5),
  mu = mu,
  sigma = matrix(1, 1, 1),
  alpha = 0.10,
  mar_rate = 0.25,
  seed = 2026
)

head(sim$data)
#>           x1         en missing label truth observed_missing latent_missing
#> 1 -1.6414707 0.04276823   FALSE     1     1            FALSE          FALSE
#> 2 -1.3326701 0.09023544   FALSE     1     1            FALSE          FALSE
#> 3  0.7890114 0.29252504   FALSE     2     2            FALSE          FALSE
#> 4  1.3010109 0.09718721   FALSE     2     2            FALSE          FALSE
#> 5 -1.6987886 0.03709502   FALSE     1     1            FALSE          FALSE
#> 6  1.2186267 0.11759460   FALSE     2     2            FALSE          FALSE
#>   missing_source    prob_mar    entropy
#> 1       observed 0.009206644 0.04276823
#> 2       observed 0.080268792 0.09023544
#> 3       observed 0.748322109 0.29252504
#> 4       observed 0.098318415 0.09718721
#> 5       observed 0.006026648 0.03709502
#> 6       observed 0.161889212 0.11759460
table(sim$data$missing_source, useNA = "ifany")
#> 
#>      mar     mcar observed 
#>       22       10       88

The simulated data contain the feature columns (x1, …, xp), the observed label (label), the complete reference class (truth, available because this is a simulation), and information about the missing-label mechanism.

Fit a mixed missingness model

The unknown-source mixed model estimates the contribution of MCAR and MAR without requiring the source of each missing label to be supplied.

x <- as.matrix(sim$data["x1"])
fit <- fit_sslfmm(
  x,
  sim$data$label,
  g = 2,
  method = "mixed",
  covariance_type = "equal",
  indicator = "unknown",
  n_starts = 3,
  seed = 2027
)

fit
#> SSLfmm fit
#>   method:             mixed
#>   covariance:         equal
#>   source indicator:    unknown
#>   components:          2
#>   log-likelihood:      -275.355
#>   convergence code:    0
#>   optimizer:           nlminb
#>   alpha:               0.062865
#>   xi:                  6.3379, 4.1985
summary(fit)
#> $method
#> [1] "mixed"
#> 
#> $covariance_type
#> [1] "equal"
#> 
#> $indicator
#> [1] "unknown"
#> 
#> $g
#> [1] 2
#> 
#> $pi
#>         1         2 
#> 0.4392433 0.5607567 
#> 
#> $mu
#>          x1
#> 1 -1.650259
#> 2  1.620010
#> 
#> $sigma
#>          [,1]
#> [1,] 1.037119
#> 
#> $xi
#>      xi0      xi1 
#> 6.337916 4.198468 
#> 
#> $alpha
#> [1] 0.06286494
#> 
#> $loglik
#> [1] -275.355
#> 
#> $convergence
#> [1] 0
#> 
#> $optimizer
#> [1] "nlminb"
#> 
#> $missing_rate
#> [1] 0.2666667
#> 
#> attr(,"class")
#> [1] "summary.SSLfmm"

For applications in which the source of missingness is recorded, use indicator = "known" together with missing_source.

Prediction and uncertainty

pred_class <- predict(fit, x)
posterior <- predict(fit, x, type = "posterior")
entropy <- predict(fit, x, type = "entropy")

head(pred_class)
#> [1] 1 1 2 2 1 2
head(posterior)
#>               1           2
#> [1,] 0.99249013 0.007509866
#> [2,] 0.98035866 0.019641338
#> [3,] 0.05842256 0.941577438
#> [4,] 0.01219688 0.987803120
#> [5,] 0.99372406 0.006275944
#> [6,] 0.01575794 0.984242059
head(entropy)
#> [1] 0.04421639 0.09663996 0.22260490 0.06586866 0.03808172 0.08103506

Posterior probabilities quantify class uncertainty. Entropy provides a compact summary of uncertainty and is also the quantity used by the entropy-dependent MAR mechanism implemented in the package.

Classification assessment

Because the simulation retains the complete class labels, prediction can be assessed directly.

perf <- classification_performance(
  sim$data$truth,
  pred_class,
  posterior
)

perf$metrics
#>          accuracy        error_rate balanced_accuracy   macro_precision 
#>        0.94166667        0.05833333        0.93855219        0.94416027 
#>      macro_recall          macro_f1          log_loss 
#>        0.93855219        0.94074074        0.12922536
perf$confusion_matrix
#>      predicted
#> truth  1  2
#>     1 49  5
#>     2  2 64

In genuinely partially labelled data, performance on observations whose labels are missing cannot usually be evaluated without an external reference set.

Included Blood Transfusion example data

The package also includes the semi-synthetic blood_transfusion data set used in the software-paper application.

data("blood_transfusion")
head(blood_transfusion)
#>   id truth observed missing_indicator    Recency  Frequency        Time
#> 1  1     2        2          observed -0.9272783  7.6182487  2.61388444
#> 2  2     2       NA               mar -1.1743323  1.2818805 -0.25770846
#> 3  3     2       NA               mar -1.0508053  1.7956401  0.02945083
#> 4  4     2        2          observed -0.9272783  2.4806529  0.43967839
#> 5  5     1       NA               mar -1.0508053  3.1656657  1.75240657
#> 6  6     1       NA               mar -0.6802243 -0.2593982 -1.24225460
table(blood_transfusion$missing_indicator)
#> 
#>      mar     mcar observed 
#>       91       75      582

The truth column is retained for evaluation of the semi-synthetic example, whereas observed is the partially observed response supplied to model-fitting functions.

Reproducibility and development

The development repository is https://github.com/wujrtudou/SSLfmm and issues can be reported at https://github.com/wujrtudou/SSLfmm/issues. The package contains automated testthat tests for its public API, simulation return contracts, fitting, prediction, input validation, and the included case-study data.