SSLfmm fits semi-supervised Gaussian finite-mixture
classifiers when class labels are observed for only part of a sample.
The package provides a common interface for complete-case
(cc), missing-completely-at-random (mcar),
entropy-dependent missing-at-random (mar), and mixed
MCAR/MAR analyses. This vignette illustrates the main simulation,
fitting, prediction, and assessment workflow using only the public
package API.
library(SSLfmm)
mu <- matrix(c(-1.5, 1.5), nrow = 1, ncol = 2)
sim <- simulate_mixed_missingness(
n = 120,
pi = c(0.5, 0.5),
mu = mu,
sigma = matrix(1, 1, 1),
alpha = 0.10,
mar_rate = 0.25,
seed = 2026
)
head(sim$data)
#> x1 en missing label truth observed_missing latent_missing
#> 1 -1.6414707 0.04276823 FALSE 1 1 FALSE FALSE
#> 2 -1.3326701 0.09023544 FALSE 1 1 FALSE FALSE
#> 3 0.7890114 0.29252504 FALSE 2 2 FALSE FALSE
#> 4 1.3010109 0.09718721 FALSE 2 2 FALSE FALSE
#> 5 -1.6987886 0.03709502 FALSE 1 1 FALSE FALSE
#> 6 1.2186267 0.11759460 FALSE 2 2 FALSE FALSE
#> missing_source prob_mar entropy
#> 1 observed 0.009206644 0.04276823
#> 2 observed 0.080268792 0.09023544
#> 3 observed 0.748322109 0.29252504
#> 4 observed 0.098318415 0.09718721
#> 5 observed 0.006026648 0.03709502
#> 6 observed 0.161889212 0.11759460
table(sim$data$missing_source, useNA = "ifany")
#>
#> mar mcar observed
#> 22 10 88The simulated data contain the feature columns (x1, …,
xp), the observed label (label), the complete
reference class (truth, available because this is a
simulation), and information about the missing-label mechanism.
The unknown-source mixed model estimates the contribution of MCAR and MAR without requiring the source of each missing label to be supplied.
x <- as.matrix(sim$data["x1"])
fit <- fit_sslfmm(
x,
sim$data$label,
g = 2,
method = "mixed",
covariance_type = "equal",
indicator = "unknown",
n_starts = 3,
seed = 2027
)
fit
#> SSLfmm fit
#> method: mixed
#> covariance: equal
#> source indicator: unknown
#> components: 2
#> log-likelihood: -275.355
#> convergence code: 0
#> optimizer: nlminb
#> alpha: 0.062865
#> xi: 6.3379, 4.1985
summary(fit)
#> $method
#> [1] "mixed"
#>
#> $covariance_type
#> [1] "equal"
#>
#> $indicator
#> [1] "unknown"
#>
#> $g
#> [1] 2
#>
#> $pi
#> 1 2
#> 0.4392433 0.5607567
#>
#> $mu
#> x1
#> 1 -1.650259
#> 2 1.620010
#>
#> $sigma
#> [,1]
#> [1,] 1.037119
#>
#> $xi
#> xi0 xi1
#> 6.337916 4.198468
#>
#> $alpha
#> [1] 0.06286494
#>
#> $loglik
#> [1] -275.355
#>
#> $convergence
#> [1] 0
#>
#> $optimizer
#> [1] "nlminb"
#>
#> $missing_rate
#> [1] 0.2666667
#>
#> attr(,"class")
#> [1] "summary.SSLfmm"For applications in which the source of missingness is recorded, use
indicator = "known" together with
missing_source.
pred_class <- predict(fit, x)
posterior <- predict(fit, x, type = "posterior")
entropy <- predict(fit, x, type = "entropy")
head(pred_class)
#> [1] 1 1 2 2 1 2
head(posterior)
#> 1 2
#> [1,] 0.99249013 0.007509866
#> [2,] 0.98035866 0.019641338
#> [3,] 0.05842256 0.941577438
#> [4,] 0.01219688 0.987803120
#> [5,] 0.99372406 0.006275944
#> [6,] 0.01575794 0.984242059
head(entropy)
#> [1] 0.04421639 0.09663996 0.22260490 0.06586866 0.03808172 0.08103506Posterior probabilities quantify class uncertainty. Entropy provides a compact summary of uncertainty and is also the quantity used by the entropy-dependent MAR mechanism implemented in the package.
Because the simulation retains the complete class labels, prediction can be assessed directly.
perf <- classification_performance(
sim$data$truth,
pred_class,
posterior
)
perf$metrics
#> accuracy error_rate balanced_accuracy macro_precision
#> 0.94166667 0.05833333 0.93855219 0.94416027
#> macro_recall macro_f1 log_loss
#> 0.93855219 0.94074074 0.12922536
perf$confusion_matrix
#> predicted
#> truth 1 2
#> 1 49 5
#> 2 2 64In genuinely partially labelled data, performance on observations whose labels are missing cannot usually be evaluated without an external reference set.
The package also includes the semi-synthetic
blood_transfusion data set used in the software-paper
application.
data("blood_transfusion")
head(blood_transfusion)
#> id truth observed missing_indicator Recency Frequency Time
#> 1 1 2 2 observed -0.9272783 7.6182487 2.61388444
#> 2 2 2 NA mar -1.1743323 1.2818805 -0.25770846
#> 3 3 2 NA mar -1.0508053 1.7956401 0.02945083
#> 4 4 2 2 observed -0.9272783 2.4806529 0.43967839
#> 5 5 1 NA mar -1.0508053 3.1656657 1.75240657
#> 6 6 1 NA mar -0.6802243 -0.2593982 -1.24225460
table(blood_transfusion$missing_indicator)
#>
#> mar mcar observed
#> 91 75 582The truth column is retained for evaluation of the
semi-synthetic example, whereas observed is the partially
observed response supplied to model-fitting functions.
The development repository is https://github.com/wujrtudou/SSLfmm and issues can be
reported at https://github.com/wujrtudou/SSLfmm/issues. The package
contains automated testthat tests for its public API,
simulation return contracts, fitting, prediction, input validation, and
the included case-study data.