Three behaviour changes affect row and column counts. All three are bug fixes, but each one changes the output on real data. If your downstream analysis depends on exact row counts, read the quantified comparison in the last section before upgrading.
Besides the behaviour changes above, several functions changed their
argument interface between 0.1.5 and 0.1.8. The old start /
end pair was replaced by a single cols
argument (character names, integer indices, or logical mask):
| Function | 0.1.5 arguments | 0.1.8 arguments |
|---|---|---|
varidele |
start, end |
cols |
obsedele |
start, end |
cols |
condextr |
start, end |
cols |
percoutl |
start, end |
cols |
optisolu |
start, end |
cols |
dataprep |
start, end |
cols |
descdata |
start, end |
cols |
descplot |
start, end |
cols |
percdata |
start, end |
cols |
percplot |
start, end |
cols |
shorvalu |
start, end |
cols |
melt |
cols (id columns) |
id / measure.vars |
A call such as varidele(data, 3, 15) was interpreted as
start = 3, end = 15 in 0.1.5, but in 0.1.8 it is parsed as
cols = 3 (the 15 is dropped or becomes
fraction). You must rewrite it as
varidele(data, cols = 3:15). The same applies to every
function in the table.
obsedele() now scans each column
independently. The 0.1.5 implementation collapsed all selected
columns into one long vector before computing missing runs; this changed
NA run boundaries and could both over-delete boundary rows
and retain rows that should have been deleted. The 0.1.8 implementation
scans each column independently: a row is deleted when any
selected column has a missing run longer than half minutes
on both sides.
half is now always in minutes, and the
boundary is inclusive. In 0.1.5, half counted grid
rows in units of by: with
by = "5 min", half = 30 the effective window was 150
minutes. In 0.1.8, half is always in minutes, independent
of by. Rows whose nearest anchor is exactly
half minutes away are retained (“within half
minutes” is a <= condition).
optisolu() no longer crashes with
cores > 16. The 0.1.5
parallel::makeCluster() path exhausted memory when the
worker processes each received a full copy of the input. The 0.1.8
implementation loads the package on each worker, exports the input data
only once per worker, and runs each (interval, times) case
in a separate task, so cores = 64 and
cores = NULL (automatic) are both safe. Note that the
optimal parameter values returned by optisolu() may differ
slightly between 0.1.5 and 0.1.8 because the underlying outlier-marking
and observation-deletion backends have changed. Re-run
optisolu() after upgrading if exact parameter values
matter.
The full description is in news(package = "dataprep").
The design reasoning behind the pipeline is in
vignette("dataprep-philosophy").
Consider five observations sampled every 10 minutes, with a valid value only at the two ends:
df <- data.frame(
date = as.POSIXct("2024-01-01 00:00:00", tz = "UTC") + 0:4 * 600,
x = c(1, NA, NA, NA, 5)
)
obsedele(df, cols = "x", half = 30)
#> date x
#> 1 2024-01-01 00:00:00 1
#> 2 2024-01-01 00:10:00 NA
#> 3 2024-01-01 00:20:00 NA
#> 4 2024-01-01 00:30:00 NA
#> 5 2024-01-01 00:40:00 5The three interior rows are each 10, 20, or 30 minutes from the
nearest valid anchor. Under 0.1.5 the third row was deleted because the
delete condition was inclusive (dl >= half); under 0.1.8
it is retained because the delete condition is strict
(dl > half && dr > half), so
dl == half satisfies “within half
minutes”.
Consider two channels with mutually exclusive missing runs:
df <- data.frame(
date = as.POSIXct("2024-01-01 00:00:00", tz = "UTC") + 0:9 * 600,
x = c(1, NA, NA, NA, NA, NA, NA, NA, NA, 5),
y = c(NA, NA, NA, NA, 2, NA, NA, NA, NA, NA)
)
nrow(obsedele(df, cols = c("x", "y"), half = 60))
#> [1] 10Under 0.1.5, the x and y NA
runs were merged before computing the run length. This changed the run
boundaries and could either over-delete rows or retain rows that should
have been deleted. Under 0.1.8 the two channels are checked
independently: a row is deleted when any selected column has a
run longer than half minutes on both sides.
The retention criterion has been stable across releases, but the implementation has improved steadily. The table below summarises the three generations:
| Release | Strategy | Complexity |
|---|---|---|
| 0.1.0 | Borrowed running mean: expand the series onto a regular grid with
tidyr::complete(), compute a 59-minute centred moving
average on a temporary column, and use its emptiness pattern to flag
long runs. |
O(grid length) per subset |
| 0.1.5 | Run-length encoding: use data.table::rleid() and
rowid() to collapse consecutive NAs into runs.
The retention threshold was half * num grid rows, where
num is the leading number in the by string:
with by = "5 min", half = 30 this was a 150-minute
window. |
O(n) time, O(n) temporary storage |
| 0.1.8 | Anchor scan: for each missing value, look up the nearest non-missing anchor on each side and compare the two time distances directly. No grid, no run-length state. | O(n) time, O(1) extra allocation per
column |
Each generation produces the same deletion decision on the same
input, but the constant factors shrink. The 0.1.8 anchor scan is the
first version that is fast enough to run interactively on full-year
data: on SMEAR I Varrio 2025 (49,422 rows × 61 numeric channels),
obsedele() now runs in 0.05 s against 11.6 s in 0.1.5 on
Ubuntu 25.10 (about 232×), and in 0.035 s against 22.5 s on Windows 11
Pro for Workstations (about 648×). See
vignette("dataprep-performance") for the full
benchmark.
On SMEAR I Varrio 2025 (49,422 rows × 61 numeric channels, 10-minute sampling), running the same pipeline with the same parameters:
| Stage | 0.1.5 | 0.1.8 | Δ |
|---|---|---|---|
varidele |
25 columns deleted | 25 columns deleted | 0 |
obsedele |
1,494 rows deleted | 1,496 rows deleted | +2 |
condextr |
1,868 rows deleted | 1,863 rows deleted | −5 |
shorvalu |
50,376 NAs filled | 50,387 NAs filled | +11 |
dataprep final |
46,060 rows | 46,063 rows | +3 |
Net change: 0.006% of the input. The six rows that differ between
versions all sit at run boundaries where the anchor distance is within
one sampling interval of half minutes.
The six differing rows fall into two groups:
Rows kept by 0.1.8, deleted by 0.1.5 (3 rows).
These are rows whose nearest anchor is exactly half minutes
away. Under the 0.1.5 delete condition (dl >= half) they
were deleted; under the 0.1.8 condition
(dl > half && dr > half) they are retained.
Each of these rows has a valid anchor within one sampling interval of
half, so retaining them is consistent with the physical
constraint described in
vignette("dataprep-philosophy").
Rows deleted by 0.1.8, kept by 0.1.5 (3 rows, overlapping
with the above). These are rows that pass the check on some
columns but fail on at least one. Under 0.1.5 the merge-columns approach
effectively widened the anchor window on these rows; under 0.1.8 each
column is checked independently, so the row is deleted. These rows would
have been interpolated across a gap longer than half
minutes in at least one channel, which contradicts the design.
The net result is 3 additional rows in the final output.
For users who only call melt() and dcast(),
or who only use dataprep for descriptive statistics
(descdata, na_diagnose, percdata,
percplot, descplot), there is no behaviour
change. Those functions have been re-implemented in 0.1.8 for speed
(melt and dcast) or reorganised internally
(data_report, dry_run), but the output on
every tested input is identical to 0.1.5 except where noted in the news
file.
The 0.1.8 release rewrites every heavy cleaning routine in C++. The table below compares against 0.1.5 on three dataset sizes from the same source (SMEAR I Varrio forest). All numbers are speed-up ratios (0.1.5 time / 0.1.8 time); a value below 1.0× means 0.1.8 is slightly slower on that cell.
| Function | 500 rows | 7,640 rows | 49,422 rows (Ubuntu 25.10) |
|---|---|---|---|
varidele |
1.1× | 1.1× | 11.6× |
obsedele |
203× | 424× | 232× |
condextr |
196× | 217× | 1146× |
optisolu |
188× | 77× | 109× |
dataprep |
185× | 228× | 247× |
On Windows 11 Pro for Workstations, the same full-year pipeline gives
obsedele ≈ 648×, condextr ≈ 839×,
shorvalu ≈ 81×, optisolu ≈ 25× (at
cores = 32), and the integrated dataprep call
≈ 173×. varidele is around 1.17× on this cell; this is
expected, since varidele is a single
colMeans(is.na(.)) in both versions and the new code path
has little room for improvement.
Note on
optisolucores. The 0.1.5 implementation could crash whencores > 16. The benchmark above usedcores = 16for both versions to keep the comparison fair. 0.1.8 loads the package on each worker, exports the input data once per worker, and runs each(interval, times)case as a separate task, socores = 64is safe. The practical speed-up on a many-core host is larger than the table above. Also note that the optimal parameter values returned byoptisolu()may differ slightly between versions; re-runoptisolu()after upgrading if exact values matter.
Benchmarks and checks were run on two reference hosts. Only the core
configuration is listed here; full details are in
vignette("dataprep-performance").
Ubuntu 25.10 — R 4.5.1, g++ 15.2.0; 2× AMD EPYC 9965 192-Core (384 physical / 768 logical cores), 1.0 TiB (16 × 64 GiB Micron, DDR5-5600, Multi-bit ECC), full AVX-512.
Windows 11 Pro for Workstations — R 4.6.1 (ucrt), GCC 14.3.0; 2× AMD EPYC 7B12 64-Core (128 physical / 128 logical cores), about 224 GiB RAM.
Design philosophy —
vignette("dataprep-philosophy"). Why the pipeline has the
shape it does.
Cleaning pipeline walkthrough —
vignette("dataprep-cleaning"). Step-by-step execution on a
real dataset.
Performance and cross-engine consistency —
vignette("dataprep-performance"). Full benchmark tables and
8-engine consistency checks.
Fast reshaping —
vignette("dataprep-melt-dcast").
sessionInfo()
#> R version 4.5.1 (2025-06-13)
#> Platform: x86_64-pc-linux-gnu
#> Running under: Ubuntu 25.10
#>
#> Matrix products: default
#> BLAS: /usr/lib/x86_64-linux-gnu/openblas-openmp/libblas.so.3
#> LAPACK: /usr/lib/x86_64-linux-gnu/openblas-openmp/libopenblasp-r0.3.30.so; LAPACK version 3.12.0
#>
#> locale:
#> [1] LC_CTYPE=zh_CN.UTF-8 LC_NUMERIC=C
#> [3] LC_TIME=zh_CN.UTF-8 LC_COLLATE=zh_CN.UTF-8
#> [5] LC_MONETARY=zh_CN.UTF-8 LC_MESSAGES=zh_CN.UTF-8
#> [7] LC_PAPER=zh_CN.UTF-8 LC_NAME=C
#> [9] LC_ADDRESS=C LC_TELEPHONE=C
#> [11] LC_MEASUREMENT=zh_CN.UTF-8 LC_IDENTIFICATION=C
#>
#> time zone: Asia/Shanghai
#> tzcode source: system (glibc)
#>
#> attached base packages:
#> [1] stats graphics grDevices utils datasets methods base
#>
#> other attached packages:
#> [1] dataprep_0.1.8
#>
#> loaded via a namespace (and not attached):
#> [1] vctrs_0.7.3 cli_3.6.6 knitr_1.52 rlang_1.3.0
#> [5] xfun_0.61 otel_0.2.0 generics_0.1.4 S7_0.2.2
#> [9] jsonlite_2.0.0 glue_1.8.1 htmltools_0.5.9 sass_0.4.10
#> [13] scales_1.4.0 rmarkdown_2.32 grid_4.5.1 tibble_3.3.1
#> [17] evaluate_1.0.5 jquerylib_0.1.4 fastmap_1.2.0 yaml_2.3.12
#> [21] lifecycle_1.0.5 compiler_4.5.1 dplyr_1.2.1 RColorBrewer_1.1-3
#> [25] pkgconfig_2.0.3 Rcpp_1.1.2 farver_2.1.2 digest_0.6.39
#> [29] R6_2.6.1 tidyselect_1.2.1 pillar_1.11.1 parallel_4.5.1
#> [33] magrittr_2.0.5 bslib_0.12.0 withr_3.0.3 tools_4.5.1
#> [37] gtable_0.3.6 ggplot2_4.0.3 cachem_1.1.0