melt() and dcast() in dataprep
0.1.8 are drop-in replacements for reshape2::melt /
reshape2::dcast, but the underlying code is written in C++
with SIMD (AVX2 / AVX-512) and optional OpenMP parallelism. Both
functions produce output identical to reshape2,
data.table, tidyr, pandas,
polars, dask, and duckdb on every
tested shape, within tol = 1e-12.
The speed-up relative to each of the seven alternatives spans:
| Operation | Range across both hosts |
|---|---|
melt |
0.6–1628.9× |
dcast |
1.9–799.8× |
The median across all tested cells and all competitors is 11.3× on
Ubuntu and 5.6× on Windows for melt, and 46.5× on Ubuntu
and 41.4× on Windows for dcast. The mean is 67.8× and 46.6×
for melt, and 90.2× and 76.6× for dcast,
respectively. The melt median is pulled down by the 1e5-row
small tables; at larger scales the speed-up is much higher. Full tables
— including mean, median, and per-competitor ranges — are in
vignette("dataprep-performance") and
README.md.
The sub-1.0× cells are polars at 1e7 rows × 10 id + 9
val on Ubuntu (0.6×), and polars at 1e5 rows × 10 id + 9
val on Windows (0.7×). Every other cell has dataprep ahead
of or on par with the fastest competitor.
Benchmarks were run on two reference hosts. Only the core
configuration is listed here; full hardware details are in
README.md.
Ubuntu 25.10 — 2× AMD EPYC 9965 192-Core (384 physical / 768 logical cores), 1.0 TiB (16 × 64 GiB Micron, DDR5-5600, Multi-bit ECC), full AVX-512; R 4.5.1, g++ 15.2.0.
Windows 11 Pro for Workstations — 2× AMD EPYC 7B12 64-Core (128 physical / 128 logical cores), about 224 GiB RAM, no AVX-512; R 4.6.1 (ucrt), GCC 14.3.0.
Software versions on both hosts: data.table 1.18.6.1,
reshape2 1.4.5, tidyr 1.3.2,
reticulate 1.47.0; Python 3.13.7 (Ubuntu) / 3.13.15
(Windows), pandas 3.0.6, polars 1.44.2
(runtime rt64), dask 2026.8.0, duckdb
1.5.5.
Reproducing the benchmarks:
melt()Non-numeric and factor columns are treated as IDs by default.
na.rmmajor = "row" vs
major = "col"Row-major ("row") is usually faster when there are few
measure columns; column-major ("col") is often faster when
there are many measure columns because bulk copies dominate.
major = NULL (the default) is equivalent to
"col"; there is no automatic switching based on the input
shape.
cores controls OpenMP.
options(dataprep.cores = ...) sets a default for the whole
session.
melt() is fastmelt_cpp uses six design choices that matter at
scale.
majormelt() produces the same long-format table as
reshape2::melt (column-major, major = "col")
or tidyr::pivot_longer (row-major,
major = "row"). The two layouts have very different memory
access patterns:
Column-major writes each output column as one
contiguous block: for measure column k, the output block
[k * n .. (k+1) * n) is a direct memcpy of the
input column. This is the fastest possible path when the number of value
columns is moderate.
Row-major writes every row as
n_meas consecutive doubles. For small n_meas
this is compact, but for large n_meas it requires a per-row
transpose.
major = NULL (the default) is equivalent to
"col"; there is no automatic switching based on the input
shape.
For large outputs, melt_cpp uses AVX-512 or AVX2
streaming stores (_mm512_stream_pd,
_mm256_stream_si256) to write directly to memory, bypassing
the CPU cache. This avoids the cache pollution that would otherwise
evict useful input data, and it is the reason the 1e8-row
melt finishes in under 0.5 s on Ubuntu (vs 2.6 s for the
next-fastest engine, polars) and in under 1.3 s on Windows.
A _mm_sfence() is issued at the end of each streaming
region to guarantee visibility.
Output vectors larger than 512 KB are allocated through
Rf_allocVector() and then hinted with
madvise(MADV_HUGEPAGE), so the kernel can back them with 2
MB pages. This reduces first-touch page faults on the 1e8-row case.
There is no custom R_allocator_t, no
MAP_POPULATE, and no free pool: those were described in
earlier drafts but are not part of the shipped 0.1.8 backend.
Two fast paths are used. Tiny inputs (n <= 2048,
n_meas <= 64, n_id <= 8) route to
melt_tiny_cpp; small inputs
(total <= 131072, n_meas <= 256) route
to melt_small_cpp. Both skip hugepage hinting, OpenMP
setup, and thread-cap detection. This is what makes melt()
the fastest engine even on 1e3-row tables, where the initialization
overhead of the other backends dominates their runtimes.
For the row-major path, the per-thread block size is computed from
the L3 cache size (read once from
/sys/devices/system/cpu/cpu0/cache/index3/size on Linux).
Each block is sized to fit in half of L3, which keeps both the input
reads and the output writes inside the cache for the duration of the
block.
The "data.frame" class tag, the "factor"
class tag, the default "variable" / "value"
column names, and the factor levels vector are constructed once per
process and reused afterwards via R_PreserveObject. This
removes a small but measurable per-call cost that shows up on the
small-input benchmarks.
dcast()If a (id, variable) pair appears more than once, pass
fun.aggregate. The default is “last occurrence wins”,
matching data.table::dcast’s
fun.aggregate = NULL behaviour.
dcast() is fastdcast_cpp performs four design choices that matter at
scale.
When the input is a canonical melt() output —
variable is periodic and every id column is
constant within one period — dcast_cpp skips the hash
tables entirely and performs a tile transpose: each tile of
TILE x period doubles is read contiguously into an L1
buffer, transposed in place, and written contiguously to the output
columns. This keeps both reads and writes sequential and enables OpenMP
parallelisation. The cost per tile is O(TILE * period),
independent of the number of levels, which is why wide-level tables
scale well.
For block-aligned input with few id columns,
dcast_cpp packs the row key into a single
uint64_t. When the combined bit budget of the
id columns exceeds 64, it falls back to a 96-bit
fingerprint (uint64_t h1 + uint32_t h2) computed by a 4-way
parallel FNV-1a and stored in a 16-byte slot table, which is more
cache-friendly than the previous 128-bit scheme.
When the input is not sorted, a 4-pass LSD radix sort over the packed
keys replaces the hash table entirely. For inputs whose rows are already
grouped — which is the common case after a melt() +
arrange() pipeline — the sort is skipped.
When the block path is taken and the id column is a
permutation of 1..n_blocks (the most common case for
canonical melt() output), a direct-index shortcut bypasses
both the hash table and the sort.
The full pipeline is described in the dcast() help
page.
The dcast 1e6 × 100 id cell is the only
case in the entire benchmark suite where a competitor reaches a
single-digit ratio: polars at 6.7× on Ubuntu and 2.6× on
Windows. Both remain behind dataprep. This is because
polars’s SIMD hash is competitive when the row key is very
wide (100 columns), while dcast_cpp’s 96-bit fingerprint
verification is O(n_id) per row in that regime.
melt() and dcast() produce output identical
to reshape2 (the reference implementation) on every tested
cell. All pairs of engines agree pairwise within
tol = 1e-12.
melt consistency| rows | n_id | n_val | engines passed | pairwise |
|---|---|---|---|---|
| 1,000 | 1 | 9 | 8/8 | all consistent |
| 100,000 | 1 | 9 | 8/8 | all consistent |
| 1,000 | 1 | 100 | 8/8 | all consistent |
| 10,000 | 10 | 10 | 8/8 | all consistent |
Engines: dataprep, reshape2,
data.table, tidyr, pandas,
polars, dask, duckdb.
dcast consistency| n_long | n_id | n_levels | engines passed | pairwise |
|---|---|---|---|---|
| 5,000 | 2 | 5 | 8/8 | all consistent |
| 50,000 | 1 | 50 | 8/8 | all consistent |
| 50,000 | 10 | 10 | 8/8 | all consistent |
| 1,000,000 | 1 | 10 | 8/8 | all consistent |
The consistency scripts are shipped under inst/:
The tables below summarise the two hosts in one place. Full per-cell
tables are in vignette("dataprep-performance").
| Operation | Min | Median | Mean | Max |
|---|---|---|---|---|
melt() |
0.6× (polars @ 1e7 × 19 × 10 × 9) | 11.3× | 67.8× | 1628.9× (dask @ 1e3 × 10001 × 1 × 10000) |
dcast() |
1.9× (reshape2 @ 1e3 × 1 × 10) | 46.5× | 90.2× | 463.8× (data.table @ 1e8 × 1 × 100) |
| Operation | Min | Median | Mean | Max |
|---|---|---|---|---|
melt() |
0.7× (polars @ 1e5 × 19 × 10 × 9) | 5.6× | 46.6× | 893.2× (dask @ 1e3 × 10001 × 1 × 10000) |
dcast() |
2.6× (polars @ 1e6 × 100 × 10) | 41.4× | 76.6× | 799.8× (duckdb @ 1e8 × 1 × 100) |
For melt(), on the largest cells (1e8 rows, 8 GB of
input), dataprep is the only engine that completes within
2.5 s: under 0.3 s on Ubuntu and under 1.0 s on Windows.
sessionInfo()
#> R version 4.5.1 (2025-06-13)
#> Platform: x86_64-pc-linux-gnu
#> Running under: Ubuntu 25.10
#>
#> Matrix products: default
#> BLAS: /usr/lib/x86_64-linux-gnu/openblas-openmp/libblas.so.3
#> LAPACK: /usr/lib/x86_64-linux-gnu/openblas-openmp/libopenblasp-r0.3.30.so; LAPACK version 3.12.0
#>
#> locale:
#> [1] LC_CTYPE=zh_CN.UTF-8 LC_NUMERIC=C
#> [3] LC_TIME=zh_CN.UTF-8 LC_COLLATE=zh_CN.UTF-8
#> [5] LC_MONETARY=zh_CN.UTF-8 LC_MESSAGES=zh_CN.UTF-8
#> [7] LC_PAPER=zh_CN.UTF-8 LC_NAME=C
#> [9] LC_ADDRESS=C LC_TELEPHONE=C
#> [11] LC_MEASUREMENT=zh_CN.UTF-8 LC_IDENTIFICATION=C
#>
#> time zone: Asia/Shanghai
#> tzcode source: system (glibc)
#>
#> attached base packages:
#> [1] stats graphics grDevices utils datasets methods base
#>
#> other attached packages:
#> [1] dataprep_0.1.8
#>
#> loaded via a namespace (and not attached):
#> [1] vctrs_0.7.3 cli_3.6.6 knitr_1.52 rlang_1.3.0
#> [5] xfun_0.61 otel_0.2.0 generics_0.1.4 S7_0.2.2
#> [9] jsonlite_2.0.0 glue_1.8.1 htmltools_0.5.9 sass_0.4.10
#> [13] scales_1.4.0 rmarkdown_2.32 grid_4.5.1 tibble_3.3.1
#> [17] evaluate_1.0.5 jquerylib_0.1.4 fastmap_1.2.0 yaml_2.3.12
#> [21] lifecycle_1.0.5 compiler_4.5.1 dplyr_1.2.1 RColorBrewer_1.1-3
#> [25] pkgconfig_2.0.3 Rcpp_1.1.2 farver_2.1.2 digest_0.6.39
#> [29] R6_2.6.1 tidyselect_1.2.1 pillar_1.11.1 parallel_4.5.1
#> [33] magrittr_2.0.5 bslib_0.12.0 withr_3.0.3 tools_4.5.1
#> [37] gtable_0.3.6 ggplot2_4.0.3 cachem_1.1.0