dataprep: performance and cross-engine consistency

library(dataprep)

Overview

dataprep 0.1.8 ships two reshaping backends, melt() and dcast(), benchmarked here against all seven major alternatives in the R and Python ecosystems:

Every cell is measured with a C++ steady-clock timer and an adaptive times rule (20 / 15 / 10 / 5 / 1 iterations based on warmup time). Two statistics are recorded per cell:

For every cell the tables also report the speed-up of dataprep relative to each competitor, so the reader can see the full gradient from “about the same” to “three orders of magnitude”.

Mean vs median

The four benchmark CSV files shipped under inst/extdata/ carry both the mean and the median of every per-cell timing sample. The scatter plot below puts them side by side: each point is one (tool, host, shape) combination, the x axis is the median in milliseconds and the y axis is the mean. Points on the 1:1 line mean the two statistics agree; points above the line mean the mean is inflated by a long right tail in the per-iteration timings.

suppressPackageStartupMessages(library(ggplot2))

read_bench <- function(fname, op) {
  p <- system.file("extdata", fname, package = "dataprep")
  d <- read.csv(p, stringsAsFactors = FALSE)
  d <- d[!d$skipped, c("tool", "mean", "median")]
  d$op <- op
  d
}

bench <- rbind(
  read_bench("bench_melt_ubuntu.csv",  "melt (Ubuntu)"),
  read_bench("bench_dcast_ubuntu.csv", "dcast (Ubuntu)"),
  read_bench("bench_melt_win.csv",     "melt (Windows)"),
  read_bench("bench_dcast_win.csv",    "dcast (Windows)")
)

ggplot(bench, aes(median, mean)) +
  geom_abline(slope = 1, intercept = 0,
              linetype = "dashed", colour = "grey50") +
  geom_point(alpha = 0.45, size = 1.4) +
  scale_x_log10() +
  scale_y_log10() +
  facet_wrap(~ tool, ncol = 4) +
  labs(x = "median (ms, log scale)",
       y = "mean (ms, log scale)") +
  theme_bw(base_size = 10)

Almost every point sits on or just above the 1:1 line. The visible exceptions are the few dask and duckdb cells in the 1e3 × 10000 shape, where a single slow iteration pulls the mean up by up to 50 %; those are also the cells where the two statistics disagree the most on the speed-up ratio. For the dataprep column itself the two statistics never differ by more than a few percent, which is why the mean-based numbers quoted throughout this vignette are representative of the steady state.

Test environment

Benchmarks were run on two reference hosts. Only the core configuration is listed here; full hardware details are in README.md.

Software versions on both hosts: data.table 1.18.6.1, reshape2 1.4.5, tidyr 1.3.2, reticulate 1.47.0; Python 3.13.7 (Ubuntu) / 3.13.15 (Windows), pandas 3.0.6, polars 1.44.2 (runtime rt64), dask 2026.8.0, duckdb 1.5.5.

How to read these numbers

The two hosts differ in core count, cache size and memory bandwidth. Two properties shape the numbers that follow:

Both effects favour dataprep in the numbers below. The relative ranking is robust; the absolute multipliers — especially the 1629× and 800× figures — should be interpreted as “best-case on a very large machine”. On a typical 8–16-core workstation the same comparisons are within 10–100×.

Cleaning pipeline (dataprep 0.1.5 → 0.1.8)

The 0.1.8 release rewrites every heavy cleaning routine in C++. The table below compares against 0.1.5 on three dataset sizes from the same source (SMEAR I Varrio forest). All numbers are speed-up ratios (0.1.5 time / 0.1.8 time) on Ubuntu 25.10.

Function 500 rows 7,640 rows 49,422 rows
varidele 1.2× 1.1× 11.6×
obsedele 203× 424× 232×
condextr 196× 217× 1146×
optisolu 188× 77× 109×
dataprep 185× 228× 247×

On Windows 11 Pro for Workstations, the same full-year pipeline gives obsedele ≈ 648×, condextr ≈ 839×, shorvalu ≈ 81×, optisolu ≈ 25× (at cores = 32), and the integrated dataprep call ≈ 173×. varidele is around 1.17× on this cell; this is expected, since varidele is a single colMeans(is.na(.)) in both versions and the new code path has little room for improvement.

Note on optisolu cores. The 0.1.5 implementation could crash when cores > 16, because its parallel::makeCluster() path gave each worker a full copy of the data. The benchmark above used cores = 16 for both versions to keep the comparison fair. 0.1.8 loads the package on each worker, exports the input data once per worker, and runs each (interval, times) case as a separate task, so cores = 64 is safe. The practical speed-up on a many-core host is larger. Also note that the optimal parameter values returned by optisolu() may differ slightly between 0.1.5 and 0.1.8.

melt() — wide to long

Input shapes are described as rows × (n_id + n_val). All numbers in the cells are means in milliseconds; the value in parentheses is dataprep’s speed-up relative to that competitor.

Vary rows, 1 id + 9 value columns

rows dataprep reshape2 data.table tidyr pandas polars dask duckdb
1e3 0.173 0.378 (2.2×) 0.257 (1.5×) 2.788 (16.1×) 2.128 (12.3×) 0.645 (3.7×) 15.786 (91.3×) 4.489 (26.0×)
1e4 0.241 0.462 (1.9×) 0.334 (1.4×) 3.261 (13.5×) 2.499 (10.4×) 0.829 (3.4×) 15.818 (65.6×) 10.850 (45.0×)
1e5 0.679 1.204 (1.8×) 1.034 (1.5×) 8.167 (12.0×) 6.659 (9.8×) 2.029 (3.0×) 17.638 (26.0×) 68.473 (100.8×)
1e6 3.474 18.797 (5.4×) 9.646 (2.8×) 80.014 (23.0×) 61.686 (17.8×) 14.671 (4.2×) 47.584 (13.7×) 648.895 (186.8×)
1e7 33.412 364.564 (10.9×) 365.065 (10.9×) 1083.987 (32.4×) 710.806 (21.3×) 139.959 (4.2×) 482.693 (14.4×) 6389.564 (191.2×)
1e8 276.295 3579.061 (13.0×) 3571.833 (12.9×) 12126.897 (43.9×) 7463.089 (27.0×) 3121.344 (11.3×) 4756.547 (17.2×) 71947.305 (260.4×)

The sub-1.0× cells are polars at 1e7 × 10 id on Ubuntu (0.6×) and polars at 1e5 × 10 id on Windows (0.7×).

Vary rows, 10 id (5 int + 5 chr)

rows dataprep reshape2 data.table tidyr pandas polars dask duckdb
1e3 0.242 0.574 (2.4×) 0.408 (1.7×) 3.013 (12.5×) 5.565 (23.0×) 0.908 (3.8×) 51.309 (212.2×) 10.359 (42.8×)
1e4 0.472 2.017 (4.3×) 1.766 (3.7×) 4.811 (10.2×) 6.123 (13.0×) 1.953 (4.1×) 51.327 (108.7×) 49.807 (105.4×)
1e5 3.680 16.603 (4.5×) 14.949 (4.1×) 22.929 (6.2×) 13.307 (3.6×) 4.299 (1.2×) 56.135 (15.3×) 472.664 (128.4×)
1e6 20.354 208.901 (10.3×) 157.119 (7.7×) 221.574 (10.9×) 90.483 (4.4×) 39.946 (2.0×) 111.196 (5.5×) 4701.936 (231.0×)
1e7 827.346 3125.168 (3.8×) 2651.650 (3.2×) 3724.859 (4.5×) 1340.982 (1.6×) 518.916 (0.6×) 963.953 (1.2×) 47964.021 (58.0×)

Vary value columns, 1e3 rows, 1 id

n_val dataprep reshape2 data.table tidyr pandas polars dask duckdb
10 0.174 0.381 (2.2×) 0.256 (1.5×) 2.803 (16.1×) 2.123 (12.2×) 0.587 (3.4×) 15.358 (88.2×) 4.829 (27.7×)
100 0.261 1.110 (4.3×) 0.371 (1.4×) 3.699 (14.2×) 6.431 (24.7×) 0.750 (2.9×) 44.906 (172.1×) 21.673 (83.1×)
1000 0.916 8.127 (8.9×) 1.291 (1.4×) 12.233 (13.4×) 48.227 (52.6×) 2.704 (3.0×) 313.498 (342.2×) 178.161 (194.5×)
10000 2.295 93.092 (40.6×) 12.415 (5.4×) 104.924 (45.7×) 495.659 (216.0×) 19.408 (8.5×) 3737.866 (1628.9×) 1919.247 (836.4×)

The 1e3 × 10000 cell is the widest gap in the entire benchmark suite: dataprep returns in 2.30 ms, dask in 3.74 s, and duckdb in 1.92 s.

Vary value columns, 1e3 rows, 10 id

n_val dataprep reshape2 data.table tidyr pandas polars dask duckdb
10 0.247 0.576 (2.3×) 0.426 (1.7×) 3.022 (12.2×) 5.662 (22.9×) 1.010 (4.1×) 48.905 (198.0×) 13.088 (53.0×)
100 0.533 2.972 (5.6×) 1.985 (3.7×) 5.397 (10.1×) 20.806 (39.0×) 2.094 (3.9×) 188.372 (353.4×) 60.608 (113.7×)
1000 4.177 25.214 (6.0×) 16.681 (4.0×) 28.636 (6.9×) 166.521 (39.9×) 6.611 (1.6×) 1784.542 (427.2×) 570.076 (136.5×)
10000 25.009 308.602 (12.3×) 182.019 (7.3×) 282.574 (11.3×) 1783.493 (71.3×) 71.605 (2.9×) 23777.051 (950.7×) 5815.567 (232.5×)

melt() on Windows 11 Pro for Workstations

The same four slices as the Ubuntu host, with no AVX-512.

Vary rows, 1 id + 9 value columns

rows dataprep reshape2 data.table tidyr pandas polars dask duckdb
1e3 0.286 0.647 (2.3×) 0.468 (1.6×) 4.048 (14.2×) 3.365 (11.8×) 0.528 (1.8×) 28.149 (98.4×) 8.201 (28.7×)
1e4 0.529 0.963 (1.8×) 0.738 (1.4×) 5.106 (9.7×) 5.550 (10.5×) 0.849 (1.6×) 28.967 (54.8×) 23.331 (44.1×)
1e5 2.690 3.366 (1.3×) 3.368 (1.3×) 16.860 (6.3×) 22.516 (8.4×) 3.344 (1.2×) 44.894 (16.7×) 180.028 (66.9×)
1e6 10.515 26.197 (2.5×) 24.209 (2.3×) 160.785 (15.3×) 214.683 (20.4×) 21.721 (2.1×) 181.218 (17.2×) 1537.791 (146.2×)
1e7 77.008 245.861 (3.2×) 247.910 (3.2×) 1561.960 (20.3×) 1925.931 (25.0×) 275.120 (3.6×) 1499.583 (19.5×) 14517.337 (188.5×)
1e8 935.148 2636.186 (2.8×) 2534.491 (2.7×) 16599.644 (17.8×) 19578.283 (20.9×) 4263.390 (4.6×) 14714.823 (15.7×) 148009.529 (158.3×)

Vary rows, 10 id (5 int + 5 chr)

rows dataprep reshape2 data.table tidyr pandas polars dask duckdb
1e3 0.570 1.206 (2.1×) 0.947 (1.7×) 4.695 (8.2×) 11.374 (20.0×) 1.266 (2.2×) 98.325 (172.6×) 21.374 (37.5×)
1e4 1.706 4.985 (2.9×) 3.630 (2.1×) 8.509 (5.0×) 14.912 (8.7×) 2.074 (1.2×) 103.953 (60.9×) 114.451 (67.1×)
1e5 13.969 38.706 (2.8×) 27.974 (2.0×) 45.778 (3.3×) 44.324 (3.2×) 9.372 (0.7×) 134.847 (9.7×) 1007.726 (72.1×)
1e6 92.041 392.050 (4.3×) 272.765 (3.0×) 444.403 (4.8×) 323.335 (3.5×) 92.341 (1.0×) 371.440 (4.0×) 9545.364 (103.7×)
1e7 857.942 3974.411 (4.6×) 2867.507 (3.3×) 4703.791 (5.5×) 3213.337 (3.7×) 965.515 (1.1×) 2710.996 (3.2×) 96232.400 (112.2×)

Vary value columns, 1e3 rows, 1 id

n_val dataprep reshape2 data.table tidyr pandas polars dask duckdb
10 0.440 0.773 (1.8×) 0.608 (1.4×) 4.256 (9.7×) 3.618 (8.2×) 0.642 (1.5×) 28.358 (64.5×) 31.298 (71.2×)
100 0.939 2.505 (2.7×) 1.241 (1.3×) 6.455 (6.9×) 16.555 (17.6×) 228.139 (242.9×) 112.824 (120.1×) 47.901 (51.0×)
1000 3.843 16.235 (4.2×) 3.787 (1.0×) 23.592 (6.1×) 142.100 (37.0×) 4.550 (1.2×) 964.746 (251.0×) 452.106 (117.6×)
10000 11.569 161.647 (14.0×) 34.383 (3.0×) 203.580 (17.6×) 1410.476 (121.9×) 36.412 (3.1×) 10333.080 (893.2×) 5310.694 (459.0×)

Vary value columns, 1e3 rows, 10 id

n_val dataprep reshape2 data.table tidyr pandas polars dask duckdb
10 0.551 1.255 (2.3×) 0.887 (1.6×) 4.577 (8.3×) 12.074 (21.9×) 1.108 (2.0×) 102.415 (185.7×) 24.440 (44.3×)
100 1.701 6.266 (3.7×) 3.721 (2.2×) 9.453 (5.6×) 58.290 (34.3×) 2.607 (1.5×) 501.132 (294.7×) 141.639 (83.3×)
1000 16.754 54.631 (3.3×) 31.430 (1.9×) 55.585 (3.3×) 567.945 (33.9×) 12.350 (0.7×) 4799.771 (286.5×) 1373.995 (82.0×)
10000 101.066 554.243 (5.5×) 346.842 (3.4×) 512.059 (5.1×) 5672.202 (56.1×) 115.524 (1.1×) 54507.720 (539.3×) 12874.917 (127.4×)

The largest Windows multiplier for melt() is 893.2× (dask at 1e3 rows, 1 id + 10000 value columns). On the Ubuntu host the corresponding cell reaches 1628.9×.

dcast() — long to wide

Input is a canonical long table with every (id, variable) pair present exactly once. All numbers in the cells are means in milliseconds; the value in parentheses is dataprep’s speed-up relative to that competitor.

Vary n_long, 1 id, 10 levels

n_long dataprep reshape2 data.table tidyr pandas polars dask duckdb
1e3 0.888 1.664 (1.9×) 1.810 (2.0×) 4.134 (4.7×) 1.905 (2.1×) 33.322 (37.5×) 8.107 (9.1×) 7.270 (8.2×)
1e4 0.940 2.611 (2.8×) 2.559 (2.7×) 4.474 (4.8×) 2.329 (2.5×) 48.157 (51.3×) 8.868 (9.4×) 10.398 (11.1×)
1e5 1.090 19.351 (17.8×) 14.634 (13.4×) 7.749 (7.1×) 7.078 (6.5×) 49.857 (45.7×) 14.741 (13.5×) 34.205 (31.4×)
1e6 1.654 151.826 (91.8×) 328.980 (198.9×) 44.728 (27.0×) 56.741 (34.3×) 105.018 (63.5×) 78.217 (47.3×) 155.880 (94.3×)
1e7 9.202 1693.368 (184.0×) 560.932 (61.0×) 716.377 (77.9×) 741.921 (80.6×) 310.878 (33.8×) 957.022 (104.0×) 1706.000 (185.4×)
1e8 91.926 21766.497 (236.8×) 16869.402 (183.5×) 10507.819 (114.3×) 10397.700 (113.1×) 2434.035 (26.5×) 13416.534 (145.9×) 17118.026 (186.2×)

Vary levels, 1 id, 1e6 rows

levels dataprep reshape2 data.table tidyr pandas polars dask duckdb
10 1.654 151.826 (91.8×) 328.980 (198.9×) 44.728 (27.0×) 56.741 (34.3×) 105.018 (63.5×) 78.217 (47.3×) 155.880 (94.3×)
100 1.415 101.438 (71.7×) 329.202 (232.6×) 42.047 (29.7×) 53.686 (37.9×) 173.478 (122.6×) 74.540 (52.7×) 178.421 (126.0×)
1000 2.107 100.949 (47.9×) 251.767 (119.5×) 45.413 (21.6×) 56.445 (26.8×) 186.575 (88.6×) 76.185 (36.2×) 195.787 (92.9×)
10000 12.285 140.127 (11.4×) 405.321 (33.0×) 57.827 (4.7×) 59.199 (4.8×) 308.858 (25.1×) 80.920 (6.6×) 493.181 (40.1×)

Vary n_long, 1 id, 100 levels

n_long dataprep reshape2 data.table tidyr pandas polars dask duckdb
1e4 1.010 2.860 (2.8×) 2.884 (2.9×) 4.640 (4.6×) 2.483 (2.5×) 45.241 (44.8×) 9.460 (9.4×) 14.845 (14.7×)
1e5 1.184 18.509 (15.6×) 8.577 (7.2×) 7.826 (6.6×) 6.861 (5.8×) 60.782 (51.3×) 14.937 (12.6×) 43.299 (36.6×)
1e6 1.415 101.438 (71.7×) 329.202 (232.6×) 42.047 (29.7×) 53.686 (37.9×) 173.478 (122.6×) 74.540 (52.7×) 178.421 (126.0×)
1e7 5.073 2016.036 (397.4×) 575.604 (113.5×) 660.455 (130.2×) 816.883 (161.0×) 506.557 (99.9×) 1002.368 (197.6×) 1669.419 (329.1×)
1e8 40.680 16963.494 (417.0×) 18866.887 (463.8×) 8100.475 (199.1×) 9529.877 (234.3×) 2460.625 (60.5×) 12764.776 (313.8×) 17616.392 (433.1×)

The 1e8 × 100 levels cell is the strongest dcast result on this host: dataprep returns in 40.68 ms, reshape2 in 16.96 s.

Vary n_id, 1e6 rows, 10 levels

n_id dataprep reshape2 data.table tidyr pandas polars dask duckdb
1 1.654 151.826 (91.8×) 328.980 (198.9×) 44.728 (27.0×) 56.741 (34.3×) 105.018 (63.5×) 78.217 (47.3×) 155.880 (94.3×)
2 1.915 218.272 (114.0×) 304.452 (159.0×) 54.255 (28.3×) 80.896 (42.2×) 122.649 (64.0×) 108.293 (56.5×) 285.377 (149.0×)
10 3.565 1247.607 (350.0×) 405.322 (113.7×) 93.897 (26.3×) 192.356 (54.0×) 119.942 (33.6×) 248.201 (69.6×) 922.863 (258.9×)
100 22.579 10394.072 (460.3×) 549.582 (24.3×) 431.569 (19.1×) 1415.001 (62.7×) 150.336 (6.7×) 1608.617 (71.2×) 8314.759 (368.3×)

The 100 id cell is the only case in the entire benchmark suite where a competitor reaches a single-digit ratio. polars is within 7.4×. It remains behind dataprep.

dcast() on Windows 11 Pro for Workstations

The same four slices as the Ubuntu host, with no AVX-512.

Vary n_long, 1 id, 10 levels

n_long dataprep reshape2 data.table tidyr pandas polars dask duckdb
1e3 0.384 2.412 (6.3×) 3.885 (10.1×) 6.465 (16.9×) 2.832 (7.4×) 5.114 (13.3×) 15.597 (40.7×) 19.382 (50.5×)
1e4 0.478 4.178 (8.7×) 7.391 (15.5×) 8.016 (16.8×) 5.078 (10.6×) 5.806 (12.1×) 18.732 (39.2×) 24.628 (51.5×)
1e5 1.029 31.267 (30.4×) 39.986 (38.9×) 13.887 (13.5×) 20.444 (19.9×) 11.069 (10.8×) 42.061 (40.9×) 73.489 (71.5×)
1e6 3.426 281.533 (82.2×) 148.638 (43.4×) 93.290 (27.2×) 251.508 (73.4×) 44.604 (13.0×) 346.422 (101.1×) 391.237 (114.2×)
1e7 23.194 2820.301 (121.6×) 989.520 (42.7×) 1481.134 (63.9×) 2698.928 (116.4×) 459.512 (19.8×) 3299.792 (142.3×) 3619.348 (156.0×)
1e8 214.305 29441.634 (137.4×) 13164.922 (61.4×) 16383.457 (76.4×) 31753.144 (148.2×) 4722.561 (22.0×) 38192.158 (178.2×) 33441.939 (156.0×)

Vary levels, 1 id, 1e6 rows

levels dataprep reshape2 data.table tidyr pandas polars dask duckdb
10 3.426 281.533 (82.2×) 148.638 (43.4×) 93.290 (27.2×) 251.508 (73.4×) 44.604 (13.0×) 346.422 (101.1×) 391.237 (114.2×)
100 4.049 173.399 (42.8×) 169.140 (41.8×) 91.050 (22.5×) 229.426 (56.7×) 56.026 (13.8×) 332.823 (82.2×) 800.504 (197.7×)
1000 5.064 189.364 (37.4×) 153.576 (30.3×) 81.644 (16.1×) 253.038 (50.0×) 178.688 (35.3×) 316.612 (62.5×) 1136.216 (224.4×)
10000 29.411 252.442 (8.6×) 191.452 (6.5×) 115.340 (3.9×) 275.314 (9.4×) 1350.575 (45.9×) 367.614 (12.5×) 13596.212 (462.3×)

Vary n_long, 1 id, 100 levels

n_long dataprep reshape2 data.table tidyr pandas polars dask duckdb
1e4 0.514 4.960 (9.7×) 8.608 (16.8×) 8.088 (15.7×) 5.519 (10.7×) 9.883 (19.2×) 18.710 (36.4×) 94.938 (184.8×)
1e5 0.838 34.405 (41.1×) 41.520 (49.6×) 17.537 (20.9×) 37.751 (45.1×) 13.291 (15.9×) 63.692 (76.0×) 303.785 (362.7×)
1e6 4.049 173.399 (42.8×) 169.140 (41.8×) 91.050 (22.5×) 229.426 (56.7×) 56.026 (13.8×) 332.823 (82.2×) 800.504 (197.7×)
1e7 38.527 3197.642 (83.0×) 1100.505 (28.6×) 1359.999 (35.3×) 2337.317 (60.7×) 814.529 (21.1×) 3014.871 (78.3×) 7495.310 (194.5×)
1e8 107.287 23690.584 (220.8×) 14715.772 (137.2×) 14245.599 (132.8×) 25843.625 (240.9×) 6222.505 (58.0×) 33050.665 (308.1×) 85804.834 (799.8×)

Vary n_id, 1e6 rows, 10 levels

n_id dataprep reshape2 data.table tidyr pandas polars dask duckdb
1 3.426 281.533 (82.2×) 148.638 (43.4×) 93.290 (27.2×) 251.508 (73.4×) 44.604 (13.0×) 346.422 (101.1×) 391.237 (114.2×)
2 10.712 401.842 (37.5×) 195.915 (18.3×) 116.768 (10.9×) 367.464 (34.3×) 57.728 (5.4×) 487.128 (45.5×) 697.317 (65.1×)
10 16.392 2687.391 (163.9×) 326.747 (19.9×) 187.891 (11.5×) 875.166 (53.4×) 66.015 (4.0×) 1109.804 (67.7×) 2012.130 (122.8×)
100 76.279 23604.925 (309.5×) 1002.287 (13.1×) 740.787 (9.7×) 6652.292 (87.2×) 198.455 (2.6×) 8254.600 (108.2×) 15421.548 (202.2×)

The largest Windows multiplier for dcast() is 799.8× (duckdb at 1e8 rows, 1 id, 100 levels). On the Ubuntu host the corresponding cell reaches 433.1×. The largest Ubuntu multiplier for dcast() overall is 463.8× (data.table at 1e8 rows, 1 id, 100 levels).

Summary of speedups

Speedup is defined as competitor mean / dataprep mean. Each table summarises every benchmark cell on that host, across all seven competitors (reshape2, data.table, tidyr, pandas, polars, dask, duckdb).

Ubuntu 25.10

Operation Min Median Mean Max
melt() 0.6× (polars @ 1e7 × 19 × 10 × 9) 11.3× 67.8× 1628.9× (dask @ 1e3 × 10001 × 1 × 10000)
dcast() 1.9× (reshape2 @ 1e3 × 1 × 10) 46.5× 90.2× 463.8× (data.table @ 1e8 × 1 × 100)

Windows 11 Pro for Workstations

Operation Min Median Mean Max
melt() 0.7× (polars @ 1e5 × 19 × 10 × 9) 5.6× 46.6× 893.2× (dask @ 1e3 × 10001 × 1 × 10000)
dcast() 2.6× (polars @ 1e6 × 100 × 10) 41.4× 76.6× 799.8× (duckdb @ 1e8 × 1 × 100)

Combined across both hosts:

Cross-engine consistency

Both melt() and dcast() produce output numerically identical to reshape2 on every tested cell. All pairs of engines agree pairwise within tol = 1e-12.

melt consistency

rows n_id n_val engines passed pairwise
1,000 1 9 8/8 all consistent
100,000 1 9 8/8 all consistent
1,000 1 100 8/8 all consistent
10,000 10 10 8/8 all consistent

dcast consistency

n_long n_id n_levels engines passed pairwise
5,000 2 5 8/8 all consistent
50,000 1 50 8/8 all consistent
50,000 10 10 8/8 all consistent
1,000,000 1 10 8/8 all consistent

Engines compared: dataprep, reshape2, data.table, tidyr, pandas, polars, dask, duckdb.

Reproducing the benchmarks

The full runner is shipped under inst/:

Scripts are disabled by default so that R CMD check does not run them. To enable:

Sys.setenv(DATAPREP_RUN_BENCHMARK = "1")
source(system.file("benchmark_melt_dcast.R", package = "dataprep"))

Session info

sessionInfo()
#> R version 4.5.1 (2025-06-13)
#> Platform: x86_64-pc-linux-gnu
#> Running under: Ubuntu 25.10
#> 
#> Matrix products: default
#> BLAS:   /usr/lib/x86_64-linux-gnu/openblas-openmp/libblas.so.3 
#> LAPACK: /usr/lib/x86_64-linux-gnu/openblas-openmp/libopenblasp-r0.3.30.so;  LAPACK version 3.12.0
#> 
#> locale:
#>  [1] LC_CTYPE=zh_CN.UTF-8       LC_NUMERIC=C              
#>  [3] LC_TIME=zh_CN.UTF-8        LC_COLLATE=zh_CN.UTF-8    
#>  [5] LC_MONETARY=zh_CN.UTF-8    LC_MESSAGES=zh_CN.UTF-8   
#>  [7] LC_PAPER=zh_CN.UTF-8       LC_NAME=C                 
#>  [9] LC_ADDRESS=C               LC_TELEPHONE=C            
#> [11] LC_MEASUREMENT=zh_CN.UTF-8 LC_IDENTIFICATION=C       
#> 
#> time zone: Asia/Shanghai
#> tzcode source: system (glibc)
#> 
#> attached base packages:
#> [1] stats     graphics  grDevices utils     datasets  methods   base     
#> 
#> other attached packages:
#> [1] ggplot2_4.0.3  dataprep_0.1.8
#> 
#> loaded via a namespace (and not attached):
#>  [1] vctrs_0.7.3        cli_3.6.6          knitr_1.52         rlang_1.3.0       
#>  [5] xfun_0.61          otel_0.2.0         generics_0.1.4     S7_0.2.2          
#>  [9] jsonlite_2.0.0     glue_1.8.1         htmltools_0.5.9    sass_0.4.10       
#> [13] scales_1.4.0       rmarkdown_2.32     grid_4.5.1         tibble_3.3.1      
#> [17] evaluate_1.0.5     jquerylib_0.1.4    fastmap_1.2.0      yaml_2.3.12       
#> [21] lifecycle_1.0.5    compiler_4.5.1     dplyr_1.2.1        RColorBrewer_1.1-3
#> [25] pkgconfig_2.0.3    Rcpp_1.1.2         farver_2.1.2       digest_0.6.39     
#> [29] R6_2.6.1           tidyselect_1.2.1   pillar_1.11.1      parallel_4.5.1    
#> [33] magrittr_2.0.5     bslib_0.12.0       withr_3.0.3        tools_4.5.1       
#> [37] gtable_0.3.6       cachem_1.1.0