Package: dataprep
Type: Package
Title: Fast, Efficient, and Versatile Data Preprocessing and Reshaping
        with 'C++', 'OpenMP' & 'SIMD'
Version: 0.1.8
Authors@R: c(
    person("Chun-Sheng", "Liang", role = c("aut", "cre"),
           email = "chun-shengliang@qq.com",
           comment = c(ORCID = "0000-0002-7386-1958")),
    person("Hao", "Wu", role = "aut"),
    person("Hai-Yan", "Li", role = "aut"),
    person("Qiang", "Zhang", role = "aut"),
    person("Zhanqing", "Li", role = "aut"),
    person("Ke-Bin", "He", role = "aut")
  )
Description: 
    Fast, efficient, and versatile preprocessing and reshaping of tabular
    and time-series data. Most heavy routines are implemented in 'C++' via
    'Rcpp', with optional 'OpenMP' parallelization and 'SIMD' acceleration
    ('AVX2' / 'AVX-512') on supported hardware. The 0.1.8 release rewrites
    the cleaning routines in 'C++' and delivers a 1.1–1146× speedup over
    0.1.5. The 'melt()' and 'dcast()' reshaping functions achieve a
    0.6×–1628.9× speedup for 'melt()' and a 1.9×–799.8× speedup for
    'dcast()' relative to every one of the seven major alternatives in
    the R and Python ecosystems, at every tested scale (from 1,000 to
    100,000,000 rows), and produce output identical to 'reshape2',
    'data.table', 'tidyr', 'pandas', 'polars', 'dask', and 'duckdb'.
    Core preprocessing steps include variable deletion by missing
    fraction, observation deletion by consecutive missing runs,
    point-by-point weighted outlier removal via conditional extremum,
    traditional percentile-based outlier removal, and linear
    interpolation within short time periods. The package also provides
    fast reshaping, descriptive statistics, missing-value diagnosis,
    multiple imputation strategies, winsorization, several outlier
    detection methods (IQR, MAD, percentile), data transformation and
    standardization, categorical encoding, duplicate removal, data
    validation, data quality reporting, and stratified sampling.
    Feature-engineering helpers cover binning, high-correlation and
    low-variance filtering, and string cleaning. Time-series tools
    cover detrending, diurnal-cycle removal, rolling statistics, lag
    creation, resampling, simple decomposition, day/night and season
    flags, log returns, drift detection, and panel balancing.
    Fit/transform-style machine-learning interfaces prevent data
    leakage during preprocessing. Methods are based on, and improved
    from: Liang, C.-S., Wu, H., Li, H.-Y., Zhang, Q., Li, Z. & He,
    K.-B. (2020) <doi:10.1016/j.scitotenv.2020.140923>. This work
    was supported by the National Natural Science Foundation of China
    (No. 12301674).
Depends: R (>= 4.0.0)
ByteCompile: TRUE
Imports: Rcpp (>= 1.0.10), ggplot2, stats, parallel
LinkingTo: Rcpp (>= 1.0.10)
Suggests: knitr, rmarkdown, testthat (>= 3.0.0), data.table, reshape2,
        tidyr, dplyr, microbenchmark, reticulate
VignetteBuilder: knitr, rmarkdown
SystemRequirements: C++17, optional OpenMP
License: GPL (>= 2)
Encoding: UTF-8
NeedsCompilation: yes
URL: https://github.com/chunshengliang/dataprep,
        https://chunshengliang.github.io/dataprep/
BugReports: https://github.com/chunshengliang/dataprep/issues
Config/testthat/edition: 3
LazyData: true
Config/roxygen2/version: 8.1.0
Packaged: 2026-10-01 09:36:47 UTC; zhike
Author: Chun-Sheng Liang [aut, cre] (ORCID:
    <https://orcid.org/0000-0002-7386-1958>),
  Hao Wu [aut],
  Hai-Yan Li [aut],
  Qiang Zhang [aut],
  Zhanqing Li [aut],
  Ke-Bin He [aut]
Maintainer: Chun-Sheng Liang <chun-shengliang@qq.com>
Repository: CRAN
Date/Publication: 2026-10-01 15:41:00 UTC
