---
title: "Starting on S3"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Starting on S3}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment  = "#>",
  eval     = FALSE
)
```

```{=html}
<link rel="stylesheet" href="../inst/vignette-setup/datom.css">
```

<div class="datom-callout datom-callout-goal" markdown="1">
**Goal:** Stand up a versioned datom project whose data lives in **Amazon S3**,
and onboard a study's files with `datom_sync()`. The steps are the ones from
[Getting Started](getting-started.html) -- one file, an update, a no-op, a
batch -- and the only difference is the store you build.
</div>

> **Want to try datom locally first?** Start with
> [Getting Started](getting-started.html). It uses a folder on your machine
> instead of S3 and needs no AWS account. The functions are the same.

You look after the data for **study001**, a clinical trial, and your team
already works in S3. You want the first extract to land in the shared bucket,
versioned from day one. datom keeps the data in S3 and the record of every
version in a git repository, so the history can be read and reproduced from any
machine.

<div class="datom-callout datom-callout-tip" markdown="1">
**Two locations, two roles:**

- **Input folder** -- where new files land before datom takes them in. Files
  here can be overwritten or deleted. Syncing a file datom already holds does
  nothing.
- **Storage** -- once `datom_sync()` takes a file in, S3 holds a versioned copy.
  The input file is no longer needed.
</div>

## Requirements

- **A GitHub account** with a personal access token (PAT) scoped to `repo`.
  datom creates the metadata repository through the GitHub API with it; the
  `gh` CLI is not needed.
- **An S3 bucket** you can read and write. datom does not create buckets:
  encryption, versioning and retention are your organization's policy. This
  article uses one bucket per study, with a folder per project inside it.
- **AWS credentials** (an access key and a secret key) for that bucket.
- **The `git2r`, `rio` and `keyring` packages.** datom uses `git2r` for the
  metadata repository and `rio` to read the files you sync; `keyring` holds
  your secrets in the OS keychain.

Store the three secrets in your keychain once:

```{r keyring-setup}
keyring::key_set(service = "GITHUB_PAT")
keyring::key_set(service = "AWS_ACCESS_KEY_ID")
keyring::key_set(service = "AWS_SECRET_ACCESS_KEY")
```

## Load your secrets

This is the only place the keychain is read. Everything after it reads the
environment with `Sys.getenv()`, so the secrets never appear in your code.

```{r secrets}
Sys.setenv(
  GITHUB_PAT            = keyring::key_get(service = "GITHUB_PAT"),
  AWS_ACCESS_KEY_ID     = keyring::key_get(service = "AWS_ACCESS_KEY_ID"),
  AWS_SECRET_ACCESS_KEY = keyring::key_get(service = "AWS_SECRET_ACCESS_KEY")
)
```

On CI or in a container, set these three environment variables through your
platform's secret store and skip this chunk.

## Settings

Every value used more than once is set here, so changing one means changing it
in one place.

```{r settings}
library(datom)

# --- Settings you control ----------------------------------------------------
bucket           <- "study001"            # one bucket per study
region           <- "us-east-1"

project_imported <- "study001-imported"   # recorded in the project's metadata
prefix_imported  <- "imported/"           # this project's folder in the bucket
repo_imported    <- "study001-imported"   # GitHub repo name

# Local working folder for the metadata repository. The data never lands here;
# it goes straight to S3.
workdir_imported <- fs::path(tempdir(), "study001-imported")
```

The project, repo and folder names match by convention only. They are separate
settings because they do not have to match.

## Where the data goes

```
s3://study001/
    imported/datom/        onboarded tables        (this article)
```

`imported` is datom's word for a table that came in from a file. datom manages
the `datom/` folder under each prefix: it reads and writes the data, metadata
and version records there, and leaves anything else in the bucket alone.

## Build the store

A **store** says where the data lives and how to reach it. Giving it a GitHub
token makes it a **writer** store: it can create the metadata repository and
record new versions.

```{r store-write}
store_write_imported <- datom_store(
  data = datom_store_s3(
    bucket     = bucket,
    prefix     = prefix_imported,
    region     = region,
    access_key = Sys.getenv("AWS_ACCESS_KEY_ID"),
    secret_key = Sys.getenv("AWS_SECRET_ACCESS_KEY")
  ),
  github_pat = Sys.getenv("GITHUB_PAT")
)
```

datom checks that it can reach the bucket as the store is built, so a wrong
credential shows up here rather than at the first sync.

## Create the repository

```{r init}
datom_init_repo(
  path         = workdir_imported,
  project_name = project_imported,
  store        = store_write_imported,
  create_repo  = TRUE,
  repo_name    = repo_imported
)
#> v Created GitHub repo ".../study001-imported".
#> v Initialized datom repository "study001-imported" at '.../study001-imported'
```

This creates the GitHub repository, clones it into `workdir_imported`, and
commits a `project.yaml` recording where the project's data lives. No data goes
to GitHub, only metadata.

There is no `mode` argument here. Leaving it out gives an ordinary repository,
one that takes files in with `datom_sync()`.

The clone also gets an `input_files/` folder. git ignores it, so nothing placed
there is ever committed. It is where `datom_sync()` looks for new files.

## Connect

```{r conn-write}
conn_write_imported <- datom_get_conn(
  path  = workdir_imported,
  store = store_write_imported
)
print(conn_write_imported)
#> 
#> -- datom connection
#> * Project: "study001-imported"
#> * Backend: "s3"
#> * Role: "developer"
#> * Data root: "study001"
#> * Data prefix: "imported/"
#> * Data region: "us-east-1"
#> * Governance: not attached
#> * Path: '.../study001-imported'
#> * Data repo: <https://github.com/.../study001-imported.git>
```

## Step 1: Sync one file

The month-1 extract of the demographics table has arrived. Put it in the input
folder:

```{r step1-write}
inputs_imported <- fs::path(workdir_imported, "input_files")

write.csv(
  x         = datom_example_data(domain = "dm", cutoff_date = "2026-01-28"),
  file      = fs::path(inputs_imported, "dm.csv"),
  row.names = FALSE
)
```

Scan the folder, then sync:

```{r step1-sync}
manifest <- datom_sync_manifest(conn = conn_write_imported)
#> i Scanned 1 file: 1 new, 0 changed, 0 unchanged.

synced <- datom_sync(conn = conn_write_imported, manifest = manifest)
#> i Syncing 1 table...
#> v Wrote "dm" (full): "153bff41"
#> v "dm" synced (new).
#> i Sync complete: 1 succeeded, 0 failed, 0 skipped.
```

The file was converted to parquet and uploaded to S3. The version record was
committed to the metadata repository and pushed to GitHub. `datom_sync()`
returns the manifest with a `result` for each file, kept here as `synced`.

```{r step1-list}
datom_list(conn = conn_write_imported)
#>   name  kind current_version current_data_sha         last_updated
#> 1   dm table        153bff41         decbafd2 2026-09-27T05:58:03Z

dm_history <- datom_history(conn = conn_write_imported, name = "dm",
                            short_hash = TRUE)
dm_history[, c("version", "timestamp", "commit_message")]
#>    version            timestamp commit_message
#> 1 153bff41 2026-09-27T05:58:03Z  Sync dm (new)
```

The input file is no longer needed. The table reads from S3:

```{r step1-delete}
fs::file_delete(fs::path(inputs_imported, "dm.csv"))
nrow(datom_read(conn = conn_write_imported, name = "dm"))
#> [1] 4
```

## Step 2: Update one file

The month-2 extract arrives with new subjects:

```{r step2}
write.csv(
  x         = datom_example_data(domain = "dm", cutoff_date = "2026-02-28"),
  file      = fs::path(inputs_imported, "dm.csv"),
  row.names = FALSE
)

manifest <- datom_sync_manifest(conn = conn_write_imported)
#> i Scanned 1 file: 0 new, 1 changed, 0 unchanged.

synced <- datom_sync(conn = conn_write_imported, manifest = manifest)
#> i Syncing 1 table...
#> v Wrote "dm" (full): "0fac26cd"
#> v "dm" synced (changed).
#> i Sync complete: 1 succeeded, 0 failed, 0 skipped.
```

Both versions stay readable. The newest is the default; an older one is read by
its version. The first 8 characters of a version are enough, as with a git
commit:

```{r step2-read}
dm_history <- datom_history(conn = conn_write_imported, name = "dm",
                            short_hash = TRUE)
dm_history[, c("version", "timestamp", "commit_message")]
#>    version            timestamp    commit_message
#> 1 0fac26cd 2026-09-27T05:58:10Z Sync dm (changed)
#> 2 153bff41 2026-09-27T05:58:03Z     Sync dm (new)

nrow(datom_read(conn = conn_write_imported, name = "dm"))
#> [1] 16

dm_version <- dm_history$version[nrow(dm_history)]   # oldest row: month 1
nrow(datom_read(conn = conn_write_imported, name = "dm", version = dm_version))
#> [1] 4
```

## Step 3: Sync again with nothing new

```{r step3}
manifest <- datom_sync_manifest(conn = conn_write_imported)
#> i Scanned 1 file: 0 new, 0 changed, 1 unchanged.

synced <- datom_sync(conn = conn_write_imported, manifest = manifest)
#> i No new or changed files. Nothing to sync.
```

A file datom already holds is not uploaded again, so the sync is safe to run on
a schedule.

## Step 4: A batch of files

The month-3 extract brings four tables at once: demographics (`dm`), dosing
(`ex`), labs (`lb`) and adverse events (`ae`).

```{r step4}
for (domain in c("dm", "ex", "lb", "ae")) {
  write.csv(
    x         = datom_example_data(domain = domain, cutoff_date = "2026-03-28"),
    file      = fs::path(inputs_imported, paste0(domain, ".csv")),
    row.names = FALSE
  )
}

manifest <- datom_sync_manifest(conn = conn_write_imported)
#> i Scanned 4 files: 3 new, 1 changed, 0 unchanged.

synced <- datom_sync(conn = conn_write_imported, manifest = manifest)
#> i Syncing 4 tables...
#> v Wrote "ae" (full): "075773e9"
#> v "ae" synced (new).
#> v Wrote "dm" (full): "773e6862"
#> v "dm" synced (changed).
#> v Wrote "ex" (full): "8dbcc9a7"
#> v "ex" synced (new).
#> v Wrote "lb" (full): "435bccb0"
#> v "lb" synced (new).
#> i Sync complete: 4 succeeded, 0 failed, 0 skipped.
```

All four tables are now versioned in S3:

```{r step4-list}
datom_list(conn = conn_write_imported)
#>   name  kind current_version current_data_sha         last_updated
#> 1   dm table        773e6862         e547f03d 2026-09-27T05:58:22Z
#> 2   ae table        075773e9         d5f8dd5a 2026-09-27T05:58:17Z
#> 3   ex table        8dbcc9a7         ab96afc3 2026-09-27T05:58:26Z
#> 4   lb table        435bccb0         5d419c60 2026-09-27T05:58:31Z
```

## Reading as a reader

A colleague who only reads needs bucket credentials and nothing else: no GitHub
token and no clone. A store without a token is a **reader** store.

```{r reader}
store_read_imported <- datom_store(
  data = datom_store_s3(
    bucket     = bucket,
    prefix     = prefix_imported,
    region     = region,
    access_key = Sys.getenv("AWS_ACCESS_KEY_ID"),
    secret_key = Sys.getenv("AWS_SECRET_ACCESS_KEY")
  )
)

conn_read_imported <- datom_get_conn(
  store        = store_read_imported,
  project_name = project_imported
)

nrow(datom_read(conn = conn_read_imported, name = "lb"))
#> [1] 205
```

The read goes straight to S3, not through GitHub.

## Where you are

- Four tables are versioned in `s3://study001/imported/datom/`.
- Their history is in a GitHub repository; the data itself never went there.
- The same sync steps work on S3 as on a local folder.

## What next

- **Keep going.** [Citing a Set of Tables](citable-sets.html) builds on this
  data: it collects these tables, and the ones you derive from them, into one
  versioned, citable set. Continue in the same R session and skip the teardown
  below for now.
- **Stop here.** Run the [teardown](#teardown).

## Governance and migration come later

Two things are deliberately left out of this article:

- **Governance**: a shared register of projects, managed reader access and
  managed teardown. It is optional, provided by the companion package
  `datomanager`, and starting on S3 does not commit you to it.
- **Migration**: moving an existing project's data from one store to another
  while keeping its history. That is also a `datomanager` workflow. You started
  on S3, so you do not need it now.

## Teardown

Delete the project's storage first, then its repository:

```{r teardown-imported}
datom_storage_delete_prefix(conn = conn_write_imported)
datom_repo_delete(conn = conn_write_imported, confirm = project_imported)
```

`datom_storage_delete_prefix()` deletes everything under `imported/datom/`, and
it does not ask first. The rest of the bucket is left alone, and so is the
bucket itself.

`datom_repo_delete()` deletes the GitHub repository and the local clone. It
needs the project name as `confirm`, and it does not touch storage.
