Backends
bakelite writes replicas to a pluggable storage backend. Five families are production-supported, plus an experimental sixth:
- Local filesystem — a directory (often a mounted volume). Simple and fully supported.
- S3 / S3-compatible — built on the
object_storecrate, so the same implementation serves AWS S3, Cloudflare R2, Backblaze B2, and MinIO. See Configuration → S3-compatible for per-provider config (endpoint, region, path-style) and the least-privilege IAM policy to attach. Supports immutable backups via Object Lock (WORM) — bakelite is append-only, so it runs against a locked bucket unchanged — and cheaper cold storage via storage classes (the cold bulk to an instant-retrieval tier, hot change-sets in the default). - Google Cloud Storage — the same
object_storelayer. Credentials resolve from a service-account key or Application Default Credentials. See Configuration → Google Cloud Storage. - Azure Blob Storage — likewise. Account + container in the config; the credential comes from the environment. See Configuration → Azure Blob Storage.
- SFTP (over SSH) — copy to any SSH host. The transport is pure Rust
(
russh+russh-sftp), so there's no systemsshbinary and no OpenSSL/libssh2 C dependency — the single static binary keeps working. Authenticate by SSH key (recommended) or password; the server's host key is verified againstknown_hosts. It's the local-filesystem backend over a network (no object-store version inspection or multipart, which are S3-only concepts). See Configuration → SFTP. - Peer-to-peer swarm (experimental) — instead of an external store, let a fleet of bakelite nodes back each other up over an encrypted iroh (QUIC) mesh: no bucket, no bill, and a lost machine restores from its peers. A durability controller holds the fleet to a replication factor — peers must prove they still hold each object, and shortfalls are repaired automatically. It's a redundancy layer on top of a normal sync primary — experimental, but built into the standard binary. See the Backup swarm guide.
GCS and Azure go through the identical ObjectStoreBackend mapping that the S3
backend uses, so they go through the same Backend-trait conformance tests the in-memory
and S3 suites run. The S3-specific version-overhead
inspection (bakelite usage noncurrent-version reporting) and multipart reclaim
are S3-only — on GCS/Azure, as on R2, bakelite usage reports that overhead as
"not inspected". The native GCS and Azure backends are exercised against real
accounts by the same opt-in, credential-gated conformance suite (run locally; see
below).
Why a compatibility matrix
"S3-compatible" is a spectrum. Services agree on the core object API but diverge at
the edges — versioning semantics, multipart-upload listing, lifecycle behavior.
We learned this the hard way: MinIO's ListMultipartUploads silently ignores the
prefix parameter (real AWS honours it), which would have made multipart reclaim
miss orphans on MinIO had we relied on server-side filtering. Emulators don't
always match real-service behavior (the MinIO prefix handling above is one
example), so bakelite is tested against the real services too, not just emulators.
How it's tested
One parameterized conformance suite (s3_conformance in
crates/bakelite-core/tests/backend_conformance.rs) runs against every target by
reading BAKELITE_TEST_S3_* env. It exercises four layers:
- the full
Backendtrait contract (incl. the streaming snapshot/multipart path); - S3 inspection — versioning + version-overhead reporting;
- reclaim end-to-end — create a dangling multipart upload, list it, age-gate it, abort it, and check it's removed (the regression guard for the prefix divergence);
- a capability probe that records raw provider behavior (it never asserts) and
emits
BAKELITE_CAPABILITY_JSON {…}, the source for the table below.
- Emulators (MinIO + LocalStack) run locally via
just minio-up && just test-s3/just localstack-up && just test-localstack. - Real providers run opt-in (AWS, R2, B2) via
just test-provider <label>with the provider'sBAKELITE_TEST_S3_*credentials in the environment. - Native GCS + Azure run opt-in the same way, through the
gcs_conformance/azure_conformanceentrypoints. These exercise layer 1 (the trait contract) against a live account; the S3-only inspection/reclaim/capability layers don't apply. - Locally:
just minio-up && just test-s3,just localstack-up && just test-localstack, orjust test-provider <label>against a real bucket (export itsAWS_*+BAKELITE_TEST_S3_{ENDPOINT,BUCKET,REGION,PATH_STYLE}first).
Compatibility matrix
| Provider | Trait conformance | Versioning inspection | Multipart | ListMultipartUploads prefix honoured | Noncurrent versions | Delete markers |
|---|---|---|---|---|---|---|
| Local filesystem | ✅ | — | — | — | — | — |
| SFTP | ✅⁴ | — | — | — | — | — |
| AWS S3 | ✅ | ✅ enabled | ✅ | ✅ | ✅ | ✅ |
| Cloudflare R2 | ✅ | ❌ (403)¹ | ✅ | ✅ | —¹ | —¹ |
| Backblaze B2 | ✅ | ✅ enabled | ✅ | ✅ | ✅ | ✅ |
| MinIO | ✅ | ✅ enabled | ✅ | ❌² | ✅ | ✅ |
| LocalStack (3.x) | ✅ | ✅ (off by default) | ✅ | ✅ | — | — |
| Google Cloud Storage | ✅³ | —³ | ✅ | — | —³ | —³ |
| Azure Blob | ✅³ | —³ | ✅ | — | —³ | —³ |
Measured by the suite's capability probe, run locally against the emulators and each real provider. The experimental peer-to-peer swarm isn't an object store, so these S3-shaped columns don't apply to it; its correctness is exercised by the durability-controller torture suite instead.
Notes:
- Cloudflare R2 answers
GetBucketVersioning/ListObjectVersionswith 403 — it doesn't expose the object-versioning APIs. bakelite tolerates this:bakelite usagereports version overhead as "not inspected" on R2, and replication, restore, and multipart reclaim all work normally (R2 does support multipart, and honours the prefix). It only means R2's hidden-version overhead can't be reported. - MinIO silently ignores the
prefixonListMultipartUploadsand returns every upload. bakelite never relies on server-side prefix filtering (it lists whole-bucket and filters client-side), so reclaim is correct here regardless — this column documents the divergence that drove that design. Every other tested service honours the prefix. - Google Cloud Storage / Azure Blob use the same
object_store-backedObjectStoreBackendas S3. Their trait conformance (✅³) is exercised against a live account by the opt-ingcs_conformance/azure_conformancesuite entrypoints, on top of the shared in-memory and S3 suites. The S3-only version-overhead inspection and multipart reclaim don't apply (—³); large snapshots still stream as multipart viaobject_store(GCS resumable uploads / Azure block blobs). Expire old data with the provider's own lifecycle/object-versioning controls. See the per-provider setup notes under Configuration → Google Cloud Storage and Azure Blob Storage. - SFTP runs the same
Backend-trait conformance suite the other backends do, against a disposableatmoz/sftpserver — locally viajust sftp-up && just test-sftp. Being a plain remote filesystem, the S3-only columns (versioning, multipart) don't apply.
Noncurrent versions / delete markers are only observable on a bucket with versioning enabled, and bakelite never deletes object versions — object versions are a separate disaster-recovery layer that bakelite leaves alone; expire noncurrent versions with a bucket lifecycle policy. See S3 storage overhead.
Request pricing (R2 and friends)
Object stores bill two things: bytes kept, and requests made. For a continuous
replicator the second one is the surprise. Cloudflare R2 charges Class A operations
— PUT, LIST, COPY, each multipart part — at $4.50 per million and makes
DELETE free, so churn shows up nowhere in your storage figure and entirely on the
request line. S3, B2, GCS and Azure price the same shapes differently, but the
ranking is the same.
Where a database's requests come from:
- One
PUTper change-set. Under continuous writes bakelite ships one everymax_batch_delay(1s by default). Compaction merges those into a single compacted segment half a minute later and retires them, so most of thosePUTs bought nothing durable. - The same churn again, per async mirror. A mirror copies what the primary currently holds, which by default includes every level-0 change-set that will be retired on the next pass.
LISTs. Listing is a Class A operation, billed per 1000-key page, and it is what an idle mirror spends its money on.- The manifest and the tip pointer, rewritten whenever they change.
The recommended topology
Put the hot path on a local disk and let an async mirror carry the offsite copy, shaped so it only receives what it would actually restore from:
[[database]]
name = "app"
path = "/var/lib/app/app.db"
[[database.backends]] # primary: local, free, every commit
type = "file"
path = "/var/backups/bakelite"
[[database.backends]] # offsite: shaped for requests
type = "s3"
bucket = "offsite-backups"
prefix = "bakelite/app"
mode = "async"
min_level = 1 # skip the L0 churn compaction retires anyway
copy_manifest = false # advisory cache; restore rebuilds it by listing
min_level = 1 means the mirror receives snapshots plus compacted segments, so its
restore points land at compaction granularity (compaction_levels[0], 30s by
default) rather than per-commit — on that mirror only. The local primary still
captures every commit. min_level = 2 coarsens it a step further. copy_manifest = false drops an object that is rewritten every few seconds and that restore never
needs. Between passes the reconciler tracks the destination's contents itself
(mirror_relist_interval), so a pass with nothing new to copy costs the bucket
nothing at all.
Rough Class A requests per month for one database committing about twice a second, at R2's list price (R2's free tier is 1M Class A / 10M Class B / 10 GB per month):
| Topology | Class A / month | ≈ cost |
|---|---|---|
| Object store as the sync destination | ~5.8M | ~$28 |
| Async mirror, defaults | ~4.5M | ~$20 |
Async mirror, min_level = 1 | ~0.2M | free tier |
Async mirror, min_level = 2 | ~0.05M | free tier |
An order-of-magnitude guide, not a quote: your own figure scales with commit rate
and compaction_levels. bakelite doctor flags the two configurations that
reliably cost more than they need to — an object store on the hot path of a
compacting database, and an unshaped object-store mirror.