Release 0.33.0 · Operations

Backing up SoliDB: checkpoints, dumps, and when to use each

Until 0.33.0 the only way to back up SoliDB was solidb-dump: re-read every document over HTTP and write it out as JSON lines. 0.33.0 adds a physical backup, POST /_api/backup, which takes a RocksDB checkpoint of the whole instance in about the time it takes to flush. The two are not interchangeable. This is what each one does, when to reach for which, and the commands.

The dump, restore and checkpoint shown here were run on a throwaway server. That build was newer than 0.33.0; where it behaves differently, this post describes 0.33.0 as its source reads, and the extra password warning the newer tools print is left out.

Two kinds of backup

All databases on a SoliDB server share one RocksDB instance, and each collection is a column family inside it. That fixes what a backup can be.

A physical backup copies the storage engine's files. RocksDB can do this without copying bytes: a checkpoint is a new directory whose immutable SST files are hard links to the ones the live database is using, plus the small files that describe them. It covers the whole instance, because that is the unit RocksDB knows about. There is no per-database checkpoint.

A logical backup asks the server for documents and writes them down. solidb-dump runs FOR doc IN c RETURN doc through the cursor API, collection by collection, and writes one JSON record per line. It can take one database or one collection, and the result is text you can read, edit and load into another version.

A checkpoint hard-links SST files; a dump writes JSONL through the API Checkpoint same disk, same inodes data/ 000031.sst 000032.sst … MANIFEST checkpoint/ 000031.sst 000032.sst … MANIFEST - - - hard link: one file, two names whole instance, point in time, lost with the volume Dump re-reads documents over HTTP server cursor API solidb-dump -d shop {"_type":"collection",…} {"_type":"index",…} {"_type":"document","doc":{…}} one database or collection, portable, bigger, slower to load back
A checkpoint gives the same SST files a second name in a new directory; a dump reads documents through the API and writes a new file of JSON lines.

Physical: POST /_api/backup

The endpoint takes one field, path: a directory on the server, since the server process is the one holding the database open. It must not exist yet. It requires admin, because the caller picks where the server writes and the result holds every database whatever the caller's grants.

terminal
curl -X POST http://localhost:6745/_api/backup \
  -u admin:"$SOLIDB_PASS" \
  -H 'Content-Type: application/json' \
  -d '{"path": "/var/backups/solidb/2026-07-27"}'

The server flushes the memtables first, so the checkpoint holds recent writes without depending on WAL replay, then asks RocksDB for the checkpoint. That work runs off the async worker threads, because hard-linking scales with the number of column families. The reply says "status": "created", echoes the path, states "scope": "instance" and carries a note repeating the warning below. A second request to the same path is refused with a 400: checkpoint target … already exists.

On the test server the hard links were easy to see. An SST file in the data directory and its namesake in the checkpoint had the same inode and a link count of 2. A document inserted after the checkpoint was absent when a server was started on a copy of it; the three customers written before it were there, and so was the index on orders.

That is the restore procedure. A checkpoint is a complete data directory, so you point a server at it:

terminal
solidb --port 6745 --data-dir /var/backups/solidb/2026-07-27

Two limits come with that speed. First, the checkpoint is not protection against losing the filesystem. The hard links point at the same blocks on the same volume as the live database, so a checkpoint guards against a bad write or a dropped collection and nothing else. Copy it somewhere else once it is taken. A plain copy writes new files, and on the test server the copied SST had its own inode and a link count of 1. Second, it covers the instance you sent it to. On a cluster, the handler checkpoints the local node's RocksDB and nothing else.

It is also all or nothing. Starting from a checkpoint rolls every database back to that moment. To get back one collection someone dropped, start the checkpoint as a second server and dump that collection out of it.

Logical: solidb-dump and solidb-restore

In 0.33.0 solidb-dump needs -d. Add -c for a single collection, and -o for a file (it writes to stdout otherwise). The shop database on the test server has two collections and one persistent index:

terminal
solidb-dump -H localhost -P 6745 -u admin -p "$SOLIDB_PASS" -d shop -o shop.jsonl
Authenticating as user: admin
Dumping database: shop
Found 2 collections
  Collection: customers
  Collection: orders
✓ Dump written to shop.jsonl

The file declares each collection, then its indexes, then its documents:

shop.jsonl (excerpt)
{"_collection":"orders","_collectionType":"document","_database":"shop","_type":"collection"}
{"_collection":"orders","_collectionType":"document","_database":"shop","_index_kind":"persistent","_type":"index","field":"customer","fields":["customer"],"name":"by_customer","unique":false}
{"_collection":"orders","_collectionType":"document","_database":"shop","_type":"document","doc":{"_created_at":"2026-09-27T15:56:25.709817613+00:00","_id":"shop:orders/o1","_key":"o1","_rev":"db28e364-ee38-4635-97dc-d42e9de5a3a2","_updated_at":"2026-09-27T15:56:25.709817613+00:00","customer":"c1","total":10}}

Restoring into a new database, created on the way:

terminal
solidb-restore -u admin -p "$SOLIDB_PASS" -i shop.jsonl -d shop_copy --create-database
Authenticating as user: admin
Restoring using streaming mode (JSONL/Mixed)...
  Created database: shop_copy
✓ Restore completed
  → 6 items imported

Afterwards shop_copy held 3 customers, 2 orders and the by_customer index.

0.33.0 made both tools safer to put in a script:

  • Exit codes mean something. A dump that hit a warning, such as an index list it could not read, now ends with Dump finished with N warning(s); output may be incomplete and a non-zero status. A restore with failed items exits non-zero, and so does one that skipped records it could not route. --allow-skipped accepts those, for dumps from older versions.
  • Documents are enveloped. Routing fields sit outside doc, as above, so a user field named _type or _collection no longer collides with the dump's own.
  • Columnar indexes travel as columnar_index records rather than schema indexed flags.
  • --overwrite upserts by _key (the import endpoint's mode=upsert) instead of failing on a key that exists. --drop is still the clean option: it drops the collections before restoring.
  • --scheme https for a server behind TLS, collection and database names encoded in URLs, and -u or -p given without the other is rejected.

What it costs: a comment in the backup handler's source puts a dump at roughly 3x the on-disk size, and a restore replays every write. A dump is also read collection by collection, so it is not a single point in time across collections.

Which one, when

Checkpointsolidb-dump
ScopeWhole instanceOne database or collection
SpeedNear-instantScales with document count
SizeHard-linked, ~0 extra at firstSeveral times the on-disk size
ConsistencyPoint-in-time across all collectionsPer collection, as read
PortableSame storage format onlyJSONL: readable, editable, portable

And by situation:

You need to…Use
Survive losing the server or its diskCheckpoint, copied off the volume
Undo a bad deploy across several collectionsCheckpoint: every collection from the same moment
Get back one dropped collectionDump it from a server started on a checkpoint, then restore it
Move data to a different SoliDB versionsolidb-dump / solidb-restore
Copy one database to staging, or clone a collectionsolidb-dump -d (and -c), restore with -d / -c
Inspect or edit the data before loading itsolidb-dump

A runbook

Nightly: a checkpoint, then a copy to another machine. The copy is the backup; the local directory is only the snapshot it was taken from.

terminal
# 1. Checkpoint the instance (admin, path on the server, must not exist)
DAY=$(date +%F)
curl -fsS -X POST http://localhost:6745/_api/backup \
  -u admin:"$SOLIDB_PASS" -H 'Content-Type: application/json' \
  -d "{\"path\": \"/var/backups/solidb/$DAY\"}"

# 2. Copy it off the volume
rsync -a /var/backups/solidb/$DAY backup-host:/srv/solidb/

Per database, where you want something portable, check the exit status:

terminal
# Non-zero on any warning: a partial dump must not look like a good one
solidb-dump -u admin -p "$SOLIDB_PASS" -d shop -o shop-$DAY.jsonl \
  && gzip shop-$DAY.jsonl \
  || echo "dump of shop failed" >&2

Restoring a whole instance: start a server on a copy of the checkpoint.

terminal
rsync -a backup-host:/srv/solidb/2026-07-27/ /var/lib/solidb-restored/
solidb --port 6745 --data-dir /var/lib/solidb-restored

Restoring one database from a dump. solidb-restore reads a file, so decompress first:

terminal
gunzip -k shop-2026-07-27.jsonl.gz
# into a fresh database, side by side with the live one
solidb-restore -u admin -p "$SOLIDB_PASS" -i shop-2026-07-27.jsonl -d shop_restored --create-database
# or over the original, dropping its collections first
solidb-restore -u admin -p "$SOLIDB_PASS" -i shop-2026-07-27.jsonl --drop

Keep dumps and checkpoints somewhere access-controlled: both hold every document in plain form. Full option tables are on the tooling page, and the 0.33.0 entry in the changelog lists the rest of the release.