SoliDB is a multi-model database in a single Rust binary.

Documents, graphs, vectors, time-series, geo and blobs share one RocksDB store and one query language, SDBQL. The figures on this page were measured while building it. Each one names the release it shipped in and links to the post that shows how it was taken.

O'Saasy License · Rust and RocksDB · HTTP, WebSocket and a MessagePack driver

A transaction that should fail, failing.
# users has a unique index on email
POST …/transaction/$TX/query
  INSERT {_key:"u1", email:"a@x.io"} INTO users   → staged
POST …/transaction/$TX/query
  INSERT {_key:"u2", email:"a@x.io"} INTO users   → staged
POST …/transaction/$TX/commit
{"code":400,"type":"InvalidDocument",
 "error":"Unique constraint violated:
   'uniq_email:046140782e696f00' is claimed
   by both 'u1' and 'u2' in this transaction"}

FOR u IN users RETURN u._key   → "result":[]
Release 2.1.0. Captured on a throwaway server. Before 2.1.0 both inserts committed; now the commit fails as a whole and writes nothing. Source: /blog/solidb-2-1-transactions-conflicts-roles
Columnar rows joined to documents, in one query.
# metrics is a columnar collection, hosts a document one
FOR m IN metrics
  FILTER m.cpu > 50
  FOR h IN hosts
    FILTER h._key == m.host
    RETURN {region: h.region, cpu: m.cpu}

→ [{"cpu": 58.0, "region": "eu-west"},
   {"cpu": 72.5, "region": "eu-west"},
   {"cpu": 90.0, "region": "us-east"}]
Release 0.33.0. Every result in the post was produced on a throwaway server. Source: /blog/columnar-collections-in-sdbql
Days with no orders still get a row.
LET per_day = COUNT_BY((FOR o IN orders RETURN o),
                       o -> DATE_TRUNC(o.at, "day"))
FOR d IN DATE_SERIES("2024-03-01", "2024-03-04", "day")
  RETURN {day: LEFT(d, 10), orders: per_day[d] || 0}

→ [{"day": "2024-03-01", "orders": 2},
   {"day": "2024-03-02", "orders": 0},
   {"day": "2024-03-03", "orders": 1},
   {"day": "2024-03-04", "orders": 0}]
Release 1.3.0. Produced on a 1.3.0 server. Source: /blog/sdbql-1-3-new-functions
A role that stops at one database.
POST /_api/auth/users/bob/roles
  {"role":"editor","database":"v"}
→ 201 {"username":"bob","role":"editor","database":"v", …}

bob reads   v                → 200
bob writes  v                → 200
bob reads   other            → 403
bob writes  other            → 403
bob creates a database       → 403
Release 2.1.0. Captured on a throwaway server. Source: /blog/solidb-2-1-transactions-conflicts-roles
Code the server runs cannot be written by name.
# ana holds the editor role
curl -X POST …/_api/database/blog/document/_scripts \
  -H "authorization: Bearer $T" \
  -d '{"name":"pwn","path":"pwn","code":"return 1"}'

→ HTTP 403
{"code": 403,
 "error": "Access denied: '_scripts' is managed by the
   server and is not writable through this API; use
   the dedicated admin endpoints",
 "type": "Forbidden"}
Release 1.1.0. Captured on a throwaway server; request abridged. Source: /blog/secure-by-default
A second node joins and serves the first one's data.
# n2 log
Sent join request to 127.0.0.1:6923
Successfully joined cluster. Received 2 peers.
Starting full sync: 1 databases, 7 documents
Full sync complete, final sequence: 5

# query on n2
FOR n IN notes RETURN {key: n._key, from: n.from}
→ [{"from": "n1", "key": "hello"}]
Release 0.34.0. Real nodes started on one host; log abridged. Source: /blog/clustering-across-machines
1 / 6

Measured, with the method attached.

Every row is copied from a release post. Where the post says a figure is a single run, a wall-clock time from a shell loop, or was not measured at all, this table says so too. The two bars in a row share one scale.

Create 10 collections

On a copy of a 1,144-collection instance, 7.4 GB. Each create used to rewrite and fsync RocksDB's whole OPTIONS file.

1.31,952 ms
2.053 ms

10 inserts that auto-create their collection

Same instance. Wall-clock HTTP calls from a shell loop: compare the rows, do not quote them as latency.

1.31,918 ms
2.056 ms

Peak memory, 4,000-row UPSERT

Rows sent as one 500 KB bind variable, which was copied into every row's context. Before the fix it hit the 30 s timeout; after, it finished in 70 ms.

1.2.215.9 GB
1.2.3178 MB

Peak RSS, 400,000 documents into 200 collections

613-collection checkpoint, prod profile, one run per variant. RocksDB's allocations moved from glibc arenas to jemalloc.

1.01,881 MB
1.11,356 MB

Server start, about 920 collections

Two full sweeps of every collection ran at each start. After a clean shutdown they are skipped. The new start time was not measured, so none is shown.

1.122.5–33.0 s
1.2not measured

The same query shape for every model.

FOR … RETURN walks a graph, ranks vectors and buckets a time series. The query text is shown without invented output: run it against your own data with the SDBQL reference.

Graph1 to 2 hops out
FOR v, e IN 1..2 OUTBOUND 'people/alice' follows
  FILTER e.since > 2020
  RETURN { person: v.name, since: e.since }
Vectorbind the query embedding as @vec
FOR doc IN articles
  LET sim = VECTOR_SIMILARITY(doc.embedding, @vec)
  FILTER sim > 0.8
  SORT sim DESC LIMIT 10
  RETURN { title: doc.title, score: sim }
Hybridvector index + fulltext field, one call
FOR r IN HYBRID_SEARCH("articles", "embedding_idx",
    "content", @vec, "machine learning",
    { vector_weight: 0.8, text_weight: 0.2, limit: 5 })
  RETURN r
Time-seriesrequests per hour
FOR e IN events
  COLLECT bucket = TIME_BUCKET(e.ts, '1h')
  AGGREGATE hits = COUNT(1)
  SORT bucket ASC
  RETURN { bucket, hits }

What it will not do for you.

Written down here so you find them before production does. Each one is also in the docs.

A checkpoint is not an off-site backup.

POST /_api/backup hard-links SST files on the same volume. It protects against a bad write or a dropped collection, not against losing the disk. On a cluster it covers the local node only. Copy it elsewhere.

Replication is eventually consistent.

Nodes replicate master to master, and writes queue for a node that is offline. Until the queue drains, two nodes can disagree about the same document.

Every query has a ceiling.

An HTTP query runs under a 30-second deadline and may materialise at most 5,000,000 intermediate rows (SOLIDB_MAX_INTERMEDIATE_ROWS). A query that needs more fails; the server keeps running.

No job queue for your application.

SoliDB schedules its own work: trigger dispatch, embedding generation, view refresh. Your background jobs and cron belong in your application framework.

One unique-index edge case remains.

Deleting the document that holds a unique value and inserting another with the same value, in one transaction, is refused. Do those in two transactions.

The 2.0 storage migration is one-way.

A 1.x binary cannot open a data directory once its collections have moved. Take a checkpoint first, and have free disk of 1.2 × your largest collection plus 1 GiB.

Build it and run it.

RocksDB is compiled from source, so the first build needs a C++ toolchain and takes a while. The server listens on port 6745 and prints a random admin password to its log on first boot.

Client libraries: Rust, Go, Python, JavaScript (Node and browser), PHP, Ruby, Elixir, Laravel Eloquent, and a mobile SDK. Or plain HTTP and curl. Clients in the docs.

# Debian / Ubuntu build dependencies
$ apt-get install build-essential clang libclang-dev \
    pkg-config libssl-dev libzstd-dev

$ git clone https://github.com/solisoft/solidb
$ cd solidb && cargo install --path .
$ solidb --port 6745 --data-dir ./data