What hundreds of collections cost, and what 1.2.0 took off the bill
In SoliDB, one collection is one RocksDB column family, and the instance pays for every one of them in places that have nothing to do with the collection you are touching: at startup, when a collection is created or dropped, when collections are listed, when the WAL fills. 1.2.0 removes most of that bill. This is where it came from and what changed.
The numbers are the ones measured for the 1.2.0 changelog. The terminal output was captured on a throwaway test server; the code paths it exercises are the ones 1.2.0 shipped.
Why a collection costs more than its data
Every collection lives in its own column family. That keeps a collection’s documents, indexes, fulltext entries and blob chunks together — the doc:, idx:, ft: and blo: prefixes all sit inside it — and makes dropping a collection a single operation. It also means RocksDB knows about every collection individually.
The expensive part is the OPTIONS file. Each create_cf and drop_cf rewrites and fsyncs the whole file, which holds one section per column family, and then re-parses what it just wrote. The cost is proportional to the instance’s total collection count, not to the collection being created. RocksDB has no setting that avoids it. While the rewrite runs it holds the database mutex, so writes elsewhere wait behind it — latency that no single query accounts for.
On an instance with a handful of collections none of this shows. On one with hundreds, spread over dozens of databases and churned by test suites, it was most of the work the server did. Each 1.2.0 change below removes one place where that total count was paid for.
Startup: two sweeps that a graceful stop made pointless
On an instance with ~920 collections, SoliDB took 22.5–33.0 s to start. RocksDB’s own open measured 1.81 s. The rest was two full sweeps of every collection:
initializerecounted everydoc:key of every collection, unconditionally. That is the right thing after a crash, when the persisted document counts may be behind the data. After a graceful stop it was wasted.Collection::newwalked theblo:prefix of every collection to count blob chunks — including the ~96% of collections that held no blob at all.
1.2.0 writes a clean-shutdown marker into _meta during a graceful shutdown, once the collection stats and RocksDB have been flushed, so the marker never vouches for data that is not on disk. At the next start the marker is read and cleared in the same step. If it was there, the persisted counts are trusted and the recount is skipped. The blob chunk count is no longer computed up front; it is resolved the first time a collection needs it.
The crash path did not change. No marker — because the process was killed, or because the marker could not be cleared — means the full recount runs, exactly as before. The server log says which path it took. After a stop with SIGTERM:
2026-09-27T15:56:28.322968Z INFO solidb::storage::engine: Clean shutdown detected — trusting persisted document counts
And after a start that followed a failed one, with no marker to find:
2026-09-27T15:56:04.371341Z INFO solidb::storage::engine: Recalculated document counts for 7 collections
Databases no longer start with two collections
Creating a database used to create _scripts and _slow_queries in it up front: two create_cf calls, two full OPTIONS rewrites, whether that database would ever run a script or log a slow query. On one dev instance that came to 43 _scripts and 42 _slow_queries column families across 46 databases, almost all empty.
Both collections were already created on first use, so 1.2.0 simply stops pre-creating them. The race the pre-creation was guarding against is handled where it happens: the slow-query logger creates the collection and then retries the lookup ten times.
Listing without the column-family lock
Whether a collection existed was answered by RocksDB’s column-family map. Reading it cloned every name in the whole instance — 963 allocations to list one database — and took a read lock that create_cf and drop_cf hold for the whole OPTIONS rewrite. So listing one database’s collections could block for hundreds of milliseconds behind a collection being created in another.
1.2.0 keeps a coll:{db}:{name} entry per collection in _meta. Listing is a prefix scan over those entries, with no column-family lock.
The column-family map stays the truth; the entries are an index over it. A startup pass adopts any column family that has no entry, so a crash between create_cf and the entry write — or a downgrade to a binary that never wrote entries — heals itself, and every listing path falls back to the map when there is no _meta to consult. A missing entry can never make a collection disappear.
The same clone-every-name call drove the embedding and TTL sweeps once per database, so each pass cost databases × total collections string allocations. Both now take a single grouped pass.
Delete, then recreate, for free
Deleting a collection called drop_cf inline. Creating one with the same name afterwards called create_cf. Two full rewrites and fsyncs of the OPTIONS file, to end up exactly where it started. That pattern is ordinary traffic: test suites drop and recreate the same collections on every run, and writing to a collection that does not exist creates it. On one dev instance, 11,241 collections were auto-created in eighteen days, across 321 distinct names — one of them 358 times. That was the dominant cost of the instance’s traffic.
In 1.2.0 a delete no longer drops the column family. It erases the collection’s data with a range tombstone, so the space is reclaimed at deletion time, and keeps the empty column family for a grace period — SOLIDB_CF_REUSE_GRACE_SECS, 300 seconds by default. A create under the same name within that window claims the shell and reuses it. Every column family is built from the same options, so a reused one is indistinguishable from a new one. A single reaper thread drops the shells nobody came back for.
Nothing visible changes. The collection is gone the instant the delete returns, as before, and its data is already erased, so the shell holds nothing while it waits. Pending drops are persisted as markers in _meta; on shutdown the reaper stops without spending rewrites, and the next startup resumes them.
Here it is on a fresh server. Writing a document to orders auto-creates it, deleting it keeps the shell, and writing again reuses it — solidb_cf_ops_total does not move, solidb_cf_reuses_total does, and the query sees only the new document:
# server started with SOLIDB_METRICS_TOKEN=tok; /metrics needs it
curl -s -o /dev/null -u admin:admin -X POST localhost:6745/_api/database/blog/document/orders \
-H 'content-type: application/json' -d '{"total":15}'
curl -s -H 'X-Metrics-Token: tok' localhost:6745/metrics | grep -E '^solidb_(cf_|collections_auto)'solidb_cf_ops_total 2 solidb_cf_reuses_total 0 solidb_collections_autocreated_total 1 solidb_cf_op_seconds_total 0.003977
curl -s -o /dev/null -u admin:admin -X DELETE localhost:6745/_api/database/blog/collection/orders
curl -s -o /dev/null -u admin:admin -X POST localhost:6745/_api/database/blog/document/orders \
-H 'content-type: application/json' -d '{"total":40}'
curl -s -H 'X-Metrics-Token: tok' localhost:6745/metrics | grep -E '^solidb_(cf_|collections_auto)'solidb_cf_ops_total 2 solidb_cf_reuses_total 1 solidb_collections_autocreated_total 2 solidb_cf_op_seconds_total 0.003977
FOR o IN orders RETURN o.total
[40]
On a fresh instance a rewrite takes milliseconds, as solidb_cf_op_seconds_total shows. The point is that it grows with every collection the instance holds, and the reuse path does not.
The WAL budget is a flush trigger
max_total_wal_size was pinned at 50 MB. It reads like a disk cap, but it is a flush trigger: when the WAL crosses it, RocksDB flushes every column family that holds data in the oldest WAL file. With many collections each holding a little recent data, that is a stampede.
Measured on a 963-collection instance: 27,904 of 27,910 flushes were WAL Full, 97.7% of them writing under 4 KB, in 39 bursts averaging ~715 column families each. That is where 3,717 SST files came from, 87.8% of them under 64 KB. And it did not even hold the line: the WAL sat at 201.8 MB.
In 1.2.0 the budget is part of the storage profile — 2 GB in prod, 256 MB under --dev — and --max-total-wal-size (or SOLIDB_MAX_TOTAL_WAL_SIZE) overrides it. Lowering it to save disk is a false economy; it brings back the thousands of tiny SSTs. To bound memory, use --memtable-budget: the prod profile now sets a global memtable budget, and that trigger flushes one column family at a time rather than all of them. Two smaller fixes rode along: arena_block_size is down to 64 KB, since the 1 MB default is reserved per column family, and RocksDB’s info log has a size bound after reaching 118 MB.
What to watch, and one switch
1.2.0 adds four counters to /metrics, so column-family churn shows up as a number instead of as unexplained latency:
| Metric | Counts |
|---|---|
solidb_cf_ops_total | Column-family creates and drops since start — each one an OPTIONS rewrite |
solidb_cf_op_seconds_total | Wall time spent inside those creates and drops |
solidb_cf_reuses_total | Deleted collections’ column families wiped and reused instead of dropped and recreated |
solidb_collections_autocreated_total | Collections created by a write to a name that did not exist |
If solidb_collections_autocreated_total climbs on an instance whose schema is managed elsewhere, every one of those is a new column family and a bigger OPTIONS file — sometimes from a typo. SOLIDB_AUTO_CREATE_COLLECTIONS=0 turns auto-creation off, and the write fails instead:
# server started with SOLIDB_AUTO_CREATE_COLLECTIONS=0
curl -s -u admin:admin -X POST localhost:6745/_api/database/blog/document/ordres \
-H 'content-type: application/json' -d '{"total":15}'{"code":404,"error":"ordres","type":"CollectionNotFound"}Auto-creation stays on by default, because applications rely on it.
None of this needs configuration to take effect: the first graceful stop on 1.2.0 writes the marker, the _meta entries are backfilled at startup, and deletes start keeping their shells. The /metrics endpoint is described in monitoring; the full list of 1.2.0 changes is in the changelog.