SoliDB 2.0: creating a collection no longer depends on how many you have
On a development instance with 1,144 collections, creating a collection took about 0.2 s — and so did every insert into a collection that did not exist yet. On a fresh instance the same operation took 9 ms. Inserting documents cost the same on both. SoliDB 2.0 changes how collections are stored so that creating one costs a few milliseconds however large the instance is.
All figures on this page were measured on copies of that instance (7.4 GB, 1,148 collections), never on the live data.
Why a create got slower as the instance grew
Up to 1.3, every SoliDB collection was its own RocksDB column family. That is a natural fit — a collection's documents, indexes and metadata sit together, and dropping the collection drops the family — but it has a cost that grows quietly: RocksDB keeps the options of every column family in one OPTIONS file, and every create_cf or drop_cf rewrites that whole file and fsyncs it, under the database mutex. The file has one section per column family, so the cost of creating one collection is proportional to the number of collections in the whole instance.
On the instance that prompted this, the OPTIONS file had grown to 6.3 MB. We compared a checkpoint copy of it with an empty server, same binary:
| Operation | Empty server | 1,144 collections |
|---|---|---|
300 inserts into an existing collection | 2,347 ms | 1,973 ms |
bulk INSERT of 20,000 rows | 166 ms | 189 ms |
create 10 collections | 94 ms | 2,132 ms |
10 inserts that auto-create their collection | 110 ms | 2,234 ms |
Documents were never the problem. Collections were — and a test suite that drops and recreates a 63-collection database on every run pays that 0.2 s sixty-three times, plus the drops. There is no RocksDB option that makes the OPTIONS rewrite cheaper; the only way out is to stop creating column families.
One column family, many keyspaces
In 2.0 every collection lives in a single column family, __keyspaces__. What used to be the column family's boundary is now a key prefix: eight bytes, the database id then the collection id, each a big-endian 32-bit number allocated in SoliDB's catalog and never reused.
Creating a collection is one atomic write: the registry entry with its keyspace id, the id counter, and the collection's type. Dropping one is the registry delete plus one DeleteRange over [prefix, prefix + 1); dropping a whole database is a single range delete over its id range. A background thread compacts dropped ranges so their disk space comes back soon after, and resumes after a restart.
Because ids are never reused, a collection dropped and recreated under the same name gets a fresh, empty range: it cannot see anything its predecessor wrote, and caches keyed by keyspace simply miss. A handle that some request still holds on the old collection now answers CollectionNotFound instead of writing into the new one.
Inside a keyspace nothing changed: documents are still under doc:, index entries under idx:, fulltext terms under ft_term:. Collection code keeps building those logical keys; one small layer adds the prefix on the way in, strips it on the way out, and bounds every iterator to its keyspace — which also fixed four scans that, in 1.x, could walk past the end of their collection.
The numbers
The same operations on two copies of the 1,144-collection instance — one left on the 1.3 layout, one migrated — each run twice:
| Operation | 1.3 | 2.0 |
|---|---|---|
create 10 collections | 1,952 ms | 53 ms |
10 inserts that auto-create | 1,918 / 1,971 ms | 56 / 57 ms |
delete 10 collections | 49 / 56 ms | 54 / 56 ms |
drop a database | 6 / 7 ms | 6 / 7 ms |
300 single inserts | 3,201 ms | 1,580 / 1,601 ms |
OPTIONS file | 6.3 MB | 43,887 bytes |
1.3’s second “create 10” run took 55 ms and is left out of the table: it recreated the same ten names within the five-minute window in which 1.x reuses a dropped collection’s column family. A name the instance had not seen recently still cost ~0.2 s. Deletes and database drops were already cheap in 1.x, which deferred the column-family drop to a background thread; they stay cheap. The create path is where the time went, and it is now flat. The timings are wall-clock times of HTTP calls from a shell loop, so they include the client; they are for comparing the two columns, not for quoting as per-operation latency.
The migration
A data directory written by 1.x is migrated automatically at the first start of 2.0, inside the storage engine's initialisation and before the server binds its port, so no request ever sees a collection half-moved. Collections move one at a time, smallest first:
# 1. record migrating_to = ks in the registry (synced) # 2. copy every key into ks, counting keys and hashing key + value # 3. re-read the copy, compare count and hash # 4. switch the registry entry to ks (synced) # 5. free the old column family's files; drop it in the background
Every step can be interrupted. A crash during the copy leaves migrating_to on the entry, and the next start wipes the target range and copies again. A crash after the switch leaves an old column family to free, which the next start does. A collection that fails verification or hits an I/O error stays in its own column family — still served, exactly as in 1.x — and is retried at the next start. Freeing each old column family's files right after its copy means the migration needs about one extra collection of disk, not a second copy of everything.
On the 7.4 GB instance: 1,148 collections in 128.6 s, 11.2 GB copied uncompressed, no failures, and five column families left alone because their database no longer existed — the remains of interrupted drops, reported rather than resurrected or deleted. The per-collection document counts matched the untouched copy everywhere except in three test databases that a test suite was writing to between the two checkpoints. The 1,147 emptied column families were then dropped in the background, one at a time, in 124.5 s — the last time this instance pays for OPTIONS rewrites.
Upgrading
The change is one-way: a 1.x binary cannot open a data directory once its collections have moved. Before upgrading, take a checkpoint — POST /_api/backup — and copy it off the volume; rolling back means restoring it. Check that free disk is at least 1.2 × your largest collection plus 1 GiB (the migration refuses to start otherwise). Then either start 2.0 normally, or run the migration on its own first:
solidb --data-dir ./data --migrate-only
A few things behave differently afterwards. A collection's disk_usage is RocksDB's approximate size of its key range, counting flushed data only. The /metrics endpoint gains solidb_keyspace_creates_total, solidb_keyspace_drops_total and solidb_keyspace_gc_compactions_total, while solidb_cf_ops_total now only counts the drops of migrated 1.x column families. The full layout is described in docs/storage-format.md, the upgrade steps in docs/BACKUP.md, and every change in the changelog.