jemalloc now serves RocksDB, and /metrics says where memory goes
A 613-collection SoliDB instance was OOM-killed at 21.7 GB RSS while holding 6.3 GB of data, and nothing it exposed could say which part of it had grown. SoliDB 1.1.0 changes two things: jemalloc now serves RocksDB's C++ allocations as well as Rust's, and /metrics reports memory per component.
The measurements quoted here are the ones recorded in the 1.1.0 changelog and source. The /metrics excerpt comes from a throwaway server built from a later release; the gauges it shows are the ones 1.1.0 added, unchanged.
The arithmetic that looks right
The kernel killed a bare solidb — prod profile, no flags — at 21.7 GB resident and 113 GB virtual, on a 31 GB machine. It held 6.3 GB on disk across 2969 SST files. Idle, the same dataset sits at about 1.5 GB, so the growth came from load: an import writing to many collections at once.
There is an explanation that fits almost too well. One collection is one RocksDB column family, and the prod profile gives each one a 64 MB write buffer. 613 collections times 64 MB is 38.3 GB, which is close enough to 21.7 GB to be believed.
It is wrong. Memtables measured 2.2 MB at idle and plateaued around 202 MB during a 400 000-document import while RSS kept climbing. The growth tracked bytes written, not collections open. Nothing the server exposed at the time could show that, and the reason the numbers did not add up was one level lower than RocksDB: which allocator was holding the memory.
Two allocators in one process
SoliDB has used jemalloc as its global allocator for a while, because glibc's per-thread arenas keep RSS at the allocation high-water mark long after the memory is freed. But tikv-jemallocator was linked with its _rjem_ prefix: 556 _rjem_* symbols in the binary and no unprefixed malloc. A prefixed jemalloc backs Rust's GlobalAlloc and nothing else.
RocksDB is C++. Its new and malloc calls — block cache, table readers, memtables, iterators, compaction buffers — went straight to glibc. On a large instance that is the bulk of the memory, and it was on the allocator the dependency had been added to avoid.
That changes what a fix can reach. The background_thread and 2-second decay tuning set in main.rs had fixed real Rust-side retention: on a dev box with 89 databases, a 3200-query burst at 32× concurrency took RSS from 662 MB to 1296 MB and was still at 954 MB five minutes later; with the tuning, the same burst came back to 625 MB. But it could not touch RocksDB. Here is the idle baseline of the 613-collection dataset, prod profile, measured on a checkpoint:
| Where | Size |
|---|---|
| RSS (cgroup) | 1540 MB |
glibc main arena, [heap] | 1094 MB |
| 4 glibc arenas of ~64 MB | 256 MB |
| 1 arena of ~59 MB | 59 MB |
| binary and the rest | ~90 MB |
| RocksDB, self-reported | 584 MB (block cache 498, table readers 84, memtables 2) |
jemalloc stats.resident | 35 MB |
About 1.4 GB was in glibc; RocksDB accounted for 584 MB of it. The rest was arena retention from startup, which counts documents by reading essentially the whole dataset. Jemalloc — the allocator everyone assumed was in charge — held 35 MB.
What 1.1.0 changed
1.1.0 enables tikv-jemallocator's unprefixed_malloc_on_supported_platforms feature. C and C++ malloc now go through jemalloc as well, and the existing tuning finally applies to RocksDB. Same 613-collection checkpoint, 400 000 documents imported across 200 collections, prod profile, one run per variant:
| glibc | jemalloc | |
|---|---|---|
peak RSS (VmHWM) | 1881 MB | 1356 MB |
| glibc arenas (~64 MB each) | 4 → 20 | 4 → 2 |
[heap] (brk) | 1094 MB | absent |
| still held afterwards | 1726 MB, swapped out and never returned | — |
One trap comes with it. The tuning is a link-time symbol read by jemalloc before main runs, and its name follows the prefix: plain malloc_conf with the feature, _rjem_malloc_conf without. Change one and not the other and the tuning is silently lost. So the server reads the effective options back at startup and logs them. This is the line to look for:
INFO solidb: jemalloc: background_thread=true, dirty_decay_ms=2000, muzzy_decay_ms=2000
If the symbol does not reach the allocator, the same function logs a warning that background_thread is off instead. There is no jemalloc on MSVC builds, so none of this applies on Windows.
Reading /metrics
The second change is attribution. /metrics gains gauges for each RocksDB consumer and six from jemalloc's own accounting. The endpoint answers 401 unless the scraper sends SOLIDB_METRICS_TOKEN, carries an admin JWT, or SOLIDB_METRICS_PUBLIC=1 is set (see Monitoring). With an admin login:
# three collections, 20 000 small documents each, prod profile
TOKEN=$(curl -s -X POST localhost:6745/auth/login -H 'content-type: application/json' \
-d '{"username":"admin","password":"admin"}' | jq -r .token)
curl -s -H "Authorization: Bearer $TOKEN" localhost:6745/metrics \
| grep -E '^solidb_(memtable|table_readers|block_cache|sst|column|cached|jemalloc)'solidb_memtable_bytes 19556352 solidb_memtable_total_bytes 19556352 solidb_table_readers_bytes 19719 solidb_block_cache_bytes 4558 solidb_block_cache_pinned_bytes 87 solidb_sst_files 8 solidb_column_families 13 solidb_cached_collection_handles 12 solidb_jemalloc_allocated_bytes 109900968 solidb_jemalloc_active_bytes 114028544 solidb_jemalloc_resident_bytes 122896384 solidb_jemalloc_mapped_bytes 176951296 solidb_jemalloc_retained_bytes 260304896 solidb_jemalloc_metadata_bytes 9571232
The process's VmRSS at that moment was 153 580 kB, so jemalloc's resident figure covers most of it. Before 1.1.0 the RocksDB half of that would have been invisible to it. Reading the gauges:
solidb_memtable_bytes is live memtable memory across every column family; solidb_memtable_total_bytes adds immutable memtables waiting to flush. The gap between them is flush backlog. Here the 60 000 fresh documents are still in memtables and nothing is waiting.
solidb_table_readers_bytes is index and filter blocks pinned per open SST, outside the block cache and never evicted, whenever --bounded-index-cache is off and --max-open-files is -1 — the prod defaults. This is the one that grows with the dataset.
solidb_block_cache_bytes and _pinned_bytes are read from the single shared LRU. The cache is one for every column family, so summing it per column family would multiply it by the collection count.
solidb_sst_files, solidb_column_families (including default and _meta) and solidb_cached_collection_handles are counts. The last one matters on large instances: each handle carries a tokio::sync::broadcast ring, so at 613 collections it is a memory figure.
The solidb_jemalloc_*_bytes gauges are jemalloc's stats.*. allocated is live data; the distance up to resident is fragmentation and pages not yet purged. retained is address space kept mapped, which is what separates a 113 GB virtual size from a leak.
All of it is computed when the endpoint is scraped, never on a timer: it is one RocksDB property read per column family, and the repository has already had to fix a stats collector that did ten such reads per collection every five seconds. The jemalloc figures need the stats feature of tikv-jemalloc-ctl, which builds the C library with its counters enabled.
The knobs that bound memory
RocksDB memory in SoliDB is dominated by per-column-family structures, and one collection is one column family, so RAM scales with the collection count unless something caps it. Two presets, and every knob they set is also a flag of its own:
| Flag | prod (default) | --dev |
|---|---|---|
--block-cache | 512 MB | 128 MB |
--write-buffer-size (per collection) | 64 MB, up to 3 memtables | 8 MB, up to 2 |
--memtable-budget (all collections) | unlimited | 128 MB |
--max-open-files | -1 (unlimited) | 512 |
--bounded-index-cache | off | on |
--max-background-jobs | 6 | 2 |
Each flag also reads an environment variable (SOLIDB_BLOCK_CACHE, SOLIDB_MEMTABLE_BUDGET, SOLIDB_WRITE_BUFFER_SIZE, SOLIDB_MAX_OPEN_FILES, SOLIDB_BOUNDED_INDEX_CACHE, SOLIDB_MAX_BACKGROUND_JOBS). A production node does not need --dev to bound its memory, and should not take its smaller cache and fewer background jobs just for that:
# prod throughput, with the three unbounded consumers capped
solidb --data-dir /var/lib/solidb \
--memtable-budget 1GB --bounded-index-cache --max-open-files 5000The three that matter are the ones prod leaves open. --memtable-budget caps total memtable memory, forcing flushes before each collection fills its own buffer. --bounded-index-cache moves index and filter blocks into the block cache, where they are counted and evicted, at some read latency. --max-open-files caps the table cache, which is what bounds pinned index and filter blocks when they stay outside the cache. The server says so at startup when they are unset: a warning that there is no global memtable budget, and one that index and filter blocks are pinned per SST with an unlimited table cache.
The dev profile was measured on a box with 1718 column families across 89 databases and 2 GB on disk. RocksDB accounted for about 135 MB there — 18 MB of memtables and 116 MB of block cache, both capped by the profile. Most of the rest of the 2.7–3.8 GB RSS was traffic-driven and on the Rust side, which is where the jemalloc tuning helped.
What is not established
Whether the allocator change alone accounts for the 21.7 GB is not known. The comparison is one run per variant, and 400 000 documents peak at 1.9 GB, not 21.7. Which consumer grew under the original load was never measured, because until 1.1.0 there was nothing to measure it with. On a large instance, set the three storage knobs above, scrape /metrics, and compare solidb_jemalloc_resident_bytes with the RocksDB gauges: the part that is growing now has a name.