A 3 MB bind variable, 20,200 rows, and a server killed at 61 GB
SoliDB works to one rule: a request may fail, but it must never take the server down. 1.1.0 closed the known ways a single request could crash or stall the process. On 26 September a bulk upsert found another way, and the kernel OOM-killed the server at about 61 GB. This post covers the rule, the incident, and the fix in 1.2.3.
The incident and before/after figures come from the 1.2.3 changelog and the fix commit. The run shown near the end was made on a 1.2.3-or-later build.
The rule
A database server is shared. One client's bad query, bad script or bad payload should come back to that client as an error, not as an outage for everyone else. In practice a request can bring the server down in two ways: by killing the process (a stack overflow, a panic, an allocation the kernel refuses) or by stalling it (a thread pool that never frees up, a network call with no deadline).
1.1.0 went through both lists. The changelog entry is titled A request may fail; it may no longer take the server with it. On the crash side it fixed SQL parser recursion, cyclic Lua tables, RANGE i64 overflow, K_PATHS recursion, and WebSocket Lua scripts that ran without a hook or a memory limit. On the stall side it fixed Lua pool spin, unbounded scatter-gather and fetch HTTP, driver payload reads, executors without a deadline, and a transaction reaper that was never scheduled. Whole-collection reads on the query path are now bounded before they allocate, COLLECT streams its aggregates, and index builds run on the blocking pool.
Some resources a single client could grow without limit now have a ceiling, each set by an environment variable. An invalid value or 0 falls back to the default:
| Variable | Default | Bounds |
|---|---|---|
SOLIDB_MAX_INTERMEDIATE_ROWS | 5,000,000 | Rows one query may materialise, including nested-FOR cartesian products that no LIMIT applies to |
SOLIDB_QUEUE_MAX_CONCURRENCY | 4 × cores | Jobs executing at once |
SOLIDB_STREAM_MAX_BUFFER_EVENTS / _BYTES | 100000 / 64 MiB | Per-stream window buffer |
SOLIDB_LUA_FETCH_MAX_BYTES | 10 MB | Lua fetch response body |
SOLIDB_LUA_RATE_LIMIT_MAX_KEYS | 100000 | Keys held by solidb.rate_limit |
SOLIDB_REPL_MAX_SESSIONS_PER_USER / SOLIDB_REPL_MAX_SESSIONS | 8 / 1000 | REPL sessions |
SOLIDB_CLUSTER_HTTP_TIMEOUT_SECS / _READ_TIMEOUT_SECS | 60 / 60 | Inter-node HTTP requests |
Queries also have a deadline: the HTTP query handler runs the executor with a 30-second timeout. Deadlines plus row ceilings looked like enough to bound what a query can cost.
16:55, 26 September
At 16:55 on 2026-09-26 the kernel OOM-killed the main SoliDB server. It had reached 17.9 GB resident and 43 GB in swap, about 61 GB in total.
The query that did it was an ordinary bulk import. It sent 20,200 rows as one bind variable of about 3 MB, and ran an UPSERT for each row:
FOR row IN @rows LET k = LOWER(row.email) UPSERT { _key: k } INSERT { _key: k, name: row.name, source: @source } UPDATE { name: row.name, source: @source } IN contacts
It is the usual shape for an import: one round trip, values passed as bind variables rather than spliced into the query text. 20,200 rows is 0.4% of the intermediate-row ceiling. 3 MB is a small request body. None of the limits above had anything to object to.
Rows × every bind variable
While a query runs, each row is evaluated against a context: a map from variable names (row, k, and so on) to values. The executor starts from an initial context and clones it once per row, so every row starts from the same bindings.
Before 1.2.3, the bind variables were copied into that initial context as @rows and @source. That meant every per-row clone also copied the entire 3 MB @rows array, including the rows the loop had already processed and the ones it hadn't reached yet. The cost was rows × the size of every bind variable. For the incident query that is 20,200 × 3 MB, about 60 GB, in line with the ~61 GB the kernel saw.
@rows was copied into every row's context. In 1.2.3 the contexts no longer carry bind variables, and @rows exists once, on the executor.The 30-second query timeout didn't help. As the changelog says, the copies happen faster than the timeout fires. A deadline limits how long a query runs, not how much it allocates in that time.
The fix
The evaluator already looked up bind variables on the executor first (self.bind_vars), so the copies in the context were redundant. 1.2.3 removes all five places that made them: the query entry point (src/sdbql/executor/execution/entry.rs), subqueries, EXPLAIN, and the two HTTP mutation paths in src/server/handlers/query.rs and src/server/transaction_handlers.rs. A top-level query now starts from an empty context. Bind variables stay on the executor and are never cloned per row.
The changelog gives measurements for a 4,000-row upsert with a 500 KB @rows:
| Peak memory | Outcome | |
|---|---|---|
| Before (1.2.2) | 15.9 GB | 30 s timeout |
| After (1.2.3) | 178 MB | 70 ms |
A regression test, test_bulk_upsert_over_large_bind_array, runs this shape: 5,000 rows upserted over 1,000 existing documents, with a second bind variable used inside the INSERT and UPDATE. Its comment explains why the row count was chosen: if the bug returns, the test gets slow instead of fatal. A companion test checks that bind variables still resolve wherever evaluation happens in a derived context: correlated subqueries, lambdas, and LETs before the first FOR.
The same shape, run again
Here is the query from the incident on a throwaway server built from 1.2.3 or later. The contacts collection holds 1,000 documents, u0 to u999. The request body carries 5,000 rows (U500 to U5499, each with a 64-character padding field) and is 589,289 bytes. The query counts the rows it processed:
FOR row IN @rows LET k = LOWER(row.email) UPSERT { _key: k } INSERT { _key: k, name: row.name, source: @source } UPDATE { name: row.name, source: @source } IN contacts COLLECT WITH COUNT INTO n RETURN n
{"result": [5000], "count": 1, "has_more": false, "cached": false,
"executionTimeMs": 58.468908,
"inserted": 4500, "updated": 500, "deleted": 0}Of the 5,000 rows, 500 matched existing keys and were updated, and 4,500 were inserted. The query ran in 58 ms and the whole HTTP request took 64 ms. Sending the same body twice more, which now updates all 5,000 rows, took 66.9 ms and 68.1 ms. The server process's peak resident size (VmHWM) for the whole run, with RocksDB included, was 223,792 kB. The UPDATE merged into the existing documents instead of replacing them:
FOR c IN contacts FILTER c._key IN ["u700", "u5499"] SORT c._key RETURN UNSET(c, "_rev", "_id", "_created_at", "_updated_at")
[
{"_key": "u5499", "name": "Nom 5499", "source": "import-2026-09"},
{"_key": "u700", "name": "Nom 700", "phone": "keep", "source": "import-2026-09"}
]For how UPSERT decides between its two branches, see mutations. For bind variables and the cursor endpoint, see the query API.
What it says about the rule
Every limit from 1.1.0 was still in place when the server died. None of them was built for this. The row ceiling counts rows, and 20,200 is far below 5,000,000. The query deadline counts seconds. The copies were cheap one at a time, and the problem was the multiplication: a small input times a moderate row count. Measuring each dimension on its own missed it.
The fix removes the multiplication instead of adding another limit. Bind variables are held once, so their cost no longer grows with the row count. On 1.2.3, send the batch as one bind variable.
If you run 1.2.2 or older and import through bind variables, upgrade. Until then, keep each batch small, because memory grows with rows times the total size of the bind variables and the timeout will not stop it. Every entry is in the changelog, and the resource limits are listed in getting started.