Rule Cascade
Reference

Performance

How fast each runtime evaluates, how the rule server behaves under load, and how to measure it on your own hardware.

How fast each runtime evaluates, how the rule server behaves under load, and how to measure it on your own hardware. The numbers below were measured on one laptop; they are a record of what one run printed, not a guarantee, and they are not checked by CI (benchmarks on shared CI runners are too noisy to gate on).

Running the benchmarks

make bench                     # every runtime, then the rule server; BENCH_SECONDS=5 by default
make bench-typescript bench-java bench-go bench-python bench-server   # one at a time
BENCH_SECONDS=20 make bench-go

Every runtime benchmark does the same thing:

  • reads conformance/bundles/acme.payments.transfer.bundle.json (conformance/bundles/acme.payments.transfer.bundle.json) with its fromBundle (no compiling from YAML);
  • evaluates the requests of tools/bench-requests.json (tools/bench-requests.json) round-robin on the server channel, on one thread: an allowed domestic transfer, a denied one with a rendered message, an international one with three findings and an effect, and one narrowed to a component;
  • warms up for BENCH_WARMUP seconds (default 2), then measures for BENCH_SECONDS;
  • times every evaluation with the monotonic clock (so each latency includes one clock read) and reports evaluations per second (count / measured time) and the p50 and p99 latency (nearest rank).
RuntimeHarness
TypeScriptpackages/typescript/bench/evaluate.mjs (packages/typescript/bench/evaluate.mjs), on dist/
JavaBenchmark.java (packages/java/src/test/java/io/github/yarlisaisolutions/rulecascade/Benchmark.java), a plain main (no JMH), compiled with javac after make java
GoBenchmarkEvaluateTransfer in bench_test.go (packages/go/bench_test.go), go test -bench with p50-ns, p99-ns and evals/s metrics
Pythonpackages/python/bench/evaluate.py (packages/python/bench/evaluate.py), timeit.default_timer per evaluation

The rule server load test, packages/server/bench/load.mjs (packages/server/bench/load.mjs), starts dist/main.js with the token on and examples/contracts as the rules, once with RULE_SERVER_WORKERS=0 and once with a worker pool, one after the other, and posts the same requests to /evaluations over keep-alive connections from load-generating worker threads. It needs no dependency. BENCH_CONNECTIONS (32), BENCH_CLIENT_THREADS (2) and BENCH_WORKERS ("0 N") change the setup. The generator runs on the same machine, so the result compares the two settings; it is not a capacity figure for a server on its own hardware.

Measured on

Apple M2, 8 cores (4 performance, 4 efficiency), 24 GB, macOS 27.0.1 (arm64); Node.js 22.22.2, OpenJDK 17.0.19 (Homebrew), Go 1.26.5, CPython 3.14.6; Rule Cascade 1.0.0-alpha.2; 3 October 2026.

The machine was not idle: other work kept the load average between 17 and 22 during both runs. Expect quieter numbers, and smaller differences between runs, on an idle machine. Two consecutive runs of make bench with BENCH_SECONDS=5, and a third later the same day with the load average between 11 and 15, verbatim:

Run 1:

typescript rule-cascade 1.0.0-alpha.2 node 22.22.2: 35862 evals/s, p50 25.0 us, p99 83.7 us, 179309 evaluations in 5.0 s
java rule-cascade 1.0.0-alpha.2 jdk 17.0.19+0: 78054 evals/s, p50 8.8 us, p99 43.5 us, 390271 evaluations in 5.0 s
BenchmarkEvaluateTransfer-8   	  320043	     22982 ns/op	     43513 evals/s	     19750 p50-ns	     69834 p99-ns
python rule-cascade 1.0.0a2 CPython 3.14.6: 6126 evals/s, p50 134.8 us, p99 492.3 us, 30630 evaluations in 5.0 s
rule server load test: node 22.22.2, 8 CPUs, 32 connections on 2 client threads, 2 s warm-up, 5 s measured
RULE_SERVER_WORKERS=0: 10186 req/s, p50 2.79 ms, p99 8.92 ms, 50930 requests, 0 not 200
RULE_SERVER_WORKERS=4: 5072 req/s, p50 4.03 ms, p99 39.82 ms, 25358 requests, 0 not 200

Run 2:

typescript rule-cascade 1.0.0-alpha.2 node 22.22.2: 27997 evals/s, p50 32.2 us, p99 112.6 us, 139987 evaluations in 5.0 s
java rule-cascade 1.0.0-alpha.2 jdk 17.0.19+0: 87202 evals/s, p50 10.5 us, p99 29.0 us, 436012 evaluations in 5.0 s
BenchmarkEvaluateTransfer-8   	  238772	     32945 ns/op	     30353 evals/s	     26041 p50-ns	    157875 p99-ns
python rule-cascade 1.0.0a2 CPython 3.14.6: 5652 evals/s, p50 159.3 us, p99 428.3 us, 28259 evaluations in 5.0 s
rule server load test: node 22.22.2, 8 CPUs, 32 connections on 2 client threads, 2 s warm-up, 5 s measured
RULE_SERVER_WORKERS=0: 8166 req/s, p50 3.50 ms, p99 9.51 ms, 40832 requests, 0 not 200
RULE_SERVER_WORKERS=4: 7808 req/s, p50 3.06 ms, p99 23.10 ms, 39039 requests, 0 not 200

Run 3 (load average 11 to 15):

typescript rule-cascade 1.0.0-alpha.2 node 22.22.2: 26890 evals/s, p50 32.8 us, p99 116.5 us, 134448 evaluations in 5.0 s
java rule-cascade 1.0.0-alpha.2 jdk 17.0.19+0: 118563 evals/s, p50 8.0 us, p99 13.9 us, 592814 evaluations in 5.0 s
BenchmarkEvaluateTransfer-8   	  254622	     23028 ns/op	     43425 evals/s	     20541 p50-ns	     77375 p99-ns
python rule-cascade 1.0.0a2 CPython 3.14.6: 8394 evals/s, p50 117.8 us, p99 191.5 us, 41973 evaluations in 5.0 s
rule server load test: node 22.22.2, 8 CPUs, 32 connections on 2 client threads, 2 s warm-up, 5 s measured
RULE_SERVER_WORKERS=0: 10041 req/s, p50 3.04 ms, p99 6.44 ms, 50203 requests, 0 not 200
RULE_SERVER_WORKERS=4: 11729 req/s, p50 2.55 ms, p99 5.82 ms, 58646 requests, 0 not 200

An earlier single run of the Java benchmark on the same machine, a little less loaded, printed 124989 evals/s, p50 7.7 us, p99 11.6 us: the spread between runs is as large as the differences you might want to read into them.

What the numbers say

  • Evaluating one request takes microseconds in every runtime: about 8 to 11 µs at the median in Java, 20 to 26 µs in Go, 25 to 33 µs in TypeScript and 118 to 160 µs in Python, for this ruleset (14 rules, decimal arithmetic, message rendering). A service that embeds a runtime adds no noticeable latency to a request.
  • The rule server spends its time on HTTP, not on rules. At about 8 000 to 12 000 requests per second, the median request takes about 2.5 to 3.5 ms end to end, of which the evaluation is some 30 µs. The rest is the HTTP stack, JSON parsing and serialisation, and the one log line per request.
  • The worker pool made no consistent difference to throughput here. In runs 1 and 2, RULE_SERVER_WORKERS=4 answered fewer or about as many requests per second as 0, with a higher p99; in run 3 it answered more, with a lower p99. The difference between the settings is smaller than the spread between runs on this machine, so these runs do not show that workers help or hurt. What is certain is the cost they add: a request crosses a thread boundary twice (structured clone of the request and of the result), and the evaluation it moves off the event loop is cheap. The pool exists to protect the event loop: an evaluation that runs too long (a pathological input, a slow custom operator) costs one worker for RULE_EVALUATION_TIMEOUT_MS and is answered 503, instead of stalling every other request. Enable it when your rulesets have evaluations that are slow relative to the HTTP work, or when the protection matters more than a few percent of throughput. To scale throughput, add replicas: the server is stateless.

Worker pool details

VariableMeaningDefault
RULE_SERVER_WORKERSWorker threads for evaluations; 0 evaluates on the event loop0
RULE_EVALUATION_TIMEOUT_MSWith workers: the longest an evaluation may run before it is answered 503 and its worker replaced; 0 for no limit5000
  • Each worker holds the current snapshot, read from the bundles of the rulesets being served. Before a worker is given a request it receives the store's current snapshot if it holds an older one, and the request's ruleset is resolved against that same snapshot, so a reload never puts a request and its rules out of step. Results are byte-for-byte the same as in process; the server tests compare them.
  • The timeout is measured from the moment a worker takes the request, not from its arrival. Waiting requests are bounded by count (256 per worker); beyond that the answer is 503 with Retry-After: 1.
  • Workers that die before answering anything (a broken installation) are not replaced in a loop: after three in a row the pool gives up, logs a fatal line, and evaluations are answered 503.
  • Custom operators cannot cross into a worker thread. An embedding host passes operatorsModule, an ES module that exports them, and still passes operators so that missingOperators() can check them. The program rule-cascade-server registers no custom operators.
  • An evaluation on the event loop (RULE_SERVER_WORKERS=0) cannot be interrupted; the timeout has no effect there and the server says so in a warning at start-up.

On this page