Performance
How fast each runtime evaluates, how the rule server behaves under load, and how to measure it on your own hardware.
How fast each runtime evaluates, how the rule server behaves under load, and how to measure it on your own hardware. The numbers below were measured on one laptop; they are a record of what one run printed, not a guarantee, and they are not checked by CI (benchmarks on shared CI runners are too noisy to gate on).
Running the benchmarks
make bench # every runtime, then the rule server; BENCH_SECONDS=5 by default
make bench-typescript bench-java bench-go bench-python bench-server # one at a time
BENCH_SECONDS=20 make bench-goEvery runtime benchmark does the same thing:
- reads
conformance/bundles/acme.payments.transfer.bundle.json(conformance/bundles/acme.payments.transfer.bundle.json) with itsfromBundle(no compiling from YAML); - evaluates the requests of
tools/bench-requests.json(tools/bench-requests.json) round-robin on the server channel, on one thread: an allowed domestic transfer, a denied one with a rendered message, an international one with three findings and an effect, and one narrowed to a component; - warms up for
BENCH_WARMUPseconds (default 2), then measures forBENCH_SECONDS; - times every evaluation with the monotonic clock (so each latency includes one clock read) and reports evaluations per second (count / measured time) and the p50 and p99 latency (nearest rank).
| Runtime | Harness |
|---|---|
| TypeScript | packages/typescript/bench/evaluate.mjs (packages/typescript/bench/evaluate.mjs), on dist/ |
| Java | Benchmark.java (packages/java/src/test/java/io/github/yarlisaisolutions/rulecascade/Benchmark.java), a plain main (no JMH), compiled with javac after make java |
| Go | BenchmarkEvaluateTransfer in bench_test.go (packages/go/bench_test.go), go test -bench with p50-ns, p99-ns and evals/s metrics |
| Python | packages/python/bench/evaluate.py (packages/python/bench/evaluate.py), timeit.default_timer per evaluation |
The rule server load test, packages/server/bench/load.mjs (packages/server/bench/load.mjs),
starts dist/main.js with the token on and examples/contracts as the rules, once with
RULE_SERVER_WORKERS=0 and once with a worker pool, one after the other, and posts the same
requests to /evaluations over keep-alive connections from load-generating worker threads. It needs
no dependency. BENCH_CONNECTIONS (32), BENCH_CLIENT_THREADS (2) and BENCH_WORKERS ("0 N")
change the setup. The generator runs on the same machine, so the result compares the two settings;
it is not a capacity figure for a server on its own hardware.
Measured on
Apple M2, 8 cores (4 performance, 4 efficiency), 24 GB, macOS 27.0.1 (arm64); Node.js 22.22.2, OpenJDK 17.0.19 (Homebrew), Go 1.26.5, CPython 3.14.6; Rule Cascade 1.0.0-alpha.2; 3 October 2026.
The machine was not idle: other work kept the load average between 17 and 22 during both runs.
Expect quieter numbers, and smaller differences between runs, on an idle machine. Two consecutive
runs of make bench with BENCH_SECONDS=5, and a third later the same day with the load average
between 11 and 15, verbatim:
Run 1:
typescript rule-cascade 1.0.0-alpha.2 node 22.22.2: 35862 evals/s, p50 25.0 us, p99 83.7 us, 179309 evaluations in 5.0 s
java rule-cascade 1.0.0-alpha.2 jdk 17.0.19+0: 78054 evals/s, p50 8.8 us, p99 43.5 us, 390271 evaluations in 5.0 s
BenchmarkEvaluateTransfer-8 320043 22982 ns/op 43513 evals/s 19750 p50-ns 69834 p99-ns
python rule-cascade 1.0.0a2 CPython 3.14.6: 6126 evals/s, p50 134.8 us, p99 492.3 us, 30630 evaluations in 5.0 s
rule server load test: node 22.22.2, 8 CPUs, 32 connections on 2 client threads, 2 s warm-up, 5 s measured
RULE_SERVER_WORKERS=0: 10186 req/s, p50 2.79 ms, p99 8.92 ms, 50930 requests, 0 not 200
RULE_SERVER_WORKERS=4: 5072 req/s, p50 4.03 ms, p99 39.82 ms, 25358 requests, 0 not 200Run 2:
typescript rule-cascade 1.0.0-alpha.2 node 22.22.2: 27997 evals/s, p50 32.2 us, p99 112.6 us, 139987 evaluations in 5.0 s
java rule-cascade 1.0.0-alpha.2 jdk 17.0.19+0: 87202 evals/s, p50 10.5 us, p99 29.0 us, 436012 evaluations in 5.0 s
BenchmarkEvaluateTransfer-8 238772 32945 ns/op 30353 evals/s 26041 p50-ns 157875 p99-ns
python rule-cascade 1.0.0a2 CPython 3.14.6: 5652 evals/s, p50 159.3 us, p99 428.3 us, 28259 evaluations in 5.0 s
rule server load test: node 22.22.2, 8 CPUs, 32 connections on 2 client threads, 2 s warm-up, 5 s measured
RULE_SERVER_WORKERS=0: 8166 req/s, p50 3.50 ms, p99 9.51 ms, 40832 requests, 0 not 200
RULE_SERVER_WORKERS=4: 7808 req/s, p50 3.06 ms, p99 23.10 ms, 39039 requests, 0 not 200Run 3 (load average 11 to 15):
typescript rule-cascade 1.0.0-alpha.2 node 22.22.2: 26890 evals/s, p50 32.8 us, p99 116.5 us, 134448 evaluations in 5.0 s
java rule-cascade 1.0.0-alpha.2 jdk 17.0.19+0: 118563 evals/s, p50 8.0 us, p99 13.9 us, 592814 evaluations in 5.0 s
BenchmarkEvaluateTransfer-8 254622 23028 ns/op 43425 evals/s 20541 p50-ns 77375 p99-ns
python rule-cascade 1.0.0a2 CPython 3.14.6: 8394 evals/s, p50 117.8 us, p99 191.5 us, 41973 evaluations in 5.0 s
rule server load test: node 22.22.2, 8 CPUs, 32 connections on 2 client threads, 2 s warm-up, 5 s measured
RULE_SERVER_WORKERS=0: 10041 req/s, p50 3.04 ms, p99 6.44 ms, 50203 requests, 0 not 200
RULE_SERVER_WORKERS=4: 11729 req/s, p50 2.55 ms, p99 5.82 ms, 58646 requests, 0 not 200An earlier single run of the Java benchmark on the same machine, a little less loaded, printed
124989 evals/s, p50 7.7 us, p99 11.6 us: the spread between runs is as large as the differences
you might want to read into them.
What the numbers say
- Evaluating one request takes microseconds in every runtime: about 8 to 11 µs at the median in Java, 20 to 26 µs in Go, 25 to 33 µs in TypeScript and 118 to 160 µs in Python, for this ruleset (14 rules, decimal arithmetic, message rendering). A service that embeds a runtime adds no noticeable latency to a request.
- The rule server spends its time on HTTP, not on rules. At about 8 000 to 12 000 requests per second, the median request takes about 2.5 to 3.5 ms end to end, of which the evaluation is some 30 µs. The rest is the HTTP stack, JSON parsing and serialisation, and the one log line per request.
- The worker pool made no consistent difference to throughput here. In runs 1 and 2,
RULE_SERVER_WORKERS=4answered fewer or about as many requests per second as0, with a higher p99; in run 3 it answered more, with a lower p99. The difference between the settings is smaller than the spread between runs on this machine, so these runs do not show that workers help or hurt. What is certain is the cost they add: a request crosses a thread boundary twice (structured clone of the request and of the result), and the evaluation it moves off the event loop is cheap. The pool exists to protect the event loop: an evaluation that runs too long (a pathological input, a slow custom operator) costs one worker forRULE_EVALUATION_TIMEOUT_MSand is answered503, instead of stalling every other request. Enable it when your rulesets have evaluations that are slow relative to the HTTP work, or when the protection matters more than a few percent of throughput. To scale throughput, add replicas: the server is stateless.
Worker pool details
| Variable | Meaning | Default |
|---|---|---|
RULE_SERVER_WORKERS | Worker threads for evaluations; 0 evaluates on the event loop | 0 |
RULE_EVALUATION_TIMEOUT_MS | With workers: the longest an evaluation may run before it is answered 503 and its worker replaced; 0 for no limit | 5000 |
- Each worker holds the current snapshot, read from the bundles of the rulesets being served. Before a worker is given a request it receives the store's current snapshot if it holds an older one, and the request's ruleset is resolved against that same snapshot, so a reload never puts a request and its rules out of step. Results are byte-for-byte the same as in process; the server tests compare them.
- The timeout is measured from the moment a worker takes the request, not from its arrival. Waiting
requests are bounded by count (256 per worker); beyond that the answer is
503withRetry-After: 1. - Workers that die before answering anything (a broken installation) are not replaced in a loop:
after three in a row the pool gives up, logs a
fatalline, and evaluations are answered503. - Custom operators cannot cross into a worker thread. An embedding host passes
operatorsModule, an ES module that exports them, and still passesoperatorsso thatmissingOperators()can check them. The programrule-cascade-serverregisters no custom operators. - An evaluation on the event loop (
RULE_SERVER_WORKERS=0) cannot be interrupted; the timeout has no effect there and the server says so in a warning at start-up.
Caching and refreshing rules
A ruleset version is immutable and identified by its checksum, so anything derived from it (a bundle, a manifest) can be cached for as long as you like. What changes is which version is current.
@yarlisaisolutions/rule-cascade
The Rule Cascade runtime for browsers and Node.js. It implements both conformance levels of the specification: it evaluates bundles and manifests (evaluator) and it loads source documents and produces bundles (compiler).