Caching and refreshing rules
A ruleset version is immutable and identified by its checksum, so anything derived from it (a bundle, a manifest) can be cached for as long as you like. What changes is which version is current.
A ruleset version is immutable and identified by its checksum, so anything derived from it (a bundle, a manifest) can be cached for as long as you like. What changes is which version is current. This page covers the three places that decide when to look for a new one:
| Where | Decides | Configured by |
|---|---|---|
| The rule server | When to reload its rules directory; how HTTP caches may keep client manifests | Environment variables; cache-policy.yaml in the rules directory |
| The TypeScript manifest client | When a browser or Node.js process asks for a manifest again | Options of createManifestClient |
| An embedded runtime (Java, Go, Python) | When a service loads its bundle again | A refreshing holder: RuleSetHolder, Holder |
None of this is part of the ruleset specification: cache policy is a deployment choice, and the same bundle can be served under different policies.
Choosing a mode
| Mode | A cached copy is used | Freshness | Use it when |
|---|---|---|---|
revalidate (default) | After the server confirms it (304) | Every request sees the current rules | You want the simplest correct behaviour; the revalidation is cheap |
ttl | For a fixed time without asking | Up to the TTL behind | Many clients, rules that change rarely, and a delay of minutes is acceptable |
permanent | For ever, by checksum | Behind until something says a new checksum exists | A CDN in front of the server, or clients that learn the checksum from elsewhere (an evaluation result, GET /rulesets, a deploy notification) |
The server still evaluates every state-changing operation against its own current rules, so a client that holds an older manifest gives older feedback, never a wrong decision.
The rule server
Scheduled refresh
| Variable | Meaning | Default |
|---|---|---|
RULES_REFRESH | manual: reload on SIGHUP only. interval or cron: also check the rules directory on a schedule | manual |
RULES_REFRESH_INTERVAL | With interval: 30s, 5m, 1h, 1500ms or a number of seconds. At least 1s | none |
RULES_REFRESH_CRON | With cron: five fields, see Cron syntax | none |
RULES_REFRESH_TZ | IANA time zone of the cron schedule | UTC |
A scheduled check computes a SHA-256 over the names and contents of every file under RULES_DIR
and reloads only when it differs from the last load attempted. Entries whose names start with a
dot are skipped, which leaves out the ..data directories of a Kubernetes ConfigMap mount.
Each reload has the semantics of SIGHUP: a complete new snapshot is built and swapped in one
step, and if anything fails to load the previous rules keep being served.
- Reloads and refused reloads are logged, with
triggerset toSIGHUP,intervalorcron. A check that finds nothing changed is not logged. - Content that was refused is not tried again until it changes, so a bad file is reported once.
SIGHUPalways reloads, whatever the schedule.- The check reads every file on the event loop, as
SIGHUPdoes. For a large directory use an interval of 10 seconds or more. - The timers do not keep the process alive and are stopped when
SIGTERMstarts the drain. - An invalid setting stops the process at start-up with a
fatallog line and exit status 1.
RULES_REFRESH=interval RULES_REFRESH_INTERVAL=30s ... # follow a mounted ConfigMap
RULES_REFRESH=cron RULES_REFRESH_CRON='0 6 * * 1-5' RULES_REFRESH_TZ=Europe/Paris ... # 06:00 on weekdaysHTTP cache policy
An optional cache-policy.yaml (or cache-policy.yml, cache-policy.json; only one of them) in
RULES_DIR sets how client manifests may be cached. It is loaded and checked with the rulesets: a
bad policy file stops the start, and a reload that finds one is refused like a bad ruleset.
default:
mode: revalidate # the default when there is no file
rulesets:
acme.payments.transfer:
mode: permanent
acme.org.base:
mode: ttl
maxAge: 300 # seconds
staleWhileRevalidate: 60 # optional
staleIfError: 86400 # optionalmodeis required in every rule.maxAge(above 0),staleWhileRevalidateandstaleIfErrorare whole seconds up to one year and belong tottlrules only.- Every key under
rulesetsmust be a ruleset being served: an override for an unknown id is refused, because a typo would silently fall back to the default. - Unknown members are refused.
The Cache-Control of GET /rulesets/{id}/manifest?channel=client:
| Mode | Without ?checksum= | With the checksum being served |
|---|---|---|
revalidate | public, max-age=0, must-revalidate | the same |
ttl | public, max-age=N[, stale-while-revalidate=S][, stale-if-error=E] | the same |
permanent | public, max-age=0, must-revalidate | public, max-age=31536000, immutable |
The ETag is the same in every mode, and If-None-Match still answers 304.
?checksum= names the version the caller wants. It is an extension of this server, not part of
spec/v1/rule-evaluation.openapi.yaml. When it is not the checksum being served the answer is
409 Conflict with Cache-Control: no-store and a problem body whose checksum member is the one
being served. 409, not 404: RFC 9110 lets caches store a 404 without explicit freshness, and a CDN
that kept one would hide the version once it is deployed; a 409 is not cacheable by default, and
no-store says so explicitly. The check applies in every mode.
Server manifests, bundles and server-only rule details are always private, no-cache, whatever the
policy says.
Discovering a new version in permanent mode
- The client asks without a checksum (
revalidateheaders), receives the manifest and learns its checksum from it or from the ETag. - From then on it asks with
?checksum=<that checksum>. A CDN or the browser answers from its cache, forever. - It learns of a new version from an evaluation result (every result carries
checksum), fromGET /rulesets, from a scheduled unversioned request, or from a deployment notification. The old checksum URL keeps answering409once the server has moved on, so a stale pin is noticed.
Behind a CDN
- Cache only
GET /rulesets/{id}/manifest?channel=client(with or withoutchecksum). Use the full query string in the cache key. - Bypass the CDN for
channel=server,/bundle,/rulesets/{id}/rules/{rule}and/evaluations. They need the bearer token, and their answers depend on it. - Do not cache error responses beyond what their headers allow; in particular do not configure
negative caching for
409. - During a rolling deployment, replicas may serve different checksums for a moment. A pinned request
that reaches an old replica gets
409; the client falls back to the unversioned URL.
The TypeScript manifest client
import { createManifestClient, localStorageAdapter } from '@yarlisaisolutions/rule-cascade';
const manifests = createManifestClient({
baseUrl: '/api/rules',
mode: 'ttl', // 'revalidate' (default) | 'ttl' | 'permanent'
ttlMs: 5 * 60_000, // ttl: use the cached copy this long; also the background refresh period
staleIfError: true, // default: serve the cached copy when the server cannot be reached
refreshCron: '*/15 * * * *', // optional: refresh every held manifest on this schedule
refreshTimeZone: 'UTC',
storage: localStorageAdapter(), // optional: survive a page reload
onUpdate: (id, manifest) => render(id, manifest),
onStale: ({ rulesetId, error }) => console.warn(rulesetId, error.message),
});
const manifest = await manifests.get('acme.payments.transfer');
await manifests.get('acme.payments.transfer', { checksum: result.checksum }); // pin a known version
manifests.close(); // stop the background timers| Option | Default | Meaning |
|---|---|---|
mode | revalidate | See Choosing a mode |
ttlMs | 60000 | ttl: the age below which get makes no request; held manifests are refreshed in the background every ttlMs |
staleIfError | true | When the request fails, or the answer is 408, 429, 5xx or not a client manifest, get returns the cached copy and calls onStale. Any other 4xx is a refusal: get throws and the cached copy is dropped |
refreshCron, refreshTimeZone | none, UTC | Revalidate every held manifest on a cron schedule |
storage | none | { get(key), set(key, value) }, sync or async. localStorageAdapter(prefix?) wraps localStorage and does nothing where it is unavailable or full |
onUpdate | none | Called when a manifest replaces a cached copy with another checksum, from get or a background refresh |
- Concurrent
getcalls for the same ruleset share one request. - In
permanentmode a held manifest is returned without a request.get(id, { checksum })fetches the URL with that checksum when the held copy has another one, and falls back to the unversioned URL on409. A server that does not know?checksum=answers with its current manifest, which the client takes. refresh(id?)revalidates now;close()stops the timers. Timers do not keep Node.js alive.
Embedded runtimes
A service that embeds a runtime loads a bundle. A holder loads it again on a schedule and swaps it in atomically; when a load fails the last good rules stay in place and a callback hears about it. A ruleset with the checksum already held is not swapped in. The first load happens in the constructor and its failure stops the start.
Java (io.github.yarlisaisolutions.rulecascade.RuleSetHolder, one daemon thread):
RuleSetHolder rules = RuleSetHolder.builder(() -> RuleSet.fromBundle(readBundle()))
.every(Duration.ofMinutes(5)) // or .cron("0 * * * *", ZoneId.of("Europe/Paris"))
.onReload(r -> { if (!r.ok()) log.warn("rules not refreshed", r.error()); })
.start();
rules.get().evaluate(request);
rules.close();Go (rulecascade.Holder, one goroutine):
rules, err := rulecascade.NewHolder(loadBundle, rulecascade.HolderOptions{
Interval: 5 * time.Minute, // or Cron: "0 * * * *", Location: paris
OnReload: func(r rulecascade.Reload) {
if r.Err != nil {
log.Printf("rules not refreshed: %v", r.Err)
}
},
})
if err != nil {
log.Fatal(err)
}
defer rules.Close()
result, err := rules.Get().Evaluate(request, "server", nil)Python (rule_cascade.RuleSetHolder, a daemon threading.Timer):
from rule_cascade import RuleSetHolder
rules = RuleSetHolder(load_bundle, interval=300, on_reload=lambda r: r.error and log.warning(r.error))
# or RuleSetHolder(load_bundle, cron="0 * * * *", tz=ZoneInfo("Europe/Paris"))
rules.get().evaluate(request)
rules.close()Each holder also has refresh() to load now, and reports when the next scheduled refresh runs
(nextRun(), NextRun(), next_run).
Cron syntax
Five fields separated by spaces or tabs: minute hour day-of-month month day-of-week.
| Field | Values |
|---|---|
| minute | 0-59 |
| hour | 0-23 |
| day of month | 1-31 |
| month | 1-12 |
| day of week | 0-7; 0 and 7 are Sunday |
Each field is *, a number, a range a-b, a step */n or a-b/n, a/n (from a to the maximum),
or a comma-separated list of those. Names (JAN, MON) and macros (@daily) are not supported.
- When day of month and day of week are both restricted, a day matches when either matches
(
0 0 13 * 5is the 13th and every Friday). A field that begins with*,*/2included, is not restricted, and then both must match. - A schedule that can never fire, such as
0 0 30 2 *, is refused. - The next fire time is strictly after the current minute; seconds are ignored.
- In a zone with daylight saving time, a wall-clock time that does not exist (spring forward) is skipped, and one that happens twice (fall back) fires once, at the first occurrence.
The TypeScript, Java, Go and Python parsers are tested against one table,
tools/cron-cases.json (tools/cron-cases.json).
| Expression | Fires |
|---|---|
*/5 * * * * | Every five minutes |
0 * * * * | On the hour |
30 2 * * * | 02:30 every day |
0 6 * * 1-5 | 06:00 Monday to Friday |
0 0 1 * * | Midnight on the first of the month |