Cache reference

Cache reference

Fetching bars from a data adapter – which may be reading a multi-terabyte market-data lake and constructing bars from ticks – is the dominant cost of a backtest, and it is repeated identically on every run. The series cache stores fetch results on disk (default ~/algolang/cache) so a repeat run asks the adapter for no bars. The adapter still starts, and coverage and instrument-info requests stay live (below); on marketfeed a futures run’s first instrument-info request scans the root’s definition files, about 43 s for ES. It is a cross-process cache: every workspace and every algo process on the machine shares one directory safely.

The cache is invisible in results by construction: a cache-off, cache-cold, and cache-warm run of the same configuration produce byte-identical fills, reports, manifests, and audit records. Only the BENCH timing/byte diagnostics and the algo: cache: ... stderr notes differ. Any corrupt, truncated, foreign, or version-skewed cache file degrades to a miss and a live fetch; no run ever fails because of cache content.

One level

The cache has one level since GLE-324. Every FetchBars call the engine makes (the main series, the warmup pre-roll, strategy-declared series) and a continuous-futures run’s token-pinned composite fetch are served through it. An entry stores the raw bars exactly as the adapter returned them, keyed by everything that determines their content:

  • adapter identity: binary base name plus the adapter’s own init-ack name, provider, and description (adapters that embed a build id in the description self-invalidate when rebuilt), and the full adapter input map (which carries the lake directory);
  • the NEGOTIATED wire protocol (pb/dbn/csv decode different field sets);
  • the wire symbol exactly as requested (equity adjustment qualifiers such as :adj=splitdiv ride the symbol string), and for an adjusted equity (adj=split|splitdiv) the adjustment anchor the data client stamps on its requests, the run end (GLE-328; the same anchor tail as the composite fetch below, without a token), since a different anchor is a different series;
  • the complete bar spec: kind, interval, thresholds, reversal multiplier, threshold mode, EWMA spans, init parameters, include-path, the named path-interval (omitted when empty, so older entries keep their hashes), close-on-boundary;
  • the session field: the trading-day container (“eth”/“rth”/“session”; empty = epoch bars), or an intraday grid boundary (“utc-day”/“eth”/“rth”/“session” with an intraday interval), threaded from the run’s per-series session by the engine – an RTH re-run can never alias a warm ETH entry;
  • for an intraday grid request, the calendar revision the capability probe read (a length-prefixed GLE207-v1 tail, present only when set), so a calendar change cannot be served from a stale entry;
  • for a continuous run’s token-pinned composite fetch, the served snapshot token and the explicit adjustment anchor as UTC RFC3339Nano (a length-prefixed GLE324-v1 tail, present only when set, so every earlier hash stands), so an entry serves one served state only; a composite fetch that pins no token (a clamped or re-anchored end) is not cached;
  • the exact request range (UnixNano start/end, encoded in the filename);
  • the on-disk format version.

Excluded, deliberately: batch size (chunking only), replication and series-id tagging (applied downstream of the seam), and equity corporate-action accounting (the events fire downstream on the raw books; the adjusted view itself is the feed’s, keyed by the symbol qualifier and the anchor above). Corporate-action, trade, coverage, and instrument-info requests are never cached – they stay live on every run.

No Level 2. Earlier builds kept a second level for a continuous run’s finished build product (the adjusted series, manifest, boundaries and segments the engine then stitched itself). Since GLE-323 the engine builds nothing: the adjusted bars, the per-span constants and the seam pairs come from the wire, and the composite fetch is the whole cost, which Level 1 covers under the snapshot token. GLE-324 deleted the Level-2 code. A continuous/ directory of .cont files left by an older cache is not read and not listed; a pattern-less algo cache flush removes it and counts its files as removed entries, while a patterned flush leaves it alone. Schedule reads stay live on every run, like corporate-action, trade, coverage and instrument-info requests, so the feed’s degradation flags and a new generation are seen even when the bars are a cache hit.

Threat model, stated plainly: this defends against corruption (disk damage, truncation, torn writes, bit rot, version skew) and against accidental or partial rewrites. It does not defend against a local actor who can run code as the user and forge a complete, internally consistent, correctly checksummed entry – such an actor could equally replace the engine binary itself. The cache directory carries the same trust as the binaries that read it.

The subset optimization

A request whose non-range key fields match an existing entry, whose END is exactly equal, and whose start is at or after the entry’s start is served by slicing the entry (keep bars with ts_close >= start); no new entry is written. Running 2015-2024 after having run 2010-2024 reads the cache, not the lake, and leaves the directory unchanged.

This applies to time bars only. Time-bar bucket boundaries are epoch-anchored and adapters follow the canonical membership rule (start <= ts_close < end) with full boundary-bucket aggregation, so a slice is bit-identical to a fresh narrower fetch. Every activity kind (volume/tick/dollar/imbalance/runs/renko) accumulates construction state from the fetch-window start, so a slice would NOT equal a cold fetch; those kinds are exact-range only.

When the cache engages

All of the following must hold, else the run behaves as cache-off:

  • the run is a backtest (never --mode live);
  • the requested end does not extend past 00:00 UTC of the current day (the window must be fully historical; a config date-to: yesterday, which expands to an exclusive end of exactly today 00:00 UTC, IS eligible). The algo run default --end 2100-01-01 therefore never caches – pass an explicit historical --end to benefit; the run prints why caching was skipped;
  • the cache mode is on (the algo run default) or refresh. Library callers and tests get a zero-value RunConfig, which is cache-off; the golden-test harness passes --cache explicitly (GOLDEN_CACHE, on by default).

--mode live with an explicit --cache on|refresh is an error; with the default it silently disables. TRADES-schema (data-plane benchmark) runs never cache.

Invalidation is manual

There is no TTL and no source-freshness check in v1 (the entry header reserves a freshness field for a future source fingerprint). The lake is assumed immutable for already-historical ranges; when that assumption breaks (a vendor revises history, the lake is rebuilt), the levers are:

  • algo cache flush [pattern...] – remove matching entries (filename or symbol globs); no pattern removes everything;
  • --cache refresh – bypass all reads for one run and overwrite its entries;
  • unlike the legacy cache, a write always OVERWRITES an existing entry, so a refresh genuinely replaces stale data.

Accepted consequences of this model, documented so nobody rediscovers them:

  • an entry ends at its recorded end forever; a run ending later is a miss and refetches (day-granularity freshness comes free from the historical-end gate);
  • --cache refresh rewrites only what the run fetches; an overlapping WIDER stale entry from an earlier run remains and may still subset-serve future requests – flush the symbol when the lake changes materially;
  • if two same-end supersets disagree (one predates a lake revision) and the newer is corrupt or mid-flush, the subset scan falls through to the older valid one rather than going live. Within the manual-invalidation contract this is by design; flush when in doubt;
  • the cached upper bound is trusted: bars are not re-filtered against end on read (the entry’s end equals the request’s by key).

Concurrency

The directory is safe under concurrent runs, including across workspaces: writes go to a unique dot-prefixed temp file in the target directory (fsync, then atomic rename over the final name – an overwrite never tears a concurrent reader, which keeps its open inode); reads take a shared flock; publishing renames hold a shared directory lock while flush holds it exclusively around its check-and-unlink, so a flush can never delete an entry that was just replaced under it. Two processes missing the same key both fetch and both write; last rename wins with identical content. Decode failures never delete files (the “corrupt” file may be a rename mid-flight).

Maintenance CLI

algo cache                 # stats: directory, entry count, total size
algo cache list [glob...]  # one line per entry (level, symbol, kind/interval,
                           # range, size, created)
algo cache flush [glob...] # remove matching entries (no pattern: everything)
algo cache --cache-dir DIR # operate on a non-default directory

ALGO_CACHE=off|on|refresh and ALGO_CACHE_DIR=... act as environment fallbacks wherever the corresponding flag was not given explicitly (the CI kill switch).

On-disk format (for tooling authors)

<dir>/bars/<symbol-prefix>_<keyhash>_<startns>_<endns>.bars. Every entry is self-describing: an ALGOSC1 magic, a length-prefixed cleartext JSON header carrying the format version, the level (1), the FULL key, the range, and record counts, then the payload (length-delimited protobuf bars). A <dir>/continuous/ directory of .cont files is a legacy Level-2 cache from before GLE-324, which this build does not read. The prefix is sanitized and cosmetic – identity lives in the hash plus the header, and readers re-validate the header key field-for-field, so hash collisions or renamed files can mislabel nothing. Filenames parse from the right (the prefix may itself contain underscores).

Session containers and intraday filtering

TWO distinct things have been called “sessions”; do not conflate them.

The trading-day bar container (implemented). The run-config key session (and --session) selects which trading-day container serves a 1d/1w series: “eth”, “rth”, or “session” (the symbol’s declared alias). The session rides GetDataRequest.session and the adapter serves session-keyed daily bars (session bars). Because bar CONTENT genuinely differs per container, the session is part of the cache key – the adapter-side branch of the design fork below, decided. No cache version moved: every historical entry encoded Session "", so epoch requests stay warm and session requests are guaranteed misses.

Intraday grids (implemented). A boundary session with an intraday interval is a different series again (intraday grids): the session and interval already key it, and the grid revision joins the key as a length-prefixed GLE207-v1 tail only when set, so every historical hash is unchanged and a new calendar revision cannot be served from a stale entry.

Intraday session-window filtering (NOT built). The old Algolang could restrict each day’s data to an hours/timezone window and restrict which days delivered data. algolang does not have this: the run-config keys session-hours / session-days are accepted but ignored with an explicit warning (engine/runconfig points at the distinct trading-day container). If it is ever built engine-side post-fetch, ITS window would stay out of the cache key (one cached entry serves every window swept); adapter-side windowing would need its own key axis. That fork remains open for the filtering feature only.

Deliberate non-goals in v1

Source-freshness fingerprinting (field reserved), TTL, eviction/size budgets, an index file (headers are the truth; list scans them), cross-process single-flight on simultaneous misses, caching of corporate actions or raw trades, and any in-memory sharing tier (in-process refcounted deduplication remains open and complementary).