Pricing a data request

Pricing a data request

Databento bills historical data per uncompressed DBN byte, but the metadata endpoints that quote a request are free — so every pull gets priced before any batch job is submitted. algodata cost walks a universe x dataset x schema matrix through metadata.get_dataset_range, metadata.list_unit_prices, metadata.get_billable_size, and metadata.get_cost, and prints the invoice: billable size, $/GB rate, and the exact dollar figure (which respects flat-rate plan discounts) per cell, with per-dataset subtotals and a grand total. It never requests time-series data, so running it costs nothing.

export DATABENTO_API_KEY=db-...   # metadata endpoints are free but authenticated

# Full matrix from a universe file, plus a machine-readable copy:
bin/algodata cost -universe scripts/universe.example.json -csv invoice.csv

# Ad-hoc single query:
bin/algodata cost -dataset GLBX.MDP3 -symbols ES.FUT,NQ.FUT -schemas tbbo,mbo

Flags (all optional unless noted):

FlagDefaultMeaning
-universe FILE—universe JSON (mutually exclusive with -dataset/-symbols)
-dataset, -symbols—ad-hoc mode: one dataset + comma-separated symbols
-schemastbbo,mbo,definition,statistics,statusschemas to price
-stypeparentstype_in symbology (parent, continuous, raw_symbol, instrument_id)
-start, -endfull availabilityISO 8601 range (end exclusive); overrides universe-file entries
-modehistoricalwhich $/GB column to display; the cost itself is mode-independent (batch and streaming bill identically)
-key$DATABENTO_API_KEYAPI key (Basic auth, key as username)

The universe file (see scripts/universe.example.json — its symbols are placeholders, not the production 63) is per-dataset symbol lists with optional overrides; precedence is dataset entry > explicit flag > file-level default > flag default:

{
  "stype_in": "parent",
  "schemas": ["tbbo", "mbo", "definition", "statistics", "status"],
  "datasets": [
    {"dataset": "GLBX.MDP3", "symbols": ["ES.FUT", "NQ.FUT"]},
    {"dataset": "XEUR.EOBI", "symbols": ["FGBL.FUT"],
     "schemas": ["tbbo"], "start": "2025-03-10", "notes": "free-text"}
  ]
}

Semantics worth knowing:

  • Ranges clamp per schema. get_dataset_range reports per-schema availability, and each cell is clamped to it (noted in the table) — e.g. GLBX.MDP3 mbo begins 2017-05-21 inside a dataset range that reaches back to 2010, because CME’s pre-2017 feed was the level-aggregated MDP 2.0. An empty intersection prices as a skip, not an error.
  • Quotes are exact for whole-day ranges (the API’s stated accuracy is 10-minute multiples; 24 h multiples for definition).
  • Parent symbology (ES.FUT) covers every expiry plus calendar spreads: book schemas quote 1.5–3x a front-month-only pull because back months are heavily quoted.
  • Datasets are priced concurrently but paced at 10 requests/second (half the documented per-IP limit), with one Retry-After-honoring retry on HTTP 429. Up to 2,000 symbols per dataset entry (the API’s per-request cap).
  • A failing cell becomes an ERROR: row and the run continues; the exit status is non-zero if any cell failed.