An example of what you build on specsolve. A capacity-expansion planner: the model is 69 lines of YAML, the job that solves and archives it 614 lines of Python, this site 1223 lines of Markdown and SQL over the parquet the job wrote, and the notebook 342 lines. All of it is in the source; nothing else is behind it.

Reading the archive

Every other page here is a client. So is a DuckDB shell, a notebook, and a BI tool pointed at the directory. None of them is privileged, because the contract is the directory, not a library — and this page is where you learn to read it.

Every result below is a real query, run by DuckDB in your browser against the same parquet the dashboard reads. The SQL is above each one, and you can change the last one.

What the solve job wrote

One archive per scenario. This is base, by shape rather than by file — every quantity is one file, every period in it:

runs/base/
├── spec.yaml                  3.0 kB   the spec, as solved
├── catalog.parquet            4.1 kB   every name, its kind, path and dimensions
├── sources.parquet            1.4 kB   (specsolve_run, source, digest)
├── sources/                 11 files   every input, as solved
└── answer/
    ├── record.parquet         5.2 kB   one row per period: status, objective
    ├── metrics.parquet        4.6 kB   one row per period: size, seconds
    ├── primal/               3 files   build, p, total
    ├── dual/                 4 files   accumulate, balance, capacity, carbon
    └── expression/           3 files   capex, emissions, opex

Three kinds of thing are in there. spec.yaml and sources/ are what was solved — the spec and every input, so the run reproduces. answer/ is what came back: primal/ per variable, dual/ per constraint, expression/ per named quantity, one file each. record.parquet, metrics.parquet and sources.parquet are the record: one row per period saying how it terminated, what it cost to build and solve, and what each input's bytes digest to.

What a query gets back

A value frame carries the model's own dimensions, a value, and the run it came from. Nothing else, and no index:

Those column names — year, generator — are the model's, not this repository's. They come from the spec that was solved, which is why a reader who has never seen the model can still group by generator, and why two quantities keyed the same way join without a mapping table.

Two rules, one query each

Every file carries specsolve_run, the archive's directory name, so any file concatenates across archives with a single glob and needs no path parsing. The site's loader calls it run, and reads the period off the record's slice:

The archive carries its own catalogue. catalog.parquet lists every quantity, its kind, its path and the dimensions that key it, written by the solve. Nothing is declared twice:

Run one yourself

The tables above are registered; edit the query and it re-runs. objective, metrics, digests, total and emissions are in scope.

The same thing, in your own tools

Neither of these imports specsolve, and neither imports this repository's warehouse.py. Both answer the four numbers the pathway page leads with, and tests/test_clients.py holds them to each other on every archive in the directory — so a drift between them fails CI rather than reaching this page.

Ten lines of polars — uv run python clients/headline.py runs/base
"""The four headline numbers off an archive, in polars, with no specsolve and no client library.

    uv run python clients/headline.py runs/base

The archive is a table per glob, so reading it needs nothing this repository
ships. `headline.sql` answers the same four questions in DuckDB, and
`tests/test_clients.py` holds the two to each other.
"""

import sys
from pathlib import Path

import polars as pl


def headline(run: Path) -> dict[str, float]:
    frames = lambda name: pl.read_parquet(run / 'answer' / f'{name}.parquet')
    at = lambda frame, year: frame.filter(pl.col('year') == year)['value'].sum()

    objective = pl.read_parquet(run / 'answer/record.parquet')
    first, last = objective['slice'].cast(pl.Int64).min(), objective['slice'].cast(pl.Int64).max()
    emissions, fleet = frames('expression/emissions'), at(frames('primal/total'), last)
    clean = pl.read_parquet(run / 'sources/rate.parquet').filter(pl.col('value') == 0)['generator'].to_list()

    return {
        'pathway_cost': objective['objective'].sum(),
        'emissions_cut': 1 - at(emissions, last) / at(emissions, first),
        'zero_carbon_share': frames('primal/total')
        .filter(pl.col('year') == last, pl.col('generator').is_in(clean))['value']
        .sum()
        / fleet,
        'carbon_price': -at(frames('dual/carbon'), last) + 0.0,
    }


if __name__ == '__main__':
    for name, value in headline(Path(sys.argv[1] if len(sys.argv) > 1 else 'runs/base')).items():
        print(f'{name:<18} {value:>15,.4f}')
One DuckDB query, every scenario at once — duckdb -c ".read clients/headline.sql"

runs/ is written by the solve job rather than checked in, and the globs are relative, so both of these matter:

uv run showcase-solve --runs runs     # once, if runs/ is not there yet
duckdb -c ".read clients/headline.sql"

Get either wrong and the query says which one to run, rather than reporting a path that does not exist.