diff --git a/README.md b/README.md index 20f0b9e..85043b5 100644 --- a/README.md +++ b/README.md @@ -1,38 +1,59 @@ # Lightcone Research Stack documentation -[![Docs](https://img.shields.io/github/actions/workflow/status/LightconeResearch/docs/docs.yml?branch=main&style=flat&label=docs&color=darkgreen)](https://docs.lightconeresearch.org) -[![License](https://img.shields.io/badge/License-BSD_3--Clause-426b78.svg?style=flat)](LICENSE) +User guides and reference material for the [Lightcone Research Stack](https://lightconeresearch.org/): the `lc` execution tools, research agent skills, and their integration with the external [ASTRA specification](https://astra-spec.org/latest/). -The **Lightcone Research Stack** is [Lightcone Research](https://lightconeresearch.org/)'s -tooling for research analyses described with [ASTRA](https://astra-spec.org/latest/) -(Agentic Schema for Transparent Research Analysis). You describe an analysis in an -`astra.yaml` specification; the stack validates it and takes care of the rest — -execution, environments, and provenance. +**[Read the documentation →](https://docs.lightconeresearch.org/)** -**→ Read the documentation at ** +## Start here -## Where to start +- [Install](https://docs.lightconeresearch.org/user/install/) — one CLI setup, with an optional agent plugin. +- [Your first analysis](https://docs.lightconeresearch.org/user/getting-started/) — run a complete local example and inspect its provenance. +- [Work with an agent](https://docs.lightconeresearch.org/user/agents/) — scope, implement, and resume a research project. +- [Meet the stack](https://docs.lightconeresearch.org/user/) — understand how the tools fit together. -- [Install](https://docs.lightconeresearch.org/user/install/) — uv, git, and the `lc` command -- [Getting started](https://docs.lightconeresearch.org/user/getting-started/) — your first analysis, from `lc init` to a published result -- [Core concepts](https://docs.lightconeresearch.org/user/concepts/) — projects, output identity, and how provenance is recorded -- [Running on a cluster](https://docs.lightconeresearch.org/user/cluster/) — SLURM, containers on HPC, and parallel filesystems -- [Troubleshooting](https://docs.lightconeresearch.org/user/troubleshooting/) — common errors and how to fix them +These docs currently target the upcoming explicit compute workflow. The installation guide pins a source revision that includes `lc compute`; the published `0.5.0rc4` release uses an earlier workflow. Keep installation instructions, tutorials, and command reference aligned when moving to a new release. -## Components +## Build locally -| Component | What it does | Repository | -| --- | --- | --- | -| **lightcone-cli** | The `lc` CLI: project scaffolding, locked environments, sandboxed execution, and the provenance layer | [LightconeResearch/lightcone-cli](https://github.com/LightconeResearch/lightcone-cli) | -| **astra-tools** | The SDK and `astra` CLI for ASTRA specifications: schema, validation, and evidence verification helpers | [LightconeResearch/astra-tools](https://github.com/LightconeResearch/astra-tools) | +With [uv](https://docs.astral.sh/uv/getting-started/installation/) installed: -## Feedback +```bash +uv sync --locked +uv run zensical serve +``` -The stack is in early alpha, and bug reports, design challenges, and use cases it -doesn't cover yet are welcome. Report a problem with a tool on that tool's -repository; report a problem with the documentation itself — a page that is wrong, -unclear, or out of date — [here](https://github.com/LightconeResearch/docs/issues). +Before submitting a change: -## License +```bash +uv run zensical build --clean --strict +``` + +Pull requests run the same strict build. Updates to `main` are deployed through GitHub Pages by [the docs workflow](.github/workflows/docs.yml). + +## Where to edit + +| Location | Purpose | +| --- | --- | +| `docs/index.md` | Stack landing page | +| `docs/user/` | Installation, tutorials, and task-oriented guides | +| `docs/cli/` | CLI command reference | +| `docs/api/`, `docs/architecture.md` | CLI implementation reference | +| `docs/contributing/`, `docs/maintainer.md` | Contributor guidance | +| `zensical.toml` | Navigation and site configuration | +| `docs/stylesheets/extra.css`, `overrides/` | Website-aligned typography, colors, and layout | + +The design follows [lightcone-website](https://github.com/LightconeResearch/lightcone-website): Quattrocento headings, Newsreader prose, Alegreya navigation, JetBrains Mono code, parchment surfaces, and antique-gold accents. The landing-page engraving is the same *Uranometria* (Bayer, 1603) asset used by the website. + +Check technical claims against the relevant source: + +- [lightcone-cli](https://github.com/LightconeResearch/lightcone-cli) — execution and provenance. +- [agent-skills](https://github.com/LightconeResearch/agent-skills) — plugin installation, skills, and hooks. +- [ASTRA documentation](https://astra-spec.org/latest/) and [astra-tools](https://github.com/LightconeResearch/astra-tools) — the external specification and validation tools. + +Preserve existing page URLs when reorganizing navigation. Keep advanced implementation details in the reference and contributor sections, and put a runnable path before optional setup. + +## Feedback and license + +Report documentation problems in [this repository's issues](https://github.com/LightconeResearch/docs/issues). Report tool behavior in the relevant tool's repository. The stack is in early alpha and feedback from real analyses is welcome. BSD 3-Clause — see [LICENSE](LICENSE). diff --git a/docs/api/assets.md b/docs/api/assets.md index b584bd9..d07dee7 100644 --- a/docs/api/assets.md +++ b/docs/api/assets.md @@ -15,7 +15,7 @@ Source: `src/lightcone/engine/assets.py`. | `classify(...)` | The one rule: `current` / `behind` / `stale`, with the why. Two callers — the worker and the read-only walk. | | `Verdict.calls_for_a_remake(refresh=)` | The one place a state becomes an action: `stale` always, `behind` only when asked. | | `data_version(path)` | Content hash of a directory or file — computed in the worker, before anything is annexed. | -| `Versions` | Per-run memo so a shared declared input hashes once, not once per dependent. | +| `Versions` | Memoizes content hashes within a classification or execution context. Separate worker processes do not share a mutable cache. | | `read(sidecar)` / `write(...)` | The manifest, `..manifest.json`. Both take the sidecar's own path, so a caller holding an output path has to say `manifest_path` out loud. | | `output_path(root, u, id, fmt)` | The output's file, guarded: any part that is not a single path component is refused, and so is a format that could not be an extension. | | `manifest_path(output)` | The sidecar beside it, named from the id alone — so it keeps its path, and its history, across a re-declared format. | diff --git a/docs/api/compute.md b/docs/api/compute.md new file mode 100644 index 0000000..322f9b9 --- /dev/null +++ b/docs/api/compute.md @@ -0,0 +1,57 @@ +# lightcone.engine.compute + +The allocation boundary shared by CLI lifecycle operations and execution. +`Compute` loads resource policy and obtains fresh native observations. +It owns no service, registry, or saved current-cluster selection. + +| Symbol | Contract | +|---|---| +| `Request.parse(...)` | Common exact/minimum CPU and memory requests, node count, walltime, startup class. | +| `Catalog.load(path)` | Ordered fixed shapes and stable connection namespaces; use the built-in local catalog only when the implicit default file is absent. | +| `Compute.plan(request)` | Select an eligible offer and freeze its native launch settings without allocation. | +| `Compute.launch(plan)` | Submit once and return a self-contained `Identity`. | +| `Compute.discover()` | Snapshots and per-connection errors, querying each authority once. | +| `Compute.status(id, wait=False, timeout=300)` | Native allocation state plus authenticated Dask readiness. | +| `Compute.down(id)` | Native termination independent of scheduler health. | +| `connect(id, timeout=10, config_path=None)` | Context manager borrowing a standard Dask client; closes the client, never the allocation. | +| `Provider` | `plan`, `launch`, `discover`, `inspect`, `connect`, `terminate`. | + +The built-in catalog exposes one `local` offer: one CPU, 1 GiB, one node, +fast startup, 30-minute default and two-hour maximum lifetime. It creates no +configuration file or allocation. Configured catalogs replace it completely. +Missing paths selected through an argument or `LC_COMPUTE_CONFIG`, unreadable +files, and invalid catalogs remain errors. Stable connection namespaces let +separate invocations discover and attach to the same local allocations. + +`local.py` and `slurm.py` implement the provider protocol. Adding an adapter means +adding one provider factory and its native mapping; `run` and `materialize` only +borrow clients through the common API. Provider settings stay behind that seam. +`runtime.py` owns private files, standard TLS material, and authenticated scheduler +identity checks. `local_runtime.py` and `slurm_bootstrap.py` compose stock Dask +components; they do not define custom workers or membership protocols. + +`Snapshot` distinguishes native allocation evidence from scheduler observations. +No live allocation size is filled from today's catalog. Connection namespaces +persist independently of offers, and IDs encode native incarnation evidence +without a UUID-to-job lookup database. Exceptions retain known cluster IDs and +submission tokens for partial/ambiguous acceptance. + +Execution submits ordinary tasks through the borrowed client's `submit` method. +Dask chooses the workers and handles dependencies; invocation-specific keys prevent +unintended reuse across commands. There is no worker-selection layer, per-worker +preflight orchestration, source fingerprinting, or login-node guard. Driver-side +preparation and the existing task runtime/sandbox checks remain in their owners. +`output.py` transports byte chunks through standard Dask events so detached +workers' output reaches the invoking CLI. Probes preserve both streams; +materialization sends recipe output to stderr to leave stdout for its report. + +Local teardown drains the allocation's validated process group rather than +assuming the owner's exit proves every child stopped. Failed unpublished launches +are cleaned up, and incomplete locator directories do not hide healthy allocations. +Cancellation and concurrent project writers are not made safe by allocation +management; callers must respect the documented execution limits. + +Tests cover deterministic selection, malformed identities and catalogs, partial +native failures, acceptance ambiguity, PID reuse, detached local lifetime, standard +Dask bootstrap, and explicit execution through borrowed clients. Slurm command +contracts are simulated; a real NERSC submission remains a deployment check. diff --git a/docs/api/container.md b/docs/api/container.md index 14a3256..bd360f6 100644 --- a/docs/api/container.md +++ b/docs/api/container.md @@ -57,7 +57,7 @@ Sources: `src/lightcone/engine/image.py`, `OCIBackend`, data-parameterized; the podman family is stated once (`_PODMAN_FAMILY`) and asked positively, so a new runtime falls outside it by default. podman-hpc adds exactly one step (`migrate`, - outside the load branch) and joins `_SHARED_STORE_RUNTIMES`. + outside the load branch). Execution verifies the prepared image on each worker. Detection order podman-hpc → podman → docker; docker's daemon is probed at detection. - **The architecture gate refuses before the load** — a wrong-arch diff --git a/docs/api/crate.md b/docs/api/crate.md index c403edb..9f93a38 100644 --- a/docs/api/crate.md +++ b/docs/api/crate.md @@ -1,9 +1,10 @@ # lightcone.engine.crate The publication view: the repository described as a Workflow Run -RO-Crate. The project *is* the crate — `ro-crate-metadata.json` sits at -the root, describes what the repository already holds, and a deposit is -`git archive`, not an export step. lc's manifests stay the canonical +RO-Crate. `ro-crate-metadata.json` sits at the root and describes the +repository's research objects. A complete deposit must include the annexed data +and results; `git archive` alone carries their Git representations, not their +content. lc's manifests stay the canonical record; the crate is the same facts in schema.org vocabulary for archives and viewers that will never run `lc`. diff --git a/docs/api/index.md b/docs/api/index.md index f3126ea..8674ea0 100644 --- a/docs/api/index.md +++ b/docs/api/index.md @@ -17,7 +17,7 @@ is responsibility and contract, not every signature. | [`assets`](assets.md) | One output: its directory, manifest, and state | pure | | [`worker`](worker.md) | Making one output; the rerun entry point | impure | | [`materialize`](materialize.md) | The driver: gates, scheduling, the save/restore loop, status | impure | -| [`venue`](venue.md) | Where a run executes: SLURM detection, the login guard | impure | +| [`compute`](compute.md) | Resource requests, native allocation lifecycle, borrowed Dask clients | impure | | [`sandbox`](sandbox.md) | The exec boundary: policy, backends, attestation, denials | mixed | | [`image` & `container`](container.md) | The container hatch: declaration → image → archive → runtime | pure / impure | | [`crate`](crate.md) | The publication view: the repo as an RO-Crate | pure | diff --git a/docs/api/materialize.md b/docs/api/materialize.md index 54b570d..f957937 100644 --- a/docs/api/materialize.md +++ b/docs/api/materialize.md @@ -7,22 +7,25 @@ its classification walk. Source: `src/lightcone/engine/materialize.py`. +Recipes are ordinary Dask tasks. Their stdout/stderr is forwarded as bytes to the +driver's stderr, independently of success or failure, leaving stdout for the report. + ## Key symbols | Symbol | Role | |---|---| -| `materialize(root, targets, *, refresh)` | The run: guards → converge → plan → fetch → schedule → save/restore loop → crate converge. | +| `materialize(root, targets, *, cluster_id, refresh)` | The run: guards → converge → plan → fetch → schedule → save/restore loop → crate converge. | | `check(root, targets, *, refresh)` | The same classification without executing, committing, or fetching. Exempt from the dirty refusal. | | `status(root)` | The report: every output's state and provenance commit, plus the mode/image/sandbox header facts. | | `MaterializeReport` / `StatusReport` | The JSON surfaces; `ok` and `up_to_date` first. | -| `cluster_for_run()` | The venue ladder, and the two-method scheduler seam (`submit`, `completed`). | +| `cluster_for_run(cluster_id)` | Borrow the cluster; the submit/completed scheduler seam (`submit`, `completed`). | | `run_record(...)` / `datalad_run_subject(...)` | The commit message `datalad rerun` replays, and the one spelling of its subject line — shared with the foreign-write comparator, because two strings here would drift. | | `_engine_requirement()` | How a record pins its engine: by version for a release, by source commit (hatch-vcs) for a dev build. | ## The run's order, and why -1. **Login guard first** — the allocation is the remedy with queue - latency, so the user submits it before fixing anything else. +1. **Explicit cluster first** — validate native allocation identity and connect + to its scheduler before preparing the project. 2. **Dirty refusal before the environment converge** — in containerized mode the converge can commit an image archive, and `dataset.save` commits the whole index; on a dirty tree the user's @@ -37,9 +40,9 @@ Source: `src/lightcone/engine/materialize.py`. — the driver commits as results arrive, so any per-task read could answer differently mid-run. Nondeterminism in a provenance field is worse than either answer. -6. **Save on `ok`, restore otherwise, `try/finally` around the loop** - — an interrupt restores whatever is still outstanding; the tree - ends as clean as it started. +6. **Save on `ok`, restore reported failures** — unreported outputs are retained + after interruption because their tasks may still be writing. Allocation + management does not provide concurrent-writer or cancellation guarantees. ## What must stay true @@ -51,9 +54,9 @@ Source: `src/lightcone/engine/materialize.py`. - **`up_to_date` is `ok and not made and not planned`** — a run where every recipe failed must not report "nothing to do", and `behind` never counts against it. -- **A read-only verb never tracebacks.** Anything `check`/`status` - cannot read classifies as "will be remade" and the real error - belongs to the recipe that follows. +- **Unreadable output state is reportable.** An unreadable manifest or input + can classify as "will be remade". An invalid spec, universe, or lock instead + raises `ProjectError`, which the CLI presents as a command error. - **The run record is genuinely re-runnable**: engine pinned by requirement, project environment rebuilt by the worker from the rerun commit's own lock, format tested *through datalad's parser* @@ -69,4 +72,4 @@ Source: `src/lightcone/engine/materialize.py`. `tests/test_materialize.py` — real repositories, real recipes, a real `LocalCluster` through the seam exactly once, real `datalad rerun` for the record's whole claim. `cluster_for_run` is the one monkeypatch -point for venue-free tests. +point for allocation-free tests. diff --git a/docs/api/sandbox.md b/docs/api/sandbox.md index d1175a3..9d0f171 100644 --- a/docs/api/sandbox.md +++ b/docs/api/sandbox.md @@ -18,10 +18,14 @@ plus `lightcone/_sandbox_exec.py`, the Landlock shim. | `Capability` | What this host can do — `detect()`'s answer, the only `sys.platform` branch. | | `Attestation` | What was actually enforced, derived from the flags applied — never from what the matrix says should have happened. | | `Backend.wrap(policy, argv)` | The pure rewrite. `contains_prefix` declares whether the uv hop rides inside (a container is a world; a host mechanism trusts host plumbing). | -| `exec_policy(...)` | The one policy: probe and recipe get the same thing. Building it is where the impurity lives (the per-run private `$HOME`); `scope()` owns its cleanup. | +| `exec_policy(...)` | Shared policy builder with a recipe output directory or probe `results/` write scope. Creates the private `$HOME`; `scope()` owns its cleanup. | | `Unavailable` | A real backend that wraps to the same argv and attests `fs: open`. Saying so is the caller's job; pretending is nobody's. | | `denial.explain()` / `denial.trailer()` | Best-guess remedies (allowed to return nothing) and the unconditional trailer on every nonzero sandboxed exit. | +An optional output receiver gets stdout/stderr byte chunks. Capturing output never +decodes or normalizes stdout; only the retained stderr tail is decoded for denial +classification. Without a receiver, stdout remains inherited. + ## What must stay true - **`wrap` stays pure** — no temp files, no FDs, no global state diff --git a/docs/api/venue.md b/docs/api/venue.md index d86735d..b639298 100644 --- a/docs/api/venue.md +++ b/docs/api/venue.md @@ -1,63 +1,10 @@ -# lightcone.engine.venue +# Venue selection has moved to compute -Where a run executes. A venue is host state, never project state — -nothing here reads the project or enters any identity. The one venue -beyond the local machine is a SLURM allocation, detected rather than -configured: the user already answered every resource question at -`salloc`, so the allocation *is* the declaration and lc's job is to -span it. +The former `lightcone.engine.venue` module has been replaced by +[`lightcone.engine.compute`](compute.md). Execution now requires an explicit +allocation ID; it does not infer a venue from Slurm environment variables or +start a local cluster automatically. -Source: `src/lightcone/engine/venue.py` (consumed by -`materialize.cluster_for_run`). - -## Key symbols - -| Symbol | Role | -|---|---| -| `slurm_client()` | The allocation branch: a scheduler in the driver process bound to `SLURMD_NODENAME`, one `srun --overlap` launching a worker per node on `sys.executable`. | -| `require_compute_node(command)` | The login guard: refuses iff a known center's marker is set and `SLURM_JOB_ID` is not, printing that center's own `salloc`/`sbatch` spellings. | -| `allocation_nodes()` | How many nodes the allocation holds; 0 outside one. | -| `_SITES` | One row per known center — name, marker, remedies, **verified against the center's documentation, never guessed**. NERSC is the seeded row. | - -## What must stay true - -- **The detection ladder lives in `cluster_for_run()` alone.** Nothing - else asks where a run executes; a future submission-model venue is - one more branch there plus only the config it genuinely needs. -- **Workers run the driver's own interpreter** (`sys.executable -m - distributed.cli.dask_worker`) — on HPC that is the tool env on the - shared filesystem, so driver and workers are the identical - installation and version skew is structurally out. Workers need no - git and no annex. -- **The worker flags are each load-bearing**: `--nthreads=` - (tasks block in `subprocess.wait()` with the GIL released), - `--no-nanny` (srun won't relaunch either), `--memory-limit 0` (the - real work is behind the exec boundary; Dask would pause workers over - phantom numbers), `--death-timeout 60` (a worker whose driver died - exits instead of holding the node), `--local-directory /tmp` - **literal** (a site prolog can scope `TMPDIR` per node or step, so a - driver-resolved path can be absent elsewhere). -- **The srun child is the one documented exception to `project._run`** - — it lives as long as the run and its stderr must reach the terminal - live. Teardown retires workers first, then wait → terminate → kill, - bounded; connection is a poll loop so a dead srun reports *its exit - code* now, not a timeout later. -- **A leak refuses loudly, never falls back silently**: `SLURM_JOB_ID` - with no srun on PATH, a non-integer count variable, an unresolvable - `SLURMD_NODENAME` — each is a named refusal. -- **The guard is materialize-scoped** (plus the rerun entry point — - the record's `cmd` is how recipes reach login nodes without `lc` in - the command line). `check`, `status` and `lc run` never call it: a - login node is exactly where "where does this stand" gets asked. -- **A containerized multi-node run requires a shared image store** — - `_SHARED_STORE_RUNTIMES` (podman-hpc), asked positively, checked in - `materialize()` before the runtime resolves so the refusal costs no - build. - -## Tests - -`tests/test_venue.py` — fakes the *host*, never the code: SLURM -variables set deliberately, a bash stub standing in for srun, and the -end-to-end tests run a real graph through the real bind/launch/teardown -on any machine. The `venue_env` autouse fixture scrubs venue variables -suite-wide (derived from `_SITES`, so a new center is one row). +For usage, see [Compute and clusters](../user/cluster.md) and the +[`lc compute` reference](../cli/compute.md). For implementation details, see +[compute internals](compute.md). diff --git a/docs/api/worker.md b/docs/api/worker.md index a54cc75..3d6e255 100644 --- a/docs/api/worker.md +++ b/docs/api/worker.md @@ -16,6 +16,10 @@ so advertising it would hand people a footgun. Source: `src/lightcone/engine/worker.py`. +Cluster execution supplies an output receiver to `materialize`/`execute`, which +passes byte chunks from the sandbox back to the invocation. Standalone reruns +retain direct terminal output. + ## Key symbols | Symbol | Role | diff --git a/docs/architecture.md b/docs/architecture.md index 32081a2..0129f82 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -1,6 +1,11 @@ -# Architecture +# Execution architecture -How lightcone-cli is put together, for someone about to change it. The +This page covers the execution layer of the Lightcone research stack. +Agent skills guide project development and reporting; `lightcone-cli` owns +execution and provenance; ASTRA supplies an external specification standard. +For repository ownership, see [Contribute to Lightcone](maintainer.md). + +Here is how lightcone-cli is put together, for someone about to change it. The [user-guide concepts page](user/concepts.md) covers what the tool promises; this page covers how the promises are kept. @@ -20,7 +25,7 @@ exit codes how outputs are made *means* - **The engine** owns everything about what a project is and how outputs get made. It raises `ProjectError`; the CLI's group class translates that into a clean error message, once, for every verb. -- **ASTRA** owns what a spec means. Scoping, `from:` references, +- **The external ASTRA project** owns what a spec means. Scoping, `from:` references, conditional outputs, universe resolution, and the recipe placeholder grammar are all answered by `astra.resolve` and checked by `astra.validation` — never re-implemented here. When the spec's @@ -35,14 +40,14 @@ imports. ## One run, end to end ```text -lc materialize - │ guard: compute node? tools? git identity? +lc materialize "$CLUSTER" + │ connect: native identity + Dask readiness + │ guard: tools? git identity? │ refuse: dirty tree - │ converge: uv.lock ⇄ .venv (and the image, containerized) │ plan: astra validate + resolve → Graph of Tasks │ fetch: git annex get (declared inputs not in this clone) - │ venue: SLURM allocation? → srun workers · else LocalCluster - ├─► workers: reset output dir → sandbox → recipe → hash → manifest + │ converge: uv.lock ⇄ .venv (and the image, containerized) + ├─► workers: reset output files → sandbox → recipe → hash → manifest │ (never raise; return ok/current/behind/failed/blocked) └─ driver: consume results in one thread ok → dataset.save (commit + run record) @@ -50,6 +55,10 @@ lc materialize finally → converge ro-crate-metadata.json (if licensed) ``` +After interruption, unreported outputs remain in place because their remote +tasks may still be writing. Stop the allocation and establish that work has +stopped before repairing those files. + The division of labor is strict and load-bearing: - **The driver owns git, alone.** Workers execute and return a @@ -110,7 +119,7 @@ git history while exiting 0 — so `lc init` sets `filter.annex.required=true`, which makes git refuse loudly instead (every filtered command, not only `git add`). Getting git-annex onto that `PATH` is the install's job, not the -repository's: `uv tool install lightcone-cli` puts it there alongside +repository's: the [tool installation](user/install.md) puts it there alongside `lc`. Each output is committed with a **run record** — a `[DATALAD RUNCMD]` @@ -120,7 +129,7 @@ worker entry point, so `datalad rerun` replays the making of an output with the gates, the sandbox, and the manifest intact. Results are committed *thin* (hard-linked to their annex object), which is safe precisely because lc never writes an output in place — the worker -resets the directory first. +removes the output's owned files first, without deleting sibling outputs. ## The exec boundary @@ -134,9 +143,10 @@ Because every backend is a pure argv rewrite, all of them are testable on a host that can't run them, and the manifest's `hermeticity` field records what was *actually* enforced — never what should have been. -There is one policy, `exec_policy`: probe and recipe get exactly the -same thing (tree read-only apart from `results/`), so "works under -`lc run`" and "works as a recipe" stay the same fact. +There is one policy builder, `exec_policy`: both probes and recipes get a +read-only project with a writable output scope. A recipe can write in the +directory its output lands in; a probe can write throughout `results/`. +The same environment and dependency rules apply to both. ## The container hatch @@ -151,16 +161,24 @@ environment sync and each recipe exec, over a read-only rootfs with the mount table as the whole policy. Execution pins the archive's config-blob id, never a tag. -## Venues - -`materialize.cluster_for_run()` is the one place that decides where a -run executes, and the seam it returns is two methods wide — -`submit(fn, *args, key=…)` and `completed(handles)`. A SLURM -allocation (detected by `SLURM_JOB_ID`) gets one worker per node via a -single `srun`, running the driver's own interpreter so driver and -workers are the identical installation. Anything else is the local -machine. Venues are detected, never configured; the only venue config -that exists is the allocation the user already requested. +## Compute allocations + +`engine.compute` owns allocation lifecycle through a small provider protocol. +A visible YAML catalog supplies ordered resource offers and stable native service +namespaces. When the implicit default file is absent, a built-in local catalog +provides one CPU and 1 GiB without setup. An explicit catalog replaces that default; +missing explicit paths and invalid files remain errors. No catalog is written and +no allocation starts until `compute launch` resolves resources and submits once. Slurm queries +and validated local OS identities are authoritative for allocations; Dask is the +authority for connected workers. Private scheduler/TLS files are connection +material, not a registry. + +`compute.connect(CLUSTER_ID)` borrows a standard Dask client and closes only that +client on exit. Both execution commands require a cluster ID. The materialization +scheduler keeps its `submit`/`completed` seam. Driver preparation and existing +task runtime/sandbox checks remain unchanged. Tasks use ordinary Dask scheduling; +there is no separate worker-selection or preflight layer, or site-marker guard. No execution command implicitly allocates compute. +See [compute internals](api/compute.md) and [deployment limits](user/cluster.md). ## The publication view @@ -190,7 +208,7 @@ src/lightcone/ # namespace — NO __init__.py ├── worker.py # making one output; the rerun entry point ├── materialize.py # the driver: gates, Dask, the save/restore loop ├── run.py # what `lc run` is - ├── venue.py # where a run executes + ├── compute/ # common allocation API, local and Slurm providers ├── sandbox/ # the exec boundary └── templates/ # the scaffold's file content, as real files ``` diff --git a/docs/assets/uranometria.webp b/docs/assets/uranometria.webp new file mode 100644 index 0000000..41e4e40 Binary files /dev/null and b/docs/assets/uranometria.webp differ diff --git a/docs/cli/compute.md b/docs/cli/compute.md new file mode 100644 index 0000000..bc104f8 --- /dev/null +++ b/docs/cli/compute.md @@ -0,0 +1,56 @@ +# lc compute + +Manage explicitly allocated Dask clusters using resource offers. +No project is required for these commands. + +This command is part of the compute-enabled preview. Install the pinned +source version from the [installation guide](../user/install.md); the older +PyPI releases do not provide `lc compute`. + +```text +lc compute [--config PATH] resources [--json] +lc compute [--config PATH] launch --cpus VALUE --memory VALUE + [--num-nodes N] [--time DURATION] [--startup fast] [--dry-run] [--json] +lc compute [--config PATH] status [CLUSTER_ID] [--wait] [--timeout SECONDS] [--json] +lc compute [--config PATH] down CLUSTER_ID [--json] +``` + +Without configuration, `resources` exposes a built-in `local` offer: one CPU, +1 GiB, one node, fast startup, and a 30-minute default lifetime (two-hour maximum). +Launch it with `lc compute launch --cpus 1 --memory 1`; execution still requires +the returned cluster ID. + +`~/lightcone-compute.yaml`, when present, replaces this built-in catalog. +`LC_COMPUTE_CONFIG` selects another file for both compute and execution commands; +`--config PATH` overrides it for this invocation. Missing explicit paths and +invalid catalogs are errors. Only a missing implicit default file enables the +built-in catalog, without writing a file or starting any compute. See the +[local and Slurm setup](../user/cluster.md) for examples. + +| Command | Behavior | +|---|---| +| `resources` | Ordered available offers, per-node shape, node limit, default/maximum time, and startup class. Free capacity remains unknown. | +| `launch` | Resolve one resource request and submit exactly once; print the opaque cluster ID on acceptance. | +| `launch --dry-run` | Show the resolved shape and native launch parameters without allocation. | +| `status` | Query each configured native authority once; retain partial discovery errors. | +| `status ID` | Inspect native state and probe Dask readiness separately. | +| `status ID --wait` | Wait for readiness, with a default deadline of 300 seconds; timeout leaves the allocation unchanged. | +| `down ID` | Request native termination even if the scheduler is unavailable. | + +CPU quantities are logical CPUs **per node**, memory is **GiB per node**, and +`--num-nodes` defaults to one. Bare quantities are exact; `4+` means at least four. +Time accepts positive whole minutes or hours, such as `30m` or `2h`. Without +`--time`, the chosen offer's default applies. `fast` is a service class, not a +queue-time promise. Limits apply to each allocation; aggregate quotas remain +with the native backend. + +The first eligible offer wins. An invalid configuration or failed submission is +an error, with no automatic resubmission elsewhere. An uncertain submission error +includes its token and any known cluster ID. Inspect existing allocations before +retrying it. + +`--json` emits versioned, allowlisted data without scheduler credentials. Status +reports phases `pending`, `active`, `stopping`, `ended`, or `unknown`; allocation +evidence is separate from Dask `observation`, `ready`, and worker count. Discovery +can partially succeed and still exit 1. Native errors, invalid requests, and +readiness timeouts also exit 1. diff --git a/docs/cli/index.md b/docs/cli/index.md index eda2f2e..a0eb2d9 100644 --- a/docs/cli/index.md +++ b/docs/cli/index.md @@ -1,22 +1,26 @@ # CLI Reference -The `lc` CLI is a thin wrapper around the engine. The user-facing -surface is small on purpose — `astra.yaml` carries the analysis -description, and the CLI is the durable, scriptable way to execute -and audit it. +The `lc` command executes and tracks your research project. Use it directly +from a terminal or through the Lightcone agent skills. Your analysis lives in +`astra.yaml`; the CLI manages its environment, compute, results, and provenance. + +This reference follows the compute-enabled preview in the +[installation guide](../user/install.md). ## Global behavior -- **The current directory is the project.** Every command except - `init` assumes it is invoked from the project root; there is no - walk-up and no global configuration. Outside a project, a command - errors cleanly. +- **Run project commands from the project root.** `lc init` can create + a directory, and `lc compute` works without a project. Other commands + inspect the current directory; they do not search parent directories. +- **Choose compute explicitly.** Launch an allocation with `lc compute`, + then pass its ID to `lc run` or `lc materialize`. Compute settings live + in an optional user catalog; a small local offer works without one. - **Nothing waits on a human.** No command prompts or opens an interactive shell — every verb runs to completion on its arguments alone, which is what makes the CLI safe to drive from scripts and agents. - **Refusals carry their remedy.** When a command refuses (a dirty - tree, a login node, a missing image), the message names the exact + tree, an unavailable cluster, a missing image), the message names the exact command that fixes it. ## Commands @@ -25,7 +29,8 @@ and audit it. |---------|---------| | [`lc init`](init.md) | Converge a directory into a Lightcone project (idempotent). | | [`lc materialize`](materialize.md) | Make the analysis's outputs; commit each one as it lands. | -| [`lc status`](status.md) | Report the state of every output. Reads only; always exits 0. | +| [`lc status`](status.md) | Inspect outputs and provenance without running recipes. | +| [`lc compute`](compute.md) | Allocate resources, inspect clusters, and end allocations. | | [`lc run`](run.md) | Run an ad-hoc command in the project environment, under isolation. | | [`lc build`](build.md) | Containerized projects: build the image and commit it. | @@ -42,12 +47,15 @@ Options: ## Exit codes - `0` — the command did what it says. -- `1` — a refusal or a failure, with the reason on stderr. For +- `1` — a refusal or a failure. For `lc materialize --check` and `lc init --check`, exit 1 means "work would be done" — the gate form scripts branch on. +- `2` — invalid CLI usage, such as a missing required argument. - `lc run` is a proxy: it exits with the command's own code (`128 + N` for a signal), so pipelines read it exactly as they would the bare command. -Every verb with a report takes `--json` for the machine-readable form; -each verb's page shows its shape. +Report commands accept `--json`; each command's page describes its shape. +`lc status` exits 0 when it can report the project, including stale results; +an unreadable or invalid project can still fail. Compute errors requested with +`--json` are returned as JSON on stdout. diff --git a/docs/cli/init.md b/docs/cli/init.md index 3005d5b..a49ba70 100644 --- a/docs/cli/init.md +++ b/docs/cli/init.md @@ -97,7 +97,7 @@ Inside `.git`, convergence sets one configuration key — reported as the That is the only thing `lc init` adds to what `git annex init` wrote. How git finds git-annex is still ordinary `PATH` resolution, which is -why `lc` should be installed with `uv tool install lightcone-cli` — it +why the [installation guide](../user/install.md) uses `uv tool install` — it puts `git-annex` on your `PATH` alongside `lc`. If your `git add` ever refuses, see [`fatal: … clean filter 'annex' failed`](../user/troubleshooting.md#fatal-clean-filter-annex-failed) @@ -130,6 +130,14 @@ cd my-analysis # decisions — and write the scripts the recipes name. uv add numpy # declare what the scripts import git add -A && git commit -m "First analysis" -lc materialize # make the outputs +lc compute launch --cpus 1 --memory 1 +# Copy the cluster ID printed above: +CLUSTER=PASTE_CLUSTER_ID_HERE +lc compute status "$CLUSTER" --wait +lc materialize "$CLUSTER" # make the outputs lc status # see where everything stands +lc compute down "$CLUSTER" # release the allocation ``` + +See [Compute and clusters](../user/cluster.md) for larger local allocations +and Slurm setup. diff --git a/docs/cli/materialize.md b/docs/cli/materialize.md index 6a7283c..4c14aaf 100644 --- a/docs/cli/materialize.md +++ b/docs/cli/materialize.md @@ -9,9 +9,15 @@ manifest, in a commit whose message is a replayable run record. ## Synopsis ```text -lc materialize [OPTIONS] [TARGETS]... +lc materialize [OPTIONS] CLUSTER_ID [TARGETS]... +lc materialize --check [OPTIONS] [TARGETS]... ``` +Execution requires the cluster ID returned by `lc compute launch`. No cluster +is chosen or started implicitly. `--check` needs no cluster. +Keep the allocation available while running this command; execution connects +to compute before preparing the project. + With no targets, everything the spec declares, across every universe. A target narrows the run to an output and whatever it depends on: @@ -21,6 +27,10 @@ A target narrows the run to an output and whatever it depends on: A target that matches nothing is an error listing what exists — quietly making nothing is the least useful thing a build tool can do. +For a first run, follow [Compute and clusters](../user/cluster.md#start-locally) +to launch compute and set `CLUSTER` to the returned ID. The allocation stays +available after materialization; release it with `lc compute down "$CLUSTER"`. + ## What gets remade An output is remade when it is `stale` — the analysis defines it @@ -36,13 +46,19 @@ never touched, under any flag. ## The run's contract -- **Starts clean, ends clean.** A dirty tree is a refusal (the message - sorts your uncommitted work from stray files under `results/`); a - failed or interrupted recipe's partial work is rolled back. +- **Starts clean.** A dirty tree is a refusal. A recipe that returns a + failure has its partial work restored. After a cluster interruption, + unreported outputs are retained because tasks may still be running. + Stop the allocation with `lc compute down CLUSTER_ID` and confirm its recipes + have stopped before cleaning results. Local containers may need separate + termination through their runtime; see [execution limits](../user/cluster.md#execution-requirements-and-limits). - **Fetches what it needs.** Declared inputs whose annexed content is not in this clone are fetched before anything hashes. - **Commits as it goes.** Each output lands in its own commit, written by the driver in one thread while other recipes keep running. +- **Forwards recipe diagnostics.** Recipe stdout and stderr reach the invoking + terminal on stderr, including failed recipes. stdout remains available for the + report, so `--json` stays machine-readable. - **Reports every independent failure.** One failing recipe doesn't abort the rest; its dependents report `blocked` and the run exits 1 with all of it listed. @@ -52,8 +68,8 @@ never touched, under any flag. On a containerized project, the run resolves the committed image first (building it as a preflight if the declaration is committed but the -image never built). Inside a SLURM allocation, the run spans every -allocated node — see [Running on a Cluster](../user/cluster.md). +image never built). Tasks use the explicitly selected cluster and the client +detaches on completion, leaving that cluster available — see [Running on a Cluster](../user/cluster.md). ## Check mode @@ -71,10 +87,9 @@ is what it is for. | `--refresh` | off | Also remake `behind` outputs. Never touches `current` ones. | | `--json` | off | Emit the report as JSON on stdout. | -There is deliberately no `--jobs` (a run takes every core; sizing -belongs to the allocation you run it in), no `--force`, and no flag to -*skip* a stale output — deleting its directory is your own file -operation, and stronger consent than a flag. +Task concurrency comes from the configured cluster. There is no `--jobs` or +`--force` option. `--refresh` includes results made under an earlier environment; +it leaves current results alone. ## The JSON report @@ -104,10 +119,10 @@ remedies are built to be pasted. ## Examples ```bash -lc materialize # everything, all universes -lc materialize fit # one output (and upstreams), every universe -lc materialize robust/fit # one universe's output +lc materialize "$CLUSTER" # everything, all universes +lc materialize "$CLUSTER" fit # one output (and upstreams), every universe +lc materialize "$CLUSTER" robust/fit # one universe's output lc materialize --check # would anything run? (exit 1 = yes) -lc materialize --refresh # also remake behind outputs +lc materialize "$CLUSTER" --refresh # also remake behind outputs lc materialize --check --json # the machine-readable gate ``` diff --git a/docs/cli/run.md b/docs/cli/run.md index 2171b98..7d89a2b 100644 --- a/docs/cli/run.md +++ b/docs/cli/run.md @@ -1,28 +1,50 @@ # lc run Run an ad-hoc command in the project environment, under isolation. -This is the probe verb: it executes exactly one command the way a -recipe would be executed — same environment, same sandbox — so "does -it work under `lc run`?" and "will it work as a recipe?" are the same -question. +Use this to check imports or try a script before adding its recipe to +`astra.yaml`. It uses the project's prepared environment and sandbox rules. +Recipes have a narrower write scope: their output directory, whereas a probe +can write throughout `results/`. ## Synopsis ```text -lc run COMMAND... +lc run CLUSTER_ID -- COMMAND... ``` -Everything after `run` is the command, verbatim — flags included. +The first argument is the cluster ID returned by `lc compute launch`. +Everything after `--` is the command, verbatim — flags included. Argv, the `docker run` / `uv run` convention: a single quoted string would be exec'd as one filename, so probe shell syntax through -`bash -c` instead. `lc run` takes no options of its own, so nothing -else needs escaping: +`bash -c` instead. [Launch compute](../user/cluster.md#start-locally) and set +`CLUSTER` to the returned cluster ID: ```bash -lc run python -c "import numpy; print(numpy.__version__)" -lc run python src/fit.py --points data/points.csv --outliers keep --output /tmp/probe +lc run "$CLUSTER" -- python -c "import numpy; print(numpy.__version__)" +lc run "$CLUSTER" -- python src/fit.py --points data/points.csv --outliers keep --output /tmp/probe ``` +The command is submitted as an ordinary task to the cluster's Dask scheduler, +which chooses a worker. The command uses the prepared project environment and +the same sandbox as a recipe. stdout/stderr are forwarded as bytes, preserving binary output and +line endings when redirected. The client detaches on completion; the allocation +stays available until `lc compute down` or its time limit. A missing cluster ID +is an error, with no implicit local execution. See [compute](compute.md). + +The command receives EOF on stdin; terminal input and pipes into `lc run` are not +forwarded. Pass input files through the project's declared inputs instead. +For direct execution, ambient environment variables come from the worker's +allocation environment. Prefixing the CLI with `NAME=value` does not forward +that variable to an existing cluster. Set command-specific values inside the +command, for example `lc run "$CLUSTER" -- env NAME=value python script.py`. +Containerized commands use the image's environment and the sandbox overlays. + +Interrupting the CLI detaches its client; the remote command may still be running. +Use `lc compute down "$CLUSTER"` to stop the allocation before working with files +the interrupted command could still be writing. Confirm that the command has +stopped; local containers may require separate termination through their runtime +(see [execution limits](../user/cluster.md#execution-requirements-and-limits)). + ## What it does - **Converges the environment first.** The probe syncs `.venv` to the @@ -41,8 +63,9 @@ lc run python src/fit.py --points data/points.csv --outliers keep --output /tmp/ an ASTRA input declaration for data, `results/` or `tempfile.mkdtemp()` for writes). -A probe has no output and writes no manifest: nothing it does is -recorded anywhere. Any uv project works — `lc run` doesn't require an +A probe writes no provenance manifest or result commit. Files it writes in +`results/` remain ordinary, untracked probe files; remove them before +materializing the project. Any uv project works — `lc run` doesn't require an `astra.yaml`, only `pyproject.toml`, `uv.lock` and `.venv` in the current directory. @@ -55,7 +78,7 @@ and the fix (declare the dependency) is the same in both places. ## Examples ```bash -lc run python -c "import scipy" # is the package in the lock? -lc run bash -c 'echo $HOME' # see the private HOME a recipe gets -lc run python src/fit.py --help # exercise a script exactly as a recipe would +lc run "$CLUSTER" -- python -c "import scipy" # is the package in the lock? +lc run "$CLUSTER" -- bash -c 'echo $HOME' # see the private HOME a recipe gets +lc run "$CLUSTER" -- python src/fit.py --help # exercise a script exactly as a recipe would ``` diff --git a/docs/cli/status.md b/docs/cli/status.md index fae5fdb..5344f93 100644 --- a/docs/cli/status.md +++ b/docs/cli/status.md @@ -2,7 +2,7 @@ Report what state each of the analysis's outputs is in. Reads only: it runs nothing, commits nothing, transfers no data, does not mind an -unclean tree, and always exits `0` — a state is not a failure. The +unclean tree, and exits `0` for a successful report, even when outputs are stale. The moment you most need to know where a project stands is when it isn't clean, so this verb works there. @@ -17,6 +17,7 @@ lc status [OPTIONS] ```text mode: direct sandbox: landlock (fs: declared, network: allowed) + crate: not maintained — declare [project].license to enable it · current baseline/fit a3f1f11 · current baseline/fit_plot a3f1f11 @@ -31,6 +32,11 @@ The header is repository facts: which mode the project executes in what enforcement a run on this host would get. No runtime and no network is needed to answer either. +The sandbox header describes the machine where you inspect the project. A +remote worker may have different capabilities; each output's manifest records +the enforcement actually used. The crate line reports publication metadata. +An invalid project or unreadable specification can still make this command fail. + Then one line per output the spec declares, in dependency order: its state, **the commit it was made at**, and — for anything not current — why. The commit column is the verb's reason to exist: "which code made @@ -67,7 +73,8 @@ eyes, check for exit codes. "mode": "direct", "image": null, "sandbox": "landlock (fs: declared, network: allowed)", - "counts": {"current": 4, "behind": 0, "stale": 0}, + "crate": "not maintained — declare [project].license to enable it", + "counts": {"current": 1, "behind": 0, "stale": 0}, "outputs": [ { "output": "baseline/fit", @@ -87,4 +94,5 @@ was materialized at and its content identity (both empty if it never was), and `foreign_write` — the sha of a hand-edit's commit when one was detected, which the prose `why` cannot carry for a machine consumer. For a containerized project, `image` is -`{"tag": ..., "state": "present" | "absent" | "unfetched"}`. +an object with `tag`, `archive`, and a `state` of `present`, `absent`, or +`unfetched`. diff --git a/docs/contributing/extending.md b/docs/contributing/extending.md index ae3312f..4bf73f1 100644 --- a/docs/contributing/extending.md +++ b/docs/contributing/extending.md @@ -16,9 +16,9 @@ use it". | When an output is remade | `engine/assets.py` (+ `test_assets.py`) | One `classify`; callers differ by one input value, never by logic. Ask first: does the change *contradict* the project (stale) or is it *circumstance* (behind)? | | How the spec becomes a graph | `engine/plan.py` (+ `test_plan.py`) | Ask `astra.resolve`; a missing answer is a PR to astra-tools; ambiguity is a `ProjectError`, never a guess. | | How a recipe runs | `engine/worker.py` (+ `test_worker.py`) | Never raises; no git; mutation-check every denial test. | -| What a run commits | `engine/materialize.py` (+ `test_materialize.py`) | The driver owns git alone; the tree ends as clean as it started. | -| Where a run executes | `engine/venue.py` + `cluster_for_run` (+ `test_venue.py`) | One detection ladder; venues detected, never configured; test by faking the host. | -| Supporting a new HPC center | `venue._SITES` | One row — marker + the center's own `salloc`/`sbatch` spellings, verified against its documentation, never guessed. | +| What a run commits | `engine/materialize.py` (+ `test_materialize.py`) | The driver owns git alone; restore reported failures, retain unreported partial files after interruption. | +| Where a run executes | `engine/compute/` + `cluster_for_run` (+ `test_compute*.py`) | Explicit allocation IDs; implement the provider protocol and register one factory. Execution borrows standard clients. | +| Supporting a new HPC center | Compute catalog | Expose resource offers with the site's native Slurm settings; there is no hostname-based placement guard. | | What a sandboxed command may touch | `sandbox/policy.py` (+ `test_sandbox_policy.py`) | Path sets only — no mechanism leaks in. | | Adding a sandbox mechanism | one module in `sandbox/` + one line in `detect()` | `wrap` pure, `attest` honest, `contains_prefix` answered. Nothing above the seam changes. | | A denial message | `sandbox/denial.py` (+ `test_sandbox_denial.py`) | Remedies copy-pasteable and real *today*; the trailer stays unconditional. | diff --git a/docs/contributing/setup.md b/docs/contributing/setup.md index 188f1cc..fecc4cb 100644 --- a/docs/contributing/setup.md +++ b/docs/contributing/setup.md @@ -1,78 +1,95 @@ -# Development Setup +# Development setup -Everything runs through [uv](https://docs.astral.sh/uv/); there is no -task runner and no other build tooling. +Choose the repository for the part of the stack you are changing. These are +contributor instructions; the [installation guide](../user/install.md) is the +shorter path for researchers. -## Clone & install +## Documentation site + +This site is built in the standalone `LightconeResearch/docs` repository: ```bash -git clone https://github.com/LightconeResearch/lightcone-cli.git -cd lightcone-cli -uv sync --group dev +git clone https://github.com/LightconeResearch/docs.git +cd docs +uv sync --locked +uv run zensical serve ``` -That resolves the engine and the dev tools (pytest, ruff, mypy, -datalad, the rocrate validator) into `.venv`. `uv run lc --version` -runs the checkout's `lc`. +Edit the Markdown in `docs/`, navigation in `zensical.toml`, and styles in +`docs/stylesheets/extra.css`. Before submitting a change, build the site with +the same command as CI: + +```bash +uv run zensical build --clean --strict +``` -You also need `git` on `PATH` (the one tool uv cannot install); -git-annex arrives as a wheel with the sync. +The build writes `site/`. CI validates pull requests and publishes changes on +`main` to GitHub Pages. Keep examples aligned with the documented CLI version; +a working example from a development branch may use commands the installed +release does not provide. -## The loop +## Agent skills ```bash -uv run pytest # the suite -uv run ruff check src/ tests/ # lint (--fix to apply) -uv run mypy src/ # strict mode +git clone https://github.com/LightconeResearch/agent-skills.git +cd agent-skills +npm run build +npm test ``` -These three are exactly what CI runs (`tests.yml`, `lint.yml`) — green -locally means green there, modulo the gated suites below. - -Most of the suite is hermetic: an autouse fixture stubs the engine's -one subprocess seam, so tests spawn nothing and touch no network. The -exceptions opt in explicitly — see [Testing](testing.md). +The generator requires Node.js 18 or newer. Edit canonical files under `skills/` +and `hooks/`, or plugin composition and tool pins in `skills.config.json`. +`npm run build` generates the packaged plugins and marketplace manifests; commit +those generated changes with their sources. Follow that repository's +[contribution guide](https://github.com/LightconeResearch/agent-skills/blob/main/CONTRIBUTING.md) +for plugin version bumps and installation smoke tests. -### The gated suites +## Lightcone CLI -Three suites answer questions only a real mechanism can, and each -skips where its mechanism is missing — with an environment variable CI -sets to turn the skip into a hard failure: +```bash +git clone https://github.com/LightconeResearch/lightcone-cli.git +cd lightcone-cli +git switch --detach 835de9e7c2df726722ffff8c6863d6f9b16ee577 +uv sync --group dev +``` -| Variable | Suite | Needs | -|---|---|---| -| `LC_SANDBOX_TESTS_REQUIRED=1` | `test_sandbox_enforcement.py` | Landlock (Linux) or Seatbelt (macOS) | -| `LC_CONTAINER_TESTS_REQUIRED=1` | `test_container_smoke.py` | podman or docker | -| `LC_CRATE_TESTS_REQUIRED=1` | `test_crate_smoke.py` | nothing beyond dev deps | +The checkout above matches this site's compute preview. Create a development +branch from that revision when making a change, or use the project's current +development branch and account for any differences from these docs. -## Building the docs +This installs the engine and its development tools into `.venv`. Run +`uv run lc --version` to use the checkout's CLI. Git must be available on `PATH`; +git-annex arrives with the Python dependencies. ```bash -uv sync --group docs -uv run zensical build # renders into site/ -uv run zensical serve # live preview +uv run pytest +uv run ruff check src/ tests/ +uv run mypy src/ ``` -The site deploys on release (`docs-deploy.yml`), so docs track the -released CLI, not `main`. A pre-release deploys nothing — the site keeps -serving the last full release. +The suite includes pure and stubbed tests alongside integration tests that use +real tools. See [Testing](testing.md) for the boundaries and fixture conventions. -## Building the wheel +### Tests that need a real mechanism + +These suites skip if their mechanism is unavailable. CI sets the corresponding +variable to turn a skip into a failure: + +| Variable | Suite | Needs | +|---|---|---| +| `LC_SANDBOX_TESTS_REQUIRED=1` | `test_sandbox_enforcement.py` | Landlock on Linux or Seatbelt on macOS | +| `LC_CONTAINER_TESTS_REQUIRED=1` | `test_container_smoke.py` | Podman or Docker | +| `LC_CRATE_TESTS_REQUIRED=1` | `test_crate_smoke.py` | The development dependencies | + +Run the relevant real suite when changing sandbox enforcement, containers, or +publication metadata. + +### Build the CLI distribution ```bash uv build ``` -CI runs this only to publish. The version comes from hatch-vcs — the -git tag for a release, tag-plus-commit for a dev build — which is also -what lets a run record pin a dev engine by its source commit. - -## Pre-PR checklist - -1. `uv run pytest` — including, if your change touches the sandbox, - containers, or the crate, the relevant gated suite on a host that - can run it. -2. `uv run ruff check src/ tests/` and `uv run mypy src/`. -3. New behavior lands with its tests, in the same PR. -4. Read [Extending](extending.md) — it says where each kind of change - belongs, and the invariants it must keep. +The CLI version comes from Git through hatch-vcs: a release tag for a release, +or a tag plus commit for a development build. Read [Extending](extending.md) +before adding engine behavior so the change uses the existing interfaces. diff --git a/docs/contributing/testing.md b/docs/contributing/testing.md index 8143fac..1e180c0 100644 --- a/docs/contributing/testing.md +++ b/docs/contributing/testing.md @@ -31,7 +31,7 @@ testing execution. | Classification | `test_assets.py` | pure | | One output, real recipe | `test_worker.py` | real boundary, real repo | | The run, the record | `test_materialize.py` | real repos; one real `LocalCluster`; real `datalad rerun` | -| Venue detection & launch | `test_venue.py` | fakes the *host* (env vars, a stub srun), never the code | +| Compute lifecycle | `test_compute*.py` | real detached local clusters; fake Slurm commands; real stock Dask bootstrap | | Policy / wrap / denial | `test_sandbox_*.py` | pure, run on every OS | | The kernel's answer | `test_sandbox_enforcement.py` | gated | | Image identity | `test_image.py` | pure — structure and ordering, never byte goldens | diff --git a/docs/index.md b/docs/index.md index e1b65bc..a14e5cd 100644 --- a/docs/index.md +++ b/docs/index.md @@ -1,62 +1,106 @@ -# Lightcone Research Stack +--- +title: The Lightcone Research Stack +description: Install Lightcone, work with a research agent, and turn an analysis into reproducible results. +hide: + - navigation + - toc +--- + +
+ +
+ +

Lightcone Research · Documentation

+ +# From research question
to reproducible result. + +

A practical toolkit for research you can inspect, reproduce, and build on. Describe your analysis, work with an AI agent, and keep your results connected to the choices that produced them.

+ +[Install the stack :lucide-arrow-right:](user/install.md){ .md-button .md-button--primary } +[Run your first analysis](user/getting-started.md){ .lc-text-link } -The **Lightcone Research Stack** is [Lightcone Research][lr]'s tooling for research -analyses described with [**ASTRA**][astra] (Agentic Schema for Transparent Research -Analysis). -You describe an analysis in an `astra.yaml` specification; the stack validates it and -takes care of the rest — execution, environments, and provenance. +

Open source. Built for researchers. Early alpha.

+Uranometria · Bayer · 1603 -!!! warning "Alpha development" - The stack is in **early alpha**. Its tools are still moving — expect breaking - changes between minor versions. Bug reports, design challenges, and use cases the - tooling doesn't yet cover are exactly what we want to hear at this stage; please - open an issue on the [repository](https://github.com/LightconeResearch) of the tool - concerned. +
-## Choose your path to the documentation +
-
+

The pieces, together

-- __I want to try it out__ – :lucide-rocket: +## One research workflow. A few focused tools. - --- +You make the scientific choices. The stack helps you record them, implement them, and follow them through to your results. - Installation instructions, step-by-step tutorial, and fast tour of the lightcone framework and its workflow capabilities. +
- [User Guide](user/index.md){ .md-button .md-button--primary } +
-- __I want to contribute__ – :lucide-cog: +01 / Describe - --- +### An inspectable analysis - In-depth tour of lightcone-cli's architecture and internals, as well as contribution instructions, aimed at - contributors and maintainers. +Use **ASTRA** to describe your inputs, outputs, methodological decisions, and supporting evidence in an `astra.yaml` file. - [Developer corner](maintainer.md){ .md-button .md-button--primary } +[Meet ASTRA, the external standard :lucide-arrow-up-right:](user/astra.md)
---- +
+ +02 / Collaborate + +### An agent that knows the workflow + +Add the **Lightcone plugin** to your coding agent for help scoping a question, writing the specification, implementing recipes, and understanding results. + +[Work with an agent :lucide-arrow-right:](user/agents.md) + +
-## Components of the stack +
-
+03 / Reproduce -- __lightcone-cli__ +### Results with a record - The library that ships the `lc` CLI: project scaffolding, locked environments, sandboxed execution, and the provenance layer. Depends on [**astra-tools**][astra-tools], the SDK for working with ASTRA analysis specifications. +Use the **`lc` command line** to run your analysis in a locked environment and record the inputs, choices, and code behind each output. - [:fontawesome-brands-github: Repository][cli]{ .md-button } +[Run your first analysis :lucide-arrow-right:](user/getting-started.md) -- __astra-tools__ +
+ +
+ +
+ +
+ +
+ +

Start with your question

+ +## A small first step. - The SDK for working with [**ASTRA**][astra] analysis specifications. This library provides the `astra` CLI which handles the [**ASTRA**][astra] lifecycle and validation process (schema, prior insights & findings, evidence verification helpers). +Install the tools, run a complete example, then bring your own data. You can work from a terminal or add an agent when you want help. - [:fontawesome-brands-github: Repository][astra-tools]{ .md-button } +[Understand how the stack fits together :lucide-arrow-right:](user/index.md)
-[lr]: https://lightconeresearch.org/ -[astra]: https://astra-spec.org/latest/ -[astra-tools]: https://github.com/LightconeResearch/astra-tools -[cli]: https://github.com/LightconeResearch/lightcone-cli +
    +
  1. 01InstallSet up the CLI and your agent.
  2. +
  3. 02Your first analysisGo from a specification to a result.
  4. +
  5. 03Work with an agentStart, resume, and refine a project.
  6. +
  7. 04Share your workPrepare results others can inspect.
  8. +
+ +
+ +
+ +**Growing with the research community.** Lightcone is in early alpha; commands and formats may change. [Report an issue](https://github.com/LightconeResearch/docs/issues), [contribute](maintainer.md), or visit [Lightcone Research](https://lightconeresearch.org/). + +
+ +
diff --git a/docs/maintainer.md b/docs/maintainer.md index fdfac4a..4e22d56 100644 --- a/docs/maintainer.md +++ b/docs/maintainer.md @@ -1,58 +1,36 @@ -# Developer corner - -`lightcone-cli` is a small engine with strong opinions: one way to -identify an output, one way to store it, one boundary to execute it -behind. This guide covers everything below the user surface — how the -engine is put together, what each module owns, and how to get a -working dev loop. - -If you're looking for the user-facing docs, the -[user guide](user/index.md) is the other half of this site. - -## What this covers - -- [Architecture](architecture.md) — the CLI/engine/ASTRA split, the - run pipeline, identity, storage, the exec boundary, and the - invariants that hold them together. -- [CLI Reference](cli/index.md) — every `lc` command: flags, JSON - report shapes, exit codes. -- [Engine Internals](api/index.md) — the `lightcone.engine.*` - modules: what each owns, its key symbols, and what must stay true - of it. -- [Contributing](contributing/setup.md) — clone, install, run the - test suite; [how the suite is shaped](contributing/testing.md); and - [where a change belongs](contributing/extending.md). - -## Get started in three commands - -!!! tip "Dev loop" - - ```bash - git clone https://github.com/LightconeResearch/lightcone-cli.git - cd lightcone-cli - uv sync --group dev && uv run pytest - ``` - -Test, lint (`uv run ruff check src/ tests/`) and type-check -(`uv run mypy src/`) are the whole loop — there is deliberately no -task runner in between. - -## The house rules - -A few conventions run through every module; changes are reviewed -against them: - -- **No dead code, no foreshadowing.** Nothing lands before the layer - that calls it, and no message names a verb or flag that doesn't - exist yet. `lc --help` advertises only what works. -- **No escape hatches around guarantees.** A feature that enforces - something ships without a flag to turn the enforcement off. -- **Literal behavior over invented convenience.** The current - directory is the project; erroring beats walking up or guessing. - Nothing prompts — a verb is run by an agent more often than a - person, and a prompt is a hang. -- **One implementation per rule.** Classification, path naming, the - run-record subject, tool resolution — each has exactly one spelling, - and a second copy is where the two start to disagree. -- **Honest reporting.** What was enforced, what was skipped, and what - a clone can't see are all recorded or said — never assumed. +# Contribute to Lightcone + +Lightcone's documentation, agent skills, and execution tools live in separate +repositories. Start with the repository that owns the behavior you want to change. + +| Contribution | Repository | Start here | +|---|---|---| +| Guides, installation, navigation, or site design | [LightconeResearch/docs](https://github.com/LightconeResearch/docs) | [Build the documentation](contributing/setup.md#documentation-site) | +| Agent instructions, plugin packaging, or validation hooks | [LightconeResearch/agent-skills](https://github.com/LightconeResearch/agent-skills) | [Work on agent skills](contributing/setup.md#agent-skills) | +| Project execution, environments, provenance, or CLI behavior | [LightconeResearch/lightcone-cli](https://github.com/LightconeResearch/lightcone-cli) | [CLI development](contributing/setup.md#lightcone-cli) | +| Analysis schema or ASTRA tooling | External [ASTRA project](https://astra-spec.org/latest/) | Follow ASTRA's own documentation and contribution process | + +For research work, start with the [user guide](user/index.md). You do not need +the developer tools below to use Lightcone. + +## CLI internals + +The [architecture guide](architecture.md) explains how the execution engine +uses ASTRA, uv, Git, git-annex, and its sandbox. The [engine reference](api/index.md) +maps responsibilities to modules; it is intended for contributors, rather than +as a supported Python API for research projects. + +Read [Testing](contributing/testing.md) before changing the engine and +[Extending](contributing/extending.md) to find where a change belongs. +The [CLI reference](cli/index.md) documents the user-facing contract. + +## Keep the stack consistent + +When changing behavior, update the matching instructions and examples. Check +installation commands against the version documented by this site and the tool +pins in the agent-skills repository. Document an external dependency as external; +ASTRA's schema and validation semantics remain the ASTRA project's responsibility. + +Keep errors actionable, record the enforcement actually used, and keep research +projects independent of the engine's Python internals. Code, tests, and their +dependencies should land together. diff --git a/docs/stylesheets/extra.css b/docs/stylesheets/extra.css index e454c12..989bc1e 100644 --- a/docs/stylesheets/extra.css +++ b/docs/stylesheets/extra.css @@ -1,57 +1,257 @@ -@import url('https://fonts.googleapis.com/css2?family=EB+Garamond:ital,wght@0,400..800;1,400..800&family=Libre+Baskerville:ital,wght@0,400;0,700;1,400&family=Inter:wght@300;400;500;600&family=JetBrains+Mono:wght@400;500&display=swap'); +@import url("https://fonts.googleapis.com/css2?family=Alegreya:ital,wght@0,400;0,500;0,600;0,700;1,400&family=JetBrains+Mono:wght@400;500&family=Newsreader:ital,opsz,wght@0,6..72,400;0,6..72,500;0,6..72,600;1,6..72,400&family=Quattrocento:wght@400;700&display=swap"); -/* ── Light mode ─────────────────────────────────────────────────────────── */ - -:root > * { - --md-primary-fg-color: #4e5a70; - --md-primary-fg-color--light: #6b7a8d; - --md-primary-fg-color--dark: #3a4456; - --md-accent-fg-color: #426b78; - --md-default-bg-color: #f8f7f3; - --md-default-bg-color--light: #ffffff; - --md-default-bg-color--dark: #f1efe9; - --md-text-font: "Libre Baskerville", Georgia, serif; - --md-code-font: "JetBrains Mono", "Fira Code", monospace; +/* Shared with lightcone-website/app/globals.css. Keep the four type voices + and the parchment / ink / antique-gold palette aligned with the website. */ +:root { + --lc-heading-font: "Quattrocento", Georgia, serif; + --lc-body-font: "Newsreader", Georgia, serif; + --lc-ui-font: "Alegreya", Georgia, serif; + --md-text-font: var(--lc-body-font); + --md-code-font: "JetBrains Mono", "Fira Code", monospace; } -body, -.md-header, -.md-main, -.md-main__inner, -.md-content, -.md-tabs, -.md-sidebar { - background-color: #f8f7f3; +[data-md-color-scheme="default"] { + --lc-paper: #f8f7f3; + --lc-surface: #f1efe9; + --lc-deep: #e8e5dd; + --lc-text: #221f20; + --lc-ink: #4e5a70; + --lc-gold: #a67c3c; + --lc-link: #3f7280; + --lc-muted: #6f645a; + --lc-rule: rgb(78 90 112 / 18%); + --lc-code: #374256; + --md-default-bg-color: var(--lc-paper); + --md-default-fg-color: var(--lc-text); + --md-default-fg-color--light: var(--lc-muted); + --md-default-fg-color--lighter: var(--lc-rule); + --md-default-fg-color--lightest: rgb(78 90 112 / 7%); + --md-primary-fg-color: var(--lc-ink); + --md-primary-bg-color: var(--lc-paper); + --md-accent-fg-color: var(--lc-link); + --md-accent-fg-color--transparent: rgb(63 114 128 / 8%); + --md-typeset-a-color: var(--lc-link); + --md-code-fg-color: var(--lc-code); + --md-code-bg-color: var(--lc-surface); + --md-footer-bg-color: var(--lc-surface); + --md-footer-bg-color--dark: var(--lc-surface); + --md-footer-fg-color: var(--lc-text); + --md-footer-fg-color--light: var(--lc-muted); + --md-footer-fg-color--lighter: var(--lc-muted); } -/* ── Dark mode ──────────────────────────────────────────────────────────── */ - [data-md-color-scheme="slate"] { - --md-default-bg-color: #221f20; - --md-hue: 219; + --lc-paper: #221f20; + --lc-surface: #2b2829; + --lc-deep: #353133; + --lc-text: #e9e5de; + --lc-ink: #b8c4d6; + --lc-gold: #d2ac70; + --lc-link: #96c5ce; + --lc-muted: #b4aaa0; + --lc-rule: rgb(224 217 205 / 18%); + --lc-code: #d3dce8; + --md-hue: 219; + --md-default-bg-color: var(--lc-paper); + --md-default-fg-color: var(--lc-text); + --md-default-fg-color--light: var(--lc-muted); + --md-default-fg-color--lighter: var(--lc-rule); + --md-default-fg-color--lightest: rgb(224 217 205 / 7%); + --md-primary-fg-color: var(--lc-ink); + --md-primary-bg-color: var(--lc-paper); + --md-accent-fg-color: var(--lc-link); + --md-accent-fg-color--transparent: rgb(150 197 206 / 10%); + --md-typeset-a-color: var(--lc-link); + --md-code-fg-color: var(--lc-code); + --md-code-bg-color: var(--lc-surface); + --md-footer-bg-color: var(--lc-surface); + --md-footer-bg-color--dark: var(--lc-surface); + --md-footer-fg-color: var(--lc-text); + --md-footer-fg-color--light: var(--lc-muted); + --md-footer-fg-color--lighter: var(--lc-muted); } -[data-md-color-scheme="slate"] body, -[data-md-color-scheme="slate"] .md-header, -[data-md-color-scheme="slate"] .md-main, -[data-md-color-scheme="slate"] .md-main__inner, -[data-md-color-scheme="slate"] .md-content, -[data-md-color-scheme="slate"] .md-tabs, -[data-md-color-scheme="slate"] .md-sidebar { - background-color: #221f20; +body { background: var(--lc-paper); } +::selection { background: var(--lc-deep); color: var(--lc-ink); } + +/* Masthead and wayfinding retain Zensical's native search, drawer and TOC. */ +.md-header, +.md-tabs { + background: var(--lc-paper); + color: var(--lc-text); + box-shadow: none; + font-family: var(--lc-ui-font); +} +.md-header { border-bottom: 1px solid var(--lc-rule); } +.md-header__inner { min-height: 3.6rem; } +.md-header__title { + font-family: var(--lc-heading-font); + font-size: 1rem; + font-weight: 400; + color: var(--lc-ink); } +.md-header__button.md-logo img { height: 1.65rem; width: auto; } +.md-tabs { border-bottom: 1px solid var(--lc-rule); } +.md-tabs__link { font-size: .8rem; opacity: 1; } +.md-tabs__link--active, +.md-tabs__item--active .md-tabs__link { color: var(--lc-ink); font-weight: 600; } +.md-tabs__item--active { border-bottom-color: var(--lc-gold); } +.md-tabs__link:hover { color: var(--lc-link); } +.md-search__form { background: var(--lc-surface); border: 1px solid var(--lc-rule); border-radius: .2rem; } +.md-search__input { font-family: var(--lc-ui-font); } +.md-search__input::placeholder { color: var(--lc-muted); } +.md-header .md-search__icon { color: var(--lc-ink); } +.md-grid { max-width: 68rem; } +.md-main__inner { margin-top: 1.8rem; } +.md-nav { font-family: var(--lc-ui-font); font-size: .78rem; line-height: 1.4; } +.md-nav__title { color: var(--lc-ink); background: var(--lc-paper); box-shadow: none; font-weight: 600; } +.md-nav__item--section > .md-nav__link { color: var(--lc-ink); font-weight: 600; } +.md-nav__link--active, +.md-nav__item .md-nav__link--active { color: var(--lc-link); font-weight: 600; } +.md-nav__link:hover { color: var(--lc-gold); } +.md-sidebar--secondary .md-nav { font-size: .72rem; } -[data-md-color-scheme="slate"] .md-nav__link--active, -[data-md-color-scheme="slate"] .md-nav__item--active > .md-nav__link { - background-color: rgba(106, 147, 160, 0.20); - color: #f8f7f3; +/* A reading page: bookish prose, a compact UI, and clearly separated code. */ +.md-typeset { font-size: .94rem; line-height: 1.65; } +.md-typeset h1, +.md-typeset h2, +.md-typeset h3, +.md-typeset h4 { + font-family: var(--lc-heading-font); + font-weight: 400; + line-height: 1.22; + letter-spacing: -.02em; +} +.md-typeset h1 { margin-bottom: .8em; color: var(--lc-gold); font-size: 2.2rem; } +.md-typeset h2 { color: var(--lc-ink); font-size: 1.5rem; margin-top: 1.7em; } +.md-typeset h3 { color: var(--lc-ink); font-size: 1.15rem; } +.md-typeset h4 { font-weight: 700; } +.md-typeset a { text-underline-offset: .18em; } +.md-typeset p a:not(.md-button):hover, +.md-typeset li a:not(.md-button):hover { text-decoration: underline; } +.md-typeset hr { border-color: var(--lc-rule); margin: 2rem 0; } +.md-typeset code { font-size: .72em; border-radius: .15rem; } +.md-typeset pre { font-size: .77rem; line-height: 1.65; } +.md-typeset pre > code { border: 1px solid var(--lc-rule); border-radius: .2rem; padding: 1rem; font-size: inherit; } +.md-typeset table:not([class]) { font-size: .83rem; box-shadow: none; border: 1px solid var(--lc-rule); border-radius: .15rem; } +.md-typeset table:not([class]) th { font-family: var(--lc-ui-font); font-weight: 600; background: var(--lc-surface); color: var(--lc-ink); } +.md-typeset table:not([class]) td { border-color: var(--lc-rule); } +.md-typeset .admonition, +.md-typeset details { font-size: .84rem; box-shadow: none; border-radius: .2rem; background: var(--lc-surface); } +.md-typeset .admonition-title, +.md-typeset summary { font-family: var(--lc-ui-font); } +.md-typeset .tabbed-labels { font-family: var(--lc-ui-font); } +.md-typeset .tabbed-labels > label { font-size: .8rem; } +.md-typeset .md-button { + font-family: var(--lc-ui-font); + font-size: .85rem; + font-weight: 500; + border: 1px solid var(--lc-ink); + border-radius: .15rem; + padding: .65em 1.15em; + transition: background-color .15s, color .15s; } +.md-typeset .md-button--primary { background: var(--lc-ink); color: var(--lc-paper); } +.md-typeset .md-button:hover { background: var(--lc-link); border-color: var(--lc-link); color: var(--lc-paper); } +a:focus-visible, +button:focus-visible, +label:focus-visible { outline: 2px solid var(--lc-link); outline-offset: 4px; } -[data-md-color-scheme="slate"] .md-content a { - color: #85c0d0; +/* Landing page: the website's engraved frontispiece and ruled sections. */ +.lc-home { max-width: 60rem; margin: 0 auto; } +.lc-hero { + position: relative; + isolation: isolate; + padding: 3rem 0 4rem; + margin: -1.8rem 0 0; + overflow: hidden; } +.lc-hero::before { + content: ""; + position: absolute; + z-index: -1; + inset: 0; + background: linear-gradient(rgb(166 124 60 / 20%), rgb(166 124 60 / 20%)), url("../assets/uranometria.webp") 85% 40% / cover; + opacity: .46; + mask-image: linear-gradient(to right, transparent 32%, rgb(0 0 0 / 8%) 48%, #000 100%); +} +.lc-home .lc-hero h1 { max-width: 20ch; font-size: clamp(2rem, 4.5vw, 3.25rem); line-height: 1.12; margin: .65rem 0 1.1rem; letter-spacing: -.035em; } +.lc-home .lc-eyebrow, +.lc-step { + font-family: var(--lc-ui-font); + font-size: .62rem; + font-weight: 600; + letter-spacing: .16em; + text-transform: uppercase; + color: var(--lc-muted); +} +.lc-home .lc-lead { max-width: 35rem; font-size: 1.12rem; line-height: 1.55; margin-bottom: 1.6rem; } +.lc-text-link { display: inline-block; margin: .6rem 0 .6rem 1.1rem; font-family: var(--lc-ui-font); font-size: .85rem; } +.lc-home .lc-hero-note { font-family: var(--lc-ui-font); font-size: .72rem; color: var(--lc-muted); margin-top: 1.2rem; } +.lc-engraving-credit { position: absolute; bottom: 1.3rem; right: 1.8rem; font-family: var(--lc-ui-font); font-size: .54rem; letter-spacing: .15em; text-transform: uppercase; color: var(--lc-muted); } +.lc-section { padding: 2.2rem 0 2.7rem; border-top: 1px solid var(--lc-rule); } +.lc-home h2 { margin: .5rem 0 .7rem; font-size: 1.8rem; } +.lc-home .lc-section > p:not(.lc-eyebrow) { max-width: 40rem; } +.lc-stack { display: grid; grid-template-columns: repeat(3, minmax(0, 1fr)); margin-top: 2rem; } +.lc-stack-item { min-width: 0; padding: 0 1.7rem; border-left: 1px solid var(--lc-rule); } +.lc-stack-item:first-child { padding-left: 0; border-left: 0; } +.lc-stack-item:last-child { padding-right: 0; } +.lc-stack-item h3 { font-size: 1.3rem; margin: .8rem 0; } +.lc-stack-item p { font-size: .91rem; } +.lc-stack-item p:last-child { font-family: var(--lc-ui-font); font-size: .78rem; margin-top: 1.2rem; } +.lc-paths { display: grid; grid-template-columns: 1fr 1.2fr; gap: 3rem; } +.lc-paths > div > p:last-child { font-family: var(--lc-ui-font); font-size: .8rem; } +.md-typeset .lc-reading-list { list-style: none; padding: 0; margin: 0; } +.md-typeset .lc-reading-list li { margin: 0; border-bottom: 1px solid var(--lc-rule); } +.md-typeset .lc-reading-list li:first-child { border-top: 1px solid var(--lc-rule); } +.md-typeset .lc-reading-list a { display: grid; grid-template-columns: 1.6rem 1fr 1rem; column-gap: .65rem; padding: .8rem 0; align-items: baseline; color: var(--lc-ink); text-decoration: none; } +.lc-reading-list span { font-family: var(--lc-ui-font); color: var(--lc-muted); font-size: .65rem; grid-row: 1 / 3; } +.lc-reading-list strong { font-family: var(--lc-heading-font); font-size: 1.1rem; font-weight: 400; } +.lc-reading-list small { grid-column: 2; color: var(--lc-muted); font-size: .8rem; } +.lc-reading-list b { grid-column: 3; grid-row: 1 / 3; align-self: center; font-weight: 400; } +.md-typeset .lc-reading-list a:hover { text-decoration: none; color: var(--lc-gold); } +.lc-colophon { border-top: 1px solid var(--lc-rule); padding: 1.3rem 0 2rem; font-size: .8rem; color: var(--lc-muted); } +[data-md-color-scheme="slate"] .lc-hero::before { opacity: .17; } + +/* Quiet footer, with a route back to the main website. */ +.lc-site-footer { border-top: 1px solid var(--lc-rule); background: var(--lc-surface); } +.lc-site-footer__inner { display: flex; justify-content: space-between; align-items: baseline; flex-wrap: wrap; gap: .6rem; padding: 1.6rem 1rem; } +.lc-site-footer__brand { font-family: var(--lc-heading-font); color: var(--lc-ink); font-size: 1.1rem; } +.lc-site-footer__inner > span { font-family: var(--lc-ui-font); font-size: .75rem; color: var(--lc-muted); } +.md-footer { font-family: var(--lc-ui-font); background: var(--lc-surface); } +.md-footer-meta { border-top: 1px solid var(--lc-rule); } +.md-footer__title { font-size: .9rem; } +.md-copyright { color: var(--lc-muted); } +.md-copyright__highlight { color: var(--lc-ink); } -[data-md-color-scheme="slate"] .md-content :not(pre) > code { - background-color: rgba(76, 63, 70, 0.15); - color: #c8dde5; +@media (max-width: 76.234375em) { + .md-nav--primary .md-nav__title { background: var(--lc-surface); color: var(--lc-ink); } + .md-nav--primary .md-nav__source { background: var(--lc-deep); color: var(--lc-text); } +} +@media (max-width: 59.984375em) { + .lc-hero { padding-top: 2rem; } + .lc-paths { gap: 1.8rem; } + .lc-stack-item { padding: 0 1rem; } +} +@media (max-width: 44.984375em) { + .md-typeset { font-size: .9rem; } + .md-typeset h1 { font-size: 1.9rem; } + .md-header__title { font-size: .83rem; } + .lc-hero { margin: -1rem 0 0; padding: 1.4rem 0 2.7rem; } + .lc-home .lc-hero h1 { font-size: 2.15rem; } + .lc-home .lc-lead { font-size: 1rem; } + .lc-hero::before { opacity: .2; mask-image: linear-gradient(to right, transparent, #000); } + .lc-engraving-credit { right: .8rem; bottom: .7rem; font-size: .5rem; } + .lc-stack, + .lc-paths { grid-template-columns: 1fr; gap: 1.7rem; } + .lc-stack-item, + .lc-stack-item:first-child { padding: 0 0 1rem; border: 0; border-bottom: 1px solid var(--lc-rule); } + .lc-stack-item:last-child { border-bottom: 0; padding-bottom: 0; } + .lc-section { padding: 1.7rem 0; } + .lc-home h2 { font-size: 1.5rem; } + .lc-text-link { margin-left: 0; margin-right: 1rem; } + .lc-hero .md-button { margin-right: 1rem; } +} +@media (prefers-reduced-motion: reduce) { + *, *::before, *::after { scroll-behavior: auto !important; transition: none !important; animation: none !important; } } diff --git a/docs/user/agents.md b/docs/user/agents.md new file mode 100644 index 0000000..b0148bb --- /dev/null +++ b/docs/user/agents.md @@ -0,0 +1,129 @@ +# Work with an agent + +Describe the research you want to do, then work with your agent to turn it +into a reviewable analysis. The Lightcone plugin gives a compatible coding +agent guidance for scoping a question, recording methodological choices, +writing recipes, running the analysis, and preparing a report. + +**Start here:** [install Lightcone and the optional agent plugin](install.md#add-an-agent-optional). +The `lightcone` plugin includes both the Lightcone and ASTRA skills, along +with their hooks. You only need that one plugin for the full workflow. + +!!! note "Using the compute preview" + The current plugin targets `lightcone-cli` version `0.5.0rc4`; these docs + follow the upcoming compute workflow. Tell your agent to use the installed + CLI's `--help`: execution now needs an allocation from `lc compute launch`, + followed by `lc run CLUSTER_ID -- COMMAND` or `lc materialize CLUSTER_ID`. + Release it with `lc compute down CLUSTER_ID` when finished. The + [first analysis](getting-started.md#4-make-the-results) shows the full sequence. + +## Start a project + +Create a project, then open your agent in that directory: + +```bash +lc init my-analysis +cd my-analysis +``` + +Invoke the Lightcone skill and describe your question. For example: + +=== "Claude Code" + + ```text + /lightcone:lightcone + I want to measure how sensitive my result is to the treatment of + outliers. My data is in data/measurements.csv. Help me define the + question, outputs, and methodological choices before implementing it. + ``` + +=== "Codex" + + ```text + $lightcone:lightcone + I want to measure how sensitive my result is to the treatment of + outliers. My data is in data/measurements.csv. Help me define the + question, outputs, and methodological choices before implementing it. + ``` + +Replace the example data path with a file you have actually provided. +The skill is designed to start with the scientific question and record the +answers in `astra.yaml`. Review the proposed outputs, decision options, +and baseline before asking the agent to implement them. + +If you already have scripts or a notebook, start with those: + +```text +Read the existing analysis in notebooks/exploration.ipynb. Identify its +inputs, outputs, and consequential methodological choices. Help me capture +them in astra.yaml and turn the first result into a reproducible recipe. +``` + +## Move from a question to results + +| Stage | What to ask for | What you can review | +| --- | --- | --- | +| Scope | Clarify the question, data, intended outputs, and plausible alternatives | `astra.yaml` and the baseline universe | +| Ground | Read relevant papers and connect evidence to methodological choices | Prior insights with checkable quotations | +| Implement | Write one recipe at a time, with decision values passed as arguments | Scripts, declared dependencies, and test runs | +| Execute | Commit the analysis and materialize the requested outputs | Result files, manifests, and `lc status` | +| Interpret | Compare universes and explain what the outputs support | Findings linked to their evidence | +| Communicate | Update the report from the analysis and prepare it for sharing | `index.md`, figures, and publication metadata | + +Tell the agent your compute limits and which universes you want to run. +A useful first request is one output in the baseline universe; you can +expand once you have inspected it. + +## Give the agent useful context + +Provide the research question, where the data lives, any existing code or +papers, and what a useful result would look like. Mention constraints such +as memory, runtime, available hardware, or methods you have already ruled out. + +For example: + +```text +Use the Lightcone skill. Start with the baseline result on my laptop. +Keep the run within 1 CPU and 1 GiB of memory. Explain any missing inputs +or dependencies before starting. Keep consequential methodological choices +in astra.yaml, and show me the outputs and their provenance when finished. +``` + +The plugin supplies workflow guidance; the tools still need to be available +where the agent runs. A remote agent environment needs its own `uv`, git, +`lc`, and access to the project and data. + +## Resume an analysis + +Open the agent in the existing project and ask it to orient itself: + +```text +Use the Lightcone skill. Read astra.yaml, AGENTS.md if present, and lc status. +Summarize the question, decisions, existing results, and unfinished work. +Then help me add a second defensible method and compare it with the baseline. +``` + +`astra.yaml` holds the scientific structure. An `AGENTS.md` file can retain +working context such as data quirks, project commands, and decisions about +scope. Results and manifests retain the execution record. Together they +help a new session continue from the project itself. + +For a newly cloned project, run `lc init` in its root to restore the local +environment and repository settings before working with it. + +## Review the handoff + +Ask the agent to finish with the result locations, what changed in the +specification, and any unresolved questions. `lc status` reports the state +of each output; `lc materialize --check` provides a check that exits +nonzero when requested outputs need work. + +The [first analysis tutorial](getting-started.md) shows the commands and +files behind this workflow. [Core concepts](concepts.md) explains how +Lightcone decides whether a result needs to run again. + +Continue with [Write a report](reporting.md) to connect the narrative to +the analysis, or [Share an analysis](sharing.md) to prepare it for others. + +Plugin source and available skills live in +[LightconeResearch/agent-skills ↗](https://github.com/LightconeResearch/agent-skills). diff --git a/docs/user/astra.md b/docs/user/astra.md new file mode 100644 index 0000000..10e647d --- /dev/null +++ b/docs/user/astra.md @@ -0,0 +1,76 @@ +# ASTRA: the analysis specification + +ASTRA is the **external, open specification** that Lightcone uses to describe +an analysis. It has its own documentation, schema, and tools. You can use ASTRA +with other execution tools; Lightcone provides one way to run it and record +the results. [Read the ASTRA documentation ↗](https://astra-spec.org/latest/). + +## What belongs in the specification + +Your project's `astra.yaml` connects the research question to the work needed +to answer it: + +| Element | What you record | +| --- | --- | +| Inputs | The datasets, files, and upstream analyses you use | +| Outputs | The metrics, figures, tables, or other artifacts you intend to produce | +| Decisions | Methodological choices, their alternatives, and the reasons for them | +| Recipes | Commands that produce outputs from declared inputs and choices | +| Prior insights | Existing claims and evidence that inform your approach | +| Findings | Claims supported by the analysis's outputs | +| Universes | A selection of decision options to evaluate together | + +Start with the inputs, outputs, and decisions you need for one result. Add +evidence as the research develops. The +[ASTRA format reference ↗](https://astra-spec.org/latest/specification/) +explains the fields and how they fit together. + +## How Lightcone uses it + +The [Lightcone agent skills](agents.md) help you draft and revise the spec. +The `astra` tools validate its structure and evidence. The `lc` CLI executes +its recipes and tracks the resulting artifacts. + +For an output that `lc` will execute, declare a `format`, a `recipe.command`, +and its input and decision dependencies. The command must write one file to +the `{output}` path supplied by `lc`. See [Your first analysis](getting-started.md) +for a complete working example. + +ASTRA also supports composing sub-analyses. This Lightcone CLI preview executes +flat analyses; use a single analysis with sibling outputs for its recipes. + +Keep the `version:` field written by `lc init`: it comes from the installed +ASTRA schema. It is a schema version, separate from the `lc` and `astra-tools` +package versions. + +## Validate and inspect + +You can run the ASTRA tools through `uvx`, which downloads and caches the +tool environment on first use. These examples use the version required by +the current Lightcone CLI and bundled agent skill: + +```bash +uvx astra-tools@0.2.18 validate +uvx astra-tools@0.2.18 info +uvx astra-tools@0.2.18 spec Output +``` + +Run these from your project directory. With no filename, `validate` checks +the project's specs and universe files. For literature-backed claims, add +`--verify-evidence` to check supporting quotations against their sources: + +```bash +uvx astra-tools@0.2.18 validate astra.yaml --verify-evidence +``` + +The agent plugin runs ASTRA tools with its own version pin and includes a +validation hook for supported agents. Installing Lightcone alone does not +put the separate `astra` command on your shell's `PATH`; `uvx` avoids needing +another installation. + +## Continue with ASTRA + +- [ASTRA getting started ↗](https://astra-spec.org/latest/getting-started/) — learn the format independently of Lightcone. +- [Specification reference ↗](https://astra-spec.org/latest/specification/) — decisions, evidence, composition, and field definitions. +- [ASTRA CLI reference ↗](https://astra-spec.org/latest/cli/) — validation, universes, and paper utilities. +- [ASTRA tools source ↗](https://github.com/LightconeResearch/astra-tools) — the Python CLI and SDK. diff --git a/docs/user/cluster.md b/docs/user/cluster.md index f422897..1e002fa 100644 --- a/docs/user/cluster.md +++ b/docs/user/cluster.md @@ -1,127 +1,169 @@ -# Running on a Cluster +# Compute and clusters -When local laptop time isn't enough, the same project runs on a SLURM -HPC system. There is no separate configuration to learn and no flag to -pass — `lc materialize` detects where it is running, and the allocation -you request *is* the resource declaration. +Use the same workflow on your workstation and on Slurm: launch an allocation, +run your project on its cluster ID, and release it when you finish. `lc status` +and `lc materialize --check` inspect the project without allocating compute. -## The big picture +These commands use the development version in the [installation guide](install.md). -`lc materialize` runs its tasks through a scheduler, and picks the venue -by looking at the environment: +## Start locally -1. **Inside a SLURM allocation** (`SLURM_JOB_ID` is set) → the run - spans every node the allocation holds: one worker per node, launched - via `srun`, using every core it was granted. -2. **Anywhere else** → the local machine, using every core. +Run these commands from a prepared project's root. Without a compute catalog, +Lightcone offers one local CPU and 1 GiB of memory for 30 minutes. You can request +up to two hours with `--time 2h`. -You already answered every sizing question at `salloc` / `sbatch` — -how many nodes, which constraint, how long — so `lc` asks none of its -own. There is no `--jobs`, no worker count, no venue config file. +```bash +lc compute resources +lc compute launch --cpus 1 --memory 1 +``` -## A typical SLURM workflow +Copy the cluster ID printed by `launch` and assign it to `CLUSTER`: -### 1. Prepare on the login node +```bash +CLUSTER=PASTE_CLUSTER_ID_HERE +lc compute status "$CLUSTER" --wait +lc run "$CLUSTER" -- python -c 'print("hello from the cluster")' +lc materialize "$CLUSTER" +lc compute down "$CLUSTER" +``` -Everything except executing recipes works on a login node — and one -verb is *for* it: +Launch returns when the allocation is accepted; `status --wait` waits until its +workers are ready. Finishing a run leaves the allocation available for another +command. `down` requests termination; the allocation also has a time limit. +Local CPU and memory settings are cooperative limits, rather than an exclusive +reservation of your workstation's hardware. + +For scripts, capture the versioned JSON report instead of copying the ID: ```bash -cd $SCRATCH/my-analysis -lc materialize --check # what would run, and why -lc status # where every output stands -lc build # containerized projects: build + commit the image +CLUSTER=$(lc compute launch --cpus 1 --memory 1 --json | uv run python -c 'import json,sys; print(json.load(sys.stdin)["id"])') ``` -### 2. Get an allocation and materialize inside it - -=== "Interactive" - ```bash - salloc --nodes=1 --constraint=cpu --qos=interactive --time=02:00:00 - # salloc drops you onto a compute node; from there: - cd $SCRATCH/my-analysis - lc materialize - ``` +## Request more resources -=== "Batch" - ```bash - cd $SCRATCH/my-analysis - sbatch --nodes=1 --constraint=cpu --qos=regular --time=02:00:00 \ - --wrap 'lc materialize' - ``` +`lc compute resources` lists the configured offers. CPU and memory quantities +are **per node**, with memory in **GiB**. Bare numbers request an exact match; +a `+` suffix accepts a larger offer. The first eligible offer wins. - (Make sure `lc` is on `PATH` in the batch environment — with a - `uv tool install`, that's `export PATH=$HOME/.local/bin:$PATH` in - the script if your shell profile doesn't already do it.) +```bash +lc compute launch --cpus 4+ --memory 8+ --time 1h --dry-run +``` -Ask for more nodes and the run uses them — independent outputs and -universes spread across the allocation with nothing else to say. +`--dry-run` shows the selected offer and launch parameters without allocating. +A real launch requires an offer that can satisfy your request; the built-in +one-CPU offer cannot satisfy this larger example. `--num-nodes` defaults to one. +`--startup fast` selects a service class, without guaranteeing a queue time. + +## Customize resource offers + +Create `~/lightcone-compute.yaml` to replace the built-in offer. This example +exposes a four-CPU local allocation: + +```yaml +version: 1 +connections: + workstation: + namespace: 22c84e48-2f0a-4cd2-90a2-30ce2e909bd1 + provider: local +offers: + - name: workstation + connection: workstation + resources: {cpus: 4, memory: 8} + max_nodes: 1 + time: {default: 30m, max: 2h} + startup: {class: fast} +``` -### 3. Guard rails on known centers +The namespace is a stable UUID identifying the connection. Keep it unchanged +while its clusters exist. A catalog replaces the default completely; stop any +built-in allocation before replacing its connection. -On centers `lc` knows (NERSC today), running `lc materialize` on a -login node refuses with the center's own allocation spellings rather -than quietly hammering a shared node: +Set `LC_COMPUTE_CONFIG` to use another file for both compute and execution: +```bash +export LC_COMPUTE_CONFIG="$HOME/my-compute.yaml" +lc compute resources ``` -Error: lc materialize executes recipes on compute nodes, and this is a -NERSC login node (NERSC_HOST is set with no SLURM allocation active). -Get an allocation and run it there: - - interactive: - salloc --nodes=1 --constraint=cpu --qos=interactive --time=02:00:00 - lc materialize +The compute group's `--config PATH` overrides this for a single invocation. +A missing explicitly selected file is an error. Only the absent default file +enables the built-in offer. + +## Configure Slurm + +Slurm setup depends on your facility. The submitting machine needs native +Slurm commands and permission to use the requested account and service. +The project and the Lightcone installation must be accessible to the workers. + +The following is an illustrative catalog, not a tested deployment recipe. +Replace the paths, account, partition, and resource shape with values for your +facility. The `python` path must point to an environment containing the same +Lightcone and Dask versions as the submitting client. + +```yaml +version: 1 +connections: + hpc: + namespace: 9d0c0fc5-9be8-407a-a3ec-f17c4110b162 + provider: slurm + launch: + python: /shared/tools/lightcone/bin/python + connection_root: /shared/home/alice/.lightcone/compute + scratch_root: /shared/scratch/alice/lightcone + task_slots_per_node: 30 + cpu_bind: threads +offers: + - name: batch + connection: hpc + resources: {cpus: 32, memory: 128} + max_nodes: 4 + time: {default: 1h, max: 12h} + startup: {class: batch} + config: + submit: sbatch + account: myproject + partition: compute +``` - batch (from the project root): - sbatch --nodes=1 --constraint=cpu --qos=regular --time=02:00:00 \ - --wrap 'lc materialize' +Then inspect and launch the allocation: -lc materialize --check, lc status and lc run work anywhere. +```bash +lc compute launch --cpus 32 --memory 128 --num-nodes 2 --time 1h --dry-run +lc compute launch --cpus 32 --memory 128 --num-nodes 2 --time 1h ``` -The read-only verbs are exempt on purpose — a login node is exactly -where "where does this project stand?" gets asked. - -## Containers on HPC - -A containerized project (one with `[tool.lightcone.image]` in its -`pyproject.toml`) works the same way, with three site realities to -know: - -- **`podman-hpc` is detected first.** Sites install it precisely - because plain podman's image store is invisible to compute nodes; - where both exist, `lc` prefers the wrapper and runs its extra - `migrate` step automatically, so the image is readable from every - node. -- **Build on a login node, once.** `lc build` builds the image and - commits it into the repository as versioned content — compute nodes - never build and need no registry access; an unfetched image arrives - through the annex like any other data. The archive records the - architecture it was built for, and a mismatched host is refused - before anything runs — so build where the architecture matches the - compute nodes (on NERSC, a login node). -- **Multi-node runs require a shared image store.** With plain podman - or docker the image exists only on the driver's node, so `lc` - refuses a multi-node containerized run unless the runtime is - `podman-hpc`. Single-node allocations work with any runtime. - -## Data on parallel filesystems - -Keep active projects on the filesystem your center recommends for job -I/O (`$SCRATCH` on NERSC), and remember scratch purge policies — the -project is a git repository, so `git push` to a remote (and -`git annex copy --to` for the bytes) is the durable copy. - -!!! warning "Early days" - HPC support is the youngest part of lightcone-cli and has not yet - been broadly validated on production systems. If something refuses, - hangs, or surprises you on your center, please - [open an issue](https://github.com/LightconeResearch/lightcone-cli/issues) - — site reports are exactly what this layer needs right now. - -## Where to next - -- [Core Concepts](concepts.md) — the model all of this rests on. -- [Troubleshooting](troubleshooting.md) — the refusals, quoted, with - remedies. +Use its returned ID with the same `status --wait`, `materialize`, and `down` +commands as the local example. You do not wrap materialization in `salloc` +or `sbatch`; the compute provider submits the allocation. + +A connection can set `context` to a native Slurm cluster name. Offers can use +`submit: salloc` with site-specific QOS or constraints for interactive services. +Their lifetime across logout or session cleanup needs checking at your site. +Batch submissions are independent of the invoking CLI. + +Check the resolved partition and walltime policy during deployment. Planning +freezes a partition and rejects unlimited overrun policies; site routing based +on QOS may need explicit configuration. Task concurrency is controlled by +`task_slots_per_node`; the scheduler also consumes allocation resources. + +## Execution requirements and limits + +- The driver and workers need the project, prepared environment, and inputs at + the **same absolute paths**, with matching Lightcone, Python major/minor, and + Dask versions. Relaunch compute after changing the worker installation. +- Containerized projects need their prepared image and runtime on every worker. +- Commands receive EOF on stdin. Direct commands inherit the workers' allocation + environment; variables added to the invoking shell after launch are not + forwarded. Use `lc run "$CLUSTER" -- env NAME=value python script.py` for a + command-specific value. +- Use one execution invocation per project at a time. An interrupted client may + leave recipes running; stop the allocation and confirm the work has stopped + before repairing partial outputs. Local containers managed by an external + runtime may also need stopping through that runtime. + +A failed or uncertain submission is not automatically retried on another offer. +If an error includes a submission token, inspect `lc compute status` and the +native scheduler before retrying; the allocation may already exist. Keep the +connection in your catalog while you need to inspect or terminate its clusters. + +The [compute reference](../cli/compute.md) lists all options and status fields. diff --git a/docs/user/concepts.md b/docs/user/concepts.md index b290397..7fb065e 100644 --- a/docs/user/concepts.md +++ b/docs/user/concepts.md @@ -1,151 +1,136 @@ -# Core Concepts +# Core concepts -The mental model behind `lc`, in one page. Nothing here is required to -follow [Getting Started](getting-started.md) — come back when you want -to know *why* the tool behaves the way it does. +A Lightcone project brings your research question, analysis choices, code, data, +and results together in one versioned directory. Agent skills help you develop +and review that project; `lc` runs the analysis and records how each result was +made. The analysis specification follows **ASTRA, an external standard** with +its own [documentation and tooling](https://astra-spec.org/latest/). -## A project is three files +## The project is the research record -A lightcone project is a directory holding an ASTRA spec and a uv -project: +| File or directory | What it holds | +|---|---| +| `astra.yaml` | Inputs, outputs, recipes, methodological decisions, and evidence | +| `universes/` | Selections of decision values to evaluate | +| `pyproject.toml` and `uv.lock` | The analysis dependencies and their resolved versions | +| `data/` | Input data, with large content stored through git-annex | +| `results/` | Generated outputs and their provenance manifests | +| `index.md` and `myst.yml` | The report and its MyST configuration | -- **`astra.yaml`** describes the analysis — inputs, outputs, recipes, - methodological decisions. It is the single source of truth: everything - `lc` does is downstream of it. -- **`pyproject.toml` + `uv.lock`** describe the environment — every - package a recipe may import, resolved to exact versions. The `.venv` - is built *from* the lock and is disposable; the lock is what's real, - and it travels in git. - -There is no global configuration, no registry, no state outside the -project. Clone the repository and you have everything except two pieces -of local machinery (`.venv` and the git-annex initialization), which -`lc init` rebuilds. - -Adding a dependency is a uv operation, not an `lc` one: +Keep your analysis code in the layout that suits your project, such as `src/`. +Add its dependencies with uv: ```bash uv add numpy ``` -That updates `pyproject.toml`, re-locks, and syncs `.venv` in one step. -Recipes import from the locked environment and nothing else — a stray -`pip install` on your machine changes nothing they can see. +The project environment is separate from the installed `lc` tool. Recipes use +packages from your lock file; installing a package elsewhere does not add it to +the analysis. After cloning a project, `lc init` reconstructs the local +environment and initializes git-annex. Annexed data may still need fetching; +materialization fetches the declared inputs it needs. + +## Decisions become universes + +A decision records a methodological choice, its defensible options, and your +reasoning. A universe selects values for those decisions. For example, `baseline` +might keep all observations while `robust` excludes outliers. Lightcone produces +separate results for each universe, so you can compare how the choice affects +your findings. + +ASTRA defines the schema, validation rules, and universe semantics. Lightcone +uses those definitions to execute recipes. See the [glossary](glossary.md) +for the individual terms. + +## Choose compute, then execute -## An output has an identity, and three facts about it +Compute is separate from the research record. Launch a local or Slurm allocation +with `lc compute`, then pass its ID to `lc materialize` or `lc run`. A small +local offer is available without configuration; a user-level catalog exposes +larger allocations and HPC services. -Every materialized output records, in its -`..manifest.json`: +```bash +lc materialize "$CLUSTER" +lc status +lc compute down "$CLUSTER" +``` + +Here `CLUSTER` is the ID returned by `lc compute launch`. A finished command +leaves the allocation available for reuse until you stop it or its lifetime +expires. [Compute and clusters](cluster.md) walks through the full lifecycle. +The project and environment must be visible at the same paths to the driver +and workers. -1. **What it is** — a hash of its recipe and the decision values that - shaped it (its *definition*). -2. **What it was made from** — a content hash of each declared input. -3. **What it ran under** — a hash of the environment (the lock, the - interpreter, the image declaration if any), plus the git commit the - run started at. +## What makes an output current? -Those three facts are deliberately not one fact, because they age -differently — and that is what the three states mean: +Every output has a manifest beside it: `..manifest.json`. It records +the recipe and selected decisions, input content hashes, the environment, and +the Git commit at the start of the run. -| state | means | what `lc materialize` does | +| State | Meaning | Next materialization | |---|---|---| -| `current` | the output is exactly what the spec asks for, made from these inputs, under this environment | nothing | -| `stale` | the output **contradicts** the project: the spec now defines it differently, or an input's content changed | remakes it | -| `behind` | the output is still exactly what the spec asks for — only the **environment** moved since it was made | reports it, leaves it alone | - -The line between `stale` and `behind` is contradiction versus -circumstance. A stale output is mislabelled — keeping it would be a -lie, so it is remade. A behind output is not wrong in any way: one -`uv add` for a plotting script rewrites the lock for the whole project, -and remaking a week of computation over that buys nothing. Its manifest -records exactly which environment and commit produced it, and that -commit's own `uv.lock` reconstructs the environment if you ever need -it. - -When you *do* want behind outputs remade — before a release, say — -that is one flag: +| `current` | Definition, inputs, and environment match the recorded result | Leaves it alone | +| `stale` | Definition or input content changed, or a result was committed outside its run record | Remakes it | +| `behind` | Definition and inputs match, but the environment changed | Leaves it alone unless you use `--refresh` | + +Input changes propagate by content. If rebuilding an upstream result produces +identical bytes, it does not force a downstream rebuild. Changing a dependency +in `uv.lock` makes existing results `behind`, so an unrelated package update +does not automatically repeat expensive work. To refresh those results: ```bash -lc materialize --refresh +lc materialize "$CLUSTER" --refresh ``` -`--refresh` only ever widens a run: a `current` output stays current -under it, and there is deliberately no flag in the other direction — -nothing suppresses the rebuild of a stale output. - -One more way an output can be stale: a hand edit. Every output is -committed by the run that made it, so a file changed by hand and -committed shows up in history under a commit that is not a run record — -and the output classifies `stale` everywhere, with `lc status` naming -the foreign commit. - -## Everything is committed, and the tree stays clean - -`lc` versions results in the project's own git repository: git carries -the history and the small files, git-annex carries the data bytes — -transparently, behind the ordinary `git add` / `git commit` you already -type. - -That model has two consequences you'll feel: - -- **A run starts from a clean tree.** Every output is committed - together with the code that produced it; a run that started from - uncommitted edits could not say what that code was. So: commit, then - materialize. -- **A run ends with a clean tree.** Each output is committed as it - lands — with its manifest, in a commit whose message is a *run - record* that `datalad rerun` can replay. A failed recipe's partial - work is rolled back. Your `git log` is the build log. - -`results/` is `lc`'s to write. Don't put files there by hand — a -hand-placed file has no manifest and no run record, and the foreign -write check above exists precisely to catch it. - -## Two modes, derived from the project - -How recipes execute is never configured — it is read off the project: - -- **Direct mode** (the default): recipes run on your machine, in the - project's `.venv`, under an OS sandbox — Landlock on Linux, Seatbelt - on macOS. The project tree is read-only except each recipe's own - directory its output lands in; undeclared tools don't execute. -- **Containerized mode**: declaring a `[tool.lightcone.image]` table in - `pyproject.toml` *is* the switch. Recipes then run inside a - content-addressed image built from that declaration — and the image - itself is saved into the repository as versioned content, so a clone - obtains the exact bytes with no registry and no credentials. - `lc status` shows the mode and the image's state. - -Either way, every manifest records what enforcement actually ran -(`hermeticity`) — a host with no sandbox mechanism runs the recipe and -says so, rather than pretending. - -## Reading and gating are different verbs - -- **`lc status`** reports. It always exits 0 — a state is not a - failure — runs nothing, and doesn't mind a dirty tree, because the - moment you most need it is when things aren't clean. It's also the - verb that shows the commit each output was made at. -- **`lc materialize --check`** gates. It classifies everything without - running anything and exits 1 if a run would do work — the thing a - script or CI job branches on. - -Both have `--json`; the first two keys of the check report, `ok` and -`up_to_date`, are the ones to branch on. - -## Publication is a license away - -Declaring a `license` under `[project]` in `pyproject.toml` is -declaring the intent to publish. From then on, every `lc materialize` -maintains `ro-crate-metadata.json` at the project root — an -[RO-Crate](https://www.researchobject.org/ro-crate/) describing the -project, its outputs, and the runs that produced them. The repository -*is* the crate; depositing it is `git archive` on something you already -have. - -## Where to next - -- [Running on a Cluster](cluster.md) — the same model on SLURM. -- [Troubleshooting](troubleshooting.md) — the refusals quoted, with - their remedies. -- [Glossary](glossary.md) — the terms, one at a time. +Source code does not automatically participate in the definition hash. Declare +scripts as ASTRA inputs when their content should trigger a rebuild. A recipe +can then refer to a script through its input placeholder. + +## Commit before running + +Materialization starts from a clean Git tree so each result has a precise code +revision. Review and commit your edits before running. The driver then commits +each successful output with its manifest and a replayable run record. Git +stores the history and small files; git-annex stores the larger data bytes. + +A recipe that reports failure has its partial result restored. After a lost +connection or interrupted command, remote recipes may still be writing. Stop +the allocation and confirm they have stopped before cleaning `results/`. +[Execution limits](cluster.md#execution-requirements-and-limits) explain this +case, including local containers that may require separate termination. + +Use one execution command per project at a time. Treat `results/` as generated +content; keep hand-written analysis and report files outside it. + +## Two execution environments + +In **direct mode**, recipes use the project's prepared environment on an +allocation worker. On supported hosts, filesystem access is constrained by +Landlock on Linux or Seatbelt on macOS. Recipes may write in their output +directory and private scratch space. + +In **containerized mode**, a `[tool.lightcone.image]` table in `pyproject.toml` +declares the system layer. `lc build` saves that image in the repository through +git-annex; Python dependencies still come from the project lock. See +[`lc build`](../cli/build.md) for the declaration and runtime requirements. + +Every manifest records the isolation actually enforced. A host without a +supported sandbox reports that fact. Network access remains allowed. + +## Inspect results without running them + +`lc status` shows each result's state and provenance commit. A successful report +exits 0 even if outputs are stale. `lc materialize --check` is the automation +gate: it exits 1 when work would be needed. Neither command needs a cluster or +fetches data, and both accept `--json`. + +## Reports and publication metadata + +Write the research narrative in the project's MyST report and reference the +outputs that support your findings. Declaring a license under `[project]` in +`pyproject.toml` also enables maintenance of `ro-crate-metadata.json`, a +machine-readable description of the project and its provenance. + +That metadata helps archives and other tools understand the research record. +A complete deposit must include the data and result bytes as well as the +metadata: `git archive` alone does not include git-annex content. diff --git a/docs/user/getting-started.md b/docs/user/getting-started.md index dd3052b..cb160ec 100644 --- a/docs/user/getting-started.md +++ b/docs/user/getting-started.md @@ -1,394 +1,234 @@ -# Getting Started +# Your first analysis -Let's go from nothing on your disk to a working, reproducible analysis. -You can read this top to bottom without running anything, or follow along — -every command is copy-paste ready. +Build a small analysis that answers one question: **how much does the choice +of estimator change a result?** You will summarize five measurements using +the mean and the median, then inspect both results and the record of how +they were made. -**What you'll build:** a small two-output analysis that fits a line to a -noisy dataset and sweeps one methodological decision — whether points far -from an initial fit are kept or clipped. The result is two universes, -`baseline` and `robust`, each with its own fitted slope and figure, and a -project that ends published as an [RO-Crate](https://www.researchobject.org/ro-crate/). +Complete the [installation](install.md) first. This example uses Python's +standard library, runs locally, and needs no container or external dataset. +If you prefer to start from your own research question with an assistant, +follow [Work with an agent](agents.md). -Make sure you've finished the [install](install.md) first. - -## 1. Create a project +## 1. Create the project ```bash -lc init line-fit-demo -cd line-fit-demo +lc init first-analysis +cd first-analysis ``` -`lc init` converges the directory to a small, opinionated layout and -stops; it doesn't ask any questions, and it's idempotent — re-running -it later only fills in whatever is missing. +Lightcone creates the project files, a managed Python environment, and a +git repository configured to store data and results. The files you will +work with are: -``` -line-fit-demo/ -├── astra.yaml # the spec, empty for now — this is where everything lives -├── pyproject.toml # the project's environment: its dependencies… -├── .python-version # …and the exact interpreter, locked by uv -├── uv.lock -├── .venv/ # built from the lock (local, never committed) -├── .git/ # a git repository, with git-annex initialized -├── .gitattributes # the storage policy: what the annex carries -├── .gitignore -├── .datalad/ # dataset identity — the project is a DataLad dataset -├── data/ # declared input data lives here -├── results/ # outputs materialize here — lc's to write, not yours -├── universes/ -│ └── baseline.yaml # one universe, selecting nothing yet -├── myst.yml # MyST report configuration -└── index.md # template report, to reference the spec from -``` +| File or directory | Your use | +| --- | --- | +| `astra.yaml` | Describe the inputs, outputs, and methodological choices | +| `universes/` | Select the choices to run | +| `data/` | Keep the input measurements | +| `src/` | Write the analysis scripts; you will create this directory | +| `pyproject.toml` and `uv.lock` | Declare and lock the project's dependencies | +| `results/` | Read the outputs produced by Lightcone | +| `index.md` and `myst.yml` | Develop a report when you are ready | -Two things are worth registering now: +Run the remaining commands from this project directory. -- **The project is a git repository.** Every - output `lc` makes is committed together with the code that produced - it; large files ride in git-annex behind the scenes, but you only - ever type ordinary `git add` and `git commit`. -- **The environment is the lock.** `pyproject.toml` + `uv.lock` define - exactly what your recipes can import, and `.venv` is built from them. - You'll add packages with `uv add` in a moment — never `pip install`. +## 2. Add the data and script -The file you'll actually work in is **`astra.yaml`** — the single source -of truth for your analysis. Inputs, outputs, methodological decisions, -recipes: everything else lightcone-cli does is downstream of this file. +Create the input file and a script that accepts the estimator as an argument: -## 2. Add the data +```bash +mkdir -p src +cat > data/measurements.txt <<'EOF' +1 +2 +3 +4 +18 +EOF -A real project starts from a dataset; ours will generate a small one — -200 points on a line, with a few outliers thrown far off it: +cat > src/estimate.py <<'EOF' +import argparse +import json +import statistics +from pathlib import Path -```bash -python3 - <<'EOF' -import random -random.seed(0) -rows = ["x,y"] -for _ in range(200): - x = random.uniform(0, 10) - y = 2.5 * x + 1.0 + random.gauss(0, 1.5) - if random.random() < 0.04: - y += random.gauss(0, 15) - rows.append(f"{x:.6f},{y:.6f}") -open("data/points.csv", "w").write("\n".join(rows) + "\n") +parser = argparse.ArgumentParser() +parser.add_argument("--data", required=True) +parser.add_argument("--method", choices=["mean", "median"], required=True) +parser.add_argument("--output", required=True) +args = parser.parse_args() + +values = [float(line) for line in Path(args.data).read_text().splitlines()] +estimator = statistics.mean if args.method == "mean" else statistics.median +result = {"method": args.method, "estimate": estimator(values), "n": len(values)} +Path(args.output).write_text(json.dumps(result, indent=2) + "\n") EOF ``` -`data/` is where declared inputs live. When you commit, the -`.gitattributes` policy routes the file's bytes into git-annex -automatically — the file stays an ordinary readable, writable file in -your tree, and the repository stays light. +For your own scripts, add packages with `uv add`, for example +`uv add numpy`. Lightcone runs recipes in the environment described by the +project's lockfile. -## 3. Write the spec +## 3. Describe the analysis -`astra.yaml` was scaffolded as an empty analysis. Fill it in with ours: +Replace the example contents of `astra.yaml` with the following. Keep the +`version:` value written by `lc init` if it differs from this example: -```yaml -version: "0.0.13" # ASTRA schema version — keep what the scaffold wrote -name: "line_fit" -description: | - Fit a straight line to a small synthetic dataset and sweep one - methodological decision: whether points far from an initial fit are - kept or clipped before the final fit. +```yaml title="astra.yaml" +version: "0.0.14" +name: Estimator comparison +description: Compare the mean and median of five measurements, including a large value. inputs: - - id: points + - id: measurements type: data - source: data/points.csv - description: "200 synthetic (x, y) points, a few of them far off the line" + source: data/measurements.txt + description: Five example measurements. + - id: estimator_script + type: data + source: src/estimate.py + description: Implementation of the two estimators. outputs: - - id: fit + - id: estimate type: metric format: json - description: "Slope and intercept of the least-squares line" - inputs: [points] - decisions: [outliers] + description: Estimated central value and number of measurements. + inputs: [measurements, estimator_script] + decisions: [estimator] recipe: - command: python src/fit.py --points {inputs.points} --outliers {decisions.outliers} --output {output} - - - id: fit_plot - type: figure - format: png - description: "The points and the fitted line" - inputs: [points, fit] - recipe: - command: python src/plot.py --points {inputs.points} --fit {inputs.fit} --output {output} + command: >- + python {inputs.estimator_script} + --data {inputs.measurements} + --method {decisions.estimator} + --output {output} decisions: - outliers: - label: "Outlier handling" - rationale: "A few points sit far off the line; keeping or clipping them shifts the slope." - default: keep + estimator: + label: Choice of estimator + rationale: The mean and median respond differently to extreme measurements. + default: mean options: - keep: - label: "Keep every point" - clip: - label: "Drop points beyond 3 sigma of an initial fit" -``` - -A few things to notice: - -- Each output declares its full dependency contract: `fit` depends on - the `points` input and the `outliers` decision; `fit_plot` depends on - `points` and on the sibling output `fit`. That contract is how `lc` - orders the build — and how it knows what to rebuild when something - changes. -- Recipes reference those dependencies through placeholders — - `{inputs.points}`, `{decisions.outliers}`, `{output}` — which are - expanded at execution time. `{output}` is the output's own file, - `results//.`; the engine creates the - directory before the recipe runs, and the recipe writes that one path. -- Each output declares a `format` — the extension its artifact is - written with. It is what names the file, so a consumer knows what an - output *is* from the spec alone, and one output is always one file. -- The decision's options aren't hardcoded anywhere in code; the scripts - will take them as command-line arguments. - -`universes/baseline.yaml` was scaffolded empty, so give it a value for -our decision: - -```yaml -id: baseline -description: "Every point kept — the decision defaults." -decisions: - outliers: keep + mean: + label: Arithmetic mean + median: + label: Median ``` -Each universe is one complete selection of decision values; its results -materialize to `results//.`. +The script is a declared input too, so changing its contents makes the +result need rebuilding. The recipe receives the selected decision and the +output filename from Lightcone. -Check the spec is well-formed: +Replace `universes/baseline.yaml` and add `universes/robust.yaml`: ```bash -astra validate astra.yaml -``` - -(`astra` is the spec-side CLI; it ships with `astra-tools`, a dependency -of lightcone-cli.) - -## 4. Write the scripts - -Two short scripts, in a `src/` directory (`mkdir src` — the scaffold -doesn't create it; where code lives is your choice, the recipes above -just happen to point there). First `src/fit.py`: - -```python -import argparse -import json -from pathlib import Path - -import numpy as np - -parser = argparse.ArgumentParser() -parser.add_argument("--points", required=True) -parser.add_argument("--outliers", choices=["keep", "clip"], required=True) -parser.add_argument("--output", required=True) -args = parser.parse_args() +cat > universes/baseline.yaml <<'EOF' +id: baseline +description: Summarize the measurements using the mean. +decisions: + estimator: mean +EOF -x, y = np.loadtxt(args.points, delimiter=",", skiprows=1, unpack=True) -if args.outliers == "clip": - slope, intercept = np.polyfit(x, y, 1) - residuals = y - (slope * x + intercept) - mask = np.abs(residuals) < 3 * residuals.std() - x, y = x[mask], y[mask] -slope, intercept = np.polyfit(x, y, 1) - -Path(args.output).write_text( - json.dumps({"slope": slope, "intercept": intercept, "n_used": len(x)}, indent=2) -) +cat > universes/robust.yaml <<'EOF' +id: robust +description: Summarize the measurements using the median. +decisions: + estimator: median +EOF ``` -Then `src/plot.py` — reads the upstream output's file, makes the -figure: - -```python -import argparse -import json -from pathlib import Path - -import matplotlib - -matplotlib.use("Agg") -import matplotlib.pyplot as plt -import numpy as np +Each file defines a **universe**: one set of methodological choices. Both +use the same data and script. Validate the spec and both universes: -parser = argparse.ArgumentParser() -parser.add_argument("--points", required=True) -parser.add_argument("--fit", required=True) -parser.add_argument("--output", required=True) -args = parser.parse_args() - -x, y = np.loadtxt(args.points, delimiter=",", skiprows=1, unpack=True) -fit = json.loads(Path(args.fit).read_text()) - -fig, ax = plt.subplots() -ax.scatter(x, y, s=12) -xs = np.linspace(x.min(), x.max(), 2) -ax.plot(xs, fit["slope"] * xs + fit["intercept"], color="C1") -ax.set_xlabel("x") -ax.set_ylabel("y") -ax.set_title(f"slope = {fit['slope']:.3f}") -fig.savefig(args.output, dpi=150) +```bash +uvx astra-tools@0.2.18 validate ``` -Both scripts import from the project's locked environment, so declare -what they need: +## 4. Make the results + +Commit the project files before running. Lightcone uses this commit to +identify the analysis that produced each result: ```bash -uv add numpy matplotlib +git add astra.yaml universes/ src/ data/ pyproject.toml uv.lock \ + .python-version .gitignore .gitattributes .datalad/ myst.yml index.md results/README.md +git commit -m "Define the estimator comparison" ``` -That one command updates `pyproject.toml`, re-locks `uv.lock`, and syncs -`.venv`. It's the only way packages reach a recipe — recipes run -sandboxed in the locked environment, so a stray `pip install` on your -machine changes nothing they can see. That's a feature: the lock *is* -the record of what your results were computed with. - -## 5. Materialize - -Commit, then build: +Request a small local allocation, wait until it is ready, and run the +analysis. The built-in offer is configured for 1 CPU and 1 GiB of memory +without a compute configuration file: ```bash -git add -A && git commit -m "Line-fit analysis" -lc materialize +CLUSTER_ID=$(lc compute launch --cpus 1 --memory 1 --json | \ + uv run --locked python -c 'import json, sys; print(json.load(sys.stdin)["id"])') +lc compute status "$CLUSTER_ID" --wait +lc materialize "$CLUSTER_ID" ``` -The commit isn't ceremony — every output is committed together with the -code that produced it, so a build refuses to start from a tree with -uncommitted edits (it wouldn't be able to say what code ran). Then: +Keep this terminal open: `CLUSTER_ID` identifies the allocation for later +commands. If you already configured a compute catalog, the request uses +that catalog's offers instead; inspect them with `lc compute resources`. -``` - ✓ made baseline/fit - ✓ made baseline/fit_plot - ! no [project].license in pyproject.toml, so no RO-Crate publication - view is maintained — declare one to enable it +Lightcone runs the recipe once for each universe and commits each output +with its provenance manifest. It chooses the result paths: -✓ Made 2 output(s) in /home/you/line-fit-demo +```text +results/ +├── baseline/ +│ ├── estimate.json +│ └── .estimate.manifest.json +└── robust/ + ├── estimate.json + └── .estimate.manifest.json ``` -(We'll come back to that license line in step 7.) Each output landed in -`results/baseline/.` next to a -`..manifest.json` — -a manifest recording the recipe, the decisions, the input hashes, the -environment, and the commit — and was committed with a run record that -`datalad rerun` can replay. Look at `git log`: the build wrote history, -not just files. +The message about a missing project license does not prevent execution. +You can add a license when you prepare the project for sharing. -Check where things stand any time: +## 5. Compare and inspect ```bash +cat results/baseline/estimate.json +cat results/robust/estimate.json lc status +lc materialize --check +git log --oneline -3 ``` -``` - mode: direct - sandbox: landlock (fs: declared, network: allowed) - crate: not maintained — declare [project].license to enable it - - · current baseline/fit a3f1f11 - · current baseline/fit_plot a3f1f11 +| Universe | Estimator | Result | +| --- | --- | --- | +| `baseline` | Mean | `5.6` | +| `robust` | Median | `3.0` | -2 current -``` +The difference shows the effect of the large measurement under these two +choices. Both results should be `current`, and `lc materialize --check` +should succeed. Each manifest records the inputs, decisions, command, +environment, and producing commit. -The commit column is the answer to "which code made this?" — for every -output, current or not. And `lc materialize` is idempotent: run it again -and it reports the project is up to date without executing anything. - -## 6. Sweep the decision - -Add the second universe — `universes/robust.yaml`: - -```yaml -id: robust -description: "Points beyond 3 sigma of an initial fit are dropped." -decisions: - outliers: clip -``` - -Commit and materialize again: +Run the same command again to see matching outputs reused, then release +the allocation: ```bash -git add -A && git commit -m "Add the robust universe" -lc materialize -``` - +lc materialize "$CLUSTER_ID" +lc compute down "$CLUSTER_ID" ``` - ✓ made robust/fit - ✓ made robust/fit_plot - · up to date baseline/fit - · up to date baseline/fit_plot -✓ Made 2 output(s) in /home/you/line-fit-demo -``` - -Only the new universe's outputs ran — `baseline` was already exactly -what the spec asks for, so it wasn't touched. Your comparison is on -disk: with this guide's synthetic dataset, clipping drops 4 points and -moves the slope from 2.414 to 2.450 — visibly closer to the true 2.5 -the data was generated with. - -If a recipe fails, `lc materialize` reports which output failed and why, -and leaves the tree as clean as it found it; fix the script or the spec, -commit, and rerun — only the affected outputs re-execute. +Release the allocation with `lc compute down "$CLUSTER_ID"` even if a recipe +fails. You can always inspect allocations with `lc compute status` if you +lose the shell variable. -## 7. Publish +To try a change, edit `data/measurements.txt` and commit that file. Launch +another allocation with the commands in step 4, materialize, and release +it when finished. Both universes will use the new data. -RO-Crate requires a license, so declaring one is how you tell `lc` the -project is meant for the outside world. Add one line under `[project]` -in `pyproject.toml`: - -```toml -license = "CC-BY-4.0" -``` - -then commit and materialize once more: - -```bash -git add -A && git commit -m "Declare a license" -lc materialize -``` +## Continue with your research -Nothing is rebuilt — but `ro-crate-metadata.json` appears at the project -root and is committed automatically. From here on, every materialize -keeps it in line with the repository: the project *is* the crate, and -depositing it is just `git archive` (or `datalad export-archive`) on a -repository you already have. - -## What just happened - -- `astra.yaml` was the only place your analysis was *described* — - inputs, outputs, the decision, and the recipes all live there. -- The scripts take decision values as plain command-line arguments, so - nothing methodological is hardcoded. -- `lc materialize` ran each recipe in the project's locked environment, - sandboxed — free to write the directory its output lands in, and - nothing else — - and committed every output with a manifest and a re-runnable run - record. -- `lc status` and `lc materialize --check` read those manifests — they - don't re-execute anything; they just classify. An output is remade - when the spec defines it differently than it was made, or when its - declared inputs changed; an output whose *environment* has since - moved is reported as `behind` and deliberately left alone — the - manifest records exactly which environment and commit produced it. - -Clone this repository on a fresh machine, run `lc init` (it rebuilds -the two pieces of local state git doesn't carry — the `.venv` and the -annex), then `lc materialize`: it reports up to date without fetching a -single data byte, because the provenance travels in git. The bytes -themselves follow with `git annex get` whenever you actually need them. - -## Where to next - -- [Core Concepts](concepts.md) — the model behind what you just did: - the three states, the commit discipline, the two execution modes. -- [Running on a Cluster](cluster.md) — take the same project to SLURM. -- [Troubleshooting](troubleshooting.md) — when something goes sideways. -- [Glossary](glossary.md) — terms like universe, decision, and manifest - in plain language. -- The [ASTRA docs](https://astra-spec.org/latest/) — the full spec: - sub-analyses, prior insights, findings, and evidence. +- [Work with an agent](agents.md) to scope a larger question or bring in existing code. +- [Core concepts](concepts.md) explains decisions, universes, and result states. +- [ASTRA](astra.md) introduces evidence and links to the external specification. +- [Write a report](reporting.md) connects the narrative to your analysis and results. +- [Share an analysis](sharing.md) explains publication metadata and transferring result files. +- [Running on a cluster](cluster.md) covers larger compute allocations. +- [Troubleshooting](troubleshooting.md) helps with installation and execution errors. diff --git a/docs/user/glossary.md b/docs/user/glossary.md index 537ba98..96cb996 100644 --- a/docs/user/glossary.md +++ b/docs/user/glossary.md @@ -1,17 +1,32 @@ # Glossary -The terms you'll see all over the docs and the `lc` command output, in -plain language. +Terms used across the Lightcone research workflow, its agent skills, and +the `lc` command output. + +## Lightcone + +The research stack that connects agent guidance, analysis implementation, +reproducible execution, and reporting. Its agent skills help you develop a +project; `lightcone-cli` executes it and records provenance. The analysis +specification follows the external ASTRA standard. + +## Agent skill + +Instructions and supporting resources that teach a compatible coding agent +a research workflow. Lightcone's `lightcone` plugin bundles the `lightcone` +and `astra` skills, together with validation hooks. Skills guide the agent; +the CLI and ASTRA tools perform execution and validation. ## ASTRA **A**gentic **S**chema for **T**ransparent **R**esearch **A**nalysis. -The schema lightcone-cli is built around. ASTRA's job is to capture an +An external standard that Lightcone uses. ASTRA's job is to capture an analysis's inputs, outputs, and methodological decisions in a single file (`astra.yaml`); lightcone-cli's job is to execute that spec reproducibly. ASTRA ships separately as the `astra-tools` package, and its `astra` CLI handles the spec itself (validation, universe -management, evidence verification). +management, evidence verification). Its authoritative documentation lives at +[astra-spec.org](https://astra-spec.org/latest/). ## astra.yaml @@ -62,6 +77,20 @@ in dependency order and commits every result as it lands. Idempotent — a second run remakes only what is `stale`, and a run with nothing to do says so and touches nothing. +## Compute allocation + +Resources launched with `lc compute launch` for a bounded lifetime, locally or +on Slurm. The returned cluster ID is required by `lc materialize` and `lc run`. +An execution command borrows the allocation; it stays available until you +release it with `lc compute down` or its lifetime expires. + +## Resource offer + +A CPU and memory shape, node limit, and time limit that your compute catalog +makes available. `lc compute resources` lists offers in selection order. CPU +and memory are per node; an offer describes a requestable shape, not live free +capacity. See [Compute and clusters](cluster.md). + ## Manifest The per-output sidecar JSON file, `..manifest.json` beside @@ -169,8 +198,9 @@ The publication view. Declare a `license` under `[project]` in `pyproject.toml` and every materialize maintains `ro-crate-metadata.json` — a machine-readable description of the project, its outputs, and the runs that produced them, following the -Provenance Run Crate profile. The repository is the crate; deposit is -`git archive`. +Provenance Run Crate profile. The metadata describes the repository. A complete +deposit also needs the actual data and output files: `git archive` alone does +not include git-annex content. ## Prior insight diff --git a/docs/user/index.md b/docs/user/index.md index c2dc24b..d3b77f3 100644 --- a/docs/user/index.md +++ b/docs/user/index.md @@ -1,73 +1,40 @@ -# Welcome to the user guide +# Meet the research stack -`lightcone-cli` is a small toolchain that turns a research question into -a reproducible analysis. You describe what you're trying to learn as a -precise specification — an `astra.yaml` file following the -[**ASTRA**][astra] schema — and the `lc` command line keeps the -resulting code, environments, decisions, and outputs in sync. +Lightcone connects a research question to a record of how you answered it. You describe the analysis, write the code, and run it; the tools keep track of the methodological choices, inputs, environment, and outputs along the way. -ASTRA specs are plain YAML, designed to be easy for both humans and AI -assistants to write. However the spec gets written, **you stay in charge -of the scientific choices** — every methodological decision is declared -in the open, and `lc` records exactly what produced every result: the -recipe, the decisions, the input data, the environment, and the commit. +Start with [installation](install.md) and [your first analysis](getting-started.md). Add the agent plugin if you want an assistant to help you work through the same workflow. -## What this guide covers +## What each piece does -- [Install](install.md) — get the `lc` command line running on your - machine or on a cluster. -- [Getting Started](getting-started.md) — create your first project, - build it end-to-end, and understand what each piece does. -- [Core Concepts](concepts.md) — the model behind the tool: what the - states mean, why everything is committed, and how the two execution - modes differ. -- [Running on a Cluster](cluster.md) — taking your analysis to a SLURM - HPC system. -- [Troubleshooting](troubleshooting.md) — common issues and how to - unstick them. -- [Glossary](glossary.md) — the terms that show up everywhere - (universe, decision, manifest, …) explained in plain language. +| Piece | What you use it for | Where it lives | +| --- | --- | --- | +| **Lightcone CLI** (`lc`) | Create a project, run recipes, inspect result status, and record provenance. | [lightcone-cli](https://github.com/LightconeResearch/lightcone-cli) | +| **Lightcone agent plugin** | Help your coding agent scope a question, maintain a specification, implement recipes, and resume work. Includes the ASTRA skill and validation hooks. | [agent-skills](https://github.com/LightconeResearch/agent-skills) | +| **ASTRA specification and tools** — external | Describe the analysis in `astra.yaml`; validate its structure and inspect its decisions and evidence with `astra`. | [ASTRA documentation ↗](https://astra-spec.org/latest/) · [astra-tools ↗](https://github.com/LightconeResearch/astra-tools) | -## What you'll do, in a handful of lines +ASTRA has its own specification, releases, and documentation. Lightcone builds on it. This site covers [how ASTRA fits into your project](astra.md); the external ASTRA documentation is the reference for the schema and its tools. -!!! tip "Quick start" +The agent plugin is optional. Your specification is ordinary YAML and your recipes run your own scripts, so you can use the same project with or without an assistant. - === "uv" - ```bash - uv tool install lightcone-cli - lc init my-analysis && cd my-analysis - # describe your analysis in astra.yaml, write your scripts, - # declare what they import (uv add numpy ...), then: - git add -A && git commit -m "First analysis" - lc materialize - ``` +## From a question to a result - === "pip" - ```bash - pip install lightcone-cli - lc init my-analysis && cd my-analysis - # describe your analysis in astra.yaml, write your scripts, - # declare what they import (uv add numpy ...), then: - git add -A && git commit -m "First analysis" - lc materialize - ``` +1. **Describe the work.** Record the question, inputs, expected outputs, and methodological decisions in `astra.yaml`. An [agent can help you scope it](agents.md). +2. **Implement the analysis.** Write scripts, add dependencies with `uv add`, and connect each output to a recipe in the specification. +3. **Run and inspect.** Commit your changes, materialize the outputs with `lc`, and inspect their status. Each recorded result carries the information needed to trace how it was produced. +4. **Compare choices.** Create a *universe* for each combination of decision options you want to study. [Core concepts](concepts.md) explains how those alternatives relate to outputs. +5. **Communicate the result.** [Write a report](reporting.md) and [prepare the project for sharing](sharing.md), keeping the result connected to its provenance. -That's the shortest possible path. The rest of the guide is the -unhurried version — and the commit is not ceremony: every output is -committed together with the code that produced it, which is why a build -starts from a clean tree. +You remain responsible for the scientific argument: which alternatives are defensible, whether a method answers the question, and what the results support. A valid specification and a reproducible run make that argument easier to inspect. -## What lightcone-cli is *not* +## Find the help you need -- **A statistics package.** It runs your code; it doesn't compute - things itself. -- **A workflow language.** Recipes in `astra.yaml` are short shell - commands, not a DSL. There's no learning curve beyond what's in - [Getting Started](getting-started.md). -- **An IDE.** `lc` is a command-line tool; write `astra.yaml` and your - analysis code with whatever editor or tooling you prefer. - -If you'd rather skim the design and architecture, the -[maintainer docs](../maintainer.md) are the other half of this site. - -[astra]: https://astra-spec.org/latest/ +| I want to… | Start here | +| --- | --- | +| Get the tools running | [Install](install.md) | +| Follow a complete example | [Your first analysis](getting-started.md) | +| Start or resume with a coding agent | [Work with an agent](agents.md) | +| Understand decisions, universes, and output status | [Core concepts](concepts.md) | +| Run on an HPC system | [Run on a cluster](cluster.md) | +| Resolve an error | [Troubleshooting](troubleshooting.md) | +| Look up a command | [CLI reference](../cli/index.md) | +| Work on the tools themselves | [Contribute](../maintainer.md) | diff --git a/docs/user/install.md b/docs/user/install.md index 48f13ed..1065aa5 100644 --- a/docs/user/install.md +++ b/docs/user/install.md @@ -1,124 +1,132 @@ -# Install +# Install Lightcone -To work on a lightcone project you need two things on your machine: -[uv](https://docs.astral.sh/uv/) and git. Everything else — Python -itself included — is installed by uv or ships with `lc`. +Install the `lc` command, then add the agent plugin if you want to work with +an assistant. You need **git and uv** on Linux or macOS. On Windows, use a +Linux environment under WSL. -!!! note "Supported platforms" - Linux (glibc 2.34+, x86_64 or aarch64) and macOS (14+ on Apple - silicon, 15+ on Intel). On Windows, use WSL. +!!! info "Preview installation" + These docs follow the upcoming explicit compute workflow. The installation + below pins [the preview source](https://github.com/LightconeResearch/lightcone-cli/tree/835de9e7c2df726722ffff8c6863d6f9b16ee577), + which includes `lc compute`. The PyPI releases, including `0.5.0rc4`, + do not yet include that command. -## 1. uv and git +## 1. Install uv -`lc` uses uv as its only environment substrate — projects are -`pyproject.toml` + `uv.lock`, and uv manages the Python interpreters -too, so there is no separate Python install step. +If `uv --version` already reports version **0.12 or newer**, skip this step. +Otherwise, run the [official uv installer](https://docs.astral.sh/uv/getting-started/installation/): -=== "macOS / Linux" - ```bash - curl -LsSf https://astral.sh/uv/install.sh | sh - ``` - - git is preinstalled on macOS; on Linux use your package manager - (`apt install git`, `dnf install git`, …). - -=== "NERSC Perlmutter" - NERSC doesn't ship `uv`, but it installs into your home directory - with a single curl: - - ```bash - curl -LsSf https://astral.sh/uv/install.sh | sh - ``` +```bash +curl -LsSf https://astral.sh/uv/install.sh | sh +``` - `uv` lands under `~/.local/bin` — make sure it's on your `PATH`. - git is already on the system. +Restart your terminal so it can find `uv`, then check that git is available: -## 2. lightcone-cli +```bash +uv --version +git --version +``` -The published name on PyPI is `lightcone-cli`; the command it provides -is `lc`. +If git is missing, follow the [git installation instructions](https://git-scm.com/install/) +for your operating system. uv installs and manages Python for you. -=== "uv" - ```bash - uv tool install lightcone-cli - ``` +## 2. Install Lightcone -=== "pip" - ```bash - python -m pip install lightcone-cli - ``` +```bash +uv tool install 'lightcone-cli @ git+https://github.com/LightconeResearch/lightcone-cli.git@835de9e7c2df726722ffff8c6863d6f9b16ee577' +lc --version +lc compute --help +``` -Get a confirmation of the proper installation by running +This installs `lc` and its dependencies in an isolated tool environment, +including the git-annex executables needed to store project data. There +is no separate Python, ASTRA, or container setup for the first analysis. - lc --version # → lc, version ... +??? info "Supported platforms" + The bundled git-annex wheels require Linux with glibc 2.34 or newer + (x86_64 or aarch64), macOS 14 or newer on Apple silicon, or macOS 15 + or newer on Intel. For Windows, run the installation and analysis + commands inside WSL. -> **Note** Some people may have already set a personal shell alias -> `lc='ls --color'`. If that's you, installing lightcone-cli will shadow -> the alias — make sure to rebind it (e.g. `alias l='ls --color'`). +If your terminal cannot find `lc`, run `uv tool update-shell` and open a +new terminal. If you already have an unrelated shell alias named `lc`, +remove or rename it first. -## 3. Tell git who you are +## 3. Check your git identity -Every output `lc` makes is committed, so git needs an identity before -the first build — `lc materialize` checks up front rather than failing -after your recipes have run: +Lightcone commits the outputs it produces. If you already make git commits +on this machine, you can skip this step. Otherwise, set your own name and +email: ```bash -git config --global user.name "Ada Lovelace" -git config --global user.email "ada@example.org" +git config --global user.name "Your Name" +git config --global user.email "you@example.org" ``` -If you already commit from this machine, you're done. +**You're ready.** Continue to [Your first analysis](getting-started.md), +or add the agent plugin below. -## 4. (Optional) Podman or Docker +## Add an agent (optional) -Only *containerized* projects need a container runtime — a project opts -in by declaring `[tool.lightcone.image]` in its `pyproject.toml`, and -until it does, recipes run directly on your machine in the project's -own locked environment. +Use an existing installation of Claude Code or Codex. Register the Lightcone +marketplace and install the **`lightcone`** plugin: -- Local machine: install [Podman](https://podman.io/) (rootless, no - daemon) or [Docker](https://docs.docker.com/get-docker/). -- HPC login node: see [Running on a Cluster](cluster.md). +=== "Claude Code" -There is nothing to configure: `lc` detects whichever runtime is -available (`podman-hpc`, then `podman`, then `docker` — skipping docker -if its daemon isn't running). + Run in your terminal: -## Sanity check + ```bash + claude plugin marketplace add LightconeResearch/agent-skills + claude plugin install lightcone@lightcone-research + ``` - lc --help - lc init --help + Start a new session and invoke `/lightcone:lightcone`. + See [Claude Code's plugin documentation](https://code.claude.com/docs/en/discover-plugins) + for plugin management. -Both should print help text. If `lc` is shadowed by an `ls` alias, -unset it (`unalias lc`) or use the full path (`$(which lc) --version`). +=== "Codex" -## Updating + Run in your terminal: -=== "uv tool" ```bash - uv tool upgrade lightcone-cli + codex plugin marketplace add LightconeResearch/agent-skills + codex plugin add lightcone@lightcone-research ``` -=== "pip" - ```bash - pip install -U lightcone-cli - ``` + Start a new session and invoke `$lightcone:lightcone`. + If your Codex version does not offer `plugin add`, use its plugin browser + to install `lightcone` from the added marketplace. See + [OpenAI's marketplace documentation](https://developers.openai.com/plugins/build/plugins#add-a-marketplace-from-the-cli) + and [plugin installation guidance](https://developers.openai.com/learn/developers-codex-plugin#install-the-plugin). -An upgrade never invalidates your results: the engine's version is -recorded in every output's manifest, but it is not part of any output's -identity, so nothing gets rebuilt just because `lc` moved. +The plugin bundles the Lightcone and ASTRA skills, including validation +hooks. **Do not install the `astra` plugin alongside it**: that skill is +already included. The plugin uses `uvx` to fetch its pinned ASTRA tools on +first use. -## Uninstalling +The current plugin was written for the earlier CLI release. For the preview's +compute commands, use the workflow in these docs and the installed CLI's +`--help`. Continue to [Work with an agent](agents.md). -=== "uv tool" - ```bash - uv tool uninstall lightcone-cli - ``` +## Optional tools for later -=== "pip" - ```bash - pip uninstall lightcone-cli - ``` +| When you need it | What to add | +| --- | --- | +| A project needs system libraries in a container | Podman or Docker; see [execution environments](concepts.md#two-execution-environments) | +| You want more local resources or a Slurm allocation | A compute catalog; see [Running on a cluster](cluster.md) | +| You want to inspect or validate ASTRA by hand | Use `uvx astra-tools@0.2.18`; see [ASTRA tools](astra.md#validate-and-inspect) | +| You want to preview an interactive report | MyST and its Node.js runtime; see [Write a report](reporting.md) | + +## Update or remove Lightcone + +This preview is pinned so everyone following the tutorial gets the same +CLI. To move to another preview commit or a published release, run +`uv tool install` with that version or source. Check its release notes +before changing the version used for an ongoing project. + +To remove the command: + +```bash +uv tool uninstall lightcone-cli +``` -Your projects are untouched — everything `lc` knows about an analysis -lives in the project's own repository, not in any global state. +Your project files and their git history remain on disk. diff --git a/docs/user/reporting.md b/docs/user/reporting.md new file mode 100644 index 0000000..b54421c --- /dev/null +++ b/docs/user/reporting.md @@ -0,0 +1,89 @@ +# Write a report + +Your Lightcone project includes a place to explain the question, methods, and +results. `lc init` creates `index.md` for the report and `myst.yml` for its +configuration. Write your argument in Markdown, and reference the analysis's +decisions and outputs so the report follows the work as it changes. + +## Preview the scaffold + +Install the [MyST CLI](https://mystmd.org/guide/installing) with a current Node.js +LTS release and npm available: + +```bash +npm install -g mystmd +myst --version +``` + +From your project directory, alongside `astra.yaml`, start the preview: + +```bash +myst start +``` + +Open the address printed in the terminal. Replace the TODO sections in +`index.md` with your introduction, methods, and results. MyST refreshes the +preview when you save Markdown. After changing `astra.yaml`, a universe, or a +result file, save a Markdown page again or restart the preview. + +The generated configuration already loads the ASTRA reporting plugin and article +theme. Keep that configuration when editing the report; there is no separate +plugin installation step. + +## Reference your analysis + +Use the ids from your own `astra.yaml` to mention an output or embed a decision. +For example, the [worked example](getting-started.md) declares `estimator` and +`estimate`: + +````markdown title="index.md" +## Methods + +:::{astra} decisions.estimator +::: + +## Results + +The {astra}`outputs.estimate` output records the result for each estimator. +```` + +For measured values, use a live value reference supported by your output format +instead of copying a number into prose. The external +[report authoring guide](https://lightconeresearch.github.io/MySTRA/authoring/) +covers values, tables, citations, and other reference forms. + +## Check the report against the results + +Before sharing, check the analysis: + +```bash +lc status +lc materialize --check +``` + +If work is needed, commit your edits and run `lc materialize CLUSTER_ID` with +your allocated cluster's id. See [compute setup](cluster.md) if you need an +allocation. The report reads result files from disk; building the report does +not run analysis recipes or retrieve missing annexed results. + +Then build and inspect the report: + +```bash +myst build +myst build --html +``` + +The HTML site is written to `_build/html/`. Read the build diagnostics and check +the preview for unresolved references and missing figures or values. A successful +build alone is insufficient: some references can fall back to plain text. After +renaming an analysis element, update its references in the Markdown too. When you +have several universes, check which one the reporting plugin resolves before +interpreting its figures and numbers. + +The scaffold uses unpinned plugin and theme URLs. For a report you need to +rebuild later, pin their versions in `myst.yml` and record the MyST version used; +the comments in the generated configuration show where to pin them. Reporting +is still evolving, so consult the linked authoring documentation for the version +you use. + +Next: [share your work](sharing.md), including the analysis and its provenance. diff --git a/docs/user/sharing.md b/docs/user/sharing.md new file mode 100644 index 0000000..1e0d96a --- /dev/null +++ b/docs/user/sharing.md @@ -0,0 +1,86 @@ +# Share your work + +A useful research handoff includes the analysis, the files it produced, and the +record of how those files were made. Lightcone can maintain a machine-readable +[RO-Crate](https://www.researchobject.org/ro-crate/) description of that project. +Creating this metadata prepares your work for sharing; depositing it in an +archive or publishing a report is a separate step. + +## Prepare the project + +Review the authorship recorded in `astra.yaml`, the explanation in your +[report](reporting.md), and the project license. To enable RO-Crate maintenance, +add the license you have chosen under the existing `[project]` section of +`pyproject.toml`. For example, for a project you intend to release under CC BY 4.0: + +```toml +license = "CC-BY-4.0" +``` + +Commit your edits, then materialize using an allocated cluster: + +```bash +git add astra.yaml pyproject.toml index.md myst.yml +git commit -m "Prepare the analysis for sharing" +lc materialize CLUSTER_ID +``` + +Replace `CLUSTER_ID` with the id returned by `lc compute launch`; an existing +allocation is fine. See [compute setup](cluster.md) to launch one. Materialize +runs any outputs that need work and updates `ro-crate-metadata.json` in a trailing +commit. If the outputs are already current, they are left in place. + +Check the result: + +```bash +uvx astra-tools@0.2.18 validate +lc materialize --check +lc status +git status --short +``` + +Resolve validation and build failures, confirm `ro-crate-metadata.json` exists, +and read any crate warnings. Review `behind` outputs explicitly: they satisfy +the current analysis definition but were produced under an earlier environment, +and the default check can pass with them present. The [status +reference](../cli/status.md) explains these states. + +## Include the actual data + +Git carries your specification, code, manifests, and history. Large files tracked +by git-annex can be represented in Git by pointers, with their contents stored +elsewhere. A Git clone or `git archive` alone does not guarantee that a recipient +has the input data, result files, or container image bytes. + +For a collaborator continuing the analysis, share the repository **and** an +accessible annex storage location. Tell them where to retrieve the data; a +successful clone only proves that the Git history was transferred. + +For a file archive, first retrieve the content you intend to include, then use +[DataLad's archive exporter](https://docs.datalad.org/en/stable/generated/man/datalad-export-archive.html). +From the project root: + +```bash +git annex get . +uvx datalad export-archive --missing-content error ../analysis.tar.gz +``` + +The exporter runs through `uvx`; DataLad does not need a separate permanent +installation. Keeping `--missing-content error` makes missing file content stop +the export. The archive is written outside the project so it does not become a +new analysis input. Inspect its contents and extract a copy to check the expected +inputs, outputs, manifests, and `ro-crate-metadata.json` before handing it off. + +This is a snapshot of dataset files. Keep the Git repository available separately +when collaborators need commit history or the recorded DataLad run commands. + +## Hand off the work + +Share the checked archive through your chosen repository or archival service, +with the report and the analysis version it describes. Hosting `_build/html/` +shares a readable report; sharing the project and its data lets someone inspect +and continue the analysis. Record the destination and citation in the project +once the deposit exists. + +No Lightcone command above uploads your files or creates a DOI. The local +RO-Crate describes the research object you are preparing to share. diff --git a/docs/user/troubleshooting.md b/docs/user/troubleshooting.md index e7e59f1..c6cfa98 100644 --- a/docs/user/troubleshooting.md +++ b/docs/user/troubleshooting.md @@ -1,199 +1,185 @@ # Troubleshooting -Common situations and how to unstick them, roughly ordered by how often -they come up. `lc`'s refusals try to carry their own remedy — this page -adds the context around them. +Start with the command's error message and `lc --version`. The installation +and examples on this site target the compute-enabled development version; +follow the [installation guide](install.md) if your CLI has a different surface. -## "uncommitted changes in …" +## "lc: command not found" or `lc` prints a directory listing -``` -Error: uncommitted changes in /home/you/my-analysis — every -materialization is committed with the code that produced it, so a run -cannot start from a tree that does not say what that code is. +Check what your shell resolves: - commit these: git add -A . && git commit -m "…" - M src/fit.py +```bash +type lc +uv tool update-shell ``` -Not an error in your project — just the order of operations: commit, -then materialize. The refusal sorts the paths it found: work you own -gets the `commit these` line, while leftover files under `results/` -(from an interrupted run of an older `lc`, or a hand write) are listed -as wreckage to discard instead — `results/` is `lc`'s to write, and -committing hand-placed files there defeats the provenance the tool -exists for. +Open a new shell after updating its configuration. If a personal alias such as +`lc='ls --color'` hides the command, remove it with `unalias lc`. -## "… is not a Lightcone project" +## "No such command 'compute'" -You're outside a project. The current directory *is* the project — `lc` -never walks up to find one, by design — so: +You have a CLI release from before explicit compute allocation. Install the +source version in the [installation guide](install.md), then verify: ```bash -cd path/to/your/project +lc compute --help ``` -or, starting fresh, `lc init my-analysis && cd my-analysis`. If you're -in a fresh clone, run `lc init` once — it rebuilds the `.venv` and the -annex, the two pieces of local state git doesn't carry. +## Missing or unavailable cluster -## "lc: command not found" or `lc` prints a directory listing +Execution requires the ID returned by `lc compute launch`. Check its state before +running the project: -Two possibilities: +```bash +lc compute status "$CLUSTER" --wait +lc materialize "$CLUSTER" +``` -1. The tool isn't on `PATH` — with `uv tool install`, that's - `~/.local/bin`; `uv tool update-shell` fixes the profile. -2. Your shell has a personal alias `lc='ls --color'` shadowing the - real command. Run `type lc` to see; `unalias lc` to remove. +Here `CLUSTER` must contain your allocation's ID. Follow +[Compute and clusters](cluster.md#start-locally) if you have not launched one. +For a read-only check, omit the ID: `lc materialize --check`. -## A recipe fails with "Permission denied" or "No module named …" +A readiness timeout leaves the allocation in place. Inspect its status before +launching another. If it has ended, launch a new allocation. Keep the catalog +connection that created it, and use the same `LC_COMPUTE_CONFIG` for launch and +execution. If submission reports a token with an uncertain result, inspect +existing allocations before retrying. -Every sandboxed failure ends with this trailer: +## "uncommitted changes in …" + +Materialization records the code revision that produced each result. Review +and commit your analysis edits before running: +```bash +git status +git add astra.yaml src/ pyproject.toml uv.lock +git commit -m "Update analysis" ``` -this ran under the lc sandbox (landlock) — a permissions or missing-file -error can mean the command reached for something outside the declared -environment + +Adjust the paths to your own edits. Files under `results/` need separate care: +those are generated outputs, not research code to commit by hand. If an earlier +run was interrupted, stop its allocation and confirm its recipes have stopped +before cleaning partial results. See +[execution limits](cluster.md#execution-requirements-and-limits). + +## "… is not a Lightcone project" + +Project commands inspect the current directory, without searching parent +folders. Move to the directory containing your project: + +```bash +cd path/to/your/project ``` -Recipes run in the project's locked environment, with the tree -read-only apart from the directory their output lands in. The common cases: +For a fresh clone, run `lc init` to reconstruct the environment and initialize +git-annex. To start a new project, use `lc init my-analysis`. + +## A recipe fails with "Permission denied" or "No module named …" -- **`ModuleNotFoundError`** — the package isn't in the project's lock. - `uv add `, commit, re-run. (Installing it on the host with - `pip` changes nothing a recipe sees — that's the point.) -- **Reading a file outside the project** — declare it as an ASTRA - input; declared inputs are readable and their content becomes part - of the output's provenance. -- **Writing outside that directory** — a recipe's product - belongs in `{output}`; for true scratch files, use - `tempfile.mkdtemp()`, which lands in the writable temp area. +Recipes use the project's locked environment and filesystem sandbox. Check +these common causes: -To probe interactively, `lc run ` runs any command under -exactly the isolation a recipe gets — if it works there, it works as a -recipe. +| Symptom | Fix | +|---|---| +| Missing Python package | Run `uv add PACKAGE`, commit the dependency changes, and rerun. | +| Cannot read data outside the project | Declare the file as an ASTRA input; workers also need access at that path. | +| Cannot write a file | Write the product to `{output}` and use `tempfile.mkdtemp()` for scratch files. | + +Probe imports or a script with `lc run "$CLUSTER" -- COMMAND`. A probe uses +the same environment and dependency rules, with a broader write scope under +`results/`. It does not forward interactive input or record result provenance. +Remove probe files before materialization. ## Everything shows `behind` after a `uv add` -Not a problem, and nothing was invalidated. `behind` means: the output -is still exactly what the spec asks for, but the environment has moved -since it was made. Environment changes deliberately don't trigger -rebuilds — the manifest records which environment and commit produced -each output, so nothing is lost by leaving it. When you do want them -remade under the current environment: +`behind` means the result still matches its definition and inputs, while the +environment changed. Lightcone keeps it with its original provenance. To rebuild +behind results under the current environment: ```bash -lc materialize --refresh +lc materialize "$CLUSTER" --refresh ``` -See [Core Concepts](concepts.md) for the `stale` / `behind` -distinction. - ## Everything shows `stale` after a spec edit -`stale` means the spec now defines the output differently than it was -made — you edited its recipe, a decision, or a declared input's -content changed. That's the invalidation model working; the next -`lc materialize` remakes exactly those outputs. +A result becomes `stale` when its recipe, selected decisions, or declared input +content changes. Materialization remakes the affected outputs. A committed manual +edit to a generated output or manifest is also stale. -One edit that deliberately does *not* invalidate: changing your -analysis code (`src/…`). The recipe *string* is the identity, so if -you want code changes to cascade, declare the source file as an ASTRA -input of the outputs it shapes — that choice is yours to make per -output. +Changing a script does not automatically invalidate its outputs. Declare the +source file as an ASTRA input when its content should trigger a rebuild. +See [Core concepts](concepts.md#what-makes-an-output-current). ## "the content is not in this clone" -``` -data/points.csv: the content is not in this clone — git-annex holds a -reference to it, not the data. Fetch it with `git annex get data/points.csv`. -``` - -The clone has the *pointer* to an annexed file but not its bytes. -`lc materialize` fetches the declared inputs it needs by itself; the -read-only verbs (`lc status`, `--check`) never transfer data, so they -report the fact instead. Fetch by hand only when you want the bytes -for your own inspection. - -## "fatal: … clean filter 'annex' failed" +Git-annex has a reference to a file whose bytes are not present locally. +Materialization fetches the declared inputs it needs. Read-only commands do not +transfer data. To inspect a file yourself, fetch it explicitly: -``` -git-annex filter-process: line 1: git-annex: command not found -error: could not read greeting from subprocess 'git-annex filter-process' -error: initialization for subprocess 'git-annex filter-process' failed -fatal: data/catalog.fits: clean filter 'annex' failed +```bash +git annex get data/points.csv ``` -Your shell's `PATH` has no `git-annex`, so git could not run the filter -that turns a large file into an annex pointer. **Nothing was staged**, -which is the point: without `filter.annex.required=true` — which -`lc init` sets — git would have exited 0 and committed the raw bytes -into history instead. +This requires a reachable source that holds the content. A Git remote alone +does not guarantee that the annexed bytes are available. -Once a project holds committed annexed content, this is not limited to -`git add`. Any command that has to run the filter over that content -stops the same way, `git status`, `git diff` and `git checkout` -included — so the whole project reads as broken until git-annex is back -on your `PATH`. That is the intended shape of the failure: a repository -you cannot use is recoverable in one command, and one that quietly -absorbed a multi-gigabyte file is not. +## "fatal: … clean filter 'annex' failed" -`git-annex` ships with `lc`, so a tool install puts both on your `PATH`: +Git cannot find or run git-annex. This can affect `git add`, `git status`, and +other commands that use the annex filter. `lc init` marks the filter as required +so a failure cannot silently put large data files into ordinary Git history. ```bash -uv tool install lightcone-cli git-annex version +uv tool update-shell ``` -If `lc` runs but `git-annex` does not, uv's tool directory is not on -your `PATH` — run `uv tool update-shell` and open a new shell. Running -`lc` through `uvx` puts nothing on your `PATH` at all, so a plain -`git add` cannot work that way. +If git-annex is missing, follow the [installation guide](install.md) to install +Lightcone as a uv tool, then open a new shell. A one-off `uvx` invocation does +not install git-annex onto your shell's `PATH`. -This failure is deliberately loud. `lc init` sets -`filter.annex.required=true` in every project precisely because -without it git handles the same situation by printing the error, -**exiting 0, and staging your data's raw bytes into git history** — -committing a multi-gigabyte dataset into git proper, silently, where -every clone carries it forever. A refused `git add` costs you one -`lc init`; the silent version costs you the repository. +## Selecting compute from a login shell -## "… and this is a NERSC login node" - -`lc materialize` executes recipes, and on centers `lc` recognizes it -refuses to do that on a shared login node. The refusal prints the -center's own `salloc` and `sbatch` spellings — copy one, run the same -command inside the allocation. `lc status`, `lc materialize --check`, -`lc build` and `lc run` work anywhere. See -[Running on a Cluster](cluster.md). +Use your facility's Slurm offers in the compute catalog. The development CLI +does not reject local allocations by inspecting login-node names or site +markers; allocation choices are explicit. Configure the appropriate offers and +follow your facility's usage policy. See [Slurm setup](cluster.md#configure-slurm). ## git doesn't know who you are -Every output is committed, so a machine that has never committed needs -an identity before the first run — `lc materialize` checks up front, -before any recipe spends time: +Set your Git identity before the first result commit: ```bash git config --global user.name "Ada Lovelace" git config --global user.email "ada@example.org" ``` +Use your own name and email. + ## Containerized projects -- **"image absent"** — the declared image hasn't been built and - committed yet: `lc build` (announced by materialize too, which - builds it as a preflight when missing). -- **No runtime found** — install [Podman](https://podman.io/) or - [Docker](https://docs.docker.com/get-docker/); detection is - automatic and there is nothing to configure. -- **Architecture mismatch** — the committed archive records the - architecture it was built for, and a host that can't execute it is - refused before the recipe would have died mid-run. Build on a - matching host (on NERSC, a login node), commit, push, and pull on - the other side. +| Problem | Next step | +|---|---| +| Image absent | Run `lc build`; materialization can also build a missing image as a preflight. | +| No container runtime | Provide Podman, Docker, or podman-hpc on the machines that need it. | +| Architecture mismatch | Build the image on a host with the architecture of the execution workers. | +| Image unavailable on a worker | Make the prepared image available on every worker; the cluster does not distribute node-local image stores automatically. | + +See [`lc build`](../cli/build.md) for the image declaration and lifecycle. + +## Agent skills or ASTRA validation + +If your agent cannot find a skill, check the plugin installation in +[Using an agent](agents.md). The `lightcone` plugin already bundles the ASTRA +skill, so a second ASTRA plugin is unnecessary. + +Schema and evidence-validation errors come from the external ASTRA tooling. +Use [Working with ASTRA](astra.md) to find the authoritative schema reference +and validation commands. ## Filing a bug -Open an issue at -[github.com/LightconeResearch/lightcone-cli/issues](https://github.com/LightconeResearch/lightcone-cli/issues). -Include the output of `lc --version`, the command you ran, and the -full message — the refusals are designed to be pasted. +For execution errors, open an issue in +[lightcone-cli](https://github.com/LightconeResearch/lightcone-cli/issues) with +the version, command, and complete error message. For plugin behavior, use +[agent-skills](https://github.com/LightconeResearch/agent-skills/issues). diff --git a/overrides/404.html b/overrides/404.html new file mode 100644 index 0000000..33377a3 --- /dev/null +++ b/overrides/404.html @@ -0,0 +1,11 @@ +{% extends "main.html" %} + +{% block content %} +

Page not found

+

Use search to find a command or guide, or continue from one of these pages.

+ +{% endblock %} diff --git a/overrides/main.html b/overrides/main.html new file mode 100644 index 0000000..0d12618 --- /dev/null +++ b/overrides/main.html @@ -0,0 +1,16 @@ +{% extends "base.html" %} + +{% block extrahead %} + + +{% endblock %} + +{% block footer %} + + {{ super() }} +{% endblock %} diff --git a/zensical.toml b/zensical.toml index 630d5e6..7528e41 100644 --- a/zensical.toml +++ b/zensical.toml @@ -1,8 +1,8 @@ [project] site_name = "Lightcone Research Stack" -site_description = "Documentation for the Lightcone Research Stack" +site_description = "Install Lightcone, work with a research agent, and build reproducible analyses with ASTRA." site_url = "https://docs.lightconeresearch.org/" -site_author = "Lightcone Research Team" +site_author = "Lightcone Research" repo_url = "https://github.com/LightconeResearch/docs" repo_name = "LightconeResearch/docs" copyright = "© 2026 Lightcone Research" @@ -11,27 +11,39 @@ extra_css = ["stylesheets/extra.css"] nav = [ {"Home" = "index.md"}, - {"User Guide" = [ - {"Welcome" = "user/index.md"}, + {"Start here" = [ + {"Meet the stack" = "user/index.md"}, {"Install" = "user/install.md"}, - {"Getting Started" = "user/getting-started.md"}, - {"Core Concepts" = "user/concepts.md"}, - {"Running on a Cluster" = "user/cluster.md"}, + {"Your first analysis" = "user/getting-started.md"}, + ]}, + {"Guides" = [ + {"Work with an agent" = "user/agents.md"}, + {"Core concepts" = "user/concepts.md"}, + {"Run on a cluster" = "user/cluster.md"}, + {"Write a report" = "user/reporting.md"}, + {"Share your work" = "user/sharing.md"}, {"Troubleshooting" = "user/troubleshooting.md"}, - {"Glossary" = "user/glossary.md"}, ]}, - {"Developer Corner" = [ - {"Welcome" = "maintainer.md"}, - {"Architecture" = "architecture.md"}, - {"CLI Reference" = [ + {"Reference" = [ + {"CLI commands" = [ {"Overview" = "cli/index.md"}, {"lc init" = "cli/init.md"}, + {"lc compute" = "cli/compute.md"}, {"lc materialize" = "cli/materialize.md"}, {"lc status" = "cli/status.md"}, {"lc run" = "cli/run.md"}, {"lc build" = "cli/build.md"}, ]}, - {"Engine Internals" = [ + {"ASTRA · external standard" = "user/astra.md"}, + {"Glossary" = "user/glossary.md"}, + ]}, + {"Contribute" = [ + {"Overview" = "maintainer.md"}, + {"Development setup" = "contributing/setup.md"}, + {"Testing" = "contributing/testing.md"}, + {"Extending the CLI" = "contributing/extending.md"}, + {"CLI architecture" = "architecture.md"}, + {"Engine internals" = [ {"Overview" = "api/index.md"}, {"project" = "api/project.md"}, {"dataset" = "api/dataset.md"}, @@ -40,28 +52,27 @@ nav = [ {"assets" = "api/assets.md"}, {"worker" = "api/worker.md"}, {"materialize" = "api/materialize.md"}, - {"venue" = "api/venue.md"}, + {"compute" = "api/compute.md"}, + {"venue (legacy)" = "api/venue.md"}, {"sandbox" = "api/sandbox.md"}, {"image & container" = "api/container.md"}, {"crate" = "api/crate.md"}, ]}, - {"Contributing" = [ - {"Development Setup" = "contributing/setup.md"}, - {"Testing" = "contributing/testing.md"}, - {"Extending" = "contributing/extending.md"}, - ]}, ]}, - {"ASTRA docs" = "https://astra-spec.org/latest/"}, + {"ASTRA docs ↗" = "https://astra-spec.org/latest/"}, ] [project.theme] variant = "modern" +custom_dir = "overrides" logo = "assets/logo.svg" favicon = "assets/favicon.svg" +font = false features = [ "navigation.tabs", "navigation.sections", "navigation.top", + "navigation.footer", "search.highlight", "content.code.copy", ] @@ -79,3 +90,16 @@ primary = "custom" accent = "custom" toggle.icon = "lucide/moon" toggle.name = "Switch to light mode" + +[project.extra] +generator = false + +[[project.extra.social]] +icon = "lucide/globe" +link = "https://lightconeresearch.org/" +name = "Lightcone Research website" + +[[project.extra.social]] +icon = "fontawesome/brands/github" +link = "https://github.com/LightconeResearch" +name = "Lightcone Research on GitHub"