From 681f156ae4a7ef8154ac724c17124b46da80a7c1 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Tue, 1 Sep 2026 10:05:11 -0400 Subject: [PATCH 01/85] feat(orchestration): the keep list, and the script that acts on it persist_exp.py copies one exposure's named PSF products off /scratch onto the persistent root and records what went, with sizes. The threat it answers is the 60-day purge, not clean_exposure: run_dir is scratch and products_dir is /project, so the only way a per-exposure product outlives its campaign is to leave the filesystem. Exempting files from reclamation would not have done it. The search is recursive beneath the PSF chain's four module output dirs, because setools writes into mask/, rand_split/, new_cat/, plot/ and stat/ rather than flat -- so the config's patterns stay plain file names and the layout stays ours. A pattern that matches nothing is a recorded warning (setools rejects sparse CCDs); nothing matching at all is a failure, since a green manifest over an empty copy is what would let reclamation delete an unsaved exposure. config.yaml's persist_exp: defaults to validation_psf-*.fits -- the psfex_interp VALIDATION catalogue, the rho/tau statistics input, the minimum. The opt-in candidates are documented there with what each buys; sizes are still to be measured. PSFEx residuals and XML are not candidates as the chain stands: the committed default.psfex sets CHECKIMAGE_TYPE NONE and WRITE_XML N. Co-Authored-By: Claude Opus 5 --- workflow/config.yaml | 59 ++++++++++++ workflow/scripts/persist_exp.py | 156 ++++++++++++++++++++++++++++++++ 2 files changed, 215 insertions(+) create mode 100644 workflow/scripts/persist_exp.py diff --git a/workflow/config.yaml b/workflow/config.yaml index 63b29413c..4f85b24e3 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -68,6 +68,65 @@ outputs: # would otherwise have to rebuild from tile headers. index_db: /project/def-mjhudson/cdaley/sp-products/smk-g6/index/run_index.sqlite +# Per-exposure PSF products to COPY onto the persistent root before the scratch +# store goes (`exp_persist`, exposure.smk). A list of plain file-name globs, +# matched recursively under the PSF chain's four module output dirs +# (/exp///output/run_sp_exp_SxSePsfPi/*/output/ — +# sextractor_runner, setools_runner, psfex_runner, psfex_interp_runner). +# Matches land flat in /exp///psf/. +# +# WHY COPY RATHER THAN EXEMPT THESE FROM CLEANUP. Reclamation is not the threat. +# run_dir is /scratch and is PURGED on a 60-day window whether or not +# clean_exposure ever ran; products_dir is /project, backed up and not purged. +# The only way a per-exposure product outlives its campaign is to leave the +# filesystem. (Ordering is free: clean_exposure takes the exp_persist manifest +# as an input, so a store is never reclaimed before its keepers are written.) +# +# EDITING THIS LIST IS CHEAP. It rides on exp_persist's `params`, so a change +# reruns the copy (seconds) and NOT exp_psf (four hours per exposure). That +# separation is the whole reason exp_persist is a rule of its own. +# +# The default is the minimum: the psfex_interp VALIDATION catalogue, one per +# CCD, which is the input to the rho/tau statistics. Without it the PSF +# diagnostics cannot be recomputed after a purge without rebuilding the exposure +# chain from VOS. +# +# OPT-IN CANDIDATES, and what each buys (sizes to be measured): +# star_split_ratio_80-*.fits setools' 80% TRAINING star sample, the set PSFEx +# actually fitted. Refit or perturb the model. +# +# star_split_ratio_20-*.fits the 20% VALIDATION sample — the positions the +# validation_psf rows correspond to, with the +# measured star shapes beside them. +# star_selection-*.fits the PRE-SPLIT selection (setools writes it under +# mask/). The only file that can answer "which +# stars were rejected, and why" — the split +# samples have already lost the rejects. +# +# star_stat-*.txt setools' per-CCD STAT block (star counts, +# stars/deg^2, FWHM mode and cuts, under stat/): +# the selection's summary without its catalogue. +# +# *.psf the PSFEx model itself. Keeping it means the PSF +# can be re-interpolated at ANY position later +# without rebuilding the exposure chain — the +# single most capability-adding entry here. +# +# psfex_cat*.cat PSFEx's own output catalogue (FITS_LDAC). +# +# PSFEx residual/check images and its XML diagnostics are NOT candidates as the +# chain stands: the committed default.psfex sets CHECKIMAGE_TYPE NONE and +# WRITE_XML N, so nothing is emitted to match. They are a config change first, +# a pattern second. +# +# NOTE ON products_dir DEFAULTING TO run_dir (a fixture or smoke test): the copy +# then lands beside the store on the same filesystem and buys nothing, and the +# manifest sits in the exposure's own manifests/ dir, which clean_exposure +# deletes wholesale — so a one-root run re-persists after every reclamation. +# Harmless, and exactly the pre-D5 behaviour a one-root run asks for. +persist_exp: + - validation_psf-*.fits + # Rolling exposure-store reclamation (D5). When true, the COMPUTE DAG grows one # `clean_exposure` job per exposure. It fires once every campaign tile that reads # that exposure has its vignets, deletes the exposure's store AND its manifests, diff --git a/workflow/scripts/persist_exp.py b/workflow/scripts/persist_exp.py new file mode 100644 index 000000000..1a17c40cb --- /dev/null +++ b/workflow/scripts/persist_exp.py @@ -0,0 +1,156 @@ +#!/usr/bin/env python3 +"""Copy ONE exposure's keepable PSF products off scratch, and record what went. + +Run as the shell of the in-DAG ``exp_persist`` rule, never by hand. + +WHY A COPY AND NOT AN EXEMPTION FROM CLEANUP. The obvious alternative — teach +``clean_exposure`` to spare these files — does not work, because reclamation is +not what threatens them. The exposure store lives on ``run_dir``, which is +/scratch: a 60-day purge takes everything there whether or not this workflow +ever cleaned it. ``products_dir`` is /project, backed up and not purged. So the +only way a per-exposure product outlives its campaign is to LEAVE THE +FILESYSTEM, and that is a copy. Reclamation ordering then falls out for free: +``clean_exposure`` takes this rule's manifest as an input, so the store is never +deleted before its keepers have been written elsewhere. + +WHY A SEPARATE RULE AND NOT A ``cp`` APPENDED TO ``exp_psf``. The list of what +to keep is a decision that will be revisited — rho statistics want one file +today, a residual study may want three tomorrow — and ``exp_psf`` is four hours +per exposure. The list rides on this rule's ``params``, so editing it makes +snakemake rerun THIS rule (seconds of cp) and leaves the PSF chain alone. Folded +into ``exp_psf``, the same edit would re-derive every PSF model in the campaign. + +WHAT IT SEARCHES. ``/output/run_sp_exp_SxSePsfPi/*/output/`` — the four +module output dirs of the PSF config (sextractor, setools, psfex, psfex_interp) +— RECURSIVELY. The recursion is not laziness: setools does not write flat, it +writes into ``mask/``, ``rand_split/``, ``new_cat/``, ``plot/`` and ``stat/`` +beneath its own output dir, so a caller who wrote ``star_split_ratio_80-*.fits`` +meaning "the training star sample" would match nothing under a non-recursive +glob. Patterns are therefore plain FILE names and the layout is ours to know, +not the config author's. + +ZERO MATCHES FOR ONE PATTERN IS A WARNING, NOT A FAILURE. setools rejects sparse +CCDs (~0.2% attrition, tolerated by exp_psf's own count floor), so per-CCD +counts are not fixed, and a pattern naming an optional diagnostic may legitimately +find nothing. ZERO FILES IN TOTAL IS A FAILURE: it means the store was not what +we think it is, and writing a green manifest over that would let +``clean_exposure`` delete an exposure whose products were never saved. + +The destination is FLAT — one ``psf/`` dir per exposure, no module subtree — +because the module a file came from is already in its name and the consumer +(rho/tau statistics) globs the directory. A name collision between two modules +is therefore a hard error rather than a silent overwrite; nothing in the current +config can produce one, and if a future one can we want to hear about it. + +The manifest is the rule's ONLY declared output, and it lives on the persistent +root beside the copies (``/exp///manifests/``), NOT in +the exposure's scratch ``manifests/`` dir which ``clean_exposure`` deletes +wholesale. It is deliberately NOT a ``directory()`` output: what was copied, and +how big each file was, is provenance we want written down, and a directory +output attests only that some directory exists. + +It carries no timestamp and is written tmp-then-``cmp``-then-``mv`` (the pattern +``exp_star_cat`` uses), so a rerun that copies the same files leaves the mtime +alone — mtime is a rerun trigger, and an unconditional rewrite would make every +downstream ``clean_exposure`` look out of date once per invocation. +""" + +import argparse +import filecmp +import json +import shutil +import sys +from pathlib import Path + +# The PSF chain's run dir (RUN_NAME in config_exp_psfex.ini). Hardcoded rather +# than passed: this rule persists the PSF stage's products and nothing else, and +# a knob here would be a knob for "persist some other stage", which is a +# different rule. +RUN_NAME = "run_sp_exp_SxSePsfPi" + + +def collect(exp_dir: Path, patterns: list) -> tuple: + """Matched files per pattern, in a stable order, plus the empty patterns.""" + root = exp_dir / "output" / RUN_NAME + found, empty = {}, [] + for pat in patterns: + # One glob per module output dir, recursive beneath it (see the module + # docstring on setools' subdirectories). sorted() over the union keeps + # the manifest byte-stable across filesystem readdir order. + hits = sorted({p for mod in sorted(root.glob("*/output")) + for p in mod.rglob(pat) if p.is_file()}) + if hits: + found[pat] = hits + else: + empty.append(pat) + return found, empty + + +def main() -> None: + p = argparse.ArgumentParser(description=__doc__) + p.add_argument("--exp-dir", required=True, type=Path, + help="the exposure's scratch store") + p.add_argument("--exp", required=True) + p.add_argument("--dest", required=True, type=Path, + help="/exp///psf") + p.add_argument("--manifest", required=True, type=Path) + p.add_argument("--pattern", action="append", default=[], + help="repeatable; a plain file-name glob") + args = p.parse_args() + + if not args.pattern: + sys.exit("persist_exp: no --pattern given (config persist_exp is empty)") + + found, empty = collect(args.exp_dir, args.pattern) + if not found: + sys.exit(f"persist_exp: {args.exp}: no file matched any of " + f"{args.pattern} under {args.exp_dir}/output/{RUN_NAME}") + + args.dest.mkdir(parents=True, exist_ok=True) + seen, files = {}, [] + for pat, hits in found.items(): + for src in hits: + dst = args.dest / src.name + if src.name in seen: + sys.exit(f"persist_exp: {args.exp}: two source files are both " + f"named {src.name} ({seen[src.name]} and {src}); the " + f"destination is flat, so this would silently overwrite") + seen[src.name] = src + # Skip a byte-identical copy: it is not just an I/O saving, it keeps + # the destination's mtimes still for anything downstream that reads + # them. + if not (dst.exists() and filecmp.cmp(src, dst, shallow=False)): + tmp = dst.with_name(dst.name + ".tmp") + shutil.copy2(src, tmp) + tmp.replace(dst) # atomic: no half-copied product + files.append({"name": src.name, "pattern": pat, + "src": str(src), "bytes": dst.stat().st_size}) + + body = { + "stage": "exp_persist", "level": "exp", "unit": args.exp, + "status": "complete", + "dest": str(args.dest), + "patterns": list(args.pattern), + # The warning the docstring argues for: named patterns that matched + # nothing. Present as a key even when empty, so a reader never has to + # wonder whether an old manifest predates the field. + "patterns_unmatched": empty, + "n_files": len(files), + "bytes": sum(f["bytes"] for f in files), + "files": sorted(files, key=lambda f: f["name"]), + } + args.manifest.parent.mkdir(parents=True, exist_ok=True) + tmp = args.manifest.with_name(args.manifest.name + ".tmp") + tmp.write_text(json.dumps(body, indent=2, sort_keys=True) + "\n") + if args.manifest.exists() and filecmp.cmp(tmp, args.manifest, shallow=False): + tmp.unlink() # unchanged: leave the mtime alone + else: + tmp.replace(args.manifest) + + warn = f" ({len(empty)} pattern(s) matched nothing: {empty})" if empty else "" + print(f"[persist_exp] {args.exp}: {len(files)} file(s), " + f"{body['bytes'] / 1e6:.1f} MB -> {args.dest}{warn}") + + +if __name__ == "__main__": + main() From de243a8f47ebfbc7d84854cdbb43880f1c45f066 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Tue, 1 Sep 2026 10:05:25 -0400 Subject: [PATCH 02/85] feat(orchestration): exp_persist, between exp_psf and clean_exposure One rule per exposure, output = ONE manifest on the persistent root at /exp///manifests/exp_persist.json. Not a directory() output: what we want written down is which files were copied and how big each was, and a directory attests only that a directory exists. Byte-stable, so a no-op rerun does not move an mtime clean_exposure reads. Its own rule rather than a cp on the end of exp_psf, and that is the whole point: the keep list rides on params, so adding a pattern reruns seconds of copying instead of four hours of PSF fitting per exposure. A localrule, by the arithmetic that made exp_star_cat one -- a few MB of cp, ~20k of them at DR6 scale, each shorter than the scheduling latency that would submit it. The mid-chain grouping constraint does not bite: its neighbours are exp_psf (too heavy to fuse) and clean_exposure (local already). clean_exposure gains the manifest as an input, so a store is never reclaimed before its keepers have left scratch -- conditional only on there being a keep list, since "keep nothing" must not become a dependency on a rule that would fail for having nothing to copy. rule all requests the persist manifests DIRECTLY, not only through clean_exposure: the purge takes the store whether or not clean: is on, so hanging the copy off reclamation alone would lose everything in a clean:false campaign. Cleaned exposures are excluded -- their exp_psf manifest is gone, so asking would rebuild the chain from VOS, and a tombstone already means the copy happened. Co-Authored-By: Claude Opus 5 --- workflow/Snakefile | 58 ++++++++++++++++++++++++++++- workflow/rules/exposure.smk | 73 ++++++++++++++++++++++++++++++++++++- 2 files changed, 128 insertions(+), 3 deletions(-) diff --git a/workflow/Snakefile b/workflow/Snakefile index bbdfeefb6..a5fa5c54a 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -263,6 +263,7 @@ EXP_DIR = str(RUN_DIR / "exp" / "{shard}" / "{exp}") # The persistent root mirrors the scratch one, shard for shard, so the two trees # read as the same campaign seen from two filesystems. PROD_TILE_DIR = str(PRODUCTS_DIR / "tiles" / "{shard}" / "{tile}") +PROD_EXP_DIR = str(PRODUCTS_DIR / "exp" / "{shard}" / "{exp}") def tile_dir(tile): return f"{RUN_DIR}/tiles/{tile[:2]}/{tile}" @@ -276,6 +277,18 @@ def tile_manifest(tile, stage): def exp_manifest(exp, stage): return f"{exp_dir(exp)}/manifests/{stage}.json" +def prod_exp_dir(exp): + """The exposure's dir on the PERSISTENT root — where exp_persist writes. + + Sharded identically to the scratch one, so the two trees read as the same + campaign seen from two filesystems, exposure side as well as tile side.""" + return f"{PRODUCTS_DIR}/exp/{exp[:2]}/{exp}" + +def prod_exp_manifest(exp, stage): + """A manifest that must SURVIVE reclamation, so it is not in the exposure's + scratch manifests/ dir (clean_exposure deletes that wholesale).""" + return f"{prod_exp_dir(exp)}/manifests/{stage}.json" + def forest_dir(tile): return f"{tile_dir(tile)}/exp_forest" @@ -314,6 +327,7 @@ SCRIPT_HASH = script_hash("completeness.py") FOREST_HASH = script_hash("build_forest.py") CLEAN_HASH = script_hash("clean_exposure.py") CLEAN_TILE_HASH = script_hash("clean_tile.py") +PERSIST_HASH = script_hash("persist_exp.py") # ngmix_range.py earns a hash for a stronger reason than the others. What it # emits is not a stale RESULT but a stale BOUNDARY, and a tile's eight chunks are # a PARTITION of its object IDs: resume a tile across an edit to the split and @@ -431,6 +445,40 @@ def clean_targets(): out.append(tombstone(exp)) return sorted(out) +# --- persisted exposure products (D5) -------------------------------------- +# The keep list is config, not a rule input, and it is READ HERE so that exactly +# one place converts it into the form the rule carries. An empty list is a +# deliberate "keep nothing" and produces no jobs at all. +PERSIST_EXP = list(config.get("persist_exp") or []) + + +def persist_targets(): + """Which exposures this invocation must copy PSF products off scratch for. + + `rule all` requests these DIRECTLY rather than reaching them only through + clean_exposure. Persistence and reclamation are different concerns — the + /scratch purge takes the store whether or not `clean:` is on — and hanging + the copy off the clean rule alone would mean a campaign run with clean:false + persists nothing and loses everything at the purge. + + Scope is the ready tiles' exposures, which `all` already builds through the + tile chain, so nothing new is pulled into the DAG by asking. + + EXCEPT A CLEANED EXPOSURE. Its exp_psf manifest was deleted by + clean_exposure, so requesting its persist manifest would make the DAG + rebuild the whole exposure chain from VOS — the avalanche tile.smk's + reclaimed-edge cut exists to prevent, arriving through a new target instead. + A tombstone means the copy already happened (clean_exposure cannot run + before exp_persist), so there is nothing to ask for. + + HEAD PROCESS ONLY, for the same reason as clean_targets() above. + """ + if not PERSIST_EXP or not workflow.is_main_process: + return [] + exps = {e for t in TILES_READY for e in tile_exposures(t)} + return sorted(prod_exp_manifest(e, "exp_persist") for e in exps + if not Path(tombstone(e)).exists()) + # --- tile reclamation (D5) -------------------------------------------------- # A separate flag from `clean:` (config.yaml carries the full # argument): exposure reclamation costs nothing but a rebuild if a tile is @@ -601,11 +649,19 @@ include: "rules/tile.smk" # localrule would: a local job cannot be fused into a submitted group. The old # star-catalogue rules were exactly that, and they are gone with the internal # mask generation.) -localrules: all, prepare_all_tiles, clean_exposure, clean_tile +# +# exp_persist joins them for the same arithmetic — one tar of a few MB per +# exposure, ~20k of them at DR6 scale, each far shorter than the scheduling +# latency that would submit it (exposure.smk argues the placement in full). It +# sits mid-chain between exp_psf and clean_exposure, but both of those are +# outside every group already (exp_psf is heavy, clean_exposure is local), so it +# adds no new grouping constraint. +localrules: all, prepare_all_tiles, clean_exposure, clean_tile, exp_persist rule all: input: [final_cat(t) for t in TILES_READY], + persist_targets(), clean_targets(), clean_tile_targets(), diff --git a/workflow/rules/exposure.smk b/workflow/rules/exposure.smk index 4e1ef45a4..0bf80b9c6 100644 --- a/workflow/rules/exposure.smk +++ b/workflow/rules/exposure.smk @@ -1,6 +1,6 @@ """Exposure chain — per exposure, keyed by exp base id (dedup is structural). - exp_get_images -> exp_split -> exp_psf + exp_get_images -> exp_split -> exp_psf -> exp_persist Each in the exposure's own sharded work dir, chained by manifests; every config reads fixed ``$SP_RUN/output/run_sp_exp_*`` INPUT_DIRs, so nothing resolves a @@ -17,6 +17,13 @@ the per-band ``MASK_`` columns on the tile side. Neither needs a rule, a star catalogue, or a network fetch — hence no ``star_catalogue`` / ``exp_star_cat`` here, and no ``exp_mask``. +``exp_persist`` is the one rule here that writes to the PERSISTENT root: it +packs the PSF products named by `persist_exp:` into one tar per exposure off +/scratch before the purge (or clean_exposure) can take them. It is a separate +rule from exp_psf precisely so that editing that list costs a re-pack and not a +four-hour refit; the full +argument is in workflow/scripts/persist_exp.py. + NO temp() anywhere in this file, ever (D5). Exposures overlap tiles by construction (~7-10 tiles each), so their consumer set closes over the CAMPAIGN, not over one invocation — reclamation here is clean_exposure's job (S5), driven @@ -106,6 +113,59 @@ rule exp_psf: sp_shell("exp_psf", f"config_exp_{PSF_MODEL}.ini") +# --- persistence (D5) ------------------------------------------------------- +# The counterpart of reclamation, and it must come first in the DAG: this copies +# the exposure's keepable PSF products onto the persistent root, and +# clean_exposure below takes its manifest as an input so the store is never +# reclaimed before the keepers have left /scratch. The purge would take them +# anyway — that, not clean_exposure, is what this rule exists for +# (persist_exp.py's docstring argues both halves, and config.yaml's +# `persist_exp:` block carries the keep list and its candidates). +# +# A LOCALRULE (declared in the Snakefile), by exactly the arithmetic that made +# exp_star_cat one: the body is `cp` of a few MB from one shared filesystem to +# another, seconds of work, and one sbatch per exposure would be ~20k +# submissions at DR6 scale for jobs shorter than the scheduling latency. The +# grouping constraint that binds mid-chain localrules (this file's docstring) +# does not bite here: exp_persist's only neighbours are exp_psf, which is too +# heavy to ever fuse, and clean_exposure, which is local itself. +# +# ONE DECLARED OUTPUT, AND IT IS A MANIFEST, NOT A directory(). The copies are +# not declared: a directory output would attest that a directory exists, where +# what we want written down is WHICH files were copied and how big each was — +# the provenance a rho-statistics run months from now needs in order to know +# what it is reading. The manifest is byte-stable, so a no-op rerun does not +# move its mtime and does not make clean_exposure look out of date. +# +# THE KEEP LIST RIDES ON params. That is the entire reason this is not three +# lines of cp appended to exp_psf's shell: `params` is a rerun trigger, so +# adding a pattern reruns the copy and leaves the PSF chain alone. +rule exp_persist: + input: + rules.exp_psf.output.manifest + output: + manifest = f"{PROD_EXP_DIR}/manifests/exp_persist.json" + # No `log:`: the script's only failure modes are "nothing matched" and a + # name collision, both of which it reports on stderr and neither of which + # has a per-CCD verdict worth a completeness record. + params: + patterns = " ".join(f"--pattern '{p}'" for p in PERSIST_EXP), + exp_dir = lambda wc: exp_dir(wc.exp), + dest = lambda wc: f"{prod_exp_dir(wc.exp)}/psf", + script_hash = PERSIST_HASH + threads: 1 + retries: 2 + resources: + mem_mb = 2000, + runtime = 10 + shell: + "set -euo pipefail\n" + f"python {SCRIPTS}/persist_exp.py" + " --exp-dir '{params.exp_dir}' --exp {wildcards.exp}" + " --dest '{params.dest}' --manifest {output.manifest}" + " {params.patterns}" + + # --- reclamation (D5) ------------------------------------------------------- # The one exception to "no reclamation in this file": clean_exposure OWNS # exposure-level deletion, and it is a real job, not temp() bookkeeping, because @@ -137,7 +197,16 @@ rule clean_exposure: # spatial neighbours. In-scope consumers keep their edge: they may run in # this DAG, so the clean must be ordered after them. lambda wc: [tile_manifest(t, "tile_vignets") - for t in clean_consumers(wc.exp) if t in READY_SET] + for t in clean_consumers(wc.exp) if t in READY_SET], + # The keepers must be off /scratch before the store goes. Unlike the + # consumer edges above, this edge does not depend on scope: it is the + # same exposure's own rule, so it drags nothing into the DAG that this + # exposure's chain did not already put there. It is conditional only on + # there being a keep list at all — with `persist_exp:` empty, "keep + # nothing" is a coherent instruction and must not become a dependency on + # a rule that would fail for having nothing to copy. + lambda wc: ([prod_exp_manifest(wc.exp, "exp_persist")] + if PERSIST_EXP else []) output: tombstone = f"{EXP_DIR}/cleaned.json" params: From 4cf639f75a7295160fbd630aad7fc7251b57f366 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Tue, 1 Sep 2026 10:05:25 -0400 Subject: [PATCH 03/85] docs(orchestration): exp_persist in the rule list, and why the report omits it run_report disk-scans the scratch run_dir, and exp_persist's manifest is the one exposure manifest that lives on products_dir instead -- the placement that makes it survive clean_exposure. Listed in EXP_STAGES it would read as "not run" for every exposure in the campaign, so it is deliberately absent, with the reason on the line. Co-Authored-By: Claude Opus 5 --- workflow/README.md | 17 ++++++++++++++++- workflow/scripts/run_report.py | 7 +++++++ 2 files changed, 23 insertions(+), 1 deletion(-) diff --git a/workflow/README.md b/workflow/README.md index 55611cedc..f4e00c2ea 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -156,7 +156,7 @@ workflow/ bin/sp committed launcher (module load + /project venv + launch code snapshot + run/report/container/cancel) rules/ prepare.smk tile get_images/uncompress/find_exposures - exposure.smk per-exposure: get_images, split, psf (no temp()) + exposure.smk per-exposure: get_images, split, psf, persist (no temp()) tile.smk per-tile: exp forest, merge_headers, detect, vignets, ngmix, merge, make_cat scripts/ sp_rule.py the thin per-unit wrapper (isolation furniture, config copy, log-sync, count check) @@ -165,6 +165,7 @@ workflow/ completeness.py the ported count table (shared by sp_rule + run_report) run_report.py standalone report (NOT a DAG node; run_report hooks call it) container.py image layers + the resolution order behind `sp container` (stdlib-only) + persist_exp.py ONE exposure's keepable PSF products -> products_dir (the exp_persist rule) clean_exposure.py ONE exposure's store + manifests + logs -> tombstone (the clean_exposure rule) profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; keep-going ``` @@ -240,6 +241,20 @@ profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; kee trigger reads that cut as a reason to rerun the very tiles it protects. Know the consequence — `--forcerun` on a tile whose `final_cat` exists will not rebuild its reclaimed exposures. Delete the `final_cat` first. +- **PSF products leave scratch before the purge does.** `exp_persist` copies + the files named by `persist_exp:` in `config.yaml` (default: the psfex_interp + `validation_psf-*.fits`, the rho/tau statistics input) from the exposure's + scratch store into `/exp///psf/`, and writes ONE + manifest beside them recording the patterns, the files and their sizes. The + threat it answers is the /scratch purge, not `clean_exposure` — the store goes + in 60 days whether or not the workflow reclaimed it — so it runs even with + `clean: false`, requested directly by `rule all`. `clean_exposure` takes its + manifest as an input, so reclamation can never overtake the copy. It is a + rule of its own rather than a `cp` on the end of `exp_psf` because the keep + list rides on `params`: adding a pattern reruns seconds of copying, not four + hours of PSF fitting per exposure. A pattern that matches nothing is a + recorded warning (setools rejects sparse CCDs); matching nothing at all is a + failure. A `localrule`, like `exp_star_cat` and for the same arithmetic. - **A dead tile can be told to stop pinning exposures.** An exposure is cleanable only once every consuming tile has its vignets, so one permanently-failed tile holds its ~80 exposures for the life of the diff --git a/workflow/scripts/run_report.py b/workflow/scripts/run_report.py index 24c07670b..11cdcffe4 100644 --- a/workflow/scripts/run_report.py +++ b/workflow/scripts/run_report.py @@ -61,6 +61,13 @@ "tile_ngmix", "tile_merge_cats", "tile_make_cat"] EXP_STAGES = ["exp_get_images", "exp_split", "exp_psf"] +# exp_persist is DELIBERATELY NOT in that list. This report disk-scans the +# scratch run_dir, and exp_persist's manifest is the one exposure manifest that +# lives on products_dir instead — that placement is what makes it survive +# clean_exposure. Listed here it would read as "not run" for every exposure in +# the campaign. Reporting on the persisted products means scanning the second +# root, which is a report this one does not yet do. + # The manifests clean_tile leaves on disk (workflow/scripts/clean_tile.py names # the mechanism that owns each). Their presence is therefore NOT evidence that a # tile's chain was rebuilt, which absorb_tombstones needs to know From 1fa64551675e3c6a6a4b88a6776ba3c1c27e95e6 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 2 Sep 2026 20:16:01 -0400 Subject: [PATCH 04/85] feat(orchestration): exp_persist packs one tar per exposure, not loose copies MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Inodes, not bytes, bind on /project (~1 M-file group quota): smk-m2 measured ~200 loose files per exposure with all candidates on — 25k for 64 tiles, ~2 M at DR6 scale, for 7 GB. persist_exp.py now writes /exp// /psf/.tar (uncompressed, flat members, deterministic: ownership zeroed, sorted, tmp-cmp-mv so a no-op rerun keeps the mtime) and the manifest lists every member. Manifest path, rule wiring and params are unchanged. config.yaml's candidate table carries the smk-m2 per-exposure sizes; psfex_cat and star_stat are marked unmeasured (no live store held them). Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_015qtLUV3bVLPV5p6un7aTFR --- workflow/README.md | 12 ++--- workflow/Snakefile | 2 +- workflow/config.yaml | 57 ++++++++++++++--------- workflow/rules/exposure.smk | 18 ++++---- workflow/scripts/persist_exp.py | 80 ++++++++++++++++++++++----------- 5 files changed, 107 insertions(+), 62 deletions(-) diff --git a/workflow/README.md b/workflow/README.md index f4e00c2ea..f2ce81024 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -165,7 +165,7 @@ workflow/ completeness.py the ported count table (shared by sp_rule + run_report) run_report.py standalone report (NOT a DAG node; run_report hooks call it) container.py image layers + the resolution order behind `sp container` (stdlib-only) - persist_exp.py ONE exposure's keepable PSF products -> products_dir (the exp_persist rule) + persist_exp.py ONE exposure's keepable PSF products -> one tar on products_dir (the exp_persist rule) clean_exposure.py ONE exposure's store + manifests + logs -> tombstone (the clean_exposure rule) profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; keep-going ``` @@ -241,17 +241,19 @@ profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; kee trigger reads that cut as a reason to rerun the very tiles it protects. Know the consequence — `--forcerun` on a tile whose `final_cat` exists will not rebuild its reclaimed exposures. Delete the `final_cat` first. -- **PSF products leave scratch before the purge does.** `exp_persist` copies +- **PSF products leave scratch before the purge does.** `exp_persist` packs the files named by `persist_exp:` in `config.yaml` (default: the psfex_interp `validation_psf-*.fits`, the rho/tau statistics input) from the exposure's - scratch store into `/exp///psf/`, and writes ONE - manifest beside them recording the patterns, the files and their sizes. The + scratch store into ONE uncompressed tar, + `/exp///psf/.tar` (inodes, not bytes, bind + on /project), and writes ONE manifest beside it recording the patterns, the + members and their sizes. The threat it answers is the /scratch purge, not `clean_exposure` — the store goes in 60 days whether or not the workflow reclaimed it — so it runs even with `clean: false`, requested directly by `rule all`. `clean_exposure` takes its manifest as an input, so reclamation can never overtake the copy. It is a rule of its own rather than a `cp` on the end of `exp_psf` because the keep - list rides on `params`: adding a pattern reruns seconds of copying, not four + list rides on `params`: adding a pattern reruns seconds of packing, not four hours of PSF fitting per exposure. A pattern that matches nothing is a recorded warning (setools rejects sparse CCDs); matching nothing at all is a failure. A `localrule`, like `exp_star_cat` and for the same arithmetic. diff --git a/workflow/Snakefile b/workflow/Snakefile index a5fa5c54a..fbba2c851 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -453,7 +453,7 @@ PERSIST_EXP = list(config.get("persist_exp") or []) def persist_targets(): - """Which exposures this invocation must copy PSF products off scratch for. + """Which exposures this invocation must pack PSF products off scratch for. `rule all` requests these DIRECTLY rather than reaching them only through clean_exposure. Persistence and reclamation are different concerns — the diff --git a/workflow/config.yaml b/workflow/config.yaml index 4f85b24e3..ce434314c 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -68,12 +68,17 @@ outputs: # would otherwise have to rebuild from tile headers. index_db: /project/def-mjhudson/cdaley/sp-products/smk-g6/index/run_index.sqlite -# Per-exposure PSF products to COPY onto the persistent root before the scratch +# Per-exposure PSF products to carry onto the persistent root before the scratch # store goes (`exp_persist`, exposure.smk). A list of plain file-name globs, # matched recursively under the PSF chain's four module output dirs # (/exp///output/run_sp_exp_SxSePsfPi/*/output/ — # sextractor_runner, setools_runner, psfex_runner, psfex_interp_runner). -# Matches land flat in /exp///psf/. +# Matches are packed, flat, into ONE uncompressed tar per exposure: +# /exp///psf/.tar, with a manifest listing the +# members beside it. One tar rather than loose copies because inodes, not bytes, +# bind on /project (~1 M-file group quota; loose copies would be ~200 files per +# exposure, ~2 M at DR6 scale). FITS members read straight from the tar: +# fits.open(io.BytesIO(tarfile.open(t).extractfile(m).read())). # # WHY COPY RATHER THAN EXEMPT THESE FROM CLEANUP. Reclamation is not the threat. # run_dir is /scratch and is PURGED on a 60-day window whether or not @@ -83,7 +88,7 @@ outputs: # as an input, so a store is never reclaimed before its keepers are written.) # # EDITING THIS LIST IS CHEAP. It rides on exp_persist's `params`, so a change -# reruns the copy (seconds) and NOT exp_psf (four hours per exposure). That +# reruns the packing (seconds) and NOT exp_psf (four hours per exposure). That # separation is the whole reason exp_persist is a rule of its own. # # The default is the minimum: the psfex_interp VALIDATION catalogue, one per @@ -91,35 +96,43 @@ outputs: # diagnostics cannot be recomputed after a purge without rebuilding the exposure # chain from VOS. # -# OPT-IN CANDIDATES, and what each buys (sizes to be measured): -# star_split_ratio_80-*.fits setools' 80% TRAINING star sample, the set PSFEx -# actually fitted. Refit or perturb the model. -# -# star_split_ratio_20-*.fits the 20% VALIDATION sample — the positions the -# validation_psf rows correspond to, with the -# measured star shapes beside them. +# OPT-IN CANDIDATES, and what each buys. Sizes are per exposure (40 CCDs), +# measured on smk-m2 (127 exposures, 64 tiles); a 64-tile campaign with all of +# the measured ones on came to 7.2 GB: +# validation_psf-*.fits (the default) 2.0 MB +# *.psf the PSFEx model itself. Keeping it means the PSF +# can be re-interpolated at ANY position later +# without rebuilding the exposure chain — the +# single most capability-adding entry here. +# 2.8 MB +# psfex_cat-*.cat PSFEx's own output catalogue (FITS_LDAC): the +# per-star FLAGS_PSF / CHI2_PSF, i.e. WHICH stars +# outlier rejection clipped. Not recoverable from +# anything else (the .psf header keeps only the +# LOADED/ACCEPTED counts). unmeasured # star_selection-*.fits the PRE-SPLIT selection (setools writes it under # mask/). The only file that can answer "which -# stars were rejected, and why" — the split -# samples have already lost the rejects. -# +# stars were rejected by the selection cuts, and +# why" — the split samples have already lost the +# rejects. 24.5 MB +# star_split_ratio_80-*.fits setools' 80% TRAINING star sample, the set PSFEx +# actually fitted. Rows duplicate star_selection. +# 19.9 MB +# star_split_ratio_20-*.fits the 20% VALIDATION sample — the positions the +# validation_psf rows correspond to. Rows +# duplicate star_selection. 7.1 MB # star_stat-*.txt setools' per-CCD STAT block (star counts, # stars/deg^2, FWHM mode and cuts, under stat/): # the selection's summary without its catalogue. -# -# *.psf the PSFEx model itself. Keeping it means the PSF -# can be re-interpolated at ANY position later -# without rebuilding the exposure chain — the -# single most capability-adding entry here. -# -# psfex_cat*.cat PSFEx's own output catalogue (FITS_LDAC). -# +# unmeasured +# A production keep list is `validation_psf` + `*.psf` + `psfex_cat` (~5 MB per +# exposure); the star_split files are only worth it if star_selection is off. # PSFEx residual/check images and its XML diagnostics are NOT candidates as the # chain stands: the committed default.psfex sets CHECKIMAGE_TYPE NONE and # WRITE_XML N, so nothing is emitted to match. They are a config change first, # a pattern second. # -# NOTE ON products_dir DEFAULTING TO run_dir (a fixture or smoke test): the copy +# NOTE ON products_dir DEFAULTING TO run_dir (a fixture or smoke test): the tar # then lands beside the store on the same filesystem and buys nothing, and the # manifest sits in the exposure's own manifests/ dir, which clean_exposure # deletes wholesale — so a one-root run re-persists after every reclamation. diff --git a/workflow/rules/exposure.smk b/workflow/rules/exposure.smk index 0bf80b9c6..47bb4ae26 100644 --- a/workflow/rules/exposure.smk +++ b/workflow/rules/exposure.smk @@ -114,8 +114,8 @@ rule exp_psf: # --- persistence (D5) ------------------------------------------------------- -# The counterpart of reclamation, and it must come first in the DAG: this copies -# the exposure's keepable PSF products onto the persistent root, and +# The counterpart of reclamation, and it must come first in the DAG: this packs +# the exposure's keepable PSF products into one tar on the persistent root, and # clean_exposure below takes its manifest as an input so the store is never # reclaimed before the keepers have left /scratch. The purge would take them # anyway — that, not clean_exposure, is what this rule exists for @@ -123,23 +123,23 @@ rule exp_psf: # `persist_exp:` block carries the keep list and its candidates). # # A LOCALRULE (declared in the Snakefile), by exactly the arithmetic that made -# exp_star_cat one: the body is `cp` of a few MB from one shared filesystem to -# another, seconds of work, and one sbatch per exposure would be ~20k +# exp_star_cat one: the body is a `tar` of a few MB from one shared filesystem +# to another, seconds of work, and one sbatch per exposure would be ~20k # submissions at DR6 scale for jobs shorter than the scheduling latency. The # grouping constraint that binds mid-chain localrules (this file's docstring) # does not bite here: exp_persist's only neighbours are exp_psf, which is too # heavy to ever fuse, and clean_exposure, which is local itself. # -# ONE DECLARED OUTPUT, AND IT IS A MANIFEST, NOT A directory(). The copies are -# not declared: a directory output would attest that a directory exists, where -# what we want written down is WHICH files were copied and how big each was — +# ONE DECLARED OUTPUT, AND IT IS A MANIFEST, NOT THE TAR OR A directory(). The +# tar is not declared: a directory output would attest that a directory exists, +# where what we want written down is WHICH files were packed and how big each was — # the provenance a rho-statistics run months from now needs in order to know # what it is reading. The manifest is byte-stable, so a no-op rerun does not # move its mtime and does not make clean_exposure look out of date. # # THE KEEP LIST RIDES ON params. That is the entire reason this is not three -# lines of cp appended to exp_psf's shell: `params` is a rerun trigger, so -# adding a pattern reruns the copy and leaves the PSF chain alone. +# lines of tar appended to exp_psf's shell: `params` is a rerun trigger, so +# adding a pattern reruns the packing and leaves the PSF chain alone. rule exp_persist: input: rules.exp_psf.output.manifest diff --git a/workflow/scripts/persist_exp.py b/workflow/scripts/persist_exp.py index 1a17c40cb..a23ab150d 100644 --- a/workflow/scripts/persist_exp.py +++ b/workflow/scripts/persist_exp.py @@ -1,5 +1,5 @@ #!/usr/bin/env python3 -"""Copy ONE exposure's keepable PSF products off scratch, and record what went. +"""Pack ONE exposure's keepable PSF products into a tar off scratch, and record what went. Run as the shell of the in-DAG ``exp_persist`` rule, never by hand. @@ -36,21 +36,39 @@ we think it is, and writing a green manifest over that would let ``clean_exposure`` delete an exposure whose products were never saved. -The destination is FLAT — one ``psf/`` dir per exposure, no module subtree — -because the module a file came from is already in its name and the consumer -(rho/tau statistics) globs the directory. A name collision between two modules -is therefore a hard error rather than a silent overwrite; nothing in the current -config can produce one, and if a future one can we want to hear about it. +The manifest lists every member (name, pattern, source path, bytes), so a reader +knows what the tar holds without opening it. + +ONE UNCOMPRESSED TAR PER EXPOSURE, ``/.tar``, NOT LOOSE COPIES. +Inodes, not bytes, are what bind on /project: the group quota is ~1 M files, +and loose per-CCD copies are ~200 per exposure with all candidates on — ~25k for +a 64-tile campaign, ~2 M at DR6 scale, against ~7 GB of bytes. A tar collapses +that to one inode per exposure and costs nothing to read: FITS members go +``tarfile.open(t).extractfile(m).read()`` -> ``fits.open(io.BytesIO(...))``, +which is why a tar rather than a multi-HDU FITS bundle (the keep list mixes +FITS, ``.psf`` and ``.txt``; a FITS container could not hold the last two). +Uncompressed because FITS barely compresses and a plain tar is seekable. + +Members are FLAT — file name only, no module subtree — because the module a +file came from is already in its name and the consumer globs member names. A +name collision between two modules is therefore a hard error rather than a +silent overwrite; nothing in the current config can produce one, and if a +future one can we want to hear about it. + +The tar is written DETERMINISTICALLY (ownership zeroed, members in sorted +order, source mtimes kept), tmp-then-``cmp``-then-``mv``: a rerun over an +unchanged store produces a byte-identical tar and leaves the existing one's +mtime alone. The manifest is the rule's ONLY declared output, and it lives on the persistent -root beside the copies (``/exp///manifests/``), NOT in +root beside the tar (``/exp///manifests/``, beside the tar's ``psf/``), NOT in the exposure's scratch ``manifests/`` dir which ``clean_exposure`` deletes wholesale. It is deliberately NOT a ``directory()`` output: what was copied, and how big each file was, is provenance we want written down, and a directory output attests only that some directory exists. It carries no timestamp and is written tmp-then-``cmp``-then-``mv`` (the pattern -``exp_star_cat`` uses), so a rerun that copies the same files leaves the mtime +``exp_star_cat`` uses), so a rerun that packs the same files leaves the mtime alone — mtime is a rerun trigger, and an unconditional rewrite would make every downstream ``clean_exposure`` look out of date once per invocation. """ @@ -58,8 +76,8 @@ import argparse import filecmp import json -import shutil import sys +import tarfile from pathlib import Path # The PSF chain's run dir (RUN_NAME in config_exp_psfex.ini). Hardcoded rather @@ -92,7 +110,8 @@ def main() -> None: help="the exposure's scratch store") p.add_argument("--exp", required=True) p.add_argument("--dest", required=True, type=Path, - help="/exp///psf") + help="/exp///psf; the tar is " + "/.tar") p.add_argument("--manifest", required=True, type=Path) p.add_argument("--pattern", action="append", default=[], help="repeatable; a plain file-name glob") @@ -107,29 +126,40 @@ def main() -> None: f"{args.pattern} under {args.exp_dir}/output/{RUN_NAME}") args.dest.mkdir(parents=True, exist_ok=True) + tar_path = args.dest / f"{args.exp}.tar" seen, files = {}, [] for pat, hits in found.items(): for src in hits: - dst = args.dest / src.name if src.name in seen: sys.exit(f"persist_exp: {args.exp}: two source files are both " - f"named {src.name} ({seen[src.name]} and {src}); the " - f"destination is flat, so this would silently overwrite") - seen[src.name] = src - # Skip a byte-identical copy: it is not just an I/O saving, it keeps - # the destination's mtimes still for anything downstream that reads - # them. - if not (dst.exists() and filecmp.cmp(src, dst, shallow=False)): - tmp = dst.with_name(dst.name + ".tmp") - shutil.copy2(src, tmp) - tmp.replace(dst) # atomic: no half-copied product + f"named {src.name} ({seen[src.name][0]} and {src}); tar " + f"members are flat, so this would silently overwrite") + seen[src.name] = (src, pat) files.append({"name": src.name, "pattern": pat, - "src": str(src), "bytes": dst.stat().st_size}) + "src": str(src), "bytes": src.stat().st_size}) + files.sort(key=lambda f: f["name"]) + + def anonymous(ti: tarfile.TarInfo) -> tarfile.TarInfo: + # Ownership is the one thing that would differ between two writes of + # the same files from different accounts/nodes; drop it. mtime stays: + # it is the product's, and it is stable while the store is. + ti.uid = ti.gid = 0 + ti.uname = ti.gname = "" + return ti + + tmp = tar_path.with_name(tar_path.name + ".tmp") + with tarfile.open(tmp, "w", format=tarfile.PAX_FORMAT) as tf: + for f in files: + tf.add(seen[f["name"]][0], arcname=f["name"], filter=anonymous) + if tar_path.exists() and filecmp.cmp(tmp, tar_path, shallow=False): + tmp.unlink() # unchanged: leave the mtime alone + else: + tmp.replace(tar_path) # atomic: no half-written archive body = { "stage": "exp_persist", "level": "exp", "unit": args.exp, "status": "complete", - "dest": str(args.dest), + "tar": str(tar_path), "patterns": list(args.pattern), # The warning the docstring argues for: named patterns that matched # nothing. Present as a key even when empty, so a reader never has to @@ -137,7 +167,7 @@ def main() -> None: "patterns_unmatched": empty, "n_files": len(files), "bytes": sum(f["bytes"] for f in files), - "files": sorted(files, key=lambda f: f["name"]), + "files": files, } args.manifest.parent.mkdir(parents=True, exist_ok=True) tmp = args.manifest.with_name(args.manifest.name + ".tmp") @@ -149,7 +179,7 @@ def main() -> None: warn = f" ({len(empty)} pattern(s) matched nothing: {empty})" if empty else "" print(f"[persist_exp] {args.exp}: {len(files)} file(s), " - f"{body['bytes'] / 1e6:.1f} MB -> {args.dest}{warn}") + f"{body['bytes'] / 1e6:.1f} MB -> {tar_path}{warn}") if __name__ == "__main__": From a14a2f36658177bec2cfc6c3ef1a41c87ef7d93f Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 2 Sep 2026 20:19:22 -0400 Subject: [PATCH 05/85] fix(orchestration): persist_exp never leaves a .tmp behind on failure MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit An orphaned tar.tmp on /project is an inode nothing revisits — the leak the tar design exists to avoid. try/finally around both tmp writes. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_015qtLUV3bVLPV5p6un7aTFR --- workflow/scripts/persist_exp.py | 33 +++++++++++++++++++++------------ 1 file changed, 21 insertions(+), 12 deletions(-) diff --git a/workflow/scripts/persist_exp.py b/workflow/scripts/persist_exp.py index a23ab150d..c13fca604 100644 --- a/workflow/scripts/persist_exp.py +++ b/workflow/scripts/persist_exp.py @@ -147,14 +147,20 @@ def anonymous(ti: tarfile.TarInfo) -> tarfile.TarInfo: ti.uname = ti.gname = "" return ti + # tmp-then-cmp-then-mv, and the tmp NEVER outlives a failure: an orphaned + # .tmp on /project is an inode nothing revisits — the leak this whole tar + # design exists to avoid, one per failed attempt at DR6 scale. tmp = tar_path.with_name(tar_path.name + ".tmp") - with tarfile.open(tmp, "w", format=tarfile.PAX_FORMAT) as tf: - for f in files: - tf.add(seen[f["name"]][0], arcname=f["name"], filter=anonymous) - if tar_path.exists() and filecmp.cmp(tmp, tar_path, shallow=False): - tmp.unlink() # unchanged: leave the mtime alone - else: - tmp.replace(tar_path) # atomic: no half-written archive + try: + with tarfile.open(tmp, "w", format=tarfile.PAX_FORMAT) as tf: + for f in files: + tf.add(seen[f["name"]][0], arcname=f["name"], filter=anonymous) + if tar_path.exists() and filecmp.cmp(tmp, tar_path, shallow=False): + tmp.unlink() # unchanged: leave the mtime alone + else: + tmp.replace(tar_path) # atomic: no half-written archive + finally: + tmp.unlink(missing_ok=True) body = { "stage": "exp_persist", "level": "exp", "unit": args.exp, @@ -171,11 +177,14 @@ def anonymous(ti: tarfile.TarInfo) -> tarfile.TarInfo: } args.manifest.parent.mkdir(parents=True, exist_ok=True) tmp = args.manifest.with_name(args.manifest.name + ".tmp") - tmp.write_text(json.dumps(body, indent=2, sort_keys=True) + "\n") - if args.manifest.exists() and filecmp.cmp(tmp, args.manifest, shallow=False): - tmp.unlink() # unchanged: leave the mtime alone - else: - tmp.replace(args.manifest) + try: + tmp.write_text(json.dumps(body, indent=2, sort_keys=True) + "\n") + if args.manifest.exists() and filecmp.cmp(tmp, args.manifest, shallow=False): + tmp.unlink() # unchanged: leave the mtime alone + else: + tmp.replace(args.manifest) + finally: + tmp.unlink(missing_ok=True) warn = f" ({len(empty)} pattern(s) matched nothing: {empty})" if empty else "" print(f"[persist_exp] {args.exp}: {len(files)} file(s), " From d91da36ce11731fc65b256aed761b010c8764499 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 08:58:57 -0400 Subject: [PATCH 06/85] docs(orchestration): drop references to the removed star-catalogue rules develop no longer has exp_star_cat / star_catalogue (PR #847); the three comments that cited exp_star_cat as the localrule precedent now cite clean_exposure, which makes the same argument. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- workflow/README.md | 2 +- workflow/rules/exposure.smk | 2 +- workflow/scripts/persist_exp.py | 2 +- 3 files changed, 3 insertions(+), 3 deletions(-) diff --git a/workflow/README.md b/workflow/README.md index f2ce81024..2b6745eb8 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -256,7 +256,7 @@ profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; kee list rides on `params`: adding a pattern reruns seconds of packing, not four hours of PSF fitting per exposure. A pattern that matches nothing is a recorded warning (setools rejects sparse CCDs); matching nothing at all is a - failure. A `localrule`, like `exp_star_cat` and for the same arithmetic. + failure. A `localrule`, by the same arithmetic as `clean_exposure`. - **A dead tile can be told to stop pinning exposures.** An exposure is cleanable only once every consuming tile has its vignets, so one permanently-failed tile holds its ~80 exposures for the life of the diff --git a/workflow/rules/exposure.smk b/workflow/rules/exposure.smk index 47bb4ae26..40a4044d4 100644 --- a/workflow/rules/exposure.smk +++ b/workflow/rules/exposure.smk @@ -123,7 +123,7 @@ rule exp_psf: # `persist_exp:` block carries the keep list and its candidates). # # A LOCALRULE (declared in the Snakefile), by exactly the arithmetic that made -# exp_star_cat one: the body is a `tar` of a few MB from one shared filesystem +# clean_exposure one: the body is a `tar` of a few MB from one shared filesystem # to another, seconds of work, and one sbatch per exposure would be ~20k # submissions at DR6 scale for jobs shorter than the scheduling latency. The # grouping constraint that binds mid-chain localrules (this file's docstring) diff --git a/workflow/scripts/persist_exp.py b/workflow/scripts/persist_exp.py index c13fca604..e78d3e9a4 100644 --- a/workflow/scripts/persist_exp.py +++ b/workflow/scripts/persist_exp.py @@ -68,7 +68,7 @@ output attests only that some directory exists. It carries no timestamp and is written tmp-then-``cmp``-then-``mv`` (the pattern -``exp_star_cat`` uses), so a rerun that packs the same files leaves the mtime +``clean_exposure`` uses), so a rerun that packs the same files leaves the mtime alone — mtime is a rerun trigger, and an unconditional rewrite would make every downstream ``clean_exposure`` look out of date once per invocation. """ From 626aabacf6f2f1f9013f2bd874366838f4d2e793 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 10:08:36 -0400 Subject: [PATCH 07/85] feat(orchestration): the two campaign-level merges MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Every rule in this workflow was per unit, and the two products downstream analysis actually opens are per campaign. So a run ended one file short on each side and both merges were a manual pass afterwards. These are the last links of chains the workflow already had. star_cat_merge stacks every exposure's every CCD's validation_psf into one /full_starcat-0000000.fits — the rho/tau statistics input, at the path sp_validation hardcodes. It reads the members straight out of the per-exposure tars exp_persist wrote (tarfile + BytesIO); unpacking ~800k files to merge them would defeat the tar's whole purpose. The stacking is MergeStarCatPSFEX, the class the old merge_starcat_runner called, so the column list keeps exactly one definition. That class gains one thing: an entry may be [fileobj, name] rather than [path], fits.open taking the first element and the CCD_NB regex the last — the same string for the runner's one-element entries. final_cat_merge collects every ready tile's final_cat into /final_cat_.hdf5: one dataset per tile under a group named for the campaign, the final_cat.param columns, an n_tiles attribute. That schema is sp_validation's reader's, so it is fixed. The COLUMN EXTRACTION reuses scripts/python/create_final_cat.py (read_param_file, read_data, copy_data) so the column grammar keeps one definition; the file is written here, because that script's own discovery walks a directory layout this workflow does not have and groups by a unit ShapePipe v2 has dropped. Two places where the reference implementation is not reproducible are pinned down at the call site rather than copied: copy_data leaves every non-requested column as uninitialised memory, and read_param_file's column order varies with the process hash seed. bin/sp now snapshots the repo's scripts/ so the campaign pins that file like everything else it runs. Neither merge puts its input paths in its shell: ~20k of them is an order of magnitude over Linux's 128 KiB MAX_ARG_STRLEN for one argv entry. Each job is handed the two small files the Snakefile itself started from — the tile list and the run index — and DERIVES the same set from them, through readers that now live in build_index.py beside the schema. The rule's params carries that set's fingerprint, which is the rerun trigger, and the equality of the two sides is what makes the trigger mean anything: a glob over products_dir would merge tiles or exposures from an earlier, larger tile list sharing the root, rows no trigger could see. star_cat_merge depends on a live exposure through its exp_persist manifest and on a RECLAIMED one through its tar, which no rule declares and which therefore requires nothing to be built. Requesting a reclaimed exposure's manifest instead rebuilds its whole chain from VOS, and ancient() does not prevent that: measured on smk-g6 with one reclaimed exposure given a manifest by hand, the dry run grew exp_get_images, exp_split, exp_psf and exp_persist jobs. Reclaimed exposures belong in the star catalogue — carrying their PSF products off scratch is what exp_persist is for. Both rebuild rather than append, so the output is a function of its input set: byte-stable on a no-op rerun (tmp-then-cmp-then-mv), rebuilt when a unit is appended. Neither is a localrule — one job over ~20k units is real work. star_cat_merge produces no job, and a parse-time warning rather than a runtime failure, when persist_exp keeps no validation_psf-shaped file. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- .../merge_starcat_package/merge_starcat.py | 21 +- workflow/README.md | 42 +++- workflow/Snakefile | 180 +++++++++++++- workflow/bin/sp | 9 +- workflow/config.yaml | 23 ++ workflow/rules/exposure.smk | 67 +++++ workflow/rules/tile.smk | 60 +++++ workflow/scripts/build_index.py | 39 +++ workflow/scripts/merge_final_cat.py | 222 +++++++++++++++++ workflow/scripts/merge_star_cat.py | 234 ++++++++++++++++++ 10 files changed, 886 insertions(+), 11 deletions(-) create mode 100644 workflow/scripts/merge_final_cat.py create mode 100644 workflow/scripts/merge_star_cat.py diff --git a/src/shapepipe/modules/merge_starcat_package/merge_starcat.py b/src/shapepipe/modules/merge_starcat_package/merge_starcat.py index 7bc0cb76b..1717825f9 100644 --- a/src/shapepipe/modules/merge_starcat_package/merge_starcat.py +++ b/src/shapepipe/modules/merge_starcat_package/merge_starcat.py @@ -524,7 +524,15 @@ class MergeStarCatPSFEX(object): Parameters ---------- input_file_list : list - Input files + Input entries. Each entry is a list, as the module runner builds them: + ``[path]`` from the file handler. An entry may also carry a name + alongside an already-open source, ``[fileobj, name]`` — ``fits.open`` + takes the first element and the CCD number is parsed from the LAST, + which is the same string in the one-element case. That is what lets a + caller merge catalogues it never wrote to disk (the Snakemake + workflow's ``star_cat_merge`` reads them out of the per-exposure tars + with ``tarfile`` + ``BytesIO``), without this class learning anything + about where they came from. output_dir : str Output directory w_log : logging.Logger @@ -569,10 +577,15 @@ def process(self): ) for name in self._input_file_list: + # The source to read and the NAME to parse the CCD number out of. + # Identical for a plain [path] entry; different only when the caller + # hands over an open file-like object plus the member name it came + # under (see the class docstring). + source, label = name[0], name[-1] try: - starcat_j = fits.open(name[0], memmap=False, ignore_missing_simple=True) + starcat_j = fits.open(source, memmap=False, ignore_missing_simple=True) except OSError as e: - print(f"Error while opening file '{name[0]}'") + print(f"Error while opening file '{label}'") #raise continue @@ -614,7 +627,7 @@ def process(self): psfex_acc += list(np.zeros_like(data_j["X"])) # CCD number - ccd_nb += [re.split(r"\-([0-9]*)\-([0-9]+)\.", name[0])[-2]] * len( + ccd_nb += [re.split(r"\-([0-9]*)\-([0-9]+)\.", label)[-2]] * len( data_j["RA"] ) diff --git a/workflow/README.md b/workflow/README.md index 2b6745eb8..824eddc73 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -107,7 +107,9 @@ and the run fails if either phase failed. ## The launch code snapshot `sp run` copies the code it is about to launch — `workflow/` (config symlinks -dereferenced), `src/` and the profile — into `/code`, records HEAD +dereferenced), `src/`, the repo's `scripts/` (`final_cat_merge` loads +`scripts/python/create_final_cat.py` by path) and the profile — into +`/code`, records HEAD plus a dirty flag in `/code/snapshot.json`, and runs the campaign entirely out of that copy. It matters because a campaign is not one process: the SLURM executor re-invokes snakemake on every job's node, so jobs re-parse the @@ -156,8 +158,8 @@ workflow/ bin/sp committed launcher (module load + /project venv + launch code snapshot + run/report/container/cancel) rules/ prepare.smk tile get_images/uncompress/find_exposures - exposure.smk per-exposure: get_images, split, psf, persist (no temp()) - tile.smk per-tile: exp forest, merge_headers, detect, vignets, ngmix, merge, make_cat + exposure.smk per-exposure: get_images, split, psf, persist (no temp()); campaign star_cat_merge + tile.smk per-tile: exp forest, merge_headers, detect, vignets, ngmix, merge, make_cat; campaign final_cat_merge scripts/ sp_rule.py the thin per-unit wrapper (isolation furniture, config copy, log-sync, count check) build_index.py prepare-phase run_index.sqlite builder (plain script) @@ -166,6 +168,8 @@ workflow/ run_report.py standalone report (NOT a DAG node; run_report hooks call it) container.py image layers + the resolution order behind `sp container` (stdlib-only) persist_exp.py ONE exposure's keepable PSF products -> one tar on products_dir (the exp_persist rule) + merge_star_cat.py ALL exposures' validation_psf, read out of the tars -> full_starcat (the star_cat_merge rule) + merge_final_cat.py ALL tiles' final_cat -> final_cat_.hdf5 (the final_cat_merge rule) clean_exposure.py ONE exposure's store + manifests + logs -> tombstone (the clean_exposure rule) profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; keep-going ``` @@ -257,6 +261,38 @@ profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; kee hours of PSF fitting per exposure. A pattern that matches nothing is a recorded warning (setools rejects sparse CCDs); matching nothing at all is a failure. A `localrule`, by the same arithmetic as `clean_exposure`. +- **The campaign ends in two merged catalogues, and the workflow now makes + both.** Everything above is per unit; the two products downstream analysis + actually opens are per *campaign*, and until these rules existed each was a + manual pass after the run. + `star_cat_merge` stacks every exposure's every CCD's `validation_psf-*.fits` + into one `/full_starcat-0000000.fits` — the rho/tau statistics + input, at the path sp_validation hardcodes. It reads the members straight out + of the per-exposure tars (`tarfile` + `BytesIO`; unpacking ~800k files to + merge them would defeat the tar's whole purpose) and stacks them with + `MergeStarCatPSFEX`, the same class the old `merge_starcat_runner` called, so + the column list has exactly one definition. Its input is the same + `exp_persist` manifest set `rule all` already requests, so it pulls nothing + new into the DAG, and it exists only when `persist_exp:` keeps a + `validation_psf-*.fits`-shaped file — otherwise no job, and a warning at parse + time rather than a failure on a node. + `final_cat_merge` collects every ready tile's `final_cat-.fits` into + `/final_cat_.hdf5`: one dataset per tile under a group + named for the campaign, the `final_cat.param` columns, an `n_tiles` attribute. + That schema is what sp_validation's reader opens, so it is fixed; the column + extraction reuses `scripts/python/create_final_cat.py` while the file is + written here, because that script's own discovery walks a directory layout + this workflow does not have. `campaign:` in `config.yaml` names the group and + defaults to the persistent root's basename. + Both rebuild from the whole persistent root rather than appending, so the + output is a function of its input set: byte-stable on a no-op rerun + (tmp-then-`cmp`-then-`mv`), and rebuilt when a tile or exposure is appended + (the input list's fingerprint rides on `params`). Neither is a `localrule` — + one job over ~20k units is real work — and neither puts its input paths in its + shell, which is not fastidiousness: ~20k paths is an order of magnitude over + Linux's 128 KiB `MAX_ARG_STRLEN` for a single argv entry, so each script + rediscovers the set under `products_dir` while the fingerprint travels on + `params`. - **A dead tile can be told to stop pinning exposures.** An exposure is cleanable only once every consuming tile has its vignets, so one permanently-failed tile holds its ~80 exposures for the life of the diff --git a/workflow/Snakefile b/workflow/Snakefile index fbba2c851..2954222bc 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -31,6 +31,7 @@ the manifest says "this stage succeeded", the log says "here is what happened" (the contract is argued in completeness.py's docstring). """ +import fnmatch import functools import hashlib import json @@ -40,6 +41,10 @@ import sys from pathlib import Path from snakemake.exceptions import WorkflowError +# Explicit rather than relying on the name snakemake injects into this +# namespace: the one place we log at parse time is a branch that only a +# non-default keep list reaches, and a NameError there would be found by a user. +from snakemake.logging import logger # Resolved relative to THIS file, not the working directory: snakemake runs with # --directory on /scratch (bin/sp) so .snakemake/ state never lands on /project @@ -92,6 +97,14 @@ RUN_DIR = Path(OUTPUTS["run_dir"]) # second path: one root, exactly the pre-D5 layout. PRODUCTS_DIR = Path(OUTPUTS.get("products_dir") or RUN_DIR) INDEX_DB = Path(OUTPUTS["index_db"]) +# The campaign's NAME — what the two campaign-level merges label their output +# with (`final_cat_.hdf5`, and the group inside it that holds the +# campaign's per-tile datasets). It +# defaults to the persistent root's basename, which is already how every +# campaign here is named (smk-g4, smk-g5, smk-g6: run_dir, products_dir and +# index all end in it), so the common case needs no key at all. Set `campaign:` +# in config.yaml when the two must differ. +CAMPAIGN = config.get("campaign") or PRODUCTS_DIR.name SCRIPTS = Path(workflow.basedir) / "scripts" # The config chain is the repo's committed directory (D2). The configs and # rules that set their environment variables must be versioned together. There @@ -328,6 +341,8 @@ FOREST_HASH = script_hash("build_forest.py") CLEAN_HASH = script_hash("clean_exposure.py") CLEAN_TILE_HASH = script_hash("clean_tile.py") PERSIST_HASH = script_hash("persist_exp.py") +MERGE_STAR_HASH = script_hash("merge_star_cat.py") +MERGE_FINAL_HASH = script_hash("merge_final_cat.py") # ngmix_range.py earns a hash for a stronger reason than the others. What it # emits is not a stale RESULT but a stale BOUNDARY, and a tile's eight chunks are # a PARTITION of its object IDs: resume a tile across an edit to the split and @@ -473,12 +488,173 @@ def persist_targets(): HEAD PROCESS ONLY, for the same reason as clean_targets() above. """ - if not PERSIST_EXP or not workflow.is_main_process: + if not workflow.is_main_process: + return [] + return persist_manifests() + + +@functools.lru_cache(maxsize=1) +def persist_manifests(): + """persist_targets() without the head-process guard, memoised. + + The guard on persist_targets() is a cost decision, not a correctness one: + `rule all` reads it at MODULE level, so a job parse would pay a whole + campaign's index walk for a target it can never schedule. star_cat_merge + reads the same list through an INPUT FUNCTION, which snakemake evaluates + only for parses that actually build that job — the head process, and the one + merge job's own re-parse under the slurm executor, which genuinely needs it. + So this half carries no guard and the memo keeps either parse to one walk. + """ + if not PERSIST_EXP: return [] exps = {e for t in TILES_READY for e in tile_exposures(t)} return sorted(prod_exp_manifest(e, "exp_persist") for e in exps if not Path(tombstone(e)).exists()) +# --- the campaign-level merges --------------------------------------------- +# Two rules, one job each per campaign, both writing to the persistent root, and +# both the LAST link of a chain whose per-unit half the workflow already had: +# the exposure side ends in one `full_starcat-0000000.fits` (every CCD's PSF +# validation catalogue, stacked — the rho/tau statistics input) and the tile side +# in one `final_cat_.hdf5` (every tile's final catalogue — the shear +# catalogue sp_validation reads). Until they existed the workflow's product set +# was two files short of what the old `combine_runs.bash` + `create_final_cat.py` +# chain delivered, and every campaign ended with a manual merge. +# +# NEITHER IS A LOCALRULE, and the arithmetic runs the opposite way from +# exp_persist's. Those rules are ~20k jobs of seconds each, so submitting them +# costs more in scheduling latency than the work; these are ONE job each over the +# whole campaign — ~800k catalogues stacked in memory, or ~20k catalogues read +# end to end at DR6 scale. That is a compute job, and it belongs on a node. +# +# NEITHER PUTS ITS INPUT PATHS IN ITS SHELL. `{input}` at DR6 scale is ~20k paths +# in a single argv entry, an order of magnitude over Linux's 128 KiB +# MAX_ARG_STRLEN, and the job would die on exec. So each rule's `input` is the +# DAG EDGE (what must exist first) and each script rediscovers the same set under +# products_dir; what travels is a FINGERPRINT of the input list, on `params`, +# which is what makes the merge rerun when the set changes and not otherwise. +# The scripts' docstrings argue the rediscovery — it is also what lets the merges +# cover exposures whose scratch stores reclamation has since taken. + + +def input_fingerprint(paths): + """A short digest of an input list, for `params`. + + `params` is a rerun trigger and a path list is not: two campaigns' worth of + manifests hashes to two different values, so appending a tile reruns the + merge, while a rerun over the same set leaves it alone. Sorted before + hashing because the list's ORDER is not part of what changed. + """ + joined = "\n".join(sorted(str(p) for p in paths)) + return f"{len(paths)}:{hashlib.md5(joined.encode()).hexdigest()[:12]}" + + +# The tar member the star merge consumes. The keep list is globs, so the test is +# "would this member be kept", not a string comparison — `validation_psf-*.fits`, +# `validation_psf*`, `*.fits` and a bare `*` all say yes, and all are things a +# user might reasonably write. +_STAR_CAT_MEMBER = "validation_psf-2605805-12.fits" +STAR_CAT_MERGE = any(fnmatch.fnmatch(_STAR_CAT_MEMBER, p) for p in PERSIST_EXP) + +# LOUD AT PARSE TIME, once, and only where it can be acted on. A keep list +# without the validation catalogues is a legitimate configuration (persist the +# PSF models alone, say) — it is not an error, so it must not become a job that +# fails on a node an hour later. It is worth SAYING, because the omission is +# silent otherwise and the missing product only surfaces when a rho-statistics +# run cannot find its input. +if PERSIST_EXP and not STAR_CAT_MERGE and workflow.is_main_process \ + and PHASE == "compute": + logger.warning( + f"star_cat_merge: no job — persist_exp {PERSIST_EXP} keeps no " + f"'{_STAR_CAT_MEMBER}'-shaped file, so there is nothing to stack into " + f"{PRODUCTS_DIR}/full_starcat-0000000.fits (the rho/tau statistics " + f"input). Add 'validation_psf-*.fits' to persist_exp: to get it.") + + +def full_starcat(): + """The campaign's merged star catalogue. The NAME is not ours to choose: + sp_validation hardcodes `full_starcat-0000000.fits` beside its data dir.""" + return f"{PRODUCTS_DIR}/full_starcat-0000000.fits" + + +def final_cat_hdf5(): + """The campaign's merged shear catalogue — sp_validation's galaxy_cat_path.""" + return f"{PRODUCTS_DIR}/final_cat_{CAMPAIGN}.hdf5" + + +def prod_exp_tar(exp): + """The tar exp_persist writes. Not a declared output of anything — see + star_cat_inputs().""" + return f"{prod_exp_dir(exp)}/psf/{exp}.tar" + + +@functools.lru_cache(maxsize=1) +def star_cat_inputs(): + """What star_cat_merge waits for: every exposure of TILES_READY whose PSF + products are on the persistent root, live and reclaimed alike. + + THE SAME SET merge_star_cat.py derives at job time, and that equality is + load-bearing — the fingerprint on `params` is taken over THIS list, so + anything the job stacked that was not in it would be rows no rerun trigger + could see. The job states the rule from its own side: same tile list, same + index, exp_persist manifest present on the persistent root. By the time it + runs, every exposure below has one. + + RECLAIMED EXPOSURES BELONG IN THE STAR CATALOGUE. Carrying their PSF + products off scratch is exactly what exp_persist is for, and a merge that + dropped them would shrink the campaign's star catalogue every time + reclamation ran. But their exp_psf manifest is gone, so REQUESTING their + exp_persist manifest rebuilds the whole exposure chain from VOS — the + avalanche persist_targets() drops them to avoid. + ancient() DOES NOT HELP: it suppresses the timestamp comparison, not the + missing input, and snakemake schedules the chain anyway. Measured on smk-g6 + with one reclaimed exposure given a manifest by hand: the dry run grew + exp_get_images, exp_split, exp_psf and exp_persist jobs. + + So a reclaimed exposure is depended on through its TAR instead. The tar is + not a declared output of any rule (exp_persist declares only its manifest, + deliberately — persist_exp.py says why), so a tar that exists is a DAG leaf: + snakemake requires it and builds nothing. A live exposure keeps its manifest + edge, which is what orders the merge after the packing; its tar does not + exist yet, so it could not serve as the edge. + + An exposure reclaimed by a workflow PREDATING exp_persist has neither tar nor + manifest and is in no set at all. Nothing short of rebuilding its chain from + VOS recovers it; the merge reports how many exposures it found. + """ + if not PERSIST_EXP: + return [] + live = set(persist_manifests()) + exps = {e for t in TILES_READY for e in tile_exposures(t)} + reclaimed = sorted(prod_exp_tar(e) for e in exps + if prod_exp_manifest(e, "exp_persist") not in live + and Path(prod_exp_manifest(e, "exp_persist")).exists() + and Path(prod_exp_tar(e)).exists()) + return sorted(live) + reclaimed + + +def star_cat_targets(): + """`full_starcat` when there is anything to stack into it, else nothing. + + Three ways to get nothing, and all three are states rather than errors: the + keep list holds no validation catalogue (warned about above), `persist_exp:` + is empty at all, or every exposure in scope is already tombstoned — a + campaign resumed after reclamation, whose exposures were cleaned by a + workflow that predates exp_persist and therefore left neither tar nor + manifest to read. A rule with an empty input list would still be a JOB, and + it would write an empty star catalogue over a good one. + """ + if not STAR_CAT_MERGE or not workflow.is_main_process: + return [] + return [full_starcat()] if star_cat_inputs() else [] + + +def final_cat_targets(): + """The merged hdf5, whenever this campaign has a tile to put in it.""" + if not workflow.is_main_process or not TILES_READY: + return [] + return [final_cat_hdf5()] + # --- tile reclamation (D5) -------------------------------------------------- # A separate flag from `clean:` (config.yaml carries the full # argument): exposure reclamation costs nothing but a rebuild if a tile is @@ -662,6 +838,8 @@ rule all: input: [final_cat(t) for t in TILES_READY], persist_targets(), + star_cat_targets(), + final_cat_targets(), clean_targets(), clean_tile_targets(), diff --git a/workflow/bin/sp b/workflow/bin/sp index fdc3d4f59..e3aec91a4 100755 --- a/workflow/bin/sp +++ b/workflow/bin/sp @@ -72,7 +72,10 @@ STATE_DIR="${SP_STATE_DIR:-${RUN_DIR}-state}"; mkdir -p "$STATE_DIR" # WHAT. `sp run` copies the code it is about to launch into $STATE_DIR/code and # runs the campaign entirely out of that copy: the Snakefile, the rules, the # scripts, the ini chain (symlinks DEREFERENCED -- workflow/config/cfis points -# into example/, and the copy must be self-contained), src/, and the profile. +# into example/, and the copy must be self-contained), src/, the repo's own +# scripts/ (final_cat_merge loads scripts/python/create_final_cat.py by path -- +# it is a script, not an installed module, and the hdf5 layout it defines must +# be pinned to the campaign like everything else here), and the profile. # Every workflow-internal path hangs off `workflow.basedir`, which IS the # snapshot, so they all follow it for free; the profile's PYTHONPATH pin is the # one that cannot (YAML splices nothing) and is rewritten below. @@ -91,10 +94,10 @@ snapshot_code() { mkdir -p "$SNAPSHOT" if command -v rsync >/dev/null 2>&1; then rsync -a --delete --copy-links --exclude '__pycache__' --exclude '*.egg-info' \ - "$HERE" "$REPO/src" "$REPO/profiles" "$SNAPSHOT/" + "$HERE" "$REPO/src" "$REPO/scripts" "$REPO/profiles" "$SNAPSHOT/" else rm -rf "$SNAPSHOT"; mkdir -p "$SNAPSHOT" - cp -rL "$HERE" "$REPO/src" "$REPO/profiles" "$SNAPSHOT/" + cp -rL "$HERE" "$REPO/src" "$REPO/scripts" "$REPO/profiles" "$SNAPSHOT/" find "$SNAPSHOT" -name __pycache__ -type d -prune -exec rm -rf {} + fi diff --git a/workflow/config.yaml b/workflow/config.yaml index ce434314c..70ace906e 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -57,6 +57,19 @@ outputs: # (-state; bin/sp explains why). products_dir: /project/def-mjhudson/cdaley/sp-products/smk-g6 +# The campaign's NAME. It labels the two campaign-level merges' output — +# /final_cat_.hdf5 and the group inside it holding that +# campaign's per-tile datasets — +# and nothing else; the per-unit stores are named by their own IDs. UNSET means +# the persistent root's basename, which is already how every campaign here is +# named (run_dir, products_dir and index_db all end in smk-g6), so this key only +# earns its place when the two must differ. +# +# It is not a rule input, so renaming a campaign mid-flight changes the merged +# catalogue's PATH and therefore builds a new one; the per-tile catalogues it +# reads are untouched. +# campaign: smk-g6 + # There is no config_src knob: the config chain is workflow/config/cfis, resolved # relative to the Snakefile. The configs interpolate $SP_RUN / $SP_UNIT_NUM / # $SP_CONFIG / $SP_EXP / $NGMIX_* and the rules export them -- configs and rules @@ -96,6 +109,16 @@ outputs: # diagnostics cannot be recomputed after a purge without rebuilding the exposure # chain from VOS. # +# THIS LIST GATES `star_cat_merge`. That campaign-level rule stacks every +# exposure's every CCD's validation_psf into ONE +# /full_starcat-0000000.fits, reading the members straight out of +# the tars. A keep list that matches no `validation_psf-*.fits` is a legitimate +# configuration (keep the PSF models alone, say) and produces NO merge job and a +# warning at parse time — not a failure on a node an hour later. Note the +# corollary: an exposure already reclaimed by a workflow that predates +# exp_persist left no tar, so it contributes nothing and cannot be recovered +# short of rebuilding its chain from VOS. +# # OPT-IN CANDIDATES, and what each buys. Sizes are per exposure (40 CCDs), # measured on smk-m2 (127 exposures, 64 tiles); a 64-tile campaign with all of # the measured ones on came to 7.2 GB: diff --git a/workflow/rules/exposure.smk b/workflow/rules/exposure.smk index 40a4044d4..9c6831549 100644 --- a/workflow/rules/exposure.smk +++ b/workflow/rules/exposure.smk @@ -220,3 +220,70 @@ rule clean_exposure: f"python {SCRIPTS}/clean_exposure.py" " --exp-dir $(dirname {output.tombstone}) --exp {wildcards.exp}" " --tombstone {output.tombstone} --consumers '{params.consumers}'" + + +# --- the campaign's star catalogue ------------------------------------------ +# ONE job per campaign: every exposure's every CCD's `validation_psf--.fits`, +# stacked into `/full_starcat-0000000.fits`. That file is the +# rho/tau statistics input and sp_validation reads it at exactly that path, +# doing no merging of its own; the old bash chain built it with +# `combine_runs.bash psf` + a `merge_starcat_runner` pass, and the workflow +# emitted neither. The stacking itself is `MergeStarCatPSFEX` — the same class +# the old runner called, reused rather than restated, so a column added to the +# module is a column added here (merge_star_cat.py argues the reuse and the +# tar-member reading). +# +# THE INPUT IS star_cat_inputs() (Snakefile): every exposure of TILES_READY whose +# PSF products are on the persistent root — the live ones through the exp_persist +# manifest edge `rule all` already requests, the RECLAIMED ones through their TAR, +# which no rule declares and which therefore requires nothing to be built. That +# asymmetry is not a flourish; requesting a reclaimed exposure's manifest +# rebuilds its whole chain from VOS, and ancient() does not prevent it (measured +# — the Snakefile carries the numbers). Nothing new enters the DAG either way. It +# is read through an INPUT FUNCTION rather than at module level so that only a +# parse which actually builds this job pays for the walk. +# +# THE PATHS DO NOT REACH THE SHELL, and that is not a style choice: ~20k manifest +# paths is an order of magnitude over Linux's 128 KiB MAX_ARG_STRLEN for a single +# argv entry, so `{input}` here would be a job that dies on exec at DR6 scale. +# The job is handed the two small files the Snakefile itself started from — the +# tile list and the index — and derives THE SAME SET from them; `params.inputs` +# carries that set's FINGERPRINT, which is the rerun trigger. The equality is +# the point: a job that stacked anything the fingerprint did not see would be +# rows no rerun trigger could notice, which is what a glob over products_dir +# would have given on a root shared with an earlier, larger tile list. +# Byte-stable output otherwise (tmp-then-cmp-then-mv), so a no-op rerun does not +# move its mtime. +# +# NOT A LOCALRULE. exp_persist is local because it is 20k jobs of seconds; this +# is one job that holds a campaign's stars in memory (~800k catalogues at DR6 +# scale). mem_mb is a guess scaled by attempt, not a measurement — the campaigns +# run so far are 127 exposures, three orders of magnitude short of the case this +# sizing is for, and the first DR6-scale run should replace this number with a +# benchmark. +# +# NO JOB AT ALL when `persist_exp:` keeps no validation catalogue, or when every +# exposure in scope is tombstoned: star_cat_targets() (Snakefile) simply does not +# request the output, and the parse says so rather than a node failing later. +rule star_cat_merge: + input: + lambda wc: star_cat_inputs() + output: + star_cat = full_starcat() + params: + products_dir = str(PRODUCTS_DIR), + tile_list = str(config["tile_list"]), + index_db = str(INDEX_DB), + inputs = lambda wc, input: input_fingerprint(input), + script_hash = MERGE_STAR_HASH + threads: 1 + resources: + mem_mb = lambda wc, attempt: 16000 * attempt, + runtime = 120 + shell: + "set -euo pipefail\n" + f"python {SCRIPTS}/merge_star_cat.py" + " --products-dir '{params.products_dir}'" + " --tile-list '{params.tile_list}' --index-db '{params.index_db}'" + " --output {output.star_cat}" + f" --psf-model {PSF_MODEL}" diff --git a/workflow/rules/tile.smk b/workflow/rules/tile.smk index 68d6da9ba..cd10ce6cc 100644 --- a/workflow/rules/tile.smk +++ b/workflow/rules/tile.smk @@ -898,3 +898,63 @@ rule clean_tile: f"python {SCRIPTS}/clean_tile.py" " --tile-dir $(dirname {output.tombstone}) --tile {wildcards.tile}" " --tombstone {output.tombstone}" + + +# --- the campaign's shear catalogue ----------------------------------------- +# ONE job per campaign, the tile-side twin of exposure.smk's star_cat_merge, and +# the same three design calls hold: the input is the list `rule all` already +# requests (every ready tile's final_cat), the paths never reach the shell +# (MAX_ARG_STRLEN), and a fingerprint on `params` is what makes it rerun when a +# tile is appended. The job derives the same set the fingerprint was taken over +# from the tile list and the index rather than globbing products_dir — on a +# products root shared with an earlier, larger tile list a glob would merge tiles +# no rerun trigger ever saw. +# +# THE OUTPUT SCHEMA IS AN INTERFACE, NOT A CHOICE. sp_validation opens this file +# as its `galaxy_cat_path`: one dataset per tile under a named group, the +# columns of workflow/config/cfis/final_cat.param, an `n_tiles` attribute on the +# root. The group is named for the CAMPAIGN, which is the only unit this +# workflow has above the tile. So the rule reuses +# scripts/python/create_final_cat.py's column extraction rather than restating +# it, and writes the file itself — merge_final_cat.py argues that split, the one +# legacy literal in the schema, and the two places where the reference +# implementation had to be pinned down to be reproducible. +# +# THE INPUT IS final_cat, NOT the tile_make_cat manifest, for the same reason +# clean_tile's is: final_cat on the persistent root IS the campaign's +# tile-finished marker (see final_cat() in the Snakefile), and it is the file +# this rule actually reads. +# +# NOT A LOCALRULE, and here the reason is IO rather than memory: the job reads +# every tile's catalogue end to end on every run — ~32-46 MB per tile, so ~2 GB +# for a 64-tile campaign and ~800 GB at DR6's 23k tiles. It rebuilds rather than +# appends because a DAG output must be a function of its input set +# (merge_final_cat.py); incremental update by hand is what +# `create_final_cat.py -s add` remains for. Memory is one tile's catalogue at a +# time plus the hdf5 write buffer, which is why mem_mb is modest where +# star_cat_merge's is not. +rule final_cat_merge: + input: + lambda wc: [final_cat(t) for t in TILES_READY] + output: + merged = final_cat_hdf5() + params: + products_dir = str(PRODUCTS_DIR), + tile_list = str(config["tile_list"]), + index_db = str(INDEX_DB), + param_file = str(CONFIG_DIR / "final_cat.param"), + campaign = CAMPAIGN, + inputs = lambda wc, input: input_fingerprint(input), + script_hash = MERGE_FINAL_HASH + threads: 1 + resources: + mem_mb = lambda wc, attempt: 8000 * attempt, + runtime = 120 + shell: + "set -euo pipefail\n" + f"python {SCRIPTS}/merge_final_cat.py" + " --products-dir '{params.products_dir}'" + " --tile-list '{params.tile_list}' --index-db '{params.index_db}'" + " --output {output.merged}" + " --campaign '{params.campaign}'" + " --param-file '{params.param_file}'" diff --git a/workflow/scripts/build_index.py b/workflow/scripts/build_index.py index 2dd191306..83e9c606b 100644 --- a/workflow/scripts/build_index.py +++ b/workflow/scripts/build_index.py @@ -152,6 +152,45 @@ def build(tile_ids: list[str], run_dir: Path, db_path: Path, "n_missing": len(missing)} +# --- reading it back, for the campaign-level merges ------------------------- +# The Snakefile loads this index into dicts at parse time and derives the +# campaign's unit sets from them (TILES_READY, and the exposures those tiles +# read). A merge JOB has to derive the same two sets, and cannot be handed them +# on its command line — ~20k paths is an order of magnitude over Linux's 128 KiB +# MAX_ARG_STRLEN for a single argv entry. So it is given the two things the +# Snakefile itself started from, the tile list and this database, and rebuilds +# the sets here. Both halves therefore read the schema through one module rather +# than two hand-written queries that could drift apart. + + +def campaign_tiles(tile_list: Path, db_path: Path) -> list[str]: + """The campaign's ready tiles: declared in the list AND indexed. + + Exactly the Snakefile's TILES_READY, computed the same way from the same two + files — a declared tile with no indexed exposure list cannot have been + computed, so it has no catalogue to merge. + """ + with open(tile_list) as f: + declared = [ln.strip() for ln in f if ln.strip()] + con = sqlite3.connect(db_path, timeout=60) + indexed = {r[0] for r in con.execute("SELECT DISTINCT tile_id FROM tile_exposures")} + con.close() + return [t for t in declared if t in indexed] + + +def campaign_exposures(tile_list: Path, db_path: Path) -> list[str]: + """Every exposure the campaign's ready tiles read, sorted. + + Exactly the set the Snakefile's persist_manifests() builds its manifest + paths from. + """ + tiles = set(campaign_tiles(tile_list, db_path)) + con = sqlite3.connect(db_path, timeout=60) + rows = con.execute("SELECT tile_id, exp_id FROM tile_exposures").fetchall() + con.close() + return sorted({e for t, e in rows if t in tiles}) + + def main() -> None: p = argparse.ArgumentParser(description=__doc__) p.add_argument("--tile-list", required=True, type=Path, diff --git a/workflow/scripts/merge_final_cat.py b/workflow/scripts/merge_final_cat.py new file mode 100644 index 000000000..2d2974fc0 --- /dev/null +++ b/workflow/scripts/merge_final_cat.py @@ -0,0 +1,222 @@ +#!/usr/bin/env python3 +"""Collect the campaign's per-tile final catalogues into ONE hdf5 file. + +Run as the shell of the campaign-level ``final_cat_merge`` rule, never by hand. + +WHAT IT PRODUCES, AND FOR WHOM. ``/final_cat_.hdf5``: +one dataset per tile, carrying the columns named by +``workflow/config/cfis/final_cat.param``, plus an ``n_tiles`` attribute on the +file root. sp_validation opens that file as its ``galaxy_cat_path`` +(``sp_validation/catalog.py``), so its SCHEMA is an interface and not a choice — +see ``SPVAL_GROUP`` below for the one legacy literal in it. + +(sp_validation's own ``merge_catalogues`` is a different layer entirely: it +works over already-calibrated ``shape_catalog_comprehensive_*.fits``. It does +not do this merge, and this does not do that one.) + +WHAT IT REUSES, AND WHAT IT DOES NOT. The column extraction is +``create_final_cat.py``'s — ``read_param_file`` for the parameter list, +``read_data`` and ``copy_data`` for pulling those columns out of one catalogue +with their FITS dtypes — so the column grammar keeps exactly one definition. +Its ``process()`` is NOT used and neither is any of its discovery: that function +walks a directory tree the workflow does not have and never will, and it groups +by a unit ShapePipe v2 no longer has. This script walks the workflow's own +products tree instead (``tiles/<2-char prefix>//final_cat-.fits``) and +writes the hdf5 itself. + +WHERE ``create_final_cat.py`` IS FOUND. Beside this workflow, at +``/scripts/python/create_final_cat.py`` — resolved relative to THIS file, +so it follows the launch code snapshot (``bin/sp``) exactly as +``workflow/scripts/*`` does, and a campaign never reads a mid-run edit. It is +loaded by path rather than imported: it is a script, not an installed module, +and the container's ``shapepipe`` install does not carry it. + +IT REBUILDS THE WHOLE FILE, IT DOES NOT APPEND. ``create_final_cat.py``'s own +``process()`` skips tiles already in the file, which is right for a hand-driven +incremental update (``-s add`` / ``-s remove`` are that tool's job). A DAG rule +wants the opposite: the output must be a pure function of the input set, so that +a no-op rerun is byte-stable and a changed set is visibly a different file. +Appending would make the result depend on the order campaigns were run in, and +would silently keep a tile whose catalogue was later rebuilt. The cost is +reading every tile's catalogue on every run of the rule — real work at DR6 scale +(~20k tiles), which is why this is not a localrule. + +BYTE-STABLE ON A NO-OP RERUN: written to a tmp path, compared, moved only if it +differs (the pattern ``persist_exp.py`` and ``clean_exposure.py`` use). Tiles +are visited in sorted ID order so the file is a function of the input set alone. +An unconditional rewrite would move the output's mtime every invocation. + +TWO PLACES WHERE THE REFERENCE IMPLEMENTATION IS NOT DETERMINISTIC, and where +this script therefore pins the behaviour down rather than copying it. Both are +in the DTYPE, and both are invisible when a human runs the tool once by hand: + + * ``copy_data`` allocates ``np.empty`` with the SOURCE catalogue's full + dtype and then fills only the requested columns, so every column NOT in + ``final_cat.param`` reaches the hdf5 file as uninitialised memory — + different bytes on every run, and meaningless data in the file besides. We + hand ``copy_data`` a dtype restricted to the requested columns, so every + field it writes is a field it fills. The file then carries exactly the + ``final_cat.param`` columns, which is what sp_validation reads and what the + parameter file is for. + * ``read_param_file`` returns ``list(set(...))``, whose order varies with the + process's string hash seed. Column ORDER in a structured dtype is part of + the file, so that alone would defeat the byte comparison. We order the + fields by the source catalogue's own column order instead. + +WHICH TILES — AND WHY THE JOB DERIVES THE SET RATHER THAN BEING TOLD IT. The set +is the CAMPAIGN's: every tile both declared in ``tile_list`` and present in the +index, which is exactly the Snakefile's TILES_READY, rebuilt here from the same +two files the Snakefile started from (``--tile-list`` and ``--index-db``, read +through ``build_index.campaign_tiles`` so there is one definition and not two +that can drift). It is derived rather than passed because at DR6 scale the set +is ~20k paths and a shell command reaches ``execve`` as a SINGLE argv entry +capped at 128 KiB by ``MAX_ARG_STRLEN``; the rule's ``input`` is the DAG edge +and its ``params`` carries a fingerprint of that same list, which is the rerun +trigger. + +THE TWO SETS ARE THE SAME SET, which is the point of deriving it this way rather +than globbing ``/tiles``: a products root shared with an earlier, +larger tile list would hand the job tiles the fingerprint never saw and no rerun +trigger would notice. A tile in the derived set whose catalogue is missing is a +hard error here, not a skip — under the DAG it cannot happen, since every one of +them is a declared input of this job. +""" + +import argparse +import filecmp +import importlib.util +import sys +from pathlib import Path + +import h5py +import numpy as np + +# Same directory; the rule invokes this file by path, so it is sys.path[0]. +import build_index + +# /scripts/python/create_final_cat.py, from /workflow/scripts/this. +CFC_PATH = (Path(__file__).resolve().parents[2] + / "scripts" / "python" / "create_final_cat.py") + + +def spval_group(campaign: str) -> str: + """The hdf5 group the campaign's per-tile datasets live under. + + ``patches/`` is a LEGACY KEY IN sp_validation's FILE SCHEMA, kept verbatim + only so its reader works unchanged (CosmoStat/sp_validation#340 tracks + removing it); it names nothing in this workflow, which has campaigns and + tiles and no other unit. This is the one place the literal appears — + everything else here says campaign. + """ + return f"patches/{campaign}" + + +def load_create_final_cat(): + """The hdf5 layout's definition, loaded by path (see the module docstring).""" + if not CFC_PATH.exists(): + sys.exit(f"merge_final_cat: {CFC_PATH} is not there — the launch code " + f"snapshot must carry scripts/python/ (see bin/sp).") + spec = importlib.util.spec_from_file_location("create_final_cat", CFC_PATH) + mod = importlib.util.module_from_spec(spec) + spec.loader.exec_module(mod) + return mod + + +def catalogues(products_dir: Path, tile_list: Path, index_db: Path) -> list: + """``(tile ID, path)`` for the campaign's tiles, in ID order. + + Not a glob over the products root: see the module docstring on why the set + is the campaign's and not the filesystem's. + """ + out, missing = [], [] + for tile in sorted(build_index.campaign_tiles(tile_list, index_db)): + path = (products_dir / "tiles" / tile[:2] / tile + / f"final_cat-{tile}.fits") + if path.exists(): + out.append((tile, path)) + else: + missing.append(tile) + if missing: + sys.exit(f"merge_final_cat: {len(missing)} campaign tile(s) have no " + f"final catalogue: {' '.join(missing[:5])}" + f"{' ...' if len(missing) > 5 else ''}") + return out + + +def main() -> None: + p = argparse.ArgumentParser(description=__doc__) + p.add_argument("--products-dir", required=True, type=Path, + help="the persistent root; per-tile catalogues are found " + "beneath it") + p.add_argument("--tile-list", required=True, type=Path, + help="the campaign's tile list (config tile_list)") + p.add_argument("--index-db", required=True, type=Path, + help="the campaign's run index (config outputs.index_db)") + p.add_argument("--output", required=True, type=Path) + p.add_argument("--campaign", required=True, + help="names the campaign's group in the output file") + p.add_argument("--param-file", required=True, type=Path, + help="workflow/config/cfis/final_cat.param — the column list") + p.add_argument("--hdu", type=int, default=1) + args = p.parse_args() + + cfc = load_create_final_cat() + param_list = cfc.read_param_file(str(args.param_file), verbose=False) + if not param_list: + sys.exit(f"merge_final_cat: no columns read from {args.param_file}") + # read_data/copy_data read their knobs out of this dict, exactly as + # create_final_cat.py's own main() builds it. + params = {"hdu_num": args.hdu, "param_list": param_list, "verbose": False} + + tiles = catalogues(args.products_dir, args.tile_list, args.index_db) + if not tiles: + # An empty hdf5 would satisfy every downstream existence check and + # produce an empty shear catalogue. + sys.exit(f"merge_final_cat: no tile in {args.tile_list} is indexed in " + f"{args.index_db}, so there is nothing to merge") + + # tmp-then-cmp-then-mv; the tmp never outlives this process. + args.output.parent.mkdir(parents=True, exist_ok=True) + tmp = args.output.with_name(args.output.name + ".tmp") + try: + tmp.unlink(missing_ok=True) # h5py "a" would reopen a stale one + with h5py.File(tmp, "w") as hdf5_file: + group = hdf5_file.create_group(spval_group(args.campaign)) + columns = None + for tile, path in tiles: + extracted, dtype = cfc.read_data(str(path), params) + # Requested columns, in the SOURCE catalogue's order (see the + # module docstring on determinism). Computed from the first + # tile and reused, so a tile whose catalogue is missing a + # column fails loudly on the assignment rather than quietly + # producing a differently-shaped dataset. + if columns is None: + columns = [c for c in dtype.names + if c in set(params["param_list"])] + missing = sorted(set(params["param_list"]) - set(columns)) + if missing: + sys.exit(f"merge_final_cat: {path} has none of the " + f"requested column(s): {' '.join(missing)}") + subset = np.dtype([(c, dtype[c]) for c in columns]) + group.create_dataset( + tile, + data=cfc.copy_data(columns, extracted, subset), + dtype=subset, + ) + # The same attribute create_final_cat.py's print_list() writes, and + # what sp_validation reads to know how many tiles it is holding. + hdf5_file.attrs["n_tiles"] = len(tiles) + + if args.output.exists() and filecmp.cmp(tmp, args.output, shallow=False): + print(f"[merge_final_cat] unchanged: {args.output}") + else: + tmp.replace(args.output) # atomic: same filesystem + print(f"[merge_final_cat] {len(tiles)} tile(s), " + f"{len(param_list)} column(s) -> {args.output} " + f"(group {spval_group(args.campaign)})") + finally: + tmp.unlink(missing_ok=True) + + +if __name__ == "__main__": + main() diff --git a/workflow/scripts/merge_star_cat.py b/workflow/scripts/merge_star_cat.py new file mode 100644 index 000000000..a4f2e3d9f --- /dev/null +++ b/workflow/scripts/merge_star_cat.py @@ -0,0 +1,234 @@ +#!/usr/bin/env python3 +"""Concatenate the campaign's per-CCD PSF validation catalogues into ONE full_starcat. + +Run as the shell of the campaign-level ``star_cat_merge`` rule, never by hand. + +WHAT IT PRODUCES, AND FOR WHOM. ``/full_starcat-0000000.fits``: +every exposure's every CCD's ``validation_psf--.fits`` row, stacked, +with a ``CCD_NB`` column recording which CCD each row came from. It is the input +to the rho/tau statistics — sp_validation reads exactly this path +(``star_cat_path`` in its ``scripts/calibration/params.py``) and does no merging +of its own. Historically it was ``combine_runs.bash psf`` + a +``merge_starcat_runner`` pass; the workflow emitted neither, so the product set +was short one file. This script is that pass, driven by the DAG instead of by +bash. + +IT DOES NOT REIMPLEMENT THE COLUMN LIST. The stacking, the column names and the +CCD_NB parse all live in ``MergeStarCatPSFEX`` +(``shapepipe.modules.merge_starcat_package.merge_starcat``), which is what the +old runner called. This script only decides WHICH catalogues that class is +handed, and where the result lands. A column added to the module is a column +added here for free — which is the entire reason for the indirection. + +IT READS THE TARS, IT DOES NOT UNPACK THEM. ``exp_persist`` packs each +exposure's keepers into one uncompressed tar on the persistent root +(``/exp///psf/.tar``) precisely because inodes, +not bytes, bind on /project. Unpacking ~20k tars × ~40 members to merge them +would materialise ~800k files on the filesystem that design exists to protect, +and then delete them. So members are read into memory +(``tarfile.extractfile(m).read()`` -> ``io.BytesIO``) one at a time and handed +to the merge class as ``[fileobj, member_name]`` pairs. The member NAME is what +the CCD_NB regex parses, which is why the pair carries it; the class takes the +name from the last element of the entry, so a plain ``[path]`` entry behaves +exactly as it always did. + +WHICH EXPOSURES — AND WHY THE JOB DERIVES THE SET RATHER THAN BEING TOLD IT. +The set is the CAMPAIGN's: every exposure read by a tile that is both declared +in ``tile_list`` and present in the index, which is the Snakefile's TILES_READY +walked one edge further. This script rebuilds it from the same two files the +Snakefile started from (``--tile-list`` and ``--index-db``, both small, both on +the persistent root, both read through ``build_index.campaign_exposures`` so +there is one query and not two that can drift), and then takes the exposures +whose ``exp_persist`` manifest is on the persistent root. + +It is derived rather than passed because at DR6 scale the set is ~20k paths, and +a shell command reaches ``execve`` as a SINGLE argv entry capped at 128 KiB by +``MAX_ARG_STRLEN``. Passing them would be a job that dies before it starts. So +the rule's ``input`` is the DAG EDGE — what must exist before this runs — and +the rule's ``params`` carries a FINGERPRINT of that same list, which is what +makes the merge rerun when the set changes. + +THE TWO SETS ARE THE SAME SET, and that equality is the point of deriving it +this way rather than globbing the tree. The rule's input is ``star_cat_inputs()`` +(Snakefile): for each exposure of TILES_READY whose PSF products are on the +persistent root, an edge — the ``exp_persist`` manifest for a live exposure, the +TAR for one whose scratch store reclamation already took (that function argues +the asymmetry, which is about not rebuilding a reclaimed chain from VOS). +Nothing at all for an exposure reclaimed before ``exp_persist`` existed, which +left neither and is unrecoverable short of that rebuild. What this script +selects is the same rule stated from the job's side: same tiles, same index, +manifest present — and by the time the job runs, every exposure with an edge has +one. A glob over ``/exp`` would NOT be the same set: it would +sweep in exposures of an earlier, larger tile list sharing the products root, +stacking rows the fingerprint never saw and no rerun trigger would notice. + +THE MANIFEST, NOT THE TAR, IS WHAT IT READS FIRST: the manifest records what was +actually packed, pattern by pattern, member by member, with sizes. Selecting +members from it means this script never guesses at tar contents, and an exposure +whose keep list did not include the validation catalogues contributes nothing +visibly rather than silently. + +BYTE-STABLE ON A NO-OP RERUN: written to a tmp path, compared, and moved only +if it differs (the pattern ``persist_exp.py`` and ``clean_exposure.py`` use). +An unconditional rewrite would move the output's mtime on every invocation. +Members are visited in sorted (exposure, member) order so the row order is a +function of the input set alone. + +PSFEX ONLY, DELIBERATELY. ``PSF_MODEL`` is ``psfex`` in every campaign the +workflow has run; ``MergeStarCatMCCD`` and ``MergeStarCatSetools`` exist beside +it and take the same constructor, so the hook is the one-line class choice in +``merge_class()`` below — an implementation, not a design, away. +""" + +import argparse +import filecmp +import io +import json +import logging +import shutil +import sys +import tarfile +import tempfile +from fnmatch import fnmatch +from pathlib import Path + +from shapepipe.modules.merge_starcat_package import merge_starcat + +# Same directory; the rule invokes this file by path, so it is sys.path[0]. +import build_index + +# The output name is not ours to choose: sp_validation hardcodes it +# (`star_cat_path = f"{data_dir}/full_starcat-0000000.fits"`), and +# MergeStarCatPSFEX writes exactly this basename into the output dir it is +# given. Kept here as the name this script promises to produce. +OUT_NAME = "full_starcat-0000000.fits" + +# The keep-list pattern whose members this merge consumes. The rule refuses to +# exist unless `persist_exp:` contains a pattern matching this shape (the +# Snakefile does that check at parse time), so by the time we get here the +# members are expected to be present. +MEMBER_PATTERN = "validation_psf-*.fits" + + +def merge_class(psf_model: str): + """The merge class for this PSF model — the one-line MCCD/setools hook.""" + try: + return {"psfex": merge_starcat.MergeStarCatPSFEX, + "mccd": merge_starcat.MergeStarCatMCCD, + "setools": merge_starcat.MergeStarCatSetools}[psf_model] + except KeyError: + sys.exit(f"merge_star_cat: unknown psf_model {psf_model!r}") + + +def manifests(products_dir: Path, tile_list: Path, index_db: Path) -> list: + """The campaign's exp_persist manifests that are on disk, in exposure order. + + Not a glob over the products root: see the module docstring on why the set + is the campaign's and not the filesystem's. + """ + out = [] + for exp in build_index.campaign_exposures(tile_list, index_db): + path = (products_dir / "exp" / exp[:2] / exp / "manifests" + / "exp_persist.json") + if path.exists(): + out.append(path) + return out + + +def entries(manifest_paths: list, pattern: str) -> tuple: + """``[fileobj, member_name]`` for every matching member, and the tar count. + + One tar is opened at a time and its members are read into memory; the tars + are never unpacked to disk (see the module docstring). The returned file + objects are BytesIO, so nothing stays open on the filesystem — at ~50 KB per + member and ~40 members per exposure this is ~2 MB per exposure held only for + as long as the merge takes to consume it, but note that the merge class + holds the whole stack in python lists regardless, which is the real memory + term the rule's mem_mb is sized against. + """ + out, n_tars, empty = [], 0, [] + for man_path in manifest_paths: + man = json.loads(man_path.read_text()) + wanted = sorted(f["name"] for f in man["files"] + if fnmatch(f["name"], pattern)) + if not wanted: + empty.append(man["unit"]) + continue + tar_path = Path(man["tar"]) + if not tar_path.exists(): + sys.exit(f"merge_star_cat: {man_path} names a tar that is not " + f"there: {tar_path}") + with tarfile.open(tar_path) as tf: + for name in wanted: + member = tf.extractfile(name) + if member is None: + sys.exit(f"merge_star_cat: {tar_path} has no member " + f"{name}, which its manifest lists") + out.append([io.BytesIO(member.read()), name]) + n_tars += 1 + return out, n_tars, empty + + +def main() -> None: + p = argparse.ArgumentParser(description=__doc__) + p.add_argument("--products-dir", required=True, type=Path, + help="the persistent root; exp_persist manifests and tars " + "are found beneath it") + p.add_argument("--tile-list", required=True, type=Path, + help="the campaign's tile list (config tile_list)") + p.add_argument("--index-db", required=True, type=Path, + help="the campaign's run index (config outputs.index_db)") + p.add_argument("--output", required=True, type=Path, + help=f"the merged catalogue; its basename is {OUT_NAME}") + p.add_argument("--psf-model", default="psfex") + p.add_argument("--pattern", default=MEMBER_PATTERN, + help="tar-member glob to merge; default %(default)s") + args = p.parse_args() + + if args.output.name != OUT_NAME: + # The merge class writes OUT_NAME into a directory it is handed; a + # differently-named declared output would silently never be produced. + sys.exit(f"merge_star_cat: --output must be named {OUT_NAME} " + f"(got {args.output.name})") + + log = logging.getLogger("merge_star_cat") + logging.basicConfig(format="[merge_star_cat] %(message)s", + level=logging.INFO, stream=sys.stdout) + + manifest_paths = manifests(args.products_dir, args.tile_list, args.index_db) + file_list, n_tars, empty = entries(manifest_paths, args.pattern) + if not file_list: + # Not a no-op: an empty star catalogue would pass every downstream + # existence check and produce meaningless rho statistics. + sys.exit(f"merge_star_cat: no member matched {args.pattern!r} in any " + f"of {len(manifest_paths)} exp_persist manifest(s) for this " + f"campaign — is '{args.pattern}' in the persist_exp keep list?") + if empty: + log.info(f"{len(empty)} exposure(s) persisted no {args.pattern}: " + f"{', '.join(sorted(empty)[:5])}" + f"{' ...' if len(empty) > 5 else ''}") + + # tmp-then-cmp-then-mv. The merge class chooses its own basename inside the + # directory it is given, so the tmp is a DIRECTORY, not a file path, and it + # never outlives this process — an orphan on /project is an inode nothing + # revisits. + args.output.parent.mkdir(parents=True, exist_ok=True) + tmp_dir = Path(tempfile.mkdtemp(dir=args.output.parent, + prefix=".star_cat_merge.")) + try: + merge_class(args.psf_model)(file_list, str(tmp_dir), log).process() + tmp = tmp_dir / OUT_NAME + if not tmp.exists(): + sys.exit(f"merge_star_cat: the merge wrote no {OUT_NAME}") + if args.output.exists() and filecmp.cmp(tmp, args.output, shallow=False): + log.info(f"unchanged: {args.output}") + else: + tmp.replace(args.output) # atomic: same filesystem + log.info(f"{len(file_list)} catalogue(s) from {n_tars} exposure(s) " + f"-> {args.output}") + finally: + shutil.rmtree(tmp_dir, ignore_errors=True) + + +if __name__ == "__main__": + main() From fc9aa3ea298374c57c2547cdbe84b8da4ab2eb4a Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 10:33:28 -0400 Subject: [PATCH 08/85] fix(orchestration): seven defects in the campaign-level merges MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit PERSISTENCE KEYED ON THE WRONG FILE, and the consequence was the avalanche it exists to prevent. persist_targets() skipped an exposure only when its SCRATCH tombstone was there, but clean_exposure is one of two ways a store disappears and the other leaves nothing behind: /scratch is purged on a 60-day window whether or not the workflow reclaimed anything. After a purge — or on any campaign with clean: false — every exposure looked live, its persist manifest was requested, its exp_psf manifest was gone, and snakemake rebuilt the whole exposure chain from VOS. The test is now exp_store_reclaimed(), used by persist_manifests() and by star_cat_merge's live/reclaimed split so the two cannot disagree, and it takes BOTH pieces of evidence that an exposure once had a store: a tombstone, or a persisted manifest with no exp_psf manifest beside it. Both are needed. The manifest clause alone regressed the tombstone case — measured on smk-g6, whose 126 exposures were cleaned before exp_persist existed and so have no manifest: the dry run grew 126 exp_get_images, exp_split, exp_psf and exp_persist jobs, a whole campaign rebuilt from VOS. The tombstone clause alone is the purge bug above. Neither file can be dropped, because the parse cannot otherwise tell a store that is GONE from one not BUILT yet, and a fresh campaign must still be asked to persist. OVERLAPPING KEEP PATTERNS FAILED EVERY EXPOSURE. persist_exp treated a file matched by two patterns as a flat-member name collision, so 'validation_psf-*.fits' alongside '*.fits' — an ordinary way to write a keep list — aborted the pack. Two DIFFERENT paths on one member name is still fatal; the same path twice is now one file, recorded under the first pattern that matched it. merge_star_cat MATERIALISED THE WHOLE CAMPAIGN before merging a row: every member's bytes, ~2 MB per exposure, ~40 GB at DR6's ~20k exposures against a rule asking for 16 GB. TarMembers hands the merge class the same entries one tar at a time, so peak memory is one member plus the class's own accumulators, which are the unavoidable term. It keeps __len__ off the manifests so the count is still logged before a tar is opened. THE STAR MERGE RERAN ON BOOKKEEPING. Its fingerprint was over input PATHS, and an exposure's edge flips from its manifest to its tar the moment its store is reclaimed — so every reclamation pass reran the merge over identical content. It is over the exposure IDS now, which move only when the set does, and which are what the job derives on its own side. final_cat_merge's is over tile ids for the same reason. merge_final_cat's MISSING-COLUMN CHECK WAS UNREACHABLE. create_final_cat's read_data wraps its column selection in a bare `except:` that prints and falls through, so a missing column left its return values unbound and the caller got UnboundLocalError from the return statement, naming nothing. The columns are checked against the catalogue's own header before read_data is called, and the message now names every missing one. A TILE LISTED TWICE killed the merge on the second create_dataset. The tile list is appended to by hand, so duplicates happen; campaign_tiles() dedupes it order-preserving, and the Snakefile's TILES does the same so the fingerprint and the job's derived set still name the same set. merge_class OFFERED MCCD AND SETOOLS while only MergeStarCatPSFEX had learned the [fileobj, name] entry shape. Both now take the entry's name from its last element like PSFEX does — unchanged for the module runner, whose entries are [path]. Setools needs one thing more before it can read a tar (it hands file_io input_file_list[0][0] as a template path), and merge_class says so rather than implying otherwise. Also: README no longer lists scripts/sp_rule.py, which does not exist. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- .../merge_starcat_package/merge_starcat.py | 17 ++- workflow/README.md | 1 - workflow/Snakefile | 123 ++++++++++++++---- workflow/rules/exposure.smk | 2 +- workflow/rules/tile.smk | 2 +- workflow/scripts/build_index.py | 11 +- workflow/scripts/merge_final_cat.py | 23 +++- workflow/scripts/merge_star_cat.py | 95 +++++++++----- workflow/scripts/persist_exp.py | 10 ++ 9 files changed, 214 insertions(+), 70 deletions(-) diff --git a/src/shapepipe/modules/merge_starcat_package/merge_starcat.py b/src/shapepipe/modules/merge_starcat_package/merge_starcat.py index 1717825f9..2b1bf8c34 100644 --- a/src/shapepipe/modules/merge_starcat_package/merge_starcat.py +++ b/src/shapepipe/modules/merge_starcat_package/merge_starcat.py @@ -239,10 +239,15 @@ def process(self): my_mask[inside_circle] = True for name in self._input_file_list: + # The source to read and the NAME to report it by; identical for a + # plain [path] entry (see MergeStarCatPSFEX's docstring on the + # [fileobj, name] form). This class takes its CCD numbers from the + # data's own CCD_ID_LIST, so the name is only ever used in messages. + source, label = name[0], name[-1] try: - starcat_j = fits.open(name[0], memmap=False, ignore_missing_simple=True) + starcat_j = fits.open(source, memmap=False, ignore_missing_simple=True) except ValueError: - print(f"Error for file {name[0]}, check FITS file integrity") + print(f"Error for file {label}, check FITS file integrity") #raise continue @@ -799,7 +804,11 @@ def process(self): ) for name in self._input_file_list: - starcat_j = fits.open(name[0], memmap=False) + # The source to read and the NAME to parse the CCD number out of; + # identical for a plain [path] entry (see MergeStarCatPSFEX's + # docstring on the [fileobj, name] form). + source, label = name[0], name[-1] + starcat_j = fits.open(source, memmap=False) data_j = starcat_j[self._hdu_table].data @@ -824,7 +833,7 @@ def process(self): snr += list(data_j["SNR_WIN"]) # CCD number - ccd_nb += [re.split(r"\-([0-9]*)\-([0-9]+)\.", name[0])[-2]] * len( + ccd_nb += [re.split(r"\-([0-9]*)\-([0-9]+)\.", label)[-2]] * len( data_j["XWIN_IMAGE"] ) diff --git a/workflow/README.md b/workflow/README.md index 824eddc73..eb801da6a 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -161,7 +161,6 @@ workflow/ exposure.smk per-exposure: get_images, split, psf, persist (no temp()); campaign star_cat_merge tile.smk per-tile: exp forest, merge_headers, detect, vignets, ngmix, merge, make_cat; campaign final_cat_merge scripts/ - sp_rule.py the thin per-unit wrapper (isolation furniture, config copy, log-sync, count check) build_index.py prepare-phase run_index.sqlite builder (plain script) build_forest.py per-tile exposure symlink forest (group-compatible shell) completeness.py the ported count table (shared by sp_rule + run_report) diff --git a/workflow/Snakefile b/workflow/Snakefile index 2954222bc..587f271b2 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -136,8 +136,14 @@ if PHASE not in ("prepare", "compute", "passthrough"): raise WorkflowError( f"SP_PHASE={PHASE!r} is not one of prepare, compute, passthrough.") +# DEDUPED, order preserved. The tile list is appended to by hand across a +# campaign, so a tile can appear twice — harmless for a per-tile target, which +# is the same path requested twice, but not for the campaign-level merges: their +# fingerprint counts what is in this list and the job derives a deduped set from +# the same file (build_index.campaign_tiles), so a duplicate line would make the +# two disagree about a set they must name identically. with open(config["tile_list"]) as f: - TILES = [ln.strip() for ln in f if ln.strip()] + TILES = list(dict.fromkeys(ln.strip() for ln in f if ln.strip())) # --- ngmix scatter (D4) ---------------------------------------------------- # Native directive: `--set-scatter ngmix=N` overrides it, N=1 degenerates to one @@ -479,12 +485,11 @@ def persist_targets(): Scope is the ready tiles' exposures, which `all` already builds through the tile chain, so nothing new is pulled into the DAG by asking. - EXCEPT A CLEANED EXPOSURE. Its exp_psf manifest was deleted by - clean_exposure, so requesting its persist manifest would make the DAG - rebuild the whole exposure chain from VOS — the avalanche tile.smk's - reclaimed-edge cut exists to prevent, arriving through a new target instead. - A tombstone means the copy already happened (clean_exposure cannot run - before exp_persist), so there is nothing to ask for. + EXCEPT AN EXPOSURE WHOSE STORE IS GONE. Its exp_psf manifest is not there, + so requesting its persist manifest would make the DAG rebuild the whole + exposure chain from VOS — the avalanche tile.smk's reclaimed-edge cut exists + to prevent, arriving through a new target instead. exp_store_reclaimed() + below is that test, and what it does NOT test is the tombstone. HEAD PROCESS ONLY, for the same reason as clean_targets() above. """ @@ -509,7 +514,46 @@ def persist_manifests(): return [] exps = {e for t in TILES_READY for e in tile_exposures(t)} return sorted(prod_exp_manifest(e, "exp_persist") for e in exps - if not Path(tombstone(e)).exists()) + if not exp_store_reclaimed(e)) + + +def exp_store_reclaimed(exp): + """True when this exposure's PSF products exist ONLY on the persistent root. + + The one condition both the persist target list and star_cat_merge's edge + choice turn on, and it deliberately does NOT read the tombstone. + + THE TOMBSTONE ALONE IS NOT THE EVIDENCE. It says clean_exposure ran, and + clean_exposure is only one of the two ways a scratch store disappears. + + WHAT THE TEST HAS TO SEPARATE is a store that is GONE from one that has not + been BUILT yet, and no single file says that. This runs at parse time, before + any exp_psf job of a fresh campaign has run, so "the exp_psf manifest is + missing" alone would skip every exposure of a new campaign and persist + nothing at all. The question is therefore: is there evidence this exposure + once had a store? Two files carry it, and either will do: + + * the TOMBSTONE — clean_exposure ran, so the store was built and reclaimed, + and whatever was going to be packed was packed before it went; + * a PERSISTED MANIFEST with no exp_psf manifest beside it — exp_persist + ran, so the store existed, and it is not there now. This is the purge + case, and it is the one the tombstone cannot see: /scratch is purged on + a 60-day window whether or not this workflow reclaimed anything, and it + leaves nothing behind. Keying on the tombstone alone meant that after a + purge, or on any campaign run with `clean: false`, every exposure looked + live, its persist manifest was requested, its exp_psf manifest was not + there, and snakemake rebuilt the entire exposure chain from VOS. + + An exposure with a LIVE store and a manifest is not reclaimed and is still + asked for, which is what lets an edit to `persist_exp:` re-pack in seconds + rather than be silently ignored — the whole reason exp_persist is a rule of + its own. An exposure with neither file is asked for too: either it has not + run yet, or it was purged having saved nothing, and only the DAG can tell + those apart by trying. + """ + return (Path(tombstone(exp)).exists() + or (Path(prod_exp_manifest(exp, "exp_persist")).exists() + and not Path(exp_manifest(exp, "exp_psf")).exists())) # --- the campaign-level merges --------------------------------------------- # Two rules, one job each per campaign, both writing to the persistent root, and @@ -530,23 +574,31 @@ def persist_manifests(): # NEITHER PUTS ITS INPUT PATHS IN ITS SHELL. `{input}` at DR6 scale is ~20k paths # in a single argv entry, an order of magnitude over Linux's 128 KiB # MAX_ARG_STRLEN, and the job would die on exec. So each rule's `input` is the -# DAG EDGE (what must exist first) and each script rediscovers the same set under -# products_dir; what travels is a FINGERPRINT of the input list, on `params`, -# which is what makes the merge rerun when the set changes and not otherwise. +# DAG EDGE (what must exist first) and each script rediscovers the same set from +# the tile list and the index; what travels is a FINGERPRINT of that set's unit +# ids, on `params`, which is what makes the merge rerun when the set changes and +# not otherwise. # The scripts' docstrings argue the rediscovery — it is also what lets the merges # cover exposures whose scratch stores reclamation has since taken. -def input_fingerprint(paths): - """A short digest of an input list, for `params`. +def unit_fingerprint(units): + """A short digest of a set of UNIT IDS, for a merge rule's `params`. + + `params` is a rerun trigger and a set of ids is not: appending a tile grows + the set, moves the digest and reruns the merge, while a rerun over the same + set leaves it alone. Sorted before hashing because the ORDER is not part of + what changed. - `params` is a rerun trigger and a path list is not: two campaigns' worth of - manifests hashes to two different values, so appending a tile reruns the - merge, while a rerun over the same set leaves it alone. Sorted before - hashing because the list's ORDER is not part of what changed. + IDS RATHER THAN THE RULE'S `input` PATHS. A path can change while the set + does not — star_cat_merge's edge for one exposure flips from its manifest to + its tar when the store is reclaimed — and a merge that reruns over identical + content on every reclamation pass is a rerun trigger firing on bookkeeping. + The ids are also exactly what the job derives on its own side, so both + halves agree on the set and on how it is named. """ - joined = "\n".join(sorted(str(p) for p in paths)) - return f"{len(paths)}:{hashlib.md5(joined.encode()).hexdigest()[:12]}" + joined = "\n".join(sorted(str(u) for u in units)) + return f"{len(units)}:{hashlib.md5(joined.encode()).hexdigest()[:12]}" # The tar member the star merge consumes. The keep list is globs, so the test is @@ -624,13 +676,32 @@ def star_cat_inputs(): """ if not PERSIST_EXP: return [] - live = set(persist_manifests()) - exps = {e for t in TILES_READY for e in tile_exposures(t)} - reclaimed = sorted(prod_exp_tar(e) for e in exps - if prod_exp_manifest(e, "exp_persist") not in live - and Path(prod_exp_manifest(e, "exp_persist")).exists() - and Path(prod_exp_tar(e)).exists()) - return sorted(live) + reclaimed + live, reclaimed = [], [] + for exp in sorted({e for t in TILES_READY for e in tile_exposures(t)}): + if not exp_store_reclaimed(exp): + live.append(prod_exp_manifest(exp, "exp_persist")) + elif Path(prod_exp_tar(exp)).exists(): + reclaimed.append(prod_exp_tar(exp)) + return live + reclaimed + + +@functools.lru_cache(maxsize=1) +def star_cat_exposures(): + """The exposure IDs star_cat_merge stacks — what its fingerprint is taken + over. + + THE IDS, NOT THE PATHS, and the difference is a rerun. An exposure's edge + FLIPS from its manifest to its tar the moment its store is reclaimed, so a + fingerprint over paths moves on every reclamation pass and reruns the merge + over content that did not change. The ids move only when the set does, which + is what the trigger is for. It is also what merge_star_cat.py derives on the + job side, so the two agree on the set AND on how it is named. + """ + if not PERSIST_EXP: + return [] + return sorted(e for e in {e for t in TILES_READY for e in tile_exposures(t)} + if Path(prod_exp_manifest(e, "exp_persist")).exists() + or not exp_store_reclaimed(e)) def star_cat_targets(): diff --git a/workflow/rules/exposure.smk b/workflow/rules/exposure.smk index 9c6831549..7b6d33882 100644 --- a/workflow/rules/exposure.smk +++ b/workflow/rules/exposure.smk @@ -274,7 +274,7 @@ rule star_cat_merge: products_dir = str(PRODUCTS_DIR), tile_list = str(config["tile_list"]), index_db = str(INDEX_DB), - inputs = lambda wc, input: input_fingerprint(input), + inputs = unit_fingerprint(star_cat_exposures()), script_hash = MERGE_STAR_HASH threads: 1 resources: diff --git a/workflow/rules/tile.smk b/workflow/rules/tile.smk index cd10ce6cc..ed5daceb4 100644 --- a/workflow/rules/tile.smk +++ b/workflow/rules/tile.smk @@ -944,7 +944,7 @@ rule final_cat_merge: index_db = str(INDEX_DB), param_file = str(CONFIG_DIR / "final_cat.param"), campaign = CAMPAIGN, - inputs = lambda wc, input: input_fingerprint(input), + inputs = unit_fingerprint(TILES_READY), script_hash = MERGE_FINAL_HASH threads: 1 resources: diff --git a/workflow/scripts/build_index.py b/workflow/scripts/build_index.py index 83e9c606b..24290b561 100644 --- a/workflow/scripts/build_index.py +++ b/workflow/scripts/build_index.py @@ -170,8 +170,17 @@ def campaign_tiles(tile_list: Path, db_path: Path) -> list[str]: files — a declared tile with no indexed exposure list cannot have been computed, so it has no catalogue to merge. """ + # DEDUPED, order preserved. The tile list is appended to by hand across a + # campaign, so a tile can appear twice; a merge would then try to write that + # tile's dataset twice and die on the second. Deduping here rather than at + # the call sites keeps the answer the same for every reader of the index. + seen, declared = set(), [] with open(tile_list) as f: - declared = [ln.strip() for ln in f if ln.strip()] + for line in f: + tile = line.strip() + if tile and tile not in seen: + seen.add(tile) + declared.append(tile) con = sqlite3.connect(db_path, timeout=60) indexed = {r[0] for r in con.execute("SELECT DISTINCT tile_id FROM tile_exposures")} con.close() diff --git a/workflow/scripts/merge_final_cat.py b/workflow/scripts/merge_final_cat.py index 2d2974fc0..c12d5fcee 100644 --- a/workflow/scripts/merge_final_cat.py +++ b/workflow/scripts/merge_final_cat.py @@ -90,6 +90,7 @@ import h5py import numpy as np +from astropy.io import fits # Same directory; the rule invokes this file by path, so it is sys.path[0]. import build_index @@ -143,6 +144,16 @@ def catalogues(products_dir: Path, tile_list: Path, index_db: Path) -> list: return out +def check_columns(path: Path, hdu: int, wanted: list) -> None: + """Fail loudly, and by name, when a catalogue lacks a requested column.""" + with fits.open(path, memmap=False) as hdu_list: + present = set(hdu_list[hdu].columns.names) + missing = sorted(c for c in wanted if c not in present) + if missing: + sys.exit(f"merge_final_cat: {path} is missing {len(missing)} of the " + f"{len(wanted)} requested column(s): {' '.join(missing)}") + + def main() -> None: p = argparse.ArgumentParser(description=__doc__) p.add_argument("--products-dir", required=True, type=Path, @@ -184,6 +195,14 @@ def main() -> None: group = hdf5_file.create_group(spval_group(args.campaign)) columns = None for tile, path in tiles: + # BEFORE read_data, and not inside it. read_data wraps its + # column selection in a bare `except:` that prints and falls + # through, so a missing column leaves its return values unbound + # and the caller sees UnboundLocalError from the return + # statement — the real name, and every other missing name, never + # reaches the caller at all. Reading the header costs nothing + # next to reading the table. + check_columns(path, args.hdu, params["param_list"]) extracted, dtype = cfc.read_data(str(path), params) # Requested columns, in the SOURCE catalogue's order (see the # module docstring on determinism). Computed from the first @@ -193,10 +212,6 @@ def main() -> None: if columns is None: columns = [c for c in dtype.names if c in set(params["param_list"])] - missing = sorted(set(params["param_list"]) - set(columns)) - if missing: - sys.exit(f"merge_final_cat: {path} has none of the " - f"requested column(s): {' '.join(missing)}") subset = np.dtype([(c, dtype[c]) for c in columns]) group.create_dataset( tile, diff --git a/workflow/scripts/merge_star_cat.py b/workflow/scripts/merge_star_cat.py index a4f2e3d9f..25c6e4c22 100644 --- a/workflow/scripts/merge_star_cat.py +++ b/workflow/scripts/merge_star_cat.py @@ -20,17 +20,18 @@ handed, and where the result lands. A column added to the module is a column added here for free — which is the entire reason for the indirection. -IT READS THE TARS, IT DOES NOT UNPACK THEM. ``exp_persist`` packs each -exposure's keepers into one uncompressed tar on the persistent root +IT READS THE TARS, IT DOES NOT UNPACK THEM, AND IT STREAMS. ``exp_persist`` +packs each exposure's keepers into one uncompressed tar on the persistent root (``/exp///psf/.tar``) precisely because inodes, not bytes, bind on /project. Unpacking ~20k tars × ~40 members to merge them would materialise ~800k files on the filesystem that design exists to protect, -and then delete them. So members are read into memory -(``tarfile.extractfile(m).read()`` -> ``io.BytesIO``) one at a time and handed -to the merge class as ``[fileobj, member_name]`` pairs. The member NAME is what -the CCD_NB regex parses, which is why the pair carries it; the class takes the -name from the last element of the entry, so a plain ``[path]`` entry behaves -exactly as it always did. +and then delete them. So members are read out of the tars in memory +(``tarfile.extractfile(m).read()`` -> ``io.BytesIO``) and handed to the merge +class as ``[fileobj, member_name]`` pairs — ONE AT A TIME, lazily, through +``TarMembers`` below, because materialising them all first is ~40 GB at DR6 +scale. The member NAME is what the CCD_NB regex parses, which is why the pair +carries it; the class takes the name from the last element of the entry, so a +plain ``[path]`` entry behaves exactly as it always did. WHICH EXPOSURES — AND WHY THE JOB DERIVES THE SET RATHER THAN BEING TOLD IT. The set is the CAMPAIGN's: every exposure read by a tile that is both declared @@ -111,7 +112,15 @@ def merge_class(psf_model: str): - """The merge class for this PSF model — the one-line MCCD/setools hook.""" + """The merge class for this PSF model — the one-line MCCD/setools hook. + + Only psfex is exercised: it is what every campaign has run. MCCD reaches the + tars unchanged (it takes its CCD numbers from the data, and it now reports + by the entry's name like the others). SETOOLS would need one more thing — + it passes ``input_file_list[0][0]`` to file_io as a template path, which a + streamed entry is not — so wiring setools to this path is a change to that + class, not a change here. + """ try: return {"psfex": merge_starcat.MergeStarCatPSFEX, "mccd": merge_starcat.MergeStarCatMCCD, @@ -135,18 +144,14 @@ def manifests(products_dir: Path, tile_list: Path, index_db: Path) -> list: return out -def entries(manifest_paths: list, pattern: str) -> tuple: - """``[fileobj, member_name]`` for every matching member, and the tar count. +def selection(manifest_paths: list, pattern: str) -> tuple: + """``[(tar path, [member names])]`` for the merge, and the empty exposures. - One tar is opened at a time and its members are read into memory; the tars - are never unpacked to disk (see the module docstring). The returned file - objects are BytesIO, so nothing stays open on the filesystem — at ~50 KB per - member and ~40 members per exposure this is ~2 MB per exposure held only for - as long as the merge takes to consume it, but note that the merge class - holds the whole stack in python lists regardless, which is the real memory - term the rule's mem_mb is sized against. + Reads the manifests only. Every tar is checked for existence HERE, so a + products root missing a file fails before a single row is stacked rather + than an hour in. """ - out, n_tars, empty = [], 0, [] + chosen, empty = [], [] for man_path in manifest_paths: man = json.loads(man_path.read_text()) wanted = sorted(f["name"] for f in man["files"] @@ -158,15 +163,40 @@ def entries(manifest_paths: list, pattern: str) -> tuple: if not tar_path.exists(): sys.exit(f"merge_star_cat: {man_path} names a tar that is not " f"there: {tar_path}") - with tarfile.open(tar_path) as tf: - for name in wanted: - member = tf.extractfile(name) - if member is None: - sys.exit(f"merge_star_cat: {tar_path} has no member " - f"{name}, which its manifest lists") - out.append([io.BytesIO(member.read()), name]) - n_tars += 1 - return out, n_tars, empty + chosen.append((tar_path, wanted)) + return chosen, empty + + +class TarMembers: + """The merge class's input list, materialised ONE TAR AT A TIME. + + ``MergeStarCatPSFEX`` wants something it can take the length of and iterate + once, handing it ``[fileobj, name]`` entries; it never indexes and never + rewinds. So it does not need a list, and a list is the one thing we cannot + afford: reading every member up front is the whole campaign in memory at + once — ~2 MB per exposure, so ~40 GB at DR6's ~20k exposures, against a + rule asking for 16 GB. Read lazily, peak memory is ONE member's bytes plus + the merge class's own accumulators, which are the real and unavoidable term. + + ``__len__`` comes from the manifests, so the class can log the count before + a single tar is opened. + """ + + def __init__(self, chosen): + self._chosen = chosen + + def __len__(self): + return sum(len(names) for _, names in self._chosen) + + def __iter__(self): + for tar_path, names in self._chosen: + with tarfile.open(tar_path) as tf: + for name in names: + member = tf.extractfile(name) + if member is None: + sys.exit(f"merge_star_cat: {tar_path} has no member " + f"{name}, which its manifest lists") + yield [io.BytesIO(member.read()), name] def main() -> None: @@ -196,8 +226,9 @@ def main() -> None: level=logging.INFO, stream=sys.stdout) manifest_paths = manifests(args.products_dir, args.tile_list, args.index_db) - file_list, n_tars, empty = entries(manifest_paths, args.pattern) - if not file_list: + chosen, empty = selection(manifest_paths, args.pattern) + file_list = TarMembers(chosen) + if not len(file_list): # Not a no-op: an empty star catalogue would pass every downstream # existence check and produce meaningless rho statistics. sys.exit(f"merge_star_cat: no member matched {args.pattern!r} in any " @@ -224,8 +255,8 @@ def main() -> None: log.info(f"unchanged: {args.output}") else: tmp.replace(args.output) # atomic: same filesystem - log.info(f"{len(file_list)} catalogue(s) from {n_tars} exposure(s) " - f"-> {args.output}") + log.info(f"{len(file_list)} catalogue(s) from {len(chosen)} " + f"exposure(s) -> {args.output}") finally: shutil.rmtree(tmp_dir, ignore_errors=True) diff --git a/workflow/scripts/persist_exp.py b/workflow/scripts/persist_exp.py index e78d3e9a4..40e8c9354 100644 --- a/workflow/scripts/persist_exp.py +++ b/workflow/scripts/persist_exp.py @@ -127,10 +127,20 @@ def main() -> None: args.dest.mkdir(parents=True, exist_ok=True) tar_path = args.dest / f"{args.exp}.tar" + # A file matched by TWO patterns is one file, not a collision. Keep lists + # overlap on purpose — `validation_psf-*.fits` alongside `*.fits` is a + # perfectly ordinary way to say "the validation catalogues, and everything + # else FITS while we are here" — and treating the second match as a name + # clash failed every exposure in the campaign. What must still be fatal is + # two DIFFERENT paths landing on one flat member name, which would silently + # overwrite; that is a same-name/different-source test, and the first + # pattern to match a file is the one recorded for it. seen, files = {}, [] for pat, hits in found.items(): for src in hits: if src.name in seen: + if seen[src.name][0] == src: + continue # same file, a second matching pattern sys.exit(f"persist_exp: {args.exp}: two source files are both " f"named {src.name} ({seen[src.name][0]} and {src}); tar " f"members are flat, so this would silently overwrite") From d774cc3663fcb31148e932a39bbae2392dae8eba Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 19:18:02 -0400 Subject: [PATCH 09/85] fix(create_final_cat): make the merged-catalogue writer reproducible MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three defects in scripts/python/create_final_cat.py, fixed where they live rather than worked around in the workflow rule that now calls it. A hand-run of the tool deserves them as much as the rule does, and files it has already written carry the first one. copy_data allocated np.empty with the SOURCE catalogue's full dtype and then filled only the requested columns, so every column NOT in the parameter file reached the hdf5 file as uninitialised memory: meaningless values, and different bytes on every run over the same inputs. It allocates the requested columns alone now, in the source catalogue's order. The parameter file says what the merged catalogue is for; those are the columns it gets. read_param_file returned list(set(...)), whose order varies with the process's string hash seed. Column order is part of a structured dtype and therefore part of the file, so two runs over the same inputs disagreed. Ordered dedup via dict.fromkeys. (The duplicate-count message also only printed for more than one duplicate, and said {n} literally.) read_data wrapped its column selection in a bare `except:` that printed and fell through, leaving its return values unbound — so a missing column surfaced to the caller as UnboundLocalError from the return statement, naming neither the file nor the column. It raises a KeyError naming the file and every missing column, in parameter-file order. process()'s own create_dataset follows the array copy_data returns rather than the source dtype, which are no longer the same thing. merge_final_cat.py drops the equivalents it had been carrying at the call site and relies on the fixed functions. The fixture hdf5 is byte-identical either way (md5 6de2d261…): the workaround and the fix produce the same file, which is the point. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- scripts/python/create_final_cat.py | 59 ++++++++++++++++++---------- workflow/scripts/merge_final_cat.py | 61 +++++------------------------ 2 files changed, 47 insertions(+), 73 deletions(-) diff --git a/scripts/python/create_final_cat.py b/scripts/python/create_final_cat.py index 2b583b857..efef24620 100755 --- a/scripts/python/create_final_cat.py +++ b/scripts/python/create_final_cat.py @@ -144,13 +144,17 @@ def read_param_file(path, verbose=False): print("No parameters read", end="") print(" into merged catalogue") - param_list_unique = list(set(param_list)) - + # Ordered dedup. list(set(...)) reordered the columns by the process's + # string hash seed, so two runs of this tool over the same inputs produced + # files whose datasets differed in column ORDER — which is part of a + # structured dtype, and therefore part of the file. + param_list_unique = list(dict.fromkeys(param_list)) + if verbose: n = len(param_list) - len(param_list_unique) - if n > 1: - print("Removed {n} duplicate entries") - + if n > 0: + print(f"Removed {n} duplicate entries") + return param_list_unique @@ -312,16 +316,20 @@ def read_data(fits_file, params): if params["param_list"] is None: params["param_list"] = [col for col in data.keys()] - try: - extracted_data = {col: data[col] for col in params["param_list"]} - dtype = data.dtype - except: - print(f"Error for ID {id}, path {fits_file}") - for col in params["param_list"]: - if col not in data: - print(col, end=" ") - print() - continue + # RAISE, do not print and fall through. The bare `except:` this replaces + # left extracted_data and dtype unbound, so the caller's own error was an + # UnboundLocalError from the return statement below, naming neither the + # file nor the column that was actually missing. + present = set(data.dtype.names or ()) + missing = [col for col in params["param_list"] if col not in present] + if missing: + raise KeyError( + f"{fits_file}: missing {len(missing)} of the " + f"{len(params['param_list'])} requested column(s): " + f"{' '.join(missing)}" + ) + extracted_data = {col: data[col] for col in params["param_list"]} + dtype = data.dtype return extracted_data, dtype @@ -330,16 +338,23 @@ def copy_data(param_list, extracted_data, dtype): """Copy Data. """ + # THE REQUESTED COLUMNS ONLY, in the SOURCE catalogue's order. Allocating + # with the source's full dtype and filling only the requested columns left + # every other column as uninitialised memory: meaningless values in the + # output file, and different bytes on every run of this tool over the same + # inputs. The parameter file says which columns the merged catalogue is + # for; those are the columns it gets. + columns = [col for col in (dtype.names or ()) if col in set(param_list)] + subset = np.dtype([(col, dtype[col]) for col in columns]) + # Initialize new data structure structured_data = np.empty( len(extracted_data[param_list[0]]), - dtype=dtype, + dtype=subset, ) # Loop over parameters - for col in param_list: - if not col in extracted_data: - print(f"Column {col} not in file with ID {id}") + for col in columns: structured_data[col] = extracted_data[col] #if isinstance(extracted_data[col][0], (np.ndarray, tuple, list)): @@ -467,12 +482,14 @@ def process(params): structured_data = copy_data(params["param_list"], extracted_data, dtype) - # Create a new dataset + # Create a new dataset. dtype comes from the array copy_data + # built, not from the source catalogue: they differ now that + # copy_data allocates the requested columns alone. try: patch_group.create_dataset( str(id), data=structured_data, - dtype=dtype, + dtype=structured_data.dtype, ) except: print(f"Error for {id}: Could not create dataset in group {patch}") diff --git a/workflow/scripts/merge_final_cat.py b/workflow/scripts/merge_final_cat.py index c12d5fcee..58357fcd5 100644 --- a/workflow/scripts/merge_final_cat.py +++ b/workflow/scripts/merge_final_cat.py @@ -18,6 +18,13 @@ ``create_final_cat.py``'s — ``read_param_file`` for the parameter list, ``read_data`` and ``copy_data`` for pulling those columns out of one catalogue with their FITS dtypes — so the column grammar keeps exactly one definition. +Those three are REPRODUCIBLE FUNCTIONS, and this PR is what made them so: the +parameter list comes back ordered rather than through a set, ``copy_data`` +allocates the requested columns alone rather than leaving every other column of +the source as uninitialised memory, and a missing column raises with its own +name instead of falling out of a bare ``except:`` as an UnboundLocalError. The +fixes are upstream, in that script, because a hand-run of it deserves them as +much as this rule does. Its ``process()`` is NOT used and neither is any of its discovery: that function walks a directory tree the workflow does not have and never will, and it groups by a unit ShapePipe v2 no longer has. This script walks the workflow's own @@ -46,23 +53,6 @@ are visited in sorted ID order so the file is a function of the input set alone. An unconditional rewrite would move the output's mtime every invocation. -TWO PLACES WHERE THE REFERENCE IMPLEMENTATION IS NOT DETERMINISTIC, and where -this script therefore pins the behaviour down rather than copying it. Both are -in the DTYPE, and both are invisible when a human runs the tool once by hand: - - * ``copy_data`` allocates ``np.empty`` with the SOURCE catalogue's full - dtype and then fills only the requested columns, so every column NOT in - ``final_cat.param`` reaches the hdf5 file as uninitialised memory — - different bytes on every run, and meaningless data in the file besides. We - hand ``copy_data`` a dtype restricted to the requested columns, so every - field it writes is a field it fills. The file then carries exactly the - ``final_cat.param`` columns, which is what sp_validation reads and what the - parameter file is for. - * ``read_param_file`` returns ``list(set(...))``, whose order varies with the - process's string hash seed. Column ORDER in a structured dtype is part of - the file, so that alone would defeat the byte comparison. We order the - fields by the source catalogue's own column order instead. - WHICH TILES — AND WHY THE JOB DERIVES THE SET RATHER THAN BEING TOLD IT. The set is the CAMPAIGN's: every tile both declared in ``tile_list`` and present in the index, which is exactly the Snakefile's TILES_READY, rebuilt here from the same @@ -89,8 +79,6 @@ from pathlib import Path import h5py -import numpy as np -from astropy.io import fits # Same directory; the rule invokes this file by path, so it is sys.path[0]. import build_index @@ -144,16 +132,6 @@ def catalogues(products_dir: Path, tile_list: Path, index_db: Path) -> list: return out -def check_columns(path: Path, hdu: int, wanted: list) -> None: - """Fail loudly, and by name, when a catalogue lacks a requested column.""" - with fits.open(path, memmap=False) as hdu_list: - present = set(hdu_list[hdu].columns.names) - missing = sorted(c for c in wanted if c not in present) - if missing: - sys.exit(f"merge_final_cat: {path} is missing {len(missing)} of the " - f"{len(wanted)} requested column(s): {' '.join(missing)}") - - def main() -> None: p = argparse.ArgumentParser(description=__doc__) p.add_argument("--products-dir", required=True, type=Path, @@ -193,31 +171,10 @@ def main() -> None: tmp.unlink(missing_ok=True) # h5py "a" would reopen a stale one with h5py.File(tmp, "w") as hdf5_file: group = hdf5_file.create_group(spval_group(args.campaign)) - columns = None for tile, path in tiles: - # BEFORE read_data, and not inside it. read_data wraps its - # column selection in a bare `except:` that prints and falls - # through, so a missing column leaves its return values unbound - # and the caller sees UnboundLocalError from the return - # statement — the real name, and every other missing name, never - # reaches the caller at all. Reading the header costs nothing - # next to reading the table. - check_columns(path, args.hdu, params["param_list"]) extracted, dtype = cfc.read_data(str(path), params) - # Requested columns, in the SOURCE catalogue's order (see the - # module docstring on determinism). Computed from the first - # tile and reused, so a tile whose catalogue is missing a - # column fails loudly on the assignment rather than quietly - # producing a differently-shaped dataset. - if columns is None: - columns = [c for c in dtype.names - if c in set(params["param_list"])] - subset = np.dtype([(c, dtype[c]) for c in columns]) - group.create_dataset( - tile, - data=cfc.copy_data(columns, extracted, subset), - dtype=subset, - ) + data = cfc.copy_data(params["param_list"], extracted, dtype) + group.create_dataset(tile, data=data, dtype=data.dtype) # The same attribute create_final_cat.py's print_list() writes, and # what sp_validation reads to know how many tiles it is holding. hdf5_file.attrs["n_tiles"] = len(tiles) From 98532ceccad08c0e5051e03a733be2f71fc1557a Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 19:22:07 -0400 Subject: [PATCH 10/85] perf(orchestration): size the two merges from the data, not from a guess MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both merges had a constant mem_mb, which is wrong by however much a campaign differs from the one it was tuned on — and these are the only two rules whose single job grows with the whole campaign. Both are now measured slopes, evaluated against the campaign's own bytes at DAG build, still * attempt. STAR SIDE, measured on this login node in the campaign container over synthetic tars, 20 and 80 exposures of 40 CCDs x 400 stars: input members peak RSS (getrusage RUSAGE_CHILDREN) 32.3 MB 383 MB 129.0 MB 1313 MB a slope of 10.1x input bytes over a ~73 MB interpreter floor. Tenfold because MergeStarCatPSFEX accumulates every column into python LISTS of python floats before building the output arrays. THE CONSEQUENCE IS A CEILING and the Snakefile says so: at ~2 MB of members per exposure a 16 GB job merges roughly 800 exposures, and DR6's ~20k would want ~400 GB. A full-survey full_starcat needs that accumulation changed to preallocated arrays or a two-pass count — a change to MergeStarCatPSFEX, not to this rule, and not in this PR. The formula is honest about the slope so the job asks for what it will use and fails at submission rather than most of the way through. TILE SIDE, measured against smk-g6's real catalogues, 2 tiles (73.9 MB in, largest 39.6 MB) and 6 tiles (235.5 MB in, largest 47.7 MB): peak RSS 129 MB and 139 MB. FLAT in the tile count, because the merge holds one catalogue at a time — so it is sized on the LARGEST tile at ~3x, not on the total. Runtime is the total, since every tile is read end to end. On smk-g6's 64 tiles the rule resolves to mem_mb=1002, runtime=51, against the flat 8000/120 it had. Sizes come from stat() on the tar or the catalogue, falling back to the measured per-unit default when a fresh campaign has not produced it yet. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- workflow/Snakefile | 72 +++++++++++++++++++++++++++++++++++++ workflow/rules/exposure.smk | 12 +++++-- workflow/rules/tile.smk | 11 ++++-- 3 files changed, 91 insertions(+), 4 deletions(-) diff --git a/workflow/Snakefile b/workflow/Snakefile index 587f271b2..8889b7f38 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -704,6 +704,78 @@ def star_cat_exposures(): or not exp_store_reclaimed(e)) +# --- sizing the two merges (D4) --------------------------------------------- +# MEASURED, not guessed, and measured as a SLOPE rather than a single number: +# these are the only two rules whose one job's footprint grows with the whole +# campaign, so a constant is wrong by however much the campaign is not the one +# it was tuned on. +# +# Both slopes were measured on this login node, inside the campaign container, +# against synthetic tars for the star side and against smk-g6's real +# catalogues for the tile side. Peak RSS is getrusage(RUSAGE_CHILDREN). +# +# STAR SIDE, and it is the alarming one. Two points, 20 and 80 exposures of 40 +# CCDs x 400 stars (1.6 MB of members per exposure, against the 2.0 MB measured +# on smk-m2): +# +# input members peak RSS +# 32.3 MB 383 MB +# 129.0 MB 1313 MB +# +# a slope of 10.1x the input bytes and an intercept of ~73 MB (interpreter, +# astropy, shapepipe). Tenfold, because MergeStarCatPSFEX accumulates every +# column into PYTHON LISTS of python floats before building the output arrays — +# 8 bytes of payload becomes a 32-byte object plus an 8-byte pointer. THE +# CONSEQUENCE IS A CEILING, and it should be said plainly: at ~2 MB per +# exposure, a 16 GB job merges roughly 800 exposures, and DR6's ~20k exposures +# would want ~400 GB. A full-survey full_starcat needs the accumulation changed +# to preallocated arrays or a two-pass count — a change to MergeStarCatPSFEX, +# not to this rule, and not in this PR. The formula below is honest about the +# slope so the job asks for what it will use and fails at submission rather +# than at 90% of the way through a campaign-length merge. +# +# TILE SIDE, and it is the reassuring one. Two points against real smk-g6 +# catalogues, 2 tiles (73.9 MB in, largest 39.6 MB) and 6 tiles (235.5 MB in, +# largest 47.7 MB): peak RSS 129 MB and 139 MB. FLAT IN THE NUMBER OF TILES — +# the merge holds one catalogue at a time — so it is sized on the LARGEST tile, +# not the total, at ~3x it plus the interpreter. +STAR_MEM_BASE_MB = 500 # interpreter + astropy + shapepipe, rounded up +STAR_MEM_FACTOR = 12 # x input bytes; 10.1 measured, rounded up +FINAL_MEM_BASE_MB = 800 +FINAL_MEM_FACTOR = 4 # x the LARGEST tile; ~3 measured +# What one unit costs when its product is not on disk yet to be stat()ed — a +# fresh campaign sizes its merge before anything has been packed or made. +# Both are the measured medians in config.yaml's persist_exp block and D5 notes. +EXP_BYTES_DEFAULT = 2_000_000 +TILE_BYTES_DEFAULT = 46_000_000 + + +def _size(path, default): + """Bytes on disk, or the documented per-unit default if it is not there.""" + try: + return Path(path).stat().st_size + except OSError: + return default + + +def star_cat_bytes(): + """Total member bytes the star merge will read. + + The TAR is what gets stat()ed, not the manifest: it is one stat per + exposure rather than a json parse, and it is present for exactly the + exposures whose products already exist — live-and-already-packed as well as + reclaimed. An exposure not yet packed contributes the measured default. + """ + return sum(_size(prod_exp_tar(e), EXP_BYTES_DEFAULT) + for e in star_cat_exposures()) + + +def final_cat_max_bytes(): + """The LARGEST tile catalogue the hdf5 merge will read — what sizes it.""" + return max([_size(final_cat(t), TILE_BYTES_DEFAULT) for t in TILES_READY] + or [TILE_BYTES_DEFAULT]) + + def star_cat_targets(): """`full_starcat` when there is anything to stack into it, else nothing. diff --git a/workflow/rules/exposure.smk b/workflow/rules/exposure.smk index 7b6d33882..789e9da1d 100644 --- a/workflow/rules/exposure.smk +++ b/workflow/rules/exposure.smk @@ -278,8 +278,16 @@ rule star_cat_merge: script_hash = MERGE_STAR_HASH threads: 1 resources: - mem_mb = lambda wc, attempt: 16000 * attempt, - runtime = 120 + # Sized on the campaign's own member bytes, slope and intercept + # measured (the Snakefile's sizing block carries both points, and the + # ceiling this rule runs into at DR6 scale). Still * attempt, because a + # measured slope on synthetic tars is not a guarantee about real ones. + mem_mb = lambda wc, attempt: attempt * ( + STAR_MEM_BASE_MB + STAR_MEM_FACTOR * star_cat_bytes() // 1_000_000), + # ~2 min per GB of members on the measurement above, doubled, over a + # floor that covers the fixed cost of opening ~40 members per exposure. + runtime = lambda wc, attempt: attempt * ( + 30 + 4 * star_cat_bytes() // 1_000_000_000) shell: "set -euo pipefail\n" f"python {SCRIPTS}/merge_star_cat.py" diff --git a/workflow/rules/tile.smk b/workflow/rules/tile.smk index ed5daceb4..4eb293b04 100644 --- a/workflow/rules/tile.smk +++ b/workflow/rules/tile.smk @@ -948,8 +948,15 @@ rule final_cat_merge: script_hash = MERGE_FINAL_HASH threads: 1 resources: - mem_mb = lambda wc, attempt: 8000 * attempt, - runtime = 120 + # Sized on the LARGEST tile, not the total: the merge holds one + # catalogue at a time, and the measurement is flat in the tile count + # (the Snakefile's sizing block carries both points). + mem_mb = lambda wc, attempt: attempt * ( + FINAL_MEM_BASE_MB + + FINAL_MEM_FACTOR * final_cat_max_bytes() // 1_000_000), + # Runtime, unlike memory, is the TOTAL: every tile is read end to end. + # ~1 min per 10 tiles on the measurement, triply generous, over a floor. + runtime = lambda wc, attempt: attempt * (30 + len(TILES_READY) // 3) shell: "set -euo pipefail\n" f"python {SCRIPTS}/merge_final_cat.py" From d0ecfdd5283455460f333e0b5d556b7c9f2d241a Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 19:25:23 -0400 Subject: [PATCH 11/85] feat(orchestration): persist_exp keeps NAMED products, not globs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Closes the readability half of CosmoStat/shapepipe#844, and makes the 2026-09-08 call's request — keep the PSF model — something you can write down as `psf_model` rather than `*.psf`. `persist_exp:` entries are now names from a catalogue in persist_exp.py, which is the single source of truth for what each one means, what it costs per exposure and what keeping it buys: psf_validation validation_psf-*.fits 2.0 MB psf_model *.psf 2.8 MB psfex_cat psfex_cat-*.cat unmeasured star_selection star_selection-*.fits 24.5 MB star_train star_split_ratio_80-*.fits 19.9 MB star_test star_split_ratio_20-*.fits 7.1 MB star_stats star_stat-*.txt unmeasured `persist_exp.py --list-products` renders it, and config.yaml's block IS that rendering rather than a second copy of it — the old block was a long comment listing globs and their sizes, maintained by hand beside the code that actually knew them. THE DEFAULT BECOMES psf_validation + psf_model, ~4.8 MB per exposure. The model is the single most capability-adding thing an exposure can keep: with it the PSF can be re-interpolated at any position later without rebuilding the chain from VOS, and without it that capability dies with the /scratch purge. A raw glob is still accepted as an escape hatch for a file the catalogue does not name yet. The test is syntactic and cheap — a glob metacharacter or a dot means glob, a bare identifier means name — so `*.psf` and `psf_model` cannot be confused. An unknown NAME is a parse-time WorkflowError listing the valid ones, not a silently empty keep or a per-exposure failure an hour in. The manifest records both: `products` as written, `patterns` resolved, and each member's own `product`. star_cat_merge's gate and its member glob resolve through the same catalogue, so adding a product cannot leave the two disagreeing, and the "add this to persist_exp" hint now names the product. Tile-side retention is explicitly out of scope and noted as such in config.yaml: final_cat is the only tile product that persists today, and it is written straight to products_dir by tile_make_cat. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- workflow/README.md | 25 ++++- workflow/Snakefile | 29 ++++-- workflow/config.yaml | 118 +++++++++++---------- workflow/scripts/merge_star_cat.py | 16 +-- workflow/scripts/persist_exp.py | 159 +++++++++++++++++++++++++++-- 5 files changed, 267 insertions(+), 80 deletions(-) diff --git a/workflow/README.md b/workflow/README.md index eb801da6a..1ad740b48 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -245,8 +245,7 @@ profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; kee Know the consequence — `--forcerun` on a tile whose `final_cat` exists will not rebuild its reclaimed exposures. Delete the `final_cat` first. - **PSF products leave scratch before the purge does.** `exp_persist` packs - the files named by `persist_exp:` in `config.yaml` (default: the psfex_interp - `validation_psf-*.fits`, the rho/tau statistics input) from the exposure's + the products named by `persist_exp:` in `config.yaml` from the exposure's scratch store into ONE uncompressed tar, `/exp///psf/.tar` (inodes, not bytes, bind on /project), and writes ONE manifest beside it recording the patterns, the @@ -260,6 +259,28 @@ profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; kee hours of PSF fitting per exposure. A pattern that matches nothing is a recorded warning (setools rejects sparse CCDs); matching nothing at all is a failure. A `localrule`, by the same arithmetic as `clean_exposure`. +- **The keep list names products, not globs.** `persist_exp:` entries are names + from a catalogue in `workflow/scripts/persist_exp.py`, which is the single + source of truth for what each one means and what keeping it buys + ([#844](https://github.com/CosmoStat/shapepipe/issues/844)); `config.yaml`'s + block is that catalogue rendered, and + `persist_exp.py --list-products` prints it. Sizes are per exposure, 40 CCDs, + measured on smk-m2. + + | product | glob | per exposure | what it buys | + |---|---|---|---| + | `psf_validation` | `validation_psf-*.fits` | 2.0 MB | the rho/tau statistics input, and `star_cat_merge`'s | + | `psf_model` | `*.psf` | 2.8 MB | re-interpolate the PSF anywhere later, no rebuild | + | `psfex_cat` | `psfex_cat-*.cat` | unmeasured | which stars PSFEx's outlier rejection clipped | + | `star_selection` | `star_selection-*.fits` | 24.5 MB | which stars the selection cuts rejected, and why | + | `star_train` | `star_split_ratio_80-*.fits` | 19.9 MB | the 80% sample PSFEx fitted | + | `star_test` | `star_split_ratio_20-*.fits` | 7.1 MB | the 20% sample `psf_validation` corresponds to | + | `star_stats` | `star_stat-*.txt` | unmeasured | setools' per-CCD counts, density and FWHM cuts | + + The default is `psf_validation` + `psf_model`. A raw glob is still accepted as + an escape hatch — anything with a glob metacharacter or a dot is read as one — + and an unknown *name* is a parse-time error listing the valid ones. The list + is exposure-side only; tile-side retention is #844 follow-up. - **The campaign ends in two merged catalogues, and the workflow now makes both.** Everything above is per unit; the two products downstream analysis actually opens are per *campaign*, and until these rules existed each was a diff --git a/workflow/Snakefile b/workflow/Snakefile index 8889b7f38..8f871df3b 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -472,6 +472,21 @@ def clean_targets(): # deliberate "keep nothing" and produces no jobs at all. PERSIST_EXP = list(config.get("persist_exp") or []) +# The keep list names PRODUCTS (`psf_model`), not globs (`*.psf`); the +# catalogue that maps one to the other lives in persist_exp.py, which is also +# what the rule runs, so there is one definition and not a copy here. +# UNKNOWN NAMES DIE AT PARSE TIME, listing the valid ones — a typo in a keep +# list would otherwise be a silently-empty keep or a per-exposure failure an +# hour into a campaign. +import persist_exp as _persist # noqa: E402 + +for _entry in PERSIST_EXP: + try: + _persist.resolve(_entry) + except KeyError as _exc: + raise WorkflowError(f"config persist_exp: {_exc.args[0]}") +PERSIST_GLOBS = [_persist.resolve(e) for e in PERSIST_EXP] + def persist_targets(): """Which exposures this invocation must pack PSF products off scratch for. @@ -601,12 +616,14 @@ def unit_fingerprint(units): return f"{len(units)}:{hashlib.md5(joined.encode()).hexdigest()[:12]}" -# The tar member the star merge consumes. The keep list is globs, so the test is -# "would this member be kept", not a string comparison — `validation_psf-*.fits`, -# `validation_psf*`, `*.fits` and a bare `*` all say yes, and all are things a -# user might reasonably write. +# The tar member the star merge consumes, and the test is "would this member be +# kept" rather than "is `psf_validation` in the list". Both spellings must count: +# the product name, and a raw glob that happens to cover it (`validation_psf*`, +# `*.fits`, a bare `*`) — all things a user might reasonably write. So the keep +# list is RESOLVED to globs first and the member matched against those. +_STAR_CAT_PRODUCT = "psf_validation" _STAR_CAT_MEMBER = "validation_psf-2605805-12.fits" -STAR_CAT_MERGE = any(fnmatch.fnmatch(_STAR_CAT_MEMBER, p) for p in PERSIST_EXP) +STAR_CAT_MERGE = any(fnmatch.fnmatch(_STAR_CAT_MEMBER, g) for g in PERSIST_GLOBS) # LOUD AT PARSE TIME, once, and only where it can be acted on. A keep list # without the validation catalogues is a legitimate configuration (persist the @@ -620,7 +637,7 @@ if PERSIST_EXP and not STAR_CAT_MERGE and workflow.is_main_process \ f"star_cat_merge: no job — persist_exp {PERSIST_EXP} keeps no " f"'{_STAR_CAT_MEMBER}'-shaped file, so there is nothing to stack into " f"{PRODUCTS_DIR}/full_starcat-0000000.fits (the rho/tau statistics " - f"input). Add 'validation_psf-*.fits' to persist_exp: to get it.") + f"input). Add '{_STAR_CAT_PRODUCT}' to persist_exp: to get it.") def full_starcat(): diff --git a/workflow/config.yaml b/workflow/config.yaml index 70ace906e..03c282113 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -82,15 +82,54 @@ outputs: index_db: /project/def-mjhudson/cdaley/sp-products/smk-g6/index/run_index.sqlite # Per-exposure PSF products to carry onto the persistent root before the scratch -# store goes (`exp_persist`, exposure.smk). A list of plain file-name globs, -# matched recursively under the PSF chain's four module output dirs -# (/exp///output/run_sp_exp_SxSePsfPi/*/output/ — -# sextractor_runner, setools_runner, psfex_runner, psfex_interp_runner). +# store goes (`exp_persist`, exposure.smk). A list of PRODUCT NAMES — not globs. +# The catalogue below is the rendering of workflow/scripts/persist_exp.py's +# PRODUCTS table, which is the single source of truth for what each name means +# and what keeping it buys (CosmoStat/shapepipe#844); print it any time with +# +# workflow/bin/sp container exec python workflow/scripts/persist_exp.py --list-products +# +# product glob size/exposure +# -------------- -------------------------- ------------- +# star_selection star_selection-*.fits 24.5 MB +# setools' PRE-SPLIT selection. The only file that answers which +# stars the selection cuts rejected and why; the split samples have +# already lost the rejects. +# star_train star_split_ratio_80-*.fits 19.9 MB +# the 80% TRAINING sample, the stars PSFEx actually fitted. Rows +# duplicate star_selection. +# star_test star_split_ratio_20-*.fits 7.1 MB +# the 20% VALIDATION sample — the positions psf_validation's rows +# correspond to. Rows duplicate star_selection. +# star_stats star_stat-*.txt unmeasured +# setools' per-CCD STAT block: star counts, stars/deg^2, FWHM mode +# and cuts. The selection's summary without its catalogue. +# psf_model *.psf 2.8 MB +# the PSFEx model itself. Keeping it means the PSF can be +# re-interpolated at ANY position later without rebuilding the +# exposure chain — the single most capability-adding entry here. +# psfex_cat psfex_cat-*.cat unmeasured +# PSFEx's own output catalogue (FITS_LDAC): the per-star FLAGS_PSF +# and CHI2_PSF, i.e. WHICH stars outlier rejection clipped. Not +# recoverable from anything else — the .psf header keeps only the +# LOADED/ACCEPTED counts. +# psf_validation validation_psf-*.fits 2.0 MB +# the psfex_interp validation catalogue, one per CCD: the input to +# the rho/tau statistics, and to the star_cat_merge rule that stacks +# them into the campaign's full_starcat. +# +# A RAW GLOB IS STILL ACCEPTED, as an escape hatch for a file the catalogue does +# not name yet: anything carrying a glob metacharacter or a dot is taken as a +# glob rather than a name (`*.psf` is a glob, `psf_model` is the name for it). +# An unknown NAME is a parse-time error listing the valid ones, never a silently +# empty keep. +# # Matches are packed, flat, into ONE uncompressed tar per exposure: # /exp///psf/.tar, with a manifest listing the -# members beside it. One tar rather than loose copies because inodes, not bytes, -# bind on /project (~1 M-file group quota; loose copies would be ~200 files per -# exposure, ~2 M at DR6 scale). FITS members read straight from the tar: +# members (and the product each came from) beside it. One tar rather than loose +# copies because inodes, not bytes, bind on /project (~1 M-file group quota; +# loose copies would be ~200 files per exposure, ~2 M at DR6 scale). FITS +# members read straight from the tar: # fits.open(io.BytesIO(tarfile.open(t).extractfile(m).read())). # # WHY COPY RATHER THAN EXEMPT THESE FROM CLEANUP. Reclamation is not the threat. @@ -104,56 +143,24 @@ outputs: # reruns the packing (seconds) and NOT exp_psf (four hours per exposure). That # separation is the whole reason exp_persist is a rule of its own. # -# The default is the minimum: the psfex_interp VALIDATION catalogue, one per -# CCD, which is the input to the rho/tau statistics. Without it the PSF -# diagnostics cannot be recomputed after a purge without rebuilding the exposure -# chain from VOS. +# THE DEFAULT is psf_validation + psf_model (~4.8 MB per exposure): the rho/tau +# statistics input, and the model that lets the PSF be re-interpolated at any +# position later without rebuilding the exposure chain from VOS. Add psfex_cat +# for a production run if you want to know which stars PSFEx clipped; the +# star_* products are for selection studies and cost an order of magnitude more. +# +# THIS LIST IS EXPOSURE-SIDE ONLY. Tile-side retention is not configurable: the +# only tile product that persists today is final_cat, written by tile_make_cat +# straight to products_dir. A tile keep list is #844 follow-up. # # THIS LIST GATES `star_cat_merge`. That campaign-level rule stacks every -# exposure's every CCD's validation_psf into ONE +# exposure's every CCD's psf_validation into ONE # /full_starcat-0000000.fits, reading the members straight out of -# the tars. A keep list that matches no `validation_psf-*.fits` is a legitimate -# configuration (keep the PSF models alone, say) and produces NO merge job and a -# warning at parse time — not a failure on a node an hour later. Note the -# corollary: an exposure already reclaimed by a workflow that predates -# exp_persist left no tar, so it contributes nothing and cannot be recovered -# short of rebuilding its chain from VOS. -# -# OPT-IN CANDIDATES, and what each buys. Sizes are per exposure (40 CCDs), -# measured on smk-m2 (127 exposures, 64 tiles); a 64-tile campaign with all of -# the measured ones on came to 7.2 GB: -# validation_psf-*.fits (the default) 2.0 MB -# *.psf the PSFEx model itself. Keeping it means the PSF -# can be re-interpolated at ANY position later -# without rebuilding the exposure chain — the -# single most capability-adding entry here. -# 2.8 MB -# psfex_cat-*.cat PSFEx's own output catalogue (FITS_LDAC): the -# per-star FLAGS_PSF / CHI2_PSF, i.e. WHICH stars -# outlier rejection clipped. Not recoverable from -# anything else (the .psf header keeps only the -# LOADED/ACCEPTED counts). unmeasured -# star_selection-*.fits the PRE-SPLIT selection (setools writes it under -# mask/). The only file that can answer "which -# stars were rejected by the selection cuts, and -# why" — the split samples have already lost the -# rejects. 24.5 MB -# star_split_ratio_80-*.fits setools' 80% TRAINING star sample, the set PSFEx -# actually fitted. Rows duplicate star_selection. -# 19.9 MB -# star_split_ratio_20-*.fits the 20% VALIDATION sample — the positions the -# validation_psf rows correspond to. Rows -# duplicate star_selection. 7.1 MB -# star_stat-*.txt setools' per-CCD STAT block (star counts, -# stars/deg^2, FWHM mode and cuts, under stat/): -# the selection's summary without its catalogue. -# unmeasured -# A production keep list is `validation_psf` + `*.psf` + `psfex_cat` (~5 MB per -# exposure); the star_split files are only worth it if star_selection is off. -# PSFEx residual/check images and its XML diagnostics are NOT candidates as the -# chain stands: the committed default.psfex sets CHECKIMAGE_TYPE NONE and -# WRITE_XML N, so nothing is emitted to match. They are a config change first, -# a pattern second. +# the tars. A keep list without psf_validation is a legitimate configuration and +# produces NO merge job and a warning at parse time — not a failure on a node an +# hour later. Note the corollary: an exposure already reclaimed by a workflow +# that predates exp_persist left no tar, so it contributes nothing and cannot be +# recovered short of rebuilding its chain from VOS. # # NOTE ON products_dir DEFAULTING TO run_dir (a fixture or smoke test): the tar # then lands beside the store on the same filesystem and buys nothing, and the @@ -161,7 +168,8 @@ outputs: # deletes wholesale — so a one-root run re-persists after every reclamation. # Harmless, and exactly the pre-D5 behaviour a one-root run asks for. persist_exp: - - validation_psf-*.fits + - psf_validation + - psf_model # Rolling exposure-store reclamation (D5). When true, the COMPUTE DAG grows one # `clean_exposure` job per exposure. It fires once every campaign tile that reads diff --git a/workflow/scripts/merge_star_cat.py b/workflow/scripts/merge_star_cat.py index 25c6e4c22..002d156ff 100644 --- a/workflow/scripts/merge_star_cat.py +++ b/workflow/scripts/merge_star_cat.py @@ -97,6 +97,7 @@ class as ``[fileobj, member_name]`` pairs — ONE AT A TIME, lazily, through # Same directory; the rule invokes this file by path, so it is sys.path[0]. import build_index +import persist_exp # The output name is not ours to choose: sp_validation hardcodes it # (`star_cat_path = f"{data_dir}/full_starcat-0000000.fits"`), and @@ -104,11 +105,13 @@ class as ``[fileobj, member_name]`` pairs — ONE AT A TIME, lazily, through # given. Kept here as the name this script promises to produce. OUT_NAME = "full_starcat-0000000.fits" -# The keep-list pattern whose members this merge consumes. The rule refuses to -# exist unless `persist_exp:` contains a pattern matching this shape (the -# Snakefile does that check at parse time), so by the time we get here the -# members are expected to be present. -MEMBER_PATTERN = "validation_psf-*.fits" +# The members this merge consumes, named as the keep list names them and +# resolved through the same catalogue persist_exp packs by — so the glob has one +# definition and adding a product cannot leave the two disagreeing. The rule +# refuses to exist unless `persist_exp:` keeps something of this shape (the +# Snakefile checks at parse time), so the members are expected here. +MEMBER_PRODUCT = "psf_validation" +MEMBER_PATTERN = persist_exp.resolve(MEMBER_PRODUCT) def merge_class(psf_model: str): @@ -233,7 +236,8 @@ def main() -> None: # existence check and produce meaningless rho statistics. sys.exit(f"merge_star_cat: no member matched {args.pattern!r} in any " f"of {len(manifest_paths)} exp_persist manifest(s) for this " - f"campaign — is '{args.pattern}' in the persist_exp keep list?") + f"campaign — is '{MEMBER_PRODUCT}' in the persist_exp keep " + f"list?") if empty: log.info(f"{len(empty)} exposure(s) persisted no {args.pattern}: " f"{', '.join(sorted(empty)[:5])}" diff --git a/workflow/scripts/persist_exp.py b/workflow/scripts/persist_exp.py index 40e8c9354..fef439b7b 100644 --- a/workflow/scripts/persist_exp.py +++ b/workflow/scripts/persist_exp.py @@ -86,40 +86,175 @@ # different rule. RUN_NAME = "run_sp_exp_SxSePsfPi" +# --- the product catalogue (CosmoStat/shapepipe#844) ------------------------ +# THE SINGLE SOURCE OF TRUTH for what an exposure can keep. `persist_exp:` in +# config.yaml names PRODUCTS, not globs: `psf_model`, not `*.psf`. The glob is +# an implementation detail of the module that writes the file, and a keep list +# written in globs is a keep list nobody can read — the argument that produced +# #844 and the 2026-09-08 call's request to keep the PSF model, which had to be +# spelled `*.psf` to be said at all. +# +# Each entry is (glob, per-exposure size, what keeping it buys). Sizes are for +# 40 CCDs, measured on smk-m2 (127 exposures, 64 tiles); "?" means not yet +# measured. `persist_exp.py --list-products` renders this table, and +# config.yaml's block is that rendering rather than a second copy of it. +# +# ORDER IS THE ORDER OF THE CHAIN — sextractor, setools, psfex, psfex_interp — +# so the table reads as the pipeline runs. +PRODUCTS = { + "star_selection": ( + "star_selection-*.fits", 24_500_000, + "setools' PRE-SPLIT selection. The only file that answers which stars " + "the selection cuts rejected and why; the split samples have already " + "lost the rejects."), + "star_train": ( + "star_split_ratio_80-*.fits", 19_900_000, + "the 80% TRAINING sample, the stars PSFEx actually fitted. Rows " + "duplicate star_selection."), + "star_test": ( + "star_split_ratio_20-*.fits", 7_100_000, + "the 20% VALIDATION sample — the positions psf_validation's rows " + "correspond to. Rows duplicate star_selection."), + "star_stats": ( + "star_stat-*.txt", None, + "setools' per-CCD STAT block: star counts, stars/deg^2, FWHM mode and " + "cuts. The selection's summary without its catalogue."), + "psf_model": ( + "*.psf", 2_800_000, + "the PSFEx model itself. Keeping it means the PSF can be " + "re-interpolated at ANY position later without rebuilding the exposure " + "chain — the single most capability-adding entry here."), + "psfex_cat": ( + "psfex_cat-*.cat", None, + "PSFEx's own output catalogue (FITS_LDAC): the per-star FLAGS_PSF and " + "CHI2_PSF, i.e. WHICH stars outlier rejection clipped. Not recoverable " + "from anything else — the .psf header keeps only the LOADED/ACCEPTED " + "counts."), + "psf_validation": ( + "validation_psf-*.fits", 2_000_000, + "the psfex_interp validation catalogue, one per CCD: the input to the " + "rho/tau statistics, and to the star_cat_merge rule that stacks them " + "into the campaign's full_starcat."), +} + +# PSFEx residual/check images and its XML diagnostics are deliberately absent: +# the committed default.psfex sets CHECKIMAGE_TYPE NONE and WRITE_XML N, so +# nothing is emitted to match. They are a config change first, a catalogue +# entry second. + +# A raw glob is still accepted, as an escape hatch for a file the catalogue does +# not name yet. The test is syntactic and deliberately cheap: a product name is +# a bare identifier, so anything carrying a glob metacharacter or a dot is a +# glob. That makes `*.psf`, `star_stat-*.txt` and `default.psfex` globs, and +# `psf_model` a name, with no ambiguity a user could stumble into. +_GLOBBY = set("*?[]. ") + + +def is_glob(entry: str) -> bool: + """True when this keep-list entry is a raw glob rather than a product name.""" + return any(ch in _GLOBBY for ch in entry) + + +def resolve(entry: str) -> str: + """The file-name glob for one keep-list entry, name or raw glob.""" + if is_glob(entry): + return entry + try: + return PRODUCTS[entry][0] + except KeyError: + raise KeyError( + f"unknown persist_exp product {entry!r}; the products are " + f"{', '.join(PRODUCTS)} (or write a raw glob such as '*.psf')" + ) from None + + +def product_of(entry: str) -> str: + """The NAME to record for an entry — the entry itself for a raw glob.""" + return entry + + +def render_products() -> str: + """The catalogue as a table, for --list-products and for config.yaml.""" + width = max(len(n) for n in PRODUCTS) + lines = [f"{'product'.ljust(width)} {'glob'.ljust(26)} size/exposure", + f"{'-' * width} {'-' * 26} -------------"] + for name, (glob, size, why) in PRODUCTS.items(): + size_s = "unmeasured" if size is None else f"{size / 1e6:.1f} MB" + lines.append(f"{name.ljust(width)} {glob.ljust(26)} {size_s}") + for i, chunk in enumerate(_wrap(why, 66)): + lines.append(f"{' ' * width} {chunk}") + return "\n".join(lines) + + +def _wrap(text: str, width: int) -> list: + out, line = [], "" + for word in text.split(): + if line and len(line) + 1 + len(word) > width: + out.append(line) + line = word + else: + line = f"{line} {word}".strip() + if line: + out.append(line) + return out + def collect(exp_dir: Path, patterns: list) -> tuple: - """Matched files per pattern, in a stable order, plus the empty patterns.""" + """Matched files per ENTRY, in a stable order, plus the entries that matched + nothing. Entries are product names or raw globs; resolve() takes either.""" root = exp_dir / "output" / RUN_NAME found, empty = {}, [] - for pat in patterns: + for entry in patterns: + pat = resolve(entry) # One glob per module output dir, recursive beneath it (see the module # docstring on setools' subdirectories). sorted() over the union keeps # the manifest byte-stable across filesystem readdir order. hits = sorted({p for mod in sorted(root.glob("*/output")) for p in mod.rglob(pat) if p.is_file()}) if hits: - found[pat] = hits + found[entry] = hits else: - empty.append(pat) + empty.append(entry) return found, empty def main() -> None: p = argparse.ArgumentParser(description=__doc__) - p.add_argument("--exp-dir", required=True, type=Path, + p.add_argument("--exp-dir", type=Path, help="the exposure's scratch store") - p.add_argument("--exp", required=True) - p.add_argument("--dest", required=True, type=Path, + p.add_argument("--exp") + p.add_argument("--dest", type=Path, help="/exp///psf; the tar is " "/.tar") - p.add_argument("--manifest", required=True, type=Path) + p.add_argument("--manifest", type=Path) p.add_argument("--pattern", action="append", default=[], - help="repeatable; a plain file-name glob") + help="repeatable; a product name (see --list-products) or a " + "raw file-name glob") + p.add_argument("--list-products", action="store_true", + help="print the product catalogue and exit") args = p.parse_args() + # --list-products is a QUERY, not a run: it answers "what can I keep?" and + # needs no exposure, so the run arguments are optional at the parser and + # required here instead. + if args.list_products: + print(render_products()) + return + missing = [f"--{n.replace('_', '-')}" for n in + ("exp_dir", "exp", "dest", "manifest") + if getattr(args, n) is None] + if missing: + p.error(f"the following arguments are required: {', '.join(missing)}") + if not args.pattern: sys.exit("persist_exp: no --pattern given (config persist_exp is empty)") + for entry in args.pattern: # loud, and before any work + try: + resolve(entry) + except KeyError as exc: + sys.exit(f"persist_exp: {exc.args[0]}") + found, empty = collect(args.exp_dir, args.pattern) if not found: sys.exit(f"persist_exp: {args.exp}: no file matched any of " @@ -145,7 +280,8 @@ def main() -> None: f"named {src.name} ({seen[src.name][0]} and {src}); tar " f"members are flat, so this would silently overwrite") seen[src.name] = (src, pat) - files.append({"name": src.name, "pattern": pat, + files.append({"name": src.name, "product": pat, + "pattern": resolve(pat), "src": str(src), "bytes": src.stat().st_size}) files.sort(key=lambda f: f["name"]) @@ -176,7 +312,8 @@ def anonymous(ti: tarfile.TarInfo) -> tarfile.TarInfo: "stage": "exp_persist", "level": "exp", "unit": args.exp, "status": "complete", "tar": str(tar_path), - "patterns": list(args.pattern), + "products": list(args.pattern), + "patterns": [resolve(e) for e in args.pattern], # The warning the docstring argues for: named patterns that matched # nothing. Present as a key even when empty, so a reader never has to # wonder whether an old manifest predates the field. From 3b8c7d59ac9de8a258007ea2ce1cb121cd78500f Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 19:32:30 -0400 Subject: [PATCH 12/85] perf(merge_starcat): accumulate arrays, not python lists of floats MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit All three merge classes built every output column by extending a python list with one value per star: `x += list(data["X"])`. Four bytes of float32 payload became a 32-byte python object plus an 8-byte pointer in an overallocating list, measured end to end at ~10x the input bytes — which put a full-survey full_starcat (~20k exposures x 40 CCDs) at ~400 GB of RAM and out of reach of any node. One array per input catalogue per column, concatenated once at the end. Same values, same order, same dtypes — np.array() over a list of numpy scalars and np.concatenate() over the arrays they came from agree on both. The stacking helper empties the list it is handed, which is half the saving: concatenate holds the chunks and the result at once, so releasing column by column peaks at one campaign plus one column rather than two campaigns. MEASURED on the same two fixture points, 20 and 80 exposures of 40 CCDs x 400 stars: input members peak RSS, before after 32.3 MB 383 MB 238 MB 129.0 MB 1313 MB 740 MB 10.1x -> 5.5x, over a ~62 MB interpreter floor. The rule's mem_mb factor follows. A 16 GB job now merges ~2200 exposures rather than ~800. THE REMAINING 5.5x IS THE OUTPUT SIDE: file_io writes every float column as FITS 1D, so float32 inputs become a float64 table astropy then buffers. That is a change to the output FORMAT, which is what sp_validation reads, and a different decision from this one. BYTE-IDENTICAL OUTPUT, both ways in. The workflow's tar path and the module runner's plain [path] path produce the same file as before the change, md5 f7caa1cf… on the fixture — the runner path checked by calling MergeStarCatPSFEX directly with [[path]] entries as merge_starcat_runner builds them. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- .../merge_starcat_package/merge_starcat.py | 247 +++++++++--------- workflow/Snakefile | 42 +-- 2 files changed, 153 insertions(+), 136 deletions(-) diff --git a/src/shapepipe/modules/merge_starcat_package/merge_starcat.py b/src/shapepipe/modules/merge_starcat_package/merge_starcat.py index 2b1bf8c34..3463ccf43 100644 --- a/src/shapepipe/modules/merge_starcat_package/merge_starcat.py +++ b/src/shapepipe/modules/merge_starcat_package/merge_starcat.py @@ -17,6 +17,36 @@ from shapepipe.pipeline import file_io +def _stack(chunks, dtype=None): + """Concatenate one column's per-catalogue arrays into a single array. + + THE COLUMN ACCUMULATORS ARE LISTS OF ARRAYS, ONE PER INPUT CATALOGUE, and + not lists of values, because these classes are the last step of a whole + campaign. ``x += list(data["X"])`` turns 4 bytes of float32 payload into a + 32-byte python object plus an 8-byte pointer in a list that overallocates — + measured at ~10x the input bytes end to end, which put a full-survey merge + (~20k exposures x 40 CCDs) at ~400 GB of RAM and made it unrunnable on any + node. One array per catalogue plus one concatenate at the end holds ~1x, and + produces the identical output: np.array() over a list of numpy scalars and + np.concatenate() over the arrays they came from agree on dtype and on order. + + IT EMPTIES THE LIST IT IS GIVEN, and that is not a side effect to tidy away + later — it is half the saving. np.concatenate holds the chunks and the + result at once, so a caller that stacks sixteen columns while all sixteen + chunk lists are still alive peaks at twice the campaign. Released column by + column, the peak is one campaign plus one column. Callers stack once, at the + end, and do not touch the accumulators afterwards. + + An empty input list is a merge over no catalogues, which the callers guard + against; it returns an empty array so the output column still exists. + """ + if not chunks: + return np.array([], dtype=dtype or np.float64) + out = np.concatenate(chunks) + del chunks[:] + return out + + class MergeStarCatMCCD(object): """Merge Star Catalogue MCCD. @@ -320,66 +350,37 @@ def process(self): model_var.append(model_var_val) model_var_size.append(model_var_val.size) + # ONE ARRAY PER CATALOGUE PER COLUMN (see _stack): the per-value + # python lists this replaces cost ~10x the input bytes. # positions - x += list( - starcat_j[self._hdu_table].data["GLOB_POSITION_IMG_LIST"][:, 0] - ) - y += list( - starcat_j[self._hdu_table].data["GLOB_POSITION_IMG_LIST"][:, 1] - ) + pos = starcat_j[self._hdu_table].data["GLOB_POSITION_IMG_LIST"] + x.append(np.asarray(pos[:, 0])) + y.append(np.asarray(pos[:, 1])) # RA and DEC positions try: - ra += list(starcat_j[self._hdu_table].data["RA_LIST"][:]) - dec += list(starcat_j[self._hdu_table].data["DEC_LIST"][:]) + ra.append(np.asarray(starcat_j[self._hdu_table].data["RA_LIST"][:])) + dec.append(np.asarray(starcat_j[self._hdu_table].data["DEC_LIST"][:])) except Exception: - ra += list( - np.zeros( - starcat_j[self._hdu_table] - .data["GLOB_POSITION_IMG_LIST"][:, 0] - .shape, - dtype=int, - ) - ) - dec += list( - np.zeros( - starcat_j[self._hdu_table] - .data["GLOB_POSITION_IMG_LIST"][:, 0] - .shape, - dtype=int, - ) - ) + ra.append(np.zeros(pos[:, 0].shape, dtype=int)) + dec.append(np.zeros(pos[:, 0].shape, dtype=int)) # shapes (convert sigmas to T = 2 sigma^2) - g1_psf += list( - starcat_j[self._hdu_table].data["PSF_MOM_LIST"][:, 0] - ) - g2_psf += list( - starcat_j[self._hdu_table].data["PSF_MOM_LIST"][:, 1] - ) - size_psf += list( - cs_size.sigma_to_T( - starcat_j[self._hdu_table].data["PSF_MOM_LIST"][:, 2] - ) - ) - g1 += list(starcat_j[self._hdu_table].data["STAR_MOM_LIST"][:, 0]) - g2 += list(starcat_j[self._hdu_table].data["STAR_MOM_LIST"][:, 1]) - size += list( - cs_size.sigma_to_T( - starcat_j[self._hdu_table].data["STAR_MOM_LIST"][:, 2] - ) - ) + psf_mom = starcat_j[self._hdu_table].data["PSF_MOM_LIST"] + star_mom = starcat_j[self._hdu_table].data["STAR_MOM_LIST"] + g1_psf.append(np.asarray(psf_mom[:, 0])) + g2_psf.append(np.asarray(psf_mom[:, 1])) + size_psf.append(np.asarray(cs_size.sigma_to_T(psf_mom[:, 2]))) + g1.append(np.asarray(star_mom[:, 0])) + g2.append(np.asarray(star_mom[:, 1])) + size.append(np.asarray(cs_size.sigma_to_T(star_mom[:, 2]))) # flags - flag_psf += list( - starcat_j[self._hdu_table].data["PSF_MOM_LIST"][:, 3] - ) - flag_star += list( - starcat_j[self._hdu_table].data["STAR_MOM_LIST"][:, 3] - ) + flag_psf.append(np.asarray(psf_mom[:, 3])) + flag_star.append(np.asarray(star_mom[:, 3])) # ccd id list - ccd_nb += list(starcat_j[self._hdu_table].data["CCD_ID_LIST"]) + ccd_nb.append(np.asarray(starcat_j[self._hdu_table].data["CCD_ID_LIST"])) starcat_j.close() @@ -452,15 +453,21 @@ def process(self): ) # Mask and transform to numpy arrays - flagmask = np.abs(np.array(flag_star) - 1) * np.abs( - np.array(flag_psf) - 1 - ) - psf_e1 = np.array(g1_psf)[flagmask.astype(bool)] - psf_e2 = np.array(g2_psf)[flagmask.astype(bool)] - psf_r2 = np.array(size_psf)[flagmask.astype(bool)] - star_e1 = np.array(g1)[flagmask.astype(bool)] - star_e2 = np.array(g2)[flagmask.astype(bool)] - star_r2 = np.array(size)[flagmask.astype(bool)] + # Concatenate once, here: everything below already wanted arrays and + # was calling np.array() on python lists to get them (see _stack). + x, y, ra, dec = _stack(x), _stack(y), _stack(ra), _stack(dec) + g1_psf, g2_psf, size_psf = _stack(g1_psf), _stack(g2_psf), _stack(size_psf) + g1, g2, size = _stack(g1), _stack(g2), _stack(size) + flag_psf, flag_star = _stack(flag_psf), _stack(flag_star) + ccd_nb = _stack(ccd_nb) + + flagmask = np.abs(flag_star - 1) * np.abs(flag_psf - 1) + psf_e1 = g1_psf[flagmask.astype(bool)] + psf_e2 = g2_psf[flagmask.astype(bool)] + psf_r2 = size_psf[flagmask.astype(bool)] + star_e1 = g1[flagmask.astype(bool)] + star_e2 = g2[flagmask.astype(bool)] + star_r2 = size[flagmask.astype(bool)] rmse, mean, std_dev = MSC.stats_calculator(star_e1, psf_e1) self._w_log.info( @@ -596,45 +603,49 @@ def process(self): data_j = starcat_j[self._hdu_table].data + # ONE ARRAY PER CATALOGUE PER COLUMN, concatenated once at the end + # (see _stack): the per-value python lists this replaces cost ~10x + # the input bytes and put a full-survey merge out of reach. # positions - x += list(data_j["X"]) - y += list(data_j["Y"]) - ra += list(data_j["RA"]) - dec += list(data_j["DEC"]) + x.append(np.asarray(data_j["X"])) + y.append(np.asarray(data_j["Y"])) + ra.append(np.asarray(data_j["RA"])) + dec.append(np.asarray(data_j["DEC"])) # shapes (size column already holds T = 2 sigma^2) - g1_psf += list(data_j["HSM_G1_PSF"]) - g2_psf += list(data_j["HSM_G2_PSF"]) - size_psf += list(data_j["HSM_T_PSF"]) - g1 += list(data_j["HSM_G1_STAR"]) - g2 += list(data_j["HSM_G2_STAR"]) - size += list(data_j["HSM_T_STAR"]) + g1_psf.append(np.asarray(data_j["HSM_G1_PSF"])) + g2_psf.append(np.asarray(data_j["HSM_G2_PSF"])) + size_psf.append(np.asarray(data_j["HSM_T_PSF"])) + g1.append(np.asarray(data_j["HSM_G1_STAR"])) + g2.append(np.asarray(data_j["HSM_G2_STAR"])) + size.append(np.asarray(data_j["HSM_T_STAR"])) # flags - flag_psf += list(data_j["HSM_FLAG_PSF"]) - flag_star += list(data_j["HSM_FLAG_STAR"]) + flag_psf.append(np.asarray(data_j["HSM_FLAG_PSF"])) + flag_star.append(np.asarray(data_j["HSM_FLAG_STAR"])) # misc # MKDEBUG: The following columns do not exist (yet) # for psf converted (pix2wcs) files. try: - mag += list(data_j["MAG"]) + mag.append(np.asarray(data_j["MAG"])) except: - mag += list(np.zeros_like(data_j["X"])) + mag.append(np.zeros_like(data_j["X"])) try: - snr += list(data_j["SNR"]) + snr.append(np.asarray(data_j["SNR"])) except: - snr += list(np.zeros_like(data_j["X"])) + snr.append(np.zeros_like(data_j["X"])) try: - psfex_acc += list(data_j["ACCEPTED"]) + psfex_acc.append(np.asarray(data_j["ACCEPTED"])) except: - psfex_acc += list(np.zeros_like(data_j["X"])) + psfex_acc.append(np.zeros_like(data_j["X"])) - # CCD number - ccd_nb += [re.split(r"\-([0-9]*)\-([0-9]+)\.", label)[-2]] * len( - data_j["RA"] - ) + # CCD number: this catalogue's one value over its own rows, as an + # array rather than a python list holding the same string N times. + ccd_nb.append(np.full( + len(data_j["RA"]), + re.split(r"\-([0-9]*)\-([0-9]+)\.", label)[-2])) # Prepare output FITS catalogue # MKDEBUG: SEx_cat=True -> False @@ -647,22 +658,22 @@ def process(self): # Collect columns (size stored as T = 2 sigma^2) data = { - "X": x, - "Y": y, - "RA": ra, - "DEC": dec, - "HSM_G1_PSF": g1_psf, - "HSM_G2_PSF": g2_psf, - "HSM_T_PSF": size_psf, - "HSM_G1_STAR": g1, - "HSM_G2_STAR": g2, - "HSM_T_STAR": size, - "HSM_FLAG_PSF": flag_psf, - "HSM_FLAG_STAR": flag_star, - "MAG": mag, - "SNR": snr, - "ACCEPTED": psfex_acc, - "CCD_NB": ccd_nb, + "X": _stack(x), + "Y": _stack(y), + "RA": _stack(ra), + "DEC": _stack(dec), + "HSM_G1_PSF": _stack(g1_psf), + "HSM_G2_PSF": _stack(g2_psf), + "HSM_T_PSF": _stack(size_psf), + "HSM_G1_STAR": _stack(g1), + "HSM_G2_STAR": _stack(g2), + "HSM_T_STAR": _stack(size), + "HSM_FLAG_PSF": _stack(flag_psf), + "HSM_FLAG_STAR": _stack(flag_star), + "MAG": _stack(mag), + "SNR": _stack(snr), + "ACCEPTED": _stack(psfex_acc), + "CCD_NB": _stack(ccd_nb, dtype="U1"), } # Write file @@ -813,29 +824,29 @@ def process(self): data_j = starcat_j[self._hdu_table].data # positions - x += list(data_j["XWIN_IMAGE"]) - y += list(data_j["YWIN_IMAGE"]) - ra += list(data_j["XWIN_WORLD"]) - dec += list(data_j["YWIN_WORLD"]) + x.append(np.asarray(data_j["XWIN_IMAGE"])) + y.append(np.asarray(data_j["YWIN_IMAGE"])) + ra.append(np.asarray(data_j["XWIN_WORLD"])) + dec.append(np.asarray(data_j["YWIN_WORLD"])) m11, m20, m02 = self.get_moments(data_j) eps1, eps2 = self.get_ellipticity(m11, m20, m02, "epsilon") chi1, chi2 = self.get_ellipticity(m11, m20, m02, "chi") - size += list(data_j["FLUX_RADIUS"]) + size.append(np.asarray(data_j["FLUX_RADIUS"])) # flags - flags += list(data_j["FLAGS_WIN"]) - flags_ext += list(data_j["IMAFLAGS_ISO"]) + flags.append(np.asarray(data_j["FLAGS_WIN"])) + flags_ext.append(np.asarray(data_j["IMAFLAGS_ISO"])) # misc - mag += list(data_j["MAG_WIN"]) - snr += list(data_j["SNR_WIN"]) + mag.append(np.asarray(data_j["MAG_WIN"])) + snr.append(np.asarray(data_j["SNR_WIN"])) # CCD number - ccd_nb += [re.split(r"\-([0-9]*)\-([0-9]+)\.", label)[-2]] * len( - data_j["XWIN_IMAGE"] - ) + ccd_nb.append(np.full( + len(data_j["XWIN_IMAGE"]), + re.split(r"\-([0-9]*)\-([0-9]+)\.", label)[-2])) # Prepare output FITS catalogue output = file_io.FITSCatalogue( @@ -847,20 +858,20 @@ def process(self): # Collect columns # convert back to sigma for consistency data = { - "X": x, - "Y": y, - "RA": ra, - "DEC": dec, + "X": _stack(x), + "Y": _stack(y), + "RA": _stack(ra), + "DEC": _stack(dec), "EPS1": eps1, "EPS2": eps2, "CHI1": chi1, "CHI2": chi2, - "SIZE": size, - "FLAGS": flags, - "FLAGS_EXT": flags_ext, - "MAG": mag, - "SNR": snr, - "CCD_NB": ccd_nb, + "SIZE": _stack(size), + "FLAGS": _stack(flags), + "FLAGS_EXT": _stack(flags_ext), + "MAG": _stack(mag), + "SNR": _stack(snr), + "CCD_NB": _stack(ccd_nb, dtype="U1"), } # Write file diff --git a/workflow/Snakefile b/workflow/Snakefile index 8f871df3b..47a8f7282 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -731,25 +731,31 @@ def star_cat_exposures(): # against synthetic tars for the star side and against smk-g6's real # catalogues for the tile side. Peak RSS is getrusage(RUSAGE_CHILDREN). # -# STAR SIDE, and it is the alarming one. Two points, 20 and 80 exposures of 40 -# CCDs x 400 stars (1.6 MB of members per exposure, against the 2.0 MB measured -# on smk-m2): +# STAR SIDE. Two points, 20 and 80 exposures of 40 CCDs x 400 stars (1.6 MB of +# members per exposure, against the 2.0 MB measured on smk-m2): # -# input members peak RSS -# 32.3 MB 383 MB -# 129.0 MB 1313 MB +# input members peak RSS, python lists peak RSS, arrays +# 32.3 MB 383 MB 238 MB +# 129.0 MB 1313 MB 740 MB # -# a slope of 10.1x the input bytes and an intercept of ~73 MB (interpreter, -# astropy, shapepipe). Tenfold, because MergeStarCatPSFEX accumulates every -# column into PYTHON LISTS of python floats before building the output arrays — -# 8 bytes of payload becomes a 32-byte object plus an 8-byte pointer. THE -# CONSEQUENCE IS A CEILING, and it should be said plainly: at ~2 MB per -# exposure, a 16 GB job merges roughly 800 exposures, and DR6's ~20k exposures -# would want ~400 GB. A full-survey full_starcat needs the accumulation changed -# to preallocated arrays or a two-pass count — a change to MergeStarCatPSFEX, -# not to this rule, and not in this PR. The formula below is honest about the -# slope so the job asks for what it will use and fails at submission rather -# than at 90% of the way through a campaign-length merge. +# a slope of 10.1x the input bytes before, 5.5x now, over a ~62 MB interpreter +# floor. The tenfold was MergeStarCat*'s accumulation of every column into +# PYTHON LISTS of python floats — 4 bytes of float32 payload becoming a 32-byte +# object plus an 8-byte pointer — and this PR replaced it with one array per +# catalogue and one concatenate at the end, output byte-identical. +# +# THE REMAINING 5.5x IS THE OUTPUT SIDE, and it is not a leak: file_io writes +# every float column as FITS 1D, so a float32 input becomes a float64 table +# that astropy then buffers to write. Halving it means changing the OUTPUT +# format, which is what sp_validation reads — a different decision from this +# one, and not ours to take here. +# +# THE CEILING MOVED BUT DID NOT GO. At ~2 MB of members per exposure a 16 GB +# job now merges ~2200 exposures rather than ~800, and DR6's ~20k would want +# ~280 GB rather than ~400 GB. A full-survey full_starcat still needs the +# output side addressed; the formula below is honest about the slope so the job +# asks for what it will use and fails at submission rather than most of the way +# through. # # TILE SIDE, and it is the reassuring one. Two points against real smk-g6 # catalogues, 2 tiles (73.9 MB in, largest 39.6 MB) and 6 tiles (235.5 MB in, @@ -757,7 +763,7 @@ def star_cat_exposures(): # the merge holds one catalogue at a time — so it is sized on the LARGEST tile, # not the total, at ~3x it plus the interpreter. STAR_MEM_BASE_MB = 500 # interpreter + astropy + shapepipe, rounded up -STAR_MEM_FACTOR = 12 # x input bytes; 10.1 measured, rounded up +STAR_MEM_FACTOR = 7 # x input bytes; 5.5 measured, rounded up FINAL_MEM_BASE_MB = 800 FINAL_MEM_FACTOR = 4 # x the LARGEST tile; ~3 measured # What one unit costs when its product is not on disk yet to be stat()ed — a From 0ad640350a390b2def72ce6f7e390e0ae3b6610b Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 19:34:47 -0400 Subject: [PATCH 13/85] feat(orchestration): the star catalogue's inputs are not a user choice MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit `persist_exp:` was doing two jobs. It decided what a campaign keeps for later — a retention question, and the user's — and it also decided whether the campaign's star catalogue could be built at all, because star_cat_merge existed only when the keep list happened to name something validation_psf-shaped. That made the survey's PSF diagnostics an opt-in, and a typo away from silently absent. exp_persist now packs psf_validation for every exposure whatever the config says. It is the merged catalogue's PROVENANCE — a full_starcat with no per-exposure inputs beside it cannot be audited, re-cut or recomputed after a purge — and it is what keeps APPENDING TILES CHEAP, since a tile added next month brings exposures whose catalogues must join the existing stack and the alternative is rebuilding their chains from VOS. ~2 MB per exposure: ~40 GB and ~40k inodes at DR6 scale against a ~1 M-inode group quota, which is the price of being able to say where the number came from. `persist_exp:` is therefore purely additive retention, defaulting to psf_model, and an EMPTY list is now a coherent instruction rather than a switch that turns persistence off: the tar holds the merge's inputs and nothing else. The keep-list gate on star_cat_merge and its parse-time warning are gone with it, as is clean_exposure's conditional edge on exp_persist — there is no configuration left under which that rule has nothing to wait for. No transient/cleanup knob: these files are kept, not staged. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- workflow/README.md | 58 ++++++++++----------------------- workflow/Snakefile | 49 +++++++++------------------- workflow/config.yaml | 56 ++++++++++++++++--------------- workflow/rules/exposure.smk | 13 ++++---- workflow/scripts/persist_exp.py | 39 +++++++++++++++++----- 5 files changed, 99 insertions(+), 116 deletions(-) diff --git a/workflow/README.md b/workflow/README.md index 1ad740b48..832a894f1 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -259,17 +259,24 @@ profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; kee hours of PSF fitting per exposure. A pattern that matches nothing is a recorded warning (setools rejects sparse CCDs); matching nothing at all is a failure. A `localrule`, by the same arithmetic as `clean_exposure`. -- **The keep list names products, not globs.** `persist_exp:` entries are names - from a catalogue in `workflow/scripts/persist_exp.py`, which is the single - source of truth for what each one means and what keeping it buys +- **The star catalogue's inputs are always kept; `persist_exp:` is what you + keep on top.** `exp_persist` packs `psf_validation` — the psfex_interp + validation catalogue, one per CCD — for every exposure whatever the config + says, because `star_cat_merge` stacks exactly those into the campaign's + `full_starcat`. They are that catalogue's provenance, and they are what keeps + appending a tile next month cheap rather than a rebuild from VOS. About 2 MB + per exposure: ~40 GB and ~40k inodes at DR6 scale, against a ~1 M-inode group + quota. `persist_exp:` is purely additive, and an empty list is legal — the tar + then holds the merge's inputs and nothing else. +- **The keep list names products, not globs.** Entries are names from a + catalogue in `workflow/scripts/persist_exp.py`, which is the single source of + truth for what each one means and what keeping it buys ([#844](https://github.com/CosmoStat/shapepipe/issues/844)); `config.yaml`'s - block is that catalogue rendered, and - `persist_exp.py --list-products` prints it. Sizes are per exposure, 40 CCDs, - measured on smk-m2. + block is that catalogue rendered, and `persist_exp.py --list-products` prints + it. Sizes are per exposure, 40 CCDs, measured on smk-m2. | product | glob | per exposure | what it buys | |---|---|---|---| - | `psf_validation` | `validation_psf-*.fits` | 2.0 MB | the rho/tau statistics input, and `star_cat_merge`'s | | `psf_model` | `*.psf` | 2.8 MB | re-interpolate the PSF anywhere later, no rebuild | | `psfex_cat` | `psfex_cat-*.cat` | unmeasured | which stars PSFEx's outlier rejection clipped | | `star_selection` | `star_selection-*.fits` | 24.5 MB | which stars the selection cuts rejected, and why | @@ -277,42 +284,11 @@ profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; kee | `star_test` | `star_split_ratio_20-*.fits` | 7.1 MB | the 20% sample `psf_validation` corresponds to | | `star_stats` | `star_stat-*.txt` | unmeasured | setools' per-CCD counts, density and FWHM cuts | - The default is `psf_validation` + `psf_model`. A raw glob is still accepted as - an escape hatch — anything with a glob metacharacter or a dot is read as one — + The default is `psf_model`. `psf_validation` is in the catalogue too but needs + no naming; naming it anyway is harmless. A raw glob is still accepted as an + escape hatch — anything with a glob metacharacter or a dot is read as one — and an unknown *name* is a parse-time error listing the valid ones. The list is exposure-side only; tile-side retention is #844 follow-up. -- **The campaign ends in two merged catalogues, and the workflow now makes - both.** Everything above is per unit; the two products downstream analysis - actually opens are per *campaign*, and until these rules existed each was a - manual pass after the run. - `star_cat_merge` stacks every exposure's every CCD's `validation_psf-*.fits` - into one `/full_starcat-0000000.fits` — the rho/tau statistics - input, at the path sp_validation hardcodes. It reads the members straight out - of the per-exposure tars (`tarfile` + `BytesIO`; unpacking ~800k files to - merge them would defeat the tar's whole purpose) and stacks them with - `MergeStarCatPSFEX`, the same class the old `merge_starcat_runner` called, so - the column list has exactly one definition. Its input is the same - `exp_persist` manifest set `rule all` already requests, so it pulls nothing - new into the DAG, and it exists only when `persist_exp:` keeps a - `validation_psf-*.fits`-shaped file — otherwise no job, and a warning at parse - time rather than a failure on a node. - `final_cat_merge` collects every ready tile's `final_cat-.fits` into - `/final_cat_.hdf5`: one dataset per tile under a group - named for the campaign, the `final_cat.param` columns, an `n_tiles` attribute. - That schema is what sp_validation's reader opens, so it is fixed; the column - extraction reuses `scripts/python/create_final_cat.py` while the file is - written here, because that script's own discovery walks a directory layout - this workflow does not have. `campaign:` in `config.yaml` names the group and - defaults to the persistent root's basename. - Both rebuild from the whole persistent root rather than appending, so the - output is a function of its input set: byte-stable on a no-op rerun - (tmp-then-`cmp`-then-`mv`), and rebuilt when a tile or exposure is appended - (the input list's fingerprint rides on `params`). Neither is a `localrule` — - one job over ~20k units is real work — and neither puts its input paths in its - shell, which is not fastidiousness: ~20k paths is an order of magnitude over - Linux's 128 KiB `MAX_ARG_STRLEN` for a single argv entry, so each script - rediscovers the set under `products_dir` while the fingerprint travels on - `params`. - **A dead tile can be told to stop pinning exposures.** An exposure is cleanable only once every consuming tile has its vignets, so one permanently-failed tile holds its ~80 exposures for the life of the diff --git a/workflow/Snakefile b/workflow/Snakefile index 47a8f7282..0473c7ac6 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -31,7 +31,6 @@ the manifest says "this stage succeeded", the log says "here is what happened" (the contract is argued in completeness.py's docstring). """ -import fnmatch import functools import hashlib import json @@ -470,6 +469,12 @@ def clean_targets(): # The keep list is config, not a rule input, and it is READ HERE so that exactly # one place converts it into the form the rule carries. An empty list is a # deliberate "keep nothing" and produces no jobs at all. +# OPTIONAL RETENTION, and only that. What star_cat_merge needs — every CCD's +# psf_validation — is packed by exp_persist whatever this list says +# (persist_exp.py's ALWAYS argues why: provenance for the merged catalogue, and +# a cheap tile append later). So an EMPTY list is a coherent instruction and not +# a switch that turns persistence off: the tar then holds the star catalogue's +# inputs and nothing else, and exp_persist still runs for every exposure. PERSIST_EXP = list(config.get("persist_exp") or []) # The keep list names PRODUCTS (`psf_model`), not globs (`*.psf`); the @@ -485,7 +490,6 @@ for _entry in PERSIST_EXP: _persist.resolve(_entry) except KeyError as _exc: raise WorkflowError(f"config persist_exp: {_exc.args[0]}") -PERSIST_GLOBS = [_persist.resolve(e) for e in PERSIST_EXP] def persist_targets(): @@ -525,8 +529,6 @@ def persist_manifests(): merge job's own re-parse under the slurm executor, which genuinely needs it. So this half carries no guard and the memo keeps either parse to one walk. """ - if not PERSIST_EXP: - return [] exps = {e for t in TILES_READY for e in tile_exposures(t)} return sorted(prod_exp_manifest(e, "exp_persist") for e in exps if not exp_store_reclaimed(e)) @@ -616,28 +618,12 @@ def unit_fingerprint(units): return f"{len(units)}:{hashlib.md5(joined.encode()).hexdigest()[:12]}" -# The tar member the star merge consumes, and the test is "would this member be -# kept" rather than "is `psf_validation` in the list". Both spellings must count: -# the product name, and a raw glob that happens to cover it (`validation_psf*`, -# `*.fits`, a bare `*`) — all things a user might reasonably write. So the keep -# list is RESOLVED to globs first and the member matched against those. -_STAR_CAT_PRODUCT = "psf_validation" -_STAR_CAT_MEMBER = "validation_psf-2605805-12.fits" -STAR_CAT_MERGE = any(fnmatch.fnmatch(_STAR_CAT_MEMBER, g) for g in PERSIST_GLOBS) - -# LOUD AT PARSE TIME, once, and only where it can be acted on. A keep list -# without the validation catalogues is a legitimate configuration (persist the -# PSF models alone, say) — it is not an error, so it must not become a job that -# fails on a node an hour later. It is worth SAYING, because the omission is -# silent otherwise and the missing product only surfaces when a rho-statistics -# run cannot find its input. -if PERSIST_EXP and not STAR_CAT_MERGE and workflow.is_main_process \ - and PHASE == "compute": - logger.warning( - f"star_cat_merge: no job — persist_exp {PERSIST_EXP} keeps no " - f"'{_STAR_CAT_MEMBER}'-shaped file, so there is nothing to stack into " - f"{PRODUCTS_DIR}/full_starcat-0000000.fits (the rho/tau statistics " - f"input). Add '{_STAR_CAT_PRODUCT}' to persist_exp: to get it.") +# NO GATE ON THE KEEP LIST. star_cat_merge used to exist only when +# `persist_exp:` named something validation_psf-shaped, which made the +# campaign's star catalogue an opt-in and a typo away from silently absent. +# exp_persist now packs psf_validation unconditionally, so the merge is +# requested whenever the campaign has a persisted exposure at all, and +# star_cat_targets() below is the only condition left. def full_starcat(): @@ -691,8 +677,6 @@ def star_cat_inputs(): manifest and is in no set at all. Nothing short of rebuilding its chain from VOS recovers it; the merge reports how many exposures it found. """ - if not PERSIST_EXP: - return [] live, reclaimed = [], [] for exp in sorted({e for t in TILES_READY for e in tile_exposures(t)}): if not exp_store_reclaimed(exp): @@ -714,8 +698,6 @@ def star_cat_exposures(): is what the trigger is for. It is also what merge_star_cat.py derives on the job side, so the two agree on the set AND on how it is named. """ - if not PERSIST_EXP: - return [] return sorted(e for e in {e for t in TILES_READY for e in tile_exposures(t)} if Path(prod_exp_manifest(e, "exp_persist")).exists() or not exp_store_reclaimed(e)) @@ -802,15 +784,14 @@ def final_cat_max_bytes(): def star_cat_targets(): """`full_starcat` when there is anything to stack into it, else nothing. - Three ways to get nothing, and all three are states rather than errors: the - keep list holds no validation catalogue (warned about above), `persist_exp:` - is empty at all, or every exposure in scope is already tombstoned — a + One way to get nothing, and it is a state rather than an error: every + exposure in scope is already tombstoned — a campaign resumed after reclamation, whose exposures were cleaned by a workflow that predates exp_persist and therefore left neither tar nor manifest to read. A rule with an empty input list would still be a JOB, and it would write an empty star catalogue over a good one. """ - if not STAR_CAT_MERGE or not workflow.is_main_process: + if not workflow.is_main_process: return [] return [full_starcat()] if star_cat_inputs() else [] diff --git a/workflow/config.yaml b/workflow/config.yaml index 03c282113..1b19784b5 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -81,11 +81,27 @@ outputs: # would otherwise have to rebuild from tile headers. index_db: /project/def-mjhudson/cdaley/sp-products/smk-g6/index/run_index.sqlite -# Per-exposure PSF products to carry onto the persistent root before the scratch -# store goes (`exp_persist`, exposure.smk). A list of PRODUCT NAMES — not globs. -# The catalogue below is the rendering of workflow/scripts/persist_exp.py's -# PRODUCTS table, which is the single source of truth for what each name means -# and what keeping it buys (CosmoStat/shapepipe#844); print it any time with +# OPTIONAL per-exposure retention: what to carry onto the persistent root ON TOP +# OF the star catalogue's own inputs, before the scratch store goes +# (`exp_persist`, exposure.smk). +# +# WHAT IS ALWAYS KEPT, AND IS NOT A CHOICE HERE: psf_validation, the psfex_interp +# validation catalogue, one per CCD. `star_cat_merge` stacks every one of them +# into the campaign's /full_starcat-0000000.fits, so they are that +# catalogue's PROVENANCE — a merged star catalogue with no per-exposure inputs +# beside it cannot be audited, re-cut, or recomputed after a purge — and they are +# what keeps APPENDING TILES CHEAP, since a tile added next month brings +# exposures whose catalogues must join the existing stack. ~2 MB per exposure: +# ~40 GB and ~40k inodes at DR6 scale, against a ~1 M-inode group quota. That is +# the price of being able to say where the number came from, and it is paid. +# +# SO THIS LIST IS PURELY ADDITIVE, and an empty one is a coherent instruction: +# the tar then holds the star catalogue's inputs and nothing else. +# +# Entries are PRODUCT NAMES — not globs. The catalogue below is the rendering of +# workflow/scripts/persist_exp.py's PRODUCTS table, which is the single source of +# truth for what each name means and what keeping it buys +# (CosmoStat/shapepipe#844); print it any time with # # workflow/bin/sp container exec python workflow/scripts/persist_exp.py --list-products # @@ -121,16 +137,13 @@ outputs: # A RAW GLOB IS STILL ACCEPTED, as an escape hatch for a file the catalogue does # not name yet: anything carrying a glob metacharacter or a dot is taken as a # glob rather than a name (`*.psf` is a glob, `psf_model` is the name for it). -# An unknown NAME is a parse-time error listing the valid ones, never a silently -# empty keep. +# An unknown NAME is a parse-time error listing the valid ones. # # Matches are packed, flat, into ONE uncompressed tar per exposure: # /exp///psf/.tar, with a manifest listing the # members (and the product each came from) beside it. One tar rather than loose -# copies because inodes, not bytes, bind on /project (~1 M-file group quota; -# loose copies would be ~200 files per exposure, ~2 M at DR6 scale). FITS -# members read straight from the tar: -# fits.open(io.BytesIO(tarfile.open(t).extractfile(m).read())). +# copies because inodes, not bytes, bind on /project. FITS members read straight +# from the tar: fits.open(io.BytesIO(tarfile.open(t).extractfile(m).read())). # # WHY COPY RATHER THAN EXEMPT THESE FROM CLEANUP. Reclamation is not the threat. # run_dir is /scratch and is PURGED on a 60-day window whether or not @@ -143,32 +156,23 @@ outputs: # reruns the packing (seconds) and NOT exp_psf (four hours per exposure). That # separation is the whole reason exp_persist is a rule of its own. # -# THE DEFAULT is psf_validation + psf_model (~4.8 MB per exposure): the rho/tau -# statistics input, and the model that lets the PSF be re-interpolated at any -# position later without rebuilding the exposure chain from VOS. Add psfex_cat -# for a production run if you want to know which stars PSFEx clipped; the -# star_* products are for selection studies and cost an order of magnitude more. +# THE DEFAULT is psf_model (2.8 MB per exposure on top of psf_validation's 2.0): +# the model that lets the PSF be re-interpolated at any position later without +# rebuilding the exposure chain from VOS, which is the single most +# capability-adding thing an exposure can keep. Add psfex_cat for a production +# run if you want to know which stars PSFEx clipped; the star_* products are for +# selection studies and cost an order of magnitude more. # # THIS LIST IS EXPOSURE-SIDE ONLY. Tile-side retention is not configurable: the # only tile product that persists today is final_cat, written by tile_make_cat # straight to products_dir. A tile keep list is #844 follow-up. # -# THIS LIST GATES `star_cat_merge`. That campaign-level rule stacks every -# exposure's every CCD's psf_validation into ONE -# /full_starcat-0000000.fits, reading the members straight out of -# the tars. A keep list without psf_validation is a legitimate configuration and -# produces NO merge job and a warning at parse time — not a failure on a node an -# hour later. Note the corollary: an exposure already reclaimed by a workflow -# that predates exp_persist left no tar, so it contributes nothing and cannot be -# recovered short of rebuilding its chain from VOS. -# # NOTE ON products_dir DEFAULTING TO run_dir (a fixture or smoke test): the tar # then lands beside the store on the same filesystem and buys nothing, and the # manifest sits in the exposure's own manifests/ dir, which clean_exposure # deletes wholesale — so a one-root run re-persists after every reclamation. # Harmless, and exactly the pre-D5 behaviour a one-root run asks for. persist_exp: - - psf_validation - psf_model # Rolling exposure-store reclamation (D5). When true, the COMPUTE DAG grows one diff --git a/workflow/rules/exposure.smk b/workflow/rules/exposure.smk index 789e9da1d..0cbafb448 100644 --- a/workflow/rules/exposure.smk +++ b/workflow/rules/exposure.smk @@ -149,6 +149,9 @@ rule exp_persist: # name collision, both of which it reports on stderr and neither of which # has a per-CCD verdict worth a completeness record. params: + # Only the OPTIONAL retention list travels: psf_validation is packed + # by persist_exp.py whatever this says. It still rides on params, so + # adding a product re-packs (seconds) rather than re-fitting the PSF. patterns = " ".join(f"--pattern '{p}'" for p in PERSIST_EXP), exp_dir = lambda wc: exp_dir(wc.exp), dest = lambda wc: f"{prod_exp_dir(wc.exp)}/psf", @@ -201,12 +204,10 @@ rule clean_exposure: # The keepers must be off /scratch before the store goes. Unlike the # consumer edges above, this edge does not depend on scope: it is the # same exposure's own rule, so it drags nothing into the DAG that this - # exposure's chain did not already put there. It is conditional only on - # there being a keep list at all — with `persist_exp:` empty, "keep - # nothing" is a coherent instruction and must not become a dependency on - # a rule that would fail for having nothing to copy. - lambda wc: ([prod_exp_manifest(wc.exp, "exp_persist")] - if PERSIST_EXP else []) + # exposure's chain did not already put there. It is UNCONDITIONAL now: + # exp_persist always packs the star catalogue's inputs, so there is no + # keep list under which this rule has nothing to wait for. + lambda wc: [prod_exp_manifest(wc.exp, "exp_persist")] output: tombstone = f"{EXP_DIR}/cleaned.json" params: diff --git a/workflow/scripts/persist_exp.py b/workflow/scripts/persist_exp.py index fef439b7b..4e2dff349 100644 --- a/workflow/scripts/persist_exp.py +++ b/workflow/scripts/persist_exp.py @@ -29,6 +29,11 @@ glob. Patterns are therefore plain FILE names and the layout is ours to know, not the config author's. +THE KEEP LIST IS WHAT THE CAMPAIGN KEEPS ON TOP OF THE MERGE'S INPUTS. +``psf_validation`` is packed unconditionally (see ALWAYS below); ``persist_exp:`` +is purely optional retention, and an EMPTY one is a coherent instruction — the +tar then holds the star catalogue's inputs and nothing else. + ZERO MATCHES FOR ONE PATTERN IS A WARNING, NOT A FAILURE. setools rejects sparse CCDs (~0.2% attrition, tolerated by exp_psf's own count floor), so per-CCD counts are not fixed, and a pattern naming an optional diagnostic may legitimately @@ -101,6 +106,19 @@ # # ORDER IS THE ORDER OF THE CHAIN — sextractor, setools, psfex, psfex_interp — # so the table reads as the pipeline runs. +# THE STAR CATALOGUE'S INPUTS ARE NOT A USER CHOICE. star_cat_merge stacks +# every CCD's psf_validation into the campaign's full_starcat, so exp_persist +# ALWAYS packs it, whatever `persist_exp:` says. Two reasons, and neither is +# about taste. It is the merged catalogue's PROVENANCE: a full_starcat with no +# per-exposure inputs beside it cannot be audited, re-cut or recomputed after a +# purge. And it is what keeps APPENDING TILES CHEAP: a tile added next month +# brings exposures whose validation catalogues must join the existing stack, and +# if the earlier ones are gone the merge either shrinks or rebuilds their chains +# from VOS. ~2 MB per exposure, so ~40 GB and ~40k inodes at DR6 scale, against +# a group quota of ~1 M inodes — the cost of being able to say where the number +# came from. +ALWAYS = "psf_validation" + PRODUCTS = { "star_selection": ( "star_selection-*.fits", 24_500_000, @@ -228,8 +246,9 @@ def main() -> None: "/.tar") p.add_argument("--manifest", type=Path) p.add_argument("--pattern", action="append", default=[], - help="repeatable; a product name (see --list-products) or a " - "raw file-name glob") + help=f"repeatable; a product name (see --list-products) or " + f"a raw file-name glob. {ALWAYS} is packed whether or " + f"not it is named — star_cat_merge needs it") p.add_argument("--list-products", action="store_true", help="print the product catalogue and exit") args = p.parse_args() @@ -246,19 +265,21 @@ def main() -> None: if missing: p.error(f"the following arguments are required: {', '.join(missing)}") - if not args.pattern: - sys.exit("persist_exp: no --pattern given (config persist_exp is empty)") + # The merge's input first and always, then whatever the campaign chose to + # keep on top of it (see ALWAYS). Deduped, so naming it explicitly in + # persist_exp: is harmless rather than a repeated pattern. + entries = [ALWAYS] + [e for e in args.pattern if e != ALWAYS] - for entry in args.pattern: # loud, and before any work + for entry in entries: # loud, and before any work try: resolve(entry) except KeyError as exc: sys.exit(f"persist_exp: {exc.args[0]}") - found, empty = collect(args.exp_dir, args.pattern) + found, empty = collect(args.exp_dir, entries) if not found: sys.exit(f"persist_exp: {args.exp}: no file matched any of " - f"{args.pattern} under {args.exp_dir}/output/{RUN_NAME}") + f"{entries} under {args.exp_dir}/output/{RUN_NAME}") args.dest.mkdir(parents=True, exist_ok=True) tar_path = args.dest / f"{args.exp}.tar" @@ -312,8 +333,8 @@ def anonymous(ti: tarfile.TarInfo) -> tarfile.TarInfo: "stage": "exp_persist", "level": "exp", "unit": args.exp, "status": "complete", "tar": str(tar_path), - "products": list(args.pattern), - "patterns": [resolve(e) for e in args.pattern], + "products": entries, + "patterns": [resolve(e) for e in entries], # The warning the docstring argues for: named patterns that matched # nothing. Present as a key even when empty, so a reader never has to # wonder whether an old manifest predates the field. From 90dfb00afb2fa42be559b61c2517faa933f8f135 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 19:43:10 -0400 Subject: [PATCH 14/85] fix(cfis): two stale columns in final_cat.param, and no mask column at all MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit final_cat_merge on smk-g6's real catalogues failed on three of the 67 columns the parameter file asks for. Two of them the file should not have been asking for. IMAFLAGS_ISO is DROPPED. The tile-side SExtractor runs with FLAG_IMAGE = False and DOT_PARAM_FILE = default_noimaflags.param (config_tile_Sx.ini), so the column is never written into a tile catalogue and asking for it could only fail. Instrument flags reach the pipeline on the EXPOSURE side, where exp_split delivers the flag image and SExtractor reads it. NGMIX_MOM_FAIL is RENAMED to NGMIX_MCAL_TYPES_FAIL, which is what f0fca23e called it in June and what the catalogues carry. NGMIX_NEIGHBOUR_FLAG STAYS. It was added to make_cat in fa6e0016 on 2026-07-12 and it is the blend flag the systematics tests need. smk-g6's catalogues do not have it — checked on the files — so final_cat_merge still fails there, and that failure is correct: the campaign's catalogues are missing a column the analysis wants, which is a fact about the data and not about this file. Its launch snapshot is gone (only .snakemake survives under smk-g6-state), so the run's HEAD cannot be read back; what remains is that its catalogues carry NGMIX_MCAL_TYPES_FAIL (June) but not NGMIX_NEIGHBOUR_FLAG (July), consistent with a snapshot taken between the two. NO MASK COLUMN REPLACES IMAFLAGS_ISO, AND THE FILE NOW SAYS WHY. The intended replacement is make_cat's per-band MASK_, queried from the healsparse maps named by MASK_EXT_PATHS — and the workflow sets none: config_tile_Mc.ini has no such entry, save_mask_ext_data is never called, no MASK_ column exists in any catalogue this workflow has produced, and smk-g6's carry none. Naming one here would fail every merge on every campaign. The merged catalogue therefore carries no mask information today; that is a CONFIG gap, and closing it is setting MASK_EXT_PATHS first and adding the column names second. No healsparse map is staged under /project/def-mjhudson yet. With NGMIX_NEIGHBOUR_FLAG set aside, the merge runs clean over all 64 of smk-g6's real catalogues: 2.50 GB read in 20 s at 151 MB peak RSS, producing a 0.94 GB hdf5 of 65 columns. That also confirms the tile-side sizing — the rule asks for 1002 MB and 51 minutes. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- workflow/config/cfis/final_cat.param | 26 +++++++++++++++++++++++--- 1 file changed, 23 insertions(+), 3 deletions(-) diff --git a/workflow/config/cfis/final_cat.param b/workflow/config/cfis/final_cat.param index 00bcb3f73..f3fa39677 100644 --- a/workflow/config/cfis/final_cat.param +++ b/workflow/config/cfis/final_cat.param @@ -8,7 +8,26 @@ TILE_ID # flags FLAGS -IMAFLAGS_ISO +# NO IMAFLAGS_ISO, AND NO MASK COLUMN AT ALL — READ THIS BEFORE ADDING ONE. +# The tile-side SExtractor runs with FLAG_IMAGE = False and DOT_PARAM_FILE = +# default_noimaflags.param (config_tile_Sx.ini), so IMAFLAGS_ISO is never +# written into a tile catalogue and asking for it here only made the merge +# fail. Instrument flags reach the pipeline on the EXPOSURE side, where +# exp_split delivers the flag image and SExtractor reads it. +# +# Its intended replacement is make_cat's per-band MASK_ columns, queried +# from the sky-fixed healsparse maps named by MASK_EXT_PATHS. THE WORKFLOW SETS +# NO SUCH PATHS: config_tile_Mc.ini has no MASK_EXT_PATHS entry, so +# save_mask_ext_data is never called, no MASK_ column exists in any tile +# catalogue this workflow has produced, and smk-g6's carry none (checked). +# Naming one here would fail every merge on every campaign. +# +# So the merged catalogue carries NO mask information today, and that is a +# CONFIG gap and not a gap in this file: turning it on is setting +# MASK_EXT_PATHS in config_tile_Mc.ini (`band:path` pairs, the same grammar as +# the commented MASK_PATHS in config_exp_psfex.ini) and adding the matching +# MASK_ names here, in that order. No healsparse map is staged under +# /project/def-mjhudson yet. NGMIX_MCAL_FLAGS # PSF ellipticity (original image PSF) @@ -113,5 +132,6 @@ NGMIX_T_PSF_ORIG_NOSHEAR # PSF size measured on reconvolved image # NGMIX_T_PSF_RECONV_NOSHEAR -# ngmix moment failure flag -NGMIX_MOM_FAIL +# ngmix metacalibration type failure flag (renamed from NGMIX_MOM_FAIL in +# f0fca23e; catalogues written before that commit carry the old name) +NGMIX_MCAL_TYPES_FAIL From a434c9aa5bed174761084cae8062cabed64fda8d Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 19:43:10 -0400 Subject: [PATCH 15/85] perf(merge_starcat): two passes, so nothing is held twice MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The array accumulation landed a commit ago took the star merge from ~10x the input bytes to ~5.5x. What was left was the accumulation itself: one array per input catalogue, then a concatenate that has to hold its inputs and its result at the same time. MergeStarCatPSFEX now makes two passes. The first reads only the FITS HEADER of every input — NAXIS2, the row count — and touches no data block; the second allocates each output column once, at its exact final length, and fills it slice by slice. There are no chunks and no concatenate, so peak memory is one output plus one input catalogue. The workflow's tar reader hands over the archive's own file object rather than a BytesIO of the whole member, so the counting pass costs a header rather than a member. Both it and a plain list of paths are iterable twice, which the two passes require; a one-shot iterable would fill nothing on the second pass, so the merge checks that the passes agree on the row count rather than writing a catalogue padded with uninitialised memory. MEASURED on the same two fixture points, 20 and 80 exposures of 40 CCDs x 400 stars: input members python lists arrays+concat two passes 32.3 MB 383 MB 238 MB 221 MB 129.0 MB 1313 MB 740 MB 661 MB slope 10.1x 5.5x 4.8x The rule's mem_mb factor follows. A 16 GB job now merges ~1300 exposures. THE REMAINING 4.8x IS THE OUTPUT SIDE: file_io writes every float column as FITS 1D, so float32 inputs become a float64 table astropy then buffers — 141 MB of table for 78 MB of payload at the 80-exposure point. What stands between here and a full-survey full_starcat is that format, not the merge. MergeStarCatMCCD and MergeStarCatSetools keep the array accumulation. Their process() computes campaign-wide statistics over the same columns, so a two-pass rewrite there is a larger change with no consumer today — psfex is what every campaign runs. BYTE-IDENTICAL OUTPUT, both ways in: the workflow's tar path and the module runner's plain [path] path both give md5 f7caa1cf… on the fixture, unchanged through both rewrites. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- .../merge_starcat_package/merge_starcat.py | 169 ++++++++++-------- workflow/Snakefile | 41 ++--- workflow/scripts/merge_star_cat.py | 12 +- 3 files changed, 127 insertions(+), 95 deletions(-) diff --git a/src/shapepipe/modules/merge_starcat_package/merge_starcat.py b/src/shapepipe/modules/merge_starcat_package/merge_starcat.py index 3463ccf43..97e25e0c2 100644 --- a/src/shapepipe/modules/merge_starcat_package/merge_starcat.py +++ b/src/shapepipe/modules/merge_starcat_package/merge_starcat.py @@ -571,81 +571,119 @@ def __init__( self._hdu_table = hdu_table self._input_cat_type = input_cat_type + # The columns this class writes, and where each comes from. Kept as data + # rather than as sixteen repeated lines, because a two-pass merge would + # otherwise state every column three times: to size it, to allocate it and + # to fill it. + _COLUMNS = ( + ("X", "X"), ("Y", "Y"), ("RA", "RA"), ("DEC", "DEC"), + ("HSM_G1_PSF", "HSM_G1_PSF"), ("HSM_G2_PSF", "HSM_G2_PSF"), + ("HSM_T_PSF", "HSM_T_PSF"), ("HSM_G1_STAR", "HSM_G1_STAR"), + ("HSM_G2_STAR", "HSM_G2_STAR"), ("HSM_T_STAR", "HSM_T_STAR"), + ("HSM_FLAG_PSF", "HSM_FLAG_PSF"), ("HSM_FLAG_STAR", "HSM_FLAG_STAR"), + ) + # Present in psfex_interp output, absent from pix2wcs-converted files + # (MKDEBUG); zero-filled when missing rather than failing the merge. + _OPTIONAL = (("MAG", "MAG"), ("SNR", "SNR"), ("ACCEPTED", "ACCEPTED")) + + def _ccd_nb(self, label): + """The CCD number this catalogue's rows carry, parsed from its name.""" + return re.split(r"\-([0-9]*)\-([0-9]+)\.", label)[-2] + def process(self): """Process. Process merging. + TWO PASSES, AND NEITHER HOLDS THE CAMPAIGN TWICE. The first reads only + the FITS HEADER of every input — NAXIS2, the row count — and never + touches a data block; the second allocates the output columns once, at + their exact final length, and fills them slice by slice. Peak memory is + therefore ONE output plus ONE input catalogue. + + What this replaces, in two steps, is instructive about the cost of the + obvious code. Accumulating each column into a python LIST OF VALUES — + ``x += list(data["X"])`` — turned 4 bytes of float32 payload into a + 32-byte object plus an 8-byte pointer, measured at ~10x the input bytes + end to end and putting a full-survey merge (~20k exposures x 40 CCDs) at + ~400 GB. Accumulating one ARRAY PER CATALOGUE and concatenating once + brought that to ~5.5x. This pass structure removes what was left of the + accumulation: there are no chunks, and no concatenate that must hold its + inputs and its result at the same time. + + ``self._input_file_list`` MUST BE ITERABLE TWICE. A list is; so is the + workflow's tar reader, whose ``__iter__`` opens the archives afresh. + A one-shot generator is not, and would silently merge nothing on the + second pass — hence the explicit length check below. """ - x, y, ra, dec = [], [], [], [] - g1_psf, g2_psf, size_psf = [], [], [] - g1, g2, size = [], [], [] - flag_psf, flag_star = [], [] - mag, snr, psfex_acc = [], [], [] - ccd_nb = [] - self._w_log.info( f"Merging {len(self._input_file_list)} star catalogues" ) + # --- pass 1: row counts and dtypes, from headers alone -------------- + counts, labels, dtypes, n_total = [], [], None, 0 for name in self._input_file_list: - # The source to read and the NAME to parse the CCD number out of. - # Identical for a plain [path] entry; different only when the caller - # hands over an open file-like object plus the member name it came - # under (see the class docstring). source, label = name[0], name[-1] try: - starcat_j = fits.open(source, memmap=False, ignore_missing_simple=True) - except OSError as e: + with fits.open(source, memmap=False, + ignore_missing_simple=True) as starcat_j: + hdu = starcat_j[self._hdu_table] + n_rows = hdu.header["NAXIS2"] + if dtypes is None: + # ColDefs.dtype describes the table without reading it. + dtypes = hdu.columns.dtype + except OSError: print(f"Error while opening file '{label}'") #raise continue - + counts.append(n_rows) + labels.append(label) + n_total += n_rows + + if dtypes is None: + raise ValueError("merge_starcat: no readable input catalogue") + + # --- allocate once, at the exact final length ----------------------- + present = set(dtypes.names) + data = {out: np.empty(n_total, dtype=dtypes[col]) + for out, col in self._COLUMNS} + for out, col in self._OPTIONAL: + data[out] = np.empty( + n_total, dtype=dtypes[col] if col in present else dtypes["X"]) + # CCD_NB is one string per catalogue, repeated over its rows; its width + # is the widest CCD number in the campaign, which pass 1 already knows. + width = max((len(self._ccd_nb(lb)) for lb in labels), default=1) + data["CCD_NB"] = np.empty(n_total, dtype=f"U{width}") + + # --- pass 2: fill --------------------------------------------------- + at = 0 + for name in self._input_file_list: + source, label = name[0], name[-1] + try: + starcat_j = fits.open(source, memmap=False, + ignore_missing_simple=True) + except OSError: + continue data_j = starcat_j[self._hdu_table].data + n_rows = len(data_j) + sl = slice(at, at + n_rows) - # ONE ARRAY PER CATALOGUE PER COLUMN, concatenated once at the end - # (see _stack): the per-value python lists this replaces cost ~10x - # the input bytes and put a full-survey merge out of reach. - # positions - x.append(np.asarray(data_j["X"])) - y.append(np.asarray(data_j["Y"])) - ra.append(np.asarray(data_j["RA"])) - dec.append(np.asarray(data_j["DEC"])) - - # shapes (size column already holds T = 2 sigma^2) - g1_psf.append(np.asarray(data_j["HSM_G1_PSF"])) - g2_psf.append(np.asarray(data_j["HSM_G2_PSF"])) - size_psf.append(np.asarray(data_j["HSM_T_PSF"])) - g1.append(np.asarray(data_j["HSM_G1_STAR"])) - g2.append(np.asarray(data_j["HSM_G2_STAR"])) - size.append(np.asarray(data_j["HSM_T_STAR"])) - - # flags - flag_psf.append(np.asarray(data_j["HSM_FLAG_PSF"])) - flag_star.append(np.asarray(data_j["HSM_FLAG_STAR"])) - - # misc + for out, col in self._COLUMNS: + data[out][sl] = data_j[col] + for out, col in self._OPTIONAL: + data[out][sl] = data_j[col] if col in present else 0 + data["CCD_NB"][sl] = self._ccd_nb(label) - # MKDEBUG: The following columns do not exist (yet) - # for psf converted (pix2wcs) files. - try: - mag.append(np.asarray(data_j["MAG"])) - except: - mag.append(np.zeros_like(data_j["X"])) - try: - snr.append(np.asarray(data_j["SNR"])) - except: - snr.append(np.zeros_like(data_j["X"])) - try: - psfex_acc.append(np.asarray(data_j["ACCEPTED"])) - except: - psfex_acc.append(np.zeros_like(data_j["X"])) + at += n_rows + starcat_j.close() - # CCD number: this catalogue's one value over its own rows, as an - # array rather than a python list holding the same string N times. - ccd_nb.append(np.full( - len(data_j["RA"]), - re.split(r"\-([0-9]*)\-([0-9]+)\.", label)[-2])) + if at != n_total: + # The two passes disagreed: an input changed under us, or the list + # was a one-shot iterable. Either way the output would be padded + # with uninitialised memory, so say so rather than write it. + raise ValueError( + f"merge_starcat: pass 1 counted {n_total} rows, pass 2 filled " + f"{at} — is the input list iterable more than once?") # Prepare output FITS catalogue # MKDEBUG: SEx_cat=True -> False @@ -656,25 +694,8 @@ def process(self): SEx_catalogue=False, ) - # Collect columns (size stored as T = 2 sigma^2) - data = { - "X": _stack(x), - "Y": _stack(y), - "RA": _stack(ra), - "DEC": _stack(dec), - "HSM_G1_PSF": _stack(g1_psf), - "HSM_G2_PSF": _stack(g2_psf), - "HSM_T_PSF": _stack(size_psf), - "HSM_G1_STAR": _stack(g1), - "HSM_G2_STAR": _stack(g2), - "HSM_T_STAR": _stack(size), - "HSM_FLAG_PSF": _stack(flag_psf), - "HSM_FLAG_STAR": _stack(flag_star), - "MAG": _stack(mag), - "SNR": _stack(snr), - "ACCEPTED": _stack(psfex_acc), - "CCD_NB": _stack(ccd_nb, dtype="U1"), - } + # `data` was built by the two passes above (size stored as T = 2 + # sigma^2); every column is already an array of its final length. # Write file # MKDEBUG for psf conv (pix2WCS) files do not write as SExtractorCat; diff --git a/workflow/Snakefile b/workflow/Snakefile index 0473c7ac6..734d11add 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -714,30 +714,31 @@ def star_cat_exposures(): # catalogues for the tile side. Peak RSS is getrusage(RUSAGE_CHILDREN). # # STAR SIDE. Two points, 20 and 80 exposures of 40 CCDs x 400 stars (1.6 MB of -# members per exposure, against the 2.0 MB measured on smk-m2): +# members per exposure, against the 2.0 MB measured on smk-m2), across the two +# rewrites this PR made to the accumulation in MergeStarCatPSFEX: # -# input members peak RSS, python lists peak RSS, arrays -# 32.3 MB 383 MB 238 MB -# 129.0 MB 1313 MB 740 MB +# input members python lists arrays+concat two passes +# 32.3 MB 383 MB 238 MB 221 MB +# 129.0 MB 1313 MB 740 MB 661 MB # -# a slope of 10.1x the input bytes before, 5.5x now, over a ~62 MB interpreter -# floor. The tenfold was MergeStarCat*'s accumulation of every column into -# PYTHON LISTS of python floats — 4 bytes of float32 payload becoming a 32-byte -# object plus an 8-byte pointer — and this PR replaced it with one array per -# catalogue and one concatenate at the end, output byte-identical. +# slope 10.1x 5.5x 4.8x # -# THE REMAINING 5.5x IS THE OUTPUT SIDE, and it is not a leak: file_io writes +# The tenfold was one python float object (32 bytes) plus a list pointer (8) +# per 4 bytes of float32 payload. Arrays per catalogue removed that; the +# two-pass structure — count rows from the FITS headers, allocate once at the +# exact length, then fill — removed what remained, so nothing is held twice. +# +# THE REMAINING 4.8x IS THE OUTPUT SIDE, and it is not a leak: file_io writes # every float column as FITS 1D, so a float32 input becomes a float64 table -# that astropy then buffers to write. Halving it means changing the OUTPUT -# format, which is what sp_validation reads — a different decision from this -# one, and not ours to take here. +# that astropy then buffers to write — 141 MB of table for 78 MB of payload at +# the 80-exposure point, plus its write copy. Halving it means changing the +# OUTPUT format, which is what sp_validation reads: a different decision from +# this one, and not ours to take here. # -# THE CEILING MOVED BUT DID NOT GO. At ~2 MB of members per exposure a 16 GB -# job now merges ~2200 exposures rather than ~800, and DR6's ~20k would want -# ~280 GB rather than ~400 GB. A full-survey full_starcat still needs the -# output side addressed; the formula below is honest about the slope so the job -# asks for what it will use and fails at submission rather than most of the way -# through. +# THE CEILING MOVED AND IS NOW ELSEWHERE. At ~2 MB of members per exposure a +# 16 GB job merges ~1300 exposures rather than ~800, and DR6's ~20k would want +# ~240 GB rather than ~400 GB. What stands between here and a full-survey +# full_starcat is the float64 output, not the merge. # # TILE SIDE, and it is the reassuring one. Two points against real smk-g6 # catalogues, 2 tiles (73.9 MB in, largest 39.6 MB) and 6 tiles (235.5 MB in, @@ -745,7 +746,7 @@ def star_cat_exposures(): # the merge holds one catalogue at a time — so it is sized on the LARGEST tile, # not the total, at ~3x it plus the interpreter. STAR_MEM_BASE_MB = 500 # interpreter + astropy + shapepipe, rounded up -STAR_MEM_FACTOR = 7 # x input bytes; 5.5 measured, rounded up +STAR_MEM_FACTOR = 6 # x input bytes; 4.8 measured, rounded up FINAL_MEM_BASE_MB = 800 FINAL_MEM_FACTOR = 4 # x the LARGEST tile; ~3 measured # What one unit costs when its product is not on disk yet to be stat()ed — a diff --git a/workflow/scripts/merge_star_cat.py b/workflow/scripts/merge_star_cat.py index 002d156ff..f39364932 100644 --- a/workflow/scripts/merge_star_cat.py +++ b/workflow/scripts/merge_star_cat.py @@ -183,6 +183,10 @@ class TarMembers: ``__len__`` comes from the manifests, so the class can log the count before a single tar is opened. + + IT IS ITERABLE MORE THAN ONCE, and must be: the merge makes two passes, one + for row counts from the headers and one to fill. Each ``__iter__`` opens the + archives afresh, so the second pass sees the same members in the same order. """ def __init__(self, chosen): @@ -199,7 +203,13 @@ def __iter__(self): if member is None: sys.exit(f"merge_star_cat: {tar_path} has no member " f"{name}, which its manifest lists") - yield [io.BytesIO(member.read()), name] + # The tar's own file object, not a BytesIO of the whole + # member: it is seekable (the archive is uncompressed by + # design) and astropy reads through it, so the merge's + # first pass costs a header rather than a member. The + # object is valid only until the next member is reached, + # which is exactly how the merge consumes it. + yield [member, name] def main() -> None: From 3c2982588738718fa028f5e581638a2d74a51f2e Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 19:47:05 -0400 Subject: [PATCH 16/85] feat(orchestration): final_cat_merge reconciles instead of rebuilding MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The rule read every tile's catalogue on every run, because a DAG output must be a function of its input set and rebuilding is the simple way to guarantee that. At DR6 scale it is also ~800 GB of IO to add one 35 MB tile. It now brings the file INTO AGREEMENT with the campaign: a tile with no dataset is added, a dataset whose tile has left the campaign is deleted, a dataset whose source catalogue CHANGED is re-read, and one that agrees with its source is left alone, unread. Each dataset records its source's size and mtime as attributes, and a mismatch is what changed means — which is also what keeps the file from drifting from its inputs the way an append-only tool does. create_final_cat.py's own process() implements only the append-only half of this, skipping any tile already present whatever the file on disk now says. WHAT IS AND IS NOT A FUNCTION OF THE INPUT SET, since this is the guarantee being traded. The file's CONTENT is: the same tiles with the same catalogues give the same datasets, the same columns and the same n_tiles, whether they arrived at once or one campaign at a time. Its BYTE LAYOUT is not, because hdf5 lays a group out in the order things were added. That is the price of not re-reading the campaign. UNTOUCHED ON A NO-OP, which is stronger than the byte comparison it replaces and cheaper to establish: reconciling is PLANNED against a read-only open, and an empty plan never opens the file for writing, so its mtime cannot move. A non-empty plan is carried out on a copy which is then moved into place, so a crash mid-merge leaves the old catalogue intact. VERIFIED on a three-tile fixture: build (3 added), no-op (unchanged, mtime identical to the nanosecond), append one tile WHILE AN EXISTING TILE'S CATALOGUE IS UNREADABLE — chmod 000, which succeeds and reports 1 added, so the existing tiles were demonstrably not read — rewrite of one catalogue (1 refreshed), and dropping two tiles from the list (2 removed, datasets gone, n_tiles 1). A from-scratch build of the same set is byte-stable across reruns. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- workflow/README.md | 33 ++++++ workflow/rules/tile.smk | 20 ++-- workflow/scripts/merge_final_cat.py | 175 ++++++++++++++++++++++------ 3 files changed, 182 insertions(+), 46 deletions(-) diff --git a/workflow/README.md b/workflow/README.md index 832a894f1..ab0f72d89 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -289,6 +289,39 @@ profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; kee escape hatch — anything with a glob metacharacter or a dot is read as one — and an unknown *name* is a parse-time error listing the valid ones. The list is exposure-side only; tile-side retention is #844 follow-up. +- **The campaign ends in two merged catalogues, and the workflow makes both.** + Everything above is per unit; the two products downstream analysis actually + opens are per *campaign*, and until these rules existed each was a manual pass + after the run. + `star_cat_merge` stacks every exposure's every CCD's `psf_validation` into one + `/full_starcat-0000000.fits` — the rho/tau statistics input, at + the path sp_validation hardcodes. It reads the members straight out of the + per-exposure tars (`tarfile`; unpacking ~800k files to merge them would defeat + the tar's whole purpose) and stacks them with `MergeStarCatPSFEX`, the same + class the old `merge_starcat_runner` called, so the column list has exactly + one definition. It exists whenever the campaign has a persisted exposure. + `final_cat_merge` collects every ready tile's `final_cat-.fits` into + `/final_cat_.hdf5`: one dataset per tile under a group + named for the campaign, the `final_cat.param` columns, an `n_tiles` attribute. + That schema is what sp_validation's reader opens, so it is fixed; the column + extraction reuses `scripts/python/create_final_cat.py` while the file is + written here, because that script's own discovery walks a directory layout + this workflow does not have. `campaign:` in `config.yaml` names the group and + defaults to the persistent root's basename. + `star_cat_merge` restacks the whole campaign, so its output is a function of + its input set and byte-stable on a no-op rerun (tmp-then-`cmp`-then-`mv`). + `final_cat_merge` RECONCILES instead — adds the tiles that have no dataset, + drops datasets whose tile left the campaign, re-reads one whose catalogue + changed (each dataset records its source's size and mtime), and leaves the + rest unread — because re-reading a campaign to add one tile is ~800 GB of IO + at DR6 scale. Its *content* is still a function of the input set; its byte + layout is not, and a no-op leaves the file untouched rather than rewritten. + Both rerun when the set changes: the unit ids' fingerprint rides on `params`. + Neither is a `localrule` — one job over ~20k units is real work — and neither + puts its input paths in its shell, which is not fastidiousness: ~20k paths is + an order of magnitude over Linux's 128 KiB `MAX_ARG_STRLEN` for a single argv + entry, so each job is handed the tile list and the run index and derives the + same set from them. - **A dead tile can be told to stop pinning exposures.** An exposure is cleanable only once every consuming tile has its vignets, so one permanently-failed tile holds its ~80 exposures for the life of the diff --git a/workflow/rules/tile.smk b/workflow/rules/tile.smk index 4eb293b04..deb90138c 100644 --- a/workflow/rules/tile.smk +++ b/workflow/rules/tile.smk @@ -925,14 +925,18 @@ rule clean_tile: # tile-finished marker (see final_cat() in the Snakefile), and it is the file # this rule actually reads. # -# NOT A LOCALRULE, and here the reason is IO rather than memory: the job reads -# every tile's catalogue end to end on every run — ~32-46 MB per tile, so ~2 GB -# for a 64-tile campaign and ~800 GB at DR6's 23k tiles. It rebuilds rather than -# appends because a DAG output must be a function of its input set -# (merge_final_cat.py); incremental update by hand is what -# `create_final_cat.py -s add` remains for. Memory is one tile's catalogue at a -# time plus the hdf5 write buffer, which is why mem_mb is modest where -# star_cat_merge's is not. +# NOT A LOCALRULE, and here the reason is IO rather than memory: a first build +# reads every tile's catalogue end to end — ~32-46 MB per tile, so ~2 GB for a +# 64-tile campaign and ~800 GB at DR6's 23k tiles. It RECONCILES rather than +# rebuilds or appends: a tile with no dataset is added, a dataset whose tile +# left the campaign is deleted, a dataset whose source catalogue changed is +# re-read, and one that agrees with its source is left alone. So an append +# reads the appended tiles and nothing else, while the file still cannot drift +# from its inputs the way an append-only tool does (merge_final_cat.py argues +# what is and is not a function of the input set here). Memory is one tile's +# catalogue at a time plus the hdf5 write buffer, which is why mem_mb is modest +# where star_cat_merge's is not — and why runtime, which is sized on the whole +# campaign, is the pessimistic first-build case. rule final_cat_merge: input: lambda wc: [final_cat(t) for t in TILES_READY] diff --git a/workflow/scripts/merge_final_cat.py b/workflow/scripts/merge_final_cat.py index 58357fcd5..601c2d0b7 100644 --- a/workflow/scripts/merge_final_cat.py +++ b/workflow/scripts/merge_final_cat.py @@ -38,20 +38,41 @@ loaded by path rather than imported: it is a script, not an installed module, and the container's ``shapepipe`` install does not carry it. -IT REBUILDS THE WHOLE FILE, IT DOES NOT APPEND. ``create_final_cat.py``'s own -``process()`` skips tiles already in the file, which is right for a hand-driven -incremental update (``-s add`` / ``-s remove`` are that tool's job). A DAG rule -wants the opposite: the output must be a pure function of the input set, so that -a no-op rerun is byte-stable and a changed set is visibly a different file. -Appending would make the result depend on the order campaigns were run in, and -would silently keep a tile whose catalogue was later rebuilt. The cost is -reading every tile's catalogue on every run of the rule — real work at DR6 scale -(~20k tiles), which is why this is not a localrule. - -BYTE-STABLE ON A NO-OP RERUN: written to a tmp path, compared, moved only if it -differs (the pattern ``persist_exp.py`` and ``clean_exposure.py`` use). Tiles -are visited in sorted ID order so the file is a function of the input set alone. -An unconditional rewrite would move the output's mtime every invocation. +IT RECONCILES, IT NEITHER REBUILDS NOR BLINDLY APPENDS. The output must be a +function of the input set — that is what makes the rule's fingerprint mean +something — but reading every tile's catalogue to add one tile is ~800 GB of IO +at DR6 scale for ~35 MB of new data. So the file is brought INTO AGREEMENT with +the campaign instead: + + * a campaign tile with no dataset is read and added; + * a dataset whose tile is no longer in the campaign is deleted; + * a dataset whose source catalogue has CHANGED is re-read. Each one records + its source's size and mtime as attributes, and a mismatch is what "changed" + means. This is the only reason a finished tile is ever read twice, and it is + the reason the file cannot drift from its inputs the way an append-only + tool does; + * a dataset that agrees with its source is left alone, unread. + +An append therefore reads exactly the appended tiles. ``create_final_cat.py``'s +own ``process()`` implements the append-only half of this — it skips a tile +already in the file, whatever the file on disk now says — which is right for a +hand-driven update and wrong for a DAG output; ``-s add`` / ``-s remove`` +remain that tool's way to do this by hand. + +WHAT IS AND IS NOT A FUNCTION OF THE INPUT SET. The file's CONTENT is: the same +tiles with the same catalogues give the same datasets, the same columns and the +same n_tiles, whether they arrived at once or one campaign at a time. Its BYTE +LAYOUT is not, because hdf5 lays out a group in the order things were added. +That is the trade for not re-reading the campaign, and it is why the no-op case +below compares actions rather than bytes. +UNTOUCHED ON A NO-OP RERUN, which is stronger than byte-stable and cheaper to +establish. Reconciling is planned before anything is written: if the plan is +empty the file is not opened for writing at all, so its mtime cannot move — and +mtime is a rerun trigger, so an unconditional rewrite would make every +invocation look like a change. When the plan is NOT empty the existing file is +copied to a tmp path, changed there and moved into place, so a crash mid-merge +leaves the old catalogue intact rather than a half-written one. The copy is a +fraction of the reading it replaces. WHICH TILES — AND WHY THE JOB DERIVES THE SET RATHER THAN BEING TOLD IT. The set is the CAMPAIGN's: every tile both declared in ``tile_list`` and present in the @@ -73,8 +94,8 @@ """ import argparse -import filecmp import importlib.util +import shutil import sys from pathlib import Path @@ -132,6 +153,97 @@ def catalogues(products_dir: Path, tile_list: Path, index_db: Path) -> list: return out +class Plan: + """What reconciling this campaign into this file requires: three tile lists. + + ``add`` and ``refresh`` are both "read the catalogue and write the dataset"; + they are separate only so the log can say which happened, because a refresh + means a finished tile's catalogue moved under us and that is worth seeing. + """ + + def __init__(self, add, refresh, remove): + self.add, self.refresh, self.remove = add, refresh, remove + + def empty(self): + return not (self.add or self.refresh or self.remove) + + def describe(self): + return (f"{len(self.add)} added, {len(self.refresh)} refreshed, " + f"{len(self.remove)} removed") + + +def stamp(path: Path) -> tuple: + """The source catalogue's identity, as recorded on its dataset. + + Size and mtime, not a checksum: the file is ~35 MB and the question is + "did this change since we read it", which mtime answers for a pipeline + that writes a catalogue once. A campaign that rewrites a final_cat in + place with identical size and mtime would defeat it, and nothing does. + """ + st = path.stat() + return st.st_size, st.st_mtime_ns + + +def reconcile_plan(output: Path, group_path: str, tiles: list) -> Plan: + """Compare the file on disk with the campaign, WITHOUT writing anything. + + Opened read-only, so a no-op invocation cannot move the output's mtime. + """ + if not output.exists(): + return Plan([t for t, _ in tiles], [], []) + + want = {tile: path for tile, path in tiles} + add, refresh = [], [] + with h5py.File(output, "r") as f: + have = dict(f[group_path].items()) if group_path in f else {} + present = set(have) + for tile, path in tiles: + if tile not in present: + add.append(tile) + continue + attrs = have[tile].attrs + if (int(attrs.get("src_bytes", -1)), + int(attrs.get("src_mtime_ns", -1))) != stamp(path): + refresh.append(tile) + return Plan(add, refresh, sorted(present - set(want))) + + +def apply_plan(output: Path, group_path: str, plan: Plan, tiles: list, + cfc, params: dict) -> None: + """Carry the plan out on a COPY, then move it into place. + + The copy is what makes a crash mid-merge leave the old catalogue intact, + and it costs a fraction of the reading it replaces — an append that copies + a 1 GB file to add one 35 MB tile still beats re-reading the campaign. + """ + paths = dict(tiles) + tmp = output.with_name(output.name + ".tmp") + try: + tmp.unlink(missing_ok=True) + if output.exists(): + shutil.copy2(output, tmp) + with h5py.File(tmp, "a") as f: + group = f[group_path] if group_path in f else f.create_group(group_path) + for tile in plan.remove: + del group[tile] + for tile in plan.refresh: + del group[tile] + for tile in plan.add + plan.refresh: + path = paths[tile] + extracted, dtype = cfc.read_data(str(path), params) + data = cfc.copy_data(params["param_list"], extracted, dtype) + dset = group.create_dataset(tile, data=data, dtype=data.dtype) + # The dataset's own record of what it was read from; this is + # what makes a later invocation able to leave it alone. + dset.attrs["src_bytes"], dset.attrs["src_mtime_ns"] = stamp(path) + # The same attribute create_final_cat.py's print_list() writes, and + # what sp_validation reads to know how many tiles it is holding. + f.attrs["n_tiles"] = len(group) + tmp.replace(output) # atomic: same filesystem + finally: + tmp.unlink(missing_ok=True) + + def main() -> None: p = argparse.ArgumentParser(description=__doc__) p.add_argument("--products-dir", required=True, type=Path, @@ -164,30 +276,17 @@ def main() -> None: sys.exit(f"merge_final_cat: no tile in {args.tile_list} is indexed in " f"{args.index_db}, so there is nothing to merge") - # tmp-then-cmp-then-mv; the tmp never outlives this process. args.output.parent.mkdir(parents=True, exist_ok=True) - tmp = args.output.with_name(args.output.name + ".tmp") - try: - tmp.unlink(missing_ok=True) # h5py "a" would reopen a stale one - with h5py.File(tmp, "w") as hdf5_file: - group = hdf5_file.create_group(spval_group(args.campaign)) - for tile, path in tiles: - extracted, dtype = cfc.read_data(str(path), params) - data = cfc.copy_data(params["param_list"], extracted, dtype) - group.create_dataset(tile, data=data, dtype=data.dtype) - # The same attribute create_final_cat.py's print_list() writes, and - # what sp_validation reads to know how many tiles it is holding. - hdf5_file.attrs["n_tiles"] = len(tiles) - - if args.output.exists() and filecmp.cmp(tmp, args.output, shallow=False): - print(f"[merge_final_cat] unchanged: {args.output}") - else: - tmp.replace(args.output) # atomic: same filesystem - print(f"[merge_final_cat] {len(tiles)} tile(s), " - f"{len(param_list)} column(s) -> {args.output} " - f"(group {spval_group(args.campaign)})") - finally: - tmp.unlink(missing_ok=True) + group_path = spval_group(args.campaign) + plan = reconcile_plan(args.output, group_path, tiles) + if not plan.empty(): + apply_plan(args.output, group_path, plan, tiles, cfc, params) + print(f"[merge_final_cat] {plan.describe()} -> {args.output} " + f"({len(tiles)} tile(s), {len(param_list)} column(s), " + f"group {group_path})") + else: + print(f"[merge_final_cat] unchanged: {args.output} " + f"({len(tiles)} tile(s))") if __name__ == "__main__": From 78d2ce9c8620e6e07d377833cdd8ba0490ad581a Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 20:03:30 -0400 Subject: [PATCH 17/85] fix(orchestration): eight defects found reviewing the merge work MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit OPTIONAL COLUMNS WERE DECIDED ONCE FOR THE WHOLE MERGE. MergeStarCatPSFEX read MAG/SNR/ACCEPTED out of the first catalogue's dtype and applied that verdict to every file behind it, so a merge over a mix of ordinary and pix2wcs-converted catalogues was wrong in both directions: ordinary first raised KeyError on the first converted file, and converted first SILENTLY ZEROED the real values of every ordinary file. The dtype now comes from any file that carries the column and pass 2 asks each file for its own schema, so only the files that actually lack a column are zero-filled. Both orderings verified on fixtures. A FAILED psfex_interp COULD GET A GREEN MANIFEST, and clean_exposure takes that manifest as its go-ahead to delete the store. With the default retention list, an exposure whose interpolation failed but whose PSFEx model landed had a non-empty match set, so the pack succeeded and the stars went with the store — unrecoverable short of rebuilding the chain from VOS. psf_validation is not optional: nothing matching it now fails the job, while the store is still on disk. Retention products that match nothing stay warnings. SHRINKING THE KEEP LIST DELETED PRODUCTS FROM /project. The list rides on `params`, so editing it reruns the pack — which rewrote the tar without what had been dropped, on the backed-up filesystem, with the scratch store it came from usually already reclaimed. RETENTION IS NOW ADDITIVE: an existing tar is a FLOOR, its members carried into the new one whatever the current list says, so a config change can only ever add. Removing a product is a deliberate act on products_dir, not a config edit. Verified: pack with [psf_model], rerun with an empty list, the .psf is still there under its own product name and the tar is byte-identical. THE COLUMN SET REACHED NO RERUN TRIGGER. final_cat_merge's reconcile keyed staleness on each source catalogue's size and mtime, and the column set is not a source catalogue: final_cat.param arrives through `params`, and the hash covered workflow/scripts/ only, not scripts/python/create_final_cat.py. So this PR's own edit to final_cat.param would have left every dataset in an existing hdf5 written to the old schema with nothing to notice. The file now carries a digest of the resolved column list on its root and refreshes every tile when it moves, and MERGE_FINAL_HASH covers all three files the rule's behaviour comes from. Verified: build, edit the parameter file, rerun -> 3 refreshed. STAR_CAT_MERGE'S MEMORY WAS SIZED ON THE TAR, which holds whatever the campaign retains, while the merge reads the psf_validation members alone. Measured on a fixture with a 3 MB PSF model kept: the tar is 92x the members it will read, and the default retention is 2.4x. It also jumped discontinuously as exposures were packed. The manifests record the product each member came from — exactly so this is answerable without opening a tar — so the sizing sums those members. NO REQUEST WAS CAPPED. A mem_mb above the partition maximum is a job SLURM never schedules and snakemake never diagnoses: it sits PENDING while the campaign looks alive. Both merge formulas grow with the campaign, so at some size they cross it. `max_mem_mb:` (default 750000, for Nibi's 766 GB standard node) caps both, with a parse-time warning naming the rule that was capped. copy_data ORDERED ITS OUTPUT BY THE SOURCE CATALOGUE, which made the merged dtype a property of the catalogue rather than of the parameter file: the ordered dedup added to read_param_file had no effect, and two tiles written by different ShapePipe versions landed in one group with two different structured dtypes, which np.concatenate refuses. It orders by param_list now. Verified on two catalogues with reversed column orders and an extra column: one dtype, concatenate works. The fixture hdf5 md5 moves with the column order, d2882294… -> 43ff946d…. RECONCILE LEAKED SPACE. It copied the file and deleted datasets in place, and HDF5 never reclaims that, so every refresh of a tile grew the file by that tile. A plan that removes or refreshes anything now builds the tmp fresh, moving the datasets it keeps across with h5py's own group copy — a dataset-level copy that never reads a row into numpy — so the result is compact; pure-append plans still copy and append. Verified: five successive full refreshes leave the file the same size, and dropping a tile shrinks it. Also: comments referring to the deleted parse-time keep-list gate are gone; `-s add` is documented as what it is (accepted by create_final_cat.py's validator, then falling through to the ordinary walk, so not a way to add one tile by hand); and three latent issues are noted where they live rather than fixed — hdu.columns.dtype ignoring TSCAL/TZERO (no validation_psf column is scaled), MergeStarCatSetools rebinding its ellipticity accumulators so only the last file's reach the output (pre-existing, setools is not wired to any workflow path), and the .tmp a SIGKILL can orphan next to the catalogue. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- scripts/python/create_final_cat.py | 20 ++-- .../merge_starcat_package/merge_starcat.py | 41 +++++-- workflow/README.md | 6 +- workflow/Snakefile | 105 +++++++++++++++--- workflow/config.yaml | 9 ++ workflow/rules/exposure.smk | 3 +- workflow/rules/tile.smk | 3 +- workflow/scripts/merge_final_cat.py | 85 ++++++++++---- workflow/scripts/merge_star_cat.py | 12 +- workflow/scripts/persist_exp.py | 77 ++++++++++++- 10 files changed, 299 insertions(+), 62 deletions(-) diff --git a/scripts/python/create_final_cat.py b/scripts/python/create_final_cat.py index efef24620..ea8a4d723 100755 --- a/scripts/python/create_final_cat.py +++ b/scripts/python/create_final_cat.py @@ -338,13 +338,19 @@ def copy_data(param_list, extracted_data, dtype): """Copy Data. """ - # THE REQUESTED COLUMNS ONLY, in the SOURCE catalogue's order. Allocating - # with the source's full dtype and filling only the requested columns left - # every other column as uninitialised memory: meaningless values in the - # output file, and different bytes on every run of this tool over the same - # inputs. The parameter file says which columns the merged catalogue is - # for; those are the columns it gets. - columns = [col for col in (dtype.names or ()) if col in set(param_list)] + # THE REQUESTED COLUMNS ONLY, IN THE PARAMETER FILE'S ORDER. Two things + # are being fixed here and they are easy to conflate. Allocating with the + # source's full dtype and filling only the requested columns left every + # other column as uninitialised memory — meaningless values, and different + # bytes on every run over the same inputs. And ordering the result by the + # SOURCE catalogue's columns made the output dtype a property of the + # catalogue rather than of the parameter file: two tiles written by + # different ShapePipe versions, whose catalogues order or extend their + # columns differently, then landed in one merged file with two different + # structured dtypes, which np.concatenate refuses. The parameter file is + # the schema; it says which columns AND in what order. + wanted = set(dtype.names or ()) + columns = [col for col in param_list if col in wanted] subset = np.dtype([(col, dtype[col]) for col in columns]) # Initialize new data structure diff --git a/src/shapepipe/modules/merge_starcat_package/merge_starcat.py b/src/shapepipe/modules/merge_starcat_package/merge_starcat.py index 97e25e0c2..7f751006c 100644 --- a/src/shapepipe/modules/merge_starcat_package/merge_starcat.py +++ b/src/shapepipe/modules/merge_starcat_package/merge_starcat.py @@ -621,7 +621,15 @@ def process(self): ) # --- pass 1: row counts and dtypes, from headers alone -------------- - counts, labels, dtypes, n_total = [], [], None, 0 + # THE OPTIONAL COLUMNS ARE A PER-FILE QUESTION, NOT A PER-MERGE ONE. + # A pix2wcs-converted catalogue has no MAG/SNR/ACCEPTED while an + # ordinary one does, and a merge can be handed both. Deciding from the + # first file alone got it wrong in both directions: converted-first + # zero-filled the real values of every ordinary file behind it, and + # ordinary-first raised KeyError on the first converted one. So the + # dtype comes from ANY file that carries the column, and pass 2 asks + # each file for itself. + labels, dtypes, opt_dtypes, n_total = [], None, {}, 0 for name in self._input_file_list: source, label = name[0], name[-1] try: @@ -629,14 +637,21 @@ def process(self): ignore_missing_simple=True) as starcat_j: hdu = starcat_j[self._hdu_table] n_rows = hdu.header["NAXIS2"] + # ColDefs.dtype describes the table without reading it. + # NOTE: it is the RAW storage dtype and ignores TSCAL/TZERO, + # so a scaled column would be allocated narrower than the + # values .data returns. Latent, not live: no validation_psf + # column is scaled. Read the dtype off .data if one ever is. + cols = hdu.columns.dtype if dtypes is None: - # ColDefs.dtype describes the table without reading it. - dtypes = hdu.columns.dtype + dtypes = cols + for _, col in self._OPTIONAL: + if col not in opt_dtypes and col in (cols.names or ()): + opt_dtypes[col] = cols[col] except OSError: print(f"Error while opening file '{label}'") #raise continue - counts.append(n_rows) labels.append(label) n_total += n_rows @@ -644,12 +659,12 @@ def process(self): raise ValueError("merge_starcat: no readable input catalogue") # --- allocate once, at the exact final length ----------------------- - present = set(dtypes.names) data = {out: np.empty(n_total, dtype=dtypes[col]) for out, col in self._COLUMNS} for out, col in self._OPTIONAL: - data[out] = np.empty( - n_total, dtype=dtypes[col] if col in present else dtypes["X"]) + # A column no file carries still gets a column, zero-filled, in the + # positional dtype the old code used for it. + data[out] = np.empty(n_total, dtype=opt_dtypes.get(col, dtypes["X"])) # CCD_NB is one string per catalogue, repeated over its rows; its width # is the widest CCD number in the campaign, which pass 1 already knows. width = max((len(self._ccd_nb(lb)) for lb in labels), default=1) @@ -668,10 +683,13 @@ def process(self): n_rows = len(data_j) sl = slice(at, at + n_rows) + have = set(data_j.dtype.names or ()) for out, col in self._COLUMNS: data[out][sl] = data_j[col] for out, col in self._OPTIONAL: - data[out][sl] = data_j[col] if col in present else 0 + # THIS file's schema, not the merge's: zero-fill only the files + # that actually lack the column. + data[out][sl] = data_j[col] if col in have else 0 data["CCD_NB"][sl] = self._ccd_nb(label) at += n_rows @@ -850,6 +868,13 @@ def process(self): ra.append(np.asarray(data_j["XWIN_WORLD"])) dec.append(np.asarray(data_j["YWIN_WORLD"])) + # PRE-EXISTING BUG, LEFT ALONE DELIBERATELY: these four REBIND the + # accumulators initialised above rather than appending to them, so + # only the LAST input file's ellipticities reach the output while + # every other column carries the whole merge. Setools is not wired + # to any workflow path today; fixing it is its own change with its + # own verification, and doing it silently inside a memory rewrite + # would bury it. m11, m20, m02 = self.get_moments(data_j) eps1, eps2 = self.get_ellipticity(m11, m20, m02, "epsilon") chi1, chi2 = self.get_ellipticity(m11, m20, m02, "chi") diff --git a/workflow/README.md b/workflow/README.md index ab0f72d89..85cf4f465 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -285,7 +285,11 @@ profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; kee | `star_stats` | `star_stat-*.txt` | unmeasured | setools' per-CCD counts, density and FWHM cuts | The default is `psf_model`. `psf_validation` is in the catalogue too but needs - no naming; naming it anyway is harmless. A raw glob is still accepted as an + no naming; naming it anyway is harmless. **Retention is additive**: an + existing tar is a floor, so shrinking the list adds nothing and removes + nothing. Dropping a product is a deliberate act on `products_dir`, not a + config edit — otherwise editing a config would delete products from the + backed-up filesystem whose scratch originals are long gone. A raw glob is still accepted as an escape hatch — anything with a glob metacharacter or a dot is read as one — and an unknown *name* is a parse-time error listing the valid ones. The list is exposure-side only; tile-side retention is #844 follow-up. diff --git a/workflow/Snakefile b/workflow/Snakefile index 734d11add..c509b770e 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -339,7 +339,21 @@ def unit_num(unit): # gets its own hash, on the forest rule only. def script_hash(name): """The 12-hex fingerprint of one script under workflow/scripts/.""" - return hashlib.md5((SCRIPTS / name).read_bytes()).hexdigest()[:12] + return path_hash(SCRIPTS / name) + + +def path_hash(path): + """The 12-hex fingerprint of any file the rules depend on but do not own. + + A rule whose behaviour comes from more than its own script needs all of it + in one trigger. final_cat_merge is the case: what it writes is decided by + scripts/python/create_final_cat.py (the column extraction) and by + config/cfis/final_cat.param (which columns), and NEITHER is under + workflow/scripts/ nor a declared input. Without them in the hash, this PR's + own edits to both would have left every finished campaign's hdf5 untouched + and nothing would have said so. + """ + return hashlib.md5(Path(path).read_bytes()).hexdigest()[:12] SCRIPT_HASH = script_hash("completeness.py") FOREST_HASH = script_hash("build_forest.py") @@ -347,7 +361,13 @@ CLEAN_HASH = script_hash("clean_exposure.py") CLEAN_TILE_HASH = script_hash("clean_tile.py") PERSIST_HASH = script_hash("persist_exp.py") MERGE_STAR_HASH = script_hash("merge_star_cat.py") -MERGE_FINAL_HASH = script_hash("merge_final_cat.py") +# Three files, one trigger: the rule's script, the column extraction it calls, +# and the parameter file that says which columns (path_hash argues why). +MERGE_FINAL_HASH = ":".join(( + script_hash("merge_final_cat.py"), + path_hash(Path(workflow.basedir).parent / "scripts" / "python" + / "create_final_cat.py"), + path_hash(CONFIG_DIR / "final_cat.param"))) # ngmix_range.py earns a hash for a stronger reason than the others. What it # emits is not a stale RESULT but a stale BOUNDARY, and a tile's eight chunks are # a PARTITION of its object IDs: resume a tile across an edit to the split and @@ -467,8 +487,8 @@ def clean_targets(): # --- persisted exposure products (D5) -------------------------------------- # The keep list is config, not a rule input, and it is READ HERE so that exactly -# one place converts it into the form the rule carries. An empty list is a -# deliberate "keep nothing" and produces no jobs at all. +# one place converts it into the form the rule carries. +# # OPTIONAL RETENTION, and only that. What star_cat_merge needs — every CCD's # psf_validation — is packed by exp_persist whatever this list says # (persist_exp.py's ALWAYS argues why: provenance for the merged catalogue, and @@ -745,14 +765,49 @@ def star_cat_exposures(): # largest 47.7 MB): peak RSS 129 MB and 139 MB. FLAT IN THE NUMBER OF TILES — # the merge holds one catalogue at a time — so it is sized on the LARGEST tile, # not the total, at ~3x it plus the interpreter. +# THE CEILING ON ANY REQUEST, and it is not a formatting nicety: a mem_mb above +# the partition maximum is a job SLURM will never schedule and snakemake will +# never diagnose — it sits PENDING with a reason nobody reads while the campaign +# looks alive. The two merge formulas grow with the campaign, so at some size +# they WILL cross it; capping turns "silently never runs" into "runs on the +# biggest node there is, and possibly dies with a diagnosable OOM". +# +# Nibi's standard compute node is 766 GB (192 cores, 4 GB/core); 750000 leaves +# room for the OS and the slurm accounting overhead. Override with `max_mem_mb:` +# for a cluster with smaller nodes, or to reserve headroom. +MAX_MEM_MB = int(config.get("max_mem_mb", 750_000)) +_capped_warned = set() + + +def capped_mem(mb, rule): + """min(mb, MAX_MEM_MB), and say so ONCE at parse time when it bites.""" + mb = int(mb) + if mb > MAX_MEM_MB: + if rule not in _capped_warned and workflow.is_main_process: + _capped_warned.add(rule) + logger.warning( + f"{rule}: sized at {mb} MB, capped to max_mem_mb={MAX_MEM_MB} " + f"(Nibi's standard node is 766 GB). The job will run with less " + f"memory than the measurement says it wants — expect an OOM, " + f"and split the campaign or fix the merge rather than raising " + f"this number past what a node has.") + return MAX_MEM_MB + return mb + + STAR_MEM_BASE_MB = 500 # interpreter + astropy + shapepipe, rounded up STAR_MEM_FACTOR = 6 # x input bytes; 4.8 measured, rounded up FINAL_MEM_BASE_MB = 800 FINAL_MEM_FACTOR = 4 # x the LARGEST tile; ~3 measured -# What one unit costs when its product is not on disk yet to be stat()ed — a -# fresh campaign sizes its merge before anything has been packed or made. -# Both are the measured medians in config.yaml's persist_exp block and D5 notes. +# What one unit costs when its product is not on disk yet to be measured — a +# fresh campaign sizes its merge before anything has been packed or made. The +# exposure figure is psf_validation's alone (the only members the star merge +# reads), not a whole tar's; both are the measured medians in config.yaml's +# persist_exp block and the D5 notes. EXP_BYTES_DEFAULT = 2_000_000 +# The product whose members star_cat_merge stacks — named once, here and in +# merge_star_cat.py, and resolved through persist_exp.py's catalogue. +STAR_CAT_PRODUCT = _persist.ALWAYS TILE_BYTES_DEFAULT = 46_000_000 @@ -765,15 +820,35 @@ def _size(path, default): def star_cat_bytes(): - """Total member bytes the star merge will read. - - The TAR is what gets stat()ed, not the manifest: it is one stat per - exposure rather than a json parse, and it is present for exactly the - exposures whose products already exist — live-and-already-packed as well as - reclaimed. An exposure not yet packed contributes the measured default. + """Total bytes of the members the star merge will actually read. + + THE TAR'S SIZE IS THE WRONG NUMBER, and increasingly wrong as the keep list + grows: the merge reads the psf_validation members and nothing else, while + the tar also holds whatever `persist_exp:` retains. With the default + retention that is 2.4x too much, and with the star_* products on it is ~36x + — a memory request that misses by more than an order of magnitude, and one + that would jump the moment an exposure got packed, since an unpacked one + contributed the per-exposure default instead. So the MANIFEST is read and + only the psf_validation members are counted; persist_exp records the product + each member came from, exactly so this is answerable without opening a tar. + + One json parse per exposure at DAG build, and only for the parse that builds + this job. An exposure not yet packed has no manifest and contributes the + measured default, which is the psf_validation figure and not the tar's. """ - return sum(_size(prod_exp_tar(e), EXP_BYTES_DEFAULT) - for e in star_cat_exposures()) + total = 0 + for exp in star_cat_exposures(): + manifest = Path(prod_exp_manifest(exp, "exp_persist")) + if not manifest.exists(): + total += EXP_BYTES_DEFAULT + continue + try: + body = json.loads(manifest.read_text()) + total += sum(f["bytes"] for f in body["files"] + if f.get("product") == STAR_CAT_PRODUCT) + except (OSError, ValueError, KeyError): + total += EXP_BYTES_DEFAULT + return total def final_cat_max_bytes(): diff --git a/workflow/config.yaml b/workflow/config.yaml index 1b19784b5..99ae1b560 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -239,6 +239,15 @@ clean_tiles: true # it is dead, not while you are still debugging it. clean_ignore_tiles: [] +# The ceiling on any rule's mem_mb. A request above the partition maximum is a +# job SLURM never schedules and snakemake never diagnoses: it sits PENDING while +# the campaign looks alive. Nibi's standard compute node is 766 GB (192 cores at +# 4 GB/core), so 750000 leaves room for the OS and slurm's own overhead. The two +# campaign-level merges size themselves from the campaign's bytes and will cross +# this at survey scale — the cap turns "never runs" into "runs on the biggest +# node there is", with a parse-time warning saying which rule was capped. +max_mem_mb: 750000 + # ngmix within-tile chunking: static N chunks (closed ID ranges computed # per-tile, in-job, from the tile's own sexcat). ngmix_chunks: 8 diff --git a/workflow/rules/exposure.smk b/workflow/rules/exposure.smk index 0cbafb448..0b6a7a715 100644 --- a/workflow/rules/exposure.smk +++ b/workflow/rules/exposure.smk @@ -283,8 +283,9 @@ rule star_cat_merge: # measured (the Snakefile's sizing block carries both points, and the # ceiling this rule runs into at DR6 scale). Still * attempt, because a # measured slope on synthetic tars is not a guarantee about real ones. - mem_mb = lambda wc, attempt: attempt * ( + mem_mb = lambda wc, attempt: capped_mem(attempt * ( STAR_MEM_BASE_MB + STAR_MEM_FACTOR * star_cat_bytes() // 1_000_000), + "star_cat_merge"), # ~2 min per GB of members on the measurement above, doubled, over a # floor that covers the fixed cost of opening ~40 members per exposure. runtime = lambda wc, attempt: attempt * ( diff --git a/workflow/rules/tile.smk b/workflow/rules/tile.smk index deb90138c..46d84f02f 100644 --- a/workflow/rules/tile.smk +++ b/workflow/rules/tile.smk @@ -955,9 +955,10 @@ rule final_cat_merge: # Sized on the LARGEST tile, not the total: the merge holds one # catalogue at a time, and the measurement is flat in the tile count # (the Snakefile's sizing block carries both points). - mem_mb = lambda wc, attempt: attempt * ( + mem_mb = lambda wc, attempt: capped_mem(attempt * ( FINAL_MEM_BASE_MB + FINAL_MEM_FACTOR * final_cat_max_bytes() // 1_000_000), + "final_cat_merge"), # Runtime, unlike memory, is the TOTAL: every tile is read end to end. # ~1 min per 10 tiles on the measurement, triply generous, over a floor. runtime = lambda wc, attempt: attempt * (30 + len(TILES_READY) // 3) diff --git a/workflow/scripts/merge_final_cat.py b/workflow/scripts/merge_final_cat.py index 601c2d0b7..4bae00f98 100644 --- a/workflow/scripts/merge_final_cat.py +++ b/workflow/scripts/merge_final_cat.py @@ -56,8 +56,10 @@ An append therefore reads exactly the appended tiles. ``create_final_cat.py``'s own ``process()`` implements the append-only half of this — it skips a tile already in the file, whatever the file on disk now says — which is right for a -hand-driven update and wrong for a DAG output; ``-s add`` / ``-s remove`` -remain that tool's way to do this by hand. +hand-driven update and wrong for a DAG output. (Its ``-s`` single-ID mode +implements ``check`` and ``remove``; ``add`` is accepted by the argument +validator and then falls through to the ordinary walk, so it is not a way to +add one tile by hand.) WHAT IS AND IS NOT A FUNCTION OF THE INPUT SET. The file's CONTENT is: the same tiles with the same catalogues give the same datasets, the same columns and the @@ -94,6 +96,7 @@ """ import argparse +import hashlib import importlib.util import shutil import sys @@ -153,6 +156,20 @@ def catalogues(products_dir: Path, tile_list: Path, index_db: Path) -> list: return out +def schema_digest(param_list: list) -> str: + """A fingerprint of the COLUMN SET the datasets were written with. + + Recorded on the file's root and compared on every reconcile, because the + column set is the one input to this merge that nothing else can see. It is + not a source catalogue, so no dataset's size/mtime stamp moves when it + changes; it reaches the job through --param-file, which is a `params` value + and not a rule input. Without this, editing final_cat.param — which this PR + itself does — would leave every dataset in an existing hdf5 written to the + OLD schema, and nothing would ever notice. + """ + return hashlib.md5("\n".join(param_list).encode()).hexdigest()[:16] + + class Plan: """What reconciling this campaign into this file requires: three tile lists. @@ -184,7 +201,8 @@ def stamp(path: Path) -> tuple: return st.st_size, st.st_mtime_ns -def reconcile_plan(output: Path, group_path: str, tiles: list) -> Plan: +def reconcile_plan(output: Path, group_path: str, tiles: list, + digest: str) -> Plan: """Compare the file on disk with the campaign, WITHOUT writing anything. Opened read-only, so a no-op invocation cannot move the output's mtime. @@ -195,39 +213,62 @@ def reconcile_plan(output: Path, group_path: str, tiles: list) -> Plan: want = {tile: path for tile, path in tiles} add, refresh = [], [] with h5py.File(output, "r") as f: + # A changed column set invalidates every dataset at once — they were + # written to the old schema and nothing about their sources moved. + stale_schema = f.attrs.get("param_digest") != digest have = dict(f[group_path].items()) if group_path in f else {} present = set(have) for tile, path in tiles: if tile not in present: add.append(tile) - continue - attrs = have[tile].attrs - if (int(attrs.get("src_bytes", -1)), - int(attrs.get("src_mtime_ns", -1))) != stamp(path): + elif stale_schema: refresh.append(tile) + else: + attrs = have[tile].attrs + if (int(attrs.get("src_bytes", -1)), + int(attrs.get("src_mtime_ns", -1))) != stamp(path): + refresh.append(tile) return Plan(add, refresh, sorted(present - set(want))) def apply_plan(output: Path, group_path: str, plan: Plan, tiles: list, - cfc, params: dict) -> None: - """Carry the plan out on a COPY, then move it into place. - - The copy is what makes a crash mid-merge leave the old catalogue intact, - and it costs a fraction of the reading it replaces — an append that copies - a 1 GB file to add one 35 MB tile still beats re-reading the campaign. + cfc, params: dict, digest: str) -> None: + """Carry the plan out on a tmp file, then move it into place. + + TWO WAYS TO BUILD THE TMP, and which one is used is about SPACE, not speed. + HDF5 never reclaims the space a deleted dataset occupied, so a file that is + copied and then edited in place grows for the life of the campaign — every + refresh of a 15 MB tile leaks 15 MB. So: + + * a plan that only ADDS copies the existing file and appends to it. There + is nothing to reclaim, and copying beats rewriting. + * a plan that removes or refreshes anything builds the tmp FRESH, moving + the datasets it keeps across with h5py's own group copy — which is a + dataset-level copy inside the library and never reads a row into numpy — + and writing only the tiles that actually changed. The result is compact. + + Either way the tmp is moved into place at the end, so a crash mid-merge + leaves the old catalogue intact rather than a half-written one. A SIGKILL + between writing the tmp and renaming it leaves the tmp behind — one file, + next to the catalogue, overwritten by the next run; the rename itself is + atomic, which is the property that matters. """ paths = dict(tiles) + rewrite = bool(plan.remove or plan.refresh) + keep = [t for t, _ in tiles if t not in set(plan.add) | set(plan.refresh)] tmp = output.with_name(output.name + ".tmp") try: tmp.unlink(missing_ok=True) - if output.exists(): + if output.exists() and not rewrite: shutil.copy2(output, tmp) with h5py.File(tmp, "a") as f: - group = f[group_path] if group_path in f else f.create_group(group_path) - for tile in plan.remove: - del group[tile] - for tile in plan.refresh: - del group[tile] + group = (f[group_path] if group_path in f + else f.create_group(group_path)) + if rewrite and output.exists(): + with h5py.File(output, "r") as src: + for tile in keep: + src[f"{group_path}/{tile}"].copy( + src[f"{group_path}/{tile}"], group, name=tile) for tile in plan.add + plan.refresh: path = paths[tile] extracted, dtype = cfc.read_data(str(path), params) @@ -239,6 +280,7 @@ def apply_plan(output: Path, group_path: str, plan: Plan, tiles: list, # The same attribute create_final_cat.py's print_list() writes, and # what sp_validation reads to know how many tiles it is holding. f.attrs["n_tiles"] = len(group) + f.attrs["param_digest"] = digest tmp.replace(output) # atomic: same filesystem finally: tmp.unlink(missing_ok=True) @@ -278,9 +320,10 @@ def main() -> None: args.output.parent.mkdir(parents=True, exist_ok=True) group_path = spval_group(args.campaign) - plan = reconcile_plan(args.output, group_path, tiles) + digest = schema_digest(param_list) + plan = reconcile_plan(args.output, group_path, tiles, digest) if not plan.empty(): - apply_plan(args.output, group_path, plan, tiles, cfc, params) + apply_plan(args.output, group_path, plan, tiles, cfc, params, digest) print(f"[merge_final_cat] {plan.describe()} -> {args.output} " f"({len(tiles)} tile(s), {len(param_list)} column(s), " f"group {group_path})") diff --git a/workflow/scripts/merge_star_cat.py b/workflow/scripts/merge_star_cat.py index f39364932..131758f57 100644 --- a/workflow/scripts/merge_star_cat.py +++ b/workflow/scripts/merge_star_cat.py @@ -107,9 +107,10 @@ class as ``[fileobj, member_name]`` pairs — ONE AT A TIME, lazily, through # The members this merge consumes, named as the keep list names them and # resolved through the same catalogue persist_exp packs by — so the glob has one -# definition and adding a product cannot leave the two disagreeing. The rule -# refuses to exist unless `persist_exp:` keeps something of this shape (the -# Snakefile checks at parse time), so the members are expected here. +# definition and adding a product cannot leave the two disagreeing. They are +# always there to find: persist_exp packs this product for every exposure +# whatever `persist_exp:` says, and fails the pack rather than writing a +# manifest without it. MEMBER_PRODUCT = "psf_validation" MEMBER_PATTERN = persist_exp.resolve(MEMBER_PRODUCT) @@ -246,8 +247,9 @@ def main() -> None: # existence check and produce meaningless rho statistics. sys.exit(f"merge_star_cat: no member matched {args.pattern!r} in any " f"of {len(manifest_paths)} exp_persist manifest(s) for this " - f"campaign — is '{MEMBER_PRODUCT}' in the persist_exp keep " - f"list?") + f"campaign. persist_exp packs {MEMBER_PRODUCT} for every " + f"exposure, so this means the manifests are not what we think " + f"they are.") if empty: log.info(f"{len(empty)} exposure(s) persisted no {args.pattern}: " f"{', '.join(sorted(empty)[:5])}" diff --git a/workflow/scripts/persist_exp.py b/workflow/scripts/persist_exp.py index 4e2dff349..24e0c1d66 100644 --- a/workflow/scripts/persist_exp.py +++ b/workflow/scripts/persist_exp.py @@ -29,6 +29,15 @@ glob. Patterns are therefore plain FILE names and the layout is ours to know, not the config author's. +RETENTION IS ADDITIVE, AND THAT IS A SAFETY PROPERTY. The keep list rides on +the rule's ``params``, so SHRINKING it reruns this script — and a naive rerun +would rewrite the tar without the products that were dropped, deleting them +from the backed-up filesystem because someone edited a config, with the scratch +store they came from usually long gone. An existing tar is therefore a FLOOR: +its members are carried into the new one whatever the current list says, and a +config change can only ever add. Removing a product is a deliberate act on +products_dir, not a config edit. + THE KEEP LIST IS WHAT THE CAMPAIGN KEEPS ON TOP OF THE MERGE'S INPUTS. ``psf_validation`` is packed unconditionally (see ALWAYS below); ``persist_exp:`` is purely optional retention, and an EMPTY one is a coherent instruction — the @@ -277,6 +286,21 @@ def main() -> None: sys.exit(f"persist_exp: {exc.args[0]}") found, empty = collect(args.exp_dir, entries) + # THE MERGE'S INPUTS ARE NOT ALLOWED TO BE MISSING, and this is a harder + # rule than "something matched". An exposure whose psfex_interp failed but + # whose PSFEx model landed has a non-empty match set under the default + # retention list, so it used to get a green manifest — and clean_exposure + # takes that manifest as its go-ahead and deletes the store, taking the + # stars with it. There is no recovering them afterwards short of rebuilding + # the chain from VOS, so a missing psf_validation fails the job here, while + # the store is still on disk. Retention products that match nothing stay + # warnings: they are optional by construction. + if ALWAYS not in found: + sys.exit(f"persist_exp: {args.exp}: nothing matched {ALWAYS} " + f"({resolve(ALWAYS)}) under {args.exp_dir}/output/{RUN_NAME}. " + f"That is the star catalogue's input and it is not optional — " + f"refusing to write a manifest that would let clean_exposure " + f"reclaim this store.") if not found: sys.exit(f"persist_exp: {args.exp}: no file matched any of " f"{entries} under {args.exp_dir}/output/{RUN_NAME}") @@ -304,6 +328,35 @@ def main() -> None: files.append({"name": src.name, "product": pat, "pattern": resolve(pat), "src": str(src), "bytes": src.stat().st_size}) + + # --- RETENTION IS ADDITIVE: an existing tar is a FLOOR, never a draft ---- + # Shrinking `persist_exp:` used to rerun this rule (the list rides on + # params, which is the whole point of the rule) and overwrite the tar with + # a smaller one — deleting products from the BACKED-UP filesystem because + # someone edited a config. The scratch store they came from is usually gone + # by then, so nothing could put them back. Whatever is already in the tar + # therefore stays in it: a config change can only ever ADD. + # + # Removing a product is consequently not a config edit. It is a deliberate + # act on products_dir, and it should look like one. + carried, prior_products = [], {} + if tar_path.exists(): + prior = args.manifest + if prior.exists(): + try: + prior_products = {f["name"]: f.get("product", "?") + for f in json.loads(prior.read_text())["files"]} + except (OSError, ValueError, KeyError): + pass # a damaged manifest loses only labels + with tarfile.open(tar_path) as tf: + for ti in tf.getmembers(): + if ti.name in seen or not ti.isfile(): + continue # a live source supersedes it + carried.append(ti.name) + files.append({"name": ti.name, + "product": prior_products.get(ti.name, "?"), + "pattern": None, "src": None, "bytes": ti.size}) + files.sort(key=lambda f: f["name"]) def anonymous(ti: tarfile.TarInfo) -> tarfile.TarInfo: @@ -319,9 +372,24 @@ def anonymous(ti: tarfile.TarInfo) -> tarfile.TarInfo: # design exists to avoid, one per failed attempt at DR6 scale. tmp = tar_path.with_name(tar_path.name + ".tmp") try: + # One pass in sorted member order, taking each member from whichever + # side has it: a live source on disk, or the existing tar. Members are + # copied across with their own TarInfo, so a carried member is + # byte-for-byte what it was and a rerun that changes nothing still + # produces an identical archive. with tarfile.open(tmp, "w", format=tarfile.PAX_FORMAT) as tf: - for f in files: - tf.add(seen[f["name"]][0], arcname=f["name"], filter=anonymous) + old_tar = (tarfile.open(tar_path) if carried else None) + try: + for f in files: + if f["name"] in seen: + tf.add(seen[f["name"]][0], arcname=f["name"], + filter=anonymous) + else: + ti = anonymous(old_tar.getmember(f["name"])) + tf.addfile(ti, old_tar.extractfile(f["name"])) + finally: + if old_tar is not None: + old_tar.close() if tar_path.exists() and filecmp.cmp(tmp, tar_path, shallow=False): tmp.unlink() # unchanged: leave the mtime alone else: @@ -354,7 +422,10 @@ def anonymous(ti: tarfile.TarInfo) -> tarfile.TarInfo: finally: tmp.unlink(missing_ok=True) - warn = f" ({len(empty)} pattern(s) matched nothing: {empty})" if empty else "" + warn = (f" ({len(empty)} retention product(s) matched nothing: {empty})" + if empty else "") + if carried: + warn += f" ({len(carried)} member(s) carried from the existing tar)" print(f"[persist_exp] {args.exp}: {len(files)} file(s), " f"{body['bytes'] / 1e6:.1f} MB -> {tar_path}{warn}") From 0fb514b85d88c85c8da5d91949bef8738500fae1 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 20:16:43 -0400 Subject: [PATCH 18/85] feat(orchestration): the star catalogue becomes hdf5, reconciled like the tile one MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The campaign's two products behaved differently for no reason anyone chose. The shear catalogue reconciled — an append read the appended tiles — while the star catalogue was one flat FITS table that had to be restacked from every exposure the campaign had ever seen to add one: ~40 GB of members at DR6 scale to add ~2 MB, held in memory while it happened. Now they are the same thing. /full_starcat_.hdf5, one dataset per exposure at exposures/, each holding that exposure's every CCD's rows with a CCD_NB column, an n_exposures root attribute and the same column digest the tile side carries. Named for the campaign exactly as the shear catalogue beside it is. CCD_NB IS AN INT: it is parsed out of the member name, where it is always digits, so a string buys nothing and costs 8 bytes a row against 4. DTYPES ARE NATIVE: float32 stays float32, where the FITS writer widened every float column to 1D, doubling both the file and the peak memory of the job that wrote it for no information. MEMORY IS NOW FLAT IN THE CAMPAIGN — one exposure at a time — so the rule is sized on the largest exposure's members rather than the campaign's, and the ~240 GB a DR6-scale flat table would have wanted is simply not a number any more. The Snakefile's sizing block keeps the measurements that got us here, because they are the argument for the format. THE RECONCILE MACHINERY IS NOW ONE MODULE, workflow/scripts/hdf5_reconcile.py, used by both merges rather than duplicated: plan against a read-only open, add/refresh/remove, refresh everything when the column digest moves, compact rewrite when anything is removed or refreshed, untouched on a no-op. Writing it twice would have been two chances to disagree about what an output owes its inputs. THE WORKFLOW NO LONGER CALLS MergeStarCat* AT ALL, so merge_star_cat.py drops the shapepipe import and the psf-model switch, and the [fileobj, name] entry shape those classes learned for it is REVERTED — with no caller it was upstream surface with nothing behind it. What stays upstream is what fixes the module runner's own problems: the two-pass allocation, and asking each file for its own optional columns instead of deciding once for the merge. The runner path is byte-identical to before all of it, md5 f7caa1cf… on the fixture. VERIFIED on the fixtures: build (2 added); no-op (unchanged, mtime identical to the nanosecond); append one exposure with the others' tars at chmod 000, which succeeds and reports 1 added, so they were demonstrably not read; remove one (1 removed, dataset gone, n_exposures 2). Every one of the 16 columns equals the FITS version's values. The tile side's compaction sequence was re-run against a fix this work exposed — the keep-what-changed path called Dataset.copy, which does not exist, and only bites when a plan both rewrites and keeps something. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- .../merge_starcat_package/merge_starcat.py | 54 +-- workflow/README.md | 17 +- workflow/Snakefile | 88 ++-- workflow/rules/exposure.smk | 45 ++- workflow/scripts/hdf5_reconcile.py | 167 ++++++++ workflow/scripts/merge_final_cat.py | 207 ++-------- workflow/scripts/merge_star_cat.py | 380 +++++++++--------- 7 files changed, 500 insertions(+), 458 deletions(-) create mode 100644 workflow/scripts/hdf5_reconcile.py diff --git a/src/shapepipe/modules/merge_starcat_package/merge_starcat.py b/src/shapepipe/modules/merge_starcat_package/merge_starcat.py index 7f751006c..1e0cc7256 100644 --- a/src/shapepipe/modules/merge_starcat_package/merge_starcat.py +++ b/src/shapepipe/modules/merge_starcat_package/merge_starcat.py @@ -269,15 +269,10 @@ def process(self): my_mask[inside_circle] = True for name in self._input_file_list: - # The source to read and the NAME to report it by; identical for a - # plain [path] entry (see MergeStarCatPSFEX's docstring on the - # [fileobj, name] form). This class takes its CCD numbers from the - # data's own CCD_ID_LIST, so the name is only ever used in messages. - source, label = name[0], name[-1] try: - starcat_j = fits.open(source, memmap=False, ignore_missing_simple=True) + starcat_j = fits.open(name[0], memmap=False, ignore_missing_simple=True) except ValueError: - print(f"Error for file {label}, check FITS file integrity") + print(f"Error for file {name[0]}, check FITS file integrity") #raise continue @@ -536,15 +531,7 @@ class MergeStarCatPSFEX(object): Parameters ---------- input_file_list : list - Input entries. Each entry is a list, as the module runner builds them: - ``[path]`` from the file handler. An entry may also carry a name - alongside an already-open source, ``[fileobj, name]`` — ``fits.open`` - takes the first element and the CCD number is parsed from the LAST, - which is the same string in the one-element case. That is what lets a - caller merge catalogues it never wrote to disk (the Snakemake - workflow's ``star_cat_merge`` reads them out of the per-exposure tars - with ``tarfile`` + ``BytesIO``), without this class learning anything - about where they came from. + Input files output_dir : str Output directory w_log : logging.Logger @@ -586,9 +573,9 @@ def __init__( # (MKDEBUG); zero-filled when missing rather than failing the merge. _OPTIONAL = (("MAG", "MAG"), ("SNR", "SNR"), ("ACCEPTED", "ACCEPTED")) - def _ccd_nb(self, label): + def _ccd_nb(self, path): """The CCD number this catalogue's rows carry, parsed from its name.""" - return re.split(r"\-([0-9]*)\-([0-9]+)\.", label)[-2] + return re.split(r"\-([0-9]*)\-([0-9]+)\.", path)[-2] def process(self): """Process. @@ -611,10 +598,9 @@ def process(self): accumulation: there are no chunks, and no concatenate that must hold its inputs and its result at the same time. - ``self._input_file_list`` MUST BE ITERABLE TWICE. A list is; so is the - workflow's tar reader, whose ``__iter__`` opens the archives afresh. - A one-shot generator is not, and would silently merge nothing on the - second pass — hence the explicit length check below. + ``self._input_file_list`` MUST BE ITERABLE TWICE, which the module + runner's list is. A one-shot generator is not, and would silently merge + nothing on the second pass — hence the explicit row-count check below. """ self._w_log.info( f"Merging {len(self._input_file_list)} star catalogues" @@ -629,11 +615,10 @@ def process(self): # ordinary-first raised KeyError on the first converted one. So the # dtype comes from ANY file that carries the column, and pass 2 asks # each file for itself. - labels, dtypes, opt_dtypes, n_total = [], None, {}, 0 + names, dtypes, opt_dtypes, n_total = [], None, {}, 0 for name in self._input_file_list: - source, label = name[0], name[-1] try: - with fits.open(source, memmap=False, + with fits.open(name[0], memmap=False, ignore_missing_simple=True) as starcat_j: hdu = starcat_j[self._hdu_table] n_rows = hdu.header["NAXIS2"] @@ -649,10 +634,10 @@ def process(self): if col not in opt_dtypes and col in (cols.names or ()): opt_dtypes[col] = cols[col] except OSError: - print(f"Error while opening file '{label}'") + print(f"Error while opening file '{name[0]}'") #raise continue - labels.append(label) + names.append(name[0]) n_total += n_rows if dtypes is None: @@ -667,15 +652,14 @@ def process(self): data[out] = np.empty(n_total, dtype=opt_dtypes.get(col, dtypes["X"])) # CCD_NB is one string per catalogue, repeated over its rows; its width # is the widest CCD number in the campaign, which pass 1 already knows. - width = max((len(self._ccd_nb(lb)) for lb in labels), default=1) + width = max((len(self._ccd_nb(n)) for n in names), default=1) data["CCD_NB"] = np.empty(n_total, dtype=f"U{width}") # --- pass 2: fill --------------------------------------------------- at = 0 for name in self._input_file_list: - source, label = name[0], name[-1] try: - starcat_j = fits.open(source, memmap=False, + starcat_j = fits.open(name[0], memmap=False, ignore_missing_simple=True) except OSError: continue @@ -690,7 +674,7 @@ def process(self): # THIS file's schema, not the merge's: zero-fill only the files # that actually lack the column. data[out][sl] = data_j[col] if col in have else 0 - data["CCD_NB"][sl] = self._ccd_nb(label) + data["CCD_NB"][sl] = self._ccd_nb(name[0]) at += n_rows starcat_j.close() @@ -854,11 +838,7 @@ def process(self): ) for name in self._input_file_list: - # The source to read and the NAME to parse the CCD number out of; - # identical for a plain [path] entry (see MergeStarCatPSFEX's - # docstring on the [fileobj, name] form). - source, label = name[0], name[-1] - starcat_j = fits.open(source, memmap=False) + starcat_j = fits.open(name[0], memmap=False) data_j = starcat_j[self._hdu_table].data @@ -892,7 +872,7 @@ def process(self): # CCD number ccd_nb.append(np.full( len(data_j["XWIN_IMAGE"]), - re.split(r"\-([0-9]*)\-([0-9]+)\.", label)[-2])) + re.split(r"\-([0-9]*)\-([0-9]+)\.", name[0])[-2])) # Prepare output FITS catalogue output = file_io.FITSCatalogue( diff --git a/workflow/README.md b/workflow/README.md index 85cf4f465..40371f2b4 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -312,14 +312,15 @@ profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; kee written here, because that script's own discovery walks a directory layout this workflow does not have. `campaign:` in `config.yaml` names the group and defaults to the persistent root's basename. - `star_cat_merge` restacks the whole campaign, so its output is a function of - its input set and byte-stable on a no-op rerun (tmp-then-`cmp`-then-`mv`). - `final_cat_merge` RECONCILES instead — adds the tiles that have no dataset, - drops datasets whose tile left the campaign, re-reads one whose catalogue - changed (each dataset records its source's size and mtime), and leaves the - rest unread — because re-reading a campaign to add one tile is ~800 GB of IO - at DR6 scale. Its *content* is still a function of the input set; its byte - layout is not, and a no-op leaves the file untouched rather than rewritten. + BOTH RECONCILE, through one shared module (`hdf5_reconcile.py`) so the + campaign's two products cannot disagree about what an output owes its inputs. + Each adds the units that have no dataset, drops datasets whose unit left the + campaign, re-reads one whose source changed (every dataset records its + source's size and mtime) or whose column set moved (a digest on the file's + root), and leaves the rest unread — because re-reading a campaign to add one + unit is ~800 GB of IO at DR6 scale. The *content* is still a function of the + input set; the byte layout is not, and a no-op leaves the file untouched + rather than rewritten. Both rerun when the set changes: the unit ids' fingerprint rides on `params`. Neither is a `localrule` — one job over ~20k units is real work — and neither puts its input paths in its shell, which is not fastidiousness: ~20k paths is diff --git a/workflow/Snakefile b/workflow/Snakefile index c509b770e..ea4129f25 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -31,6 +31,7 @@ the manifest says "this stage succeeded", the log says "here is what happened" (the contract is argued in completeness.py's docstring). """ +import fnmatch import functools import hashlib import json @@ -647,9 +648,14 @@ def unit_fingerprint(units): def full_starcat(): - """The campaign's merged star catalogue. The NAME is not ours to choose: - sp_validation hardcodes `full_starcat-0000000.fits` beside its data dir.""" - return f"{PRODUCTS_DIR}/full_starcat-0000000.fits" + """The campaign's merged star catalogue — the rho/tau statistics input. + + hdf5, one dataset per exposure, named for the campaign exactly as the shear + catalogue beside it is. The old flat FITS table it replaces was called + `full_starcat-0000000.fits` and sp_validation still opens that name; + CosmoStat/sp_validation#340 moves its readers to this file, the same + migration that retires the `patches/` key on the galaxy side.""" + return f"{PRODUCTS_DIR}/full_starcat_{CAMPAIGN}.hdf5" def final_cat_hdf5(): @@ -733,32 +739,31 @@ def star_cat_exposures(): # against synthetic tars for the star side and against smk-g6's real # catalogues for the tile side. Peak RSS is getrusage(RUSAGE_CHILDREN). # -# STAR SIDE. Two points, 20 and 80 exposures of 40 CCDs x 400 stars (1.6 MB of -# members per exposure, against the 2.0 MB measured on smk-m2), across the two -# rewrites this PR made to the accumulation in MergeStarCatPSFEX: +# STAR SIDE, AND IT IS FLAT IN THE CAMPAIGN. The merge writes one hdf5 dataset +# per exposure and reads one exposure at a time, so it is sized on the LARGEST +# exposure's members — ~2 MB — not on the campaign's. What follows is the +# history of how that came to be true, because the numbers are the argument. +# +# The FITS full_starcat this replaced was one flat table, so the job held the +# whole campaign. Two fixture points, 20 and 80 exposures of 40 CCDs x 400 stars +# (1.6 MB of members per exposure, against the 2.0 MB measured on smk-m2), +# across the rewrites this PR made to MergeStarCatPSFEX: # # input members python lists arrays+concat two passes # 32.3 MB 383 MB 238 MB 221 MB # 129.0 MB 1313 MB 740 MB 661 MB -# # slope 10.1x 5.5x 4.8x # -# The tenfold was one python float object (32 bytes) plus a list pointer (8) -# per 4 bytes of float32 payload. Arrays per catalogue removed that; the -# two-pass structure — count rows from the FITS headers, allocate once at the -# exact length, then fill — removed what remained, so nothing is held twice. +# The tenfold was one python float object (32 bytes) plus a list pointer (8) per +# 4 bytes of float32 payload. Arrays per catalogue removed that; counting rows +# from the headers and filling a preallocated array removed the rest. What +# remained at 4.8x was the OUTPUT: file_io writes every float column as FITS 1D, +# so float32 became a float64 table astropy then buffered. # -# THE REMAINING 4.8x IS THE OUTPUT SIDE, and it is not a leak: file_io writes -# every float column as FITS 1D, so a float32 input becomes a float64 table -# that astropy then buffers to write — 141 MB of table for 78 MB of payload at -# the 80-exposure point, plus its write copy. Halving it means changing the -# OUTPUT format, which is what sp_validation reads: a different decision from -# this one, and not ours to take here. -# -# THE CEILING MOVED AND IS NOW ELSEWHERE. At ~2 MB of members per exposure a -# 16 GB job merges ~1300 exposures rather than ~800, and DR6's ~20k would want -# ~240 GB rather than ~400 GB. What stands between here and a full-survey -# full_starcat is the float64 output, not the merge. +# Per-exposure hdf5 removes the term entirely rather than shrinking it — and +# with it the ~240 GB a DR6-scale flat table would have wanted. Those +# improvements stay upstream regardless: the module runner still merges to one +# FITS table, and they are its fix. # # TILE SIDE, and it is the reassuring one. Two points against real smk-g6 # catalogues, 2 tiles (73.9 MB in, largest 39.6 MB) and 6 tiles (235.5 MB in, @@ -795,8 +800,8 @@ def capped_mem(mb, rule): return mb -STAR_MEM_BASE_MB = 500 # interpreter + astropy + shapepipe, rounded up -STAR_MEM_FACTOR = 6 # x input bytes; 4.8 measured, rounded up +STAR_MEM_BASE_MB = 500 # interpreter + astropy + h5py, rounded up +STAR_MEM_FACTOR = 6 # x the LARGEST exposure's members FINAL_MEM_BASE_MB = 800 FINAL_MEM_FACTOR = 4 # x the LARGEST tile; ~3 measured # What one unit costs when its product is not on disk yet to be measured — a @@ -808,6 +813,7 @@ EXP_BYTES_DEFAULT = 2_000_000 # The product whose members star_cat_merge stacks — named once, here and in # merge_star_cat.py, and resolved through persist_exp.py's catalogue. STAR_CAT_PRODUCT = _persist.ALWAYS +STAR_CAT_PATTERN = _persist.resolve(_persist.ALWAYS) TILE_BYTES_DEFAULT = 46_000_000 @@ -819,6 +825,17 @@ def _size(path, default): return default +def star_cat_max_bytes(): + """The LARGEST exposure's psf_validation members — what sizes the merge. + + The star merge holds ONE exposure at a time now that its output is hdf5 + with a dataset per exposure, so its memory is flat in the campaign exactly + as the tile side's is. Sizing on the total would ask a node for a campaign's + worth of memory to hold ~2 MB. + """ + return max(_star_cat_exposure_bytes() or [EXP_BYTES_DEFAULT]) + + def star_cat_bytes(): """Total bytes of the members the star merge will actually read. @@ -836,19 +853,30 @@ def star_cat_bytes(): this job. An exposure not yet packed has no manifest and contributes the measured default, which is the psf_validation figure and not the tar's. """ - total = 0 + return sum(_star_cat_exposure_bytes()) + + +@functools.lru_cache(maxsize=1) +def _star_cat_exposure_bytes(): + """Per exposure, the bytes of the members the star merge will read.""" + out = [] for exp in star_cat_exposures(): manifest = Path(prod_exp_manifest(exp, "exp_persist")) if not manifest.exists(): - total += EXP_BYTES_DEFAULT + out.append(EXP_BYTES_DEFAULT) continue try: body = json.loads(manifest.read_text()) - total += sum(f["bytes"] for f in body["files"] - if f.get("product") == STAR_CAT_PRODUCT) + # By product name or, for a manifest written before that field + # existed or by a raw-glob keep list, by file name — the same test + # merge_star_cat.is_member() applies, so the sizing counts exactly + # the members the job will read. + out.append(sum(f["bytes"] for f in body["files"] + if f.get("product") == STAR_CAT_PRODUCT + or fnmatch.fnmatch(f["name"], STAR_CAT_PATTERN))) except (OSError, ValueError, KeyError): - total += EXP_BYTES_DEFAULT - return total + out.append(EXP_BYTES_DEFAULT) + return out def final_cat_max_bytes(): diff --git a/workflow/rules/exposure.smk b/workflow/rules/exposure.smk index 0b6a7a715..5845adb8d 100644 --- a/workflow/rules/exposure.smk +++ b/workflow/rules/exposure.smk @@ -225,14 +225,20 @@ rule clean_exposure: # --- the campaign's star catalogue ------------------------------------------ # ONE job per campaign: every exposure's every CCD's `validation_psf--.fits`, -# stacked into `/full_starcat-0000000.fits`. That file is the -# rho/tau statistics input and sp_validation reads it at exactly that path, -# doing no merging of its own; the old bash chain built it with -# `combine_runs.bash psf` + a `merge_starcat_runner` pass, and the workflow -# emitted neither. The stacking itself is `MergeStarCatPSFEX` — the same class -# the old runner called, reused rather than restated, so a column added to the -# module is a column added here (merge_star_cat.py argues the reuse and the -# tar-member reading). +# collected into `/full_starcat_.hdf5`, one dataset per +# exposure. That file is the rho/tau statistics input; the old bash chain built +# a flat FITS table with `combine_runs.bash psf` + a `merge_starcat_runner` +# pass, and the workflow emitted neither. sp_validation still opens the FITS +# name today — CosmoStat/sp_validation#340 moves its readers to this file, the +# same migration that retires the `patches/` key on the tile side. +# +# ONE DATASET PER EXPOSURE, NOT ONE TABLE, and it is the same decision as the +# tile side's: it makes the file RECONCILABLE. A flat table had to be restacked +# from every exposure the campaign had ever seen to add one — ~40 GB of members +# at DR6 scale to add ~2 MB — and held the whole campaign in memory while it did +# so. Reconciled, an append reads the appended exposures and nothing else, and +# the job holds one exposure at a time. hdf5_reconcile.py is the shared +# machinery; merge_star_cat.py argues the format and the tar reading. # # THE INPUT IS star_cat_inputs() (Snakefile): every exposure of TILES_READY whose # PSF products are on the persistent root — the live ones through the exp_persist @@ -253,19 +259,14 @@ rule clean_exposure: # the point: a job that stacked anything the fingerprint did not see would be # rows no rerun trigger could notice, which is what a glob over products_dir # would have given on a root shared with an earlier, larger tile list. -# Byte-stable output otherwise (tmp-then-cmp-then-mv), so a no-op rerun does not -# move its mtime. # # NOT A LOCALRULE. exp_persist is local because it is 20k jobs of seconds; this -# is one job that holds a campaign's stars in memory (~800k catalogues at DR6 -# scale). mem_mb is a guess scaled by attempt, not a measurement — the campaigns -# run so far are 127 exposures, three orders of magnitude short of the case this -# sizing is for, and the first DR6-scale run should replace this number with a -# benchmark. +# is one job that reads the campaign's tars end to end. Its MEMORY is flat in +# the campaign (one exposure at a time) and sized on the largest exposure; its +# RUNTIME is the total. # -# NO JOB AT ALL when `persist_exp:` keeps no validation catalogue, or when every -# exposure in scope is tombstoned: star_cat_targets() (Snakefile) simply does not -# request the output, and the parse says so rather than a node failing later. +# NO JOB AT ALL when every exposure in scope is tombstoned with no tar left +# behind: star_cat_targets() (Snakefile) simply does not request the output. rule star_cat_merge: input: lambda wc: star_cat_inputs() @@ -275,6 +276,7 @@ rule star_cat_merge: products_dir = str(PRODUCTS_DIR), tile_list = str(config["tile_list"]), index_db = str(INDEX_DB), + campaign = CAMPAIGN, inputs = unit_fingerprint(star_cat_exposures()), script_hash = MERGE_STAR_HASH threads: 1 @@ -283,8 +285,11 @@ rule star_cat_merge: # measured (the Snakefile's sizing block carries both points, and the # ceiling this rule runs into at DR6 scale). Still * attempt, because a # measured slope on synthetic tars is not a guarantee about real ones. + # Sized on the LARGEST exposure, not the total: the merge holds one + # exposure at a time (the Snakefile's sizing block carries the history). mem_mb = lambda wc, attempt: capped_mem(attempt * ( - STAR_MEM_BASE_MB + STAR_MEM_FACTOR * star_cat_bytes() // 1_000_000), + STAR_MEM_BASE_MB + + STAR_MEM_FACTOR * star_cat_max_bytes() // 1_000_000), "star_cat_merge"), # ~2 min per GB of members on the measurement above, doubled, over a # floor that covers the fixed cost of opening ~40 members per exposure. @@ -296,4 +301,4 @@ rule star_cat_merge: " --products-dir '{params.products_dir}'" " --tile-list '{params.tile_list}' --index-db '{params.index_db}'" " --output {output.star_cat}" - f" --psf-model {PSF_MODEL}" + " --campaign '{params.campaign}'" diff --git a/workflow/scripts/hdf5_reconcile.py b/workflow/scripts/hdf5_reconcile.py new file mode 100644 index 000000000..0635b5976 --- /dev/null +++ b/workflow/scripts/hdf5_reconcile.py @@ -0,0 +1,167 @@ +#!/usr/bin/env python3 +"""Bring an hdf5 catalogue into agreement with a campaign, one dataset per unit. + +Shared by the two campaign-level merges — ``merge_final_cat.py`` (one dataset +per tile) and ``merge_star_cat.py`` (one per exposure) — because they want +exactly the same thing of their output and disagreeing about it would be a bug +waiting to happen rather than a difference worth having. + +WHY RECONCILE RATHER THAN REBUILD. The output must be a function of the input +set — that is what makes the rules' fingerprints mean anything — but reading +every unit to add one is ~800 GB of IO at DR6 scale for a few tens of MB of new +data. So the file is brought INTO AGREEMENT with the campaign instead: + + * a unit with no dataset is read and added; + * a dataset whose unit has left the campaign is deleted; + * a dataset whose SOURCE has changed is re-read. Each records its source's + size and mtime as attributes, and a mismatch is what changed means. This is + the only reason a finished unit is read twice, and it is why the file + cannot drift from its inputs the way an append-only tool does; + * a dataset whose column set was written under a DIFFERENT SCHEMA is re-read. + The column set is the one input nothing else can see: it is not a source + file, so no stamp moves when it changes. It travels as a digest on the + file's root. + * a dataset that agrees with its source and its schema is left alone, unread. + +An append therefore reads exactly the appended units. + +WHAT IS AND IS NOT A FUNCTION OF THE INPUT SET. The file's CONTENT is: the same +units with the same sources give the same datasets, the same columns and the +same count attribute, whether they arrived at once or one batch at a time. Its +BYTE LAYOUT is not, because hdf5 lays a group out in the order things were +added. That is the trade for not re-reading the campaign, and it is why the +no-op case compares ACTIONS rather than bytes. + +UNTOUCHED ON A NO-OP, which is stronger than byte-stable and cheaper to +establish. Reconciling is PLANNED against a read-only open; an empty plan never +opens the file for writing, so its mtime cannot move — and mtime is a rerun +trigger, so an unconditional rewrite would make every invocation look like a +change. +""" + +import hashlib +import shutil +from pathlib import Path + +import h5py + + +def schema_digest(columns) -> str: + """A fingerprint of the COLUMN SET the datasets were written with.""" + return hashlib.md5("\n".join(columns).encode()).hexdigest()[:16] + + +def stamp(path: Path) -> tuple: + """A source's identity, as recorded on the dataset built from it. + + Size and mtime, not a checksum: the question is "did this change since we + read it", which mtime answers for a pipeline that writes a file once. A + campaign that rewrote a source in place with identical size and mtime would + defeat it, and nothing does. + """ + st = Path(path).stat() + return st.st_size, st.st_mtime_ns + + +class Plan: + """What reconciling requires: three unit lists. + + ``add`` and ``refresh`` are both "read the source and write the dataset"; + they are separate only so the log can say which happened, because a refresh + means a finished unit's source moved under us and that is worth seeing. + """ + + def __init__(self, add, refresh, remove): + self.add, self.refresh, self.remove = add, refresh, remove + + def empty(self): + return not (self.add or self.refresh or self.remove) + + def describe(self): + return (f"{len(self.add)} added, {len(self.refresh)} refreshed, " + f"{len(self.remove)} removed") + + +def plan(output: Path, group_path: str, units: list, digest: str) -> Plan: + """Compare the file on disk with the campaign, WITHOUT writing anything.""" + if not output.exists(): + return Plan([u for u, _ in units], [], []) + + want = {unit for unit, _ in units} + add, refresh = [], [] + with h5py.File(output, "r") as f: + stale_schema = f.attrs.get("param_digest") != digest + have = dict(f[group_path].items()) if group_path in f else {} + present = set(have) + for unit, source in units: + if unit not in present: + add.append(unit) + elif stale_schema: + refresh.append(unit) + else: + attrs = have[unit].attrs + if (int(attrs.get("src_bytes", -1)), + int(attrs.get("src_mtime_ns", -1))) != stamp(source): + refresh.append(unit) + return Plan(add, refresh, sorted(present - want)) + + +def apply(output: Path, group_path: str, todo: Plan, units: list, read, + digest: str, count_attr: str) -> None: + """Carry the plan out on a tmp file, then move it into place. + + ``read(unit, source)`` returns the structured array for one unit; it is + called only for the units the plan names, which is what makes an append + cheap. + + TWO WAYS TO BUILD THE TMP, and which one is used is about SPACE, not speed. + HDF5 never reclaims the space a deleted dataset occupied, so a file that is + copied and then edited in place grows for the life of the campaign — every + refresh of a unit leaks that unit. So: + + * a plan that only ADDS copies the existing file and appends to it. There + is nothing to reclaim, and copying beats rewriting. + * a plan that removes or refreshes anything builds the tmp FRESH, moving + the datasets it keeps across with h5py's own group copy — a + dataset-level copy inside the library that never reads a row into numpy + — and writing only the units that actually changed. The result is + compact. + + Either way the tmp is moved into place at the end, so a crash mid-merge + leaves the old catalogue intact rather than a half-written one. A SIGKILL + between writing the tmp and renaming it leaves the tmp behind — one file, + beside the catalogue, overwritten by the next run; the rename itself is + atomic, which is the property that matters. + """ + sources = dict(units) + rewrite = bool(todo.remove or todo.refresh) + written = set(todo.add) | set(todo.refresh) + keep = [u for u, _ in units if u not in written] + tmp = output.with_name(output.name + ".tmp") + try: + tmp.unlink(missing_ok=True) + if output.exists() and not rewrite: + shutil.copy2(output, tmp) + with h5py.File(tmp, "a") as f: + group = (f[group_path] if group_path in f + else f.create_group(group_path)) + if rewrite and output.exists(): + with h5py.File(output, "r") as src: + for unit in keep: + # File.copy, not Dataset.copy — the latter does not + # exist, and the difference only shows when a plan both + # rewrites and keeps something. + src.copy(f"{group_path}/{unit}", group, name=unit) + for unit in todo.add + todo.refresh: + source = sources[unit] + data = read(unit, source) + dset = group.create_dataset(unit, data=data, dtype=data.dtype) + # The dataset's own record of what it was read from; this is + # what lets a later invocation leave it alone. + dset.attrs["src_bytes"], dset.attrs["src_mtime_ns"] = \ + stamp(source) + f.attrs[count_attr] = len(group) + f.attrs["param_digest"] = digest + tmp.replace(output) # atomic: same filesystem + finally: + tmp.unlink(missing_ok=True) diff --git a/workflow/scripts/merge_final_cat.py b/workflow/scripts/merge_final_cat.py index 4bae00f98..2db61e7e0 100644 --- a/workflow/scripts/merge_final_cat.py +++ b/workflow/scripts/merge_final_cat.py @@ -38,43 +38,18 @@ loaded by path rather than imported: it is a script, not an installed module, and the container's ``shapepipe`` install does not carry it. -IT RECONCILES, IT NEITHER REBUILDS NOR BLINDLY APPENDS. The output must be a -function of the input set — that is what makes the rule's fingerprint mean -something — but reading every tile's catalogue to add one tile is ~800 GB of IO -at DR6 scale for ~35 MB of new data. So the file is brought INTO AGREEMENT with -the campaign instead: - - * a campaign tile with no dataset is read and added; - * a dataset whose tile is no longer in the campaign is deleted; - * a dataset whose source catalogue has CHANGED is re-read. Each one records - its source's size and mtime as attributes, and a mismatch is what "changed" - means. This is the only reason a finished tile is ever read twice, and it is - the reason the file cannot drift from its inputs the way an append-only - tool does; - * a dataset that agrees with its source is left alone, unread. - -An append therefore reads exactly the appended tiles. ``create_final_cat.py``'s -own ``process()`` implements the append-only half of this — it skips a tile -already in the file, whatever the file on disk now says — which is right for a -hand-driven update and wrong for a DAG output. (Its ``-s`` single-ID mode -implements ``check`` and ``remove``; ``add`` is accepted by the argument -validator and then falls through to the ordinary walk, so it is not a way to -add one tile by hand.) - -WHAT IS AND IS NOT A FUNCTION OF THE INPUT SET. The file's CONTENT is: the same -tiles with the same catalogues give the same datasets, the same columns and the -same n_tiles, whether they arrived at once or one campaign at a time. Its BYTE -LAYOUT is not, because hdf5 lays out a group in the order things were added. -That is the trade for not re-reading the campaign, and it is why the no-op case -below compares actions rather than bytes. -UNTOUCHED ON A NO-OP RERUN, which is stronger than byte-stable and cheaper to -establish. Reconciling is planned before anything is written: if the plan is -empty the file is not opened for writing at all, so its mtime cannot move — and -mtime is a rerun trigger, so an unconditional rewrite would make every -invocation look like a change. When the plan is NOT empty the existing file is -copied to a tmp path, changed there and moved into place, so a crash mid-merge -leaves the old catalogue intact rather than a half-written one. The copy is a -fraction of the reading it replaces. +IT RECONCILES, IT NEITHER REBUILDS NOR BLINDLY APPENDS, and the machinery for +that is ``hdf5_reconcile.py``, shared with the star side so the campaign's two +products cannot disagree about what an output owes its inputs. That module +carries the argument in full: an append reads the appended tiles, a source that +changed is re-read, a tile that left the campaign is deleted, a column-set +change refreshes everything, and a no-op leaves the file untouched. +``create_final_cat.py``'s own ``process()`` implements only the append-only half +— it skips a tile already in the file, whatever the file on disk now says — +which is right for a hand-driven update and wrong for a DAG output. (Its ``-s`` +single-ID mode implements ``check`` and ``remove``; ``add`` is accepted by the +argument validator and then falls through to the ordinary walk, so it is not a +way to add one tile by hand.) WHICH TILES — AND WHY THE JOB DERIVES THE SET RATHER THAN BEING TOLD IT. The set is the CAMPAIGN's: every tile both declared in ``tile_list`` and present in the @@ -96,16 +71,13 @@ """ import argparse -import hashlib import importlib.util -import shutil import sys from pathlib import Path -import h5py - # Same directory; the rule invokes this file by path, so it is sys.path[0]. import build_index +import hdf5_reconcile # /scripts/python/create_final_cat.py, from /workflow/scripts/this. CFC_PATH = (Path(__file__).resolve().parents[2] @@ -156,136 +128,6 @@ def catalogues(products_dir: Path, tile_list: Path, index_db: Path) -> list: return out -def schema_digest(param_list: list) -> str: - """A fingerprint of the COLUMN SET the datasets were written with. - - Recorded on the file's root and compared on every reconcile, because the - column set is the one input to this merge that nothing else can see. It is - not a source catalogue, so no dataset's size/mtime stamp moves when it - changes; it reaches the job through --param-file, which is a `params` value - and not a rule input. Without this, editing final_cat.param — which this PR - itself does — would leave every dataset in an existing hdf5 written to the - OLD schema, and nothing would ever notice. - """ - return hashlib.md5("\n".join(param_list).encode()).hexdigest()[:16] - - -class Plan: - """What reconciling this campaign into this file requires: three tile lists. - - ``add`` and ``refresh`` are both "read the catalogue and write the dataset"; - they are separate only so the log can say which happened, because a refresh - means a finished tile's catalogue moved under us and that is worth seeing. - """ - - def __init__(self, add, refresh, remove): - self.add, self.refresh, self.remove = add, refresh, remove - - def empty(self): - return not (self.add or self.refresh or self.remove) - - def describe(self): - return (f"{len(self.add)} added, {len(self.refresh)} refreshed, " - f"{len(self.remove)} removed") - - -def stamp(path: Path) -> tuple: - """The source catalogue's identity, as recorded on its dataset. - - Size and mtime, not a checksum: the file is ~35 MB and the question is - "did this change since we read it", which mtime answers for a pipeline - that writes a catalogue once. A campaign that rewrites a final_cat in - place with identical size and mtime would defeat it, and nothing does. - """ - st = path.stat() - return st.st_size, st.st_mtime_ns - - -def reconcile_plan(output: Path, group_path: str, tiles: list, - digest: str) -> Plan: - """Compare the file on disk with the campaign, WITHOUT writing anything. - - Opened read-only, so a no-op invocation cannot move the output's mtime. - """ - if not output.exists(): - return Plan([t for t, _ in tiles], [], []) - - want = {tile: path for tile, path in tiles} - add, refresh = [], [] - with h5py.File(output, "r") as f: - # A changed column set invalidates every dataset at once — they were - # written to the old schema and nothing about their sources moved. - stale_schema = f.attrs.get("param_digest") != digest - have = dict(f[group_path].items()) if group_path in f else {} - present = set(have) - for tile, path in tiles: - if tile not in present: - add.append(tile) - elif stale_schema: - refresh.append(tile) - else: - attrs = have[tile].attrs - if (int(attrs.get("src_bytes", -1)), - int(attrs.get("src_mtime_ns", -1))) != stamp(path): - refresh.append(tile) - return Plan(add, refresh, sorted(present - set(want))) - - -def apply_plan(output: Path, group_path: str, plan: Plan, tiles: list, - cfc, params: dict, digest: str) -> None: - """Carry the plan out on a tmp file, then move it into place. - - TWO WAYS TO BUILD THE TMP, and which one is used is about SPACE, not speed. - HDF5 never reclaims the space a deleted dataset occupied, so a file that is - copied and then edited in place grows for the life of the campaign — every - refresh of a 15 MB tile leaks 15 MB. So: - - * a plan that only ADDS copies the existing file and appends to it. There - is nothing to reclaim, and copying beats rewriting. - * a plan that removes or refreshes anything builds the tmp FRESH, moving - the datasets it keeps across with h5py's own group copy — which is a - dataset-level copy inside the library and never reads a row into numpy — - and writing only the tiles that actually changed. The result is compact. - - Either way the tmp is moved into place at the end, so a crash mid-merge - leaves the old catalogue intact rather than a half-written one. A SIGKILL - between writing the tmp and renaming it leaves the tmp behind — one file, - next to the catalogue, overwritten by the next run; the rename itself is - atomic, which is the property that matters. - """ - paths = dict(tiles) - rewrite = bool(plan.remove or plan.refresh) - keep = [t for t, _ in tiles if t not in set(plan.add) | set(plan.refresh)] - tmp = output.with_name(output.name + ".tmp") - try: - tmp.unlink(missing_ok=True) - if output.exists() and not rewrite: - shutil.copy2(output, tmp) - with h5py.File(tmp, "a") as f: - group = (f[group_path] if group_path in f - else f.create_group(group_path)) - if rewrite and output.exists(): - with h5py.File(output, "r") as src: - for tile in keep: - src[f"{group_path}/{tile}"].copy( - src[f"{group_path}/{tile}"], group, name=tile) - for tile in plan.add + plan.refresh: - path = paths[tile] - extracted, dtype = cfc.read_data(str(path), params) - data = cfc.copy_data(params["param_list"], extracted, dtype) - dset = group.create_dataset(tile, data=data, dtype=data.dtype) - # The dataset's own record of what it was read from; this is - # what makes a later invocation able to leave it alone. - dset.attrs["src_bytes"], dset.attrs["src_mtime_ns"] = stamp(path) - # The same attribute create_final_cat.py's print_list() writes, and - # what sp_validation reads to know how many tiles it is holding. - f.attrs["n_tiles"] = len(group) - f.attrs["param_digest"] = digest - tmp.replace(output) # atomic: same filesystem - finally: - tmp.unlink(missing_ok=True) - - def main() -> None: p = argparse.ArgumentParser(description=__doc__) p.add_argument("--products-dir", required=True, type=Path, @@ -320,16 +162,23 @@ def main() -> None: args.output.parent.mkdir(parents=True, exist_ok=True) group_path = spval_group(args.campaign) - digest = schema_digest(param_list) - plan = reconcile_plan(args.output, group_path, tiles, digest) - if not plan.empty(): - apply_plan(args.output, group_path, plan, tiles, cfc, params, digest) - print(f"[merge_final_cat] {plan.describe()} -> {args.output} " - f"({len(tiles)} tile(s), {len(param_list)} column(s), " - f"group {group_path})") - else: + digest = hdf5_reconcile.schema_digest(param_list) + + def read_tile(tile, path): + """One tile's requested columns, via create_final_cat.py's own reader.""" + extracted, dtype = cfc.read_data(str(path), params) + return cfc.copy_data(params["param_list"], extracted, dtype) + + todo = hdf5_reconcile.plan(args.output, group_path, tiles, digest) + if todo.empty(): print(f"[merge_final_cat] unchanged: {args.output} " f"({len(tiles)} tile(s))") + return + hdf5_reconcile.apply(args.output, group_path, todo, tiles, read_tile, + digest, "n_tiles") + print(f"[merge_final_cat] {todo.describe()} -> {args.output} " + f"({len(tiles)} tile(s), {len(param_list)} column(s), " + f"group {group_path})") if __name__ == "__main__": diff --git a/workflow/scripts/merge_star_cat.py b/workflow/scripts/merge_star_cat.py index 131758f57..25c9db0e1 100644 --- a/workflow/scripts/merge_star_cat.py +++ b/workflow/scripts/merge_star_cat.py @@ -1,37 +1,52 @@ #!/usr/bin/env python3 -"""Concatenate the campaign's per-CCD PSF validation catalogues into ONE full_starcat. +"""Collect the campaign's per-CCD PSF validation catalogues into ONE hdf5 file. Run as the shell of the campaign-level ``star_cat_merge`` rule, never by hand. -WHAT IT PRODUCES, AND FOR WHOM. ``/full_starcat-0000000.fits``: -every exposure's every CCD's ``validation_psf--.fits`` row, stacked, -with a ``CCD_NB`` column recording which CCD each row came from. It is the input -to the rho/tau statistics — sp_validation reads exactly this path -(``star_cat_path`` in its ``scripts/calibration/params.py``) and does no merging -of its own. Historically it was ``combine_runs.bash psf`` + a -``merge_starcat_runner`` pass; the workflow emitted neither, so the product set -was short one file. This script is that pass, driven by the DAG instead of by -bash. - -IT DOES NOT REIMPLEMENT THE COLUMN LIST. The stacking, the column names and the -CCD_NB parse all live in ``MergeStarCatPSFEX`` -(``shapepipe.modules.merge_starcat_package.merge_starcat``), which is what the -old runner called. This script only decides WHICH catalogues that class is -handed, and where the result lands. A column added to the module is a column -added here for free — which is the entire reason for the indirection. - -IT READS THE TARS, IT DOES NOT UNPACK THEM, AND IT STREAMS. ``exp_persist`` -packs each exposure's keepers into one uncompressed tar on the persistent root +WHAT IT PRODUCES, AND FOR WHOM. ``/full_starcat_.hdf5``: +one dataset per exposure at ``exposures/``, holding that exposure's every +CCD's ``validation_psf--.fits`` rows stacked, with a ``CCD_NB`` column +recording which CCD each row came from. It is the input to the rho/tau +statistics. Historically this was ``combine_runs.bash psf`` plus a +``merge_starcat_runner`` pass producing one flat FITS table, +``full_starcat-0000000.fits``, and sp_validation still opens that name today; +its readers move to this hdf5 under CosmoStat/sp_validation#340, the same +migration that retires the ``patches/`` key on the galaxy side. + +WHY HDF5, AND WHY ONE DATASET PER EXPOSURE. The campaign's two products should +behave the same way, and one flat table cannot: appending a tile meant +restacking every exposure the campaign had ever seen — ~40 GB of members at DR6 +scale to add ~2 MB. Per-exposure datasets make the file RECONCILABLE +(hdf5_reconcile.py carries that argument, and merge_final_cat.py is the same +machinery on the tile side), so an append reads the appended exposures and +nothing else while the file still cannot drift from its inputs. Memory follows: +one exposure at a time, not one campaign. + +NATIVE DTYPES. Columns are written as the validation_psf files store them — +float32 stays float32. The FITS writer this replaces widened every float column +to ``1D``, doubling both the file and the peak memory of the job that wrote it, +for no information. + +CCD_NB IS AN INTEGER. It is parsed out of the member name +(``validation_psf--.fits``), where it is always digits, so a string +buys nothing — and an int column costs 4 bytes a row against the 8 a +two-character fixed-width string does. + +IT READS THE TARS, IT DOES NOT UNPACK THEM. ``exp_persist`` packs each +exposure's keepers into one uncompressed tar on the persistent root (``/exp///psf/.tar``) precisely because inodes, -not bytes, bind on /project. Unpacking ~20k tars × ~40 members to merge them -would materialise ~800k files on the filesystem that design exists to protect, -and then delete them. So members are read out of the tars in memory -(``tarfile.extractfile(m).read()`` -> ``io.BytesIO``) and handed to the merge -class as ``[fileobj, member_name]`` pairs — ONE AT A TIME, lazily, through -``TarMembers`` below, because materialising them all first is ~40 GB at DR6 -scale. The member NAME is what the CCD_NB regex parses, which is why the pair -carries it; the class takes the name from the last element of the entry, so a -plain ``[path]`` entry behaves exactly as it always did. +not bytes, bind on /project. Unpacking ~20k tars x ~40 members to merge them +would materialise ~800k files on the filesystem that design exists to protect. +Members are read through the archive's own file object — seekable, the tar +being uncompressed by design — so the counting pass costs a header rather than +a member. + +THE OPTIONAL COLUMNS ARE A PER-FILE QUESTION. A pix2wcs-converted catalogue has +no MAG/SNR/ACCEPTED where an ordinary one does, and a campaign can hold both. +Deciding once for the merge is wrong in both directions: it either fails on the +first converted file or silently zeroes the real values of every ordinary one. +Each file is asked for its own schema, and only the files that lack a column are +zero-filled. WHICH EXPOSURES — AND WHY THE JOB DERIVES THE SET RATHER THAN BEING TOLD IT. The set is the CAMPAIGN's: every exposure read by a tile that is both declared @@ -42,99 +57,96 @@ class as ``[fileobj, member_name]`` pairs — ONE AT A TIME, lazily, through there is one query and not two that can drift), and then takes the exposures whose ``exp_persist`` manifest is on the persistent root. -It is derived rather than passed because at DR6 scale the set is ~20k paths, and +It is derived rather than passed because at DR6 scale the set is ~20k paths and a shell command reaches ``execve`` as a SINGLE argv entry capped at 128 KiB by -``MAX_ARG_STRLEN``. Passing them would be a job that dies before it starts. So -the rule's ``input`` is the DAG EDGE — what must exist before this runs — and -the rule's ``params`` carries a FINGERPRINT of that same list, which is what -makes the merge rerun when the set changes. - -THE TWO SETS ARE THE SAME SET, and that equality is the point of deriving it -this way rather than globbing the tree. The rule's input is ``star_cat_inputs()`` -(Snakefile): for each exposure of TILES_READY whose PSF products are on the -persistent root, an edge — the ``exp_persist`` manifest for a live exposure, the -TAR for one whose scratch store reclamation already took (that function argues -the asymmetry, which is about not rebuilding a reclaimed chain from VOS). -Nothing at all for an exposure reclaimed before ``exp_persist`` existed, which -left neither and is unrecoverable short of that rebuild. What this script -selects is the same rule stated from the job's side: same tiles, same index, -manifest present — and by the time the job runs, every exposure with an edge has -one. A glob over ``/exp`` would NOT be the same set: it would -sweep in exposures of an earlier, larger tile list sharing the products root, -stacking rows the fingerprint never saw and no rerun trigger would notice. +``MAX_ARG_STRLEN``. So the rule's ``input`` is the DAG EDGE — what must exist +before this runs — and its ``params`` carries a FINGERPRINT of the same set, +which is what makes the merge rerun when the set changes. A glob over +``/exp`` would NOT be the same set: it would sweep in exposures of +an earlier, larger tile list sharing the products root, stacking rows the +fingerprint never saw and no rerun trigger would notice. THE MANIFEST, NOT THE TAR, IS WHAT IT READS FIRST: the manifest records what was -actually packed, pattern by pattern, member by member, with sizes. Selecting -members from it means this script never guesses at tar contents, and an exposure -whose keep list did not include the validation catalogues contributes nothing -visibly rather than silently. - -BYTE-STABLE ON A NO-OP RERUN: written to a tmp path, compared, and moved only -if it differs (the pattern ``persist_exp.py`` and ``clean_exposure.py`` use). -An unconditional rewrite would move the output's mtime on every invocation. -Members are visited in sorted (exposure, member) order so the row order is a -function of the input set alone. - -PSFEX ONLY, DELIBERATELY. ``PSF_MODEL`` is ``psfex`` in every campaign the -workflow has run; ``MergeStarCatMCCD`` and ``MergeStarCatSetools`` exist beside -it and take the same constructor, so the hook is the one-line class choice in -``merge_class()`` below — an implementation, not a design, away. +actually packed, member by member, with sizes and the product each came from, so +this script never guesses at tar contents. """ import argparse -import filecmp -import io import json -import logging -import shutil import sys import tarfile -import tempfile from fnmatch import fnmatch from pathlib import Path -from shapepipe.modules.merge_starcat_package import merge_starcat +import numpy as np +from astropy.io import fits # Same directory; the rule invokes this file by path, so it is sys.path[0]. import build_index +import hdf5_reconcile import persist_exp -# The output name is not ours to choose: sp_validation hardcodes it -# (`star_cat_path = f"{data_dir}/full_starcat-0000000.fits"`), and -# MergeStarCatPSFEX writes exactly this basename into the output dir it is -# given. Kept here as the name this script promises to produce. -OUT_NAME = "full_starcat-0000000.fits" - # The members this merge consumes, named as the keep list names them and # resolved through the same catalogue persist_exp packs by — so the glob has one # definition and adding a product cannot leave the two disagreeing. They are # always there to find: persist_exp packs this product for every exposure # whatever `persist_exp:` says, and fails the pack rather than writing a # manifest without it. -MEMBER_PRODUCT = "psf_validation" +MEMBER_PRODUCT = persist_exp.ALWAYS MEMBER_PATTERN = persist_exp.resolve(MEMBER_PRODUCT) +# The group holding the per-exposure datasets. Unlike the galaxy side's +# `patches/`, this name is ours and says what it holds. +GROUP = "exposures" + +# The validation_psf table's HDU: what MergeStarCatPSFEX defaulted to and what +# psfex_interp writes — a SExtractor-style file, empty primary, header-carrying +# image extension, then the table. +HDU = 2 + +# The columns, in the order the FITS full_starcat carried them, which is the +# order every consumer has seen. The optional three are zero-filled per file. +COLUMNS = ("X", "Y", "RA", "DEC", + "HSM_G1_PSF", "HSM_G2_PSF", "HSM_T_PSF", + "HSM_G1_STAR", "HSM_G2_STAR", "HSM_T_STAR", + "HSM_FLAG_PSF", "HSM_FLAG_STAR") +OPTIONAL = ("MAG", "SNR", "ACCEPTED") +CCD_COLUMN = "CCD_NB" +ALL_COLUMNS = COLUMNS + OPTIONAL + (CCD_COLUMN,) -def merge_class(psf_model: str): - """The merge class for this PSF model — the one-line MCCD/setools hook. - Only psfex is exercised: it is what every campaign has run. MCCD reaches the - tars unchanged (it takes its CCD numbers from the data, and it now reports - by the entry's name like the others). SETOOLS would need one more thing — - it passes ``input_file_list[0][0]`` to file_io as a template path, which a - streamed entry is not — so wiring setools to this path is a change to that - class, not a change here. +def ccd_number(member_name: str) -> int: + """The CCD this member's rows belong to: ``validation_psf--.fits``. + + Always digits, which is why the column is an int; a member name that does + not carry one is a tar we do not understand, and saying so beats writing a + sentinel into the catalogue. + """ + ccd = member_name.rsplit(".", 1)[0].rsplit("-", 1)[-1] + if not ccd.isdigit(): + sys.exit(f"merge_star_cat: cannot read a CCD number out of member " + f"name {member_name!r}") + return int(ccd) + + +def is_member(entry: dict) -> bool: + """Is this manifest entry one of the members this merge reads? + + BY PRODUCT NAME, OR FAILING THAT BY FILE NAME. persist_exp records the + product every member came from and always packs psf_validation, so the name + is the answer for anything it writes today. The glob is the fallback, and it + earns its place twice over: a tar packed before the product field existed + has no label at all, and a keep list written as a raw glob + (`validation_psf-*.fits` rather than `psf_validation`) labels its members + with the glob. Neither should make the campaign's star catalogue silently + empty. """ - try: - return {"psfex": merge_starcat.MergeStarCatPSFEX, - "mccd": merge_starcat.MergeStarCatMCCD, - "setools": merge_starcat.MergeStarCatSetools}[psf_model] - except KeyError: - sys.exit(f"merge_star_cat: unknown psf_model {psf_model!r}") + return (entry.get("product") == MEMBER_PRODUCT + or fnmatch(entry["name"], MEMBER_PATTERN)) def manifests(products_dir: Path, tile_list: Path, index_db: Path) -> list: - """The campaign's exp_persist manifests that are on disk, in exposure order. + """``(exposure, manifest path)`` for the campaign's packed exposures. Not a glob over the products root: see the module docstring on why the set is the campaign's and not the filesystem's. @@ -144,73 +156,94 @@ def manifests(products_dir: Path, tile_list: Path, index_db: Path) -> list: path = (products_dir / "exp" / exp[:2] / exp / "manifests" / "exp_persist.json") if path.exists(): - out.append(path) + out.append((exp, path)) return out -def selection(manifest_paths: list, pattern: str) -> tuple: - """``[(tar path, [member names])]`` for the merge, and the empty exposures. +def tars(manifest_paths: list) -> tuple: + """``[(exposure, tar path)]`` for the merge, and the exposures with nothing. - Reads the manifests only. Every tar is checked for existence HERE, so a - products root missing a file fails before a single row is stacked rather - than an hour in. + Every tar is checked for existence HERE, so a products root missing a file + fails before a single row is read rather than an hour in. The tar is also + the unit's SOURCE for reconciling: its size and mtime are what a later + invocation compares against to decide whether this exposure changed. """ chosen, empty = [], [] - for man_path in manifest_paths: + for exp, man_path in manifest_paths: man = json.loads(man_path.read_text()) - wanted = sorted(f["name"] for f in man["files"] - if fnmatch(f["name"], pattern)) - if not wanted: - empty.append(man["unit"]) + if not any(is_member(f) for f in man["files"]): + empty.append(exp) continue tar_path = Path(man["tar"]) if not tar_path.exists(): sys.exit(f"merge_star_cat: {man_path} names a tar that is not " f"there: {tar_path}") - chosen.append((tar_path, wanted)) + chosen.append((exp, tar_path)) return chosen, empty -class TarMembers: - """The merge class's input list, materialised ONE TAR AT A TIME. - - ``MergeStarCatPSFEX`` wants something it can take the length of and iterate - once, handing it ``[fileobj, name]`` entries; it never indexes and never - rewinds. So it does not need a list, and a list is the one thing we cannot - afford: reading every member up front is the whole campaign in memory at - once — ~2 MB per exposure, so ~40 GB at DR6's ~20k exposures, against a - rule asking for 16 GB. Read lazily, peak memory is ONE member's bytes plus - the merge class's own accumulators, which are the real and unavoidable term. +def read_exposure(exp: str, tar_path: Path) -> np.ndarray: + """One exposure's every CCD, stacked, as a structured array. - ``__len__`` comes from the manifests, so the class can log the count before - a single tar is opened. - - IT IS ITERABLE MORE THAN ONCE, and must be: the merge makes two passes, one - for row counts from the headers and one to fill. Each ``__iter__`` opens the - archives afresh, so the second pass sees the same members in the same order. + TWO PASSES over the tar's members, and neither holds the exposure twice: + the first reads only each member's FITS HEADER — NAXIS2, the row count — + and the second allocates the columns once at their exact final length and + fills them slice by slice. Members are visited in sorted name order, so the + row order is a function of the tar's contents alone. """ - - def __init__(self, chosen): - self._chosen = chosen - - def __len__(self): - return sum(len(names) for _, names in self._chosen) - - def __iter__(self): - for tar_path, names in self._chosen: - with tarfile.open(tar_path) as tf: - for name in names: - member = tf.extractfile(name) - if member is None: - sys.exit(f"merge_star_cat: {tar_path} has no member " - f"{name}, which its manifest lists") - # The tar's own file object, not a BytesIO of the whole - # member: it is seekable (the archive is uncompressed by - # design) and astropy reads through it, so the merge's - # first pass costs a header rather than a member. The - # object is valid only until the next member is reached, - # which is exactly how the merge consumes it. - yield [member, name] + with tarfile.open(tar_path) as tf: + names = sorted(n for n in tf.getnames() + if Path(n).match(MEMBER_PATTERN)) + if not names: + sys.exit(f"merge_star_cat: {tar_path} holds no {MEMBER_PATTERN}") + + # --- pass 1: row counts and dtypes, from headers alone -------------- + counts, dtypes, opt_dtypes, n_total = [], None, {}, 0 + for name in names: + with fits.open(tf.extractfile(name), memmap=False, + ignore_missing_simple=True) as hdul: + hdu = hdul[HDU] + counts.append(hdu.header["NAXIS2"]) + # ColDefs.dtype describes the table without reading it. NOTE: + # it is the RAW storage dtype and ignores TSCAL/TZERO, so a + # scaled column would be allocated narrower than the values + # .data returns. Latent, not live: no validation_psf column is + # scaled. Read the dtype off .data if one ever is. + cols = hdu.columns.dtype + if dtypes is None: + dtypes = cols + for col in OPTIONAL: + if col not in opt_dtypes and col in (cols.names or ()): + opt_dtypes[col] = cols[col] + n_total += counts[-1] + + fields = [(c, dtypes[c]) for c in COLUMNS] + # A column no file of this exposure carries still gets a column, + # zero-filled, in the dtype the positional column X uses. + fields += [(c, opt_dtypes.get(c, dtypes["X"])) for c in OPTIONAL] + fields += [(CCD_COLUMN, np.int32)] + data = np.empty(n_total, dtype=np.dtype(fields)) + + # --- pass 2: fill --------------------------------------------------- + at = 0 + for name, n_rows in zip(names, counts): + with fits.open(tf.extractfile(name), memmap=False, + ignore_missing_simple=True) as hdul: + rows = hdul[HDU].data + have = set(rows.dtype.names or ()) + sl = slice(at, at + n_rows) + for col in COLUMNS: + data[col][sl] = rows[col] + for col in OPTIONAL: + # THIS file's schema, not the exposure's. + data[col][sl] = rows[col] if col in have else 0 + data[CCD_COLUMN][sl] = ccd_number(name) + at += n_rows + + if at != n_total: + raise ValueError(f"merge_star_cat: {tar_path}: pass 1 counted " + f"{n_total} rows, pass 2 filled {at}") + return data def main() -> None: @@ -222,59 +255,38 @@ def main() -> None: help="the campaign's tile list (config tile_list)") p.add_argument("--index-db", required=True, type=Path, help="the campaign's run index (config outputs.index_db)") - p.add_argument("--output", required=True, type=Path, - help=f"the merged catalogue; its basename is {OUT_NAME}") - p.add_argument("--psf-model", default="psfex") - p.add_argument("--pattern", default=MEMBER_PATTERN, - help="tar-member glob to merge; default %(default)s") + p.add_argument("--output", required=True, type=Path) + p.add_argument("--campaign", required=True, + help="named in the log; the group name is fixed") args = p.parse_args() - if args.output.name != OUT_NAME: - # The merge class writes OUT_NAME into a directory it is handed; a - # differently-named declared output would silently never be produced. - sys.exit(f"merge_star_cat: --output must be named {OUT_NAME} " - f"(got {args.output.name})") - - log = logging.getLogger("merge_star_cat") - logging.basicConfig(format="[merge_star_cat] %(message)s", - level=logging.INFO, stream=sys.stdout) - manifest_paths = manifests(args.products_dir, args.tile_list, args.index_db) - chosen, empty = selection(manifest_paths, args.pattern) - file_list = TarMembers(chosen) - if not len(file_list): + chosen, empty = tars(manifest_paths) + if not chosen: # Not a no-op: an empty star catalogue would pass every downstream # existence check and produce meaningless rho statistics. - sys.exit(f"merge_star_cat: no member matched {args.pattern!r} in any " - f"of {len(manifest_paths)} exp_persist manifest(s) for this " + sys.exit(f"merge_star_cat: no {MEMBER_PRODUCT} member in any of " + f"{len(manifest_paths)} exp_persist manifest(s) for this " f"campaign. persist_exp packs {MEMBER_PRODUCT} for every " f"exposure, so this means the manifests are not what we think " f"they are.") if empty: - log.info(f"{len(empty)} exposure(s) persisted no {args.pattern}: " - f"{', '.join(sorted(empty)[:5])}" - f"{' ...' if len(empty) > 5 else ''}") - - # tmp-then-cmp-then-mv. The merge class chooses its own basename inside the - # directory it is given, so the tmp is a DIRECTORY, not a file path, and it - # never outlives this process — an orphan on /project is an inode nothing - # revisits. + print(f"[merge_star_cat] {len(empty)} exposure(s) persisted no " + f"{MEMBER_PRODUCT}: {', '.join(sorted(empty)[:5])}" + f"{' ...' if len(empty) > 5 else ''}") + args.output.parent.mkdir(parents=True, exist_ok=True) - tmp_dir = Path(tempfile.mkdtemp(dir=args.output.parent, - prefix=".star_cat_merge.")) - try: - merge_class(args.psf_model)(file_list, str(tmp_dir), log).process() - tmp = tmp_dir / OUT_NAME - if not tmp.exists(): - sys.exit(f"merge_star_cat: the merge wrote no {OUT_NAME}") - if args.output.exists() and filecmp.cmp(tmp, args.output, shallow=False): - log.info(f"unchanged: {args.output}") - else: - tmp.replace(args.output) # atomic: same filesystem - log.info(f"{len(file_list)} catalogue(s) from {len(chosen)} " - f"exposure(s) -> {args.output}") - finally: - shutil.rmtree(tmp_dir, ignore_errors=True) + digest = hdf5_reconcile.schema_digest(ALL_COLUMNS) + todo = hdf5_reconcile.plan(args.output, GROUP, chosen, digest) + if todo.empty(): + print(f"[merge_star_cat] unchanged: {args.output} " + f"({len(chosen)} exposure(s))") + return + hdf5_reconcile.apply(args.output, GROUP, todo, chosen, read_exposure, + digest, "n_exposures") + print(f"[merge_star_cat] {todo.describe()} -> {args.output} " + f"({len(chosen)} exposure(s), {len(ALL_COLUMNS)} column(s), " + f"campaign {args.campaign})") if __name__ == "__main__": From 4ed9fb0d8ccb0f5912513519678e3458691ea354 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 20:27:06 -0400 Subject: [PATCH 19/85] fix(orchestration): nine findings from the third review MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit path_hash KILLED EVERY INVOCATION over a file one rule needs. It runs at module level, so a snapshot without scripts/python/ raised FileNotFoundError during the parse — taking `sp --unlock`, `sp report` and every dry run with it, and taking them with a bare traceback rather than the diagnosis merge_final_cat.py already carries for exactly this case, which the job could never reach. The hash degrades to a sentinel and warns once; the parse survives, the rule still exists, and the job prints the message written for it. Verified both halves with the file moved aside. THE STAR MERGE'S OPTIONAL COLUMNS HAD NO CANONICAL DTYPE. MAG/SNR/ACCEPTED took the dtype of whichever file carried them, falling back to X's float32 when none did — so an exposure whose files are all pix2wcs-converted got a float32 ACCEPTED while its neighbours got int32, and datasets under exposures/ differed in dtype. np.concatenate refuses that, and no digest can repair it because nothing about the schema CHANGED. The three are pinned (int32, float32, float32) and cast. Verified on two exposures, one carrying them and one not: identical dtypes, concatenate works. A MISLABELED MANIFEST ABORTED THE CAMPAIGN. is_member() accepts a member by product label OR by file name — right for "did this exposure keep the product" — but read_exposure() selects by name alone, so an entry labelled psf_validation whose name did not match put the exposure in the merge and then killed the whole job when the tar held nothing selectable. Membership is now the name on both sides, with one line saying an entry was labelled and skipped. RENAMING `campaign:` WOULD HAVE HALF-UPDATED THE FILE. The tile hdf5 carries the campaign in its GROUP, so a rename pointed the rule at a new group inside the same file: a second group beside the first, the first frozen and stale, and n_tiles describing one of them. One file is one campaign — apply refuses and names what is already there. APPEND IS CHEAP IN READS, NOT IN WRITES, and the docstrings said otherwise. The existing file is copied so the result can be moved into place atomically: one pass over it and, briefly, twice its size on disk. Corrected, and apply now refuses when the filesystem cannot hold it rather than filling /project and leaving a truncated tmp beside a catalogue people trust. A CORRUPT TAR RAISED A RAW ReadError, on both sides. persist_exp now refuses to write a new tar and says the old one is untouched and may hold products nothing else has; merge_star_cat names the tar and says not to delete it. THE 16-COLUMN SCHEMA IS DEFINED TWICE and nothing held the two together. MergeStarCatPSFEX writes the flat FITS table the module runner emits; merge_star_cat.py writes the hdf5. Separate implementations are right — only one of them reads tars, keeps native dtypes and reconciles — but a column added to one writer would simply be missing from the other's product, found by whoever next computed rho statistics from the wrong one. tests/unit/test_star_cat_columns.py asserts the names and their order agree; verified passing, and verified failing when one list is changed. Also noted where it lives: adding a retention product re-packs the tar and moves its mtime, so the star merge refreshes those exposures although their validation members are byte-for-byte unchanged — seconds per exposure against per-member bookkeeping on every exposure, which is not a trade worth making. Stale docs updated to the hdf5 product: the Snakefile's merges header, the README's star_cat_merge paragraph and scripts list (the hdf5 paragraph written last round never landed — its edit script aborted before writing), and config.yaml's psf_validation block. The README now says there are two writers and names the test that keeps their schema together. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- tests/unit/test_star_cat_columns.py | 95 +++++++++++++++++++++++++++++ workflow/README.md | 27 +++++--- workflow/Snakefile | 25 ++++++-- workflow/config.yaml | 2 +- workflow/scripts/hdf5_reconcile.py | 63 ++++++++++++++++++- workflow/scripts/merge_star_cat.py | 60 +++++++++++++----- workflow/scripts/persist_exp.py | 11 +++- 7 files changed, 253 insertions(+), 30 deletions(-) create mode 100644 tests/unit/test_star_cat_columns.py diff --git a/tests/unit/test_star_cat_columns.py b/tests/unit/test_star_cat_columns.py new file mode 100644 index 000000000..20975815d --- /dev/null +++ b/tests/unit/test_star_cat_columns.py @@ -0,0 +1,95 @@ +"""The star catalogue's 16 columns are defined twice, and must not drift. + +Two writers emit a full_starcat, for two consumers that have to agree about it: + + * ``MergeStarCatPSFEX`` (``src/shapepipe/modules/merge_starcat_package``), + which the ``merge_starcat`` MODULE RUNNER calls, writing the flat FITS table + sp_validation opens today; + * ``workflow/scripts/merge_star_cat.py``, the Snakemake workflow's + ``star_cat_merge`` rule, writing the per-exposure hdf5 that replaces it + (CosmoStat/sp_validation#340 moves the readers). + +They were one definition until the workflow stopped calling the module class: +the rule reads validation_psf members out of the per-exposure tars, keeps their +native dtypes and reconciles its output, none of which the class does or should +do. Two implementations is the right answer for the behaviour; two COLUMN LISTS +is not, and nothing else would notice them diverging — a column added to one +writer would simply be absent from the other's product, discovered by whoever +next tried to compute rho statistics from the wrong one. + +Hence this module, which asserts the one thing they must share. It does NOT +assert the dtypes: the whole point of the hdf5 writer is that they differ (the +FITS one widens every float to 1D). Only the names, and their order. + +Deliberately import-light on the workflow side: merge_star_cat.py pulls in h5py +and astropy, which the class does too, so a container-free run is not on offer +here and is not worth contorting for. +""" + +import importlib.util +import sys +from pathlib import Path + +import pytest + + +REPO_ROOT = Path(__file__).resolve().parents[2] +SCRIPTS = REPO_ROOT / "workflow" / "scripts" +SCRIPT = SCRIPTS / "merge_star_cat.py" + + +def _load_workflow_merge(): + """Import the rule's script by path — ``workflow/scripts`` is not a package. + + Its own imports (build_index, hdf5_reconcile, persist_exp) are siblings it + reaches through ``sys.path[0]``, which is how the rule invokes it, so the + directory goes on the path here too. + """ + assert SCRIPT.exists(), f"{SCRIPT} not found; the rule calls it by path" + sys.path.insert(0, str(SCRIPTS)) + try: + spec = importlib.util.spec_from_file_location("_merge_star_cat", SCRIPT) + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + finally: + sys.path.remove(str(SCRIPTS)) + return module + + +@pytest.fixture(scope="module") +def writers(): + """The two column lists, each in the order its writer emits them.""" + h5py = pytest.importorskip("h5py") # noqa: F841 - workflow dep + pytest.importorskip("astropy") + workflow = _load_workflow_merge() + from shapepipe.modules.merge_starcat_package.merge_starcat import ( + MergeStarCatPSFEX, + ) + # The class carries (output name, source column) pairs plus its optional + # set and appends CCD_NB last; the script carries output names throughout. + module_columns = ( + tuple(out for out, _ in MergeStarCatPSFEX._COLUMNS) + + tuple(out for out, _ in MergeStarCatPSFEX._OPTIONAL) + + ("CCD_NB",) + ) + return module_columns, tuple(workflow.ALL_COLUMNS) + + +def test_column_names_and_order_agree(writers): + """Same names, same order — the schema both products promise.""" + module_columns, workflow_columns = writers + assert workflow_columns == module_columns + + +def test_sixteen_columns(writers): + """The count is itself the documented contract (README, config.yaml).""" + module_columns, workflow_columns = writers + assert len(module_columns) == 16 + assert len(workflow_columns) == 16 + + +def test_ccd_nb_is_last(writers): + """CCD_NB is appended per input file rather than read from one, in both.""" + module_columns, workflow_columns = writers + assert module_columns[-1] == "CCD_NB" + assert workflow_columns[-1] == "CCD_NB" diff --git a/workflow/README.md b/workflow/README.md index 40371f2b4..c00d66c18 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -167,7 +167,8 @@ workflow/ run_report.py standalone report (NOT a DAG node; run_report hooks call it) container.py image layers + the resolution order behind `sp container` (stdlib-only) persist_exp.py ONE exposure's keepable PSF products -> one tar on products_dir (the exp_persist rule) - merge_star_cat.py ALL exposures' validation_psf, read out of the tars -> full_starcat (the star_cat_merge rule) + hdf5_reconcile.py bring an hdf5 catalogue into agreement with a campaign (shared by both merges) + merge_star_cat.py ALL exposures' validation_psf, out of the tars -> full_starcat_.hdf5 merge_final_cat.py ALL tiles' final_cat -> final_cat_.hdf5 (the final_cat_merge rule) clean_exposure.py ONE exposure's store + manifests + logs -> tombstone (the clean_exposure rule) profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; keep-going @@ -297,13 +298,23 @@ profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; kee Everything above is per unit; the two products downstream analysis actually opens are per *campaign*, and until these rules existed each was a manual pass after the run. - `star_cat_merge` stacks every exposure's every CCD's `psf_validation` into one - `/full_starcat-0000000.fits` — the rho/tau statistics input, at - the path sp_validation hardcodes. It reads the members straight out of the - per-exposure tars (`tarfile`; unpacking ~800k files to merge them would defeat - the tar's whole purpose) and stacks them with `MergeStarCatPSFEX`, the same - class the old `merge_starcat_runner` called, so the column list has exactly - one definition. It exists whenever the campaign has a persisted exposure. + `star_cat_merge` collects every exposure's every CCD's `psf_validation` into + `/full_starcat_.hdf5`, one dataset per exposure at + `exposures/` — the rho/tau statistics input. It reads the members + straight out of the per-exposure tars (`tarfile`; unpacking ~800k files to + merge them would defeat the tar's whole purpose), keeps their native dtypes, + and stores `CCD_NB` as an int. sp_validation still opens the old flat FITS + name, `full_starcat-0000000.fits`; its readers move to this file under + [sp_validation#340](https://github.com/CosmoStat/sp_validation/issues/340), + the same migration that retires the `patches/` key on the tile side. The rule + exists whenever the campaign has a persisted exposure. + **Two writers, one schema.** The module runner still emits the flat FITS + table through `MergeStarCatPSFEX`, and this rule emits the hdf5; they are + separate implementations on purpose, because only one of them reads tars, + keeps native dtypes and reconciles. Their 16 COLUMN NAMES must not drift + apart, and nothing else would notice if they did — a column added to one + writer would just be missing from the other's product. `tests/unit/` + `test_star_cat_columns.py` is what holds them together. `final_cat_merge` collects every ready tile's `final_cat-.fits` into `/final_cat_.hdf5`: one dataset per tile under a group named for the campaign, the `final_cat.param` columns, an `n_tiles` attribute. diff --git a/workflow/Snakefile b/workflow/Snakefile index ea4129f25..4317e871f 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -353,8 +353,25 @@ def path_hash(path): workflow/scripts/ nor a declared input. Without them in the hash, this PR's own edits to both would have left every finished campaign's hdf5 untouched and nothing would have said so. + + A MISSING FILE IS NOT A PARSE ERROR. This runs at module level, so raising + here kills EVERY invocation of the workflow — `sp --unlock`, `sp report`, + a dry run — over a file only one rule needs. Worse, it kills them with a + bare FileNotFoundError, which is exactly the diagnosis merge_final_cat.py + already carries and would print if the job were allowed to reach it. So the + hash degrades to a sentinel and says so once; the parse survives, the rule + still exists, and the job fails with the message written for it. """ - return hashlib.md5(Path(path).read_bytes()).hexdigest()[:12] + path = Path(path) + try: + return hashlib.md5(path.read_bytes()).hexdigest()[:12] + except OSError: + if workflow.is_main_process: + logger.warning( + f"missing: {path} — it is part of a rule's rerun trigger, so " + f"that rule cannot tell whether it is out of date. The job " + f"that needs the file will say so when it runs.") + return "missing" SCRIPT_HASH = script_hash("completeness.py") FOREST_HASH = script_hash("build_forest.py") @@ -596,9 +613,9 @@ def exp_store_reclaimed(exp): # --- the campaign-level merges --------------------------------------------- # Two rules, one job each per campaign, both writing to the persistent root, and # both the LAST link of a chain whose per-unit half the workflow already had: -# the exposure side ends in one `full_starcat-0000000.fits` (every CCD's PSF -# validation catalogue, stacked — the rho/tau statistics input) and the tile side -# in one `final_cat_.hdf5` (every tile's final catalogue — the shear +# the exposure side ends in one `full_starcat_.hdf5` (every CCD's PSF +# validation catalogue — the rho/tau statistics input) and the tile side in one +# `final_cat_.hdf5` (every tile's final catalogue — the shear # catalogue sp_validation reads). Until they existed the workflow's product set # was two files short of what the old `combine_runs.bash` + `create_final_cat.py` # chain delivered, and every campaign ended with a manual merge. diff --git a/workflow/config.yaml b/workflow/config.yaml index 99ae1b560..38979b2d0 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -87,7 +87,7 @@ outputs: # # WHAT IS ALWAYS KEPT, AND IS NOT A CHOICE HERE: psf_validation, the psfex_interp # validation catalogue, one per CCD. `star_cat_merge` stacks every one of them -# into the campaign's /full_starcat-0000000.fits, so they are that +# into /full_starcat_.hdf5, so they are that # catalogue's PROVENANCE — a merged star catalogue with no per-exposure inputs # beside it cannot be audited, re-cut, or recomputed after a purge — and they are # what keeps APPENDING TILES CHEAP, since a tile added next month brings diff --git a/workflow/scripts/hdf5_reconcile.py b/workflow/scripts/hdf5_reconcile.py index 0635b5976..f2e3ff325 100644 --- a/workflow/scripts/hdf5_reconcile.py +++ b/workflow/scripts/hdf5_reconcile.py @@ -23,7 +23,12 @@ file's root. * a dataset that agrees with its source and its schema is left alone, unread. -An append therefore reads exactly the appended units. +An append therefore READS exactly the appended units. It still WRITES the whole +file: the existing one is copied so the result can be moved into place +atomically, which costs one pass over it and, briefly, twice its size on disk. +That is the cheap half by orders of magnitude — copying a 1 GB hdf5 against +re-reading 800 GB of catalogues — but it is not free, and `apply` refuses rather +than filling the filesystem when the free space is not there. WHAT IS AND IS NOT A FUNCTION OF THE INPUT SET. The file's CONTENT is: the same units with the same sources give the same datasets, the same columns and the @@ -41,6 +46,7 @@ import hashlib import shutil +import sys from pathlib import Path import h5py @@ -106,6 +112,56 @@ def plan(output: Path, group_path: str, units: list, digest: str) -> Plan: return Plan(add, refresh, sorted(present - want)) +# Twice the file, plus a tenth of it again: the copy and the original coexist, +# and hdf5 is not a format to run to the last byte of a filesystem on. +FREE_SPACE_MARGIN = 2.1 + + +def check_free_space(output: Path) -> None: + """Refuse to start a rewrite the filesystem cannot hold. + + A merge that fills /project does not just fail: it fails everything else + writing there at the same time, and it can leave a truncated tmp beside a + catalogue people trust. Cheaper to say so first. + """ + if not output.exists(): + return + size = output.stat().st_size + free = shutil.disk_usage(output.parent).free + if free < size * FREE_SPACE_MARGIN: + sys.exit( + f"hdf5_reconcile: {output.parent} has {free / 1e9:.1f} GB free and " + f"this merge needs about {size * FREE_SPACE_MARGIN / 1e9:.1f} GB — " + f"it rewrites {output.name} ({size / 1e9:.1f} GB) through a tmp " + f"copy beside it. Free space or move products_dir; the existing " + f"catalogue is untouched.") + + +def check_sole_group(output: Path, group_path: str) -> None: + """One file, one campaign — refuse to half-update a file holding two. + + Renaming `campaign:` mid-flight points the rule at a NEW group inside the + SAME file (the path carries the campaign only on the tile side, where the + group does). Reconciling would then add a second group beside the first, + leave the first frozen and stale, and set a count attribute describing only + one of them. Nothing downstream reads such a file correctly, and no rule + here means to produce one. Say what is there and stop. + """ + if not output.exists() or "/" not in group_path: + return + parent, leaf = group_path.rsplit("/", 1) + with h5py.File(output, "r") as f: + if parent not in f: + return + others = sorted(k for k in f[parent] if k != leaf) + if others: + sys.exit( + f"hdf5_reconcile: {output} already holds {parent}/" + f"{', '.join(others)} beside {group_path}. One file is one " + f"campaign: reconciling would freeze the other group and count " + f"only this one. Point `campaign:` back, or write to a new path.") + + def apply(output: Path, group_path: str, todo: Plan, units: list, read, digest: str, count_attr: str) -> None: """Carry the plan out on a tmp file, then move it into place. @@ -120,7 +176,8 @@ def apply(output: Path, group_path: str, todo: Plan, units: list, read, refresh of a unit leaks that unit. So: * a plan that only ADDS copies the existing file and appends to it. There - is nothing to reclaim, and copying beats rewriting. + is nothing to reclaim, and copying beats rewriting. It is still a pass + over the whole file — an append is cheap in READS, not in writes. * a plan that removes or refreshes anything builds the tmp FRESH, moving the datasets it keeps across with h5py's own group copy — a dataset-level copy inside the library that never reads a row into numpy @@ -135,6 +192,8 @@ def apply(output: Path, group_path: str, todo: Plan, units: list, read, """ sources = dict(units) rewrite = bool(todo.remove or todo.refresh) + check_free_space(output) + check_sole_group(output, group_path) written = set(todo.add) | set(todo.refresh) keep = [u for u, _ in units if u not in written] tmp = output.with_name(output.name + ".tmp") diff --git a/workflow/scripts/merge_star_cat.py b/workflow/scripts/merge_star_cat.py index 25c9db0e1..84ba633d1 100644 --- a/workflow/scripts/merge_star_cat.py +++ b/workflow/scripts/merge_star_cat.py @@ -110,9 +110,16 @@ "HSM_G1_PSF", "HSM_G2_PSF", "HSM_T_PSF", "HSM_G1_STAR", "HSM_G2_STAR", "HSM_T_STAR", "HSM_FLAG_PSF", "HSM_FLAG_STAR") -OPTIONAL = ("MAG", "SNR", "ACCEPTED") +# CANONICAL DTYPES, not whatever the first file that carries the column happens +# to use. These three are absent from pix2wcs-converted catalogues, so an +# exposure whose files all lack them would otherwise be allocated a fallback +# dtype while its neighbours got the real one — and datasets under exposures/* +# would then differ in dtype, which np.concatenate refuses and no digest can +# repair, since nothing about the schema CHANGED. Pinning the dtype here is what +# makes every exposure's dataset the same shape whatever its files carry. +OPTIONAL = {"MAG": np.float32, "SNR": np.float32, "ACCEPTED": np.int32} CCD_COLUMN = "CCD_NB" -ALL_COLUMNS = COLUMNS + OPTIONAL + (CCD_COLUMN,) +ALL_COLUMNS = COLUMNS + tuple(OPTIONAL) + (CCD_COLUMN,) def ccd_number(member_name: str) -> int: @@ -171,7 +178,21 @@ def tars(manifest_paths: list) -> tuple: chosen, empty = [], [] for exp, man_path in manifest_paths: man = json.loads(man_path.read_text()) - if not any(is_member(f) for f in man["files"]): + # MEMBERSHIP IS THE MEMBER NAME, and only the member name. is_member() + # will also accept a manifest's own product LABEL, which is the right + # test for "did this exposure keep the product" — but a label is not + # what read_exposure() selects on, and a mislabeled entry whose name + # does not match would put this exposure in the merge and then abort + # the whole campaign when the tar turned out to hold nothing selectable. + # So the two agree by construction: both ask the name. + if not any(fnmatch(f["name"], MEMBER_PATTERN) for f in man["files"]): + if any(is_member(f) for f in man["files"]): + # Labelled as the product, named as something else. Worth one + # line — it means a manifest we did not write, or a keep list + # whose glob does not match the member it matched. + print(f"[merge_star_cat] {exp}: manifest labels a " + f"{MEMBER_PRODUCT} member whose name does not match " + f"{MEMBER_PATTERN}; not merging it") empty.append(exp) continue tar_path = Path(man["tar"]) @@ -190,15 +211,29 @@ def read_exposure(exp: str, tar_path: Path) -> np.ndarray: and the second allocates the columns once at their exact final length and fills them slice by slice. Members are visited in sorted name order, so the row order is a function of the tar's contents alone. + + NOTE ON WHEN THIS IS CALLED AGAIN. The unit's source is the TAR, so adding a + retention product re-packs it, moves its mtime, and refreshes this exposure + even though its validation members are byte-for-byte what they were. Reading + one exposure is seconds and the alternative — stamping the members rather + than the archive — buys a rarely-taken shortcut for a per-member bookkeeping + cost on every exposure. Not worth it. """ - with tarfile.open(tar_path) as tf: + try: + tf = tarfile.open(tar_path) + except tarfile.TarError as exc: + sys.exit(f"merge_star_cat: cannot read {tar_path}: {exc}. That tar is " + f"this exposure's only copy of its PSF products — do not " + f"delete it; re-pack the exposure if its scratch store is " + f"still there, and treat the exposure as lost if it is not.") + with tf: names = sorted(n for n in tf.getnames() - if Path(n).match(MEMBER_PATTERN)) + if fnmatch(n, MEMBER_PATTERN)) if not names: sys.exit(f"merge_star_cat: {tar_path} holds no {MEMBER_PATTERN}") # --- pass 1: row counts and dtypes, from headers alone -------------- - counts, dtypes, opt_dtypes, n_total = [], None, {}, 0 + counts, dtypes, n_total = [], None, 0 for name in names: with fits.open(tf.extractfile(name), memmap=False, ignore_missing_simple=True) as hdul: @@ -209,18 +244,15 @@ def read_exposure(exp: str, tar_path: Path) -> np.ndarray: # scaled column would be allocated narrower than the values # .data returns. Latent, not live: no validation_psf column is # scaled. Read the dtype off .data if one ever is. - cols = hdu.columns.dtype if dtypes is None: - dtypes = cols - for col in OPTIONAL: - if col not in opt_dtypes and col in (cols.names or ()): - opt_dtypes[col] = cols[col] + dtypes = hdu.columns.dtype n_total += counts[-1] fields = [(c, dtypes[c]) for c in COLUMNS] - # A column no file of this exposure carries still gets a column, - # zero-filled, in the dtype the positional column X uses. - fields += [(c, opt_dtypes.get(c, dtypes["X"])) for c in OPTIONAL] + # The optional three take their CANONICAL dtype, not one file's (see + # OPTIONAL): every exposure's dataset must have the same dtype whether + # or not its files carry the column. + fields += list(OPTIONAL.items()) fields += [(CCD_COLUMN, np.int32)] data = np.empty(n_total, dtype=np.dtype(fields)) diff --git a/workflow/scripts/persist_exp.py b/workflow/scripts/persist_exp.py index 24e0c1d66..bab1a5f0a 100644 --- a/workflow/scripts/persist_exp.py +++ b/workflow/scripts/persist_exp.py @@ -348,7 +348,15 @@ def main() -> None: for f in json.loads(prior.read_text())["files"]} except (OSError, ValueError, KeyError): pass # a damaged manifest loses only labels - with tarfile.open(tar_path) as tf: + try: + old_read = tarfile.open(tar_path) + except tarfile.TarError as exc: + sys.exit(f"persist_exp: {args.exp}: cannot read the existing " + f"{tar_path}: {exc}. Refusing to write a new one — the " + f"old tar is left exactly as it is, and it may still hold " + f"products nothing else has. Move it aside deliberately " + f"if you have decided it is lost.") + with old_read as tf: for ti in tf.getmembers(): if ti.name in seen or not ti.isfile(): continue # a live source supersedes it @@ -378,6 +386,7 @@ def anonymous(ti: tarfile.TarInfo) -> tarfile.TarInfo: # byte-for-byte what it was and a rerun that changes nothing still # produces an identical archive. with tarfile.open(tmp, "w", format=tarfile.PAX_FORMAT) as tf: + # Already proven readable above, where the members were listed. old_tar = (tarfile.open(tar_path) if carried else None) try: for f in files: From e86d8c8d1149abab965d26976be284e63dae9044 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 20:41:35 -0400 Subject: [PATCH 20/85] test(unit): property-based state machines for reconcile and persist_exp MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two rules carry state across invocations and are correct only over SEQUENCES: hdf5_reconcile brings a catalogue into agreement with a campaign that changes under it, and persist_exp packs a tar whose existing members are a floor. Neither is a claim the example tests can finish making, so each gets a hypothesis state machine that walks a random sequence of campaign edits and asserts the model after every one. hdf5_reconcile: units added, refreshed and removed, and the column set flipped, against a real file. After every step the datasets are the campaign's units with their sources' content, the count and digest attributes agree, every dataset shares one dtype, and a no-op leaves the mtime alone. Source mtimes are set explicitly, so a same-size rewrite inside one filesystem tick cannot masquerade as a refresh. Compaction is asserted where the module actually claims it — the rebuild path — and stated as "does not grow with history": a rebuild costs ~1.4 kB more than a from-scratch build (h5py's group copy writes more metadata than create_dataset does) and that overhead is constant, which test_repeated_refresh_does_not_grow_the_file pins directly. Plus a crash injected at the rename, which must leave the previous file byte-identical. persist_exp: random keep lists of product names, raw globs and overlapping mixtures over a store that gains and loses products. Members are additive across packs, never duplicated, and the manifest agrees with the tar down to the product labels; a missing psf_validation fails without writing a manifest, an unknown product name is refused before any work, a corrupt tar is left exactly as it is, and two different sources with one member name are still fatal. Both files were checked against five mutants (never rebuild, never refresh, non-additive retention, optional psf_validation, tolerated collision); each is caught. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar --- tests/unit/test_hdf5_reconcile_props.py | 336 ++++++++++++++++++++++ tests/unit/test_persist_exp_props.py | 360 ++++++++++++++++++++++++ 2 files changed, 696 insertions(+) create mode 100644 tests/unit/test_hdf5_reconcile_props.py create mode 100644 tests/unit/test_persist_exp_props.py diff --git a/tests/unit/test_hdf5_reconcile_props.py b/tests/unit/test_hdf5_reconcile_props.py new file mode 100644 index 000000000..a990c3f89 --- /dev/null +++ b/tests/unit/test_hdf5_reconcile_props.py @@ -0,0 +1,336 @@ +"""Property-based state machine over ``workflow/scripts/hdf5_reconcile.py``. + +The module's contract is that an hdf5 catalogue reconciled against a campaign +is a FUNCTION OF ITS INPUT SET — the same units with the same sources give the +same datasets, the same dtypes and the same count attribute, however they got +there. That is a claim about every reachable sequence of appends, refreshes and +removals, not about the three the unit tests happen to walk, so it is tested +here against a model: a random sequence of campaign edits, each followed by a +real plan/apply against a real file on disk, with the model asserted after +every step. + +The operations are the four things a campaign can do between invocations — +add a unit, change a unit's source, drop a unit, change the column set — plus +a no-op, which is the one that must leave the file's mtime alone. + +Source mtimes are set EXPLICITLY with ``os.utime`` rather than left to the +clock. ``stamp()`` is (size, mtime_ns), so a test that rewrote a file with the +same length inside one filesystem tick would silently exercise "nothing +changed" while believing it exercised a refresh. +""" + +import importlib.util +import os +import sys +from pathlib import Path + +import numpy as np +import pytest +from hypothesis import HealthCheck, settings +from hypothesis import strategies as st +from hypothesis.stateful import ( + RuleBasedStateMachine, + initialize, + invariant, + precondition, + rule, +) + +h5py = pytest.importorskip("h5py") + +REPO_ROOT = Path(__file__).resolve().parents[2] +SCRIPTS = REPO_ROOT / "workflow" / "scripts" + + +def _load(name): + path = SCRIPTS / f"{name}.py" + assert path.exists(), f"{path} not found; the rules call it by path" + sys.path.insert(0, str(SCRIPTS)) + try: + spec = importlib.util.spec_from_file_location(f"_{name}", path) + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + finally: + sys.path.remove(str(SCRIPTS)) + return module + + +reconcile = _load("hdf5_reconcile") + +GROUP = "cat/campaign_a" +COUNT_ATTR = "n_units" +UNITS = ["u0", "u1", "u2", "u3"] +# Two column sets, so a schema change is a real change of dtype and width. +COLUMN_SETS = [("RA", "DEC", "E1"), ("RA", "DEC", "E1", "FWHM")] +# What a rebuild may cost over a from-scratch build of the same campaign: the +# metadata h5py's group copy writes for a moved dataset. Measured at ~1.4 kB +# and constant in the number of rebuilds; the allowance is generous because the +# property being defended is "does not grow with history", not an exact size. +COPY_SLACK = 8192 + + +def _array(columns, rows, seed): + rng = np.random.default_rng(seed) + dtype = [(c, " None: + np.save(path, array, allow_pickle=False) + os.utime(path, ns=(mtime_ns, mtime_ns)) + + +def _read(unit, source): + return np.load(source, allow_pickle=False) + + +def _build(output: Path, sources: dict, columns): + """Plan and apply once, exactly as the two merge rules do; return the plan.""" + units = sorted(sources.items()) + digest = reconcile.schema_digest(columns) + todo = reconcile.plan(output, GROUP, units, digest) + if todo.empty(): + return todo + reconcile.apply(output, GROUP, todo, units, _read, digest, COUNT_ATTR) + return todo + + +class ReconcileMachine(RuleBasedStateMachine): + """A campaign that changes under a catalogue that must keep up with it.""" + + @initialize() + def setup(self): + self.dir = Path( + __import__("tempfile").mkdtemp(prefix="reconcile-props-") + ) + self.output = self.dir / "cat.h5" + self.columns = COLUMN_SETS[0] + self.sources = {} # unit -> source path + self.expected = {} # unit -> array as last written + self.clock = 1_000_000_000_000_000_000 + # Compaction is only claimed of the rebuild path (a plan that removes + # or refreshes). An add-only plan copies the file and appends, so its + # layout carries whatever the previous writes left behind. + self.rebuilt = False + + def teardown(self): + __import__("shutil").rmtree(self.dir, ignore_errors=True) + + # --- the campaign's moves ------------------------------------------- + def _tick(self): + self.clock += 1_000_000_000 + return self.clock + + def _step(self, changed): + before = (self.output.stat().st_mtime_ns + if self.output.exists() else None) + todo = _build(self.output, self.sources, self.columns) + self.rebuilt = bool(todo.remove or todo.refresh) + if not changed and before is not None: + assert self.output.stat().st_mtime_ns == before, ( + "a no-op reconcile rewrote the file; mtime is a rerun trigger" + ) + + @rule(pick=st.integers(0, 2**16), rows=st.integers(1, 5), + seed=st.integers(0, 2**16)) + @precondition(lambda self: len(self.sources) < len(UNITS)) + def add_unit(self, pick, rows, seed): + free = sorted(set(UNITS) - set(self.sources)) + unit = free[pick % len(free)] + path = self.dir / f"{unit}.npy" + array = _array(self.columns, rows, seed) + _write_source(path, array, self._tick()) + self.sources[unit] = path + self.expected[unit] = array + self._step(changed=True) + + @rule(pick=st.integers(0, 2**16), rows=st.integers(1, 5), + seed=st.integers(0, 2**16), resize=st.booleans()) + @precondition(lambda self: bool(self.sources)) + def modify_source(self, pick, rows, seed, resize): + unit = sorted(self.sources)[pick % len(self.sources)] + old = self.expected[unit] + rows = rows if resize else len(old) + array = _array(self.columns, rows, seed) + _write_source(self.sources[unit], array, self._tick()) + self.expected[unit] = array + self._step(changed=True) + + @rule(pick=st.integers(0, 2**16)) + @precondition(lambda self: bool(self.sources)) + def remove_unit(self, pick): + unit = sorted(self.sources)[pick % len(self.sources)] + self.sources.pop(unit).unlink() + self.expected.pop(unit) + self._step(changed=True) + + @rule() + def change_columns(self): + """Flip to the other column set — a digest change, so every unit + refreshes.""" + columns = next(c for c in COLUMN_SETS if c != self.columns) + self.columns = columns + # A schema change is a change to how the SOURCES are read, so the + # sources are rewritten under the new column set as the campaign would. + for i, (unit, path) in enumerate(sorted(self.sources.items())): + array = _array(columns, len(self.expected[unit]), 4242 + i) + _write_source(path, array, self._tick()) + self.expected[unit] = array + self._step(changed=True) + + @rule() + def no_op(self): + self._step(changed=False) + + # --- what must be true after every step ------------------------------ + @invariant() + def file_matches_campaign(self): + if not self.expected: + return + assert self.output.exists() + with h5py.File(self.output, "r") as f: + assert set(f[GROUP]) == set(self.expected), ( + "datasets and campaign units disagree") + assert f.attrs[COUNT_ATTR] == len(self.expected) + assert (f.attrs["param_digest"] + == reconcile.schema_digest(self.columns)) + dtypes = set() + for unit, want in self.expected.items(): + got = f[GROUP][unit][...] + assert got.dtype.names == want.dtype.names + np.testing.assert_array_equal(got, want) + dtypes.add(got.dtype) + stamp = reconcile.stamp(self.sources[unit]) + assert (int(f[GROUP][unit].attrs["src_bytes"]), + int(f[GROUP][unit].attrs["src_mtime_ns"])) == stamp + assert len(dtypes) == 1, ( + "sources share a column list; datasets must share a dtype") + + @invariant() + def compact(self): + """A rebuild does not carry the old file's dead space forward. + + HDF5 never reclaims a deleted dataset's space, which is why ``apply`` + builds the tmp FRESH whenever a plan removes or refreshes anything + instead of copying and editing in place. If that path stopped firing, + a long-lived campaign would grow by one unit per refresh forever. + + The bound is a from-scratch build of the same campaign plus a fixed + allowance: moving a dataset across with h5py's group copy costs a + little more metadata than creating it from an array does, measured at + ~1.4 kB here and — see the cycle test below — independent of how many + times the file has been rebuilt. What must never hold is growth that + tracks the history. + """ + if not self.rebuilt or not self.output.exists(): + return + fresh = self.dir / "fresh.h5" + fresh.unlink(missing_ok=True) + try: + _build(fresh, self.sources, self.columns) + if not fresh.exists(): + return + assert (self.output.stat().st_size + <= fresh.stat().st_size + COPY_SLACK), ( + "a rebuilt file is carrying dead space: " + f"{self.output.stat().st_size} bytes against " + f"{fresh.stat().st_size} from scratch") + finally: + fresh.unlink(missing_ok=True) + + +ReconcileMachine.TestCase.settings = settings( + max_examples=150, + stateful_step_count=14, + deadline=None, + suppress_health_check=[HealthCheck.too_slow, HealthCheck.data_too_large], +) +TestReconcileMachine = ReconcileMachine.TestCase + + +def test_crash_between_tmp_and_replace_leaves_the_file_untouched(): + """A failed rename must leave the previous catalogue byte-identical. + + Not a hypothesis case: the interesting axis is the crash point, and there + is one. ``os.replace`` is made to raise where the tmp is moved into place. + """ + import shutil + import tempfile + + work = Path(tempfile.mkdtemp(prefix="reconcile-crash-")) + try: + output = work / "cat.h5" + columns = COLUMN_SETS[0] + sources = {} + for i, unit in enumerate(UNITS[:2]): + path = work / f"{unit}.npy" + _write_source(path, _array(columns, 3, i), + 1_000_000_000_000_000_000 + i) + sources[unit] = path + _build(output, sources, columns) + before = output.read_bytes() + before_mtime = output.stat().st_mtime_ns + + # A third unit arrives, and the rename fails. + path = work / "u2.npy" + _write_source(path, _array(columns, 3, 99), 1_000_000_000_000_000_099) + sources["u2"] = path + units = sorted(sources.items()) + digest = reconcile.schema_digest(columns) + todo = reconcile.plan(output, GROUP, units, digest) + assert todo.add == ["u2"] + + real_replace = Path.replace + + def boom(self, target): + raise OSError("simulated crash between write and rename") + + Path.replace = boom + try: + with pytest.raises(OSError): + reconcile.apply(output, GROUP, todo, units, _read, digest, + COUNT_ATTR) + finally: + Path.replace = real_replace + + assert output.read_bytes() == before, "the old catalogue was modified" + assert output.stat().st_mtime_ns == before_mtime + assert not (work / "cat.h5.tmp").exists(), "tmp outlived the failure" + finally: + shutil.rmtree(work, ignore_errors=True) + + +def test_repeated_refresh_does_not_grow_the_file(): + """The leak the rebuild path exists to prevent, asserted directly. + + Twelve refreshes of one unit in a two-unit campaign. If ``apply`` ever + copied the file and edited it in place, each would strand the previous + dataset's bytes and the size would climb monotonically. + """ + import shutil + import tempfile + + work = Path(tempfile.mkdtemp(prefix="reconcile-growth-")) + try: + output = work / "cat.h5" + columns = COLUMN_SETS[0] + sources = {} + for i, unit in enumerate(("u0", "u1")): + path = work / f"{unit}.npy" + _write_source(path, _array(columns, 4, i), 10**18 + i) + sources[unit] = path + _build(output, sources, columns) + + sizes = [] + for k in range(12): + _write_source(sources["u0"], _array(columns, 4, 100 + k), + 10**18 + 100 + k) + todo = _build(output, sources, columns) + assert todo.refresh == ["u0"], todo.describe() + sizes.append(output.stat().st_size) + assert len(set(sizes)) == 1, f"file size drifted across refreshes: {sizes}" + finally: + shutil.rmtree(work, ignore_errors=True) diff --git a/tests/unit/test_persist_exp_props.py b/tests/unit/test_persist_exp_props.py new file mode 100644 index 000000000..4fc0bc039 --- /dev/null +++ b/tests/unit/test_persist_exp_props.py @@ -0,0 +1,360 @@ +"""Property-based state machine over ``workflow/scripts/persist_exp.py``. + +``exp_persist`` packs one exposure's keepable PSF products into a tar on +/project and writes a manifest describing it, and its central promise is that +RETENTION IS ADDITIVE: an existing tar is a floor, so shrinking the campaign's +keep list can never delete a product from the backed-up filesystem. That is a +claim about every sequence of keep lists and store states the campaign can +walk through, so it is tested here against a model — random keep lists over a +random set of present products, packed repeatedly, with the tar and the +manifest asserted after every pack. + +The keep lists mix product NAMES (``psf_model``), RAW GLOBS (``*.fits``) and +overlapping combinations of the two, because overlap is the case that once +failed every exposure in a campaign: two patterns matching one file is one +file, not a name collision. A genuine collision — two DIFFERENT source paths +landing on one flat member name — must still be fatal, and has its own test. + +The script is driven through ``main()`` with a patched ``sys.argv`` rather than +a subprocess: the rule invokes it as a script, but a subprocess per hypothesis +step would put this file out of reach of a login node's time budget. +""" + +import fnmatch +import hashlib +import importlib.util +import json +import shutil +import sys +import tarfile +import tempfile +from pathlib import Path + +import pytest +from hypothesis import HealthCheck, given, settings +from hypothesis import strategies as st +from hypothesis.stateful import ( + RuleBasedStateMachine, + initialize, + precondition, + rule, +) + +REPO_ROOT = Path(__file__).resolve().parents[2] +SCRIPTS = REPO_ROOT / "workflow" / "scripts" + + +def _load(name): + path = SCRIPTS / f"{name}.py" + assert path.exists(), f"{path} not found; the rule calls it by path" + sys.path.insert(0, str(SCRIPTS)) + try: + spec = importlib.util.spec_from_file_location(f"_{name}", path) + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + finally: + sys.path.remove(str(SCRIPTS)) + return module + + +persist = _load("persist_exp") +ALWAYS = persist.ALWAYS + +# One concrete file name per catalogued product, in the module output dir the +# real chain writes it to. The names are shaped like the campaign's (module +# tag, exposure, CCD) so the catalogue's globs match them for the same reason +# they match the real thing. +LAYOUT = { + "star_selection": ("setools", "mask", "star_selection-2079614-5.fits"), + "star_train": ("setools", "rand_split", + "star_split_ratio_80-2079614-5.fits"), + "star_test": ("setools", "rand_split", + "star_split_ratio_20-2079614-5.fits"), + "star_stats": ("setools", "stat", "star_stat-2079614-5.txt"), + "psf_model": ("psfex", "", "star_split_ratio_80-2079614-5.psf"), + "psfex_cat": ("psfex", "", "psfex_cat-2079614-5.cat"), + "psf_validation": ("psfex_interp", "", "validation_psf-2079614-5.fits"), +} +OPTIONAL = sorted(set(LAYOUT) - {ALWAYS}) +# What a campaign can write in `persist_exp:` — names, raw globs, and one name +# the catalogue does not know, which must be refused before any work happens. +ENTRIES = OPTIONAL + ["*.fits", "*.psf", "star_*", "validation_psf-*.fits"] +UNKNOWN = "psf_residuals" + +EXP = "2079614" + + +def _md5(path: Path) -> str: + return hashlib.md5(path.read_bytes()).hexdigest() + + +def _members(tar: Path) -> list: + with tarfile.open(tar) as tf: + return [ti.name for ti in tf.getmembers() if ti.isfile()] + + +class Store: + """One exposure's scratch store, its destination, and how to pack it.""" + + def __init__(self): + self.root = Path(tempfile.mkdtemp(prefix="persist-exp-props-")) + self.exp_dir = self.root / "exp" / EXP + self.dest = self.root / "products" / "psf" + self.manifest = self.root / "products" / "manifests" / f"{EXP}.json" + self.tar = self.dest / f"{EXP}.tar" + + def close(self): + shutil.rmtree(self.root, ignore_errors=True) + + def path_of(self, product: str) -> Path: + module, sub, name = LAYOUT[product] + base = (self.exp_dir / "output" / persist.RUN_NAME + / f"run_sp_{module}" / "output") + return (base / sub / name) if sub else (base / name) + + def write(self, product: str, payload: bytes) -> None: + path = self.path_of(product) + path.parent.mkdir(parents=True, exist_ok=True) + path.write_bytes(payload) + + def drop(self, product: str) -> None: + self.path_of(product).unlink(missing_ok=True) + + def pack(self, keep: list) -> int: + """Run the script's ``main`` as the rule does. 0 on success.""" + argv = ["persist_exp.py", "--exp-dir", str(self.exp_dir), + "--exp", EXP, "--dest", str(self.dest), + "--manifest", str(self.manifest)] + for entry in keep: + argv += ["--pattern", entry] + old = sys.argv + sys.argv = argv + try: + persist.main() + return 0 + except SystemExit as exc: + return 1 if exc.code not in (0, None) else 0 + finally: + sys.argv = old + + +class PersistExpMachine(RuleBasedStateMachine): + """A store that gains and loses products under a keep list that changes.""" + + @initialize() + def setup(self): + self.store = Store() + self.present = set() + self.prior_members = [] # members of the last tar written + self.prior_labels = {} # member -> product recorded for it + + def teardown(self): + self.store.close() + + # --- the store and the config move ---------------------------------- + @rule(product=st.sampled_from(sorted(LAYOUT)), size=st.integers(1, 64)) + def add_product(self, product, size): + self.store.write(product, bytes([len(product) % 251]) * size) + self.present.add(product) + + @rule(product=st.sampled_from(sorted(LAYOUT))) + def drop_product(self, product): + self.store.drop(product) + self.present.discard(product) + + @rule(keep=st.lists(st.sampled_from(ENTRIES), max_size=4, unique=True)) + def pack(self, keep): + self._pack_and_check(keep) + + @rule(keep=st.lists(st.sampled_from(ENTRIES), max_size=3, unique=True)) + def pack_with_unknown_product(self, keep): + """An unknown product name is refused before anything is written.""" + before = (_md5(self.store.tar) if self.store.tar.exists() else None) + code = self.store.pack(keep + [UNKNOWN]) + assert code != 0, "an unknown product name was accepted" + after = (_md5(self.store.tar) if self.store.tar.exists() else None) + assert after == before, "a refused keep list still touched the tar" + + @rule() + @precondition(lambda self: ALWAYS in self.present) + def pack_twice_unchanged(self): + """A rerun over an unchanged store must not move a single byte.""" + keep = sorted(OPTIONAL)[:2] + self._pack_and_check(keep) + tar_md5, man_md5 = _md5(self.store.tar), _md5(self.store.manifest) + tar_mtime = self.store.tar.stat().st_mtime_ns + man_mtime = self.store.manifest.stat().st_mtime_ns + assert self.store.pack(keep) == 0 + assert _md5(self.store.tar) == tar_md5, "the tar is not byte-stable" + assert _md5(self.store.manifest) == man_md5, "the manifest is not byte-stable" + assert self.store.tar.stat().st_mtime_ns == tar_mtime, ( + "an unchanged rerun rewrote the tar; mtime is a rerun trigger") + assert self.store.manifest.stat().st_mtime_ns == man_mtime, ( + "an unchanged rerun rewrote the manifest") + + # --- what a pack must leave behind ----------------------------------- + def _pack_and_check(self, keep): + had_tar = self.store.tar.exists() + tar_before = _md5(self.store.tar) if had_tar else None + man_before = (_md5(self.store.manifest) + if self.store.manifest.exists() else None) + code = self.store.pack(keep) + + if ALWAYS not in self.present: + # The star catalogue's input is not optional: the job fails and + # nothing downstream may be told the store is safe to reclaim. + assert code != 0, ( + f"{ALWAYS} is missing and the pack still succeeded") + assert (_md5(self.store.tar) if self.store.tar.exists() + else None) == tar_before, "a failed pack touched the tar" + assert (_md5(self.store.manifest) + if self.store.manifest.exists() + else None) == man_before, ( + "a failed pack wrote a manifest; clean_exposure would take " + "that as permission to delete the store") + return + + assert code == 0, f"pack failed with {ALWAYS} present and keep={keep}" + assert self.store.tar.exists() and self.store.manifest.exists() + members = _members(self.store.tar) + assert len(members) == len(set(members)), ( + f"duplicate member names in the tar: {members}") + + # ADDITIVE: an existing tar is a floor. + assert set(members) >= set(self.prior_members), ( + "members vanished from the tar: " + f"{sorted(set(self.prior_members) - set(members))}") + + body = json.loads(self.store.manifest.read_text()) + listed = {f["name"] for f in body["files"]} + assert listed == set(members), ( + "manifest and tar disagree about what was packed: " + f"{sorted(listed ^ set(members))}") + assert body["n_files"] == len(members) + assert body["unit"] == EXP and body["status"] == "complete" + + entries = [ALWAYS] + [e for e in keep if e != ALWAYS] + assert body["products"] == entries + for f in body["files"]: + if f["src"] is None: # carried from the previous tar + assert f["name"] in self.prior_members + assert f["product"] == self.prior_labels.get(f["name"], "?") + continue + assert f["product"] in entries, ( + f"{f['name']} labelled {f['product']!r}, not in the keep list") + assert fnmatch.fnmatch(f["name"], persist.resolve(f["product"])), ( + f"{f['name']} does not match {f['product']!r}'s glob") + assert Path(f["src"]).exists() + assert f["bytes"] == Path(f["src"]).stat().st_size + + # Every present product the keep list asks for is in there. + for entry in entries: + glob = persist.resolve(entry) + for product in self.present: + if fnmatch.fnmatch(LAYOUT[product][2], glob): + assert LAYOUT[product][2] in listed, ( + f"{product} matched {entry!r} but was not packed") + + self.prior_members = members + self.prior_labels = {f["name"]: f["product"] for f in body["files"]} + + +PersistExpMachine.TestCase.settings = settings( + max_examples=120, + stateful_step_count=12, + deadline=None, + suppress_health_check=[HealthCheck.too_slow, HealthCheck.data_too_large], +) +TestPersistExpMachine = PersistExpMachine.TestCase + + +@pytest.fixture() +def store(): + s = Store() + yield s + s.close() + + +def _seed(store, products=(ALWAYS,)): + for i, product in enumerate(products): + store.write(product, bytes([i + 1]) * (16 + i)) + + +def test_corrupt_existing_tar_is_refused_and_left_alone(store): + """A tar that cannot be read may still hold the only copy of something.""" + _seed(store, (ALWAYS, "psf_model")) + assert store.pack(["psf_model"]) == 0 + store.tar.write_bytes(b"not a tar at all, not even close" * 8) + corrupt = store.tar.read_bytes() + man_before = _md5(store.manifest) + + assert store.pack(["psf_model"]) != 0, "a corrupt tar was overwritten" + assert store.tar.read_bytes() == corrupt, "the corrupt tar was modified" + assert _md5(store.manifest) == man_before, ( + "a manifest was written over a tar that could not be read") + assert not store.tar.with_name(store.tar.name + ".tmp").exists() + + +def test_two_sources_with_one_member_name_is_fatal(store): + """Members are flat, so a real name clash would silently overwrite.""" + _seed(store, (ALWAYS,)) + # The same file name under a second module output dir. + clash = (store.exp_dir / "output" / persist.RUN_NAME / "run_sp_setools" + / "output" / "new_cat" / LAYOUT[ALWAYS][2]) + clash.parent.mkdir(parents=True, exist_ok=True) + clash.write_bytes(b"a different file with the same name") + + assert store.pack([]) != 0, "two different sources shared a member name" + assert not store.manifest.exists() + assert not store.tar.exists() + + +def test_shrinking_the_keep_list_cannot_delete_a_product(store): + """The property the additive rule exists for, stated end to end.""" + _seed(store, (ALWAYS, "psf_model", "star_train")) + assert store.pack(["psf_model", "star_train"]) == 0 + wide = set(_members(store.tar)) + assert LAYOUT["psf_model"][2] in wide + + # The campaign changes its mind, and the scratch store is gone. + for product in ("psf_model", "star_train"): + store.drop(product) + assert store.pack([]) == 0 + assert set(_members(store.tar)) == wide, ( + "shrinking persist_exp: deleted products from the backed-up tar") + body = json.loads(store.manifest.read_text()) + carried = {f["name"] for f in body["files"] if f["src"] is None} + assert LAYOUT["psf_model"][2] in carried + assert {f["name"]: f["product"] for f in body["files"]}[ + LAYOUT["psf_model"][2]] == "psf_model", ( + "a carried member lost the product label the old manifest had") + + +@settings(max_examples=80, deadline=None, + suppress_health_check=[HealthCheck.too_slow]) +@given(st.lists(st.sampled_from( + [ALWAYS, "*.fits", "validation_psf-*.fits", "psf_validation", + "star_*", "*.psf", "psf_model", "star_train"]), + min_size=1, max_size=5)) +def test_overlapping_patterns_never_fail(keep): + """Two patterns matching one file is one file, not a name collision. + + Overlap is ordinary — ``validation_psf-*.fits`` beside ``*.fits`` is a + perfectly reasonable way to say "the validation catalogues, and everything + else FITS while we are here" — and treating the second match as a clash + once failed every exposure in a campaign. + + A fresh store per example, because the additive rule makes packing + stateful and this property is about ONE pack. + """ + s = Store() + try: + _seed(s, tuple(LAYOUT)) + assert s.pack(keep) == 0, f"overlapping keep list failed: {keep}" + members = _members(s.tar) + assert len(members) == len(set(members)), members + # Every product present matched something, so all seven are packed. + assert set(members) == {name for _, _, name in LAYOUT.values()} & set( + members) + finally: + s.close() From 1a343e6a8b4e4e2944c0b85397ffe13e3e9b6f00 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Wed, 9 Sep 2026 22:37:28 -0400 Subject: [PATCH 21/85] fix(persist-exp): the PSF run dir is run_sp_exp_SxSePsf persist_exp.py carries its own copy of the exposure PSF stage's run-dir name, and it still held the pre-98bc0857 spelling. exp_persist would have searched run_sp_exp_SxSePsfPi, found nothing, and tarred an empty product set -- a silent loss rather than a failure, since an exposure with no keepable products is a legitimate state. This half of the rename lives here rather than in the develop hotfix because persist_exp.py does not exist on develop; it arrives with this branch. tests/unit/test_workflow_run_names.py (in the hotfix) checks this file when it is present and skips the check when it is not, so the guard travels with whichever branch has something to guard. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar (cherry picked from commit 5c1d41fe31537c6d22628670de4be8083ded5beb) --- workflow/scripts/persist_exp.py | 14 ++++++++------ 1 file changed, 8 insertions(+), 6 deletions(-) diff --git a/workflow/scripts/persist_exp.py b/workflow/scripts/persist_exp.py index bab1a5f0a..c3e14633d 100644 --- a/workflow/scripts/persist_exp.py +++ b/workflow/scripts/persist_exp.py @@ -20,7 +20,7 @@ snakemake rerun THIS rule (seconds of cp) and leaves the PSF chain alone. Folded into ``exp_psf``, the same edit would re-derive every PSF model in the campaign. -WHAT IT SEARCHES. ``/output/run_sp_exp_SxSePsfPi/*/output/`` — the four +WHAT IT SEARCHES. ``/output/run_sp_exp_SxSePsf/*/output/`` — the four module output dirs of the PSF config (sextractor, setools, psfex, psfex_interp) — RECURSIVELY. The recursion is not laziness: setools does not write flat, it writes into ``mask/``, ``rand_split/``, ``new_cat/``, ``plot/`` and ``stat/`` @@ -94,11 +94,13 @@ import tarfile from pathlib import Path -# The PSF chain's run dir (RUN_NAME in config_exp_psfex.ini). Hardcoded rather -# than passed: this rule persists the PSF stage's products and nothing else, and -# a knob here would be a knob for "persist some other stage", which is a -# different rule. -RUN_NAME = "run_sp_exp_SxSePsfPi" +# The PSF chain's run dir: RUN_NAME in config_exp_psfex.ini AND in +# config_exp_mccd.ini, which carry the same name on purpose so nothing +# downstream of exp_psf branches on the PSF model. Hardcoded rather than passed: +# this rule persists the PSF stage's products and nothing else, and a knob here +# would be a knob for "persist some other stage", which is a different rule. +# tests/unit/test_workflow_run_names.py holds this equal to the configs. +RUN_NAME = "run_sp_exp_SxSePsf" # --- the product catalogue (CosmoStat/shapepipe#844) ------------------------ # THE SINGLE SOURCE OF TRUTH for what an exposure can keep. `persist_exp:` in From 5b88f1711c2645b1e33c9a4e6ac4ced0002bc57e Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Fri, 11 Sep 2026 14:38:45 +0200 Subject: [PATCH 22/85] workflow: image simulations as an input mode of the real-data workflow Image-sim m-bias ran through the legacy bash runner and example/cfis_image_sims, a config chain frozen at an older pipeline state (no background vignets, no bkg-rms ngmix, no neighbour/mom-fail flags; its default.* symlinks dangle since e32c4fb9). An m measured that way does not calibrate the pipeline that makes the real catalogue. This makes the simulations an input of the same workflow: - input_type: data | image_sims selects $SP_CONFIG. config/cfis_image_sims is an overlay: real files only for the ingestion INIs whose input naming differs (tile Git/Uz/Fe, exposure Gie/Sp -- the sim get_images writes the real-data output names, so every later stage reads identical patterns), everything else symlinked into config/cfis. - psf_model: fake -- the simulations' true PSF. The exposure stage runs only SExtractor (background maps for the vignets); tile_vignets runs fake_interp_runner, which writes galaxy_psf from `psf_dict` under the name the shared configs read (${SP_PSF}_interp_runner). Star-bearing sims can run psfex/mccd unchanged. - unit_pre writes dashed tile numbers for sims (get_images substitutes them verbatim) and exports PSF_DICT when set. A data run's prologue is byte-identical (checked by diffing dry-run shell commands against the base), so no finished data unit is rerun by the params trigger. - bin/sp: SP_PROFILE selects profiles/; SP_RUN_CONFIG replaces (not layers on) workflow/config.yaml and is snapshotted with the code; the venv is optional when snakemake is on PATH. container.py follows both. - profiles/candide: SLURM profile for candide (node-local /tmp bound as /local/scratch for the tile store). Note: completeness.py gains the `fake` tables, and its content hash is a param on every rule, so this lands at a campaign boundary like any completeness edit. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01KNJYtm9z6VeoChxtDLYn4W --- profiles/candide/config.yaml | 34 ++++ src/shapepipe/modules/fake_interp_runner.py | 44 +++++ workflow/README.md | 28 +++ workflow/Snakefile | 62 ++++++- workflow/bin/sp | 40 ++++- workflow/config.yaml | 10 +- .../config/cfis_image_sims/config_MCCD.ini | 1 + .../config/cfis_image_sims/config_exp_Gie.ini | 104 +++++++++++ .../config/cfis_image_sims/config_exp_Sp.ini | 81 +++++++++ .../cfis_image_sims/config_exp_fake.ini | 129 ++++++++++++++ .../cfis_image_sims/config_exp_mccd.ini | 1 + .../cfis_image_sims/config_exp_psfex.ini | 1 + .../config/cfis_image_sims/config_tile_Fe.ini | 80 +++++++++ .../cfis_image_sims/config_tile_Git.ini | 98 ++++++++++ .../config/cfis_image_sims/config_tile_Mc.ini | 1 + .../cfis_image_sims/config_tile_Mh_exp.ini | 1 + .../config/cfis_image_sims/config_tile_Ms.ini | 1 + .../config_tile_Ng_template.ini | 1 + .../config_tile_PiViVi_fake.ini | 167 ++++++++++++++++++ .../config_tile_PiViVi_mccd.ini | 1 + .../config_tile_PiViVi_psfex.ini | 1 + .../config/cfis_image_sims/config_tile_Sx.ini | 1 + .../config/cfis_image_sims/config_tile_Uz.ini | 77 ++++++++ workflow/config/cfis_image_sims/default.conv | 1 + workflow/config/cfis_image_sims/default.param | 1 + workflow/config/cfis_image_sims/default.psfex | 1 + .../config/cfis_image_sims/default_exp.sex | 1 + .../cfis_image_sims/default_noimaflags.param | 1 + .../config/cfis_image_sims/default_tile.sex | 1 + .../config/cfis_image_sims/final_cat.param | 1 + .../cfis_image_sims/star_selection.setools | 1 + workflow/scripts/completeness.py | 16 +- workflow/scripts/container.py | 8 +- 33 files changed, 981 insertions(+), 15 deletions(-) create mode 100644 profiles/candide/config.yaml create mode 100644 src/shapepipe/modules/fake_interp_runner.py create mode 120000 workflow/config/cfis_image_sims/config_MCCD.ini create mode 100644 workflow/config/cfis_image_sims/config_exp_Gie.ini create mode 100644 workflow/config/cfis_image_sims/config_exp_Sp.ini create mode 100644 workflow/config/cfis_image_sims/config_exp_fake.ini create mode 120000 workflow/config/cfis_image_sims/config_exp_mccd.ini create mode 120000 workflow/config/cfis_image_sims/config_exp_psfex.ini create mode 100644 workflow/config/cfis_image_sims/config_tile_Fe.ini create mode 100644 workflow/config/cfis_image_sims/config_tile_Git.ini create mode 120000 workflow/config/cfis_image_sims/config_tile_Mc.ini create mode 120000 workflow/config/cfis_image_sims/config_tile_Mh_exp.ini create mode 120000 workflow/config/cfis_image_sims/config_tile_Ms.ini create mode 120000 workflow/config/cfis_image_sims/config_tile_Ng_template.ini create mode 100644 workflow/config/cfis_image_sims/config_tile_PiViVi_fake.ini create mode 120000 workflow/config/cfis_image_sims/config_tile_PiViVi_mccd.ini create mode 120000 workflow/config/cfis_image_sims/config_tile_PiViVi_psfex.ini create mode 120000 workflow/config/cfis_image_sims/config_tile_Sx.ini create mode 100644 workflow/config/cfis_image_sims/config_tile_Uz.ini create mode 120000 workflow/config/cfis_image_sims/default.conv create mode 120000 workflow/config/cfis_image_sims/default.param create mode 120000 workflow/config/cfis_image_sims/default.psfex create mode 120000 workflow/config/cfis_image_sims/default_exp.sex create mode 120000 workflow/config/cfis_image_sims/default_noimaflags.param create mode 120000 workflow/config/cfis_image_sims/default_tile.sex create mode 120000 workflow/config/cfis_image_sims/final_cat.param create mode 120000 workflow/config/cfis_image_sims/star_selection.setools diff --git a/profiles/candide/config.yaml b/profiles/candide/config.yaml new file mode 100644 index 000000000..67ea1bfe5 --- /dev/null +++ b/profiles/candide/config.yaml @@ -0,0 +1,34 @@ +# Snakemake profile for the ShapePipe workflow on candide (IAP). Selected by +# `SP_PROFILE=candide workflow/bin/sp run`. Mirrors profiles/nibi/config.yaml; +# only the cluster facts differ, and each is stated where it is set. +executor: slurm +# candide is a shared ~25-node cluster: cap concurrent submissions well below +# what nibi takes. +jobs: 150 +default-resources: + mem_mb: 2000 + runtime: 120 # minutes + slurm_account: cusers + # comp nodes have TmpDisk=0, so the tile_shape group (--tmp, see tile.smk's + # TILE_SLURM_EXTRA) lands on pscomp; everything else may use either. + slurm_partition: "comp,pscomp" + # n09/n17/n36: excluded in the sp_validation candide profile; n23: slurm jobs + # failed there (sp_validation 2b40c99). + slurm_extra: "'--exclude=n09,n17,n23,n36'" +software-deployment-method: [apptainer] +# /local/scratch is where tile.smk puts the node-local tile store (TILE_LOCAL, +# ~5.6 GB per real-data tile). candide has no /local, and its node-local /scratch +# is writable only per job (/scratch/$USER/$SLURM_JOB_ID, made by the prolog), +# which this line cannot name: snakemake passes it through literally, and +# --cleanenv drops SLURM_JOB_ID inside the container. So the node's /tmp (local +# disk, ~29 GB free, sticky) is bound there instead. That holds ~5 concurrent +# tile stores per node; if a campaign packs more tile_shape groups onto one node, +# this is the line to revisit. The PYTHONPATH pin is rewritten by bin/sp to the +# launch snapshot's src/; the value here only has to be a checkout's src/. +apptainer-args: "--cleanenv --env OMP_NUM_THREADS=1 --env MALLOC_ARENA_MAX=2 --env PYTHONPATH=/n17data/mkilbing/astro/repositories/github/shapepipe/src --bind /home,/automnt,/n17data,/n23data1,/n09data --bind /tmp:/local/scratch" +latency-wait: 60 # NFS: wait for outputs to appear after a job +keep-going: true # a failed job poisons only its cone; siblings run on +rerun-incomplete: true # re-do jobs left incomplete by an unclean death +show-failed-logs: true +printshellcmds: true +rerun-triggers: [mtime, params, code, software-env] diff --git a/src/shapepipe/modules/fake_interp_runner.py b/src/shapepipe/modules/fake_interp_runner.py new file mode 100644 index 000000000..bc167f924 --- /dev/null +++ b/src/shapepipe/modules/fake_interp_runner.py @@ -0,0 +1,44 @@ +"""FAKE INTERP RUNNER. + +Module runner for ``fake_psf`` under the ``_interp_runner`` name. + +The Snakemake workflow's configs read the galaxy PSF from +``${SP_PSF}_interp_runner`` for every PSF model, so with ``psf_model: fake`` +(image simulations, true PSF) this runner stands where ``psfex_interp_runner`` +and ``mccd_interp_runner`` stand for the real data. It writes the same +``galaxy_psf`` SqliteDict, taken from the simulation's PSF dictionary instead +of a fitted model. ``fake_psf_runner`` is the same module under its original +name, used by the legacy bash job scripts. + +:Author: Martin Kilbinger + +""" + +from shapepipe.modules.fake_psf_package import fake_psf +from shapepipe.modules.module_decorator import module_runner + + +@module_runner( + version="1.0", + file_pattern=["sexcat"], + file_ext=".fits", + depends=["numpy", "astropy", "sqlitedict"], + numbering_scheme="-000-000", +) +def fake_interp_runner( + input_file_list, + run_dirs, + file_number_string, + config, + module_config_sec, + w_log, +): + """Define The Fake Interp Runner.""" + sexcat_path = input_file_list[0] + psf_dict_path = config.getexpanded(module_config_sec, "PSF_DICT_PATH") + output_path = f'{run_dirs["output"]}/galaxy_psf{file_number_string}.sqlite' + + inst = fake_psf.FakePsf(sexcat_path, psf_dict_path, output_path, w_log) + inst.process() + + return None, None diff --git a/workflow/README.md b/workflow/README.md index 55611cedc..97510c901 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -45,6 +45,34 @@ Anything other than `run`, `report`, `container`, `cancel` passes straight throu snakemake with the workflow's profile and state dir — the direct command path for `sp --unlock`, `sp --dag`, `sp exp_psf ...`. +## Image simulations + +The same workflow runs the SKiLLS image simulations used to measure the shear +multiplicative bias, so that m calibrates the pipeline that makes the real +catalogue rather than a frozen copy of it. A simulation run sets +`input_type: image_sims`, which points `$SP_CONFIG` at +`config/cfis_image_sims/`. That directory holds real files only for the stages +whose input naming differs (tile Git/Uz/Fe, exposure Gie/Sp) and for the true-PSF +model; everything else is a symlink into `config/cfis/`, so a change to the +real-data chain reaches the simulations with no second edit. Keep the diff of +each overlay file to its `cfis/` original confined to input naming. + +`psf_model: fake` is the simulations' true PSF: the exposure stage runs only +SExtractor (for the background maps the vignets read), and `tile_vignets` runs +`fake_interp_runner`, which writes the `galaxy_psf` product from `psf_dict`. +Simulations that contain stars can run `psfex` or `mccd` exactly as the data do. + +One campaign per shear branch, each with its own run config: + +```bash +SP_PROFILE=candide SP_RUN_CONFIG=/path/run_1p2z_grid_1.yaml workflow/bin/sp run +``` + +`SP_RUN_CONFIG` replaces `workflow/config.yaml` (it is not layered on it, so no +key falls back to the committed run's paths) and is snapshotted with the code; +`SP_PROFILE` picks `profiles//`. sp_validation's image-simulation workflow +drives these campaigns and measures m from their final catalogues. + ## The container image `sp container` owns which image the jobs run inside. Two layers, and the second diff --git a/workflow/Snakefile b/workflow/Snakefile index bbdfeefb6..94e52f7f8 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -44,7 +44,13 @@ from snakemake.exceptions import WorkflowError # Resolved relative to THIS file, not the working directory: snakemake runs with # --directory on /scratch (bin/sp) so .snakemake/ state never lands on /project # (group quota is a hard 27/27 TiB — a metadata write mid-run died on it live). -configfile: str(Path(workflow.snakefile).parent / "config.yaml") +# +# SP_RUN_CONFIG (set by bin/sp) REPLACES the committed config.yaml rather than +# layering on it: a run driven from outside -- one image-simulation shear branch, +# say -- must not inherit the committed run's paths for any key it omits. +_RUN_CONFIG = os.environ.get("SP_RUN_CONFIG") or str( + Path(workflow.snakefile).parent / "config.yaml") +configfile: _RUN_CONFIG # Every job's shell runs inside this container (apptainer software-deployment in # the profile); the user never types apptainer. WHICH image is the one resolution @@ -71,12 +77,40 @@ if _kind == "none": container: _image -PSF_MODELS = {"psfex", "mccd"} +# --- input type and PSF model ---------------------------------------------- +# input_type selects WHERE the pixels come from; everything downstream of the +# ingestion stages is the same chain. `image_sims` runs the SKiLLS image +# simulations through the real-data configs: the overlay config dir +# (config/cfis_image_sims) holds real files only for the ingestion INIs whose +# input naming differs, and symlinks the rest into config/cfis, so an m-bias +# measured on the sims calibrates the pipeline that makes the real catalogue. +INPUT_TYPES = {"data", "image_sims"} +INPUT_TYPE = config.get("input_type", "data") +if INPUT_TYPE not in INPUT_TYPES: + raise WorkflowError( + f"Invalid input_type={INPUT_TYPE!r}; expected one of " + f"{sorted(INPUT_TYPES)}.") + +# `fake` is the image-simulation true PSF: no exposure PSF fit, and +# fake_interp_runner writes the tile's galaxy_psf from the simulation's PSF +# dictionary (config `psf_dict`) -- named so that the configs' shared +# `${SP_PSF}_interp_runner` reads it exactly as it reads psfex_interp_runner. Simulations that contain stars can instead run +# psfex or mccd exactly as the real data does. +PSF_MODELS = {"psfex", "mccd", "fake"} PSF_MODEL = config.get("psf_model", "psfex") if PSF_MODEL not in PSF_MODELS: raise WorkflowError( f"Invalid psf_model={PSF_MODEL!r}; expected one of " f"{sorted(PSF_MODELS)}.") +if PSF_MODEL == "fake" and INPUT_TYPE != "image_sims": + raise WorkflowError( + "psf_model=fake is the image-simulation true PSF; it needs " + "input_type=image_sims.") +PSF_DICT = config.get("psf_dict") or "" +if PSF_MODEL == "fake" and not PSF_DICT: + raise WorkflowError("psf_model=fake needs `psf_dict:` (the simulation's " + "pickled PSF dictionary) in the run config.") + # --- paths ----------------------------------------------------------------- # RUN_DIR is the scratch root: bulk intermediates, sized so a batch finishes @@ -95,8 +129,9 @@ INDEX_DB = Path(OUTPUTS["index_db"]) SCRIPTS = Path(workflow.basedir) / "scripts" # The config chain is the repo's committed directory (D2). The configs and # rules that set their environment variables must be versioned together. There -# is no `config_src` knob. -CONFIG_DIR = Path(workflow.basedir) / "config" / "cfis" +# is no `config_src` knob; input_type picks between the two committed dirs. +CONFIG_DIR = Path(workflow.basedir) / "config" / { + "data": "cfis", "image_sims": "cfis_image_sims"}[INPUT_TYPE] sys.path.insert(0, str(SCRIPTS)) import build_index # noqa: E402 @@ -510,7 +545,19 @@ def unit_pre(stage, unit, *, exp_name=None, forest=None, env=None, It also exports the configured input roots as ``SP_INPUT_TILES`` and ``SP_INPUT_EXPOSURES`` and the PSF choice as ``SP_PSF`` for the committed - ini chain. + ini chain (plus ``PSF_DICT`` for the image-simulation true PSF, which the + configs read as ``${SP_PSF}_interp_runner`` = ``fake_interp_runner``). + + Every line here is part of each rule's ``params.pre`` and so of the + ``params`` rerun trigger: a line added for every rule reruns every finished + unit of a campaign on its next ``sp run``. The image-simulation lines are + therefore conditional, and a data run's prologue is byte-identical to what + it was before input_type existed. + + ``tile_numbers.txt`` carries the number in the form the INPUT files are + named with, because get_images substitutes it verbatim into + INPUT_FILE_PATTERN: dot format for the survey tiles (``CFIS.210.282.r``), + dash format for the simulations (``CFIS_simu_image-210-282``). Finally it ``rm -rf``s this stage's own fixed run dir — ShapePipe's FileHandler raises on an existing run dir, and it is how a rerun never sees @@ -535,13 +582,16 @@ def unit_pre(stage, unit, *, exp_name=None, forest=None, env=None, f"export {_THREAD_CAPS}", 'mkdir -p "$SP_RUN/output" "$SP_RUN/manifests" "$SP_RUN/logs"', ] + if PSF_DICT: + lines.append(f"export PSF_DICT='{PSF_DICT}'") if forest: lines.append(f"export SP_EXP='{forest}'") for k, v in (env or {}).items(): lines.append(f"export {k}='{v}'") if level == "tile": - lines.append(f"printf '%s\\n' '{unit}' > \"$SP_RUN/tile_numbers.txt\"") + number = unit.replace(".", "-") if INPUT_TYPE == "image_sims" else unit + lines.append(f"printf '%s\\n' '{number}' > \"$SP_RUN/tile_numbers.txt\"") else: fe = "$SP_RUN/output/run_sp_tile_Fe/find_exposures_runner/output" lines += [f'mkdir -p "{fe}"', diff --git a/workflow/bin/sp b/workflow/bin/sp index fdc3d4f59..37eefa1e6 100755 --- a/workflow/bin/sp +++ b/workflow/bin/sp @@ -24,17 +24,37 @@ # executor re-invokes it inside jobs, so it cannot live on a node-local path). # One entry point, so a fresh tmux or a restart after a crash always launches # with the right state. +# +# Three environment variables make it drivable off nibi and from outside: +# +# SP_PROFILE profiles// to launch with (default: nibi). +# SP_SNAKEMAKE_ENV venv to activate; skipped when it does not exist and a +# snakemake is already on PATH (e.g. candide). +# SP_RUN_CONFIG a run config that REPLACES workflow/config.yaml (it is not +# layered on it, so no key falls back to the committed run's +# paths). Snapshotted with the code by `sp run`. This is how +# sp_validation drives one campaign per image-simulation +# shear branch. set -euo pipefail HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" # workflow/ REPO="$(dirname "$HERE")" VENV="${SP_SNAKEMAKE_ENV:-/project/def-mjhudson/cdaley/snakemake-env}" SCRIPTS="$HERE/scripts" -CONFIG="$HERE/config.yaml" +CONFIG="${SP_RUN_CONFIG:-$HERE/config.yaml}" +PROFILE="${SP_PROFILE:-nibi}" +[ -f "$CONFIG" ] || { echo "sp: run config $CONFIG does not exist" >&2; exit 2; } +[ -d "$REPO/profiles/$PROFILE" ] || { + echo "sp: no profile $REPO/profiles/$PROFILE (SP_PROFILE=$PROFILE)" >&2; exit 2; } module load apptainer/1.4.5 2>/dev/null || true -# shellcheck disable=SC1091 -source "$VENV/bin/activate" +if [ -f "$VENV/bin/activate" ]; then + # shellcheck disable=SC1091 + source "$VENV/bin/activate" +elif ! command -v snakemake >/dev/null 2>&1; then + echo "sp: no venv at $VENV and no snakemake on PATH (set SP_SNAKEMAKE_ENV)" >&2 + exit 2 +fi # Scalar reader for workflow/config.yaml. The venv is active by this point and # snakemake depends on PyYAML, so this parses the file rather than pattern-matching @@ -104,7 +124,7 @@ snapshot_code() { # rewritten HERE, in the snapshot's copy, and `sm` points --profile at that copy. # The checked-in profile keeps its literal checkout path and stays the source of # truth for the FLAGS; exactly one path is substituted, by one line of sed-work. - python - "$SNAPSHOT/profiles/nibi/config.yaml" "$SNAPSHOT/src" <<'PY' + python - "$SNAPSHOT/profiles/$PROFILE/config.yaml" "$SNAPSHOT/src" <<'PY' import pathlib, re, sys f, src = pathlib.Path(sys.argv[1]), sys.argv[2] text, n = re.subn(r'(--env PYTHONPATH=)\S+', lambda m: m.group(1) + src, @@ -137,6 +157,11 @@ out.write_text(json.dumps({ "dirty_files": status.splitlines() if status else [], }, indent=2) + "\n") PY + # An external run config is part of what is launched: snapshot it with the + # code, so a job re-parsing hours later reads the config the run started with. + if [ -n "${SP_RUN_CONFIG:-}" ]; then + cp "$SP_RUN_CONFIG" "$SNAPSHOT/run_config.yaml" + fi echo "sp: code snapshot refreshed at $SNAPSHOT" >&2 } @@ -164,9 +189,14 @@ code_root() { [ -f "$SNAPSHOT/workflow/Snakefile" ] && echo "$SNAPSHOT" || echo # index without building it and schedules no side effects. sm() { local root; root="$(code_root)" + # The run config the Snakefile reads: the snapshot's copy when this run has + # one, so every parse of the campaign -- head and jobs -- sees the same file. + if [ -n "${SP_RUN_CONFIG:-}" ] && [ -f "$root/run_config.yaml" ]; then + export SP_RUN_CONFIG="$root/run_config.yaml" + fi SP_PHASE="${SP_PHASE:-passthrough}" \ snakemake --snakefile "$root/workflow/Snakefile" \ - --profile "$root/profiles/nibi" \ + --profile "$root/profiles/$PROFILE" \ --directory "$STATE_DIR" "$@" } diff --git a/workflow/config.yaml b/workflow/config.yaml index 63b29413c..336cda5d6 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -28,9 +28,17 @@ inputs: # The container every job runs inside (apptainer software-deployment in the profile). container: /project/def-mjhudson/cdaley/containers/shapepipe-develop-runtime.sif -# PSF model used by the exposure and tile interpolation stages. +# PSF model used by the exposure and tile interpolation stages: psfex, mccd, or +# fake (image simulations only: the true PSF from `psf_dict`, no exposure fit). psf_model: psfex +# Where the pixels come from: `data` (survey tiles/exposures, config/cfis) or +# `image_sims` (SKiLLS simulations, config/cfis_image_sims -- an overlay that +# differs from config/cfis only in the ingestion stages). An image_sims run +# normally sets `psf_model: fake` and `psf_dict:`, and is driven with its own run +# config through SP_RUN_CONFIG (see README, "Image simulations"). +input_type: data + # THE TWO ROOTS (D5). # # run_dir is the SCRATCH root and the $SP_RUN every config interpolates: bulk diff --git a/workflow/config/cfis_image_sims/config_MCCD.ini b/workflow/config/cfis_image_sims/config_MCCD.ini new file mode 120000 index 000000000..3ba64d26e --- /dev/null +++ b/workflow/config/cfis_image_sims/config_MCCD.ini @@ -0,0 +1 @@ +../cfis/config_MCCD.ini \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/config_exp_Gie.ini b/workflow/config/cfis_image_sims/config_exp_Gie.ini new file mode 100644 index 000000000..d2c782672 --- /dev/null +++ b/workflow/config/cfis_image_sims/config_exp_Gie.ini @@ -0,0 +1,104 @@ +# IMAGE SIMULATIONS overlay of ../cfis/config_exp_Gie.ini: identical except the lines +# marked below. Keep the diff to ../cfis/config_exp_Gie.ini confined to input naming. +# ShapePipe configuration file for: get images + + +## Default ShapePipe options +[DEFAULT] + +# verbose mode (optional), default: True, print messages on terminal +VERBOSE = False + +# Name of run (optional) default: shapepipe_run +RUN_NAME = run_sp_exp_Gie + +# Add date and time to RUN_NAME, optional, default: False +RUN_DATETIME = False + + +## ShapePipe execution options +[EXECUTION] + +# Module name, single string or comma-separated list of valid module runner names +MODULE = get_images_runner + +# Parallel processing mode, SMP or MPI +MODE = SMP + + +## ShapePipe file handling options +[FILE] + +# Log file master name, optional, default: shapepipe +LOG_NAME = log_sp + +# Runner log file name, optional, default: shapepipe_runs +RUN_LOG_NAME = log_run_sp + +# Input directory, containing input files, single string or list of names +INPUT_DIR = $SP_RUN + +# Output directory +OUTPUT_DIR = $SP_RUN/output + + +## ShapePipe job handling options +[JOB] + +# Batch size of parallel processing (optional), default is 1, i.e. run all jobs in serial +SMP_BATCH_SIZE = 1 + +# Timeout value (optional), default is None, i.e. no timeout limit applied +TIMEOUT = 96:00:00 + + +## Module options + +# Get exposures +[GET_IMAGES_RUNNER] + +INPUT_DIR = $SP_RUN/output/run_sp_tile_Fe/find_exposures_runner/output + +FILE_PATTERN = exp_numbers + +FILE_EXT = .txt + +# NUMBERING_SCHEME (optional) string with numbering pattern for input files +NUMBERING_SCHEME = -000-000 + + +# Paths + +# Output path (optional, default is [FILE]:OUTPUT_DIR +# OUTPUT_PATH = input_images + +# Input path where original images are stored. Can be local path or vos url. +# Single string or list of strings +INPUT_PATH = $SP_INPUT_EXPOSURES, $SP_INPUT_EXPOSURES, $SP_INPUT_EXPOSURES + +# Input file pattern including tile number as dummy template +# sims +INPUT_FILE_PATTERN = simu_image-0000000, simu_weight-0000000, simu_flag-0000000 + +# Input file extensions +# sims +INPUT_FILE_EXT = .fits, .fits, .fits + +# Input numbering scheme, python regexp +# sims: 7-digit exposure ids +INPUT_NUMBERING = \d{7} + +# Output file pattern without number +OUTPUT_FILE_PATTERN = image-, weight-, flag- + +# Method to retrieve images, one in 'vos', 'symlink' +RETRIEVE = symlink + +# If RETRIEVE=vos, number of attempts to download +# Optional, default=3 +N_TRY = 3 + +# Retrieve command options, optional +RETRIEVE_OPTIONS = --certfile=$HOME/.ssl/cadcproxy.pem + +#CHECK_EXISTING_DIR = $SP_RUN/output/run_sp_Gie_prev diff --git a/workflow/config/cfis_image_sims/config_exp_Sp.ini b/workflow/config/cfis_image_sims/config_exp_Sp.ini new file mode 100644 index 000000000..9783189f3 --- /dev/null +++ b/workflow/config/cfis_image_sims/config_exp_Sp.ini @@ -0,0 +1,81 @@ +# IMAGE SIMULATIONS overlay of ../cfis/config_exp_Sp.ini: identical except the lines +# marked below. Keep the diff to ../cfis/config_exp_Sp.ini confined to input naming. +# ShapePipe configuration file for single-exposures, +# split images + + +## Default ShapePipe options +[DEFAULT] + +# verbose mode (optional), default: True, print messages on terminal +VERBOSE = True + +# Name of run (optional) default: shapepipe_run +RUN_NAME = run_sp_exp_Sp + +# Add date and time to RUN_NAME, optional, default: True +RUN_DATETIME = False + + +## ShapePipe execution options +[EXECUTION] + +# Module name, single string or comma-separated list of valid module runner names +MODULE = split_exp_runner + +# Run mode, SMP or MPI +MODE = SMP + + +## ShapePipe file handling options +[FILE] + +# Log file master name, optional, default: shapepipe +LOG_NAME = log_sp + +# Runner log file name, optional, default: shapepipe_runs +RUN_LOG_NAME = log_run_sp + +# NUMBER_LIST selects this unit; the workflow sets SP_UNIT_NUM to the +# dashed exposure ID (exp split is a "tile-scheme" stage per sp_rule.py). +NUMBER_LIST = $SP_UNIT_NUM + +# Input directory, containing input files, single string or list of names with length matching FILE_PATTERN +INPUT_DIR = . + +# Output directory +OUTPUT_DIR = $SP_RUN/output + + +## ShapePipe job handling options +[JOB] + +# Batch size of parallel processing (optional), default is 1, i.e. run all jobs in serial +SMP_BATCH_SIZE = 8 + +# Timeout value (optional), default is None, i.e. no timeout limit applied +TIMEOUT = 96:00:00 + + +## Module options + +[SPLIT_EXP_RUNNER] + +INPUT_DIR = $SP_RUN/output/run_sp_exp_Gie/get_images_runner/output + +FILE_PATTERN = image, weight, flag + +# Matches compressed single-exposure files +# sims: uncompressed exposures +FILE_EXT = .fits, .fits, .fits + +NUMBERING_SCHEME = -0000000 + +# OUTPUT_SUFFIX, actually file name prefixes. +# Expected keyword "flag" will lead to a behavior where the data are saved as int. +# The code also expects the image data to use the "image" suffix +# (default value in the pipeline). +OUTPUT_SUFFIX = image, weight, flag + +# Number of HDUs/CCDs of mosaic +N_HDU = 40 diff --git a/workflow/config/cfis_image_sims/config_exp_fake.ini b/workflow/config/cfis_image_sims/config_exp_fake.ini new file mode 100644 index 000000000..f09a4054b --- /dev/null +++ b/workflow/config/cfis_image_sims/config_exp_fake.ini @@ -0,0 +1,129 @@ +# ShapePipe configuration file for single-exposures. IMAGE SIMULATIONS, true PSF +# (psf_model: fake). No PSF fit: the simulations' PSF is known, and the tile's +# fake_interp_runner reads it from the PSF dictionary. What remains is the +# SExtractor pass, whose background and background_rms checkimages the tile +# vignets read. Derived from ../cfis/config_exp_psfex.ini: the [SEXTRACTOR_RUNNER] +# section is kept identical to it, and RUN_NAME stays the exp_psf stage's run dir +# (workflow/scripts/completeness.py STAGE_DIR). + + +## Default ShapePipe options +[DEFAULT] + +# verbose mode (optional), default: True, print messages on terminal +VERBOSE = True + +# Name of run (optional) default: shapepipe_run +RUN_NAME = run_sp_exp_SxSePsf + +# Add date and time to RUN_NAME, optional, default: True +RUN_DATETIME = False + + +## ShapePipe execution options +[EXECUTION] + +# Module name, single string or comma-separated list of valid module runner names +MODULE = sextractor_runner + +# Run mode, SMP or MPI +MODE = SMP + + +## ShapePipe file handling options +[FILE] + +# Log file master name, optional, default: shapepipe +LOG_NAME = log_sp + +# Runner log file name, optional, default: shapepipe_runs +RUN_LOG_NAME = log_run_sp + +# Input directory, containing input files, single string or list of names with length matching FILE_PATTERN +INPUT_DIR = $SP_RUN/output + +# Output directory +OUTPUT_DIR = $SP_RUN/output + + +## ShapePipe job handling options +[JOB] + +# Batch size of parallel processing (optional), default is 1, i.e. run all jobs in serial +SMP_BATCH_SIZE = 8 + +# Timeout value (optional), default is None, i.e. no timeout limit applied +TIMEOUT = 96:00:00 + + +## Module options + +[SEXTRACTOR_RUNNER] + +# The split CCDs, and nothing else: ShapePipe generates no masks +INPUT_DIR = $SP_RUN/output/run_sp_exp_Sp/split_exp_runner/output + +# Read the instrument flag image split_exp wrote per CCD +FILE_PATTERN = image, weight, flag + +# Explicit extensions: a 3-entry FILE_PATTERN override must not fall back on +# the decorator's 4-entry FILE_EXT default (length check fails at startup) +FILE_EXT = .fits, .fits, .fits + +NUMBERING_SCHEME = -0000000-0 + +# SExtractor executable path +EXEC_PATH = source-extractor + +# SExtractor configuration files +DOT_SEX_FILE = $SP_CONFIG/default_exp.sex +DOT_PARAM_FILE = $SP_CONFIG//default.param +DOT_CONV_FILE = $SP_CONFIG/default.conv + +# Use input weight image if True +WEIGHT_IMAGE = True + +# Use input flag image if True +FLAG_IMAGE = True + +# Use input PSF file if True +PSF_FILE = False + +# Use distinct image for detection (SExtractor in +# dual-image mode) if True. +DETECTION_IMAGE = False + +# Distinct weight image for detection (SExtractor +# in dual-image mode) +DETECTION_WEIGHT = False + +# True if photometry zero-point is to be read from exposure image header +ZP_FROM_HEADER = True + +# If ZP_FROM_HEADER is True, zero-point key name +ZP_KEY = PHOTZP + +# Background information from image header. +# If BKG_FROM_HEADER is True, background value will be read from header. +# In that case, the value of BACK_TYPE will be set atomatically to MANUAL. +# This is used e.g. for the LSB images. +BKG_FROM_HEADER = False +# LSB images: +# BKG_FROM_HEADER = True + +# If BKG_FROM_HEADER is True, background value key name +# LSB images: +#BKG_KEY = IMMODE + +# Type of image check (optional), default not used, can be a list of +# BACKGROUND, BACKGROUND_RMS, INIBACKGROUND, MINIBACK_RMS, -BACKGROUND, +# FILTERED, OBJECTS, -OBJECTS, SEGMENTATION, APERTURES +CHECKIMAGE = BACKGROUND, BACKGROUND_RMS + +# File name suffix for the output sextractor files (optional) SUFFIX = tile +SUFFIX = sexcat + +## Post-processing + +# Not required for single exposures +MAKE_POST_PROCESS = FALSE diff --git a/workflow/config/cfis_image_sims/config_exp_mccd.ini b/workflow/config/cfis_image_sims/config_exp_mccd.ini new file mode 120000 index 000000000..58166d49c --- /dev/null +++ b/workflow/config/cfis_image_sims/config_exp_mccd.ini @@ -0,0 +1 @@ +../cfis/config_exp_mccd.ini \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/config_exp_psfex.ini b/workflow/config/cfis_image_sims/config_exp_psfex.ini new file mode 120000 index 000000000..ea86cb50f --- /dev/null +++ b/workflow/config/cfis_image_sims/config_exp_psfex.ini @@ -0,0 +1 @@ +../cfis/config_exp_psfex.ini \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/config_tile_Fe.ini b/workflow/config/cfis_image_sims/config_tile_Fe.ini new file mode 100644 index 000000000..2a003ea9c --- /dev/null +++ b/workflow/config/cfis_image_sims/config_tile_Fe.ini @@ -0,0 +1,80 @@ +# IMAGE SIMULATIONS overlay of ../cfis/config_tile_Fe.ini: identical except the lines +# marked below. Keep the diff to ../cfis/config_tile_Fe.ini confined to input naming. +# ShapePipe configuration file for: find exposures + + +## Default ShapePipe options +[DEFAULT] + +# verbose mode (optional), default: True, print messages on terminal +VERBOSE = False + +# Name of run (optional) default: shapepipe_run +RUN_NAME = run_sp_tile_Fe + +# Add date and time to RUN_NAME, optional, default: False +RUN_DATETIME = False + + +## ShapePipe execution options +[EXECUTION] + +# Module name, single string or comma-separated list of valid module runner names +MODULE = find_exposures_runner + +# Parallel processing mode, SMP or MPI +MODE = SMP + + +## ShapePipe file handling options +[FILE] + +# Log file master name, optional, default: shapepipe +LOG_NAME = log_sp + +# Runner log file name, optional, default: shapepipe_runs +RUN_LOG_NAME = log_run_sp + +# NUMBER_LIST selects this unit; the workflow sets SP_UNIT_NUM to the +# dashed tile ID (e.g. -210-282). +NUMBER_LIST = $SP_UNIT_NUM + +# Input directory, containing input files, single string or list of names +INPUT_DIR = $SP_RUN + +# Output directory +OUTPUT_DIR = $SP_RUN/output + + +## ShapePipe job handling options +[JOB] + +# Batch size of parallel processing (optional), default is 1, i.e. run all jobs in serial +SMP_BATCH_SIZE = 1 + +# Timeout value (optional), default is None, i.e. no timeout limit applied +TIMEOUT = 96:00:00 + + +## Module options + +# Get tiles +[FIND_EXPOSURES_RUNNER] + +INPUT_DIR = $SP_RUN/output/run_sp_tile_Git/get_images_runner/output + +FILE_PATTERN = CFIS_image + +FILE_EXT = .fits + +# NUMBERING_SCHEME (optional) string with numbering pattern for input files +NUMBERING_SCHEME = -000-000 + +# Column number of exposure name in FITS header +# sims: HISTORY layout of the simulated tiles +COLNUM = 2 + +# Prefix to remove from exposure name +# sims +EXP_PREFIX = simu_image- + diff --git a/workflow/config/cfis_image_sims/config_tile_Git.ini b/workflow/config/cfis_image_sims/config_tile_Git.ini new file mode 100644 index 000000000..69fcff540 --- /dev/null +++ b/workflow/config/cfis_image_sims/config_tile_Git.ini @@ -0,0 +1,98 @@ +# IMAGE SIMULATIONS overlay of ../cfis/config_tile_Git.ini: identical except the lines +# marked below. Keep the diff to ../cfis/config_tile_Git.ini confined to input naming. +# ShapePipe configuration file for: get tile images + + +## Default ShapePipe options +[DEFAULT] + +# verbose mode (optional), default: True, print messages on terminal +VERBOSE = False + +# Name of run (optional) default: shapepipe_run +RUN_NAME = run_sp_tile_Git + +# Add date and time to RUN_NAME, optional, default: False +RUN_DATETIME = False + + +## ShapePipe execution options +[EXECUTION] + +# Module name, single string or comma-separated list of valid module runner names +MODULE = get_images_runner + +# Parallel processing mode, SMP or MPI +MODE = SMP + + +## ShapePipe file handling options +[FILE] + +# Log file master name, optional, default: shapepipe +LOG_NAME = log_sp + +# Runner log file name, optional, default: shapepipe_runs +RUN_LOG_NAME = log_run_sp + +# Input directory, containing input files, single string or list of names +INPUT_DIR = $SP_RUN + +# Output directory +OUTPUT_DIR = $SP_RUN/output + + +## ShapePipe job handling options +[JOB] + +# Batch size of parallel processing (optional), default is 1, i.e. run all jobs in serial +SMP_BATCH_SIZE = 1 + +# Timeout value (optional), default is None, i.e. no timeout limit applied +TIMEOUT = 96:00:00 + + +## Module options + +# Get tiles +[GET_IMAGES_RUNNER] + +FILE_PATTERN = tile_numbers + +FILE_EXT = .txt + +# NUMBERING_SCHEME (optional) string with numbering pattern for input files +NUMBERING_SCHEME = + +# Paths + +# Input path where original images are stored. Can be local path or vos url. +# Single string or list of strings +INPUT_PATH = $SP_INPUT_TILES, $SP_INPUT_TILES + +# Input file pattern including tile number as dummy template +# sims +INPUT_FILE_PATTERN = CFIS_simu_image-000-000, CFIS_simu_weight-000-000 + +# Input file extensions +# sims: weight uncompressed +INPUT_FILE_EXT = .fits, .fits + +# Input numbering scheme, python regexp +# sims: dash-numbered tiles +INPUT_NUMBERING = \d{3}-\d{3} + +# Output file pattern without number +OUTPUT_FILE_PATTERN = CFIS_image-, CFIS_weight- + +# Copy/download method, one in 'vos', 'symlink' +RETRIEVE = symlink + +# If RETRIEVE=vos, number of attempts to download +# Optional, default=3 +N_TRY = 3 + +# Copy command options, optional +RETRIEVE_OPTIONS = --certfile=$HOME/.ssl/cadcproxy.pem + +#CHECK_EXISTING_DIR = $SP_RUN/data_tiles diff --git a/workflow/config/cfis_image_sims/config_tile_Mc.ini b/workflow/config/cfis_image_sims/config_tile_Mc.ini new file mode 120000 index 000000000..eeba9f53e --- /dev/null +++ b/workflow/config/cfis_image_sims/config_tile_Mc.ini @@ -0,0 +1 @@ +../cfis/config_tile_Mc.ini \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/config_tile_Mh_exp.ini b/workflow/config/cfis_image_sims/config_tile_Mh_exp.ini new file mode 120000 index 000000000..5e8e91716 --- /dev/null +++ b/workflow/config/cfis_image_sims/config_tile_Mh_exp.ini @@ -0,0 +1 @@ +../cfis/config_tile_Mh_exp.ini \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/config_tile_Ms.ini b/workflow/config/cfis_image_sims/config_tile_Ms.ini new file mode 120000 index 000000000..f301e8673 --- /dev/null +++ b/workflow/config/cfis_image_sims/config_tile_Ms.ini @@ -0,0 +1 @@ +../cfis/config_tile_Ms.ini \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/config_tile_Ng_template.ini b/workflow/config/cfis_image_sims/config_tile_Ng_template.ini new file mode 120000 index 000000000..c00711e7b --- /dev/null +++ b/workflow/config/cfis_image_sims/config_tile_Ng_template.ini @@ -0,0 +1 @@ +../cfis/config_tile_Ng_template.ini \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/config_tile_PiViVi_fake.ini b/workflow/config/cfis_image_sims/config_tile_PiViVi_fake.ini new file mode 100644 index 000000000..4651c4e7c --- /dev/null +++ b/workflow/config/cfis_image_sims/config_tile_PiViVi_fake.ini @@ -0,0 +1,167 @@ +# ShapePipe configuration file for tile, from detection up to shape measurement. +# IMAGE SIMULATIONS, true PSF (psf_model: fake). Derived from +# ../cfis/config_tile_PiViVi_psfex.ini: only the PSF section differs. + + +## Default ShapePipe options +[DEFAULT] + +# verbose mode (optional), default: True, print messages on terminal +VERBOSE = True + +# Name of run (optional) default: shapepipe_run +RUN_NAME = run_sp_tile_PiViVi + +# Add date and time to RUN_NAME, optional, default: False +RUN_DATETIME = False + + +## ShapePipe execution options +[EXECUTION] + +# Module name, single string or comma-separated list of valid module runner names +#MODULE = psfex_interp_runner, + +MODULE = ${SP_PSF}_interp_runner, vignetmaker_runner, vignetmaker_runner + +# Parallel processing mode, SMP or MPI +MODE = SMP + + +## ShapePipe file handling options +[FILE] + +# Log file master name, optional, default: shapepipe +LOG_NAME = log_sp + +# Runner log file name, optional, default: shapepipe_runs +RUN_LOG_NAME = log_run_sp + +# NUMBER_LIST selects this unit; the workflow sets SP_UNIT_NUM to the +# dashed tile ID (e.g. -210-282). +NUMBER_LIST = $SP_UNIT_NUM + +# Input directory, containing input files, single string or list of names +INPUT_DIR = . + +# Output directory. +# +# NODE-LOCAL, not $SP_RUN/output. This run produces the tile's ~5.6 GB vignette +# store, which is read only by ngmix and make_cat -- both of which run in the +# same fused group job on the same node (workflow/rules/tile.smk, TILE_GROUP). +# Writing it to the node's NVMe instead of NFS scratch is the whole point: it +# removes the store from the per-tile scratch high-water AND removes ~163 GB of +# small random NFS reads per tile. +# +# $SP_VIGNET_OUT is exported by every member of that group (TILE_LOCAL) as +# $SLURM_TMPDIR/sp-tile/output. Set it to $SP_RUN/output to keep the store on +# shared storage. It is NOT optional: ShapePipe's config expansion is strict +# (pipeline/config.py::_expandvars_strict), so an unset value is a loud error. +# +# The INPUT_DIRs below stay on $SP_RUN -- they are Sx / Mh_exp / Fe products on +# scratch, and every path here is absolute, so nothing resolves through a run +# log that this split would break. +OUTPUT_DIR = $SP_VIGNET_OUT + + +## ShapePipe job handling options +[JOB] + +# Batch size of parallel processing (optional), default is 1, i.e. run all jobs in serial +SMP_BATCH_SIZE = 16 + +# Timeout value (optional), default is None, i.e. no timeout limit applied +TIMEOUT = 96:00:00 + + +## Module options + +[FAKE_INTERP_RUNNER] + +# The tile's multi-epoch SExtractor catalogue: its EPOCH HDUs carry, per object, +# the exposure and CCD each epoch fell on, which key the PSF dictionary. +INPUT_DIR = $SP_RUN/output/run_sp_tile_Sx/sextractor_runner/output + +FILE_PATTERN = sexcat + +FILE_EXT = .fits + +# NUMBERING_SCHEME (optional) string with numbering pattern for input files +NUMBERING_SCHEME = -000-000 + +# The simulation's pickled PSF dictionary ("-" -> PSF), exported by +# the workflow from the run config's `psf_dict:` +PSF_DICT_PATH = $PSF_DICT + + +# Create vignets for tiles weights +[VIGNETMAKER_RUNNER_RUN_1] + +INPUT_DIR = $SP_RUN/output/run_sp_tile_Sx/sextractor_runner/output, $SP_RUN/output/run_sp_tile_Uz/uncompress_fits_runner/output + +FILE_PATTERN = sexcat, CFIS_weight + +FILE_EXT = .fits, .fits + +# NUMBERING_SCHEME (optional) string with numbering pattern for input files +NUMBERING_SCHEME = -000-000 + +MASKING = False +MASK_VALUE = 0 + +# Run mode for psfex interpolation: +# CLASSIC: 'classical' run, interpolate to object positions +# MULTI-EPOCH: interpolate for multi-epoch images +# VALIDATION: validation for single-epoch images +MODE = CLASSIC + +# Coordinate frame type, one in PIX (pixel frame), SPHE (spherical coordinates) +COORD = PIX +POSITION_PARAMS = XWIN_IMAGE,YWIN_IMAGE + +# Vignet size in pixels +STAMP_SIZE = 51 + +# Output file name prefix, file name is _vignet.fits +PREFIX = weight + + +[VIGNETMAKER_RUNNER_RUN_2] + +# Create multi-epoch vignets for tiles corresponding to +# positions on single-exposures + +INPUT_DIR = $SP_RUN/output/run_sp_tile_Sx/sextractor_runner/output, $SP_RUN/output/run_sp_tile_Mh_exp/merge_headers_runner/output, $SP_RUN/output/run_sp_tile_Fe/find_exposures_runner/output + +FILE_PATTERN = sexcat, log_exp_headers, exp_numbers + +FILE_EXT = .fits, .sqlite, .txt + +# NUMBERING_SCHEME (optional) string with numbering pattern for input files +NUMBERING_SCHEME = -000-000 + +MASKING = False +MASK_VALUE = 0 + +# Run mode for psfex interpolation: +# CLASSIC: 'classical' run, interpolate to object positions +# MULTI-EPOCH: interpolate for multi-epoch images +# VALIDATION: validation for single-epoch images +MODE = MULTI-EPOCH + +# Coordinate frame type, one in PIX (pixel frame), SPHE (spherical coordinates) +COORD = SPHE +POSITION_PARAMS = XWIN_WORLD,YWIN_WORLD + +# Vignet size in pixels +STAMP_SIZE = 51 + +# Output file name prefix, file name is vignet.fits +PREFIX = + +# Additional parameters for path and file pattern corresponding to single-exposure +# run outputs. ME_IMAGE_EXP_DIR/ME_IMAGE_EXP_RUNNERS replace ME_IMAGE_DIR for +# the v2.0 per-exposure pipeline; output dirs are discovered by scanning $SP_EXP. +ME_IMAGE_EXP_DIR = $SP_EXP +ME_IMAGE_EXP_RUNNERS = split_exp_runner, split_exp_runner, split_exp_runner, sextractor_runner, sextractor_runner +ME_IMAGE_PATTERN = flag, image, weight, background, background_rms diff --git a/workflow/config/cfis_image_sims/config_tile_PiViVi_mccd.ini b/workflow/config/cfis_image_sims/config_tile_PiViVi_mccd.ini new file mode 120000 index 000000000..6271fb312 --- /dev/null +++ b/workflow/config/cfis_image_sims/config_tile_PiViVi_mccd.ini @@ -0,0 +1 @@ +../cfis/config_tile_PiViVi_mccd.ini \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/config_tile_PiViVi_psfex.ini b/workflow/config/cfis_image_sims/config_tile_PiViVi_psfex.ini new file mode 120000 index 000000000..5f120bca9 --- /dev/null +++ b/workflow/config/cfis_image_sims/config_tile_PiViVi_psfex.ini @@ -0,0 +1 @@ +../cfis/config_tile_PiViVi_psfex.ini \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/config_tile_Sx.ini b/workflow/config/cfis_image_sims/config_tile_Sx.ini new file mode 120000 index 000000000..81637de66 --- /dev/null +++ b/workflow/config/cfis_image_sims/config_tile_Sx.ini @@ -0,0 +1 @@ +../cfis/config_tile_Sx.ini \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/config_tile_Uz.ini b/workflow/config/cfis_image_sims/config_tile_Uz.ini new file mode 100644 index 000000000..18a13e6eb --- /dev/null +++ b/workflow/config/cfis_image_sims/config_tile_Uz.ini @@ -0,0 +1,77 @@ +# IMAGE SIMULATIONS overlay of ../cfis/config_tile_Uz.ini: identical except the lines +# marked below. Keep the diff to ../cfis/config_tile_Uz.ini confined to input naming. +# ShapePipe configuration file for: uncompress FITS image + + +## Default ShapePipe options +[DEFAULT] + +# verbose mode (optional), default: True, print messages on terminal +VERBOSE = True + +# Name of run (optional) default: shapepipe_run +RUN_NAME = run_sp_tile_Uz + +# Add date and time to RUN_NAME, optional, default: False +RUN_DATETIME = False + + +## ShapePipe execution options +[EXECUTION] + +# Module name, single string or comma-separated list of valid module runner names +MODULE = uncompress_fits_runner + +# Parallel processing mode, SMP or MPI +MODE = SMP + + +## ShapePipe file handling options +[FILE] + +# Log file master name, optional, default: shapepipe +LOG_NAME = log_sp + +# Runner log file name, optional, default: shapepipe_runs +RUN_LOG_NAME = log_run_sp + +# NUMBER_LIST selects this unit; the workflow sets SP_UNIT_NUM to the +# dashed tile ID (e.g. -210-282). +NUMBER_LIST = $SP_UNIT_NUM + +# Input directory, containing input files, single string or list of names +INPUT_DIR = . + +# Output directory +OUTPUT_DIR = $SP_RUN/output + + +## ShapePipe job handling options +[JOB] + +# Batch size of parallel processing (optional), default is 1, i.e. run all jobs in serial +SMP_BATCH_SIZE = 16 + +# Timeout value (optional), default is None, i.e. no timeout limit applied +TIMEOUT = 96:00:00 + + +## Module options +[UNCOMPRESS_FITS_RUNNER] + +INPUT_DIR = $SP_RUN/output/run_sp_tile_Git/get_images_runner/output + +FILE_PATTERN = CFIS_weight + +# sims: the weight is not compressed +FILE_EXT = .fits + +# NUMBERING_SCHEME (optional) string with numbering pattern for input files +NUMBERING_SCHEME = -000-000 + +# Input HDU of image data, optional, default=0 +# sims: plain FITS, data in the primary HDU +HDU_DATA = 0 + +# Output file pattern +OUTPUT_PATTERN = CFIS_weight diff --git a/workflow/config/cfis_image_sims/default.conv b/workflow/config/cfis_image_sims/default.conv new file mode 120000 index 000000000..bd71df850 --- /dev/null +++ b/workflow/config/cfis_image_sims/default.conv @@ -0,0 +1 @@ +../cfis/default.conv \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/default.param b/workflow/config/cfis_image_sims/default.param new file mode 120000 index 000000000..49e000314 --- /dev/null +++ b/workflow/config/cfis_image_sims/default.param @@ -0,0 +1 @@ +../cfis/default.param \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/default.psfex b/workflow/config/cfis_image_sims/default.psfex new file mode 120000 index 000000000..65dd716d1 --- /dev/null +++ b/workflow/config/cfis_image_sims/default.psfex @@ -0,0 +1 @@ +../cfis/default.psfex \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/default_exp.sex b/workflow/config/cfis_image_sims/default_exp.sex new file mode 120000 index 000000000..4c83dea8e --- /dev/null +++ b/workflow/config/cfis_image_sims/default_exp.sex @@ -0,0 +1 @@ +../cfis/default_exp.sex \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/default_noimaflags.param b/workflow/config/cfis_image_sims/default_noimaflags.param new file mode 120000 index 000000000..75451801e --- /dev/null +++ b/workflow/config/cfis_image_sims/default_noimaflags.param @@ -0,0 +1 @@ +../cfis/default_noimaflags.param \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/default_tile.sex b/workflow/config/cfis_image_sims/default_tile.sex new file mode 120000 index 000000000..8770da87e --- /dev/null +++ b/workflow/config/cfis_image_sims/default_tile.sex @@ -0,0 +1 @@ +../cfis/default_tile.sex \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/final_cat.param b/workflow/config/cfis_image_sims/final_cat.param new file mode 120000 index 000000000..c286947d0 --- /dev/null +++ b/workflow/config/cfis_image_sims/final_cat.param @@ -0,0 +1 @@ +../cfis/final_cat.param \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/star_selection.setools b/workflow/config/cfis_image_sims/star_selection.setools new file mode 120000 index 000000000..11a570187 --- /dev/null +++ b/workflow/config/cfis_image_sims/star_selection.setools @@ -0,0 +1 @@ +../cfis/star_selection.setools \ No newline at end of file diff --git a/workflow/scripts/completeness.py b/workflow/scripts/completeness.py index ff776e6e7..0a22179f5 100644 --- a/workflow/scripts/completeness.py +++ b/workflow/scripts/completeness.py @@ -118,6 +118,13 @@ # config_exp_mccd enables the ten meanshape and six histogram plots. "mccd_plots_runner": dict(expect=16, warn=True), }, + # Image simulations with the true PSF (psf_model: fake): no PSF fit on + # the exposures, only the SExtractor pass whose background/background_rms + # checkimages the tile vignets read (config_exp_fake.ini in + # config/cfis_image_sims). Same 40 CCDs x (sexcat, background, rms). + "fake": { + "sextractor_runner": dict(expect=120), + }, }, # --- tile post --- @@ -138,6 +145,13 @@ "vignetmaker_runner_run_1": dict(expect=1, warn=True), "vignetmaker_runner_run_2": dict(expect=5, warn=True), }, + # Image simulations: fake_interp_runner writes the same galaxy_psf + # sqlite psfex_interp_runner writes, from the simulation's PSF dictionary. + "fake": { + "fake_interp_runner": dict(expect=1), + "vignetmaker_runner_run_1": dict(expect=1), + "vignetmaker_runner_run_2": dict(expect=5), + }, }, # One check runs inside run_sp_tile_ngmix_Ng${SP_NGMIX_CHUNK}u per chunk, # so expect=1 is the correct per-chunk count. @@ -188,7 +202,7 @@ def check_counts(stage, run_dir): table = table[psf_model] except KeyError as exc: raise ValueError( - f"Invalid SP_PSF={psf_model!r}; expected one of psfex, mccd." + f"Invalid SP_PSF={psf_model!r}; expected one of {sorted(table)}." ) from exc details, ok = [], True for runner, spec in table.items(): diff --git a/workflow/scripts/container.py b/workflow/scripts/container.py index afc7019ed..c35f91ad8 100644 --- a/workflow/scripts/container.py +++ b/workflow/scripts/container.py @@ -59,14 +59,18 @@ class ContainerError(Exception): # The single source of truth for the fallback image: the workflow's own # `container:` key, which is also what the Snakefile reads. Written down once, # here, so the CLI and the workflow cannot disagree about the default. -CONFIG_FILE = Path(__file__).resolve().parents[1] / "config.yaml" +# SP_RUN_CONFIG replaces it for a run driven with its own config (bin/sp). +CONFIG_FILE = Path(os.environ.get("SP_RUN_CONFIG") + or Path(__file__).resolve().parents[1] / "config.yaml") CONFIG_KEY = "container" # The profile whose `apptainer-args:` every workflow job runs under. `exec` reads # it at runtime rather than restating it, so a one-off `sp container exec` and a # job see the same environment (the PYTHONPATH pin above all: a divergence there # means the one-off imports a different src/ than the workflow does). -PROFILE_FILE = Path(__file__).resolve().parents[2] / "profiles" / "nibi" / "config.yaml" +# SP_PROFILE names it (profiles//, default nibi), as it does for bin/sp. +PROFILE_FILE = (Path(__file__).resolve().parents[2] / "profiles" + / os.environ.get("SP_PROFILE", "nibi") / "config.yaml") # ~/.cache/shapepipe by default; SP_CACHE_DIR moves the whole cache (e.g. onto # a filesystem with room), XDG_CACHE_HOME moves it with everything else. From 4e17feee3e1246a4fdcdd8ebdb9a085ddefaf3ad Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Fri, 11 Sep 2026 21:03:07 +0200 Subject: [PATCH 23/85] workflow: image-sims merge column list, and the merge reads the workflow layout create_final_cat.py -I globbed run_sp_tile_Mc_*, the legacy bash runner's dated run dir; the workflow writes an undated run_sp_tile_Mc. Accept both. The overlay's final_cat.param was a symlink to the real-data list, which names columns make_cat does not write (IMAFLAGS_ISO: the tile SExtractor has no flag image; NGMIX_MOM_FAIL: written as NGMIX_MCAL_TYPES_FAIL) and omits NUMBER and NGMIX_MCAL_TYPES_FAIL, which sp_validation's image-sim extract reads. It is now a real file: the legacy example/cfis_image_sims list plus NGMIX_NEIGHBOUR_FLAG, checked column by column against a smoke-run final_cat. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01RSjZrhgsGCLAJCLb5Zjmc3 --- scripts/python/create_final_cat.py | 4 +- workflow/Snakefile | 3 +- .../config/cfis_image_sims/final_cat.param | 114 +++++++++++++++++- 3 files changed, 118 insertions(+), 3 deletions(-) mode change 120000 => 100644 workflow/config/cfis_image_sims/final_cat.param diff --git a/scripts/python/create_final_cat.py b/scripts/python/create_final_cat.py index 2b583b857..d98232818 100755 --- a/scripts/python/create_final_cat.py +++ b/scripts/python/create_final_cat.py @@ -391,7 +391,9 @@ def process(params): if params["image_sims"]: patch_name = params["patch"] - run_prefix = "run_sp_tile_Mc_*" + # The Snakemake workflow writes an undated run_sp_tile_Mc; the legacy + # bash runner wrote run_sp_tile_Mc_. + run_prefix = "run_sp_tile_Mc*" else: patch_name = rf"P{params['patch']}" run_prefix = "run_sp_tile_Mc_*" diff --git a/workflow/Snakefile b/workflow/Snakefile index 94e52f7f8..3162276b7 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -82,7 +82,8 @@ container: _image # ingestion stages is the same chain. `image_sims` runs the SKiLLS image # simulations through the real-data configs: the overlay config dir # (config/cfis_image_sims) holds real files only for the ingestion INIs whose -# input naming differs, and symlinks the rest into config/cfis, so an m-bias +# input naming differs, the fake-PSF INIs, and the merge column list +# (final_cat.param), and symlinks the rest into config/cfis, so an m-bias # measured on the sims calibrates the pipeline that makes the real catalogue. INPUT_TYPES = {"data", "image_sims"} INPUT_TYPE = config.get("input_type", "data") diff --git a/workflow/config/cfis_image_sims/final_cat.param b/workflow/config/cfis_image_sims/final_cat.param deleted file mode 120000 index c286947d0..000000000 --- a/workflow/config/cfis_image_sims/final_cat.param +++ /dev/null @@ -1 +0,0 @@ -../cfis/final_cat.param \ No newline at end of file diff --git a/workflow/config/cfis_image_sims/final_cat.param b/workflow/config/cfis_image_sims/final_cat.param new file mode 100644 index 000000000..5bb0bdd17 --- /dev/null +++ b/workflow/config/cfis_image_sims/final_cat.param @@ -0,0 +1,113 @@ +# Final-catalogue column selection for the image simulations. +# +# create_final_cat.py -I selects exactly these columns from each simulated +# tile's make_cat output (tiles///output/run_sp_tile_Mc/...) into +# final_cat_{sim}.hdf5, and sp_validation's extract step reads the same list; +# every entry must exist in that catalogue. +# +# A real file in this overlay dir, not a symlink into ../cfis/, because the +# real-data list names columns make_cat does not write for the simulations: +# - IMAFLAGS_ISO : the tile SExtractor runs without a flag image +# (config_tile_Sx.ini: FLAG_IMAGE = False) +# - NGMIX_MOM_FAIL : make_cat writes it as NGMIX_MCAL_TYPES_FAIL +# and it omits two that sp_validation's image-sim extract step reads: NUMBER and +# NGMIX_MCAL_TYPES_FAIL (the metacal moments-failure cut). + +# coordinates +XWIN_WORLD +YWIN_WORLD + +# tile ID, for plot of tile-dependent additive bias. +TILE_ID + +# SExtractor number +NUMBER + +# flags +FLAGS +NGMIX_MCAL_FLAGS +NGMIX_MCAL_TYPES_FAIL +NGMIX_NEIGHBOUR_FLAG + +# PSF ellipticity (original image PSF) +NGMIX_G1_PSF_ORIG_NOSHEAR +NGMIX_G2_PSF_ORIG_NOSHEAR + +# Number of epochs (exposures) +N_EPOCH +NGMIX_N_EPOCH + +## Shape measurement outputs +## Ngmix: model fitting + +# galaxy ellipticity +NGMIX_G1_1M +NGMIX_G2_1M +NGMIX_G1_1P +NGMIX_G2_1P +NGMIX_G1_2M +NGMIX_G2_2M +NGMIX_G1_2P +NGMIX_G2_2P +NGMIX_G1_NOSHEAR +NGMIX_G2_NOSHEAR +NGMIX_G1_ERR_NOSHEAR +NGMIX_G2_ERR_NOSHEAR + +# flags +NGMIX_FLAGS_1M +NGMIX_FLAGS_1P +NGMIX_FLAGS_2M +NGMIX_FLAGS_2P +NGMIX_FLAGS_NOSHEAR + +# size and error +NGMIX_T_1M +NGMIX_T_1P +NGMIX_T_2M +NGMIX_T_2P +NGMIX_T_NOSHEAR +NGMIX_T_ERR_1M +NGMIX_T_ERR_1P +NGMIX_T_ERR_2M +NGMIX_T_ERR_2P +NGMIX_T_ERR_NOSHEAR + +# reconvolved PSF size (used by the metacal size cut / Tpsf) +NGMIX_T_PSF_RECONV_1M +NGMIX_T_PSF_RECONV_1P +NGMIX_T_PSF_RECONV_2M +NGMIX_T_PSF_RECONV_2P +NGMIX_T_PSF_RECONV_NOSHEAR + +# flux and error +NGMIX_FLUX_1M +NGMIX_FLUX_1P +NGMIX_FLUX_2M +NGMIX_FLUX_2P +NGMIX_FLUX_NOSHEAR +NGMIX_FLUX_ERR_1M +NGMIX_FLUX_ERR_1P +NGMIX_FLUX_ERR_2M +NGMIX_FLUX_ERR_2P +NGMIX_FLUX_ERR_NOSHEAR + +# magnitudes +MAG_AUTO +MAGERR_AUTO +MAG_WIN +MAGERR_WIN +FLUX_AUTO +FLUXERR_AUTO +FLUX_APER +FLUXERR_APER +FLUX_RADIUS + +# SNR from SExtractor +SNR_WIN + +FWHM_IMAGE +FWHM_WORLD + +# PSF size measured on original image +NGMIX_T_PSF_ORIG_NOSHEAR From 716b160798362b75ee447709646a9fe4ae823d38 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Fri, 11 Sep 2026 21:22:37 +0200 Subject: [PATCH 24/85] workflow: campaign-unique node-local tile store for image sims; candide excludes The node-local tile store was /local/scratch/sp-, unique within a campaign but not across concurrent ones. The image-simulation shear branches are concurrent campaigns over the same tile IDs, and on candide /local/scratch is the node's shared /tmp: two branches' tile_shape jobs on n09 shared one store, one tile_vignets wiped it under the other's ngmix, and both failed. The failure can also be silent -- ngmix reading the other branch's vignets. image_sims now prefixes the store name with a hash of the run dir; data campaigns keep the bare name, so their shell commands (a rerun trigger) are unchanged. The tile_shape group's own slurm_extra (the --tmp floor) replaced the profile default and dropped candide's node excludes; its jobs ran on n09 and n17. profiles/candide restates both for the four members via set-resources (slurm_extra only; the rules keep their threads and attempt-scaled mem_mb, checked in a dry run). Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01RSjZrhgsGCLAJCLb5Zjmc3 --- profiles/candide/config.yaml | 14 ++++++++++++++ workflow/rules/tile.smk | 17 +++++++++++++++-- 2 files changed, 29 insertions(+), 2 deletions(-) diff --git a/profiles/candide/config.yaml b/profiles/candide/config.yaml index 67ea1bfe5..b4b1d63c2 100644 --- a/profiles/candide/config.yaml +++ b/profiles/candide/config.yaml @@ -32,3 +32,17 @@ rerun-incomplete: true # re-do jobs left incomplete by an unclean deat show-failed-logs: true printshellcmds: true rerun-triggers: [mtime, params, code, software-env] +# The tile_shape group (tile.smk) sets its own slurm_extra (TILE_SLURM_EXTRA, +# the --tmp floor), which REPLACES the default above and so drops the node +# excludes: its jobs landed on n09 and n17. Restate both, identically on all +# four members (a group's string resources must agree). Only slurm_extra is +# overridden; the rules keep their own threads and attempt-scaled mem_mb. +set-resources: + tile_vignets: + slurm_extra: "'--tmp=16000 --exclude=n09,n17,n23,n36'" + tile_ngmix: + slurm_extra: "'--tmp=16000 --exclude=n09,n17,n23,n36'" + tile_merge_cats: + slurm_extra: "'--tmp=16000 --exclude=n09,n17,n23,n36'" + tile_make_cat: + slurm_extra: "'--tmp=16000 --exclude=n09,n17,n23,n36'" diff --git a/workflow/rules/tile.smk b/workflow/rules/tile.smk index 68d6da9ba..8e3cf5171 100644 --- a/workflow/rules/tile.smk +++ b/workflow/rules/tile.smk @@ -130,6 +130,11 @@ is what downstream selections cut on. # `--rerun-triggers mtime code software-env`, accepting that clean_exposure's # consumer-set staleness detection (which rides on params) is off for that # invocation. +# Per-campaign prefix of the node-local store name; see tile_local(). +LOCAL_TAG = (hashlib.sha1(str(RUN_DIR).encode()).hexdigest()[:8] + "-" + if INPUT_TYPE == "image_sims" else "") + + def tile_local(tile): """The node-local prologue, as bash, for one tile. @@ -144,7 +149,15 @@ def tile_local(tile): the container (probe job 20798618) -- and the tile id is a wildcard snakemake substitutes at DAG time, so the shell string carries a concrete path with no `$` left for anything to escape. One group job per tile means - the name cannot collide; the sticky bit means nobody else can remove it. + the name cannot collide within a campaign; the sticky bit means nobody else + can remove it. Across campaigns it can: the image-simulation shear branches + are concurrent campaigns over the SAME tile IDs, and on candide + `/local/scratch` is the node's shared `/tmp`. Two branches' fused jobs on one + node then share one store -- a second `tile_vignets` wipes and rewrites it + under the first branch's ngmix, and the first `tile_make_cat`'s EXIT trap + deletes it under the second (seen on candide n09). So image_sims prefixes + the name with LOCAL_TAG, a hash of the run dir. Data campaigns keep the + bare tile name, so their shell commands -- a rerun trigger -- are unchanged. What we give up is Slurm's own cleanup of `$SLURM_TMPDIR`. TILE_VIGNET_FRESH reclaims a stale directory on the next attempt for the same tile, the trap @@ -201,7 +214,7 @@ if [ ! -d /local/scratch ]; then echo " profiles/nibi/config.yaml apptainer-args must carry --bind /local" >&2 exit 1 fi -export SP_LOCAL="/local/scratch/sp-{tile}" +export SP_LOCAL="/local/scratch/sp-{LOCAL_TAG}{tile}" export SP_VIGNET_OUT="$SP_LOCAL/output" export NGMIX_VIGNET_DIR="$SP_LOCAL/output/run_sp_tile_PiViVi" export SP_WCS_DIR="$SP_LOCAL/wcs" From fdc9db408b10e6444648cf435482e4af76826149 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Sat, 12 Sep 2026 00:32:05 +0200 Subject: [PATCH 25/85] workflow: port the MCCD exposure chain to the workflow grammar config_exp_mccd.ini was never ported: it read split_exp and mask_runner outputs through INPUT_MODULE (mask_runner is gone since #847, and with it the pipeline_flag files), had no mask_query, chained setools and MCCD through last:, and left RUN_DATETIME at its default, so the run dir was dated and nothing downstream could find it. psf_model: mccd could not run. It is now the PSFEx chain up to the star selection -- same split-CCD inputs, flag image, mask_query and setools, so both models fit the same stars -- followed by mccd_preprocessing (all CCDs of the exposure into one training and one test catalogue) and mccd_fit_val (fitted_model-.npy, which the tiles' mccd_interp reads, plus validation_psf-.fits). merge_starcat and mccd_plots are dropped: they are campaign-level diagnostics, not per-exposure products. The completeness table's MCCD branch had preprocessing at 80 (per CCD; it is 2 per exposure), listed merge_starcat and the plots, and was all warning-only. It now carries the counts this chain produces, mandatory as for PSFEx: an exposure without a model has no PSF on any tile it overlaps. Checked on SKiLLS star sim 1z2z (exposure 2086792): 120/40/80/2 through preprocessing. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01VSU9pPZjpBjzWorAL88aKQ --- workflow/config/cfis/config_exp_mccd.ini | 155 ++++++++++------------- workflow/scripts/completeness.py | 26 ++-- 2 files changed, 80 insertions(+), 101 deletions(-) diff --git a/workflow/config/cfis/config_exp_mccd.ini b/workflow/config/cfis/config_exp_mccd.ini index 635952ca7..9a6a6f7bf 100644 --- a/workflow/config/cfis/config_exp_mccd.ini +++ b/workflow/config/cfis/config_exp_mccd.ini @@ -1,5 +1,8 @@ -# ShapePipe configuration file for single-exposures, MCCD PSF model. -# Process exposures after masking, from star detection to PSF model. +# ShapePipe configuration file for single-exposures. MCCD PSF model. +# Process exposures after splitting, from star detection to PSF model. +# ShapePipe generates no masks: SExtractor reads the instrument flag image +# delivered with the exposure, and mask_query flags detections against the +# external healsparse maps (see [MASK_QUERY_RUNNER] below). ## Default ShapePipe options @@ -12,16 +15,15 @@ VERBOSE = True RUN_NAME = run_sp_exp_SxSePsf # Add date and time to RUN_NAME, optional, default: True -; RUN_DATETIME = False +RUN_DATETIME = False ## ShapePipe execution options [EXECUTION] # Module name, single string or comma-separated list of valid module runner names -MODULE = sextractor_runner, setools_runner, - mccd_preprocessing_runner, mccd_fit_val_runner, - merge_starcat_runner, mccd_plots_runner +MODULE = sextractor_runner, mask_query_runner, setools_runner, + mccd_preprocessing_runner, mccd_fit_val_runner # Run mode, SMP or MPI MODE = SMP @@ -47,7 +49,7 @@ OUTPUT_DIR = $SP_RUN/output [JOB] # Batch size of parallel processing (optional), default is 1, i.e. run all jobs in serial -SMP_BATCH_SIZE = 1 +SMP_BATCH_SIZE = 8 # Timeout value (optional), default is None, i.e. no timeout limit applied TIMEOUT = 96:00:00 @@ -57,16 +59,15 @@ TIMEOUT = 96:00:00 [SEXTRACTOR_RUNNER] -# Somehow this works but not -# - omitting -# - $SP_RUN/output -#INPUT_DIR = . +# The split CCDs, and nothing else: ShapePipe generates no masks +INPUT_DIR = $SP_RUN/output/run_sp_exp_Sp/split_exp_runner/output -# Input from two modules -INPUT_MODULE = split_exp_runner, mask_runner +# Read the instrument flag image split_exp wrote per CCD +FILE_PATTERN = image, weight, flag -# Read pipeline flag files created by mask module -FILE_PATTERN = image, weight, pipeline_flag +# Explicit extensions: a 3-entry FILE_PATTERN override must not fall back on +# the decorator's 4-entry FILE_EXT default (length check fails at startup) +FILE_EXT = .fits, .fits, .fits NUMBERING_SCHEME = -0000000-0 @@ -127,52 +128,72 @@ SUFFIX = sexcat MAKE_POST_PROCESS = FALSE -[SETOOLS_RUNNER] +[MASK_QUERY_RUNNER] -INPUT_MODULE = last:sextractor_runner +INPUT_DIR = $SP_RUN/output/run_sp_exp_SxSePsf/sextractor_runner/output -# Note: Make sure this doe not match the SExtractor background images +# Note: Make sure this does not match the SExtractor background images # (sexcat_background*) -FILE_PATTERN = sexcat_sexcat +FILE_PATTERN = sexcat NUMBERING_SCHEME = -0000000-0 -# SETools config file -SETOOLS_CONFIG_PATH = $SP_CONFIG/star_selection.setools - - -[MCCD_PREPROCESSING_RUNNER] +# The PSF-star diet, and it is deliberately NARROW: instrument flags (read by +# SExtractor as IMAFLAGS_ISO) plus the healsparse star-body map (UNIONS bit 2), +# and nothing else. Halo bits 0 and 1 are excluded on purpose — halos flag +# objects for the final catalogue, they do not reject PSF stars (mask-force +# telecon, 2026-07-21) — and MaxiMask is not in the diet either. Widen it by +# adding paths here; every map that is True (boolean) or nonzero (integer) at a +# detection sets MASK_EXT. Comma-separated. +# +# UNSET BY DEFAULT, and the module is a strict no-op without it: the catalogue +# is passed through with no MASK_EXT column, the same gating make_cat gives +# MASK_EXT_PATHS. mask_query stays in the MODULE chain either way, so turning +# the query on is uncommenting one line and never editing the chain. +# +# Even with it set, NOTHING CUTS ON IT: the star selection ships permissive +# (instrument flags only) and the column is carried for transparency and +# measurement — see star_selection.setools' header for the one-line change +# that would impose it. +; MASK_PATHS = $SP_CONFIG/mask_ugriz_nside131072_n4.hsp + +# Optional: restrict integer maps to these bits (value & MASK_BITS). Absent, +# any nonzero value flags. Boolean maps — the UNIONS per-bit products, one map +# per bit — ignore it, which is why the diet above is a path list and not a +# bit mask. +; MASK_BITS = 4 -# Path to MCCD config file -CONFIG_PATH = $SP_CONFIG/config_MCCD.ini -MODE = FIT_VALIDATION +[SETOOLS_RUNNER] -VERBOSE = False +INPUT_DIR = $SP_RUN/output/run_sp_exp_SxSePsf/mask_query_runner/output -INPUT_DIR = last:setools_runner +FILE_PATTERN = sexcat_ext -# Input are individual CCDs, thus single-exposure single-HDU images NUMBERING_SCHEME = -0000000-0 -FILE_PATTERN = star_split_ratio_80, star_split_ratio_20 - -FILE_EXT = .fits, .fits - - -[MCCD_FIT_VAL_RUNNER] +# SETools config file +SETOOLS_CONFIG_PATH = $SP_CONFIG/star_selection.setools -# Path to MCCD config file -CONFIG_PATH = $SP_CONFIG/config_MCCD.ini -MODE = FIT_VALIDATION +# The MCCD chain differs from the PSFEx one only from here on; everything up +# to the star selection is the same, so the two PSF models are fitted on the +# same stars. Every module below chains from the one before it in this run +# (the decorators' input_module), so none needs an INPUT_DIR. +# +# Not in this chain: merge_starcat_runner and mccd_plots_runner. Both are +# campaign-level diagnostics over all exposures' validation stars, not +# per-exposure products; the tiles need only the fitted model. -VERBOSE = False +[MCCD_PREPROCESSING_RUNNER] -NUMBERING_SCHEME = -0000000 +# Serial runner: it receives every CCD's star catalogues of this exposure and +# writes one focal-plane training (80%) and one test (20%) catalogue for it. +FILE_PATTERN = star_split_ratio_80, star_split_ratio_20 +FILE_EXT = .fits, .fits -[MCCD_MERGE_STARCAT_RUNNER] +NUMBERING_SCHEME = -0000000-0 # Path to MCCD config file CONFIG_PATH = $SP_CONFIG/config_MCCD.ini @@ -181,10 +202,13 @@ MODE = FIT_VALIDATION VERBOSE = False -NUMBERING_SCHEME = -0000000 +[MCCD_FIT_VAL_RUNNER] -[MCCD_PLOTS_RUNNER] +# One focal-plane model per exposure: fitted_model-.npy, which the tiles' +# mccd_interp_runner reads (config_tile_PiViVi_mccd.ini, ME_DOT_PSF_PATTERN), +# and validation_psf-.fits, the model at the 20% test stars. +NUMBERING_SCHEME = -0000000 # Path to MCCD config file CONFIG_PATH = $SP_CONFIG/config_MCCD.ini @@ -192,46 +216,3 @@ CONFIG_PATH = $SP_CONFIG/config_MCCD.ini MODE = FIT_VALIDATION VERBOSE = False - -# Now MCCD has created a focal-plane PSF model, including all CCDS per images, -# thus single-exposure files -NUMBERING_SCHEME = -0000000 - -PLOT_MEANSHAPES = True - -# X_GRID, Y_GRID: correspond to the number of bins in each direction of each -# CCD from the focal plane. Ex: each CCD will be binned in 5x10 regular grids. -X_GRID = 5 -Y_GRID = 10 - -PLOT_HISTOGRAMS = True - -# REMOVE_OUTLIERS: Remove validated stars that are outliers in terms of shape -# before drawing the plots. -REMOVE_OUTLIERS = False - - - -[MCCD_INTERP_RUNNER] - -# MODE: Define the way the MCCD interpolation will run. -# CLASSIC for classical run. -# MULTI-EPOCH for multi epoch. -MODE = CLASSIC - -# Position parameter names -# for multi-epoch XWIN_WORLD,YWIN_WORLD -# For classical XWIN_IMAGE,YWIN_IMAGE: -POSITION_PARAMS = XWIN_IMAGE,YWIN_IMAGE - -# Get PSF shapes calculated and saved on the output dict -GET_SHAPES = True - -# Directory with PSF models -PSF_MODEL_DIR = $SP_RUN/output - -# PSF model patterns -PSF_MODEL_PATTERN = fitted_model - -# PSF model separator -PSF_MODEL_SEPARATOR = - diff --git a/workflow/scripts/completeness.py b/workflow/scripts/completeness.py index 0a22179f5..00a9d09b0 100644 --- a/workflow/scripts/completeness.py +++ b/workflow/scripts/completeness.py @@ -102,21 +102,19 @@ "psfex_runner": dict(expect=80), "psfex_interp_runner": dict(expect=40, warn=True), }, - # MCCD counts are derived from config_exp_mccd.ini and its per-CCD - # runners, but this chain has not been exercised through this workflow. - # Keep the expected counts visible while making the unverified branch - # warning-only until a real campaign validates its counts. + # MCCD shares the chain up to setools with PSFEx, then fits one + # focal-plane model per exposure. Preprocessing is a serial runner that + # merges the 40 CCDs' split catalogues into one training and one test + # catalogue; fit_val writes the model (fitted_model-.npy, what the + # tiles interpolate) and its validation catalogue. Both are + # exposure-wide and all-or-nothing: an exposure without a model has no + # PSF on any tile it overlaps. "mccd": { - "sextractor_runner": dict(expect=120, warn=True), - "setools_runner": dict(expect=80, warn=True, - subpath="rand_split"), - "mccd_preprocessing_runner": dict(expect=80, warn=True), - # Fit/validation is exposure-wide: one model and one validation - # catalogue, unlike the per-CCD preprocessing outputs. - "mccd_fit_val_runner": dict(expect=2, warn=True), - "merge_starcat_runner": dict(expect=1, warn=True), - # config_exp_mccd enables the ten meanshape and six histogram plots. - "mccd_plots_runner": dict(expect=16, warn=True), + "sextractor_runner": dict(expect=120), + "mask_query_runner": dict(expect=40), + "setools_runner": dict(expect=80, subpath="rand_split"), + "mccd_preprocessing_runner": dict(expect=2), + "mccd_fit_val_runner": dict(expect=2), }, # Image simulations with the true PSF (psf_model: fake): no PSF fit on # the exposures, only the SExtractor pass whose background/background_rms From 836f91dcad7eaf7c14235f85321e68356b35d529 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Sat, 12 Sep 2026 02:19:29 +0200 Subject: [PATCH 26/85] workflow: exp_psf reserves 2 cores, not 8, for the MCCD chain MCCD fits one focal-plane model per exposure after the per-CCD stages: 85 min single-threaded for ~2500 stars and 8.3 GB (SKiLLS star sim 1z2z_1, exposure 2086792), and 48 min on 8 BLAS threads -- 1.75x for 8x the cores, which the thread caps forbid anyway. With threads: 8 the fit leaves 7 reserved cores idle for its whole length. The mccd chain now takes 2: the per-CCD stages run 2 wide (minutes), the fit is unchanged, and an exposure reserves about a quarter of the core-hours. PSFEx and fake keep 8. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01VSU9pPZjpBjzWorAL88aKQ --- workflow/rules/exposure.smk | 9 ++++++++- 1 file changed, 8 insertions(+), 1 deletion(-) diff --git a/workflow/rules/exposure.smk b/workflow/rules/exposure.smk index 4e1ef45a4..bc35e70b9 100644 --- a/workflow/rules/exposure.smk +++ b/workflow/rules/exposure.smk @@ -81,6 +81,13 @@ rule exp_split: # -> psfex_interp, per CCD. # setools may reject a sparse CCD (~0.2% attrition) — tolerated by the floor's # :warn on psfex_interp_runner. +# +# MCCD instead fits ONE focal-plane model per exposure after the per-CCD +# stages: ~85 min single-threaded for ~2500 stars (SKiLLS star sim, 8.3 GB), +# and only 1.75x faster on 8 BLAS threads (which the thread caps forbid anyway). +# With 8 cores reserved the fit would idle 7 of them for its whole length, so +# the mccd chain takes 2: the per-CCD stages run 2 wide (minutes), the fit is +# unchanged, and the exposure reserves a quarter of the core-hours. rule exp_psf: input: rules.exp_split.output.manifest @@ -91,7 +98,7 @@ rule exp_psf: params: pre = lambda wc: unit_pre("exp_psf", wc.exp), script_hash = SCRIPT_HASH - threads: 8 + threads: 2 if PSF_MODEL == "mccd" else 8 retries: 2 benchmark: # BESIDE manifests/, not inside it: clean_exposure deletes manifests/ From ad80f23fc0ac474c24cb48503663f48ed11c7ec8 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Sat, 12 Sep 2026 02:58:03 +0200 Subject: [PATCH 27/85] mccd_interp: test for N_EPOCH among the column names Multi-epoch MCCD interpolation checked `"N_EPOCH" not in cat.get_data()`, a membership test against the FITS record array, which current numpy rejects ("Cannot compare structured or void to non-void arrays"): every tile failed in tile_vignets with psf_model: mccd. psfex_interp already tests `.dtype.names`; so does this now. With it, tile 233.293 of SKiLLS star sim 1z2z_1 interpolates the seven exposures' MCCD models (a few sparse CCDs lose an epoch, as intended). Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01VSU9pPZjpBjzWorAL88aKQ --- src/shapepipe/modules/mccd_package/mccd_interpolation_script.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/src/shapepipe/modules/mccd_package/mccd_interpolation_script.py b/src/shapepipe/modules/mccd_package/mccd_interpolation_script.py index b6183f4e1..0942d93ac 100644 --- a/src/shapepipe/modules/mccd_package/mccd_interpolation_script.py +++ b/src/shapepipe/modules/mccd_package/mccd_interpolation_script.py @@ -442,7 +442,7 @@ def _interpolate_me(self): all_id = np.copy(cat.get_data()["NUMBER"]) key_ne = "N_EPOCH" - if key_ne not in cat.get_data(): + if key_ne not in cat.get_data().dtype.names: raise KeyError( f"Key {key_ne} not found in input galaxy catalogue, needed for" + " PSF interpolation to multi-epoch data; run previous module" From 2082dcf8d23fe3fb6123eeeec59e7d5e106a5668 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Sat, 12 Sep 2026 03:57:46 +0200 Subject: [PATCH 28/85] workflow: MCCD validated through the full chain; drop the warn-only flags With the ported exposure chain, the mccd fixes and the mccd_interp fix, psf_model: mccd ran SKiLLS star sim 1z2z_1 tile 233.293 end to end: seven exposures' focal-plane models (120/40/80/2/2 each), mccd_interp's galaxy_psf store, vignets, ngmix and a final_cat of 21857 objects. tile_vignets' mccd counts are now mandatory as for psfex, and the README no longer calls mccd "wired but unvalidated". Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01VSU9pPZjpBjzWorAL88aKQ --- workflow/README.md | 3 ++- workflow/scripts/completeness.py | 10 +++++----- 2 files changed, 7 insertions(+), 6 deletions(-) diff --git a/workflow/README.md b/workflow/README.md index 97510c901..d404d1b16 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -25,7 +25,8 @@ uv pip install 'snakemake>=9,<10' 'snakemake-executor-plugin-slurm>=2.7,<3' # Edit workflow/config.yaml: tile_list, inputs.tiles/exposures, outputs.run_dir, # outputs.products_dir/index_db, and container. -# `psf_model` is `psfex` or `mccd`; mccd is wired but unvalidated here, while psfex is exercised by smk-g4 through smk-g6. +# `psf_model` is `psfex` or `mccd`. psfex is exercised by smk-g4 through smk-g6; mccd has run the full chain on +# an image-sim star tile (one focal-plane model per exposure, ~1.5 CPU-hours each). # The committed launcher loads apptainer/1.4.5 + the /project venv, so a # fresh shell always has the right state. diff --git a/workflow/scripts/completeness.py b/workflow/scripts/completeness.py index 00a9d09b0..e400e14c6 100644 --- a/workflow/scripts/completeness.py +++ b/workflow/scripts/completeness.py @@ -136,12 +136,12 @@ # v2.0's 4 was the canfar flavor. every vignette feeds ngmix, so the expected count is all-or-nothing. "vignetmaker_runner_run_2": dict(expect=5), }, - # MCCD is wired but unvalidated here; retain the expected runner names - # and counts as warnings until a workflow campaign exercises them. + # As psfex: mccd_interp writes the tile's galaxy_psf store from the + # exposures' focal-plane models (SKiLLS star sim 1z2z_1, 233.293). "mccd": { - "mccd_interp_runner": dict(expect=1, warn=True), - "vignetmaker_runner_run_1": dict(expect=1, warn=True), - "vignetmaker_runner_run_2": dict(expect=5, warn=True), + "mccd_interp_runner": dict(expect=1), + "vignetmaker_runner_run_1": dict(expect=1), + "vignetmaker_runner_run_2": dict(expect=5), }, # Image simulations: fake_interp_runner writes the same galaxy_psf # sqlite psfex_interp_runner writes, from the simulation's PSF dictionary. From 9085cd0b877abf9c4c8a61f87fb702aff30b37bf Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Sat, 12 Sep 2026 10:21:03 +0200 Subject: [PATCH 29/85] workflow: tile_store_root run-config key moves the tile store's bind, per campaign bin/sp rewrites the source of the `/local/scratch` bind in the launch snapshot's profile (appending one when the profile has none) and creates the directory. tile_local() is untouched: its output is a params rerun trigger, a bind is not, so the key can be set on a resume. candide's 31 GB node /tmp cannot hold dense image-sim tile stores; image-sim campaigns there point this at NFS. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01Qv5MwRiVKa8wnH665dWBia --- profiles/candide/config.yaml | 5 +++-- workflow/README.md | 6 ++++++ workflow/bin/sp | 36 ++++++++++++++++++++++++++++++++++++ workflow/config.yaml | 9 +++++++++ workflow/rules/tile.smk | 4 +++- 5 files changed, 57 insertions(+), 3 deletions(-) diff --git a/profiles/candide/config.yaml b/profiles/candide/config.yaml index b4b1d63c2..bacf48071 100644 --- a/profiles/candide/config.yaml +++ b/profiles/candide/config.yaml @@ -22,8 +22,9 @@ software-deployment-method: [apptainer] # which this line cannot name: snakemake passes it through literally, and # --cleanenv drops SLURM_JOB_ID inside the container. So the node's /tmp (local # disk, ~29 GB free, sticky) is bound there instead. That holds ~5 concurrent -# tile stores per node; if a campaign packs more tile_shape groups onto one node, -# this is the line to revisit. The PYTHONPATH pin is rewritten by bin/sp to the +# tile stores per node, too few for dense image-sim tiles: such a campaign sets +# `tile_store_root:` in its run config, and bin/sp replaces this bind's source in +# the launch snapshot (see workflow/config.yaml). The PYTHONPATH pin is rewritten by bin/sp to the # launch snapshot's src/; the value here only has to be a checkout's src/. apptainer-args: "--cleanenv --env OMP_NUM_THREADS=1 --env MALLOC_ARENA_MAX=2 --env PYTHONPATH=/n17data/mkilbing/astro/repositories/github/shapepipe/src --bind /home,/automnt,/n17data,/n23data1,/n09data --bind /tmp:/local/scratch" latency-wait: 60 # NFS: wait for outputs to appear after a job diff --git a/workflow/README.md b/workflow/README.md index d404d1b16..43a7d018d 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -74,6 +74,12 @@ key falls back to the committed run's paths) and is snapshotted with the code; `SP_PROFILE` picks `profiles//`. sp_validation's image-simulation workflow drives these campaigns and measures m from their final catalogues. +On candide, the node-local tile store (bound from the node's 31 GB `/tmp`) does not +hold several dense image-sim tiles at once. Set `tile_store_root:` in the run +config to a shared directory; `sp run` binds it to `/local/scratch` for that +campaign. The store names carry a per-campaign hash, so all branches can share one +root. + ## The container image `sp container` owns which image the jobs run inside. Two layers, and the second diff --git a/workflow/bin/sp b/workflow/bin/sp index 37eefa1e6..b64905181 100755 --- a/workflow/bin/sp +++ b/workflow/bin/sp @@ -135,6 +135,42 @@ if not n: f.write_text(text) PY + # `tile_store_root:` in the run config moves the node-local tile store + # (tile.smk's tile_local(): the vignette + WCS stores of each tile_shape group) + # off the profile's `/local/scratch` bind, for this campaign only. It is the + # same substitution as the pin above -- the BIND SOURCE of `/local/scratch` in + # the snapshot's profile -- and deliberately not a path in tile_local(): that + # function's output is a params rerun trigger, so a path there would reschedule + # every finished tile of a resumed campaign (its warning block says how badly), + # while a bind is not a trigger at all (a dry run after changing it schedules + # only the missing tiles). Why a campaign wants it: candide binds the node's + # 31 GB /tmp, which a few dense image-sim tile stores or failed groups' orphans + # fill, after which every later group on that node fails in seconds. A shared + # directory (NFS) trades ngmix read latency for room; the campaign-unique store + # names (LOCAL_TAG) keep concurrent campaigns apart in one root. Profiles with + # no `X:/local/scratch` bind (nibi's `--bind /local`) get one appended. + python - "$SNAPSHOT/profiles/$PROFILE/config.yaml" "$CONFIG" <<'PY' +import os, pathlib, re, sys, yaml +f, run_config = pathlib.Path(sys.argv[1]), sys.argv[2] +root = (yaml.safe_load(open(run_config)) or {}).get("tile_store_root") +if root: + root = str(root) + if not os.path.isabs(root): + sys.exit(f"sp: tile_store_root must be an absolute path, got {root!r}") + os.makedirs(root, exist_ok=True) + bind = f"--bind {root}:/local/scratch" + text, n = re.subn(r'--bind \S+:/local/scratch(?=[\s"])', lambda m: bind, + f.read_text(), count=1) + if not n: + text, n = re.subn(r'(apptainer-args:\s*"[^"]*)"', + lambda m: m.group(1) + " " + bind + '"', text, count=1) + if not n: + sys.exit("sp: tile_store_root is set but the profile has no quoted " + "apptainer-args to bind it with") + f.write_text(text) + print(f"sp: tile store root {root} (bound to /local/scratch)", file=sys.stderr) +PY + python - "$SNAPSHOT/snapshot.json" "$REPO" <<'PY' import json, pathlib, subprocess, sys, time out, repo = pathlib.Path(sys.argv[1]), sys.argv[2] diff --git a/workflow/config.yaml b/workflow/config.yaml index 336cda5d6..58dde3d5b 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -144,6 +144,15 @@ clean_ignore_tiles: [] # per-tile, in-job, from the tile's own sexcat). ngmix_chunks: 8 +# Where the node-local tile store lives (tile.smk's tile_local(): each tile_shape +# group's vignette and WCS stores). Unset means the profile's `/local/scratch` +# bind: nibi's /local NVMe, candide's 31 GB node /tmp. Set it to an absolute +# directory to bind that instead, for this campaign only; bin/sp rewrites the +# snapshot's profile and creates the directory. Not a rerun trigger, so it may be +# set on a resume. Image-sim campaigns on candide want a shared directory (dense +# tiles and failed groups' orphans fill /tmp); the cost is ngmix read latency. +# tile_store_root: /n23data1//shapepipe/tilestore + # The container's installed shapepipe is overridden by the prod worktree via # --env PYTHONPATH in profiles/nibi (settled call 3); no config knob here. # build_index.py's missing-tile fraction threshold is passed by workflow/bin/sp diff --git a/workflow/rules/tile.smk b/workflow/rules/tile.smk index 8e3cf5171..ccbc83fbf 100644 --- a/workflow/rules/tile.smk +++ b/workflow/rules/tile.smk @@ -91,7 +91,9 @@ is what downstream selections cut on. # # HOW THE PATH GETS IN HERE. Not through the environment: profiles/nibi passes # --bind /local and tile_local() below DERIVES the path from the tile wildcard. -# Why nothing can be communicated instead is on that profile line. +# Why nothing can be communicated instead is on that profile line. A campaign +# that needs the store elsewhere (candide's 31 GB /tmp) moves the BIND, via the +# run config's `tile_store_root` (bin/sp), never the path here -- see below. # # THE COST WE ACCEPT: a failure anywhere in the tile re-runs the WHOLE tile, # not one chunk, because the store dies with the job. At ~1 h per fused tile From 7637b390527a5763366f1078b387630e8079500f Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Tue, 15 Sep 2026 10:36:48 +0200 Subject: [PATCH 30/85] completeness.py: atomic write for stage log/manifest JSON write_if_changed() used Path.write_text(), which truncates before writing. run_report.py reads every logs/*.json right after compute finishes (the onsuccess hook) and can catch one mid-write as an empty file -- observed during the grid_3 image-sims acceptance test (exp_get_images.json for one exposure came back "Expecting value: line 1 column 1"). Write to a same-directory temp file and os.replace() it into place instead, which is atomic and closes the race. --- workflow/scripts/completeness.py | 17 +++++++++++++++-- 1 file changed, 15 insertions(+), 2 deletions(-) diff --git a/workflow/scripts/completeness.py b/workflow/scripts/completeness.py index e400e14c6..bc8975064 100644 --- a/workflow/scripts/completeness.py +++ b/workflow/scripts/completeness.py @@ -354,10 +354,23 @@ def build_manifest(stage, run_dir, unit, stage_subdir=None): def write_if_changed(path: Path, text: str) -> None: - """Write only when the bytes differ — see the module docstring on mtime.""" + """Write only when the bytes differ — see the module docstring on mtime. + + Via a same-directory temp file + ``os.replace``, not ``path.write_text``: + the latter truncates before writing, so a reader (``run_report.py``, run + automatically at compute's end) can catch the file empty mid-write. Same + directory keeps the replace on one filesystem, which is what makes it + atomic; a losing writer's temp file is unlinked rather than left behind. + """ path.parent.mkdir(parents=True, exist_ok=True) if not path.exists() or path.read_text() != text: - path.write_text(text) + tmp = path.with_name(f".{path.name}.tmp{os.getpid()}") + try: + tmp.write_text(text) + os.replace(tmp, path) + except BaseException: + tmp.unlink(missing_ok=True) + raise def _unit_from_run_dir(run_dir): From 7aababf93ad70942fe5ba9f5c6392c8bf9764767 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Tue, 15 Sep 2026 16:25:41 +0200 Subject: [PATCH 31/85] workflow/config.yaml: universal template, no cluster-specific paths The committed default carried one person's real nibi paths (tile_list, inputs, container, outputs) baked in. That means `sp run` was not actually the same command on every machine: anyone who forgot to set SP_RUN_CONFIG on a different cluster would either hit a confusing filesystem error somewhere downstream, or silently point at paths that happen to also resolve on their machine but belong to someone else's campaign. config.yaml now carries a REPLACE_ME sentinel for every machine- or campaign-specific path, with a top-of-file explanation that a real run always supplies its own via SP_RUN_CONFIG. The Snakefile checks the keys that have no safe fallback (tile_list, inputs.tiles/exposures, outputs.run_dir/index_db) right after the config file is read and raises one clear, identical error on any cluster if one is still the placeholder, rather than failing confusingly later or not at all. `container:` is left out of that check: a local sandbox or cached SIF is resolved before it, so a placeholder there is often harmless. Also cleans up several stale comments left over from earlier campaign states (smk-g4/g5 narrative, `clean`/`clean_tiles` comments describing "OFF" for a value that had already been changed to `true`) and two "knob" -> "setting" wording fixes. --- workflow/Snakefile | 25 ++++++++++++++ workflow/config.yaml | 78 +++++++++++++++++++++++++------------------- 2 files changed, 70 insertions(+), 33 deletions(-) diff --git a/workflow/Snakefile b/workflow/Snakefile index 3162276b7..a6a51b92e 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -52,6 +52,31 @@ _RUN_CONFIG = os.environ.get("SP_RUN_CONFIG") or str( Path(workflow.snakefile).parent / "config.yaml") configfile: _RUN_CONFIG +# The committed config.yaml is a universal template (no machine- or +# campaign-specific path in it), so the same `sp run` invocation works on +# every cluster: whoever runs it always supplies their own paths through +# SP_RUN_CONFIG (see config.yaml's own top comment and the README). Left +# unset, the template's REPLACE_ME sentinel would otherwise read as a real, +# oddly-named path -- silently missing tiles, or a job dying hours in on a +# node that can't see "/REPLACE_ME" -- rather than the one-line, same-message- +# everywhere error this check gives instead. `container:` is checked +# separately (container.py resolves a local sandbox or cached SIF first, so a +# placeholder there is often harmless) and so is left out here. +_PLACEHOLDER = "REPLACE_ME" +for _key, _value in { + "tile_list": config.get("tile_list"), + "inputs.tiles": (config.get("inputs") or {}).get("tiles"), + "inputs.exposures": (config.get("inputs") or {}).get("exposures"), + "outputs.run_dir": (config.get("outputs") or {}).get("run_dir"), + "outputs.index_db": (config.get("outputs") or {}).get("index_db"), +}.items(): + if _value == _PLACEHOLDER: + raise WorkflowError( + f"{_key!r} in {_RUN_CONFIG} is still the template placeholder " + f"{_PLACEHOLDER!r}. workflow/config.yaml ships as a universal " + f"template with no real paths in it; point SP_RUN_CONFIG at your " + f"own run config instead (see the README, 'Run configuration').") + # Every job's shell runs inside this container (apptainer software-deployment in # the profile); the user never types apptainer. WHICH image is the one resolution # order `sp container` exposes — this user's writable sandbox if they built one, diff --git a/workflow/config.yaml b/workflow/config.yaml index 58dde3d5b..31c45c038 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -1,32 +1,41 @@ # Run configuration for the ShapePipe Snakemake workflow. # +# THIS FILE IS A UNIVERSAL TEMPLATE, NOT A LIVE CAMPAIGN'S CONFIG. It carries +# no machine- or campaign-specific path, on purpose: `sp run` must be the same +# command on nibi, candide, or anywhere else, and a committed default that +# happened to bake in one person's real paths would silently run a stranger's +# campaign against the wrong cluster's filesystem the moment they forgot to +# override it. Every path below is REPLACE_ME; the Snakefile refuses to parse +# past one of the ones it cannot otherwise fail clearly on later (tile_list, +# inputs.tiles, inputs.exposures, outputs.run_dir, outputs.index_db -- see the +# check near the top of Snakefile). `container:` is exempt from that check: a +# local sandbox or cached SIF (`sp container status`) is resolved first and +# makes this key irrelevant for anyone who has one. +# +# A real run always supplies its own paths via SP_RUN_CONFIG, which REPLACES +# this file wholesale rather than layering over it (workflow/bin/sp; see the +# README's "Run configuration" section): +# +# SP_PROFILE= SP_RUN_CONFIG=/path/to/your_run.yaml \ +# workflow/bin/sp run +# # A "run" is declared by a tile list plus the paths below. Everything here is # read at parse time; none of it is a rule input, so editing it (e.g. appending # tiles) never invalidates completed work — it only changes which jobs exist. -# The tile list that scopes this run (one "IDra.IDdec" per line). -# The campaign grows by appending to this file — parse-time config, so -# completed work is never invalidated. Current contents: smk-g5, ONE tile -# (186.307), the paired control for the two fixes landed after smk-g4 — the -# node-local WCS sqlite and tile_ngmix's mem_mb 14000 -> 5000. That tile ran -# ALONE as smk-g4 job 20799387 (elapsed 8063 s, chunk median 6813 s, chunk mean -# load 59%, MaxRSS 26.6 GiB), so the baseline carries no concurrency confound -# and the contrast is same tile, same solitude, two commits apart. -# The 34-tile smk-g4 set is preserved at smk-g4/tiles34.txt. -# (The 210/211 quad used earlier is unusable: its tile images are symlinks into -# anaennis' moved processed_tiles tree — ~8.5k of the 10.3k staged tiles are -# broken.) -tile_list: /project/def-mjhudson/cdaley/sp-products/smk-g6/tiles.txt +# The tile list that scopes this run (one "IDra.IDdec" per line). The +# campaign grows by appending to this file — parse-time config, so completed +# work is never invalidated. +tile_list: REPLACE_ME # Inputs are the only site-specific data paths; all committed configs consume these # through SP_INPUT_TILES and SP_INPUT_EXPOSURES. inputs: - # Pre-staged P3 data (get_images RETRIEVE=symlink). - tiles: /project/def-mjhudson/unions-wl/tiles - exposures: /project/def-mjhudson/unions-wl/exposures + tiles: REPLACE_ME + exposures: REPLACE_ME # The container every job runs inside (apptainer software-deployment in the profile). -container: /project/def-mjhudson/cdaley/containers/shapepipe-develop-runtime.sif +container: REPLACE_ME # PSF model used by the exposure and tile interpolation stages: psfex, mccd, or # fake (image simulations only: the true PSF from `psf_dict`, no exposure fit). @@ -46,7 +55,7 @@ input_type: data # sharded per-unit stores live under it: # /tiles/<2-char prefix>// and /exp/// outputs: - run_dir: /scratch/cdaley/shapepipe-output/smk-g6 + run_dir: REPLACE_ME # products_dir is the PERSISTENT root: the durable, low-volume products — the # final catalogues (/tiles///final_cat-.fits, @@ -59,31 +68,34 @@ outputs: # tile re-declare inputs against exposure stores reclamation already deleted. # # Unset means "one root": products land under run_dir, exactly the pre-D5 -# layout, which is what a fixture or smoke test wants. +# layout, which is what a fixture or smoke test wants. Delete this whole key +# (not REPLACE_ME) if that's what you want -- it's the one path in this file +# that's genuinely optional rather than templated. # # Snakemake's own state is the one durable-looking thing that stays on scratch # (-state; bin/sp explains why). - products_dir: /project/def-mjhudson/cdaley/sp-products/smk-g6 + products_dir: REPLACE_ME -# There is no config_src knob: the config chain is workflow/config/cfis, resolved -# relative to the Snakefile. The configs interpolate $SP_RUN / $SP_UNIT_NUM / -# $SP_CONFIG / $SP_EXP / $NGMIX_* and the rules export them -- configs and rules -# are one artefact and must version together, so the dir is fixed by construction. +# There is no separate config_src setting: the config chain is +# workflow/config/cfis, resolved relative to the Snakefile. The configs +# interpolate $SP_RUN / $SP_UNIT_NUM / $SP_CONFIG / $SP_EXP / $NGMIX_* and the +# rules export them -- configs and rules are one artefact and must version +# together, so the dir is fixed by construction. # The run index, and — sharing its directory — missing.json and run_report.json. # On the persistent root with the catalogues (D5): the index is the record of # which tile reads which exposure, so it is what a post-purge reconstruction # would otherwise have to rebuild from tile headers. - index_db: /project/def-mjhudson/cdaley/sp-products/smk-g6/index/run_index.sqlite + index_db: REPLACE_ME # Rolling exposure-store reclamation (D5). When true, the COMPUTE DAG grows one # `clean_exposure` job per exposure. It fires once every campaign tile that reads # that exposure has its vignets, deletes the exposure's store AND its manifests, # and leaves `cleaned.json`, which absorbs the manifests — `sp report` reads them -# back out of the tombstone and reports the exposure as `cleaned`. OFF for smk-g5: -# this is a one-tile control, its ~8 exposure stores cost ~54 GiB against 979 GiB -# free, reclamation is already measured exactly by smk-g4, and keeping the stores -# means a re-measurement re-runs only the shape chain instead of the whole tile. +# back out of the tombstone and reports the exposure as `cleaned`. Set it false +# for a small control run whose exposure stores you want to keep around (e.g. to +# re-measure only the shape chain rather than rebuild the whole tile) -- true is +# the right default for anything that isn't a deliberately-kept fixture. clean: true # Rolling TILE-store reclamation (D5). When true, the COMPUTE DAG grows one @@ -122,8 +134,8 @@ clean: true # against a 1M quota. Reclaimed, a tile costs 10 inodes and 12.9 KB — 231k inodes # and ~300 MB for the whole of DR6. # -# OFF here for the same reason `clean:` is: smk-g5 is a one-tile control, and its -# store is what a re-measurement would read. +# Set it false for the same kind of small control run `clean:` above would be +# false for, when the per-tile store is what a re-measurement needs to read. clean_tiles: true # Tiles that may NOT pin an exposure store (default: empty). @@ -153,7 +165,7 @@ ngmix_chunks: 8 # tiles and failed groups' orphans fill /tmp); the cost is ngmix read latency. # tile_store_root: /n23data1//shapepipe/tilestore -# The container's installed shapepipe is overridden by the prod worktree via -# --env PYTHONPATH in profiles/nibi (settled call 3); no config knob here. +# The container's installed shapepipe is overridden by the launching worktree +# via each profile's own --env PYTHONPATH; there is no setting for it here. # build_index.py's missing-tile fraction threshold is passed by workflow/bin/sp # (SP_MISSING_THRESHOLD, default 0.0 = any missing tile is fatal). From d64f88ed40afb1a56d509c33a58d3185f51bbe2f Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Tue, 15 Sep 2026 17:02:38 +0200 Subject: [PATCH 32/85] workflow: retrieve mode (symlink|vos) is a run-config key, not fixed in the ini RETRIEVE=symlink was hard-coded in config_tile_Git.ini and config_exp_Gie.ini (both config/cfis and config/cfis_image_sims), which assumes a pre-staged local mirror. That's nibi's layout; candide's real-data ingestion fetches tiles/exposures from VOSpace (RETRIEVE=vos) and has no local mirror. Add `retrieve: symlink|vos` to the run config, validated at parse time, exported as SP_RETRIEVE by unit_pre() and interpolated into the four ini files. input_type=image_sims always forces symlink regardless of the run config's value, since simulation output is generated locally and never lives in VOSpace -- so a `data` run config's `retrieve: vos` reused as a sims campaign's base can't leak into the sims ingestion stages. Verified with a dry run against both input_type=data (retrieve: vos) and the existing image_sims acceptance-test config (still forces symlink) that SP_RETRIEVE renders correctly in each case. --- workflow/Snakefile | 28 +++++++++++++++++-- workflow/config.yaml | 9 ++++++ workflow/config/cfis/config_exp_Gie.ini | 7 +++-- workflow/config/cfis/config_tile_Git.ini | 6 ++-- .../config/cfis_image_sims/config_exp_Gie.ini | 7 +++-- .../cfis_image_sims/config_tile_Git.ini | 7 +++-- 6 files changed, 53 insertions(+), 11 deletions(-) diff --git a/workflow/Snakefile b/workflow/Snakefile index a6a51b92e..78672861d 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -137,6 +137,25 @@ if PSF_MODEL == "fake" and not PSF_DICT: raise WorkflowError("psf_model=fake needs `psf_dict:` (the simulation's " "pickled PSF dictionary) in the run config.") +# How tile/exposure images reach the run: `symlink` (a pre-staged local +# mirror -- nibi's layout) or `vos` (fetched from VOSpace per unit -- candide's +# real-data layout, no local mirror needed). A per-cluster choice, so it is a +# run-config key (`retrieve:`) rather than fixed in the committed ini files; +# `unit_pre` exports it as SP_RETRIEVE for config_tile_Git.ini / +# config_exp_Gie.ini. `image_sims` always overrides it to `symlink` below, +# regardless of what the run config says: simulation output is generated +# locally and never lives in VOSpace, so a `data` run config's `retrieve: vos` +# reused as a sims campaign's base must not leak into the sims ingestion +# stages. +RETRIEVE_MODES = {"symlink", "vos"} +RETRIEVE_MODE = config.get("retrieve", "symlink") +if RETRIEVE_MODE not in RETRIEVE_MODES: + raise WorkflowError( + f"Invalid retrieve={RETRIEVE_MODE!r}; expected one of " + f"{sorted(RETRIEVE_MODES)}.") +if INPUT_TYPE == "image_sims": + RETRIEVE_MODE = "symlink" + # --- paths ----------------------------------------------------------------- # RUN_DIR is the scratch root: bulk intermediates, sized so a batch finishes @@ -570,9 +589,11 @@ def unit_pre(stage, unit, *, exp_name=None, forest=None, env=None, the committed config dir. It also exports the configured input roots as ``SP_INPUT_TILES`` and - ``SP_INPUT_EXPOSURES`` and the PSF choice as ``SP_PSF`` for the committed - ini chain (plus ``PSF_DICT`` for the image-simulation true PSF, which the - configs read as ``${SP_PSF}_interp_runner`` = ``fake_interp_runner``). + ``SP_INPUT_EXPOSURES``, how to fetch them as ``SP_RETRIEVE`` (``symlink`` + or ``vos``; forced to ``symlink`` for ``image_sims``, see ``RETRIEVE_MODE`` + above), and the PSF choice as ``SP_PSF`` for the committed ini chain (plus + ``PSF_DICT`` for the image-simulation true PSF, which the configs read as + ``${SP_PSF}_interp_runner`` = ``fake_interp_runner``). Every line here is part of each rule's ``params.pre`` and so of the ``params`` rerun trigger: a line added for every rule reruns every finished @@ -602,6 +623,7 @@ def unit_pre(stage, unit, *, exp_name=None, forest=None, env=None, f"export SP_CONFIG='{CONFIG_DIR}'", f"export SP_INPUT_TILES='{INPUTS['tiles']}'", f"export SP_INPUT_EXPOSURES='{INPUTS['exposures']}'", + f"export SP_RETRIEVE='{RETRIEVE_MODE}'", f"export SP_PSF='{PSF_MODEL}'", # Also set via apptainer-args in the profile; kept here so a hand-run of # this same line outside snakemake behaves identically. diff --git a/workflow/config.yaml b/workflow/config.yaml index 31c45c038..0ca3b239f 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -34,6 +34,15 @@ inputs: tiles: REPLACE_ME exposures: REPLACE_ME +# How the tile/exposure images in `inputs:` above reach the run: `symlink` +# (a pre-staged local mirror) or `vos` (fetched from VOSpace per unit, no +# local mirror). This is cluster infrastructure, not a campaign choice -- +# nibi mirrors locally, candide fetches from VOSpace -- so pick the one that +# matches wherever `inputs.tiles`/`inputs.exposures` actually point. Ignored +# for `input_type: image_sims`, which is always `symlink` (simulation output +# is generated locally and never lives in VOSpace). +retrieve: symlink + # The container every job runs inside (apptainer software-deployment in the profile). container: REPLACE_ME diff --git a/workflow/config/cfis/config_exp_Gie.ini b/workflow/config/cfis/config_exp_Gie.ini index b0cce78e8..bfd65b84e 100644 --- a/workflow/config/cfis/config_exp_Gie.ini +++ b/workflow/config/cfis/config_exp_Gie.ini @@ -86,8 +86,11 @@ INPUT_NUMBERING = \d{6} # Output file pattern without number OUTPUT_FILE_PATTERN = image-, weight-, flag- -# Method to retrieve images, one in 'vos', 'symlink' -RETRIEVE = symlink +# Method to retrieve images, one in 'vos', 'symlink' -- set by the run +# config's `retrieve:` key (unit_pre exports SP_RETRIEVE), not fixed here: +# nibi mirrors the exposures locally (symlink), candide fetches them from +# VOSpace (vos). +RETRIEVE = $SP_RETRIEVE # If RETRIEVE=vos, number of attempts to download # Optional, default=3 diff --git a/workflow/config/cfis/config_tile_Git.ini b/workflow/config/cfis/config_tile_Git.ini index cccf0a8b9..87a189f07 100644 --- a/workflow/config/cfis/config_tile_Git.ini +++ b/workflow/config/cfis/config_tile_Git.ini @@ -80,8 +80,10 @@ INPUT_NUMBERING = \d{3}\.\d{3} # Output file pattern without number OUTPUT_FILE_PATTERN = CFIS_image-, CFIS_weight- -# Copy/download method, one in 'vos', 'symlink' -RETRIEVE = symlink +# Copy/download method, one in 'vos', 'symlink' -- set by the run config's +# `retrieve:` key (unit_pre exports SP_RETRIEVE), not fixed here: nibi mirrors +# the tiles locally (symlink), candide fetches them from VOSpace (vos). +RETRIEVE = $SP_RETRIEVE # If RETRIEVE=vos, number of attempts to download # Optional, default=3 diff --git a/workflow/config/cfis_image_sims/config_exp_Gie.ini b/workflow/config/cfis_image_sims/config_exp_Gie.ini index d2c782672..c5eee0f4b 100644 --- a/workflow/config/cfis_image_sims/config_exp_Gie.ini +++ b/workflow/config/cfis_image_sims/config_exp_Gie.ini @@ -91,8 +91,11 @@ INPUT_NUMBERING = \d{7} # Output file pattern without number OUTPUT_FILE_PATTERN = image-, weight-, flag- -# Method to retrieve images, one in 'vos', 'symlink' -RETRIEVE = symlink +# Method to retrieve images, one in 'vos', 'symlink'. Not a sims/data +# difference (unlike the fields marked above): the Snakefile always exports +# SP_RETRIEVE=symlink for image_sims regardless of the run config's +# `retrieve:` key, since simulation output is always local, never VOSpace. +RETRIEVE = $SP_RETRIEVE # If RETRIEVE=vos, number of attempts to download # Optional, default=3 diff --git a/workflow/config/cfis_image_sims/config_tile_Git.ini b/workflow/config/cfis_image_sims/config_tile_Git.ini index 69fcff540..4489788cb 100644 --- a/workflow/config/cfis_image_sims/config_tile_Git.ini +++ b/workflow/config/cfis_image_sims/config_tile_Git.ini @@ -85,8 +85,11 @@ INPUT_NUMBERING = \d{3}-\d{3} # Output file pattern without number OUTPUT_FILE_PATTERN = CFIS_image-, CFIS_weight- -# Copy/download method, one in 'vos', 'symlink' -RETRIEVE = symlink +# Copy/download method, one in 'vos', 'symlink'. Not a sims/data difference +# (unlike the fields marked above): the Snakefile always exports +# SP_RETRIEVE=symlink for image_sims regardless of the run config's +# `retrieve:` key, since simulation output is always local, never VOSpace. +RETRIEVE = $SP_RETRIEVE # If RETRIEVE=vos, number of attempts to download # Optional, default=3 From c7bcfddd3d9d7d273f662c73a83af7d1cbb00e74 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Tue, 15 Sep 2026 17:29:51 +0200 Subject: [PATCH 33/85] config.yaml: restore real nibi defaults, keep the placeholder check as a guard REPLACE_ME sentinels made the committed config unrunnable as-is on any cluster, including nibi -- its own actual mainline user. Restore the real, working nibi values (tile_list, inputs, container, outputs) so `sp run` with no other setup does something real and correct there, the same way it always used to. The REPLACE_ME check in the Snakefile stays: it now guards against a future edit that strips these defaults without supplying real ones, rather than gatekeeping every run. Running the committed default from candide now fails honestly on nibi's real, unreachable /project path (FileNotFoundError) rather than on the placeholder check -- correct, since candide cannot see nibi's filesystem and must supply its own paths via SP_RUN_CONFIG regardless. --- workflow/config.yaml | 58 +++++++++++++++++++++++--------------------- 1 file changed, 31 insertions(+), 27 deletions(-) diff --git a/workflow/config.yaml b/workflow/config.yaml index 0ca3b239f..2db875bab 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -1,38 +1,43 @@ # Run configuration for the ShapePipe Snakemake workflow. # -# THIS FILE IS A UNIVERSAL TEMPLATE, NOT A LIVE CAMPAIGN'S CONFIG. It carries -# no machine- or campaign-specific path, on purpose: `sp run` must be the same -# command on nibi, candide, or anywhere else, and a committed default that -# happened to bake in one person's real paths would silently run a stranger's -# campaign against the wrong cluster's filesystem the moment they forgot to -# override it. Every path below is REPLACE_ME; the Snakefile refuses to parse -# past one of the ones it cannot otherwise fail clearly on later (tile_list, -# inputs.tiles, inputs.exposures, outputs.run_dir, outputs.index_db -- see the -# check near the top of Snakefile). `container:` is exempt from that check: a -# local sandbox or cached SIF (`sp container status`) is resolved first and -# makes this key irrelevant for anyone who has one. -# -# A real run always supplies its own paths via SP_RUN_CONFIG, which REPLACES -# this file wholesale rather than layering over it (workflow/bin/sp; see the -# README's "Run configuration" section): +# This file is nibi's own real, working default (Cail's real-data campaign, +# the checkout's current mainline user) -- `sp run` with no other setup does +# something real and correct on nibi. It is NOT the right config for another +# cluster or a different campaign: `SP_RUN_CONFIG` points `sp run` at a run +# config that REPLACES this file wholesale rather than layering over it +# (workflow/bin/sp; see the README's "Run configuration" section), which is +# how a candide run, an image-sims branch, or a second nibi campaign supplies +# its own paths without touching this one: # # SP_PROFILE= SP_RUN_CONFIG=/path/to/your_run.yaml \ # workflow/bin/sp run # +# Every path a run genuinely cannot do without (tile_list, inputs.tiles, +# inputs.exposures, outputs.run_dir, outputs.index_db) is checked at parse +# time against the sentinel value REPLACE_ME (see the check near the top of +# Snakefile) -- so a config that strips these defaults without supplying real +# ones fails with one clear message instead of a confusing one downstream, or +# silently running against nibi's paths from a machine that can't see them. +# `container:` is exempt from that check: a local sandbox or cached SIF +# (`sp container status`) is resolved first and makes this key irrelevant for +# anyone who has one. +# # A "run" is declared by a tile list plus the paths below. Everything here is # read at parse time; none of it is a rule input, so editing it (e.g. appending # tiles) never invalidates completed work — it only changes which jobs exist. # The tile list that scopes this run (one "IDra.IDdec" per line). The # campaign grows by appending to this file — parse-time config, so completed -# work is never invalidated. -tile_list: REPLACE_ME +# work is never invalidated. This is the real, currently-growing list for +# nibi's smk-gN campaign; a different run always supplies its own. +tile_list: /project/def-mjhudson/cdaley/sp-products/smk-g6/tiles.txt # Inputs are the only site-specific data paths; all committed configs consume these -# through SP_INPUT_TILES and SP_INPUT_EXPOSURES. +# through SP_INPUT_TILES and SP_INPUT_EXPOSURES. Pre-staged P3 data on nibi +# (get_images RETRIEVE=symlink, below). inputs: - tiles: REPLACE_ME - exposures: REPLACE_ME + tiles: /project/def-mjhudson/unions-wl/tiles + exposures: /project/def-mjhudson/unions-wl/exposures # How the tile/exposure images in `inputs:` above reach the run: `symlink` # (a pre-staged local mirror) or `vos` (fetched from VOSpace per unit, no @@ -44,7 +49,7 @@ inputs: retrieve: symlink # The container every job runs inside (apptainer software-deployment in the profile). -container: REPLACE_ME +container: /project/def-mjhudson/cdaley/containers/shapepipe-develop-runtime.sif # PSF model used by the exposure and tile interpolation stages: psfex, mccd, or # fake (image simulations only: the true PSF from `psf_dict`, no exposure fit). @@ -64,7 +69,7 @@ input_type: data # sharded per-unit stores live under it: # /tiles/<2-char prefix>// and /exp/// outputs: - run_dir: REPLACE_ME + run_dir: /scratch/cdaley/shapepipe-output/smk-g6 # products_dir is the PERSISTENT root: the durable, low-volume products — the # final catalogues (/tiles///final_cat-.fits, @@ -77,13 +82,12 @@ outputs: # tile re-declare inputs against exposure stores reclamation already deleted. # # Unset means "one root": products land under run_dir, exactly the pre-D5 -# layout, which is what a fixture or smoke test wants. Delete this whole key -# (not REPLACE_ME) if that's what you want -- it's the one path in this file -# that's genuinely optional rather than templated. +# layout, which is what a fixture or smoke test wants -- delete this whole +# key, rather than pointing it at run_dir, if that's what you want. # # Snakemake's own state is the one durable-looking thing that stays on scratch # (-state; bin/sp explains why). - products_dir: REPLACE_ME + products_dir: /project/def-mjhudson/cdaley/sp-products/smk-g6 # There is no separate config_src setting: the config chain is # workflow/config/cfis, resolved relative to the Snakefile. The configs @@ -95,7 +99,7 @@ outputs: # On the persistent root with the catalogues (D5): the index is the record of # which tile reads which exposure, so it is what a post-purge reconstruction # would otherwise have to rebuild from tile headers. - index_db: REPLACE_ME + index_db: /project/def-mjhudson/cdaley/sp-products/smk-g6/index/run_index.sqlite # Rolling exposure-store reclamation (D5). When true, the COMPUTE DAG grows one # `clean_exposure` job per exposure. It fires once every campaign tile that reads From 5be4c431a0578d1e7147a85b48c4b0bc8bec2952 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Tue, 15 Sep 2026 18:19:58 +0200 Subject: [PATCH 34/85] Snakefile: check machine: against SP_PROFILE at parse time machine: (the run config) and SP_PROFILE (the env var bin/sp reads to pick profiles//, default nibi) are two independent switches for the same fact -- which cluster this is -- set in two different places at two different times, with nothing stopping them from disagreeing: export SP_PROFILE=candide but leave a copied run config's `machine: nibi` unedited, and jobs run under candide's SLURM settings against nibi's data paths. Raise a clear WorkflowError at parse time if they don't match, skipped only when `machine:` is genuinely absent (a config predating this key). No code default for `machine:` -- guessing one on its behalf would defeat the point of the check. Also fixes a comment above the REPLACE_ME check left stale by the previous commit: it still called config.yaml "a universal template with no real paths in it" after nibi's real defaults were restored. --- workflow/Snakefile | 45 +++++++++++++++++++++++++++++++------------- workflow/config.yaml | 8 ++++++++ 2 files changed, 40 insertions(+), 13 deletions(-) diff --git a/workflow/Snakefile b/workflow/Snakefile index 78672861d..9a0951bf2 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -52,16 +52,16 @@ _RUN_CONFIG = os.environ.get("SP_RUN_CONFIG") or str( Path(workflow.snakefile).parent / "config.yaml") configfile: _RUN_CONFIG -# The committed config.yaml is a universal template (no machine- or -# campaign-specific path in it), so the same `sp run` invocation works on -# every cluster: whoever runs it always supplies their own paths through -# SP_RUN_CONFIG (see config.yaml's own top comment and the README). Left -# unset, the template's REPLACE_ME sentinel would otherwise read as a real, -# oddly-named path -- silently missing tiles, or a job dying hours in on a -# node that can't see "/REPLACE_ME" -- rather than the one-line, same-message- -# everywhere error this check gives instead. `container:` is checked -# separately (container.py resolves a local sandbox or cached SIF first, so a -# placeholder there is often harmless) and so is left out here. +# config.yaml ships nibi's own real, working paths (the checkout's current +# mainline campaign) so `sp run` does something correct there with no other +# setup; a different cluster or campaign supplies its own via SP_RUN_CONFIG +# (see config.yaml's own top comment and the README) rather than editing this +# file. This check guards against a FUTURE edit that strips these defaults +# without supplying real ones -- one clear, identical-everywhere error instead +# of a confusing failure downstream, or (worse) a job dying hours in on a node +# that can't see "/REPLACE_ME". `container:` is checked separately +# (container.py resolves a local sandbox or cached SIF first, so a placeholder +# there is often harmless) and so is left out here. _PLACEHOLDER = "REPLACE_ME" for _key, _value in { "tile_list": config.get("tile_list"), @@ -73,9 +73,28 @@ for _key, _value in { if _value == _PLACEHOLDER: raise WorkflowError( f"{_key!r} in {_RUN_CONFIG} is still the template placeholder " - f"{_PLACEHOLDER!r}. workflow/config.yaml ships as a universal " - f"template with no real paths in it; point SP_RUN_CONFIG at your " - f"own run config instead (see the README, 'Run configuration').") + f"{_PLACEHOLDER!r}. Point SP_RUN_CONFIG at your own run config " + f"instead of editing the committed one (see the README, " + f"'Run configuration').") + +# `machine:` (the run config) and SP_PROFILE (the environment variable +# bin/sp reads to pick profiles//, default "nibi") are two independent +# switches for the same fact -- which cluster this is -- set in two different +# places at two different times, and nothing stops them disagreeing: export +# SP_PROFILE=candide but leave a copied run config's `machine: nibi` unedited, +# and jobs run under candide's SLURM settings against nibi's data paths. +# `machine:` has no code default (a wrong SLURM profile is exactly the kind +# of mistake this check exists to catch, so guessing one on its behalf would +# defeat the point); the check is skipped only if `machine:` truly is unset +# (a config predating this key), never silently. +_SP_PROFILE = os.environ.get("SP_PROFILE", "nibi") +_MACHINE = config.get("machine") +if _MACHINE is not None and _MACHINE != _SP_PROFILE: + raise WorkflowError( + f"machine={_MACHINE!r} in {_RUN_CONFIG} does not match " + f"SP_PROFILE={_SP_PROFILE!r} (env, defaults to 'nibi' when unset). " + f"Set SP_PROFILE={_MACHINE} to launch on the machine this config " + f"declares, or fix `machine:` if the config is the stale one.") # Every job's shell runs inside this container (apptainer software-deployment in # the profile); the user never types apptainer. WHICH image is the one resolution diff --git a/workflow/config.yaml b/workflow/config.yaml index 2db875bab..a9f90332a 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -26,6 +26,14 @@ # read at parse time; none of it is a rule input, so editing it (e.g. appending # tiles) never invalidates completed work — it only changes which jobs exist. +# Which cluster this run config is for. Must match SP_PROFILE (the +# environment variable that picks profiles//, default "nibi") -- the +# Snakefile raises a clear error at parse time if they disagree, since that +# combination means a job would run under one cluster's SLURM settings +# against another cluster's data paths. +machine: nibi +#machine: candide + # The tile list that scopes this run (one "IDra.IDdec" per line). The # campaign grows by appending to this file — parse-time config, so completed # work is never invalidated. This is the real, currently-growing list for From 29442d3b631d2e656f7990db4ce19e685bc936c6 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Tue, 15 Sep 2026 18:45:13 +0200 Subject: [PATCH 35/85] config.yaml: machines: table drives per-machine (and per-input_type) defaults `machine: nibi` selects config.yaml's own real defaults; the same key now also indexes a `machines:` table carrying tile_list, retrieve, inputs, outputs and container for both nibi and candide, so switching `machine:` (kept consistent with SP_PROFILE by the existing check) is enough to point a bare `sp run` at the right cluster's paths. On candide, where the real survey (VOSpace, RETRIEVE=vos) and the SKiLLS simulations (local disk, RETRIEVE=symlink, a different container) need different values, a machine entry nests them under `data:`/ `image_sims:` instead of stating them flat. `$base_dir` in a string is substituted with that machine's own `base_dir:`, so nibi's block states its one common root once instead of four times. A value may be the sentinel `TBD` -- known to be needed, not yet known what it is (candide's real-data ingestion: no local mirror exists there, the real vos: paths and output roots aren't confirmed yet) -- checked at parse time with the same one-clear-error mechanism REPLACE_ME used to provide. Container resolution needed a specific bridge: container.py reads `container:` by regex off the run config FILE, independent of the Snakefile's `config` dict, so a machine-derived container (which only exists after this Python-side merge) would otherwise be invisible to it. Reused SP_CONTAINER (already the documented way to point resolve_image() at a specific image) instead of teaching container.py about `machines:` too -- set only when neither an explicit top-level `container:` nor a user-set SP_CONTAINER already exists, preserving the same override precedence as every other machine default. Also fixes SP_RUN_CONFIG to LAYER on top of config.yaml rather than replacing it wholesale (workflow.configfile(), called conditionally -- it's the plain method the `configfile:` directive itself compiles to, so it works inside an `if`; merges recursively per snakemake.utils.update_config). Necessary for `machines:` to reach the common case at all: every image-sims run and any one-off cluster override already goes through SP_RUN_CONFIG, and under the old wholesale-replace behaviour none of them would ever have seen the table. Safe against the original replace-wholesale concern (an external config silently inheriting an unrelated real campaign's paths) because `machines:` entries are well-scoped defaults gated on `machine:`/`input_type:` and independently checked against SP_PROFILE, not one campaign's actual values reused as another's fallback. Verified: the committed nibi default still fails honestly (on nibi's real, candide-unreachable path) rather than on a placeholder; a machine:/SP_PROFILE mismatch still raises before that; a candide+data run config (all TBD) raises the new clear placeholder error; and the grid_3 image-sims acceptance-test config, with only `machine: candide` added and its explicit retrieve:/container: removed, dry-runs identically to before except for the retrieve: addition already committed separately -- confirming the machines: table now supplies what that config used to have to state itself, without disturbing the already-completed run's resume state. --- workflow/Snakefile | 176 ++++++++++++++++++++++++++++++------------- workflow/config.yaml | 161 ++++++++++++++++++++++----------------- 2 files changed, 214 insertions(+), 123 deletions(-) diff --git a/workflow/Snakefile b/workflow/Snakefile index 9a0951bf2..0a81548c9 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -45,37 +45,42 @@ from snakemake.exceptions import WorkflowError # --directory on /scratch (bin/sp) so .snakemake/ state never lands on /project # (group quota is a hard 27/27 TiB — a metadata write mid-run died on it live). # -# SP_RUN_CONFIG (set by bin/sp) REPLACES the committed config.yaml rather than -# layering on it: a run driven from outside -- one image-simulation shear branch, -# say -- must not inherit the committed run's paths for any key it omits. -_RUN_CONFIG = os.environ.get("SP_RUN_CONFIG") or str( - Path(workflow.snakefile).parent / "config.yaml") -configfile: _RUN_CONFIG - -# config.yaml ships nibi's own real, working paths (the checkout's current -# mainline campaign) so `sp run` does something correct there with no other -# setup; a different cluster or campaign supplies its own via SP_RUN_CONFIG -# (see config.yaml's own top comment and the README) rather than editing this -# file. This check guards against a FUTURE edit that strips these defaults -# without supplying real ones -- one clear, identical-everywhere error instead -# of a confusing failure downstream, or (worse) a job dying hours in on a node -# that can't see "/REPLACE_ME". `container:` is checked separately -# (container.py resolves a local sandbox or cached SIF first, so a placeholder -# there is often harmless) and so is left out here. -_PLACEHOLDER = "REPLACE_ME" -for _key, _value in { - "tile_list": config.get("tile_list"), - "inputs.tiles": (config.get("inputs") or {}).get("tiles"), - "inputs.exposures": (config.get("inputs") or {}).get("exposures"), - "outputs.run_dir": (config.get("outputs") or {}).get("run_dir"), - "outputs.index_db": (config.get("outputs") or {}).get("index_db"), -}.items(): - if _value == _PLACEHOLDER: - raise WorkflowError( - f"{_key!r} in {_RUN_CONFIG} is still the template placeholder " - f"{_PLACEHOLDER!r}. Point SP_RUN_CONFIG at your own run config " - f"instead of editing the committed one (see the README, " - f"'Run configuration').") +# config.yaml's defaults -- including the per-machine `machines:` table -- +# are ALWAYS loaded first; SP_RUN_CONFIG (set by bin/sp), when given, is +# layered on top via a second, conditional load (workflow.configfile() is a +# plain method -- what the `configfile:` directive itself compiles to -- +# so it can be called from inside an `if`; it merges recursively via +# snakemake.utils.update_config, not by replacement). This matters +# specifically for `machines:`: an external run config (every image-sims +# run, any one-off nibi/candide override) must still see the per-machine +# defaults it exists to provide, or they would only ever apply to the one +# case nobody actually runs -- the bare committed config with no override at +# all. A key given in SP_RUN_CONFIG still wins over the same key from +# config.yaml (later configfile: wins in the merge), so an override remains +# exactly that -- this is the "layer, don't inherit blind" version of what +# used to be a wholesale replace: `machines:` supplies well-scoped defaults +# (gated on `machine:`/`input_type:`, checked against SP_PROFILE below), not +# one campaign's actual paths reused as another's fallback. +_CONFIG_YAML = str(Path(workflow.snakefile).parent / "config.yaml") +_RUN_CONFIG = os.environ.get("SP_RUN_CONFIG") or _CONFIG_YAML +configfile: _CONFIG_YAML +if os.environ.get("SP_RUN_CONFIG"): + workflow.configfile(os.environ["SP_RUN_CONFIG"]) + +# --- input type (needed below, before PSF model) ---------------------------- +# input_type selects WHERE the pixels come from; everything downstream of the +# ingestion stages is the same chain. `image_sims` runs the SKiLLS image +# simulations through the real-data configs: the overlay config dir +# (config/cfis_image_sims) holds real files only for the ingestion INIs whose +# input naming differs, the fake-PSF INIs, and the merge column list +# (final_cat.param), and symlinks the rest into config/cfis, so an m-bias +# measured on the sims calibrates the pipeline that makes the real catalogue. +INPUT_TYPES = {"data", "image_sims"} +INPUT_TYPE = config.get("input_type", "data") +if INPUT_TYPE not in INPUT_TYPES: + raise WorkflowError( + f"Invalid input_type={INPUT_TYPE!r}; expected one of " + f"{sorted(INPUT_TYPES)}.") # `machine:` (the run config) and SP_PROFILE (the environment variable # bin/sp reads to pick profiles//, default "nibi") are two independent @@ -87,15 +92,78 @@ for _key, _value in { # of mistake this check exists to catch, so guessing one on its behalf would # defeat the point); the check is skipped only if `machine:` truly is unset # (a config predating this key), never silently. -_SP_PROFILE = os.environ.get("SP_PROFILE", "nibi") -_MACHINE = config.get("machine") -if _MACHINE is not None and _MACHINE != _SP_PROFILE: +SP_PROFILE = os.environ.get("SP_PROFILE", "nibi") +MACHINE = config.get("machine") +if MACHINE is not None and MACHINE != SP_PROFILE: raise WorkflowError( - f"machine={_MACHINE!r} in {_RUN_CONFIG} does not match " - f"SP_PROFILE={_SP_PROFILE!r} (env, defaults to 'nibi' when unset). " - f"Set SP_PROFILE={_MACHINE} to launch on the machine this config " + f"machine={MACHINE!r} in {_RUN_CONFIG} does not match " + f"SP_PROFILE={SP_PROFILE!r} (env, defaults to 'nibi' when unset). " + f"Set SP_PROFILE={MACHINE} to launch on the machine this config " f"declares, or fix `machine:` if the config is the stale one.") +# --- per-machine defaults (config.yaml's `machines:` table) ----------------- +# tile_list / retrieve / inputs / outputs / container differ by machine, and +# on candide also by input_type (the real survey and the SKiLLS simulations +# live in entirely different places, fetched differently) -- `machines:` in +# the run config carries one entry per machine, itself carrying either those +# keys directly (uniform across input_type, e.g. nibi) or nested one level +# under `data:`/`image_sims:` (e.g. candide). Any of these keys given at the +# TOP LEVEL of the run config (outside `machines:`) overrides its machine +# default -- same precedence SP_RUN_CONFIG already has over config.yaml, one +# level down. `$base_dir` inside a string is replaced with that machine +# entry's own `base_dir:`, so nibi's block states the one path that actually +# is common to its tiles/exposures/products/container once, instead of four +# times. A machine default may itself be the sentinel `TBD` (see the check +# below) -- known to be needed, not yet known what it is. +def _expand_base_dir(value, base_dir): + if isinstance(value, str): + return value.replace("$base_dir", base_dir) if base_dir else value + if isinstance(value, dict): + return {k: _expand_base_dir(v, base_dir) for k, v in value.items()} + return value + +_mconf = (config.get("machines") or {}).get(MACHINE) or {} +_sub = _mconf.get(INPUT_TYPE) or {} +_base_dir = _mconf.get("base_dir", "") +_MACHINE_CONTAINER = None +for _key in ("tile_list", "retrieve"): + if _key not in config: + _val = _sub.get(_key, _mconf.get(_key)) + if _val is not None: + config[_key] = _expand_base_dir(_val, _base_dir) +_MACHINE_CONTAINER = _sub.get("container", _mconf.get("container")) +if _MACHINE_CONTAINER is not None: + _MACHINE_CONTAINER = _expand_base_dir(_MACHINE_CONTAINER, _base_dir) +for _section in ("inputs", "outputs"): + _val = _sub.get(_section, _mconf.get(_section)) + if _val: + _merged = _expand_base_dir(dict(_val), _base_dir) + _merged.update(config.get(_section) or {}) + config[_section] = _merged + +# Every path a run genuinely cannot do without (tile_list, inputs.tiles, +# inputs.exposures, outputs.run_dir, outputs.index_db) is checked, after the +# machine defaults above are folded in, against the sentinel value `TBD` -- +# known to be needed, not yet known what it is (typically: this machine +# entry's real value hasn't been confirmed yet, see config.yaml). One clear, +# identical-everywhere error instead of a confusing failure downstream, or +# (worse) a job dying hours in on a node that can't see a path literally +# named "TBD". `container:` is checked separately, below its own resolution. +_PLACEHOLDER = "TBD" +for _key, _value in { + "tile_list": config.get("tile_list"), + "inputs.tiles": (config.get("inputs") or {}).get("tiles"), + "inputs.exposures": (config.get("inputs") or {}).get("exposures"), + "outputs.run_dir": (config.get("outputs") or {}).get("run_dir"), + "outputs.index_db": (config.get("outputs") or {}).get("index_db"), +}.items(): + if _value == _PLACEHOLDER: + raise WorkflowError( + f"{_key!r} resolves to the placeholder {_PLACEHOLDER!r} for " + f"machine={MACHINE!r}, input_type={INPUT_TYPE!r} -- this value " + f"is not known yet (see config.yaml's `machines:` table). Supply " + f"it explicitly in your run config until it is.") + # Every job's shell runs inside this container (apptainer software-deployment in # the profile); the user never types apptainer. WHICH image is the one resolution # order `sp container` exposes — this user's writable sandbox if they built one, @@ -105,6 +173,21 @@ if _MACHINE is not None and _MACHINE != _SP_PROFILE: sys.path.insert(0, str(Path(workflow.snakefile).parent / "scripts")) import container as _container # noqa: E402 +# container.py's `configured_default()` reads `container:` by REGEX off the +# run config FILE, independent of the `config` dict above (it is deliberately +# stdlib-only, no yaml import) -- so a machine-derived container, which only +# exists after the Python-side merge above, is invisible to it. SP_CONTAINER +# is already the documented way to point resolve_image() at a specific image +# instead of the cache (`local_sif()` reads it), so reuse that rather than +# teach container.py about `machines:` too. Set only when nothing more +# specific already exists: an explicit top-level `container:` in the run +# config (which STILL reaches container.py, by file, the normal way) or an +# SP_CONTAINER the user set themselves both take precedence, matching every +# other machine default above. +if _MACHINE_CONTAINER and not os.environ.get("SP_CONTAINER") \ + and not config.get("container"): + os.environ["SP_CONTAINER"] = _MACHINE_CONTAINER + # Loud at PARSE time, both ways: a broken SP_CONTAINER override, or nothing to # resolve at all (empty cache AND no `container:` key). The empty string the # resolver returns for "none" is a valid-looking container directive that passes @@ -121,21 +204,7 @@ if _kind == "none": container: _image -# --- input type and PSF model ---------------------------------------------- -# input_type selects WHERE the pixels come from; everything downstream of the -# ingestion stages is the same chain. `image_sims` runs the SKiLLS image -# simulations through the real-data configs: the overlay config dir -# (config/cfis_image_sims) holds real files only for the ingestion INIs whose -# input naming differs, the fake-PSF INIs, and the merge column list -# (final_cat.param), and symlinks the rest into config/cfis, so an m-bias -# measured on the sims calibrates the pipeline that makes the real catalogue. -INPUT_TYPES = {"data", "image_sims"} -INPUT_TYPE = config.get("input_type", "data") -if INPUT_TYPE not in INPUT_TYPES: - raise WorkflowError( - f"Invalid input_type={INPUT_TYPE!r}; expected one of " - f"{sorted(INPUT_TYPES)}.") - +# --- PSF model --------------------------------------------------------------- # `fake` is the image-simulation true PSF: no exposure PSF fit, and # fake_interp_runner writes the tile's galaxy_psf from the simulation's PSF # dictionary (config `psf_dict`) -- named so that the configs' shared @@ -193,7 +262,8 @@ INDEX_DB = Path(OUTPUTS["index_db"]) SCRIPTS = Path(workflow.basedir) / "scripts" # The config chain is the repo's committed directory (D2). The configs and # rules that set their environment variables must be versioned together. There -# is no `config_src` knob; input_type picks between the two committed dirs. +# is no separate `config_src` setting; input_type picks between the two +# committed dirs. CONFIG_DIR = Path(workflow.basedir) / "config" / { "data": "cfis", "image_sims": "cfis_image_sims"}[INPUT_TYPE] diff --git a/workflow/config.yaml b/workflow/config.yaml index a9f90332a..a3dc1d1d7 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -12,16 +12,6 @@ # SP_PROFILE= SP_RUN_CONFIG=/path/to/your_run.yaml \ # workflow/bin/sp run # -# Every path a run genuinely cannot do without (tile_list, inputs.tiles, -# inputs.exposures, outputs.run_dir, outputs.index_db) is checked at parse -# time against the sentinel value REPLACE_ME (see the check near the top of -# Snakefile) -- so a config that strips these defaults without supplying real -# ones fails with one clear message instead of a confusing one downstream, or -# silently running against nibi's paths from a machine that can't see them. -# `container:` is exempt from that check: a local sandbox or cached SIF -# (`sp container status`) is resolved first and makes this key irrelevant for -# anyone who has one. -# # A "run" is declared by a tile list plus the paths below. Everything here is # read at parse time; none of it is a rule input, so editing it (e.g. appending # tiles) never invalidates completed work — it only changes which jobs exist. @@ -34,80 +24,111 @@ machine: nibi #machine: candide -# The tile list that scopes this run (one "IDra.IDdec" per line). The -# campaign grows by appending to this file — parse-time config, so completed -# work is never invalidated. This is the real, currently-growing list for -# nibi's smk-gN campaign; a different run always supplies its own. -tile_list: /project/def-mjhudson/cdaley/sp-products/smk-g6/tiles.txt - -# Inputs are the only site-specific data paths; all committed configs consume these -# through SP_INPUT_TILES and SP_INPUT_EXPOSURES. Pre-staged P3 data on nibi -# (get_images RETRIEVE=symlink, below). -inputs: - tiles: /project/def-mjhudson/unions-wl/tiles - exposures: /project/def-mjhudson/unions-wl/exposures - -# How the tile/exposure images in `inputs:` above reach the run: `symlink` -# (a pre-staged local mirror) or `vos` (fetched from VOSpace per unit, no -# local mirror). This is cluster infrastructure, not a campaign choice -- -# nibi mirrors locally, candide fetches from VOSpace -- so pick the one that -# matches wherever `inputs.tiles`/`inputs.exposures` actually point. Ignored -# for `input_type: image_sims`, which is always `symlink` (simulation output -# is generated locally and never lives in VOSpace). -retrieve: symlink - -# The container every job runs inside (apptainer software-deployment in the profile). -container: /project/def-mjhudson/cdaley/containers/shapepipe-develop-runtime.sif - -# PSF model used by the exposure and tile interpolation stages: psfex, mccd, or -# fake (image simulations only: the true PSF from `psf_dict`, no exposure fit). -psf_model: psfex - # Where the pixels come from: `data` (survey tiles/exposures, config/cfis) or # `image_sims` (SKiLLS simulations, config/cfis_image_sims -- an overlay that # differs from config/cfis only in the ingestion stages). An image_sims run -# normally sets `psf_model: fake` and `psf_dict:`, and is driven with its own run -# config through SP_RUN_CONFIG (see README, "Image simulations"). +# normally sets `psf_model: fake` and `psf_dict:`, and is driven with its own +# run config through SP_RUN_CONFIG. There is no separate `config_src` +# setting: input_type alone picks between the two committed ini directories, +# which must version together with the rules that set their environment +# variables. input_type: data -# THE TWO ROOTS (D5). +# PSF model used by the exposure and tile interpolation stages: psfex, mccd, or +# fake (image simulations only: the true PSF from `psf_dict`, no exposure fit). +psf_model: psfex + +# Per-(machine, and on candide also input_type) settings: tile_list, retrieve, +# inputs, outputs, container -- see the Snakefile for the exact merge and the +# `$base_dir` substitution rule. Any of these keys given at the TOP LEVEL of +# this file (outside `machines:`) overrides its machine default; `#retrieve:` +# and `#container:` below are exactly that override point, commented out. A +# value of `TBD` means: known to be needed, not yet known what it is -- +# parsing refuses to proceed past one. +# +# THE TWO OUTPUT ROOTS EVERY MACHINE ENTRY DECLARES (D5). # # run_dir is the SCRATCH root and the $SP_RUN every config interpolates: bulk # intermediates, sized so a batch finishes inside the 60-day purge window. The # sharded per-unit stores live under it: # /tiles/<2-char prefix>// and /exp/// -outputs: - run_dir: /scratch/cdaley/shapepipe-output/smk-g6 - +# # products_dir is the PERSISTENT root: the durable, low-volume products — the # final catalogues (/tiles///final_cat-.fits, # mirroring the scratch tree shard for shard), the index, the report. ~32-46 MB -# per tile, so a full DR6 campaign is a few hundred GB. +# per tile, so a full DR6 campaign is a few hundred GB. The final catalogue is +# also the tile-finished MARKER that cuts a finished tile's exposure edges +# (D5), which is the second reason it cannot sit on scratch: a purge would not +# merely lose a product, it would make every finished tile re-declare inputs +# against exposure stores reclamation already deleted. Unset means "one +# root": products land under run_dir, exactly the pre-D5 layout, which is +# what a fixture or smoke test wants -- omit the key, rather than pointing it +# at run_dir, if that's what you want. # -# The final catalogue is also the tile-finished MARKER that cuts a finished -# tile's exposure edges (D5), which is the second reason it cannot sit on -# scratch: a purge would not merely lose a product, it would make every finished -# tile re-declare inputs against exposure stores reclamation already deleted. -# -# Unset means "one root": products land under run_dir, exactly the pre-D5 -# layout, which is what a fixture or smoke test wants -- delete this whole -# key, rather than pointing it at run_dir, if that's what you want. -# -# Snakemake's own state is the one durable-looking thing that stays on scratch -# (-state; bin/sp explains why). - products_dir: /project/def-mjhudson/cdaley/sp-products/smk-g6 - -# There is no separate config_src setting: the config chain is -# workflow/config/cfis, resolved relative to the Snakefile. The configs -# interpolate $SP_RUN / $SP_UNIT_NUM / $SP_CONFIG / $SP_EXP / $NGMIX_* and the -# rules export them -- configs and rules are one artefact and must version -# together, so the dir is fixed by construction. - -# The run index, and — sharing its directory — missing.json and run_report.json. -# On the persistent root with the catalogues (D5): the index is the record of -# which tile reads which exposure, so it is what a post-purge reconstruction -# would otherwise have to rebuild from tile headers. - index_db: /project/def-mjhudson/cdaley/sp-products/smk-g6/index/run_index.sqlite +# index_db is the run index, and — sharing its directory — missing.json and +# run_report.json. On the persistent root with the catalogues (D5): the index +# is the record of which tile reads which exposure, so it is what a +# post-purge reconstruction would otherwise have to rebuild from tile +# headers. Snakemake's own state is the one durable-looking thing that stays +# on scratch instead (-state; bin/sp explains why). +machines: + nibi: + base_dir: /project/def-mjhudson + + # The real, currently-growing tile list for nibi's smk-gN campaign; a + # different run always supplies its own. + tile_list: $base_dir/cdaley/sp-products/smk-g6/tiles.txt + + # Pre-staged P3 data (get_images RETRIEVE=symlink). + retrieve: symlink + inputs: + tiles: $base_dir/unions-wl/tiles + exposures: $base_dir/unions-wl/exposures + + outputs: + run_dir: /scratch/cdaley/shapepipe-output/smk-g6 + products_dir: $base_dir/cdaley/sp-products/smk-g6 + index_db: $base_dir/cdaley/sp-products/smk-g6/index/run_index.sqlite + + container: $base_dir/cdaley/containers/shapepipe-develop-runtime.sif + + candide: + base_dir: /n17data/UNIONS/WL + + data: + # candide has no pre-staged local mirror of the survey (unlike nibi): + # images are fetched from VOSpace per unit instead. + retrieve: vos + tile_list: TBD + inputs: + tiles: TBD # a vos: URL, once known + exposures: TBD + outputs: + run_dir: TBD + container: /n17data/cdaley/containers/shapepipe_develop-runtime-20260718.sif + + image_sims: + # SKiLLS simulation output already lives on local disk. + retrieve: symlink + container: /n17data/cdaley/containers/shapepipe_im_sims-runtime.sif + # tile_list, inputs and outputs are intentionally absent here: each + # SKiLLS grid/shear-branch campaign supplies its own via SP_RUN_CONFIG + # (e.g. inputs.tiles: /n09data/hervas/skills_out//images/SP_tiles) + # -- this entry only saves such a run from restating retrieve/container + # too. + +# How the tile/exposure images in `inputs:` above reach the run: `symlink` +# (a pre-staged local mirror) or `vos` (fetched from VOSpace per unit, no +# local mirror). Set per machine (and, on candide, per input_type) above; +# ignored either way for `input_type: image_sims`, which is always `symlink` +# (simulation output is generated locally and never lives in VOSpace). Give +# it here directly only to override the machine default for one run. +#retrieve: symlink + +# The container every job runs inside (apptainer software-deployment in the +# profile). Set per machine (and, on candide, per input_type) above; give it +# here directly only to override that default for one run. +#container: ... # Rolling exposure-store reclamation (D5). When true, the COMPUTE DAG grows one # `clean_exposure` job per exposure. It fires once every campaign tile that reads From 7293e3b2ef513dfd56b4c25b9b3fbc6c8aca4f3f Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Tue, 15 Sep 2026 19:00:43 +0200 Subject: [PATCH 36/85] workflow: one run-config resolver for Snakefile, bin/sp and container.py The machines: table was only understood by the Snakefile. bin/sp read outputs.run_dir/index_db and tile_store_root straight from the config file, so the committed config (outputs now under machines.nibi) broke `sp run` on nibi with a KeyError, and container.py only saw a top-level container: line. scripts/run_config.py now does the whole resolution (config.yaml, SP_RUN_CONFIG merged on top, then machines[machine] [input_type] for anything unset) and all three use it. - container.resolve_image() takes the resolved container explicitly; drops the SP_CONTAINER workaround, which had let a machine default outrank a user's cached SIF, against the documented order. - nibi's machine entry is nested under data: like candide's, so a nibi image_sims run cannot inherit real-data paths. - Unset required keys are reported like TBD, in one error. - README gains a "Run configuration" section; config.yaml, Snakefile and the ini comments are cut down. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_013NzWDtdQbsK1VhvYbDnmTo --- workflow/README.md | 31 ++- workflow/Snakefile | 202 ++++------------- workflow/bin/sp | 30 ++- workflow/config.yaml | 213 ++++-------------- workflow/config/cfis/config_exp_Gie.ini | 5 +- workflow/config/cfis/config_tile_Git.ini | 4 +- .../config/cfis_image_sims/config_exp_Gie.ini | 5 +- .../cfis_image_sims/config_tile_Git.ini | 5 +- workflow/scripts/completeness.py | 9 +- workflow/scripts/container.py | 38 ++-- workflow/scripts/run_config.py | 83 +++++++ 11 files changed, 234 insertions(+), 391 deletions(-) create mode 100644 workflow/scripts/run_config.py diff --git a/workflow/README.md b/workflow/README.md index 43a7d018d..c0652c1eb 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -69,10 +69,33 @@ One campaign per shear branch, each with its own run config: SP_PROFILE=candide SP_RUN_CONFIG=/path/run_1p2z_grid_1.yaml workflow/bin/sp run ``` -`SP_RUN_CONFIG` replaces `workflow/config.yaml` (it is not layered on it, so no -key falls back to the committed run's paths) and is snapshotted with the code; -`SP_PROFILE` picks `profiles//`. sp_validation's image-simulation workflow -drives these campaigns and measures m from their final catalogues. +## Run configuration + +`SP_RUN_CONFIG` is merged on top of `workflow/config.yaml` and snapshotted with +the code. `machine:` (which must match `SP_PROFILE`, default `nibi`) and +`input_type:` then select an entry of the `machines:` table, which supplies +`tile_list`, `retrieve` (`symlink` or `vos`), `inputs`, `outputs` and +`container` for any of these the run config leaves unset (`$base_dir` expands +to that machine's `base_dir`). A value of `TBD` stops the run at parse time +until it is set. A run config therefore only needs what differs, e.g. for one +SKiLLS shear branch on candide: + +```yaml +machine: candide +input_type: image_sims +psf_model: fake +psf_dict: /home/hervas/fhervas/workdir_skills/input/psf_files/Full_psf_dict.pickle +tile_list: /path/to/tiles.txt +inputs: + tiles: /n09data/hervas/skills_out/1z2z_grid_3/images/SP_tiles + exposures: /n09data/hervas/skills_out/1z2z_grid_3/images/SP_exp +outputs: + run_dir: /path/to/run + index_db: /path/to/run/index.sqlite +``` + +sp_validation's image-simulation workflow drives these campaigns and measures m +from their final catalogues. On candide, the node-local tile store (bound from the node's 31 GB `/tmp`) does not hold several dense image-sim tiles at once. Set `tile_store_root:` in the run diff --git a/workflow/Snakefile b/workflow/Snakefile index 0a81548c9..b3612ed24 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -42,39 +42,22 @@ from pathlib import Path from snakemake.exceptions import WorkflowError # Resolved relative to THIS file, not the working directory: snakemake runs with -# --directory on /scratch (bin/sp) so .snakemake/ state never lands on /project -# (group quota is a hard 27/27 TiB — a metadata write mid-run died on it live). +# --directory on /scratch (bin/sp) so .snakemake/ state never lands on /project. # -# config.yaml's defaults -- including the per-machine `machines:` table -- -# are ALWAYS loaded first; SP_RUN_CONFIG (set by bin/sp), when given, is -# layered on top via a second, conditional load (workflow.configfile() is a -# plain method -- what the `configfile:` directive itself compiles to -- -# so it can be called from inside an `if`; it merges recursively via -# snakemake.utils.update_config, not by replacement). This matters -# specifically for `machines:`: an external run config (every image-sims -# run, any one-off nibi/candide override) must still see the per-machine -# defaults it exists to provide, or they would only ever apply to the one -# case nobody actually runs -- the bare committed config with no override at -# all. A key given in SP_RUN_CONFIG still wins over the same key from -# config.yaml (later configfile: wins in the merge), so an override remains -# exactly that -- this is the "layer, don't inherit blind" version of what -# used to be a wholesale replace: `machines:` supplies well-scoped defaults -# (gated on `machine:`/`input_type:`, checked against SP_PROFILE below), not -# one campaign's actual paths reused as another's fallback. -_CONFIG_YAML = str(Path(workflow.snakefile).parent / "config.yaml") -_RUN_CONFIG = os.environ.get("SP_RUN_CONFIG") or _CONFIG_YAML -configfile: _CONFIG_YAML +# config.yaml is always loaded; SP_RUN_CONFIG (bin/sp) is merged on top, then +# the `machines:` entry for (machine, input_type) fills whatever is still unset +# (scripts/run_config.py, shared with bin/sp and container.py). +sys.path.insert(0, str(Path(workflow.snakefile).parent / "scripts")) +import run_config # noqa: E402 + +_RUN_CONFIG = os.environ.get("SP_RUN_CONFIG") or str( + Path(workflow.snakefile).parent / "config.yaml") +configfile: str(Path(workflow.snakefile).parent / "config.yaml") if os.environ.get("SP_RUN_CONFIG"): workflow.configfile(os.environ["SP_RUN_CONFIG"]) -# --- input type (needed below, before PSF model) ---------------------------- -# input_type selects WHERE the pixels come from; everything downstream of the -# ingestion stages is the same chain. `image_sims` runs the SKiLLS image -# simulations through the real-data configs: the overlay config dir -# (config/cfis_image_sims) holds real files only for the ingestion INIs whose -# input naming differs, the fake-PSF INIs, and the merge column list -# (final_cat.param), and symlinks the rest into config/cfis, so an m-bias -# measured on the sims calibrates the pipeline that makes the real catalogue. +# input_type picks where the pixels come from (and config/cfis vs +# config/cfis_image_sims); everything after ingestion is the same chain. INPUT_TYPES = {"data", "image_sims"} INPUT_TYPE = config.get("input_type", "data") if INPUT_TYPE not in INPUT_TYPES: @@ -82,134 +65,44 @@ if INPUT_TYPE not in INPUT_TYPES: f"Invalid input_type={INPUT_TYPE!r}; expected one of " f"{sorted(INPUT_TYPES)}.") -# `machine:` (the run config) and SP_PROFILE (the environment variable -# bin/sp reads to pick profiles//, default "nibi") are two independent -# switches for the same fact -- which cluster this is -- set in two different -# places at two different times, and nothing stops them disagreeing: export -# SP_PROFILE=candide but leave a copied run config's `machine: nibi` unedited, -# and jobs run under candide's SLURM settings against nibi's data paths. -# `machine:` has no code default (a wrong SLURM profile is exactly the kind -# of mistake this check exists to catch, so guessing one on its behalf would -# defeat the point); the check is skipped only if `machine:` truly is unset -# (a config predating this key), never silently. +# `machine:` must agree with SP_PROFILE (bin/sp's profile choice, default +# nibi): otherwise jobs run under one cluster's SLURM profile against another +# cluster's paths. Skipped only if `machine:` is unset. SP_PROFILE = os.environ.get("SP_PROFILE", "nibi") MACHINE = config.get("machine") if MACHINE is not None and MACHINE != SP_PROFILE: raise WorkflowError( - f"machine={MACHINE!r} in {_RUN_CONFIG} does not match " - f"SP_PROFILE={SP_PROFILE!r} (env, defaults to 'nibi' when unset). " - f"Set SP_PROFILE={MACHINE} to launch on the machine this config " - f"declares, or fix `machine:` if the config is the stale one.") - -# --- per-machine defaults (config.yaml's `machines:` table) ----------------- -# tile_list / retrieve / inputs / outputs / container differ by machine, and -# on candide also by input_type (the real survey and the SKiLLS simulations -# live in entirely different places, fetched differently) -- `machines:` in -# the run config carries one entry per machine, itself carrying either those -# keys directly (uniform across input_type, e.g. nibi) or nested one level -# under `data:`/`image_sims:` (e.g. candide). Any of these keys given at the -# TOP LEVEL of the run config (outside `machines:`) overrides its machine -# default -- same precedence SP_RUN_CONFIG already has over config.yaml, one -# level down. `$base_dir` inside a string is replaced with that machine -# entry's own `base_dir:`, so nibi's block states the one path that actually -# is common to its tiles/exposures/products/container once, instead of four -# times. A machine default may itself be the sentinel `TBD` (see the check -# below) -- known to be needed, not yet known what it is. -def _expand_base_dir(value, base_dir): - if isinstance(value, str): - return value.replace("$base_dir", base_dir) if base_dir else value - if isinstance(value, dict): - return {k: _expand_base_dir(v, base_dir) for k, v in value.items()} - return value - -_mconf = (config.get("machines") or {}).get(MACHINE) or {} -_sub = _mconf.get(INPUT_TYPE) or {} -_base_dir = _mconf.get("base_dir", "") -_MACHINE_CONTAINER = None -for _key in ("tile_list", "retrieve"): - if _key not in config: - _val = _sub.get(_key, _mconf.get(_key)) - if _val is not None: - config[_key] = _expand_base_dir(_val, _base_dir) -_MACHINE_CONTAINER = _sub.get("container", _mconf.get("container")) -if _MACHINE_CONTAINER is not None: - _MACHINE_CONTAINER = _expand_base_dir(_MACHINE_CONTAINER, _base_dir) -for _section in ("inputs", "outputs"): - _val = _sub.get(_section, _mconf.get(_section)) - if _val: - _merged = _expand_base_dir(dict(_val), _base_dir) - _merged.update(config.get(_section) or {}) - config[_section] = _merged - -# Every path a run genuinely cannot do without (tile_list, inputs.tiles, -# inputs.exposures, outputs.run_dir, outputs.index_db) is checked, after the -# machine defaults above are folded in, against the sentinel value `TBD` -- -# known to be needed, not yet known what it is (typically: this machine -# entry's real value hasn't been confirmed yet, see config.yaml). One clear, -# identical-everywhere error instead of a confusing failure downstream, or -# (worse) a job dying hours in on a node that can't see a path literally -# named "TBD". `container:` is checked separately, below its own resolution. -_PLACEHOLDER = "TBD" -for _key, _value in { - "tile_list": config.get("tile_list"), - "inputs.tiles": (config.get("inputs") or {}).get("tiles"), - "inputs.exposures": (config.get("inputs") or {}).get("exposures"), - "outputs.run_dir": (config.get("outputs") or {}).get("run_dir"), - "outputs.index_db": (config.get("outputs") or {}).get("index_db"), -}.items(): - if _value == _PLACEHOLDER: - raise WorkflowError( - f"{_key!r} resolves to the placeholder {_PLACEHOLDER!r} for " - f"machine={MACHINE!r}, input_type={INPUT_TYPE!r} -- this value " - f"is not known yet (see config.yaml's `machines:` table). Supply " - f"it explicitly in your run config until it is.") - -# Every job's shell runs inside this container (apptainer software-deployment in -# the profile); the user never types apptainer. WHICH image is the one resolution -# order `sp container` exposes — this user's writable sandbox if they built one, -# else their cached SIF if they pulled one, else `container:` above. An empty -# cache therefore lands on exactly the shared /project image the workflow has -# always run, and a package installed into a sandbox reaches the jobs too. -sys.path.insert(0, str(Path(workflow.snakefile).parent / "scripts")) + f"machine={MACHINE!r} ({_RUN_CONFIG}) does not match " + f"SP_PROFILE={SP_PROFILE!r}. Set SP_PROFILE={MACHINE}, or fix " + f"`machine:`.") + +run_config.apply_machine_defaults(config) +_unresolved = run_config.unresolved(config) +if _unresolved: + raise WorkflowError( + f"Unset or {run_config.PLACEHOLDER!r} for machine={MACHINE!r}, " + f"input_type={INPUT_TYPE!r}: {', '.join(_unresolved)}. Set them in " + f"your run config (SP_RUN_CONFIG).") + +# The image every job runs in: this user's sandbox, else their cached SIF, +# else the run config's `container:`. Checked at parse time, since an empty +# container directive passes a dry run and only fails on a node. import container as _container # noqa: E402 -# container.py's `configured_default()` reads `container:` by REGEX off the -# run config FILE, independent of the `config` dict above (it is deliberately -# stdlib-only, no yaml import) -- so a machine-derived container, which only -# exists after the Python-side merge above, is invisible to it. SP_CONTAINER -# is already the documented way to point resolve_image() at a specific image -# instead of the cache (`local_sif()` reads it), so reuse that rather than -# teach container.py about `machines:` too. Set only when nothing more -# specific already exists: an explicit top-level `container:` in the run -# config (which STILL reaches container.py, by file, the normal way) or an -# SP_CONTAINER the user set themselves both take precedence, matching every -# other machine default above. -if _MACHINE_CONTAINER and not os.environ.get("SP_CONTAINER") \ - and not config.get("container"): - os.environ["SP_CONTAINER"] = _MACHINE_CONTAINER - -# Loud at PARSE time, both ways: a broken SP_CONTAINER override, or nothing to -# resolve at all (empty cache AND no `container:` key). The empty string the -# resolver returns for "none" is a valid-looking container directive that passes -# dry-run and only fails once jobs reach a node. try: - _image, _kind = _container.resolve_image() + _image, _kind = _container.resolve_image(config.get("container")) except _container.ContainerError as _exc: raise WorkflowError(str(_exc)) if _kind == "none": raise WorkflowError( - f"No container image resolved: no sandbox, no cached SIF, and no " - f"'{_container.CONFIG_KEY}:' key in {_container.CONFIG_FILE}. " - f"Run `sp container pull`, or restore the key.") + "No container image resolved: no sandbox, no cached SIF, and no " + "`container:` in the run config. Run `sp container pull`, or set it.") container: _image -# --- PSF model --------------------------------------------------------------- -# `fake` is the image-simulation true PSF: no exposure PSF fit, and -# fake_interp_runner writes the tile's galaxy_psf from the simulation's PSF -# dictionary (config `psf_dict`) -- named so that the configs' shared -# `${SP_PSF}_interp_runner` reads it exactly as it reads psfex_interp_runner. Simulations that contain stars can instead run -# psfex or mccd exactly as the real data does. +# `fake` is the image-simulation true PSF: no exposure PSF fit; the tile's +# galaxy_psf comes from `psf_dict` via fake_interp_runner (read by the +# configs as ${SP_PSF}_interp_runner). Sims with stars may use psfex/mccd. PSF_MODELS = {"psfex", "mccd", "fake"} PSF_MODEL = config.get("psf_model", "psfex") if PSF_MODEL not in PSF_MODELS: @@ -225,16 +118,8 @@ if PSF_MODEL == "fake" and not PSF_DICT: raise WorkflowError("psf_model=fake needs `psf_dict:` (the simulation's " "pickled PSF dictionary) in the run config.") -# How tile/exposure images reach the run: `symlink` (a pre-staged local -# mirror -- nibi's layout) or `vos` (fetched from VOSpace per unit -- candide's -# real-data layout, no local mirror needed). A per-cluster choice, so it is a -# run-config key (`retrieve:`) rather than fixed in the committed ini files; -# `unit_pre` exports it as SP_RETRIEVE for config_tile_Git.ini / -# config_exp_Gie.ini. `image_sims` always overrides it to `symlink` below, -# regardless of what the run config says: simulation output is generated -# locally and never lives in VOSpace, so a `data` run config's `retrieve: vos` -# reused as a sims campaign's base must not leak into the sims ingestion -# stages. +# How images reach the run (exported as SP_RETRIEVE for the Git/Gie configs): +# `symlink` to a local mirror, or `vos` download. Sims are always local. RETRIEVE_MODES = {"symlink", "vos"} RETRIEVE_MODE = config.get("retrieve", "symlink") if RETRIEVE_MODE not in RETRIEVE_MODES: @@ -262,8 +147,7 @@ INDEX_DB = Path(OUTPUTS["index_db"]) SCRIPTS = Path(workflow.basedir) / "scripts" # The config chain is the repo's committed directory (D2). The configs and # rules that set their environment variables must be versioned together. There -# is no separate `config_src` setting; input_type picks between the two -# committed dirs. +# is no separate `config_src` setting; input_type picks between the two dirs. CONFIG_DIR = Path(workflow.basedir) / "config" / { "data": "cfis", "image_sims": "cfis_image_sims"}[INPUT_TYPE] @@ -678,11 +562,9 @@ def unit_pre(stage, unit, *, exp_name=None, forest=None, env=None, the committed config dir. It also exports the configured input roots as ``SP_INPUT_TILES`` and - ``SP_INPUT_EXPOSURES``, how to fetch them as ``SP_RETRIEVE`` (``symlink`` - or ``vos``; forced to ``symlink`` for ``image_sims``, see ``RETRIEVE_MODE`` - above), and the PSF choice as ``SP_PSF`` for the committed ini chain (plus - ``PSF_DICT`` for the image-simulation true PSF, which the configs read as - ``${SP_PSF}_interp_runner`` = ``fake_interp_runner``). + ``SP_INPUT_EXPOSURES``, the retrieve mode as ``SP_RETRIEVE`` and the PSF + choice as ``SP_PSF`` for the committed ini chain (plus ``PSF_DICT`` for the + image-simulation true PSF). Every line here is part of each rule's ``params.pre`` and so of the ``params`` rerun trigger: a line added for every rule reruns every finished diff --git a/workflow/bin/sp b/workflow/bin/sp index b64905181..11ce102da 100755 --- a/workflow/bin/sp +++ b/workflow/bin/sp @@ -30,11 +30,10 @@ # SP_PROFILE profiles// to launch with (default: nibi). # SP_SNAKEMAKE_ENV venv to activate; skipped when it does not exist and a # snakemake is already on PATH (e.g. candide). -# SP_RUN_CONFIG a run config that REPLACES workflow/config.yaml (it is not -# layered on it, so no key falls back to the committed run's -# paths). Snapshotted with the code by `sp run`. This is how -# sp_validation drives one campaign per image-simulation -# shear branch. +# SP_RUN_CONFIG a run config merged on top of workflow/config.yaml +# (scripts/run_config.py). Snapshotted with the code by +# `sp run`. This is how sp_validation drives one campaign +# per image-simulation shear branch. set -euo pipefail HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" # workflow/ @@ -56,15 +55,13 @@ elif ! command -v snakemake >/dev/null 2>&1; then exit 2 fi -# Scalar reader for workflow/config.yaml. The venv is active by this point and -# snakemake depends on PyYAML, so this parses the file rather than pattern-matching -# it -- quotes, inline comments and nesting are the parser's problem, not ours. -cfg() { python -c 'import sys, yaml -value = yaml.safe_load(open(sys.argv[1])) -for key in sys.argv[2].split("."): - value = value[key] -print(value or "")' "$CONFIG" "$1"; } +# Resolved run-config value (config.yaml + SP_RUN_CONFIG + machine defaults), +# the same resolution the Snakefile uses. +cfg() { python "$SCRIPTS/run_config.py" "$HERE/config.yaml" "${SP_RUN_CONFIG:-}" "$1"; } RUN_DIR="$(cfg outputs.run_dir)"; INDEX_DB="$(cfg outputs.index_db)" +if [ "${1:-}" != container ] && { [ -z "$RUN_DIR" ] || [ "$RUN_DIR" = TBD ]; }; then + echo "sp: outputs.run_dir is unset or TBD for this machine/input_type" >&2; exit 2 +fi # Snakemake state (.snakemake: metadata, locks, incomplete markers) lives NEXT TO # THE RUN on /scratch, never on the persistent root — the one exception to D5's @@ -149,10 +146,9 @@ PY # directory (NFS) trades ngmix read latency for room; the campaign-unique store # names (LOCAL_TAG) keep concurrent campaigns apart in one root. Profiles with # no `X:/local/scratch` bind (nibi's `--bind /local`) get one appended. - python - "$SNAPSHOT/profiles/$PROFILE/config.yaml" "$CONFIG" <<'PY' -import os, pathlib, re, sys, yaml -f, run_config = pathlib.Path(sys.argv[1]), sys.argv[2] -root = (yaml.safe_load(open(run_config)) or {}).get("tile_store_root") + python - "$SNAPSHOT/profiles/$PROFILE/config.yaml" "$(cfg tile_store_root)" <<'PY' +import os, pathlib, re, sys +f, root = pathlib.Path(sys.argv[1]), sys.argv[2] if root: root = str(root) if not os.path.isabs(root): diff --git a/workflow/config.yaml b/workflow/config.yaml index a3dc1d1d7..ab60bd20c 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -1,213 +1,80 @@ # Run configuration for the ShapePipe Snakemake workflow. # -# This file is nibi's own real, working default (Cail's real-data campaign, -# the checkout's current mainline user) -- `sp run` with no other setup does -# something real and correct on nibi. It is NOT the right config for another -# cluster or a different campaign: `SP_RUN_CONFIG` points `sp run` at a run -# config that REPLACES this file wholesale rather than layering over it -# (workflow/bin/sp; see the README's "Run configuration" section), which is -# how a candide run, an image-sims branch, or a second nibi campaign supplies -# its own paths without touching this one: +# A run config given via SP_RUN_CONFIG is merged on top of this file, then +# the `machines:` entry for (machine, input_type) fills anything still unset: # -# SP_PROFILE= SP_RUN_CONFIG=/path/to/your_run.yaml \ -# workflow/bin/sp run +# SP_PROFILE= SP_RUN_CONFIG=/path/to/run.yaml workflow/bin/sp run # -# A "run" is declared by a tile list plus the paths below. Everything here is -# read at parse time; none of it is a rule input, so editing it (e.g. appending -# tiles) never invalidates completed work — it only changes which jobs exist. +# Everything is read at parse time; editing it (e.g. appending tiles) never +# invalidates completed work. -# Which cluster this run config is for. Must match SP_PROFILE (the -# environment variable that picks profiles//, default "nibi") -- the -# Snakefile raises a clear error at parse time if they disagree, since that -# combination means a job would run under one cluster's SLURM settings -# against another cluster's data paths. +# Cluster; must match SP_PROFILE (default nibi). machine: nibi #machine: candide -# Where the pixels come from: `data` (survey tiles/exposures, config/cfis) or -# `image_sims` (SKiLLS simulations, config/cfis_image_sims -- an overlay that -# differs from config/cfis only in the ingestion stages). An image_sims run -# normally sets `psf_model: fake` and `psf_dict:`, and is driven with its own -# run config through SP_RUN_CONFIG. There is no separate `config_src` -# setting: input_type alone picks between the two committed ini directories, -# which must version together with the rules that set their environment -# variables. +# data (survey, config/cfis) or image_sims (SKiLLS, config/cfis_image_sims). input_type: data -# PSF model used by the exposure and tile interpolation stages: psfex, mccd, or -# fake (image simulations only: the true PSF from `psf_dict`, no exposure fit). +# psfex, mccd, or fake (image_sims only: true PSF from `psf_dict`). psf_model: psfex -# Per-(machine, and on candide also input_type) settings: tile_list, retrieve, -# inputs, outputs, container -- see the Snakefile for the exact merge and the -# `$base_dir` substitution rule. Any of these keys given at the TOP LEVEL of -# this file (outside `machines:`) overrides its machine default; `#retrieve:` -# and `#container:` below are exactly that override point, commented out. A -# value of `TBD` means: known to be needed, not yet known what it is -- -# parsing refuses to proceed past one. +# Per machine and input_type: tile_list, retrieve (symlink|vos), inputs, +# outputs, container. A key set at top level (or in SP_RUN_CONFIG) overrides +# it. `$base_dir` expands to the machine's base_dir. TBD = not known yet; the +# run refuses to start until it is set. # -# THE TWO OUTPUT ROOTS EVERY MACHINE ENTRY DECLARES (D5). -# -# run_dir is the SCRATCH root and the $SP_RUN every config interpolates: bulk -# intermediates, sized so a batch finishes inside the 60-day purge window. The -# sharded per-unit stores live under it: -# /tiles/<2-char prefix>// and /exp/// -# -# products_dir is the PERSISTENT root: the durable, low-volume products — the -# final catalogues (/tiles///final_cat-.fits, -# mirroring the scratch tree shard for shard), the index, the report. ~32-46 MB -# per tile, so a full DR6 campaign is a few hundred GB. The final catalogue is -# also the tile-finished MARKER that cuts a finished tile's exposure edges -# (D5), which is the second reason it cannot sit on scratch: a purge would not -# merely lose a product, it would make every finished tile re-declare inputs -# against exposure stores reclamation already deleted. Unset means "one -# root": products land under run_dir, exactly the pre-D5 layout, which is -# what a fixture or smoke test wants -- omit the key, rather than pointing it -# at run_dir, if that's what you want. -# -# index_db is the run index, and — sharing its directory — missing.json and -# run_report.json. On the persistent root with the catalogues (D5): the index -# is the record of which tile reads which exposure, so it is what a -# post-purge reconstruction would otherwise have to rebuild from tile -# headers. Snakemake's own state is the one durable-looking thing that stays -# on scratch instead (-state; bin/sp explains why). +# outputs: run_dir = scratch root for intermediates; products_dir = persistent +# root for final catalogues, index and report (omit to use run_dir); +# index_db = tile/exposure index. machines: nibi: base_dir: /project/def-mjhudson - - # The real, currently-growing tile list for nibi's smk-gN campaign; a - # different run always supplies its own. - tile_list: $base_dir/cdaley/sp-products/smk-g6/tiles.txt - - # Pre-staged P3 data (get_images RETRIEVE=symlink). - retrieve: symlink - inputs: - tiles: $base_dir/unions-wl/tiles - exposures: $base_dir/unions-wl/exposures - - outputs: - run_dir: /scratch/cdaley/shapepipe-output/smk-g6 - products_dir: $base_dir/cdaley/sp-products/smk-g6 - index_db: $base_dir/cdaley/sp-products/smk-g6/index/run_index.sqlite - - container: $base_dir/cdaley/containers/shapepipe-develop-runtime.sif + data: + tile_list: $base_dir/cdaley/sp-products/smk-g6/tiles.txt + retrieve: symlink + inputs: + tiles: $base_dir/unions-wl/tiles + exposures: $base_dir/unions-wl/exposures + outputs: + run_dir: /scratch/cdaley/shapepipe-output/smk-g6 + products_dir: $base_dir/cdaley/sp-products/smk-g6 + index_db: $base_dir/cdaley/sp-products/smk-g6/index/run_index.sqlite + container: $base_dir/cdaley/containers/shapepipe-develop-runtime.sif candide: base_dir: /n17data/UNIONS/WL - data: - # candide has no pre-staged local mirror of the survey (unlike nibi): - # images are fetched from VOSpace per unit instead. retrieve: vos tile_list: TBD inputs: - tiles: TBD # a vos: URL, once known + tiles: TBD # vos: URL exposures: TBD outputs: run_dir: TBD + index_db: TBD container: /n17data/cdaley/containers/shapepipe_develop-runtime-20260718.sif - image_sims: - # SKiLLS simulation output already lives on local disk. + # tile_list, inputs, outputs: per SKiLLS campaign, in its own run config, + # e.g. inputs.tiles: /n09data/hervas/skills_out//images/SP_tiles retrieve: symlink container: /n17data/cdaley/containers/shapepipe_im_sims-runtime.sif - # tile_list, inputs and outputs are intentionally absent here: each - # SKiLLS grid/shear-branch campaign supplies its own via SP_RUN_CONFIG - # (e.g. inputs.tiles: /n09data/hervas/skills_out//images/SP_tiles) - # -- this entry only saves such a run from restating retrieve/container - # too. - -# How the tile/exposure images in `inputs:` above reach the run: `symlink` -# (a pre-staged local mirror) or `vos` (fetched from VOSpace per unit, no -# local mirror). Set per machine (and, on candide, per input_type) above; -# ignored either way for `input_type: image_sims`, which is always `symlink` -# (simulation output is generated locally and never lives in VOSpace). Give -# it here directly only to override the machine default for one run. -#retrieve: symlink -# The container every job runs inside (apptainer software-deployment in the -# profile). Set per machine (and, on candide, per input_type) above; give it -# here directly only to override that default for one run. -#container: ... - -# Rolling exposure-store reclamation (D5). When true, the COMPUTE DAG grows one -# `clean_exposure` job per exposure. It fires once every campaign tile that reads -# that exposure has its vignets, deletes the exposure's store AND its manifests, -# and leaves `cleaned.json`, which absorbs the manifests — `sp report` reads them -# back out of the tombstone and reports the exposure as `cleaned`. Set it false -# for a small control run whose exposure stores you want to keep around (e.g. to -# re-measure only the shape chain rather than rebuild the whole tile) -- true is -# the right default for anything that isn't a deliberately-kept fixture. +# Per-exposure store reclamation once every tile reading it has its vignets. clean: true -# Rolling TILE-store reclamation (D5). When true, the COMPUTE DAG grows one -# `clean_tile` job per in-scope tile, ordered after that tile's own final_cat, so -# a tile's scratch store goes as soon as its chain lands rather than at the end -# of the batch. It deletes the tile's whole /tiles/// tree -# except four things another mechanism owns — `cleaned.json`, -# `manifests/tile_vignets.json` (clean_exposure's eligibility currency), -# `manifests/tile_find_exposures.json` (prepare_all_tiles' target) and the Fe -# exposure list under output/ (build_index re-reads it at every compute parse). -# clean_tile.py argues each one. -# -# A SEPARATE SWITCH FROM `clean:` ABOVE, AND NOT AN OVERSIGHT. Exposure -# reclamation is reversible in the only sense that matters: a tile appended later -# rebuilds the exposure chain from VOS, expensively but completely. TILE -# reclamation DESTROYS THE PER-TILE AUDIT TRAIL and nothing rebuilds it. -# TWO tools read it: sp-products/tools/sp_tilecost.py and sp_costmodel.py. Both -# attribute the fused tile_shape group job's cost per tile from that tile's -# scratch store — the SExtractor catalogue (NAXIS2 of the sexcat is the object -# count, the cost model's independent variable; sp_costmodel also reads its -# EPOCH_k extensions for the geometric epoch count) and the eight tile_ngmix -# benchmark TSVs — and that model is how every mem_mb and runtime number in -# tile.smk was derived. The sexcat is ~380 MB per tile and cannot be kept; the -# benchmark ROWS are absorbed into the tombstone, so the measurements survive -# but not at the paths the tools read. -# -# (The third artifact they read, run_sp_tile_ngmix_Ngu/, is NOT lost to this -# flag: tile_ngmix declares its chunk dir temp(), so snakemake reclaims it as -# soon as tile_merge_cats runs. It is already absent from every finished tile, -# clean_tiles or not.) -# -# The honest one-line summary: REQUIRED FOR ANY BATCH OVER ~850 TILES, AND IT -# TURNS OFF PER-TILE COST ATTRIBUTION. The arithmetic: a finished tile leaves -# 1.19 GiB across 137 inodes (measured on 186.307), so a 1 TiB scratch quota is -# full at 859 tiles, and DR6's 23,114 tiles would want 26.9 TiB and 3.17M inodes -# against a 1M quota. Reclaimed, a tile costs 10 inodes and 12.9 KB — 231k inodes -# and ~300 MB for the whole of DR6. -# -# Set it false for the same kind of small control run `clean:` above would be -# false for, when the per-tile store is what a re-measurement needs to read. +# Per-tile store reclamation after the tile's final_cat. Needed above ~850 +# tiles (scratch quota); deletes the per-tile data sp_tilecost/sp_costmodel +# read. Set both false for a small control run you want to re-measure. clean_tiles: true -# Tiles that may NOT pin an exposure store (default: empty). -# -# An exposure is eligible for cleaning only once EVERY consuming tile has its -# vignets. One permanently-failed tile therefore holds all its exposures -# (~8 measured on P3: 7.9 exposures/tile) for the life of the campaign. A tile listed here is dropped from the -# consumer sets, and its exposures become eligible. -# -# READ THIS BEFORE ADDING A TILE. Ignoring a tile is a decision to give up its -# exposures' stores. If you later retry that tile, those exposure chains are -# gone and will be REBUILT from scratch — get_images, split, psf, per -# exposure. That is correct, and expensive. Ignore a tile when you have decided -# it is dead, not while you are still debugging it. +# Failed tiles that should not keep their exposures from being cleaned. +# Their exposures are rebuilt from scratch if the tile is retried. clean_ignore_tiles: [] -# ngmix within-tile chunking: static N chunks (closed ID ranges computed -# per-tile, in-job, from the tile's own sexcat). +# ngmix chunks per tile. ngmix_chunks: 8 -# Where the node-local tile store lives (tile.smk's tile_local(): each tile_shape -# group's vignette and WCS stores). Unset means the profile's `/local/scratch` -# bind: nibi's /local NVMe, candide's 31 GB node /tmp. Set it to an absolute -# directory to bind that instead, for this campaign only; bin/sp rewrites the -# snapshot's profile and creates the directory. Not a rerun trigger, so it may be -# set on a resume. Image-sim campaigns on candide want a shared directory (dense -# tiles and failed groups' orphans fill /tmp); the cost is ngmix read latency. +# Node-local tile store (vignets, WCS). Unset: the profile's /local/scratch +# bind (nibi NVMe, candide node /tmp). Dense image-sim campaigns on candide +# want a shared directory instead. # tile_store_root: /n23data1//shapepipe/tilestore - -# The container's installed shapepipe is overridden by the launching worktree -# via each profile's own --env PYTHONPATH; there is no setting for it here. -# build_index.py's missing-tile fraction threshold is passed by workflow/bin/sp -# (SP_MISSING_THRESHOLD, default 0.0 = any missing tile is fatal). diff --git a/workflow/config/cfis/config_exp_Gie.ini b/workflow/config/cfis/config_exp_Gie.ini index bfd65b84e..9f69fbaee 100644 --- a/workflow/config/cfis/config_exp_Gie.ini +++ b/workflow/config/cfis/config_exp_Gie.ini @@ -86,10 +86,7 @@ INPUT_NUMBERING = \d{6} # Output file pattern without number OUTPUT_FILE_PATTERN = image-, weight-, flag- -# Method to retrieve images, one in 'vos', 'symlink' -- set by the run -# config's `retrieve:` key (unit_pre exports SP_RETRIEVE), not fixed here: -# nibi mirrors the exposures locally (symlink), candide fetches them from -# VOSpace (vos). +# 'vos' or 'symlink', from the run config's `retrieve:` RETRIEVE = $SP_RETRIEVE # If RETRIEVE=vos, number of attempts to download diff --git a/workflow/config/cfis/config_tile_Git.ini b/workflow/config/cfis/config_tile_Git.ini index 87a189f07..8aad4193b 100644 --- a/workflow/config/cfis/config_tile_Git.ini +++ b/workflow/config/cfis/config_tile_Git.ini @@ -80,9 +80,7 @@ INPUT_NUMBERING = \d{3}\.\d{3} # Output file pattern without number OUTPUT_FILE_PATTERN = CFIS_image-, CFIS_weight- -# Copy/download method, one in 'vos', 'symlink' -- set by the run config's -# `retrieve:` key (unit_pre exports SP_RETRIEVE), not fixed here: nibi mirrors -# the tiles locally (symlink), candide fetches them from VOSpace (vos). +# 'vos' or 'symlink', from the run config's `retrieve:` RETRIEVE = $SP_RETRIEVE # If RETRIEVE=vos, number of attempts to download diff --git a/workflow/config/cfis_image_sims/config_exp_Gie.ini b/workflow/config/cfis_image_sims/config_exp_Gie.ini index c5eee0f4b..bbf1b98af 100644 --- a/workflow/config/cfis_image_sims/config_exp_Gie.ini +++ b/workflow/config/cfis_image_sims/config_exp_Gie.ini @@ -91,10 +91,7 @@ INPUT_NUMBERING = \d{7} # Output file pattern without number OUTPUT_FILE_PATTERN = image-, weight-, flag- -# Method to retrieve images, one in 'vos', 'symlink'. Not a sims/data -# difference (unlike the fields marked above): the Snakefile always exports -# SP_RETRIEVE=symlink for image_sims regardless of the run config's -# `retrieve:` key, since simulation output is always local, never VOSpace. +# 'vos' or 'symlink', from the run config's `retrieve:` RETRIEVE = $SP_RETRIEVE # If RETRIEVE=vos, number of attempts to download diff --git a/workflow/config/cfis_image_sims/config_tile_Git.ini b/workflow/config/cfis_image_sims/config_tile_Git.ini index 4489788cb..d638ff559 100644 --- a/workflow/config/cfis_image_sims/config_tile_Git.ini +++ b/workflow/config/cfis_image_sims/config_tile_Git.ini @@ -85,10 +85,7 @@ INPUT_NUMBERING = \d{3}-\d{3} # Output file pattern without number OUTPUT_FILE_PATTERN = CFIS_image-, CFIS_weight- -# Copy/download method, one in 'vos', 'symlink'. Not a sims/data difference -# (unlike the fields marked above): the Snakefile always exports -# SP_RETRIEVE=symlink for image_sims regardless of the run config's -# `retrieve:` key, since simulation output is always local, never VOSpace. +# 'vos' or 'symlink', from the run config's `retrieve:` RETRIEVE = $SP_RETRIEVE # If RETRIEVE=vos, number of attempts to download diff --git a/workflow/scripts/completeness.py b/workflow/scripts/completeness.py index bc8975064..14a68b7ae 100644 --- a/workflow/scripts/completeness.py +++ b/workflow/scripts/completeness.py @@ -354,13 +354,10 @@ def build_manifest(stage, run_dir, unit, stage_subdir=None): def write_if_changed(path: Path, text: str) -> None: - """Write only when the bytes differ — see the module docstring on mtime. + """Write only when the bytes differ (see the module docstring on mtime). - Via a same-directory temp file + ``os.replace``, not ``path.write_text``: - the latter truncates before writing, so a reader (``run_report.py``, run - automatically at compute's end) can catch the file empty mid-write. Same - directory keeps the replace on one filesystem, which is what makes it - atomic; a losing writer's temp file is unlinked rather than left behind. + Atomic (temp file + os.replace): a reader such as run_report.py must never + see the file truncated mid-write. """ path.parent.mkdir(parents=True, exist_ok=True) if not path.exists() or path.read_text() != text: diff --git a/workflow/scripts/container.py b/workflow/scripts/container.py index c35f91ad8..defafb495 100644 --- a/workflow/scripts/container.py +++ b/workflow/scripts/container.py @@ -56,12 +56,10 @@ class ContainerError(Exception): # the slim variant the workflow runs. CONTAINER_URI = "docker://ghcr.io/cosmostat/shapepipe:develop-runtime" -# The single source of truth for the fallback image: the workflow's own -# `container:` key, which is also what the Snakefile reads. Written down once, -# here, so the CLI and the workflow cannot disagree about the default. -# SP_RUN_CONFIG replaces it for a run driven with its own config (bin/sp). -CONFIG_FILE = Path(os.environ.get("SP_RUN_CONFIG") - or Path(__file__).resolve().parents[1] / "config.yaml") +# The fallback image is the run config's resolved `container:` (config.yaml, +# SP_RUN_CONFIG on top, then the machine default; see run_config.py). +BASE_CONFIG = Path(__file__).resolve().parents[1] / "config.yaml" +CONFIG_FILE = Path(os.environ.get("SP_RUN_CONFIG") or BASE_CONFIG) CONFIG_KEY = "container" # The profile whose `apptainer-args:` every workflow job runs under. `exec` reads @@ -92,17 +90,22 @@ class ContainerError(Exception): def configured_default(): - """Return the ``container:`` path from workflow/config.yaml, or ``None``. + """Return the resolved ``container:`` of the run config, or ``None``. - A deliberately minimal scalar read (the same one bin/sp does in sed): this - module is stdlib-only, so there is no yaml to import. + Via run_config.py when PyYAML is available; on a bare host without it, + a top-level ``container:`` line in the run config file is still read. """ try: - text = CONFIG_FILE.read_text() - except OSError: - return None - match = re.search(rf"^{CONFIG_KEY}:[ \t]*(\S+)", text, re.MULTILINE) - return match.group(1) if match else None + import run_config + except ImportError: + try: + text = CONFIG_FILE.read_text() + except OSError: + return None + match = re.search(rf"^{CONFIG_KEY}:[ \t]*(\S+)", text, re.MULTILINE) + return match.group(1) if match else None + run = os.environ.get("SP_RUN_CONFIG") + return run_config.load(BASE_CONFIG, run).get(CONFIG_KEY) or None def profile_apptainer_args(): @@ -131,9 +134,12 @@ def local_sandbox(): return (Path(override) if override else DEFAULT_SANDBOX).expanduser() -def resolve_image(): +def resolve_image(configured=None): """Return ``(path, kind)`` for the image everything should run. + ``configured`` is the run config's resolved ``container:`` when the + caller already has it (the Snakefile); otherwise it is looked up. + ``kind`` is ``"sandbox"``, ``"sif"``, ``"configured"`` (the shared /project image named in config.yaml -- the default when the cache is empty) or ``"none"``. @@ -152,7 +158,7 @@ def resolve_image(): f"SP_CONTAINER={os.environ['SP_CONTAINER']} does not exist " f"(resolved to {sif}). Unset it, or point it at an image that does." ) - default = configured_default() + default = configured or configured_default() if default: return default, "configured" return "", "none" diff --git a/workflow/scripts/run_config.py b/workflow/scripts/run_config.py new file mode 100644 index 000000000..5daa69382 --- /dev/null +++ b/workflow/scripts/run_config.py @@ -0,0 +1,83 @@ +#!/usr/bin/env python3 +"""Resolve a run config: config.yaml, then SP_RUN_CONFIG merged on top, then +the `machines:` entry for (machine, input_type) filling whatever is still +unset. One definition shared by the Snakefile, bin/sp and container.py. + +CLI (used by bin/sp): run_config.py CONFIG_YAML RUN_CONFIG KEY[.SUBKEY] +prints the resolved value, or an empty line if unset. RUN_CONFIG may be "". +""" + +import sys + +import yaml + +PLACEHOLDER = "TBD" +MACHINE_KEYS = ("tile_list", "retrieve", "container", "inputs", "outputs") +REQUIRED = ("tile_list", "inputs.tiles", "inputs.exposures", + "outputs.run_dir", "outputs.index_db") + + +def merge(base, over): + """Recursive dict merge; `over` wins (as snakemake's update_config).""" + out = dict(base) + for key, value in over.items(): + if isinstance(value, dict) and isinstance(out.get(key), dict): + out[key] = merge(out[key], value) + else: + out[key] = value + return out + + +def _expand(value, base_dir): + if isinstance(value, str): + return value.replace("$base_dir", base_dir) + if isinstance(value, dict): + return {k: _expand(v, base_dir) for k, v in value.items()} + return value + + +def apply_machine_defaults(config): + """Fill unset MACHINE_KEYS from machines[machine][input_type], in place. + + A key already in `config` wins; for `inputs`/`outputs` the merge is per + sub-key. `$base_dir` expands to machines[machine].base_dir. + """ + entry = (config.get("machines") or {}).get(config.get("machine")) or {} + defaults = entry.get(config.get("input_type", "data")) or {} + base_dir = str(entry.get("base_dir", "")) + for key in MACHINE_KEYS: + if key not in defaults: + continue + value = _expand(defaults[key], base_dir) + if isinstance(value, dict): + config[key] = merge(value, config.get(key) or {}) + else: + config.setdefault(key, value) + return config + + +def get(config, dotted): + value = config + for part in dotted.split("."): + value = value.get(part) if isinstance(value, dict) else None + return value + + +def unresolved(config): + """REQUIRED keys that are unset or still the placeholder.""" + return [k for k in REQUIRED if get(config, k) in (None, "", PLACEHOLDER)] + + +def load(config_yaml, run_config=None): + with open(config_yaml) as f: + config = yaml.safe_load(f) or {} + if run_config: + with open(run_config) as f: + config = merge(config, yaml.safe_load(f) or {}) + return apply_machine_defaults(config) + + +if __name__ == "__main__": + config_yaml, run_config, key = sys.argv[1:4] + value = get(load(config_yaml, run_config or None), key) + print("" if value is None else value) From 084e305a6090f9ca65e2054eba1879601e471e99 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Tue, 15 Sep 2026 19:17:33 +0200 Subject: [PATCH 37/85] run_config: `run:` name, expanded as $run in run-config paths config.yaml gains `run: smk-g6`, and nibi's paths use $run instead of the literal campaign name. $base_dir and $run now expand in values set in the run config itself too, not only in machine defaults, so a sims run config can name its output directories by `run:`. A required path left holding an unexpanded $variable is reported like TBD. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_013NzWDtdQbsK1VhvYbDnmTo --- workflow/README.md | 2 +- workflow/config.yaml | 15 +++++++++------ workflow/scripts/run_config.py | 35 ++++++++++++++++++++-------------- 3 files changed, 31 insertions(+), 21 deletions(-) diff --git a/workflow/README.md b/workflow/README.md index c0652c1eb..0ab08c0d4 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -76,7 +76,7 @@ the code. `machine:` (which must match `SP_PROFILE`, default `nibi`) and `input_type:` then select an entry of the `machines:` table, which supplies `tile_list`, `retrieve` (`symlink` or `vos`), `inputs`, `outputs` and `container` for any of these the run config leaves unset (`$base_dir` expands -to that machine's `base_dir`). A value of `TBD` stops the run at parse time +to that machine's `base_dir`, `$run` to the run config's `run:`). A value of `TBD` stops the run at parse time until it is set. A run config therefore only needs what differs, e.g. for one SKiLLS shear branch on candide: diff --git a/workflow/config.yaml b/workflow/config.yaml index ab60bd20c..59307fb2b 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -18,10 +18,13 @@ input_type: data # psfex, mccd, or fake (image_sims only: true PSF from `psf_dict`). psf_model: psfex +# Run name, available as `$run` in the paths below. +run: smk-g6 + # Per machine and input_type: tile_list, retrieve (symlink|vos), inputs, # outputs, container. A key set at top level (or in SP_RUN_CONFIG) overrides -# it. `$base_dir` expands to the machine's base_dir. TBD = not known yet; the -# run refuses to start until it is set. +# it. `$base_dir` expands to the machine's base_dir, `$run` to `run:`. +# TBD = not known yet; the run refuses to start until it is set. # # outputs: run_dir = scratch root for intermediates; products_dir = persistent # root for final catalogues, index and report (omit to use run_dir); @@ -30,15 +33,15 @@ machines: nibi: base_dir: /project/def-mjhudson data: - tile_list: $base_dir/cdaley/sp-products/smk-g6/tiles.txt + tile_list: $base_dir/cdaley/sp-products/$run/tiles.txt retrieve: symlink inputs: tiles: $base_dir/unions-wl/tiles exposures: $base_dir/unions-wl/exposures outputs: - run_dir: /scratch/cdaley/shapepipe-output/smk-g6 - products_dir: $base_dir/cdaley/sp-products/smk-g6 - index_db: $base_dir/cdaley/sp-products/smk-g6/index/run_index.sqlite + run_dir: /scratch/cdaley/shapepipe-output/$run + products_dir: $base_dir/cdaley/sp-products/$run + index_db: $base_dir/cdaley/sp-products/$run/index/run_index.sqlite container: $base_dir/cdaley/containers/shapepipe-develop-runtime.sif candide: diff --git a/workflow/scripts/run_config.py b/workflow/scripts/run_config.py index 5daa69382..481219f6e 100644 --- a/workflow/scripts/run_config.py +++ b/workflow/scripts/run_config.py @@ -7,6 +7,7 @@ prints the resolved value, or an empty line if unset. RUN_CONFIG may be "". """ +import re import sys import yaml @@ -28,11 +29,14 @@ def merge(base, over): return out -def _expand(value, base_dir): +def _expand(value, variables): + """Replace $name for each set name in `variables` (base_dir, run).""" if isinstance(value, str): - return value.replace("$base_dir", base_dir) + return re.sub(r"\$(\w+)", + lambda m: str(variables.get(m.group(1)) or m.group(0)), + value) if isinstance(value, dict): - return {k: _expand(v, base_dir) for k, v in value.items()} + return {k: _expand(v, variables) for k, v in value.items()} return value @@ -40,19 +44,20 @@ def apply_machine_defaults(config): """Fill unset MACHINE_KEYS from machines[machine][input_type], in place. A key already in `config` wins; for `inputs`/`outputs` the merge is per - sub-key. `$base_dir` expands to machines[machine].base_dir. + sub-key. In all of these, `$base_dir` expands to machines[machine].base_dir + and `$run` to the top-level `run:`. """ entry = (config.get("machines") or {}).get(config.get("machine")) or {} defaults = entry.get(config.get("input_type", "data")) or {} - base_dir = str(entry.get("base_dir", "")) + variables = {"base_dir": entry.get("base_dir"), "run": config.get("run")} for key in MACHINE_KEYS: - if key not in defaults: - continue - value = _expand(defaults[key], base_dir) - if isinstance(value, dict): - config[key] = merge(value, config.get(key) or {}) - else: - config.setdefault(key, value) + default = defaults.get(key) + if isinstance(default, dict): + config[key] = merge(default, config.get(key) or {}) + elif default is not None: + config.setdefault(key, default) + if key in config: + config[key] = _expand(config[key], variables) return config @@ -64,8 +69,10 @@ def get(config, dotted): def unresolved(config): - """REQUIRED keys that are unset or still the placeholder.""" - return [k for k in REQUIRED if get(config, k) in (None, "", PLACEHOLDER)] + """REQUIRED keys that are unset, the placeholder, or hold an unexpanded + $variable (e.g. `$run` with no `run:` set).""" + return [k for k in REQUIRED + if get(config, k) in (None, "", PLACEHOLDER) or "$" in str(get(config, k))] def load(config_yaml, run_config=None): From 78cc0ca2dd23ebd3e480c907dfa1f8bd829705c4 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Thu, 17 Sep 2026 07:18:07 +0200 Subject: [PATCH 38/85] added user run config template --- scripts/python/create_final_cat.py | 2 -- workflow/README.md | 4 +-- workflow/Snakefile | 10 +++---- workflow/config.yaml | 45 +++++++++++++++++++----------- workflow/run_template.yaml | 18 ++++++++++++ workflow/scripts/run_config.py | 7 ++++- 6 files changed, 60 insertions(+), 26 deletions(-) create mode 100644 workflow/run_template.yaml diff --git a/scripts/python/create_final_cat.py b/scripts/python/create_final_cat.py index d98232818..dc14a3324 100755 --- a/scripts/python/create_final_cat.py +++ b/scripts/python/create_final_cat.py @@ -391,8 +391,6 @@ def process(params): if params["image_sims"]: patch_name = params["patch"] - # The Snakemake workflow writes an undated run_sp_tile_Mc; the legacy - # bash runner wrote run_sp_tile_Mc_. run_prefix = "run_sp_tile_Mc*" else: patch_name = rf"P{params['patch']}" diff --git a/workflow/README.md b/workflow/README.md index 0ab08c0d4..c23bb2fab 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -72,8 +72,8 @@ SP_PROFILE=candide SP_RUN_CONFIG=/path/run_1p2z_grid_1.yaml workflow/bin/sp run ## Run configuration `SP_RUN_CONFIG` is merged on top of `workflow/config.yaml` and snapshotted with -the code. `machine:` (which must match `SP_PROFILE`, default `nibi`) and -`input_type:` then select an entry of the `machines:` table, which supplies +the code. `SP_PROFILE` (default `nibi`, or `machine:` in the run config, which must +agree with it) and `input_type:` then select an entry of the `machines:` table, which supplies `tile_list`, `retrieve` (`symlink` or `vos`), `inputs`, `outputs` and `container` for any of these the run config leaves unset (`$base_dir` expands to that machine's `base_dir`, `$run` to the run config's `run:`). A value of `TBD` stops the run at parse time diff --git a/workflow/Snakefile b/workflow/Snakefile index b3612ed24..9288053ec 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -65,12 +65,12 @@ if INPUT_TYPE not in INPUT_TYPES: f"Invalid input_type={INPUT_TYPE!r}; expected one of " f"{sorted(INPUT_TYPES)}.") -# `machine:` must agree with SP_PROFILE (bin/sp's profile choice, default -# nibi): otherwise jobs run under one cluster's SLURM profile against another -# cluster's paths. Skipped only if `machine:` is unset. +# The machine is SP_PROFILE (bin/sp's profile choice, default nibi). A run +# config may also state `machine:`, and must then agree with it: otherwise +# jobs run under one cluster's SLURM profile against another's paths. SP_PROFILE = os.environ.get("SP_PROFILE", "nibi") -MACHINE = config.get("machine") -if MACHINE is not None and MACHINE != SP_PROFILE: +MACHINE = config.get("machine") or SP_PROFILE +if config.get("machine") is not None and config["machine"] != SP_PROFILE: raise WorkflowError( f"machine={MACHINE!r} ({_RUN_CONFIG}) does not match " f"SP_PROFILE={SP_PROFILE!r}. Set SP_PROFILE={MACHINE}, or fix " diff --git a/workflow/config.yaml b/workflow/config.yaml index 59307fb2b..d5f31f039 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -1,5 +1,14 @@ -# Run configuration for the ShapePipe Snakemake workflow. -# +# Main configuration for the ShapePipe Snakemake workflow. + +# To run, first set a profile (machine/architecture): + +# SP_PROFILE=nibi|candide + +# Optional, set the path to your personal config file, to override default +# settings or add variables + +# SP_RUN_CONFIG=/path/to/my_run.yaml + # A run config given via SP_RUN_CONFIG is merged on top of this file, then # the `machines:` entry for (machine, input_type) fills anything still unset: # @@ -8,10 +17,6 @@ # Everything is read at parse time; editing it (e.g. appending tiles) never # invalidates completed work. -# Cluster; must match SP_PROFILE (default nibi). -machine: nibi -#machine: candide - # data (survey, config/cfis) or image_sims (SKiLLS, config/cfis_image_sims). input_type: data @@ -21,14 +26,17 @@ psf_model: psfex # Run name, available as `$run` in the paths below. run: smk-g6 -# Per machine and input_type: tile_list, retrieve (symlink|vos), inputs, -# outputs, container. A key set at top level (or in SP_RUN_CONFIG) overrides -# it. `$base_dir` expands to the machine's base_dir, `$run` to `run:`. -# TBD = not known yet; the run refuses to start until it is set. -# -# outputs: run_dir = scratch root for intermediates; products_dir = persistent -# root for final catalogues, index and report (omit to use run_dir); -# index_db = tile/exposure index. +# Entries per machine (SP_PROFILE, default nibi; implemented: nibi, candide) +# and input_type. A run config may instead state `machine:`, which must +# then agree with SP_PROFILE. +# - base_dir +# - tile_list: text file with tile IDs +# - retrieve: method to get input data, allowed are symlink, vos +# - inputs.tiles,, .exposures: path for input tile and exposure +# - outputs.run_dir: (scratch) run directory, where tmp files will be stored +# - outputs.products_dir: path to final products +# - outputs.index_db: path to bookkeeping index sqlite file + machines: nibi: base_dir: /project/def-mjhudson @@ -57,9 +65,14 @@ machines: index_db: TBD container: /n17data/cdaley/containers/shapepipe_develop-runtime-20260718.sif image_sims: - # tile_list, inputs, outputs: per SKiLLS campaign, in its own run config, - # e.g. inputs.tiles: /n09data/hervas/skills_out//images/SP_tiles retrieve: symlink + tile_list: TBD + inputs: + # Need to be pointed to appropriate grid + tiles: /n09data/hervas/skills_out/1z2z_grid_1/images/SP_tiles + exposures: /n09data/hervas/skills_out/1z2z_grid_1/images/SP_exp + # Dictionary of PSF model stamps + psf_dict: /home/hervas/fhervas/workdir_skills/input/psf_files/Full_psf_dict.pickle container: /n17data/cdaley/containers/shapepipe_im_sims-runtime.sif # Per-exposure store reclamation once every tile reading it has its vignets. diff --git a/workflow/run_template.yaml b/workflow/run_template.yaml new file mode 100644 index 000000000..c75bdb736 --- /dev/null +++ b/workflow/run_template.yaml @@ -0,0 +1,18 @@ +# Run config template file, can be specified along with main config.yaml for +# ShapePipe run. + +my_base_dir: + +tile_list: $my_base_dir/tiles.txt + +sim: 1z2z_grid_1 + +inputs: + tiles: /n09data/hervas/skills_out/$sim/images/SP_tiles + exposures: /n09data/hervas/skills_out/$sim/images/SP_exp + +outputs: + run_dir: $my_base_dir/$sim/run + products_dir: $my_base_dir/$sim/product + +shapepipe_repo: diff --git a/workflow/scripts/run_config.py b/workflow/scripts/run_config.py index 481219f6e..c6ed4f3d9 100644 --- a/workflow/scripts/run_config.py +++ b/workflow/scripts/run_config.py @@ -7,6 +7,7 @@ prints the resolved value, or an empty line if unset. RUN_CONFIG may be "". """ +import os import re import sys @@ -43,11 +44,15 @@ def _expand(value, variables): def apply_machine_defaults(config): """Fill unset MACHINE_KEYS from machines[machine][input_type], in place. + The machine is `machine:` when the run config states one, else SP_PROFILE + (default nibi) -- the same value bin/sp picks the SLURM profile with. + A key already in `config` wins; for `inputs`/`outputs` the merge is per sub-key. In all of these, `$base_dir` expands to machines[machine].base_dir and `$run` to the top-level `run:`. """ - entry = (config.get("machines") or {}).get(config.get("machine")) or {} + machine = config.get("machine") or os.environ.get("SP_PROFILE", "nibi") + entry = (config.get("machines") or {}).get(machine) or {} defaults = entry.get(config.get("input_type", "data")) or {} variables = {"base_dir": entry.get("base_dir"), "run": config.get("run")} for key in MACHINE_KEYS: From d3b088a8d02e920dd8226562af38afc91242f332 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Thu, 17 Sep 2026 07:46:58 +0200 Subject: [PATCH 39/85] bin/sp: -c/--config-file for the run config, instead of SP_RUN_CONFIG The run config was only settable through an exported variable: invisible in the command, absent from shell history and job scripts, and silently inherited by later commands. `sp -c my_run.yaml` now takes it on any verb; sp consumes the flag, makes the path absolute and exports SP_RUN_CONFIG itself, which is what the Snakefile, container.py and the snapshot read. An already-set SP_RUN_CONFIG is still honoured, and -c wins. -c is sp's own flag, not snakemake's --cores; cores pass through as --cores/-j. Closes the `sp run --config-file` half of #891 step 1. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_013NzWDtdQbsK1VhvYbDnmTo --- workflow/README.md | 8 +++++--- workflow/bin/sp | 48 +++++++++++++++++++++++++++++++++++--------- workflow/config.yaml | 13 +++++------- 3 files changed, 49 insertions(+), 20 deletions(-) diff --git a/workflow/README.md b/workflow/README.md index c23bb2fab..16795aa32 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -66,13 +66,15 @@ Simulations that contain stars can run `psfex` or `mccd` exactly as the data do. One campaign per shear branch, each with its own run config: ```bash -SP_PROFILE=candide SP_RUN_CONFIG=/path/run_1p2z_grid_1.yaml workflow/bin/sp run +SP_PROFILE=candide workflow/bin/sp run -c /path/run_1p2z_grid_1.yaml ``` ## Run configuration -`SP_RUN_CONFIG` is merged on top of `workflow/config.yaml` and snapshotted with -the code. `SP_PROFILE` (default `nibi`, or `machine:` in the run config, which must +A run config passed with `-c/--config-file` is merged on top of +`workflow/config.yaml` and snapshotted with the code. (`-c` is `sp`'s own flag; +pass snakemake's cores as `--cores`/`-j`. `SP_RUN_CONFIG` still works and is what +the jobs read.) `SP_PROFILE` (default `nibi`, or `machine:` in the run config, which must agree with it) and `input_type:` then select an entry of the `machines:` table, which supplies `tile_list`, `retrieve` (`symlink` or `vos`), `inputs`, `outputs` and `container` for any of these the run config leaves unset (`$base_dir` expands diff --git a/workflow/bin/sp b/workflow/bin/sp index 11ce102da..bee176e26 100755 --- a/workflow/bin/sp +++ b/workflow/bin/sp @@ -3,7 +3,8 @@ # # Three verbs, nothing else: # -# sp run [ARGS...] bring the products on disk up to date with the tile list. +# sp run [-c FILE] [ARGS...] +# bring the products on disk up to date with the tile list. # Snapshots the code into the state dir first (see "the # launch code snapshot" below) and runs out of the copy, so # editing the checkout mid-campaign cannot reach the jobs. @@ -25,24 +26,53 @@ # One entry point, so a fresh tmux or a restart after a crash always launches # with the right state. # -# Three environment variables make it drivable off nibi and from outside: +# THE RUN CONFIG: -c/--config-file , on any verb. It is merged on top of +# workflow/config.yaml (scripts/run_config.py), so it states only what differs, +# and `sp run` snapshots it with the code. This is how one campaign per +# image-simulation shear branch is driven. NOTE: -c here is sp's own flag, not +# snakemake's --cores; pass cores as --cores/-j, which sp forwards. +# +# Environment: # # SP_PROFILE profiles// to launch with (default: nibi). # SP_SNAKEMAKE_ENV venv to activate; skipped when it does not exist and a -# snakemake is already on PATH (e.g. candide). -# SP_RUN_CONFIG a run config merged on top of workflow/config.yaml -# (scripts/run_config.py). Snapshotted with the code by -# `sp run`. This is how sp_validation drives one campaign -# per image-simulation shear branch. +# snakemake is already on PATH (e.g. candide). A site +# detail, not something a run needs to set. +# SP_RUN_CONFIG what -c sets, and what the jobs read. Honoured when +# already set, so older callers keep working; -c wins. set -euo pipefail HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" # workflow/ REPO="$(dirname "$HERE")" VENV="${SP_SNAKEMAKE_ENV:-/project/def-mjhudson/cdaley/snakemake-env}" SCRIPTS="$HERE/scripts" -CONFIG="${SP_RUN_CONFIG:-$HERE/config.yaml}" PROFILE="${SP_PROFILE:-nibi}" -[ -f "$CONFIG" ] || { echo "sp: run config $CONFIG does not exist" >&2; exit 2; } + +# -c/--config-file is consumed HERE and never forwarded: snakemake has its own +# --configfile with different merge semantics, and this one also has to reach +# the jobs (as SP_RUN_CONFIG) and the snapshot. Everything else passes through. +RUN_CONFIG="${SP_RUN_CONFIG:-}" +_sp_args=() +while [ $# -gt 0 ]; do + case "$1" in + -c|--config|--config-file) + [ $# -ge 2 ] || { echo "sp: $1 needs a file" >&2; exit 2; } + RUN_CONFIG="$2"; shift 2 ;; + -c=*|--config=*|--config-file=*) + RUN_CONFIG="${1#*=}"; shift ;; + *) _sp_args+=("$1"); shift ;; + esac +done +set -- ${_sp_args[@]+"${_sp_args[@]}"} + +if [ -n "$RUN_CONFIG" ]; then + [ -f "$RUN_CONFIG" ] || { + echo "sp: run config $RUN_CONFIG does not exist" >&2; exit 2; } + # Absolute: the jobs and the snapshot resolve it from other directories. + RUN_CONFIG="$(cd "$(dirname "$RUN_CONFIG")" && pwd)/$(basename "$RUN_CONFIG")" + export SP_RUN_CONFIG="$RUN_CONFIG" +fi + [ -d "$REPO/profiles/$PROFILE" ] || { echo "sp: no profile $REPO/profiles/$PROFILE (SP_PROFILE=$PROFILE)" >&2; exit 2; } diff --git a/workflow/config.yaml b/workflow/config.yaml index d5f31f039..2ac109136 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -4,15 +4,12 @@ # SP_PROFILE=nibi|candide -# Optional, set the path to your personal config file, to override default -# settings or add variables - -# SP_RUN_CONFIG=/path/to/my_run.yaml - -# A run config given via SP_RUN_CONFIG is merged on top of this file, then -# the `machines:` entry for (machine, input_type) fills anything still unset: +# Your own run config, passed with -c, is merged on top of this file; then the +# `machines:` entry for (machine, input_type) fills anything still unset: +# +# SP_PROFILE= workflow/bin/sp run -c /path/to/my_run.yaml # -# SP_PROFILE= SP_RUN_CONFIG=/path/to/run.yaml workflow/bin/sp run +# It therefore states only what differs from the defaults below. # # Everything is read at parse time; editing it (e.g. appending tiles) never # invalidates completed work. From 746ca9c3c4da62d8ddc5f81b37829e9fa9352ca9 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Thu, 17 Sep 2026 07:53:21 +0200 Subject: [PATCH 40/85] bin/sp: document -c, the two-file merge and the environment in the header Typical invocations, the order the run config and config.yaml are read in (and that the machines: table fills the rest), the run_config.py one-liner for inspecting a resolved value, why -c is not snakemake's --cores, and the remaining SP_* variables including SP_PHASE, which sp sets itself. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_013NzWDtdQbsK1VhvYbDnmTo --- workflow/bin/sp | 40 ++++++++++++++++++++++++++++++++++------ 1 file changed, 34 insertions(+), 6 deletions(-) diff --git a/workflow/bin/sp b/workflow/bin/sp index bee176e26..dc3049ee3 100755 --- a/workflow/bin/sp +++ b/workflow/bin/sp @@ -26,20 +26,48 @@ # One entry point, so a fresh tmux or a restart after a crash always launches # with the right state. # -# THE RUN CONFIG: -c/--config-file , on any verb. It is merged on top of -# workflow/config.yaml (scripts/run_config.py), so it states only what differs, -# and `sp run` snapshots it with the code. This is how one campaign per -# image-simulation shear branch is driven. NOTE: -c here is sp's own flag, not -# snakemake's --cores; pass cores as --cores/-j, which sp forwards. +# Typical use, from the repo root: +# +# SP_PROFILE=candide workflow/bin/sp run -c ~/my_run.yaml # a campaign +# SP_PROFILE=candide workflow/bin/sp run -c ~/my_run.yaml -n # dry run +# SP_PROFILE=candide workflow/bin/sp report -c ~/my_run.yaml # status now +# workflow/bin/sp run # config.yaml only +# +# THE RUN CONFIG: -c/--config-file , accepted on any verb (also +# --config, and the =FILE forms). Two files are always read, in this order: +# +# 1. workflow/config.yaml committed defaults, including the `machines:` +# table (per machine and input_type) +# 2. the -c file merged on top, so it states ONLY what differs +# +# then anything still unset is filled from the `machines:` entry for this +# machine (SP_PROFILE, or `machine:`) and `input_type`. scripts/run_config.py +# does that resolution and is shared with container.py; to see what a pair of +# files resolves to without launching: +# +# python scripts/run_config.py workflow/config.yaml ~/my_run.yaml outputs.run_dir +# +# sp consumes -c, makes it absolute, exports it as SP_RUN_CONFIG (what the +# jobs and the snapshot read) and never forwards it: snakemake has its own +# --configfile with different merge semantics. +# +# NOTE: -c is sp's own flag, NOT snakemake's --cores. Cores pass through as +# --cores/-j, like every other snakemake argument. # # Environment: # -# SP_PROFILE profiles// to launch with (default: nibi). +# SP_PROFILE profiles// to launch with (default: nibi). Also +# selects the `machines:` entry, and must agree with a +# `machine:` key if the run config states one. # SP_SNAKEMAKE_ENV venv to activate; skipped when it does not exist and a # snakemake is already on PATH (e.g. candide). A site # detail, not something a run needs to set. # SP_RUN_CONFIG what -c sets, and what the jobs read. Honoured when # already set, so older callers keep working; -c wins. +# SP_STATE_DIR snakemake state + code snapshot (default -state). +# SP_CONTAINER image to run in, ahead of the run config's `container:`. +# SP_PHASE set by sp itself (prepare/compute/passthrough); the +# Snakefile refuses to parse without it. Never set it. set -euo pipefail HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" # workflow/ From 4e0bedcb6106c525a8c4c24901f8029ffdda86f7 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Thu, 17 Sep 2026 07:58:31 +0200 Subject: [PATCH 41/85] improved commeents --- workflow/bin/sp | 7 +++++-- 1 file changed, 5 insertions(+), 2 deletions(-) diff --git a/workflow/bin/sp b/workflow/bin/sp index dc3049ee3..189ac89ff 100755 --- a/workflow/bin/sp +++ b/workflow/bin/sp @@ -1,7 +1,7 @@ #!/usr/bin/env bash -# sp — the committed launcher for the ShapePipe Snakemake workflow (PRD #848 D1). +# Main ShapePipe launch script for Snakemake workflow # -# Three verbs, nothing else: +# Usage: # # sp run [-c FILE] [ARGS...] # bring the products on disk up to date with the tile list. @@ -12,6 +12,9 @@ # 1. PREPARE snakemake prepare_all_tiles # 2. COMPUTE snakemake all <- its PARSE builds the index # ARGS (--jobs, -n, --forcerun, ...) pass through to BOTH. +# FILE is a user-specified config file that can be used +# to add and/or overwrite entries in the default config +# file workflow/config.yaml (read automatically) # sp report [ARGS...] emit run_report.json now (mid-run is fine). # sp container VERB manage the image every job runs inside: pull, status, # sandbox, exec, resolve. `sp container --help` documents From 3e269c452d5f635132f002792aa4c2ba831d5577 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Thu, 17 Sep 2026 13:21:24 +0200 Subject: [PATCH 42/85] final_cat_merge, star_cat_merge: record code provenance in the HDF5 attributes MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit sp_validation opens the campaign's merged catalogues without any link back to the code that built them; a `final_cat_.hdf5` or `full_starcat_.hdf5` produced under one commit is indistinguishable from one produced under another. `sp run` already snapshots that information per campaign (bin/sp's `$STATE_DIR/code/snapshot.json`), so this wires it into the two merge rules: hdf5_reconcile.code_provenance() reads the snapshot (or says "unknown" when the workflow runs outside `sp run`) and apply() stamps code_head/code_branch/code_dirty/code_snapshot_at, plus a newline-joined code_dirty_files when the snapshot was dirty, onto the output file's root — written only when the file is actually rewritten, matching the existing no-op-leaves-the-mtime-alone contract. Co-Authored-By: Claude Fable 5.1 --- workflow/Snakefile | 5 ++++ workflow/rules/exposure.smk | 2 ++ workflow/rules/tile.smk | 2 ++ workflow/scripts/hdf5_reconcile.py | 41 ++++++++++++++++++++++++++++- workflow/scripts/merge_final_cat.py | 6 ++++- workflow/scripts/merge_star_cat.py | 6 ++++- 6 files changed, 59 insertions(+), 3 deletions(-) diff --git a/workflow/Snakefile b/workflow/Snakefile index 4317e871f..1a45598ed 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -110,6 +110,11 @@ SCRIPTS = Path(workflow.basedir) / "scripts" # rules that set their environment variables must be versioned together. There # is no `config_src` knob. CONFIG_DIR = Path(workflow.basedir) / "config" / "cfis" +# `sp run`'s code snapshot (bin/sp), a sibling of workflow.basedir inside +# $STATE_DIR/code. Read by the campaign-level merges so the products they write +# carry the code that produced them; absent when this Snakefile is driven +# outside `sp run`, which the merges tolerate (hdf5_reconcile.code_provenance). +SNAPSHOT_JSON = Path(workflow.basedir).parent / "snapshot.json" sys.path.insert(0, str(SCRIPTS)) import build_index # noqa: E402 diff --git a/workflow/rules/exposure.smk b/workflow/rules/exposure.smk index 5845adb8d..708840fdf 100644 --- a/workflow/rules/exposure.smk +++ b/workflow/rules/exposure.smk @@ -277,6 +277,7 @@ rule star_cat_merge: tile_list = str(config["tile_list"]), index_db = str(INDEX_DB), campaign = CAMPAIGN, + snapshot = str(SNAPSHOT_JSON), inputs = unit_fingerprint(star_cat_exposures()), script_hash = MERGE_STAR_HASH threads: 1 @@ -302,3 +303,4 @@ rule star_cat_merge: " --tile-list '{params.tile_list}' --index-db '{params.index_db}'" " --output {output.star_cat}" " --campaign '{params.campaign}'" + " --snapshot-json '{params.snapshot}'" diff --git a/workflow/rules/tile.smk b/workflow/rules/tile.smk index 46d84f02f..51e1ff3f5 100644 --- a/workflow/rules/tile.smk +++ b/workflow/rules/tile.smk @@ -948,6 +948,7 @@ rule final_cat_merge: index_db = str(INDEX_DB), param_file = str(CONFIG_DIR / "final_cat.param"), campaign = CAMPAIGN, + snapshot = str(SNAPSHOT_JSON), inputs = unit_fingerprint(TILES_READY), script_hash = MERGE_FINAL_HASH threads: 1 @@ -970,3 +971,4 @@ rule final_cat_merge: " --output {output.merged}" " --campaign '{params.campaign}'" " --param-file '{params.param_file}'" + " --snapshot-json '{params.snapshot}'" diff --git a/workflow/scripts/hdf5_reconcile.py b/workflow/scripts/hdf5_reconcile.py index f2e3ff325..c99101860 100644 --- a/workflow/scripts/hdf5_reconcile.py +++ b/workflow/scripts/hdf5_reconcile.py @@ -45,6 +45,7 @@ """ import hashlib +import json import shutil import sys from pathlib import Path @@ -57,6 +58,25 @@ def schema_digest(columns) -> str: return hashlib.md5("\n".join(columns).encode()).hexdigest()[:16] +def code_provenance(snapshot_json) -> dict: + """The launch code's identity, to stamp onto the merged file's root. + + ``snapshot_json`` is ``sp run``'s code snapshot (``bin/sp``'s + ``$STATE_DIR/code/snapshot.json``), passed through by the calling rule. A + workflow driven outside ``sp run`` has no such file — the merge still + succeeds, and the caller writes ``code_head = "unknown"`` rather than + failing an otherwise-good build. + """ + if not snapshot_json or not Path(snapshot_json).exists(): + return {"head": "unknown"} + data = json.loads(Path(snapshot_json).read_text()) + out = {k: data[k] for k in ("head", "branch", "dirty", "taken_at") + if k in data} + if data.get("dirty") and data.get("dirty_files"): + out["dirty_files"] = data["dirty_files"] + return out + + def stamp(path: Path) -> tuple: """A source's identity, as recorded on the dataset built from it. @@ -163,13 +183,22 @@ def check_sole_group(output: Path, group_path: str) -> None: def apply(output: Path, group_path: str, todo: Plan, units: list, read, - digest: str, count_attr: str) -> None: + digest: str, count_attr: str, provenance: dict | None = None) -> None: """Carry the plan out on a tmp file, then move it into place. ``read(unit, source)`` returns the structured array for one unit; it is called only for the units the plan names, which is what makes an append cheap. + ``provenance`` (``code_provenance()``'s return) is stamped onto the file's + root as ``code_head``/``code_branch``/``code_dirty``/``code_snapshot_at``, + plus ``code_dirty_files`` (newline-joined) when the snapshot was dirty. It + is written here, alongside ``count_attr`` and ``param_digest``, rather than + on every no-op invocation: reconciling is planned against a read-only open, + and an empty plan must leave the file's mtime alone (see the module + docstring), so a run that changes no data never touches the file even if + the code that would have produced it has moved on. + TWO WAYS TO BUILD THE TMP, and which one is used is about SPACE, not speed. HDF5 never reclaims the space a deleted dataset occupied, so a file that is copied and then edited in place grows for the life of the campaign — every @@ -221,6 +250,16 @@ def apply(output: Path, group_path: str, todo: Plan, units: list, read, stamp(source) f.attrs[count_attr] = len(group) f.attrs["param_digest"] = digest + if provenance: + f.attrs["code_head"] = provenance.get("head", "unknown") + for key, attr in (("branch", "code_branch"), + ("dirty", "code_dirty"), + ("taken_at", "code_snapshot_at")): + if key in provenance: + f.attrs[attr] = provenance[key] + if provenance.get("dirty_files"): + f.attrs["code_dirty_files"] = \ + "\n".join(provenance["dirty_files"]) tmp.replace(output) # atomic: same filesystem finally: tmp.unlink(missing_ok=True) diff --git a/workflow/scripts/merge_final_cat.py b/workflow/scripts/merge_final_cat.py index 2db61e7e0..aaca2298c 100644 --- a/workflow/scripts/merge_final_cat.py +++ b/workflow/scripts/merge_final_cat.py @@ -143,6 +143,9 @@ def main() -> None: p.add_argument("--param-file", required=True, type=Path, help="workflow/config/cfis/final_cat.param — the column list") p.add_argument("--hdu", type=int, default=1) + p.add_argument("--snapshot-json", type=Path, default=None, + help="sp run's code snapshot (bin/sp's " + "$STATE_DIR/code/snapshot.json); absent outside sp run") args = p.parse_args() cfc = load_create_final_cat() @@ -175,7 +178,8 @@ def read_tile(tile, path): f"({len(tiles)} tile(s))") return hdf5_reconcile.apply(args.output, group_path, todo, tiles, read_tile, - digest, "n_tiles") + digest, "n_tiles", + hdf5_reconcile.code_provenance(args.snapshot_json)) print(f"[merge_final_cat] {todo.describe()} -> {args.output} " f"({len(tiles)} tile(s), {len(param_list)} column(s), " f"group {group_path})") diff --git a/workflow/scripts/merge_star_cat.py b/workflow/scripts/merge_star_cat.py index 84ba633d1..c9b4e45f1 100644 --- a/workflow/scripts/merge_star_cat.py +++ b/workflow/scripts/merge_star_cat.py @@ -290,6 +290,9 @@ def main() -> None: p.add_argument("--output", required=True, type=Path) p.add_argument("--campaign", required=True, help="named in the log; the group name is fixed") + p.add_argument("--snapshot-json", type=Path, default=None, + help="sp run's code snapshot (bin/sp's " + "$STATE_DIR/code/snapshot.json); absent outside sp run") args = p.parse_args() manifest_paths = manifests(args.products_dir, args.tile_list, args.index_db) @@ -315,7 +318,8 @@ def main() -> None: f"({len(chosen)} exposure(s))") return hdf5_reconcile.apply(args.output, GROUP, todo, chosen, read_exposure, - digest, "n_exposures") + digest, "n_exposures", + hdf5_reconcile.code_provenance(args.snapshot_json)) print(f"[merge_star_cat] {todo.describe()} -> {args.output} " f"({len(chosen)} exposure(s), {len(ALL_COLUMNS)} column(s), " f"campaign {args.campaign})") From fbfdf3ba0e0f43c02143cb03f480af0d3e8e1c97 Mon Sep 17 00:00:00 2001 From: martinkilbinger Date: Wed, 23 Sep 2026 14:03:48 +0200 Subject: [PATCH 43/85] workflow smoothed; running until hdf5 file --- profiles/candide/config.yaml | 9 ++ scripts/python/create_final_cat.py | 110 ++++++++++++++++++--- src/shapepipe/modules/get_images_runner.py | 7 +- workflow/Snakefile | 50 ++++++++++ workflow/bin/sp | 72 +++++++------- workflow/config.yaml | 46 ++++----- workflow/scripts/run_config.py | 10 +- 7 files changed, 224 insertions(+), 80 deletions(-) diff --git a/profiles/candide/config.yaml b/profiles/candide/config.yaml index bacf48071..2932e64eb 100644 --- a/profiles/candide/config.yaml +++ b/profiles/candide/config.yaml @@ -29,6 +29,15 @@ software-deployment-method: [apptainer] apptainer-args: "--cleanenv --env OMP_NUM_THREADS=1 --env MALLOC_ARENA_MAX=2 --env PYTHONPATH=/n17data/mkilbing/astro/repositories/github/shapepipe/src --bind /home,/automnt,/n17data,/n23data1,/n09data --bind /tmp:/local/scratch" latency-wait: 60 # NFS: wait for outputs to appear after a job keep-going: true # a failed job poisons only its cone; siblings run on +# Transient node-level failures are the common case on a shared cluster: +# a fused tile_shape group died at joblib's fork_exec with BlockingIOError +# [Errno 11] because the node was out of process slots, killing one tile of +# 39 outright. Without this, attempt is always 1, so every +# `mem_mb = lambda wc, attempt: N * attempt` in tile.smk is dead code and a +# one-off OOM or EAGAIN is fatal. Safe for the fused group: a retry +# resubmits the same member set, tile_vignets included (tile.smk's +# "one sharp edge" note). tile_ngmix keeps its own rule-level retries: 2. +retries: 2 rerun-incomplete: true # re-do jobs left incomplete by an unclean death show-failed-logs: true printshellcmds: true diff --git a/scripts/python/create_final_cat.py b/scripts/python/create_final_cat.py index dc14a3324..7571cb052 100755 --- a/scripts/python/create_final_cat.py +++ b/scripts/python/create_final_cat.py @@ -27,6 +27,56 @@ from cs_util import logging +def params_from_run_config(params, defaults): + """Fill unset paths from a workflow run config. + + The workflow already knows where a campaign writes, so a manual merge + should not have to restate it. Resolution goes through the workflow's own + resolver (workflow/scripts/run_config.py), layering the run config on + workflow/config.yaml and then the machines: table, so what lands here is + what the rules would have used. + + Only values still at their default are filled -- an explicit flag always + wins. Nothing is derived for the data path: its patch naming differs and + is not this function's business. + """ + repo = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) + sys.path.insert(0, os.path.join(repo, "workflow", "scripts")) + import run_config as _rc + + cfg = _rc.load(os.path.join(repo, "workflow", "config.yaml"), + params["run_config"]) + if cfg.get("input_type") != "image_sims": + raise ValueError( + f"run config {params['run_config']} has input_type=" + f"{cfg.get('input_type', 'data')!r}; -c derives paths for " + "image_sims only") + + outputs = cfg.get("outputs") or {} + products = outputs.get("products_dir") or outputs.get("run_dir") + if not products: + raise ValueError(f"run config {params['run_config']} sets neither " + "outputs.products_dir nor outputs.run_dir") + + # The same derivation as the workflow's merge_final_cats rule: the patch + # dir is the branch dir holding product/tiles, and -i is its parent. + patch_dir = os.path.dirname(os.path.normpath(products)) + patch = os.path.basename(patch_dir) + derived = { + "image_sims": True, + "input_root_dir": os.path.dirname(patch_dir), + "patch": patch, + "merged_cat_path": os.path.join(products, f"final_cat_{patch}.hdf5"), + "output_summary": os.path.join(products, "n_tiles_final.txt"), + "param_path": os.path.join(repo, "workflow", "config", + "cfis_image_sims", "final_cat.param"), + } + for key, value in derived.items(): + if params.get(key) == defaults.get(key): + params[key] = value + return params + + def set_params_from_command_line(args): """Set Params From Command Line. @@ -48,6 +98,12 @@ def set_params_from_command_line(args): _params[key] = options[key] del options + + # A run config fills in whatever is still at its default + # (see the docstring); explicit flags always win. + if _params.get("run_config"): + _defaults, _, _, _ = params_default() + _params = params_from_run_config(_params, _defaults) # Save calling command logging.log_command(args) @@ -72,6 +128,7 @@ def params_default(): "ID": None, "single_op": None, "image_sims": False, + "run_config": None, } _short_options = { "input_root_dir": "-i", @@ -82,6 +139,7 @@ def params_default(): "output_summary": "-o", "single_op": "-s", "image_sims": "-I", + "run_config": "-c", } _types = { "hdu_num": "int", @@ -98,6 +156,9 @@ def params_default(): "ID": "ID for single-ID operation, default={}", "single_op": "single ID operation, allowed are 'check', 'add', 'remove'; default={}", "image_sims": "image simulations mode (different dir layout and run prefix), default={}", + "run_config": "workflow run config (e.g. sp_1p2z_grid_1.yaml); fills" + " in the paths below that were not given explicitly," + " default={}", } return _params, _short_options, _types, _help_strings @@ -373,9 +434,17 @@ def collect_tile_ids_image_sims(patch_path): list of (tile_id, tile_path) tuples """ id_pattern = re.compile(r"^\d+\.\d+$") - tiles_root = os.path.join(patch_path, "tiles") + # Two layouts. The Gen-2 / native runs keep tiles under the patch dir + # itself; the unified workflow (shapepipe #891) publishes to a separate + # products root, and clean_tile deletes the run-dir copies -- so on a + # reclaimed campaign /product/tiles holds the ONLY catalogues. result = [] - if not os.path.isdir(tiles_root): + for sub in ("tiles", os.path.join("product", "tiles"), + os.path.join("products", "tiles")): + tiles_root = os.path.join(patch_path, sub) + if os.path.isdir(tiles_root): + break + else: return result for prefix in os.listdir(tiles_root): prefix_path = os.path.join(tiles_root, prefix) @@ -387,6 +456,28 @@ def collect_tile_ids_image_sims(patch_path): return result +def find_final_cat(id, id_path, run_prefix): + """Path of this tile's final catalogue, or None. + + Flat layout first: the unified workflow copies one catalogue per tile to + /tiles///final_cat-.fits, in DOT form and with + no run sub-tree. Then the legacy layout, where the catalogue sits under the + tile's own shapepipe run dir in DASH form, newest run wins. + """ + flat = os.path.join(id_path, f"final_cat-{id}.fits") + if os.path.exists(flat): + return flat + + base_pattern = os.path.join(id_path, "output", run_prefix) + all_matches = [d for d in glob.glob(base_pattern) if os.path.isdir(d)] + if not all_matches: + return None + newest_dir = max(all_matches, key=os.path.getmtime) + id_dash = re.sub(r"\.", "-", id) + legacy = f"{newest_dir}/make_cat_runner/output/final_cat-{id_dash}.fits" + return legacy if os.path.exists(legacy) else None + + def process(params): if params["image_sims"]: @@ -446,22 +537,11 @@ def process(params): print(f"Skipping {id} (already processed)") continue - base_pattern = os.path.join(id_path, "output", run_prefix) - all_matches = [d for d in glob.glob(base_pattern) if os.path.isdir(d)] - if not all_matches: + fits_file = find_final_cat(id, id_path, run_prefix) + if fits_file is None: if params["verbose"]: print(f"Final cat for {id} not found, continuing") continue - newest_dir = max(all_matches, key=os.path.getmtime) - - id_dash = re.sub(r"\.", "-", id) - fits_file = f"{newest_dir}/make_cat_runner/output/final_cat-{id_dash}.fits" - - # Exclude unsuccessful run without output FITS file - if not os.path.exists(fits_file): - if params["verbose"]: - print(f"Run without output file found for {id}, skipping") - continue extracted_data, dtype = read_data(fits_file, params) diff --git a/src/shapepipe/modules/get_images_runner.py b/src/shapepipe/modules/get_images_runner.py index b1a03b15a..28e81b7d5 100644 --- a/src/shapepipe/modules/get_images_runner.py +++ b/src/shapepipe/modules/get_images_runner.py @@ -27,8 +27,11 @@ def get_images_runner( """Define The Get Images Runner.""" # Read config file section - # Copy/download method - retrieve_method = config.get(module_config_sec, "RETRIEVE") + # Copy/download method. getexpanded, not get: the workflow's ini sets + # RETRIEVE = $SP_RETRIEVE (a run-config key since d64f88ed), and plain get + # returns the literal "$SP_RETRIEVE", which then fails the check below. + # RETRIEVE_OPTIONS a few lines down already reads the same way. + retrieve_method = config.getexpanded(module_config_sec, "RETRIEVE") retrieve_ok = ["vos", "symlink"] if retrieve_method not in retrieve_ok: raise ValueError( diff --git a/workflow/Snakefile b/workflow/Snakefile index 9288053ec..2befdb5ec 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -672,9 +672,59 @@ include: "rules/tile.smk" # mask generation.) localrules: all, prepare_all_tiles, clean_exposure, clean_tile +# --- merged catalogue (image sims only) ------------------------------------ +# The sp_validation side starts from ONE hdf5 per branch, not 39 FITS tiles, so +# the campaign is not finished until they are merged. Image sims only: the data +# path's patches are merged separately, on a different naming convention. +# +# Paths are derived from products_dir rather than from `run:`, so they hold +# whatever the run config calls things. create_final_cat.py takes a root to +# scan (-i) and a patch name (-P) that is BOTH the directory to match under it +# and the hdf5 group, so the patch directory is the branch dir and the root is +# its parent -- exactly the `-i .. -P ` the old sp_validation rule ran +# from inside the branch. +MERGE_PATCH_DIR = PRODUCTS_DIR.parent # the branch dir, holding product/tiles +MERGE_ROOT = MERGE_PATCH_DIR.parent # -i +MERGE_PATCH = MERGE_PATCH_DIR.name # -P, and the hdf5 group +MERGED_CAT = PRODUCTS_DIR / f"final_cat_{MERGE_PATCH}.hdf5" +CREATE_FINAL_CAT = Path(workflow.basedir).parent / "scripts" / "python" / "create_final_cat.py" + + +def merge_targets(): + """The merged catalogue, once every ready tile has published.""" + return [str(MERGED_CAT)] if INPUT_TYPE == "image_sims" and TILES_READY else [] + + +rule merge_final_cats: + """Merge the per-tile final catalogues into one hdf5 (image sims). + + Reads the PRODUCTS root, not the run dir: clean_tile deletes the run-dir + copies, so on a reclaimed campaign the published ones are all that is left. + Ordering needs no edge to clean_tile for the same reason. + """ + input: + [final_cat(t) for t in TILES_READY], + output: + cat = str(MERGED_CAT), + n_tiles = str(PRODUCTS_DIR / "n_tiles_final.txt"), + params: + root = str(MERGE_ROOT), + patch = MERGE_PATCH, + param_file = str(CONFIG_DIR / "final_cat.param"), + threads: 1 + resources: + mem_mb = lambda wc, attempt: 8000 * attempt, + runtime = 120 + shell: + f"python {CREATE_FINAL_CAT} -I" + " -m {output.cat} -i {params.root} -p {params.param_file}" + " -P {params.patch} -o {output.n_tiles} -v" + + rule all: input: [final_cat(t) for t in TILES_READY], + merge_targets(), clean_targets(), clean_tile_targets(), diff --git a/workflow/bin/sp b/workflow/bin/sp index 189ac89ff..c4dc50995 100755 --- a/workflow/bin/sp +++ b/workflow/bin/sp @@ -36,41 +36,24 @@ # SP_PROFILE=candide workflow/bin/sp report -c ~/my_run.yaml # status now # workflow/bin/sp run # config.yaml only # -# THE RUN CONFIG: -c/--config-file , accepted on any verb (also -# --config, and the =FILE forms). Two files are always read, in this order: -# -# 1. workflow/config.yaml committed defaults, including the `machines:` -# table (per machine and input_type) -# 2. the -c file merged on top, so it states ONLY what differs -# -# then anything still unset is filled from the `machines:` entry for this -# machine (SP_PROFILE, or `machine:`) and `input_type`. scripts/run_config.py -# does that resolution and is shared with container.py; to see what a pair of -# files resolves to without launching: +# Run config: -c FILE, also --config-file or --config. Read after +# workflow/config.yaml and merged on top; anything still unset comes from the +# machines: entry for SP_PROFILE and input_type (scripts/run_config.py). +# To check a resolved value without running: # # python scripts/run_config.py workflow/config.yaml ~/my_run.yaml outputs.run_dir # -# sp consumes -c, makes it absolute, exports it as SP_RUN_CONFIG (what the -# jobs and the snapshot read) and never forwards it: snakemake has its own -# --configfile with different merge semantics. -# -# NOTE: -c is sp's own flag, NOT snakemake's --cores. Cores pass through as -# --cores/-j, like every other snakemake argument. +# -c is sp's flag, not snakemake's --cores; pass cores as --cores or -j. # # Environment: # -# SP_PROFILE profiles// to launch with (default: nibi). Also -# selects the `machines:` entry, and must agree with a -# `machine:` key if the run config states one. -# SP_SNAKEMAKE_ENV venv to activate; skipped when it does not exist and a -# snakemake is already on PATH (e.g. candide). A site -# detail, not something a run needs to set. -# SP_RUN_CONFIG what -c sets, and what the jobs read. Honoured when -# already set, so older callers keep working; -c wins. -# SP_STATE_DIR snakemake state + code snapshot (default -state). -# SP_CONTAINER image to run in, ahead of the run config's `container:`. -# SP_PHASE set by sp itself (prepare/compute/passthrough); the -# Snakefile refuses to parse without it. Never set it. +# SP_PROFILE profiles/ to use (default nibi); also selects the +# machines: entry, and must match `machine:` if set +# SP_SNAKEMAKE_ENV venv to activate; skipped if snakemake is on PATH +# SP_RUN_CONFIG same as -c, and what the jobs read; -c wins +# SP_STATE_DIR snakemake state and code snapshot (default -state) +# SP_CONTAINER image to use, before the config's container: +# SP_PHASE set by sp (prepare/compute/passthrough); do not set set -euo pipefail HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" # workflow/ @@ -79,9 +62,8 @@ VENV="${SP_SNAKEMAKE_ENV:-/project/def-mjhudson/cdaley/snakemake-env}" SCRIPTS="$HERE/scripts" PROFILE="${SP_PROFILE:-nibi}" -# -c/--config-file is consumed HERE and never forwarded: snakemake has its own -# --configfile with different merge semantics, and this one also has to reach -# the jobs (as SP_RUN_CONFIG) and the snapshot. Everything else passes through. +# -c is consumed here, not forwarded: snakemake's own --configfile merges +# differently, and the jobs and the snapshot need it as SP_RUN_CONFIG. RUN_CONFIG="${SP_RUN_CONFIG:-}" _sp_args=() while [ $# -gt 0 ]; do @@ -99,7 +81,7 @@ set -- ${_sp_args[@]+"${_sp_args[@]}"} if [ -n "$RUN_CONFIG" ]; then [ -f "$RUN_CONFIG" ] || { echo "sp: run config $RUN_CONFIG does not exist" >&2; exit 2; } - # Absolute: the jobs and the snapshot resolve it from other directories. + # Absolute: jobs and the snapshot run from other directories. RUN_CONFIG="$(cd "$(dirname "$RUN_CONFIG")" && pwd)/$(basename "$RUN_CONFIG")" export SP_RUN_CONFIG="$RUN_CONFIG" fi @@ -107,7 +89,18 @@ fi [ -d "$REPO/profiles/$PROFILE" ] || { echo "sp: no profile $REPO/profiles/$PROFILE (SP_PROFILE=$PROFILE)" >&2; exit 2; } -module load apptainer/1.4.5 2>/dev/null || true +# snakemake resolves `apptainer` via PATH at job runtime, so all this has to +# do is put it there. candide already has it in /usr/bin and ships no +# modulefile for it, so probing for one only produces a scary ERROR line; +# skip the module entirely when apptainer is already usable. Both streams are +# dropped because the message arrives on stderr under environment-modules and +# can arrive on stdout under Lmod. +if ! command -v apptainer >/dev/null 2>&1; then + module load apptainer/1.4.5 >/dev/null 2>&1 || true + command -v apptainer >/dev/null 2>&1 || { + echo "sp: apptainer is not on PATH and 'module load apptainer/1.4.5' failed" >&2 + exit 2; } +fi if [ -f "$VENV/bin/activate" ]; then # shellcheck disable=SC1091 source "$VENV/bin/activate" @@ -116,8 +109,8 @@ elif ! command -v snakemake >/dev/null 2>&1; then exit 2 fi -# Resolved run-config value (config.yaml + SP_RUN_CONFIG + machine defaults), -# the same resolution the Snakefile uses. +# Resolved config value (config.yaml + run config + machine defaults), as the +# Snakefile resolves it. cfg() { python "$SCRIPTS/run_config.py" "$HERE/config.yaml" "${SP_RUN_CONFIG:-}" "$1"; } RUN_DIR="$(cfg outputs.run_dir)"; INDEX_DB="$(cfg outputs.index_db)" if [ "${1:-}" != container ] && { [ -z "$RUN_DIR" ] || [ "$RUN_DIR" = TBD ]; }; then @@ -150,7 +143,8 @@ STATE_DIR="${SP_STATE_DIR:-${RUN_DIR}-state}"; mkdir -p "$STATE_DIR" # WHAT. `sp run` copies the code it is about to launch into $STATE_DIR/code and # runs the campaign entirely out of that copy: the Snakefile, the rules, the # scripts, the ini chain (symlinks DEREFERENCED -- workflow/config/cfis points -# into example/, and the copy must be self-contained), src/, and the profile. +# into example/, and the copy must be self-contained), src/, the profile, and +# the repo's top-level scripts/ (576 KB), which merge_final_cats runs out of. # Every workflow-internal path hangs off `workflow.basedir`, which IS the # snapshot, so they all follow it for free; the profile's PYTHONPATH pin is the # one that cannot (YAML splices nothing) and is rewritten below. @@ -169,10 +163,10 @@ snapshot_code() { mkdir -p "$SNAPSHOT" if command -v rsync >/dev/null 2>&1; then rsync -a --delete --copy-links --exclude '__pycache__' --exclude '*.egg-info' \ - "$HERE" "$REPO/src" "$REPO/profiles" "$SNAPSHOT/" + "$HERE" "$REPO/src" "$REPO/profiles" "$REPO/scripts" "$SNAPSHOT/" else rm -rf "$SNAPSHOT"; mkdir -p "$SNAPSHOT" - cp -rL "$HERE" "$REPO/src" "$REPO/profiles" "$SNAPSHOT/" + cp -rL "$HERE" "$REPO/src" "$REPO/profiles" "$REPO/scripts" "$SNAPSHOT/" find "$SNAPSHOT" -name __pycache__ -type d -prune -exec rm -rf {} + fi diff --git a/workflow/config.yaml b/workflow/config.yaml index 2ac109136..76a7ee3d9 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -1,24 +1,25 @@ # Main configuration for the ShapePipe Snakemake workflow. -# To run, first set a profile (machine/architecture): +# To run, first set the profile env variable to the desired machine/architecture: # SP_PROFILE=nibi|candide -# Your own run config, passed with -c, is merged on top of this file; then the -# `machines:` entry for (machine, input_type) fills anything still unset: -# -# SP_PROFILE= workflow/bin/sp run -c /path/to/my_run.yaml -# -# It therefore states only what differs from the defaults below. +# Then, run in /path/to/shapepipe: + +# ./workflow/bin/sp run -c /path/to/my_run.yaml # -# Everything is read at parse time; editing it (e.g. appending tiles) never -# invalidates completed work. +# The config file my_run.yaml can contain user- or run-specific entries +# that overwrite the main config file workflow/config.yaml . -# data (survey, config/cfis) or image_sims (SKiLLS, config/cfis_image_sims). +# Main keys, global for all machines + +# input_tyle: allowed are data or image_sims input_type: data -# psfex, mccd, or fake (image_sims only: true PSF from `psf_dict`). -psf_model: psfex +# psf_model (psfex, mccd, or fake) and psf_dict are set per (machine, +# input_type) in the machines: table below, NOT here: a top-level key +# shadows that table. The Snakefile falls back to psfex if no entry +# supplies one. # Run name, available as `$run` in the paths below. run: smk-g6 @@ -33,11 +34,13 @@ run: smk-g6 # - outputs.run_dir: (scratch) run directory, where tmp files will be stored # - outputs.products_dir: path to final products # - outputs.index_db: path to bookkeeping index sqlite file +# - container: TBD machines: nibi: base_dir: /project/def-mjhudson data: + psf_model: psfex tile_list: $base_dir/cdaley/sp-products/$run/tiles.txt retrieve: symlink inputs: @@ -52,22 +55,19 @@ machines: candide: base_dir: /n17data/UNIONS/WL data: + psf_model: psfex retrieve: vos - tile_list: TBD - inputs: - tiles: TBD # vos: URL - exposures: TBD - outputs: - run_dir: TBD - index_db: TBD container: /n17data/cdaley/containers/shapepipe_develop-runtime-20260718.sif image_sims: retrieve: symlink - tile_list: TBD + # True simulation PSF: no exposure PSF fit; fake_interp_runner + # reads psf_dict below. + psf_model: fake inputs: - # Need to be pointed to appropriate grid - tiles: /n09data/hervas/skills_out/1z2z_grid_1/images/SP_tiles - exposures: /n09data/hervas/skills_out/1z2z_grid_1/images/SP_exp + # $run is the branch dir (e.g. 1p2z_grid_1), so a run config + # picks the grid by setting `run:` alone. + tiles: /n09data/hervas/skills_out/$run/images/SP_tiles + exposures: /n09data/hervas/skills_out/$run/images/SP_exp # Dictionary of PSF model stamps psf_dict: /home/hervas/fhervas/workdir_skills/input/psf_files/Full_psf_dict.pickle container: /n17data/cdaley/containers/shapepipe_im_sims-runtime.sif diff --git a/workflow/scripts/run_config.py b/workflow/scripts/run_config.py index c6ed4f3d9..2925bfca5 100644 --- a/workflow/scripts/run_config.py +++ b/workflow/scripts/run_config.py @@ -14,7 +14,15 @@ import yaml PLACEHOLDER = "TBD" -MACHINE_KEYS = ("tile_list", "retrieve", "container", "inputs", "outputs") +# Keys the machines: table may default, per (machine, input_type). +# psf_model/psf_dict belong here because they are per-input_type facts, +# not per-run ones: psf_model=fake is only legal with +# input_type=image_sims, and psf_dict is the sim PSF it reads. A key +# also present at the TOP level of config.yaml shadows the table (the +# setdefault below only fires when the key is absent), so a key listed +# here must not carry a top-level default as well. +MACHINE_KEYS = ("tile_list", "retrieve", "container", "inputs", "outputs", + "psf_model", "psf_dict") REQUIRED = ("tile_list", "inputs.tiles", "inputs.exposures", "outputs.run_dir", "outputs.index_db") From 0b5fae7cdcf1ef32d48a9b24ad80a1840a722797 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 01:06:52 +0200 Subject: [PATCH 44/85] bin/sp: don't swallow -c/--config meant for the delegated command The argv scan for sp's own -c/--config-file took every -c/--config/ --config-file anywhere in argv as the run config, including ones that belong to the command sp delegates to. This broke the README's own `sp container exec python -c ...` (container.py already reads SP_RUN_CONFIG, so its own -c is never sp's run config) and `sp run --config clean=false` (the override flag() in the Snakefile is written for). Stop the scan at `--` and at the command after `container exec`, and drop the `--config` alias so snakemake's own `--config k=v` passes through untouched. Co-Authored-By: Claude Opus 5.5 --- workflow/bin/sp | 26 ++++++++++++++++++++------ 1 file changed, 20 insertions(+), 6 deletions(-) diff --git a/workflow/bin/sp b/workflow/bin/sp index c4dc50995..7d10b988d 100755 --- a/workflow/bin/sp +++ b/workflow/bin/sp @@ -36,9 +36,9 @@ # SP_PROFILE=candide workflow/bin/sp report -c ~/my_run.yaml # status now # workflow/bin/sp run # config.yaml only # -# Run config: -c FILE, also --config-file or --config. Read after -# workflow/config.yaml and merged on top; anything still unset comes from the -# machines: entry for SP_PROFILE and input_type (scripts/run_config.py). +# Run config: -c FILE, also --config-file. Read after workflow/config.yaml and +# merged on top; anything still unset comes from the machines: entry for +# SP_PROFILE and input_type (scripts/run_config.py). # To check a resolved value without running: # # python scripts/run_config.py workflow/config.yaml ~/my_run.yaml outputs.run_dir @@ -63,16 +63,30 @@ SCRIPTS="$HERE/scripts" PROFILE="${SP_PROFILE:-nibi}" # -c is consumed here, not forwarded: snakemake's own --configfile merges -# differently, and the jobs and the snapshot need it as SP_RUN_CONFIG. +# differently, and the jobs and the snapshot need it as SP_RUN_CONFIG. The scan +# stops at `--` and at the command after `container exec`, so a -c meant for +# the delegated command (`sp container exec python -c ...`) or for snakemake's +# own escape hatch (`sp -- ...`) is left alone. `--config` is NOT sp's: keeping +# it free lets snakemake's own `--config k=v` (what the Snakefile's flag() +# reads) pass through untouched. RUN_CONFIG="${SP_RUN_CONFIG:-}" _sp_args=() while [ $# -gt 0 ]; do case "$1" in - -c|--config|--config-file) + --) + _sp_args+=("$@"); break ;; + -c|--config-file) [ $# -ge 2 ] || { echo "sp: $1 needs a file" >&2; exit 2; } RUN_CONFIG="$2"; shift 2 ;; - -c=*|--config=*|--config-file=*) + -c=*|--config-file=*) RUN_CONFIG="${1#*=}"; shift ;; + container) + # Everything from `exec`'s own command on is the delegated command's, + # not sp's: `container.py` reads SP_RUN_CONFIG, so its own -c is never + # sp's run config. + _sp_args+=("$1"); shift + [ "${1:-}" = exec ] && { _sp_args+=("$@"); break; } + ;; *) _sp_args+=("$1"); shift ;; esac done From d22c4a0b8e295eb2b03145adde9e3cb524325e9a Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 01:06:57 +0200 Subject: [PATCH 45/85] workflow/README: add products_dir to the image-sims run-config example Without products_dir, merge_final_cats falls back to run_dir and create_final_cat.py finds no matching tile directory under it, so the rule writes an empty hdf5 and exits 0 -- a failure that only surfaces later, in sp_validation, as "No data found". Giving the example the template's /run + /product layout avoids that trap; hardening the rule itself is left to #879, which replaces merge_final_cats for sims. Co-Authored-By: Claude Opus 5.5 --- workflow/README.md | 1 + 1 file changed, 1 insertion(+) diff --git a/workflow/README.md b/workflow/README.md index 16795aa32..e3081ccca 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -93,6 +93,7 @@ inputs: exposures: /n09data/hervas/skills_out/1z2z_grid_3/images/SP_exp outputs: run_dir: /path/to/run + products_dir: /path/to/product index_db: /path/to/run/index.sqlite ``` From 5ded2ca88071e3370b6548ecd99138f22cfa3135 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 01:11:23 +0200 Subject: [PATCH 46/85] run_template.yaml: make it resolve Three problems kept this from resolving: run_config.py only expands $base_dir and $run, so the template's own $my_base_dir and $sim stayed literal and the Snakefile rejected the paths; with no input_type, the run took the data branch (cfis configs, psfex, VOS retrieval) instead of image_sims; and outputs.index_db, a required key, was missing. shapepipe_repo: is also dropped -- nothing reads it. Puts the sim name in run: and uses $run throughout, which run_config.py already expands; the remaining placeholders are for hand-editing, same as before. Co-Authored-By: Claude Opus 5.5 --- workflow/run_template.yaml | 18 +++++++++--------- 1 file changed, 9 insertions(+), 9 deletions(-) diff --git a/workflow/run_template.yaml b/workflow/run_template.yaml index c75bdb736..5516aa58f 100644 --- a/workflow/run_template.yaml +++ b/workflow/run_template.yaml @@ -1,18 +1,18 @@ # Run config template file, can be specified along with main config.yaml for # ShapePipe run. -my_base_dir: +input_type: image_sims -tile_list: $my_base_dir/tiles.txt +# Run name: the SKiLLS grid to use. Available as $run in the paths below. +run: 1z2z_grid_1 -sim: 1z2z_grid_1 +tile_list: /tiles.txt inputs: - tiles: /n09data/hervas/skills_out/$sim/images/SP_tiles - exposures: /n09data/hervas/skills_out/$sim/images/SP_exp + tiles: /n09data/hervas/skills_out/$run/images/SP_tiles + exposures: /n09data/hervas/skills_out/$run/images/SP_exp outputs: - run_dir: $my_base_dir/$sim/run - products_dir: $my_base_dir/$sim/product - -shapepipe_repo: + run_dir: /$run/run + products_dir: /$run/product + index_db: /$run/run/index.sqlite From dae9e1bc4ea95209a3f5fddd526489e026d31a66 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sun, 30 Aug 2026 20:25:38 -0400 Subject: [PATCH 47/85] docs(astra): record the pipeline's scientific decisions in astra.yaml MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ShapePipe's scientific choices — detection thresholds, masking geometry, star selection, PSF model, ngmix priors and seeding, flag semantics, completeness floors — live in code and committed configs with their reasoning nowhere, or spread across PRs, papers and comments. astra.yaml gathers them: 50 decisions across eight sub-analyses, each with its rationale, the alternatives that were rejected and why, and a greppable anchor back to the code or config that implements it. universes/committed.yaml pins the option this branch selects for every one. The record is ASTRA (astra-tools; `uvx astra-tools@0.2.17 guide`), applied here at codebase level rather than to a single analysis. Conventions are stated in the file's header: anchors as `path::symbol` / `path#SECTION.KEY` and never line numbers, [HARDCODED] for a scientific value with no config exposure, [LINT] for a place where the record and the code — or the code and itself — disagree, [PENDING #NNN] for state not yet on develop. Authoring it surfaced nine such lints, two of which #873 fixes, and mapped ten places where the published Guinot+22 / Farrens+22 descriptions have drifted from the code since publication; 16 decisions carry verbatim paper quotes as prior insights. CLAUDE.md gains the standing instruction: a scientific change is not finished until the record is, amended in the same PR. The membership test is whether a different defensible choice would change which objects enter the shear catalogue, or the numbers attached to them. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01Y2muA2sRojbxRNxU2SKQeP --- CLAUDE.md | 42 +- astra.yaml | 1837 ++++++++++++++++++++++++++++++++++++++ universes/committed.yaml | 70 ++ 3 files changed, 1948 insertions(+), 1 deletion(-) create mode 100644 astra.yaml create mode 100644 universes/committed.yaml diff --git a/CLAUDE.md b/CLAUDE.md index 7cfeda96f..53263db00 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -116,4 +116,44 @@ keep in their own stores outside it. A `.felt/` directory (a markdown "fiber" no store used with the `felt` CLI) is **not tracked here**: it's gitignored, and where it exists it's a machine-local symlink into a private, separately git-synced store, so a fresh clone won't have one. Record durable decisions in the PR, issue, or docs -where the change lives. +where the change lives — and *scientific* decisions in `astra.yaml`, below. + +## Scientific decisions live in `astra.yaml` + +`astra.yaml` at the repo root is the pipeline's decision record: every +consequential scientific choice embedded in the code and the committed configs, +each with its rationale, the alternatives that were considered and why they were +rejected, and an anchor back to the code or config that implements it. +`universes/committed.yaml` pins the option this branch's configuration +selects for every decision. The format +is ASTRA; `uvx astra-tools@0.2.17 guide` is the briefing and +`uvx astra-tools@0.2.17 spec` the field reference. + +**A scientific change is not finished until the record is.** When a change moves +what the pipeline measures, amend `astra.yaml` in the same PR — add the decision +if it is new, or edit its rationale, options and anchors if it moved — pin the +selected option in `universes/committed.yaml`, and say so in the PR description. +Purely technical changes (refactors, performance, packaging, I/O) leave it alone, +except where they move a value the record carries: the completeness floors in +`workflow/scripts/completeness.py` are orchestration code holding a scientific +decision. + +The membership test is whether *a different defensible choice would change which +objects enter the shear catalogue, or the numbers attached to them.* Detection +threshold and deblending contrast, masking geometry, star-selection cuts, PSF +model degree, ngmix priors and seeding, flag semantics, completeness floors — in. +Manifest sentinels, chunk sizes, allocation strategy, directory layout — out; +those live in the PR and the PRD. + +The file's own header states the conventions it follows. In short: every +rationale ends with a greppable `Anchor: path::symbol; path#SECTION.KEY` +sentence whose refs never cite line numbers; `[HARDCODED]` marks a scientific value +with no config exposure; `[LINT]` marks a place where the record and the code, or +the code and itself, disagree. Validate before committing: + +```bash +uvx astra-tools@0.2.17 validate +``` + +The record was authored against this branch's workflow configs; entries marked +`[PENDING #NNN]` describe state that has not yet reached `develop`. diff --git a/astra.yaml b/astra.yaml new file mode 100644 index 000000000..70b25b554 --- /dev/null +++ b/astra.yaml @@ -0,0 +1,1837 @@ +# ASTRA record for ShapePipe: the scientific decisions embedded in the code and +# the committed configs, with their reasoning and the alternatives that were +# rejected. It is the place scientific decisions are written down — see the +# "Scientific decisions" section of CLAUDE.md for when and how to amend it. +# +# The record describes the pipeline as orchestrated by workflow/Snakefile +# (PRD CosmoStat/shapepipe#848, PR #852). Conventions: +# +# * Decisions anchor to code, not recipes. Every rationale ends with one +# sentence "Anchor: ; ; ..." in a strict, greppable grammar. +# Each ref is a path relative to the shapepipe repo root, in one of three +# forms: CODE `path::symbol`, CONFIG `path#SECTION.KEY` (or `path#KEY` for +# sectionless .sex/.psfex/.ww/.param files), FILE `path` for a whole file +# or package. No line numbers — they rot; a line-level fact names its +# enclosing symbol. The analysis-ASTRA rule "never hardcode; reference via +# {decisions.x}" cannot hold in a codebase — the committed configs ARE the +# values. [HARDCODED] marks a scientific value living in code with no +# config exposure: the silent defaults the record exists to surface. +# * The default universe IS the committed configuration (universes/committed). +# Alternatives are excluded-with-reasons or genuinely open forks. +# * Sub-analyses follow the pipeline's methodological units — masking, +# detection, preparation, star selection + PSF, shape measurement, PSF +# diagnostics, survey geometry, catalogue assembly — not its ~20 Snakemake +# rules. Cross-cutting decisions stay top-level. A prior_insight repeated +# inside a sub-analysis carries a `_local` suffix: ids are scoped, and the +# duplicate keeps the sub-analysis readable on its own. +# * Outputs are representative product FAMILIES (one final_cat per tile), +# not enumerable artifacts; no recipes — the executor is the Snakemake +# workflow. +# * [LINT] marks places where this record and the code already disagree, or +# where the code disagrees with itself — found while authoring this file. +# * [PENDING #NNN] marks state that is live on feat/snakemake-orchestration +# — and therefore in smk-g4, the 34-tile validation campaign run under +# this branch — but not yet merged to develop. The record follows the +# branch and names the open PR. +# * A `path#KEY` anchor names the key's position in the file, not its +# activation: where the decision is "this is deliberately off", the key +# it points at may be commented out (e.g. final_cat.param#SPREAD_CLASS). + +version: "0.0.14" +name: ShapePipe scientific decisions +description: >- + Codebase-level decision record for the ShapePipe weak-lensing pipeline + (UNIONS/CFIS). Membership test: "a different defensible choice would change + which objects enter the shear catalogue, or the numbers attached to them." + Workflow mechanics that reproduce identical numbers (manifest sentinels, + clean-cascade cut, directory() outputs, allocation strategy, chunking under + position seeding) are deliberately absent; they live in the PRD and code. +tags: [shapepipe, weak-lensing, unions, codebase-record] +container: shapepipe-develop-runtime.sif + +inputs: + - id: tile_images + type: data + source: CADC-staged CFIS/UNIONS r-band tile stacks + exposure triplets (workflow/config.yaml) + description: >- + Pre-staged P3 tiles and single-exposure image/weight/flag triplets on + /project; get_images runs with RETRIEVE=symlink against this store. + - id: gsc_star_catalogue + type: data + source: GSC 2.3 (Vizier I/305/out) cone queries — scripts/python/create_star_cat.py + description: >- + Reference star catalogue driving bright-star masking. Catalogue choice, + query geometry, and magnitude handling are decisions in the masking + sub-analysis. + +outputs: + - id: final_cat + type: data + format: fits + description: >- + Per-tile shear catalogue family, the terminal science product (one per + campaign tile; make_cat_runner). Column selection and failure sentinels + are decisions in catalogue_assembly. + inputs: [tile_images] + decisions: [per_unit_count_floor, postage_stamp_size, photometric_zeropoint] + +decisions: + + # ── cross-cutting ──────────────────────────────────────────────────────── + + per_unit_count_floor: + label: Per-unit completeness policy under partial failure + rationale: >- + A 40-CCD stage where some CCDs legitimately produce nothing (sparse CCD, + setools rejects everything) cannot be all-or-nothing. The field's + converged answer (DES PSF blacklist, Rubin quantum registry) is per-unit + outcome records gated on a quality floor: record the attrition, fail + loud only below the floor, continue the survey. The floor VALUES are the + scientific content — how much silent per-CCD attrition can enter the + catalogue. The COMPLETENESS table holds them (exp_split + expect=121/floor=41, exp_mask expect=40/floor=1, psfex expect=80/floor=2, + psfex_interp floor=0 warn-only). Related leak the floor does not cover: + merge_sep_cats warns-and-skips a missing ngmix chunk, silently shrinking + a tile's shape catalogue below the floor's radar; and make_cat's own 10% + size-shortfall guard is commented out (see + catalogue_assembly.shape_catalogue_shortfall_guard). + Anchor: workflow/scripts/completeness.py::COMPLETENESS; + src/shapepipe/modules/merge_sep_cats_package/merge_sep_cats.py::MergeSep.process. + default: count_floor + options: + count_floor: + label: Count-floor table (expect/floor per runner; fail below floor) + insights: [des_psf_blacklist, guinot22_star_floor_22] + all_or_nothing: + label: Every expected sub-product required + excluded: true + excluded_reason: >- + Legitimately-absent CCDs would fail whole exposures and poison their + downstream cone; Snakemake has no optional-output primitive; field + precedent is tolerated, recorded attrition. + no_floor: + label: Accept whatever is produced, no gate + excluded: true + excluded_reason: >- + Silent attrition — a stage producing 2 of 40 CCDs would flow into + the catalogue unremarked. + + postage_stamp_size: + label: Postage-stamp size, 51 px everywhere + rationale: >- + One number pins three coupled apertures: the SExtractor vignet cut + around each detection (VIGNET(51,51) in default_noimaflags.param / + default.param, VIGNET_SIZE=51 in the dormant external-catalogue path, + example/cfis/config_tile_Uc.ini), the vignetmaker + stamps that feed ngmix (STAMP_SIZE=51 in config_tile_PiViVi.ini, both + runs; nearest-pixel centring, no sub-pixel interpolation in + VignetMaker._get_stamp), and the PSFEx model stamp (PSF_SIZE 51,51 in + default.psfex). The stamp IS the pixel data ngmix fits: it bounds + measurable galaxy size and truncates the wings of large galaxies. + Rationale for 51 not recorded in code. + Anchor: workflow/config/cfis/default_noimaflags.param#VIGNET; + example/cfis/config_tile_Uc.ini#READ_EXT_SEXCAT_RUNNER.VIGNET_SIZE; + workflow/config/cfis/config_tile_PiViVi.ini#VIGNETMAKER_RUNNER_RUN_1.STAMP_SIZE; + workflow/config/cfis/default.psfex#PSF_SIZE; + src/shapepipe/modules/vignetmaker_package/vignetmaker.py::VignetMaker._get_stamp. + default: px_51 + options: + px_51: + label: 51x51 px (~9.5 arcsec at 0.187"/px) + larger_adaptive: + label: Larger or size-adaptive stamps + excluded: true + excluded_reason: >- + Not wired; would need coupled changes in three places (a change in + any one alone desynchronises galaxy stamp, PSF stamp, and vignet). + + photometric_zeropoint: + label: Magnitude zero-point convention, fixed 30.0 on tiles + rationale: >- + Tiles use a hard-coded MAG_ZEROPOINT 30.0 for every tile + (default_tile.sex; ZP_FROM_HEADER=False in config_tile_Sx.ini), and + ngmix repeats it (MAG_ZP=30.0 in config_tile_Ng_template.ini). + Exposures instead read the per-image header zero-point + (ZP_FROM_HEADER=True, ZP_KEY=PHOTZP in config_exp_psfex.ini). The tile + convention leans on MegaPipe's calibrated stacks; the star-selection + magnitude window (18-22) and mask magnitude limits inherit whichever + convention their stage uses. SExtractorCaller.get_zero_point is the + header-reading path, unused on tiles. + Anchor: workflow/config/cfis/default_tile.sex#MAG_ZEROPOINT; + workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.ZP_FROM_HEADER; + workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.MAG_ZP; + workflow/config/cfis/config_exp_psfex.ini#SEXTRACTOR_RUNNER.ZP_KEY; + src/shapepipe/modules/sextractor_package/sextractor_script.py::SExtractorCaller.get_zero_point. + default: fixed_30_tiles_header_exposures + options: + fixed_30_tiles_header_exposures: + label: Tiles fixed 30.0; exposures from header PHOTZP + header_everywhere: + label: Per-image header zero-points on tiles too + excluded: true + excluded_reason: >- + MegaPipe stacks are calibrated to ZP 30 by construction; per-tile + header reads add a failure path for no expected numerical change. + (If that claim is wrong, this is a real fork — verify.) + + baseline_validation_criterion: + label: Validation criterion against the v2.0 bash baseline + rationale: >- + Because shape_measurement.ngmix_seed_mode deliberately changes noise + streams, P1 validation against v2.0 is statistical parity + (population-level agreement), not bit parity. Everything upstream of + ngmix (through PSFEx) validated bit-exactly (P0: 4/4 PASS). This + defines the evidence standard for "the same pipeline" — surfaced to the + collaboration as open Q5 in PRD #848. + [PENDING #873] Run-to-run determinism, which is a different property + from parity with v2.0, is now complete. With the setools star split + seeded (star_selection_psf.psf_train_validation_split) the last unseeded + draw in the science chain is gone: two runs of this code over the same + inputs now produce the same PSF star sample, the same PSF models and + the same shapes, which they did not before. That also settles a tension + this record carried — the bit-parity claim above sat next to an + unseeded star split that could not have been bit-reproducible, and the + P0 exposure-stage comparison did see PSF-validation CCD attrition + differ between the two sides. Statistical rather than bit parity is + therefore demanded only against the v2.0 baseline, not between runs of + the current pipeline. + Anchor: workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.SEED_FROM_POSITION; + src/shapepipe/modules/ngmix_package/ngmix.py::position_seed; + src/shapepipe/modules/setools_package/setools.py::SETools._make_rand_split. + default: statistical_parity + options: + statistical_parity: + label: Population-level agreement in shear observables + bit_parity: + label: Bit-identical catalogues + excluded: true + excluded_reason: >- + Impossible by construction once the seed mode changed; requiring it + would freeze the chunk-dependent v2.0 RNG forever. + +prior_insights: + des_psf_blacklist: + claim: >- + DES enters a CCD's PSF model into a blacklist rather than failing the + exposure - in Y3, any CCD with fewer than 25 stars surviving outlier + rejection is blacklisted and excluded downstream (~2% of data removed), + and processing proceeds. + created_at: "2026-07-16T00:00:00Z" + evidence: + - id: ev_jarvis_y3 + doi: "10.48550/arXiv.2011.03409" + quote: + exact: "we enter it into a" + suffix: " \u201cblacklist\u201d and exclude this CCD" + location: { page: 10 } + guinot22_star_floor_22: + claim: >- + The published ShapePipe/UNIONS analysis applies a per-CCD quality floor + rather than failing whole exposures: a CCD with fewer than 22 selected + stars is discarded for PSF estimation and contributes no epoch to the + shape measurement, while processing continues. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_star_floor + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'The dashed line represents the cut at 22 stars/CCD below which the CCD is discarded for the PSF estimation.' + location: { page: 4 } + +findings: + orchestration_parity: + claim: >- + The Snakemake orchestration reproduces the bash baseline bit-exactly + through PSFEx (P0 validation, 4/4 PASS on the 186/187 quad). + Read it with two caveats. It is a statement about the pre-#873 code: + both #873 changes move products (a different realised star split, a + stricter science-path star gate), so re-establishing parity would mean + regenerating the baseline under the current branch. And the parity is + bit-exact in the products compared, not everywhere: the P0 + exposure-stage comparison did see PSF-validation CCD attrition differ + between the two sides, which the then-unseeded star split explains + (see baseline_validation_criterion). + created_at: "2026-08-19T00:00:00Z" + evidence: + - id: ev_final_cat + artifact: final_cat + record_authoring_found_defects: + claim: >- + Nine places in the code disagree with themselves or with their + documentation, each carried as a [LINT] mark: the 22-vs-20 + STAR_THRESH mismatch between PSF validation and science interpolation; + the unseeded train/validation rand_split (setools.py:664 — the star + sample entering the PSF model is irreproducible run-to-run); additive + (non-bitwise) mask-plane combination, safe today only because the + committed flag values are disjoint; the dead MESSIER_PIXEL_SCALE config + key; final_cat.param requesting IMAFLAGS_ISO that the merged catalogue + never receives; the centroid_source default disagreement (runner "wcs" + vs module "hsm", latent for direct callers); setools logging a FWHM cut + (mode +- 0.1 px in arcsec) half the applied one (mode +- 0.2 px), and + mixing pixel scales 0.187/0.186 within one file; TILE_LIST + overlap-flagging documented but never implemented; the mccd_plots + module docstring advertising rho statistics that live downstream now. + Status: the first two are fixed. CosmoStat/shapepipe#873 seeds the + rand_split and raises the science-path STAR_THRESH to 22, and commit + 90782098 mirrors that threshold into the workflow's own committed + config fork. #873 is OPEN against develop; both fixes are live on + feat/snakemake-orchestration only, and the 34-tile smk-g4 campaign is + the first run under them. The other seven stand, including the + mccd_plots docstring that still advertises rho statistics the package + no longer computes. + created_at: "2026-08-29T00:00:00Z" + derived: true + evidence: + - id: ev_final_cat_defects + artifact: final_cat + code_paper_divergence: + claim: >- + The two ShapePipe papers state roughly 17 of this record's 50 decisions + (now carried as prior_insights with verbatim quotes), have drifted from + the code on 10 of them since publication, and are silent on the rest. + Among them: DETECT_MINAREA 10 -> 5; DEBLEND_MINCONT 0.001 -> + 0.0005 on tiles; tile background AUTO -> MANUAL 0; in-line spread-model + star/galaxy classification -> disabled and deferred downstream; HSM + moment initialisation -> WCS centroids and prior-based guesses; GSC 2.2 + via cdsclient -> GSC 2.3 via astroquery; PSF acceptance 22 stars/CCD + published for the science path vs 20 committed there — closed since by + #873 + 90782098, which put the science path on 22, so nine of the ten + drifts remain open on the orchestration branch. + created_at: "2026-08-29T00:00:00Z" + derived: true + evidence: + - id: ev_final_cat_divergence + artifact: final_cat + +analyses: + + # ═════════════════════════════════════════════════════════════════════════ + masking: + description: >- + Which pixels are excluded before anything is measured. Modules: + src/shapepipe/modules/mask_package/mask.py (halo/spike/DSO/border + builders, WeightWatcher driver), scripts/python/create_star_cat.py + (star-catalogue fetch), configs config_exp_Ma.ini + + config_onthefly.mask / config_tile_onthefly.mask + mask_default/. + [LINT] MESSIER_PIXEL_SCALE is set in config_tile_onthefly.mask but + never read — mask_dso takes pixel scale from the WCS. [LINT] + _build_final_mask combines mask planes by ADDITION (mask.py:1141+), + not bitwise OR; the committed flag values (2/4/16/32/128) are disjoint + so no live collision exists, but any future duplicate value corrupts + the flag semantics silently. (FLAG_OUTFLAGS 2 in default.ww is inert: + no input flag image is passed to WeightWatcher — mask.py:1047-1074.) + inputs: + - id: ccd_images + type: data + source: split per-CCD exposure images + weights + CFIS flag maps (exp_split family) + - id: star_catalogue + type: data + source: GSC 2.3 per-exposure catalogues (exp_star_cat cache) + outputs: + - id: exposure_mask + type: data + format: fits + description: Per-CCD pipeline flag maps (run_sp_exp_Ma family). + decisions: + [star_catalogue_query, star_magnitude_definition, + bright_star_mask_geometry, deep_sky_object_masking, + border_mask_width, pixel_threshold_flags, external_flag_usage] + decisions: + star_catalogue_query: + label: Reference star catalogue and query for bright-star masking + rationale: >- + GSC 2.3 (Vizier I/305/out), columns GSC2.3/RAJ2000/DEJ2000/Fmag/ + jmag/Vmag/Nmag/Class, cone radius covering the full CCD mosaic, no + magnitude cut at query time. GSC 2.2 rejected in a code comment + beside the catalogue ID ("does not have Fmag"). + Provenance hazard: the star-cat cache is not keyed by script + version — a semantic change to this query reruns the rule but takes + the skip-if-exists branch; clear the cache by hand for the change + to reach the data (workflow/config.yaml star_cats comment). + Query geometry: search radius = half the image diagonal about the + field centre (Mask._get_image_radius); source precedence: with + CDSCLIENT_PATH set in the .mask configs the online-query branch + wins unless an external star cat is passed (USE_EXT_STAR=True in + config_exp_Ma.ini routes the exp_star_cat cache in). + Published description (Farrens+22 p.2): cdsclient downloads GSC 2.2 + (with cdsclient 3.84 pinned in its Table A.1); current code: GSC 2.3 + queried through astroquery — two things drifted, the catalogue + version (Fmag is needed for the magnitude cut) and the query + transport, since cdsclient is never invoked yet survives as a + required-but-unused CDSCLIENT_PATH still set to the stale + findgsc2.2 in config_tile_onthefly.mask. + Anchor: scripts/python/create_star_cat.py::CDS_CAT_ID; + src/shapepipe/modules/mask_package/mask.py::Mask._CDS_cat_ID; + src/shapepipe/modules/mask_package/mask.py::Mask._cds_keys; + workflow/config.yaml. + default: gsc_23_vizier + options: + gsc_23_vizier: + label: GSC 2.3 cone queries, all bands, no query-time mag cut + insights: [farrens22_star_cat_on_disk] + gaia: + label: Gaia-based star catalogue + excluded: true + excluded_reason: >- + Not wired. Deeper and better photometry; switching changes mask + geometry and hence the selection function — a real DR-level fork. + star_magnitude_definition: + label: Per-star magnitude for mask scaling + rationale: >- + mag = unweighted mean of the finite GSC bands among F, j, V, N; + stars with no finite band are logged and not masked; only Class==0 + objects masked. Comment records why not a naive mean: NaN bands + would NaN-poison the mag < mag_limit test, leaving exactly the + bright stars with incomplete photometry unmasked. + Anchor: src/shapepipe/modules/mask_package/mask.py::Mask._create_mask. + default: mean_finite_bands + options: + mean_finite_bands: { label: Mean of finite F/j/V/N; Class==0 only } + single_band: + label: Single-band (Fmag) magnitude + excluded: true + excluded_reason: Drops stars with missing Fmag from masking entirely. + bright_star_mask_geometry: + label: Halo + diffraction-spike mask geometry and magnitude scaling + rationale: >- + DS9 polygon templates scaled linearly with magnitude about a pivot: + halo HALO_MAG_LIM=13, HALO_SCALE_FACTOR=0.05, HALO_MAG_PIVOT=13.8 + (halo_mask.reg, ~270 px); spike SPIKE_MAG_LIM=18, + SPIKE_SCALE_FACTOR=0.3, SPIKE_MAG_PIVOT=13.8 + (MEGAPRIME_star_i_13.8.reg); scaling = 1 - factor*(mag-pivot), + floored at 0.1 by Mask._scaling_min. + Identical in exposure and tile configs. Template filename encodes + provenance (MegaPrime i-band mag-13.8 star); numeric rationale not + recorded. Note the 5-mag gap: stars in 13-18 get spikes but no halo. + Anchor: workflow/config/cfis/config_onthefly.mask#HALO_PARAMETERS.HALO_MAG_LIM; + workflow/config/cfis/config_onthefly.mask#SPIKE_PARAMETERS.SPIKE_MAG_LIM; + workflow/config/cfis/mask_default/halo_mask.reg; + workflow/config/cfis/mask_default/MEGAPRIME_star_i_13.8.reg; + src/shapepipe/modules/mask_package/mask.py::Mask._create_mask; + src/shapepipe/modules/mask_package/mask.py::Mask._scaling_min. + default: megaprime_polygon_linear_scaling + options: + megaprime_polygon_linear_scaling: + label: Fixed MegaPrime templates, linear mag scaling, floor 0.1 + radial_profile_fit: + label: Per-star radial-profile-driven mask size + excluded: true + excluded_reason: Not wired; the survey precedent is template-based. + deep_sky_object_masking: + label: Messier + NGC objects masked as circles, no enlargement + rationale: >- + Circles of radius max(size_X, size_Y), MESSIER_SIZE_PLUS=0, + NGC_SIZE_PLUS=0 (function default is 0.1 — the 0 is a choice); + flags 16/32. A comment records the overlap-test fix (corner-only + test missed small interior objects). + Anchor: workflow/config/cfis/config_onthefly.mask#MESSIER_PARAMETERS.MESSIER_SIZE_PLUS; + workflow/config/cfis/config_tile_onthefly.mask#NGC_PARAMETERS.NGC_SIZE_PLUS; + src/shapepipe/modules/mask_package/mask.py::Mask.mask_dso. + default: circles_no_padding + options: + circles_no_padding: + label: "size_plus = 0: mask exactly the catalogued extent" + insights: [farrens22_messier_mask] + padded_circles: + label: size_plus > 0 (code default 0.1) + excluded: true + excluded_reason: >- + Rationale for dropping the padding not recorded; flagged as a + question rather than an endorsed exclusion. + border_mask_width: + label: CCD border mask, 50 px on exposures, none on tiles + rationale: >- + Exposures BORDER_WIDTH=50 (flag 4); tiles BORDER_MAKE=False. + Mask.mask_border's own default is 100 — the committed 50 is a + choice, unrecorded. Trims CCD edges where PSF and astrometry + degrade; changes the effective footprint. + Anchor: workflow/config/cfis/config_onthefly.mask#BORDER_PARAMETERS.BORDER_WIDTH; + workflow/config/cfis/config_tile_onthefly.mask#BORDER_PARAMETERS.BORDER_MAKE; + src/shapepipe/modules/mask_package/mask.py::Mask.mask_border. + default: px50_exposures_only + options: + px50_exposures_only: + label: 50 px exposure borders; tiles unmasked + insights: [farrens22_border_mask] + px100: + label: 100 px (module default) + excluded: true + excluded_reason: Halves usable edge area for no recorded gain. + pixel_threshold_flags: + label: WeightWatcher weight/flag thresholds into mask bits + rationale: >- + WEIGHT_MIN 0, WEIGHT_MAX 1000, WEIGHT_OUTFLAGS 1; FLAG_MASKS 0x01, + FLAG_OUTFLAGS 2; POLY_OUTWEIGHTS 0. Zero-weight and externally + flagged pixels excluded on these thresholds. Values are stock, not + derived from the CFIS weight distribution; rationale not recorded. + The FLAG_* keys are inert in the committed invocation — no flag + image is passed to WeightWatcher by Mask._exec_WW. + Anchor: workflow/config/cfis/mask_default/default.ww#WEIGHT_MIN; + workflow/config/cfis/mask_default/default.ww#FLAG_MASKS; + src/shapepipe/modules/mask_package/mask.py::Mask._exec_WW. + default: stock_ww_thresholds + options: + stock_ww_thresholds: + label: Stock WeightWatcher thresholds + insights: [farrens22_weightwatcher] + external_flag_usage: + label: CFIS external flag maps folded into exposure masks + rationale: >- + USE_EXT_FLAG=True on exposures (imports CADC-provided bad-pixel / + cosmic-ray / trail flags); EF_MAKE=False on tiles. The external + plane enters via Mask._build_final_mask's path_external_flag branch. + Anchor: workflow/config/cfis/config_exp_Ma.ini#MASK_RUNNER.USE_EXT_FLAG; + workflow/config/cfis/config_tile_onthefly.mask#EXTERNAL_FLAG.EF_MAKE; + src/shapepipe/modules/mask_package/mask.py::Mask._build_final_mask. + default: exposures_only + options: + exposures_only: { label: "External flags on exposures, not tiles" } + ignore_external: + label: Pipeline-generated masks only + excluded: true + excluded_reason: Discards upstream knowledge of bad pixels. + prior_insights: + farrens22_star_cat_on_disk: + claim: >- + The ShapePipe release paper documents an on-disk star catalogue, in + GSC format, as a supported substitute for the online query, + motivated by compute nodes without internet access. + created_at: "2022-06-01T00:00:00Z" + evidence: + - id: ev_farrens22_star_cat_disk + doi: "10.48550/arXiv.2206.14689" + quote: + exact: 'Alternatively, a star catalogue available on disk (with the same format as the GSC) can also be used' + location: { page: 2 } + farrens22_messier_mask: + claim: >- + Messier objects are named in the published masking procedure as one + of the object classes ShapePipe masks. + created_at: "2022-06-01T00:00:00Z" + evidence: + - id: ev_farrens22_messier + doi: "10.48550/arXiv.2206.14689" + quote: + exact: 'Messier objects, and border regions.' + location: { page: 2 } + farrens22_border_mask: + claim: >- + CCD border regions are named in the published masking procedure as + one of the regions ShapePipe masks. + created_at: "2022-06-01T00:00:00Z" + evidence: + - id: ev_farrens22_border + doi: "10.48550/arXiv.2206.14689" + quote: + exact: 'Messier objects, and border regions.' + location: { page: 2 } + farrens22_weightwatcher: + claim: >- + The published pipeline generates the mask image itself with + WeightWatcher (Marmo & Bertin 2008), fixing the tool but none of its + threshold values. + created_at: "2022-06-01T00:00:00Z" + evidence: + - id: ev_farrens22_ww + doi: "10.48550/arXiv.2206.14689" + location: { page: 2 } + + # ═════════════════════════════════════════════════════════════════════════ + detection: + description: >- + Object detection on r-band tiles (single-image mode) and exposures (for + star finding). Module: src/shapepipe/modules/sextractor_package/ + sextractor_script.py (config assembly, ZP/background overrides, + post-processing that assigns per-epoch CCD membership). Configs: + config_tile_Sx.ini + default_tile.sex + default.conv + + default_noimaflags.param (tiles); default_exp.sex (exposures — same + thresholds, but DEBLEND_MINCONT 0.001 vs tile 0.0005 and BACK_TYPE AUTO + vs tile MANUAL 0, both deliberate and unexplained divergences). + [LINT] final_cat.param (consumed by the post-proc merge_final_cat, + not by make_cat) requests IMAFLAGS_ISO, but the tile chain never + produces it (FLAG_IMAGE=False, default_noimaflags.param); the + exposure-side IMAFLAGS_ISO stays exposure-side (merge_starcat.py:807 + only). The merged catalogue never receives the column. + inputs: + - id: tile_stack + type: data + source: MegaPipe r-band tile stack + weight (uncompressed, merged headers) + outputs: + - id: tile_sexcat + type: data + format: fits + description: Per-tile SExtractor LDAC catalogue with per-epoch CCD membership. + decisions: + [detection_threshold_policy, deblending_policy, background_model, + weighting_and_interpolation, detection_source_mode, + epoch_membership_ccd_bounds, photometry_parameters, + cleaning_and_neighbour_masking] + decisions: + photometry_parameters: + label: Photometric aperture definitions — Kron parameters, apertures, half-light fraction + rationale: >- + PHOT_AUTOPARAMS 2.5,3.5 (Kron factor / minimum radius), + PHOT_APERTURES 5 px, PHOT_FLUXFRAC 0.5, BACKPHOTO_TYPE GLOBAL — + identical in both .sex files. MAG_AUTO is the axis of the + star-selection magnitude box AND the catalogue magnitude; FLUX_AUTO + is PSFEx's photometric normalisation (default.psfex PHOTFLUX_KEY). + A different Kron factor shifts magnitudes systematically, moving + which stars build the PSF model and every magnitude-based + downstream cut. Rationale not recorded (stock values). Anchor: + workflow/config/cfis/default_tile.sex#PHOT_AUTOPARAMS; + workflow/config/cfis/default_exp.sex#PHOT_AUTOPARAMS; + workflow/config/cfis/default.psfex#PHOTFLUX_KEY. + default: kron_25_35 + options: + kron_25_35: { label: "Kron 2.5/3.5, aperture 5 px, FLUXFRAC 0.5, global background" } + cleaning_and_neighbour_masking: + label: Spurious-detection cleaning and neighbour-pixel correction + rationale: >- + CLEAN Y with CLEAN_PARAM 1.0 deletes detections consistent with + being wings of a brighter neighbour — a post-deblend change to the + object list; MASK_TYPE CORRECT replaces neighbour pixels during + photometry (vs BLANK/NONE), changing fluxes and windowed moments + of blends. Identical in both .sex files; rationale not recorded. + Anchor: workflow/config/cfis/default_tile.sex#CLEAN; + workflow/config/cfis/default_tile.sex#MASK_TYPE; + workflow/config/cfis/default_exp.sex#CLEAN. + default: clean_1_correct + options: + clean_1_correct: { label: "CLEAN 1.0 + MASK_TYPE CORRECT" } + detection_threshold_policy: + label: Detection significance, minimum area, matched filter + rationale: >- + DETECT_THRESH 1.5 sigma RELATIVE, ANALYSIS_THRESH 1.5, + DETECT_MINAREA 5, FILTER default.conv (3x3 pyramid kernel, "all + ground, FWHM = 2 pixels" — vs CFIS seeing ~0.65 arcsec = 3.5 px at + 0.187"/px, so the filter is not matched to the survey PSF). + Sets the faint end of the source sample. Rationale not recorded + (stock EB 2017 header). + Published description (Guinot+22 p.5, Table 2): DETECT_MINAREA 10; + current code: 5, in default_exp.sex as well as default_tile.sex — + the small-object end has been loosened since publication on both the + star-detection and tile-detection passes, while DETECT_THRESH 1.5 + RELATIVE and the 3x3 FWHM=2 px kernel still match. + Anchor: workflow/config/cfis/default_tile.sex#DETECT_THRESH; + workflow/config/cfis/default_tile.sex#DETECT_MINAREA; + workflow/config/cfis/default.conv. + default: thresh_1p5_minarea5_fwhm2px_filter + options: + thresh_1p5_minarea5_fwhm2px_filter: + label: 1.5 sigma, minarea 5, FWHM=2px kernel + seeing_matched_filter: + label: Kernel matched to CFIS seeing (~3.5 px) + excluded: true + excluded_reason: >- + Not wired; would change depth and the faint-end selection + function — a real fork, excluded only as not-the-committed-path. + deblending_policy: + label: Deblending sub-thresholds and contrast + rationale: >- + DEBLEND_NTHRESH 32, DEBLEND_MINCONT 0.0005 on tiles (2x more + aggressive splitting than the exposure 0.001 and 10x more than the + SExtractor default 0.005 — divergences not documented), CLEAN Y + PARAM 1.0. Controls object count, centroids, and blend + contamination in shapes. + Published description (Guinot+22 p.5, Table 2, galaxy detection on + the stacked tiles): DEBLEND_MINCONT 0.001; current code: 0.0005 on + tiles, with only the exposure side still carrying 0.001 — the + divergence lands on precisely the configuration the paper documents. + NTHRESH 32 matches. + Anchor: workflow/config/cfis/default_tile.sex#DEBLEND_MINCONT; + workflow/config/cfis/default_exp.sex#DEBLEND_MINCONT. + default: mincont_5em4_tiles + options: + mincont_5em4_tiles: { label: "MINCONT 0.0005 tiles / 0.001 exposures" } + background_model: + label: Tile background fixed to zero, not estimated + rationale: >- + BACK_TYPE MANUAL, BACK_VALUE 0.0, BKG_FROM_HEADER=False on tiles — + trusts MegaPipe stack background removal; exposures use BACK_TYPE + AUTO (64/3 mesh). Residual sky offsets propagate into thresholds, + fluxes, completeness. Divergence deliberate, unexplained. + Published description (Guinot+22 p.5): Table 2's caption asserts all + non-tabulated SExtractor parameters keep their defaults, i.e. + BACK_TYPE AUTO, and the paper never mentions the background choice + at all; current code: BACK_TYPE MANUAL with BACK_VALUE 0.0 on tiles, + identically in workflow/ and example/ — the standing tile + configuration, not a one-off, diverging silently from the published + parametrisation. + Anchor: workflow/config/cfis/default_tile.sex#BACK_TYPE; + workflow/config/cfis/default_exp.sex#BACK_TYPE; + workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.BKG_FROM_HEADER; + src/shapepipe/modules/sextractor_package/sextractor_script.py::SExtractorCaller.get_background. + default: manual_zero_tiles_auto_exposures + options: + manual_zero_tiles_auto_exposures: + label: Tiles trust the stack (0.0); exposures estimate + auto_everywhere: + label: SExtractor AUTO background on tiles too + excluded: true + excluded_reason: >- + Double-subtracts if MegaPipe already removed it; if MegaPipe + residuals are nonzero this exclusion is wrong — verify. + weighting_and_interpolation: + label: Weight-map usage and zero-weight pixel interpolation + rationale: >- + Two settings depart from stock SExtractor: WEIGHT_TYPE MAP_WEIGHT + (default NONE) and INTERP_TYPE ALL (default NONE — SExtractor + invents flux across zero-weight pixels). The accompanying + RESCALE_WEIGHTS Y, WEIGHT_GAIN Y, MASK_TYPE CORRECT and + INTERP_MAXXLAG/INTERP_MAXYLAG 16 are the SExtractor defaults, so + they are settings the configs restate rather than choices. The + variance policy sets effective per-pixel SNR and thus the detection + set; INTERP_TYPE ALL alters pixel data feeding measurements. + Rationale not recorded. + Published description (Guinot+22 p.5): Table 2's "all other + parameters are kept to their default values" silently covers both + non-default settings; current code: MAP_WEIGHT + INTERP_TYPE ALL in + default_tile.sex and default_exp.sex alike — the paper gives no hint + that the weight map or the zero-weight interpolation is in play. + Anchor: workflow/config/cfis/default_tile.sex#WEIGHT_TYPE; + workflow/config/cfis/default_tile.sex#INTERP_TYPE; + src/shapepipe/modules/sextractor_package/sextractor_script.py::SExtractorCaller.set_input_files. + default: map_weight_interp_all + options: + map_weight_interp_all: { label: MAP_WEIGHT + INTERP ALL + MASK CORRECT } + no_interpolation: + label: INTERP_TYPE NONE + excluded: true + excluded_reason: >- + Changes photometry near masks; the committed choice is itself + unjustified in code — flagged as a question, not an endorsement. + detection_source_mode: + label: Single-image detection on the r-band tile + rationale: >- + DETECTION_IMAGE=False, FLAG_IMAGE=False at detection, + param file default_noimaflags.param. No dual-image mode, no + detection coadd, no flag propagation at detection time. The + sx_nomask variant is the committed chain because it matches the + validated bash baseline. STATUS: DELIBERATELY UNDECIDED (Cail, + 2026-08-29) — whether DR6 detects masked or unmasked is punted to + the planned masking-unification rework (not yet tracked in an issue); + the default records baseline + parity, not a settled methodological choice. The masked variant is + one config + one rule + a tile-side star-cat analogue away. + Anchor: workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.DETECTION_IMAGE; + workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.FLAG_IMAGE; + workflow/config/cfis/default_noimaflags.param; + workflow/rules/tile.smk. + default: sx_nomask_single_image + options: + sx_nomask_single_image: + label: Unmasked single-image r-band detection + insights: [guinot22_stacked_detection] + sx_masked: + label: Detection on the masked tile + dual_image_coadd: + label: Dual-image mode with a detection coadd + excluded: true + excluded_reason: No detection coadd exists in UNIONS r-band processing. + epoch_membership_ccd_bounds: + label: Which exposure CCDs an object belongs to (N_EPOCH) + rationale: >- + CCD_SIZE = 33,2080,1,4612 with strict inequalities — the 33-px left + trim silently discards a CCD strip from epoch membership; WCS + inversion failures skip the CCD ("no epoch recorded"), changing + N_EPOCH. Sets how many exposures contribute to each galaxy's + multi-epoch fit. Rationale beyond "number of pixels in a CCD" not + recorded. + Anchor: workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.CCD_SIZE; + src/shapepipe/modules/sextractor_package/sextractor_script.py::make_post_process; + src/shapepipe/modules/sextractor_package/sextractor_script.py::ccd_candidate_mask. + default: trimmed_bounds_33_2080 + options: + trimmed_bounds_33_2080: { label: "x in (33,2080), y in (1,4612), strict" } + full_ccd: + label: Full 1-2048 x-range, inclusive bounds + excluded: true + excluded_reason: >- + The trim presumably excludes a bad edge region, but nothing in + code says so — flagged as a question. + prior_insights: + guinot22_stacked_detection: + claim: >- + Source extraction in the published analysis is performed on the + stacked tile images, for signal-to-noise and because artefacts are + suppressed relative to single exposures. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_stacked_detection + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'We do the extraction on stacked images which provide a better signal-to-noise ratio, and most artifacts have a reduced amplitude with respect to single exposures' + location: { page: 5 } + + # ═════════════════════════════════════════════════════════════════════════ + preparation: + description: >- + How pixels, WCS, and epoch membership are prepared before anything is + measured. Modules: + src/shapepipe/modules/split_exp_package/split_exp.py, + merge_headers_package/merge_headers.py, + find_exposures_package/find_exposures.py, + vignetmaker_package/vignetmaker.py. + inputs: + - id: exposure_files + type: data + source: delivered CFIS exposure triplets (image/weight/flag MEF) + tile stacks + outputs: + - id: epoch_stamps + type: data + format: fits + description: Per-object multi-epoch vignets + per-CCD WCS log feeding ngmix. + decisions: + [astrometric_solution_source, ccd_split_extent, + epoch_provenance_from_tile_history, object_position_columns, + stamp_positioning_and_padding, epoch_flag_source] + decisions: + astrometric_solution_source: + label: Astrometry taken verbatim from delivered per-CCD headers + rationale: >- + split_exp builds WCS(h) from each raw CCD header at split time, + pickles it, and merge_headers writes the lot into + log_exp_headers.sqlite; every downstream world<->pixel transform + (stamp positioning, epoch membership, position seeding) uses that + stored solution. No re-derivation, no astrometric refinement — the + survey's delivered astrometry IS the pipeline's astrometry. + Alternative (a joint astrometric re-fit a la DES/Rubin) would move + every stamp centre and every position seed. Anchor: + src/shapepipe/modules/split_exp_package/split_exp.py::SplitExposures.create_hdus; + src/shapepipe/modules/merge_headers_package/merge_headers.py::merge_headers; + src/shapepipe/modules/vignetmaker_package/vignetmaker.py::VignetMaker._get_stamp_me. + default: delivered_headers + options: + delivered_headers: + label: "WCS(header) verbatim, stored at split time" + insights: [guinot22_gaia_astrometry] + astrometric_refit: + label: Joint astrometric re-solution + excluded: true + excluded_reason: Not wired; CFIS delivered astrometry is trusted. + ccd_split_extent: + label: All 40 MegaCam HDUs split and carried as candidate epochs + rationale: >- + N_HDU=40 with a hard check (any other HDU count raises) — every + CCD including the ear CCDs 36-39 is a candidate epoch wherever the + WCS lands it. The MegaCamFlip special-casing of 36/37 shows the + ears flow through shape measurement. Alternative: exclude ear CCDs + (different optical path/orientation history). Anchor: + workflow/config/cfis/config_exp_Sp.ini#SPLIT_EXP_RUNNER.N_HDU; + src/shapepipe/modules/split_exp_package/split_exp.py::SplitExposures.create_hdus. + default: all_40_hdus + options: + all_40_hdus: + label: "40 HDUs, hard-fail on any other count" + insights: [guinot22_forty_chips] + epoch_provenance_from_tile_history: + label: Epoch sets parsed from tile FITS HISTORY cards + rationale: >- + A tile's contributing exposures are recovered by parsing column 3 + of each HISTORY line, stripping prefix "p", deduplicating — the + coadd's own provenance record is trusted as the epoch list. The + LSB s-prefix rename is present but commented out. A mis-parse + changes N_EPOCH and which exposures are fit. Anchor: + workflow/config/cfis/config_tile_Fe.ini#FIND_EXPOSURES_RUNNER.COLNUM; + src/shapepipe/modules/find_exposures_package/find_exposures.py::FindExposures.get_exposure_list. + default: history_parse + options: + history_parse: { label: "HISTORY column 3, prefix p, dedup" } + object_position_columns: + label: Windowed centroids (XWIN/YWIN) define every position + rationale: >- + PSF interpolation sites, tile stamp centres, multi-epoch stamp + centres, and the catalogue sky position all use SExtractor's + windowed centroid — XWIN_WORLD/YWIN_WORLD on the tile side (SPHE), + XWIN_IMAGE/YWIN_IMAGE exposure-side (PIX). Windowed vs isophotal + vs model centroids differ systematically for blends and asymmetric + galaxies, and the centroid definition feeds the position seed. + Anchor: workflow/config/cfis/config_tile_PiViVi.ini#PSFEX_INTERP_RUNNER.POSITION_PARAMS; + workflow/config/cfis/config_tile_PiViVi.ini#VIGNETMAKER_RUNNER_RUN_2.POSITION_PARAMS; + workflow/config/cfis/config_exp_psfex.ini#POSITION_PARAMS. + default: xwin_windowed + options: + xwin_windowed: { label: Windowed centroids everywhere } + stamp_positioning_and_padding: + label: Nearest-pixel stamp centring; edge stamps zero-padded + rationale: >- + Multi-epoch stamps are placed by round-tripping the tile world + position through the stored per-CCD WCS, then rounding to the + nearest pixel (no sub-pixel interpolation — the residual sub-pixel + offset is absorbed by the fit's centroid prior, cen sigma = 1 + pixel). Objects whose stamp overruns a CCD or tile edge are KEPT, + out-of-image pixels zero-filled (sf_tools FetchStamps + pad_mode='constant'); no boundary rejection exists — zero-padded + pixels enter the fit as data with whatever weight the padded + weight stamp carries. Anchor: + src/shapepipe/modules/vignetmaker_package/vignetmaker.py::VignetMaker._get_stamp; + src/shapepipe/modules/vignetmaker_package/vignetmaker.py::VignetMaker._get_stamp_me. + default: round_and_zero_pad + options: + round_and_zero_pad: { label: "Nearest-pixel + zero padding, no edge rejection" } + epoch_flag_source: + label: Per-epoch flag stamps come from RAW CFIS flags, not the pipeline mask + rationale: >- + The multi-epoch vignet run reads its flag stamps from + split_exp_runner output — the delivered instrumental flags — + while mask_runner's pipeline_flag (halos, spikes, DSOs, borders) + feeds only the exposure-side star finding + (config_exp_psfex.ini FILE_PATTERN pipeline_flag). Combined with + unmasked tile detection (detection.detection_source_mode), the + consequence is stark: THE BRIGHT-STAR MASKS CURRENTLY AFFECT ONLY + PSF-STAR SELECTION — neither the galaxy sample (no tile mask, no + IMAFLAGS cut possible) nor the pixels ngmix fits (raw flags only) + see them. Whether that is intended belongs to the + planned masking-unification rework, whose object-level half is + sp_validation's IMAFLAGS_ISO cut on a column this chain never + produces; this is the pixel-level half. Anchor: + workflow/config/cfis/config_tile_PiViVi.ini#VIGNETMAKER_RUNNER_RUN_2.ME_IMAGE_EXP_RUNNERS; + workflow/config/cfis/config_exp_psfex.ini#SEXTRACTOR_RUNNER.FILE_PATTERN; + workflow/config/cfis/config_exp_Ma.ini#MASK_RUNNER.PREFIX. + default: raw_flags + options: + raw_flags: { label: split_exp raw flags gate epoch pixels } + pipeline_flags: + label: pipeline_flag (incl. bright-star masks) gates epoch pixels + prior_insights: + guinot22_gaia_astrometry: + claim: >- + The astrometric solution the analysis relies on is the upstream + MegaPipe/Gaia DR2 calibration, accurate to within 20 mas; no + astrometric re-fit inside the pipeline is described. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_astrometry + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'An astrometric calibration within 20 mas was achieved using the Gaia DR2 observations' + location: { page: 2 } + guinot22_forty_chips: + claim: >- + Star selection and PSF modelling are carried out independently on + each of the 40 MegaCam chips, with no chip excluded. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_forty_chips + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'is performed independently on each of the 40 chips that constitute the MegaCAM' + location: { page: 3 } + + # ═════════════════════════════════════════════════════════════════════════ + star_selection_psf: + description: >- + Which objects constrain the PSF, and the PSF model itself. Modules: + src/shapepipe/pipeline/str_handler.py (_mode — the iterative + histogram-zoom FWHM mode estimator centring the star box; median + fallback below N=20), src/shapepipe/modules/setools_package/setools.py + (_make_rand_split), src/shapepipe/modules/psfex_interp_package/ + psfex_interp.py (acceptance gates, HSM shapes). Configs: + star_selection.setools, default.psfex, config_exp_psfex.ini. + [LINT] star_stat logs the FWHM cut as mode +- 0.1*0.187 while the mask + applies mode +- 0.2 px — the run's own log misstates the selection. + [LINT] pixel scale appears as 0.187 (load-bearing) and 0.186 + (plot-only) in the same setools file. + inputs: + - id: exposure_sexcat + type: data + source: per-CCD exposure SExtractor catalogues (default_exp.sex run) + outputs: + - id: psf_model + type: data + format: psf + description: >- + Per-CCD PSFEx models + interpolated PSFs at object positions + (run_sp_exp_SxSePsfPi family), with HSM shape diagnostics. + decisions: + [star_selection_box, psf_train_validation_split, + psfex_candidate_vetting, psf_modelling_software, + psf_model_complexity, psf_acceptance_thresholds] + decisions: + star_selection_box: + label: Stellar-locus selection — mag window + FWHM window around the mode + rationale: >- + 18 < MAG_AUTO < 22, |FWHM - mode| <= 0.2 px, FLAGS==0, + IMAFLAGS_ISO==0; the mode is computed on a preselection + (MAG_AUTO<21, 0.3-1.5 arcsec at 0.187"/px) via the iterative + histogram-zoom estimator (str_handler.py::_mode, eps=0.001; median + fallback for N<20, -1 for N=0 — small-N behaviour changes selection + on sparse CCDs). PSFEx's automatic FWHM-range selection is off + (SAMPLE_AUTOSELECT N) and bad-pixel filtering is off, but PSFEx's + compiled-in fixed sample cuts still apply on top of this box — + see psfex_candidate_vetting. + Anchor: workflow/config/cfis/star_selection.setools#MASK:star_selection.MAG_AUTO; + workflow/config/cfis/star_selection.setools#MASK:preselect.MAG_AUTO; + workflow/config/cfis/default.psfex#SAMPLE_AUTOSELECT; + src/shapepipe/pipeline/str_handler.py::_mode. + default: mode_centred_box + options: + mode_centred_box: + label: FWHM-mode-centred box, +-0.2 px, mag 18-22, setools-only vetting + insights: [guinot22_star_box] + size_mag_locus_fit: + label: Fitted size-magnitude stellar locus + excluded: true + excluded_reason: Not wired; the mode-box is the validated v2.0 selection. + psfex_autoselect: + label: PSFEx SAMPLE_AUTOSELECT vetting on top + excluded: true + excluded_reason: >- + Deliberately disabled so selection lives in one place; rationale + not recorded in code. + psf_train_validation_split: + label: Random 80/20 star split — model fit vs held-out validation + rationale: >- + RAND_SPLIT ratio 20: star_split_ratio_80 fits the PSFEx model + (config_exp_psfex.ini FILE_PATTERN, and the tile multi-epoch + interpolation ME_DOT_PSF_PATTERN in config_tile_PiViVi.ini); + star_split_ratio_20 is the independent PSF-residual diagnostic + (PSFEX_INTERP MODE=VALIDATION). Trades model precision (fewer + training stars per CCD, interacting with the STAR_THRESH gate) + against an independent residual test. + [PENDING #873] The split is DETERMINISTIC: _make_rand_split takes + np.random.RandomState(seed).permutation(cat_size), the seed being + the digits of the unit's file number mod 2^32 — a pure function of + the input catalogue, fixed per CCD and independent of processing + order, the same philosophy as shape_measurement.ngmix_seed_mode's + SEED_FROM_POSITION. Before this the split drew from unseeded + np.random.randint, so the star sample entering the PSF model — and + therefore every shape downstream of it — differed between + identical runs; it was the one unseeded draw the position-seed work + left uncovered. One-off cost: the realised 80/20 membership changes + once (it is one further draw, now frozen), so PSF models and shapes + shift by that draw relative to every earlier product. + Anchor: workflow/config/cfis/star_selection.setools#RAND_SPLIT:star_split.RATIO; + src/shapepipe/modules/setools_package/setools.py::SETools._make_rand_split; + workflow/config/cfis/config_exp_psfex.ini#PSFEX_RUNNER.FILE_PATTERN; + workflow/config/cfis/config_tile_PiViVi.ini#PSFEX_INTERP_RUNNER.ME_DOT_PSF_PATTERN. + default: split_80_20_seeded + options: + split_80_20_seeded: + label: 80% train / 20% validation, seeded from the file number + insights: [guinot22_star_split] + split_80_20_unseeded: + label: Same split, unseeded np.random (pre-#873) + excluded: true + excluded_reason: >- + Retired by #873: it made the PSF star sample — and every shape + downstream of it — irreproducible run-to-run, the single + remaining unseeded draw in the science chain. Kept on the + record because every UNIONS product built before the smk-g4 + campaign was produced under it. + no_holdout: + label: 100% of stars in the model, no held-out diagnostic + excluded: true + excluded_reason: Loses the independent rho-statistic input. + psfex_candidate_vetting: + label: PSFEx-side candidate vetting — built-in defaults, unpinned + rationale: >- + default.psfex sets only SAMPLE_AUTOSELECT N; SAMPLE_MINSN, + SAMPLE_MAXELLIP, SAMPLE_FWHMRANGE, SAMPLE_VARIABILITY are absent, + so PSFEx's compiled-in defaults apply silently (MINSN 20, + MAXELLIP 0.3, FWHMRANGE 2-10 px, VARIABILITY 0.2) — a second star + selection nobody's config records, and one that changes if the + PSFEx binary version changes. BADPIXEL_FILTER N + PSF_RECENTER N: + star vignets with flagged/sentinel pixels are accepted unfiltered + and candidates are not recentred (CENTER_KEYS XWIN). The setools + box is therefore not the whole selection. [HARDCODED] (in the + PSFEx binary). + Anchor: workflow/config/cfis/default.psfex#SAMPLE_AUTOSELECT; + workflow/config/cfis/default.psfex#BADPIXEL_FILTER; + workflow/config/cfis/default.psfex#PSF_RECENTER. + default: builtin_defaults + options: + builtin_defaults: + label: "Compiled-in MINSN 20 / MAXELLIP 0.3 / FWHMRANGE 2-10, no bad-pixel filter" + insights: [guinot22_psfex_preselection_off] + pinned_explicit: + label: Write the SAMPLE_* values explicitly into default.psfex + psf_modelling_software: + label: PSF modelling software — PSFEx per-CCD vs MCCD focal-plane + rationale: >- + The committed chain fits PSFEx independently per CCD. MCCD + (Liaudat+2021) is a maintained in-tree alternative: a focal-plane + model fit across all 40 CCDs at once with a hybrid local+global + decomposition (src/shapepipe/modules/mccd_package/ + six + mccd_*_runner.py; knobs in example/cfis/config_MCCD.ini — + N_COMP_LOC=8, D_COMP_GLOB=8, LOC_MODEL=hybrid, MIN_N_STARS=20, + RMSE_THRESH=1.25). Unwired in workflow/config/cfis/ (needs the + MCCD config adapted, and config_exp_mccd.ini carries a stale + hardcoded PSF_MODEL_DIR path). The image-simulation path + substitutes PSF modelling entirely: fake_psf_runner injects the + true input PSF from a SKiLLS dictionary in psfex_interp's output + format. + Anchor: src/shapepipe/modules/mccd_package; + src/shapepipe/modules/fake_psf_package; + example/cfis/config_MCCD.ini#INSTANCE.N_COMP_LOC; + example/cfis/config_MCCD.ini#INPUTS.MIN_N_STARS; + example/cfis/config_exp_mccd.ini. + default: psfex + options: + psfex: + label: PSFEx, independent per-CCD models + insights: [guinot22_psfex_software, farrens22_two_psf_methods] + mccd_focal_plane: + label: MCCD hybrid local+global focal-plane model + true_input_psf: + label: fake_psf injection of the simulation's true PSF + excluded: true + excluded_reason: >- + Only meaningful on simulated images where the true PSF exists; + not a data-analysis option. + psf_model_complexity: + label: PSFEx model — pixel basis, degree-2 spatial polynomial per CCD + rationale: >- + BASIS_TYPE PIXEL, BASIS_NUMBER 20, PSF_SIZE 51,51, PSF_SAMPLING 1, + PSFVAR_DEGREES 2 in XWIN,YWIN per CCD (MEF_TYPE INDEPENDENT, + STABILITY_TYPE EXPOSURE), PSF_RECENTER N. Model flexibility sets + the PSF-leakage/overfitting balance — the dominant additive + systematic in cosmic shear. Values are the stock EB 2017 header; + rationale not recorded in code. + Anchor: workflow/config/cfis/default.psfex#BASIS_TYPE; + workflow/config/cfis/default.psfex#PSFVAR_DEGREES; + workflow/config/cfis/default.psfex#PSF_SIZE. + default: pixel_basis_deg2_per_ccd + options: + pixel_basis_deg2_per_ccd: + label: "PIXEL basis, degree 2, per-CCD" + insights: [guinot22_psf_no_oversampling] + deg3: + label: Degree-3 spatial variation + excluded: true + excluded_reason: >- + More flexibility per CCD needs more stars per CCD than the + count-floor world guarantees; not validated. + psf_acceptance_thresholds: + label: Per-CCD PSF-model quality gate (min stars, max chi2) + rationale: >- + A CCD whose model has ACCEPTED < STAR_THRESH or CHI2 > 2 is not + interpolated — its galaxies drop from the shear catalogue: direct + footprint selection, the in-code analogue of the DES blacklist. + [PENDING #873] Both passes now gate at 22 stars: the VALIDATION-mode + exposure config always did (config_exp_psfex.ini), and #873 raised + the MULTI-EPOCH science path 20 -> 22 in example/cfis + (config_tile_PiViVi_canfar_{sx,uc}.ini), with commit 90782098 + mirroring it into workflow/config/cfis/config_tile_PiViVi.ini — the + committed config fork this workflow actually reads (#848 D2). + Provenance of the retired 20, which is what makes this a fix rather + than a preference: commit fdc86553 (Kilbinger, 2020-06-30, "Forgot + to update new star number threshold for 80% of stars") deliberately + bumped 20 -> 22 to account for the 80/20 split, but only in the + validation config; the tile config kept the pre-split 20, so for + five years the SCIENCE path gated on the value that 2020 fix meant + to retire. (20 is also the psfex_interp function default, so the + stale-value reading rested on the commit provenance rather than on + the config alone.) + Published description (Guinot+22 p.4, Fig. 3): 22 stars/CCD, applied + to exactly the CCDs feeding multi-epoch shape measurement — the + number now agrees. Two gaps remain: an undocumented CHI2_THRESH=2 in + both configs, and the mechanism — interpsfex tests the PSFEx header + ACCEPTED/CHI2 at interpolation time and drops that epoch for objects + on the CCD, rather than excluding the CCD from PSF modelling as the + paper describes. + Anchor: src/shapepipe/modules/psfex_interp_package/psfex_interp.py::PSFExInterpolator.interpsfex; + workflow/config/cfis/config_exp_psfex.ini#PSFEX_INTERP_RUNNER.STAR_THRESH; + workflow/config/cfis/config_tile_PiViVi.ini#PSFEX_INTERP_RUNNER.STAR_THRESH; + example/cfis/config_tile_PiViVi_canfar_sx.ini#PSFEX_INTERP_RUNNER.STAR_THRESH; + example/cfis/config_tile_PiViVi_canfar_uc.ini#PSFEX_INTERP_RUNNER.STAR_THRESH. + default: stars22_chi2_2 + options: + stars22_chi2_2: + label: ">= 22 stars on both passes, chi2 <= 2" + insights: [des_psf_blacklist_local, guinot22_star_floor_22_local] + stars20_chi2_2: + label: ">= 20 stars on the science path, 22 in validation (pre-#873)" + excluded: true + excluded_reason: >- + Retired by #873 + 90782098. It was never a chosen value: it is + the pre-split threshold fdc86553 raised to 22 in 2020 for the + validation config and forgot on the science path, leaving the + science gate below both the published floor (Guinot+22 Fig. 3) + and the pipeline's own intent. Every UNIONS product built before + the smk-g4 campaign carries it. + des_25: + label: DES Y3 threshold (25 stars) + excluded: true + excluded_reason: >- + Not adopted; CFIS CCDs are smaller than DECam's — the right + number is survey-specific. + prior_insights: + des_psf_blacklist_local: + claim: >- + DES Y3 blacklists any CCD with fewer than 25 stars surviving + outlier rejection in the PSF fit. + created_at: "2026-07-16T00:00:00Z" + evidence: + - id: ev_jarvis_y3_local + doi: "10.48550/arXiv.2011.03409" + quote: + exact: "fewer than 25 stars survived the outlier rejection" + location: { page: 10 } + guinot22_star_floor_22_local: + claim: >- + The published ShapePipe/UNIONS analysis discards a CCD from the PSF + estimation when fewer than 22 stars are selected on it — the floor + the science-path PSF-interpolation gate now applies. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_star_floor_local + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'The dashed line represents the cut at 22 stars/CCD below which the CCD is discarded for the PSF estimation.' + location: { page: 4 } + guinot22_star_box: + claim: >- + The published star selection keeps objects whose FWHM lies within + 0.04 arcsec of the mode of a size preselection, restricted to the + magnitude range 18 < r < 22. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_star_box + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'From this pre-selection we keep objects for which the FWHM is within 0.04 arcsec of the mode. In addition to these size cuts, we only use star candidates in the magnitude range 18 < r < 22.' + location: { page: 4 } + guinot22_star_split: + claim: >- + The published analysis randomly splits the star sample in two, 80% + building the PSF model and 20% held out for the validation tests. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_star_split + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'To be able to perform these tests properly, our star sample has been randomly divided in two:' + location: { page: 7 } + guinot22_psfex_preselection_off: + claim: >- + PSFEx's internal pre-selection is deliberately disabled so that the + pipeline's own star selection is the only one, the paper describing + the model as fit on the entire star sample. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_preselection_off + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'Since we carry out our own star selection (see Sect. 4.1), we disable the internal PSFEx pre-selection, and the PSF is thus obtained using the entire star sample.' + location: { page: 4 } + guinot22_psfex_software: + claim: >- + The published UNIONS shear catalogue uses PSFEx for PSF modelling, + with MCCD named as upcoming rather than current work. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_psfex_software + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'We make use of the PSFEx software package' + location: { page: 4 } + farrens22_two_psf_methods: + claim: >- + ShapePipe ships two PSF modelling methods, PSFEx and MCCD, either or + both of which may be run and used for the galaxy shape measurement. + created_at: "2022-06-01T00:00:00Z" + evidence: + - id: ev_farrens22_two_psf + doi: "10.48550/arXiv.2206.14689" + quote: + exact: 'ShapePipe allows either or both methods to be run and subsequently used for the galaxy shape measurement.' + location: { page: 3 } + guinot22_psf_no_oversampling: + claim: >- + The PSFEx parametrisation is tabulated in the paper (pixel basis, + degree-2 spatial variation), with the deliberate choice not to + over-sample the PSF models. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_psf_complexity + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'The PSFEx parameters we used are presented in Table 1. We have chosen not to over-sample the PSF models.' + location: { page: 5 } + + # ═════════════════════════════════════════════════════════════════════════ + shape_measurement: + description: >- + Galaxy shape estimation: ngmix single-Gaussian fits with metacalibration. + Modules: src/shapepipe/modules/ngmix_package/ngmix.py (priors, metacal + setup, epoch handling, postage-stamp prep), ngmix_runner.py (config + exposure). Config: config_tile_Ng_template.ini. Most values here are + HARDCODED — scientific choices living in code with no config exposure; + this sub-analysis is where the silent-default risk concentrates. + [LINT] centroid_source default disagrees between the runner ("wcs", + production; always passed explicitly, ngmix_runner.py:170) and every + module-level signature ("hsm") — unreachable in the pipeline path, but + direct callers (tests, notebooks) silently get the other choice. + [LINT] pixel scale is 0.186 here (config PIXEL_SCALE) vs 0.187 in the + setools/masking configs — and star_selection.setools itself mixes + 0.187 (cuts) with 0.186 (SCATTER stat, :75). + inputs: + - id: vignets + type: data + source: 51x51 galaxy/weight/background-RMS vignets + interpolated PSFs (vignetmaker, psfex_interp) + outputs: + - id: ngmix_cat + type: data + format: fits + description: Per-tile metacal shear catalogue chunks (ngmix_runner family). + decisions: + [ngmix_seed_mode, galaxy_model, fit_priors, metacal_scheme, + centroid_source, epoch_quality_and_weighting, noise_model, + psf_epoch_loss_policy, megacam_ccd_flip] + decisions: + ngmix_seed_mode: + label: ngmix per-object RNG seeding + rationale: >- + Production historically seeded one RandomState from the tile ID, + consumed in object order — results depended on chunk boundaries. + The position seed (3-arcsec sky boxes + CCD offsets, zig-zag fold + + Cantor pairing mod 2^32; ngmix.py::position_seed) makes every + stream a function of sky position: chunk-invariant, + bit-reproducible, and metacal fixnoise counter-noise cancels across + image-simulation branches (ngmix#796). Consequence: chunking is + demoted to a pure throughput knob (reverting this decision + re-promotes it). Cost: noise streams change vs v2.0 — see + top-level baseline_validation_criterion. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::position_seed; + workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.SEED_FROM_POSITION; + src/shapepipe/modules/ngmix_runner.py::ngmix_runner. + default: position_seed + options: + position_seed: + label: Per-object seed from (ra, dec, ccd), 3-arcsec boxes + tile_seed: + label: Tile-wide RandomState (v2.0) + excluded: true + excluded_reason: >- + Chunk-dependent; retired outright (SEED_FROM_POSITION=False now + raises — ngmix_runner.py:110-116). + galaxy_model: + label: Galaxy and PSF model — single Gaussian [HARDCODED] + rationale: >- + ngmix.fitting.Fitter(model='gauss') for both galaxy and PSF + (ngmix.py::make_runners); guessers TPSFFluxAndPriorGuesser / + TFluxGuesser with T=0.25 and catalogue-flux guess, Runner ntry=5, + PSFRunner ntry=2 — with a non-convex likelihood, guess and retries + decide which objects converge (failed fits are NaN-filled with + flags, not raised). Rationale not recorded. Under metacal, model + bias largely cancels in the response, which is the standard defense + of 'gauss'; not stated in code. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::make_runners. + default: gauss + options: + gauss: + label: "Single Gaussian, T guess 0.25, ntry 5/2" + insights: [guinot22_gaussian_model] + exp_or_bdf: + label: exp / bdf galaxy models + excluded: true + excluded_reason: >- + Slower, and metacal makes the gain marginal; not validated on + CFIS. + fit_priors: + label: ngmix joint prior — GPriorBA(0.4), cen sigma = pixel scale, flat T/F [HARDCODED] + rationale: >- + Ellipticity GPriorBA sigma=0.4; centroid CenPrior sigma = one pixel + scale (0.186 arcsec, config PIXEL_SCALE — the coupling + sigma=pixel_scale is itself the hardcoded choice); flat T in + [-1, 1e3], flat F in [-100, 1e9] with negative support (bounds + decide which noisy fits survive vs rail). get_prior takes T/F range + arguments but no caller passes them. Prior width drives noise bias; + rationale not recorded. + Published description (Guinot+22 p.7): centroid sigma = pixel scale + ~0.187 arcsec, flat F in [-1e4, 1e9], flat half-light radius r50 in + [-10, 1e6] arcsec, ellipticity prior from Bernstein & Armstrong + (2014); current code: PIXEL_SCALE 0.186, flat F in [-100, 1e9], and + a flat prior on ngmix's second-moment size T in [-1, 1e3] rather + than on r50 — the prior families agree, the flux bound and pixel + scale have drifted, and the size prior is a different + parameterisation rather than a changed number. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::get_prior; + workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.PIXEL_SCALE. + default: gpriorba04_flat + options: + gpriorba04_flat: { label: "GPriorBA 0.4 + flat T/F with negative support" } + nonneg_informative: + label: Non-negative or informative T/F priors + excluded: true + excluded_reason: >- + Truncating negative support biases the noshear ensemble mean; + metacal wants symmetric noise response. + metacal_scheme: + label: Metacalibration — 5 types, step 0.01, fitgauss reconv, fixnoise [HARDCODED] + rationale: >- + types [noshear,1p,1m,2p,2m], step 0.01, psf='fitgauss' (runner + default; moves the metacal response directly — alternatives gauss/ + dilate/azgauss listed in the docstring), fixnoise=True, + use_noise_image=True, MetacalBootstrapper(ignore_failed_psf=True) + (changes which epochs enter the fit). No *_psf sheared types, so no + mcal_R_psf PSF-response term in the catalogue. fixnoise rationale + appears only in the position_seed docstring (counter-noise + cancellation). + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::do_ngmix_metacal. + default: five_types_step001_fitgauss + options: + five_types_step001_fitgauss: + label: "noshear+1p/1m/2p/2m, step 0.01, fitgauss, fixnoise" + insights: [guinot22_metacal_five_images] + with_psf_response: + label: Add sheared-PSF types for R_psf + excluded: true + excluded_reason: >- + Not wired; leakage is instead diagnosed via PSF_ORIG columns + + rho statistics downstream. + centroid_source: + label: Jacobian origin from WCS astrometry, not HSM moments + rationale: >- + Production runner default "wcs"; hsm is "legacy... noisy for stars + and flagged as incorrect by Fabian — see #767" (runner comment; a + rare recorded rationale). Moves the centroid-prior centre per + object. The runner reads an optional CENTROID_SOURCE config option + that no committed CFIS config sets. [LINT] module-level default is + still "hsm" — see this sub-analysis's description. + Published description (Guinot+22 p.7): HSM adaptive moments, run on + each sheared version, supplied the whole initial guess vector + (centroid, r50, flux) for the least-squares fit; current code: that + initialisation is gone — guesses come from ngmix's + TPSFFluxAndPriorGuesser with fixed T=0.25 and a catalogue flux, and + the only surviving HSM role is the optional stamp re-centering that + sets the Jacobian origin. So the drift is wider than a swapped + centroid source. (The paper's other HSM use, PSF/star shape + diagnostics, is unaffected.) + Anchor: src/shapepipe/modules/ngmix_runner.py::ngmix_runner; + src/shapepipe/modules/ngmix_package/ngmix.py::make_ngmix_observation. + default: wcs + options: + wcs: { label: WCS-projected catalogue position } + hsm: + label: HSM adaptive-moment centroid + excluded: true + excluded_reason: Noisy for stars; flagged incorrect (shapepipe#767). + epoch_quality_and_weighting: + label: Epoch admission, masking cut, and multi-epoch combination [HARDCODED] + rationale: >- + An epoch is dropped if >1/3 of its stamp is masked + (prepare_postage_stamps; the comment says "objects", the code drops + epochs — an object with zero surviving epochs drops out); failed + PSF fits drop epochs (flags != 0); fluxes rescaled by header FSCALE + (gal*Fscale, weight/Fscale^2); the diagnostic PSF is averaged over + epochs weighted by obs.weight.sum(). Joint multi-epoch fit over + survivors. Rationale for 1/3 and for the weight choice not + recorded. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::prepare_postage_stamps; + src/shapepipe/modules/ngmix_package/ngmix.py::rescale_epoch_fluxes; + src/shapepipe/modules/ngmix_package/ngmix.py::_average_psf_fits. + default: third_masked_cut + options: + third_masked_cut: { label: "Drop epoch if >1/3 masked; FSCALE rescale; weight-sum PSF average" } + noise_model: + label: Per-pixel inverse variance from background-RMS vignets + rationale: >- + BKG_RMS_VIGNET_PATH set in the CFIS template: weight = + 1/bkg_rms^2 per pixel (all-or-nothing; missing file errors); + fallback scalar 1/sigma_mad^2. Masked pixels filled with Gaussian + noise at sig_noise. A scalar sigma "mis-reports errors and erodes + the inverse-variance advantage whenever the RMS map actually + varies" (recorded rationale, fixnoise bookkeeping). PSF observation + gets a flat weight from PSF_NOISE=1e-5 — hardcoded module constant, + validated 1e-4..1e-6 on the digital twin (#749/#774 comment); + without it the g-prior swamps the PSF likelihood. Per-epoch + background subtraction BKG_SUB=True (off only for sims). + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::prepare_ngmix_weights; + src/shapepipe/modules/ngmix_package/ngmix.py::PSF_NOISE; + src/shapepipe/modules/ngmix_package/ngmix.py::background_subtract; + workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.BKG_RMS_VIGNET_PATH. + default: rms_vignet_weights + options: + rms_vignet_weights: { label: Per-pixel RMS-map weights + PSF_NOISE 1e-5 } + scalar_sigma_mad: + label: Scalar sigma_mad per epoch + excluded: true + excluded_reason: Mis-reports errors where the RMS map varies (recorded). + psf_epoch_loss_policy: + label: Object-level policy when CCDs fail PSF interpolation + rationale: >- + When k of ~40 CCDs fail the PSF acceptance gate (~5-6% attrition, + per-exposure clustered, matches the bash baseline — but measured + with the science gate at 20 stars, so [PENDING #873] at 22 it can + only rise, and smk-g4 is the first campaign to re-measure it), the + pipeline applies NO further quality gate: tiles complete, each + object records NGMIX_N_EPOCH, and sp report surfaces per-tile + epoch loss. + Object-level protection is delegated entirely to the validation + stage's epoch-count cut (sp_validation's galaxy selection cuts on N_EPOCH >= 1). Rationale (Cail, + 2026-08-29, PRD walk): epoch loss is a per-object depth effect + already recorded in the catalogue; gating at pipeline level would + fail whole tiles for a versionable catalogue decision. PRD #848's + open-questions section was removed accordingly. Per-tile epoch loss + is surfaced by the run report; NGMIX_N_EPOCH is the per-object + record. + Anchor: workflow/scripts/run_report.py; + workflow/config/cfis/final_cat.param#NGMIX_N_EPOCH. + default: record_and_delegate + options: + record_and_delegate: + label: Record NGMIX_N_EPOCH, report attrition, no pipeline gate + pipeline_epoch_floor: + label: Fail tiles below a minimum surviving-epoch fraction + excluded: true + excluded_reason: >- + Fails whole tiles for what is a versionable per-object + catalogue decision; the depth effect is already recorded. + megacam_ccd_flip: + label: 180-degree tile-vignet rotation for MegaCam CCDs <18 and 36-37 [HARDCODED] + rationale: >- + "MegaPipe has CCDs that are upside down" (docstring) — the tile + vignet is rotated to register with epoch stamps; a wrong flip + mis-registers the tile mask against the epoch, changing flagged + pixels and the 1/3-masked cut. Carries its own recorded caveat: + "will give incorrect results when used with THELI ccds. Fix this." + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::Ngmix.MegaCamFlip. + default: megapipe_flip + options: + megapipe_flip: { label: Flip CCDs <18 and 36/37 (MegaPipe orientation) } + prior_insights: + guinot22_gaussian_model: + claim: >- + The published shape measurement models galaxies with a single + Gaussian profile, arguing the resulting model bias is small and + largely absorbed by metacalibration. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_gaussian + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'Despite being very simple, the model bias (Kacprzak et al.' + suffix: ' 2014) is small.' + location: { page: 7 } + guinot22_metacal_five_images: + claim: >- + Metacalibration in the published analysis creates four sheared + images for calibration plus one for measurement, with a shear step + of 0.01 and a 90-degree-rotated noise image to cancel noise + correlations. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_metacal + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'This method creates four images used for the calibration, and one for the measurement.' + location: { page: 6 } + + # ═════════════════════════════════════════════════════════════════════════ + psf_diagnostics: + description: >- + The PSF-fidelity diagnostic chain: merge the held-out (20%) validation + stars into one catalogue, bin PSF shapes and residuals over the focal + plane. DORMANT in the committed snakemake chain — no rule runs it. + Modules: src/shapepipe/modules/merge_starcat_runner.py (+ per-model + merge classes in merge_starcat.py), mccd_plots_runner.py (serves both + PSF models despite its name). Boundary note: the module docstring + (mccd_package/__init__.py:157) still advertises rho-statistics plots, + but no rho/treecorr code remains in shapepipe — rho/tau statistics + moved downstream to sp_validation (rho_tau.py via + shear_psf_leakage.RhoStat/TauStat; treecorr min_sep/max_sep/nbins, + jackknife patch numbers hardcoded per survey with a "TODO to yaml"). + The diagnostic decision chain thus crosses the repo boundary into + sp_validation. [LINT] the module docstring still advertises rho + statistics this package no longer computes. + inputs: + - id: validation_star_cats + type: data + source: per-CCD star_split_ratio_20 catalogues with PSF/star HSM shapes (psfex_interp VALIDATION mode) + outputs: + - id: merged_star_catalogue + type: data + format: fits + description: >- + One full_starcat over the run — the input rho/tau statistics and + leakage diagnostics consume downstream. + decisions: [starcat_merge_source, meanshape_binning] + decisions: + starcat_merge_source: + label: Which PSF model's validation output feeds the merged star catalogue + rationale: >- + merge_starcat_runner dispatches on PSF_MODEL in {psfex, mccd, + setools} to per-model merge classes (different HDU conventions: + mccd HDU 1, psfex/setools HDU 2). Follows star_selection_psf. + psf_modelling_software; recorded separately because the merge can + also consume raw setools output (pre-model diagnostics). + Anchor: src/shapepipe/modules/merge_starcat_runner.py::merge_starcat_runner; + src/shapepipe/modules/merge_starcat_package/merge_starcat.py. + default: psfex + options: + psfex: + label: PSFEx validation catalogues (HDU 2) + insights: [guinot22_psfex_for_v1] + mccd: { label: MCCD validation catalogues (HDU 1) } + setools: { label: Raw setools star catalogues } + meanshape_binning: + label: Focal-plane mean-shape binning and outlier handling + rationale: >- + PSF ellipticity/size and residuals binned per CCD over the focal + plane: X_GRID=5, Y_GRID=10 bins per CCD, colour scales MAX_E=0.05, + MAX_DE=0.005, REMOVE_OUTLIERS=False (example/cfis config; no + committed workflow config exists). Grid resolution sets which + spatial PSF-residual structure is visible; outlier removal changes + what the diagnostic hides. + Published description (Guinot+22 p.8): focal-plane residual maps + averaged in ~20 arcsec cells per CCD; current code: no committed + workflow config picks a grid at all — example/cfis carries both the + 5x10 grid recorded here (~77x86 arcsec) and, in + config_valjoint_Pl_mccd.ini, a 20x40 grid (~19x22 arcsec) that + reproduces the paper. This is therefore an undetermined knob with + two committed precedents rather than a value that drifted. The + uniform REMOVE_OUTLIERS=False is a genuine paper-silence gap. + Anchor: example/cfis/config_MsPl_psfex.ini#MCCD_PLOTS_RUNNER.X_GRID; + example/cfis/config_MsPl_psfex.ini#MCCD_PLOTS_RUNNER.REMOVE_OUTLIERS; + src/shapepipe/modules/mccd_plots_runner.py; + src/shapepipe/modules/mccd_package/mccd_plot_utilities.py::plot_meanshapes. + default: grid_5x10 + options: + grid_5x10: { label: "5x10 per CCD, outliers kept" } + prior_insights: + guinot22_psfex_for_v1: + claim: >- + PSFEx is the PSF model behind the published UNIONS v1 catalogue, so + the merged validation star catalogue and its diagnostics are fed by + PSFEx output. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_psfex_v1 + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'We make use of the PSFEx software package' + location: { page: 4 } + + # ═════════════════════════════════════════════════════════════════════════ + survey_geometry: + description: >- + Effective survey area and mask geometry for two-point estimators. + DORMANT — no committed workflow rule. Module: + src/shapepipe/modules/random_cat_package/random_cat.py (+ runner): + uniform randoms over each tile rejected against the pipeline mask, + yielding effective area (overlap- and mask-corrected) and optionally + the mask itself as a HEALPix map (save_as_healpix). This is the + in-repo ancestor of the planned healsparse external-mask rework — whichever way that rework lands, this + sub-analysis is where its geometry decisions belong. Bitrot risk: + healpy is imported but absent from pyproject.toml dependencies. + inputs: + - id: tile_masks + type: data + source: per-tile pipeline flag maps + final catalogues + outputs: + - id: random_catalogue + type: data + format: fits + description: Per-tile random catalogue + effective area (+ optional HEALPix mask). + decisions: [random_sampling, healpix_mask_export] + decisions: + random_sampling: + label: Random-point density for area estimation + rationale: >- + N_RANDOM=50000 with DENSITY=True (per square degree; False = total + per tile) in the example config; no committed workflow value. + Sampling density sets the Monte Carlo noise floor on effective + area, which propagates to two-point normalisation. + Anchor: example/cfis/config_Rc.ini#RANDOM_CAT_RUNNER.N_RANDOM; + example/cfis/config_Rc.ini#RANDOM_CAT_RUNNER.DENSITY; + src/shapepipe/modules/random_cat_runner.py::random_cat_runner. + default: per_sqdeg_50k + options: + per_sqdeg_50k: { label: "50000 per sq deg" } + healpix_mask_export: + label: HEALPix export resolution for the pipeline mask + rationale: >- + SAVE_MASK_AS_HEALPIX=True, HEALPIX_OUT_NSIDE=1024 (~3.4 arcmin + pixels) in the example config — coarser than the arcsecond-scale + mask features it rasterises; the resolution choice decides what the + exported mask can represent. Supersession candidate under + masking-unification (healsparse). + Anchor: example/cfis/config_Rc.ini#RANDOM_CAT_RUNNER.SAVE_MASK_AS_HEALPIX; + example/cfis/config_Rc.ini#RANDOM_CAT_RUNNER.HEALPIX_OUT_NSIDE; + src/shapepipe/modules/random_cat_package/random_cat.py::RandomCat.save_as_healpix. + default: nside_1024 + options: + nside_1024: { label: nside 1024 } + + # ═════════════════════════════════════════════════════════════════════════ + catalogue_assembly: + description: >- + Final per-tile catalogue: merging shape chunks, attaching photometry + and PSF diagnostics, classification, sentinels. Modules: + src/shapepipe/modules/make_cat_package/make_cat.py (+ runner), + merge_sep_cats.py, vignetmaker_package (stamps, see top-level + postage_stamp_size), find_exposures_package (epoch list from tile + HISTORY cards). Configs: config_tile_Mc.ini, config_tile_PiViVi.ini, + final_cat.param. + inputs: + - id: ngmix_chunks + type: data + source: per-tile ngmix catalogue chunks + tile sexcat + PSF diagnostics + outputs: + - id: tile_final_cat + type: data + format: fits + description: The assembled per-tile science catalogue (final_cat family). + decisions: + [star_galaxy_classification, tile_overlap_handling, + column_selection, failure_sentinels, postproc_run_provenance, + shape_catalogue_shortfall_guard] + decisions: + star_galaxy_classification: + label: Star/galaxy separation — deferred out of the pipeline + rationale: >- + Production sets SM_DO_CLASSIFICATION=False (config_tile_Mc.ini) and + wires no spread-model input: SPREAD_MODEL/SPREADERR_MODEL are + sentinel 99, no SPREAD_CLASS column, and the SPREAD_* entries in + final_cat.param are commented out. The dormant machinery + (make_cat.py::save_sm_data) classifies on class = sm + 2*sm_err + with star |class|<0.003, galaxy class>0.01 — thresholds hardcoded + in the function signature. Reactivation is a 4-line config diff + (run spread_model_runner after psfex_interp+vignetmaker, add its + output to make_cat inputs, flip the switch — the exact diff + between example/cfis/config_make_cat_psfex.ini and _nosm.ini; the + defunct tile wiring config_tile_PiViSmVi.ini is the reference). + Separation therefore happens entirely downstream (sp_validation); + the pipeline ships everything. Rationale for deferring not + recorded. + Published description (Guinot+22 p.5-6): galaxies are selected + inside the pipeline with the spread model, at s + 2*sigma_s > + 0.0003 together with s > 0 and 20 < MAG_AUTO < 26; current code: + classification is disabled entirely and separation deferred to + sp_validation, with the dormant make_cat thresholds putting the + like-for-like galaxy boundary at 0.01, some thirty times the + published cut (0.003 is its separate star-side bound). Even + reactivated, the code implements only the spread-model test — the + paper's companion cuts have no in-pipeline counterpart. + Anchor: workflow/config/cfis/config_tile_Mc.ini#MAKE_CAT_RUNNER.SM_DO_CLASSIFICATION; + src/shapepipe/modules/make_cat_package/make_cat.py::save_sm_data; + workflow/config/cfis/final_cat.param#SPREAD_CLASS; + example/cfis/config_make_cat_psfex_nosm.ini; + example/cfis/defunct/config_tile_PiViSmVi.ini. + default: deferred_downstream + options: + deferred_downstream: + label: No in-pipeline classification; catalogue ships all objects + spread_model_inline: + label: spread_model classification in make_cat (0.003/0.01) + excluded: true + excluded_reason: >- + Machinery present but unwired in production; reactivating it + changes which objects downstream sees as galaxies. + tile_overlap_handling: + label: Tile-overlap duplicates — neither removed nor flagged + rationale: >- + Adjacent tiles overlap; objects in the overlap are measured in + both. make_cat attaches only TILE_ID (parsed from the sexcat + filename); no unique-object rule, no overlap flag. [LINT] the + documented config key TILE_LIST ("used to flag objects in areas of + overlap between tiles", in the make_cat package docstring) is + implemented nowhere — grep across src/ and workflow/ finds only + the docstring. VERIFIED downstream: sp_validation dedups at + classification time (galaxy.py::classification_galaxy_overlap_ra_dec + RA/Dec cuts to non-overlapping tile areas, and the WCS-based + mask_overlap variant; applied as cut_overlap in + classification_galaxy_base) — so this is today's division of labour, + and the pipeline's contract is "ship duplicates, TILE_ID is the + handle"; flagging overlaps in the catalogue stays an open option. The dead TILE_LIST docstring remains a lint. + Anchor: src/shapepipe/modules/make_cat_package/make_cat.py::save_sextractor_data; + src/shapepipe/modules/make_cat_package/__init__.py. + default: no_dedup_in_pipeline + options: + no_dedup_in_pipeline: + label: Ship duplicates; TILE_ID is the only handle + overlap_flagging: + label: Implement the documented TILE_LIST overlap flag + nearest_tile_centre: + label: Keep each object only in its nearest tile + excluded: true + excluded_reason: >- + Requires cross-tile coordination at assembly time, which the + per-tile DAG deliberately avoids; dedup belongs downstream if + anywhere. + column_selection: + label: Which columns survive into the science catalogue + rationale: >- + final_cat.param: positions XWIN/YWIN_WORLD, TILE_ID, flags + (FLAGS, IMAFLAGS_ISO, NGMIX_MCAL_FLAGS), PSF ellipticity from + PSF_ORIG only, all five metacal branches of G1/G2/T/FLUX/FLAGS, + but shear errors only for NOSHEAR (sheared-branch error columns + commented out — downstream response-weighted estimators cannot + propagate per-branch errors), SExtractor photometry (MAG_AUTO, + FLUX_APER, FLUX_RADIUS, SNR_WIN, FWHM_*), N_EPOCH/NGMIX_N_EPOCH, + NGMIX_MOM_FAIL. Doesn't change membership, but determines which + numbers exist for downstream cuts and calibration. Note + final_cat.param is read by scripts/python/create_final_cat.py in + post-processing, outside the per-tile DAG. [LINT] see detection: + IMAFLAGS_ISO is requested but never reaches the merged catalogue. + Anchor: workflow/config/cfis/final_cat.param; + scripts/python/create_final_cat.py. + default: committed_param_list + options: + committed_param_list: { label: The committed final_cat.param set } + failure_sentinels: + label: Objects without shape measurements kept, with sentinel values [HARDCODED] + rationale: >- + Unmatched objects (no ngmix row) stay in the catalogue with + sentinels: sizes/fluxes/flags 0, error fluxes/mags -1, + ellipticities -10, T_ERR 1e30. The sentinel choice defines what a + downstream cut must exclude — a naive G1 > -1 cut silently changes + the sample. Flag-0-for-failure is the sharpest hazard: a failed + object's NGMIX_MCAL_FLAGS reads as success. Rationale not recorded. + Anchor: src/shapepipe/modules/make_cat_package/make_cat.py::SaveCatalogue._save_ngmix_data. + default: sentinel_values + options: + sentinel_values: { label: "Keep with sentinels (flags 0, e -10, T_ERR 1e30)" } + drop_unmatched: + label: Drop objects without shapes + excluded: true + excluded_reason: >- + Loses the photometry-only population and hides attrition from + the completeness accounting. + postproc_run_provenance: + label: Post-proc run selection — newest mtime wins, merged patches never refresh + rationale: >- + create_final_cat picks each tile's make_cat run by newest + directory mtime (skipping runs without an output FITS), and the + merged patch catalogue is incremental: a tile already present is + never refreshed — a reprocessed tile reaches the patch only via + an explicit single-ID remove+add. mtime is filesystem state, not + provenance: a touched old run can outrank a newer one. Outside + the per-tile DAG (scripts/, not workflow/). Anchor: + scripts/python/create_final_cat.py::process. + default: newest_mtime_incremental + options: + newest_mtime_incremental: { label: "Newest mtime, incremental merge, manual refresh" } + shape_catalogue_shortfall_guard: + label: 10% shape-shortfall guard — logged, not enforced [HARDCODED] + rationale: >- + If the merged shape catalogue covers <10% of the detection + catalogue, make_cat logs an error but the enforcement (return + + raise) is commented out in both the module and its runner: a tile + whose shapes are 90% missing from a processing error is written and + looks normal. The comment distinguishes the two causes (measurement + failure = ok; premature merge = error) but not why enforcement is + off. Interacts with top-level per_unit_count_floor, which floors + tile_ngmix at 1 file and cannot see intra-file attrition. + Anchor: src/shapepipe/modules/make_cat_package/make_cat.py::SaveCatalogue._save_ngmix_data; + src/shapepipe/modules/make_cat_runner.py::make_cat_runner. + default: log_only + options: + log_only: { label: "Log the shortfall, write the tile anyway" } + enforce_10pct: + label: Fail the tile below 10% coverage + excluded: true + excluded_reason: >- + Was the coded intent, then disabled — reason unrecorded; + flagged as a question, not an endorsement. + +# ── Dormant science paths, surveyed but not yet recorded as sub-analyses ──── +# Candidates for future passes (each carries real scientific knobs): +# * External photometry match — match_external_package (TOLERANCE=0.3 +# arcsec against UNIONS ugriz; the external catalogue path is hardcoded +# to an IAP/candide location). +# * Image-simulation validation wiring — example/cfis_image_sims/: +# same chain over SKiLLS images with fake_psf substitution; bash-shaped, +# not yet ported to snakemake. diff --git a/universes/committed.yaml b/universes/committed.yaml new file mode 100644 index 000000000..5530226cf --- /dev/null +++ b/universes/committed.yaml @@ -0,0 +1,70 @@ +id: committed +description: The committed configuration on feat/snakemake-orchestration. +decisions: + per_unit_count_floor: count_floor + postage_stamp_size: px_51 + photometric_zeropoint: fixed_30_tiles_header_exposures + baseline_validation_criterion: statistical_parity +analyses: + masking: + decisions: + star_catalogue_query: gsc_23_vizier + star_magnitude_definition: mean_finite_bands + bright_star_mask_geometry: megaprime_polygon_linear_scaling + deep_sky_object_masking: circles_no_padding + border_mask_width: px50_exposures_only + pixel_threshold_flags: stock_ww_thresholds + external_flag_usage: exposures_only + detection: + decisions: + detection_threshold_policy: thresh_1p5_minarea5_fwhm2px_filter + deblending_policy: mincont_5em4_tiles + background_model: manual_zero_tiles_auto_exposures + weighting_and_interpolation: map_weight_interp_all + detection_source_mode: sx_nomask_single_image + epoch_membership_ccd_bounds: trimmed_bounds_33_2080 + photometry_parameters: kron_25_35 + cleaning_and_neighbour_masking: clean_1_correct + preparation: + decisions: + astrometric_solution_source: delivered_headers + ccd_split_extent: all_40_hdus + epoch_provenance_from_tile_history: history_parse + object_position_columns: xwin_windowed + stamp_positioning_and_padding: round_and_zero_pad + epoch_flag_source: raw_flags + star_selection_psf: + decisions: + star_selection_box: mode_centred_box + psfex_candidate_vetting: builtin_defaults + psf_modelling_software: psfex + psf_train_validation_split: split_80_20_seeded + psf_model_complexity: pixel_basis_deg2_per_ccd + psf_acceptance_thresholds: stars22_chi2_2 + shape_measurement: + decisions: + ngmix_seed_mode: position_seed + galaxy_model: gauss + fit_priors: gpriorba04_flat + metacal_scheme: five_types_step001_fitgauss + centroid_source: wcs + epoch_quality_and_weighting: third_masked_cut + noise_model: rms_vignet_weights + psf_epoch_loss_policy: record_and_delegate + megacam_ccd_flip: megapipe_flip + psf_diagnostics: + decisions: + starcat_merge_source: psfex + meanshape_binning: grid_5x10 + survey_geometry: + decisions: + random_sampling: per_sqdeg_50k + healpix_mask_export: nside_1024 + catalogue_assembly: + decisions: + star_galaxy_classification: deferred_downstream + tile_overlap_handling: no_dedup_in_pipeline + column_selection: committed_param_list + failure_sentinels: sentinel_values + postproc_run_provenance: newest_mtime_incremental + shape_catalogue_shortfall_guard: log_only From e901cc9c278eaa0e7900aa7f00a9b113136ad659 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 02:34:43 +0200 Subject: [PATCH 48/85] test(astra): validate decision anchors and universe pins --- tests/helpers/astra_record.py | 284 +++++++++++++++++++++++++++++++ tests/unit/test_astra_anchors.py | 50 ++++++ 2 files changed, 334 insertions(+) create mode 100644 tests/helpers/astra_record.py create mode 100644 tests/unit/test_astra_anchors.py diff --git a/tests/helpers/astra_record.py b/tests/helpers/astra_record.py new file mode 100644 index 000000000..2b8f70ed9 --- /dev/null +++ b/tests/helpers/astra_record.py @@ -0,0 +1,284 @@ +"""Reusable parsing and resolution helpers for ShapePipe's ASTRA record.""" + +import ast +import configparser +from dataclasses import dataclass +from pathlib import Path +import re + +import yaml + + +@dataclass(frozen=True) +class Anchor: + """An anchor sentence found in a YAML value.""" + + location: str + references: tuple[str, ...] + error: str | None = None + + +def load_yaml(path): + """Load YAML with PyYAML's safe loader.""" + + return yaml.safe_load(Path(path).read_text(encoding="utf-8")) + + +def _walk(value, location=""): + if isinstance(value, dict): + for key, child in value.items(): + path = f"{location}.{key}" if location else str(key) + yield from _walk(child, path) + elif isinstance(value, list): + for index, child in enumerate(value): + yield from _walk(child, f"{location}[{index}]") + else: + yield location, value + + +def _rationales(document): + for location, value in _walk(document): + if location.endswith(".rationale"): + yield location, value + + +def extract_anchors(document): + """Parse all ``Anchor:`` sentences and check every rationale has one.""" + + anchors = [] + for location, value in _walk(document): + if not isinstance(value, str) or "Anchor:" not in value: + continue + tail = value.split("Anchor:", 1)[1].strip() + error = None + refs = () + if value.count("Anchor:") != 1: + count = value.count("Anchor:") + error = f"expected one Anchor: marker, found {count}" + elif not tail.endswith("."): + error = "anchor sentence must end with a period" + else: + refs = tuple(part.strip() for part in tail[:-1].split(";")) + if not refs or any(not ref for ref in refs): + error = "anchor sentence contains an empty ref" + anchors.append(Anchor(location, refs, error)) + + for location, value in _rationales(document): + if ( + not isinstance(value, str) + or value.count("Anchor:") != 1 + or not value.rstrip().endswith(".") + ): + anchors.append( + Anchor( + location, + (), + "rationale must end with exactly one Anchor: sentence", + ) + ) + return anchors + + +def _parse_reference(reference): + if "::" in reference: + path, symbol = reference.split("::", 1) + return "code", path, symbol + if "#" in reference: + path, key = reference.split("#", 1) + return "config", path, key + return "path", reference, "" + + +def resolve_anchor(root, reference): + """Return ``None`` if a reference resolves, otherwise a diagnostic.""" + + kind, relative, selector = _parse_reference(reference) + path = Path(relative) + if path.is_absolute() or ".." in path.parts: + return "path must be relative to the repository root" + target = Path(root) / path + if not target.exists(): + return "path does not exist" + if kind == "path": + return None + if not target.is_file(): + return "code/config refs must name a file" + + try: + text = target.read_text(encoding="utf-8") + except (OSError, UnicodeError) as error: + return f"cannot read file: {error}" + + if kind == "code": + if target.suffix != ".py": + return "code-symbol refs must name a .py file" + try: + tree = ast.parse(text, filename=str(target)) + except SyntaxError as error: + return f"cannot parse Python file: {error}" + if not _has_symbol(tree, selector): + return f"no def/class/assignment target named {selector!r}" + return None + + suffix = target.suffix.lower() + if suffix == ".ini": + return _ini_key(text, selector) + if suffix == ".setools": + return _setools_key(text, selector) + if suffix in {".sex", ".psfex", ".ww", ".param", ".conf"}: + key = selector.rsplit(".", 1)[-1] + pattern = re.compile(rf"^\s*(?:#\s*)?{re.escape(key)}(?=$|\s|=|\()") + if any(pattern.search(line) for line in text.splitlines()): + return None + return f"no line starts with key {key!r} (commented keys are allowed)" + return f"unsupported config-key file type {suffix or '(no extension)'}" + + +def _ini_key(text, selector): + if "." not in selector: + return "INI config ref needs SECTION.KEY" + section, key = selector.rsplit(".", 1) + parser = configparser.ConfigParser( + interpolation=None, strict=False, allow_no_value=True + ) + try: + parser.read_string(text) + except configparser.Error as error: + return f"cannot parse INI file: {error}" + if not parser.has_section(section): + return f"INI section {section!r} is missing" + if not parser.has_option(section, key): + return f"INI key {key!r} is missing from section {section!r}" + return None + + +def _setools_key(text, selector): + if "." not in selector: + return "SETools config ref needs SECTION.KEY" + section, key = selector.rsplit(".", 1) + pattern = re.compile(rf"^\s*(?:#\s*)?{re.escape(key)}(?=$|\s|=|<|>)") + active = False + for line in text.splitlines(): + stripped = line.strip() + if stripped.startswith("[") and stripped.endswith("]"): + active = stripped[1:-1].strip() == section + elif active and pattern.search(line): + return None + return f"SETools key {key!r} is missing from section {section!r}" + + +def _target_names(target): + if isinstance(target, ast.Name): + return [target.id] + if isinstance(target, ast.Attribute): + return [target.attr] + if isinstance(target, (ast.Tuple, ast.List)): + return [name for item in target.elts for name in _target_names(item)] + if isinstance(target, ast.Starred): + return _target_names(target.value) + return [] + + +def _bindings(scope): + """Collect declarations and assignment targets in one lexical scope.""" + + result = {} + + def visit(node): + if isinstance( + node, + (ast.FunctionDef, ast.AsyncFunctionDef, ast.ClassDef), + ): + result[node.name] = node + return + if isinstance(node, ast.Lambda): + return + if isinstance(node, ast.Assign): + targets = node.targets + elif isinstance(node, (ast.AnnAssign, ast.AugAssign, ast.NamedExpr)): + targets = [node.target] + elif isinstance(node, (ast.For, ast.AsyncFor)): + targets = [node.target] + elif isinstance(node, (ast.With, ast.AsyncWith)): + targets = [item.optional_vars for item in node.items] + else: + targets = [] + for target in targets: + if target is not None: + result.update(dict.fromkeys(_target_names(target), node)) + if isinstance(node, ast.ExceptHandler) and node.name: + result[node.name] = node + for child in ast.iter_child_nodes(node): + visit(child) + + for statement in scope.body: + visit(statement) + return result + + +def _has_symbol(tree, symbol): + scope = tree + parts = symbol.split(".") + for index, part in enumerate(parts): + declaration = _bindings(scope).get(part) + if declaration is None: + return False + if index == len(parts) - 1: + return True + if not isinstance( + declaration, + (ast.ClassDef, ast.FunctionDef, ast.AsyncFunctionDef), + ): + return False + scope = declaration + return False + + +def universe_errors(record, universe): + """Check scoped decision IDs and options against the ASTRA record.""" + + record_decisions = _decisions(record) + pinned = _decisions(universe) + errors = [] + for location in sorted(pinned.keys() - record_decisions.keys()): + errors.append( + f"{location}: universe decision is absent from astra.yaml" + ) + for location in sorted(record_decisions.keys() - pinned.keys()): + errors.append( + f"{location}: astra.yaml decision is not pinned in the universe" + ) + for location in sorted(record_decisions.keys() & pinned.keys()): + definition = record_decisions[location] + options = ( + definition.get("options", {}) + if isinstance(definition, dict) + else {} + ) + if not isinstance(options, dict) or pinned[location] not in options: + errors.append( + f"{location}: pinned option {pinned[location]!r} is not in " + "ASTRA options" + ) + return errors + + +def _decisions(document, location=""): + scope = document if isinstance(document, dict) else {} + result = {} + for decision_id, definition in (scope.get("decisions") or {}).items(): + key = ( + f"{location}.decisions.{decision_id}" + if location + else f"decisions.{decision_id}" + ) + result[key] = definition + for analysis_id, analysis in (scope.get("analyses") or {}).items(): + child = ( + f"{location}.analyses.{analysis_id}" + if location + else f"analyses.{analysis_id}" + ) + if isinstance(analysis, dict): + result.update(_decisions(analysis, child)) + return result diff --git a/tests/unit/test_astra_anchors.py b/tests/unit/test_astra_anchors.py new file mode 100644 index 000000000..b4d95c2d1 --- /dev/null +++ b/tests/unit/test_astra_anchors.py @@ -0,0 +1,50 @@ +"""Keep ASTRA decision anchors and the committed universe resolvable.""" + +from pathlib import Path + +from tests.helpers.astra_record import ( + extract_anchors, + load_yaml, + resolve_anchor, + universe_errors, +) + + +REPO_ROOT = Path(__file__).resolve().parents[2] + + +def test_every_astra_anchor_resolves(): + record = load_yaml(REPO_ROOT / "astra.yaml") + anchors = extract_anchors(record) + errors = [] + + assert anchors, "astra.yaml contains no Anchor: sentences" + for anchor in anchors: + if anchor.error: + errors.append(f"{anchor.location}: {anchor.error}") + continue + for reference in anchor.references: + problem = resolve_anchor(REPO_ROOT, reference) + if problem: + errors.append( + f"{anchor.location}: {reference}: {problem}" + ) + + message = ( + "Unresolved ASTRA anchors or rationales:\n - " + + "\n - ".join(errors) + ) + assert not errors, message + + +def test_committed_universe_matches_astra_decisions(): + record = load_yaml(REPO_ROOT / "astra.yaml") + universe = load_yaml(REPO_ROOT / "universes" / "committed.yaml") + + errors = universe_errors(record, universe) + + message = ( + "ASTRA / committed universe mismatch:\n - " + + "\n - ".join(errors) + ) + assert not errors, message From 4b1e868213f9ede041609bf98a455eefb33dd3af Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 02:40:13 +0200 Subject: [PATCH 49/85] workflow: final_cat_merge is the one merger, for data and image sims #894's merge_final_cats (image sims only) and #879's final_cat_merge wrote the same product: /final_cat_.hdf5, one dataset per tile under patches/, the CONFIG_DIR final_cat.param columns. Keep final_cat_merge, which reconciles instead of appending and already reads CONFIG_DIR / "final_cat.param", so an image-sims campaign picks up config/cfis_image_sims/final_cat.param with no further wiring. Removed: the merge_final_cats rule, merge_targets(), MERGE_PATCH_DIR, MERGE_ROOT, MERGE_PATCH, MERGED_CAT, CREATE_FINAL_CAT. What changes for an image-sims campaign: the group name comes from the campaign name rather than the basename of products_dir's parent (the two agree under run_template.yaml's /product layout once the campaign is `run:`), and n_tiles_final.txt is no longer written; the tile count is the file's n_tiles attribute. sp_validation's image-sims rules do not read n_tiles_final.txt. Comments that named config/cfis/final_cat.param as the only column list, or merge_final_cats as the source of create_final_cat.py -c's path derivation, now say what holds. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_011xYGyn53XoyL83uuKPN7RQ --- scripts/python/create_final_cat.py | 5 +- workflow/README.md | 6 ++- workflow/Snakefile | 50 +------------------ .../config/cfis_image_sims/final_cat.param | 7 ++- workflow/rules/tile.smk | 6 ++- workflow/scripts/merge_final_cat.py | 7 +-- 6 files changed, 20 insertions(+), 61 deletions(-) diff --git a/scripts/python/create_final_cat.py b/scripts/python/create_final_cat.py index 39c71445a..3fe1dcf06 100755 --- a/scripts/python/create_final_cat.py +++ b/scripts/python/create_final_cat.py @@ -58,8 +58,9 @@ def params_from_run_config(params, defaults): raise ValueError(f"run config {params['run_config']} sets neither " "outputs.products_dir nor outputs.run_dir") - # The same derivation as the workflow's merge_final_cats rule: the patch - # dir is the branch dir holding product/tiles, and -i is its parent. + # The patch dir is the branch dir holding product/tiles, and -i is its + # parent. With the run template's layout (products_dir = /product) + # this is the path and group the workflow's final_cat_merge writes. patch_dir = os.path.dirname(os.path.normpath(products)) patch = os.path.basename(patch_dir) derived = { diff --git a/workflow/README.md b/workflow/README.md index 746a28bf2..bd719b575 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -98,7 +98,11 @@ outputs: ``` sp_validation's image-simulation workflow drives these campaigns and measures m -from their final catalogues. +from their final catalogues. A simulation campaign ends in the same merged +catalogue as a data campaign, written by the same `final_cat_merge` rule: +`/final_cat_.hdf5`, with the columns of +`config/cfis_image_sims/final_cat.param` and the tile count as the file's +`n_tiles` attribute (there is no `n_tiles_final.txt`). On candide, the node-local tile store (bound from the node's 31 GB `/tmp`) does not hold several dense image-sim tiles at once. Set `tile_store_root:` in the run diff --git a/workflow/Snakefile b/workflow/Snakefile index 6abd10896..cb0c0aa70 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -407,7 +407,7 @@ def path_hash(path): A rule whose behaviour comes from more than its own script needs all of it in one trigger. final_cat_merge is the case: what it writes is decided by scripts/python/create_final_cat.py (the column extraction) and by - config/cfis/final_cat.param (which columns), and NEITHER is under + CONFIG_DIR's final_cat.param (which columns), and NEITHER is under workflow/scripts/ nor a declared input. Without them in the hash, this PR's own edits to both would have left every finished campaign's hdf5 untouched and nothing would have said so. @@ -1176,54 +1176,6 @@ include: "rules/tile.smk" # adds no new grouping constraint. localrules: all, prepare_all_tiles, clean_exposure, clean_tile, exp_persist -# --- merged catalogue (image sims only) ------------------------------------ -# The sp_validation side starts from ONE hdf5 per branch, not 39 FITS tiles, so -# the campaign is not finished until they are merged. Image sims only: the data -# path's patches are merged separately, on a different naming convention. -# -# Paths are derived from products_dir rather than from `run:`, so they hold -# whatever the run config calls things. create_final_cat.py takes a root to -# scan (-i) and a patch name (-P) that is BOTH the directory to match under it -# and the hdf5 group, so the patch directory is the branch dir and the root is -# its parent -- exactly the `-i .. -P ` the old sp_validation rule ran -# from inside the branch. -MERGE_PATCH_DIR = PRODUCTS_DIR.parent # the branch dir, holding product/tiles -MERGE_ROOT = MERGE_PATCH_DIR.parent # -i -MERGE_PATCH = MERGE_PATCH_DIR.name # -P, and the hdf5 group -MERGED_CAT = PRODUCTS_DIR / f"final_cat_{MERGE_PATCH}.hdf5" -CREATE_FINAL_CAT = Path(workflow.basedir).parent / "scripts" / "python" / "create_final_cat.py" - - -def merge_targets(): - """The merged catalogue, once every ready tile has published.""" - return [str(MERGED_CAT)] if INPUT_TYPE == "image_sims" and TILES_READY else [] - - -rule merge_final_cats: - """Merge the per-tile final catalogues into one hdf5 (image sims). - - Reads the PRODUCTS root, not the run dir: clean_tile deletes the run-dir - copies, so on a reclaimed campaign the published ones are all that is left. - Ordering needs no edge to clean_tile for the same reason. - """ - input: - [final_cat(t) for t in TILES_READY], - output: - cat = str(MERGED_CAT), - n_tiles = str(PRODUCTS_DIR / "n_tiles_final.txt"), - params: - root = str(MERGE_ROOT), - patch = MERGE_PATCH, - param_file = str(CONFIG_DIR / "final_cat.param"), - threads: 1 - resources: - mem_mb = lambda wc, attempt: 8000 * attempt, - runtime = 120 - shell: - f"python {CREATE_FINAL_CAT} -I" - " -m {output.cat} -i {params.root} -p {params.param_file}" - " -P {params.patch} -o {output.n_tiles} -v" - rule all: input: diff --git a/workflow/config/cfis_image_sims/final_cat.param b/workflow/config/cfis_image_sims/final_cat.param index 5bb0bdd17..40c7c9faa 100644 --- a/workflow/config/cfis_image_sims/final_cat.param +++ b/workflow/config/cfis_image_sims/final_cat.param @@ -1,9 +1,8 @@ # Final-catalogue column selection for the image simulations. # -# create_final_cat.py -I selects exactly these columns from each simulated -# tile's make_cat output (tiles///output/run_sp_tile_Mc/...) into -# final_cat_{sim}.hdf5, and sp_validation's extract step reads the same list; -# every entry must exist in that catalogue. +# final_cat_merge selects exactly these columns from each simulated tile's +# final catalogue into /final_cat_.hdf5, and sp_validation's +# extract step reads the same list; every entry must exist in that catalogue. # # A real file in this overlay dir, not a symlink into ../cfis/, because the # real-data list names columns make_cat does not write for the simulations: diff --git a/workflow/rules/tile.smk b/workflow/rules/tile.smk index 0029f9bd1..8f2e21f24 100644 --- a/workflow/rules/tile.smk +++ b/workflow/rules/tile.smk @@ -927,9 +927,11 @@ rule clean_tile: # # THE OUTPUT SCHEMA IS AN INTERFACE, NOT A CHOICE. sp_validation opens this file # as its `galaxy_cat_path`: one dataset per tile under a named group, the -# columns of workflow/config/cfis/final_cat.param, an `n_tiles` attribute on the +# columns of CONFIG_DIR's final_cat.param, an `n_tiles` attribute on the # root. The group is named for the CAMPAIGN, which is the only unit this -# workflow has above the tile. So the rule reuses +# workflow has above the tile. It is the merger for both input types: an +# image-sims campaign reads config/cfis_image_sims/final_cat.param, whose +# columns are what sp_validation's image-sims extract step reads. So the rule reuses # scripts/python/create_final_cat.py's column extraction rather than restating # it, and writes the file itself — merge_final_cat.py argues that split, the one # legacy literal in the schema, and the two places where the reference diff --git a/workflow/scripts/merge_final_cat.py b/workflow/scripts/merge_final_cat.py index aaca2298c..0500e2609 100644 --- a/workflow/scripts/merge_final_cat.py +++ b/workflow/scripts/merge_final_cat.py @@ -4,8 +4,9 @@ Run as the shell of the campaign-level ``final_cat_merge`` rule, never by hand. WHAT IT PRODUCES, AND FOR WHOM. ``/final_cat_.hdf5``: -one dataset per tile, carrying the columns named by -``workflow/config/cfis/final_cat.param``, plus an ``n_tiles`` attribute on the +one dataset per tile, carrying the columns named by the input type's +``final_cat.param`` (``workflow/config/cfis/`` for data, +``workflow/config/cfis_image_sims/`` for image sims), plus an ``n_tiles`` attribute on the file root. sp_validation opens that file as its ``galaxy_cat_path`` (``sp_validation/catalog.py``), so its SCHEMA is an interface and not a choice — see ``SPVAL_GROUP`` below for the one legacy literal in it. @@ -141,7 +142,7 @@ def main() -> None: p.add_argument("--campaign", required=True, help="names the campaign's group in the output file") p.add_argument("--param-file", required=True, type=Path, - help="workflow/config/cfis/final_cat.param — the column list") + help="the input type's final_cat.param — the column list") p.add_argument("--hdu", type=int, default=1) p.add_argument("--snapshot-json", type=Path, default=None, help="sp run's code snapshot (bin/sp's " From d7ce47f9f5c44a16c9e35c6084f473c6eaf9e22e Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 02:42:24 +0200 Subject: [PATCH 50/85] workflow: the campaign name is `run:`, and `run:` is required #879 named the campaign `campaign:`, defaulting to products_dir's basename; #894 introduced `run:`, which `$run` expands to in the run config's paths. One name: CAMPAIGN = config["run"], and the `campaign:` key is gone from the Snakefile, config.yaml, README and the reconcile messages. The merged catalogues are final_cat_.hdf5 and full_starcat_.hdf5. run_config.REQUIRED gains `run`, so a run config whose paths never mention `$run` still refuses at parse time rather than naming the campaign's products after nothing. The shipped config.yaml no longer sets `run:` (it stays as a commented example); the run config does, and run_template.yaml already does. bin/sp: with `run:` unset, outputs.run_dir resolves to a path holding a literal `$run`, and sp created "<...>/$run-state" before the Snakefile could refuse. It now refuses first, naming `run:`, for every verb that needs a campaign; `sp container` and `sp cancel` need none and neither check nor create the state dir. The "config.yaml only" usage line goes. tests/unit/test_run_config.py pins the requirement and the shipped config's reliance on the run config. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_011xYGyn53XoyL83uuKPN7RQ --- tests/unit/test_run_config.py | 56 +++++++++++++++++++++++++++++ workflow/README.md | 25 +++++++------ workflow/Snakefile | 18 +++++----- workflow/bin/sp | 17 ++++++--- workflow/config.yaml | 7 ++-- workflow/rules/exposure.smk | 2 +- workflow/scripts/hdf5_reconcile.py | 4 +-- workflow/scripts/merge_final_cat.py | 2 +- workflow/scripts/merge_star_cat.py | 2 +- workflow/scripts/run_config.py | 5 ++- 10 files changed, 104 insertions(+), 34 deletions(-) create mode 100644 tests/unit/test_run_config.py diff --git a/tests/unit/test_run_config.py b/tests/unit/test_run_config.py new file mode 100644 index 000000000..058229097 --- /dev/null +++ b/tests/unit/test_run_config.py @@ -0,0 +1,56 @@ +"""``workflow/scripts/run_config.py``: `run:` is a required key. + +`run:` names the campaign's merged catalogues (``final_cat_.hdf5``, +``full_starcat_.hdf5``). ``unresolved()`` already reports a ``$run`` left +unexpanded in a path; these tests pin that a run config whose paths never +mention ``$run`` is refused too, and that the shipped ``config.yaml`` leaves the +name to the run config. +""" + +import importlib.util +from pathlib import Path + +import yaml + +REPO_ROOT = Path(__file__).resolve().parents[2] +SCRIPT = REPO_ROOT / "workflow" / "scripts" / "run_config.py" +CONFIG_YAML = REPO_ROOT / "workflow" / "config.yaml" + +_spec = importlib.util.spec_from_file_location("_run_config", SCRIPT) +run_config = importlib.util.module_from_spec(_spec) +_spec.loader.exec_module(run_config) + +# Every other REQUIRED key, as literal paths with no `$run` in them. +LITERAL = { + "tile_list": "/data/tiles.txt", + "inputs": {"tiles": "/data/tiles", "exposures": "/data/exp"}, + "outputs": {"run_dir": "/scratch/run", "index_db": "/data/index.sqlite"}, +} + + +def test_run_is_required(): + assert "run" in run_config.REQUIRED + + +def test_literal_paths_without_run_are_refused(): + assert run_config.unresolved(dict(LITERAL)) == ["run"] + + +def test_literal_paths_with_run_resolve(): + assert run_config.unresolved({**LITERAL, "run": "smk-g6"}) == [] + + +def test_shipped_config_leaves_run_to_the_run_config(tmp_path, monkeypatch): + monkeypatch.setenv("SP_PROFILE", "nibi") + assert "run" not in (yaml.safe_load(CONFIG_YAML.read_text()) or {}) + + no_run = tmp_path / "no_run.yaml" + no_run.write_text(yaml.safe_dump(LITERAL)) + assert "run" in run_config.unresolved( + run_config.load(CONFIG_YAML, no_run)) + + with_run = tmp_path / "with_run.yaml" + with_run.write_text(yaml.safe_dump({"run": "smk-test"})) + cfg = run_config.load(CONFIG_YAML, with_run) + assert run_config.unresolved(cfg) == [] + assert cfg["outputs"]["run_dir"].endswith("/smk-test") diff --git a/workflow/README.md b/workflow/README.md index bd719b575..9da89f8b4 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -22,16 +22,16 @@ uv venv /project/def-mjhudson/cdaley/snakemake-env --python 3.12 source /project/def-mjhudson/cdaley/snakemake-env/bin/activate uv pip install 'snakemake>=9,<10' 'snakemake-executor-plugin-slurm>=2.7,<3' -# Edit workflow/config.yaml: tile_list, inputs.tiles/exposures, outputs.run_dir, -# outputs.products_dir/index_db, and container. +# Write a run config (see Run configuration below) that sets at least `run:`, +# the campaign's name; workflow/config.yaml's machines: table supplies the rest. # `psf_model` is `psfex` or `mccd`. psfex is exercised by smk-g4 through smk-g6; mccd has run the full chain on # an image-sim star tile (one focal-plane model per exposure, ~1.5 CPU-hours each). # The committed launcher loads apptainer/1.4.5 + the /project venv, so a # fresh shell always has the right state. -workflow/bin/sp run # bring products on disk up to date with the tile list -workflow/bin/sp report # emit run_report.json now (mid-run is fine) +workflow/bin/sp run -c my_run.yaml # bring products on disk up to date with the tile list +workflow/bin/sp report -c my_run.yaml # emit run_report.json now (mid-run is fine) workflow/bin/sp cancel # scancel this workflow's jobs workflow/bin/sp container status # which image the jobs will run ``` @@ -78,13 +78,16 @@ the jobs read.) `SP_PROFILE` (default `nibi`, or `machine:` in the run config, w agree with it) and `input_type:` then select an entry of the `machines:` table, which supplies `tile_list`, `retrieve` (`symlink` or `vos`), `inputs`, `outputs` and `container` for any of these the run config leaves unset (`$base_dir` expands -to that machine's `base_dir`, `$run` to the run config's `run:`). A value of `TBD` stops the run at parse time +to that machine's `base_dir`, `$run` to the run config's `run:`). `run:` is +required: it also names the campaign's merged catalogues, and config.yaml leaves +it unset. An unset required key, or a value of `TBD`, stops the run at parse time until it is set. A run config therefore only needs what differs, e.g. for one SKiLLS shear branch on candide: ```yaml machine: candide input_type: image_sims +run: 1z2z_grid_3 psf_model: fake psf_dict: /home/hervas/fhervas/workdir_skills/input/psf_files/Full_psf_dict.pickle tile_list: /path/to/tiles.txt @@ -233,8 +236,8 @@ workflow/ container.py image layers + the resolution order behind `sp container` (stdlib-only) persist_exp.py ONE exposure's keepable PSF products -> one tar on products_dir (the exp_persist rule) hdf5_reconcile.py bring an hdf5 catalogue into agreement with a campaign (shared by both merges) - merge_star_cat.py ALL exposures' validation_psf, out of the tars -> full_starcat_.hdf5 - merge_final_cat.py ALL tiles' final_cat -> final_cat_.hdf5 (the final_cat_merge rule) + merge_star_cat.py ALL exposures' validation_psf, out of the tars -> full_starcat_.hdf5 + merge_final_cat.py ALL tiles' final_cat -> final_cat_.hdf5 (the final_cat_merge rule) clean_exposure.py ONE exposure's store + manifests + logs -> tombstone (the clean_exposure rule) profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; keep-going ``` @@ -364,7 +367,7 @@ profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; kee opens are per *campaign*, and until these rules existed each was a manual pass after the run. `star_cat_merge` collects every exposure's every CCD's `psf_validation` into - `/full_starcat_.hdf5`, one dataset per exposure at + `/full_starcat_.hdf5`, one dataset per exposure at `exposures/` — the rho/tau statistics input. It reads the members straight out of the per-exposure tars (`tarfile`; unpacking ~800k files to merge them would defeat the tar's whole purpose), keeps their native dtypes, @@ -381,13 +384,13 @@ profiles/nibi/config.yaml SLURM executor; apptainer SDM; per-user jobs cap; kee writer would just be missing from the other's product. `tests/unit/` `test_star_cat_columns.py` is what holds them together. `final_cat_merge` collects every ready tile's `final_cat-.fits` into - `/final_cat_.hdf5`: one dataset per tile under a group + `/final_cat_.hdf5`: one dataset per tile under a group named for the campaign, the `final_cat.param` columns, an `n_tiles` attribute. That schema is what sp_validation's reader opens, so it is fixed; the column extraction reuses `scripts/python/create_final_cat.py` while the file is written here, because that script's own discovery walks a directory layout - this workflow does not have. `campaign:` in `config.yaml` names the group and - defaults to the persistent root's basename. + this workflow does not have. The run config's `run:` names both files and the + group. BOTH RECONCILE, through one shared module (`hdf5_reconcile.py`) so the campaign's two products cannot disagree about what an output owes its inputs. Each adds the units that have no dataset, drops datasets whose unit left the diff --git a/workflow/Snakefile b/workflow/Snakefile index cb0c0aa70..9de25dbed 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -149,14 +149,12 @@ RUN_DIR = Path(OUTPUTS["run_dir"]) # second path: one root, exactly the pre-D5 layout. PRODUCTS_DIR = Path(OUTPUTS.get("products_dir") or RUN_DIR) INDEX_DB = Path(OUTPUTS["index_db"]) -# The campaign's NAME — what the two campaign-level merges label their output -# with (`final_cat_.hdf5`, and the group inside it that holds the -# campaign's per-tile datasets). It -# defaults to the persistent root's basename, which is already how every -# campaign here is named (smk-g4, smk-g5, smk-g6: run_dir, products_dir and -# index all end in it), so the common case needs no key at all. Set `campaign:` -# in config.yaml when the two must differ. -CAMPAIGN = config.get("campaign") or PRODUCTS_DIR.name +# The campaign's NAME is the run config's `run:` — the same name `$run` expands +# to in the paths above. The two campaign-level merges label their output with +# it (`final_cat_.hdf5` and the group inside it holding the per-tile +# datasets, `full_starcat_.hdf5`). REQUIRED in run_config.py, so an unset +# `run:` has already refused above. +CAMPAIGN = config["run"] SCRIPTS = Path(workflow.basedir) / "scripts" # The config chain is the repo's committed directory (D2). The configs and # rules that set their environment variables must be versioned together. There @@ -671,9 +669,9 @@ def exp_store_reclaimed(exp): # --- the campaign-level merges --------------------------------------------- # Two rules, one job each per campaign, both writing to the persistent root, and # both the LAST link of a chain whose per-unit half the workflow already had: -# the exposure side ends in one `full_starcat_.hdf5` (every CCD's PSF +# the exposure side ends in one `full_starcat_.hdf5` (every CCD's PSF # validation catalogue — the rho/tau statistics input) and the tile side in one -# `final_cat_.hdf5` (every tile's final catalogue — the shear +# `final_cat_.hdf5` (every tile's final catalogue — the shear # catalogue sp_validation reads). Until they existed the workflow's product set # was two files short of what the old `combine_runs.bash` + `create_final_cat.py` # chain delivered, and every campaign ended with a manual merge. diff --git a/workflow/bin/sp b/workflow/bin/sp index 80d3fcb8d..74ab37e44 100755 --- a/workflow/bin/sp +++ b/workflow/bin/sp @@ -34,7 +34,6 @@ # SP_PROFILE=candide workflow/bin/sp run -c ~/my_run.yaml # a campaign # SP_PROFILE=candide workflow/bin/sp run -c ~/my_run.yaml -n # dry run # SP_PROFILE=candide workflow/bin/sp report -c ~/my_run.yaml # status now -# workflow/bin/sp run # config.yaml only # # Run config: -c FILE, also --config-file. Read after workflow/config.yaml and # merged on top; anything still unset comes from the machines: entry for @@ -127,8 +126,17 @@ fi # Snakefile resolves it. cfg() { python "$SCRIPTS/run_config.py" "$HERE/config.yaml" "${SP_RUN_CONFIG:-}" "$1"; } RUN_DIR="$(cfg outputs.run_dir)"; INDEX_DB="$(cfg outputs.index_db)" -if [ "${1:-}" != container ] && { [ -z "$RUN_DIR" ] || [ "$RUN_DIR" = TBD ]; }; then - echo "sp: outputs.run_dir is unset or TBD for this machine/input_type" >&2; exit 2 +# `sp container` and `sp cancel` need neither a campaign nor its state dir; +# every other verb does. Checked here, before the state dir is created from +# RUN_DIR: with `run:` unset, RUN_DIR still holds a literal `$run`. +case "${1:-}" in container|cancel) NEEDS_CAMPAIGN=0 ;; *) NEEDS_CAMPAIGN=1 ;; esac +if [ "$NEEDS_CAMPAIGN" = 1 ]; then + if [ -z "$(cfg run)" ]; then + echo "sp: \`run:\` is unset; the run config (-c FILE) names the campaign" >&2; exit 2 + fi + if [ -z "$RUN_DIR" ] || [ "$RUN_DIR" = TBD ] || [[ "$RUN_DIR" == *'$'* ]]; then + echo "sp: outputs.run_dir is unset, TBD or unexpanded ($RUN_DIR) for this machine/input_type" >&2; exit 2 + fi fi # Snakemake state (.snakemake: metadata, locks, incomplete markers) lives NEXT TO @@ -141,7 +149,8 @@ fi # the placement is right on its own merits.) # --directory only moves state: all data paths are absolute, and the Snakefile # resolves its own configfile. -STATE_DIR="${SP_STATE_DIR:-${RUN_DIR}-state}"; mkdir -p "$STATE_DIR" +STATE_DIR="${SP_STATE_DIR:-${RUN_DIR}-state}" +[ "$NEEDS_CAMPAIGN" = 0 ] || mkdir -p "$STATE_DIR" # --- the launch code snapshot ---------------------------------------------- # THE ONE HOME for this concept; everything else points here. diff --git a/workflow/config.yaml b/workflow/config.yaml index c66ba6414..94b9e79c5 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -21,8 +21,9 @@ input_type: data # shadows that table. The Snakefile falls back to psfex if no entry # supplies one. -# Run name, available as `$run` in the paths below. -run: smk-g6 +# Run name, available as `$run` in the paths below; it also names the +# campaign's merged catalogues. Required, and set by the run config. +# run: smk-g6 # Entries per machine (SP_PROFILE, default nibi; implemented: nibi, candide) # and input_type. A run config may instead state `machine:`, which must @@ -78,7 +79,7 @@ machines: # # WHAT IS ALWAYS KEPT, AND IS NOT A CHOICE HERE: psf_validation, the psfex_interp # validation catalogue, one per CCD. `star_cat_merge` stacks every one of them -# into /full_starcat_.hdf5, so they are that +# into /full_starcat_.hdf5, so they are that # catalogue's PROVENANCE — a merged star catalogue with no per-exposure inputs # beside it cannot be audited, re-cut, or recomputed after a purge — and they are # what keeps APPENDING TILES CHEAP, since a tile added next month brings diff --git a/workflow/rules/exposure.smk b/workflow/rules/exposure.smk index ed3c211c1..8ff8df100 100644 --- a/workflow/rules/exposure.smk +++ b/workflow/rules/exposure.smk @@ -232,7 +232,7 @@ rule clean_exposure: # --- the campaign's star catalogue ------------------------------------------ # ONE job per campaign: every exposure's every CCD's `validation_psf--.fits`, -# collected into `/full_starcat_.hdf5`, one dataset per +# collected into `/full_starcat_.hdf5`, one dataset per # exposure. That file is the rho/tau statistics input; the old bash chain built # a flat FITS table with `combine_runs.bash psf` + a `merge_starcat_runner` # pass, and the workflow emitted neither. sp_validation still opens the FITS diff --git a/workflow/scripts/hdf5_reconcile.py b/workflow/scripts/hdf5_reconcile.py index c99101860..c57135cc4 100644 --- a/workflow/scripts/hdf5_reconcile.py +++ b/workflow/scripts/hdf5_reconcile.py @@ -160,7 +160,7 @@ def check_free_space(output: Path) -> None: def check_sole_group(output: Path, group_path: str) -> None: """One file, one campaign — refuse to half-update a file holding two. - Renaming `campaign:` mid-flight points the rule at a NEW group inside the + Renaming `run:` mid-flight points the rule at a NEW group inside the SAME file (the path carries the campaign only on the tile side, where the group does). Reconciling would then add a second group beside the first, leave the first frozen and stale, and set a count attribute describing only @@ -179,7 +179,7 @@ def check_sole_group(output: Path, group_path: str) -> None: f"hdf5_reconcile: {output} already holds {parent}/" f"{', '.join(others)} beside {group_path}. One file is one " f"campaign: reconciling would freeze the other group and count " - f"only this one. Point `campaign:` back, or write to a new path.") + f"only this one. Point `run:` back, or write to a new path.") def apply(output: Path, group_path: str, todo: Plan, units: list, read, diff --git a/workflow/scripts/merge_final_cat.py b/workflow/scripts/merge_final_cat.py index 0500e2609..624f5526e 100644 --- a/workflow/scripts/merge_final_cat.py +++ b/workflow/scripts/merge_final_cat.py @@ -3,7 +3,7 @@ Run as the shell of the campaign-level ``final_cat_merge`` rule, never by hand. -WHAT IT PRODUCES, AND FOR WHOM. ``/final_cat_.hdf5``: +WHAT IT PRODUCES, AND FOR WHOM. ``/final_cat_.hdf5``: one dataset per tile, carrying the columns named by the input type's ``final_cat.param`` (``workflow/config/cfis/`` for data, ``workflow/config/cfis_image_sims/`` for image sims), plus an ``n_tiles`` attribute on the diff --git a/workflow/scripts/merge_star_cat.py b/workflow/scripts/merge_star_cat.py index c9b4e45f1..3b0e32b38 100644 --- a/workflow/scripts/merge_star_cat.py +++ b/workflow/scripts/merge_star_cat.py @@ -3,7 +3,7 @@ Run as the shell of the campaign-level ``star_cat_merge`` rule, never by hand. -WHAT IT PRODUCES, AND FOR WHOM. ``/full_starcat_.hdf5``: +WHAT IT PRODUCES, AND FOR WHOM. ``/full_starcat_.hdf5``: one dataset per exposure at ``exposures/``, holding that exposure's every CCD's ``validation_psf--.fits`` rows stacked, with a ``CCD_NB`` column recording which CCD each row came from. It is the input to the rho/tau diff --git a/workflow/scripts/run_config.py b/workflow/scripts/run_config.py index 2925bfca5..5d97a9975 100644 --- a/workflow/scripts/run_config.py +++ b/workflow/scripts/run_config.py @@ -23,7 +23,10 @@ # here must not carry a top-level default as well. MACHINE_KEYS = ("tile_list", "retrieve", "container", "inputs", "outputs", "psf_model", "psf_dict") -REQUIRED = ("tile_list", "inputs.tiles", "inputs.exposures", +# `run` is required in its own right, not only through `$run` in the paths: +# it names the campaign's merged catalogues, so a run config whose paths +# never mention `$run` must still set it. +REQUIRED = ("run", "tile_list", "inputs.tiles", "inputs.exposures", "outputs.run_dir", "outputs.index_db") From 33863471a89d5997160a606b54aa3d7df649ed1e Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 02:43:05 +0200 Subject: [PATCH 51/85] workflow: no PSF persistence under psf_model: fake With the image-simulation true PSF, exp_psf runs SExtractor only (config_exp_fake.ini): there is no PSFEx model and no validation_psf-*.fits. exp_persist would exit on "nothing matched psf_validation" for every exposure, star_cat_merge would wait on those manifests, and clean_exposure, which takes the exp_persist manifest as an input, could never run. PERSISTS_PSF = PSF_MODEL != "fake" gates it in one place: psf_exposures(), the exposure set that persist_manifests(), star_cat_inputs() and star_cat_exposures() all draw from, is empty under fake. So persist_targets() and star_cat_targets() request nothing, and clean_exposure drops its exp_persist input. final_cat_merge still runs. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_011xYGyn53XoyL83uuKPN7RQ --- workflow/README.md | 2 ++ workflow/Snakefile | 28 ++++++++++++++++++++++------ workflow/rules/exposure.smk | 10 ++++++---- 3 files changed, 30 insertions(+), 10 deletions(-) diff --git a/workflow/README.md b/workflow/README.md index 9da89f8b4..915a15c39 100644 --- a/workflow/README.md +++ b/workflow/README.md @@ -61,6 +61,8 @@ each overlay file to its `cfis/` original confined to input naming. `psf_model: fake` is the simulations' true PSF: the exposure stage runs only SExtractor (for the background maps the vignets read), and `tile_vignets` runs `fake_interp_runner`, which writes the `galaxy_psf` product from `psf_dict`. +With no PSF model there is nothing to persist per exposure, so `exp_persist` and +`star_cat_merge` do not run and `clean_exposure` does not wait on them. Simulations that contain stars can run `psfex` or `mccd` exactly as the data do. One campaign per shear branch, each with its own run config: diff --git a/workflow/Snakefile b/workflow/Snakefile index 9de25dbed..fcbf68e82 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -571,6 +571,14 @@ def clean_targets(): # inputs and nothing else, and exp_persist still runs for every exposure. PERSIST_EXP = list(config.get("persist_exp") or []) +# Whether exposures have PSF products to persist at all. psf_model=fake (the +# image-simulation true PSF) fits no model: exp_psf runs SExtractor only and +# writes no validation_psf-*.fits, so exp_persist would find nothing to pack +# and star_cat_merge nothing to stack. Under fake both drop out of the DAG, +# and clean_exposure stops waiting on exp_persist; final_cat_merge is +# unaffected. psf_exposures() below is where this gate acts. +PERSISTS_PSF = PSF_MODEL != "fake" + # The keep list names PRODUCTS (`psf_model`), not globs (`*.psf`); the # catalogue that maps one to the other lives in persist_exp.py, which is also # what the rule runs, so there is one definition and not a copy here. @@ -623,11 +631,19 @@ def persist_manifests(): merge job's own re-parse under the slurm executor, which genuinely needs it. So this half carries no guard and the memo keeps either parse to one walk. """ - exps = {e for t in TILES_READY for e in tile_exposures(t)} - return sorted(prod_exp_manifest(e, "exp_persist") for e in exps + return sorted(prod_exp_manifest(e, "exp_persist") for e in psf_exposures() if not exp_store_reclaimed(e)) +def psf_exposures(): + """The ready tiles' exposures that have PSF products to persist and stack: + all of them, or none under psf_model=fake (PERSISTS_PSF). The one set + exp_persist's targets and star_cat_merge's inputs are drawn from.""" + if not PERSISTS_PSF: + return [] + return sorted({e for t in TILES_READY for e in tile_exposures(t)}) + + def exp_store_reclaimed(exp): """True when this exposure's PSF products exist ONLY on the persistent root. @@ -777,7 +793,7 @@ def star_cat_inputs(): VOS recovers it; the merge reports how many exposures it found. """ live, reclaimed = [], [] - for exp in sorted({e for t in TILES_READY for e in tile_exposures(t)}): + for exp in psf_exposures(): if not exp_store_reclaimed(exp): live.append(prod_exp_manifest(exp, "exp_persist")) elif Path(prod_exp_tar(exp)).exists(): @@ -797,9 +813,9 @@ def star_cat_exposures(): is what the trigger is for. It is also what merge_star_cat.py derives on the job side, so the two agree on the set AND on how it is named. """ - return sorted(e for e in {e for t in TILES_READY for e in tile_exposures(t)} - if Path(prod_exp_manifest(e, "exp_persist")).exists() - or not exp_store_reclaimed(e)) + return [e for e in psf_exposures() + if Path(prod_exp_manifest(e, "exp_persist")).exists() + or not exp_store_reclaimed(e)] # --- sizing the two merges (D4) --------------------------------------------- diff --git a/workflow/rules/exposure.smk b/workflow/rules/exposure.smk index 8ff8df100..6c3a38713 100644 --- a/workflow/rules/exposure.smk +++ b/workflow/rules/exposure.smk @@ -211,10 +211,12 @@ rule clean_exposure: # The keepers must be off /scratch before the store goes. Unlike the # consumer edges above, this edge does not depend on scope: it is the # same exposure's own rule, so it drags nothing into the DAG that this - # exposure's chain did not already put there. It is UNCONDITIONAL now: - # exp_persist always packs the star catalogue's inputs, so there is no - # keep list under which this rule has nothing to wait for. - lambda wc: [prod_exp_manifest(wc.exp, "exp_persist")] + # exposure's chain did not already put there. No keep list removes it: + # exp_persist always packs the star catalogue's inputs. Only + # psf_model=fake does, which has no PSF products to keep + # (PERSISTS_PSF, Snakefile). + lambda wc: ([prod_exp_manifest(wc.exp, "exp_persist")] + if PERSISTS_PSF else []) output: tombstone = f"{EXP_DIR}/cleaned.json" params: From 89f00b5a34f4d584c7cba2b88ec012c6161b7e9a Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 02:48:32 +0200 Subject: [PATCH 52/85] test(grammar): check the image-sims final_cat.param too final_cat_merge now reads workflow/config/cfis_image_sims/final_cat.param for image-sims campaigns, so it is a consumer contract like the data one: every NGMIX token it names must be a column make_cat's writer produces. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_011xYGyn53XoyL83uuKPN7RQ --- tests/module/test_psf_grammar_properties.py | 11 +++++++---- 1 file changed, 7 insertions(+), 4 deletions(-) diff --git a/tests/module/test_psf_grammar_properties.py b/tests/module/test_psf_grammar_properties.py index 6bf9e4713..d1b16f84a 100644 --- a/tests/module/test_psf_grammar_properties.py +++ b/tests/module/test_psf_grammar_properties.py @@ -182,10 +182,12 @@ def _run_save_ngmix(ngmix_path, obj_ids, cat_size_target=None): ) # The shipped final-catalogue param files, two levels up from tests/module/. -# Both are consumer contracts updated to the new grammar, so both are checked. +# All are consumer contracts updated to the new grammar, so all are checked; +# final_cat_merge reads the cfis_image_sims one for image-sims campaigns. _ROOT = Path(__file__).resolve().parents[2] PARAM_PATHS = [ _ROOT / "workflow" / "config" / "cfis" / "final_cat.param", + _ROOT / "workflow" / "config" / "cfis_image_sims" / "final_cat.param", _ROOT / "example" / "unions_800" / "cat_matched.param", ] @@ -357,11 +359,12 @@ def test_emitted_column_names_match_grammar(obj_ids, tmp_path_factory): def test_param_file_ngmix_tokens_are_producible(param_path, obj_ids): """Every NGMIX_* token the param file names is a column the writer produces. - Each shipped final-catalogue param file (``workflow/config/cfis/final_cat.param`` - and ``example/unions_800/cat_matched.param``) is a consumer contract for the + Each shipped final-catalogue param file (``final_cat.param`` under + ``workflow/config/cfis/`` and ``cfis_image_sims/``, and + ``example/unions_800/cat_matched.param``) is a consumer contract for the final catalogue; ``create_final_cat`` keeps only the listed columns, so a token it names that the writer cannot emit is a silent, empty column - downstream. Both files are checked so a future divergence in either (a + downstream. Every file is checked so a future divergence in any (a typo'd or stale NGMIX token) cannot escape the consistency check. The one known exception — ``NGMIX_MOM_FAIL``, the moments-failure flag set by a different path — is excluded BY NAME, and we assert it is genuinely outside From 09edd5e71bb740f9216be255b8ee1d6e640de388 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sun, 30 Aug 2026 20:25:38 -0400 Subject: [PATCH 53/85] docs(astra): record the pipeline's scientific decisions in astra.yaml MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ShapePipe's scientific choices — detection thresholds, masking geometry, star selection, PSF model, ngmix priors and seeding, flag semantics, completeness floors — live in code and committed configs with their reasoning nowhere, or spread across PRs, papers and comments. astra.yaml gathers them: 50 decisions across eight sub-analyses, each with its rationale, the alternatives that were rejected and why, and a greppable anchor back to the code or config that implements it. universes/committed.yaml pins the option this branch selects for every one. The record is ASTRA (astra-tools; `uvx astra-tools@0.2.17 guide`), applied here at codebase level rather than to a single analysis. Conventions are stated in the file's header: anchors as `path::symbol` / `path#SECTION.KEY` and never line numbers, [HARDCODED] for a scientific value with no config exposure, [LINT] for a place where the record and the code — or the code and itself — disagree, [PENDING #NNN] for state not yet on develop. Authoring it surfaced nine such lints, two of which #873 fixes, and mapped ten places where the published Guinot+22 / Farrens+22 descriptions have drifted from the code since publication; 16 decisions carry verbatim paper quotes as prior insights. CLAUDE.md gains the standing instruction: a scientific change is not finished until the record is, amended in the same PR. The membership test is whether a different defensible choice would change which objects enter the shear catalogue, or the numbers attached to them. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01Y2muA2sRojbxRNxU2SKQeP --- CLAUDE.md | 42 +- astra.yaml | 1837 ++++++++++++++++++++++++++++++++++++++ universes/committed.yaml | 70 ++ 3 files changed, 1948 insertions(+), 1 deletion(-) create mode 100644 astra.yaml create mode 100644 universes/committed.yaml diff --git a/CLAUDE.md b/CLAUDE.md index 01deff1fc..56579f03c 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -118,4 +118,44 @@ keep in their own stores outside it. A `.felt/` directory (a markdown "fiber" no store used with the `felt` CLI) is **not tracked here**: it's gitignored, and where it exists it's a machine-local symlink into a private, separately git-synced store, so a fresh clone won't have one. Record durable decisions in the PR, issue, or docs -where the change lives. +where the change lives — and *scientific* decisions in `astra.yaml`, below. + +## Scientific decisions live in `astra.yaml` + +`astra.yaml` at the repo root is the pipeline's decision record: every +consequential scientific choice embedded in the code and the committed configs, +each with its rationale, the alternatives that were considered and why they were +rejected, and an anchor back to the code or config that implements it. +`universes/committed.yaml` pins the option this branch's configuration +selects for every decision. The format +is ASTRA; `uvx astra-tools@0.2.17 guide` is the briefing and +`uvx astra-tools@0.2.17 spec` the field reference. + +**A scientific change is not finished until the record is.** When a change moves +what the pipeline measures, amend `astra.yaml` in the same PR — add the decision +if it is new, or edit its rationale, options and anchors if it moved — pin the +selected option in `universes/committed.yaml`, and say so in the PR description. +Purely technical changes (refactors, performance, packaging, I/O) leave it alone, +except where they move a value the record carries: the completeness floors in +`workflow/scripts/completeness.py` are orchestration code holding a scientific +decision. + +The membership test is whether *a different defensible choice would change which +objects enter the shear catalogue, or the numbers attached to them.* Detection +threshold and deblending contrast, masking geometry, star-selection cuts, PSF +model degree, ngmix priors and seeding, flag semantics, completeness floors — in. +Manifest sentinels, chunk sizes, allocation strategy, directory layout — out; +those live in the PR and the PRD. + +The file's own header states the conventions it follows. In short: every +rationale ends with a greppable `Anchor: path::symbol; path#SECTION.KEY` +sentence whose refs never cite line numbers; `[HARDCODED]` marks a scientific value +with no config exposure; `[LINT]` marks a place where the record and the code, or +the code and itself, disagree. Validate before committing: + +```bash +uvx astra-tools@0.2.17 validate +``` + +The record was authored against this branch's workflow configs; entries marked +`[PENDING #NNN]` describe state that has not yet reached `develop`. diff --git a/astra.yaml b/astra.yaml new file mode 100644 index 000000000..70b25b554 --- /dev/null +++ b/astra.yaml @@ -0,0 +1,1837 @@ +# ASTRA record for ShapePipe: the scientific decisions embedded in the code and +# the committed configs, with their reasoning and the alternatives that were +# rejected. It is the place scientific decisions are written down — see the +# "Scientific decisions" section of CLAUDE.md for when and how to amend it. +# +# The record describes the pipeline as orchestrated by workflow/Snakefile +# (PRD CosmoStat/shapepipe#848, PR #852). Conventions: +# +# * Decisions anchor to code, not recipes. Every rationale ends with one +# sentence "Anchor: ; ; ..." in a strict, greppable grammar. +# Each ref is a path relative to the shapepipe repo root, in one of three +# forms: CODE `path::symbol`, CONFIG `path#SECTION.KEY` (or `path#KEY` for +# sectionless .sex/.psfex/.ww/.param files), FILE `path` for a whole file +# or package. No line numbers — they rot; a line-level fact names its +# enclosing symbol. The analysis-ASTRA rule "never hardcode; reference via +# {decisions.x}" cannot hold in a codebase — the committed configs ARE the +# values. [HARDCODED] marks a scientific value living in code with no +# config exposure: the silent defaults the record exists to surface. +# * The default universe IS the committed configuration (universes/committed). +# Alternatives are excluded-with-reasons or genuinely open forks. +# * Sub-analyses follow the pipeline's methodological units — masking, +# detection, preparation, star selection + PSF, shape measurement, PSF +# diagnostics, survey geometry, catalogue assembly — not its ~20 Snakemake +# rules. Cross-cutting decisions stay top-level. A prior_insight repeated +# inside a sub-analysis carries a `_local` suffix: ids are scoped, and the +# duplicate keeps the sub-analysis readable on its own. +# * Outputs are representative product FAMILIES (one final_cat per tile), +# not enumerable artifacts; no recipes — the executor is the Snakemake +# workflow. +# * [LINT] marks places where this record and the code already disagree, or +# where the code disagrees with itself — found while authoring this file. +# * [PENDING #NNN] marks state that is live on feat/snakemake-orchestration +# — and therefore in smk-g4, the 34-tile validation campaign run under +# this branch — but not yet merged to develop. The record follows the +# branch and names the open PR. +# * A `path#KEY` anchor names the key's position in the file, not its +# activation: where the decision is "this is deliberately off", the key +# it points at may be commented out (e.g. final_cat.param#SPREAD_CLASS). + +version: "0.0.14" +name: ShapePipe scientific decisions +description: >- + Codebase-level decision record for the ShapePipe weak-lensing pipeline + (UNIONS/CFIS). Membership test: "a different defensible choice would change + which objects enter the shear catalogue, or the numbers attached to them." + Workflow mechanics that reproduce identical numbers (manifest sentinels, + clean-cascade cut, directory() outputs, allocation strategy, chunking under + position seeding) are deliberately absent; they live in the PRD and code. +tags: [shapepipe, weak-lensing, unions, codebase-record] +container: shapepipe-develop-runtime.sif + +inputs: + - id: tile_images + type: data + source: CADC-staged CFIS/UNIONS r-band tile stacks + exposure triplets (workflow/config.yaml) + description: >- + Pre-staged P3 tiles and single-exposure image/weight/flag triplets on + /project; get_images runs with RETRIEVE=symlink against this store. + - id: gsc_star_catalogue + type: data + source: GSC 2.3 (Vizier I/305/out) cone queries — scripts/python/create_star_cat.py + description: >- + Reference star catalogue driving bright-star masking. Catalogue choice, + query geometry, and magnitude handling are decisions in the masking + sub-analysis. + +outputs: + - id: final_cat + type: data + format: fits + description: >- + Per-tile shear catalogue family, the terminal science product (one per + campaign tile; make_cat_runner). Column selection and failure sentinels + are decisions in catalogue_assembly. + inputs: [tile_images] + decisions: [per_unit_count_floor, postage_stamp_size, photometric_zeropoint] + +decisions: + + # ── cross-cutting ──────────────────────────────────────────────────────── + + per_unit_count_floor: + label: Per-unit completeness policy under partial failure + rationale: >- + A 40-CCD stage where some CCDs legitimately produce nothing (sparse CCD, + setools rejects everything) cannot be all-or-nothing. The field's + converged answer (DES PSF blacklist, Rubin quantum registry) is per-unit + outcome records gated on a quality floor: record the attrition, fail + loud only below the floor, continue the survey. The floor VALUES are the + scientific content — how much silent per-CCD attrition can enter the + catalogue. The COMPLETENESS table holds them (exp_split + expect=121/floor=41, exp_mask expect=40/floor=1, psfex expect=80/floor=2, + psfex_interp floor=0 warn-only). Related leak the floor does not cover: + merge_sep_cats warns-and-skips a missing ngmix chunk, silently shrinking + a tile's shape catalogue below the floor's radar; and make_cat's own 10% + size-shortfall guard is commented out (see + catalogue_assembly.shape_catalogue_shortfall_guard). + Anchor: workflow/scripts/completeness.py::COMPLETENESS; + src/shapepipe/modules/merge_sep_cats_package/merge_sep_cats.py::MergeSep.process. + default: count_floor + options: + count_floor: + label: Count-floor table (expect/floor per runner; fail below floor) + insights: [des_psf_blacklist, guinot22_star_floor_22] + all_or_nothing: + label: Every expected sub-product required + excluded: true + excluded_reason: >- + Legitimately-absent CCDs would fail whole exposures and poison their + downstream cone; Snakemake has no optional-output primitive; field + precedent is tolerated, recorded attrition. + no_floor: + label: Accept whatever is produced, no gate + excluded: true + excluded_reason: >- + Silent attrition — a stage producing 2 of 40 CCDs would flow into + the catalogue unremarked. + + postage_stamp_size: + label: Postage-stamp size, 51 px everywhere + rationale: >- + One number pins three coupled apertures: the SExtractor vignet cut + around each detection (VIGNET(51,51) in default_noimaflags.param / + default.param, VIGNET_SIZE=51 in the dormant external-catalogue path, + example/cfis/config_tile_Uc.ini), the vignetmaker + stamps that feed ngmix (STAMP_SIZE=51 in config_tile_PiViVi.ini, both + runs; nearest-pixel centring, no sub-pixel interpolation in + VignetMaker._get_stamp), and the PSFEx model stamp (PSF_SIZE 51,51 in + default.psfex). The stamp IS the pixel data ngmix fits: it bounds + measurable galaxy size and truncates the wings of large galaxies. + Rationale for 51 not recorded in code. + Anchor: workflow/config/cfis/default_noimaflags.param#VIGNET; + example/cfis/config_tile_Uc.ini#READ_EXT_SEXCAT_RUNNER.VIGNET_SIZE; + workflow/config/cfis/config_tile_PiViVi.ini#VIGNETMAKER_RUNNER_RUN_1.STAMP_SIZE; + workflow/config/cfis/default.psfex#PSF_SIZE; + src/shapepipe/modules/vignetmaker_package/vignetmaker.py::VignetMaker._get_stamp. + default: px_51 + options: + px_51: + label: 51x51 px (~9.5 arcsec at 0.187"/px) + larger_adaptive: + label: Larger or size-adaptive stamps + excluded: true + excluded_reason: >- + Not wired; would need coupled changes in three places (a change in + any one alone desynchronises galaxy stamp, PSF stamp, and vignet). + + photometric_zeropoint: + label: Magnitude zero-point convention, fixed 30.0 on tiles + rationale: >- + Tiles use a hard-coded MAG_ZEROPOINT 30.0 for every tile + (default_tile.sex; ZP_FROM_HEADER=False in config_tile_Sx.ini), and + ngmix repeats it (MAG_ZP=30.0 in config_tile_Ng_template.ini). + Exposures instead read the per-image header zero-point + (ZP_FROM_HEADER=True, ZP_KEY=PHOTZP in config_exp_psfex.ini). The tile + convention leans on MegaPipe's calibrated stacks; the star-selection + magnitude window (18-22) and mask magnitude limits inherit whichever + convention their stage uses. SExtractorCaller.get_zero_point is the + header-reading path, unused on tiles. + Anchor: workflow/config/cfis/default_tile.sex#MAG_ZEROPOINT; + workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.ZP_FROM_HEADER; + workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.MAG_ZP; + workflow/config/cfis/config_exp_psfex.ini#SEXTRACTOR_RUNNER.ZP_KEY; + src/shapepipe/modules/sextractor_package/sextractor_script.py::SExtractorCaller.get_zero_point. + default: fixed_30_tiles_header_exposures + options: + fixed_30_tiles_header_exposures: + label: Tiles fixed 30.0; exposures from header PHOTZP + header_everywhere: + label: Per-image header zero-points on tiles too + excluded: true + excluded_reason: >- + MegaPipe stacks are calibrated to ZP 30 by construction; per-tile + header reads add a failure path for no expected numerical change. + (If that claim is wrong, this is a real fork — verify.) + + baseline_validation_criterion: + label: Validation criterion against the v2.0 bash baseline + rationale: >- + Because shape_measurement.ngmix_seed_mode deliberately changes noise + streams, P1 validation against v2.0 is statistical parity + (population-level agreement), not bit parity. Everything upstream of + ngmix (through PSFEx) validated bit-exactly (P0: 4/4 PASS). This + defines the evidence standard for "the same pipeline" — surfaced to the + collaboration as open Q5 in PRD #848. + [PENDING #873] Run-to-run determinism, which is a different property + from parity with v2.0, is now complete. With the setools star split + seeded (star_selection_psf.psf_train_validation_split) the last unseeded + draw in the science chain is gone: two runs of this code over the same + inputs now produce the same PSF star sample, the same PSF models and + the same shapes, which they did not before. That also settles a tension + this record carried — the bit-parity claim above sat next to an + unseeded star split that could not have been bit-reproducible, and the + P0 exposure-stage comparison did see PSF-validation CCD attrition + differ between the two sides. Statistical rather than bit parity is + therefore demanded only against the v2.0 baseline, not between runs of + the current pipeline. + Anchor: workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.SEED_FROM_POSITION; + src/shapepipe/modules/ngmix_package/ngmix.py::position_seed; + src/shapepipe/modules/setools_package/setools.py::SETools._make_rand_split. + default: statistical_parity + options: + statistical_parity: + label: Population-level agreement in shear observables + bit_parity: + label: Bit-identical catalogues + excluded: true + excluded_reason: >- + Impossible by construction once the seed mode changed; requiring it + would freeze the chunk-dependent v2.0 RNG forever. + +prior_insights: + des_psf_blacklist: + claim: >- + DES enters a CCD's PSF model into a blacklist rather than failing the + exposure - in Y3, any CCD with fewer than 25 stars surviving outlier + rejection is blacklisted and excluded downstream (~2% of data removed), + and processing proceeds. + created_at: "2026-07-16T00:00:00Z" + evidence: + - id: ev_jarvis_y3 + doi: "10.48550/arXiv.2011.03409" + quote: + exact: "we enter it into a" + suffix: " \u201cblacklist\u201d and exclude this CCD" + location: { page: 10 } + guinot22_star_floor_22: + claim: >- + The published ShapePipe/UNIONS analysis applies a per-CCD quality floor + rather than failing whole exposures: a CCD with fewer than 22 selected + stars is discarded for PSF estimation and contributes no epoch to the + shape measurement, while processing continues. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_star_floor + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'The dashed line represents the cut at 22 stars/CCD below which the CCD is discarded for the PSF estimation.' + location: { page: 4 } + +findings: + orchestration_parity: + claim: >- + The Snakemake orchestration reproduces the bash baseline bit-exactly + through PSFEx (P0 validation, 4/4 PASS on the 186/187 quad). + Read it with two caveats. It is a statement about the pre-#873 code: + both #873 changes move products (a different realised star split, a + stricter science-path star gate), so re-establishing parity would mean + regenerating the baseline under the current branch. And the parity is + bit-exact in the products compared, not everywhere: the P0 + exposure-stage comparison did see PSF-validation CCD attrition differ + between the two sides, which the then-unseeded star split explains + (see baseline_validation_criterion). + created_at: "2026-08-19T00:00:00Z" + evidence: + - id: ev_final_cat + artifact: final_cat + record_authoring_found_defects: + claim: >- + Nine places in the code disagree with themselves or with their + documentation, each carried as a [LINT] mark: the 22-vs-20 + STAR_THRESH mismatch between PSF validation and science interpolation; + the unseeded train/validation rand_split (setools.py:664 — the star + sample entering the PSF model is irreproducible run-to-run); additive + (non-bitwise) mask-plane combination, safe today only because the + committed flag values are disjoint; the dead MESSIER_PIXEL_SCALE config + key; final_cat.param requesting IMAFLAGS_ISO that the merged catalogue + never receives; the centroid_source default disagreement (runner "wcs" + vs module "hsm", latent for direct callers); setools logging a FWHM cut + (mode +- 0.1 px in arcsec) half the applied one (mode +- 0.2 px), and + mixing pixel scales 0.187/0.186 within one file; TILE_LIST + overlap-flagging documented but never implemented; the mccd_plots + module docstring advertising rho statistics that live downstream now. + Status: the first two are fixed. CosmoStat/shapepipe#873 seeds the + rand_split and raises the science-path STAR_THRESH to 22, and commit + 90782098 mirrors that threshold into the workflow's own committed + config fork. #873 is OPEN against develop; both fixes are live on + feat/snakemake-orchestration only, and the 34-tile smk-g4 campaign is + the first run under them. The other seven stand, including the + mccd_plots docstring that still advertises rho statistics the package + no longer computes. + created_at: "2026-08-29T00:00:00Z" + derived: true + evidence: + - id: ev_final_cat_defects + artifact: final_cat + code_paper_divergence: + claim: >- + The two ShapePipe papers state roughly 17 of this record's 50 decisions + (now carried as prior_insights with verbatim quotes), have drifted from + the code on 10 of them since publication, and are silent on the rest. + Among them: DETECT_MINAREA 10 -> 5; DEBLEND_MINCONT 0.001 -> + 0.0005 on tiles; tile background AUTO -> MANUAL 0; in-line spread-model + star/galaxy classification -> disabled and deferred downstream; HSM + moment initialisation -> WCS centroids and prior-based guesses; GSC 2.2 + via cdsclient -> GSC 2.3 via astroquery; PSF acceptance 22 stars/CCD + published for the science path vs 20 committed there — closed since by + #873 + 90782098, which put the science path on 22, so nine of the ten + drifts remain open on the orchestration branch. + created_at: "2026-08-29T00:00:00Z" + derived: true + evidence: + - id: ev_final_cat_divergence + artifact: final_cat + +analyses: + + # ═════════════════════════════════════════════════════════════════════════ + masking: + description: >- + Which pixels are excluded before anything is measured. Modules: + src/shapepipe/modules/mask_package/mask.py (halo/spike/DSO/border + builders, WeightWatcher driver), scripts/python/create_star_cat.py + (star-catalogue fetch), configs config_exp_Ma.ini + + config_onthefly.mask / config_tile_onthefly.mask + mask_default/. + [LINT] MESSIER_PIXEL_SCALE is set in config_tile_onthefly.mask but + never read — mask_dso takes pixel scale from the WCS. [LINT] + _build_final_mask combines mask planes by ADDITION (mask.py:1141+), + not bitwise OR; the committed flag values (2/4/16/32/128) are disjoint + so no live collision exists, but any future duplicate value corrupts + the flag semantics silently. (FLAG_OUTFLAGS 2 in default.ww is inert: + no input flag image is passed to WeightWatcher — mask.py:1047-1074.) + inputs: + - id: ccd_images + type: data + source: split per-CCD exposure images + weights + CFIS flag maps (exp_split family) + - id: star_catalogue + type: data + source: GSC 2.3 per-exposure catalogues (exp_star_cat cache) + outputs: + - id: exposure_mask + type: data + format: fits + description: Per-CCD pipeline flag maps (run_sp_exp_Ma family). + decisions: + [star_catalogue_query, star_magnitude_definition, + bright_star_mask_geometry, deep_sky_object_masking, + border_mask_width, pixel_threshold_flags, external_flag_usage] + decisions: + star_catalogue_query: + label: Reference star catalogue and query for bright-star masking + rationale: >- + GSC 2.3 (Vizier I/305/out), columns GSC2.3/RAJ2000/DEJ2000/Fmag/ + jmag/Vmag/Nmag/Class, cone radius covering the full CCD mosaic, no + magnitude cut at query time. GSC 2.2 rejected in a code comment + beside the catalogue ID ("does not have Fmag"). + Provenance hazard: the star-cat cache is not keyed by script + version — a semantic change to this query reruns the rule but takes + the skip-if-exists branch; clear the cache by hand for the change + to reach the data (workflow/config.yaml star_cats comment). + Query geometry: search radius = half the image diagonal about the + field centre (Mask._get_image_radius); source precedence: with + CDSCLIENT_PATH set in the .mask configs the online-query branch + wins unless an external star cat is passed (USE_EXT_STAR=True in + config_exp_Ma.ini routes the exp_star_cat cache in). + Published description (Farrens+22 p.2): cdsclient downloads GSC 2.2 + (with cdsclient 3.84 pinned in its Table A.1); current code: GSC 2.3 + queried through astroquery — two things drifted, the catalogue + version (Fmag is needed for the magnitude cut) and the query + transport, since cdsclient is never invoked yet survives as a + required-but-unused CDSCLIENT_PATH still set to the stale + findgsc2.2 in config_tile_onthefly.mask. + Anchor: scripts/python/create_star_cat.py::CDS_CAT_ID; + src/shapepipe/modules/mask_package/mask.py::Mask._CDS_cat_ID; + src/shapepipe/modules/mask_package/mask.py::Mask._cds_keys; + workflow/config.yaml. + default: gsc_23_vizier + options: + gsc_23_vizier: + label: GSC 2.3 cone queries, all bands, no query-time mag cut + insights: [farrens22_star_cat_on_disk] + gaia: + label: Gaia-based star catalogue + excluded: true + excluded_reason: >- + Not wired. Deeper and better photometry; switching changes mask + geometry and hence the selection function — a real DR-level fork. + star_magnitude_definition: + label: Per-star magnitude for mask scaling + rationale: >- + mag = unweighted mean of the finite GSC bands among F, j, V, N; + stars with no finite band are logged and not masked; only Class==0 + objects masked. Comment records why not a naive mean: NaN bands + would NaN-poison the mag < mag_limit test, leaving exactly the + bright stars with incomplete photometry unmasked. + Anchor: src/shapepipe/modules/mask_package/mask.py::Mask._create_mask. + default: mean_finite_bands + options: + mean_finite_bands: { label: Mean of finite F/j/V/N; Class==0 only } + single_band: + label: Single-band (Fmag) magnitude + excluded: true + excluded_reason: Drops stars with missing Fmag from masking entirely. + bright_star_mask_geometry: + label: Halo + diffraction-spike mask geometry and magnitude scaling + rationale: >- + DS9 polygon templates scaled linearly with magnitude about a pivot: + halo HALO_MAG_LIM=13, HALO_SCALE_FACTOR=0.05, HALO_MAG_PIVOT=13.8 + (halo_mask.reg, ~270 px); spike SPIKE_MAG_LIM=18, + SPIKE_SCALE_FACTOR=0.3, SPIKE_MAG_PIVOT=13.8 + (MEGAPRIME_star_i_13.8.reg); scaling = 1 - factor*(mag-pivot), + floored at 0.1 by Mask._scaling_min. + Identical in exposure and tile configs. Template filename encodes + provenance (MegaPrime i-band mag-13.8 star); numeric rationale not + recorded. Note the 5-mag gap: stars in 13-18 get spikes but no halo. + Anchor: workflow/config/cfis/config_onthefly.mask#HALO_PARAMETERS.HALO_MAG_LIM; + workflow/config/cfis/config_onthefly.mask#SPIKE_PARAMETERS.SPIKE_MAG_LIM; + workflow/config/cfis/mask_default/halo_mask.reg; + workflow/config/cfis/mask_default/MEGAPRIME_star_i_13.8.reg; + src/shapepipe/modules/mask_package/mask.py::Mask._create_mask; + src/shapepipe/modules/mask_package/mask.py::Mask._scaling_min. + default: megaprime_polygon_linear_scaling + options: + megaprime_polygon_linear_scaling: + label: Fixed MegaPrime templates, linear mag scaling, floor 0.1 + radial_profile_fit: + label: Per-star radial-profile-driven mask size + excluded: true + excluded_reason: Not wired; the survey precedent is template-based. + deep_sky_object_masking: + label: Messier + NGC objects masked as circles, no enlargement + rationale: >- + Circles of radius max(size_X, size_Y), MESSIER_SIZE_PLUS=0, + NGC_SIZE_PLUS=0 (function default is 0.1 — the 0 is a choice); + flags 16/32. A comment records the overlap-test fix (corner-only + test missed small interior objects). + Anchor: workflow/config/cfis/config_onthefly.mask#MESSIER_PARAMETERS.MESSIER_SIZE_PLUS; + workflow/config/cfis/config_tile_onthefly.mask#NGC_PARAMETERS.NGC_SIZE_PLUS; + src/shapepipe/modules/mask_package/mask.py::Mask.mask_dso. + default: circles_no_padding + options: + circles_no_padding: + label: "size_plus = 0: mask exactly the catalogued extent" + insights: [farrens22_messier_mask] + padded_circles: + label: size_plus > 0 (code default 0.1) + excluded: true + excluded_reason: >- + Rationale for dropping the padding not recorded; flagged as a + question rather than an endorsed exclusion. + border_mask_width: + label: CCD border mask, 50 px on exposures, none on tiles + rationale: >- + Exposures BORDER_WIDTH=50 (flag 4); tiles BORDER_MAKE=False. + Mask.mask_border's own default is 100 — the committed 50 is a + choice, unrecorded. Trims CCD edges where PSF and astrometry + degrade; changes the effective footprint. + Anchor: workflow/config/cfis/config_onthefly.mask#BORDER_PARAMETERS.BORDER_WIDTH; + workflow/config/cfis/config_tile_onthefly.mask#BORDER_PARAMETERS.BORDER_MAKE; + src/shapepipe/modules/mask_package/mask.py::Mask.mask_border. + default: px50_exposures_only + options: + px50_exposures_only: + label: 50 px exposure borders; tiles unmasked + insights: [farrens22_border_mask] + px100: + label: 100 px (module default) + excluded: true + excluded_reason: Halves usable edge area for no recorded gain. + pixel_threshold_flags: + label: WeightWatcher weight/flag thresholds into mask bits + rationale: >- + WEIGHT_MIN 0, WEIGHT_MAX 1000, WEIGHT_OUTFLAGS 1; FLAG_MASKS 0x01, + FLAG_OUTFLAGS 2; POLY_OUTWEIGHTS 0. Zero-weight and externally + flagged pixels excluded on these thresholds. Values are stock, not + derived from the CFIS weight distribution; rationale not recorded. + The FLAG_* keys are inert in the committed invocation — no flag + image is passed to WeightWatcher by Mask._exec_WW. + Anchor: workflow/config/cfis/mask_default/default.ww#WEIGHT_MIN; + workflow/config/cfis/mask_default/default.ww#FLAG_MASKS; + src/shapepipe/modules/mask_package/mask.py::Mask._exec_WW. + default: stock_ww_thresholds + options: + stock_ww_thresholds: + label: Stock WeightWatcher thresholds + insights: [farrens22_weightwatcher] + external_flag_usage: + label: CFIS external flag maps folded into exposure masks + rationale: >- + USE_EXT_FLAG=True on exposures (imports CADC-provided bad-pixel / + cosmic-ray / trail flags); EF_MAKE=False on tiles. The external + plane enters via Mask._build_final_mask's path_external_flag branch. + Anchor: workflow/config/cfis/config_exp_Ma.ini#MASK_RUNNER.USE_EXT_FLAG; + workflow/config/cfis/config_tile_onthefly.mask#EXTERNAL_FLAG.EF_MAKE; + src/shapepipe/modules/mask_package/mask.py::Mask._build_final_mask. + default: exposures_only + options: + exposures_only: { label: "External flags on exposures, not tiles" } + ignore_external: + label: Pipeline-generated masks only + excluded: true + excluded_reason: Discards upstream knowledge of bad pixels. + prior_insights: + farrens22_star_cat_on_disk: + claim: >- + The ShapePipe release paper documents an on-disk star catalogue, in + GSC format, as a supported substitute for the online query, + motivated by compute nodes without internet access. + created_at: "2022-06-01T00:00:00Z" + evidence: + - id: ev_farrens22_star_cat_disk + doi: "10.48550/arXiv.2206.14689" + quote: + exact: 'Alternatively, a star catalogue available on disk (with the same format as the GSC) can also be used' + location: { page: 2 } + farrens22_messier_mask: + claim: >- + Messier objects are named in the published masking procedure as one + of the object classes ShapePipe masks. + created_at: "2022-06-01T00:00:00Z" + evidence: + - id: ev_farrens22_messier + doi: "10.48550/arXiv.2206.14689" + quote: + exact: 'Messier objects, and border regions.' + location: { page: 2 } + farrens22_border_mask: + claim: >- + CCD border regions are named in the published masking procedure as + one of the regions ShapePipe masks. + created_at: "2022-06-01T00:00:00Z" + evidence: + - id: ev_farrens22_border + doi: "10.48550/arXiv.2206.14689" + quote: + exact: 'Messier objects, and border regions.' + location: { page: 2 } + farrens22_weightwatcher: + claim: >- + The published pipeline generates the mask image itself with + WeightWatcher (Marmo & Bertin 2008), fixing the tool but none of its + threshold values. + created_at: "2022-06-01T00:00:00Z" + evidence: + - id: ev_farrens22_ww + doi: "10.48550/arXiv.2206.14689" + location: { page: 2 } + + # ═════════════════════════════════════════════════════════════════════════ + detection: + description: >- + Object detection on r-band tiles (single-image mode) and exposures (for + star finding). Module: src/shapepipe/modules/sextractor_package/ + sextractor_script.py (config assembly, ZP/background overrides, + post-processing that assigns per-epoch CCD membership). Configs: + config_tile_Sx.ini + default_tile.sex + default.conv + + default_noimaflags.param (tiles); default_exp.sex (exposures — same + thresholds, but DEBLEND_MINCONT 0.001 vs tile 0.0005 and BACK_TYPE AUTO + vs tile MANUAL 0, both deliberate and unexplained divergences). + [LINT] final_cat.param (consumed by the post-proc merge_final_cat, + not by make_cat) requests IMAFLAGS_ISO, but the tile chain never + produces it (FLAG_IMAGE=False, default_noimaflags.param); the + exposure-side IMAFLAGS_ISO stays exposure-side (merge_starcat.py:807 + only). The merged catalogue never receives the column. + inputs: + - id: tile_stack + type: data + source: MegaPipe r-band tile stack + weight (uncompressed, merged headers) + outputs: + - id: tile_sexcat + type: data + format: fits + description: Per-tile SExtractor LDAC catalogue with per-epoch CCD membership. + decisions: + [detection_threshold_policy, deblending_policy, background_model, + weighting_and_interpolation, detection_source_mode, + epoch_membership_ccd_bounds, photometry_parameters, + cleaning_and_neighbour_masking] + decisions: + photometry_parameters: + label: Photometric aperture definitions — Kron parameters, apertures, half-light fraction + rationale: >- + PHOT_AUTOPARAMS 2.5,3.5 (Kron factor / minimum radius), + PHOT_APERTURES 5 px, PHOT_FLUXFRAC 0.5, BACKPHOTO_TYPE GLOBAL — + identical in both .sex files. MAG_AUTO is the axis of the + star-selection magnitude box AND the catalogue magnitude; FLUX_AUTO + is PSFEx's photometric normalisation (default.psfex PHOTFLUX_KEY). + A different Kron factor shifts magnitudes systematically, moving + which stars build the PSF model and every magnitude-based + downstream cut. Rationale not recorded (stock values). Anchor: + workflow/config/cfis/default_tile.sex#PHOT_AUTOPARAMS; + workflow/config/cfis/default_exp.sex#PHOT_AUTOPARAMS; + workflow/config/cfis/default.psfex#PHOTFLUX_KEY. + default: kron_25_35 + options: + kron_25_35: { label: "Kron 2.5/3.5, aperture 5 px, FLUXFRAC 0.5, global background" } + cleaning_and_neighbour_masking: + label: Spurious-detection cleaning and neighbour-pixel correction + rationale: >- + CLEAN Y with CLEAN_PARAM 1.0 deletes detections consistent with + being wings of a brighter neighbour — a post-deblend change to the + object list; MASK_TYPE CORRECT replaces neighbour pixels during + photometry (vs BLANK/NONE), changing fluxes and windowed moments + of blends. Identical in both .sex files; rationale not recorded. + Anchor: workflow/config/cfis/default_tile.sex#CLEAN; + workflow/config/cfis/default_tile.sex#MASK_TYPE; + workflow/config/cfis/default_exp.sex#CLEAN. + default: clean_1_correct + options: + clean_1_correct: { label: "CLEAN 1.0 + MASK_TYPE CORRECT" } + detection_threshold_policy: + label: Detection significance, minimum area, matched filter + rationale: >- + DETECT_THRESH 1.5 sigma RELATIVE, ANALYSIS_THRESH 1.5, + DETECT_MINAREA 5, FILTER default.conv (3x3 pyramid kernel, "all + ground, FWHM = 2 pixels" — vs CFIS seeing ~0.65 arcsec = 3.5 px at + 0.187"/px, so the filter is not matched to the survey PSF). + Sets the faint end of the source sample. Rationale not recorded + (stock EB 2017 header). + Published description (Guinot+22 p.5, Table 2): DETECT_MINAREA 10; + current code: 5, in default_exp.sex as well as default_tile.sex — + the small-object end has been loosened since publication on both the + star-detection and tile-detection passes, while DETECT_THRESH 1.5 + RELATIVE and the 3x3 FWHM=2 px kernel still match. + Anchor: workflow/config/cfis/default_tile.sex#DETECT_THRESH; + workflow/config/cfis/default_tile.sex#DETECT_MINAREA; + workflow/config/cfis/default.conv. + default: thresh_1p5_minarea5_fwhm2px_filter + options: + thresh_1p5_minarea5_fwhm2px_filter: + label: 1.5 sigma, minarea 5, FWHM=2px kernel + seeing_matched_filter: + label: Kernel matched to CFIS seeing (~3.5 px) + excluded: true + excluded_reason: >- + Not wired; would change depth and the faint-end selection + function — a real fork, excluded only as not-the-committed-path. + deblending_policy: + label: Deblending sub-thresholds and contrast + rationale: >- + DEBLEND_NTHRESH 32, DEBLEND_MINCONT 0.0005 on tiles (2x more + aggressive splitting than the exposure 0.001 and 10x more than the + SExtractor default 0.005 — divergences not documented), CLEAN Y + PARAM 1.0. Controls object count, centroids, and blend + contamination in shapes. + Published description (Guinot+22 p.5, Table 2, galaxy detection on + the stacked tiles): DEBLEND_MINCONT 0.001; current code: 0.0005 on + tiles, with only the exposure side still carrying 0.001 — the + divergence lands on precisely the configuration the paper documents. + NTHRESH 32 matches. + Anchor: workflow/config/cfis/default_tile.sex#DEBLEND_MINCONT; + workflow/config/cfis/default_exp.sex#DEBLEND_MINCONT. + default: mincont_5em4_tiles + options: + mincont_5em4_tiles: { label: "MINCONT 0.0005 tiles / 0.001 exposures" } + background_model: + label: Tile background fixed to zero, not estimated + rationale: >- + BACK_TYPE MANUAL, BACK_VALUE 0.0, BKG_FROM_HEADER=False on tiles — + trusts MegaPipe stack background removal; exposures use BACK_TYPE + AUTO (64/3 mesh). Residual sky offsets propagate into thresholds, + fluxes, completeness. Divergence deliberate, unexplained. + Published description (Guinot+22 p.5): Table 2's caption asserts all + non-tabulated SExtractor parameters keep their defaults, i.e. + BACK_TYPE AUTO, and the paper never mentions the background choice + at all; current code: BACK_TYPE MANUAL with BACK_VALUE 0.0 on tiles, + identically in workflow/ and example/ — the standing tile + configuration, not a one-off, diverging silently from the published + parametrisation. + Anchor: workflow/config/cfis/default_tile.sex#BACK_TYPE; + workflow/config/cfis/default_exp.sex#BACK_TYPE; + workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.BKG_FROM_HEADER; + src/shapepipe/modules/sextractor_package/sextractor_script.py::SExtractorCaller.get_background. + default: manual_zero_tiles_auto_exposures + options: + manual_zero_tiles_auto_exposures: + label: Tiles trust the stack (0.0); exposures estimate + auto_everywhere: + label: SExtractor AUTO background on tiles too + excluded: true + excluded_reason: >- + Double-subtracts if MegaPipe already removed it; if MegaPipe + residuals are nonzero this exclusion is wrong — verify. + weighting_and_interpolation: + label: Weight-map usage and zero-weight pixel interpolation + rationale: >- + Two settings depart from stock SExtractor: WEIGHT_TYPE MAP_WEIGHT + (default NONE) and INTERP_TYPE ALL (default NONE — SExtractor + invents flux across zero-weight pixels). The accompanying + RESCALE_WEIGHTS Y, WEIGHT_GAIN Y, MASK_TYPE CORRECT and + INTERP_MAXXLAG/INTERP_MAXYLAG 16 are the SExtractor defaults, so + they are settings the configs restate rather than choices. The + variance policy sets effective per-pixel SNR and thus the detection + set; INTERP_TYPE ALL alters pixel data feeding measurements. + Rationale not recorded. + Published description (Guinot+22 p.5): Table 2's "all other + parameters are kept to their default values" silently covers both + non-default settings; current code: MAP_WEIGHT + INTERP_TYPE ALL in + default_tile.sex and default_exp.sex alike — the paper gives no hint + that the weight map or the zero-weight interpolation is in play. + Anchor: workflow/config/cfis/default_tile.sex#WEIGHT_TYPE; + workflow/config/cfis/default_tile.sex#INTERP_TYPE; + src/shapepipe/modules/sextractor_package/sextractor_script.py::SExtractorCaller.set_input_files. + default: map_weight_interp_all + options: + map_weight_interp_all: { label: MAP_WEIGHT + INTERP ALL + MASK CORRECT } + no_interpolation: + label: INTERP_TYPE NONE + excluded: true + excluded_reason: >- + Changes photometry near masks; the committed choice is itself + unjustified in code — flagged as a question, not an endorsement. + detection_source_mode: + label: Single-image detection on the r-band tile + rationale: >- + DETECTION_IMAGE=False, FLAG_IMAGE=False at detection, + param file default_noimaflags.param. No dual-image mode, no + detection coadd, no flag propagation at detection time. The + sx_nomask variant is the committed chain because it matches the + validated bash baseline. STATUS: DELIBERATELY UNDECIDED (Cail, + 2026-08-29) — whether DR6 detects masked or unmasked is punted to + the planned masking-unification rework (not yet tracked in an issue); + the default records baseline + parity, not a settled methodological choice. The masked variant is + one config + one rule + a tile-side star-cat analogue away. + Anchor: workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.DETECTION_IMAGE; + workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.FLAG_IMAGE; + workflow/config/cfis/default_noimaflags.param; + workflow/rules/tile.smk. + default: sx_nomask_single_image + options: + sx_nomask_single_image: + label: Unmasked single-image r-band detection + insights: [guinot22_stacked_detection] + sx_masked: + label: Detection on the masked tile + dual_image_coadd: + label: Dual-image mode with a detection coadd + excluded: true + excluded_reason: No detection coadd exists in UNIONS r-band processing. + epoch_membership_ccd_bounds: + label: Which exposure CCDs an object belongs to (N_EPOCH) + rationale: >- + CCD_SIZE = 33,2080,1,4612 with strict inequalities — the 33-px left + trim silently discards a CCD strip from epoch membership; WCS + inversion failures skip the CCD ("no epoch recorded"), changing + N_EPOCH. Sets how many exposures contribute to each galaxy's + multi-epoch fit. Rationale beyond "number of pixels in a CCD" not + recorded. + Anchor: workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.CCD_SIZE; + src/shapepipe/modules/sextractor_package/sextractor_script.py::make_post_process; + src/shapepipe/modules/sextractor_package/sextractor_script.py::ccd_candidate_mask. + default: trimmed_bounds_33_2080 + options: + trimmed_bounds_33_2080: { label: "x in (33,2080), y in (1,4612), strict" } + full_ccd: + label: Full 1-2048 x-range, inclusive bounds + excluded: true + excluded_reason: >- + The trim presumably excludes a bad edge region, but nothing in + code says so — flagged as a question. + prior_insights: + guinot22_stacked_detection: + claim: >- + Source extraction in the published analysis is performed on the + stacked tile images, for signal-to-noise and because artefacts are + suppressed relative to single exposures. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_stacked_detection + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'We do the extraction on stacked images which provide a better signal-to-noise ratio, and most artifacts have a reduced amplitude with respect to single exposures' + location: { page: 5 } + + # ═════════════════════════════════════════════════════════════════════════ + preparation: + description: >- + How pixels, WCS, and epoch membership are prepared before anything is + measured. Modules: + src/shapepipe/modules/split_exp_package/split_exp.py, + merge_headers_package/merge_headers.py, + find_exposures_package/find_exposures.py, + vignetmaker_package/vignetmaker.py. + inputs: + - id: exposure_files + type: data + source: delivered CFIS exposure triplets (image/weight/flag MEF) + tile stacks + outputs: + - id: epoch_stamps + type: data + format: fits + description: Per-object multi-epoch vignets + per-CCD WCS log feeding ngmix. + decisions: + [astrometric_solution_source, ccd_split_extent, + epoch_provenance_from_tile_history, object_position_columns, + stamp_positioning_and_padding, epoch_flag_source] + decisions: + astrometric_solution_source: + label: Astrometry taken verbatim from delivered per-CCD headers + rationale: >- + split_exp builds WCS(h) from each raw CCD header at split time, + pickles it, and merge_headers writes the lot into + log_exp_headers.sqlite; every downstream world<->pixel transform + (stamp positioning, epoch membership, position seeding) uses that + stored solution. No re-derivation, no astrometric refinement — the + survey's delivered astrometry IS the pipeline's astrometry. + Alternative (a joint astrometric re-fit a la DES/Rubin) would move + every stamp centre and every position seed. Anchor: + src/shapepipe/modules/split_exp_package/split_exp.py::SplitExposures.create_hdus; + src/shapepipe/modules/merge_headers_package/merge_headers.py::merge_headers; + src/shapepipe/modules/vignetmaker_package/vignetmaker.py::VignetMaker._get_stamp_me. + default: delivered_headers + options: + delivered_headers: + label: "WCS(header) verbatim, stored at split time" + insights: [guinot22_gaia_astrometry] + astrometric_refit: + label: Joint astrometric re-solution + excluded: true + excluded_reason: Not wired; CFIS delivered astrometry is trusted. + ccd_split_extent: + label: All 40 MegaCam HDUs split and carried as candidate epochs + rationale: >- + N_HDU=40 with a hard check (any other HDU count raises) — every + CCD including the ear CCDs 36-39 is a candidate epoch wherever the + WCS lands it. The MegaCamFlip special-casing of 36/37 shows the + ears flow through shape measurement. Alternative: exclude ear CCDs + (different optical path/orientation history). Anchor: + workflow/config/cfis/config_exp_Sp.ini#SPLIT_EXP_RUNNER.N_HDU; + src/shapepipe/modules/split_exp_package/split_exp.py::SplitExposures.create_hdus. + default: all_40_hdus + options: + all_40_hdus: + label: "40 HDUs, hard-fail on any other count" + insights: [guinot22_forty_chips] + epoch_provenance_from_tile_history: + label: Epoch sets parsed from tile FITS HISTORY cards + rationale: >- + A tile's contributing exposures are recovered by parsing column 3 + of each HISTORY line, stripping prefix "p", deduplicating — the + coadd's own provenance record is trusted as the epoch list. The + LSB s-prefix rename is present but commented out. A mis-parse + changes N_EPOCH and which exposures are fit. Anchor: + workflow/config/cfis/config_tile_Fe.ini#FIND_EXPOSURES_RUNNER.COLNUM; + src/shapepipe/modules/find_exposures_package/find_exposures.py::FindExposures.get_exposure_list. + default: history_parse + options: + history_parse: { label: "HISTORY column 3, prefix p, dedup" } + object_position_columns: + label: Windowed centroids (XWIN/YWIN) define every position + rationale: >- + PSF interpolation sites, tile stamp centres, multi-epoch stamp + centres, and the catalogue sky position all use SExtractor's + windowed centroid — XWIN_WORLD/YWIN_WORLD on the tile side (SPHE), + XWIN_IMAGE/YWIN_IMAGE exposure-side (PIX). Windowed vs isophotal + vs model centroids differ systematically for blends and asymmetric + galaxies, and the centroid definition feeds the position seed. + Anchor: workflow/config/cfis/config_tile_PiViVi.ini#PSFEX_INTERP_RUNNER.POSITION_PARAMS; + workflow/config/cfis/config_tile_PiViVi.ini#VIGNETMAKER_RUNNER_RUN_2.POSITION_PARAMS; + workflow/config/cfis/config_exp_psfex.ini#POSITION_PARAMS. + default: xwin_windowed + options: + xwin_windowed: { label: Windowed centroids everywhere } + stamp_positioning_and_padding: + label: Nearest-pixel stamp centring; edge stamps zero-padded + rationale: >- + Multi-epoch stamps are placed by round-tripping the tile world + position through the stored per-CCD WCS, then rounding to the + nearest pixel (no sub-pixel interpolation — the residual sub-pixel + offset is absorbed by the fit's centroid prior, cen sigma = 1 + pixel). Objects whose stamp overruns a CCD or tile edge are KEPT, + out-of-image pixels zero-filled (sf_tools FetchStamps + pad_mode='constant'); no boundary rejection exists — zero-padded + pixels enter the fit as data with whatever weight the padded + weight stamp carries. Anchor: + src/shapepipe/modules/vignetmaker_package/vignetmaker.py::VignetMaker._get_stamp; + src/shapepipe/modules/vignetmaker_package/vignetmaker.py::VignetMaker._get_stamp_me. + default: round_and_zero_pad + options: + round_and_zero_pad: { label: "Nearest-pixel + zero padding, no edge rejection" } + epoch_flag_source: + label: Per-epoch flag stamps come from RAW CFIS flags, not the pipeline mask + rationale: >- + The multi-epoch vignet run reads its flag stamps from + split_exp_runner output — the delivered instrumental flags — + while mask_runner's pipeline_flag (halos, spikes, DSOs, borders) + feeds only the exposure-side star finding + (config_exp_psfex.ini FILE_PATTERN pipeline_flag). Combined with + unmasked tile detection (detection.detection_source_mode), the + consequence is stark: THE BRIGHT-STAR MASKS CURRENTLY AFFECT ONLY + PSF-STAR SELECTION — neither the galaxy sample (no tile mask, no + IMAFLAGS cut possible) nor the pixels ngmix fits (raw flags only) + see them. Whether that is intended belongs to the + planned masking-unification rework, whose object-level half is + sp_validation's IMAFLAGS_ISO cut on a column this chain never + produces; this is the pixel-level half. Anchor: + workflow/config/cfis/config_tile_PiViVi.ini#VIGNETMAKER_RUNNER_RUN_2.ME_IMAGE_EXP_RUNNERS; + workflow/config/cfis/config_exp_psfex.ini#SEXTRACTOR_RUNNER.FILE_PATTERN; + workflow/config/cfis/config_exp_Ma.ini#MASK_RUNNER.PREFIX. + default: raw_flags + options: + raw_flags: { label: split_exp raw flags gate epoch pixels } + pipeline_flags: + label: pipeline_flag (incl. bright-star masks) gates epoch pixels + prior_insights: + guinot22_gaia_astrometry: + claim: >- + The astrometric solution the analysis relies on is the upstream + MegaPipe/Gaia DR2 calibration, accurate to within 20 mas; no + astrometric re-fit inside the pipeline is described. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_astrometry + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'An astrometric calibration within 20 mas was achieved using the Gaia DR2 observations' + location: { page: 2 } + guinot22_forty_chips: + claim: >- + Star selection and PSF modelling are carried out independently on + each of the 40 MegaCam chips, with no chip excluded. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_forty_chips + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'is performed independently on each of the 40 chips that constitute the MegaCAM' + location: { page: 3 } + + # ═════════════════════════════════════════════════════════════════════════ + star_selection_psf: + description: >- + Which objects constrain the PSF, and the PSF model itself. Modules: + src/shapepipe/pipeline/str_handler.py (_mode — the iterative + histogram-zoom FWHM mode estimator centring the star box; median + fallback below N=20), src/shapepipe/modules/setools_package/setools.py + (_make_rand_split), src/shapepipe/modules/psfex_interp_package/ + psfex_interp.py (acceptance gates, HSM shapes). Configs: + star_selection.setools, default.psfex, config_exp_psfex.ini. + [LINT] star_stat logs the FWHM cut as mode +- 0.1*0.187 while the mask + applies mode +- 0.2 px — the run's own log misstates the selection. + [LINT] pixel scale appears as 0.187 (load-bearing) and 0.186 + (plot-only) in the same setools file. + inputs: + - id: exposure_sexcat + type: data + source: per-CCD exposure SExtractor catalogues (default_exp.sex run) + outputs: + - id: psf_model + type: data + format: psf + description: >- + Per-CCD PSFEx models + interpolated PSFs at object positions + (run_sp_exp_SxSePsfPi family), with HSM shape diagnostics. + decisions: + [star_selection_box, psf_train_validation_split, + psfex_candidate_vetting, psf_modelling_software, + psf_model_complexity, psf_acceptance_thresholds] + decisions: + star_selection_box: + label: Stellar-locus selection — mag window + FWHM window around the mode + rationale: >- + 18 < MAG_AUTO < 22, |FWHM - mode| <= 0.2 px, FLAGS==0, + IMAFLAGS_ISO==0; the mode is computed on a preselection + (MAG_AUTO<21, 0.3-1.5 arcsec at 0.187"/px) via the iterative + histogram-zoom estimator (str_handler.py::_mode, eps=0.001; median + fallback for N<20, -1 for N=0 — small-N behaviour changes selection + on sparse CCDs). PSFEx's automatic FWHM-range selection is off + (SAMPLE_AUTOSELECT N) and bad-pixel filtering is off, but PSFEx's + compiled-in fixed sample cuts still apply on top of this box — + see psfex_candidate_vetting. + Anchor: workflow/config/cfis/star_selection.setools#MASK:star_selection.MAG_AUTO; + workflow/config/cfis/star_selection.setools#MASK:preselect.MAG_AUTO; + workflow/config/cfis/default.psfex#SAMPLE_AUTOSELECT; + src/shapepipe/pipeline/str_handler.py::_mode. + default: mode_centred_box + options: + mode_centred_box: + label: FWHM-mode-centred box, +-0.2 px, mag 18-22, setools-only vetting + insights: [guinot22_star_box] + size_mag_locus_fit: + label: Fitted size-magnitude stellar locus + excluded: true + excluded_reason: Not wired; the mode-box is the validated v2.0 selection. + psfex_autoselect: + label: PSFEx SAMPLE_AUTOSELECT vetting on top + excluded: true + excluded_reason: >- + Deliberately disabled so selection lives in one place; rationale + not recorded in code. + psf_train_validation_split: + label: Random 80/20 star split — model fit vs held-out validation + rationale: >- + RAND_SPLIT ratio 20: star_split_ratio_80 fits the PSFEx model + (config_exp_psfex.ini FILE_PATTERN, and the tile multi-epoch + interpolation ME_DOT_PSF_PATTERN in config_tile_PiViVi.ini); + star_split_ratio_20 is the independent PSF-residual diagnostic + (PSFEX_INTERP MODE=VALIDATION). Trades model precision (fewer + training stars per CCD, interacting with the STAR_THRESH gate) + against an independent residual test. + [PENDING #873] The split is DETERMINISTIC: _make_rand_split takes + np.random.RandomState(seed).permutation(cat_size), the seed being + the digits of the unit's file number mod 2^32 — a pure function of + the input catalogue, fixed per CCD and independent of processing + order, the same philosophy as shape_measurement.ngmix_seed_mode's + SEED_FROM_POSITION. Before this the split drew from unseeded + np.random.randint, so the star sample entering the PSF model — and + therefore every shape downstream of it — differed between + identical runs; it was the one unseeded draw the position-seed work + left uncovered. One-off cost: the realised 80/20 membership changes + once (it is one further draw, now frozen), so PSF models and shapes + shift by that draw relative to every earlier product. + Anchor: workflow/config/cfis/star_selection.setools#RAND_SPLIT:star_split.RATIO; + src/shapepipe/modules/setools_package/setools.py::SETools._make_rand_split; + workflow/config/cfis/config_exp_psfex.ini#PSFEX_RUNNER.FILE_PATTERN; + workflow/config/cfis/config_tile_PiViVi.ini#PSFEX_INTERP_RUNNER.ME_DOT_PSF_PATTERN. + default: split_80_20_seeded + options: + split_80_20_seeded: + label: 80% train / 20% validation, seeded from the file number + insights: [guinot22_star_split] + split_80_20_unseeded: + label: Same split, unseeded np.random (pre-#873) + excluded: true + excluded_reason: >- + Retired by #873: it made the PSF star sample — and every shape + downstream of it — irreproducible run-to-run, the single + remaining unseeded draw in the science chain. Kept on the + record because every UNIONS product built before the smk-g4 + campaign was produced under it. + no_holdout: + label: 100% of stars in the model, no held-out diagnostic + excluded: true + excluded_reason: Loses the independent rho-statistic input. + psfex_candidate_vetting: + label: PSFEx-side candidate vetting — built-in defaults, unpinned + rationale: >- + default.psfex sets only SAMPLE_AUTOSELECT N; SAMPLE_MINSN, + SAMPLE_MAXELLIP, SAMPLE_FWHMRANGE, SAMPLE_VARIABILITY are absent, + so PSFEx's compiled-in defaults apply silently (MINSN 20, + MAXELLIP 0.3, FWHMRANGE 2-10 px, VARIABILITY 0.2) — a second star + selection nobody's config records, and one that changes if the + PSFEx binary version changes. BADPIXEL_FILTER N + PSF_RECENTER N: + star vignets with flagged/sentinel pixels are accepted unfiltered + and candidates are not recentred (CENTER_KEYS XWIN). The setools + box is therefore not the whole selection. [HARDCODED] (in the + PSFEx binary). + Anchor: workflow/config/cfis/default.psfex#SAMPLE_AUTOSELECT; + workflow/config/cfis/default.psfex#BADPIXEL_FILTER; + workflow/config/cfis/default.psfex#PSF_RECENTER. + default: builtin_defaults + options: + builtin_defaults: + label: "Compiled-in MINSN 20 / MAXELLIP 0.3 / FWHMRANGE 2-10, no bad-pixel filter" + insights: [guinot22_psfex_preselection_off] + pinned_explicit: + label: Write the SAMPLE_* values explicitly into default.psfex + psf_modelling_software: + label: PSF modelling software — PSFEx per-CCD vs MCCD focal-plane + rationale: >- + The committed chain fits PSFEx independently per CCD. MCCD + (Liaudat+2021) is a maintained in-tree alternative: a focal-plane + model fit across all 40 CCDs at once with a hybrid local+global + decomposition (src/shapepipe/modules/mccd_package/ + six + mccd_*_runner.py; knobs in example/cfis/config_MCCD.ini — + N_COMP_LOC=8, D_COMP_GLOB=8, LOC_MODEL=hybrid, MIN_N_STARS=20, + RMSE_THRESH=1.25). Unwired in workflow/config/cfis/ (needs the + MCCD config adapted, and config_exp_mccd.ini carries a stale + hardcoded PSF_MODEL_DIR path). The image-simulation path + substitutes PSF modelling entirely: fake_psf_runner injects the + true input PSF from a SKiLLS dictionary in psfex_interp's output + format. + Anchor: src/shapepipe/modules/mccd_package; + src/shapepipe/modules/fake_psf_package; + example/cfis/config_MCCD.ini#INSTANCE.N_COMP_LOC; + example/cfis/config_MCCD.ini#INPUTS.MIN_N_STARS; + example/cfis/config_exp_mccd.ini. + default: psfex + options: + psfex: + label: PSFEx, independent per-CCD models + insights: [guinot22_psfex_software, farrens22_two_psf_methods] + mccd_focal_plane: + label: MCCD hybrid local+global focal-plane model + true_input_psf: + label: fake_psf injection of the simulation's true PSF + excluded: true + excluded_reason: >- + Only meaningful on simulated images where the true PSF exists; + not a data-analysis option. + psf_model_complexity: + label: PSFEx model — pixel basis, degree-2 spatial polynomial per CCD + rationale: >- + BASIS_TYPE PIXEL, BASIS_NUMBER 20, PSF_SIZE 51,51, PSF_SAMPLING 1, + PSFVAR_DEGREES 2 in XWIN,YWIN per CCD (MEF_TYPE INDEPENDENT, + STABILITY_TYPE EXPOSURE), PSF_RECENTER N. Model flexibility sets + the PSF-leakage/overfitting balance — the dominant additive + systematic in cosmic shear. Values are the stock EB 2017 header; + rationale not recorded in code. + Anchor: workflow/config/cfis/default.psfex#BASIS_TYPE; + workflow/config/cfis/default.psfex#PSFVAR_DEGREES; + workflow/config/cfis/default.psfex#PSF_SIZE. + default: pixel_basis_deg2_per_ccd + options: + pixel_basis_deg2_per_ccd: + label: "PIXEL basis, degree 2, per-CCD" + insights: [guinot22_psf_no_oversampling] + deg3: + label: Degree-3 spatial variation + excluded: true + excluded_reason: >- + More flexibility per CCD needs more stars per CCD than the + count-floor world guarantees; not validated. + psf_acceptance_thresholds: + label: Per-CCD PSF-model quality gate (min stars, max chi2) + rationale: >- + A CCD whose model has ACCEPTED < STAR_THRESH or CHI2 > 2 is not + interpolated — its galaxies drop from the shear catalogue: direct + footprint selection, the in-code analogue of the DES blacklist. + [PENDING #873] Both passes now gate at 22 stars: the VALIDATION-mode + exposure config always did (config_exp_psfex.ini), and #873 raised + the MULTI-EPOCH science path 20 -> 22 in example/cfis + (config_tile_PiViVi_canfar_{sx,uc}.ini), with commit 90782098 + mirroring it into workflow/config/cfis/config_tile_PiViVi.ini — the + committed config fork this workflow actually reads (#848 D2). + Provenance of the retired 20, which is what makes this a fix rather + than a preference: commit fdc86553 (Kilbinger, 2020-06-30, "Forgot + to update new star number threshold for 80% of stars") deliberately + bumped 20 -> 22 to account for the 80/20 split, but only in the + validation config; the tile config kept the pre-split 20, so for + five years the SCIENCE path gated on the value that 2020 fix meant + to retire. (20 is also the psfex_interp function default, so the + stale-value reading rested on the commit provenance rather than on + the config alone.) + Published description (Guinot+22 p.4, Fig. 3): 22 stars/CCD, applied + to exactly the CCDs feeding multi-epoch shape measurement — the + number now agrees. Two gaps remain: an undocumented CHI2_THRESH=2 in + both configs, and the mechanism — interpsfex tests the PSFEx header + ACCEPTED/CHI2 at interpolation time and drops that epoch for objects + on the CCD, rather than excluding the CCD from PSF modelling as the + paper describes. + Anchor: src/shapepipe/modules/psfex_interp_package/psfex_interp.py::PSFExInterpolator.interpsfex; + workflow/config/cfis/config_exp_psfex.ini#PSFEX_INTERP_RUNNER.STAR_THRESH; + workflow/config/cfis/config_tile_PiViVi.ini#PSFEX_INTERP_RUNNER.STAR_THRESH; + example/cfis/config_tile_PiViVi_canfar_sx.ini#PSFEX_INTERP_RUNNER.STAR_THRESH; + example/cfis/config_tile_PiViVi_canfar_uc.ini#PSFEX_INTERP_RUNNER.STAR_THRESH. + default: stars22_chi2_2 + options: + stars22_chi2_2: + label: ">= 22 stars on both passes, chi2 <= 2" + insights: [des_psf_blacklist_local, guinot22_star_floor_22_local] + stars20_chi2_2: + label: ">= 20 stars on the science path, 22 in validation (pre-#873)" + excluded: true + excluded_reason: >- + Retired by #873 + 90782098. It was never a chosen value: it is + the pre-split threshold fdc86553 raised to 22 in 2020 for the + validation config and forgot on the science path, leaving the + science gate below both the published floor (Guinot+22 Fig. 3) + and the pipeline's own intent. Every UNIONS product built before + the smk-g4 campaign carries it. + des_25: + label: DES Y3 threshold (25 stars) + excluded: true + excluded_reason: >- + Not adopted; CFIS CCDs are smaller than DECam's — the right + number is survey-specific. + prior_insights: + des_psf_blacklist_local: + claim: >- + DES Y3 blacklists any CCD with fewer than 25 stars surviving + outlier rejection in the PSF fit. + created_at: "2026-07-16T00:00:00Z" + evidence: + - id: ev_jarvis_y3_local + doi: "10.48550/arXiv.2011.03409" + quote: + exact: "fewer than 25 stars survived the outlier rejection" + location: { page: 10 } + guinot22_star_floor_22_local: + claim: >- + The published ShapePipe/UNIONS analysis discards a CCD from the PSF + estimation when fewer than 22 stars are selected on it — the floor + the science-path PSF-interpolation gate now applies. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_star_floor_local + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'The dashed line represents the cut at 22 stars/CCD below which the CCD is discarded for the PSF estimation.' + location: { page: 4 } + guinot22_star_box: + claim: >- + The published star selection keeps objects whose FWHM lies within + 0.04 arcsec of the mode of a size preselection, restricted to the + magnitude range 18 < r < 22. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_star_box + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'From this pre-selection we keep objects for which the FWHM is within 0.04 arcsec of the mode. In addition to these size cuts, we only use star candidates in the magnitude range 18 < r < 22.' + location: { page: 4 } + guinot22_star_split: + claim: >- + The published analysis randomly splits the star sample in two, 80% + building the PSF model and 20% held out for the validation tests. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_star_split + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'To be able to perform these tests properly, our star sample has been randomly divided in two:' + location: { page: 7 } + guinot22_psfex_preselection_off: + claim: >- + PSFEx's internal pre-selection is deliberately disabled so that the + pipeline's own star selection is the only one, the paper describing + the model as fit on the entire star sample. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_preselection_off + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'Since we carry out our own star selection (see Sect. 4.1), we disable the internal PSFEx pre-selection, and the PSF is thus obtained using the entire star sample.' + location: { page: 4 } + guinot22_psfex_software: + claim: >- + The published UNIONS shear catalogue uses PSFEx for PSF modelling, + with MCCD named as upcoming rather than current work. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_psfex_software + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'We make use of the PSFEx software package' + location: { page: 4 } + farrens22_two_psf_methods: + claim: >- + ShapePipe ships two PSF modelling methods, PSFEx and MCCD, either or + both of which may be run and used for the galaxy shape measurement. + created_at: "2022-06-01T00:00:00Z" + evidence: + - id: ev_farrens22_two_psf + doi: "10.48550/arXiv.2206.14689" + quote: + exact: 'ShapePipe allows either or both methods to be run and subsequently used for the galaxy shape measurement.' + location: { page: 3 } + guinot22_psf_no_oversampling: + claim: >- + The PSFEx parametrisation is tabulated in the paper (pixel basis, + degree-2 spatial variation), with the deliberate choice not to + over-sample the PSF models. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_psf_complexity + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'The PSFEx parameters we used are presented in Table 1. We have chosen not to over-sample the PSF models.' + location: { page: 5 } + + # ═════════════════════════════════════════════════════════════════════════ + shape_measurement: + description: >- + Galaxy shape estimation: ngmix single-Gaussian fits with metacalibration. + Modules: src/shapepipe/modules/ngmix_package/ngmix.py (priors, metacal + setup, epoch handling, postage-stamp prep), ngmix_runner.py (config + exposure). Config: config_tile_Ng_template.ini. Most values here are + HARDCODED — scientific choices living in code with no config exposure; + this sub-analysis is where the silent-default risk concentrates. + [LINT] centroid_source default disagrees between the runner ("wcs", + production; always passed explicitly, ngmix_runner.py:170) and every + module-level signature ("hsm") — unreachable in the pipeline path, but + direct callers (tests, notebooks) silently get the other choice. + [LINT] pixel scale is 0.186 here (config PIXEL_SCALE) vs 0.187 in the + setools/masking configs — and star_selection.setools itself mixes + 0.187 (cuts) with 0.186 (SCATTER stat, :75). + inputs: + - id: vignets + type: data + source: 51x51 galaxy/weight/background-RMS vignets + interpolated PSFs (vignetmaker, psfex_interp) + outputs: + - id: ngmix_cat + type: data + format: fits + description: Per-tile metacal shear catalogue chunks (ngmix_runner family). + decisions: + [ngmix_seed_mode, galaxy_model, fit_priors, metacal_scheme, + centroid_source, epoch_quality_and_weighting, noise_model, + psf_epoch_loss_policy, megacam_ccd_flip] + decisions: + ngmix_seed_mode: + label: ngmix per-object RNG seeding + rationale: >- + Production historically seeded one RandomState from the tile ID, + consumed in object order — results depended on chunk boundaries. + The position seed (3-arcsec sky boxes + CCD offsets, zig-zag fold + + Cantor pairing mod 2^32; ngmix.py::position_seed) makes every + stream a function of sky position: chunk-invariant, + bit-reproducible, and metacal fixnoise counter-noise cancels across + image-simulation branches (ngmix#796). Consequence: chunking is + demoted to a pure throughput knob (reverting this decision + re-promotes it). Cost: noise streams change vs v2.0 — see + top-level baseline_validation_criterion. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::position_seed; + workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.SEED_FROM_POSITION; + src/shapepipe/modules/ngmix_runner.py::ngmix_runner. + default: position_seed + options: + position_seed: + label: Per-object seed from (ra, dec, ccd), 3-arcsec boxes + tile_seed: + label: Tile-wide RandomState (v2.0) + excluded: true + excluded_reason: >- + Chunk-dependent; retired outright (SEED_FROM_POSITION=False now + raises — ngmix_runner.py:110-116). + galaxy_model: + label: Galaxy and PSF model — single Gaussian [HARDCODED] + rationale: >- + ngmix.fitting.Fitter(model='gauss') for both galaxy and PSF + (ngmix.py::make_runners); guessers TPSFFluxAndPriorGuesser / + TFluxGuesser with T=0.25 and catalogue-flux guess, Runner ntry=5, + PSFRunner ntry=2 — with a non-convex likelihood, guess and retries + decide which objects converge (failed fits are NaN-filled with + flags, not raised). Rationale not recorded. Under metacal, model + bias largely cancels in the response, which is the standard defense + of 'gauss'; not stated in code. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::make_runners. + default: gauss + options: + gauss: + label: "Single Gaussian, T guess 0.25, ntry 5/2" + insights: [guinot22_gaussian_model] + exp_or_bdf: + label: exp / bdf galaxy models + excluded: true + excluded_reason: >- + Slower, and metacal makes the gain marginal; not validated on + CFIS. + fit_priors: + label: ngmix joint prior — GPriorBA(0.4), cen sigma = pixel scale, flat T/F [HARDCODED] + rationale: >- + Ellipticity GPriorBA sigma=0.4; centroid CenPrior sigma = one pixel + scale (0.186 arcsec, config PIXEL_SCALE — the coupling + sigma=pixel_scale is itself the hardcoded choice); flat T in + [-1, 1e3], flat F in [-100, 1e9] with negative support (bounds + decide which noisy fits survive vs rail). get_prior takes T/F range + arguments but no caller passes them. Prior width drives noise bias; + rationale not recorded. + Published description (Guinot+22 p.7): centroid sigma = pixel scale + ~0.187 arcsec, flat F in [-1e4, 1e9], flat half-light radius r50 in + [-10, 1e6] arcsec, ellipticity prior from Bernstein & Armstrong + (2014); current code: PIXEL_SCALE 0.186, flat F in [-100, 1e9], and + a flat prior on ngmix's second-moment size T in [-1, 1e3] rather + than on r50 — the prior families agree, the flux bound and pixel + scale have drifted, and the size prior is a different + parameterisation rather than a changed number. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::get_prior; + workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.PIXEL_SCALE. + default: gpriorba04_flat + options: + gpriorba04_flat: { label: "GPriorBA 0.4 + flat T/F with negative support" } + nonneg_informative: + label: Non-negative or informative T/F priors + excluded: true + excluded_reason: >- + Truncating negative support biases the noshear ensemble mean; + metacal wants symmetric noise response. + metacal_scheme: + label: Metacalibration — 5 types, step 0.01, fitgauss reconv, fixnoise [HARDCODED] + rationale: >- + types [noshear,1p,1m,2p,2m], step 0.01, psf='fitgauss' (runner + default; moves the metacal response directly — alternatives gauss/ + dilate/azgauss listed in the docstring), fixnoise=True, + use_noise_image=True, MetacalBootstrapper(ignore_failed_psf=True) + (changes which epochs enter the fit). No *_psf sheared types, so no + mcal_R_psf PSF-response term in the catalogue. fixnoise rationale + appears only in the position_seed docstring (counter-noise + cancellation). + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::do_ngmix_metacal. + default: five_types_step001_fitgauss + options: + five_types_step001_fitgauss: + label: "noshear+1p/1m/2p/2m, step 0.01, fitgauss, fixnoise" + insights: [guinot22_metacal_five_images] + with_psf_response: + label: Add sheared-PSF types for R_psf + excluded: true + excluded_reason: >- + Not wired; leakage is instead diagnosed via PSF_ORIG columns + + rho statistics downstream. + centroid_source: + label: Jacobian origin from WCS astrometry, not HSM moments + rationale: >- + Production runner default "wcs"; hsm is "legacy... noisy for stars + and flagged as incorrect by Fabian — see #767" (runner comment; a + rare recorded rationale). Moves the centroid-prior centre per + object. The runner reads an optional CENTROID_SOURCE config option + that no committed CFIS config sets. [LINT] module-level default is + still "hsm" — see this sub-analysis's description. + Published description (Guinot+22 p.7): HSM adaptive moments, run on + each sheared version, supplied the whole initial guess vector + (centroid, r50, flux) for the least-squares fit; current code: that + initialisation is gone — guesses come from ngmix's + TPSFFluxAndPriorGuesser with fixed T=0.25 and a catalogue flux, and + the only surviving HSM role is the optional stamp re-centering that + sets the Jacobian origin. So the drift is wider than a swapped + centroid source. (The paper's other HSM use, PSF/star shape + diagnostics, is unaffected.) + Anchor: src/shapepipe/modules/ngmix_runner.py::ngmix_runner; + src/shapepipe/modules/ngmix_package/ngmix.py::make_ngmix_observation. + default: wcs + options: + wcs: { label: WCS-projected catalogue position } + hsm: + label: HSM adaptive-moment centroid + excluded: true + excluded_reason: Noisy for stars; flagged incorrect (shapepipe#767). + epoch_quality_and_weighting: + label: Epoch admission, masking cut, and multi-epoch combination [HARDCODED] + rationale: >- + An epoch is dropped if >1/3 of its stamp is masked + (prepare_postage_stamps; the comment says "objects", the code drops + epochs — an object with zero surviving epochs drops out); failed + PSF fits drop epochs (flags != 0); fluxes rescaled by header FSCALE + (gal*Fscale, weight/Fscale^2); the diagnostic PSF is averaged over + epochs weighted by obs.weight.sum(). Joint multi-epoch fit over + survivors. Rationale for 1/3 and for the weight choice not + recorded. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::prepare_postage_stamps; + src/shapepipe/modules/ngmix_package/ngmix.py::rescale_epoch_fluxes; + src/shapepipe/modules/ngmix_package/ngmix.py::_average_psf_fits. + default: third_masked_cut + options: + third_masked_cut: { label: "Drop epoch if >1/3 masked; FSCALE rescale; weight-sum PSF average" } + noise_model: + label: Per-pixel inverse variance from background-RMS vignets + rationale: >- + BKG_RMS_VIGNET_PATH set in the CFIS template: weight = + 1/bkg_rms^2 per pixel (all-or-nothing; missing file errors); + fallback scalar 1/sigma_mad^2. Masked pixels filled with Gaussian + noise at sig_noise. A scalar sigma "mis-reports errors and erodes + the inverse-variance advantage whenever the RMS map actually + varies" (recorded rationale, fixnoise bookkeeping). PSF observation + gets a flat weight from PSF_NOISE=1e-5 — hardcoded module constant, + validated 1e-4..1e-6 on the digital twin (#749/#774 comment); + without it the g-prior swamps the PSF likelihood. Per-epoch + background subtraction BKG_SUB=True (off only for sims). + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::prepare_ngmix_weights; + src/shapepipe/modules/ngmix_package/ngmix.py::PSF_NOISE; + src/shapepipe/modules/ngmix_package/ngmix.py::background_subtract; + workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.BKG_RMS_VIGNET_PATH. + default: rms_vignet_weights + options: + rms_vignet_weights: { label: Per-pixel RMS-map weights + PSF_NOISE 1e-5 } + scalar_sigma_mad: + label: Scalar sigma_mad per epoch + excluded: true + excluded_reason: Mis-reports errors where the RMS map varies (recorded). + psf_epoch_loss_policy: + label: Object-level policy when CCDs fail PSF interpolation + rationale: >- + When k of ~40 CCDs fail the PSF acceptance gate (~5-6% attrition, + per-exposure clustered, matches the bash baseline — but measured + with the science gate at 20 stars, so [PENDING #873] at 22 it can + only rise, and smk-g4 is the first campaign to re-measure it), the + pipeline applies NO further quality gate: tiles complete, each + object records NGMIX_N_EPOCH, and sp report surfaces per-tile + epoch loss. + Object-level protection is delegated entirely to the validation + stage's epoch-count cut (sp_validation's galaxy selection cuts on N_EPOCH >= 1). Rationale (Cail, + 2026-08-29, PRD walk): epoch loss is a per-object depth effect + already recorded in the catalogue; gating at pipeline level would + fail whole tiles for a versionable catalogue decision. PRD #848's + open-questions section was removed accordingly. Per-tile epoch loss + is surfaced by the run report; NGMIX_N_EPOCH is the per-object + record. + Anchor: workflow/scripts/run_report.py; + workflow/config/cfis/final_cat.param#NGMIX_N_EPOCH. + default: record_and_delegate + options: + record_and_delegate: + label: Record NGMIX_N_EPOCH, report attrition, no pipeline gate + pipeline_epoch_floor: + label: Fail tiles below a minimum surviving-epoch fraction + excluded: true + excluded_reason: >- + Fails whole tiles for what is a versionable per-object + catalogue decision; the depth effect is already recorded. + megacam_ccd_flip: + label: 180-degree tile-vignet rotation for MegaCam CCDs <18 and 36-37 [HARDCODED] + rationale: >- + "MegaPipe has CCDs that are upside down" (docstring) — the tile + vignet is rotated to register with epoch stamps; a wrong flip + mis-registers the tile mask against the epoch, changing flagged + pixels and the 1/3-masked cut. Carries its own recorded caveat: + "will give incorrect results when used with THELI ccds. Fix this." + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::Ngmix.MegaCamFlip. + default: megapipe_flip + options: + megapipe_flip: { label: Flip CCDs <18 and 36/37 (MegaPipe orientation) } + prior_insights: + guinot22_gaussian_model: + claim: >- + The published shape measurement models galaxies with a single + Gaussian profile, arguing the resulting model bias is small and + largely absorbed by metacalibration. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_gaussian + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'Despite being very simple, the model bias (Kacprzak et al.' + suffix: ' 2014) is small.' + location: { page: 7 } + guinot22_metacal_five_images: + claim: >- + Metacalibration in the published analysis creates four sheared + images for calibration plus one for measurement, with a shear step + of 0.01 and a 90-degree-rotated noise image to cancel noise + correlations. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_metacal + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'This method creates four images used for the calibration, and one for the measurement.' + location: { page: 6 } + + # ═════════════════════════════════════════════════════════════════════════ + psf_diagnostics: + description: >- + The PSF-fidelity diagnostic chain: merge the held-out (20%) validation + stars into one catalogue, bin PSF shapes and residuals over the focal + plane. DORMANT in the committed snakemake chain — no rule runs it. + Modules: src/shapepipe/modules/merge_starcat_runner.py (+ per-model + merge classes in merge_starcat.py), mccd_plots_runner.py (serves both + PSF models despite its name). Boundary note: the module docstring + (mccd_package/__init__.py:157) still advertises rho-statistics plots, + but no rho/treecorr code remains in shapepipe — rho/tau statistics + moved downstream to sp_validation (rho_tau.py via + shear_psf_leakage.RhoStat/TauStat; treecorr min_sep/max_sep/nbins, + jackknife patch numbers hardcoded per survey with a "TODO to yaml"). + The diagnostic decision chain thus crosses the repo boundary into + sp_validation. [LINT] the module docstring still advertises rho + statistics this package no longer computes. + inputs: + - id: validation_star_cats + type: data + source: per-CCD star_split_ratio_20 catalogues with PSF/star HSM shapes (psfex_interp VALIDATION mode) + outputs: + - id: merged_star_catalogue + type: data + format: fits + description: >- + One full_starcat over the run — the input rho/tau statistics and + leakage diagnostics consume downstream. + decisions: [starcat_merge_source, meanshape_binning] + decisions: + starcat_merge_source: + label: Which PSF model's validation output feeds the merged star catalogue + rationale: >- + merge_starcat_runner dispatches on PSF_MODEL in {psfex, mccd, + setools} to per-model merge classes (different HDU conventions: + mccd HDU 1, psfex/setools HDU 2). Follows star_selection_psf. + psf_modelling_software; recorded separately because the merge can + also consume raw setools output (pre-model diagnostics). + Anchor: src/shapepipe/modules/merge_starcat_runner.py::merge_starcat_runner; + src/shapepipe/modules/merge_starcat_package/merge_starcat.py. + default: psfex + options: + psfex: + label: PSFEx validation catalogues (HDU 2) + insights: [guinot22_psfex_for_v1] + mccd: { label: MCCD validation catalogues (HDU 1) } + setools: { label: Raw setools star catalogues } + meanshape_binning: + label: Focal-plane mean-shape binning and outlier handling + rationale: >- + PSF ellipticity/size and residuals binned per CCD over the focal + plane: X_GRID=5, Y_GRID=10 bins per CCD, colour scales MAX_E=0.05, + MAX_DE=0.005, REMOVE_OUTLIERS=False (example/cfis config; no + committed workflow config exists). Grid resolution sets which + spatial PSF-residual structure is visible; outlier removal changes + what the diagnostic hides. + Published description (Guinot+22 p.8): focal-plane residual maps + averaged in ~20 arcsec cells per CCD; current code: no committed + workflow config picks a grid at all — example/cfis carries both the + 5x10 grid recorded here (~77x86 arcsec) and, in + config_valjoint_Pl_mccd.ini, a 20x40 grid (~19x22 arcsec) that + reproduces the paper. This is therefore an undetermined knob with + two committed precedents rather than a value that drifted. The + uniform REMOVE_OUTLIERS=False is a genuine paper-silence gap. + Anchor: example/cfis/config_MsPl_psfex.ini#MCCD_PLOTS_RUNNER.X_GRID; + example/cfis/config_MsPl_psfex.ini#MCCD_PLOTS_RUNNER.REMOVE_OUTLIERS; + src/shapepipe/modules/mccd_plots_runner.py; + src/shapepipe/modules/mccd_package/mccd_plot_utilities.py::plot_meanshapes. + default: grid_5x10 + options: + grid_5x10: { label: "5x10 per CCD, outliers kept" } + prior_insights: + guinot22_psfex_for_v1: + claim: >- + PSFEx is the PSF model behind the published UNIONS v1 catalogue, so + the merged validation star catalogue and its diagnostics are fed by + PSFEx output. + created_at: "2022-04-01T00:00:00Z" + evidence: + - id: ev_guinot22_psfex_v1 + doi: "10.48550/arXiv.2204.04798" + quote: + exact: 'We make use of the PSFEx software package' + location: { page: 4 } + + # ═════════════════════════════════════════════════════════════════════════ + survey_geometry: + description: >- + Effective survey area and mask geometry for two-point estimators. + DORMANT — no committed workflow rule. Module: + src/shapepipe/modules/random_cat_package/random_cat.py (+ runner): + uniform randoms over each tile rejected against the pipeline mask, + yielding effective area (overlap- and mask-corrected) and optionally + the mask itself as a HEALPix map (save_as_healpix). This is the + in-repo ancestor of the planned healsparse external-mask rework — whichever way that rework lands, this + sub-analysis is where its geometry decisions belong. Bitrot risk: + healpy is imported but absent from pyproject.toml dependencies. + inputs: + - id: tile_masks + type: data + source: per-tile pipeline flag maps + final catalogues + outputs: + - id: random_catalogue + type: data + format: fits + description: Per-tile random catalogue + effective area (+ optional HEALPix mask). + decisions: [random_sampling, healpix_mask_export] + decisions: + random_sampling: + label: Random-point density for area estimation + rationale: >- + N_RANDOM=50000 with DENSITY=True (per square degree; False = total + per tile) in the example config; no committed workflow value. + Sampling density sets the Monte Carlo noise floor on effective + area, which propagates to two-point normalisation. + Anchor: example/cfis/config_Rc.ini#RANDOM_CAT_RUNNER.N_RANDOM; + example/cfis/config_Rc.ini#RANDOM_CAT_RUNNER.DENSITY; + src/shapepipe/modules/random_cat_runner.py::random_cat_runner. + default: per_sqdeg_50k + options: + per_sqdeg_50k: { label: "50000 per sq deg" } + healpix_mask_export: + label: HEALPix export resolution for the pipeline mask + rationale: >- + SAVE_MASK_AS_HEALPIX=True, HEALPIX_OUT_NSIDE=1024 (~3.4 arcmin + pixels) in the example config — coarser than the arcsecond-scale + mask features it rasterises; the resolution choice decides what the + exported mask can represent. Supersession candidate under + masking-unification (healsparse). + Anchor: example/cfis/config_Rc.ini#RANDOM_CAT_RUNNER.SAVE_MASK_AS_HEALPIX; + example/cfis/config_Rc.ini#RANDOM_CAT_RUNNER.HEALPIX_OUT_NSIDE; + src/shapepipe/modules/random_cat_package/random_cat.py::RandomCat.save_as_healpix. + default: nside_1024 + options: + nside_1024: { label: nside 1024 } + + # ═════════════════════════════════════════════════════════════════════════ + catalogue_assembly: + description: >- + Final per-tile catalogue: merging shape chunks, attaching photometry + and PSF diagnostics, classification, sentinels. Modules: + src/shapepipe/modules/make_cat_package/make_cat.py (+ runner), + merge_sep_cats.py, vignetmaker_package (stamps, see top-level + postage_stamp_size), find_exposures_package (epoch list from tile + HISTORY cards). Configs: config_tile_Mc.ini, config_tile_PiViVi.ini, + final_cat.param. + inputs: + - id: ngmix_chunks + type: data + source: per-tile ngmix catalogue chunks + tile sexcat + PSF diagnostics + outputs: + - id: tile_final_cat + type: data + format: fits + description: The assembled per-tile science catalogue (final_cat family). + decisions: + [star_galaxy_classification, tile_overlap_handling, + column_selection, failure_sentinels, postproc_run_provenance, + shape_catalogue_shortfall_guard] + decisions: + star_galaxy_classification: + label: Star/galaxy separation — deferred out of the pipeline + rationale: >- + Production sets SM_DO_CLASSIFICATION=False (config_tile_Mc.ini) and + wires no spread-model input: SPREAD_MODEL/SPREADERR_MODEL are + sentinel 99, no SPREAD_CLASS column, and the SPREAD_* entries in + final_cat.param are commented out. The dormant machinery + (make_cat.py::save_sm_data) classifies on class = sm + 2*sm_err + with star |class|<0.003, galaxy class>0.01 — thresholds hardcoded + in the function signature. Reactivation is a 4-line config diff + (run spread_model_runner after psfex_interp+vignetmaker, add its + output to make_cat inputs, flip the switch — the exact diff + between example/cfis/config_make_cat_psfex.ini and _nosm.ini; the + defunct tile wiring config_tile_PiViSmVi.ini is the reference). + Separation therefore happens entirely downstream (sp_validation); + the pipeline ships everything. Rationale for deferring not + recorded. + Published description (Guinot+22 p.5-6): galaxies are selected + inside the pipeline with the spread model, at s + 2*sigma_s > + 0.0003 together with s > 0 and 20 < MAG_AUTO < 26; current code: + classification is disabled entirely and separation deferred to + sp_validation, with the dormant make_cat thresholds putting the + like-for-like galaxy boundary at 0.01, some thirty times the + published cut (0.003 is its separate star-side bound). Even + reactivated, the code implements only the spread-model test — the + paper's companion cuts have no in-pipeline counterpart. + Anchor: workflow/config/cfis/config_tile_Mc.ini#MAKE_CAT_RUNNER.SM_DO_CLASSIFICATION; + src/shapepipe/modules/make_cat_package/make_cat.py::save_sm_data; + workflow/config/cfis/final_cat.param#SPREAD_CLASS; + example/cfis/config_make_cat_psfex_nosm.ini; + example/cfis/defunct/config_tile_PiViSmVi.ini. + default: deferred_downstream + options: + deferred_downstream: + label: No in-pipeline classification; catalogue ships all objects + spread_model_inline: + label: spread_model classification in make_cat (0.003/0.01) + excluded: true + excluded_reason: >- + Machinery present but unwired in production; reactivating it + changes which objects downstream sees as galaxies. + tile_overlap_handling: + label: Tile-overlap duplicates — neither removed nor flagged + rationale: >- + Adjacent tiles overlap; objects in the overlap are measured in + both. make_cat attaches only TILE_ID (parsed from the sexcat + filename); no unique-object rule, no overlap flag. [LINT] the + documented config key TILE_LIST ("used to flag objects in areas of + overlap between tiles", in the make_cat package docstring) is + implemented nowhere — grep across src/ and workflow/ finds only + the docstring. VERIFIED downstream: sp_validation dedups at + classification time (galaxy.py::classification_galaxy_overlap_ra_dec + RA/Dec cuts to non-overlapping tile areas, and the WCS-based + mask_overlap variant; applied as cut_overlap in + classification_galaxy_base) — so this is today's division of labour, + and the pipeline's contract is "ship duplicates, TILE_ID is the + handle"; flagging overlaps in the catalogue stays an open option. The dead TILE_LIST docstring remains a lint. + Anchor: src/shapepipe/modules/make_cat_package/make_cat.py::save_sextractor_data; + src/shapepipe/modules/make_cat_package/__init__.py. + default: no_dedup_in_pipeline + options: + no_dedup_in_pipeline: + label: Ship duplicates; TILE_ID is the only handle + overlap_flagging: + label: Implement the documented TILE_LIST overlap flag + nearest_tile_centre: + label: Keep each object only in its nearest tile + excluded: true + excluded_reason: >- + Requires cross-tile coordination at assembly time, which the + per-tile DAG deliberately avoids; dedup belongs downstream if + anywhere. + column_selection: + label: Which columns survive into the science catalogue + rationale: >- + final_cat.param: positions XWIN/YWIN_WORLD, TILE_ID, flags + (FLAGS, IMAFLAGS_ISO, NGMIX_MCAL_FLAGS), PSF ellipticity from + PSF_ORIG only, all five metacal branches of G1/G2/T/FLUX/FLAGS, + but shear errors only for NOSHEAR (sheared-branch error columns + commented out — downstream response-weighted estimators cannot + propagate per-branch errors), SExtractor photometry (MAG_AUTO, + FLUX_APER, FLUX_RADIUS, SNR_WIN, FWHM_*), N_EPOCH/NGMIX_N_EPOCH, + NGMIX_MOM_FAIL. Doesn't change membership, but determines which + numbers exist for downstream cuts and calibration. Note + final_cat.param is read by scripts/python/create_final_cat.py in + post-processing, outside the per-tile DAG. [LINT] see detection: + IMAFLAGS_ISO is requested but never reaches the merged catalogue. + Anchor: workflow/config/cfis/final_cat.param; + scripts/python/create_final_cat.py. + default: committed_param_list + options: + committed_param_list: { label: The committed final_cat.param set } + failure_sentinels: + label: Objects without shape measurements kept, with sentinel values [HARDCODED] + rationale: >- + Unmatched objects (no ngmix row) stay in the catalogue with + sentinels: sizes/fluxes/flags 0, error fluxes/mags -1, + ellipticities -10, T_ERR 1e30. The sentinel choice defines what a + downstream cut must exclude — a naive G1 > -1 cut silently changes + the sample. Flag-0-for-failure is the sharpest hazard: a failed + object's NGMIX_MCAL_FLAGS reads as success. Rationale not recorded. + Anchor: src/shapepipe/modules/make_cat_package/make_cat.py::SaveCatalogue._save_ngmix_data. + default: sentinel_values + options: + sentinel_values: { label: "Keep with sentinels (flags 0, e -10, T_ERR 1e30)" } + drop_unmatched: + label: Drop objects without shapes + excluded: true + excluded_reason: >- + Loses the photometry-only population and hides attrition from + the completeness accounting. + postproc_run_provenance: + label: Post-proc run selection — newest mtime wins, merged patches never refresh + rationale: >- + create_final_cat picks each tile's make_cat run by newest + directory mtime (skipping runs without an output FITS), and the + merged patch catalogue is incremental: a tile already present is + never refreshed — a reprocessed tile reaches the patch only via + an explicit single-ID remove+add. mtime is filesystem state, not + provenance: a touched old run can outrank a newer one. Outside + the per-tile DAG (scripts/, not workflow/). Anchor: + scripts/python/create_final_cat.py::process. + default: newest_mtime_incremental + options: + newest_mtime_incremental: { label: "Newest mtime, incremental merge, manual refresh" } + shape_catalogue_shortfall_guard: + label: 10% shape-shortfall guard — logged, not enforced [HARDCODED] + rationale: >- + If the merged shape catalogue covers <10% of the detection + catalogue, make_cat logs an error but the enforcement (return + + raise) is commented out in both the module and its runner: a tile + whose shapes are 90% missing from a processing error is written and + looks normal. The comment distinguishes the two causes (measurement + failure = ok; premature merge = error) but not why enforcement is + off. Interacts with top-level per_unit_count_floor, which floors + tile_ngmix at 1 file and cannot see intra-file attrition. + Anchor: src/shapepipe/modules/make_cat_package/make_cat.py::SaveCatalogue._save_ngmix_data; + src/shapepipe/modules/make_cat_runner.py::make_cat_runner. + default: log_only + options: + log_only: { label: "Log the shortfall, write the tile anyway" } + enforce_10pct: + label: Fail the tile below 10% coverage + excluded: true + excluded_reason: >- + Was the coded intent, then disabled — reason unrecorded; + flagged as a question, not an endorsement. + +# ── Dormant science paths, surveyed but not yet recorded as sub-analyses ──── +# Candidates for future passes (each carries real scientific knobs): +# * External photometry match — match_external_package (TOLERANCE=0.3 +# arcsec against UNIONS ugriz; the external catalogue path is hardcoded +# to an IAP/candide location). +# * Image-simulation validation wiring — example/cfis_image_sims/: +# same chain over SKiLLS images with fake_psf substitution; bash-shaped, +# not yet ported to snakemake. diff --git a/universes/committed.yaml b/universes/committed.yaml new file mode 100644 index 000000000..5530226cf --- /dev/null +++ b/universes/committed.yaml @@ -0,0 +1,70 @@ +id: committed +description: The committed configuration on feat/snakemake-orchestration. +decisions: + per_unit_count_floor: count_floor + postage_stamp_size: px_51 + photometric_zeropoint: fixed_30_tiles_header_exposures + baseline_validation_criterion: statistical_parity +analyses: + masking: + decisions: + star_catalogue_query: gsc_23_vizier + star_magnitude_definition: mean_finite_bands + bright_star_mask_geometry: megaprime_polygon_linear_scaling + deep_sky_object_masking: circles_no_padding + border_mask_width: px50_exposures_only + pixel_threshold_flags: stock_ww_thresholds + external_flag_usage: exposures_only + detection: + decisions: + detection_threshold_policy: thresh_1p5_minarea5_fwhm2px_filter + deblending_policy: mincont_5em4_tiles + background_model: manual_zero_tiles_auto_exposures + weighting_and_interpolation: map_weight_interp_all + detection_source_mode: sx_nomask_single_image + epoch_membership_ccd_bounds: trimmed_bounds_33_2080 + photometry_parameters: kron_25_35 + cleaning_and_neighbour_masking: clean_1_correct + preparation: + decisions: + astrometric_solution_source: delivered_headers + ccd_split_extent: all_40_hdus + epoch_provenance_from_tile_history: history_parse + object_position_columns: xwin_windowed + stamp_positioning_and_padding: round_and_zero_pad + epoch_flag_source: raw_flags + star_selection_psf: + decisions: + star_selection_box: mode_centred_box + psfex_candidate_vetting: builtin_defaults + psf_modelling_software: psfex + psf_train_validation_split: split_80_20_seeded + psf_model_complexity: pixel_basis_deg2_per_ccd + psf_acceptance_thresholds: stars22_chi2_2 + shape_measurement: + decisions: + ngmix_seed_mode: position_seed + galaxy_model: gauss + fit_priors: gpriorba04_flat + metacal_scheme: five_types_step001_fitgauss + centroid_source: wcs + epoch_quality_and_weighting: third_masked_cut + noise_model: rms_vignet_weights + psf_epoch_loss_policy: record_and_delegate + megacam_ccd_flip: megapipe_flip + psf_diagnostics: + decisions: + starcat_merge_source: psfex + meanshape_binning: grid_5x10 + survey_geometry: + decisions: + random_sampling: per_sqdeg_50k + healpix_mask_export: nside_1024 + catalogue_assembly: + decisions: + star_galaxy_classification: deferred_downstream + tile_overlap_handling: no_dedup_in_pipeline + column_selection: committed_param_list + failure_sentinels: sentinel_values + postproc_run_provenance: newest_mtime_incremental + shape_catalogue_shortfall_guard: log_only From 206b48d6c4818578ed5ecfff4e730cfcf80efcc2 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 02:34:43 +0200 Subject: [PATCH 54/85] test(astra): validate decision anchors and universe pins --- tests/helpers/astra_record.py | 284 +++++++++++++++++++++++++++++++ tests/unit/test_astra_anchors.py | 50 ++++++ 2 files changed, 334 insertions(+) create mode 100644 tests/helpers/astra_record.py create mode 100644 tests/unit/test_astra_anchors.py diff --git a/tests/helpers/astra_record.py b/tests/helpers/astra_record.py new file mode 100644 index 000000000..2b8f70ed9 --- /dev/null +++ b/tests/helpers/astra_record.py @@ -0,0 +1,284 @@ +"""Reusable parsing and resolution helpers for ShapePipe's ASTRA record.""" + +import ast +import configparser +from dataclasses import dataclass +from pathlib import Path +import re + +import yaml + + +@dataclass(frozen=True) +class Anchor: + """An anchor sentence found in a YAML value.""" + + location: str + references: tuple[str, ...] + error: str | None = None + + +def load_yaml(path): + """Load YAML with PyYAML's safe loader.""" + + return yaml.safe_load(Path(path).read_text(encoding="utf-8")) + + +def _walk(value, location=""): + if isinstance(value, dict): + for key, child in value.items(): + path = f"{location}.{key}" if location else str(key) + yield from _walk(child, path) + elif isinstance(value, list): + for index, child in enumerate(value): + yield from _walk(child, f"{location}[{index}]") + else: + yield location, value + + +def _rationales(document): + for location, value in _walk(document): + if location.endswith(".rationale"): + yield location, value + + +def extract_anchors(document): + """Parse all ``Anchor:`` sentences and check every rationale has one.""" + + anchors = [] + for location, value in _walk(document): + if not isinstance(value, str) or "Anchor:" not in value: + continue + tail = value.split("Anchor:", 1)[1].strip() + error = None + refs = () + if value.count("Anchor:") != 1: + count = value.count("Anchor:") + error = f"expected one Anchor: marker, found {count}" + elif not tail.endswith("."): + error = "anchor sentence must end with a period" + else: + refs = tuple(part.strip() for part in tail[:-1].split(";")) + if not refs or any(not ref for ref in refs): + error = "anchor sentence contains an empty ref" + anchors.append(Anchor(location, refs, error)) + + for location, value in _rationales(document): + if ( + not isinstance(value, str) + or value.count("Anchor:") != 1 + or not value.rstrip().endswith(".") + ): + anchors.append( + Anchor( + location, + (), + "rationale must end with exactly one Anchor: sentence", + ) + ) + return anchors + + +def _parse_reference(reference): + if "::" in reference: + path, symbol = reference.split("::", 1) + return "code", path, symbol + if "#" in reference: + path, key = reference.split("#", 1) + return "config", path, key + return "path", reference, "" + + +def resolve_anchor(root, reference): + """Return ``None`` if a reference resolves, otherwise a diagnostic.""" + + kind, relative, selector = _parse_reference(reference) + path = Path(relative) + if path.is_absolute() or ".." in path.parts: + return "path must be relative to the repository root" + target = Path(root) / path + if not target.exists(): + return "path does not exist" + if kind == "path": + return None + if not target.is_file(): + return "code/config refs must name a file" + + try: + text = target.read_text(encoding="utf-8") + except (OSError, UnicodeError) as error: + return f"cannot read file: {error}" + + if kind == "code": + if target.suffix != ".py": + return "code-symbol refs must name a .py file" + try: + tree = ast.parse(text, filename=str(target)) + except SyntaxError as error: + return f"cannot parse Python file: {error}" + if not _has_symbol(tree, selector): + return f"no def/class/assignment target named {selector!r}" + return None + + suffix = target.suffix.lower() + if suffix == ".ini": + return _ini_key(text, selector) + if suffix == ".setools": + return _setools_key(text, selector) + if suffix in {".sex", ".psfex", ".ww", ".param", ".conf"}: + key = selector.rsplit(".", 1)[-1] + pattern = re.compile(rf"^\s*(?:#\s*)?{re.escape(key)}(?=$|\s|=|\()") + if any(pattern.search(line) for line in text.splitlines()): + return None + return f"no line starts with key {key!r} (commented keys are allowed)" + return f"unsupported config-key file type {suffix or '(no extension)'}" + + +def _ini_key(text, selector): + if "." not in selector: + return "INI config ref needs SECTION.KEY" + section, key = selector.rsplit(".", 1) + parser = configparser.ConfigParser( + interpolation=None, strict=False, allow_no_value=True + ) + try: + parser.read_string(text) + except configparser.Error as error: + return f"cannot parse INI file: {error}" + if not parser.has_section(section): + return f"INI section {section!r} is missing" + if not parser.has_option(section, key): + return f"INI key {key!r} is missing from section {section!r}" + return None + + +def _setools_key(text, selector): + if "." not in selector: + return "SETools config ref needs SECTION.KEY" + section, key = selector.rsplit(".", 1) + pattern = re.compile(rf"^\s*(?:#\s*)?{re.escape(key)}(?=$|\s|=|<|>)") + active = False + for line in text.splitlines(): + stripped = line.strip() + if stripped.startswith("[") and stripped.endswith("]"): + active = stripped[1:-1].strip() == section + elif active and pattern.search(line): + return None + return f"SETools key {key!r} is missing from section {section!r}" + + +def _target_names(target): + if isinstance(target, ast.Name): + return [target.id] + if isinstance(target, ast.Attribute): + return [target.attr] + if isinstance(target, (ast.Tuple, ast.List)): + return [name for item in target.elts for name in _target_names(item)] + if isinstance(target, ast.Starred): + return _target_names(target.value) + return [] + + +def _bindings(scope): + """Collect declarations and assignment targets in one lexical scope.""" + + result = {} + + def visit(node): + if isinstance( + node, + (ast.FunctionDef, ast.AsyncFunctionDef, ast.ClassDef), + ): + result[node.name] = node + return + if isinstance(node, ast.Lambda): + return + if isinstance(node, ast.Assign): + targets = node.targets + elif isinstance(node, (ast.AnnAssign, ast.AugAssign, ast.NamedExpr)): + targets = [node.target] + elif isinstance(node, (ast.For, ast.AsyncFor)): + targets = [node.target] + elif isinstance(node, (ast.With, ast.AsyncWith)): + targets = [item.optional_vars for item in node.items] + else: + targets = [] + for target in targets: + if target is not None: + result.update(dict.fromkeys(_target_names(target), node)) + if isinstance(node, ast.ExceptHandler) and node.name: + result[node.name] = node + for child in ast.iter_child_nodes(node): + visit(child) + + for statement in scope.body: + visit(statement) + return result + + +def _has_symbol(tree, symbol): + scope = tree + parts = symbol.split(".") + for index, part in enumerate(parts): + declaration = _bindings(scope).get(part) + if declaration is None: + return False + if index == len(parts) - 1: + return True + if not isinstance( + declaration, + (ast.ClassDef, ast.FunctionDef, ast.AsyncFunctionDef), + ): + return False + scope = declaration + return False + + +def universe_errors(record, universe): + """Check scoped decision IDs and options against the ASTRA record.""" + + record_decisions = _decisions(record) + pinned = _decisions(universe) + errors = [] + for location in sorted(pinned.keys() - record_decisions.keys()): + errors.append( + f"{location}: universe decision is absent from astra.yaml" + ) + for location in sorted(record_decisions.keys() - pinned.keys()): + errors.append( + f"{location}: astra.yaml decision is not pinned in the universe" + ) + for location in sorted(record_decisions.keys() & pinned.keys()): + definition = record_decisions[location] + options = ( + definition.get("options", {}) + if isinstance(definition, dict) + else {} + ) + if not isinstance(options, dict) or pinned[location] not in options: + errors.append( + f"{location}: pinned option {pinned[location]!r} is not in " + "ASTRA options" + ) + return errors + + +def _decisions(document, location=""): + scope = document if isinstance(document, dict) else {} + result = {} + for decision_id, definition in (scope.get("decisions") or {}).items(): + key = ( + f"{location}.decisions.{decision_id}" + if location + else f"decisions.{decision_id}" + ) + result[key] = definition + for analysis_id, analysis in (scope.get("analyses") or {}).items(): + child = ( + f"{location}.analyses.{analysis_id}" + if location + else f"analyses.{analysis_id}" + ) + if isinstance(analysis, dict): + result.update(_decisions(analysis, child)) + return result diff --git a/tests/unit/test_astra_anchors.py b/tests/unit/test_astra_anchors.py new file mode 100644 index 000000000..b4d95c2d1 --- /dev/null +++ b/tests/unit/test_astra_anchors.py @@ -0,0 +1,50 @@ +"""Keep ASTRA decision anchors and the committed universe resolvable.""" + +from pathlib import Path + +from tests.helpers.astra_record import ( + extract_anchors, + load_yaml, + resolve_anchor, + universe_errors, +) + + +REPO_ROOT = Path(__file__).resolve().parents[2] + + +def test_every_astra_anchor_resolves(): + record = load_yaml(REPO_ROOT / "astra.yaml") + anchors = extract_anchors(record) + errors = [] + + assert anchors, "astra.yaml contains no Anchor: sentences" + for anchor in anchors: + if anchor.error: + errors.append(f"{anchor.location}: {anchor.error}") + continue + for reference in anchor.references: + problem = resolve_anchor(REPO_ROOT, reference) + if problem: + errors.append( + f"{anchor.location}: {reference}: {problem}" + ) + + message = ( + "Unresolved ASTRA anchors or rationales:\n - " + + "\n - ".join(errors) + ) + assert not errors, message + + +def test_committed_universe_matches_astra_decisions(): + record = load_yaml(REPO_ROOT / "astra.yaml") + universe = load_yaml(REPO_ROOT / "universes" / "committed.yaml") + + errors = universe_errors(record, universe) + + message = ( + "ASTRA / committed universe mismatch:\n - " + + "\n - ".join(errors) + ) + assert not errors, message From 9f9ec5ea182bb24a991f82365b19ce82f8a1bef4 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 02:54:27 +0200 Subject: [PATCH 55/85] test(astra): resolve Snakemake rule anchors; JSON report mode Co-Authored-By: Claude Fable 5.1 --- tests/helpers/astra_record.py | 102 +++++++++++++++++++++++++++++++ tests/unit/test_astra_anchors.py | 25 ++++++++ 2 files changed, 127 insertions(+) diff --git a/tests/helpers/astra_record.py b/tests/helpers/astra_record.py index 2b8f70ed9..1a7123c7d 100644 --- a/tests/helpers/astra_record.py +++ b/tests/helpers/astra_record.py @@ -1,10 +1,14 @@ """Reusable parsing and resolution helpers for ShapePipe's ASTRA record.""" +import argparse import ast import configparser from dataclasses import dataclass +import json from pathlib import Path import re +import subprocess +import sys import yaml @@ -89,6 +93,26 @@ def _parse_reference(reference): return "path", reference, "" +_SNAKEMAKE_SUFFIXES = {".smk"} +_SNAKEFILE_NAMES = {"Snakefile"} + + +def _is_snakemake_file(target): + return target.suffix in _SNAKEMAKE_SUFFIXES or target.name in _SNAKEFILE_NAMES + + +def _snakemake_symbol(text, symbol): + rule_pattern = re.compile( + rf"^\s*(?:rule|checkpoint)\s+{re.escape(symbol)}\s*:", re.MULTILINE + ) + if rule_pattern.search(text): + return None + def_pattern = re.compile(rf"^\s*def\s+{re.escape(symbol)}\(", re.MULTILINE) + if def_pattern.search(text): + return None + return f"no rule/checkpoint/def named {symbol!r}" + + def resolve_anchor(root, reference): """Return ``None`` if a reference resolves, otherwise a diagnostic.""" @@ -110,6 +134,8 @@ def resolve_anchor(root, reference): return f"cannot read file: {error}" if kind == "code": + if _is_snakemake_file(target): + return _snakemake_symbol(text, selector) if target.suffix != ".py": return "code-symbol refs must name a .py file" try: @@ -282,3 +308,79 @@ def _decisions(document, location=""): if isinstance(analysis, dict): result.update(_decisions(analysis, child)) return result + + +def _git_sha(root): + try: + return subprocess.run( + ["git", "rev-parse", "HEAD"], + cwd=root, + capture_output=True, + check=True, + text=True, + ).stdout.strip() + except (OSError, subprocess.CalledProcessError): + return None + + +def build_report(root): + """Resolve every anchor and universe pin under ``root`` into a report dict.""" + + root = Path(root) + astra_yaml = root / "astra.yaml" + record = load_yaml(astra_yaml) + anchors = extract_anchors(record) + + unresolved = [] + for anchor in anchors: + if anchor.error: + unresolved.append( + {"location": anchor.location, "ref": None, "problem": anchor.error} + ) + continue + for reference in anchor.references: + problem = resolve_anchor(root, reference) + if problem: + unresolved.append( + { + "location": anchor.location, + "ref": reference, + "problem": problem, + } + ) + + universe = load_yaml(root / "universes" / "committed.yaml") + errors = universe_errors(record, universe) + + return { + "astra_yaml": str(astra_yaml), + "git_sha": _git_sha(root), + "anchors_total": len(anchors), + "unresolved": unresolved, + "universe_errors": errors, + "ok": not unresolved and not errors, + } + + +def main(argv=None): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument( + "--report", required=True, help="path to write the JSON report to" + ) + parser.add_argument( + "--root", + default=Path(__file__).resolve().parents[2], + help="repository root (default: repo root inferred from this file)", + ) + args = parser.parse_args(argv) + + report = build_report(args.root) + report_path = Path(args.report) + report_path.parent.mkdir(parents=True, exist_ok=True) + report_path.write_text(json.dumps(report, indent=2) + "\n", encoding="utf-8") + + return 0 if report["ok"] else 1 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/tests/unit/test_astra_anchors.py b/tests/unit/test_astra_anchors.py index b4d95c2d1..38b334f12 100644 --- a/tests/unit/test_astra_anchors.py +++ b/tests/unit/test_astra_anchors.py @@ -13,6 +13,31 @@ REPO_ROOT = Path(__file__).resolve().parents[2] +def test_snakemake_rule_and_function_anchors_resolve(tmp_path): + rules_dir = tmp_path / "workflow" / "rules" + rules_dir.mkdir(parents=True) + (rules_dir / "example.smk").write_text( + "def tile_local(tile):\n return tile\n\n\n" + "rule tile_detect:\n input: 'a'\n output: 'b'\n", + encoding="utf-8", + ) + + assert resolve_anchor(tmp_path, "workflow/rules/example.smk::tile_detect") is None + assert resolve_anchor(tmp_path, "workflow/rules/example.smk::tile_local") is None + + problem = resolve_anchor(tmp_path, "workflow/rules/example.smk::no_such_rule") + assert problem == "no rule/checkpoint/def named 'no_such_rule'" + + +def test_snakefile_rule_anchor_resolves(tmp_path): + (tmp_path / "workflow").mkdir() + (tmp_path / "workflow" / "Snakefile").write_text( + "checkpoint plan:\n input: 'a'\n", encoding="utf-8" + ) + + assert resolve_anchor(tmp_path, "workflow/Snakefile::plan") is None + + def test_every_astra_anchor_resolves(): record = load_yaml(REPO_ROOT / "astra.yaml") anchors = extract_anchors(record) From 5368d9df7e981caabc2369665fd18f60c52c6d96 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 03:04:34 +0200 Subject: [PATCH 56/85] docs(astra): rewrite the decision record against develop - masking: describe healsparse queries (mask_query MASK_EXT on exposures, make_cat MASK_ on tiles) and the instrument flag image as the only pixel mask, replacing the deleted in-house mask generation - detection: tiles follow the MegaPipe (Gwyn) SExtractor parameters (#896); option ids no longer encode the retired values - shape_measurement: import defect_fill, blend_handling and epoch_masked_fraction_cut from the digital twin with their literature insights; defaults are what the committed code selects - prune to the membership test: drop psf_diagnostics, survey_geometry, the workflow-policy decisions and the root findings; split compound decisions; reserve excluded for considered-and-rejected; strip chronology - re-point anchors to the current configs; the anchor test passes Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_014bvNTrAmZxcfb1ee83ApPK --- astra.yaml | 2397 ++++++++++++++++++-------------------- universes/committed.yaml | 54 +- 2 files changed, 1187 insertions(+), 1264 deletions(-) diff --git a/astra.yaml b/astra.yaml index 70b25b554..1b27db982 100644 --- a/astra.yaml +++ b/astra.yaml @@ -1,162 +1,129 @@ -# ASTRA record for ShapePipe: the scientific decisions embedded in the code and -# the committed configs, with their reasoning and the alternatives that were -# rejected. It is the place scientific decisions are written down — see the -# "Scientific decisions" section of CLAUDE.md for when and how to amend it. +# ASTRA record of ShapePipe's scientific decisions: the choices embedded in +# the code and the committed workflow configs (workflow/config/cfis/), why +# they stand, and the alternatives. CLAUDE.md says when to amend it; the +# anchor test tests/unit/test_astra_anchors.py keeps it resolvable. # -# The record describes the pipeline as orchestrated by workflow/Snakefile -# (PRD CosmoStat/shapepipe#848, PR #852). Conventions: -# -# * Decisions anchor to code, not recipes. Every rationale ends with one -# sentence "Anchor: ; ; ..." in a strict, greppable grammar. -# Each ref is a path relative to the shapepipe repo root, in one of three -# forms: CODE `path::symbol`, CONFIG `path#SECTION.KEY` (or `path#KEY` for -# sectionless .sex/.psfex/.ww/.param files), FILE `path` for a whole file -# or package. No line numbers — they rot; a line-level fact names its -# enclosing symbol. The analysis-ASTRA rule "never hardcode; reference via -# {decisions.x}" cannot hold in a codebase — the committed configs ARE the -# values. [HARDCODED] marks a scientific value living in code with no -# config exposure: the silent defaults the record exists to surface. -# * The default universe IS the committed configuration (universes/committed). -# Alternatives are excluded-with-reasons or genuinely open forks. -# * Sub-analyses follow the pipeline's methodological units — masking, -# detection, preparation, star selection + PSF, shape measurement, PSF -# diagnostics, survey geometry, catalogue assembly — not its ~20 Snakemake -# rules. Cross-cutting decisions stay top-level. A prior_insight repeated -# inside a sub-analysis carries a `_local` suffix: ids are scoped, and the -# duplicate keeps the sub-analysis readable on its own. -# * Outputs are representative product FAMILIES (one final_cat per tile), -# not enumerable artifacts; no recipes — the executor is the Snakemake -# workflow. -# * [LINT] marks places where this record and the code already disagree, or -# where the code disagrees with itself — found while authoring this file. -# * [PENDING #NNN] marks state that is live on feat/snakemake-orchestration -# — and therefore in smk-g4, the 34-tile validation campaign run under -# this branch — but not yet merged to develop. The record follows the -# branch and names the open PR. -# * A `path#KEY` anchor names the key's position in the file, not its -# activation: where the decision is "this is deliberately off", the key -# it points at may be commented out (e.g. final_cat.param#SPREAD_CLASS). +# Conventions: +# * Every rationale ends with one sentence "Anchor: ; ." Each ref +# is a path relative to the repo root: CODE `path::Symbol` (a def, class +# or assignment target, dotted for nesting), CONFIG `path#SECTION.KEY` +# (`path#KEY` for sectionless .sex/.psfex/.param files, where a +# commented-out key still resolves), or FILE `path`. No line numbers. +# Numeric values are stated once, in the rationale, next to the anchor +# that holds them. +# * A decision's `default` is the option the committed code and configs +# select; universes/committed.yaml pins it. +# * `excluded: true` means considered and rejected. An option the code does +# not implement says so in its description and is not excluded. +# * [HARDCODED] in a rationale marks a committed choice fixed in code with +# no config key to change it. +# * [LINT] in a rationale marks a place where the code disagrees with itself +# or with its own documentation. +# * A prior insight repeated inside a sub-analysis carries a `_local` +# suffix, because insight ids are scoped. version: "0.0.14" name: ShapePipe scientific decisions description: >- Codebase-level decision record for the ShapePipe weak-lensing pipeline - (UNIONS/CFIS). Membership test: "a different defensible choice would change - which objects enter the shear catalogue, or the numbers attached to them." - Workflow mechanics that reproduce identical numbers (manifest sentinels, - clean-cascade cut, directory() outputs, allocation strategy, chunking under - position seeding) are deliberately absent; they live in the PRD and code. + (UNIONS/CFIS) as orchestrated by workflow/Snakefile. Membership test: a + different defensible choice would change which objects enter the shear + catalogue, or the numbers attached to them. tags: [shapepipe, weak-lensing, unions, codebase-record] -container: shapepipe-develop-runtime.sif +container: ghcr.io/cosmostat/shapepipe:develop inputs: - id: tile_images type: data - source: CADC-staged CFIS/UNIONS r-band tile stacks + exposure triplets (workflow/config.yaml) + source: CFIS/UNIONS r-band MegaPipe tile stacks and single-exposure image/weight/flag triplets (workflow/config.yaml) description: >- - Pre-staged P3 tiles and single-exposure image/weight/flag triplets on - /project; get_images runs with RETRIEVE=symlink against this store. - - id: gsc_star_catalogue + Tiles and the exposures that built them; each exposure carries its + instrument flag image. + - id: healsparse_masks type: data - source: GSC 2.3 (Vizier I/305/out) cone queries — scripts/python/create_star_cat.py + source: UNIONS healsparse mask products (e.g. mask_ugriz_nside131072_n4.hsp) description: >- - Reference star catalogue driving bright-star masking. Catalogue choice, - query geometry, and magnitude handling are decisions in the masking - sub-analysis. + Sky-fixed mask maps, built outside ShapePipe and queried at object + positions; how they are used is the masking sub-analysis. outputs: - id: final_cat type: data format: fits description: >- - Per-tile shear catalogue family, the terminal science product (one per - campaign tile; make_cat_runner). Column selection and failure sentinels - are decisions in catalogue_assembly. - inputs: [tile_images] - decisions: [per_unit_count_floor, postage_stamp_size, photometric_zeropoint] + Per-tile shear catalogue family, the terminal science product + (make_cat_runner). + inputs: [tile_images, healsparse_masks] + decisions: [per_unit_completeness, postage_stamp_size, photometric_zeropoint] decisions: # ── cross-cutting ──────────────────────────────────────────────────────── - per_unit_count_floor: - label: Per-unit completeness policy under partial failure + per_unit_completeness: + label: Per-unit completeness gate rationale: >- - A 40-CCD stage where some CCDs legitimately produce nothing (sparse CCD, - setools rejects everything) cannot be all-or-nothing. The field's - converged answer (DES PSF blacklist, Rubin quantum registry) is per-unit - outcome records gated on a quality floor: record the attrition, fail - loud only below the floor, continue the survey. The floor VALUES are the - scientific content — how much silent per-CCD attrition can enter the - catalogue. The COMPLETENESS table holds them (exp_split - expect=121/floor=41, exp_mask expect=40/floor=1, psfex expect=80/floor=2, - psfex_interp floor=0 warn-only). Related leak the floor does not cover: - merge_sep_cats warns-and-skips a missing ngmix chunk, silently shrinking - a tile's shape catalogue below the floor's radar; and make_cat's own 10% - size-shortfall guard is commented out (see - catalogue_assembly.shape_catalogue_shortfall_guard). - Anchor: workflow/scripts/completeness.py::COMPLETENESS; - src/shapepipe/modules/merge_sep_cats_package/merge_sep_cats.py::MergeSep.process. - default: count_floor + Every rule checks its products against a nominal per-runner count: a + runner below its count fails the unit (an exposure or a tile), so a + partial unit never enters the catalogue; the missing unit's objects do + not appear at all. The one tolerated shortfall is the exposure-side + psfex_interp VALIDATION output, where a CCD whose model fails the + acceptance gate (star_selection_psf.psf_acceptance_thresholds) produces + nothing and the unit only warns; the MCCD chain, never run in a + campaign, warns on every runner. Science-path PSF rejection does not go + through this table: psfex_interp drops the epoch per object inside the + tile run. + Anchor: workflow/scripts/completeness.py::COMPLETENESS. + default: exact_counts options: - count_floor: - label: Count-floor table (expect/floor per runner; fail below floor) + exact_counts: + label: Nominal count per runner; psfex_interp validation shortfall warns insights: [des_psf_blacklist, guinot22_star_floor_22] - all_or_nothing: - label: Every expected sub-product required + count_floor: + label: Tolerate recorded attrition down to a per-runner floor excluded: true excluded_reason: >- - Legitimately-absent CCDs would fail whole exposures and poison their - downstream cone; Snakemake has no optional-output primitive; field - precedent is tolerated, recorded attrition. - no_floor: - label: Accept whatever is produced, no gate + The floors had no basis: across a 127-exposure, 64-tile campaign + every non-warning runner produced exactly its nominal count, so a + floor below it only admits failed units unremarked. + no_gate: + label: Accept whatever is produced excluded: true excluded_reason: >- - Silent attrition — a stage producing 2 of 40 CCDs would flow into - the catalogue unremarked. + A stage producing 2 of 40 CCDs would flow into the catalogue + unremarked. postage_stamp_size: - label: Postage-stamp size, 51 px everywhere + label: Postage-stamp size shared by vignets, ngmix stamps and PSF models rationale: >- - One number pins three coupled apertures: the SExtractor vignet cut - around each detection (VIGNET(51,51) in default_noimaflags.param / - default.param, VIGNET_SIZE=51 in the dormant external-catalogue path, - example/cfis/config_tile_Uc.ini), the vignetmaker - stamps that feed ngmix (STAMP_SIZE=51 in config_tile_PiViVi.ini, both - runs; nearest-pixel centring, no sub-pixel interpolation in - VignetMaker._get_stamp), and the PSFEx model stamp (PSF_SIZE 51,51 in - default.psfex). The stamp IS the pixel data ngmix fits: it bounds - measurable galaxy size and truncates the wings of large galaxies. - Rationale for 51 not recorded in code. + One number, 51 px (about 9.5 arcsec), pins three coupled apertures: + the SExtractor VIGNET(51,51) around each detection, the vignetmaker + stamps that feed ngmix (STAMP_SIZE in both vignetmaker runs), and the + PSFEx model stamp (PSF_SIZE 51,51). The stamp is the pixel data ngmix + fits: it bounds the measurable galaxy size and truncates the wings of + large galaxies. No rationale for 51 is recorded. Anchor: workflow/config/cfis/default_noimaflags.param#VIGNET; - example/cfis/config_tile_Uc.ini#READ_EXT_SEXCAT_RUNNER.VIGNET_SIZE; - workflow/config/cfis/config_tile_PiViVi.ini#VIGNETMAKER_RUNNER_RUN_1.STAMP_SIZE; - workflow/config/cfis/default.psfex#PSF_SIZE; - src/shapepipe/modules/vignetmaker_package/vignetmaker.py::VignetMaker._get_stamp. + workflow/config/cfis/config_tile_PiViVi_psfex.ini#VIGNETMAKER_RUNNER_RUN_1.STAMP_SIZE; + workflow/config/cfis/config_tile_PiViVi_psfex.ini#VIGNETMAKER_RUNNER_RUN_2.STAMP_SIZE; + workflow/config/cfis/default.psfex#PSF_SIZE. default: px_51 options: px_51: - label: 51x51 px (~9.5 arcsec at 0.187"/px) + label: 51x51 px everywhere larger_adaptive: label: Larger or size-adaptive stamps - excluded: true - excluded_reason: >- - Not wired; would need coupled changes in three places (a change in - any one alone desynchronises galaxy stamp, PSF stamp, and vignet). + description: >- + Not implemented; needs the vignet, stamp and PSF sizes changed + together, since changing one alone desynchronises them. photometric_zeropoint: - label: Magnitude zero-point convention, fixed 30.0 on tiles + label: Magnitude zero-point convention rationale: >- - Tiles use a hard-coded MAG_ZEROPOINT 30.0 for every tile - (default_tile.sex; ZP_FROM_HEADER=False in config_tile_Sx.ini), and - ngmix repeats it (MAG_ZP=30.0 in config_tile_Ng_template.ini). - Exposures instead read the per-image header zero-point - (ZP_FROM_HEADER=True, ZP_KEY=PHOTZP in config_exp_psfex.ini). The tile - convention leans on MegaPipe's calibrated stacks; the star-selection - magnitude window (18-22) and mask magnitude limits inherit whichever - convention their stage uses. SExtractorCaller.get_zero_point is the - header-reading path, unused on tiles. + Tiles use a fixed MAG_ZEROPOINT 30.0 (ZP_FROM_HEADER=False), and ngmix + repeats it in MAG_ZP; exposures read the per-image header PHOTZP. The + tile convention relies on MegaPipe stacks being calibrated to 30 by + construction. Magnitude cuts (the star-selection window, downstream + galaxy cuts) inherit whichever convention their stage uses. Anchor: workflow/config/cfis/default_tile.sex#MAG_ZEROPOINT; workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.ZP_FROM_HEADER; workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.MAG_ZP; @@ -170,44 +137,9 @@ decisions: label: Per-image header zero-points on tiles too excluded: true excluded_reason: >- - MegaPipe stacks are calibrated to ZP 30 by construction; per-tile - header reads add a failure path for no expected numerical change. - (If that claim is wrong, this is a real fork — verify.) - - baseline_validation_criterion: - label: Validation criterion against the v2.0 bash baseline - rationale: >- - Because shape_measurement.ngmix_seed_mode deliberately changes noise - streams, P1 validation against v2.0 is statistical parity - (population-level agreement), not bit parity. Everything upstream of - ngmix (through PSFEx) validated bit-exactly (P0: 4/4 PASS). This - defines the evidence standard for "the same pipeline" — surfaced to the - collaboration as open Q5 in PRD #848. - [PENDING #873] Run-to-run determinism, which is a different property - from parity with v2.0, is now complete. With the setools star split - seeded (star_selection_psf.psf_train_validation_split) the last unseeded - draw in the science chain is gone: two runs of this code over the same - inputs now produce the same PSF star sample, the same PSF models and - the same shapes, which they did not before. That also settles a tension - this record carried — the bit-parity claim above sat next to an - unseeded star split that could not have been bit-reproducible, and the - P0 exposure-stage comparison did see PSF-validation CCD attrition - differ between the two sides. Statistical rather than bit parity is - therefore demanded only against the v2.0 baseline, not between runs of - the current pipeline. - Anchor: workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.SEED_FROM_POSITION; - src/shapepipe/modules/ngmix_package/ngmix.py::position_seed; - src/shapepipe/modules/setools_package/setools.py::SETools._make_rand_split. - default: statistical_parity - options: - statistical_parity: - label: Population-level agreement in shear observables - bit_parity: - label: Bit-identical catalogues - excluded: true - excluded_reason: >- - Impossible by construction once the seed mode changed; requiring it - would freeze the chunk-dependent v2.0 RNG forever. + MegaPipe stacks are calibrated to 30, so a header read adds a failure + path for no numerical change; it becomes a real fork only if tiles + ever carry a different PHOTZP. prior_insights: des_psf_blacklist: @@ -222,7 +154,7 @@ prior_insights: doi: "10.48550/arXiv.2011.03409" quote: exact: "we enter it into a" - suffix: " \u201cblacklist\u201d and exclude this CCD" + suffix: " “blacklist” and exclude this CCD" location: { page: 10 } guinot22_star_floor_22: claim: >- @@ -238,320 +170,147 @@ prior_insights: exact: 'The dashed line represents the cut at 22 stars/CCD below which the CCD is discarded for the PSF estimation.' location: { page: 4 } -findings: - orchestration_parity: - claim: >- - The Snakemake orchestration reproduces the bash baseline bit-exactly - through PSFEx (P0 validation, 4/4 PASS on the 186/187 quad). - Read it with two caveats. It is a statement about the pre-#873 code: - both #873 changes move products (a different realised star split, a - stricter science-path star gate), so re-establishing parity would mean - regenerating the baseline under the current branch. And the parity is - bit-exact in the products compared, not everywhere: the P0 - exposure-stage comparison did see PSF-validation CCD attrition differ - between the two sides, which the then-unseeded star split explains - (see baseline_validation_criterion). - created_at: "2026-08-19T00:00:00Z" - evidence: - - id: ev_final_cat - artifact: final_cat - record_authoring_found_defects: - claim: >- - Nine places in the code disagree with themselves or with their - documentation, each carried as a [LINT] mark: the 22-vs-20 - STAR_THRESH mismatch between PSF validation and science interpolation; - the unseeded train/validation rand_split (setools.py:664 — the star - sample entering the PSF model is irreproducible run-to-run); additive - (non-bitwise) mask-plane combination, safe today only because the - committed flag values are disjoint; the dead MESSIER_PIXEL_SCALE config - key; final_cat.param requesting IMAFLAGS_ISO that the merged catalogue - never receives; the centroid_source default disagreement (runner "wcs" - vs module "hsm", latent for direct callers); setools logging a FWHM cut - (mode +- 0.1 px in arcsec) half the applied one (mode +- 0.2 px), and - mixing pixel scales 0.187/0.186 within one file; TILE_LIST - overlap-flagging documented but never implemented; the mccd_plots - module docstring advertising rho statistics that live downstream now. - Status: the first two are fixed. CosmoStat/shapepipe#873 seeds the - rand_split and raises the science-path STAR_THRESH to 22, and commit - 90782098 mirrors that threshold into the workflow's own committed - config fork. #873 is OPEN against develop; both fixes are live on - feat/snakemake-orchestration only, and the 34-tile smk-g4 campaign is - the first run under them. The other seven stand, including the - mccd_plots docstring that still advertises rho statistics the package - no longer computes. - created_at: "2026-08-29T00:00:00Z" - derived: true - evidence: - - id: ev_final_cat_defects - artifact: final_cat - code_paper_divergence: - claim: >- - The two ShapePipe papers state roughly 17 of this record's 50 decisions - (now carried as prior_insights with verbatim quotes), have drifted from - the code on 10 of them since publication, and are silent on the rest. - Among them: DETECT_MINAREA 10 -> 5; DEBLEND_MINCONT 0.001 -> - 0.0005 on tiles; tile background AUTO -> MANUAL 0; in-line spread-model - star/galaxy classification -> disabled and deferred downstream; HSM - moment initialisation -> WCS centroids and prior-based guesses; GSC 2.2 - via cdsclient -> GSC 2.3 via astroquery; PSF acceptance 22 stars/CCD - published for the science path vs 20 committed there — closed since by - #873 + 90782098, which put the science path on 22, so nine of the ten - drifts remain open on the orchestration branch. - created_at: "2026-08-29T00:00:00Z" - derived: true - evidence: - - id: ev_final_cat_divergence - artifact: final_cat - analyses: # ═════════════════════════════════════════════════════════════════════════ masking: description: >- - Which pixels are excluded before anything is measured. Modules: - src/shapepipe/modules/mask_package/mask.py (halo/spike/DSO/border - builders, WeightWatcher driver), scripts/python/create_star_cat.py - (star-catalogue fetch), configs config_exp_Ma.ini + - config_onthefly.mask / config_tile_onthefly.mask + mask_default/. - [LINT] MESSIER_PIXEL_SCALE is set in config_tile_onthefly.mask but - never read — mask_dso takes pixel scale from the WCS. [LINT] - _build_final_mask combines mask planes by ADDITION (mask.py:1141+), - not bitwise OR; the committed flag values (2/4/16/32/128) are disjoint - so no live collision exists, but any future duplicate value corrupts - the flag semantics silently. (FLAG_OUTFLAGS 2 in default.ww is inert: - no input flag image is passed to WeightWatcher — mask.py:1047-1074.) + Which masks reach the measurement, and where. ShapePipe generates no + masks. The instrument flag image delivered with each exposure is the + only mask that reaches pixels. Sky-fixed masks (star halos and bodies, + manual regions, missing bands) are healsparse maps built outside + ShapePipe; their geometry is decided there. Inside ShapePipe they are + only queried at object positions into catalogue columns, and no stage + cuts on those columns. inputs: - - id: ccd_images + - id: exposure_flags type: data - source: split per-CCD exposure images + weights + CFIS flag maps (exp_split family) - - id: star_catalogue + source: per-CCD instrument flag images split from each exposure (exp_split family) + - id: sky_masks type: data - source: GSC 2.3 per-exposure catalogues (exp_star_cat cache) + source: UNIONS healsparse mask maps outputs: - - id: exposure_mask + - id: masked_measurement_inputs type: data format: fits - description: Per-CCD pipeline flag maps (run_sp_exp_Ma family). - decisions: - [star_catalogue_query, star_magnitude_definition, - bright_star_mask_geometry, deep_sky_object_masking, - border_mask_width, pixel_threshold_flags, external_flag_usage] + description: >- + Exposure detection catalogues carrying IMAFLAGS_ISO (and MASK_EXT when + maps are configured), the flag stamps ngmix reads, and the final + catalogue's optional MASK_ columns. + inputs: [exposure_flags, sky_masks] + decisions: [pixel_mask_source, psf_star_mask_veto, sky_mask_application] decisions: - star_catalogue_query: - label: Reference star catalogue and query for bright-star masking + pixel_mask_source: + label: Only the instrument flag image reaches pixels rationale: >- - GSC 2.3 (Vizier I/305/out), columns GSC2.3/RAJ2000/DEJ2000/Fmag/ - jmag/Vmag/Nmag/Class, cone radius covering the full CCD mosaic, no - magnitude cut at query time. GSC 2.2 rejected in a code comment - beside the catalogue ID ("does not have Fmag"). - Provenance hazard: the star-cat cache is not keyed by script - version — a semantic change to this query reruns the rule but takes - the skip-if-exists branch; clear the cache by hand for the change - to reach the data (workflow/config.yaml star_cats comment). - Query geometry: search radius = half the image diagonal about the - field centre (Mask._get_image_radius); source precedence: with - CDSCLIENT_PATH set in the .mask configs the online-query branch - wins unless an external star cat is passed (USE_EXT_STAR=True in - config_exp_Ma.ini routes the exp_star_cat cache in). - Published description (Farrens+22 p.2): cdsclient downloads GSC 2.2 - (with cdsclient 3.84 pinned in its Table A.1); current code: GSC 2.3 - queried through astroquery — two things drifted, the catalogue - version (Fmag is needed for the magnitude cut) and the query - transport, since cdsclient is never invoked yet survives as a - required-but-unused CDSCLIENT_PATH still set to the stale - findgsc2.2 in config_tile_onthefly.mask. - Anchor: scripts/python/create_star_cat.py::CDS_CAT_ID; - src/shapepipe/modules/mask_package/mask.py::Mask._CDS_cat_ID; - src/shapepipe/modules/mask_package/mask.py::Mask._cds_keys; - workflow/config.yaml. - default: gsc_23_vizier + On exposures SExtractor reads the split flag image (FLAG_IMAGE=True), + producing IMAFLAGS_ISO, which the PSF star selection requires to be + zero. The multi-epoch vignet run cuts flag stamps from the same split + flag image, and ngmix gives weight 0 to every flagged pixel and drops + epochs that are mostly flagged (shape_measurement.defect_fill, + shape_measurement.epoch_masked_fraction_cut). Tiles have no flag + image, so tile detection runs unflagged + (detection.detection_source_mode). Sky-fixed masks never touch + pixels: an object inside a star halo is measured from the same + pixels as one outside it. + Anchor: workflow/config/cfis/config_exp_psfex.ini#SEXTRACTOR_RUNNER.FLAG_IMAGE; + workflow/config/cfis/config_tile_PiViVi_psfex.ini#VIGNETMAKER_RUNNER_RUN_2.ME_IMAGE_PATTERN; + src/shapepipe/modules/ngmix_package/ngmix.py::prepare_ngmix_weights. + default: instrument_flags_only options: - gsc_23_vizier: - label: GSC 2.3 cone queries, all bands, no query-time mag cut - insights: [farrens22_star_cat_on_disk] - gaia: - label: Gaia-based star catalogue + instrument_flags_only: + label: Instrument flags gate pixels; sky masks stay at catalogue level + rasterised_sky_masks: + label: Rasterise healsparse masks into the pixel flags + insights: [farrens22_pipeline_masks] excluded: true excluded_reason: >- - Not wired. Deeper and better photometry; switching changes mask - geometry and hence the selection function — a real DR-level fork. - star_magnitude_definition: - label: Per-star magnitude for mask scaling + Sky-fixed masks say where an object sits, not that its pixels are + corrupted, so what to do about them is an analysis decision. + Rasterising them would bake one mask version into every shape; + catalogue columns leave the choice downstream. + psf_star_mask_veto: + label: PSF-star candidates rejected on instrument flags only rationale: >- - mag = unweighted mean of the finite GSC bands among F, j, V, N; - stars with no finite band are logged and not masked; only Class==0 - objects masked. Comment records why not a naive mean: NaN bands - would NaN-poison the mag < mag_limit test, leaving exactly the - bright stars with incomplete photometry unmasked. - Anchor: src/shapepipe/modules/mask_package/mask.py::Mask._create_mask. - default: mean_finite_bands + The star selection cuts IMAFLAGS_ISO == 0 and nothing else from the + masks. mask_query sits in the exposure module chain between + SExtractor and setools: when MASK_PATHS names maps it writes MASK_EXT + (0 clean, nonzero flagged; off-coverage counts as clean) onto each + CCD's catalogue. MASK_PATHS ships commented out, so the committed + module passes the catalogue through with no MASK_EXT column. The + intended map is the UNIONS star-body product (bit 2); halo bits 0 + and 1 are excluded because halos say nothing about whether a star is + a good PSF sample. Imposing the veto is one line per mask block in + star_selection.setools (MASK_EXT == 0). [LINT] the + workflow/rules/exposure.smk docstring names the column FLAG_EXT; the + code writes MASK_EXT. + Anchor: workflow/config/cfis/config_exp_psfex.ini; + workflow/config/cfis/star_selection.setools#MASK:star_selection.IMAFLAGS_ISO; + src/shapepipe/modules/mask_query_runner.py::mask_query_runner; + src/shapepipe/utilities/mask_query.py::flag_positions. + default: instrument_flags_only options: - mean_finite_bands: { label: Mean of finite F/j/V/N; Class==0 only } - single_band: - label: Single-band (Fmag) magnitude - excluded: true - excluded_reason: Drops stars with missing Fmag from masking entirely. - bright_star_mask_geometry: - label: Halo + diffraction-spike mask geometry and magnitude scaling - rationale: >- - DS9 polygon templates scaled linearly with magnitude about a pivot: - halo HALO_MAG_LIM=13, HALO_SCALE_FACTOR=0.05, HALO_MAG_PIVOT=13.8 - (halo_mask.reg, ~270 px); spike SPIKE_MAG_LIM=18, - SPIKE_SCALE_FACTOR=0.3, SPIKE_MAG_PIVOT=13.8 - (MEGAPRIME_star_i_13.8.reg); scaling = 1 - factor*(mag-pivot), - floored at 0.1 by Mask._scaling_min. - Identical in exposure and tile configs. Template filename encodes - provenance (MegaPrime i-band mag-13.8 star); numeric rationale not - recorded. Note the 5-mag gap: stars in 13-18 get spikes but no halo. - Anchor: workflow/config/cfis/config_onthefly.mask#HALO_PARAMETERS.HALO_MAG_LIM; - workflow/config/cfis/config_onthefly.mask#SPIKE_PARAMETERS.SPIKE_MAG_LIM; - workflow/config/cfis/mask_default/halo_mask.reg; - workflow/config/cfis/mask_default/MEGAPRIME_star_i_13.8.reg; - src/shapepipe/modules/mask_package/mask.py::Mask._create_mask; - src/shapepipe/modules/mask_package/mask.py::Mask._scaling_min. - default: megaprime_polygon_linear_scaling - options: - megaprime_polygon_linear_scaling: - label: Fixed MegaPrime templates, linear mag scaling, floor 0.1 - radial_profile_fit: - label: Per-star radial-profile-driven mask size - excluded: true - excluded_reason: Not wired; the survey precedent is template-based. - deep_sky_object_masking: - label: Messier + NGC objects masked as circles, no enlargement - rationale: >- - Circles of radius max(size_X, size_Y), MESSIER_SIZE_PLUS=0, - NGC_SIZE_PLUS=0 (function default is 0.1 — the 0 is a choice); - flags 16/32. A comment records the overlap-test fix (corner-only - test missed small interior objects). - Anchor: workflow/config/cfis/config_onthefly.mask#MESSIER_PARAMETERS.MESSIER_SIZE_PLUS; - workflow/config/cfis/config_tile_onthefly.mask#NGC_PARAMETERS.NGC_SIZE_PLUS; - src/shapepipe/modules/mask_package/mask.py::Mask.mask_dso. - default: circles_no_padding - options: - circles_no_padding: - label: "size_plus = 0: mask exactly the catalogued extent" - insights: [farrens22_messier_mask] - padded_circles: - label: size_plus > 0 (code default 0.1) + instrument_flags_only: + label: IMAFLAGS_ISO == 0; no sky map queried + star_body_veto: + label: Also reject candidates on the star-body map (MASK_EXT == 0) + description: >- + Set MASK_PATHS to the star-body map and add MASK_EXT == 0 beside + each IMAFLAGS_ISO cut. Querying without the cut records MASK_EXT + and changes no star. + star_body_and_halo_veto: + label: Also reject candidates inside star halos excluded: true excluded_reason: >- - Rationale for dropping the padding not recorded; flagged as a - question rather than an endorsed exclusion. - border_mask_width: - label: CCD border mask, 50 px on exposures, none on tiles - rationale: >- - Exposures BORDER_WIDTH=50 (flag 4); tiles BORDER_MAKE=False. - Mask.mask_border's own default is 100 — the committed 50 is a - choice, unrecorded. Trims CCD edges where PSF and astrometry - degrade; changes the effective footprint. - Anchor: workflow/config/cfis/config_onthefly.mask#BORDER_PARAMETERS.BORDER_WIDTH; - workflow/config/cfis/config_tile_onthefly.mask#BORDER_PARAMETERS.BORDER_MAKE; - src/shapepipe/modules/mask_package/mask.py::Mask.mask_border. - default: px50_exposures_only - options: - px50_exposures_only: - label: 50 px exposure borders; tiles unmasked - insights: [farrens22_border_mask] - px100: - label: 100 px (module default) - excluded: true - excluded_reason: Halves usable edge area for no recorded gain. - pixel_threshold_flags: - label: WeightWatcher weight/flag thresholds into mask bits - rationale: >- - WEIGHT_MIN 0, WEIGHT_MAX 1000, WEIGHT_OUTFLAGS 1; FLAG_MASKS 0x01, - FLAG_OUTFLAGS 2; POLY_OUTWEIGHTS 0. Zero-weight and externally - flagged pixels excluded on these thresholds. Values are stock, not - derived from the CFIS weight distribution; rationale not recorded. - The FLAG_* keys are inert in the committed invocation — no flag - image is passed to WeightWatcher by Mask._exec_WW. - Anchor: workflow/config/cfis/mask_default/default.ww#WEIGHT_MIN; - workflow/config/cfis/mask_default/default.ww#FLAG_MASKS; - src/shapepipe/modules/mask_package/mask.py::Mask._exec_WW. - default: stock_ww_thresholds - options: - stock_ww_thresholds: - label: Stock WeightWatcher thresholds - insights: [farrens22_weightwatcher] - external_flag_usage: - label: CFIS external flag maps folded into exposure masks + Halos flag objects for the final catalogue; a star inside another + star's halo is not thereby a bad PSF sample. + sky_mask_application: + label: Object-level sky masking deferred downstream rationale: >- - USE_EXT_FLAG=True on exposures (imports CADC-provided bad-pixel / - cosmic-ray / trail flags); EF_MAKE=False on tiles. The external - plane enters via Mask._build_final_mask's path_external_flag branch. - Anchor: workflow/config/cfis/config_exp_Ma.ini#MASK_RUNNER.USE_EXT_FLAG; - workflow/config/cfis/config_tile_onthefly.mask#EXTERNAL_FLAG.EF_MAKE; - src/shapepipe/modules/mask_package/mask.py::Mask._build_final_mask. - default: exposures_only + The final catalogue ships every detected object. When MASK_EXT_PATHS + lists band:path pairs, make_cat queries each healsparse map at the + object's windowed position and writes one MASK_ column holding + the map value verbatim; an object off a map's coverage gets that + map's sentinel (False for boolean maps, which reads as unmasked; + typically -1 for integer maps). The committed make_cat config sets no + MASK_EXT_PATHS, so no mask column is written and every mask cut + happens downstream against the maps themselves. + Anchor: src/shapepipe/modules/make_cat_runner.py::make_cat_runner; + src/shapepipe/modules/make_cat_package/make_cat.py::save_mask_ext_data; + src/shapepipe/utilities/mask_query.py::query_map. + default: deferred_downstream options: - exposures_only: { label: "External flags on exposures, not tiles" } - ignore_external: - label: Pipeline-generated masks only + deferred_downstream: + label: No mask columns; all objects shipped + catalogue_columns: + label: Per-band MASK_ columns from MASK_EXT_PATHS, no cut + pipeline_cut: + label: Drop masked objects inside the pipeline excluded: true - excluded_reason: Discards upstream knowledge of bad pixels. + excluded_reason: >- + Location flags are analysis decisions; a pipeline cut would fix + one mask version into the catalogue. prior_insights: - farrens22_star_cat_on_disk: + farrens22_pipeline_masks: claim: >- - The ShapePipe release paper documents an on-disk star catalogue, in - GSC format, as a supported substitute for the online query, - motivated by compute nodes without internet access. + The published ShapePipe pipeline generated its own masks, including + Messier objects and CCD borders, and applied them to the images; the + current pipeline generates none. created_at: "2022-06-01T00:00:00Z" evidence: - - id: ev_farrens22_star_cat_disk - doi: "10.48550/arXiv.2206.14689" - quote: - exact: 'Alternatively, a star catalogue available on disk (with the same format as the GSC) can also be used' - location: { page: 2 } - farrens22_messier_mask: - claim: >- - Messier objects are named in the published masking procedure as one - of the object classes ShapePipe masks. - created_at: "2022-06-01T00:00:00Z" - evidence: - - id: ev_farrens22_messier + - id: ev_farrens22_masks doi: "10.48550/arXiv.2206.14689" quote: exact: 'Messier objects, and border regions.' location: { page: 2 } - farrens22_border_mask: - claim: >- - CCD border regions are named in the published masking procedure as - one of the regions ShapePipe masks. - created_at: "2022-06-01T00:00:00Z" - evidence: - - id: ev_farrens22_border - doi: "10.48550/arXiv.2206.14689" - quote: - exact: 'Messier objects, and border regions.' - location: { page: 2 } - farrens22_weightwatcher: - claim: >- - The published pipeline generates the mask image itself with - WeightWatcher (Marmo & Bertin 2008), fixing the tool but none of its - threshold values. - created_at: "2022-06-01T00:00:00Z" - evidence: - - id: ev_farrens22_ww - doi: "10.48550/arXiv.2206.14689" - location: { page: 2 } # ═════════════════════════════════════════════════════════════════════════ detection: description: >- - Object detection on r-band tiles (single-image mode) and exposures (for - star finding). Module: src/shapepipe/modules/sextractor_package/ - sextractor_script.py (config assembly, ZP/background overrides, - post-processing that assigns per-epoch CCD membership). Configs: - config_tile_Sx.ini + default_tile.sex + default.conv + - default_noimaflags.param (tiles); default_exp.sex (exposures — same - thresholds, but DEBLEND_MINCONT 0.001 vs tile 0.0005 and BACK_TYPE AUTO - vs tile MANUAL 0, both deliberate and unexplained divergences). - [LINT] final_cat.param (consumed by the post-proc merge_final_cat, - not by make_cat) requests IMAFLAGS_ISO, but the tile chain never - produces it (FLAG_IMAGE=False, default_noimaflags.param); the - exposure-side IMAFLAGS_ISO stays exposure-side (merge_starcat.py:807 - only). The merged catalogue never receives the column. + Object detection with SExtractor on r-band tiles (the galaxy sample) and + on single-exposure CCDs (PSF-star candidates). Tiles follow the MegaPipe + parameters of Gwyn's UNIONS tile catalogue; exposures keep ShapePipe's + stock values. inputs: - id: tile_stack type: data @@ -561,194 +320,204 @@ analyses: type: data format: fits description: Per-tile SExtractor LDAC catalogue with per-epoch CCD membership. + inputs: [tile_stack] decisions: [detection_threshold_policy, deblending_policy, background_model, - weighting_and_interpolation, detection_source_mode, + weight_map_usage, zero_weight_interpolation, detection_source_mode, epoch_membership_ccd_bounds, photometry_parameters, - cleaning_and_neighbour_masking] + spurious_detection_cleaning, blend_photometry_mask_type] decisions: - photometry_parameters: - label: Photometric aperture definitions — Kron parameters, apertures, half-light fraction - rationale: >- - PHOT_AUTOPARAMS 2.5,3.5 (Kron factor / minimum radius), - PHOT_APERTURES 5 px, PHOT_FLUXFRAC 0.5, BACKPHOTO_TYPE GLOBAL — - identical in both .sex files. MAG_AUTO is the axis of the - star-selection magnitude box AND the catalogue magnitude; FLUX_AUTO - is PSFEx's photometric normalisation (default.psfex PHOTFLUX_KEY). - A different Kron factor shifts magnitudes systematically, moving - which stars build the PSF model and every magnitude-based - downstream cut. Rationale not recorded (stock values). Anchor: - workflow/config/cfis/default_tile.sex#PHOT_AUTOPARAMS; - workflow/config/cfis/default_exp.sex#PHOT_AUTOPARAMS; - workflow/config/cfis/default.psfex#PHOTFLUX_KEY. - default: kron_25_35 - options: - kron_25_35: { label: "Kron 2.5/3.5, aperture 5 px, FLUXFRAC 0.5, global background" } - cleaning_and_neighbour_masking: - label: Spurious-detection cleaning and neighbour-pixel correction - rationale: >- - CLEAN Y with CLEAN_PARAM 1.0 deletes detections consistent with - being wings of a brighter neighbour — a post-deblend change to the - object list; MASK_TYPE CORRECT replaces neighbour pixels during - photometry (vs BLANK/NONE), changing fluxes and windowed moments - of blends. Identical in both .sex files; rationale not recorded. - Anchor: workflow/config/cfis/default_tile.sex#CLEAN; - workflow/config/cfis/default_tile.sex#MASK_TYPE; - workflow/config/cfis/default_exp.sex#CLEAN. - default: clean_1_correct - options: - clean_1_correct: { label: "CLEAN 1.0 + MASK_TYPE CORRECT" } detection_threshold_policy: label: Detection significance, minimum area, matched filter rationale: >- - DETECT_THRESH 1.5 sigma RELATIVE, ANALYSIS_THRESH 1.5, - DETECT_MINAREA 5, FILTER default.conv (3x3 pyramid kernel, "all - ground, FWHM = 2 pixels" — vs CFIS seeing ~0.65 arcsec = 3.5 px at - 0.187"/px, so the filter is not matched to the survey PSF). - Sets the faint end of the source sample. Rationale not recorded - (stock EB 2017 header). - Published description (Guinot+22 p.5, Table 2): DETECT_MINAREA 10; - current code: 5, in default_exp.sex as well as default_tile.sex — - the small-object end has been loosened since publication on both the - star-detection and tile-detection passes, while DETECT_THRESH 1.5 - RELATIVE and the 3x3 FWHM=2 px kernel still match. + Tiles: DETECT_THRESH and ANALYSIS_THRESH 1.0 sigma, DETECT_MINAREA + 3, filtered with a 7x7 Gaussian of FWHM 3 px (gauss_3.0_7x7.conv, + close to the ~3.5 px CFIS seeing). These are the parameters of the + MegaPipe tile catalogue, so ShapePipe's galaxy sample matches the + catalogue UNIONS adopts. Exposures: 1.5 sigma, minarea 5, the 3x3 + FWHM 2 px kernel (default.conv); they only feed star selection. + Guinot+22 Table 2 lists minarea 10 at 1.5 sigma with the FWHM 2 px + kernel, so the tiles differ from the paper on all three. Anchor: workflow/config/cfis/default_tile.sex#DETECT_THRESH; workflow/config/cfis/default_tile.sex#DETECT_MINAREA; - workflow/config/cfis/default.conv. - default: thresh_1p5_minarea5_fwhm2px_filter + workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.DOT_CONV_FILE; + workflow/config/cfis/gauss_3.0_7x7.conv; + workflow/config/cfis/default_exp.sex#DETECT_THRESH. + default: megapipe_tiles options: - thresh_1p5_minarea5_fwhm2px_filter: - label: 1.5 sigma, minarea 5, FWHM=2px kernel - seeing_matched_filter: - label: Kernel matched to CFIS seeing (~3.5 px) + megapipe_tiles: + label: MegaPipe values on tiles; stock values on exposures + stock_tiles: + label: Stock ShapePipe values on tiles (1.5 sigma, minarea 5, FWHM 2 px) excluded: true excluded_reason: >- - Not wired; would change depth and the faint-end selection - function — a real fork, excluded only as not-the-committed-path. + Triggers spuriously on about 10% of grid-placed Sersic galaxies + in image simulations, and does not match the MegaPipe tile + catalogue. deblending_policy: - label: Deblending sub-thresholds and contrast + label: Deblending contrast rationale: >- - DEBLEND_NTHRESH 32, DEBLEND_MINCONT 0.0005 on tiles (2x more - aggressive splitting than the exposure 0.001 and 10x more than the - SExtractor default 0.005 — divergences not documented), CLEAN Y - PARAM 1.0. Controls object count, centroids, and blend - contamination in shapes. - Published description (Guinot+22 p.5, Table 2, galaxy detection on - the stacked tiles): DEBLEND_MINCONT 0.001; current code: 0.0005 on - tiles, with only the exposure side still carrying 0.001 — the - divergence lands on precisely the configuration the paper documents. - NTHRESH 32 matches. + DEBLEND_NTHRESH 32 on both passes; DEBLEND_MINCONT 0.002 on tiles + (the MegaPipe value) and 0.001 on exposures. Contrast sets object + count, centroids, and blend contamination in shapes. Guinot+22 Table + 2 lists 0.001 for tile detection. Anchor: workflow/config/cfis/default_tile.sex#DEBLEND_MINCONT; workflow/config/cfis/default_exp.sex#DEBLEND_MINCONT. - default: mincont_5em4_tiles + default: megapipe_tiles options: - mincont_5em4_tiles: { label: "MINCONT 0.0005 tiles / 0.001 exposures" } + megapipe_tiles: + label: MINCONT 0.002 tiles / 0.001 exposures + mincont_5em4_tiles: + label: MINCONT 0.0005 on tiles + excluded: true + excluded_reason: >- + Part of the stock tile parameter set rejected in favour of the + MegaPipe values (see detection_threshold_policy). background_model: - label: Tile background fixed to zero, not estimated + label: Background estimation and photometric background rationale: >- - BACK_TYPE MANUAL, BACK_VALUE 0.0, BKG_FROM_HEADER=False on tiles — - trusts MegaPipe stack background removal; exposures use BACK_TYPE - AUTO (64/3 mesh). Residual sky offsets propagate into thresholds, - fluxes, completeness. Divergence deliberate, unexplained. - Published description (Guinot+22 p.5): Table 2's caption asserts all - non-tabulated SExtractor parameters keep their defaults, i.e. - BACK_TYPE AUTO, and the paper never mentions the background choice - at all; current code: BACK_TYPE MANUAL with BACK_VALUE 0.0 on tiles, - identically in workflow/ and example/ — the standing tile - configuration, not a one-off, diverging silently from the published - parametrisation. - Anchor: workflow/config/cfis/default_tile.sex#BACK_TYPE; - workflow/config/cfis/default_exp.sex#BACK_TYPE; + Both passes estimate the background with SExtractor AUTO. Tiles use + the MegaPipe mesh BACK_SIZE 512 with BACK_FILTERSIZE 9 and a LOCAL + photometric background (annulus BACKPHOTO_THICK 30); exposures use + mesh 64, filter 3 and a GLOBAL photometric background. The exposure + BACKGROUND and BACKGROUND_RMS maps are also what ngmix subtracts from + each epoch and weights its pixels by + (shape_measurement.galaxy_pixel_weights). The header background path + is off (BKG_FROM_HEADER=False). Residual sky offsets propagate into + thresholds, fluxes, completeness and shapes. + Anchor: workflow/config/cfis/default_tile.sex#BACK_SIZE; + workflow/config/cfis/default_tile.sex#BACKPHOTO_TYPE; + workflow/config/cfis/default_exp.sex#BACK_SIZE; workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.BKG_FROM_HEADER; src/shapepipe/modules/sextractor_package/sextractor_script.py::SExtractorCaller.get_background. - default: manual_zero_tiles_auto_exposures + default: auto_megapipe_tiles options: - manual_zero_tiles_auto_exposures: - label: Tiles trust the stack (0.0); exposures estimate - auto_everywhere: - label: SExtractor AUTO background on tiles too + auto_megapipe_tiles: + label: AUTO everywhere; MegaPipe mesh and LOCAL photometry on tiles + manual_zero_tiles: + label: Tile background fixed to 0, trusting the stack subtraction excluded: true excluded_reason: >- - Double-subtracts if MegaPipe already removed it; if MegaPipe - residuals are nonzero this exclusion is wrong — verify. - weighting_and_interpolation: - label: Weight-map usage and zero-weight pixel interpolation + Part of the stock tile parameter set rejected in favour of the + MegaPipe values (see detection_threshold_policy). + weight_map_usage: + label: Weight map as inverse variance for detection rationale: >- - Two settings depart from stock SExtractor: WEIGHT_TYPE MAP_WEIGHT - (default NONE) and INTERP_TYPE ALL (default NONE — SExtractor - invents flux across zero-weight pixels). The accompanying - RESCALE_WEIGHTS Y, WEIGHT_GAIN Y, MASK_TYPE CORRECT and - INTERP_MAXXLAG/INTERP_MAXYLAG 16 are the SExtractor defaults, so - they are settings the configs restate rather than choices. The - variance policy sets effective per-pixel SNR and thus the detection - set; INTERP_TYPE ALL alters pixel data feeding measurements. - Rationale not recorded. - Published description (Guinot+22 p.5): Table 2's "all other - parameters are kept to their default values" silently covers both - non-default settings; current code: MAP_WEIGHT + INTERP_TYPE ALL in - default_tile.sex and default_exp.sex alike — the paper gives no hint - that the weight map or the zero-weight interpolation is in play. + WEIGHT_TYPE MAP_WEIGHT on both passes (SExtractor default NONE): the + per-pixel variance sets the effective SNR and so the detection set. + RESCALE_WEIGHTS and WEIGHT_GAIN are SExtractor defaults. Guinot+22 + says all non-tabulated parameters keep their defaults, which would + mean no weight map. Anchor: workflow/config/cfis/default_tile.sex#WEIGHT_TYPE; - workflow/config/cfis/default_tile.sex#INTERP_TYPE; + workflow/config/cfis/default_exp.sex#WEIGHT_TYPE; src/shapepipe/modules/sextractor_package/sextractor_script.py::SExtractorCaller.set_input_files. - default: map_weight_interp_all + default: map_weight + options: + map_weight: + label: MAP_WEIGHT + no_weight: + label: No weight map (SExtractor default) + zero_weight_interpolation: + label: Interpolation across zero-weight pixels + rationale: >- + INTERP_TYPE ALL on both passes (SExtractor default NONE), with + INTERP_MAXXLAG/INTERP_MAXYLAG 16: SExtractor invents flux across + zero-weight pixels, which changes detections and photometry near + masked regions. No rationale is recorded. + Anchor: workflow/config/cfis/default_tile.sex#INTERP_TYPE; + workflow/config/cfis/default_exp.sex#INTERP_TYPE. + default: interp_all options: - map_weight_interp_all: { label: MAP_WEIGHT + INTERP ALL + MASK CORRECT } + interp_all: + label: INTERP_TYPE ALL no_interpolation: label: INTERP_TYPE NONE - excluded: true - excluded_reason: >- - Changes photometry near masks; the committed choice is itself - unjustified in code — flagged as a question, not an endorsement. + spurious_detection_cleaning: + label: Cleaning of spurious detections + rationale: >- + CLEAN Y with CLEAN_PARAM 1.0 on both passes deletes detections + consistent with being wings of a brighter neighbour, a post-deblend + change to the object list. Stock value; no rationale recorded. + Anchor: workflow/config/cfis/default_tile.sex#CLEAN_PARAM; + workflow/config/cfis/default_exp.sex#CLEAN_PARAM. + default: clean_1 + options: + clean_1: + label: CLEAN Y, CLEAN_PARAM 1.0 + blend_photometry_mask_type: + label: Neighbour pixels in blend photometry + rationale: >- + MASK_TYPE CORRECT on both passes replaces pixels belonging to a + neighbour by their mirror across the object centre during + photometry, changing fluxes and windowed moments of blends. Stock + value; no rationale recorded. + Anchor: workflow/config/cfis/default_tile.sex#MASK_TYPE; + workflow/config/cfis/default_exp.sex#MASK_TYPE. + default: correct + options: + correct: + label: MASK_TYPE CORRECT + blank: + label: MASK_TYPE BLANK + photometry_parameters: + label: Kron and aperture photometry definitions + rationale: >- + PHOT_AUTOPARAMS 2.5,3.5 (Kron factor, minimum radius), PHOT_APERTURES + 5 px and PHOT_FLUXFRAC 0.5 on both passes. MAG_AUTO is the axis of + the star-selection magnitude window and the catalogue magnitude; + FLUX_AUTO is PSFEx's photometric normalisation. A different Kron + factor shifts magnitudes and so every magnitude-based cut. Stock + values; no rationale recorded. + Anchor: workflow/config/cfis/default_tile.sex#PHOT_AUTOPARAMS; + workflow/config/cfis/default_exp.sex#PHOT_AUTOPARAMS; + workflow/config/cfis/default.psfex#PHOTFLUX_KEY. + default: kron_25_35 + options: + kron_25_35: + label: Kron 2.5/3.5, aperture 5 px, FLUXFRAC 0.5 detection_source_mode: - label: Single-image detection on the r-band tile + label: Single-image, unflagged detection on the r-band tile rationale: >- - DETECTION_IMAGE=False, FLAG_IMAGE=False at detection, - param file default_noimaflags.param. No dual-image mode, no - detection coadd, no flag propagation at detection time. The - sx_nomask variant is the committed chain because it matches the - validated bash baseline. STATUS: DELIBERATELY UNDECIDED (Cail, - 2026-08-29) — whether DR6 detects masked or unmasked is punted to - the planned masking-unification rework (not yet tracked in an issue); - the default records baseline - parity, not a settled methodological choice. The masked variant is - one config + one rule + a tile-side star-cat analogue away. + DETECTION_IMAGE=False and FLAG_IMAGE=False with the + default_noimaflags.param column list: tiles have no instrument flag + image and no detection coadd exists, so detection sees every tile + pixel and the tile catalogue carries no IMAFLAGS_ISO. [LINT] + final_cat.param, read by the post-processing merge, requests + IMAFLAGS_ISO, which the tile chain never produces. Anchor: workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.DETECTION_IMAGE; workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.FLAG_IMAGE; workflow/config/cfis/default_noimaflags.param; - workflow/rules/tile.smk. + workflow/config/cfis/final_cat.param#IMAFLAGS_ISO. default: sx_nomask_single_image options: sx_nomask_single_image: - label: Unmasked single-image r-band detection + label: Unflagged single-image r-band detection insights: [guinot22_stacked_detection] - sx_masked: - label: Detection on the masked tile + masked_tile: + label: Detection on a tile masked by rasterised sky masks + description: >- + Not implemented; no tile pixel mask exists (see + masking.pixel_mask_source). dual_image_coadd: label: Dual-image mode with a detection coadd - excluded: true - excluded_reason: No detection coadd exists in UNIONS r-band processing. + description: Not implemented; no UNIONS detection coadd exists. epoch_membership_ccd_bounds: label: Which exposure CCDs an object belongs to (N_EPOCH) rationale: >- - CCD_SIZE = 33,2080,1,4612 with strict inequalities — the 33-px left - trim silently discards a CCD strip from epoch membership; WCS - inversion failures skip the CCD ("no epoch recorded"), changing - N_EPOCH. Sets how many exposures contribute to each galaxy's - multi-epoch fit. Rationale beyond "number of pixels in a CCD" not - recorded. + CCD_SIZE = 33,2080,1,4612 with strict inequalities: the 33-px left + trim removes a CCD strip from epoch membership, and a WCS inversion + failure skips the CCD, lowering N_EPOCH. This sets how many exposures + enter each galaxy's multi-epoch fit. The trim is unexplained beyond + "number of pixels in a CCD". Anchor: workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.CCD_SIZE; src/shapepipe/modules/sextractor_package/sextractor_script.py::make_post_process; src/shapepipe/modules/sextractor_package/sextractor_script.py::ccd_candidate_mask. default: trimmed_bounds_33_2080 options: - trimmed_bounds_33_2080: { label: "x in (33,2080), y in (1,4612), strict" } + trimmed_bounds_33_2080: + label: "x in (33,2080), y in (1,4612), strict" full_ccd: label: Full 1-2048 x-range, inclusive bounds - excluded: true - excluded_reason: >- - The trim presumably excludes a bad edge region, but nothing in - code says so — flagged as a question. prior_insights: guinot22_stacked_detection: claim: >- @@ -766,12 +535,8 @@ analyses: # ═════════════════════════════════════════════════════════════════════════ preparation: description: >- - How pixels, WCS, and epoch membership are prepared before anything is - measured. Modules: - src/shapepipe/modules/split_exp_package/split_exp.py, - merge_headers_package/merge_headers.py, - find_exposures_package/find_exposures.py, - vignetmaker_package/vignetmaker.py. + How exposures, astrometry, epoch lists and stamps are prepared before + anything is measured. inputs: - id: exposure_files type: data @@ -779,120 +544,95 @@ analyses: outputs: - id: epoch_stamps type: data - format: fits + format: sqlite description: Per-object multi-epoch vignets + per-CCD WCS log feeding ngmix. + inputs: [exposure_files] decisions: [astrometric_solution_source, ccd_split_extent, epoch_provenance_from_tile_history, object_position_columns, - stamp_positioning_and_padding, epoch_flag_source] + stamp_positioning_and_padding] decisions: astrometric_solution_source: label: Astrometry taken verbatim from delivered per-CCD headers rationale: >- - split_exp builds WCS(h) from each raw CCD header at split time, - pickles it, and merge_headers writes the lot into - log_exp_headers.sqlite; every downstream world<->pixel transform - (stamp positioning, epoch membership, position seeding) uses that - stored solution. No re-derivation, no astrometric refinement — the - survey's delivered astrometry IS the pipeline's astrometry. - Alternative (a joint astrometric re-fit a la DES/Rubin) would move - every stamp centre and every position seed. Anchor: - src/shapepipe/modules/split_exp_package/split_exp.py::SplitExposures.create_hdus; - src/shapepipe/modules/merge_headers_package/merge_headers.py::merge_headers; - src/shapepipe/modules/vignetmaker_package/vignetmaker.py::VignetMaker._get_stamp_me. + split_exp builds WCS(header) from each raw CCD header, and + merge_headers stores the lot; every downstream world-to-pixel + transform (stamp placement, epoch membership, position seeding) uses + that solution. [HARDCODED] no re-derivation or astrometric refinement + exists. A joint re-fit would move every stamp centre and position + seed. + Anchor: src/shapepipe/modules/split_exp_package/split_exp.py::SplitExposures.create_hdus; + src/shapepipe/modules/merge_headers_package/merge_headers.py::merge_headers. default: delivered_headers options: delivered_headers: - label: "WCS(header) verbatim, stored at split time" + label: WCS(header) verbatim, stored at split time insights: [guinot22_gaia_astrometry] astrometric_refit: label: Joint astrometric re-solution - excluded: true - excluded_reason: Not wired; CFIS delivered astrometry is trusted. + description: Not implemented. ccd_split_extent: label: All 40 MegaCam HDUs split and carried as candidate epochs rationale: >- - N_HDU=40 with a hard check (any other HDU count raises) — every - CCD including the ear CCDs 36-39 is a candidate epoch wherever the - WCS lands it. The MegaCamFlip special-casing of 36/37 shows the - ears flow through shape measurement. Alternative: exclude ear CCDs - (different optical path/orientation history). Anchor: - workflow/config/cfis/config_exp_Sp.ini#SPLIT_EXP_RUNNER.N_HDU; + N_HDU=40, and any other HDU count raises: every CCD including the + ear CCDs 36-39 is a candidate epoch wherever the WCS lands it. + Excluding the ear CCDs (a different optical path) is the alternative. + Anchor: workflow/config/cfis/config_exp_Sp.ini#SPLIT_EXP_RUNNER.N_HDU; src/shapepipe/modules/split_exp_package/split_exp.py::SplitExposures.create_hdus. default: all_40_hdus options: all_40_hdus: - label: "40 HDUs, hard-fail on any other count" + label: 40 HDUs, hard-fail on any other count insights: [guinot22_forty_chips] + exclude_ear_ccds: + label: Drop CCDs 36-39 as epochs + description: Not implemented. epoch_provenance_from_tile_history: label: Epoch sets parsed from tile FITS HISTORY cards rationale: >- - A tile's contributing exposures are recovered by parsing column 3 - of each HISTORY line, stripping prefix "p", deduplicating — the - coadd's own provenance record is trusted as the epoch list. The - LSB s-prefix rename is present but commented out. A mis-parse - changes N_EPOCH and which exposures are fit. Anchor: - workflow/config/cfis/config_tile_Fe.ini#FIND_EXPOSURES_RUNNER.COLNUM; + A tile's contributing exposures are column 3 (COLNUM) of each HISTORY + line, with prefix p stripped and duplicates removed: the coadd's own + provenance is trusted as the epoch list. A mis-parse changes N_EPOCH + and which exposures are fit. + Anchor: workflow/config/cfis/config_tile_Fe.ini#FIND_EXPOSURES_RUNNER.COLNUM; src/shapepipe/modules/find_exposures_package/find_exposures.py::FindExposures.get_exposure_list. default: history_parse options: - history_parse: { label: "HISTORY column 3, prefix p, dedup" } + history_parse: + label: HISTORY column 3, prefix p, deduplicated object_position_columns: label: Windowed centroids (XWIN/YWIN) define every position rationale: >- - PSF interpolation sites, tile stamp centres, multi-epoch stamp - centres, and the catalogue sky position all use SExtractor's - windowed centroid — XWIN_WORLD/YWIN_WORLD on the tile side (SPHE), - XWIN_IMAGE/YWIN_IMAGE exposure-side (PIX). Windowed vs isophotal - vs model centroids differ systematically for blends and asymmetric - galaxies, and the centroid definition feeds the position seed. - Anchor: workflow/config/cfis/config_tile_PiViVi.ini#PSFEX_INTERP_RUNNER.POSITION_PARAMS; - workflow/config/cfis/config_tile_PiViVi.ini#VIGNETMAKER_RUNNER_RUN_2.POSITION_PARAMS; - workflow/config/cfis/config_exp_psfex.ini#POSITION_PARAMS. + PSF interpolation sites, tile and multi-epoch stamp centres, and the + catalogue position all use SExtractor's windowed centroid: + XWIN_WORLD/YWIN_WORLD on the tile side, XWIN_IMAGE/YWIN_IMAGE on + exposures. Windowed, isophotal and model centroids differ + systematically for blends and asymmetric galaxies, and the centroid + feeds the position seed and the centroid prior. + Anchor: workflow/config/cfis/config_tile_PiViVi_psfex.ini#PSFEX_INTERP_RUNNER.POSITION_PARAMS; + workflow/config/cfis/config_tile_PiViVi_psfex.ini#VIGNETMAKER_RUNNER_RUN_2.POSITION_PARAMS; + workflow/config/cfis/config_exp_psfex.ini#PSFEX_INTERP_RUNNER.POSITION_PARAMS. default: xwin_windowed options: - xwin_windowed: { label: Windowed centroids everywhere } + xwin_windowed: + label: Windowed centroids everywhere stamp_positioning_and_padding: - label: Nearest-pixel stamp centring; edge stamps zero-padded + label: Nearest-pixel stamp extraction with zero padding rationale: >- - Multi-epoch stamps are placed by round-tripping the tile world - position through the stored per-CCD WCS, then rounding to the - nearest pixel (no sub-pixel interpolation — the residual sub-pixel - offset is absorbed by the fit's centroid prior, cen sigma = 1 - pixel). Objects whose stamp overruns a CCD or tile edge are KEPT, - out-of-image pixels zero-filled (sf_tools FetchStamps - pad_mode='constant'); no boundary rejection exists — zero-padded - pixels enter the fit as data with whatever weight the padded - weight stamp carries. Anchor: - src/shapepipe/modules/vignetmaker_package/vignetmaker.py::VignetMaker._get_stamp; + [HARDCODED] stamps are cut around the pixel nearest the object's + position, with no sub-pixel interpolation; the sub-pixel remainder is + stored as the stamp's OFFSET, which ngmix uses as the Jacobian origin + (shape_measurement.centroid_source), so extraction and centroid + prior share one rounding. Multi-epoch stamps take the position from + the tile world coordinate through the stored per-CCD WCS. Objects + whose stamp overruns an image edge are kept, with out-of-image pixels + zero-filled; there is no boundary rejection. + Anchor: src/shapepipe/modules/vignetmaker_package/vignetmaker.py::get_stamps; src/shapepipe/modules/vignetmaker_package/vignetmaker.py::VignetMaker._get_stamp_me. default: round_and_zero_pad options: - round_and_zero_pad: { label: "Nearest-pixel + zero padding, no edge rejection" } - epoch_flag_source: - label: Per-epoch flag stamps come from RAW CFIS flags, not the pipeline mask - rationale: >- - The multi-epoch vignet run reads its flag stamps from - split_exp_runner output — the delivered instrumental flags — - while mask_runner's pipeline_flag (halos, spikes, DSOs, borders) - feeds only the exposure-side star finding - (config_exp_psfex.ini FILE_PATTERN pipeline_flag). Combined with - unmasked tile detection (detection.detection_source_mode), the - consequence is stark: THE BRIGHT-STAR MASKS CURRENTLY AFFECT ONLY - PSF-STAR SELECTION — neither the galaxy sample (no tile mask, no - IMAFLAGS cut possible) nor the pixels ngmix fits (raw flags only) - see them. Whether that is intended belongs to the - planned masking-unification rework, whose object-level half is - sp_validation's IMAFLAGS_ISO cut on a column this chain never - produces; this is the pixel-level half. Anchor: - workflow/config/cfis/config_tile_PiViVi.ini#VIGNETMAKER_RUNNER_RUN_2.ME_IMAGE_EXP_RUNNERS; - workflow/config/cfis/config_exp_psfex.ini#SEXTRACTOR_RUNNER.FILE_PATTERN; - workflow/config/cfis/config_exp_Ma.ini#MASK_RUNNER.PREFIX. - default: raw_flags - options: - raw_flags: { label: split_exp raw flags gate epoch pixels } - pipeline_flags: - label: pipeline_flag (incl. bright-star masks) gates epoch pixels + round_and_zero_pad: + label: Nearest-pixel extraction, zero padding, no edge rejection prior_insights: guinot22_gaia_astrometry: claim: >- @@ -921,17 +661,8 @@ analyses: # ═════════════════════════════════════════════════════════════════════════ star_selection_psf: description: >- - Which objects constrain the PSF, and the PSF model itself. Modules: - src/shapepipe/pipeline/str_handler.py (_mode — the iterative - histogram-zoom FWHM mode estimator centring the star box; median - fallback below N=20), src/shapepipe/modules/setools_package/setools.py - (_make_rand_split), src/shapepipe/modules/psfex_interp_package/ - psfex_interp.py (acceptance gates, HSM shapes). Configs: - star_selection.setools, default.psfex, config_exp_psfex.ini. - [LINT] star_stat logs the FWHM cut as mode +- 0.1*0.187 while the mask - applies mode +- 0.2 px — the run's own log misstates the selection. - [LINT] pixel scale appears as 0.187 (load-bearing) and 0.186 - (plot-only) in the same setools file. + Which objects constrain the PSF, the PSF model itself, and which CCD + models are good enough to use. inputs: - id: exposure_sexcat type: data @@ -941,131 +672,109 @@ analyses: type: data format: psf description: >- - Per-CCD PSFEx models + interpolated PSFs at object positions - (run_sp_exp_SxSePsfPi family), with HSM shape diagnostics. + Per-CCD PSF models and the PSFs interpolated at object positions, + with HSM shape diagnostics. + inputs: [exposure_sexcat] decisions: [star_selection_box, psf_train_validation_split, psfex_candidate_vetting, psf_modelling_software, psf_model_complexity, psf_acceptance_thresholds] decisions: star_selection_box: - label: Stellar-locus selection — mag window + FWHM window around the mode + label: Stellar-locus selection, magnitude window and FWHM window around the mode rationale: >- - 18 < MAG_AUTO < 22, |FWHM - mode| <= 0.2 px, FLAGS==0, - IMAFLAGS_ISO==0; the mode is computed on a preselection - (MAG_AUTO<21, 0.3-1.5 arcsec at 0.187"/px) via the iterative - histogram-zoom estimator (str_handler.py::_mode, eps=0.001; median - fallback for N<20, -1 for N=0 — small-N behaviour changes selection - on sparse CCDs). PSFEx's automatic FWHM-range selection is off - (SAMPLE_AUTOSELECT N) and bad-pixel filtering is off, but PSFEx's - compiled-in fixed sample cuts still apply on top of this box — - see psfex_candidate_vetting. + 18 < MAG_AUTO < 22, |FWHM - mode| <= 0.2 px, FLAGS == 0 and + IMAFLAGS_ISO == 0. The mode is computed on a preselection (MAG_AUTO < + 21, FWHM 0.3-1.5 arcsec at 0.187 arcsec/px) by an iterative + histogram-zoom estimator that falls back to the median below 20 + objects, so small-N behaviour changes selection on sparse CCDs. + PSFEx's own selection is off (SAMPLE_AUTOSELECT N), but its + compiled-in cuts still apply (psfex_candidate_vetting). [LINT] the + file's statistics log the FWHM cut as mode +- 0.1 px and its plot + uses 0.186 arcsec/px, while the applied cut is +- 0.2 px at 0.187. Anchor: workflow/config/cfis/star_selection.setools#MASK:star_selection.MAG_AUTO; - workflow/config/cfis/star_selection.setools#MASK:preselect.MAG_AUTO; + workflow/config/cfis/star_selection.setools#MASK:preselect.FWHM_IMAGE; workflow/config/cfis/default.psfex#SAMPLE_AUTOSELECT; - src/shapepipe/pipeline/str_handler.py::_mode. + src/shapepipe/pipeline/str_handler.py::StrInterpreter._mode. default: mode_centred_box options: mode_centred_box: - label: FWHM-mode-centred box, +-0.2 px, mag 18-22, setools-only vetting + label: FWHM-mode-centred box, +-0.2 px, mag 18-22 insights: [guinot22_star_box] size_mag_locus_fit: label: Fitted size-magnitude stellar locus - excluded: true - excluded_reason: Not wired; the mode-box is the validated v2.0 selection. + description: Not implemented. psfex_autoselect: label: PSFEx SAMPLE_AUTOSELECT vetting on top + insights: [guinot22_psfex_preselection_off] excluded: true excluded_reason: >- - Deliberately disabled so selection lives in one place; rationale - not recorded in code. + Disabled so that the pipeline's own star selection is the only + one. psf_train_validation_split: - label: Random 80/20 star split — model fit vs held-out validation + label: Seeded 80/20 star split, model fit vs held-out validation rationale: >- - RAND_SPLIT ratio 20: star_split_ratio_80 fits the PSFEx model - (config_exp_psfex.ini FILE_PATTERN, and the tile multi-epoch - interpolation ME_DOT_PSF_PATTERN in config_tile_PiViVi.ini); - star_split_ratio_20 is the independent PSF-residual diagnostic - (PSFEX_INTERP MODE=VALIDATION). Trades model precision (fewer - training stars per CCD, interacting with the STAR_THRESH gate) - against an independent residual test. - [PENDING #873] The split is DETERMINISTIC: _make_rand_split takes - np.random.RandomState(seed).permutation(cat_size), the seed being - the digits of the unit's file number mod 2^32 — a pure function of - the input catalogue, fixed per CCD and independent of processing - order, the same philosophy as shape_measurement.ngmix_seed_mode's - SEED_FROM_POSITION. Before this the split drew from unseeded - np.random.randint, so the star sample entering the PSF model — and - therefore every shape downstream of it — differed between - identical runs; it was the one unseeded draw the position-seed work - left uncovered. One-off cost: the realised 80/20 membership changes - once (it is one further draw, now frozen), so PSF models and shapes - shift by that draw relative to every earlier product. + RAND_SPLIT RATIO 20: the 80% sample fits the PSFEx model and feeds the + tile multi-epoch interpolation (ME_DOT_PSF_PATTERN); the 20% sample + is the independent residual diagnostic (psfex_interp VALIDATION + mode). The split trades training stars per CCD, which interacts with + the acceptance gate, against an independent residual test. It is + deterministic: a permutation seeded from the unit's file number, so + the PSF star sample is a pure function of the input catalogue. Anchor: workflow/config/cfis/star_selection.setools#RAND_SPLIT:star_split.RATIO; src/shapepipe/modules/setools_package/setools.py::SETools._make_rand_split; workflow/config/cfis/config_exp_psfex.ini#PSFEX_RUNNER.FILE_PATTERN; - workflow/config/cfis/config_tile_PiViVi.ini#PSFEX_INTERP_RUNNER.ME_DOT_PSF_PATTERN. + workflow/config/cfis/config_tile_PiViVi_psfex.ini#PSFEX_INTERP_RUNNER.ME_DOT_PSF_PATTERN. default: split_80_20_seeded options: split_80_20_seeded: label: 80% train / 20% validation, seeded from the file number insights: [guinot22_star_split] split_80_20_unseeded: - label: Same split, unseeded np.random (pre-#873) + label: Same split from an unseeded random draw excluded: true excluded_reason: >- - Retired by #873: it made the PSF star sample — and every shape - downstream of it — irreproducible run-to-run, the single - remaining unseeded draw in the science chain. Kept on the - record because every UNIONS product built before the smk-g4 - campaign was produced under it. + Makes the PSF star sample, and every shape downstream of it, + irreproducible run-to-run. no_holdout: - label: 100% of stars in the model, no held-out diagnostic + label: All stars in the model, no held-out diagnostic excluded: true - excluded_reason: Loses the independent rho-statistic input. + excluded_reason: Loses the independent residual and rho-statistic input. psfex_candidate_vetting: - label: PSFEx-side candidate vetting — built-in defaults, unpinned + label: PSFEx built-in candidate cuts, unpinned rationale: >- - default.psfex sets only SAMPLE_AUTOSELECT N; SAMPLE_MINSN, - SAMPLE_MAXELLIP, SAMPLE_FWHMRANGE, SAMPLE_VARIABILITY are absent, - so PSFEx's compiled-in defaults apply silently (MINSN 20, - MAXELLIP 0.3, FWHMRANGE 2-10 px, VARIABILITY 0.2) — a second star - selection nobody's config records, and one that changes if the - PSFEx binary version changes. BADPIXEL_FILTER N + PSF_RECENTER N: - star vignets with flagged/sentinel pixels are accepted unfiltered - and candidates are not recentred (CENTER_KEYS XWIN). The setools - box is therefore not the whole selection. [HARDCODED] (in the - PSFEx binary). + default.psfex sets SAMPLE_AUTOSELECT N but omits SAMPLE_MINSN, + SAMPLE_MAXELLIP, SAMPLE_FWHMRANGE and SAMPLE_VARIABILITY, so + [HARDCODED] PSFEx's compiled-in defaults apply (MINSN 20, MAXELLIP + 0.3, FWHMRANGE 2-10 px, VARIABILITY 0.2): a second star selection no + config records, which changes with the PSFEx version. + BADPIXEL_FILTER N and PSF_RECENTER N accept flagged star vignets + unfiltered and do not recentre candidates. Anchor: workflow/config/cfis/default.psfex#SAMPLE_AUTOSELECT; workflow/config/cfis/default.psfex#BADPIXEL_FILTER; workflow/config/cfis/default.psfex#PSF_RECENTER. default: builtin_defaults options: builtin_defaults: - label: "Compiled-in MINSN 20 / MAXELLIP 0.3 / FWHMRANGE 2-10, no bad-pixel filter" - insights: [guinot22_psfex_preselection_off] + label: Compiled-in SAMPLE_* defaults, no bad-pixel filter pinned_explicit: label: Write the SAMPLE_* values explicitly into default.psfex psf_modelling_software: - label: PSF modelling software — PSFEx per-CCD vs MCCD focal-plane + label: PSF model, PSFEx per CCD or MCCD over the focal plane rationale: >- - The committed chain fits PSFEx independently per CCD. MCCD - (Liaudat+2021) is a maintained in-tree alternative: a focal-plane - model fit across all 40 CCDs at once with a hybrid local+global - decomposition (src/shapepipe/modules/mccd_package/ + six - mccd_*_runner.py; knobs in example/cfis/config_MCCD.ini — - N_COMP_LOC=8, D_COMP_GLOB=8, LOC_MODEL=hybrid, MIN_N_STARS=20, - RMSE_THRESH=1.25). Unwired in workflow/config/cfis/ (needs the - MCCD config adapted, and config_exp_mccd.ini carries a stale - hardcoded PSF_MODEL_DIR path). The image-simulation path - substitutes PSF modelling entirely: fake_psf_runner injects the - true input PSF from a SKiLLS dictionary in psfex_interp's output - format. - Anchor: src/shapepipe/modules/mccd_package; - src/shapepipe/modules/fake_psf_package; - example/cfis/config_MCCD.ini#INSTANCE.N_COMP_LOC; - example/cfis/config_MCCD.ini#INPUTS.MIN_N_STARS; - example/cfis/config_exp_mccd.ini. + workflow/config.yaml psf_model selects the exposure and tile config + pair; the committed value is psfex, which fits each CCD + independently. MCCD (Liaudat+2021) fits all 40 CCDs at once with a + hybrid local+global model (N_COMP_LOC 8, D_COMP_GLOB 8, MIN_N_STARS + 20, RMSE_THRESH 1.25); the completeness table treats its counts as + warnings because no campaign has run it. [LINT] config_exp_mccd.ini + still reads pipeline_flag images from mask_runner, which no longer + exists, so the MCCD exposure chain cannot run as committed. + Anchor: workflow/config.yaml; + workflow/config/cfis/config_MCCD.ini#INSTANCE.N_COMP_LOC; + workflow/config/cfis/config_MCCD.ini#INPUTS.MIN_N_STARS; + workflow/config/cfis/config_exp_mccd.ini#SEXTRACTOR_RUNNER.INPUT_MODULE; + src/shapepipe/modules/mccd_package. default: psfex options: psfex: @@ -1073,89 +782,57 @@ analyses: insights: [guinot22_psfex_software, farrens22_two_psf_methods] mccd_focal_plane: label: MCCD hybrid local+global focal-plane model - true_input_psf: - label: fake_psf injection of the simulation's true PSF - excluded: true - excluded_reason: >- - Only meaningful on simulated images where the true PSF exists; - not a data-analysis option. + insights: [farrens22_two_psf_methods] psf_model_complexity: - label: PSFEx model — pixel basis, degree-2 spatial polynomial per CCD + label: PSFEx pixel basis with degree-2 spatial variation per CCD rationale: >- - BASIS_TYPE PIXEL, BASIS_NUMBER 20, PSF_SIZE 51,51, PSF_SAMPLING 1, - PSFVAR_DEGREES 2 in XWIN,YWIN per CCD (MEF_TYPE INDEPENDENT, - STABILITY_TYPE EXPOSURE), PSF_RECENTER N. Model flexibility sets - the PSF-leakage/overfitting balance — the dominant additive - systematic in cosmic shear. Values are the stock EB 2017 header; - rationale not recorded in code. + BASIS_TYPE PIXEL, BASIS_NUMBER 20, PSF_SAMPLING 1, PSFVAR_DEGREES 2 + in XWIN/YWIN per CCD. Model flexibility sets the balance between PSF + leakage and overfitting, the dominant additive systematic in cosmic + shear. Stock values; no rationale recorded. Anchor: workflow/config/cfis/default.psfex#BASIS_TYPE; - workflow/config/cfis/default.psfex#PSFVAR_DEGREES; - workflow/config/cfis/default.psfex#PSF_SIZE. + workflow/config/cfis/default.psfex#BASIS_NUMBER; + workflow/config/cfis/default.psfex#PSFVAR_DEGREES. default: pixel_basis_deg2_per_ccd options: pixel_basis_deg2_per_ccd: - label: "PIXEL basis, degree 2, per-CCD" + label: PIXEL basis, degree 2, per CCD insights: [guinot22_psf_no_oversampling] deg3: label: Degree-3 spatial variation - excluded: true - excluded_reason: >- - More flexibility per CCD needs more stars per CCD than the - count-floor world guarantees; not validated. + description: Needs more stars per CCD than the acceptance gate guarantees. psf_acceptance_thresholds: - label: Per-CCD PSF-model quality gate (min stars, max chi2) + label: Per-CCD PSF-model quality gate rationale: >- - A CCD whose model has ACCEPTED < STAR_THRESH or CHI2 > 2 is not - interpolated — its galaxies drop from the shear catalogue: direct - footprint selection, the in-code analogue of the DES blacklist. - [PENDING #873] Both passes now gate at 22 stars: the VALIDATION-mode - exposure config always did (config_exp_psfex.ini), and #873 raised - the MULTI-EPOCH science path 20 -> 22 in example/cfis - (config_tile_PiViVi_canfar_{sx,uc}.ini), with commit 90782098 - mirroring it into workflow/config/cfis/config_tile_PiViVi.ini — the - committed config fork this workflow actually reads (#848 D2). - Provenance of the retired 20, which is what makes this a fix rather - than a preference: commit fdc86553 (Kilbinger, 2020-06-30, "Forgot - to update new star number threshold for 80% of stars") deliberately - bumped 20 -> 22 to account for the 80/20 split, but only in the - validation config; the tile config kept the pre-split 20, so for - five years the SCIENCE path gated on the value that 2020 fix meant - to retire. (20 is also the psfex_interp function default, so the - stale-value reading rested on the commit provenance rather than on - the config alone.) - Published description (Guinot+22 p.4, Fig. 3): 22 stars/CCD, applied - to exactly the CCDs feeding multi-epoch shape measurement — the - number now agrees. Two gaps remain: an undocumented CHI2_THRESH=2 in - both configs, and the mechanism — interpsfex tests the PSFEx header - ACCEPTED/CHI2 at interpolation time and drops that epoch for objects - on the CCD, rather than excluding the CCD from PSF modelling as the - paper describes. + A CCD whose model has ACCEPTED < STAR_THRESH = 22 or CHI2 > + CHI2_THRESH = 2 is not interpolated, on both the validation and the + multi-epoch pass; 22 applies the published floor to the 80% training + sample. In the science path the CCD's epoch is dropped for every + object on it; an object left with no epoch has no shape. There is no + minimum-epoch floor in the pipeline: NGMIX_N_EPOCH records what + survived and epoch-count cuts happen downstream. Guinot+22 describes + excluding the CCD from PSF modelling rather than gating at + interpolation, and does not state the chi2 cut. Anchor: src/shapepipe/modules/psfex_interp_package/psfex_interp.py::PSFExInterpolator.interpsfex; workflow/config/cfis/config_exp_psfex.ini#PSFEX_INTERP_RUNNER.STAR_THRESH; - workflow/config/cfis/config_tile_PiViVi.ini#PSFEX_INTERP_RUNNER.STAR_THRESH; - example/cfis/config_tile_PiViVi_canfar_sx.ini#PSFEX_INTERP_RUNNER.STAR_THRESH; - example/cfis/config_tile_PiViVi_canfar_uc.ini#PSFEX_INTERP_RUNNER.STAR_THRESH. + workflow/config/cfis/config_tile_PiViVi_psfex.ini#PSFEX_INTERP_RUNNER.STAR_THRESH; + workflow/config/cfis/config_tile_PiViVi_psfex.ini#PSFEX_INTERP_RUNNER.CHI2_THRESH. default: stars22_chi2_2 options: stars22_chi2_2: - label: ">= 22 stars on both passes, chi2 <= 2" + label: ">= 22 stars and chi2 <= 2 on both passes" insights: [des_psf_blacklist_local, guinot22_star_floor_22_local] - stars20_chi2_2: - label: ">= 20 stars on the science path, 22 in validation (pre-#873)" + stars20_science_path: + label: ">= 20 stars on the science path" excluded: true excluded_reason: >- - Retired by #873 + 90782098. It was never a chosen value: it is - the pre-split threshold fdc86553 raised to 22 in 2020 for the - validation config and forgot on the science path, leaving the - science gate below both the published floor (Guinot+22 Fig. 3) - and the pipeline's own intent. Every UNIONS product built before - the smk-g4 campaign carries it. + 20 is the pre-split floor; with 80% of stars in the model it gates + below both the published floor and the validation pass. des_25: label: DES Y3 threshold (25 stars) excluded: true excluded_reason: >- - Not adopted; CFIS CCDs are smaller than DECam's — the right - number is survey-specific. + CFIS CCDs are smaller than DECam's; the floor is survey-specific. prior_insights: des_psf_blacklist_local: claim: >- @@ -1171,8 +848,7 @@ analyses: guinot22_star_floor_22_local: claim: >- The published ShapePipe/UNIONS analysis discards a CCD from the PSF - estimation when fewer than 22 stars are selected on it — the floor - the science-path PSF-interpolation gate now applies. + estimation when fewer than 22 stars are selected on it. created_at: "2022-04-01T00:00:00Z" evidence: - id: ev_guinot22_star_floor_local @@ -1253,244 +929,431 @@ analyses: # ═════════════════════════════════════════════════════════════════════════ shape_measurement: description: >- - Galaxy shape estimation: ngmix single-Gaussian fits with metacalibration. - Modules: src/shapepipe/modules/ngmix_package/ngmix.py (priors, metacal - setup, epoch handling, postage-stamp prep), ngmix_runner.py (config - exposure). Config: config_tile_Ng_template.ini. Most values here are - HARDCODED — scientific choices living in code with no config exposure; - this sub-analysis is where the silent-default risk concentrates. - [LINT] centroid_source default disagrees between the runner ("wcs", - production; always passed explicitly, ngmix_runner.py:170) and every - module-level signature ("hsm") — unreachable in the pipeline path, but - direct callers (tests, notebooks) silently get the other choice. - [LINT] pixel scale is 0.186 here (config PIXEL_SCALE) vs 0.187 in the - setools/masking configs — and star_selection.setools itself mixes - 0.187 (cuts) with 0.186 (SCATTER stat, :75). + Galaxy shape estimation: joint multi-epoch ngmix Gaussian fits with + metacalibration. Most choices here are fixed in code; this is where the + silent-default risk concentrates. inputs: - id: vignets type: data - source: 51x51 galaxy/weight/background-RMS vignets + interpolated PSFs (vignetmaker, psfex_interp) + source: multi-epoch galaxy, weight, flag, background and background-RMS vignets + interpolated PSFs outputs: - id: ngmix_cat type: data format: fits description: Per-tile metacal shear catalogue chunks (ngmix_runner family). + inputs: [vignets] decisions: - [ngmix_seed_mode, galaxy_model, fit_priors, metacal_scheme, - centroid_source, epoch_quality_and_weighting, noise_model, - psf_epoch_loss_policy, megacam_ccd_flip] + [ngmix_seed_mode, galaxy_model, fit_initialisation, fit_priors, + metacal_scheme, centroid_source, epoch_flux_rescaling, + psf_epoch_averaging, galaxy_pixel_weights, psf_likelihood_noise, + megacam_ccd_flip, defect_fill, blend_handling, + epoch_masked_fraction_cut] decisions: ngmix_seed_mode: - label: ngmix per-object RNG seeding + label: Per-object RNG seeded from sky position rationale: >- - Production historically seeded one RandomState from the tile ID, - consumed in object order — results depended on chunk boundaries. - The position seed (3-arcsec sky boxes + CCD offsets, zig-zag fold + - Cantor pairing mod 2^32; ngmix.py::position_seed) makes every - stream a function of sky position: chunk-invariant, - bit-reproducible, and metacal fixnoise counter-noise cancels across - image-simulation branches (ngmix#796). Consequence: chunking is - demoted to a pure throughput knob (reverting this decision - re-promotes it). Cost: noise streams change vs v2.0 — see - top-level baseline_validation_criterion. + Each object's RNG (noise realisations, guesses, priors) is seeded + from its position: [HARDCODED] 3-arcsec sky boxes offset by the + first epoch's CCD number, folded and Cantor-paired mod 2^32. Every + stream is a function of sky position, so results do not depend on + how a tile is chunked, and metacal's fixnoise counter-noise cancels + across image-simulation branches that share an object's box. Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::position_seed; - workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.SEED_FROM_POSITION; - src/shapepipe/modules/ngmix_runner.py::ngmix_runner. + src/shapepipe/modules/ngmix_package/ngmix.py::Ngmix.process. default: position_seed options: position_seed: label: Per-object seed from (ra, dec, ccd), 3-arcsec boxes tile_seed: - label: Tile-wide RandomState (v2.0) + label: One tile-wide RandomState consumed in object order excluded: true excluded_reason: >- - Chunk-dependent; retired outright (SEED_FROM_POSITION=False now - raises — ngmix_runner.py:110-116). + Makes every noise stream depend on chunk boundaries and object + order. galaxy_model: - label: Galaxy and PSF model — single Gaussian [HARDCODED] + label: Single-Gaussian galaxy and PSF models rationale: >- - ngmix.fitting.Fitter(model='gauss') for both galaxy and PSF - (ngmix.py::make_runners); guessers TPSFFluxAndPriorGuesser / - TFluxGuesser with T=0.25 and catalogue-flux guess, Runner ntry=5, - PSFRunner ntry=2 — with a non-convex likelihood, guess and retries - decide which objects converge (failed fits are NaN-filled with - flags, not raised). Rationale not recorded. Under metacal, model - bias largely cancels in the response, which is the standard defense - of 'gauss'; not stated in code. + [HARDCODED] ngmix Fitter(model='gauss') for both the galaxy and the + PSF. Under metacalibration, model bias largely cancels in the + response, which is the standard defence of the Gaussian; the code + does not state it. Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::make_runners. default: gauss options: gauss: - label: "Single Gaussian, T guess 0.25, ntry 5/2" + label: Single Gaussian insights: [guinot22_gaussian_model] exp_or_bdf: label: exp / bdf galaxy models - excluded: true - excluded_reason: >- - Slower, and metacal makes the gain marginal; not validated on - CFIS. + description: Not exposed; make_runners fixes the model. + fit_initialisation: + label: Fit guesses and retries + rationale: >- + [HARDCODED] TPSFFluxAndPriorGuesser (galaxy) and TFluxGuesser (PSF) + start from T = 0.25 and the catalogue flux; the galaxy runner retries + 5 times, the PSF runner twice. With a non-convex likelihood the guess + and retries decide which objects converge; a failed fit is flagged, + not raised. Guinot+22 initialised the whole guess vector from HSM + adaptive moments on each sheared image; the code no longer does. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::make_runners. + default: prior_guess_t025_ntry5_2 + options: + prior_guess_t025_ntry5_2: + label: T guess 0.25 + catalogue flux, ntry 5 / 2 + hsm_initialisation: + label: Guesses from HSM adaptive moments (Guinot+22) + description: Not implemented. fit_priors: - label: ngmix joint prior — GPriorBA(0.4), cen sigma = pixel scale, flat T/F [HARDCODED] + label: ngmix joint prior rationale: >- - Ellipticity GPriorBA sigma=0.4; centroid CenPrior sigma = one pixel - scale (0.186 arcsec, config PIXEL_SCALE — the coupling - sigma=pixel_scale is itself the hardcoded choice); flat T in - [-1, 1e3], flat F in [-100, 1e9] with negative support (bounds - decide which noisy fits survive vs rail). get_prior takes T/F range - arguments but no caller passes them. Prior width drives noise bias; - rationale not recorded. - Published description (Guinot+22 p.7): centroid sigma = pixel scale - ~0.187 arcsec, flat F in [-1e4, 1e9], flat half-light radius r50 in - [-10, 1e6] arcsec, ellipticity prior from Bernstein & Armstrong - (2014); current code: PIXEL_SCALE 0.186, flat F in [-100, 1e9], and - a flat prior on ngmix's second-moment size T in [-1, 1e3] rather - than on r50 — the prior families agree, the flux bound and pixel - scale have drifted, and the size prior is a different - parameterisation rather than a changed number. + [HARDCODED] ellipticity GPriorBA with sigma 0.4; flat T in [-1, 1e3] + and flat F in [-100, 1e9], with negative support (the bounds decide + which noisy fits survive and which rail); a centroid prior of width + one pixel scale, PIXEL_SCALE 0.186 arcsec (derived from the WCS when + the key is absent). Prior width drives noise bias; no rationale + recorded. Guinot+22 states a flat F in [-1e4, 1e9] and a flat r50 + prior rather than T. [LINT] the epoch stamps are exposure pixels, and + star selection uses 0.187 arcsec/px. Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::get_prior; workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.PIXEL_SCALE. default: gpriorba04_flat options: - gpriorba04_flat: { label: "GPriorBA 0.4 + flat T/F with negative support" } + gpriorba04_flat: + label: GPriorBA 0.4 + flat T/F with negative support nonneg_informative: label: Non-negative or informative T/F priors excluded: true excluded_reason: >- Truncating negative support biases the noshear ensemble mean; - metacal wants symmetric noise response. + metacal wants a symmetric noise response. metacal_scheme: - label: Metacalibration — 5 types, step 0.01, fitgauss reconv, fixnoise [HARDCODED] + label: Metacalibration protocol rationale: >- - types [noshear,1p,1m,2p,2m], step 0.01, psf='fitgauss' (runner - default; moves the metacal response directly — alternatives gauss/ - dilate/azgauss listed in the docstring), fixnoise=True, - use_noise_image=True, MetacalBootstrapper(ignore_failed_psf=True) - (changes which epochs enter the fit). No *_psf sheared types, so no - mcal_R_psf PSF-response term in the catalogue. fixnoise rationale - appears only in the position_seed docstring (counter-noise - cancellation). - Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::do_ngmix_metacal. + [HARDCODED] types noshear, 1p, 1m, 2p, 2m with step 0.01, fixnoise + with the noise image, and ignore_failed_psf (an epoch whose PSF fit + fails is dropped). The reconvolution kernel is METACAL_PSF, default + fitgauss, not set in the committed config; it moves the response + directly. No sheared-PSF types run, so the catalogue has no PSF + response term. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::do_ngmix_metacal; + src/shapepipe/modules/ngmix_runner.py::ngmix_runner. default: five_types_step001_fitgauss options: five_types_step001_fitgauss: - label: "noshear+1p/1m/2p/2m, step 0.01, fitgauss, fixnoise" + label: noshear+1p/1m/2p/2m, step 0.01, fitgauss, fixnoise insights: [guinot22_metacal_five_images] with_psf_response: label: Add sheared-PSF types for R_psf - excluded: true - excluded_reason: >- - Not wired; leakage is instead diagnosed via PSF_ORIG columns + - rho statistics downstream. + description: Not implemented. centroid_source: - label: Jacobian origin from WCS astrometry, not HSM moments + label: Jacobian origin at the coadd centroid rationale: >- - Production runner default "wcs"; hsm is "legacy... noisy for stars - and flagged as incorrect by Fabian — see #767" (runner comment; a - rare recorded rationale). Moves the centroid-prior centre per - object. The runner reads an optional CENTROID_SOURCE config option - that no committed CFIS config sets. [LINT] module-level default is - still "hsm" — see this sub-analysis's description. - Published description (Guinot+22 p.7): HSM adaptive moments, run on - each sheared version, supplied the whole initial guess vector - (centroid, r50, flux) for the least-squares fit; current code: that - initialisation is gone — guesses come from ngmix's - TPSFFluxAndPriorGuesser with fixed T=0.25 and a catalogue flux, and - the only surviving HSM role is the optional stamp re-centering that - sets the Jacobian origin. So the drift is wider than a swapped - centroid source. (The paper's other HSM use, PSF/star shape - diagnostics, is unaffected.) + The Jacobian origin, where the centroid prior centres, is the + sub-pixel offset the stamp extractor stored when it cut the stamp + ("wcs", the default at every level; the committed config sets no + CENTROID_SOURCE). One projection and one rounding serve both + extraction and prior, so they cannot disagree near a rounding tie. + "hsm" re-centres on adaptive moments measured from the stamp. Anchor: src/shapepipe/modules/ngmix_runner.py::ngmix_runner; src/shapepipe/modules/ngmix_package/ngmix.py::make_ngmix_observation. default: wcs options: - wcs: { label: WCS-projected catalogue position } + wcs: + label: Coadd-centroid offset from the stamp extractor hsm: label: HSM adaptive-moment centroid excluded: true - excluded_reason: Noisy for stars; flagged incorrect (shapepipe#767). - epoch_quality_and_weighting: - label: Epoch admission, masking cut, and multi-epoch combination [HARDCODED] + excluded_reason: >- + Noisy, notably for stars, and follows the light rather than the + astrometry that placed the stamp. + epoch_flux_rescaling: + label: Epochs put on a common flux scale by header FSCALE rationale: >- - An epoch is dropped if >1/3 of its stamp is masked - (prepare_postage_stamps; the comment says "objects", the code drops - epochs — an object with zero surviving epochs drops out); failed - PSF fits drop epochs (flags != 0); fluxes rescaled by header FSCALE - (gal*Fscale, weight/Fscale^2); the diagnostic PSF is averaged over - epochs weighted by obs.weight.sum(). Joint multi-epoch fit over - survivors. Rationale for 1/3 and for the weight choice not - recorded. - Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::prepare_postage_stamps; - src/shapepipe/modules/ngmix_package/ngmix.py::rescale_epoch_fluxes; - src/shapepipe/modules/ngmix_package/ngmix.py::_average_psf_fits. - default: third_masked_cut + [HARDCODED] each epoch's image is multiplied by its header FSCALE and + its weight divided by FSCALE squared (background RMS scaled with the + image) before the joint fit, so the epochs share the tile's + zero-point. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::rescale_epoch_fluxes. + default: fscale options: - third_masked_cut: { label: "Drop epoch if >1/3 masked; FSCALE rescale; weight-sum PSF average" } - noise_model: + fscale: + label: Rescale by header FSCALE + psf_epoch_averaging: + label: Catalogue PSF quantities averaged over epochs by galaxy weight + rationale: >- + [HARDCODED] the PSF shape and size written to the catalogue (the + original image PSF and the metacal reconvolution kernel) are averages + over epochs weighted by the summed galaxy inverse variance of each + epoch; epochs whose PSF fit failed are left out. These columns feed + PSF-leakage estimates downstream. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::_average_psf_fits; + src/shapepipe/modules/ngmix_package/ngmix.py::average_original_psf. + default: galaxy_weight_sum + options: + galaxy_weight_sum: + label: Weighted by summed galaxy inverse variance + galaxy_pixel_weights: label: Per-pixel inverse variance from background-RMS vignets rationale: >- - BKG_RMS_VIGNET_PATH set in the CFIS template: weight = - 1/bkg_rms^2 per pixel (all-or-nothing; missing file errors); - fallback scalar 1/sigma_mad^2. Masked pixels filled with Gaussian - noise at sig_noise. A scalar sigma "mis-reports errors and erodes - the inverse-variance advantage whenever the RMS map actually - varies" (recorded rationale, fixnoise bookkeeping). PSF observation - gets a flat weight from PSF_NOISE=1e-5 — hardcoded module constant, - validated 1e-4..1e-6 on the digital twin (#749/#774 comment); - without it the g-prior swamps the PSF likelihood. Per-epoch - background subtraction BKG_SUB=True (off only for sims). + With BKG_RMS_VIGNET_PATH set, each pixel's weight is 1/rms^2 from the + SExtractor background-RMS map (all or nothing; a missing file + raises), and the noise realisations use the same per-pixel RMS; the + fallback is a scalar 1/sigma_mad^2. A scalar sigma mis-reports errors + wherever the RMS varies. Each epoch is background-subtracted with the + SExtractor background vignet (BKG_SUB, on unless the key is set + False). Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::prepare_ngmix_weights; - src/shapepipe/modules/ngmix_package/ngmix.py::PSF_NOISE; src/shapepipe/modules/ngmix_package/ngmix.py::background_subtract; workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.BKG_RMS_VIGNET_PATH. default: rms_vignet_weights options: - rms_vignet_weights: { label: Per-pixel RMS-map weights + PSF_NOISE 1e-5 } + rms_vignet_weights: + label: Per-pixel background-RMS weights scalar_sigma_mad: label: Scalar sigma_mad per epoch excluded: true - excluded_reason: Mis-reports errors where the RMS map varies (recorded). - psf_epoch_loss_policy: - label: Object-level policy when CCDs fail PSF interpolation + excluded_reason: Mis-reports errors wherever the RMS map varies. + psf_likelihood_noise: + label: Flat PSF-observation weight rationale: >- - When k of ~40 CCDs fail the PSF acceptance gate (~5-6% attrition, - per-exposure clustered, matches the bash baseline — but measured - with the science gate at 20 stars, so [PENDING #873] at 22 it can - only rise, and smk-g4 is the first campaign to re-measure it), the - pipeline applies NO further quality gate: tiles complete, each - object records NGMIX_N_EPOCH, and sp report surfaces per-tile - epoch loss. - Object-level protection is delegated entirely to the validation - stage's epoch-count cut (sp_validation's galaxy selection cuts on N_EPOCH >= 1). Rationale (Cail, - 2026-08-29, PRD walk): epoch loss is a per-object depth effect - already recorded in the catalogue; gating at pipeline level would - fail whole tiles for a versionable catalogue decision. PRD #848's - open-questions section was removed accordingly. Per-tile epoch loss - is surfaced by the run report; NGMIX_N_EPOCH is the per-object - record. - Anchor: workflow/scripts/run_report.py; - workflow/config/cfis/final_cat.param#NGMIX_N_EPOCH. - default: record_and_delegate + [HARDCODED] the PSF observation carries a flat weight 1/PSF_NOISE^2 + with PSF_NOISE 1e-5; without it the g-prior swamps the PSF + likelihood. The recovered PSF shape and size are flat across 1e-4 to + 1e-6 on the digital twin. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::PSF_NOISE; + src/shapepipe/modules/ngmix_package/ngmix.py::make_ngmix_observation. + default: psf_noise_1em5 options: - record_and_delegate: - label: Record NGMIX_N_EPOCH, report attrition, no pipeline gate - pipeline_epoch_floor: - label: Fail tiles below a minimum surviving-epoch fraction - excluded: true - excluded_reason: >- - Fails whole tiles for what is a versionable per-object - catalogue decision; the depth effect is already recorded. + psf_noise_1em5: + label: PSF_NOISE 1e-5 megacam_ccd_flip: - label: 180-degree tile-vignet rotation for MegaCam CCDs <18 and 36-37 [HARDCODED] + label: 180-degree rotation of tile vignets for CCDs below 18 and 36-37 rationale: >- - "MegaPipe has CCDs that are upside down" (docstring) — the tile - vignet is rotated to register with epoch stamps; a wrong flip - mis-registers the tile mask against the epoch, changing flagged - pixels and the 1/3-masked cut. Carries its own recorded caveat: - "will give incorrect results when used with THELI ccds. Fix this." + [HARDCODED] "MegaPipe has CCDs that are upside down": the tile vignet + and segmentation stamp are rotated to register with the epoch stamp. + A wrong flip mis-registers the tile coverage flag against the epoch, + changing flagged pixels and the masked-fraction cut. The docstring + warns it gives incorrect results for THELI CCDs. Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::Ngmix.MegaCamFlip. default: megapipe_flip options: - megapipe_flip: { label: Flip CCDs <18 and 36/37 (MegaPipe orientation) } + megapipe_flip: + label: Flip CCDs < 18 and 36/37 (MegaPipe orientation) + defect_fill: + label: Image content of flagged / zero-weight pixels before metacal + rationale: >- + prepare_ngmix_weights gives weight 0 to every pixel with a nonzero + instrument flag, zero exposure weight or invalid background RMS. What + the IMAGE holds in those pixels still matters, because metacal never + looks at weights: ngmix builds a galsim InterpolatedImage from the + whole observation image, deconvolves, shears and reconvolves it, and + copies the weight map through unchanged. Whatever sits in a + zero-weight pixel is therefore spread into the weighted pixels within + about a PSF width, with ringing at sharp features. DES's own + corrector (ngmixer) says why it fills: "it may be important for codes + that take moments or use FFTs". The fill is coupled to the neighbour + treatment through the single BLEND_HANDLING key, which the committed + config leaves at its default. Under noisefill, masked pixels get an + independent noise realisation at the per-pixel RMS. Under uberseg, + the fill is skipped, so raw bad columns, bleeds, cosmic rays and + bright-star light enter metacal. [LINT] the prepare_ngmix_weights + docstring says noisefill keeps the weight of filled pixels (the code + zeroes it), and the ngmix_runner comment says noisefill fills + neighbour pixels (it fills flagged pixels and leaves neighbours + untouched). No DES metacal pipeline passed raw defects through + metacal: Y1 dropped every epoch with a masked pixel + (max_zero_weight_frac 0.0), Y3 filled symmetrized defects with the + best-fit central model (both candidate Y3 configs do this; which one + was production is not recorded), and Y6 interpolated symmetrized + defects in the image and in every noise image. The residual cost of + any fill is anisotropy. A filled bad column that crosses the galaxy + removes or misplaces light along one detector axis. The + reconvolution spreads that into an additive e1-type term, coherent on + the sky because CFHT/MegaCam, like DECam, has a fixed sky orientation + (inferred from CFHT's equatorial mount, no derotator). Sheldon & Huff + 2017 saw a large additive e1 even with model fill, removed by a + 90-degree compensating mask. Every DES pipeline symmetrized the + defect mask, and ShapePipe does not. The fill should be set + independently of blend_handling; the recommended option is + symmetrized_noise. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::prepare_ngmix_weights; + src/shapepipe/modules/ngmix_package/ngmix.py::do_ngmix_metacal; + src/shapepipe/modules/ngmix_runner.py::ngmix_runner. + default: noise + options: + noise: + label: Noise fill (BLEND_HANDLING = noisefill) + description: >- + Masked pixels are replaced by an independent noise realisation at + the per-pixel background RMS and keep weight 0. Consistent with + metacal's fixnoise noise image, which covers every pixel. Removes + defects, but leaves an unsymmetrized hole in the galaxy light + wherever a defect crosses the object, which can give an e1-type + additive term. + insights: [mask_metacal_acts_on_whole_stamp, mask_bad_column_symmetrize] + raw: + label: "No fill: raw defect values (BLEND_HANDLING = uberseg)" + description: >- + Masked pixels keep weight 0, but their raw values (bad columns, + saturation, bleeds, cosmic rays, bright-star light) stay in the + image that metacal deconvolves, shears and reconvolves. + excluded: true + excluded_reason: >- + Metacal acts on every pixel regardless of weight, so raw defects + leak into the weighted pixels. No published metacal pipeline does + this: DES dropped, model-filled or interpolated defects. The code + reaches this option only as a side effect of BLEND_HANDLING = + uberseg. + insights: [mask_metacal_acts_on_whole_stamp, mask_des_defect_practice] + symmetrized_noise: + label: 90-degree-symmetrized mask, then noise fill + description: >- + Not implemented. OR the defect mask with its 90-degree rotation + about the stamp centre, zero the weight on the union, and + noise-fill the union as noisefill does. This mirrors the DES + mask symmetrization (Y1/Y3 ngmixer symmetrize_weight; Y6 + symmetrize_masking) and cancels the leading column-aligned + additive term. It roughly doubles the masked area, so the epoch + cut must be applied after symmetrizing. ShapePipe stamps are + square, so rot90 is well defined. The noise image needs no change. + insights: [mask_bad_column_symmetrize, mask_des_defect_practice, mask_fixed_orientation] + interpolate: + label: Symmetrize, then interpolate image and noise image (DES Y6) + description: >- + Not implemented. Symmetrize as above, then fill the union by 2D + Clough-Tocher interpolation (scipy) of the image and, identically, + of the fixnoise noise image. This restores galaxy light across + narrow defects instead of leaving a hole. It is the DES Y6 and + Rubin metadetect practice. Poor for large holes (star masks), + which Y6 zeroes with apodized edges. + insights: [mask_interpolate_with_noise, mask_bad_column_symmetrize, mask_sharp_edges_ring] + model: + label: Symmetrize, then fill with the best-fit central model (DES Y3) + description: >- + Not implemented. Fill symmetrized defects with the PSF-convolved + best-fit model of the central object from a pre-metacal fit + (ngmix v1.3.9 replace_masked_pixels). The uberseg-only Y3 config + adds no noise (add_noise=False); the MOF-corrector config adds + it. This restores galaxy light; without noise it leaves + noise-free patches that the full-stamp fixnoise noise image does + not mirror. + insights: [mask_bad_column_symmetrize, mask_des_defect_practice] + blend_handling: + label: Neighbour treatment before metacal + rationale: >- + This decision covers how pixels shared with a neighbour are treated, + and only that; defect fill is the separate decision above. The two + are coupled through BLEND_HANDLING: noisefill (the default; the + committed config sets no key) leaves neighbours fully weighted and + untouched, while uberseg zeroes the weight of pixels nearer a + neighbour's coadd segmentation footprint than the target's + (DILATE_NEIGHBOUR, default 1) and leaves the image untouched. + Official uberseg is weight-only: esheldon/meds get_uberseg returns a + weight map and never modifies the image (a nearest-segment-pixel + Voronoi split). DES Y1's fiducial metacal ran on uberseg-weighted + stamps with the raw neighbour light still in the image. The last Y3 + config does the same, though an earlier Y3 config subtracted MOF + neighbours first. DES Y6 does not mask neighbours before the shear + step and uses uberseg only as the weight of the fit after metacal. + None of the DES or Rubin metacal/metadetect pipelines noise-fills the + neighbour side. Leaving neighbour light raw is consistent with + metacal: it is real sky, and the artificial shear shears it along + with the target, as the real shear does. The reconvolution spreads it + slightly further across the Voronoi boundary than the PSF already + had. The known residual is about +2% m for uberseg-only against MOF + subtraction in DES Y1 simulations. Separately, Sheldon et al. 2020 + find that the blending bias of per-stamp metacal is dominated by + shear-dependent detection, which no pixel treatment fixes and which + DES Y3 calibrated with simulations. Noise-filling the neighbour side + would instead cut the target's own light along an unsheared + boundary, a sharp edge that rings in the FFTs. The recommended + comparison arm is uberseg (weight-only), with defect_fill held equal + across arms. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::uberseg_weight; + src/shapepipe/modules/ngmix_package/ngmix.py::prepare_ngmix_weights; + src/shapepipe/modules/ngmix_runner.py::ngmix_runner. + default: none + options: + none: + label: No neighbour treatment (BLEND_HANDLING = noisefill) + description: >- + Neighbour pixels keep their full weight and image values. The fit + sees all neighbour light in the stamp, which is the configuration + Jarvis et al. 2016 found biased toward neighbours (worse than + their segmentation-only mask). It is a candidate cause of the + FLAGS=2 B-modes investigated in #814. + insights: [mask_uberseg_neighbour_bias] + uberseg: + label: UberSeg, weight-only + description: >- + Zero the weight of pixels nearer a neighbour footprint than the + target, as DES Y1/Y3 did. The image is untouched there, so the + neighbour's light is sheared coherently with the target. Needs + the coadd segmentation stamp (SEG_VIGNET_PATH). DILATE_NEIGHBOUR=1 + absorbs the coadd-vs-epoch overlay offset, because ShapePipe + reuses one coadd seg stamp for every epoch where MEDS reprojects + it. + insights: [mask_uberseg_weight_only, mask_uberseg_neighbour_bias, mask_des_y1_uberseg_only, mask_blend_bias_detection] + uberseg_fill: + label: UberSeg plus noise fill of neighbour-side pixels + description: >- + Not implemented. Additionally replace the neighbour-side pixels + with noise. This removes neighbour light from metacal, but cuts + the target's own wings along an unsheared Voronoi boundary: a + sharp edge (FFT ringing) that does not respond to the artificial + shear as sky does. No published pipeline does this. Useful only as + a diagnostic arm; a smooth (apodized) taper would be the less + damaging variant. + insights: [mask_sharp_edges_ring, mask_blend_bias_detection] + mof_subtract: + label: Subtract neighbour models, then UberSeg (DES Y1 alternative) + description: >- + Not implemented. Subtract multi-object-fit models of the + neighbours from the stamp before metacal and keep uberseg + weights. This removed the ~2% uberseg-only bias in DES Y1 + simulations, although Sheldon et al. 2020 found similar blend + biases with and without MOF. + insights: [mask_des_y1_uberseg_only, mask_blend_bias_detection] + epoch_masked_fraction_cut: + label: Per-epoch masked-fraction cut + rationale: >- + [HARDCODED] an epoch whose stamp has more than 1/3 of its pixels + flagged (any nonzero flag bit, including the tile-coverage bit 2**10 + set where the tile vignet is off-image) is dropped from the + multi-epoch fit; an object with no surviving epoch has no shape. DES + was stricter. Y1 rejected any epoch with a masked or zero-weight + pixel, and any whose central 4-pixel region was masked. Y3 cut at 10% + of raw zero-weight pixels. Y6 dropped images more than 10% missing + and cut objects at mfrac < 0.1, which its simulations show avoids + calibration bias. Sheldon & Huff 2017 recommend dropping problematic + epochs when many are available. The cut interacts with defect_fill: + symmetrizing roughly doubles the masked fraction, so the cut should + be applied after symmetrizing. UNIONS has fewer epochs than DES, so + the cost in effective number density has to be measured, not + assumed. A central-region veto (drop the epoch if a defect lies + within a few pixels of the centre) is a cheap refinement. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::prepare_postage_stamps. + default: one_third + options: + one_third: + label: 1/3 of the stamp flagged + description: >- + Drop an epoch only if more than 1/3 of the stamp pixels carry a + nonzero flag. + ten_percent: + label: 10% (DES Y3 / Y6) + description: >- + Not implemented. Drop an epoch if more than 10% of the + (symmetrized) stamp is masked, matching DES Y3 + max_zero_weight_frac and Y6 max_masked_fraction. + insights: [mask_multi_epoch_drop] + any_masked: + label: Any masked pixel (DES Y1) + description: >- + Not implemented. Drop an epoch if any stamp pixel is masked, so + that no fill is ever needed. DES could afford this with about 10 + epochs per band; UNIONS likely cannot. + insights: [mask_multi_epoch_drop, mask_des_defect_practice] prior_insights: guinot22_gaussian_model: claim: >- @@ -1518,153 +1381,313 @@ analyses: quote: exact: 'This method creates four images used for the calibration, and one for the measurement.' location: { page: 6 } - - # ═════════════════════════════════════════════════════════════════════════ - psf_diagnostics: - description: >- - The PSF-fidelity diagnostic chain: merge the held-out (20%) validation - stars into one catalogue, bin PSF shapes and residuals over the focal - plane. DORMANT in the committed snakemake chain — no rule runs it. - Modules: src/shapepipe/modules/merge_starcat_runner.py (+ per-model - merge classes in merge_starcat.py), mccd_plots_runner.py (serves both - PSF models despite its name). Boundary note: the module docstring - (mccd_package/__init__.py:157) still advertises rho-statistics plots, - but no rho/treecorr code remains in shapepipe — rho/tau statistics - moved downstream to sp_validation (rho_tau.py via - shear_psf_leakage.RhoStat/TauStat; treecorr min_sep/max_sep/nbins, - jackknife patch numbers hardcoded per survey with a "TODO to yaml"). - The diagnostic decision chain thus crosses the repo boundary into - sp_validation. [LINT] the module docstring still advertises rho - statistics this package no longer computes. - inputs: - - id: validation_star_cats - type: data - source: per-CCD star_split_ratio_20 catalogues with PSF/star HSM shapes (psfex_interp VALIDATION mode) - outputs: - - id: merged_star_catalogue - type: data - format: fits - description: >- - One full_starcat over the run — the input rho/tau statistics and - leakage diagnostics consume downstream. - decisions: [starcat_merge_source, meanshape_binning] - decisions: - starcat_merge_source: - label: Which PSF model's validation output feeds the merged star catalogue - rationale: >- - merge_starcat_runner dispatches on PSF_MODEL in {psfex, mccd, - setools} to per-model merge classes (different HDU conventions: - mccd HDU 1, psfex/setools HDU 2). Follows star_selection_psf. - psf_modelling_software; recorded separately because the merge can - also consume raw setools output (pre-model diagnostics). - Anchor: src/shapepipe/modules/merge_starcat_runner.py::merge_starcat_runner; - src/shapepipe/modules/merge_starcat_package/merge_starcat.py. - default: psfex - options: - psfex: - label: PSFEx validation catalogues (HDU 2) - insights: [guinot22_psfex_for_v1] - mccd: { label: MCCD validation catalogues (HDU 1) } - setools: { label: Raw setools star catalogues } - meanshape_binning: - label: Focal-plane mean-shape binning and outlier handling - rationale: >- - PSF ellipticity/size and residuals binned per CCD over the focal - plane: X_GRID=5, Y_GRID=10 bins per CCD, colour scales MAX_E=0.05, - MAX_DE=0.005, REMOVE_OUTLIERS=False (example/cfis config; no - committed workflow config exists). Grid resolution sets which - spatial PSF-residual structure is visible; outlier removal changes - what the diagnostic hides. - Published description (Guinot+22 p.8): focal-plane residual maps - averaged in ~20 arcsec cells per CCD; current code: no committed - workflow config picks a grid at all — example/cfis carries both the - 5x10 grid recorded here (~77x86 arcsec) and, in - config_valjoint_Pl_mccd.ini, a 20x40 grid (~19x22 arcsec) that - reproduces the paper. This is therefore an undetermined knob with - two committed precedents rather than a value that drifted. The - uniform REMOVE_OUTLIERS=False is a genuine paper-silence gap. - Anchor: example/cfis/config_MsPl_psfex.ini#MCCD_PLOTS_RUNNER.X_GRID; - example/cfis/config_MsPl_psfex.ini#MCCD_PLOTS_RUNNER.REMOVE_OUTLIERS; - src/shapepipe/modules/mccd_plots_runner.py; - src/shapepipe/modules/mccd_package/mccd_plot_utilities.py::plot_meanshapes. - default: grid_5x10 - options: - grid_5x10: { label: "5x10 per CCD, outliers kept" } - prior_insights: - guinot22_psfex_for_v1: + mask_uberseg_weight_only: + label: UberSeg is a weight-map operation (Jarvis 2016) claim: >- - PSFEx is the PSF model behind the published UNIONS v1 catalogue, so - the merged validation star catalogue and its diagnostics are fed by - PSFEx output. - created_at: "2022-04-01T00:00:00Z" + UberSeg, introduced for DES SV, zeroes the WEIGHT of pixels that + belong to another object's coadd segmentation footprint or lie + closer to another object than to the target; MEDS weights are also + zeroed wherever a mask flag is set. The image is not modified: in + the forward-model fits it was designed for, a zero-weight pixel + drops out of the likelihood exactly. esheldon/meds get_uberseg + matches, which returns a weight map. + created_at: "2026-09-26T00:00:00Z" + derived: true evidence: - - id: ev_guinot22_psfex_v1 - doi: "10.48550/arXiv.2204.04798" + - id: ev1 + doi: 10.48550/arXiv.1507.05603 + version: 3 quote: - exact: 'We make use of the PSFEx software package' - location: { page: 4 } - - # ═════════════════════════════════════════════════════════════════════════ - survey_geometry: - description: >- - Effective survey area and mask geometry for two-point estimators. - DORMANT — no committed workflow rule. Module: - src/shapepipe/modules/random_cat_package/random_cat.py (+ runner): - uniform randoms over each tile rejected against the pipeline mask, - yielding effective area (overlap- and mask-corrected) and optionally - the mask itself as a HEALPix map (save_as_healpix). This is the - in-repo ancestor of the planned healsparse external-mask rework — whichever way that rework lands, this - sub-analysis is where its geometry decisions belong. Bitrot risk: - healpy is imported but absent from pyproject.toml dependencies. - inputs: - - id: tile_masks - type: data - source: per-tile pipeline flag maps + final catalogues - outputs: - - id: random_catalogue - type: data - format: fits - description: Per-tile random catalogue + effective area (+ optional HEALPix mask). - decisions: [random_sampling, healpix_mask_export] - decisions: - random_sampling: - label: Random-point density for area estimation - rationale: >- - N_RANDOM=50000 with DENSITY=True (per square degree; False = total - per tile) in the example config; no committed workflow value. - Sampling density sets the Monte Carlo noise floor on effective - area, which propagates to two-point normalisation. - Anchor: example/cfis/config_Rc.ini#RANDOM_CAT_RUNNER.N_RANDOM; - example/cfis/config_Rc.ini#RANDOM_CAT_RUNNER.DENSITY; - src/shapepipe/modules/random_cat_runner.py::random_cat_runner. - default: per_sqdeg_50k - options: - per_sqdeg_50k: { label: "50000 per sq deg" } - healpix_mask_export: - label: HEALPix export resolution for the pipeline mask - rationale: >- - SAVE_MASK_AS_HEALPIX=True, HEALPIX_OUT_NSIDE=1024 (~3.4 arcmin - pixels) in the example config — coarser than the arcsecond-scale - mask features it rasterises; the resolution choice decides what the - exported mask can represent. Supersession candidate under - masking-unification (healsparse). - Anchor: example/cfis/config_Rc.ini#RANDOM_CAT_RUNNER.SAVE_MASK_AS_HEALPIX; - example/cfis/config_Rc.ini#RANDOM_CAT_RUNNER.HEALPIX_OUT_NSIDE; - src/shapepipe/modules/random_cat_package/random_cat.py::RandomCat.save_as_healpix. - default: nside_1024 - options: - nside_1024: { label: nside 1024 } + exact: We then set pixels in the weight map to zero if they were either associated with other objects in the segmentation map or were closer to any other object than to the object of interest. + - id: ev2 + doi: 10.48550/arXiv.1507.05603 + version: 3 + quote: + exact: are both of the model-fitting variety + mask_uberseg_neighbour_bias: + label: Neighbour light biases shapes toward neighbours (Jarvis 2016) + claim: >- + With a plain segmentation-map mask, light from a bright neighbour + just outside its footprint entered the fit and biased the shape + toward the neighbour; UberSeg made that bias undetectable in + end-to-end simulations. ShapePipe's default (no neighbour treatment) + masks less than even the plain segmentation map. + created_at: "2026-09-26T00:00:00Z" + derived: false + evidence: + - id: ev1 + doi: 10.48550/arXiv.1507.05603 + version: 3 + quote: + exact: when using ordinary segmentation maps we found a significant bias of the galaxy shape in the direction toward neighbors. + - id: ev2 + doi: 10.48550/arXiv.1507.05603 + version: 3 + quote: + exact: we no longer detected any systematic bias in the shape estimates due to unmasked flux from neighboring objects + mask_metacal_acts_on_whole_stamp: + label: Metacal transforms every stamp pixel, weights unseen + claim: >- + Metacal builds an interpolated image of the whole postage stamp, + deconvolves, shears and reconvolves it; its Fourier transforms + cannot accommodate missing data. So the image values of zero-weight + pixels feed the sheared images. The method papers leave masking + unaddressed. ngmix implements exactly this: an InterpolatedImage of + obs.image, with the weights copied through. + created_at: "2026-09-26T00:00:00Z" + derived: true + evidence: + - id: ev1 + doi: 10.48550/arXiv.1702.02600 + version: 1 + quote: + exact: For each galaxy and PSF postage stamp, we first create an + - id: ev2 + doi: 10.48550/arXiv.1702.02600 + version: 1 + quote: + exact: We have made no attempt to deal with the effects of masked pixels or blending + - id: ev3 + doi: 10.48550/arXiv.1702.02601 + version: 2 + quote: + exact: convolutions cannot accommodate missing data + mask_bad_column_symmetrize: + label: Filled bad columns give additive e1; symmetrize the mask + claim: >- + Sheldon & Huff 2017 found that filling bad columns (with the + best-fit model, not even noise) gave a large additive e1 bias and a + few-per-mille multiplicative bias. Both vanished when a compensating + column rotated by 90 degrees about the stamp centre was added. + Sheldon et al. 2020 recommend the same compensating mask, and DES Y6 + OR-s each bad-pixel mask with its 90-degree rotation, as previous + DES pipelines did, to cancel additive biases. + created_at: "2026-09-26T00:00:00Z" + derived: true + evidence: + - id: ev1 + doi: 10.48550/arXiv.1702.02601 + version: 2 + quote: + exact: as well as a multiplicative bias of a few parts in a thousand + - id: ev2 + doi: 10.48550/arXiv.1702.02601 + version: 2 + quote: + exact: at 90 degree rotation about the center of the image, restoring symmetry to the image + - id: ev3 + doi: 10.48550/arXiv.1911.02505 + version: 2 + quote: + exact: An additional compensating mask, rotated at right angles to the real mask, can be used to restore symmetry to the image + - id: ev4 + doi: 10.48550/arXiv.2501.05665 + version: 2 + quote: + exact: we rotate the bad pixel mask by 90 degrees and apply it via a logical + - id: ev5 + doi: 10.48550/arXiv.2501.05665 + version: 2 + quote: + exact: This process helps to cancel additive biases in the final shape measurement. + mask_des_defect_practice: + label: DES never passed raw defects through metacal + claim: >- + Every DES metacal generation removed defects from the image before + metacal: Y1 dropped any epoch with a masked pixel, Y3 filled + symmetrized defects with the best-fit central model (ngmix-y1-config + / ngmix-y3-config, see the defect_fill rationale), and Y6 + interpolated symmetrized defects following "previous DES shear + measurement pipelines". None left raw defect values in a zero-weight + pixel, which is what ShapePipe's uberseg path does today. + created_at: "2026-09-26T00:00:00Z" + derived: true + evidence: + - id: ev1 + doi: 10.48550/arXiv.2501.05665 + version: 2 + quote: + exact: Following previous DES shear measurement pipelines + - id: ev2 + doi: 10.48550/arXiv.2501.05665 + version: 2 + quote: + exact: Then, we apply a two-dimensional Clough-Tocher interpolation + mask_interpolate_with_noise: + label: Interpolate defects, and the noise image identically + claim: >- + The metadetect papers interpolate masked regions (bad columns, + cosmic rays, saturation) before the shear step, taking care not to + create a spurious shear, and pass the noise image used for metacal's + correlated-noise correction through the same interpolation. + created_at: "2026-09-26T00:00:00Z" + derived: true + evidence: + - id: ev1 + doi: 10.48550/arXiv.1911.02505 + version: 2 + quote: + exact: The FFT does not permit missing data, so the masked regions must be interpolated in some way. + - id: ev2 + doi: 10.48550/arXiv.1911.02505 + version: 2 + quote: + exact: so the noise field used for correcting correlated noise effects must also be propagated through the same coadding and interpolation + - id: ev3 + doi: 10.48550/arXiv.2303.03947 + version: 2 + quote: + exact: Before warping, we interpolated the simulated artifacts in each image such as cosmic rays, bad columns, and saturated pixels. + - id: ev4 + doi: 10.48550/arXiv.2303.03947 + version: 2 + quote: + exact: we must also run the noise image through the same procedures + mask_sharp_edges_ring: + label: Sharp mask edges ring through metacal; apodize large masks + claim: >- + Masks with sharp edges cause ringing in the metacal FFTs, so + bright-star regions are set to zero with an apodized (smoothly + tapered) edge, and the noise image gets the same masking. The same + reasoning argues against a hard noise fill cut along a neighbour's + Voronoi boundary. + created_at: "2026-09-26T00:00:00Z" + derived: true + evidence: + - id: ev1 + doi: 10.48550/arXiv.2303.03947 + version: 2 + quote: + exact: Masks with sharp features can cause ringing in the FFTs used by the + - id: ev2 + doi: 10.48550/arXiv.2501.05665 + version: 2 + quote: + exact: we further mask bright stars with apodization to avoid FFT artifacts during the deconvolution, shearing, and reconvolution processes + - id: ev3 + doi: 10.48550/arXiv.2501.05665 + version: 2 + quote: + exact: we also applied the same masking and apodization to the + mask_multi_epoch_drop: + label: Drop heavily masked epochs (10% in DES) + claim: >- + With many dithered epochs, problematic data can simply be dropped + (Sheldon & Huff 2017). DES Y6 drops input images more than 10% + missing and cuts objects at masked fraction mfrac < 0.1, which its + image simulations show is enough to avoid shear calibration bias. + ShapePipe's cut is 1/3. + created_at: "2026-09-26T00:00:00Z" + derived: true + evidence: + - id: ev1 + doi: 10.48550/arXiv.1702.02601 + version: 2 + quote: + exact: data deemed problematic can simply be left out of the fit + - id: ev2 + doi: 10.48550/arXiv.2501.05665 + version: 2 + quote: + exact: Images with a missing pixel fraction higher than 10 + - id: ev3 + doi: 10.48550/arXiv.2501.05665 + version: 2 + quote: + exact: is sufficient not to introduce shear calibration biases + mask_des_y1_uberseg_only: + label: 'DES Y1 metacal: uberseg-only fiducial, ~2% m residual' + claim: >- + DES Y1 metacal's fiducial catalogue handled neighbours with uberseg + only (raw neighbour light in the image). In dense deblending + simulations it carried m of about +2% (2.18 +/- 0.16 per cent at S/N + > 10), removed by subtracting MOF models of the neighbours. In data + the relative uberseg-vs-MOF m was 0.023 +/- 0.009. + created_at: "2026-09-26T00:00:00Z" + derived: false + evidence: + - id: ev1 + doi: 10.48550/arXiv.1708.01533 + version: 2 + quote: + exact: masks pixels close to neighbours rather than assigning a fraction of the light in each pixel to them + - id: ev2 + doi: 10.48550/arXiv.1708.01533 + version: 2 + quote: + exact: We detect no bias when subtracting the light from neighbours. + mask_blend_bias_detection: + label: Stamp-metacal blend bias is mostly shear-dependent detection + claim: >- + Sheldon et al. 2020 find that the few-percent blending bias of + per-stamp metacal comes from shear-dependent detection, not from + blended light itself: it is similar with and without neighbour + subtraction, and in metacal the space between objects is sheared + coherently. DES Y3 kept Y1's approach (Gatti et al. list no change + to neighbour or mask handling) and calibrated the resulting 2-3% + with image simulations. + created_at: "2026-09-26T00:00:00Z" + derived: true + evidence: + - id: ev1 + doi: 10.48550/arXiv.1911.02505 + version: 2 + quote: + exact: We show that this bias is not due to blending itself, but rather to shear-dependent object detection. + - id: ev2 + doi: 10.48550/arXiv.1911.02505 + version: 2 + quote: + exact: we see similar biases when no deblending is performed + - id: ev3 + doi: 10.48550/arXiv.1911.02505 + version: 2 + quote: + exact: the space between objects is sheared coherently + - id: ev4 + doi: 10.48550/arXiv.2011.03408 + version: 3 + quote: + exact: shape catalogue differs from DES Y1 in the following ways + - id: ev5 + doi: 10.48550/arXiv.2011.03408 + version: 3 + quote: + exact: which affects the shear estimates when objects are blended + mask_fixed_orientation: + label: Fixed-orientation cameras make mask anisotropy coherent + claim: >- + Masks have a preferred direction (columns, bleeds, spikes). On a + camera with a fixed sky orientation (DECam, and CFHT/MegaCam on its + equatorial mount), mask-induced shape errors therefore add up + coherently; with Rubin's camera rotation, unmasked trails averaged + away. DES Y3 has an unexplained mean e1 of 3.5e-4 that its + simulations, which include the real bad-pixel masks, do not + reproduce; whether masks cause it is not established. + created_at: "2026-09-26T00:00:00Z" + derived: true + evidence: + - id: ev1 + doi: 10.48550/arXiv.1708.01533 + version: 2 + quote: + exact: The DES focal plane does not rotate, so these effects always correspond to the same orientation in sky coordinates. + - id: ev2 + doi: 10.48550/arXiv.2303.03947 + version: 2 + quote: + exact: We find that these unmasked trails do not cause a shear bias, which we attribute to the camera rotations that randomize the direction of the trail on the sky. + - id: ev3 + doi: 10.48550/arXiv.2011.03408 + version: 3 + quote: + exact: Our shear catalogue is characterized by a non-null mean shear in one of the two components, whose origin is unknown. # ═════════════════════════════════════════════════════════════════════════ catalogue_assembly: description: >- - Final per-tile catalogue: merging shape chunks, attaching photometry - and PSF diagnostics, classification, sentinels. Modules: - src/shapepipe/modules/make_cat_package/make_cat.py (+ runner), - merge_sep_cats.py, vignetmaker_package (stamps, see top-level - postage_stamp_size), find_exposures_package (epoch list from tile - HISTORY cards). Configs: config_tile_Mc.ini, config_tile_PiViVi.ini, - final_cat.param. + The per-tile science catalogue: shape chunks merged and joined to the + detection catalogue and PSF quantities. inputs: - id: ngmix_chunks type: data @@ -1674,68 +1697,43 @@ analyses: type: data format: fits description: The assembled per-tile science catalogue (final_cat family). + inputs: [ngmix_chunks] decisions: [star_galaxy_classification, tile_overlap_handling, - column_selection, failure_sentinels, postproc_run_provenance, - shape_catalogue_shortfall_guard] + failure_sentinels] decisions: star_galaxy_classification: - label: Star/galaxy separation — deferred out of the pipeline + label: Star/galaxy separation deferred out of the pipeline rationale: >- - Production sets SM_DO_CLASSIFICATION=False (config_tile_Mc.ini) and - wires no spread-model input: SPREAD_MODEL/SPREADERR_MODEL are - sentinel 99, no SPREAD_CLASS column, and the SPREAD_* entries in - final_cat.param are commented out. The dormant machinery - (make_cat.py::save_sm_data) classifies on class = sm + 2*sm_err - with star |class|<0.003, galaxy class>0.01 — thresholds hardcoded - in the function signature. Reactivation is a 4-line config diff - (run spread_model_runner after psfex_interp+vignetmaker, add its - output to make_cat inputs, flip the switch — the exact diff - between example/cfis/config_make_cat_psfex.ini and _nosm.ini; the - defunct tile wiring config_tile_PiViSmVi.ini is the reference). - Separation therefore happens entirely downstream (sp_validation); - the pipeline ships everything. Rationale for deferring not - recorded. - Published description (Guinot+22 p.5-6): galaxies are selected - inside the pipeline with the spread model, at s + 2*sigma_s > - 0.0003 together with s > 0 and 20 < MAG_AUTO < 26; current code: - classification is disabled entirely and separation deferred to - sp_validation, with the dormant make_cat thresholds putting the - like-for-like galaxy boundary at 0.01, some thirty times the - published cut (0.003 is its separate star-side bound). Even - reactivated, the code implements only the spread-model test — the - paper's companion cuts have no in-pipeline counterpart. + SM_DO_CLASSIFICATION=False and no spread-model input is wired, so the + catalogue ships every object and separation happens downstream. The + dormant make_cat classifier uses class = sm + 2 sm_err with stars at + |class| < 0.003 and galaxies at class > 0.01, thresholds fixed in the + function signature. Guinot+22 selected galaxies in the pipeline at s + + 2 sigma_s > 0.0003 together with s > 0 and 20 < MAG_AUTO < 26; the + dormant code implements only the spread-model test, at a galaxy + boundary thirty times higher. Anchor: workflow/config/cfis/config_tile_Mc.ini#MAKE_CAT_RUNNER.SM_DO_CLASSIFICATION; src/shapepipe/modules/make_cat_package/make_cat.py::save_sm_data; - workflow/config/cfis/final_cat.param#SPREAD_CLASS; - example/cfis/config_make_cat_psfex_nosm.ini; - example/cfis/defunct/config_tile_PiViSmVi.ini. + workflow/config/cfis/final_cat.param#SPREAD_CLASS. default: deferred_downstream options: deferred_downstream: label: No in-pipeline classification; catalogue ships all objects spread_model_inline: - label: spread_model classification in make_cat (0.003/0.01) - excluded: true - excluded_reason: >- - Machinery present but unwired in production; reactivating it - changes which objects downstream sees as galaxies. + label: spread_model classification in make_cat (0.003 / 0.01) + description: >- + Not wired in the workflow; needs spread_model_runner after the + PSF interpolation and its output added to make_cat's inputs. tile_overlap_handling: - label: Tile-overlap duplicates — neither removed nor flagged + label: Tile-overlap duplicates neither removed nor flagged rationale: >- - Adjacent tiles overlap; objects in the overlap are measured in - both. make_cat attaches only TILE_ID (parsed from the sexcat - filename); no unique-object rule, no overlap flag. [LINT] the - documented config key TILE_LIST ("used to flag objects in areas of - overlap between tiles", in the make_cat package docstring) is - implemented nowhere — grep across src/ and workflow/ finds only - the docstring. VERIFIED downstream: sp_validation dedups at - classification time (galaxy.py::classification_galaxy_overlap_ra_dec - RA/Dec cuts to non-overlapping tile areas, and the WCS-based - mask_overlap variant; applied as cut_overlap in - classification_galaxy_base) — so this is today's division of labour, - and the pipeline's contract is "ship duplicates, TILE_ID is the - handle"; flagging overlaps in the catalogue stays an open option. The dead TILE_LIST docstring remains a lint. + Adjacent tiles overlap, and objects in the overlap are measured in + both. [HARDCODED] make_cat attaches only TILE_ID; there is no unique-object rule + and no overlap flag, and sp_validation removes duplicates by cutting + each tile to its non-overlapping area. [LINT] the make_cat package + docstring documents a TILE_LIST key that flags overlap objects; + nothing implements it. Anchor: src/shapepipe/modules/make_cat_package/make_cat.py::save_sextractor_data; src/shapepipe/modules/make_cat_package/__init__.py. default: no_dedup_in_pipeline @@ -1743,95 +1741,30 @@ analyses: no_dedup_in_pipeline: label: Ship duplicates; TILE_ID is the only handle overlap_flagging: - label: Implement the documented TILE_LIST overlap flag + label: Flag overlap objects in the catalogue + description: Not implemented. nearest_tile_centre: label: Keep each object only in its nearest tile excluded: true excluded_reason: >- Requires cross-tile coordination at assembly time, which the - per-tile DAG deliberately avoids; dedup belongs downstream if - anywhere. - column_selection: - label: Which columns survive into the science catalogue - rationale: >- - final_cat.param: positions XWIN/YWIN_WORLD, TILE_ID, flags - (FLAGS, IMAFLAGS_ISO, NGMIX_MCAL_FLAGS), PSF ellipticity from - PSF_ORIG only, all five metacal branches of G1/G2/T/FLUX/FLAGS, - but shear errors only for NOSHEAR (sheared-branch error columns - commented out — downstream response-weighted estimators cannot - propagate per-branch errors), SExtractor photometry (MAG_AUTO, - FLUX_APER, FLUX_RADIUS, SNR_WIN, FWHM_*), N_EPOCH/NGMIX_N_EPOCH, - NGMIX_MOM_FAIL. Doesn't change membership, but determines which - numbers exist for downstream cuts and calibration. Note - final_cat.param is read by scripts/python/create_final_cat.py in - post-processing, outside the per-tile DAG. [LINT] see detection: - IMAFLAGS_ISO is requested but never reaches the merged catalogue. - Anchor: workflow/config/cfis/final_cat.param; - scripts/python/create_final_cat.py. - default: committed_param_list - options: - committed_param_list: { label: The committed final_cat.param set } + per-tile DAG deliberately avoids. failure_sentinels: - label: Objects without shape measurements kept, with sentinel values [HARDCODED] + label: Objects without shape measurements kept, with sentinel values rationale: >- - Unmatched objects (no ngmix row) stay in the catalogue with - sentinels: sizes/fluxes/flags 0, error fluxes/mags -1, - ellipticities -10, T_ERR 1e30. The sentinel choice defines what a - downstream cut must exclude — a naive G1 > -1 cut silently changes - the sample. Flag-0-for-failure is the sharpest hazard: a failed - object's NGMIX_MCAL_FLAGS reads as success. Rationale not recorded. + [HARDCODED] detections with no ngmix row stay in the catalogue with + sentinels: sizes, fluxes, magnitudes and flags 0, flux and magnitude + errors -1, ellipticities -10, size errors 1e30. The sentinels define + what a downstream cut must exclude: a failed object's + NGMIX_MCAL_FLAGS reads 0, the success value, so a cut on flags alone + keeps it. Anchor: src/shapepipe/modules/make_cat_package/make_cat.py::SaveCatalogue._save_ngmix_data. default: sentinel_values options: - sentinel_values: { label: "Keep with sentinels (flags 0, e -10, T_ERR 1e30)" } + sentinel_values: + label: Keep with sentinels (flags 0, e -10, T_ERR 1e30) drop_unmatched: label: Drop objects without shapes excluded: true excluded_reason: >- - Loses the photometry-only population and hides attrition from - the completeness accounting. - postproc_run_provenance: - label: Post-proc run selection — newest mtime wins, merged patches never refresh - rationale: >- - create_final_cat picks each tile's make_cat run by newest - directory mtime (skipping runs without an output FITS), and the - merged patch catalogue is incremental: a tile already present is - never refreshed — a reprocessed tile reaches the patch only via - an explicit single-ID remove+add. mtime is filesystem state, not - provenance: a touched old run can outrank a newer one. Outside - the per-tile DAG (scripts/, not workflow/). Anchor: - scripts/python/create_final_cat.py::process. - default: newest_mtime_incremental - options: - newest_mtime_incremental: { label: "Newest mtime, incremental merge, manual refresh" } - shape_catalogue_shortfall_guard: - label: 10% shape-shortfall guard — logged, not enforced [HARDCODED] - rationale: >- - If the merged shape catalogue covers <10% of the detection - catalogue, make_cat logs an error but the enforcement (return + - raise) is commented out in both the module and its runner: a tile - whose shapes are 90% missing from a processing error is written and - looks normal. The comment distinguishes the two causes (measurement - failure = ok; premature merge = error) but not why enforcement is - off. Interacts with top-level per_unit_count_floor, which floors - tile_ngmix at 1 file and cannot see intra-file attrition. - Anchor: src/shapepipe/modules/make_cat_package/make_cat.py::SaveCatalogue._save_ngmix_data; - src/shapepipe/modules/make_cat_runner.py::make_cat_runner. - default: log_only - options: - log_only: { label: "Log the shortfall, write the tile anyway" } - enforce_10pct: - label: Fail the tile below 10% coverage - excluded: true - excluded_reason: >- - Was the coded intent, then disabled — reason unrecorded; - flagged as a question, not an endorsement. - -# ── Dormant science paths, surveyed but not yet recorded as sub-analyses ──── -# Candidates for future passes (each carries real scientific knobs): -# * External photometry match — match_external_package (TOLERANCE=0.3 -# arcsec against UNIONS ugriz; the external catalogue path is hardcoded -# to an IAP/candide location). -# * Image-simulation validation wiring — example/cfis_image_sims/: -# same chain over SKiLLS images with fake_psf substitution; bash-shaped, -# not yet ported to snakemake. + Loses the photometry-only population and hides attrition. diff --git a/universes/committed.yaml b/universes/committed.yaml index 5530226cf..b250eef95 100644 --- a/universes/committed.yaml +++ b/universes/committed.yaml @@ -1,30 +1,27 @@ id: committed -description: The committed configuration on feat/snakemake-orchestration. +description: The committed configuration on develop. decisions: - per_unit_count_floor: count_floor + per_unit_completeness: exact_counts postage_stamp_size: px_51 photometric_zeropoint: fixed_30_tiles_header_exposures - baseline_validation_criterion: statistical_parity analyses: masking: decisions: - star_catalogue_query: gsc_23_vizier - star_magnitude_definition: mean_finite_bands - bright_star_mask_geometry: megaprime_polygon_linear_scaling - deep_sky_object_masking: circles_no_padding - border_mask_width: px50_exposures_only - pixel_threshold_flags: stock_ww_thresholds - external_flag_usage: exposures_only + pixel_mask_source: instrument_flags_only + psf_star_mask_veto: instrument_flags_only + sky_mask_application: deferred_downstream detection: decisions: - detection_threshold_policy: thresh_1p5_minarea5_fwhm2px_filter - deblending_policy: mincont_5em4_tiles - background_model: manual_zero_tiles_auto_exposures - weighting_and_interpolation: map_weight_interp_all + detection_threshold_policy: megapipe_tiles + deblending_policy: megapipe_tiles + background_model: auto_megapipe_tiles + weight_map_usage: map_weight + zero_weight_interpolation: interp_all + spurious_detection_cleaning: clean_1 + blend_photometry_mask_type: correct + photometry_parameters: kron_25_35 detection_source_mode: sx_nomask_single_image epoch_membership_ccd_bounds: trimmed_bounds_33_2080 - photometry_parameters: kron_25_35 - cleaning_and_neighbour_masking: clean_1_correct preparation: decisions: astrometric_solution_source: delivered_headers @@ -32,39 +29,32 @@ analyses: epoch_provenance_from_tile_history: history_parse object_position_columns: xwin_windowed stamp_positioning_and_padding: round_and_zero_pad - epoch_flag_source: raw_flags star_selection_psf: decisions: star_selection_box: mode_centred_box + psf_train_validation_split: split_80_20_seeded psfex_candidate_vetting: builtin_defaults psf_modelling_software: psfex - psf_train_validation_split: split_80_20_seeded psf_model_complexity: pixel_basis_deg2_per_ccd psf_acceptance_thresholds: stars22_chi2_2 shape_measurement: decisions: ngmix_seed_mode: position_seed galaxy_model: gauss + fit_initialisation: prior_guess_t025_ntry5_2 fit_priors: gpriorba04_flat metacal_scheme: five_types_step001_fitgauss centroid_source: wcs - epoch_quality_and_weighting: third_masked_cut - noise_model: rms_vignet_weights - psf_epoch_loss_policy: record_and_delegate + epoch_flux_rescaling: fscale + psf_epoch_averaging: galaxy_weight_sum + galaxy_pixel_weights: rms_vignet_weights + psf_likelihood_noise: psf_noise_1em5 megacam_ccd_flip: megapipe_flip - psf_diagnostics: - decisions: - starcat_merge_source: psfex - meanshape_binning: grid_5x10 - survey_geometry: - decisions: - random_sampling: per_sqdeg_50k - healpix_mask_export: nside_1024 + defect_fill: noise + blend_handling: none + epoch_masked_fraction_cut: one_third catalogue_assembly: decisions: star_galaxy_classification: deferred_downstream tile_overlap_handling: no_dedup_in_pipeline - column_selection: committed_param_list failure_sentinels: sentinel_values - postproc_run_provenance: newest_mtime_incremental - shape_catalogue_shortfall_guard: log_only From a6eff5e6831a918141b4fafcb4c700d796e1727a Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 03:04:34 +0200 Subject: [PATCH 57/85] docs(claude): point the scientific-decisions section at the anchor test Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_014bvNTrAmZxcfb1ee83ApPK --- CLAUDE.md | 47 ++++++++++++++++++----------------------------- 1 file changed, 18 insertions(+), 29 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 56579f03c..b5140cac6 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -124,38 +124,27 @@ where the change lives — and *scientific* decisions in `astra.yaml`, below. `astra.yaml` at the repo root is the pipeline's decision record: every consequential scientific choice embedded in the code and the committed configs, -each with its rationale, the alternatives that were considered and why they were -rejected, and an anchor back to the code or config that implements it. -`universes/committed.yaml` pins the option this branch's configuration -selects for every decision. The format -is ASTRA; `uvx astra-tools@0.2.17 guide` is the briefing and -`uvx astra-tools@0.2.17 spec` the field reference. +with its rationale, its alternatives, and an anchor to the code or config that +implements it. `universes/committed.yaml` pins the option the committed +configuration selects for every decision. The format is ASTRA; +`uvx astra-tools@0.2.17 guide` is the briefing and `uvx astra-tools@0.2.17 spec` +the field reference. The file's header states its conventions (anchor grammar, +`[HARDCODED]`, `[LINT]`). + +Membership test: a different defensible choice would change which objects enter +the shear catalogue, or the numbers attached to them. Detection thresholds, +masking, star-selection cuts, PSF model degree, ngmix priors and seeding, flag +semantics, completeness gates: in. Workflow policy (manifests, chunking, +allocation, failure reporting, provenance) is out; it lives in the PR and in the +PRD, CosmoStat/shapepipe#848. **A scientific change is not finished until the record is.** When a change moves -what the pipeline measures, amend `astra.yaml` in the same PR — add the decision -if it is new, or edit its rationale, options and anchors if it moved — pin the -selected option in `universes/committed.yaml`, and say so in the PR description. -Purely technical changes (refactors, performance, packaging, I/O) leave it alone, -except where they move a value the record carries: the completeness floors in -`workflow/scripts/completeness.py` are orchestration code holding a scientific -decision. - -The membership test is whether *a different defensible choice would change which -objects enter the shear catalogue, or the numbers attached to them.* Detection -threshold and deblending contrast, masking geometry, star-selection cuts, PSF -model degree, ngmix priors and seeding, flag semantics, completeness floors — in. -Manifest sentinels, chunk sizes, allocation strategy, directory layout — out; -those live in the PR and the PRD. - -The file's own header states the conventions it follows. In short: every -rationale ends with a greppable `Anchor: path::symbol; path#SECTION.KEY` -sentence whose refs never cite line numbers; `[HARDCODED]` marks a scientific value -with no config exposure; `[LINT]` marks a place where the record and the code, or -the code and itself, disagree. Validate before committing: +what the pipeline measures, amend `astra.yaml` in the same PR (add the decision, +or edit its rationale, options and anchors), pin the selected option in +`universes/committed.yaml`, and say so in the PR description. The anchor test +`tests/unit/test_astra_anchors.py` runs in CI; a scientific change that breaks it +or leaves the record stale is unfinished. Before committing: ```bash uvx astra-tools@0.2.17 validate ``` - -The record was authored against this branch's workflow configs; entries marked -`[PENDING #NNN]` describe state that has not yet reached `develop`. From 62a2699ebb424c34941fc296ecf18a0f1837b241 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 03:06:46 +0200 Subject: [PATCH 58/85] feat(make_cat): fixed per-epoch slot count via N_EPOCH_SLOTS _save_psf_data writes its per-epoch column families (HSM_*_PSF_n, EXP_ID_n, CCD_n) for n = 1..max(N_EPOCH)+1 of the tile, so the column count varies tile to tile (5-7 across the 64-tile smk-g7 campaign). The campaign merge needs one schema across all tiles, so per-epoch columns cannot survive it while the slot count floats. Add an optional [MAKE_CAT_RUNNER] N_EPOCH_SLOTS that fixes the number of slots written for every family, with unfilled slots holding each family's existing sentinel. An object with more epochs than slots raises with its id, epoch count and the slot count rather than being silently truncated. Unset, the slot count is the tile's max(N_EPOCH)+1 as before. The slot allocation is driven by one family -> (sentinel, dtype) table, and each object's galaxy_psf entry is read from the sqlite store once instead of once per epoch. Co-Authored-By: Claude Fable 5.1 --- .../modules/make_cat_package/__init__.py | 6 + .../modules/make_cat_package/make_cat.py | 97 ++++----- src/shapepipe/modules/make_cat_runner.py | 8 +- tests/module/test_make_cat.py | 190 +++++++++++++++++- 4 files changed, 250 insertions(+), 51 deletions(-) diff --git a/src/shapepipe/modules/make_cat_package/__init__.py b/src/shapepipe/modules/make_cat_package/__init__.py index 4527a2ba6..77c76072c 100644 --- a/src/shapepipe/modules/make_cat_package/__init__.py +++ b/src/shapepipe/modules/make_cat_package/__init__.py @@ -47,6 +47,12 @@ retained as the extension point for a future estimator family) SAVE_PSF_DATA : bool, optional Save PSF information if ``True``; default value is ``False`` +N_EPOCH_SLOTS : int, optional + Number of slots written for each per-epoch PSF column family + (``HSM_*_PSF_n``, ``EXP_ID_n``, ``CCD_n``) when ``SAVE_PSF_DATA`` is + ``True``; unfilled slots hold the family's sentinel, and an object with + more epochs than slots raises an error. A fixed count gives every tile + the same schema. Default is the tile's maximum ``N_EPOCH`` plus one TILE_LIST : str, optional Path to list of all tile IDs, used to flag objects in areas of overlap between tiles diff --git a/src/shapepipe/modules/make_cat_package/make_cat.py b/src/shapepipe/modules/make_cat_package/make_cat.py index 062ac5719..392db6b4e 100644 --- a/src/shapepipe/modules/make_cat_package/make_cat.py +++ b/src/shapepipe/modules/make_cat_package/make_cat.py @@ -299,6 +299,7 @@ def process( mode="", cat_path=None, moments=False, + n_epoch_slots=None, ): """Process Catalogue. @@ -310,6 +311,9 @@ def process( Path to input catalogue moments : bool Option to run ``ngmix`` mode with moments + n_epoch_slots : int, optional + Number of per-epoch slots in ``psf`` mode; if ``None``, the + tile's ``max(N_EPOCH) + 1`` Returns -------- @@ -326,7 +330,7 @@ def process( if mode == "ngmix": err_msg = self._save_ngmix_data(cat_path, moments) elif mode == "psf": - self._save_psf_data(cat_path) + self._save_psf_data(cat_path, n_epoch_slots) else: err_msg = ( f"Invalid process mode ({mode}) for " @@ -614,73 +618,70 @@ def _save_ngmix_data(self, ngmix_cat_path, moments=False): return None - def _save_psf_data(self, galaxy_psf_path): + def _save_psf_data(self, galaxy_psf_path, n_epoch_slots=None): """Save PSF data. - Save the PSF catalogue into the final one. + Save the PSF catalogue into the final one, as per-epoch column + families ``HSM_G1_PSF_n``, ``HSM_G2_PSF_n``, ``HSM_T_PSF_n``, + ``HSM_FLAG_PSF_n``, ``EXP_ID_n`` and ``CCD_n``. Slot ``n`` of every + family refers to the same epoch; slots no epoch fills keep the + family's sentinel. Parameters ---------- galaxy_psf_path : str Path to the PSF catalogue to save + n_epoch_slots : int, optional + Number of slots written per family; if ``None``, the tile's + ``max(N_EPOCH) + 1`` + + Raises + ------ + ValueError + If an object has more epochs than ``n_epoch_slots`` """ galaxy_psf_cat = SqliteDict(galaxy_psf_path) - max_epoch = np.max(self._final_cat_file.get_data()["N_EPOCH"]) + 1 - - self._output_dict = { - f"HSM_G1_PSF_{idx + 1}": np.ones(len(self._obj_id)) * -10.0 - for idx in range(max_epoch) - } - self._output_dict = { - **self._output_dict, - **{ - f"HSM_G2_PSF_{idx + 1}": np.ones(len(self._obj_id)) * -10.0 - for idx in range(max_epoch) - }, - } - self._output_dict = { - **self._output_dict, - **{ - f"HSM_T_PSF_{idx + 1}": np.zeros(len(self._obj_id)) - for idx in range(max_epoch) - }, - } - self._output_dict = { - **self._output_dict, - **{ - f"HSM_FLAG_PSF_{idx + 1}": np.ones( - len(self._obj_id), dtype="int16" - ) - for idx in range(max_epoch) - }, - } - # Per-epoch exposure ID and CCD number, slot-aligned with HSM_*_PSF_n; - # -1 marks an empty slot (the CCD_N sentinel convention). - self._output_dict = { - **self._output_dict, - **{ - f"EXP_ID_{idx + 1}": np.ones(len(self._obj_id), dtype="int32") * -1 - for idx in range(max_epoch) - }, + n_epoch = self._final_cat_file.get_data()["N_EPOCH"] + if n_epoch_slots is None: + n_slots = np.max(n_epoch) + 1 + else: + n_slots = n_epoch_slots + + # Empty-slot sentinel and dtype per family. EXP_ID/CCD reuse the + # CCD_N convention: -1 is no exposure ID or CCD number. + slot_families = { + "HSM_G1_PSF": (-10.0, "float64"), + "HSM_G2_PSF": (-10.0, "float64"), + "HSM_T_PSF": (0.0, "float64"), + "HSM_FLAG_PSF": (1, "int16"), + "EXP_ID": (-1, "int32"), + "CCD": (-1, "int32"), } self._output_dict = { - **self._output_dict, - **{ - f"CCD_{idx + 1}": np.ones(len(self._obj_id), dtype="int32") * -1 - for idx in range(max_epoch) - }, + f"{family}_{slot + 1}": np.full( + len(self._obj_id), sentinel, dtype=dtype + ) + for family, (sentinel, dtype) in slot_families.items() + for slot in range(n_slots) } for idx, id_tmp in enumerate(self._obj_id): - if galaxy_psf_cat[str(id_tmp)] == "empty": + obj_epochs = galaxy_psf_cat[str(id_tmp)] + if obj_epochs == "empty": continue - for epoch, key in enumerate(galaxy_psf_cat[str(id_tmp)].keys()): + if len(obj_epochs) > n_slots: + galaxy_psf_cat.close() + raise ValueError( + f"Object {id_tmp} has {len(obj_epochs)} PSF epochs" + + f" (N_EPOCH={n_epoch[idx]}), more than the {n_slots}" + + f" per-epoch slots (N_EPOCH_SLOTS={n_epoch_slots})" + ) - gpc_data = galaxy_psf_cat[str(id_tmp)][key] + for epoch, (key, gpc_data) in enumerate(obj_epochs.items()): # `key` is "-"; reading it in the enumeration that # assigns `epoch` aligns EXP_ID_n/CCD_n with HSM_*_PSF_n by diff --git a/src/shapepipe/modules/make_cat_runner.py b/src/shapepipe/modules/make_cat_runner.py index 176341757..d40f6dc08 100644 --- a/src/shapepipe/modules/make_cat_runner.py +++ b/src/shapepipe/modules/make_cat_runner.py @@ -82,6 +82,10 @@ def make_cat_runner( save_psf = config.getboolean(module_config_sec, "SAVE_PSF_DATA") else: save_psf = False + if config.has_option(module_config_sec, "N_EPOCH_SLOTS"): + n_epoch_slots = config.getint(module_config_sec, "N_EPOCH_SLOTS") + else: + n_epoch_slots = None # Set final output file final_cat_file = make_cat.prepare_final_cat_file( @@ -135,7 +139,9 @@ def make_cat_runner( w_log.info(err_msg) if save_psf: - err_msg = sc_inst.process("psf", galaxy_psf_path) + err_msg = sc_inst.process( + "psf", galaxy_psf_path, n_epoch_slots=n_epoch_slots + ) # Optional per-band external healsparse mask lookup (UNIONS-WL/spherex#38): # add one MASK_ column per band, queried at each object's world diff --git a/tests/module/test_make_cat.py b/tests/module/test_make_cat.py index 34a1958e8..4bc0e7ee6 100644 --- a/tests/module/test_make_cat.py +++ b/tests/module/test_make_cat.py @@ -16,6 +16,7 @@ import numpy as np import numpy.testing as npt +import pytest from astropy.io import fits from sqlitedict import SqliteDict @@ -349,14 +350,14 @@ def get_data(self): return {"N_EPOCH": self._n_epoch} -def _run_save_psf(galaxy_psf_path, obj_id, n_epoch): +def _run_save_psf(galaxy_psf_path, obj_id, n_epoch, n_epoch_slots=None): """Drive ``_save_psf_data`` and return its populated output dict.""" inst = object.__new__(SaveCatalogue) inst._obj_id = np.asarray(obj_id) inst._output_dict = {} inst._final_cat_file = _FinalCatStub(n_epoch) - inst._save_psf_data(str(galaxy_psf_path)) + inst._save_psf_data(str(galaxy_psf_path), n_epoch_slots=n_epoch_slots) return inst._output_dict @@ -448,3 +449,188 @@ def test_save_psf_data_fills_sentinel_for_absent_epochs(tmp_path): assert out[col][0] == -1, col for col in ("EXP_ID_1", "CCD_1", "EXP_ID_2", "CCD_2", "EXP_ID_3", "CCD_3"): assert out[col][1] == -1, col + + +# --- _save_psf_data: fixed per-epoch slot count (N_EPOCH_SLOTS) --- + +# Per-family empty-slot sentinel: what a slot holds when no epoch fills it. +_PSF_SLOT_SENTINELS = { + "HSM_G1_PSF": -10.0, + "HSM_G2_PSF": -10.0, + "HSM_T_PSF": 0.0, + "HSM_FLAG_PSF": 1, + "EXP_ID": -1, + "CCD": -1, +} + + +class _ProcessCatStub(_FinalCatStub): + """FITSCatalogue stand-in for ``SaveCatalogue.process``; records add_col.""" + + def __init__(self, obj_id, n_epoch): + super().__init__(n_epoch) + self._number = np.asarray(obj_id) + self.cols = {} + + def open(self): + pass + + def close(self): + pass + + def get_data(self): + return {"NUMBER": self._number, "N_EPOCH": self._n_epoch} + + def add_col(self, name, data): + self.cols[name] = data + + +def _slot_numbers(out, family): + """The slot numbers ``n`` present in ``out`` for ``_n`` columns.""" + prefix = f"{family}_" + return sorted( + int(col[len(prefix):]) + for col in out + if col.startswith(prefix) and col[len(prefix):].isdigit() + ) + + +def test_save_psf_data_fixed_slots_pad_every_family(tmp_path): + """N_EPOCH_SLOTS fixes the slot count for every family, sentinel-padded. + + The campaign merge needs one schema across tiles, so a tile whose + objects all have far fewer epochs than N_EPOCH_SLOTS still writes + exactly slots 1..N_EPOCH_SLOTS for each per-epoch family, and every + slot no epoch fills holds that family's own sentinel. + """ + galaxy_psf_path = tmp_path / "galaxy_psf.sqlite" + per_obj = { + 101: {"2113864-7": _psf_epoch(0.01, 0.02, 0.5)}, + 202: { + "2113864-9": _psf_epoch(0.05, 0.06, 0.7), + "2358123-21": _psf_epoch(0.07, 0.08, 0.9), + }, + 303: "empty", + } + _write_galaxy_psf_cat(galaxy_psf_path, per_obj) + + # Driven through ``process``, the entry point the runner calls. + n_slots = 7 + cat = _ProcessCatStub([101, 202, 303], n_epoch=[1, 2, 0]) + sc = SaveCatalogue(cat, 3, _NullLogger()) + assert sc.process("psf", str(galaxy_psf_path), n_epoch_slots=n_slots) is None + out = cat.cols + + n_filled = [1, 2, 0] + for family, sentinel in _PSF_SLOT_SENTINELS.items(): + assert _slot_numbers(out, family) == list(range(1, n_slots + 1)), family + for row, filled in enumerate(n_filled): + for n in range(filled + 1, n_slots + 1): + assert out[f"{family}_{n}"][row] == sentinel, (family, row, n) + + +def test_save_psf_data_more_epochs_than_slots_raises(tmp_path): + """An object with more epochs than N_EPOCH_SLOTS raises, never truncates. + + Dropping the extra epochs would silently lose per-epoch PSF data; the + error names the object, its N_EPOCH and the slot count so the config + can be fixed. + """ + galaxy_psf_path = tmp_path / "galaxy_psf.sqlite" + per_obj = { + 101: {"2113864-7": _psf_epoch(0.01, 0.02, 0.5)}, + 404: { + "2113864-7": _psf_epoch(0.01, 0.02, 0.5), + "2229900-13": _psf_epoch(0.03, 0.04, 0.6), + "2358123-21": _psf_epoch(0.07, 0.08, 0.9), + }, + } + _write_galaxy_psf_cat(galaxy_psf_path, per_obj) + + with pytest.raises(ValueError) as excinfo: + _run_save_psf( + galaxy_psf_path, [101, 404], n_epoch=[1, 3], n_epoch_slots=2 + ) + msg = str(excinfo.value) + assert "404" in msg + assert "N_EPOCH=3" in msg + assert "N_EPOCH_SLOTS=2" in msg + + +def test_save_psf_data_exactly_slots_epochs_fits(tmp_path): + """An object whose epochs exactly fill N_EPOCH_SLOTS is written, not raised.""" + galaxy_psf_path = tmp_path / "galaxy_psf.sqlite" + per_obj = { + 404: { + "2113864-7": _psf_epoch(0.01, 0.02, 0.5), + "2229900-13": _psf_epoch(0.03, 0.04, 0.6), + }, + } + _write_galaxy_psf_cat(galaxy_psf_path, per_obj) + + out = _run_save_psf(galaxy_psf_path, [404], n_epoch=[2], n_epoch_slots=2) + + assert _slot_numbers(out, "EXP_ID") == [1, 2] + assert out["EXP_ID_2"][0] == 2229900 + npt.assert_allclose(out["HSM_G1_PSF_2"], [0.03]) + + +def test_save_psf_data_unset_slots_uses_tile_max_n_epoch_plus_one(tmp_path): + """Without N_EPOCH_SLOTS the slot count is the tile's max(N_EPOCH) + 1.""" + galaxy_psf_path = tmp_path / "galaxy_psf.sqlite" + per_obj = { + 101: {"2113864-7": _psf_epoch(0.01, 0.02, 0.5)}, + 202: { + "2113864-9": _psf_epoch(0.05, 0.06, 0.7), + "2229900-13": _psf_epoch(0.03, 0.04, 0.6), + "2358123-21": _psf_epoch(0.07, 0.08, 0.9), + }, + } + _write_galaxy_psf_cat(galaxy_psf_path, per_obj) + + out = _run_save_psf(galaxy_psf_path, [101, 202], n_epoch=[1, 3]) + + for family in _PSF_SLOT_SENTINELS: + assert _slot_numbers(out, family) == [1, 2, 3, 4], family + + +def test_save_psf_data_fixed_slots_keep_epoch_alignment(tmp_path): + """Under padding, slot n of every family still names the same epoch. + + Epochs fill slots 1..k in the galaxy_psf key order and padding sits + only in slots k+1..N_EPOCH_SLOTS, for every family alike; a flagged + epoch keeps its identity in its own slot while its HSM columns stay + at the sentinel. + """ + galaxy_psf_path = tmp_path / "galaxy_psf.sqlite" + epochs = [ + # (key, g1, g2, t, flag) in deliberately unsorted key order + ("2358123-21", 0.07, 0.08, 0.9, 0), + ("2113864-9", -10.0, -10.0, 0.0, 5), + ("2229900-13", 0.03, 0.04, 0.6, 0), + ] + per_obj = { + 505: {key: _psf_epoch(g1, g2, t, flag) for key, g1, g2, t, flag in epochs}, + } + _write_galaxy_psf_cat(galaxy_psf_path, per_obj) + + n_slots = 6 + out = _run_save_psf( + galaxy_psf_path, [505], n_epoch=[3], n_epoch_slots=n_slots + ) + + for n, (key, g1, g2, t, flag) in enumerate(epochs, start=1): + exp_id, ccd = (int(part) for part in key.split("-")) + assert out[f"EXP_ID_{n}"][0] == exp_id, n + assert out[f"CCD_{n}"][0] == ccd, n + if flag == 0: + npt.assert_allclose(out[f"HSM_G1_PSF_{n}"], [g1]) + npt.assert_allclose(out[f"HSM_G2_PSF_{n}"], [g2]) + npt.assert_allclose(out[f"HSM_T_PSF_{n}"], [t]) + assert out[f"HSM_FLAG_PSF_{n}"][0] == 0, n + else: + npt.assert_allclose(out[f"HSM_G1_PSF_{n}"], [-10.0]) + assert out[f"HSM_FLAG_PSF_{n}"][0] == 1, n + for n in range(len(epochs) + 1, n_slots + 1): + for family, sentinel in _PSF_SLOT_SENTINELS.items(): + assert out[f"{family}_{n}"][0] == sentinel, (family, n) From 512e7bf3c5712757e85dd50882d0d7bdccde4732 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 03:06:53 +0200 Subject: [PATCH 59/85] feat(workflow): save fixed-slot per-epoch PSF data in tile make_cat No workflow config sets SAVE_PSF_DATA, so no workflow catalogue carries the per-epoch HSM_*_PSF_n / EXP_ID_n / CCD_n columns. Enable it in the tile make_cat with N_EPOCH_SLOTS = 12 so every tile writes the same columns and the campaign merge can keep them. Co-Authored-By: Claude Fable 5.1 --- workflow/config/cfis/config_tile_Mc.ini | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/workflow/config/cfis/config_tile_Mc.ini b/workflow/config/cfis/config_tile_Mc.ini index 5920d0a91..1d2a3892e 100644 --- a/workflow/config/cfis/config_tile_Mc.ini +++ b/workflow/config/cfis/config_tile_Mc.ini @@ -82,3 +82,11 @@ NUMBERING_SCHEME = -000-000 SM_DO_CLASSIFICATION = False SHAPE_MEASUREMENT_TYPE = ngmix + +# Save per-epoch PSF shapes and epoch identity (HSM_*_PSF_n, EXP_ID_n, +# CCD_n). The campaign merge needs one column schema across all tiles, so +# the slot count is fixed rather than per-tile max(N_EPOCH)+1. It must +# exceed the survey's maximum epochs per object with margin (smk-g7: max 6 +# over 8 inspected tiles); an object with more epochs than slots is an error. +SAVE_PSF_DATA = True +N_EPOCH_SLOTS = 12 From b2a4d199b18d16873469a6899125a77171abb381 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 03:08:45 +0200 Subject: [PATCH 60/85] fix(workflow): five review findings on the #894 merge - Snakefile unit_pre docstring: every prologue line is a params rerun trigger, so lines are added only at a campaign boundary; the image-sims lines are conditional so a data prologue does not carry them. Drops the false "byte-identical" claim (SP_RETRIEVE is exported unconditionally). - create_final_cat.py -c: the patch/group is `run:`, as final_cat_merge names it, and a products_dir not laid out as //product refuses with a pointer to -i/-P. n_tiles_final.txt is written only when -o is given; the n_tiles attribute is always set. Without -c, -o behaves as before. - hdf5_reconcile.check_sole_group: the path carries `run:`, so a rename makes a new file; the guard is for a file at this path already holding another campaign's group. - config.yaml max_mem_mb: nibi's ceiling; a candide run never reaches it. - bin/sp: header names workflow/scripts/run_config.py; bare `sp`, -h and --help print snakemake's usage instead of refusing on an unset `run:`, and create no state dir. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_011xYGyn53XoyL83uuKPN7RQ --- scripts/python/create_final_cat.py | 27 +++++++++++++++++++-------- workflow/Snakefile | 6 +++--- workflow/bin/sp | 14 +++++++++----- workflow/config.yaml | 9 +++++---- workflow/scripts/hdf5_reconcile.py | 11 ++++++----- 5 files changed, 42 insertions(+), 25 deletions(-) diff --git a/scripts/python/create_final_cat.py b/scripts/python/create_final_cat.py index 3fe1dcf06..cbfcc8ff3 100755 --- a/scripts/python/create_final_cat.py +++ b/scripts/python/create_final_cat.py @@ -58,23 +58,32 @@ def params_from_run_config(params, defaults): raise ValueError(f"run config {params['run_config']} sets neither " "outputs.products_dir nor outputs.run_dir") - # The patch dir is the branch dir holding product/tiles, and -i is its - # parent. With the run template's layout (products_dir = /product) - # this is the path and group the workflow's final_cat_merge writes. + # The group and file are named for `run:`, as the workflow's + # final_cat_merge names them. -P is also the directory the -I walk + # matches under -i, so this needs the run template's layout, + # products_dir = //product, and -i is . + patch = cfg["run"] patch_dir = os.path.dirname(os.path.normpath(products)) - patch = os.path.basename(patch_dir) + if os.path.basename(patch_dir) != patch: + raise ValueError( + f"run config {params['run_config']}: products_dir {products} is " + f"not /{patch}/product, so -c cannot locate the tiles; " + "pass -i and -P explicitly") derived = { "image_sims": True, "input_root_dir": os.path.dirname(patch_dir), "patch": patch, "merged_cat_path": os.path.join(products, f"final_cat_{patch}.hdf5"), - "output_summary": os.path.join(products, "n_tiles_final.txt"), "param_path": os.path.join(repo, "workflow", "config", "cfis_image_sims", "final_cat.param"), } for key, value in derived.items(): if params.get(key) == defaults.get(key): params[key] = value + # The tile count is the file's n_tiles attribute; the text summary is + # written only when -o asks for it. + if params.get("output_summary") == defaults.get("output_summary"): + params["output_summary"] = None return params @@ -153,7 +162,8 @@ def params_default(): "param_path": "parameter file path, if not given use all columns, default={}", "patch": "patch number (data) or grid subdir (image_sims), default={}", "list_only": "print list of patches and IDs only, default={}", - "output_summary": "output file for numbre of tiles, default={}", + "output_summary": "output file for number of tiles (with -c, written" + " only if given), default={}", "ID": "ID for single-ID operation, default={}", "single_op": "single ID operation, allowed are 'check', 'add', 'remove'; default={}", "image_sims": "image simulations mode (different dir layout and run prefix), default={}", @@ -322,8 +332,9 @@ def print_list(params): if verbose: print(f"Total: {n_tiles} tiles") - with open(params["output_summary"], "w") as f_out: - print(n_tiles, file=f_out) + if params["output_summary"]: + with open(params["output_summary"], "w") as f_out: + print(n_tiles, file=f_out) # Write n_tiles to HDF5 file header with h5py.File(params["merged_cat_path"], "a") as hdf5_file: diff --git a/workflow/Snakefile b/workflow/Snakefile index fcbf68e82..53645422a 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -1079,9 +1079,9 @@ def unit_pre(stage, unit, *, exp_name=None, forest=None, env=None, Every line here is part of each rule's ``params.pre`` and so of the ``params`` rerun trigger: a line added for every rule reruns every finished - unit of a campaign on its next ``sp run``. The image-simulation lines are - therefore conditional, and a data run's prologue is byte-identical to what - it was before input_type existed. + unit of a campaign on its next ``sp run``, so lines are added only at a + campaign boundary. The image-simulation lines are conditional, so a data + prologue does not carry them. ``tile_numbers.txt`` carries the number in the form the INPUT files are named with, because get_images substitutes it verbatim into diff --git a/workflow/bin/sp b/workflow/bin/sp index 74ab37e44..86aecd75b 100755 --- a/workflow/bin/sp +++ b/workflow/bin/sp @@ -37,10 +37,10 @@ # # Run config: -c FILE, also --config-file. Read after workflow/config.yaml and # merged on top; anything still unset comes from the machines: entry for -# SP_PROFILE and input_type (scripts/run_config.py). +# SP_PROFILE and input_type (workflow/scripts/run_config.py). # To check a resolved value without running: # -# python scripts/run_config.py workflow/config.yaml ~/my_run.yaml outputs.run_dir +# python workflow/scripts/run_config.py workflow/config.yaml ~/my_run.yaml outputs.run_dir # # -c is sp's flag, not snakemake's --cores; pass cores as --cores or -j. # @@ -126,10 +126,10 @@ fi # Snakefile resolves it. cfg() { python "$SCRIPTS/run_config.py" "$HERE/config.yaml" "${SP_RUN_CONFIG:-}" "$1"; } RUN_DIR="$(cfg outputs.run_dir)"; INDEX_DB="$(cfg outputs.index_db)" -# `sp container` and `sp cancel` need neither a campaign nor its state dir; -# every other verb does. Checked here, before the state dir is created from +# `sp container`, `sp cancel` and the help forms (bare `sp`, -h, --help) need +# neither a campaign nor its state dir; every other verb does. Checked here, before the state dir is created from # RUN_DIR: with `run:` unset, RUN_DIR still holds a literal `$run`. -case "${1:-}" in container|cancel) NEEDS_CAMPAIGN=0 ;; *) NEEDS_CAMPAIGN=1 ;; esac +case "${1:-}" in container|cancel|""|-h|--help) NEEDS_CAMPAIGN=0 ;; *) NEEDS_CAMPAIGN=1 ;; esac if [ "$NEEDS_CAMPAIGN" = 1 ]; then if [ -z "$(cfg run)" ]; then echo "sp: \`run:\` is unset; the run config (-c FILE) names the campaign" >&2; exit 2 @@ -373,6 +373,10 @@ case "$cmd" in | awk -v r="$run" '$2 ~ r {print $1}' | xargs -r scancel echo "cancelled jobs matching '$run'; safe to --unlock / rerun now" ;; + ""|-h|--help) + # snakemake's own usage: parsing the Snakefile would need a campaign. + snakemake --help + ;; *) sm "$@" ;; diff --git a/workflow/config.yaml b/workflow/config.yaml index 94b9e79c5..43372855e 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -179,10 +179,11 @@ clean_tiles: true # Their exposures are rebuilt from scratch if the tile is retried. clean_ignore_tiles: [] -# The ceiling on any rule's mem_mb. A request above the partition maximum is a -# job SLURM never schedules and snakemake never diagnoses: it sits PENDING while -# the campaign looks alive. Nibi's standard compute node is 766 GB (192 cores at -# 4 GB/core), so 750000 leaves room for the OS and slurm's own overhead. The two +# The ceiling on any rule's mem_mb: nibi's, whose standard compute node is +# 766 GB (192 cores at 4 GB/core), so 750000 leaves room for the OS and slurm's +# own overhead. A candide run never reaches it. A request above the partition +# maximum is a job SLURM never schedules and snakemake never diagnoses: it sits +# PENDING while the campaign looks alive. The two # campaign-level merges size themselves from the campaign's bytes and will cross # this at survey scale — the cap turns "never runs" into "runs on the biggest # node there is", with a parse-time warning saying which rule was capped. diff --git a/workflow/scripts/hdf5_reconcile.py b/workflow/scripts/hdf5_reconcile.py index c57135cc4..aaafa513a 100644 --- a/workflow/scripts/hdf5_reconcile.py +++ b/workflow/scripts/hdf5_reconcile.py @@ -160,11 +160,12 @@ def check_free_space(output: Path) -> None: def check_sole_group(output: Path, group_path: str) -> None: """One file, one campaign — refuse to half-update a file holding two. - Renaming `run:` mid-flight points the rule at a NEW group inside the - SAME file (the path carries the campaign only on the tile side, where the - group does). Reconciling would then add a second group beside the first, - leave the first frozen and stale, and set a count attribute describing only - one of them. Nothing downstream reads such a file correctly, and no rule + The output path carries `run:`, so renaming a campaign produces a new + file, not a second group in this one. What this guards is a file already + at this path that holds ANOTHER campaign's group (a hand merge, or a copy). + Reconciling would then add a second group beside the first, leave the + first frozen and stale, and set a count attribute describing only one of + them. Nothing downstream reads such a file correctly, and no rule here means to produce one. Say what is there and stop. """ if not output.exists() or "/" not in group_path: From 4958d66157e5a1d219b8bff6cd312e58ef6e948a Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 03:15:05 +0200 Subject: [PATCH 61/85] workflow: scientific contracts for the campaign products Seven @sc contracts at the declarations that carry a workflow decision, so an agent editing them sees the constraint before it breaks it (sc-list prints every contract governing a path). The Snakefile and rules/*.smk carry theirs in workflow/CONTRACTS because sc-list does not parse Snakemake; the scripts carry theirs in their docstrings. The contracts cover: the campaign name's one source (`run:`), the final_cat.param allow-list and the reader's raise on a missing column, the never-fit rows the merge passes through (consumers cut on NGMIX_N_EPOCH > 0, not MCAL_FLAGS), the unit_pre rerun hazard, clean_exposure's persist edge under PERSISTS_PSF, and persist_exp's additive keep list. Co-Authored-By: Claude Fable 5.1 --- scripts/python/create_final_cat.py | 10 +++++-- workflow/CONTRACTS | 41 +++++++++++++++++++++++++++++ workflow/scripts/merge_final_cat.py | 8 ++++++ workflow/scripts/persist_exp.py | 9 +++++++ 4 files changed, 66 insertions(+), 2 deletions(-) create mode 100644 workflow/CONTRACTS diff --git a/scripts/python/create_final_cat.py b/scripts/python/create_final_cat.py index cbfcc8ff3..c35216e5e 100755 --- a/scripts/python/create_final_cat.py +++ b/scripts/python/create_final_cat.py @@ -375,8 +375,14 @@ def get_patch_group(hdf5_file, patch, verbose=False): def read_data(fits_file, params): - """Read Data. - + """Read the parameter list's columns out of one catalogue. + + @sc [label:schema] read-data-raises-on-missing-column + A requested column the catalogue lacks raises `KeyError` naming it; it is + never skipped or filled. `copy_data` keeps only columns present in the + source, so this raise is the one place a missing name stops a merge, and + without it a tile short a per-epoch slot would land in the merged file + silently narrower, with that slot's exposure identity gone. """ with fits.open(fits_file) as hdu_list: try: diff --git a/workflow/CONTRACTS b/workflow/CONTRACTS new file mode 100644 index 000000000..2faac3ca2 --- /dev/null +++ b/workflow/CONTRACTS @@ -0,0 +1,41 @@ +Scientific contracts governing workflow/ and everything beneath it. + +Format: `@sc [meta] id`, then one paragraph of prose. `sc-list ` prints +every contract governing a path; the Snakefile and rules/*.smk carry theirs here +because sc-list does not parse Snakemake. Where a check enforces a contract, the +prose names it. + +@sc [label:provenance] campaign-name-is-run +A campaign's name is the run config's `run:` and nothing else. The Snakefile +binds it once, `CAMPAIGN = config["run"]`, and every product that carries a +campaign name takes it from there: `final_cat_.hdf5` and its +`patches/` group, `full_starcat_.hdf5`, and the `$run` in the +machine defaults' `products_dir`/`index_db`. No rule or script reads a +`campaign` key or derives a name from a directory (`PRODUCTS_DIR.name`); two +sources that can disagree would file one campaign's merge under another's name. +Every campaign product path is rooted in `PRODUCTS_DIR`. + +@sc [label:schema] final-cat-param-is-exact-allow-list +Each input type's `config/*/final_cat.param` is the merged catalogue's exact +column list, in order: `final_cat_merge` writes those columns and no others, +and the reader raises on a listed column a tile lacks. So a name belongs there +only if `make_cat` writes it on EVERY tile. A per-epoch family (`EXP_ID_n`, +`CCD_n`, `HSM_*_PSF_n`) qualifies only with a fixed slot count, since +`make_cat` otherwise sizes it from the tile's own maximum `N_EPOCH`. Enforced +by tests/module/test_psf_grammar_properties.py (the shipped names). + +@sc [label:hazard] unit-pre-changes-at-campaign-boundary +Every line `unit_pre()` emits is part of each rule's `params.pre`, and both +profiles run with the `params` rerun trigger, so any change to its output +replans every finished unit of the campaign. Mid-campaign that rerun is +unsatisfiable for tiles whose exposure stores were reclaimed, and the failed +tile group's cleanup deletes those finished tiles' `final_cat`. Change +`unit_pre` output, or anything else in a rule's `params`, only at a campaign +boundary on a fresh root. + +@sc [label:custody] clean-exposure-waits-on-persist-iff-psf +`clean_exposure` takes the exposure's `exp_persist` manifest as input exactly +when `PERSISTS_PSF` (`psf_model != "fake"`). Dropping the edge under a real PSF +model lets reclamation delete the scratch store before its PSF products reach +`products_dir`; keeping it under `fake` makes every clean wait on a rule that +is not in the DAG. diff --git a/workflow/scripts/merge_final_cat.py b/workflow/scripts/merge_final_cat.py index 624f5526e..1c4be85c9 100644 --- a/workflow/scripts/merge_final_cat.py +++ b/workflow/scripts/merge_final_cat.py @@ -69,6 +69,14 @@ trigger would notice. A tile in the derived set whose catalogue is missing is a hard error here, not a skip — under the DAG it cannot happen, since every one of them is a declared input of this job. + +@sc [label:selection] never-fit-rows-pass-through +Every row of every tile catalogue reaches the merged file, unchanged, +including objects ngmix never fit. Those carry `NGMIX_N_EPOCH == 0` with +sentinel values (`NGMIX_MCAL_FLAGS == 0`, ellipticities `-10`, `T == 0`), so +`NGMIX_MCAL_FLAGS == 0` is not a validity cut: consumers select fitted objects +with `NGMIX_N_EPOCH > 0`. The merge neither fills these rows nor drops them; +that selection belongs to the consumer. """ import argparse diff --git a/workflow/scripts/persist_exp.py b/workflow/scripts/persist_exp.py index c3e14633d..997bc4489 100644 --- a/workflow/scripts/persist_exp.py +++ b/workflow/scripts/persist_exp.py @@ -85,6 +85,15 @@ ``clean_exposure`` uses), so a rerun that packs the same files leaves the mtime alone — mtime is a rerun trigger, and an unconditional rewrite would make every downstream ``clean_exposure`` look out of date once per invocation. + +@sc [label:safety] persist-exp-additive-and-always-validation +`persist_exp:` only ever adds: an existing tar is a floor whose members are +carried into the rewrite whatever the current keep list says, and +`psf_validation` is packed on every run whether or not the list names it. A +keep-list edit reruns this script, so a subtractive rewrite would delete +products from `products_dir` after their scratch store is gone, and a pack +without `psf_validation` would leave `star_cat_merge` short an exposure. +Enforced by tests/unit/test_persist_exp_props.py. """ import argparse From 5fcc99e1dae6cda09ec4dc03063188a7e56cfe74 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 03:16:41 +0200 Subject: [PATCH 62/85] test(workflow): the campaign name has one source, and products one root A config-lineage check for campaign-name-is-run. The campaign name is bound in the Snakefile, which the generic lineage checker cannot parse, so this reads the workflow sources directly: CAMPAIGN is bound once from config["run"], nothing reads a `campaign` key or names a campaign after a root directory, every merge rule's --campaign is that binding, and the machine defaults put `$run` and the index under the persistent root. The product-path helpers are lifted from the Snakefile and evaluated against sentinel roots, so the check is on the paths they return. Mutation-checked: a second binding, a campaign key, a hardcoded --campaign, a product or template moved to the scratch root, a dropped campaign name, and a literal campaign or off-root index in config.yaml each fail it. Co-Authored-By: Claude Fable 5.1 --- tests/unit/test_campaign_lineage.py | 149 ++++++++++++++++++++++++++++ workflow/CONTRACTS | 3 +- 2 files changed, 151 insertions(+), 1 deletion(-) create mode 100644 tests/unit/test_campaign_lineage.py diff --git a/tests/unit/test_campaign_lineage.py b/tests/unit/test_campaign_lineage.py new file mode 100644 index 000000000..492ec0ec8 --- /dev/null +++ b/tests/unit/test_campaign_lineage.py @@ -0,0 +1,149 @@ +"""Config lineage of the campaign name and the campaign product paths. + +Enforces workflow/CONTRACTS ``campaign-name-is-run``: the campaign's name has +one source, the run config's ``run:``, and every campaign product path is +rooted in ``PRODUCTS_DIR``. + +The skill-level lineage checker scans ``.py`` only, and the name is bound in +the Snakefile, so this reads the workflow's sources directly. The path helpers +are lifted out of the Snakefile by name and evaluated against sentinel roots, +which tests what they RETURN rather than how they are spelled. +""" + +import re +from pathlib import Path + +import pytest +import yaml + +REPO_ROOT = Path(__file__).resolve().parents[2] +WORKFLOW = REPO_ROOT / "workflow" +SNAKEFILE = WORKFLOW / "Snakefile" +RULES = sorted((WORKFLOW / "rules").glob("*.smk")) +SCRIPTS = sorted((WORKFLOW / "scripts").glob("*.py")) +LAUNCHER = WORKFLOW / "bin" / "sp" + +SNAKEMAKE_SOURCES = [SNAKEFILE, *RULES] +ALL_SOURCES = [*SNAKEMAKE_SOURCES, *SCRIPTS, LAUNCHER] + +PRODUCTS = "/sentinel/products" +SCRATCH = "/sentinel/scratch" +CAMPAIGN = "campaign-sentinel" + +# Paths that hold the campaign's durable products, and whether each one +# carries the campaign name. Called with a tile/exposure ID where they take one. +PRODUCT_HELPERS = { + "final_cat": (("210.282",), False), + "final_cat_hdf5": ((), True), + "full_starcat": ((), True), + "prod_exp_dir": (("2605805",), False), + "prod_exp_manifest": (("2605805", "exp_persist"), False), + "prod_exp_tar": (("2605805",), False), +} +PRODUCT_TEMPLATES = ("PROD_TILE_DIR", "PROD_EXP_DIR") + + +def _code_lines(path): + """Source lines with comment-only lines dropped.""" + return [(n, line) for n, line in + enumerate(path.read_text().splitlines(), 1) + if not line.lstrip().startswith("#")] + + +def _snakefile_def(name): + """The text of one top-level ``def`` in the Snakefile.""" + text = SNAKEFILE.read_text() + m = re.search(rf"^def {name}\(.*?(?=^\S)", text, re.M | re.S) + assert m, f"Snakefile no longer defines {name}(); update PRODUCT_HELPERS" + return m.group(0) + + +def _snakefile_assignment(name): + m = re.search(rf"^{name}\s*=\s*(.+)$", SNAKEFILE.read_text(), re.M) + assert m, f"Snakefile no longer assigns {name}; update PRODUCT_TEMPLATES" + return m.group(1) + + +@pytest.fixture(scope="module") +def helpers(): + """The Snakefile's product-path helpers, bound to sentinel roots.""" + ns = {"Path": Path, "PRODUCTS_DIR": Path(PRODUCTS), + "RUN_DIR": Path(SCRATCH), "CAMPAIGN": CAMPAIGN} + for name in PRODUCT_HELPERS: + exec(_snakefile_def(name), ns) + for name in PRODUCT_TEMPLATES: + ns[name] = eval(_snakefile_assignment(name), ns) + return ns + + +def test_campaign_is_bound_once_from_run(): + """``CAMPAIGN`` has exactly one binding, and it reads ``config["run"]``.""" + bindings = [(p.name, n, line.strip()) for p in SNAKEMAKE_SOURCES + for n, line in _code_lines(p) + if re.match(r"\s*CAMPAIGN\s*=", line)] + assert [b[2] for b in bindings] == ['CAMPAIGN = config["run"]'], bindings + + +def test_no_second_campaign_source(): + """No source reads a ``campaign`` config key or names a campaign after a + root directory; either could disagree with ``run:``.""" + forbidden = re.compile( + r"""\[\s*['"]campaign['"]\s*\]""" + r"""|\.get\(\s*['"]campaign['"]""" + r"|\b(?:PRODUCTS_DIR|RUN_DIR|products_dir)\b[\w)\]]*" + r"(?:\.parent)*\.(?:name|stem)\b") + hits = [f"{p.relative_to(REPO_ROOT)}:{n}: {line.strip()}" + for p in ALL_SOURCES for n, line in _code_lines(p) + if forbidden.search(line)] + assert not hits, "second campaign-name source:\n" + "\n".join(hits) + + +def test_merge_rules_pass_campaign_from_campaign(): + """Every rule param named ``campaign`` is ``CAMPAIGN``, and every + ``--campaign`` on a shell line is that param.""" + params, flags = [], [] + for p in RULES: + for n, line in _code_lines(p): + m = re.match(r"\s*campaign\s*=\s*(.+?),?\s*$", line) + if m: + params.append((p.name, n, m.group(1))) + for arg in re.findall(r"--campaign\s+(\S+)", line): + flags.append((p.name, n, arg)) + assert params, "no rule passes a campaign param; update this test" + assert all(v == "CAMPAIGN" for *_, v in params), params + assert flags and all(a.strip("'\"") == "{params.campaign}" + for *_, a in flags), flags + + +@pytest.mark.parametrize("name", sorted(PRODUCT_HELPERS)) +def test_product_paths_are_rooted_in_products_dir(helpers, name): + args, named = PRODUCT_HELPERS[name] + path = str(helpers[name](*args)) + assert path.startswith(PRODUCTS + "/"), path + assert (CAMPAIGN in path) == named, path + + +@pytest.mark.parametrize("name", PRODUCT_TEMPLATES) +def test_product_templates_are_rooted_in_products_dir(helpers, name): + assert str(helpers[name]).startswith(PRODUCTS + "/"), helpers[name] + + +def _machine_outputs(): + config = yaml.safe_load((WORKFLOW / "config.yaml").read_text()) + for machine, entry in (config.get("machines") or {}).items(): + for input_type, defaults in entry.items(): + if isinstance(defaults, dict) and "outputs" in defaults: + yield f"{machine}.{input_type}", defaults["outputs"] + + +@pytest.mark.parametrize("where,outputs", list(_machine_outputs())) +def test_machine_defaults_name_the_campaign_by_run(where, outputs): + """Where a machine default sets the persistent root, the campaign in it + is ``$run``, and the index lives beneath it.""" + products = outputs.get("products_dir") + if products is None: + pytest.skip(f"{where} sets no products_dir") + assert products.rstrip("/").endswith("/$run"), products + index = outputs.get("index_db") + if index is not None: + assert index.startswith(products.rstrip("/") + "/"), (index, products) diff --git a/workflow/CONTRACTS b/workflow/CONTRACTS index 2faac3ca2..ac8bf35d4 100644 --- a/workflow/CONTRACTS +++ b/workflow/CONTRACTS @@ -13,7 +13,8 @@ campaign name takes it from there: `final_cat_.hdf5` and its machine defaults' `products_dir`/`index_db`. No rule or script reads a `campaign` key or derives a name from a directory (`PRODUCTS_DIR.name`); two sources that can disagree would file one campaign's merge under another's name. -Every campaign product path is rooted in `PRODUCTS_DIR`. +Every campaign product path is rooted in `PRODUCTS_DIR`. Enforced by +tests/unit/test_campaign_lineage.py. @sc [label:schema] final-cat-param-is-exact-allow-list Each input type's `config/*/final_cat.param` is the merged catalogue's exact From e768908de9175496cd02244a36547a160b4bf576 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 03:19:25 +0200 Subject: [PATCH 63/85] test(final_cat_merge): columns, per-epoch slots, never-fit rows, missing names Invariants for the three contracts on the galaxy-side merge, run through merge_final_cat.py by path as the rule runs it, over three tiny tile catalogues whose columns are shuffled, padded with unlisted names, and include objects ngmix never fit. - the merged datasets carry exactly the param file's columns, in its order; - every object's slot-n tuple (EXP_ID_n, CCD_n, HSM_G1/G2_PSF_n) is unchanged, matched on NUMBER; - never-fit rows (NGMIX_N_EPOCH == 0) arrive with their sentinels intact; - a tile short a listed per-epoch column stops the merge, naming it, and writes no output. That raise is the seam that would otherwise land a tile silently narrower, its slot's exposure identity gone. Each was checked against its failure mode by editing the code under test: dropping a listed column, keeping unlisted ones, swapping one family's slots, reordering one family's rows, dropping or NaN-filling never-fit rows, and turning the missing-column raise into a skip. Every one fails a test. Co-Authored-By: Claude Fable 5.1 --- scripts/python/create_final_cat.py | 3 +- tests/unit/test_final_cat_merge_invariants.py | 185 ++++++++++++++++++ workflow/CONTRACTS | 3 +- workflow/scripts/merge_final_cat.py | 3 +- 4 files changed, 191 insertions(+), 3 deletions(-) create mode 100644 tests/unit/test_final_cat_merge_invariants.py diff --git a/scripts/python/create_final_cat.py b/scripts/python/create_final_cat.py index c35216e5e..d68a20eda 100755 --- a/scripts/python/create_final_cat.py +++ b/scripts/python/create_final_cat.py @@ -382,7 +382,8 @@ def read_data(fits_file, params): never skipped or filled. `copy_data` keeps only columns present in the source, so this raise is the one place a missing name stops a merge, and without it a tile short a per-epoch slot would land in the merged file - silently narrower, with that slot's exposure identity gone. + silently narrower, with that slot's exposure identity gone. Enforced by + tests/unit/test_final_cat_merge_invariants.py. """ with fits.open(fits_file) as hdu_list: try: diff --git a/tests/unit/test_final_cat_merge_invariants.py b/tests/unit/test_final_cat_merge_invariants.py new file mode 100644 index 000000000..eebef1629 --- /dev/null +++ b/tests/unit/test_final_cat_merge_invariants.py @@ -0,0 +1,185 @@ +"""What ``final_cat_merge`` owes the catalogues it merges. + +Enforces three contracts on a small campaign of tile catalogues, run through +``workflow/scripts/merge_final_cat.py`` as the rule's shell runs it: + +* ``final-cat-param-is-exact-allow-list`` (workflow/CONTRACTS): the merged + datasets carry the param file's columns, in its order, and nothing else; +* ``read-data-raises-on-missing-column`` (create_final_cat.py): a tile short a + listed column stops the merge, naming the column; +* ``never-fit-rows-pass-through`` (merge_final_cat.py): objects ngmix never fit + reach the merged file unchanged. + +Plus the property the per-epoch families exist for: an object's slot-n tuple +(``EXP_ID_n``, ``CCD_n``, ``HSM_*_PSF_n``) is the same after the merge as in its +tile catalogue. Objects are matched on ``NUMBER``, so a merge may reorder rows +but not move one field of a row without the others. + +Failure modes each test was checked against, by editing the code under test: +a listed column dropped or an unlisted one kept (columns); one family's slots +swapped, or one family's rows reordered (per-epoch); never-fit rows dropped or +their sentinels filled (never-fit); the missing-column raise turned into a skip +(missing column). +""" + +import sqlite3 +import subprocess +import sys +from pathlib import Path + +import numpy as np +import pytest +from numpy.lib.recfunctions import repack_fields + +h5py = pytest.importorskip("h5py") +fits = pytest.importorskip("astropy.io.fits") + +REPO_ROOT = Path(__file__).resolve().parents[2] +SCRIPT = REPO_ROOT / "workflow" / "scripts" / "merge_final_cat.py" + +CAMPAIGN = "fixture-campaign" +TILES = ("210.282", "211.282", "212.283") +N_SLOTS = 3 +EPOCH_FAMILIES = ("EXP_ID", "CCD", "HSM_G1_PSF", "HSM_G2_PSF") +SCALARS = ( + ("NUMBER", "i4"), ("XWIN_WORLD", "f8"), ("NGMIX_N_EPOCH", "i4"), + ("NGMIX_MCAL_FLAGS", "i4"), ("NGMIX_G1_NOSHEAR", "f8"), + ("NGMIX_T_NOSHEAR", "f8"), +) +EPOCH_DTYPE = {"EXP_ID": "i4", "CCD": "i4", + "HSM_G1_PSF": "f8", "HSM_G2_PSF": "f8"} +# In each tile but not in the param file: the merge must leave these behind. +UNLISTED = (("MAG_UNLISTED", "f4"), ("EXP_ID_4", "i4"), ("CCD_4", "i4")) + +PARAM_LIST = [name for name, _ in SCALARS] + [ + f"{fam}_{n}" for n in range(1, N_SLOTS + 1) for fam in EPOCH_FAMILIES] + + +def _slot(fam, n): + return f"{fam}_{n}" + + +def _tile_catalogue(seed, rows=40, n_unfit=5): + """One tile's catalogue: listed columns, unlisted ones, never-fit rows.""" + rng = np.random.default_rng(seed) + columns = list(SCALARS) + [ + (_slot(fam, n), EPOCH_DTYPE[fam]) + for n in range(1, N_SLOTS + 1) for fam in EPOCH_FAMILIES + ] + list(UNLISTED) + # The catalogue's own column order is not the param file's. + columns = [columns[i] for i in rng.permutation(len(columns))] + cat = np.zeros(rows, dtype=columns) + cat["NUMBER"] = rng.permutation(rows) + 1 + cat["XWIN_WORLD"] = rng.uniform(0, 360, rows) + cat["MAG_UNLISTED"] = rng.uniform(18, 25, rows) + cat["EXP_ID_4"] = cat["CCD_4"] = -1 + n_epoch = rng.integers(1, N_SLOTS + 1, rows) + unfit = rng.choice(rows, n_unfit, replace=False) + n_epoch[unfit] = 0 + cat["NGMIX_N_EPOCH"] = n_epoch + cat["NGMIX_G1_NOSHEAR"] = np.where(n_epoch > 0, + rng.normal(0, 0.3, rows), -10.0) + cat["NGMIX_T_NOSHEAR"] = np.where(n_epoch > 0, + rng.uniform(0.1, 1, rows), 0.0) + for n in range(1, N_SLOTS + 1): + used = n_epoch >= n + # Distinct values across slots and rows, so any exchange shows. + cat[_slot("EXP_ID", n)] = np.where( + used, rng.choice(10**7, rows, replace=False), -1) + cat[_slot("CCD", n)] = np.where(used, rng.integers(0, 40, rows), -1) + cat[_slot("HSM_G1_PSF", n)] = np.where( + used, rng.normal(0, 0.05, rows), -10.0) + cat[_slot("HSM_G2_PSF", n)] = np.where( + used, rng.normal(0, 0.05, rows), -10.0) + return cat + + +def _campaign(root: Path, drop=None): + """Lay out a campaign the way the workflow does; return (argv, sources). + + ``drop`` removes one listed column from the last tile. + """ + products = root / "products" + sources = {} + for i, tile in enumerate(TILES): + cat = _tile_catalogue(seed=1000 + i) + if drop and tile == TILES[-1]: + keep = [c for c in cat.dtype.names if c != drop] + cat = repack_fields(cat[keep]) + path = products / "tiles" / tile[:2] / tile / f"final_cat-{tile}.fits" + path.parent.mkdir(parents=True) + fits.HDUList([fits.PrimaryHDU(), fits.BinTableHDU(cat)]).writeto(path) + sources[tile] = fits.getdata(path, 1) + tile_list = root / "tiles.txt" + tile_list.write_text("\n".join(TILES) + "\n") + index = root / "index.sqlite" + con = sqlite3.connect(index) + con.execute("CREATE TABLE tile_exposures(tile_id TEXT, exp_id TEXT)") + con.executemany("INSERT INTO tile_exposures VALUES (?, ?)", + [(t, "2605805") for t in TILES]) + con.commit() + con.close() + param = root / "final_cat.param" + param.write_text("# fixture schema\nNUMBER\n\n# per-epoch\n" + + "\n".join(PARAM_LIST[1:]) + "\n") + output = products / f"final_cat_{CAMPAIGN}.hdf5" + argv = [sys.executable, str(SCRIPT), "--products-dir", str(products), + "--tile-list", str(tile_list), "--index-db", str(index), + "--output", str(output), "--campaign", CAMPAIGN, + "--param-file", str(param)] + return argv, output, sources + + +@pytest.fixture(scope="module") +def merged(tmp_path_factory): + argv, output, sources = _campaign(tmp_path_factory.mktemp("campaign")) + run = subprocess.run(argv, capture_output=True, text=True) + assert run.returncode == 0, run.stderr + with h5py.File(output, "r") as f: + group = f[f"patches/{CAMPAIGN}"] + out = {tile: group[tile][()] for tile in group} + return out, sources + + +def _by_number(arr): + return arr[np.argsort(arr["NUMBER"])] + + +def test_columns_are_exactly_the_param_list(merged): + out, _ = merged + assert sorted(out) == sorted(TILES) + for tile, data in out.items(): + assert list(data.dtype.names) == PARAM_LIST, tile + + +def test_per_epoch_tuples_survive_the_merge(merged): + out, sources = merged + for tile, data in out.items(): + got, want = _by_number(data), _by_number(sources[tile]) + assert np.array_equal(got["NUMBER"], want["NUMBER"]), tile + for n in range(1, N_SLOTS + 1): + for fam in EPOCH_FAMILIES: + col = _slot(fam, n) + assert np.array_equal(got[col], want[col]), (tile, col) + + +def test_never_fit_rows_pass_through(merged): + out, sources = merged + for tile, data in out.items(): + src = sources[tile] + assert len(data) == len(src), tile + got = _by_number(data[data["NGMIX_N_EPOCH"] == 0]) + want = _by_number(src[src["NGMIX_N_EPOCH"] == 0]) + assert len(want) > 0, "fixture lost its never-fit rows" + for col in ("NUMBER", "NGMIX_G1_NOSHEAR", "NGMIX_T_NOSHEAR", + "NGMIX_MCAL_FLAGS"): + assert np.array_equal(got[col], want[col]), (tile, col) + + +def test_a_tile_missing_a_listed_column_stops_the_merge(tmp_path): + missing = _slot("HSM_G1_PSF", N_SLOTS) + argv, output, _ = _campaign(tmp_path, drop=missing) + run = subprocess.run(argv, capture_output=True, text=True) + assert run.returncode != 0, "merge succeeded over a tile short a column" + assert missing in run.stderr + assert not output.exists() diff --git a/workflow/CONTRACTS b/workflow/CONTRACTS index ac8bf35d4..04b7ebb5b 100644 --- a/workflow/CONTRACTS +++ b/workflow/CONTRACTS @@ -23,7 +23,8 @@ and the reader raises on a listed column a tile lacks. So a name belongs there only if `make_cat` writes it on EVERY tile. A per-epoch family (`EXP_ID_n`, `CCD_n`, `HSM_*_PSF_n`) qualifies only with a fixed slot count, since `make_cat` otherwise sizes it from the tile's own maximum `N_EPOCH`. Enforced -by tests/module/test_psf_grammar_properties.py (the shipped names). +by tests/unit/test_final_cat_merge_invariants.py (the merge) and +tests/module/test_psf_grammar_properties.py (the shipped names). @sc [label:hazard] unit-pre-changes-at-campaign-boundary Every line `unit_pre()` emits is part of each rule's `params.pre`, and both diff --git a/workflow/scripts/merge_final_cat.py b/workflow/scripts/merge_final_cat.py index 1c4be85c9..1ad6fbdff 100644 --- a/workflow/scripts/merge_final_cat.py +++ b/workflow/scripts/merge_final_cat.py @@ -76,7 +76,8 @@ sentinel values (`NGMIX_MCAL_FLAGS == 0`, ellipticities `-10`, `T == 0`), so `NGMIX_MCAL_FLAGS == 0` is not a validity cut: consumers select fitted objects with `NGMIX_N_EPOCH > 0`. The merge neither fills these rows nor drops them; -that selection belongs to the consumer. +that selection belongs to the consumer. Enforced by +tests/unit/test_final_cat_merge_invariants.py. """ import argparse From e3f6f4adbc838032b5bc2d81460c312670252a31 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 03:22:45 +0200 Subject: [PATCH 64/85] docs(astra): correct seven rationale claims against the code - epoch_provenance: names keep their trailing p; EXP_PREFIX is a no-op [LINT] - fit_initialisation: only the PSF guesser takes the catalogue flux; an exception in Ngmix.process drops the object with no row - star_galaxy_classification: thresholds come from SM_STAR_THRESH / SM_GAL_THRESH, which the committed config does not set - psf_train_validation_split: seeded from the unit's file number - stamp_positioning: an out-of-image stamp centre raises - object_position_columns: tile stamps are cut at XWIN_IMAGE (COORD=PIX) - mark the PSFEx built-in SAMPLE_* behaviour and the 33-px trim unverified - record the galaxy prior reused for PSF fits and the silent epoch drops before the 1/3 cut; carry stale completeness, exposure.smk, _mode and pixel-scale comments as [LINT] Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_014bvNTrAmZxcfb1ee83ApPK --- astra.yaml | 290 ++++++++++++++++++++++++++++++++--------------------- 1 file changed, 173 insertions(+), 117 deletions(-) diff --git a/astra.yaml b/astra.yaml index 1b27db982..7ad57e61f 100644 --- a/astra.yaml +++ b/astra.yaml @@ -63,16 +63,21 @@ decisions: per_unit_completeness: label: Per-unit completeness gate rationale: >- - Every rule checks its products against a nominal per-runner count: a - runner below its count fails the unit (an exposure or a tile), so a - partial unit never enters the catalogue; the missing unit's objects do - not appear at all. The one tolerated shortfall is the exposure-side + Every rule that runs shapepipe_run checks its products against a + nominal per-runner count (tile_ngmix per chunk): a runner below + its count fails the unit (an exposure or a tile), so a partial + unit never enters the catalogue; the missing unit's objects do not + appear at all. The one tolerated shortfall is the exposure-side psfex_interp VALIDATION output, where a CCD whose model fails the - acceptance gate (star_selection_psf.psf_acceptance_thresholds) produces - nothing and the unit only warns; the MCCD chain, never run in a - campaign, warns on every runner. Science-path PSF rejection does not go - through this table: psfex_interp drops the epoch per object inside the - tile run. + acceptance gate (star_selection_psf.psf_acceptance_thresholds) + produces nothing and the unit only warns; the MCCD chain, never + run in a campaign, warns on every runner. Science-path PSF + rejection does not go through this table: psfex_interp drops the + epoch per object inside the tile run. [LINT] the completeness.py + docstring gives tile psfex_interp as its warn example, but that + runner is mandatory, and a workflow/rules/exposure.smk comment + says a floor's :warn tolerates setools rejecting a sparse CCD, but + setools has a mandatory count and no floor exists. Anchor: workflow/scripts/completeness.py::COMPLETENESS. default: exact_counts options: @@ -232,18 +237,20 @@ analyses: psf_star_mask_veto: label: PSF-star candidates rejected on instrument flags only rationale: >- - The star selection cuts IMAFLAGS_ISO == 0 and nothing else from the - masks. mask_query sits in the exposure module chain between - SExtractor and setools: when MASK_PATHS names maps it writes MASK_EXT - (0 clean, nonzero flagged; off-coverage counts as clean) onto each - CCD's catalogue. MASK_PATHS ships commented out, so the committed - module passes the catalogue through with no MASK_EXT column. The - intended map is the UNIONS star-body product (bit 2); halo bits 0 - and 1 are excluded because halos say nothing about whether a star is - a good PSF sample. Imposing the veto is one line per mask block in - star_selection.setools (MASK_EXT == 0). [LINT] the - workflow/rules/exposure.smk docstring names the column FLAG_EXT; the - code writes MASK_EXT. + The star selection cuts IMAFLAGS_ISO == 0 and nothing else + from the masks. mask_query sits in the exposure module + chain between SExtractor and setools: when MASK_PATHS + names maps it writes MASK_EXT (0 clean, nonzero flagged; + off-coverage counts as clean) onto each CCD's catalogue. + MASK_PATHS ships commented out, so the committed module + passes the catalogue through with no MASK_EXT column. The + intended map is the UNIONS star-body product (bit 2); halo + bits 0 and 1 are excluded because halos say nothing about + whether a star is a good PSF sample. Imposing the veto is + one line per mask block in star_selection.setools + (MASK_EXT == 0). [LINT] the workflow/rules/exposure.smk + docstring names the column FLAG_EXT and says setools cuts + on it; the code writes MASK_EXT and nothing cuts on it. Anchor: workflow/config/cfis/config_exp_psfex.ini; workflow/config/cfis/star_selection.setools#MASK:star_selection.IMAFLAGS_ISO; src/shapepipe/modules/mask_query_runner.py::mask_query_runner; @@ -504,11 +511,12 @@ analyses: epoch_membership_ccd_bounds: label: Which exposure CCDs an object belongs to (N_EPOCH) rationale: >- - CCD_SIZE = 33,2080,1,4612 with strict inequalities: the 33-px left - trim removes a CCD strip from epoch membership, and a WCS inversion - failure skips the CCD, lowering N_EPOCH. This sets how many exposures - enter each galaxy's multi-epoch fit. The trim is unexplained beyond - "number of pixels in a CCD". + CCD_SIZE = 33,2080,1,4612 with strict inequalities: + positions with x at or below 33 are outside the bounds, + and a WCS inversion failure skips the CCD, lowering + N_EPOCH. This sets how many exposures enter each galaxy's + multi-epoch fit. The x range spans 2048 px; whether the + excluded strip is prescan or science pixels is unverified. Anchor: workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.CCD_SIZE; src/shapepipe/modules/sextractor_package/sextractor_script.py::make_post_process; src/shapepipe/modules/sextractor_package/sextractor_script.py::ccd_candidate_mask. @@ -517,7 +525,7 @@ analyses: trimmed_bounds_33_2080: label: "x in (33,2080), y in (1,4612), strict" full_ccd: - label: Full 1-2048 x-range, inclusive bounds + label: Untrimmed x-range, inclusive bounds prior_insights: guinot22_stacked_detection: claim: >- @@ -590,11 +598,17 @@ analyses: epoch_provenance_from_tile_history: label: Epoch sets parsed from tile FITS HISTORY cards rationale: >- - A tile's contributing exposures are column 3 (COLNUM) of each HISTORY - line, with prefix p stripped and duplicates removed: the coadd's own - provenance is trusted as the epoch list. A mis-parse changes N_EPOCH - and which exposures are fit. + A tile's contributing exposures are the file names in + column 3 (COLNUM) of each HISTORY line, stripped of their + extension and deduplicated: the coadd's own provenance is + trusted as the epoch list. Names keep their trailing p + (2243881p); downstream code drops it when it needs the + bare exposure ID. A mis-parse changes N_EPOCH and which + exposures are fit. [LINT] EXP_PREFIX = p is passed to + removeprefix, which does nothing to names where p is a + suffix, so the key has no effect. Anchor: workflow/config/cfis/config_tile_Fe.ini#FIND_EXPOSURES_RUNNER.COLNUM; + workflow/config/cfis/config_tile_Fe.ini#FIND_EXPOSURES_RUNNER.EXP_PREFIX; src/shapepipe/modules/find_exposures_package/find_exposures.py::FindExposures.get_exposure_list. default: history_parse options: @@ -603,13 +617,18 @@ analyses: object_position_columns: label: Windowed centroids (XWIN/YWIN) define every position rationale: >- - PSF interpolation sites, tile and multi-epoch stamp centres, and the - catalogue position all use SExtractor's windowed centroid: - XWIN_WORLD/YWIN_WORLD on the tile side, XWIN_IMAGE/YWIN_IMAGE on - exposures. Windowed, isophotal and model centroids differ - systematically for blends and asymmetric galaxies, and the centroid - feeds the position seed and the centroid prior. - Anchor: workflow/config/cfis/config_tile_PiViVi_psfex.ini#PSFEX_INTERP_RUNNER.POSITION_PARAMS; + PSF interpolation sites, tile and multi-epoch stamp + centres, and the catalogue position all use SExtractor's + windowed centroid. Tile stamps are cut at + XWIN_IMAGE/YWIN_IMAGE in tile pixels (COORD = PIX); PSF + interpolation and multi-epoch stamps use + XWIN_WORLD/YWIN_WORLD; exposure-side PSF validation uses + XWIN_IMAGE/YWIN_IMAGE. Windowed, isophotal and model + centroids differ systematically for blends and asymmetric + galaxies, and the centroid feeds the position seed and the + centroid prior. + Anchor: workflow/config/cfis/config_tile_PiViVi_psfex.ini#VIGNETMAKER_RUNNER_RUN_1.POSITION_PARAMS; + workflow/config/cfis/config_tile_PiViVi_psfex.ini#PSFEX_INTERP_RUNNER.POSITION_PARAMS; workflow/config/cfis/config_tile_PiViVi_psfex.ini#VIGNETMAKER_RUNNER_RUN_2.POSITION_PARAMS; workflow/config/cfis/config_exp_psfex.ini#PSFEX_INTERP_RUNNER.POSITION_PARAMS. default: xwin_windowed @@ -619,14 +638,17 @@ analyses: stamp_positioning_and_padding: label: Nearest-pixel stamp extraction with zero padding rationale: >- - [HARDCODED] stamps are cut around the pixel nearest the object's - position, with no sub-pixel interpolation; the sub-pixel remainder is - stored as the stamp's OFFSET, which ngmix uses as the Jacobian origin - (shape_measurement.centroid_source), so extraction and centroid - prior share one rounding. Multi-epoch stamps take the position from - the tile world coordinate through the stored per-CCD WCS. Objects - whose stamp overruns an image edge are kept, with out-of-image pixels - zero-filled; there is no boundary rejection. + [HARDCODED] stamps are cut around the pixel nearest the + object's position, with no sub-pixel interpolation; the + sub-pixel remainder is stored as the stamp's OFFSET, which + ngmix uses as the Jacobian origin + (shape_measurement.centroid_source), so extraction and + centroid prior share one rounding. Multi-epoch stamps take + the position from the tile world coordinate through the + stored per-CCD WCS. Objects whose stamp overruns an image + edge are kept, with out-of-image pixels zero-filled. A + stamp centre that rounds outside the image raises, which + fails the vignet run for the whole tile. Anchor: src/shapepipe/modules/vignetmaker_package/vignetmaker.py::get_stamps; src/shapepipe/modules/vignetmaker_package/vignetmaker.py::VignetMaker._get_stamp_me. default: round_and_zero_pad @@ -683,17 +705,20 @@ analyses: star_selection_box: label: Stellar-locus selection, magnitude window and FWHM window around the mode rationale: >- - 18 < MAG_AUTO < 22, |FWHM - mode| <= 0.2 px, FLAGS == 0 and - IMAFLAGS_ISO == 0. The mode is computed on a preselection (MAG_AUTO < - 21, FWHM 0.3-1.5 arcsec at 0.187 arcsec/px) by an iterative - histogram-zoom estimator that falls back to the median below 20 - objects, so small-N behaviour changes selection on sparse CCDs. - PSFEx's own selection is off (SAMPLE_AUTOSELECT N), but its - compiled-in cuts still apply (psfex_candidate_vetting). [LINT] the - file's statistics log the FWHM cut as mode +- 0.1 px and its plot - uses 0.186 arcsec/px, while the applied cut is +- 0.2 px at 0.187. + 18 < MAG_AUTO < 22, |FWHM - mode| <= 0.2 px, FLAGS == 0 + and IMAFLAGS_ISO == 0. The mode is computed on a + preselection (MAG_AUTO < 21, FWHM 0.3-1.5 arcsec at 0.187 + arcsec/px) by an iterative histogram-zoom estimator that + falls back to the median below 20 objects, so small-N + behaviour changes selection on sparse CCDs. PSFEx's own + selection is off (SAMPLE_AUTOSELECT N); see + psfex_candidate_vetting for what PSFEx may still apply. + [LINT] the file's statistics log the FWHM cut as mode +- + 0.1 px and its plot uses 0.186 arcsec/px, while the + applied cut is +- 0.2 px at 0.187; the _mode docstring + puts the median fallback at 10 objects, the code at 20. Anchor: workflow/config/cfis/star_selection.setools#MASK:star_selection.MAG_AUTO; - workflow/config/cfis/star_selection.setools#MASK:preselect.FWHM_IMAGE; + workflow/config/cfis/star_selection.setools#MASK:preselect.MAG_AUTO; workflow/config/cfis/default.psfex#SAMPLE_AUTOSELECT; src/shapepipe/pipeline/str_handler.py::StrInterpreter._mode. default: mode_centred_box @@ -714,13 +739,15 @@ analyses: psf_train_validation_split: label: Seeded 80/20 star split, model fit vs held-out validation rationale: >- - RAND_SPLIT RATIO 20: the 80% sample fits the PSFEx model and feeds the - tile multi-epoch interpolation (ME_DOT_PSF_PATTERN); the 20% sample - is the independent residual diagnostic (psfex_interp VALIDATION - mode). The split trades training stars per CCD, which interacts with - the acceptance gate, against an independent residual test. It is - deterministic: a permutation seeded from the unit's file number, so - the PSF star sample is a pure function of the input catalogue. + RAND_SPLIT RATIO 20: the 80% sample fits the PSFEx model + and feeds the tile multi-epoch interpolation + (ME_DOT_PSF_PATTERN); the 20% sample is the independent + residual diagnostic (psfex_interp VALIDATION mode). The + split trades training stars per CCD, which interacts with + the acceptance gate, against an independent residual test. + It is deterministic: a permutation seeded from the digits + of the unit's file number, so a given CCD gets the same + split on every run. Anchor: workflow/config/cfis/star_selection.setools#RAND_SPLIT:star_split.RATIO; src/shapepipe/modules/setools_package/setools.py::SETools._make_rand_split; workflow/config/cfis/config_exp_psfex.ini#PSFEX_RUNNER.FILE_PATTERN; @@ -743,20 +770,21 @@ analyses: psfex_candidate_vetting: label: PSFEx built-in candidate cuts, unpinned rationale: >- - default.psfex sets SAMPLE_AUTOSELECT N but omits SAMPLE_MINSN, - SAMPLE_MAXELLIP, SAMPLE_FWHMRANGE and SAMPLE_VARIABILITY, so - [HARDCODED] PSFEx's compiled-in defaults apply (MINSN 20, MAXELLIP - 0.3, FWHMRANGE 2-10 px, VARIABILITY 0.2): a second star selection no - config records, which changes with the PSFEx version. - BADPIXEL_FILTER N and PSF_RECENTER N accept flagged star vignets - unfiltered and do not recentre candidates. + default.psfex sets SAMPLE_AUTOSELECT N but omits + SAMPLE_MINSN, SAMPLE_MAXELLIP, SAMPLE_FWHMRANGE and + SAMPLE_VARIABILITY, so [HARDCODED] whatever PSFEx compiles + in for them governs, changing with the PSFEx version; + which of these cuts still act with SAMPLE_AUTOSELECT N is + unverified. BADPIXEL_FILTER N and PSF_RECENTER N accept + flagged star vignets unfiltered and do not recentre + candidates. Anchor: workflow/config/cfis/default.psfex#SAMPLE_AUTOSELECT; workflow/config/cfis/default.psfex#BADPIXEL_FILTER; workflow/config/cfis/default.psfex#PSF_RECENTER. default: builtin_defaults options: builtin_defaults: - label: Compiled-in SAMPLE_* defaults, no bad-pixel filter + label: PSFEx built-in SAMPLE_* values, no bad-pixel filter pinned_explicit: label: Write the SAMPLE_* values explicitly into default.psfex psf_modelling_software: @@ -989,13 +1017,19 @@ analyses: fit_initialisation: label: Fit guesses and retries rationale: >- - [HARDCODED] TPSFFluxAndPriorGuesser (galaxy) and TFluxGuesser (PSF) - start from T = 0.25 and the catalogue flux; the galaxy runner retries - 5 times, the PSF runner twice. With a non-convex likelihood the guess - and retries decide which objects converge; a failed fit is flagged, - not raised. Guinot+22 initialised the whole guess vector from HSM - adaptive moments on each sheared image; the code no longer does. - Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::make_runners. + [HARDCODED] the galaxy guesser (TPSFFluxAndPriorGuesser) + starts from T = 0.25 with a flux taken from a PSF-flux + fit; the PSF guesser (TFluxGuesser) starts from T = 0.25 + and the catalogue flux. The galaxy runner retries 5 times, + the PSF runner twice. With a non-convex likelihood the + guess and retries decide which objects converge. A failed + ngmix fit is flagged; any exception during an object's fit + drops it with no ngmix row, leaving it to the catalogue + sentinels (catalogue_assembly.failure_sentinels). + Guinot+22 initialised the whole guess vector from HSM + adaptive moments on each sheared image; the code does not. + Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::make_runners; + src/shapepipe/modules/ngmix_package/ngmix.py::Ngmix.process. default: prior_guess_t025_ntry5_2 options: prior_guess_t025_ntry5_2: @@ -1006,15 +1040,21 @@ analyses: fit_priors: label: ngmix joint prior rationale: >- - [HARDCODED] ellipticity GPriorBA with sigma 0.4; flat T in [-1, 1e3] - and flat F in [-100, 1e9], with negative support (the bounds decide - which noisy fits survive and which rail); a centroid prior of width - one pixel scale, PIXEL_SCALE 0.186 arcsec (derived from the WCS when - the key is absent). Prior width drives noise bias; no rationale - recorded. Guinot+22 states a flat F in [-1e4, 1e9] and a flat r50 - prior rather than T. [LINT] the epoch stamps are exposure pixels, and - star selection uses 0.187 arcsec/px. + [HARDCODED] ellipticity GPriorBA with sigma 0.4; flat T in + [-1, 1e3] and flat F in [-100, 1e9], with negative support + (the bounds decide which noisy fits survive and which + rail); a centroid prior of width one pixel scale, + PIXEL_SCALE 0.186 arcsec (derived from the WCS when the + key is absent). The same joint prior also constrains the + PSF fits: the PSF fitter is built with the galaxy prior. + Prior width drives noise bias; no rationale recorded. + Guinot+22 states a flat F in [-1e4, 1e9] and a flat r50 + prior rather than T. [LINT] the epoch stamps are exposure + pixels, and star selection uses 0.187 arcsec/px; the + ngmix_runner comment says pixel scale also sets a noise + window, but get_noise is never called. Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::get_prior; + src/shapepipe/modules/ngmix_package/ngmix.py::make_runners; workflow/config/cfis/config_tile_Ng_template.ini#NGMIX_RUNNER.PIXEL_SCALE. default: gpriorba04_flat options: @@ -1317,21 +1357,28 @@ analyses: epoch_masked_fraction_cut: label: Per-epoch masked-fraction cut rationale: >- - [HARDCODED] an epoch whose stamp has more than 1/3 of its pixels - flagged (any nonzero flag bit, including the tile-coverage bit 2**10 - set where the tile vignet is off-image) is dropped from the - multi-epoch fit; an object with no surviving epoch has no shape. DES - was stricter. Y1 rejected any epoch with a masked or zero-weight - pixel, and any whose central 4-pixel region was masked. Y3 cut at 10% - of raw zero-weight pixels. Y6 dropped images more than 10% missing - and cut objects at mfrac < 0.1, which its simulations show avoids - calibration bias. Sheldon & Huff 2017 recommend dropping problematic - epochs when many are available. The cut interacts with defect_fill: - symmetrizing roughly doubles the masked fraction, so the cut should - be applied after symmetrizing. UNIONS has fewer epochs than DES, so - the cost in effective number density has to be measured, not - assumed. A central-region veto (drop the epoch if a defect lies - within a few pixels of the centre) is a cheap refinement. + [HARDCODED] an epoch whose stamp has more than 1/3 of its + pixels flagged (any nonzero flag bit, including the + tile-coverage bit 2**10 set where the tile vignet is + off-image) is dropped from the multi-epoch fit; an object + with no surviving epoch has no shape. The cut counts flag + pixels only, not zero-weight or invalid-RMS pixels. Before + it, an epoch is dropped silently if its galaxy stamp is + all zeros or its background-subtracted noise estimate + (sigma_mad) is not positive. DES was stricter. Y1 rejected + any epoch with a masked or zero-weight pixel, and any + whose central 4-pixel region was masked. Y3 cut at 10% of + raw zero-weight pixels. Y6 dropped images more than 10% + missing and cut objects at mfrac < 0.1, which its + simulations show avoids calibration bias. Sheldon & Huff + 2017 recommend dropping problematic epochs when many are + available. The cut interacts with defect_fill: + symmetrizing roughly doubles the masked fraction, so the + cut should be applied after symmetrizing. UNIONS has fewer + epochs than DES, so the cost in effective number density + has to be measured, not assumed. A central-region veto + (drop the epoch if a defect lies within a few pixels of + the centre) is a cheap refinement. Anchor: src/shapepipe/modules/ngmix_package/ngmix.py::prepare_postage_stamps. default: one_third options: @@ -1705,15 +1752,20 @@ analyses: star_galaxy_classification: label: Star/galaxy separation deferred out of the pipeline rationale: >- - SM_DO_CLASSIFICATION=False and no spread-model input is wired, so the - catalogue ships every object and separation happens downstream. The - dormant make_cat classifier uses class = sm + 2 sm_err with stars at - |class| < 0.003 and galaxies at class > 0.01, thresholds fixed in the - function signature. Guinot+22 selected galaxies in the pipeline at s - + 2 sigma_s > 0.0003 together with s > 0 and 20 < MAG_AUTO < 26; the - dormant code implements only the spread-model test, at a galaxy - boundary thirty times higher. + SM_DO_CLASSIFICATION=False and no spread-model input is + wired, so the catalogue ships every object and separation + happens downstream. The dormant make_cat classifier uses + class = sm + 2 sm_err with stars at |class| < + SM_STAR_THRESH and galaxies at class > SM_GAL_THRESH, read + from config when classification is on; the function + defaults are 0.003 and 0.01, and the committed config sets + neither key, so enabling classification alone raises. + Guinot+22 selected galaxies in the pipeline at s + 2 + sigma_s > 0.0003 together with s > 0 and 20 < MAG_AUTO < + 26; the dormant code implements only the spread-model + test. Anchor: workflow/config/cfis/config_tile_Mc.ini#MAKE_CAT_RUNNER.SM_DO_CLASSIFICATION; + src/shapepipe/modules/make_cat_runner.py::make_cat_runner; src/shapepipe/modules/make_cat_package/make_cat.py::save_sm_data; workflow/config/cfis/final_cat.param#SPREAD_CLASS. default: deferred_downstream @@ -1721,10 +1773,11 @@ analyses: deferred_downstream: label: No in-pipeline classification; catalogue ships all objects spread_model_inline: - label: spread_model classification in make_cat (0.003 / 0.01) + label: spread_model classification in make_cat description: >- Not wired in the workflow; needs spread_model_runner after the - PSF interpolation and its output added to make_cat's inputs. + PSF interpolation, its output added to make_cat's inputs, and + SM_STAR_THRESH / SM_GAL_THRESH set. tile_overlap_handling: label: Tile-overlap duplicates neither removed nor flagged rationale: >- @@ -1752,12 +1805,15 @@ analyses: failure_sentinels: label: Objects without shape measurements kept, with sentinel values rationale: >- - [HARDCODED] detections with no ngmix row stay in the catalogue with - sentinels: sizes, fluxes, magnitudes and flags 0, flux and magnitude - errors -1, ellipticities -10, size errors 1e30. The sentinels define - what a downstream cut must exclude: a failed object's - NGMIX_MCAL_FLAGS reads 0, the success value, so a cut on flags alone - keeps it. + [HARDCODED] detections with no ngmix row (no surviving + epoch, or an exception during the fit) stay in the + catalogue with sentinels: sizes, fluxes, magnitudes and + flags 0, flux and magnitude errors -1, ellipticities and + their errors -10, size errors 1e30, and NGMIX_N_EPOCH 0. + The sentinels define what a downstream cut must exclude: a + failed object's NGMIX_MCAL_FLAGS reads 0, the success + value, so a cut on flags alone keeps it; a cut on + NGMIX_N_EPOCH > 0 removes it. Anchor: src/shapepipe/modules/make_cat_package/make_cat.py::SaveCatalogue._save_ngmix_data. default: sentinel_values options: From 3358105b5b02b537d0dd23882c76c02440a74ba7 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 03:23:46 +0200 Subject: [PATCH 65/85] config(cfis): final_cat.param carries the 12 per-epoch slots (EXP_ID, CCD, HSM_*_PSF) make_cat writes these six families with N_EPOCH_SLOTS = 12 slots on every tile (#905), so the exact allow-list can name all 72 without failing a tile that saw fewer epochs. final_cat_merge then keeps the per-epoch identity and PSF shape through the campaign merge (#890, #903). Co-Authored-By: Claude Fable 5.1 --- workflow/config/cfis/final_cat.param | 77 ++++++++++++++++++++++++++++ 1 file changed, 77 insertions(+) diff --git a/workflow/config/cfis/final_cat.param b/workflow/config/cfis/final_cat.param index f3fa39677..0d5653ef8 100644 --- a/workflow/config/cfis/final_cat.param +++ b/workflow/config/cfis/final_cat.param @@ -45,6 +45,83 @@ NGMIX_G2_PSF_ORIG_NOSHEAR N_EPOCH NGMIX_N_EPOCH +# Per-epoch identity and PSF shape, slot n = 1..12 (config_tile_Mc.ini's +# N_EPOCH_SLOTS): EXP_ID_n/CCD_n name the exposure and CCD of epoch n, aligned +# with HSM_*_PSF_n; empty slots carry -1 (ids), -10 (g), 0 (T), 1 (flag). The +# count is fixed at the source so every tile shares this schema. +EXP_ID_1 +EXP_ID_2 +EXP_ID_3 +EXP_ID_4 +EXP_ID_5 +EXP_ID_6 +EXP_ID_7 +EXP_ID_8 +EXP_ID_9 +EXP_ID_10 +EXP_ID_11 +EXP_ID_12 +CCD_1 +CCD_2 +CCD_3 +CCD_4 +CCD_5 +CCD_6 +CCD_7 +CCD_8 +CCD_9 +CCD_10 +CCD_11 +CCD_12 +HSM_G1_PSF_1 +HSM_G1_PSF_2 +HSM_G1_PSF_3 +HSM_G1_PSF_4 +HSM_G1_PSF_5 +HSM_G1_PSF_6 +HSM_G1_PSF_7 +HSM_G1_PSF_8 +HSM_G1_PSF_9 +HSM_G1_PSF_10 +HSM_G1_PSF_11 +HSM_G1_PSF_12 +HSM_G2_PSF_1 +HSM_G2_PSF_2 +HSM_G2_PSF_3 +HSM_G2_PSF_4 +HSM_G2_PSF_5 +HSM_G2_PSF_6 +HSM_G2_PSF_7 +HSM_G2_PSF_8 +HSM_G2_PSF_9 +HSM_G2_PSF_10 +HSM_G2_PSF_11 +HSM_G2_PSF_12 +HSM_T_PSF_1 +HSM_T_PSF_2 +HSM_T_PSF_3 +HSM_T_PSF_4 +HSM_T_PSF_5 +HSM_T_PSF_6 +HSM_T_PSF_7 +HSM_T_PSF_8 +HSM_T_PSF_9 +HSM_T_PSF_10 +HSM_T_PSF_11 +HSM_T_PSF_12 +HSM_FLAG_PSF_1 +HSM_FLAG_PSF_2 +HSM_FLAG_PSF_3 +HSM_FLAG_PSF_4 +HSM_FLAG_PSF_5 +HSM_FLAG_PSF_6 +HSM_FLAG_PSF_7 +HSM_FLAG_PSF_8 +HSM_FLAG_PSF_9 +HSM_FLAG_PSF_10 +HSM_FLAG_PSF_11 +HSM_FLAG_PSF_12 + # Blend flag: coadd seg stamp held a non-central footprint (shapepipe#776) NGMIX_NEIGHBOUR_FLAG From bac6eea042ace454a3bd5ae15fbe8c3598388908 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 03:27:40 +0200 Subject: [PATCH 66/85] =?UTF-8?q?docs(astra):=20two=20lints=20this=20stack?= =?UTF-8?q?=20resolves=20=E2=80=94=20IMAFLAGS=5FISO=20is=20not=20merged,?= =?UTF-8?q?=20the=20MCCD=20chain=20runs?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit final_cat.param no longer requests IMAFLAGS_ISO (the tile chain never produced it), and #894's MCCD port reads exp_split's flag files through mask_query_runner, so both [LINT] sentences and their dead anchors go. Co-Authored-By: Claude Fable 5.1 --- astra.yaml | 15 +++++++-------- 1 file changed, 7 insertions(+), 8 deletions(-) diff --git a/astra.yaml b/astra.yaml index 7ad57e61f..c342c6b13 100644 --- a/astra.yaml +++ b/astra.yaml @@ -488,13 +488,12 @@ analyses: DETECTION_IMAGE=False and FLAG_IMAGE=False with the default_noimaflags.param column list: tiles have no instrument flag image and no detection coadd exists, so detection sees every tile - pixel and the tile catalogue carries no IMAFLAGS_ISO. [LINT] - final_cat.param, read by the post-processing merge, requests - IMAFLAGS_ISO, which the tile chain never produces. + pixel and the tile catalogue carries no IMAFLAGS_ISO; final_cat.param, + the merge's exact allow-list, does not request it. Anchor: workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.DETECTION_IMAGE; workflow/config/cfis/config_tile_Sx.ini#SEXTRACTOR_RUNNER.FLAG_IMAGE; workflow/config/cfis/default_noimaflags.param; - workflow/config/cfis/final_cat.param#IMAFLAGS_ISO. + workflow/config/cfis/final_cat.param. default: sx_nomask_single_image options: sx_nomask_single_image: @@ -795,13 +794,13 @@ analyses: independently. MCCD (Liaudat+2021) fits all 40 CCDs at once with a hybrid local+global model (N_COMP_LOC 8, D_COMP_GLOB 8, MIN_N_STARS 20, RMSE_THRESH 1.25); the completeness table treats its counts as - warnings because no campaign has run it. [LINT] config_exp_mccd.ini - still reads pipeline_flag images from mask_runner, which no longer - exists, so the MCCD exposure chain cannot run as committed. + warnings because no campaign has run it. The MCCD exposure chain + reads exp_split's image/weight/flag files and queries the sky masks + with mask_query_runner, like the PSFEx chain. Anchor: workflow/config.yaml; workflow/config/cfis/config_MCCD.ini#INSTANCE.N_COMP_LOC; workflow/config/cfis/config_MCCD.ini#INPUTS.MIN_N_STARS; - workflow/config/cfis/config_exp_mccd.ini#SEXTRACTOR_RUNNER.INPUT_MODULE; + workflow/config/cfis/config_exp_mccd.ini#SEXTRACTOR_RUNNER.FILE_PATTERN; src/shapepipe/modules/mccd_package. default: psfex options: From df20cf4d7c6736cf221238a3d159b51f6c93e517 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 04:10:19 +0200 Subject: [PATCH 67/85] config(cfis): final_cat.param keeps NUMBER sp_validation's extraction requires the SExtractor object number (scripts/calibration/params.py) and every tile catalogue carries it; the exact allow-list dropped it, so the merged campaign catalogue could not be extracted. Co-Authored-By: Claude Fable 5.1 --- workflow/config/cfis/final_cat.param | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/workflow/config/cfis/final_cat.param b/workflow/config/cfis/final_cat.param index 0d5653ef8..ea82f7ee3 100644 --- a/workflow/config/cfis/final_cat.param +++ b/workflow/config/cfis/final_cat.param @@ -6,6 +6,10 @@ YWIN_WORLD # Can maybe be removed. TILE_ID +# SExtractor object number within the tile: sp_validation's extraction joins +# on it, and the per-tile catalogue always carries it. +NUMBER + # flags FLAGS # NO IMAFLAGS_ISO, AND NO MASK COLUMN AT ALL — READ THIS BEFORE ADDING ONE. From 05e73818242f79aa15158a177247c0a61a535137 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 10:50:54 +0200 Subject: [PATCH 68/85] =?UTF-8?q?fix(workflow):=20the=20sims=20overlay=20r?= =?UTF-8?q?esolves=20gauss=5F3.0=5F7x7.conv;=20`sp=20--=20=E2=80=A6`=20dro?= =?UTF-8?q?ps=20its=20separator?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit config_tile_Sx.ini (shared by symlink) names $SP_CONFIG/gauss_3.0_7x7.conv, and the cfis_image_sims overlay had no such file: the DAG built and Source Extractor died on every sims tile. bin/sp kept the `--` in the arguments it forwards, so snakemake read the flags after it as target names ("No rule to produce --cores"). Co-Authored-By: Claude Fable 5.1 --- workflow/bin/sp | 2 +- workflow/config/cfis_image_sims/gauss_3.0_7x7.conv | 1 + 2 files changed, 2 insertions(+), 1 deletion(-) create mode 120000 workflow/config/cfis_image_sims/gauss_3.0_7x7.conv diff --git a/workflow/bin/sp b/workflow/bin/sp index 86aecd75b..944595172 100755 --- a/workflow/bin/sp +++ b/workflow/bin/sp @@ -73,7 +73,7 @@ _sp_args=() while [ $# -gt 0 ]; do case "$1" in --) - _sp_args+=("$@"); break ;; + shift; _sp_args+=("$@"); break ;; -c|--config-file) [ $# -ge 2 ] || { echo "sp: $1 needs a file" >&2; exit 2; } RUN_CONFIG="$2"; shift 2 ;; diff --git a/workflow/config/cfis_image_sims/gauss_3.0_7x7.conv b/workflow/config/cfis_image_sims/gauss_3.0_7x7.conv new file mode 120000 index 000000000..25e80a4b4 --- /dev/null +++ b/workflow/config/cfis_image_sims/gauss_3.0_7x7.conv @@ -0,0 +1 @@ +../cfis/gauss_3.0_7x7.conv \ No newline at end of file From 45a9f1b5c9c5223c9713a5386008fde6f6c1720b Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 10:52:42 +0200 Subject: [PATCH 69/85] fix(workflow): refuse psf_model=mccd at parse time PERSISTS_PSF includes mccd, but persist_exp.py maps psf_model to *.psf and merge_star_cat.py reads the per-CCD validation table in HDU 2; MCCD writes fitted_model-.npy and an exposure-wide table in HDU 1. A real MCCD campaign therefore ran every exposure and died at star_cat_merge. The Snakefile now refuses mccd beside PERSISTS_PSF, naming both readers, until an MCCD-aware persist table and reader exist. CONTRACTS' persist-edge prose says psfex. Test: tests/unit/test_parse_guards.py lifts the guard and checks the parse calls it. A dry run with psf_model: mccd stops with the message. Co-Authored-By: Claude Opus 5.5 --- tests/unit/test_parse_guards.py | 54 +++++++++++++++++++++++++++++++++ workflow/CONTRACTS | 5 +-- workflow/Snakefile | 23 +++++++++++++- 3 files changed, 79 insertions(+), 3 deletions(-) create mode 100644 tests/unit/test_parse_guards.py diff --git a/tests/unit/test_parse_guards.py b/tests/unit/test_parse_guards.py new file mode 100644 index 000000000..85e50c53e --- /dev/null +++ b/tests/unit/test_parse_guards.py @@ -0,0 +1,54 @@ +"""Parse-time refusals in the Snakefile for campaigns that would otherwise fail +late or destroy products. + +Each guard is a top-level ``def`` in the Snakefile, called once at parse time. +The tests lift the ``def`` by name (as ``test_campaign_lineage.py`` does for the +path helpers), exercise it against a stub ``WorkflowError``, and check that the +Snakefile calls it on the live values. +""" + +import re +from pathlib import Path + +import pytest + +REPO_ROOT = Path(__file__).resolve().parents[2] +SNAKEFILE = REPO_ROOT / "workflow" / "Snakefile" + + +class WorkflowError(Exception): + pass + + +def _lift(name): + text = SNAKEFILE.read_text() + m = re.search(rf"^def {name}\(.*?(?=^\S)", text, re.M | re.S) + assert m, f"Snakefile no longer defines {name}()" + ns = {"Path": Path, "WorkflowError": WorkflowError} + exec(m.group(0), ns) + return ns[name] + + +def _called(call): + """The Snakefile calls ``call`` at top level (not only inside a def).""" + return re.search(rf"^{re.escape(call)}\s*$", SNAKEFILE.read_text(), re.M) + + +# --- MCCD: persistence and the star-catalogue merge read PSFEx products only -- + +def test_mccd_is_refused_naming_the_psfex_only_readers(): + guard = _lift("refuse_unpersistable_psf") + with pytest.raises(WorkflowError) as exc: + guard("mccd") + msg = str(exc.value) + assert "persist_exp.py" in msg and "merge_star_cat.py" in msg + assert "PSFEx" in msg + + +@pytest.mark.parametrize("model", ["psfex", "fake"]) +def test_psfex_and_fake_pass(model): + _lift("refuse_unpersistable_psf")(model) + + +def test_psf_guard_runs_on_the_parsed_model(): + assert _called("refuse_unpersistable_psf(PSF_MODEL)") diff --git a/workflow/CONTRACTS b/workflow/CONTRACTS index 04b7ebb5b..f87728b3e 100644 --- a/workflow/CONTRACTS +++ b/workflow/CONTRACTS @@ -37,7 +37,8 @@ boundary on a fresh root. @sc [label:custody] clean-exposure-waits-on-persist-iff-psf `clean_exposure` takes the exposure's `exp_persist` manifest as input exactly -when `PERSISTS_PSF` (`psf_model != "fake"`). Dropping the edge under a real PSF -model lets reclamation delete the scratch store before its PSF products reach +when `PERSISTS_PSF` (`psf_model != "fake"`, which is `psfex`: `mccd` is refused +at parse time by `refuse_unpersistable_psf`). Dropping the edge under psfex +lets reclamation delete the scratch store before its PSF products reach `products_dir`; keeping it under `fake` makes every clean wait on a rule that is not in the DAG. diff --git a/workflow/Snakefile b/workflow/Snakefile index 53645422a..4d6adb1cf 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -107,7 +107,8 @@ container: _image # `fake` is the image-simulation true PSF: no exposure PSF fit; the tile's # galaxy_psf comes from `psf_dict` via fake_interp_runner (read by the -# configs as ${SP_PSF}_interp_runner). Sims with stars may use psfex/mccd. +# configs as ${SP_PSF}_interp_runner). Sims with stars may use psfex. mccd is +# a valid chain but refused by refuse_unpersistable_psf below. PSF_MODELS = {"psfex", "mccd", "fake"} PSF_MODEL = config.get("psf_model", "psfex") if PSF_MODEL not in PSF_MODELS: @@ -579,6 +580,26 @@ PERSIST_EXP = list(config.get("persist_exp") or []) # unaffected. psf_exposures() below is where this gate acts. PERSISTS_PSF = PSF_MODEL != "fake" + +def refuse_unpersistable_psf(psf_model): + """Refuse a PSF model whose products the persistence path cannot read. + + exp_persist packs PSFEx products (persist_exp.py maps psf_model to + *.psf) and star_cat_merge reads PSFEx's per-CCD validation table in HDU 2 + (merge_star_cat.py). MCCD writes fitted_model-.npy and one + exposure-wide table in HDU 1, so an MCCD campaign would run every exposure + and then fail at star_cat_merge. Refusing at parse time costs nothing. + """ + if psf_model == "mccd": + raise WorkflowError( + "psf_model=mccd: PSF persistence (workflow/scripts/persist_exp.py) " + "and the star-catalogue merge (workflow/scripts/merge_star_cat.py) " + "read PSFEx products only. An MCCD campaign needs an MCCD-aware " + "persist table and star-catalogue reader first (tracked in the PR).") + + +refuse_unpersistable_psf(PSF_MODEL) + # The keep list names PRODUCTS (`psf_model`), not globs (`*.psf`); the # catalogue that maps one to the other lives in persist_exp.py, which is also # what the rule runs, so there is one definition and not a copy here. From 16df7ae750b4b405db156d46fa2e1eae9ef46868 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 10:54:05 +0200 Subject: [PATCH 70/85] fix(workflow): refuse reclamation when products_dir is run_dir products_dir defaults to run_dir, and with one root the final_cat lives in the tile store clean_tile prunes and the exp_persist manifests in the exposure tree clean_exposure deletes. The shipped config.yaml turns both cleaners on, so a one-root run silently reclaimed its own products. The Snakefile now refuses a resolved-equal pair unless clean and clean_tiles are both false. It reads the raw flags rather than CLEAN/CLEAN_TILES so the prepare parse refuses too. The products_dir comments in config.yaml and the Snakefile say a one-root run is for fixtures and smoke tests with both cleaners off. Test: tests/unit/test_parse_guards.py. Dry runs of the smoke fixture without products_dir: refused with clean: true; builds 28 jobs with clean: false and clean_tiles: false. Co-Authored-By: Claude Opus 5.5 --- tests/unit/test_parse_guards.py | 39 ++++++++++++++++++++++++++++++++- workflow/Snakefile | 27 ++++++++++++++++++++++- workflow/config.yaml | 12 +++++----- 3 files changed, 71 insertions(+), 7 deletions(-) diff --git a/tests/unit/test_parse_guards.py b/tests/unit/test_parse_guards.py index 85e50c53e..ba7397a01 100644 --- a/tests/unit/test_parse_guards.py +++ b/tests/unit/test_parse_guards.py @@ -31,7 +31,7 @@ def _lift(name): def _called(call): """The Snakefile calls ``call`` at top level (not only inside a def).""" - return re.search(rf"^{re.escape(call)}\s*$", SNAKEFILE.read_text(), re.M) + return re.search(rf"^{re.escape(call)}", SNAKEFILE.read_text(), re.M) # --- MCCD: persistence and the star-catalogue merge read PSFEx products only -- @@ -52,3 +52,40 @@ def test_psfex_and_fake_pass(model): def test_psf_guard_runs_on_the_parsed_model(): assert _called("refuse_unpersistable_psf(PSF_MODEL)") + + +# --- one root: the cleaners would reclaim the products ------------------------ + +@pytest.mark.parametrize("clean,clean_tiles", + [(True, False), (False, True), (True, True)]) +def test_one_root_with_a_cleaner_is_refused(tmp_path, clean, clean_tiles): + guard = _lift("refuse_one_root_cleaners") + with pytest.raises(WorkflowError) as exc: + guard(tmp_path, tmp_path, clean, clean_tiles) + msg = str(exc.value) + assert "clean_tile" in msg and "clean_exposure" in msg + assert "clean: false" in msg and "clean_tiles: false" in msg + + +def test_one_root_is_compared_resolved(tmp_path): + (tmp_path / "run").mkdir() + (tmp_path / "alias").symlink_to(tmp_path / "run") + with pytest.raises(WorkflowError): + _lift("refuse_one_root_cleaners")( + tmp_path / "alias", tmp_path / "run", True, False) + + +def test_one_root_without_cleaners_passes(tmp_path): + _lift("refuse_one_root_cleaners")(tmp_path, tmp_path, False, False) + + +def test_two_roots_with_cleaners_pass(tmp_path): + _lift("refuse_one_root_cleaners")( + tmp_path / "products", tmp_path / "run", True, True) + + +def test_root_guard_runs_on_the_raw_flags(): + """The raw config flags, not CLEAN/CLEAN_TILES: those are off outside the + compute phase, and the prepare parse should already refuse.""" + assert _called('refuse_one_root_cleaners(PRODUCTS_DIR, RUN_DIR, ' + 'flag(config.get("clean")),') diff --git a/workflow/Snakefile b/workflow/Snakefile index 4d6adb1cf..c66b7b80c 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -147,7 +147,8 @@ INPUTS = config["inputs"] OUTPUTS = config["outputs"] RUN_DIR = Path(OUTPUTS["run_dir"]) # Defaults to RUN_DIR so a scratch-only run (a fixture, a smoke test) needs no -# second path: one root, exactly the pre-D5 layout. +# second path. Such a one-root run must set clean: false and clean_tiles: +# false; refuse_one_root_cleaners enforces it. PRODUCTS_DIR = Path(OUTPUTS.get("products_dir") or RUN_DIR) INDEX_DB = Path(OUTPUTS["index_db"]) # The campaign's NAME is the run config's `run:` — the same name `$run` expands @@ -1027,6 +1028,30 @@ def final_cat_targets(): CLEAN_TILES = flag(config.get("clean_tiles", False)) and PHASE == "compute" +def refuse_one_root_cleaners(products_dir, run_dir, clean, clean_tiles): + """Refuse reclamation when products and scratch share one root. + + With products_dir == run_dir the per-tile final_cat sits inside the tile + store clean_tile prunes, and the exp_persist manifests sit inside the + exposure tree clean_exposure deletes, so reclamation would destroy the + products it exists to protect. A one-root run is for fixtures and smoke + tests, which keep everything anyway. + """ + if Path(products_dir).resolve() != Path(run_dir).resolve(): + return + if clean or clean_tiles: + raise WorkflowError( + f"outputs.products_dir is outputs.run_dir ({run_dir}): a one-root " + "run is for fixtures and smoke tests only, and must run with " + "clean: false and clean_tiles: false, since clean_tile would " + "delete final_cat and clean_exposure the persisted manifests. Set " + "both false, or give products_dir its own root.") + + +refuse_one_root_cleaners(PRODUCTS_DIR, RUN_DIR, flag(config.get("clean")), + flag(config.get("clean_tiles"))) + + def tile_tombstone(tile): """The clean_tile output. Beside manifests/ and logs/, not inside either — same placement and same reason as the exposure tombstone() above.""" diff --git a/workflow/config.yaml b/workflow/config.yaml index 43372855e..f46b05106 100644 --- a/workflow/config.yaml +++ b/workflow/config.yaml @@ -33,7 +33,8 @@ input_type: data # - retrieve: method to get input data, allowed are symlink, vos # - inputs.tiles,, .exposures: path for input tile and exposure # - outputs.run_dir: (scratch) run directory, where tmp files will be stored -# - outputs.products_dir: path to final products +# - outputs.products_dir: path to final products (default run_dir; one root +# needs clean: false and clean_tiles: false) # - outputs.index_db: path to bookkeeping index sqlite file # - container: TBD @@ -160,10 +161,11 @@ machines: # straight to products_dir. A tile keep list is #844 follow-up. # # NOTE ON products_dir DEFAULTING TO run_dir (a fixture or smoke test): the tar -# then lands beside the store on the same filesystem and buys nothing, and the -# manifest sits in the exposure's own manifests/ dir, which clean_exposure -# deletes wholesale — so a one-root run re-persists after every reclamation. -# Harmless, and exactly the pre-D5 behaviour a one-root run asks for. +# then lands beside the store on the same filesystem and buys nothing, and +# final_cat and the persist manifests sit inside the trees clean_tile and +# clean_exposure delete. A one-root run is for fixtures and smoke tests only and +# must set clean: false and clean_tiles: false; the Snakefile refuses it +# otherwise. persist_exp: - psf_model From 96abec89669f93a3958d956a8ee96d7aed084796 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 10:55:37 +0200 Subject: [PATCH 71/85] fix(workflow): run_config expands to a fixed point and reports any leftover $ Expansion ran one pass, so a base_dir holding $run left a literal $run in every path built from it, and unresolved() checked REQUIRED keys only, so an optional path (products_dir, inputs.masks, container) with a stray $variable passed validation and failed later in a job's shell ("run: unbound variable"). Expansion now repeats until stable (at most three passes), and unresolved() also reports every MACHINE_KEYS value, recursively, that still holds a $. The Snakefile's refusal message names that case. Tests: tests/unit/test_run_config.py, a base_dir with $run and optional paths with an unknown $var. A dry run with products_dir: /tmp/$nope/product is refused at parse time. Co-Authored-By: Claude Opus 5.5 --- tests/unit/test_run_config.py | 31 +++++++++++++++++++++++++-- workflow/Snakefile | 3 ++- workflow/scripts/run_config.py | 39 ++++++++++++++++++++++++++++------ 3 files changed, 63 insertions(+), 10 deletions(-) diff --git a/tests/unit/test_run_config.py b/tests/unit/test_run_config.py index 058229097..da3281d05 100644 --- a/tests/unit/test_run_config.py +++ b/tests/unit/test_run_config.py @@ -1,10 +1,12 @@ -"""``workflow/scripts/run_config.py``: `run:` is a required key. +"""``workflow/scripts/run_config.py``: `run:` is a required key, and every +machine-key path expands fully or is reported. `run:` names the campaign's merged catalogues (``final_cat_.hdf5``, ``full_starcat_.hdf5``). ``unresolved()`` already reports a ``$run`` left unexpanded in a path; these tests pin that a run config whose paths never mention ``$run`` is refused too, and that the shipped ``config.yaml`` leaves the -name to the run config. +name to the run config. A ``$base_dir`` whose value holds ``$run`` expands +through both, and an optional path left holding a ``$`` is reported. """ import importlib.util @@ -54,3 +56,28 @@ def test_shipped_config_leaves_run_to_the_run_config(tmp_path, monkeypatch): cfg = run_config.load(CONFIG_YAML, with_run) assert run_config.unresolved(cfg) == [] assert cfg["outputs"]["run_dir"].endswith("/smk-test") + + +def _machine_config(base_dir, **over): + return {"run": "smk-g6", "machine": "m", "input_type": "data", + "machines": {"m": {"base_dir": base_dir, "data": { + "tile_list": "$base_dir/tiles.txt", + "inputs": {"tiles": "$base_dir/tiles", + "exposures": "$base_dir/exp"}, + "outputs": {"run_dir": "$base_dir/run", + "index_db": "$base_dir/index.sqlite"}}}}, + **over} + + +def test_base_dir_holding_run_expands_fully(): + cfg = run_config.apply_machine_defaults(_machine_config("/b/$run")) + assert cfg["outputs"]["run_dir"] == "/b/smk-g6/run" + assert run_config.unresolved(cfg) == [] + + +def test_optional_path_with_an_unknown_variable_is_reported(): + cfg = run_config.apply_machine_defaults(_machine_config( + "/b", outputs={"products_dir": "/p/$nope/products"}, + inputs={"masks": "/m/$nope"}, container="/c/$nope.sif")) + assert set(run_config.unresolved(cfg)) == { + "outputs.products_dir", "inputs.masks", "container"} diff --git a/workflow/Snakefile b/workflow/Snakefile index c66b7b80c..153765450 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -85,7 +85,8 @@ run_config.apply_machine_defaults(config) _unresolved = run_config.unresolved(config) if _unresolved: raise WorkflowError( - f"Unset or {run_config.PLACEHOLDER!r} for machine={MACHINE!r}, " + f"Unset, {run_config.PLACEHOLDER!r}, or holding an unexpanded " + f"$variable for machine={MACHINE!r}, " f"input_type={INPUT_TYPE!r}: {', '.join(_unresolved)}. Set them in " f"your run config (SP_RUN_CONFIG).") diff --git a/workflow/scripts/run_config.py b/workflow/scripts/run_config.py index 5d97a9975..0c779de8b 100644 --- a/workflow/scripts/run_config.py +++ b/workflow/scripts/run_config.py @@ -41,14 +41,28 @@ def merge(base, over): return out -def _expand(value, variables): - """Replace $name for each set name in `variables` (base_dir, run).""" +# A variable's value may itself hold a variable (`base_dir: /x/$run`), so +# expansion repeats until nothing changes; the bound stops a self-reference. +EXPAND_PASSES = 3 + + +def _expand_once(value, variables): if isinstance(value, str): return re.sub(r"\$(\w+)", lambda m: str(variables.get(m.group(1)) or m.group(0)), value) if isinstance(value, dict): - return {k: _expand(v, variables) for k, v in value.items()} + return {k: _expand_once(v, variables) for k, v in value.items()} + return value + + +def _expand(value, variables): + """Replace $name for each set name in `variables` (base_dir, run).""" + for _ in range(EXPAND_PASSES): + expanded = _expand_once(value, variables) + if expanded == value: + break + value = expanded return value @@ -84,11 +98,22 @@ def get(config, dotted): return value +def _dollar_keys(value, prefix): + """Dotted keys under `value` whose string still holds a `$`.""" + if isinstance(value, dict): + return [k for key, sub in value.items() + for k in _dollar_keys(sub, f"{prefix}.{key}")] + return [prefix] if isinstance(value, str) and "$" in value else [] + + def unresolved(config): - """REQUIRED keys that are unset, the placeholder, or hold an unexpanded - $variable (e.g. `$run` with no `run:` set).""" - return [k for k in REQUIRED - if get(config, k) in (None, "", PLACEHOLDER) or "$" in str(get(config, k))] + """REQUIRED keys that are unset or the placeholder, then every MACHINE_KEYS + value (recursively through inputs/outputs) that still holds an unexpanded + $variable (e.g. `$run` with no `run:` set, or a misspelt name).""" + missing = [k for k in REQUIRED + if get(config, k) in (None, "", PLACEHOLDER)] + dollar = [k for key in MACHINE_KEYS for k in _dollar_keys(config.get(key), key)] + return missing + [k for k in dollar if k not in missing] def load(config_yaml, run_config=None): From 072b705e34429d78fef58f631b67cddd2c7d056d Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 10:58:15 +0200 Subject: [PATCH 72/85] fix(persist_exp): the manifest records each member's sha256 The manifest listed members and sizes only, so a PSF refit that changed values but no sizes rewrote the tar while the manifest, the DAG edge star_cat_merge waits on, kept its bytes and mtime: the merge never reran. Each member's sha256 is now read back from the published tar into the manifest. Unchanged bytes give an unchanged manifest, so a no-op rerun stays byte-stable. The merge side needs no change: hdf5_reconcile stamps each exposure by its tar's size and mtime, and persist_exp replaces the tar exactly when its bytes differ, so once the merge runs it refreshes that exposure. Tests: test_persist_exp_props (same size, new bytes changes the manifest; a rerun does not) and test_star_cat_refresh (real persist and merge: the manifest changes and the merge reports the exposure refreshed with the new values). Co-Authored-By: Claude Opus 5.5 --- tests/unit/test_persist_exp_props.py | 17 ++++++ tests/unit/test_star_cat_refresh.py | 91 ++++++++++++++++++++++++++++ workflow/scripts/merge_star_cat.py | 4 +- workflow/scripts/persist_exp.py | 13 +++- 4 files changed, 121 insertions(+), 4 deletions(-) create mode 100644 tests/unit/test_star_cat_refresh.py diff --git a/tests/unit/test_persist_exp_props.py b/tests/unit/test_persist_exp_props.py index 4fc0bc039..97d58d72c 100644 --- a/tests/unit/test_persist_exp_props.py +++ b/tests/unit/test_persist_exp_props.py @@ -358,3 +358,20 @@ def test_overlapping_patterns_never_fail(keep): members) finally: s.close() + + +def test_same_size_new_bytes_changes_the_manifest(store): + """The manifest is the DAG edge star_cat_merge waits on, so a refit that + keeps every member's size must still change it; a rerun over the same + bytes must not.""" + store.write(ALWAYS, b"\x01" * 32) + assert store.pack([]) == 0 + before = store.manifest.read_bytes() + assert store.pack([]) == 0 + assert store.manifest.read_bytes() == before, "a no-op rerun moved it" + + store.write(ALWAYS, b"\x02" * 32) + assert store.pack([]) == 0 + assert store.manifest.read_bytes() != before + (entry,) = json.loads(store.manifest.read_text())["files"] + assert entry["sha256"] == hashlib.sha256(b"\x02" * 32).hexdigest() diff --git a/tests/unit/test_star_cat_refresh.py b/tests/unit/test_star_cat_refresh.py new file mode 100644 index 000000000..eb42cda71 --- /dev/null +++ b/tests/unit/test_star_cat_refresh.py @@ -0,0 +1,91 @@ +"""A PSF refit reaches the campaign star catalogue. + +Real ``persist_exp.py`` and ``merge_star_cat.py``, run as the rules run them. +A refit that changes a value but no size rewrites the exposure's tar; its +manifest (the edge ``star_cat_merge`` waits on) must change with it, and the +next merge must refresh that exposure's dataset. +""" + +import importlib.util +import json +import sqlite3 +import subprocess +import sys +from pathlib import Path + +import h5py +import numpy as np +from astropy.io import fits + +REPO_ROOT = Path(__file__).resolve().parents[2] +SCRIPTS = REPO_ROOT / "workflow" / "scripts" +EXP = "2605805" + + +def _load(name): + sys.path.insert(0, str(SCRIPTS)) + try: + spec = importlib.util.spec_from_file_location(f"_{name}", + SCRIPTS / f"{name}.py") + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + finally: + sys.path.remove(str(SCRIPTS)) + return module + + +merge = _load("merge_star_cat") +persist = _load("persist_exp") + + +def _run(name, *args): + run = subprocess.run([sys.executable, str(SCRIPTS / f"{name}.py"), + *map(str, args)], capture_output=True, text=True) + assert run.returncode == 0, run.stderr + return run.stdout + + +def _validation(path, value): + """A PSFEx-shaped validation table: the rows live in HDU 2.""" + cols = [fits.Column(name=c, format="I" if "FLAG" in c else "E", + array=[value, value]) for c in merge.COLUMNS] + path.parent.mkdir(parents=True, exist_ok=True) + fits.HDUList([fits.PrimaryHDU(), fits.ImageHDU(), + fits.BinTableHDU.from_columns(cols)]).writeto( + path, overwrite=True) + + +def test_refit_with_equal_sizes_refreshes_the_exposure(tmp_path): + store = tmp_path / "scratch" / EXP + src = (store / "output" / persist.RUN_NAME / "psfex_interp_runner" + / "output" / f"validation_psf-{EXP}-3.fits") + dest = tmp_path / "exp" / EXP[:2] / EXP / "psf" + manifest = dest.parent / "manifests" / "exp_persist.json" + persist_args = ("--exp-dir", store, "--exp", EXP, "--dest", dest, + "--manifest", manifest) + + tiles = tmp_path / "tiles.txt" + tiles.write_text("210.282\n") + db = tmp_path / "index.sqlite" + with sqlite3.connect(db) as con: + con.execute("CREATE TABLE tile_exposures(tile_id TEXT, exp_id TEXT)") + con.execute("INSERT INTO tile_exposures VALUES ('210.282', ?)", (EXP,)) + out = tmp_path / "stars.h5" + merge_args = ("--products-dir", tmp_path, "--tile-list", tiles, + "--index-db", db, "--output", out, "--campaign", "t") + + _validation(src, 0.25) + _run("persist_exp", *persist_args) + _run("merge_star_cat", *merge_args) + size, before = src.stat().st_size, manifest.read_bytes() + + _validation(src, 0.75) + assert src.stat().st_size == size, "fixture must keep the size" + _run("persist_exp", *persist_args) + assert manifest.read_bytes() != before, ( + "the tar changed but the manifest star_cat_merge waits on did not") + + assert "1 refreshed" in _run("merge_star_cat", *merge_args) + with h5py.File(out) as f: + rows = f[f"exposures/{EXP}"][:] + np.testing.assert_allclose(rows["X"], [0.75, 0.75]) diff --git a/workflow/scripts/merge_star_cat.py b/workflow/scripts/merge_star_cat.py index 3b0e32b38..d41055adc 100644 --- a/workflow/scripts/merge_star_cat.py +++ b/workflow/scripts/merge_star_cat.py @@ -67,8 +67,8 @@ fingerprint never saw and no rerun trigger would notice. THE MANIFEST, NOT THE TAR, IS WHAT IT READS FIRST: the manifest records what was -actually packed, member by member, with sizes and the product each came from, so -this script never guesses at tar contents. +actually packed, member by member, with sizes, sha256 digests and the product +each came from, so this script never guesses at tar contents. """ import argparse diff --git a/workflow/scripts/persist_exp.py b/workflow/scripts/persist_exp.py index 997bc4489..3e6267552 100644 --- a/workflow/scripts/persist_exp.py +++ b/workflow/scripts/persist_exp.py @@ -50,8 +50,8 @@ we think it is, and writing a green manifest over that would let ``clean_exposure`` delete an exposure whose products were never saved. -The manifest lists every member (name, pattern, source path, bytes), so a reader -knows what the tar holds without opening it. +The manifest lists every member (name, pattern, source path, bytes, sha256), +so a reader knows what the tar holds without opening it. ONE UNCOMPRESSED TAR PER EXPOSURE, ``/.tar``, NOT LOOSE COPIES. Inodes, not bytes, are what bind on /project: the group quota is ~1 M files, @@ -98,6 +98,7 @@ import argparse import filecmp +import hashlib import json import sys import tarfile @@ -417,6 +418,14 @@ def anonymous(ti: tarfile.TarInfo) -> tarfile.TarInfo: finally: tmp.unlink(missing_ok=True) + # Each member's sha256, read back from the tar as published. The manifest + # is the DAG edge star_cat_merge waits on: a refit that changes values but + # no sizes must change it, and a rerun over the same bytes must not. + with tarfile.open(tar_path) as tf: + for f in files: + f["sha256"] = hashlib.file_digest( + tf.extractfile(f["name"]), "sha256").hexdigest() + body = { "stage": "exp_persist", "level": "exp", "unit": args.exp, "status": "complete", From d179f2b5771ac2b44f8152ea6deda08da8cdde55 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 11:01:41 +0200 Subject: [PATCH 73/85] fix(hdf5_reconcile): one type per column across a campaign The schema digest covers column names only, so a tile refreshed with another type for a column (a logical mask as i1 beside neighbours' f8) was accepted, and reading the campaign promoted it silently. The digest cannot carry types: no merge knows them before reading, since they come from the sources. So apply checks them as it reads. A unit whose column types differ from the datasets it would keep means the reader changed, and every unit is re-read. Kept datasets that already disagree are re-read the same way. Sources that still disagree are refused, naming the column and both types, with the file untouched. Byte order is not a type: FITS sources are big-endian. The standalone create_final_cat.py main loop is left as it is; the workflow merge goes through hdf5_reconcile. Tests: test_final_cat_merge_invariants (real merge: same-type rewrite refreshes one tile; one retyped tile is refused naming NGMIX_T_NOSHEAR, float32 and float64) and test_hdf5_reconcile_props (a reader type change on an add re-reads every unit). Co-Authored-By: Claude Opus 5.5 --- tests/unit/test_final_cat_merge_invariants.py | 36 ++++++++++ tests/unit/test_hdf5_reconcile_props.py | 29 ++++++++ workflow/scripts/hdf5_reconcile.py | 67 +++++++++++++++++++ 3 files changed, 132 insertions(+) diff --git a/tests/unit/test_final_cat_merge_invariants.py b/tests/unit/test_final_cat_merge_invariants.py index eebef1629..29c8b0634 100644 --- a/tests/unit/test_final_cat_merge_invariants.py +++ b/tests/unit/test_final_cat_merge_invariants.py @@ -183,3 +183,39 @@ def test_a_tile_missing_a_listed_column_stops_the_merge(tmp_path): assert run.returncode != 0, "merge succeeded over a tile short a column" assert missing in run.stderr assert not output.exists() + + +def _rewrite_tile(tile_path, retype=None): + """Rewrite one tile's catalogue, optionally narrowing one column to f4.""" + cat = fits.getdata(tile_path, 1) + arr = np.array(cat) + if retype: + dtype = [(n, "f4" if n == retype else arr.dtype[n]) + for n in arr.dtype.names] + arr = arr.astype(dtype) + fits.HDUList([fits.PrimaryHDU(), fits.BinTableHDU(arr)]).writeto( + tile_path, overwrite=True) + + +def test_one_column_type_per_campaign(tmp_path): + """A tile rewritten with the same types refreshes alone (FITS byte order + is not a type change); a tile whose column changes type beside tiles that + kept the old one is refused, naming the column and both dtypes, and the + published catalogue is left as it was.""" + argv, output, _ = _campaign(tmp_path) + assert subprocess.run(argv, capture_output=True).returncode == 0 + tile = lambda t: (output.parent / "tiles" / t[:2] / t + / f"final_cat-{t}.fits") + + _rewrite_tile(tile(TILES[0])) + run = subprocess.run(argv, capture_output=True, text=True) + assert run.returncode == 0, run.stderr + assert "0 added, 1 refreshed" in run.stdout, run.stdout + + before = output.read_bytes() + _rewrite_tile(tile(TILES[-1]), retype="NGMIX_T_NOSHEAR") + run = subprocess.run(argv, capture_output=True, text=True) + assert run.returncode != 0, "a campaign with two dtypes for one column" + assert "NGMIX_T_NOSHEAR" in run.stderr + assert "float32" in run.stderr and "float64" in run.stderr, run.stderr + assert output.read_bytes() == before diff --git a/tests/unit/test_hdf5_reconcile_props.py b/tests/unit/test_hdf5_reconcile_props.py index a990c3f89..d12bcf9ee 100644 --- a/tests/unit/test_hdf5_reconcile_props.py +++ b/tests/unit/test_hdf5_reconcile_props.py @@ -334,3 +334,32 @@ def test_repeated_refresh_does_not_grow_the_file(): assert len(set(sizes)) == 1, f"file size drifted across refreshes: {sizes}" finally: shutil.rmtree(work, ignore_errors=True) + + +def test_a_type_change_in_the_reader_refreshes_every_unit(tmp_path): + """The column list is unchanged, so the digest is too; a reader that now + yields another dtype must still leave one dtype in the file, by re-reading + the units it would otherwise keep.""" + columns = COLUMN_SETS[0] + sources = {} + for i, unit in enumerate(UNITS[:2]): + sources[unit] = tmp_path / f"{unit}.npy" + _write_source(sources[unit], _array(columns, 3, i), 10**18 + i) + _build(tmp_path / "cat.h5", sources, columns) + + sources["u2"] = tmp_path / "u2.npy" + _write_source(sources["u2"], _array(columns, 3, 2), 10**18 + 2) + units = sorted(sources.items()) + digest = reconcile.schema_digest(columns) + todo = reconcile.plan(tmp_path / "cat.h5", GROUP, units, digest) + assert (todo.add, todo.refresh) == (["u2"], []) + + narrow = lambda unit, source: _read(unit, source).astype( + [(c, " tuple: return st.st_size, st.st_mtime_ns +def column_types(dtype) -> dict: + """``{column: (kind, itemsize, shape)}``: a dataset's schema, as compared + across units. Byte order is left out: FITS sources are big-endian, hdf5 + may hand them back either way, and neither changes a value.""" + return {n: (dtype[n].base.kind, dtype[n].base.itemsize, dtype[n].shape) + for n in dtype.names} + + +def type_conflict(a_unit, a_dtype, b_unit, b_dtype) -> str: + """Name the first column whose type differs between two units' dtypes.""" + a, b = column_types(a_dtype), column_types(b_dtype) + for col in a: + if a[col] != b.get(col): + got = b_dtype[col].base.name if col in b else "absent" + return (f"column {col} is {a_dtype[col].base.name} in {a_unit} " + f"but {got} in {b_unit}") + return f"{b_unit} carries columns {a_unit} does not" + + +class _Retyped(Exception): + """A read unit's types differ from the datasets `apply` would keep.""" + + class Plan: """What reconciling requires: three unit lists. @@ -185,6 +215,25 @@ def check_sole_group(output: Path, group_path: str) -> None: def apply(output: Path, group_path: str, todo: Plan, units: list, read, digest: str, count_attr: str, provenance: dict | None = None) -> None: + """Carry the plan out; re-read every unit if the column types changed. + + See ``_apply``. When a unit read under the plan has other column types + than the datasets the plan would keep, the plan widens to refresh every + kept unit, so the file keeps one type per column (module docstring). + """ + try: + _apply(output, group_path, todo, units, read, digest, count_attr, + provenance) + except _Retyped as exc: + every = Plan(todo.add, [u for u, _ in units if u not in todo.add], + todo.remove) + print(f"[hdf5_reconcile] {exc}; re-reading all {len(units)} unit(s)") + _apply(output, group_path, every, units, read, digest, count_attr, + provenance) + + +def _apply(output: Path, group_path: str, todo: Plan, units: list, read, + digest: str, count_attr: str, provenance: dict | None) -> None: """Carry the plan out on a tmp file, then move it into place. ``read(unit, source)`` returns the structured array for one unit; it is @@ -241,9 +290,27 @@ def apply(output: Path, group_path: str, todo: Plan, units: list, read, # exist, and the difference only shows when a plan both # rewrites and keeps something. src.copy(f"{group_path}/{unit}", group, name=unit) + # The kept datasets' types are the reference a read unit must + # match; with nothing kept, the first unit read is. Kept datasets + # that already disagree (a file an older merge left) re-read too. + kept = [(u, group[u].dtype) for u in keep if u in group] + ref = kept[0] if kept else None + for unit, dtype in kept[1:]: + if column_types(dtype) != column_types(ref[1]): + raise _Retyped(type_conflict(*ref, unit, dtype)) for unit in todo.add + todo.refresh: source = sources[unit] data = read(unit, source) + if ref is None: + ref = (unit, data.dtype) + elif column_types(data.dtype) != column_types(ref[1]): + conflict = type_conflict(*ref, unit, data.dtype) + if keep: + raise _Retyped(conflict) + sys.exit(f"hdf5_reconcile: {conflict}. One catalogue " + f"holds one type per column; remake the units " + f"whose sources are stale. {output} is " + f"untouched.") dset = group.create_dataset(unit, data=data, dtype=data.dtype) # The dataset's own record of what it was read from; this is # what lets a later invocation leave it alone. From 187968e2a83f1ef0751ca9408cd0f4b3a1310f72 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 11:02:18 +0200 Subject: [PATCH 74/85] test(hdf5_reconcile): isolate schema invalidation from source stamps ReconcileMachine.change_columns rewrites every source as it flips the column set, so the stamps alone refresh every unit and disabling schema invalidation survives the state machine. The new test holds every stamp fixed and changes only the requested columns, so the digest is the only thing that can refresh a unit. With the digest comparison in plan() forced off, this test fails and the machine still passes. Co-Authored-By: Claude Opus 5.5 --- tests/unit/test_hdf5_reconcile_props.py | 33 +++++++++++++++++++++++++ 1 file changed, 33 insertions(+) diff --git a/tests/unit/test_hdf5_reconcile_props.py b/tests/unit/test_hdf5_reconcile_props.py index d12bcf9ee..a9c0226dd 100644 --- a/tests/unit/test_hdf5_reconcile_props.py +++ b/tests/unit/test_hdf5_reconcile_props.py @@ -363,3 +363,36 @@ def test_a_type_change_in_the_reader_refreshes_every_unit(tmp_path): for unit, path in sources.items(): np.testing.assert_array_equal(f[GROUP][unit][...], narrow(unit, path)) + + +def test_a_schema_only_change_refreshes_every_unit(tmp_path): + """The machine's change_columns also rewrites every source, so its stamps + alone would refresh everything. Here no stamp moves: the sources hold the + wider column set throughout and only the requested columns change, so + schema invalidation is the only thing that can refresh a unit.""" + wide = COLUMN_SETS[1] + sources = {} + for i, unit in enumerate(UNITS[:3]): + sources[unit] = tmp_path / f"{unit}.npy" + _write_source(sources[unit], _array(wide, 3, i), 10**18 + i) + units = sorted(sources.items()) + output = tmp_path / "cat.h5" + + def build(columns): + read = lambda unit, source: np.ascontiguousarray( + _read(unit, source)[list(columns)]).astype( + [(c, " Date: Sat, 26 Sep 2026 11:03:54 +0200 Subject: [PATCH 75/85] fix(build_index): readiness is the tiles row, not the retained edges The index keeps a tile's tile_exposures edges after its exposure list goes missing, because clean_exposure's consumer sets must still see it. Readiness was read off those same edges, so a tile named in missing.json stayed ready for the merges (campaign_tiles) and for the Snakefile's TILES_READY. A build now deletes a missing tile's tiles row and keeps its edges. ready_tiles() reads the tiles table alone, which is cheap enough for every job's parse, and both campaign_tiles and TILES_READY go through it. The consumer sets still read the edges. Test: tests/unit/test_build_index_readiness.py (the probe's missing tile is not ready, its edge survives, and it returns when its list does). Co-Authored-By: Claude Opus 5.5 --- tests/unit/test_build_index_readiness.py | 71 +++++++++++++++++++ tests/unit/test_final_cat_merge_invariants.py | 4 ++ tests/unit/test_star_cat_refresh.py | 3 + workflow/Snakefile | 11 ++- workflow/scripts/build_index.py | 30 ++++++-- 5 files changed, 110 insertions(+), 9 deletions(-) create mode 100644 tests/unit/test_build_index_readiness.py diff --git a/tests/unit/test_build_index_readiness.py b/tests/unit/test_build_index_readiness.py new file mode 100644 index 000000000..9757f8872 --- /dev/null +++ b/tests/unit/test_build_index_readiness.py @@ -0,0 +1,71 @@ +"""``build_index``: a tile is ready only while its exposure list exists. + +The index keeps a tile's edges after its exposure list disappears, because +clean_exposure's consumer sets must still see the tile that read an exposure. +Readiness must not ride on those edges: a tile named in ``missing.json`` is not +ready, for the merges (``campaign_tiles``) or for the Snakefile's +``TILES_READY``, and comes back when its list does. +""" + +import importlib.util +import json +import re +import sqlite3 +import sys +from pathlib import Path + +REPO_ROOT = Path(__file__).resolve().parents[2] +SCRIPTS = REPO_ROOT / "workflow" / "scripts" +TILE, OTHER = "210.282", "211.282" + + +def _load(): + sys.path.insert(0, str(SCRIPTS)) + try: + spec = importlib.util.spec_from_file_location( + "_build_index", SCRIPTS / "build_index.py") + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + finally: + sys.path.remove(str(SCRIPTS)) + return module + + +bi = _load() + + +def _list(run, tile, names): + path = bi.exp_list_path(run, tile) + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text("".join(f"{n}\n" for n in names)) + return path + + +def test_a_missing_tile_is_not_ready_but_keeps_its_edges(tmp_path): + run, db = tmp_path / "run", tmp_path / "index.sqlite" + listed = _list(run, TILE, ["2605805p"]) + _list(run, OTHER, ["2605806p"]) + tiles = tmp_path / "tiles.txt" + tiles.write_text(f"{TILE}\n{OTHER}\n") + bi.build([TILE, OTHER], run, db) + assert bi.campaign_tiles(tiles, db) == [TILE, OTHER] + + listed.unlink() + bi.build([TILE, OTHER], run, db, missing_threshold=1.0) + assert json.loads((db.parent / "missing.json").read_text()) == [TILE] + assert bi.campaign_tiles(tiles, db) == [OTHER] + assert bi.campaign_exposures(tiles, db) == ["2605806"] + with sqlite3.connect(db) as con: + edges = con.execute("SELECT exp_id FROM tile_exposures " + "WHERE tile_id = ?", (TILE,)).fetchall() + assert edges == [("2605805",)], "the cleanup consumer edge was dropped" + + _list(run, TILE, ["2605805p"]) + bi.build([TILE, OTHER], run, db) + assert bi.campaign_tiles(tiles, db) == [TILE, OTHER] + + +def test_the_snakefile_reads_readiness_from_build_index(): + snakefile = (REPO_ROOT / "workflow" / "Snakefile").read_text() + line = re.search(r"^TILES_READY = .*$", snakefile, re.M).group(0) + assert "build_index.ready_tiles(" in snakefile and "_READY_INDEXED" in line diff --git a/tests/unit/test_final_cat_merge_invariants.py b/tests/unit/test_final_cat_merge_invariants.py index 29c8b0634..7dcfc6183 100644 --- a/tests/unit/test_final_cat_merge_invariants.py +++ b/tests/unit/test_final_cat_merge_invariants.py @@ -114,7 +114,11 @@ def _campaign(root: Path, drop=None): tile_list.write_text("\n".join(TILES) + "\n") index = root / "index.sqlite" con = sqlite3.connect(index) + con.execute("CREATE TABLE tiles(tile_id TEXT PRIMARY KEY, ra_dir TEXT, " + "n_exp INTEGER)") con.execute("CREATE TABLE tile_exposures(tile_id TEXT, exp_id TEXT)") + con.executemany("INSERT INTO tiles VALUES (?, ?, 1)", + [(t, t.split(".")[0]) for t in TILES]) con.executemany("INSERT INTO tile_exposures VALUES (?, ?)", [(t, "2605805") for t in TILES]) con.commit() diff --git a/tests/unit/test_star_cat_refresh.py b/tests/unit/test_star_cat_refresh.py index eb42cda71..2e5e02461 100644 --- a/tests/unit/test_star_cat_refresh.py +++ b/tests/unit/test_star_cat_refresh.py @@ -68,7 +68,10 @@ def test_refit_with_equal_sizes_refreshes_the_exposure(tmp_path): tiles.write_text("210.282\n") db = tmp_path / "index.sqlite" with sqlite3.connect(db) as con: + con.execute("CREATE TABLE tiles(tile_id TEXT PRIMARY KEY, " + "ra_dir TEXT, n_exp INTEGER)") con.execute("CREATE TABLE tile_exposures(tile_id TEXT, exp_id TEXT)") + con.execute("INSERT INTO tiles VALUES ('210.282', '210', 1)") con.execute("INSERT INTO tile_exposures VALUES ('210.282', ?)", (EXP,)) out = tmp_path / "stars.h5" merge_args = ("--products-dir", tmp_path, "--tile-list", tiles, diff --git a/workflow/Snakefile b/workflow/Snakefile index 153765450..913b84543 100644 --- a/workflow/Snakefile +++ b/workflow/Snakefile @@ -311,9 +311,14 @@ else: "SELECT tile_id FROM tile_exposures WHERE exp_id = ? ORDER BY rowid", (exp,))] -# Tiles this run can actually compute: declared AND indexed. The index spans the -# campaign, so it is intersected with the declared list, not used as it. -TILES_READY = [t for t in TILES if tile_exposures(t)] +# Tiles this run can actually compute: declared AND ready in the index +# (build_index.ready_tiles, as the merges read it). The index spans the +# campaign, so it is intersected with the declared list, not used as it. A tile +# whose exposure list went missing keeps its edges for clean_consumers but is +# not ready. +_READY_INDEXED = (build_index.ready_tiles(INDEX_DB) if INDEX_DB.exists() + else set()) +TILES_READY = [t for t in TILES if t in _READY_INDEXED] READY_SET = set(TILES_READY) # A compute invocation with nothing to compute is never a success. Without this, diff --git a/workflow/scripts/build_index.py b/workflow/scripts/build_index.py index 24290b561..ea83e9048 100644 --- a/workflow/scripts/build_index.py +++ b/workflow/scripts/build_index.py @@ -125,6 +125,10 @@ def build(tile_ids: list[str], run_dir: Path, db_path: Path, all_exposures: set[tuple[str, str]] = set() for tile_id in tile_ids: if tile_id in missing_set: + # Not ready (ready_tiles reads the tiles table), but its edges stay: + # clean_exposure's consumer sets must still see a tile that read an + # exposure. + con.execute("DELETE FROM tiles WHERE tile_id = ?", (tile_id,)) continue ra_dir = tile_id.split(".")[0] exp_pairs = read_exposure_list(exp_list_path(run_dir, tile_id)) @@ -163,11 +167,27 @@ def build(tile_ids: list[str], run_dir: Path, db_path: Path, # than two hand-written queries that could drift apart. +def ready_tiles(db_path: Path) -> set[str]: + """Tiles whose exposure list the last build over them found non-empty. + + Readiness is the ``tiles`` row, which a build removes when the list goes + missing; the edges in ``tile_exposures`` outlive it as cleanup consumers, + so they alone do not make a tile ready. ``n_exp`` is the number of edges + the build wrote beside the row. Only the small ``tiles`` table is read, + since every job's parse of the Snakefile calls this. + """ + con = sqlite3.connect(db_path, timeout=60) + ready = {r[0] for r in con.execute( + "SELECT tile_id FROM tiles WHERE n_exp > 0")} + con.close() + return ready + + def campaign_tiles(tile_list: Path, db_path: Path) -> list[str]: - """The campaign's ready tiles: declared in the list AND indexed. + """The campaign's ready tiles: declared in the list AND ready_tiles(). Exactly the Snakefile's TILES_READY, computed the same way from the same two - files — a declared tile with no indexed exposure list cannot have been + files — a declared tile with no current exposure list cannot have been computed, so it has no catalogue to merge. """ # DEDUPED, order preserved. The tile list is appended to by hand across a @@ -181,10 +201,8 @@ def campaign_tiles(tile_list: Path, db_path: Path) -> list[str]: if tile and tile not in seen: seen.add(tile) declared.append(tile) - con = sqlite3.connect(db_path, timeout=60) - indexed = {r[0] for r in con.execute("SELECT DISTINCT tile_id FROM tile_exposures")} - con.close() - return [t for t in declared if t in indexed] + ready = ready_tiles(db_path) + return [t for t in declared if t in ready] def campaign_exposures(tile_list: Path, db_path: Path) -> list[str]: From 5a61aafcac406237f00cd10f3e266fb1b1b6f241 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 11:05:27 +0200 Subject: [PATCH 76/85] fix(hdf5_reconcile): delete an abandoned tmp before the free-space check A rewrite killed between writing its tmp and renaming it leaves the tmp beside the catalogue. The next run deleted it only after check_free_space, so the orphan's bytes counted against the margin: 2.5x the file free became 1.5x, below the 2.1x check, and the retry refused itself. The single writer per output owns that tmp, so it is now removed first. Test: test_hdf5_reconcile_props, with injected disk accounting as in the review probe. Co-Authored-By: Claude Opus 5.5 --- tests/unit/test_hdf5_reconcile_props.py | 33 +++++++++++++++++++++++++ workflow/scripts/hdf5_reconcile.py | 11 ++++++--- 2 files changed, 40 insertions(+), 4 deletions(-) diff --git a/tests/unit/test_hdf5_reconcile_props.py b/tests/unit/test_hdf5_reconcile_props.py index a9c0226dd..20200c444 100644 --- a/tests/unit/test_hdf5_reconcile_props.py +++ b/tests/unit/test_hdf5_reconcile_props.py @@ -396,3 +396,36 @@ def build(columns): assert (todo.add, sorted(todo.refresh)) == ([], sorted(sources)) with h5py.File(output, "r") as f: assert {f[GROUP][u].dtype.names for u in f[GROUP]} == {wide} + + +def test_an_abandoned_tmp_does_not_count_against_free_space(tmp_path, + monkeypatch): + """A killed rewrite leaves its tmp copy behind. The next run owns it (one + writer per output), so the copy is deleted before free space is judged: + 2.5x the file free counting the orphan clears the 2.1x margin; 1.5x, as + if the orphan still stood, would not.""" + import shutil + + columns = COLUMN_SETS[0] + sources = {"u0": tmp_path / "u0.npy"} + _write_source(sources["u0"], _array(columns, 3, 0), 10**18) + output = tmp_path / "cat.h5" + _build(output, sources, columns) + size = output.stat().st_size + orphan = output.with_name(output.name + ".tmp") + shutil.copy2(output, orphan) + + usage = shutil.disk_usage(tmp_path) + + def disk_usage(_): + free = int(2.5 * size) - (orphan.stat().st_size + if orphan.exists() else 0) + return type(usage)(10 * size, 10 * size - free, free) + + monkeypatch.setattr(reconcile.shutil, "disk_usage", disk_usage) + sources["u1"] = tmp_path / "u1.npy" + _write_source(sources["u1"], _array(columns, 3, 1), 10**18 + 1) + assert _build(output, sources, columns).add == ["u1"] + assert not orphan.exists() + with h5py.File(output, "r") as f: + assert set(f[GROUP]) == {"u0", "u1"} diff --git a/workflow/scripts/hdf5_reconcile.py b/workflow/scripts/hdf5_reconcile.py index 787fc25b8..444f94258 100644 --- a/workflow/scripts/hdf5_reconcile.py +++ b/workflow/scripts/hdf5_reconcile.py @@ -266,18 +266,21 @@ def _apply(output: Path, group_path: str, todo: Plan, units: list, read, Either way the tmp is moved into place at the end, so a crash mid-merge leaves the old catalogue intact rather than a half-written one. A SIGKILL between writing the tmp and renaming it leaves the tmp behind — one file, - beside the catalogue, overwritten by the next run; the rename itself is - atomic, which is the property that matters. + beside the catalogue, deleted by the next run before its space check; the + rename itself is atomic, which is the property that matters. """ sources = dict(units) rewrite = bool(todo.remove or todo.refresh) + tmp = output.with_name(output.name + ".tmp") + # A tmp left by a killed run is this run's to delete (one writer per + # output), and deleting it first keeps it from counting against the space + # this run needs. + tmp.unlink(missing_ok=True) check_free_space(output) check_sole_group(output, group_path) written = set(todo.add) | set(todo.refresh) keep = [u for u, _ in units if u not in written] - tmp = output.with_name(output.name + ".tmp") try: - tmp.unlink(missing_ok=True) if output.exists() and not rewrite: shutil.copy2(output, tmp) with h5py.File(tmp, "a") as f: From 47b2ec52952f0c1a0375ab1768b39f8aaeb509ff Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 11:07:59 +0200 Subject: [PATCH 77/85] fix(hdf5_reconcile): each provenance record replaces the last An add-only merge copies the existing file, attributes included, and then stamped only the provenance keys the new record carried. A merge without a snapshot wrote code_head=unknown beside the previous snapshot's branch, dirty flag, dirty files and time. Every code_* attribute is now cleared before the new record is stamped. Test: test_hdf5_reconcile_props (a full snapshot, then an add with head=unknown, leaves code_head alone). Co-Authored-By: Claude Opus 5.5 --- tests/unit/test_hdf5_reconcile_props.py | 27 +++++++++++++++++++++++++ workflow/scripts/hdf5_reconcile.py | 5 +++++ 2 files changed, 32 insertions(+) diff --git a/tests/unit/test_hdf5_reconcile_props.py b/tests/unit/test_hdf5_reconcile_props.py index 20200c444..8ea6c9949 100644 --- a/tests/unit/test_hdf5_reconcile_props.py +++ b/tests/unit/test_hdf5_reconcile_props.py @@ -429,3 +429,30 @@ def disk_usage(_): assert not orphan.exists() with h5py.File(output, "r") as f: assert set(f[GROUP]) == {"u0", "u1"} + + +def test_each_provenance_record_replaces_the_last(tmp_path): + """An add-only merge copies the file, attributes and all. A merge run + without a snapshot records code_head=unknown, and must not keep the last + snapshot's branch, dirty flag, dirty files or time beside it.""" + columns = COLUMN_SETS[0] + output = tmp_path / "cat.h5" + digest = reconcile.schema_digest(columns) + sources = {} + + def add(unit, provenance): + sources[unit] = tmp_path / f"{unit}.npy" + _write_source(sources[unit], _array(columns, 3, len(sources)), + 10**18 + len(sources)) + units = sorted(sources.items()) + todo = reconcile.plan(output, GROUP, units, digest) + assert todo.add == [unit] and not (todo.refresh or todo.remove) + reconcile.apply(output, GROUP, todo, units, _read, digest, + COUNT_ATTR, provenance) + + add("u0", {"head": "abc123", "branch": "old", "dirty": True, + "taken_at": "yesterday", "dirty_files": ["old.py"]}) + add("u1", {"head": "unknown"}) + with h5py.File(output, "r") as f: + code = {k: f.attrs[k] for k in f.attrs if k.startswith("code_")} + assert code == {"code_head": "unknown"} diff --git a/workflow/scripts/hdf5_reconcile.py b/workflow/scripts/hdf5_reconcile.py index 444f94258..7bf51c6ae 100644 --- a/workflow/scripts/hdf5_reconcile.py +++ b/workflow/scripts/hdf5_reconcile.py @@ -322,6 +322,11 @@ def _apply(output: Path, group_path: str, todo: Plan, units: list, read, f.attrs[count_attr] = len(group) f.attrs["param_digest"] = digest if provenance: + # One record at a time: an add-only merge copied the last + # one's attributes, and a snapshot-less record must not keep + # its branch, dirty flag, dirty files or time. + for attr in [a for a in f.attrs if a.startswith("code_")]: + del f.attrs[attr] f.attrs["code_head"] = provenance.get("head", "unknown") for key, attr in (("branch", "code_branch"), ("dirty", "code_dirty"), From b15d442790ea6d5dedf8e23fb8f1a23d8abdef55 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 11:10:52 +0200 Subject: [PATCH 78/85] fix(workflow): a reclaimed exposure's clean waits on its tar, not the manifest clean_exposure named the exp_persist manifest for every exposure. For a reclaimed one, a persist_exp: edit changes exp_persist's params, so the manifest reruns, and it sits behind exp_psf's manifest, which went with the store: snakemake scheduled exp_get_images, exp_split, exp_psf, exp_persist and clean_exposure again for every reclaimed exposure. The edge now takes the leaf treatment the defect edge uses on #887: a live store is asked for the manifest, a reclaimed one for its tar if it exists (no rule's output, so a leaf), else nothing. One lambda per edge is kept. CONTRACTS' persist edge says so. Test: the review's retention dry run (a completed, reclaimed two-exposure fixture with recorded metadata, then persist_exp: [psf_model, star_stats]) scheduled 13 jobs including the full exposure chain for both exposures; it now schedules final_cat_merge, star_cat_merge and all (3). Harness: /automnt/n17data/cdaley/scratch/smoke-879-mine/retention/test_retention.py. Co-Authored-By: Claude Opus 5.5 --- workflow/CONTRACTS | 14 ++++++++------ workflow/rules/exposure.smk | 26 ++++++++++++++++++-------- 2 files changed, 26 insertions(+), 14 deletions(-) diff --git a/workflow/CONTRACTS b/workflow/CONTRACTS index f87728b3e..c72c2e5d3 100644 --- a/workflow/CONTRACTS +++ b/workflow/CONTRACTS @@ -36,9 +36,11 @@ tile group's cleanup deletes those finished tiles' `final_cat`. Change boundary on a fresh root. @sc [label:custody] clean-exposure-waits-on-persist-iff-psf -`clean_exposure` takes the exposure's `exp_persist` manifest as input exactly -when `PERSISTS_PSF` (`psf_model != "fake"`, which is `psfex`: `mccd` is refused -at parse time by `refuse_unpersistable_psf`). Dropping the edge under psfex -lets reclamation delete the scratch store before its PSF products reach -`products_dir`; keeping it under `fake` makes every clean wait on a rule that -is not in the DAG. +`clean_exposure` waits on the exposure's persisted PSF products exactly when +`PERSISTS_PSF` (`psf_model != "fake"`, which is `psfex`: `mccd` is refused at +parse time by `refuse_unpersistable_psf`): on the `exp_persist` manifest for a +live store, and on the tar (a leaf), or nothing, for a reclaimed one. Dropping +the edge under psfex lets reclamation delete the scratch store before its PSF +products reach `products_dir`; keeping it under `fake` makes every clean wait on +a rule that is not in the DAG; naming the manifest for a reclaimed store lets a +`persist_exp:` edit rebuild the exposure from VOS. diff --git a/workflow/rules/exposure.smk b/workflow/rules/exposure.smk index 6c3a38713..b66f9811d 100644 --- a/workflow/rules/exposure.smk +++ b/workflow/rules/exposure.smk @@ -208,15 +208,25 @@ rule clean_exposure: # this DAG, so the clean must be ordered after them. lambda wc: [tile_manifest(t, "tile_vignets") for t in clean_consumers(wc.exp) if t in READY_SET], - # The keepers must be off /scratch before the store goes. Unlike the - # consumer edges above, this edge does not depend on scope: it is the - # same exposure's own rule, so it drags nothing into the DAG that this - # exposure's chain did not already put there. No keep list removes it: - # exp_persist always packs the star catalogue's inputs. Only - # psf_model=fake does, which has no PSF products to keep + # The keepers must be off /scratch before the store goes. No keep list + # removes this edge: exp_persist always packs the star catalogue's + # inputs. Only psf_model=fake does, which has no PSF products to keep # (PERSISTS_PSF, Snakefile). - lambda wc: ([prod_exp_manifest(wc.exp, "exp_persist")] - if PERSISTS_PSF else []) + # + # A LIVE exposure is asked for its exp_persist manifest: the thing to + # build, and what orders this rule after the pack. A RECLAIMED one + # (exp_store_reclaimed, Snakefile) is asked for its TAR if it has one — + # on the persistent root, no rule's declared output, hence a leaf that + # requires nothing — and for nothing if it has none. Naming the + # manifest there reopens the reclaimed chain: a `persist_exp:` edit + # changes exp_persist's params, the manifest reruns, and it sits behind + # exp_psf's manifest, which went with the store, so snakemake rebuilds + # the exposure from VOS. + lambda wc: ([] if not PERSISTS_PSF + else [prod_exp_manifest(wc.exp, "exp_persist")] + if not exp_store_reclaimed(wc.exp) + else [prod_exp_tar(wc.exp)] + if Path(prod_exp_tar(wc.exp)).exists() else []) output: tombstone = f"{EXP_DIR}/cleaned.json" params: From 5b56a681188dca81d984120c96c50b1e000f0398 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 11:14:28 +0200 Subject: [PATCH 79/85] test(conftest): the candide hostname test matches the whole bare name The prefix match took nibi's login node (c6.nibi.sharcnet) for candide, so the two candide-data cluster tests errored there. Co-Authored-By: Claude Fable 5.1 --- conftest.py | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/conftest.py b/conftest.py index 5889c46f9..e3936b30b 100644 --- a/conftest.py +++ b/conftest.py @@ -36,7 +36,7 @@ # host this suite is most often driven from is ``c03``. We match the candide # node-name families rather than a fixed list so new nodes are covered, and # allow an explicit override for CI or odd hostnames. -_CANDIDE_HOST_RE = re.compile(r"^(c\d|n\d{2})", re.IGNORECASE) +_CANDIDE_HOST_RE = re.compile(r"^(c\d{2}|n\d{2})$", re.IGNORECASE) def on_candide(): @@ -44,7 +44,8 @@ def on_candide(): The check is, in order: an explicit ``SHAPEPIPE_ON_CANDIDE`` override (``1``/``0``), then the hostname against the candide node-name families - (``c0x`` login, ``nXX`` compute). Cheap, import-safe, no cluster calls. + (``c0x`` login, ``nXX`` compute; whole bare hostname, so ``c6.nibi.sharcnet`` + does not match). Cheap, import-safe, no cluster calls. """ override = os.environ.get("SHAPEPIPE_ON_CANDIDE") if override is not None: From 33391a4cd3d71d224582394f9910bb3a464d5d88 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 03:59:48 +0200 Subject: [PATCH 80/85] test(workflow): add an isolated Snakemake DAG driver Resolve disposable campaigns through SnakemakeApi so workflow contracts can inspect actual jobs rather than reconstructed source snippets. Seed campaign history with build_index and let the compute parse index two ready tiles, including shared and out-of-scope exposure consumers. Keep state, image-resolution sentinels, and source caches under tmp_path; restore imports, environment, cwd, and Snakemake globals between parses. The existing tests discovery root includes this tier. Co-Authored-By: GPT-6 Astra --- tests/README.md | 1 + tests/workflow/__init__.py | 1 + tests/workflow/conftest.py | 24 ++++ tests/workflow/harness.py | 266 +++++++++++++++++++++++++++++++++++++ 4 files changed, 292 insertions(+) create mode 100644 tests/workflow/__init__.py create mode 100644 tests/workflow/conftest.py create mode 100644 tests/workflow/harness.py diff --git a/tests/README.md b/tests/README.md index 0b0df1524..fed39f8bb 100644 --- a/tests/README.md +++ b/tests/README.md @@ -11,6 +11,7 @@ is driven by `pytest` from the repo root (in the dev container — see the proje |----------|-------|----------| | `tests/module/` | **module-unit tests** — the fitter, file handler, split-exp, vignetmaker, ngmix internals, the GalSim weight-validation suite | per-module unit/property/integration tests; import package internals directly. (Relocated from `src/shapepipe/tests/` so the suite has one home.) | | `tests/unit/` | **structural tests** — every submodule imports, configs parse, shell scripts lint, runner metadata is well-formed, console entry points respond to `-h` | suite-level checks on the *tree*, not any one module | +| `tests/workflow/` | **Snakemake DAG checks** — isolated campaigns resolved through the Python API, without executing jobs | checks per-job dependencies, input modes, and campaign product paths | | `tests/science/` | **fast scientific guardrails** — controlled simulations with a known answer, runnable in the inner loop with nothing from the cluster | scientific correctness that must stay green on every commit | | `tests/cluster/` | **candide guardrails** — read real on-disk catalogs / submit cluster jobs | need the cluster + real data; marked and auto-skipped off it | | `tests/helpers/` | shared, non-test library code (cluster submission, artifact emission, the star-response R-function) | imported by tests as `tests.helpers.*`; not collected as tests | diff --git a/tests/workflow/__init__.py b/tests/workflow/__init__.py new file mode 100644 index 000000000..5ebccbe7c --- /dev/null +++ b/tests/workflow/__init__.py @@ -0,0 +1 @@ +"""Planning-only checks of the ShapePipe Snakemake workflow.""" diff --git a/tests/workflow/conftest.py b/tests/workflow/conftest.py new file mode 100644 index 000000000..32591b851 --- /dev/null +++ b/tests/workflow/conftest.py @@ -0,0 +1,24 @@ +"""Isolated campaign fixtures for planning-only Snakemake tests.""" + +import pytest + +from tests.workflow.harness import MODES, Campaign, resolve + + +@pytest.fixture(params=MODES, ids=[f"{mode}+{psf}" for mode, psf in MODES]) +def campaign(request, tmp_path): + """Create a disposable campaign for each supported input/PSF pair.""" + return Campaign(tmp_path / "campaign", *request.param) + + +@pytest.fixture +def resolve_dag(monkeypatch): + """Expose the resolver so tests can also assert parse-time failures.""" + return lambda campaign: resolve(campaign, monkeypatch) + + +@pytest.fixture +def dag(campaign, resolve_dag): + """Keep the API and its jobs alive for the duration of one check.""" + with resolve_dag(campaign) as resolved: + yield resolved diff --git a/tests/workflow/harness.py b/tests/workflow/harness.py new file mode 100644 index 000000000..6b86c7543 --- /dev/null +++ b/tests/workflow/harness.py @@ -0,0 +1,266 @@ +"""Build isolated campaigns and inspect Snakemake's resolved jobs. + +No executor runs: the API's ``printdag`` operation resolves input functions +and job wildcards, including the compute parse's real SQLite index build. +Only the final access to ``WorkflowApi._workflow`` is private; Snakemake's +public API exposes graph printing but not the resolved Job objects. +""" + +import io +import os +import sys +from contextlib import contextmanager, redirect_stdout +from dataclasses import dataclass +from pathlib import Path + +import yaml +from snakemake.api import SnakemakeApi +from snakemake.settings.enums import RerunTrigger +from snakemake.settings.types import ( + DAGSettings, + DeploymentSettings, + ResourceSettings, + WorkflowSettings, +) +from snakemake_interface_executor_plugins.settings import DeploymentMethod +from workflow.scripts import build_index + +REPO = Path(__file__).resolve().parents[2] +MODES = [("data", "psfex"), ("data", "mccd"), ("image_sims", "fake")] + + +@dataclass +class Campaign: + """A campaign with two ready tiles and campaign-wide consumer edges.""" + + root: Path + input_type: str + psf_model: str + name: str = "dag-campaign" + + def __post_init__(self): + """Write exposure lists, prepared manifests, and campaign history.""" + self.root.mkdir(parents=True, exist_ok=True) + self.ready = { + "123.456": ("2243881", "2243882"), + "124.456": ("2243882",), + } + self.exposures = ("2243881", "2243882") + self.unready = "125.456" + self.outside = "126.456" + self.ignored = "127.456" + self.run_dir = self.root / "scratch" / self.name + # The basename deliberately differs from run: to catch name sniffing. + self.products_dir = self.root / "persistent" / self.name / "products" + self.index_db = self.products_dir / "index" / "run_index.sqlite" + self.tile_list = self.root / "tiles.txt" + self.config_path = self.root / "run.yaml" + self.state_dir = self.root / "state" + self.image = self.root / "planning-only.sif" + # Image resolution checks existence; DAG inspection never opens it. + self.image.touch() + self.state_dir.mkdir() + self.tile_list.write_text("\n".join([ + *self.ready, self.unready, next(iter(self.ready)), "", + ])) + for tile, exposures in { + **self.ready, self.outside: ("2243881",), + self.ignored: ("2243882",), + }.items(): + path = build_index.exp_list_path(self.run_dir, tile) + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text("".join(f"{exp}p\n" for exp in exposures)) + for tile in self.ready: + for stage in ("tile_get_images", "tile_uncompress", + "tile_find_exposures"): + self._manifest(tile, stage) + self._manifest(self.outside, "tile_vignets") + # Seed out-of-scope history; the compute parse indexes ready tiles. + build_index.build( + [self.outside, self.ignored], self.run_dir, self.index_db, + ) + machine_defaults = { + "tile_list": "$base_dir/tiles.txt", + "retrieve": "symlink", + "container": str(self.image), + "inputs": { + "tiles": "$base_dir/inputs/$run/tiles", + "exposures": "$base_dir/inputs/$run/exposures", + }, + "outputs": { + "run_dir": "$base_dir/scratch/$run", + "products_dir": "$base_dir/persistent/$run/products", + "index_db": ( + "$base_dir/persistent/$run/products/index/run_index.sqlite" + ), + }, + } + self.config = { + "run": self.name, + "machine": "candide", + "input_type": self.input_type, + "psf_model": self.psf_model, + "psf_dict": str(self.root / "psf_dict.pickle") + if self.psf_model == "fake" else "", + "clean": True, + "clean_tiles": True, + "clean_ignore_tiles": [self.ignored], + "machines": {"candide": { + "base_dir": str(self.root), + "data": machine_defaults, + "image_sims": machine_defaults, + }}, + } + self.write_config() + + def _manifest(self, tile, stage): + path = self.tile_manifest(tile, stage) + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text("{}\n") + + def write_config(self): + """Write the fixture's run configuration.""" + self.config_path.write_text(yaml.safe_dump(self.config)) + + def omit_run(self): + """Omit run: while keeping every path explicit and resolvable.""" + self.config.pop("run") + for mode in ("data", "image_sims"): + defaults = self.config["machines"]["candide"][mode] + # Resolve only fixture templates, not the production config reader. + text = yaml.safe_dump(defaults) + text = text.replace("$base_dir", str(self.root)) + text = text.replace("$run", self.name) + self.config["machines"]["candide"][mode] = yaml.safe_load(text) + self.write_config() + + def tile_manifest(self, tile, stage): + """Return an independently specified scratch manifest path.""" + return (self.run_dir / "tiles" / tile[:2] / tile + / "manifests" / f"{stage}.json") + + def final_cat(self, tile): + """Return the expected persistent catalogue for one tile.""" + return (self.products_dir / "tiles" / tile[:2] / tile + / f"final_cat-{tile}.fits") + + def persist_manifest(self, exp): + """Return the expected persistent manifest for one exposure.""" + return (self.products_dir / "exp" / exp[:2] / exp + / "manifests" / "exp_persist.json") + + +@dataclass +class ResolvedDAG: + """A live API context's workflow, campaign, and graph rendering.""" + + workflow: object + campaign: Campaign + dot: str + + @property + def namespace(self): + """Return the parsed Snakefile's namespace.""" + return self.workflow.globals + + @property + def jobs(self): + """Return resolved jobs, including already-satisfied dependencies.""" + return tuple(self.workflow.dag.jobs) + + @property + def rule_names(self): + """Return only rules with jobs in this campaign's compute DAG.""" + return {job.rule.name for job in self.jobs} + + @property + def declared_rule_names(self): + """Return every parsed rule, including rules without requested jobs.""" + return {rule.name for rule in self.workflow.rules} + + def jobs_for(self, rule): + """Return jobs of a named rule in stable wildcard order.""" + return sorted( + (job for job in self.jobs if job.rule.name == rule), + key=lambda job: sorted(job.wildcards_dict.items()), + ) + + +def load_profile(name): + """Read the committed cluster profile without invoking its executor.""" + path = REPO / "profiles" / name / "config.yaml" + return yaml.safe_load(path.read_text()) + + +@contextmanager +def resolve(campaign, monkeypatch): + """Resolve ``all`` in an isolated state directory, without running jobs.""" + from snakemake import workflow as sm_workflow + + scripts = REPO / "workflow" / "scripts" + module_names = {path.stem for path in scripts.glob("*.py")} + # Snakefile declarations share Snakemake's module-global namespace. + namespace = sm_workflow.__dict__ + original_namespace = dict(namespace) + with monkeypatch.context() as patch: + patch.syspath_prepend(str(scripts)) + for name in module_names: + patch.delitem(sys.modules, name, raising=False) + for name in list(os.environ): + if name.startswith("SP_") or name == "SNAKEMAKE_PROFILE": + patch.delenv(name) + for name, value in { + "SP_PHASE": "compute", + "SP_PROFILE": "candide", + "SP_RUN_CONFIG": campaign.config_path, + "SP_STATE_DIR": campaign.state_dir, + "SP_CONTAINER": campaign.image, + "SP_CACHE_DIR": campaign.root / "cache", + "SP_SANDBOX": campaign.root / "no-sandbox", + "SP_MISSING_THRESHOLD": "0.34", + "XDG_CACHE_HOME": campaign.root / "cache", + }.items(): + patch.setenv(name, str(value)) + profile = load_profile("candide") + try: + with SnakemakeApi() as api: + workflow_api = api.workflow( + snakefile=REPO / "workflow" / "Snakefile", + workdir=campaign.state_dir, + resource_settings=ResourceSettings( + cores=4, + default_resources=profile["default-resources"], + overwrite_resources=profile.get("set-resources", {}), + ), + deployment_settings=DeploymentSettings( + deployment_method={DeploymentMethod.APPTAINER}, + apptainer_args=profile["apptainer-args"], + ), + workflow_settings=WorkflowSettings( + runtime_source_cache_path=( + campaign.root / "source-cache" + ), + ), + ) + dag_api = workflow_api.dag(DAGSettings( + targets={"all"}, + # Snakemake 9's graph renderer compares the CLI string. + print_dag_as="dot", + rerun_triggers=RerunTrigger.parse_choices_set( + profile["rerun-triggers"], + ), + )) + stream = io.StringIO() + with redirect_stdout(stream): + dag_api.printdag() + yield ResolvedDAG( + workflow_api._workflow, campaign, stream.getvalue(), + ) + finally: + # Bare script imports carry module-level environment constants. + # Remove this parse's copies before restoring any caller's modules. + for name in module_names: + sys.modules.pop(name, None) + for name in namespace.keys() - original_namespace.keys(): + del namespace[name] + namespace.update(original_namespace) From 7f5ecf2fc6f5f65b0fe46b30de9f414ad881becc Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 04:13:55 +0200 Subject: [PATCH 81/85] test(workflow): enforce campaign DAG scope and product custody Inspect resolved job inputs for data+psfex, data+mccd, and image_sims+fake. Reclamation must wait on PSF persistence exactly when PSF products exist, and on every in-scope vignets consumer. Campaign merges must read every ready tile and keep product paths and labels tied to products_dir and run:. Missing run: must fail during parsing even with literal paths. Keep final_cat_merge for image simulations, as the workflow explicitly specifies; only exposure persistence and the star merge depend on a fitted PSF. Mutation probes cover gates, missing and out-of-scope edges, renamed products, scratch routing, directory-derived campaign names, and removal of the required run key. Co-Authored-By: GPT-6 Astra --- tests/workflow/test_dag.py | 107 +++++++++++++++++++++++++++++++++++++ workflow/CONTRACTS | 8 ++- 2 files changed, 113 insertions(+), 2 deletions(-) create mode 100644 tests/workflow/test_dag.py diff --git a/tests/workflow/test_dag.py b/tests/workflow/test_dag.py new file mode 100644 index 000000000..9d3b3b9a6 --- /dev/null +++ b/tests/workflow/test_dag.py @@ -0,0 +1,107 @@ +"""Resolved-job checks for campaign scope, product paths, and PSF custody.""" + +from collections import Counter +from pathlib import Path + +import pytest +from snakemake.exceptions import WorkflowError + +BASE_RULES = { + "all", "tile_get_images", "tile_uncompress", "tile_find_exposures", + "exp_get_images", "exp_split", "exp_psf", "clean_exposure", + "tile_exp_forest", "tile_merge_headers", "tile_detect", "tile_vignets", + "tile_ngmix", "tile_merge_cats", "tile_make_cat", "clean_tile", + "final_cat_merge", +} +PSF_RULES = {"exp_persist", "star_cat_merge"} + + +def test_rule_set_matches_input_mode(campaign, dag): + """A PSF gate cannot remove real-PSF products or add them to fake PSFs.""" + expected = BASE_RULES.copy() + if campaign.psf_model != "fake": + expected |= PSF_RULES + assert dag.rule_names == expected + assert "merge_final_cats" not in dag.declared_rule_names + + +def test_clean_exposure_waits_on_persist_iff_psf(campaign, dag): + """Reclamation waits for persistence and exactly its in-scope readers.""" + jobs = dag.jobs_for("clean_exposure") + assert {job.wildcards.exp for job in jobs} == set(campaign.exposures) + for job in jobs: + exp = job.wildcards.exp + expected = [ + campaign.tile_manifest(tile, "tile_vignets") + for tile, exposures in campaign.ready.items() if exp in exposures + ] + if campaign.psf_model != "fake": + expected.append(campaign.persist_manifest(exp)) + assert Counter(map(str, job.input)) == Counter(map(str, expected)), ( + "clean-exposure-waits-on-persist-iff-psf", exp, list(job.input) + ) + + +def test_final_cat_merge_reads_every_ready_tile(campaign, dag): + """A merge cannot drop a ready tile or pull one from outside this batch.""" + jobs = dag.jobs_for("final_cat_merge") + assert len(jobs) == 1 + assert Counter(map(str, jobs[0].input)) == Counter( + str(campaign.final_cat(tile)) for tile in campaign.ready + ) + + +def test_products_use_products_dir_and_run_name(campaign, dag): + """Neither scratch nor a directory basename can name durable products.""" + expected = { + "tile_make_cat": { + campaign.final_cat(tile) for tile in campaign.ready + }, + "final_cat_merge": { + campaign.products_dir / f"final_cat_{campaign.name}.hdf5" + }, + "star_cat_merge": set(), + "exp_persist": set(), + } + if campaign.psf_model != "fake": + expected["star_cat_merge"] = { + campaign.products_dir / f"full_starcat_{campaign.name}.hdf5" + } + expected["exp_persist"] = { + campaign.persist_manifest(exp) for exp in campaign.exposures + } + output_names = { + "tile_make_cat": "final_cat", "final_cat_merge": "merged", + "star_cat_merge": "star_cat", "exp_persist": "manifest", + } + for rule, paths in expected.items(): + actual = { + Path(getattr(job.output, output_names[rule])) + for job in dag.jobs_for(rule) + } + assert actual == paths, (rule, actual, paths) + assert all( + path.is_relative_to(campaign.products_dir) for path in actual + ) + for job in dag.jobs_for("exp_persist"): + assert Path(job.params.dest) == ( + campaign.products_dir / "exp" / job.wildcards.exp[:2] + / job.wildcards.exp / "psf" + ) + for rule in ("final_cat_merge", "star_cat_merge"): + for job in dag.jobs_for(rule): + assert job.params.campaign == campaign.name + assert f"--campaign '{campaign.name}'" in job.shellcmd + assert dag.namespace["CAMPAIGN"] == campaign.name + assert Path(dag.namespace["INDEX_DB"]) == campaign.index_db + + +def test_missing_run_fails_during_parse(campaign, resolve_dag): + """Explicit paths cannot bypass the required campaign name diagnostic.""" + campaign.omit_run() + with pytest.raises(WorkflowError, match=( + r"Unset or 'TBD' for machine='candide', input_type=" + r".*: run\. Set them in your run config \(SP_RUN_CONFIG\)\." + )): + with resolve_dag(campaign): + pytest.fail("a campaign without run: must fail at parse time") diff --git a/workflow/CONTRACTS b/workflow/CONTRACTS index c72c2e5d3..9a8df3e38 100644 --- a/workflow/CONTRACTS +++ b/workflow/CONTRACTS @@ -14,7 +14,8 @@ machine defaults' `products_dir`/`index_db`. No rule or script reads a `campaign` key or derives a name from a directory (`PRODUCTS_DIR.name`); two sources that can disagree would file one campaign's merge under another's name. Every campaign product path is rooted in `PRODUCTS_DIR`. Enforced by -tests/unit/test_campaign_lineage.py. +tests/unit/test_campaign_lineage.py and +tests/workflow/test_dag.py::test_products_use_products_dir_and_run_name. @sc [label:schema] final-cat-param-is-exact-allow-list Each input type's `config/*/final_cat.param` is the merged catalogue's exact @@ -43,4 +44,7 @@ live store, and on the tar (a leaf), or nothing, for a reclaimed one. Dropping the edge under psfex lets reclamation delete the scratch store before its PSF products reach `products_dir`; keeping it under `fake` makes every clean wait on a rule that is not in the DAG; naming the manifest for a reclaimed store lets a -`persist_exp:` edit rebuild the exposure from VOS. +`persist_exp:` edit rebuild the exposure from VOS. Enforced by +tests/workflow/test_dag.py::test_clean_exposure_waits_on_persist_iff_psf, +which also checks that the consumer edges are exactly the in-scope vignets +manifests. From 0e279d10cd97b10a8a1d61bde852b650ee4d248d Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 04:25:32 +0200 Subject: [PATCH 82/85] test(workflow): pin prologues and shells at campaign boundaries Changing unit_pre or a rule's params.pre can rerun finished tiles against reclaimed exposure stores. Pin the normalized data+psfex prologues and rendered shells so that change requires explicit review and --update-params-pin at a fresh campaign boundary. Keep the rules' full thread counts when rendering, and require params in both profiles' rerun triggers. Hash every stage and declared rule, including all resolved wildcard instances, and retain normalized payloads in the temporary fixture for diagnosis. Check path independence across distinct roots and document API access, scope, regeneration, and mutation probes. Workflow contracts now name their enforcing checks. Co-Authored-By: GPT-6 Astra --- tests/workflow/README.md | 76 +++++++++++++++++++++++++++++++ tests/workflow/conftest.py | 23 +++++++++- tests/workflow/harness.py | 3 +- tests/workflow/params.py | 57 +++++++++++++++++++++++ tests/workflow/params_pin.json | 41 +++++++++++++++++ tests/workflow/test_params_pin.py | 38 ++++++++++++++++ workflow/CONTRACTS | 7 ++- 7 files changed, 242 insertions(+), 3 deletions(-) create mode 100644 tests/workflow/README.md create mode 100644 tests/workflow/params.py create mode 100644 tests/workflow/params_pin.json create mode 100644 tests/workflow/test_params_pin.py diff --git a/tests/workflow/README.md b/tests/workflow/README.md new file mode 100644 index 000000000..4c2173e71 --- /dev/null +++ b/tests/workflow/README.md @@ -0,0 +1,76 @@ +# Workflow DAG checks + +Run these checks inside the development container: + +```bash +python -m pytest tests/workflow -o addopts='' -q -p no:cacheprovider +``` + +`testpaths = ["tests"]` includes this directory in the full suite and the image-build CI run. +These tests check planning, not job execution, container validity, or SLURM group execution. +They need no survey files, cluster access, or nested Apptainer process. + +## Campaign and API + +`harness.Campaign` writes run YAML, a tile list, exposure lists, and empty prepared-stage manifests under `tmp_path`. +Two ready tiles use two exposures, one of them shared; the declared list also contains a duplicate and an unready tile. +An out-of-scope consumer has finished its vignets, and another consumer is explicitly ignored. +`build_index.build()` seeds those two consumers in the SQLite history; the Snakefile's compute parse builds the current batch's index from the exposure lists. +A prepare parse does not build the index, so no prepare invocation is necessary. + +The run configuration exercises `machine: candide` and `$base_dir`/`$run` expansion for both input types. +Scratch and persistent roots differ, and the products directory's basename deliberately differs from `run:`. +The three modes are `data+psfex`, `data+mccd`, and `image_sims+fake`. +`final_cat_merge` is present in all three: it merges galaxy catalogues regardless of the PSF source. +Only `exp_persist` and `star_cat_merge` disappear for fake PSFs. + +`resolve()` uses the same `SP_PHASE`, `SP_PROFILE`, `SP_RUN_CONFIG`, image selection, and state-directory conventions as `workflow/bin/sp`. +It isolates the environment, source cache, bare script imports, and Snakemake's shared global namespace. +An empty image sentinel satisfies parse-time image resolution without deploying software. +The candide profile supplies the resource defaults, resource overrides, Apptainer arguments, and rerun triggers; no executor starts. + +The API sequence is `SnakemakeApi()` → `api.workflow(...)` → `workflow_api.dag(DAGSettings(targets={"all"}, ...))` → `dag_api.printdag()`. +`print_dag_as="dot"` matches Snakemake 9's renderer, which compares the CLI string rather than its enum default. +The resolved jobs come from `workflow_api._workflow.dag.jobs`; this last access is private because the public API exposes graph printing but not Job objects. +`ResolvedDAG` exposes the active rule set, declared rule set, namespace, and `jobs_for(rule)` for input/output, wildcard, params, and rendered-shell inspection. +Jobs retain the rules' thread counts rather than a local `--cores` cap. +The API context remains open while tests inspect jobs and closes before the fixture restores process state. + +## Campaign-boundary pin + +`params_pin.json` pins SHA-256 digests of: + +- `unit_pre()` rendered for every stage; +- every rule's shell template; +- every resolved job's `params.pre` and formatted shell, including all ngmix chunks. + +Rules without `params.pre` use null; aggregation-only rules have no shell. +The pin includes per-stage and per-rule digests to identify which strings change. +Fixture and checkout roots become `` and `` before hashing; two independently located campaigns must give the same digest. +The normalized strings are also written to the fixture's `params_rendered.json` for inspection. +Script fingerprints and non-pre params that do not appear in a shell are outside this pin's scope. +Separate checks require `params` in both profiles' rerun triggers. + +A pin change requires a campaign boundary on a fresh root. +Review the rendered strings, then regenerate and commit the pin with the intentional workflow change: + +```bash +python -m pytest tests/workflow/test_params_pin.py --update-params-pin \ + -o addopts='' -q -p no:cacheprovider +``` + +## Failure modes and mutation probes + +Each check has a mutation that must make it fail. +Apply mutations only to a disposable checkout, run the named test without `--update-params-pin`, and require assertion failures rather than collection/setup errors. + +| Check | Mutation probes | +|---|---| +| `test_rule_set_matches_input_mode` | Invert `PERSISTS_PSF`; omit `star_cat_targets()`; rename `final_cat_merge` to `merge_final_cats`. | +| `test_clean_exposure_waits_on_persist_iff_psf` | Drop the persist edge; make it unconditional under fake PSFs; drop vignets consumers; remove the in-scope consumer filter. | +| `test_final_cat_merge_reads_every_ready_tile` | Drop one ready tile; append an out-of-scope tile. | +| `test_products_use_products_dir_and_run_name` | Rename either merged catalogue or the persist manifest; route products to scratch; derive `CAMPAIGN` from the products directory's basename. | +| `test_missing_run_fails_during_parse` | Remove `run` from `run_config.REQUIRED`; literal paths must still receive the required-key diagnostic, not a later `KeyError`. | +| `test_unit_pre_changes_at_campaign_boundary` | Append a line to `unit_pre`; change one rule's `params.pre`; change one shell; change a rendered thread count. | +| `test_params_pin_ignores_fixture_root` | Remove fixture-root normalization. | +| `test_params_is_a_rerun_trigger_under_both_profiles` | Remove `params` from candide or nibi's `rerun-triggers`. | diff --git a/tests/workflow/conftest.py b/tests/workflow/conftest.py index 32591b851..134a9c62a 100644 --- a/tests/workflow/conftest.py +++ b/tests/workflow/conftest.py @@ -2,7 +2,15 @@ import pytest -from tests.workflow.harness import MODES, Campaign, resolve +from tests.workflow.harness import MODES, Campaign, load_profile, resolve + + +def pytest_addoption(parser): + """Expose the explicit campaign-boundary pin update switch.""" + parser.addoption( + "--update-params-pin", action="store_true", default=False, + help="Update the reviewed params.pre/shell pin at a campaign boundary", + ) @pytest.fixture(params=MODES, ids=[f"{mode}+{psf}" for mode, psf in MODES]) @@ -22,3 +30,16 @@ def dag(campaign, resolve_dag): """Keep the API and its jobs alive for the duration of one check.""" with resolve_dag(campaign) as resolved: yield resolved + + +@pytest.fixture +def psfex_dag(tmp_path, resolve_dag): + """Resolve the canonical data+psfex campaign for the prologue pin.""" + with resolve_dag(Campaign(tmp_path / "campaign", "data", "psfex")) as dag: + yield dag + + +@pytest.fixture(params=["candide", "nibi"]) +def profile(request): + """Read each profile's actual rerun-trigger policy.""" + return load_profile(request.param) diff --git a/tests/workflow/harness.py b/tests/workflow/harness.py index 6b86c7543..cab0fed9a 100644 --- a/tests/workflow/harness.py +++ b/tests/workflow/harness.py @@ -228,7 +228,8 @@ def resolve(campaign, monkeypatch): snakefile=REPO / "workflow" / "Snakefile", workdir=campaign.state_dir, resource_settings=ResourceSettings( - cores=4, + # Cluster jobs retain each rule's full thread count. + nodes=profile["jobs"], default_resources=profile["default-resources"], overwrite_resources=profile.get("set-resources", {}), ), diff --git a/tests/workflow/params.py b/tests/workflow/params.py new file mode 100644 index 000000000..bb0baf86f --- /dev/null +++ b/tests/workflow/params.py @@ -0,0 +1,57 @@ +"""Fingerprint rendered prologues and shells without checkout or tmp paths.""" + +import hashlib +import json + +from tests.workflow.harness import REPO + + +def params_pin(dag): + """Return SHA-256 digests for unit_pre and every declared rule's shells. + + Each rule includes all resolved wildcard instances, its shell template, + and its rendered ``params.pre`` (null for rules without that parameter). + The normalized payload is saved beside the fixture for review on failure. + Script hashes and other non-pre params are outside this pin's scope. + """ + campaign = dag.campaign + + def normalize(value): + if value is None: + return None + return (value.replace(str(campaign.root), "") + .replace(str(REPO), "")) + + def digest(value): + text = json.dumps(value, sort_keys=True, separators=(",", ":")) + return hashlib.sha256(text.encode()).hexdigest() + + namespace = dag.namespace + unit_pre = {} + for stage, (level, _) in sorted(namespace["STAGE_DIR"].items()): + unit = next(iter(campaign.ready)) if level == "tile" else "2243881" + unit_pre[stage] = normalize(namespace["unit_pre"](stage, unit)) + rules = {} + for rule in sorted(dag.workflow.rules, key=lambda rule: rule.name): + jobs = dag.jobs_for(rule.name) + # Aggregation-only targets have no shell or pre to expand. + assert jobs or (not rule.shellcmd and not rule.params), rule.name + rules[rule.name] = { + "shell_template": normalize(rule.shellcmd), + "jobs": [{ + "wildcards": dict(sorted(job.wildcards_dict.items())), + "pre": normalize(getattr(job.params, "pre", None)), + "shell": normalize(job.shellcmd), + } for job in jobs], + } + payload = {"unit_pre": unit_pre, "rules": rules} + (campaign.root / "params_rendered.json").write_text( + json.dumps(payload, indent=2, sort_keys=True) + "\n", + ) + return { + "schema": 1, + "algorithm": "sha256", + "sha256": digest(payload), + "unit_pre": {stage: digest(pre) for stage, pre in unit_pre.items()}, + "rules": {name: digest(rule) for name, rule in rules.items()}, + } diff --git a/tests/workflow/params_pin.json b/tests/workflow/params_pin.json new file mode 100644 index 000000000..1924265dc --- /dev/null +++ b/tests/workflow/params_pin.json @@ -0,0 +1,41 @@ +{ + "algorithm": "sha256", + "rules": { + "all": "572122d8d1901e12ff591b30adf405f8920be8183129befececbb819f3392ed8", + "clean_exposure": "22cb76b13a5205d20a02a9bd3b8c8bea5ea2801555e24dd2b84b11145f7e79d9", + "clean_tile": "a5c07b0461526ed407df36a291deb866d4181524b4c8e0fbc3dd047fd9d28479", + "exp_get_images": "71e76ff7f96af1c5d2b85697cc5819e2271911f177a2075253f3b8cfa268c1a9", + "exp_persist": "302e2837542bc1102430c27c81c600b7cda32e8bddcb5fd60d33950987609fff", + "exp_psf": "2c4f6d00f1939ccbf05b4982a0202a0ff92a4373a727f4e00f3aaab4eba03352", + "exp_split": "6e954f8f3d06f44d3f164675912ce27d6216648855d04968f9168bd7d0f2c4fa", + "final_cat_merge": "e7f46859c4503a2220713d7bb2507555515d0a9632d780b20f14c59e32210023", + "prepare_all_tiles": "b8f872a22adf014e25a7fa5198f49b71a6fe9e56042ed82b682bc8763970a844", + "star_cat_merge": "6277450958474af5270982fa35360f2f237a29f7533c526ee9265dfd5acc07a0", + "tile_detect": "457456f4f71a500e4166a63933b88d14eed34a019a9814f44546e137b0e71d7a", + "tile_exp_forest": "7447ab4a1049de5f0b5c81e5f9ed2a8644c7bdab85cb0a060bfde81261a89b28", + "tile_find_exposures": "8704317871744996c44351c2836fcb222d7986a602e9e046a90d684ba7b3c184", + "tile_get_images": "331a67e747f211ebf4c14b946a7af7f9fc9f55243d69d2791d74aecc3ca228c3", + "tile_make_cat": "a57518b04c11f70bf41b320532fddffdd1115fdac604d7dd3928eb61cca28f23", + "tile_merge_cats": "ff21216ea804dccc2d2c290d2b2499d5d05f0c34c0a56993c233f43fe3c06bdb", + "tile_merge_headers": "7a344849d62936e2f5598dc8731a2c4947eff2c4c7218b1a731dd2a9577e7111", + "tile_ngmix": "6ed10a3a4d3658ba0303fb0c23ad5100ec8cbdc638ca36c3b25875087b5c3415", + "tile_uncompress": "1e2b01acbf9708e0371070fb01c5b9568d7efb5f89d1835fc6bf91e2c8b60cb3", + "tile_vignets": "9d4ae0d99c18217f2185f245281312454c8a219ec1628176e08c71a5efc4dc91" + }, + "schema": 1, + "sha256": "74adbcd14bc304f82c3f2d1a50374bc607e5313fa5f3f8b3db6302653bad67ea", + "unit_pre": { + "exp_get_images": "8dec850af212879f225fcf27a5f1281e1a075264c7b97d38c2214395d360168c", + "exp_psf": "f2358ddf7385918dc5033d10b37f6dc97a15d02b071a3ea0a4619a5f7e6f5bec", + "exp_split": "358fa8bbe59680d4f9839e007b343dd25fed733cf3156dccd30d305d33ac9480", + "tile_detect": "adad5d671fa65dd04433e1c82a725b845635e9d2e368a1e70833ed990df2d55a", + "tile_find_exposures": "c5922fb507f6fd040a179b53fba0818697661016c9688984b6ce249186dc6986", + "tile_get_images": "45b47c44bfeb34ac973b89d4e752c8c82e78028f9f0da379c44627005b45e279", + "tile_make_cat": "7579a52e32e76523c0b77d75c468a47d7f2e2ca0c55865f40e9628a7481cd79d", + "tile_merge_cats": "69cfe94ba941d2ba7b1ce24883961b47b191b80aff381da1865c6bc38de0c354", + "tile_merge_headers": "8590cfa3281c88d43eb8c4fe5760176a13e928a419cbea9bea3a8f88b0634b42", + "tile_ngmix": "b6b75b553a62df54aea6c337dcde1e0cdcd0fe30000e86880c1a4689e42c4a72", + "tile_uncompress": "91a15538491e53ee2b0d52b0472909374270918e4b4c80880c4f5a8329e40161", + "tile_vignets": "a6111cef708aa3c9fb0144b49de9f781eef84d2096ba4f8a3de9bb45b1a3774d" + } +} diff --git a/tests/workflow/test_params_pin.py b/tests/workflow/test_params_pin.py new file mode 100644 index 000000000..faeb43158 --- /dev/null +++ b/tests/workflow/test_params_pin.py @@ -0,0 +1,38 @@ +"""Campaign-boundary checks on the shell and params.pre rerun surface.""" + +import json +from pathlib import Path + +from tests.workflow.harness import Campaign +from tests.workflow.params import params_pin + +PIN = Path(__file__).with_name("params_pin.json") +BOUNDARY_MESSAGE = ( + "params.pre is a rerun trigger under both profiles; this change reruns " + "every finished unit of a resumed campaign — land it at a campaign " + "boundary and update the pin" +) + + +def test_unit_pre_changes_at_campaign_boundary(psfex_dag, pytestconfig): + """A shared prologue or per-rule shell edit must move the reviewed pin.""" + actual = params_pin(psfex_dag) + if pytestconfig.getoption("--update-params-pin"): + PIN.write_text(json.dumps(actual, indent=2, sort_keys=True) + "\n") + assert PIN.is_file(), BOUNDARY_MESSAGE + assert actual == json.loads(PIN.read_text()), BOUNDARY_MESSAGE + + +def test_params_pin_ignores_fixture_root(tmp_path, resolve_dag): + """A different temporary campaign directory cannot require a new pin.""" + pins = [] + for directory in ("first-root", "another-root"): + campaign = Campaign(tmp_path / directory, "data", "psfex") + with resolve_dag(campaign) as dag: + pins.append(params_pin(dag)) + assert pins[0] == pins[1], "normalise fixture paths before hashing" + + +def test_params_is_a_rerun_trigger_under_both_profiles(profile): + """Neither cluster profile can silently disable the protected trigger.""" + assert "params" in profile["rerun-triggers"], BOUNDARY_MESSAGE diff --git a/workflow/CONTRACTS b/workflow/CONTRACTS index 9a8df3e38..9538e78cf 100644 --- a/workflow/CONTRACTS +++ b/workflow/CONTRACTS @@ -34,7 +34,12 @@ replans every finished unit of the campaign. Mid-campaign that rerun is unsatisfiable for tiles whose exposure stores were reclaimed, and the failed tile group's cleanup deletes those finished tiles' `final_cat`. Change `unit_pre` output, or anything else in a rule's `params`, only at a campaign -boundary on a fresh root. +boundary on a fresh root. The rendered `unit_pre`, per-job `params.pre`, and +shell strings are pinned by +tests/workflow/test_params_pin.py::test_unit_pre_changes_at_campaign_boundary; +test_params_is_a_rerun_trigger_under_both_profiles checks the profile premise. +Regenerate tests/workflow/params_pin.json only at that boundary with +`pytest tests/workflow/test_params_pin.py --update-params-pin`. @sc [label:custody] clean-exposure-waits-on-persist-iff-psf `clean_exposure` waits on the exposure's persisted PSF products exactly when From a846e088bae46e914bc98db96c38bbcbecba9f19 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Sat, 26 Sep 2026 11:18:22 +0200 Subject: [PATCH 83/85] test(workflow): the DAG harness follows the parse-time guards psf_model=mccd is refused before a DAG exists, so it leaves the resolved modes and gets a refusal test; the required-run: message names unexpanded variables too. Co-Authored-By: Claude Fable 5.1 --- tests/workflow/harness.py | 4 +++- tests/workflow/test_dag.py | 12 +++++++++++- 2 files changed, 14 insertions(+), 2 deletions(-) diff --git a/tests/workflow/harness.py b/tests/workflow/harness.py index cab0fed9a..e3e487b28 100644 --- a/tests/workflow/harness.py +++ b/tests/workflow/harness.py @@ -26,7 +26,9 @@ from workflow.scripts import build_index REPO = Path(__file__).resolve().parents[2] -MODES = [("data", "psfex"), ("data", "mccd"), ("image_sims", "fake")] +MODES = [("data", "psfex"), ("image_sims", "fake")] +# psf_model=mccd is refused at parse time (refuse_unpersistable_psf), so it +# has no DAG to resolve; test_mccd_is_refused_during_parse pins the refusal. @dataclass diff --git a/tests/workflow/test_dag.py b/tests/workflow/test_dag.py index 9d3b3b9a6..e582138bc 100644 --- a/tests/workflow/test_dag.py +++ b/tests/workflow/test_dag.py @@ -6,6 +6,8 @@ import pytest from snakemake.exceptions import WorkflowError +from tests.workflow.harness import Campaign + BASE_RULES = { "all", "tile_get_images", "tile_uncompress", "tile_find_exposures", "exp_get_images", "exp_split", "exp_psf", "clean_exposure", @@ -100,8 +102,16 @@ def test_missing_run_fails_during_parse(campaign, resolve_dag): """Explicit paths cannot bypass the required campaign name diagnostic.""" campaign.omit_run() with pytest.raises(WorkflowError, match=( - r"Unset or 'TBD' for machine='candide', input_type=" + r"for machine='candide', input_type=" r".*: run\. Set them in your run config \(SP_RUN_CONFIG\)\." )): with resolve_dag(campaign): pytest.fail("a campaign without run: must fail at parse time") + + +def test_mccd_is_refused_during_parse(tmp_path, resolve_dag): + """MCCD products are unreadable to persistence and the star merge.""" + campaign = Campaign(tmp_path / "campaign", "data", "mccd") + with pytest.raises(WorkflowError, match=r"psf_model=mccd: PSF persistence"): + with resolve_dag(campaign): + pytest.fail("psf_model=mccd must be refused at parse time") From 2d833fd5fc93018912e148262537b65e98d98b77 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Mon, 28 Sep 2026 15:59:47 +0200 Subject: [PATCH 84/85] test: seams and run_config follow the develop merge test_hsm_column_seams reads MergeStarCatPSFEX's read set from its _COLUMNS table (the two-pass merge reads by table, not by literal); test_run_config pins ${name} shorthands, their nesting, run written back, and a run left holding $var reported. Co-Authored-By: Claude Opus 5.5 --- tests/module/test_hsm_column_seams.py | 39 ++++++++++++--------------- tests/unit/test_run_config.py | 15 +++++++++++ 2 files changed, 32 insertions(+), 22 deletions(-) diff --git a/tests/module/test_hsm_column_seams.py b/tests/module/test_hsm_column_seams.py index b05c52b69..8459003b3 100644 --- a/tests/module/test_hsm_column_seams.py +++ b/tests/module/test_hsm_column_seams.py @@ -4,8 +4,9 @@ (psfex_interp ↔ merge_starcat) and ``psfex-me-shapes-columns`` / ``psf-epoch-slot-columns`` (psfex_interp ↔ make_cat). The producer side is ``_hsm_columns`` (``hsm-column-grammar``), which every psfex_interp writer -goes through; the consumers read column names as string literals with no -fallback, collected from each function's AST. +goes through; make_cat reads column names as string literals with no fallback, +collected from the function's AST, and merge_starcat reads its ``_COLUMNS`` +table. """ import ast @@ -24,25 +25,10 @@ ) -def _hsm_literals(func, subscript_of=None): - """Set of ``HSM_*`` string literals in ``func``'s source. - - With ``subscript_of``, only literals used as ``["HSM_..."]`` reads - count, so a column that is still *written* under the same name cannot - mask a dropped read. - """ +def _hsm_literals(func): + """Set of ``HSM_*`` string literals in ``func``'s source.""" tree = ast.parse(textwrap.dedent(inspect.getsource(func))) - if subscript_of is None: - nodes = (n for n in ast.walk(tree) if isinstance(n, ast.Constant)) - else: - nodes = ( - n.slice - for n in ast.walk(tree) - if isinstance(n, ast.Subscript) - and isinstance(n.value, ast.Name) - and n.value.id == subscript_of - and isinstance(n.slice, ast.Constant) - ) + nodes = (n for n in ast.walk(tree) if isinstance(n, ast.Constant)) return { n.value for n in nodes @@ -55,9 +41,18 @@ def _written(obj): def test_merge_starcat_reads_exactly_what_psfex_validation_writes(): - """psfex-validation-hsm-columns == psfex-starcat-columns-strict.""" + """psfex-validation-hsm-columns == psfex-starcat-columns-strict. + + ``process`` reads every ``_COLUMNS`` source by name with no fallback (only + ``_OPTIONAL`` is zero-filled), so the table is the read set. + """ written = _written("PSF") | _written("STAR") - read = _hsm_literals(MergeStarCatPSFEX.process, subscript_of="data_j") + read = { + col for _, col in MergeStarCatPSFEX._COLUMNS if col.startswith("HSM_") + } + assert not any( + col.startswith("HSM_") for _, col in MergeStarCatPSFEX._OPTIONAL + ) assert written == read, { "written_not_read": sorted(written - read), "read_not_written": sorted(read - written), diff --git a/tests/unit/test_run_config.py b/tests/unit/test_run_config.py index da3281d05..3290f9fc3 100644 --- a/tests/unit/test_run_config.py +++ b/tests/unit/test_run_config.py @@ -75,6 +75,21 @@ def test_base_dir_holding_run_expands_fully(): assert run_config.unresolved(cfg) == [] +def test_run_config_shorthands_nest_and_run_is_written_back(): + """Top-level scalars are variables, ${name} delimits, nesting resolves.""" + cfg = run_config.apply_machine_defaults(_machine_config( + "/b", shear="1p2z", grid="grid_2", run="${shear}_${grid}", + outputs={"run_dir": "/o/$run/scratch"})) + assert cfg["run"] == "1p2z_grid_2" + assert cfg["outputs"]["run_dir"] == "/o/1p2z_grid_2/scratch" + assert run_config.unresolved(cfg) == [] + + +def test_run_holding_an_unknown_variable_is_reported(): + cfg = run_config.apply_machine_defaults(_machine_config("/b", run="$nope")) + assert "run" in run_config.unresolved(cfg) + + def test_optional_path_with_an_unknown_variable_is_reported(): cfg = run_config.apply_machine_defaults(_machine_config( "/b", outputs={"products_dir": "/p/$nope/products"}, From aa458b21ca28c23392d4d68d970aaa534a33e0d3 Mon Sep 17 00:00:00 2001 From: Cail Daley Date: Mon, 28 Sep 2026 16:17:32 +0200 Subject: [PATCH 85/85] test(workflow): snakemake-driven checks run in their own CI step Develop dropped snakemake from the image (10f9c535: it is a host tool that wraps each job in the container). The root conftest leaves tests/workflow out of collection where snakemake is absent, and deploy-image.yml runs that directory as its own step after installing snakemake>=9,<10 into a throwaway container, so the image stays snakemake-free and the DAG checks still gate CI. Co-Authored-By: Claude Opus 5.5 --- .github/workflows/deploy-image.yml | 11 +++++++++++ conftest.py | 10 ++++++++++ tests/workflow/README.md | 9 +++++++-- 3 files changed, 28 insertions(+), 2 deletions(-) diff --git a/.github/workflows/deploy-image.yml b/.github/workflows/deploy-image.yml index 32e76ca44..2bb4238a7 100644 --- a/.github/workflows/deploy-image.yml +++ b/.github/workflows/deploy-image.yml @@ -115,6 +115,17 @@ jobs: IMAGE=$(echo "${{ steps.meta.outputs.tags }}" | head -n1) docker run --rm -e HYPOTHESIS_PROFILE=ci -e SHAPEPIPE_ON_CANDIDE=0 "$IMAGE" pytest -rX + # tests/workflow drives the Snakefile through snakemake's API. Snakemake + # is a host tool and stays out of the image (it wraps each job in the + # container), so the suite above leaves that directory out + # (the root conftest.py); here it is installed into this throwaway container + # at the host pin range (workflow/README.md) and the directory runs alone. + - name: Test — workflow DAG (snakemake) + run: | + IMAGE=$(echo "${{ steps.meta.outputs.tags }}" | head -n1) + docker run --rm -e SHAPEPIPE_ON_CANDIDE=0 "$IMAGE" bash -c \ + "uv pip install 'snakemake>=9,<10' && pytest -rX --no-cov tests/workflow" + # ---------------------------------------------------------------- # Publish (push events only — never on pull_request, incl. forks). # Fires on any branch; the image is tagged with the branch name. diff --git a/conftest.py b/conftest.py index e3936b30b..f481d46d1 100644 --- a/conftest.py +++ b/conftest.py @@ -28,6 +28,16 @@ settings.load_profile(os.environ.get("HYPOTHESIS_PROFILE", "ci")) +# ``tests/workflow/`` drives the Snakefile through snakemake's API. Snakemake is +# a host tool, not part of the image (it wraps each job in the container), so +# where it is absent the directory is left out of collection; CI runs it in its +# own step after installing snakemake (deploy-image.yml). +try: + import snakemake # noqa: F401 +except ModuleNotFoundError: + collect_ignore = ["tests/workflow"] + + # --------------------------------------------------------------------------- # # Candide detection # --------------------------------------------------------------------------- # diff --git a/tests/workflow/README.md b/tests/workflow/README.md index 4c2173e71..062b4ce49 100644 --- a/tests/workflow/README.md +++ b/tests/workflow/README.md @@ -1,12 +1,17 @@ # Workflow DAG checks -Run these checks inside the development container: +These checks need snakemake, which is a host tool and not in the image. Run them +in the development container with snakemake installed on top (a writable sandbox, +or a throwaway container), at the host pin range from `workflow/README.md`: ```bash +uv pip install 'snakemake>=9,<10' python -m pytest tests/workflow -o addopts='' -q -p no:cacheprovider ``` -`testpaths = ["tests"]` includes this directory in the full suite and the image-build CI run. +Where snakemake is absent, the root `conftest.py` leaves this directory out of +collection, so the in-image suite stays green. The image-build CI runs it as its +own step, `Test — workflow DAG (snakemake)`, which installs snakemake first. These tests check planning, not job execution, container validity, or SLURM group execution. They need no survey files, cluster access, or nested Apptainer process.