Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
95 commits
Select commit Hold shift + click to select a range
681f156
feat(orchestration): the keep list, and the script that acts on it
Sep 1, 2026
de243a8
feat(orchestration): exp_persist, between exp_psf and clean_exposure
Sep 1, 2026
4cf639f
docs(orchestration): exp_persist in the rule list, and why the report…
Sep 1, 2026
1fa6455
feat(orchestration): exp_persist packs one tar per exposure, not loos…
cailmdaley Sep 3, 2026
a14a2f3
fix(orchestration): persist_exp never leaves a .tmp behind on failure
cailmdaley Sep 3, 2026
d91da36
docs(orchestration): drop references to the removed star-catalogue rules
cailmdaley Sep 9, 2026
626aaba
feat(orchestration): the two campaign-level merges
cailmdaley Sep 9, 2026
fc9aa3e
fix(orchestration): seven defects in the campaign-level merges
cailmdaley Sep 9, 2026
d774cc3
fix(create_final_cat): make the merged-catalogue writer reproducible
cailmdaley Sep 9, 2026
98532ce
perf(orchestration): size the two merges from the data, not from a guess
cailmdaley Sep 9, 2026
d0ecfdd
feat(orchestration): persist_exp keeps NAMED products, not globs
cailmdaley Sep 9, 2026
3b8c7d5
perf(merge_starcat): accumulate arrays, not python lists of floats
cailmdaley Sep 9, 2026
0ad6403
feat(orchestration): the star catalogue's inputs are not a user choice
cailmdaley Sep 9, 2026
90dfb00
fix(cfis): two stale columns in final_cat.param, and no mask column a…
cailmdaley Sep 9, 2026
a434c9a
perf(merge_starcat): two passes, so nothing is held twice
cailmdaley Sep 9, 2026
3c29825
feat(orchestration): final_cat_merge reconciles instead of rebuilding
cailmdaley Sep 9, 2026
78d2ce9
fix(orchestration): eight defects found reviewing the merge work
cailmdaley Sep 10, 2026
0fb514b
feat(orchestration): the star catalogue becomes hdf5, reconciled like…
cailmdaley Sep 10, 2026
4ed9fb0
fix(orchestration): nine findings from the third review
cailmdaley Sep 10, 2026
e86d8c8
test(unit): property-based state machines for reconcile and persist_exp
cailmdaley Sep 10, 2026
1a343e6
fix(persist-exp): the PSF run dir is run_sp_exp_SxSePsf
cailmdaley Sep 10, 2026
5b88f17
workflow: image simulations as an input mode of the real-data workflow
Sep 11, 2026
4e17fee
workflow: image-sims merge column list, and the merge reads the workf…
Sep 11, 2026
716b160
workflow: campaign-unique node-local tile store for image sims; candi…
Sep 11, 2026
fdc9db4
workflow: port the MCCD exposure chain to the workflow grammar
Sep 11, 2026
836f91d
workflow: exp_psf reserves 2 cores, not 8, for the MCCD chain
Sep 12, 2026
ad80f23
mccd_interp: test for N_EPOCH among the column names
Sep 12, 2026
2082dcf
workflow: MCCD validated through the full chain; drop the warn-only f…
Sep 12, 2026
9085cd0
workflow: tile_store_root run-config key moves the tile store's bind,…
Sep 12, 2026
7637b39
completeness.py: atomic write for stage log/manifest JSON
Sep 15, 2026
7aababf
workflow/config.yaml: universal template, no cluster-specific paths
Sep 15, 2026
d64f88e
workflow: retrieve mode (symlink|vos) is a run-config key, not fixed …
Sep 15, 2026
c7bcfdd
config.yaml: restore real nibi defaults, keep the placeholder check a…
Sep 15, 2026
5be4c43
Snakefile: check machine: against SP_PROFILE at parse time
Sep 15, 2026
29442d3
config.yaml: machines: table drives per-machine (and per-input_type) …
Sep 15, 2026
7293e3b
workflow: one run-config resolver for Snakefile, bin/sp and container.py
Sep 15, 2026
084e305
run_config: `run:` name, expanded as $run in run-config paths
Sep 15, 2026
78cc0ca
added user run config template
Sep 17, 2026
d3b088a
bin/sp: -c/--config-file for the run config, instead of SP_RUN_CONFIG
Sep 17, 2026
746ca9c
bin/sp: document -c, the two-file merge and the environment in the he…
Sep 17, 2026
4e0bedc
improved commeents
Sep 17, 2026
3e269c4
final_cat_merge, star_cat_merge: record code provenance in the HDF5 a…
cailmdaley Sep 17, 2026
fbfdf3b
workflow smoothed; running until hdf5 file
Sep 23, 2026
0b5fae7
bin/sp: don't swallow -c/--config meant for the delegated command
cailmdaley Sep 25, 2026
d22c4a0
workflow/README: add products_dir to the image-sims run-config example
cailmdaley Sep 25, 2026
5ded2ca
run_template.yaml: make it resolve
cailmdaley Sep 25, 2026
dae9e1b
docs(astra): record the pipeline's scientific decisions in astra.yaml
cailmdaley Aug 31, 2026
e901cc9
test(astra): validate decision anchors and universe pins
cailmdaley Sep 26, 2026
8e40991
Merge remote-tracking branch 'origin/develop' into feat/persist-exp-p…
cailmdaley Sep 26, 2026
d84822f
merge: feat/workflow-image-sims (#894 @5ded2ca8) + develop into feat/…
cailmdaley Sep 26, 2026
4b1e868
workflow: final_cat_merge is the one merger, for data and image sims
cailmdaley Sep 26, 2026
d7ce47f
workflow: the campaign name is `run:`, and `run:` is required
cailmdaley Sep 26, 2026
3386347
workflow: no PSF persistence under psf_model: fake
cailmdaley Sep 26, 2026
89f00b5
test(grammar): check the image-sims final_cat.param too
cailmdaley Sep 26, 2026
1685d65
Merge remote-tracking branch 'origin/docs/astra-decision-record' into…
cailmdaley Sep 26, 2026
09edd5e
docs(astra): record the pipeline's scientific decisions in astra.yaml
cailmdaley Aug 31, 2026
206b48d
test(astra): validate decision anchors and universe pins
cailmdaley Sep 26, 2026
9f9ec5e
test(astra): resolve Snakemake rule anchors; JSON report mode
cailmdaley Sep 26, 2026
5368d9d
docs(astra): rewrite the decision record against develop
cailmdaley Sep 26, 2026
a6eff5e
docs(claude): point the scientific-decisions section at the anchor test
cailmdaley Sep 26, 2026
62a2699
feat(make_cat): fixed per-epoch slot count via N_EPOCH_SLOTS
cailmdaley Sep 26, 2026
512e7bf
feat(workflow): save fixed-slot per-epoch PSF data in tile make_cat
cailmdaley Sep 26, 2026
b2a4d19
fix(workflow): five review findings on the #894 merge
cailmdaley Sep 26, 2026
4958d66
workflow: scientific contracts for the campaign products
cailmdaley Sep 26, 2026
5fcc99e
test(workflow): the campaign name has one source, and products one root
cailmdaley Sep 26, 2026
e768908
test(final_cat_merge): columns, per-epoch slots, never-fit rows, miss…
cailmdaley Sep 26, 2026
e3f6f4a
docs(astra): correct seven rationale claims against the code
cailmdaley Sep 26, 2026
27b59f7
Merge remote-tracking branch 'origin/docs/astra-decision-record' into…
cailmdaley Sep 26, 2026
eb05d3a
Merge remote-tracking branch 'origin/feat/make-cat-fixed-epoch-slots'…
cailmdaley Sep 26, 2026
d214966
Merge remote-tracking branch 'origin/docs/astra-decision-record' into…
cailmdaley Sep 26, 2026
3358105
config(cfis): final_cat.param carries the 12 per-epoch slots (EXP_ID,…
cailmdaley Sep 26, 2026
bac6eea
docs(astra): two lints this stack resolves — IMAFLAGS_ISO is not merg…
cailmdaley Sep 26, 2026
df20cf4
config(cfis): final_cat.param keeps NUMBER
cailmdaley Sep 26, 2026
05e7381
fix(workflow): the sims overlay resolves gauss_3.0_7x7.conv; `sp -- ……
cailmdaley Sep 26, 2026
45a9f1b
fix(workflow): refuse psf_model=mccd at parse time
cailmdaley Sep 26, 2026
16df7ae
fix(workflow): refuse reclamation when products_dir is run_dir
cailmdaley Sep 26, 2026
96abec8
fix(workflow): run_config expands to a fixed point and reports any le…
cailmdaley Sep 26, 2026
072b705
fix(persist_exp): the manifest records each member's sha256
cailmdaley Sep 26, 2026
d179f2b
fix(hdf5_reconcile): one type per column across a campaign
cailmdaley Sep 26, 2026
187968e
test(hdf5_reconcile): isolate schema invalidation from source stamps
cailmdaley Sep 26, 2026
70d7a20
fix(build_index): readiness is the tiles row, not the retained edges
cailmdaley Sep 26, 2026
5a61aaf
fix(hdf5_reconcile): delete an abandoned tmp before the free-space check
cailmdaley Sep 26, 2026
47b2ec5
fix(hdf5_reconcile): each provenance record replaces the last
cailmdaley Sep 26, 2026
b15d442
fix(workflow): a reclaimed exposure's clean waits on its tar, not the…
cailmdaley Sep 26, 2026
5b56a68
test(conftest): the candide hostname test matches the whole bare name
cailmdaley Sep 26, 2026
33391a4
test(workflow): add an isolated Snakemake DAG driver
cailmdaley Sep 26, 2026
7f5ecf2
test(workflow): enforce campaign DAG scope and product custody
cailmdaley Sep 26, 2026
0e279d1
test(workflow): pin prologues and shells at campaign boundaries
cailmdaley Sep 26, 2026
a846e08
test(workflow): the DAG harness follows the parse-time guards
cailmdaley Sep 26, 2026
8209e26
Merge origin/develop into feat/make-cat-fixed-epoch-slots
cailmdaley Sep 28, 2026
fedbf9c
Merge origin/develop into feat/persist-exp-products
cailmdaley Sep 28, 2026
a2b37a8
Merge origin/feat/make-cat-fixed-epoch-slots into feat/persist-exp-pr…
cailmdaley Sep 28, 2026
2d833fd
test: seams and run_config follow the develop merge
cailmdaley Sep 28, 2026
aa458b2
test(workflow): snakemake-driven checks run in their own CI step
cailmdaley Sep 28, 2026
5891cf6
Merge origin/develop into feat/persist-exp-products
cailmdaley Sep 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions .github/workflows/deploy-image.yml
Original file line number Diff line number Diff line change
Expand Up @@ -115,6 +115,17 @@ jobs:
IMAGE=$(echo "${{ steps.meta.outputs.tags }}" | head -n1)
docker run --rm -e HYPOTHESIS_PROFILE=ci -e SHAPEPIPE_ON_CANDIDE=0 "$IMAGE" pytest -rX

# tests/workflow drives the Snakefile through snakemake's API. Snakemake
# is a host tool and stays out of the image (it wraps each job in the
# container), so the suite above leaves that directory out
# (the root conftest.py); here it is installed into this throwaway container
# at the host pin range (workflow/README.md) and the directory runs alone.
- name: Test — workflow DAG (snakemake)
run: |
IMAGE=$(echo "${{ steps.meta.outputs.tags }}" | head -n1)
docker run --rm -e SHAPEPIPE_ON_CANDIDE=0 "$IMAGE" bash -c \
"uv pip install 'snakemake>=9,<10' && pytest -rX --no-cov tests/workflow"

# ----------------------------------------------------------------
# Publish (push events only — never on pull_request, incl. forks).
# Fires on any branch; the image is tagged with the branch name.
Expand Down
5 changes: 2 additions & 3 deletions astra.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -572,9 +572,8 @@ analyses:
No detection image, no flag image, and the default_noimaflags.param
column list: tiles have no instrument flag image and no detection
coadd exists, so detection sees every tile pixel and the tile
catalogue carries no IMAFLAGS_ISO. [LINT]
final_cat.param, read by the post-processing merge, requests
IMAFLAGS_ISO, which the tile chain never produces (issue #912).
catalogue carries no IMAFLAGS_ISO; final_cat.param, the merge's
exact allow-list, does not request it.
Values:
SEXTRACTOR_RUNNER.DETECTION_IMAGE = False;
SEXTRACTOR_RUNNER.FLAG_IMAGE = False.
Expand Down
15 changes: 13 additions & 2 deletions conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,16 @@
settings.load_profile(os.environ.get("HYPOTHESIS_PROFILE", "ci"))


# ``tests/workflow/`` drives the Snakefile through snakemake's API. Snakemake is
# a host tool, not part of the image (it wraps each job in the container), so
# where it is absent the directory is left out of collection; CI runs it in its
# own step after installing snakemake (deploy-image.yml).
try:
import snakemake # noqa: F401
except ModuleNotFoundError:
collect_ignore = ["tests/workflow"]


# --------------------------------------------------------------------------- #
# Candide detection
# --------------------------------------------------------------------------- #
Expand All @@ -36,15 +46,16 @@
# host this suite is most often driven from is ``c03``. We match the candide
# node-name families rather than a fixed list so new nodes are covered, and
# allow an explicit override for CI or odd hostnames.
_CANDIDE_HOST_RE = re.compile(r"^(c\d|n\d{2})", re.IGNORECASE)
_CANDIDE_HOST_RE = re.compile(r"^(c\d{2}|n\d{2})$", re.IGNORECASE)


def on_candide():
"""Return True when running on a candide node.

The check is, in order: an explicit ``SHAPEPIPE_ON_CANDIDE`` override
(``1``/``0``), then the hostname against the candide node-name families
(``c0x`` login, ``nXX`` compute). Cheap, import-safe, no cluster calls.
(``c0x`` login, ``nXX`` compute; whole bare hostname, so ``c6.nibi.sharcnet``
does not match). Cheap, import-safe, no cluster calls.
"""
override = os.environ.get("SHAPEPIPE_ON_CANDIDE")
if override is not None:
Expand Down
102 changes: 72 additions & 30 deletions scripts/python/create_final_cat.py
Original file line number Diff line number Diff line change
Expand Up @@ -58,22 +58,32 @@ def params_from_run_config(params, defaults):
raise ValueError(f"run config {params['run_config']} sets neither "
"outputs.products_dir nor outputs.run_dir")

# The same derivation as the workflow's merge_final_cats rule: the patch
# dir is the branch dir holding product/tiles, and -i is its parent.
# The group and file are named for `run:`, as the workflow's
# final_cat_merge names them. -P is also the directory the -I walk
# matches under -i, so this needs the run template's layout,
# products_dir = <root>/<run>/product, and -i is <root>.
patch = cfg["run"]
patch_dir = os.path.dirname(os.path.normpath(products))
patch = os.path.basename(patch_dir)
if os.path.basename(patch_dir) != patch:
raise ValueError(
f"run config {params['run_config']}: products_dir {products} is "
f"not <root>/{patch}/product, so -c cannot locate the tiles; "
"pass -i and -P explicitly")
derived = {
"image_sims": True,
"input_root_dir": os.path.dirname(patch_dir),
"patch": patch,
"merged_cat_path": os.path.join(products, f"final_cat_{patch}.hdf5"),
"output_summary": os.path.join(products, "n_tiles_final.txt"),
"param_path": os.path.join(repo, "workflow", "config",
"cfis_image_sims", "final_cat.param"),
}
for key, value in derived.items():
if params.get(key) == defaults.get(key):
params[key] = value
# The tile count is the file's n_tiles attribute; the text summary is
# written only when -o asks for it.
if params.get("output_summary") == defaults.get("output_summary"):
params["output_summary"] = None
return params


Expand Down Expand Up @@ -152,7 +162,8 @@ def params_default():
"param_path": "parameter file path, if not given use all columns, default={}",
"patch": "patch number (data) or grid subdir (image_sims), default={}",
"list_only": "print list of patches and IDs only, default={}",
"output_summary": "output file for numbre of tiles, default={}",
"output_summary": "output file for number of tiles (with -c, written"
" only if given), default={}",
"ID": "ID for single-ID operation, default={}",
"single_op": "single ID operation, allowed are 'check', 'add', 'remove'; default={}",
"image_sims": "image simulations mode (different dir layout and run prefix), default={}",
Expand Down Expand Up @@ -205,13 +216,17 @@ def read_param_file(path, verbose=False):
print("No parameters read", end="")
print(" into merged catalogue")

param_list_unique = list(set(param_list))

# Ordered dedup. list(set(...)) reordered the columns by the process's
# string hash seed, so two runs of this tool over the same inputs produced
# files whose datasets differed in column ORDER — which is part of a
# structured dtype, and therefore part of the file.
param_list_unique = list(dict.fromkeys(param_list))

if verbose:
n = len(param_list) - len(param_list_unique)
if n > 1:
print("Removed {n} duplicate entries")
if n > 0:
print(f"Removed {n} duplicate entries")

return param_list_unique


Expand Down Expand Up @@ -317,8 +332,9 @@ def print_list(params):
if verbose:
print(f"Total: {n_tiles} tiles")

with open(params["output_summary"], "w") as f_out:
print(n_tiles, file=f_out)
if params["output_summary"]:
with open(params["output_summary"], "w") as f_out:
print(n_tiles, file=f_out)

# Write n_tiles to HDF5 file header
with h5py.File(params["merged_cat_path"], "a") as hdf5_file:
Expand Down Expand Up @@ -359,8 +375,15 @@ def get_patch_group(hdf5_file, patch, verbose=False):


def read_data(fits_file, params):
"""Read Data.

"""Read the parameter list's columns out of one catalogue.

@sc [label:schema] read-data-raises-on-missing-column
A requested column the catalogue lacks raises `KeyError` naming it; it is
never skipped or filled. `copy_data` keeps only columns present in the
source, so this raise is the one place a missing name stops a merge, and
without it a tile short a per-epoch slot would land in the merged file
silently narrower, with that slot's exposure identity gone. Enforced by
tests/unit/test_final_cat_merge_invariants.py.
"""
with fits.open(fits_file) as hdu_list:
try:
Expand All @@ -373,16 +396,20 @@ def read_data(fits_file, params):
if params["param_list"] is None:
params["param_list"] = [col for col in data.keys()]

try:
extracted_data = {col: data[col] for col in params["param_list"]}
dtype = data.dtype
except:
print(f"Error for ID {id}, path {fits_file}")
for col in params["param_list"]:
if col not in data:
print(col, end=" ")
print()
continue
# RAISE, do not print and fall through. The bare `except:` this replaces
# left extracted_data and dtype unbound, so the caller's own error was an
# UnboundLocalError from the return statement below, naming neither the
# file nor the column that was actually missing.
present = set(data.dtype.names or ())
missing = [col for col in params["param_list"] if col not in present]
if missing:
raise KeyError(
f"{fits_file}: missing {len(missing)} of the "
f"{len(params['param_list'])} requested column(s): "
f"{' '.join(missing)}"
)
extracted_data = {col: data[col] for col in params["param_list"]}
dtype = data.dtype

return extracted_data, dtype

Expand All @@ -391,16 +418,29 @@ def copy_data(param_list, extracted_data, dtype):
"""Copy Data.

"""
# THE REQUESTED COLUMNS ONLY, IN THE PARAMETER FILE'S ORDER. Two things
# are being fixed here and they are easy to conflate. Allocating with the
# source's full dtype and filling only the requested columns left every
# other column as uninitialised memory — meaningless values, and different
# bytes on every run over the same inputs. And ordering the result by the
# SOURCE catalogue's columns made the output dtype a property of the
# catalogue rather than of the parameter file: two tiles written by
# different ShapePipe versions, whose catalogues order or extend their
# columns differently, then landed in one merged file with two different
# structured dtypes, which np.concatenate refuses. The parameter file is
# the schema; it says which columns AND in what order.
wanted = set(dtype.names or ())
columns = [col for col in param_list if col in wanted]
subset = np.dtype([(col, dtype[col]) for col in columns])

# Initialize new data structure
structured_data = np.empty(
len(extracted_data[param_list[0]]),
dtype=dtype,
dtype=subset,
)

# Loop over parameters
for col in param_list:
if not col in extracted_data:
print(f"Column {col} not in file with ID {id}")
for col in columns:
structured_data[col] = extracted_data[col]

#if isinstance(extracted_data[col][0], (np.ndarray, tuple, list)):
Expand Down Expand Up @@ -547,12 +587,14 @@ def process(params):

structured_data = copy_data(params["param_list"], extracted_data, dtype)

# Create a new dataset
# Create a new dataset. dtype comes from the array copy_data
# built, not from the source catalogue: they differ now that
# copy_data allocates the requested columns alone.
try:
patch_group.create_dataset(
str(id),
data=structured_data,
dtype=dtype,
dtype=structured_data.dtype,
)
except:
print(f"Error for {id}: Could not create dataset in group {patch}")
Expand Down
Loading
Loading