Exposure-level HealSparse maps: instrument defects and exposure count - #887
Open
cailmdaley wants to merge 37 commits into
Open
cailmdaley wants to merge 37 commits into
cailmdaley wants to merge 37 commits into
Conversation
PR #847 removed ShapePipe's mask generation and left the sky-fixed UNIONS masks to be queried per object by make_cat via MASK_EXT_PATHS. The committed config never set it, so no campaign wrote a mask column and the merged shear catalogue carried no mask information at all. It does now. The DR6 Aug-2026 ugriz bit ladder is staged group-readably at /project/def-mjhudson/unions-wl/masks/dr6-2026-08 (11 boolean healsparse maps, nside 131072, True = masked), reached as a third input root SP_INPUT_MASKS beside tiles and exposures. config_tile_Mc.ini names all 11 in MASK_EXT_PATHS and carries the bit table; final_cat.param carries the matching MASK_n1 .. MASK_n2048, replacing the IMAFLAGS_ISO note. Two caveats live with the config, because a cut written without them is wrong: the August regeneration swapped the content of bits 0 and 1, so n1|n2 is only safe combined; and n2048 is 1 where there is no Pan-STARRS z2 data, so an OR over every column masks the whole survey. Nothing cuts on these here. PSF-star selection is untouched: MASK_PATHS in config_exp_psfex.ini stays commented out. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…878) The exposures' instrument flag images — bad columns, saturated pixels, bleed trails — are the one masking input this campaign has that never leaves the pixel domain: exp_split splits them per CCD, SExtractor reads them as IMAFLAGS_ISO, and that is the end of it. The survey footprint is built from the CCD corner WCS in the headers, so it cannot subtract them and would silently include defective pixels. The lost area is percent-level, but it carries exactly the thin small-scale geometry an accurate window function needs, and only ShapePipe ever opens these files. Two rules, both in exposure.smk. exp_defect_map rasterizes ONE exposure's flag splits into a boolean healsparse fragment on the persistent root — nside_sparse 131072 over nside_coverage 128, bit-packed, True = masked, the same form as every map in the UNIONS ladder, so it drops into that ladder without a resolution change. It hangs off exp_split rather than exp_psf, so re-rasterizing at a different fidelity never touches the four-hour PSF chain, and clean_exposure takes its manifest as an input, so reclamation cannot overtake the copy. The WCS comes from the image split beside each flag file (headers only, never pixels): the flag mosaic's HDUs carry no WCS at all — checked on a real exposure — and headers-<num>.npy is a pickled array of astropy WCS instances a second consumer should not inherit. Rasterization samples each flagged CCD pixel across its full extent, corners included, and is CONSERVATIVE: a healpix pixel is masked if any part of the flagged region touches it, because the centre-based alternative erases a one-pixel bad column, which is the geometry #878 exists to keep. `oversample` rides on params with the measured convergence table behind its default of 3. defect_map_merge unions the campaign's fragments into <products_dir>/defect_map_<campaign>.hsp. It RECONCILES like final_cat_merge — a new exposure is OR-ed in on the spot, an exposure that left the campaign or a fragment that changed forces a rebuild (a union cannot be un-OR-ed), a no-op leaves the file untouched — against a sidecar that records which exposures are already in it, declared as a second output because the two are only meaningful together. Memory is flat in the exposure count: each fragment is read, reduced to its pixel ids and dropped, so the job holds one accumulator (the campaign's footprint) and one 2 MB fragment whether the campaign is 127 exposures or 20k. The map is a campaign PRODUCT, not an input: nothing here reads it back. Measured on 2079612p inside the container, 40 CCDs, 4.4% of pixels flagged: 37 s and 0.62 GB peak RSS at oversample 3, a 2.0 MB fragment of 359858 healpix pixels over 13 coverage pixels (95 s / 1.29 GB / 360840 pixels at oversample 5, so the default is within 0.3% of it at a third of the cost). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…he ladder config_tile_Mc.ini does NOT gain a MASK_EXT_PATHS entry, and the header now says why: every path in that ladder exists before the run starts, and the defect map exists only after it, so wiring it would name a file the first run of a fresh campaign cannot have and every tile would fail on a missing map. What the header gains instead is the recipe — where the map lands, copying or symlinking it into the inputs.masks root, the one MASK_EXT_PATHS entry and the matching final_cat.param line — plus the two things to know before doing it: the map is the producing campaign's own footprint, so a catalogue built from a different tile list reads False (unknown, not clean) outside it, and the rasterization is conservative, widening a one-pixel bad column to the 1.61" healpix resolution. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Three defects in the two scripts, found reviewing them. A CCD WHOSE FLAG SPLIT IS MISSING was silently dropped. ccd_files globbed flag-*.fits and rasterized whatever came back, so an incompletely materialised split dir — an age-based scratch purge deleting files one at a time, a store copied or restored half way, a truncated rsync — produced a fragment covering half the exposure, written with "status": "complete", and nothing downstream able to notice. Reproduced on a fixture: remove one flag split of two and the script exited 0 with n_ccds 1. The guard it had was on the opposite, unreachable direction (a flag without its image), which split_exp cannot produce. The loop is now driven by the EXPECTED CCD count, which comes in on params from config_exp_Sp.ini's own N_HDU, and a short split dir is a hard error. A FULLY FLAGGED CCD blew the rule's memory budget. A MegaCam exposure routinely carries a dead or saturated chip: 2048 x 4612 = 9.4M flagged pixels, 85M samples at oversample 3, held in one shot several GB of coordinate arrays and astropy temporaries against a 2 GB request — so the job OOMed on all three attempts, the exposure never got a fragment, and because clean_exposure waits on that manifest it could never be reclaimed either. The 0.62 GB measured on a 4.1%-flagged exposure said nothing about that case. The flagged-pixel loop is now batched at CHUNK source pixels and np.unique-accumulated, so peak RSS is set by the batch and not by how bad the CCD is. Measured on exactly that case, a fully flagged chip against 2079612p's real WCS: 0.74 GB and 18.3 s at 500k, 1.13 GB and 18.9 s at 1M — time is flat in the batch size and memory is linear in it, so 500k is free. The real exposure is unchanged bit for bit (359858 healpix pixels, 2.0 MB) at 34 s. THE SIDECAR WENT STALE ON A NO-OP. main() returned on an empty plan without touching it, but two of its fields describe the CAMPAIGN and not the map: add tiles whose exposures all lack fragments and campaign_exposures and exposures_without_fragment are wrong on disk while the union is untouched and correct. That is exactly the record the docstring promises is "recorded rather than merely printed". The record is now built apart from writing the map and rewritten alone when it differs; write_stable keeps the map's mtime where it is. Also: the stamp is IMPORTED from hdf5_reconcile rather than re-derived, so the two halves of a campaign's reconciliation answer "did this source change" the same way; and the docstring now says plainly, in hdf5_reconcile's own terms, that the map's CONTENT is a function of the input set and its BYTES are not (measured: 1,586,880 B rebuilt vs 1,589,760 B appended, identical valid_pixels). tests/unit/test_defect_map_reconcile.py pins the four reconcile branches, the sidecar refresh and the nside mismatch — fragments made with make_empty, since the merge half never opens a FITS image or a WCS. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…or 7 GB Four defects in the DAG wiring and the sizing. A RERUN TRIGGER RE-AVALANCHED THE CAMPAIGN. defect_map_inputs() named the manifest of any exposure that had one, reclaimed or not, but exp_defect_map's input is exp_split's manifest and that went with the scratch store. So any trigger on the rule — and params is live, carrying the script hash, the nside and the oversampling this file advertises as a minute's work to change — would schedule exp_get_images and exp_split from VOS for every reclaimed exposure: a four-hour chain each, campaign-wide, on a one-character edit. It now mirrors star_cat_inputs(): a live exposure through its manifest, a reclaimed one through its FRAGMENT, which is not a declared output of any rule and is therefore a true leaf. That also gives prod_exp_fragment() its only caller, so the fragment's location is spelled once here and once in merge_defect_map.fragment_path() rather than twice with nothing joining them. Verified on a two-exposure compute fixture: with 2079612 reclaimed and holding a fragment, requesting the map builds 4 jobs (one exp chain, the live exposure's) where both-live builds 7. THE FIRST MERGE OF A LARGE CAMPAIGN ASKED FOR ~52 GB. Before there is a sidecar the coverage-pixel count was estimated as 13 per exposure capped at the FULL SKY, and 12 * 128^2 = 196608 is reached at ~15k exposures — so a DR6-scale campaign estimated the whole sky, 25.8 GB of accumulator and ~52 GB of request (~104 GB on the retry) for a job this file's own comment measures at ~3 GB. Exposures overlap almost completely; the prior has to be a FOOTPRINT, so it is now capped at the ~23k coverage pixels the UNIONS ugriz maps measure over the same sky. A small campaign is still sized on its own exposures, where the per-exposure figure is the honest one: the two-exposure fixture asks for 806 MB. THE MERGE'S RUNTIME REACHED 11.4 H WITH NO RETRIES AND NO PARTIAL STATE, so its mem_mb/runtime attempt scaling was dead code and one walltime overrun eleven hours in lost the whole thing. retries: 1 is declared, the formula is corrected to the ~1 s per exposure actually measured, and it is capped below the walltime Alliance policy lets a job run without checkpointing — with the comment saying that a campaign needing longer needs a resumable accumulator, not a bigger number. SILENCE WAS THE ANSWER when no exposure can contribute. A campaign whose exposures were all reclaimed by a workflow predating this rule builds no defect job at all, and the operator saw nothing rather than "no map is possible here". It now says so once at parse time, like the missing-index note. Also: a comment pointed at DEFECT_COV_BYTES, which does not exist. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…n one The README and config.yaml carried the pre-batching numbers and an inode estimate a third of the real cost. The DR6 inode figure was ~20k against a ~1M group quota; it is ~60k, because each exposure gets a defect/ DIRECTORY, the fragment in it and a manifest under manifests/ — three inodes, not one. The fragment was also described as sitting "beside its PSF tar", which is not the layout; the path is spelled out. The timing is the re-measured 34 s (2079612p, 40 CCDs, 4.1% of pixels flagged, oversample 3, inside the campaign container), and both files now carry the worst-case memory bound the batching exists for — a fully flagged chip at 0.74 GB — rather than only the average exposure's 0.62 GB. The README also records what the reconcile does and does not promise (content, not bytes) and points at the new unit test. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
… run clean_exposure's exp_defect_map edge was unconditional, which reopened the avalanche defect_map_inputs() is written to avoid. An exposure whose scratch store went to the 60-day /scratch purge, or to a `clean: false` run, has NO tombstone — exp_store_reclaimed()'s docstring names that case — so clean_targets() still asks for its tombstone, the missing exp_defect_map manifest is scheduled, and its own input (exp_split's manifest) went with the store: exp_get_images and exp_split from VOS, four hours per exposure, campaign-wide, on the first run of this branch. Reproduced on a fixture campaign (tile 000.000, exposures 2079612 live and 2079613 purged-without-tombstone, tile_vignets present): a dry run of 2079613's tombstone scheduled 13 jobs including exp_get_images and exp_split; with the edge made conditional it schedules one, clean_exposure. The live exposure still gets exp_defect_map ordered before its clean, and a reclaimed exposure that already has a fragment depends on the fragment, a leaf that builds nothing. ccd_files' short-split error now says what to do about the store it pins: a damaged split dir is still a hard error rather than a half-footprint fragment marked complete, but the operator is told that deleting the scratch exp_split manifest is what rebuilds the split and unblocks the clean. defect_map_merge's mem_mb also goes through capped_mem(), as star_cat_merge and final_cat_merge already do: defect_map_cov_bytes() scales as nside^2 and defect_map.nside is an advertised knob, so one ladder change turns the request into a job SLURM never schedules and snakemake never diagnoses. Verified on the fixture with max_mem_mb=700: "defect_map_merge: sized at 806 MB, capped to max_mem_mb=700" at parse time, mem_mb=700 on the job. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
A rebuild set pixels into a fresh accumulator fragment by fragment, so every fragment that touched an unseen coverage pixel made healsparse grow — and copy — the sparse array. union_coverage() now reads each fragment's COVERAGE TABLE first (HealSparseCoverage.read, kilobytes, never the sparse array) and seeds make_empty(cov_pixels=...) with the union, so the final array is allocated once. MEASURED rather than assumed, on synthetic fragments at the campaign's own resolution (nside 131072 / coverage 128), 400 fragments over 5131 coverage pixels, a 656 MB accumulator: accumulation 8.4 s unseeded, 7.1 s seeded, with a 0.7 s coverage pre-pass. The reallocations are ~16% of the accumulation, not the dominant term a naive reading predicts — healsparse grows the array in blocks — so the review's tens-of-TB estimate does not hold, but the pre-pass is linear where the copying is not, and the win grows with the accumulator. Only the rebuild path seeds: an append starts from the map on disk, whose coverage is already most of the footprint. A fragment at the wrong coverage resolution abandons the seed so accumulate()'s named error still stands; the existing nside-mismatch test pins that. tests/unit/test_defect_map_reconcile.py passes (7/7, run in the container), and an end-to-end fixture merge reproduces rebuild == OR of fragments, no-op untouched, append, and removal-rebuild. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
This was referenced Sep 10, 2026
PR #847 removed ShapePipe's mask generation and left the sky-fixed UNIONS masks to be queried per object by make_cat via MASK_EXT_PATHS. The committed config never set it, so no campaign wrote a mask column and the merged shear catalogue carried no mask information at all. It does now. The DR6 Aug-2026 ugriz bit ladder is staged group-readably at /project/def-mjhudson/unions-wl/masks/dr6-2026-08 (11 boolean healsparse maps, nside 131072, True = masked), reached as a third input root SP_INPUT_MASKS beside tiles and exposures. config_tile_Mc.ini names all 11 in MASK_EXT_PATHS and carries the bit table; final_cat.param carries the matching MASK_n1 .. MASK_n2048, replacing the IMAFLAGS_ISO note. Two caveats live with the config, because a cut written without them is wrong: the August regeneration swapped the content of bits 0 and 1, so n1|n2 is only safe combined; and n2048 is 1 where there is no Pan-STARRS z2 data, so an OR over every column masks the whole survey. Nothing cuts on these here. PSF-star selection is untouched: MASK_PATHS in config_exp_psfex.ini stays commented out. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…n type
FITSCatalogue._get_fits_col_type checked `type(col_data[0]) is bool`, which
only matches Python's builtin bool. A numpy bool array's first element is
np.bool_, so it fell through to the float branch and every boolean column
was silently written as float64 {0.0, 1.0}. This hit make_cat's MASK_n*
columns, queried from boolean healsparse maps via mask_query.query_map,
even though the maps, the docs (config_tile_Mc.ini, final_cat.param,
workflow/README.md), and query_map's own docstring all promise a boolean
value (including the off-map sentinel False). Recognizing np.bool_ writes
the FITS 'L' (logical) format instead, matching the promised dtype; no
other code path depended on the float64 widening.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…rument-defect-map Conflict resolution, per file: - workflow/README.md: both bullets kept, the defect-map bullet first; the external-masks bullet takes the incoming "(data runs)" / machines-table wording. - workflow/Snakefile: prologue docstring and export list take the incoming side (SP_RETRIEVE; SP_INPUT_MASKS exported for input_type: data only). - workflow/config.yaml: top-level inputs: gives way to input_type: plus the machines: table (masks live there); the defect_map: block is kept whole ahead of the incoming, shorter clean: comment. - workflow/config/cfis/config_tile_Mc.ini: #887's defect-map header note sits after the MASK_EXT_PATHS commentary and before the block; its path recipe names <campaign> as `run:` in the run config. - workflow/rules/exposure.smk: clean_exposure waits on exp_persist only under PERSISTS_PSF and on the exp_defect_map manifest (or fragment) as #887 wrote it; star_cat_merge keeps --snapshot-json and defect_map_merge follows it. Image simulations are gated off the defect map in the next commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011xYGyn53XoyL83uuKPN7RQ
The simulations ship simu_flag-<exp>.fits per exposure, split and read by SExtractor like the survey's, but every pixel is zero: measured on 2079614 (1m2z_1, 1m2z_grid_2) and 2079616 (1m2z_3), 377,815,040 pixels each, none set. Rasterizing them would cost a job per exposure for an empty union. MAPS_DEFECTS = input_type == data sits beside PERSISTS_PSF and acts in defect_map_exposures(), which every defect-map target, input and sizing reads through; clean_exposure's defect edge checks it the same way its persist edge checks PERSISTS_PSF. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011xYGyn53XoyL83uuKPN7RQ
PR #847 removed ShapePipe's mask generation and left the sky-fixed UNIONS masks to be queried per object by make_cat via MASK_EXT_PATHS. The committed config never set it, so no campaign wrote a mask column and the merged shear catalogue carried no mask information at all. It does now. The DR6 Aug-2026 ugriz bit ladder is staged group-readably at /project/def-mjhudson/unions-wl/masks/dr6-2026-08 (11 boolean healsparse maps, nside 131072, True = masked), reached as a third input root SP_INPUT_MASKS beside tiles and exposures. config_tile_Mc.ini names all 11 in MASK_EXT_PATHS and carries the bit table; final_cat.param carries the matching MASK_n1 .. MASK_n2048, replacing the IMAFLAGS_ISO note. Two caveats live with the config, because a cut written without them is wrong: the August regeneration swapped the content of bits 0 and 1, so n1|n2 is only safe combined; and n2048 is 1 where there is no Pan-STARRS z2 data, so an OR over every column masks the whole survey. Nothing cuts on these here. PSF-star selection is untouched: MASK_PATHS in config_exp_psfex.ini stays commented out. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…n type
FITSCatalogue._get_fits_col_type checked `type(col_data[0]) is bool`, which
only matches Python's builtin bool. A numpy bool array's first element is
np.bool_, so it fell through to the float branch and every boolean column
was silently written as float64 {0.0, 1.0}. This hit make_cat's MASK_n*
columns, queried from boolean healsparse maps via mask_query.query_map,
even though the maps, the docs (config_tile_Mc.ini, final_cat.param,
workflow/README.md), and query_map's own docstring all promise a boolean
value (including the off-map sentinel False). Recognizing np.bool_ writes
the FITS 'L' (logical) format instead, matching the promised dtype; no
other code path depended on the float64 widening.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…t-map-v2 # Conflicts: # workflow/README.md # workflow/config/cfis/config_tile_Mc.ini
…fect map The flag images are the one masking input that never leaves the pixel domain, and a footprint built from CCD-corner WCS silently includes the defective pixels. #887 rasterizes them into a healsparse fragment per exposure and unions the campaign's fragments, so the choice of how a flagged CCD pixel becomes a masked healpix pixel is now a scientific decision. It sits beside epoch_flag_source because it is a second use of the same flag source, not a different source. Centre sampling and no map at all are recorded as rejected, with the measured oversample convergence in the rationale. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011xYGyn53XoyL83uuKPN7RQ
A flagged CCD pixel that lands in an unmasked healpix pixel is a hole in the footprint nothing downstream can see, so the rasterizer's contract is containment. The test builds a small flag image (a one-pixel bad column, a saturated blob, edge pixels and a lattice of hot pixels spaced wider than a healpix pixel) on a rotated 0.187"/pixel TAN WCS, rasterizes it with rasterize_ccd at the configured nside and oversample, and checks that the centre and four corners of every flagged pixel, taken through astropy's 0-based convention, fall in masked pixels. A clean image yields nothing. Centre-only sampling, both off-by-one pixel conventions, dropping the last row and column, dropping the last batch and swapping the axes each fail it. The hot-pixel lattice is what catches centre sampling: without it every straddled corner borrows a neighbour's centre. The contracts sit on rasterize_ccd and on merge_defect_map's reconcile_plan, linked to the defect_map_from_flags decision. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011xYGyn53XoyL83uuKPN7RQ
PR #847 removed ShapePipe's mask generation and left the sky-fixed UNIONS masks to be queried per object by make_cat via MASK_EXT_PATHS. The committed config never set it, so no campaign wrote a mask column and the merged shear catalogue carried no mask information at all. It does now. The DR6 Aug-2026 ugriz bit ladder is staged group-readably at /project/def-mjhudson/unions-wl/masks/dr6-2026-08 (11 boolean healsparse maps, nside 131072, True = masked), reached as a third input root SP_INPUT_MASKS beside tiles and exposures. config_tile_Mc.ini names all 11 in MASK_EXT_PATHS and carries the bit table; final_cat.param carries the matching MASK_n1 .. MASK_n2048, replacing the IMAFLAGS_ISO note. Two caveats live with the config, because a cut written without them is wrong: the August regeneration swapped the content of bits 0 and 1, so n1|n2 is only safe combined; and n2048 is 1 where there is no Pan-STARRS z2 data, so an OR over every column masks the whole survey. Nothing cuts on these here. PSF-star selection is untouched: MASK_PATHS in config_exp_psfex.ini stays commented out. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…n type
FITSCatalogue._get_fits_col_type checked `type(col_data[0]) is bool`, which
only matches Python's builtin bool. A numpy bool array's first element is
np.bool_, so it fell through to the float branch and every boolean column
was silently written as float64 {0.0, 1.0}. This hit make_cat's MASK_n*
columns, queried from boolean healsparse maps via mask_query.query_map,
even though the maps, the docs (config_tile_Mc.ini, final_cat.param,
workflow/README.md), and query_map's own docstring all promise a boolean
value (including the off-map sentinel False). Recognizing np.bool_ writes
the FITS 'L' (logical) format instead, matching the promised dtype; no
other code path depended on the float64 widening.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
…ask bits make_cat writes eleven MASK_n<bit> columns and cuts nothing, so which of them form the default cut is a consumer decision the record has to carry. n1|n2|n4|n8|n64|n1024 reproduces the published 2D-cosmic-shear mask pixel-exactly on whole interior granules of the Feb-2025 r ladder, so the catalogue's footprint matches the 2D analysis'. The defects-only OR misses the r-footprint bit (9% of the masked area); OR-ing every column masks the regions with no u or Pan-STARRS z2 data. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…e six-bit OR reproduces mask_r The contract mask-ext-ladder-columns on save_mask_ext_data states the coupling: column names are the ladder's labels verbatim, final_cat.param must list them line for line or every merge fails, and the default cut is decision mask_default_cut. The unit test parses the committed ini with parse_mask_ext_paths and compares it to final_cat.param, so a renamed label fails in CI. The candide test checks the decision's evidence: the OR of the six r-ladder maps equals mask_r pixel-exactly on three whole interior granules, and skips when the maps are unreadable. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…-map-v2 The defect_map output and defect_map_from_flags decision move into the rewritten record's masking sub-analysis (the rewrite has no epoch_flag_source; the flag images are its exposure_flags input).
cailmdaley
force-pushed
the
feat/wire-external-masks
branch
from
September 26, 2026 01:38
da1725b to
03fd684
Compare
…e's record The fragment manifest recorded only per-CCD healpix counts, so a re-rasterization that moved a defect while keeping the counts rewrote the fragment but left the manifest byte-identical and untouched. The manifest is the DAG edge defect_map_merge waits on, so the change never reached it, and even when the merge ran its size/mtime stamp could miss a same-size rewrite. exp_defect_map now records the fragment's SHA-256 in its manifest. The merge reads that digest back per exposure (hashing the fragment only when a manifest lacks it), records it in the sidecar, and treats a digest change as a changed fragment, which forces a rebuild. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… no-op Rerunning the merge with a new --nside over unchanged fragments produced an empty plan: the map stayed at the old resolution and the no-op path rewrote the sidecar to claim the new one, so the sidecar disagreed with the map beside it. reconcile_plan now reads the map's nside_sparse and nside_coverage from its coverage table and plans a rebuild when they differ from the requested ones. The rebuild reads the fragments, and accumulate refuses any at another resolution before either file is written. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The map was renamed into place before the sidecar was written, so a kill between the two left a new map under the old sidecar. Returning to the campaign that old sidecar described then produced an empty plan over a map still carrying the other campaign's pixels. Both files are now written to temporaries before either rename. Each merge stamps a generation id into the map's healsparse metadata (DMAPGEN) and into the sidecar. The id is a digest of the resolution and every exposure's fragment digest, so it is a function of the input set and an unchanged campaign rewrites nothing. reconcile_plan reads the map's primary header and rebuilds when the two ids disagree, which covers the one window two renames cannot close. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
test_no_op_still_refreshes_a_stale_sidecar performed the refresh itself (build_record + write_sidecar) and stayed green with main() broken. The rest of the file shared the pattern through a helper that re-implemented main's plan-then-apply, so none of them would notice a main() that always rebuilt, never refreshed, or never returned early on a no-op. Every case now runs main() over a one-tile index and asserts on what it prints and writes. The append test spies on accumulate, so it measures that only the new fragment is read, as its name claims. The stamp-based changed-fragment test is folded into the digest one, whose change also drops a pixel so only a rebuild gives the right union. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cailmdaley
added a commit
that referenced
this pull request
Sep 26, 2026
… manifest clean_exposure named the exp_persist manifest for every exposure. For a reclaimed one, a persist_exp: edit changes exp_persist's params, so the manifest reruns, and it sits behind exp_psf's manifest, which went with the store: snakemake scheduled exp_get_images, exp_split, exp_psf, exp_persist and clean_exposure again for every reclaimed exposure. The edge now takes the leaf treatment the defect edge uses on #887: a live store is asked for the manifest, a reclaimed one for its tar if it exists (no rule's output, so a leaf), else nothing. One lambda per edge is kept. CONTRACTS' persist edge says so. Test: the review's retention dry run (a completed, reclaimed two-exposure fixture with recorded metadata, then persist_exp: [psf_model, star_stats]) scheduled 13 jobs including the full exposure chain for both exposures; it now schedules final_cat_merge, star_cat_merge and all (3). Harness: /automnt/n17data/cdaley/scratch/smoke-879-mine/retention/test_retention.py. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cailmdaley
force-pushed
the
feat/wire-external-masks
branch
from
September 26, 2026 09:18
03fd684 to
94e2dc1
Compare
…strument-defect-map-v2 clean_exposure keeps one lambda per edge: #879's persist edge (tar as a leaf once reclaimed) and the defect edge.
… fixture carries a tiles table exp_defect_map and defect_map_merge join the data rule set and the defect manifest is one of clean_exposure's edges; build_index reads the tiles row for readiness, so the one-tile index fixture needs one. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
cailmdaley
added a commit
that referenced
this pull request
Sep 26, 2026
…when the Snakefile has one Beside #887 clean_exposure carries a third edge (MAPS_DEFECTS); the test binds it when present and expects it under a fitted PSF. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…t-map-v2 # Conflicts: # tests/workflow/params_pin.json
…nt-defect-map Resolved by taking the updated #886 tree and re-applying this branch's own changes (dc11dab..5f7332d): - astra.yaml: defect_map_from_flags drops its Anchor sentence for the #875 format: @sc tags on exp_defect_map, defect_map_merge and config.yaml's defect_map block (Values pin nside, nside_coverage, oversample), the two script contracts re-cite masking.defect_map_from_flags, and the fragment-contains-flags test carries its decision marker. The masked_measurement_inputs output keeps #886's mask_default_cut. - exposure.smk docstring: the defect-map export paragraph, with develop's MASK_EXT description of mask_query (FLAG_EXT is gone). - config_tile_Mc.ini: the defect-map note sits above the MASK_EXT_PATHS tag. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The merged map and its sidecar move to defect_map/defect_map_<run>.{hsp,json},
the per-product subdirectory coverage uses. campaign-name-is-run lists the
map, and test_campaign_lineage pins both paths.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DksuyF9YrZAwHQeAMXsxp4
cailmdaley
added a commit
that referenced
this pull request
Sep 28, 2026
# Conflicts: # tests/unit/test_campaign_lineage.py # tests/workflow/params_pin.json # workflow/CONTRACTS # workflow/README.md
The download_headers -> extract_field_corners -> build_coverage_map chain (VOSpace header download, exp_ra_dec.txt, the CoverageMapBuilder CLI) and build_and_plot_coverage_maps.sh go. What the exposure-count map needs stays as plain functions: ccd_footprint (_image_shape, _ccd_corners) and coverage_map_builder (build_map, unwrap_ra and the pole guard, check_nside, median_filter). plot_coverage_map stays. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BeQNTBcjTqu6TPktxysovQ
Both exposure-derived HealSparse maps now share one resolution block, config.yaml's exposure_maps: (nside 131072 / nside_coverage 128, the mask ladder's), with defect: and nexp: sub-blocks. - exp_footprint (localrule, after exp_persist): the sky corners of every CCD with a valid PSF model, read off exp_persist.json's validation_psf members and headers-<exp>.npy, to <products_dir>/exp/<shard>/<exp>/manifests/ exp_footprint.json. Requested by rule all whatever the map switch says; clean_exposure waits on it for a live store. - nexp_map (campaign, off unless exposure_maps.nexp.enabled): stamps every footprint record on the products root into <products_dir>/nexp_map/nexp_map_<run>.hsp, a uint16 count of exposures with a valid PSF per pixel, beside nexp_map_<run>.json. Memory is sized on the campaign footprint like defect_map_merge (map_cov_pixels()). - astra.yaml: decision masking.nexp_map_valid_psf_ccds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BeQNTBcjTqu6TPktxysovQ
nexp_map follows the defect map's rule: built whenever the campaign can support it (MAPS_NEXP = a fitted PSF), skipped without error for psf_model: fake. exposure_maps.nexp.enabled defaults to true and stays as an opt-out, since the map is rebuilt whole. The pin gains nexp_map's job. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BeQNTBcjTqu6TPktxysovQ
nexp_map_targets() mirrors defect_map_targets(): an in-scope exposure with no footprint record (reclaimed before exp_footprint existed) is counted in a parse-time warning, since the map undercounts there, and with no record at all the map is not requested, instead of a job that fails every invocation. docs: an Exposure-level maps page with both chains. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BeQNTBcjTqu6TPktxysovQ
A run config still carrying the retired blocks parsed silently, so coverage.enabled: false no longer turned the count map off. The error names the replacement keys under exposure_maps. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#897 renames run_config.apply_machine_defaults to apply_defaults; the ordering assertion now holds on both sides of that merge. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #886. Closes #878.
The exposures yield two HealSparse maps the footprint needs and nothing else can build. Both are per-exposure records on the persistent root combined by one campaign job, at one resolution (
exposure_maps.nside131072 /nside_coverage128, the mask ladder's). Each is built whenever its input supports it, with no configuration: the defect map for data (sim flag images are blank), the count map for a fitted PSF (notpsf_model: fake). The default sims config builds neither.Defect map (data runs). The flag images (bad columns, saturated pixels, bleed trails) never leave the pixel domain, so the CCD-corner footprint includes defective pixels.
exp_defect_map(afterexp_split) samples each flagged pixel on a 3×3 grid including its corners, maps the samples through the CCD WCS, and masks every healpix pixel they touch.True= masked. One exposure: 34 s, 0.62 GB, ~2 MB fragment.defect_map_mergeORs the fragments intodefect_map/defect_map_<run>.hspbeside a sidecar of fragment digests; a changed or removed fragment or a resolution change rebuilds. Memory scales with footprint.Exposure-count map (fitted-PSF runs). Per pixel, the number of exposures with a valid PSF model — the count sp_validation's
npoint >= 3cut reads.exp_footprint(localrule, afterexp_persist): sky corners of every CCD with avalidation_psfmember inexp_persist.json(exact: psfex_interp writes it only on success), WCS fromheaders-<exp>.npy. Always written, because the purge takes the headers.nexp_mapstamps every footprint on the products root intonexp_map/nexp_map_<run>.hsp(uint16) besidenexp_map_<run>.json. Campaign-cumulative and rebuilt whole, soexposure_maps.nexp.enabled: falseopts a campaign out (e.g. to build once at the end of many small batches). Exposures reclaimed beforeexp_footprintexisted have no record: parse warns that the map undercounts, and with no records at all no map is requested. ~6 ms per CCD (~80 min for DR6's ~800k CCDs, synthetic benchmark); peak memory ≈ the map, ~48 GB over DR6, requested at 2×.build_coverage_mapchain; its geometry (_image_shape,_ccd_corners, RA-seam and pole guards, nside validation) is kept as a library.plot_coverage_mapstays, out of the DAG.Nothing in the workflow reads either map;
config_tile_Mc.ini's header says how to add the defect map toMASK_EXT_PATHS.clean_exposurewaits for a live exposure's fragment and footprint. User docs:docs/source/exposure_maps.md.Checks: decisions
masking.defect_map_from_flagsandmasking.nexp_map_valid_psf_ccds.test_defect_map_contains_flags.py,test_defect_map_reconcile.py;test_exp_footprint.py(CCD-index invariant),test_footprint_edges.py(reclaimed stores name no footprint; missing records warn, none requests no map),test_nexp_map_gates.py(on for fitted PSFs, skipped for fake, opt-out honoured),test_nexp_map.py,test_coverage.py. Default configs resolve for data (both maps) and sims (neither). Run configs still carrying the retiredcoverage:ordefect_map:block are refused at parse time, naming theexposure_maps.*replacement (run_config.retired()). Params pin gainsexp_footprintandnexp_map; no existing rule's digest moves. 516 passed, 1 skipped in the dev image.Not yet done: build the count map on a real campaign and compare to Martin's v1.6
coverage.hspover the overlap.— Claude (Opus) on behalf of Cail