Skip to content

Read ShapePipe v1 and v2 products side by side - #343

Open
cailmdaley wants to merge 36 commits into
developfrom
feat/v2-catalogue-readers
Open

cailmdaley wants to merge 36 commits into
developfrom
feat/v2-catalogue-readers

Conversation

@cailmdaley

@cailmdaley cailmdaley commented Sep 10, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #340. Closes #342. Closes #361.

sp_validation now reads both generations of ShapePipe products through one code path written against the v2 grammar:

  • v2: the Snakemake campaign products (final_cat_<campaign>.hdf5, full_starcat_<campaign>.hdf5, MASK_n* columns, no patches).
  • v1: every existing catalogue, v1.3–v1.6, including the fiducial v1.4.6.3.

Readers (v2). catalog.read_campaign_catalogue / iter_campaign_tiles / read_star_catalogue open campaign files with a flat or nested tile layout. They validate columns on every tile, check n_tiles, and promote dtypes across tiles. Peak memory is the merged array plus one tile. Patches P1–P7 are retired, along with five scripts that only addressed that layout.

Adapter (v1). grammar.adapt presents a v1 table in v2 names and units, and is the only code that knows the generation. It reads the generation from the file's columns.

  • It renames E1_PSF_HSM → HSM_G1_PSF, NGMIX_ELL_* → NGMIX_G1/G2_*, and so on.
  • It converts SIGMA_*_HSM = σ → HSM_T_* = 2σ².
  • It corrects one v1 defect: the no-shear reconvolved-PSF size is taken from the 1P column. Remove vestigial FHP/MK NOSHEAR←1P Tpsf substitution (guaranteed no-op) #267 removed this substitution as a no-op, which holds only for v2 products.
  • It works lazily over numpy, FITS_rec and h5py. A row selection on h5py reads each dataset once.
  • The map is in docs/ngmix_psf_column_migration.md.

Masks. v1 and v2 share one mask vocabulary.

  • ApplyHspMasks writes MASK_n{b}, and the adapter renames existing data_ext columns (4_Stars → MASK_n4, …).
  • CalibrateCat.read_cat returns one table (data joined with data_ext).
  • Every mask config is a single dat: list, and a config with a dat_ext: section raises an error.
  • galaxy.mask_cut ORs the six reason bits and drops NaN rows with a warning, as the config path does.
  • v1 configs keep their IMAFLAGS_ISO == 0 cut, whose bits mean something different from MASK_n*.

ρ/τ reads catalogues through the adapter and passes tables to shear_psf_leakage (#44 there; uv.lock bumped to include it, and the seeded jackknife patches from #43). cat_config column names were corrected to match their files, and a candide test now checks every entry. Dropping a scalar mask= argument fixes the DES and jackknife ρ/τ paths, which built empty catalogues.

Breaking changes for existing scripts and notebooks:

  • read_cat() returns one table, not (dat, dat_ext).
  • get_masks_from_config(config, dat, ...) takes that table.
  • Mask configs name MASK_n{b}.
  • extract_info no longer runs on v1 per-patch final_cat files.
  • The nine SExtractor-only columns DR6 catalogue mode lacks (shapepipe #924) are no longer passed through.

Verification.

  • Release regression: a candide test recalibrates rows 200M–201M of the v1.4 comprehensive catalogue and matches released v1.4.6.3 exactly: the same 131,526 objects in the same order, and 17 per-object columns bit-identical. On a separate 5M-row / 240-tile subset, the selection is identical (684,385 objects) and so is every per-object column. The weights and leakage correction are not bit-identical to the release, because the release fed its weight and leakage binning the uncorrected PSF size while this branch uses the corrected size throughout.
  • ρ/τ on SP_v1.4.6.3 through the adapter is identical to the earlier run on a hand-translated PSF file (≤1e-13σ with the same patch seed).
  • v2: exercised on the smk-g7 campaign (64 tiles; see comment).
  • Unit tests: 415 passed. The one failure is the pre-existing missing-paths test (covariance and SP_LFmask files absent on disk).

Stacked on #354 (int32 metacal flags); merge that first.

— Claude (Opus) on behalf of Cail

🤖 Generated with Claude Code

https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ

@github-actions

github-actions Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

✅ ruff is clean — nothing to fix here.

@cailmdaley

Copy link
Copy Markdown
Collaborator Author

Exercised against the first real v2 products (smk-g7, 64 tiles): galaxy reader 1,851,100 rows / 77 columns in 11 s at 1.2 GB RSS; star reader 53,264 stars over 127 exposures; default mask cut keeps 72.98%; full selection 1,105,851 rows with mean e1 = −0.00020 ± 0.00019, e2 = −0.00015 ± 0.00019. Three follow-up commits from that run: an explicit NGMIX_N_EPOCH > 0 guard (CosmoStat/shapepipe#889 — 1.03% of objects are never fit but carry MCAL_FLAGS = 0 with −10 sentinels; the old != −10 PSF comparison happened to catch them), a NaN-safe mask cut, and n_exposures validation + an EXPID column on the star reader. One decision left open on purpose: config/calibration/mask_v2.0.yaml has no NGMIX_MCAL_FLAGS cut — adding one changes row counts, so it is Cail's call.

— Claude (Fable) on behalf of Cail

cailmdaley and others added 22 commits September 28, 2026 06:17
… key

NGMIX_MCAL_FLAGS carries bit 30 (shapepipe#854's absent-measurement flag)
alongside the native fitter bits. Format I truncates it to int16 and
silently zeroes bit 30 on write, in both the FITS and the HDF5 branch of
write_shape_catalog (the HDF5 path reuses the FITS column's array).

The per-type NGMIX_FLAGS_{1P,1M,2P,2M,NOSHEAR} columns had the same
problem from the other direction: their format entry was keyed as
"FLAGS_{suffix}", which never matches the real column name
"{prefix}_FLAGS_{suffix}", so the writer's float64 default masked the
dead key rather than narrowing anything. Both parameter files now key
and format all six metacal bitmasks the same way, at K, so they round-
trip exactly and survive JointCat's optional memory-reduction pass
(which only downcasts int32/float64, not int64).

NGMIX_MCAL_TYPES_FAIL stays at I: it is a count in [0, 5], not a bitmask.

Adds a round-trip test parametrized over both parameter files, FITS and
HDF5, and reduce_mem on/off, checking 0, bit 30 alone, and bit 30 with a
native bit together.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ShapePipe sets bit 2**30 in the metacal flags for a missing measurement,
which needs 32 bits. The parameter files now write NGMIX_MCAL_FLAGS and
the five NGMIX_FLAGS_* columns as FITS J (int32).

JointCat.dtype_out with reduce_mem narrowed every int32 column outside a
keep-list to int8, silently wrapping any value above 127. It now only
narrows float64 to float32 (RA/Dec excepted) and leaves integer columns
alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace read_hdf5_file's hardcoded patches/<name>/<tile-ID> lookup with
find_dataset_group(), which descends from the file root through single
container groups until it reaches the per-unit datasets. This reads the
legacy patches/<campaign>/ layout that ShapePipe still writes as a
compatibility shim, a future flat tiles/ layout, and the exposures/<exp>
layout of full_starcat_<campaign>.hdf5 with the same code.

read_star_catalogue() keeps the FITS path for files ending in .fits.
Requested columns missing from the data now raise a clear KeyError.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
ShapePipe v2 drops IMAFLAGS_ISO for eleven boolean MASK_n* columns.
galaxy.mask_cut() ORs a configurable list of them (default MASK_n4,
MASK_n1, MASK_n2, MASK_n8, MASK_n1024 — stars, star halos, manual galaxy
mask, MaxiMask) and returns the keep mask; a catalogue missing any of the
requested columns raises a KeyError naming them.

classification_galaxy_base takes mask_columns; extract_info.py passes the
params.py mask_columns list and uses it for the star-sample cut too, and
now reads both catalogues through the new readers. Column lists in
params.py and masking.py's SPATIAL_CUTS updated to the MASK_n* names.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
ShapePipe v2 processes a campaign (a tile list); P1-P7 no longer exist.

- catalog_builders.JointCat: get_patches()/get_n_obj() and the per-patch
  FITS merge are replaced by a merge over a list of campaign hdf5 files
  (-i final_cat_A.hdf5+final_cat_B.hdf5), read through
  read_campaign_catalogue. The 'patch' int8 column becomes a 'campaign'
  string column; the hdf5 root attr becomes 'campaigns'.
- survey.get_footprint() deleted: it was a lookup table of P1-P7 (plus W3)
  RA/Dec boundaries, meaningless for a campaign. Its only caller,
  catalog.check_matching, used it as an optional pre-filter via a 'name'
  argument that every caller passed as None; the argument goes too, along
  with the test_survey test that exercised P5.
- merge_psf_cat.py, combine_results.py, stats_tile_id_gal_counts.py,
  compute_area.py: patch vocabulary and v1/v1.5/v1.6 P-name shortcuts
  generalised to an explicit list of campaigns.
- params.py: 'name = "P7"' becomes 'campaign = None'.

Deleted (only ever meaningful for the P1-P7 era):
- scripts/prepare_patch_for_spval.sh: symlinks a v1 per-patch run tree
  (~/psfex/${patch}/output/run_sp_Ms/.../full_starcat-0000000.fits,
  tiles_${patch}.txt) into a working dir; neither the layout nor the file
  names exist in v2.
- scripts/plot_rho_stats_patches.py: globs P* directories and reads
  P*/output/run_sp_Pl/mccd_plots_runner/output/rho_stats_id.fits, one
  curve per patch. No campaign analogue.
- scripts/survey_stats_all.sh: hardcoded 'for patch in P1 ... P7' over v1
  run-directory bookkeeping.
- scripts/star_match_stats.py: sums stats_file.txt over the seven patches.
- scripts/check_tile_IDs_SP_LF.py: compares per-patch ShapePipe tile IDs
  against CFIS3500_THELI_P<n>.list Lensfit files; both sides P-named.

The word 'patch' survives only in catalog.py's comment naming the legacy
hdf5 group, and in the treecorr jackknife sense (npatch/patch_number).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Seventeen unit tests on tiny synthetic hdf5 fixtures built in a temp dir:
both campaign layouts (legacy patches/<campaign>/<tile-ID> and flat
tiles/<tile-ID>) read identically, param-list restriction, missing-column
and ambiguous-layout errors; the star reader on exposures/<exp> hdf5 and
on FITS; galaxy.mask_cut defaults, explicit list, empty list, missing
column, and a v1 IMAFLAGS_ISO-only catalogue; JointCat.merge_catalogues
across two campaigns of different layouts, plus its error paths.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
- concatenate_datasets preallocates the output and fills it column by
  column, so peak memory is the packed output plus one tile instead of the
  full-width uncut catalogue; the verbose estimate uses the real itemsize.
- validate the requested columns against every tile dataset, not only the
  first, and name the offending dataset in the error.
- read_campaign_catalogue checks the root n_tiles attribute and refuses a
  truncated file; campaign_shape reports row count and dtype from metadata.
- JointCat.merge_catalogues preallocates the merged array from that first
  pass (no more accumulate-then-concatenate, which doubled peak memory),
  promotes each column's dtype across all campaigns so a wider string or
  integer column in a later file is no longer silently truncated, and
  refuses multi-dimensional columns explicitly.
- reduce_mem reduces int32 to int16 (int8 wrapped N_EPOCH/CCD_NB values)
  and every assignment is range-checked, raising instead of wrapping.
- check_matching drops the identity index over d1 that the retired
  footprint prefilter left behind, and extract_info applies mask_cut to
  the matched subset instead of the whole catalogue.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Add config/calibration/mask_v2.0.yaml: the v1.X.11 cut set with
IMAFLAGS_ISO and the v1 post-processing masks replaced by the boolean
MASK_n<bit> columns (True = masked, hence kind: equal, value: False), the
coverage bits and MASK_n2048 listed but commented out. The v1.X configs are
left untouched: each describes a legacy catalogue that really has
IMAFLAGS_ISO, and rewriting them would break reproducing published versions.

plots.sky_plots no longer hardcodes the v1 label set; it combines whichever
of IMAFLAGS_ISO / MASK_n* / npoint3 / 1024_Maximask the config declared.
The comprehensive-to-minimal demo asks for the v2 mask labels.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
- params.py interpolated the removed 'name' into two paths (NameError on
  import); use 'campaign', give it a real default, and point star_cat_path
  at full_starcat_<campaign>.hdf5 (hdu_star_cat now documented as legacy
  FITS only). params_im_sim.py renames 'name' to 'campaign' likewise.
- combine_results.get_area matches both the campaign and the legacy patch
  wording, and raises on a missing file or unmatched pattern instead of
  returning None / a 1 deg^2 placeholder that silently rescales densities.
- merge_psf_cat writes the campaign *name* as a string column (FITS 'A<n>'),
  matching JointCat; the old 1-based ordinal depended on -p argument order.
- compute_m_bias_image_sims counts tiles via the n_tiles attribute or
  find_dataset_group, not a hardcoded legacy group inside a bare except.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Tests: assert the exact concatenation instead of sorted values, add a
fixture whose keys are inserted out of order, and cover the truncated
n_tiles file, a column missing from a later tile, dtype promotion across
campaigns in both argument orders, reduce_mem overflow, and a
multi-dimensional column.

Docs: CLAUDE.md, scripts/calibration/README.md, homogenize_cat_extended.py
and the catalog_builders docstrings now say campaign.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Three defects in the campaign reading path, found in review:

- concatenate_datasets and campaign_shape both took the output dtype from
  the first dataset alone, so a campaign whose tiles differ (S7 next to
  S12 tile IDs, i2 next to i4, f4 next to f8 after a partial
  reprocessing) had the later tiles silently truncated, downcast or
  wrapped. np.concatenate, which this code replaced, promoted. Both now
  build the dtype with group_dtype(), promoting every column across every
  dataset; catalog_builders._promote becomes an alias of the shared
  catalog.promote_dtypes rather than a second copy of it.

- merge_catalogues still held a whole campaign in memory next to the
  preallocated output, the very thing its comment claimed the rewrite
  avoided -- and with one campaign per merge in v2, that is the normal
  case, ~2x the merged catalogue at DR6 scale. It now fills the output
  tile by tile through the new catalog.iter_campaign_tiles(), so peak
  memory is the output plus a single tile.

- write_hdf5_file wrote the merged array twice, create_dataset(data=...)
  followed by an immediate dset[:] = dat_all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
mask_v2.0.yaml dropped v1.X's '64_r' r-band imaging cut without a
replacement, so a v2 calibration run admitted objects outside the r-band
footprint that v1 rejected -- a silent change of effective area, n(z) and
galaxy-density normalisation. Enable MASK_n64, v1's '64_r' equivalent.

v1's other coverage cut, npoint3 >= 3, came from an external
post-processing catalogue and has no v2 counterpart; say so in the config
rather than leaving its absence unexplained.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
extract_info fancy-indexed the full-width catalogue, dd[ind_star], only
to read ~5 boolean mask columns from it -- a copy of every column for
every matched star (~GB at DR6 scale). Mask first, index the resulting
bool array.

plot_leakage still labelled its curves "all", "P1" ... "P7": a P-named
survivor of the #340 retirement, and a fixed length that silently
mismatched the number of input files. Labels and colours now follow the
input files.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
The v2 branch inferred from v1's '64_r' naming that MASK_n64 flags r-band
imaging coverage. It does not: bit 64 is an undocumented *reason* bit of
the r-band default bitmask, and OR{n1,n2,n4,n8,n64,n1024} reproduces
mask_r, the v1 r-band mask, exactly on the P3 region.

Add MASK_n64 to DEFAULT_MASK_COLUMNS so the default galaxy cut is exactly
that set, documented as reproducing mask_r, and mirror it in the
calibration params, the v2 mask config and the minimal-catalogue demo.

Document n16/n32/n128/n256 as the u/g/i/z coverage flags (no r flag: the
catalogue is r-selected) and n2048 as absent Pan-STARRS z2. The faint vs
bright assignment of n1/n2 is unconfirmed for the Aug-2026 products, so
the labels no longer claim one.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
ShapePipe's make_cat pre-fills the NGMIX_* columns with sentinels
(G1/G2 = -10, T/FLUX = 0) and overwrites them only for objects present
in the ngmix output, so an object ngmix never fit keeps
NGMIX_MCAL_FLAGS == 0 and passes a flag-only cut. In
final_cat_smk-g7.hdf5 that is 18,983 of 1,851,100 objects (1.03%);
admitting them drags mean e1 to -0.096 (std 0.98) from +0.0001.

classification_galaxy_ngmix already rejected all 18,983 via the
NGMIX_G1_PSF_ORIG_NOSHEAR != -10 guard, so the production selection was
never affected -- verified on the real file, which gives 1,105,851 rows
out with and without the new cut. But that protection was incidental:
it is an exact float equality against a sentinel ShapePipe may change,
and the coadd N_EPOCH >= 2 cut in classification_galaxy_base does not
substitute for it (18,750 of the 18,983 have N_EPOCH >= 1). Cut on
NGMIX_N_EPOCH > 0 explicitly so the guarantee is stated, not inferred.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
mask_cut decided on truthiness via astype(bool). ShapePipe writes the
MASK_n* columns as float64 {0, 1} rather than bool (being fixed
upstream), and astype(bool) reads NaN as True, so an incomplete mask
column would have silently deleted sky. Decide on the value instead
(masked iff > 0.5), accept bool, int and float alike, and treat NaN as
"no verdict recorded" -- keep the object, but count and warn, since a
nonzero count means the product is defective. final_cat_smk-g7.hdf5
carries no NaNs and only exact 0.0/1.0, so this is defensive: the real
file gives 1,105,851 rows out before and after.

read_star_catalogue silently accepted a truncated file and threw away
exposure provenance. It now validates the n_exposures root attribute
against the datasets found, as the galaxy reader validates n_tiles
(check_n_tiles generalised to check_n_units), and adds an EXPID column
carrying the exposure number each star came from. The datasets are
named by that number and concatenating them discarded it, leaving no
way to group stars by exposure downstream. Names may be bare
("2086324", as smk-g7 writes them) or carry the CFIS suffix
("2110000p"), so EXPID takes the leading digits. On the real star
catalogue: 53,264 stars over 127 exposures, matching n_exposures.

Also note in group_dtype that ShapePipe writes TILE_ID as f8, so the
string-promotion branch is for a future string-valued TILE_ID, with a
TODO recording that as an open schema decision. No behaviour change.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
mask_v2.0.yaml, applied downstream to the comprehensive catalogue, cut
on neither NGMIX_MCAL_FLAGS nor any epoch column. It rejected the
never-fit objects only through its NGMIX_G1/G2_PSF_ORIG_NOSHEAR != -10
cuts, and only because make_cat happens to fill the PSF columns from
the same -10 literal it uses for the galaxy ellipticities. That is the
same accidental immunity just removed from
classification_galaxy_ngmix, one stage further downstream.

Add NGMIX_N_EPOCH >= 1. The column is already carried into the
comprehensive catalogue via add_cols_pre_cal in params.py. On
final_cat_smk-g7.hdf5 the cut keeps 1,832,117 of 1,851,100 objects,
removing exactly the 18,983 (1.03%) never-fit rows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VXzqmMYw7Kp8QMtoiVHjVq
sp_validation reads only the v2 column grammar, while every catalogue
on disk (v1.3.x-v1.6.x) is v1, so e.g. rho/tau raised KeyError on the
v1.4.a PSF file. New module sp_validation.grammar is the one place that
knows the difference:

- V1_RULES: the v1->v2 map of the shape-measurement columns as data (HSM
  renames, SIGMA -> T = 2 sigma^2 via cs_util.size.sigma_to_T, 2-vector or
  flattened NGMIX_ELL* split into G1/G2, PSFo/Tpsf/MOM_FAIL renames),
  applied to tables detect_generation calls v1. Mapping is by name, so v1
  values reach the names the code reads.
- MASK_RULES: the healsparse mask bits {b}_{label} (the names
  ApplyHspMasks gave them in the comprehensive HDF5's data_ext) ->
  MASK_n{b}, applied whatever the generation: they are the same bits of
  the same UNIONS bitmask ShapePipe v2 writes as MASK_n{b}
  (MASK_LABELS; 512 = outside the tile's unique region). IMAFLAGS_ISO is
  not mapped: its v1 bits mean different things.
- adapt(table, *tables): tables with nothing to rename come back
  unchanged; otherwise a V2View, a lazy column view over numpy, FITS_rec
  or h5py that joins row-aligned tables (data + data_ext), computes derived
  columns on access, and composes row selections as ranges or selected
  indices, reading only the window of rows they span.
- read_catalogue(path, hdu), materialise, v2_names, read_column_names.

rho_tau.get_rho_tau / get_jackknife_cov / get_theory_cov read each
catalogue once through read_catalogue and hand the table to
shear_psf_leakage (needs its loaded-catalogue support).

Tests: v1 twins adapt to their v2 twins for numpy (vector and flattened
ELL), FITS_rec and h5py; row selection commutes; dtype matches every
column read; data_ext mask names rename with or without a generation and
conflict with their new names; h5py selections read only their window;
the psf_size_error field and full get_rho_tau outputs agree between a v1
PSF catalogue and its v2 twin (and a rename-only sigma is caught).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cailmdaley and others added 11 commits September 28, 2026 06:17
test_cat_config_columns_exist_on_candide reads each cat_config
catalogue's FITS header through grammar.read_column_names and checks
the psf block's declared columns and the shear block's *_col columns
are among the names the file presents. Known gaps are listed with
reasons and fail the test once healed.

Config fixes it surfaced:
- psf dec_col Dec -> DEC: every PSF file (and ShapePipe v2) names it
  DEC; only FITS_rec's case-insensitive lookup hid the mismatch.
- SP_v1.4.12.3 / SP_v1.4.13.3 psf blocks named v1 columns and the
  retired square_size flag; now the v2 names like every other entry.
- SP_axel_v0.0 / SP_v1.4.5.A shear: declare their lowercase ra/dec;
  SP_v1.4.5.A's PSF shape columns are psf_g1/psf_g2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
build_catalog indexes each column with ``mask``; a scalar True only
adds an axis treecorr reshapes away (no stars are cut), and a scalar
False selects nothing, so treecorr raises "Input arrays have zero
length". That broke get_rho_tau for DES and get_jackknife_cov for
every catalogue. No flag cut was ever applied, so dropping the
argument leaves the non-DES results unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every reader on the calibration path now presents its tables in the v2
column grammar, so a ShapePipe v1 product runs through the same code and
the same configs as a v2 one:

- catalog.read_campaign_catalogue, campaign_shape, iter_campaign_tiles and
  the JointCat merge adapt each tile (param_list names v2 columns);
  read_star_catalogue adapts FITS and HDF5 star catalogues.
- CalibrateCat.read_cat returns one table: the comprehensive HDF5's data
  and data_ext joined by grammar.adapt, so mask columns read the same
  whether they sit in data (v2) or data_ext (post-processed v1), and
  galaxy.mask_cut works on either.
- get_masks_from_config takes that one table, and each mask config is one
  `dat` cut list: the v1 configs' dat_ext cuts move into it under their
  MASK_n{b} names (the selections are unchanged; IMAFLAGS_ISO stays in the
  v1 configs). The image-sim overlay and scripts/masking.py follow.
- ApplyHspMasks writes mask bit b as MASK_n{b}, from grammar.MASK_LABELS.
- cosmo_val/compute_theory_cov.py hands CovTauTh loaded, adapted tables.

Tests: a v1 comprehensive HDF5 (v1 data + data_ext with the old mask
names) and its v2 twin give identical mask_cut and per-config mask
selections for every config in config/calibration, and identical metacal
inputs and response; the campaign and star readers present v1 files in v2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…builders, seeded patches)

Pinned below develop's tip 372980f, whose scipy>=1.18 requirement cannot
resolve against cs_util<0.3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
ShapePipe v1 wrote a wrong NGMIX_Tpsf_NOSHEAR: it differs by ~2% for
nearly every object (v1.4, v1.5, v1.6 comprehensive) from the
reconvolution kernel metacal applied, which NGMIX_Tpsf_{1P,1M,2P,2M}
record (they agree to ~1e-5). The v1.4.6.3 release used the 1P value
in its metacal size cut and wrote it as the no-shear column; #267
dropped that substitution, which is a no-op only on the v2 stack, so
the branch's v1 calibration selected ~3% more objects than the release.

The adapter now presents NGMIX_T_PSF_RECONV_NOSHEAR from NGMIX_Tpsf_1P
when the table has it, falling back to NGMIX_Tpsf_NOSHEAR (a cut
catalogue such as v1.4.6.3's already holds the 1P value there), and
hides the raw column. Rules gain a fallback source; a derived column is
presented where the first of its sources sits. The module docstring
and the migration doc state the rule: name mapping, except this one
documented v1 defect.

Every consumer now sees the 1P value, including the w_des and
PSF-leakage size-ratio binning, where the release used the raw value;
those two columns therefore do not reproduce the release bit for bit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
metacal reads ~40 columns of data[mask]. Over h5py datasets the lazy
view read each column over the whole span of the selection, so a
full-sky v1 calibration re-read the 238 GB comprehensive file once per
column. Selecting rows of a view over h5py Datasets now reads the
selected rows of every column in one pass per dataset (whole rows, in
256 MB blocks over the selection's span, skipping empty blocks), as
indexing the Dataset did on develop, and wraps them in a view over the
in-memory arrays so renames and derived columns still apply. Views
over in-memory tables stay lazy. to_structured reads each dataset once
for all requested fields.

On a cold 10M-row window of v1.4.c (4.8M rows selected):
h5py Dataset[mask] 19.2 s; adapt(data, data_ext)[mask] plus the 40
metacal columns 23.5 s; the previous per-column path took 16.5 s for
its first column alone, i.e. one full re-read per column.

dtype is computed once per view and passed to row selections
(group_dtype asked for it per column, quadratically). view[()] and
view[...] select every row, as for an h5py Dataset; other tuples raise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
A column missing from some tiles (several v1 tiles carry no
SPREAD_MODEL) surfaced as numpy's bare "no field of name" KeyError.
group_dtype now checks every table first and raises a KeyError naming
how many tables lack which columns, and which ones; callers pass the
tables keyed by dataset name so the message names tiles.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
Mask configs are one dat: list over the joined data + data_ext table,
but get_masks_from_config and scripts/masking.py read only config["dat"],
so an older config with a dat_ext: list (several sit under v1.4.x and
v1.5.x) silently lost those cuts. masks.catalogue_cuts returns the dat
list and raises on dat_ext, saying to merge it into dat with {b}_{label}
renamed MASK_n{b}; the calibration, footprint and demo scripts all read
cuts through it.

scripts/masking.py's footprint cuts had dropped IMAFLAGS_ISO, which the
v1 configs still cut on and which is spatial (halo, border, Messier,
NGC, spike bits); it is back in SPATIAL_CUTS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
…s do

mask_cut kept objects whose mask column is NaN (with a warning), while
the config-driven cut that calibration runs (kind: equal, value: False)
drops them, so the two selections disagreed on a defective product.
mask_cut now drops them too, still counting and warning. Keeping them
was chosen to stop astype(bool) from dropping NaN rows silently; the
warning covers that, and dropping an object with no mask verdict is
the conservative choice for a shear catalogue.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
Drop wording that goes stale ("being fixed upstream", "unconfirmed for
the Aug-2026 products", "legacy" layouts and files) in favour of what
is true now, and ApplyHspMasks.write_hdf5_header's documented but
nonexistent campaigns parameter.

test_configured_paths_exist_on_candide resolved a calibration config's
relative params.input_path against config/calibration; such configs
run as config_mask.yaml in their run directory, where the relative
path names a file, so the guard skips those.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
…, slow)

Runs scripts/calibration/calibrate_comprehensive_cat.py with
mask_v1.X.6.yaml on rows 200M..201M of the v1.4.c comprehensive HDF5
and checks the released cut catalogue's rows from that window, matched
by (RA, Dec) and bracketed by rows from outside it: same objects, same
order, and identical per-object columns, including the no-shear
reconvolved-PSF size against the release's NGMIX_Tpsf_NOSHEAR. Globally
calibrated columns (e1, e2, w_des, leakage-corrected) are not compared.
Skipped where the release products are absent.

Passes in ~30 s on n09 (131,526 objects); before the no-shear
reconvolved-PSF correction it fails, selecting 134,936.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
@cailmdaley
cailmdaley force-pushed the feat/v2-catalogue-readers branch from 9ec5502 to 330ac8d Compare September 28, 2026 10:24
@cailmdaley cailmdaley changed the title Read ShapePipe v2 campaign products: hdf5 catalogues, MASK_n* cut, patches retired Read ShapePipe v1 and v2 products side by side Sep 28, 2026

@martinkilbinger martinkilbinger left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very nice, I like the campaign tag superceding the patch IDs. And the adopt mechanism.


# Stars
- col_name: 4_Stars
- col_name: MASK_n4

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is the reason to remove the mask type from the column name and just keep the bit number?

Comment thread docs/ngmix_psf_column_migration.md Outdated
| 1, 2 | `1_Faint_star_halos`, `2_Bright_star_halos` | star halos |
| 4 | `4_Stars` | star mask |
| 8 | `8_Manual` | manual mask (large galaxies) |
| 16–256 | `16_u`, `32_g`, `64_r`, `128_i`, `256_z` | per-band coverage |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
| 16–256 | `16_u`, `32_g`, `64_r`, `128_i`, `256_z` | per-band coverage |
| 16–256 | `16_u`, `32_g`, `64_r`, `128_i`, `256_z` | per-band coverage, where 256_z is the HSC z-band |

Comment thread docs/ngmix_psf_column_migration.md Outdated
| 16–256 | `16_u`, `32_g`, `64_r`, `128_i`, `256_z` | per-band coverage |
| 512 | `512_Tile_RA_DEC_cut` | outside the tile's unique region |
| 1024 | `1024_Maximask` | MaxiMask |
| 2048 | `2048_z2` | no Pan-STARRS z2 |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
| 2048 | `2048_z2` | no Pan-STARRS z2 |
| 2048 | `2048_z2` | per-band coverage, with 2048_z is the Pan-STARRS z-band (`z2`) |

cailmdaley and others added 3 commits September 29, 2026 16:25
ShapePipe v2 catalogues in DR6 catalogue mode no longer carry MAG_WIN,
MAGERR_WIN, SNR_WIN, FLUX_AUTO, FLUXERR_AUTO, FLUX_APER, FLUXERR_APER,
FWHM_IMAGE, FWHM_WORLD (shapepipe #924). Stop passing them through in
extract_info (params.add_cols) and calibrate_comprehensive_cat, and drop
the SExtractor SNR_WIN curve from the galaxy SNR histogram. The v1.4.6.3
release regression compares the remaining per-object columns.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BeQNTBcjTqu6TPktxysovQ

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants