Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
25 commits
Select commit Hold shift + click to select a range
598a87d
Add the UNIONS tile-catalogue inis to the workflow config chain
cailmdaley Sep 17, 2026
8702483
Add tile_detection: unions_catalogue to the workflow
cailmdaley Sep 17, 2026
379be48
Test the UNIONS catalogue conversion and its workflow wiring
cailmdaley Sep 17, 2026
affcdc0
Fix TILE_UNIQUE_ID merge, NUMBER contiguity, and FITS int64 dtype test
cailmdaley Sep 17, 2026
930b5af
Build the UberSeg segmentation map for the UNIONS catalogue path
cailmdaley Sep 17, 2026
57b7525
Default tile_detection to the UNIONS catalogue
cailmdaley Sep 17, 2026
36e96b3
Build TILE_UNIQUE_ID once, in make_cat, from one helper
cailmdaley Sep 28, 2026
4e76ccc
Select ngmix chunks by catalogue row, not NUMBER value
cailmdaley Sep 28, 2026
ba332f4
Merge develop into feat/unions-catalogue-tile-detection
cailmdaley Sep 28, 2026
7c91679
Merge refactor/tile-unique-id: TILE_UNIQUE_ID in make_cat, ngmix row …
cailmdaley Sep 28, 2026
e34b275
Segmentation run uses the tile detection's MegaPipe filter
cailmdaley Sep 28, 2026
bfd2f87
Merge origin/develop into feat/unions-catalogue-tile-detection
cailmdaley Sep 28, 2026
2e09027
Move UberSeg segmentation to its own branch
cailmdaley Sep 28, 2026
5511597
Default tile_detection and psf_model per input type, not per machine
cailmdaley Sep 28, 2026
9d1e023
Record tile_detection in astra.yaml
cailmdaley Sep 28, 2026
21e7c4e
Point inputs.catalogues at vos:cfis/tiles_DR6 on both machines
cailmdaley Sep 28, 2026
0a0396a
Merge origin/develop into feat/unions-catalogue-tile-detection
cailmdaley Sep 28, 2026
9f034f8
Leave the retired standalone merge_final_cat.py to develop
cailmdaley Sep 28, 2026
2905a88
final_cat_merge: request only the columns the tile's detection mode c…
cailmdaley Sep 29, 2026
8da837b
Mask neighbours in catalogue-mode VIGNETs from the DR6 segmentation map
cailmdaley Sep 29, 2026
a4da900
completeness: expect the seg-stamp vignetmaker run under blend_handli…
cailmdaley Sep 29, 2026
97c11a7
Revert "completeness: expect the seg-stamp vignetmaker run under blen…
cailmdaley Sep 29, 2026
a8facf5
config.yaml: plainer input_types comment (review suggestion)
cailmdaley Sep 29, 2026
bced5a3
Drop the relabelled seg-map file output
cailmdaley Sep 29, 2026
0ce4060
completeness: the catalogue converter writes one file
cailmdaley Sep 29, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
75 changes: 65 additions & 10 deletions astra.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -329,10 +329,11 @@ analyses:
# ═════════════════════════════════════════════════════════════════════════
detection:
description: >-
Object detection with SExtractor on r-band tiles (the galaxy sample) and
on single-exposure CCDs (PSF-star candidates). Tiles follow the MegaPipe
parameters of Gwyn's UNIONS tile catalogue; exposures keep ShapePipe's
stock values.
Object detection on r-band tiles (the galaxy sample) and on
single-exposure CCDs (PSF-star candidates, always SExtractor). The tile
sample is Gwyn's UNIONS tile catalogue or ShapePipe's own SExtractor run
(tile_detection); the latter follows the MegaPipe parameters of that
catalogue, while exposures keep ShapePipe's stock values.
inputs:
- id: tile_stack
type: data
Expand All @@ -344,12 +345,34 @@ analyses:
description: Per-tile SExtractor LDAC catalogue with per-epoch CCD membership.
inputs: [tile_stack]
decisions:
[detection_threshold_policy, deblending_policy, background_model,
[tile_detection, detection_threshold_policy, deblending_policy,
background_model,
weight_map_usage, zero_weight_interpolation, detection_source_mode,
epoch_membership_ccd_bounds, photometry_parameters,
epoch_membership_ccd_bounds, catalogue_neighbour_marking,
photometry_parameters,
spurious_detection_cleaning, blend_photometry_mask_type,
saturation_level]
decisions:
tile_detection:
label: Source of the tile galaxy sample
rationale: >-
Real data takes its tile sample from Gwyn's UNIONS per-tile
catalogue, converted to the sexcat the chain reads with its own
NUMBER kept, so shape, photometry and photo-z catalogues share one
object list and one ID (TILE_UNIQUE_ID). PSF stars still come from
ShapePipe's exposure-level SExtractor run, because external star
catalogues are too shallow. detection_source_mode and the tile
SExtractor settings of the other decisions here govern only the
sextractor option; epoch_membership_ccd_bounds applies to both.
Image simulations, which have no UNIONS catalogue, use sextractor.
Values:
input_types.data.tile_detection = unions_catalogue.
default: unions_catalogue
options:
unions_catalogue:
label: Gwyn's UNIONS per-tile catalogue, converted in place
sextractor:
label: SExtractor on the tile image
detection_threshold_policy:
label: Detection significance, minimum area, matched filter
rationale: >-
Expand Down Expand Up @@ -599,15 +622,48 @@ analyses:
admits its upper endpoint, column 2080. A WCS inversion failure skips
that CCD, lowering N_EPOCH; this sets how many exposures enter each
galaxy's multi-epoch fit.
Both tile_detection options run this post-processing.
Values:
SEXTRACTOR_RUNNER.CCD_SIZE = 33,2080,1,4612;
SEXTRACTOR_RUNNER.MAKE_POST_PROCESS = True.
SEXTRACTOR_RUNNER.MAKE_POST_PROCESS = True;
READ_EXT_SEXCAT_RUNNER.CCD_SIZE = 33,2080,1,4612;
READ_EXT_SEXCAT_RUNNER.MAKE_POST_PROCESS = True.
default: trimmed_bounds_33_2080
options:
trimmed_bounds_33_2080:
label: "x in (33,2080), y in (1,4612), strict"
inclusive_bounds:
label: Same bounds, inclusive
catalogue_neighbour_marking:
label: Neighbour pixels in the catalogue path's VIGNET
rationale: >-
ngmix masks a neighbour only where the tile VIGNET is -1e30 (flag
2**10: zero weight, noise-filled under noisefill), and SExtractor
writes -1e30 on neighbours' footprints and off the image. The
unions_catalogue converter reproduces that from the catalogue's own
r-band segmentation map (CFIS.<tile>.r.seg.fits.fz, on the tile's
pixel grid), so both tile_detection options mask the same way.
Segmentation labels are not the catalogue NUMBER; each object claims
the footprint under its centre pixel (the VIGNET centre pixel), which
on 202.301 matches 36,064 of 36,065 objects with none shared. The map
is relabelled to NUMBER (unclaimed footprints -1; an object on sky or
on a claimed footprint gets a 3-pixel-radius disc of its own), and
VIGNET pixels whose relabelled value is neither 0 nor the object's
NUMBER become -1e30. Without it, catalogue-mode ngmix fits neighbour
light: in pilot smk-g10, NGMIX_MAG_NOSHEAR was over 1 mag brighter
than MAG_AUTO for 6.1% of objects (2.7% with SExtractor, g9).
Values:
READ_EXT_SEXCAT_RUNNER.SEGMENTATION = True.
default: segmentation_map
options:
segmentation_map:
label: -1e30 on other objects' segmentation footprints, as SExtractor
none:
label: Image pixels only; ngmix masks no neighbours
excluded: true
excluded_reason: >-
Leaves neighbour light in the fit, unlike the sextractor option
it replaces; the cause of the smk-g10 magnitude outliers.
prior_insights:
guinot22_sextractor_params:
claim: >-
Expand Down Expand Up @@ -913,9 +969,8 @@ analyses:
psf_modelling_software:
label: PSF model, PSFEx per CCD or MCCD over the focal plane
rationale: >-
psf_model in workflow/config.yaml's machines: table (per machine and
input type; image sims use fake) selects the exposure and tile
config pair; the committed value is psfex, which fits each CCD
psf_model in workflow/config.yaml's input_types: table (image sims
use fake) selects the exposure and tile config pair; the committed value is psfex, which fits each CCD
independently. MCCD (Liaudat+2021) fits one hybrid local+global
model over the focal plane; the completeness table treats its
counts as warnings because no campaign has run it. Up to the model,
Expand Down
4 changes: 2 additions & 2 deletions scripts/python/create_final_cat.py
Original file line number Diff line number Diff line change
Expand Up @@ -33,8 +33,8 @@ def params_from_run_config(params, defaults):
The workflow already knows where a campaign writes, so a manual merge
should not have to restate it. Resolution goes through the workflow's own
resolver (workflow/scripts/run_config.py), layering the run config on
workflow/config.yaml and then the machines: table, so what lands here is
what the rules would have used.
workflow/config.yaml and then the input_types: and machines: tables, so
what lands here is what the rules would have used.

Only values still at their default are filled -- an explicit flag always
wins. Nothing is derived for the data path: its patch naming differs and
Expand Down
5 changes: 3 additions & 2 deletions src/shapepipe/modules/fake_psf_package/fake_psf.py
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,8 @@ def process(self):
raise

try:
n_gal = len(sex[3].data.field("NUMBER"))
numbers = np.asarray(sex[3].data.field("NUMBER"))
n_gal = len(numbers)
except Exception as e:
self._w_log.error(
f"Error reading catalogue data from HDU 3 in {self._sexcat_path}: {e}"
Expand Down Expand Up @@ -91,7 +92,7 @@ def process(self):
output_file = SqliteDict(self._output_path)
missing = 0
for idx, gal_row in enumerate(masked):
galaxy_number = idx + 1 # 1-based, matches NUMBER field
galaxy_number = int(numbers[idx])
gal_dict = {}
for exp_ccd in gal_row.compressed():
if exp_ccd not in psf_dict:
Expand Down
5 changes: 4 additions & 1 deletion src/shapepipe/modules/make_cat_package/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,10 @@
basic measurement parameters, the PSF model at galaxy positions, and the
shape measurement. Every detected object is kept: the catalogue carries no
star/galaxy classification, which is done downstream. Each object is
tagged with its source tile via the ``TILE_ID`` column; objects duplicated
tagged with its source tile via the ``TILE_ID`` column and carries the
survey-wide object ID ``TILE_UNIQUE_ID = tile_id * 10**6 + NUMBER``
(``tile_id = RRR * 1000 + DDD``, see
:func:`shapepipe.utilities.cfis.get_tile_unique_id`); objects duplicated
across overlapping tiles are not deduplicated here, and are left to
downstream selection.

Expand Down
25 changes: 13 additions & 12 deletions src/shapepipe/modules/make_cat_package/make_cat.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@
get_type_flags,
)
from shapepipe.pipeline import file_io
from shapepipe.utilities import mask_query
from shapepipe.utilities import cfis, mask_query


def get_output_name(output_dir, file_number_string):
Expand Down Expand Up @@ -100,7 +100,11 @@ def remove_field_name(arr, name):
def save_sextractor_data(final_cat_file, sexcat_path, remove_vignet=True):
"""Save SExtractor Data.

Save the SExtractor catalogue into the final one.
Save the SExtractor catalogue into the final one, adding the tile as
``TILE_ID`` (float ``RRR.DDD``) and the survey-wide object ID
``TILE_UNIQUE_ID`` (:func:`shapepipe.utilities.cfis.get_tile_unique_id`
of the tile and ``NUMBER``). The tile is read from the catalogue's file
name, e.g. ``sexcat-301-279.fits``.

Parameters
----------
Expand All @@ -123,22 +127,19 @@ def save_sextractor_data(final_cat_file, sexcat_path, remove_vignet=True):
data = np.copy(sexcat_file.get_data())
if remove_vignet:
data = remove_field_name(data, "VIGNET")

final_cat_file.save_as_fits(data, ext_name="RESULTS")

cat_size = len(data)

tile_id = float(
".".join(
re.split("-", os.path.splitext(os.path.split(sexcat_path)[1])[0])[
1:
]
)
tile_name = os.path.basename(sexcat_path)
nix, niy = cfis.get_tile_number(tile_name)
tile_id_array = np.full(cat_size, float(f"{nix}.{niy}"))
unique_id = cfis.get_tile_unique_id(
cfis.get_tile_id(tile_name), data["NUMBER"]
)
tile_id_array = np.ones(cat_size) * tile_id

final_cat_file.save_as_fits(data, ext_name="RESULTS")
final_cat_file.open()
final_cat_file.add_col("TILE_ID", tile_id_array)
final_cat_file.add_col("TILE_UNIQUE_ID", unique_id)

sexcat_file.close()

Expand Down
13 changes: 7 additions & 6 deletions src/shapepipe/modules/ngmix_package/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -42,13 +42,14 @@
Save the output catalogue in batches of this size; default is ``-1``
(no batch saving)
ID_OBJ_MIN : int
ID of first galaxy object to be processed; not used if set to ``-1``
(default). Environment variables are expanded, so an orchestrator can
set the object range per chunk, for example
``ID_OBJ_MIN = $SP_NGMIX_ID_OBJ_MIN``.
First tile-catalogue row to process, as a 1-based row position (not a
``NUMBER`` value); not used if set to ``-1`` (default). Environment
variables are expanded, so an orchestrator can set the row range per
chunk, for example ``ID_OBJ_MIN = $NGMIX_ROW_MIN``.
ID_OBJ_MAX : int
ID of last galaxy object to be processed; not used if set to ``-1``
(default). Environment variables are expanded, as for ``ID_OBJ_MIN``.
Last tile-catalogue row to process (1-based, inclusive); not used if
set to ``-1`` (default). Environment variables are expanded, as for
``ID_OBJ_MIN``.
BKG_RMS_VIGNET_PATH : str, optional
Path to a ``background_rms_vignet*.sqlite`` file produced by
``vignetmaker_runner``. The string may contain
Expand Down
47 changes: 37 additions & 10 deletions src/shapepipe/modules/ngmix_package/ngmix.py
Original file line number Diff line number Diff line change
Expand Up @@ -420,13 +420,40 @@ def get_prior(pixel_scale, rng, T_range=None, F_range=None):
)


def chunk_rows(n_obj, row_min, row_max):
"""Catalogue rows of one ngmix chunk.

A chunk is a closed range of 1-based row positions in the tile
catalogue, independent of the ``NUMBER`` values those rows carry, so a
partition of ``1..n_obj`` covers every object once however ``NUMBER``
is ordered or spaced. A bound ``<= 0`` is unbounded on that side.

Parameters
----------
n_obj : int
Number of rows in the tile catalogue
row_min, row_max : int
First and last row of the chunk (1-based, inclusive)

Returns
-------
range
0-based row indices of the chunk

"""
start = row_min - 1 if row_min > 0 else 0
stop = min(row_max, n_obj) if row_max > 0 else n_obj
return range(start, max(start, stop))


def position_seed(ra, dec, ccd):
"""Deterministic RNG seed from an object's sky position (ngmix#796).

Position seeding gives the same object the same RNG stream in each image
branch, provided its sky position falls in the same seed box. It also makes
the result independent of how the tile is split into
``ID_OBJ_MIN``/``ID_OBJ_MAX`` chunks, which is why it is now the only mode.
``ID_OBJ_MIN``/``ID_OBJ_MAX`` row chunks, which is why it is now the only
mode.

Box math (kept exactly as Fabian's issue #796)::

Expand Down Expand Up @@ -711,11 +738,11 @@ class Ngmix(object):
Save output catalogue in batches of this size; detaul is ``-1`` (no
batch save)
id_obj_min : int, optional
First galaxy ID to process, not used if the value is set to ``-1``;
the default is ``-1``
First catalogue row to process (1-based, see :func:`chunk_rows`),
not used if the value is set to ``-1``; the default is ``-1``
id_obj_max : int, optional
Last galaxy ID to process, not used if the value is set to ``-1``;
the default is ``-1``
Last catalogue row to process (1-based, inclusive), not used if the
value is set to ``-1``; the default is ``-1``
centroid_source : {"wcs", "hsm"}, optional
How to place the galaxy Jacobian origin for the centroid prior. The
default ``"wcs"`` places it at the coadd centroid: the sub-pixel
Expand Down Expand Up @@ -1256,11 +1283,11 @@ def process(self):
count_batch = 0
saved_batch_cumul = 0

for i_tile, obj_id in enumerate(tile_cat.obj_id):
if self._id_obj_min > 0 and obj_id < self._id_obj_min:
continue
if self._id_obj_max > 0 and obj_id > self._id_obj_max:
continue
rows = chunk_rows(
len(tile_cat.obj_id), self._id_obj_min, self._id_obj_max
)
for i_tile in rows:
obj_id = tile_cat.obj_id[i_tile]
if id_first == -1:
id_first = obj_id
id_last = obj_id
Expand Down
8 changes: 4 additions & 4 deletions src/shapepipe/modules/ngmix_runner.py
Original file line number Diff line number Diff line change
Expand Up @@ -122,10 +122,10 @@ def ngmix_runner(
# No batch saving
save_batch = -1

# First and last galaxy ID to process. Read via ``getexpanded`` so an
# orchestrator can drive the chunk bounds from environment variables
# (``$SP_NGMIX_ID_OBJ_MIN`` and friends); ``getexpanded`` is the only
# accessor in ShapePipe's config that expands ``$VAR``.
# First and last catalogue row (1-based) to process. Read via
# ``getexpanded`` so an orchestrator can drive the chunk bounds from
# environment variables (``$NGMIX_ROW_MIN`` and friends); ``getexpanded``
# is the only accessor in ShapePipe's config that expands ``$VAR``.
id_obj_min = int(config.getexpanded(module_config_sec, "ID_OBJ_MIN"))
id_obj_max = int(config.getexpanded(module_config_sec, "ID_OBJ_MAX"))

Expand Down
Loading
Loading