Skip to content

NOMAI 2.0: REGALADE host photo-z, lower duration cutoff, Rainbow wavelength fix, new feature - #716

Merged
JulienPeloton merged 20 commits into
astrolabsoftware:masterfrom
erusseil:regalade-photoz
Sep 17, 2026
Merged

JulienPeloton merged 20 commits into
astrolabsoftware:masterfrom
erusseil:regalade-photoz

Conversation

@erusseil

@erusseil erusseil commented Jul 29, 2026 •

Copy link
Copy Markdown
Collaborator

Linked to issue(s): #715

If this is a new release, did you issue the corresponding schema in fink-client?

No schema change: superluminous_score's output type/contract is unchanged.

What changes were proposed in this pull request?

This is the full "NOMAI 1.0 -> 2.0" rework of the superluminous supernovae
(SLSN) classifier, covering 10 commits:

  • Lower the age cutoff from 30 to 20 days (kernel.min_duration), so
    the classifier scores sources earlier in their evolution.
  • Fix the Rainbow fit's ZTF effective wavelengths
    (kernel.band_wave_aa): {1: 4770.0, 2: 6231.0} ->
    {1: 4746.48, 2: 6366.38} (correct SVO Filter Profile Service values for
    the g/r bands).
  • New ntrends feature (slsn_classifier.ntrend_changes): counts
    abrupt trend reversals in the light curve (a real SLSN should rise/decay
    smoothly), meant to flag noisy/bogus-looking light curves.
    -Extra post-hoc cuts Using n_trends and duration. These cuts were already used inside fink-broker. They are not directly merged in the main pipeline
  • New brightness persistence and shape_irregularity features.
    Simple statistical features, cheap to compute and very informative to find SLSNe.
    Discovered with symbolic regression.
  • New shape_irregularity feature : a simple statistical feature, cheap to compute
    very informative to distinguish SLSNe. Discovered through symbolic regression.
  • New ntrends feature (slsn_classifier.ntrend_changes): counts
    abrupt trend reversals in the light curve (a real SLSN should rise/decay
    smoothly), meant to flag noisy/bogus-looking light curves.
  • Replace the SDSS host photo-z with a REGALADE crossmatch (see below)
    -- the main change in this PR.
  • Correct light curves for Milky Way extinction instead of using ebv
    as a raw feature
    : previously ebv was passed directly to the
    classifier, hoping it would implicitly learn extinction effects from it.
    New slsn_classifier.deredden_lightcurve corrects cflux/csigflux/
    cmagpsf for Milky Way extinction, band by band, before Rainbow/salt
    fitting and statistical features, so temperature, color and peak
    magnitude features are computed on the intrinsic light curve directly --
    no feature needed, and light curves become directly comparable across
    different lines of sight. extract_features no longer outputs an ebv
    column. abs_peak's signature is left untouched, since it's called
    elsewhere (e.g. fink-broker/bin/ztf/archive_slsn_candidates.py) with a
    real ebv value on non-dereddened magnitudes; processor.py now passes
    ebv=0 since peak_mag_g/peak_mag_r are already corrected by the
    time they reach it.
  • abs_peak optimization: cosmo.luminosity_distance() was being
    called once per band per redshift offset (up to 6 calls); it doesn't
    depend on band, so it's now computed once for the 3 redshift offsets
    and reused across bands. Purely a speed change.
  • Various docstring/doctest improvements across kernel.py,
    processor.py and slsn_classifier.py, and two ruff-formatting-only
    commits.
  • The trained model itself (superluminous_classifier.joblib ->
    NOMAI.joblib, kernel.classifier_path updated accordingly), retrained
    twice: once after the initial rework, and again after the Milky Way
    extinction fix above, on a fresh feature extraction.

REGALADE host photo-z (the main change)

The classifier uses a host galaxy photo-z to compute a physical upper
bound on absolute magnitude (slsn_classifier.abs_peak), and forces the
SLSN probability to 0 when a source can't plausibly be superluminous even
in the brightest case. This photo-z previously came from a live query to
the SDSS SkyServer (get_sdss_photoz). This PR replaces it with an
offline ellipse-based crossmatch against the REGALADE galaxy catalog
(Tranin et al., https://github.com/htranin/regalade), which has much wider
sky coverage and a more accurate host association method (directional
light radius ellipse matching) than a simple nearest-neighbor SDSS query.

On our own labeled SLSN candidate sample, REGALADE finds a host for 80.6%
of objects vs. 52.1% for SDSS, and among objects matched by both, ~20%
differ by more than 0.05 in redshift (some by close to 1, i.e. clearly a
different galaxy) -- so this isn't just a coverage improvement, it also
fixes host misassociations that were silently corrupting the physical cut
in both directions (wrongly suppressing true SLSNe and wrongly letting
non-SLSNe through).

Concretely:

  • slsn_classifier.get_sdss_photoz/add_all_photoz no longer make a
    network call to SDSS. add_all_photoz keeps the same signature/contract
    (still adds photoz/photozerr columns), so processor.py's call site
    is unchanged.
  • New slsn_classifier.get_regalade_photoz(ra, dec) crossmatches against
    the REGALADE catalog using ellipse_xmatch.crossmatch_ellipses (new
    dependency, declared in setup.py), with the exact same matching
    parameters (dlr_factor=1.25, nbins=20) used to build the training
    set, so inference-time photo-z is produced the same way as training-time
    photo-z. The catalog is lazily loaded and cached at module level so a
    Spark worker only pays the FITS read once, not once per micro-batch.
  • New kernel.regalade_path/regalade_dlr_factor/regalade_nbins
    constants. See fink_science/data/catalogs/README.md for what the
    catalog file is and how it's derived.

How was this patch tested?

  • All doctests in slsn_classifier.py, kernel.py and processor.py
    pass locally, including the new/changed ones (get_regalade_photoz,
    add_all_photoz, abs_peak, deredden_lightcurve, ntrend_changes,
    statistical_features), including the full Spark-based extract_features
    doctest.
  • get_regalade_photoz verified to reproduce the training pipeline's
    photo-z values exactly (to floating point) on known objects.
  • abs_peak's optimization verified against the pre-optimization
    implementation over 2000 random trials (max diff ~1e-13), including NaN
    propagation.
  • ruff format --check / ruff check clean on every commit.
  • The full training pipeline (feature extraction, REGALADE crossmatch,
    hyperparameter search, training) was re-run end to end after the Milky
    Way extinction fix, on a fresh feature extraction.
  • Not yet run through the full Spark test suite in this environment
    (needs the REGALADE catalog file and ellipse-xmatch installed, see
    above)

erusseil and others added 15 commits July 29, 2026 15:43
Add missing docstrings/doctests across kernel.py, processor.py and
slsn_classifier.py (fit_rainbow, fit_salt, statistical_features,
ntrend_changes, quiet_model, get_sdss_photoz happy path, add_all_photoz
empty-input case), clarify parameter/return descriptions, and fix a
couple of doc bugs (stale "30 days" age cut, reversed abs_peak
redshift-uncertainty ordering, missing ntrends entry in
statistical_features' Returns). Also make array-output doctests
immune to numpy print-formatting differences by comparing values with
np.testing.assert_allclose instead of matching printed reprs.

Adds ntrend_changes as a new statistical feature (number of abrupt
trend reversals in the light curve), which was not present in the
previous version.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Reformats pd.DataFrame({...}) / lambda-wrapped list-comprehension calls
to ruff's preview "hug brackets" style, and normalizes remaining
single-quoted strings to double quotes, so both files pass
`ruff format --preview --check`.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The linter workflow (.github/workflows/linter.yml) runs
`ruff format --check .` without --preview. The previous commit
formatted with --preview, which applies not-yet-stable rules (e.g.
"hug brackets" for single-arg calls wrapping a list/dict literal) that
the stable formatter disagrees with, so it failed CI. Reformat with
plain `ruff format` to match what CI actually checks.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
luminosity_distance() was called once per band per redshift offset
(up to 6 calls); it doesn't depend on band, so compute it once for
the 3 redshift offsets and reuse across bands.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Fixes ruff format --check failure flagged by CI on the previous commit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Removes get_sdss_photoz/add_all_photoz's live SDSS SkyServer query and
replaces it with an ellipse-based crossmatch against the REGALADE galaxy
catalog, using the same matching logic and parameters (dlr_factor=1.25,
nbins=20) as the training pipeline's create_photoz_table.py -- so the
photo-z used at inference time is produced the same way as at training
time. The matching logic is vendored from ellipse_xmatch
(https://github.com/htranin/ellipse_xmatch) rather than added as a
dependency, since fink-science runs in production.

The catalog itself (fink_science/data/catalogs/regalade_minimal_ZTF.fits,
~1.5GB) is not committed -- it exceeds GitHub's 100MB push limit and
can't be bundled in the pip package like the other catalogs here. See
fink_science/data/catalogs/README.md; distribution story still TODO.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Replaces the vendored _ellipse_separation()/manual R1-binning loop in
get_regalade_photoz with a direct call to
ellipse_xmatch.crossmatch_ellipses, avoiding a second copy of the same
matching logic to keep in sync with https://github.com/htranin/ellipse_xmatch.
Declared as a git dependency in setup.py's install_requires.

Note: the fink-deps-sentinel-ztf/-rubin Docker images used by CI
(.github/workflows/run_test.yml) are built/maintained outside this repo,
so they will also need ellipse-xmatch installed for the test suite to
import slsn_classifier successfully.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
… a feature

Previously ebv was passed to the classifier as a raw feature, hoping it
would learn extinction effects from it. This corrects cflux/csigflux/
cmagpsf for Milky Way extinction (deredden_lightcurve) before Rainbow/
salt fitting and statistical features, so temperature, color and
amplitude features are computed on the intrinsic light curve directly --
no feature needed, and light curves become directly comparable across
different lines of sight.

extract_features no longer outputs an ebv column. abs_peak's signature
is left untouched (it's called elsewhere, e.g.
fink-broker/bin/ztf/archive_slsn_candidates.py, with a real ebv value on
non-dereddened magnitudes); processor.py now passes ebv=0 since
peak_mag_g/peak_mag_r are already corrected by the time they reach it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
REGALADE is now crossmatched against alerts upstream in Fink's own
pipeline (fink-broker's apply_all_xmatch, a circular/nearest-neighbour
match via xmatch_cds). get_regalade_photoz no longer searches a local
~1.5GB REGALADE catalog copy -- it refines that single candidate with a
DLR-ellipse separation test instead, using the same gnomonic-projection
formula as fink_science.ztf.xmatch.processor.xmatch_regalade.

Drops the ellipse-xmatch runtime dependency and kernel.regalade_path/
regalade_nbins (catalog loading is gone); kernel.regalade_dlr_factor is
kept, still needed to scale the ellipse test.

Note: since the ellipse test can only reject a circular-match candidate,
never recover one the circular search missed, REGALADE-derived photoz
coverage drops sharply (~100% -> ~22% on the training set, measured
locally) compared to the previous full-catalog DLR crossmatch used to
build photz_table.parquet. NOMAI.joblib will need retraining against the
new, coarser matching method.
Retrained against the coarse 20'' circular + DLR-ellipse refinement
REGALADE method (76.9% host photo-z coverage on the training set, vs
81.0% under the old full-catalog precise crossmatch -- honest bootstrap
CV F1 drops from 0.7202 to 0.7131, a small, expected difference).

Also picks up brightness_persistence/shape_irregularity, two features
validated in 2.0-experimentation and already live in
slsn_classifier.extract_features, but missing from the local training
pipeline's cached feature set until now.
@erusseil

Copy link
Copy Markdown
Collaborator Author

Update (2026-09-14): REGALADE now sourced from the alert, not a local catalog

The REGALADE integration described above has changed since this PR was opened. REGALADE was merged into fink-broker's own crossmatch pipeline (xmatch_cds, a live circular/nearest-neighbour match against VizieR, the same mechanism used for Simbad/Gaia/VSX/etc.) rather than the offline full-catalog approach this PR originally proposed. As a result:

get_regalade_photoz no longer reads a local REGALADE catalog file. It now refines the REGALADE candidate fink-broker's apply_all_xmatch already attached to the alert (regalade_ra, regalade_dec, R1, R2, PA, z, ezin), applying the same DLR-ellipse separation test as before, just against that single candidate instead of searching the full catalog.
No more local regalade_minimal_ZTF.fits dependency, and the ellipse-xmatch runtime dependency is dropped from setup.py.

NOMAI.joblib retrained on the new method (honest bootstrap CV F1: 0.7131, vs. 0.7202 under the old method — a ~1% relative difference).
Depends on fink-broker's REGALADE xmatch_cds block already being live, and ideally on fink-broker#1239 merging (20″ radius) for the coverage numbers above to hold; at the original 1.2″ radius

Alongside the existing not-bright-enough veto: probability is now also
forced to 0 for sources with too many trend reversals (too variable
photometry) or too long a duration (likely AGN/bad photometry). Both are
tuned on the real alert stream, not learned from training, since the
already-classified training sample doesn't exhibit this contamination.
# Conflicts:
#	.gitignore
#	fink_science/ztf/superluminous/kernel.py
#	fink_science/ztf/superluminous/slsn_classifier.py
Replaces the migrated (2.1.4 -> 3.4.1 JSON-roundtrip) model with one
trained directly under xgboost 3.4.1, matching production. Same data,
same features; F1 unchanged (0.7134 vs 0.7131 migrated -- noise-level
difference).

@JulienPeloton JulienPeloton left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @erusseil -- code is clean :-)

For future reference:

  • Add SLSN candidates in the test sample
  • Add call to xmatch_cds before performing classifiation, to fully test the pipeline with real values.

@JulienPeloton
JulienPeloton merged commit ad55d63 into astrolabsoftware:master Sep 17, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants