NOMAI 2.0: REGALADE host photo-z, lower duration cutoff, Rainbow wavelength fix, new feature - #716
Conversation
Add missing docstrings/doctests across kernel.py, processor.py and slsn_classifier.py (fit_rainbow, fit_salt, statistical_features, ntrend_changes, quiet_model, get_sdss_photoz happy path, add_all_photoz empty-input case), clarify parameter/return descriptions, and fix a couple of doc bugs (stale "30 days" age cut, reversed abs_peak redshift-uncertainty ordering, missing ntrends entry in statistical_features' Returns). Also make array-output doctests immune to numpy print-formatting differences by comparing values with np.testing.assert_allclose instead of matching printed reprs. Adds ntrend_changes as a new statistical feature (number of abrupt trend reversals in the light curve), which was not present in the previous version. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Reformats pd.DataFrame({...}) / lambda-wrapped list-comprehension calls
to ruff's preview "hug brackets" style, and normalizes remaining
single-quoted strings to double quotes, so both files pass
`ruff format --preview --check`.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The linter workflow (.github/workflows/linter.yml) runs `ruff format --check .` without --preview. The previous commit formatted with --preview, which applies not-yet-stable rules (e.g. "hug brackets" for single-arg calls wrapping a list/dict literal) that the stable formatter disagrees with, so it failed CI. Reformat with plain `ruff format` to match what CI actually checks. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
luminosity_distance() was called once per band per redshift offset (up to 6 calls); it doesn't depend on band, so compute it once for the 3 redshift offsets and reuse across bands. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Fixes ruff format --check failure flagged by CI on the previous commit. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Removes get_sdss_photoz/add_all_photoz's live SDSS SkyServer query and replaces it with an ellipse-based crossmatch against the REGALADE galaxy catalog, using the same matching logic and parameters (dlr_factor=1.25, nbins=20) as the training pipeline's create_photoz_table.py -- so the photo-z used at inference time is produced the same way as at training time. The matching logic is vendored from ellipse_xmatch (https://github.com/htranin/ellipse_xmatch) rather than added as a dependency, since fink-science runs in production. The catalog itself (fink_science/data/catalogs/regalade_minimal_ZTF.fits, ~1.5GB) is not committed -- it exceeds GitHub's 100MB push limit and can't be bundled in the pip package like the other catalogs here. See fink_science/data/catalogs/README.md; distribution story still TODO. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Replaces the vendored _ellipse_separation()/manual R1-binning loop in get_regalade_photoz with a direct call to ellipse_xmatch.crossmatch_ellipses, avoiding a second copy of the same matching logic to keep in sync with https://github.com/htranin/ellipse_xmatch. Declared as a git dependency in setup.py's install_requires. Note: the fink-deps-sentinel-ztf/-rubin Docker images used by CI (.github/workflows/run_test.yml) are built/maintained outside this repo, so they will also need ellipse-xmatch installed for the test suite to import slsn_classifier successfully. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
… a feature Previously ebv was passed to the classifier as a raw feature, hoping it would learn extinction effects from it. This corrects cflux/csigflux/ cmagpsf for Milky Way extinction (deredden_lightcurve) before Rainbow/ salt fitting and statistical features, so temperature, color and amplitude features are computed on the intrinsic light curve directly -- no feature needed, and light curves become directly comparable across different lines of sight. extract_features no longer outputs an ebv column. abs_peak's signature is left untouched (it's called elsewhere, e.g. fink-broker/bin/ztf/archive_slsn_candidates.py, with a real ebv value on non-dereddened magnitudes); processor.py now passes ebv=0 since peak_mag_g/peak_mag_r are already corrected by the time they reach it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
REGALADE is now crossmatched against alerts upstream in Fink's own pipeline (fink-broker's apply_all_xmatch, a circular/nearest-neighbour match via xmatch_cds). get_regalade_photoz no longer searches a local ~1.5GB REGALADE catalog copy -- it refines that single candidate with a DLR-ellipse separation test instead, using the same gnomonic-projection formula as fink_science.ztf.xmatch.processor.xmatch_regalade. Drops the ellipse-xmatch runtime dependency and kernel.regalade_path/ regalade_nbins (catalog loading is gone); kernel.regalade_dlr_factor is kept, still needed to scale the ellipse test. Note: since the ellipse test can only reject a circular-match candidate, never recover one the circular search missed, REGALADE-derived photoz coverage drops sharply (~100% -> ~22% on the training set, measured locally) compared to the previous full-catalog DLR crossmatch used to build photz_table.parquet. NOMAI.joblib will need retraining against the new, coarser matching method.
Retrained against the coarse 20'' circular + DLR-ellipse refinement REGALADE method (76.9% host photo-z coverage on the training set, vs 81.0% under the old full-catalog precise crossmatch -- honest bootstrap CV F1 drops from 0.7202 to 0.7131, a small, expected difference). Also picks up brightness_persistence/shape_irregularity, two features validated in 2.0-experimentation and already live in slsn_classifier.extract_features, but missing from the local training pipeline's cached feature set until now.
|
Update (2026-09-14): REGALADE now sourced from the alert, not a local catalog The REGALADE integration described above has changed since this PR was opened. REGALADE was merged into fink-broker's own crossmatch pipeline (xmatch_cds, a live circular/nearest-neighbour match against VizieR, the same mechanism used for Simbad/Gaia/VSX/etc.) rather than the offline full-catalog approach this PR originally proposed. As a result: get_regalade_photoz no longer reads a local REGALADE catalog file. It now refines the REGALADE candidate fink-broker's apply_all_xmatch already attached to the alert (regalade_ra, regalade_dec, R1, R2, PA, z, ezin), applying the same DLR-ellipse separation test as before, just against that single candidate instead of searching the full catalog. NOMAI.joblib retrained on the new method (honest bootstrap CV F1: 0.7131, vs. 0.7202 under the old method — a ~1% relative difference). |
Alongside the existing not-bright-enough veto: probability is now also forced to 0 for sources with too many trend reversals (too variable photometry) or too long a duration (likely AGN/bad photometry). Both are tuned on the real alert stream, not learned from training, since the already-classified training sample doesn't exhibit this contamination.
# Conflicts: # .gitignore # fink_science/ztf/superluminous/kernel.py # fink_science/ztf/superluminous/slsn_classifier.py
Replaces the migrated (2.1.4 -> 3.4.1 JSON-roundtrip) model with one trained directly under xgboost 3.4.1, matching production. Same data, same features; F1 unchanged (0.7134 vs 0.7131 migrated -- noise-level difference).
JulienPeloton
left a comment
There was a problem hiding this comment.
Thanks @erusseil -- code is clean :-)
For future reference:
- Add SLSN candidates in the test sample
- Add call to
xmatch_cdsbefore performing classifiation, to fully test the pipeline with real values.
Linked to issue(s): #715
If this is a new release, did you issue the corresponding schema in fink-client?
No schema change:
superluminous_score's output type/contract is unchanged.What changes were proposed in this pull request?
This is the full "NOMAI 1.0 -> 2.0" rework of the superluminous supernovae
(SLSN) classifier, covering 10 commits:
kernel.min_duration), sothe classifier scores sources earlier in their evolution.
(
kernel.band_wave_aa):{1: 4770.0, 2: 6231.0}->{1: 4746.48, 2: 6366.38}(correct SVO Filter Profile Service values forthe g/r bands).
ntrendsfeature (slsn_classifier.ntrend_changes): countsabrupt trend reversals in the light curve (a real SLSN should rise/decay
smoothly), meant to flag noisy/bogus-looking light curves.
-Extra post-hoc cuts Using n_trends and duration. These cuts were already used inside fink-broker. They are not directly merged in the main pipeline
brightness persistenceandshape_irregularityfeatures.Simple statistical features, cheap to compute and very informative to find SLSNe.
Discovered with symbolic regression.
shape_irregularityfeature : a simple statistical feature, cheap to computevery informative to distinguish SLSNe. Discovered through symbolic regression.
ntrendsfeature (slsn_classifier.ntrend_changes): countsabrupt trend reversals in the light curve (a real SLSN should rise/decay
smoothly), meant to flag noisy/bogus-looking light curves.
-- the main change in this PR.
ebvas a raw feature: previously
ebvwas passed directly to theclassifier, hoping it would implicitly learn extinction effects from it.
New
slsn_classifier.deredden_lightcurvecorrectscflux/csigflux/cmagpsffor Milky Way extinction, band by band, before Rainbow/saltfitting and statistical features, so temperature, color and peak
magnitude features are computed on the intrinsic light curve directly --
no feature needed, and light curves become directly comparable across
different lines of sight.
extract_featuresno longer outputs anebvcolumn.
abs_peak's signature is left untouched, since it's calledelsewhere (e.g.
fink-broker/bin/ztf/archive_slsn_candidates.py) with areal
ebvvalue on non-dereddened magnitudes;processor.pynow passesebv=0sincepeak_mag_g/peak_mag_rare already corrected by thetime they reach it.
abs_peakoptimization:cosmo.luminosity_distance()was beingcalled once per band per redshift offset (up to 6 calls); it doesn't
depend on band, so it's now computed once for the 3 redshift offsets
and reused across bands. Purely a speed change.
kernel.py,processor.pyandslsn_classifier.py, and two ruff-formatting-onlycommits.
superluminous_classifier.joblib->NOMAI.joblib,kernel.classifier_pathupdated accordingly), retrainedtwice: once after the initial rework, and again after the Milky Way
extinction fix above, on a fresh feature extraction.
REGALADE host photo-z (the main change)
The classifier uses a host galaxy photo-z to compute a physical upper
bound on absolute magnitude (
slsn_classifier.abs_peak), and forces theSLSN probability to 0 when a source can't plausibly be superluminous even
in the brightest case. This photo-z previously came from a live query to
the SDSS SkyServer (
get_sdss_photoz). This PR replaces it with anoffline ellipse-based crossmatch against the REGALADE galaxy catalog
(Tranin et al., https://github.com/htranin/regalade), which has much wider
sky coverage and a more accurate host association method (directional
light radius ellipse matching) than a simple nearest-neighbor SDSS query.
On our own labeled SLSN candidate sample, REGALADE finds a host for 80.6%
of objects vs. 52.1% for SDSS, and among objects matched by both, ~20%
differ by more than 0.05 in redshift (some by close to 1, i.e. clearly a
different galaxy) -- so this isn't just a coverage improvement, it also
fixes host misassociations that were silently corrupting the physical cut
in both directions (wrongly suppressing true SLSNe and wrongly letting
non-SLSNe through).
Concretely:
slsn_classifier.get_sdss_photoz/add_all_photozno longer make anetwork call to SDSS.
add_all_photozkeeps the same signature/contract(still adds
photoz/photozerrcolumns), soprocessor.py's call siteis unchanged.
slsn_classifier.get_regalade_photoz(ra, dec)crossmatches againstthe REGALADE catalog using
ellipse_xmatch.crossmatch_ellipses(newdependency, declared in
setup.py), with the exact same matchingparameters (
dlr_factor=1.25,nbins=20) used to build the trainingset, so inference-time photo-z is produced the same way as training-time
photo-z. The catalog is lazily loaded and cached at module level so a
Spark worker only pays the FITS read once, not once per micro-batch.
kernel.regalade_path/regalade_dlr_factor/regalade_nbinsconstants. See
fink_science/data/catalogs/README.mdfor what thecatalog file is and how it's derived.
How was this patch tested?
slsn_classifier.py,kernel.pyandprocessor.pypass locally, including the new/changed ones (
get_regalade_photoz,add_all_photoz,abs_peak,deredden_lightcurve,ntrend_changes,statistical_features), including the full Spark-basedextract_featuresdoctest.
get_regalade_photozverified to reproduce the training pipeline'sphoto-z values exactly (to floating point) on known objects.
abs_peak's optimization verified against the pre-optimizationimplementation over 2000 random trials (max diff ~1e-13), including NaN
propagation.
ruff format --check/ruff checkclean on every commit.hyperparameter search, training) was re-run end to end after the Milky
Way extinction fix, on a fresh feature extraction.
(needs the REGALADE catalog file and
ellipse-xmatchinstalled, seeabove)