Skip to content

Add sparse.COO support to harmonize_missing_values and infer_feature_types - #306

Draft
sueoglu wants to merge 8 commits into
mainfrom
fix/harmonize-missing-sparse
Draft

sueoglu wants to merge 8 commits into
mainfrom
fix/harmonize-missing-sparse

Conversation

@sueoglu

@sueoglu sueoglu commented Sep 18, 2026 •

Copy link
Copy Markdown
Collaborator

closes #305

Both functions previously either crashed or silently no-op'd on a sparse.COO X/layer. This adds proper support to both, without ever densifying the array.

infer_feature_types

  • sparse.COO is always numeric under the binsparse spec this project follows, so classification per variable only needs to decide "is this column exactly {0, 1}?" which is answerable directly from X.coords/X.data/X.fill_value
  • _detect_feature_types_sparse_coo helper implements this; infer_feature_types branches to it for sparse.COO input and keeps the existing dataframe-based path unchanged for everything else.

harmonize_missing_values

  • For sparse.COO, the implicit 0 fill value is now treated as missing (swapped to NaN) by default (done by flipping the array's fill_value from 0 to NaN, so the array stays sparse)
  • vars argument: variable names whose 0 is a real measured value, not missing. Previously-implicit zeros of these columns are materialized as explicit stored 0 entries (so they dont get changed to the new NaN fill value); every other column stays sparse
  • Left untouched: boolean arrays, scipy.sparse

Tests: added coverage in test_feature_types.py for 2D/3D sparse.COO harmonization, vars exclusion, an unknown-var error, scipy.sparse being unaffected, and sparse.COO feature-type inference (including binary detection and the all-NaN-column error). Two pre-existing IO round-trip tests (test_h5ed.py/test_zarr.py) updated to pass harmonize_missing_values=False, since they test binsparse encoding fidelity specifically and predate this behavior change.

@sueoglu
sueoglu requested a review from eroell September 18, 2026 08:49

@eroell eroell left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please add a benchmark with compute time and memory consumption in the PR comment, where a large dense 3D and a large sparse 3D, and also the 2D sparse and dense arrays are tested.

The sparse COO version should run for a ~500k x 1000k x 90 tensor of about 1% data density.

Comment thread src/ehrdata/_feature_types.py Outdated
Comment thread src/ehrdata/_feature_types.py Outdated
Comment thread src/ehrdata/_feature_types.py Outdated
Comment thread tests/test_feature_types.py
Comment thread src/ehrdata/_feature_types.py
@sueoglu

sueoglu commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Benchmark with compute time and memory usage

Used time.perf_counter() for wall-clock time, tracemalloc for peak Python-allocator memory. I wasnt able to login to cluster and couldnt run it with a large-scale case representing EHR dimensions (~500k × 1000k × 90).

case shape build (s) build peak (MB) infer_feature_types (s) infer peak (MB) harmonize_missing_values (s) harmonize peak (MB) nnz
dense_2d 2000×3000 0.09 72.0 30.75 6.2 0.00 0.0 6,000,000
sparse_2d 2000×3000 @1% 0.02 4.6 0.03 0.4 0.00 1.6 59,712
dense_3d 500×800×90 0.41 432.0 181.75 150.1 0.00 0.0 36,000,000
sparse_3d 500×800×90 @1% 0.10 36.3 0.03 1.4 0.01 9.3 358,173
  • infer_feature_types on dense 3D takes ~182s vs ~0.03s for sparse.COO at identical shape

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

sparse.COO support for harmonize_missing_values and infer_feature_types

2 participants