Describe the bug / problem
Alert and Provider objects expose no status conditions (and essentially no status at all) that operators can use to answer:
- Was the last event delivered?
- When was the last successful dispatch?
- What was the last provider error (401, 429, network, etc.)?
Live observation on notification-controller v1.9.2 (API notification.toolkit.fluxcd.io/v1beta3):
Alert.status is completely empty: status: {} (no keys).
Provider.status is empty the same way.
- Events are sometimes delivered to the chat provider, so the pipeline is not dead — but the CRs themselves give zero operational signal.
kubectl get alert effectively shows Age only — no Ready column, no lastSentAt, no lastError.
Debugging provider failures (Telegram 401 auth, 429 rate limit, DNS flaps, etc.) requires scraping notification-controller logs. That is a poor operator experience for a component whose whole job is reliability of outbound notifications.
This is distinct from #1370 (flood volume / transient DNS cascades / exclusionList). That issue is about how many events fire. This issue is about observability of the Alert and Provider objects themselves.
Steps to reproduce
- Configure a Provider (e.g.
type: telegram) and an Alert v1beta3 with eventSeverity: error and broad eventSources (GitRepository / Kustomization / HelmRelease / HelmRepository / OCIRepository / ImageUpdateAutomation wildcards are enough to generate traffic).
- Confirm some events reach the provider (or intentionally break credentials / hit rate limits).
- Inspect the CRs:
kubectl get alert -A
# Age only; no Ready / last error columns
kubectl get alert <name> -n <ns> -o yaml
# status: {}
kubectl get provider <name> -n <ns> -o yaml
# status empty / no conditions
kubectl describe alert does not surface last Telegram (or other provider) error either — logs remain the only source of truth.
Expected behavior
Status that answers “is notification path healthy?” without log diving, for example:
| Field / condition |
Purpose |
status.conditions Ready |
Provider config valid; controller can authenticate / reach endpoint |
lastHandledEvent (or similar) |
Identity of last event considered (name/kind/message hash) |
lastDispatchTime |
When the last successful send completed |
lastDispatchError |
Last failure message (redacted — no tokens, no full chat IDs, no secrets) |
Printer columns for Ready + last error age would make kubectl get alert,provider useful for triage.
Even a minimal Ready condition (config OK vs last dispatch failed) would be a large step up from status: {}.
Environment
- Kubernetes: single-node (generic)
- Flux CLI: v2.9.4
- Distribution: flux-v2.9.1
- notification-controller: v1.9.2
- Alert API: v1beta3,
eventSeverity: error
- Provider: telegram (channel ID omitted)
Alert.spec.exclusionList already populated with generic regexes for connection refused / dial tcp lookup / empty tag DB / scan failed patterns — still no status reflecting exclusions vs sends
Additional context / non-goals
Code of Conduct
Describe the bug / problem
AlertandProviderobjects expose no status conditions (and essentially no status at all) that operators can use to answer:Live observation on notification-controller v1.9.2 (API
notification.toolkit.fluxcd.io/v1beta3):Alert.statusis completely empty:status: {}(no keys).Provider.statusis empty the same way.kubectl get alerteffectively shows Age only — no Ready column, no lastSentAt, no lastError.Debugging provider failures (Telegram 401 auth, 429 rate limit, DNS flaps, etc.) requires scraping notification-controller logs. That is a poor operator experience for a component whose whole job is reliability of outbound notifications.
This is distinct from #1370 (flood volume / transient DNS cascades / exclusionList). That issue is about how many events fire. This issue is about observability of the Alert and Provider objects themselves.
Steps to reproduce
type: telegram) and an Alert v1beta3 witheventSeverity: errorand broadeventSources(GitRepository / Kustomization / HelmRelease / HelmRepository / OCIRepository / ImageUpdateAutomation wildcards are enough to generate traffic).kubectl describe alertdoes not surface last Telegram (or other provider) error either — logs remain the only source of truth.Expected behavior
Status that answers “is notification path healthy?” without log diving, for example:
status.conditionsReadylastHandledEvent(or similar)lastDispatchTimelastDispatchErrorPrinter columns for Ready + last error age would make
kubectl get alert,provideruseful for triage.Even a minimal
Readycondition (config OK vs last dispatch failed) would be a large step up fromstatus: {}.Environment
eventSeverity: errorAlert.spec.exclusionListalready populated with generic regexes for connection refused / dial tcp lookup / empty tag DB / scan failed patterns — still no status reflecting exclusions vs sendsAdditional context / non-goals
exclusionListvs actually sent — that metric would also help tune regexes safely.Code of Conduct