Skip to content

Alert and Provider have no status conditions — cannot see last dispatch or last error #1371

Description

@RaviTharuma

Describe the bug / problem

Alert and Provider objects expose no status conditions (and essentially no status at all) that operators can use to answer:

  • Was the last event delivered?
  • When was the last successful dispatch?
  • What was the last provider error (401, 429, network, etc.)?

Live observation on notification-controller v1.9.2 (API notification.toolkit.fluxcd.io/v1beta3):

  • Alert.status is completely empty: status: {} (no keys).
  • Provider.status is empty the same way.
  • Events are sometimes delivered to the chat provider, so the pipeline is not dead — but the CRs themselves give zero operational signal.
  • kubectl get alert effectively shows Age only — no Ready column, no lastSentAt, no lastError.

Debugging provider failures (Telegram 401 auth, 429 rate limit, DNS flaps, etc.) requires scraping notification-controller logs. That is a poor operator experience for a component whose whole job is reliability of outbound notifications.

This is distinct from #1370 (flood volume / transient DNS cascades / exclusionList). That issue is about how many events fire. This issue is about observability of the Alert and Provider objects themselves.

Steps to reproduce

  1. Configure a Provider (e.g. type: telegram) and an Alert v1beta3 with eventSeverity: error and broad eventSources (GitRepository / Kustomization / HelmRelease / HelmRepository / OCIRepository / ImageUpdateAutomation wildcards are enough to generate traffic).
  2. Confirm some events reach the provider (or intentionally break credentials / hit rate limits).
  3. Inspect the CRs:
kubectl get alert -A
# Age only; no Ready / last error columns

kubectl get alert <name> -n <ns> -o yaml
# status: {}

kubectl get provider <name> -n <ns> -o yaml
# status empty / no conditions
  1. kubectl describe alert does not surface last Telegram (or other provider) error either — logs remain the only source of truth.

Expected behavior

Status that answers “is notification path healthy?” without log diving, for example:

Field / condition Purpose
status.conditions Ready Provider config valid; controller can authenticate / reach endpoint
lastHandledEvent (or similar) Identity of last event considered (name/kind/message hash)
lastDispatchTime When the last successful send completed
lastDispatchError Last failure message (redacted — no tokens, no full chat IDs, no secrets)

Printer columns for Ready + last error age would make kubectl get alert,provider useful for triage.

Even a minimal Ready condition (config OK vs last dispatch failed) would be a large step up from status: {}.

Environment

  • Kubernetes: single-node (generic)
  • Flux CLI: v2.9.4
  • Distribution: flux-v2.9.1
  • notification-controller: v1.9.2
  • Alert API: v1beta3, eventSeverity: error
  • Provider: telegram (channel ID omitted)
  • Alert.spec.exclusionList already populated with generic regexes for connection refused / dial tcp lookup / empty tag DB / scan failed patterns — still no status reflecting exclusions vs sends

Additional context / non-goals

Code of Conduct

  • I agree to follow this project's Code of Conduct

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions