Skip to content

Latest commit

 

History

History
140 lines (119 loc) · 7.72 KB

File metadata and controls

140 lines (119 loc) · 7.72 KB

ReconGuard — API

FastAPI app: backend/app/api/main.py. Title: ReconGuard API, version 1.0.0. Run with uvicorn backend.app.api.main:app --reload (or via docker compose up). Interactive docs at /docs (Swagger UI, auto-generated by FastAPI).

CORS is scoped to local dev origins only (http://localhost:3000, http://127.0.0.1:3000) — appropriate for a local demonstration project, not a deployed multi-tenant service. There is no authentication layer.

Endpoints

Method Path Purpose
GET /batches List all batches, most recent first
POST /batches Submit ledger + settlement records; runs the full pipeline synchronously and returns the batch summary
GET /batches/{batch_id} Batch status and summary
GET /batches/{batch_id}/decisions All reconciliation decisions in a batch
GET /decisions/{decision_id} Full confidence-card payload for one decision (probability, decision, risk flags, evidence, raw ledger/settlement records, recomputed feature values, risk-flag explanations)
GET /decisions/{decision_id}/audit A decision's full merged audit history — decision-level events plus any human-review events, which live under a separate REVIEW entity namespace
GET /reviews?status=OPEN&batch_id=... Review queue, optionally filtered by status and/or batch
GET /reviews/{review_id} Review task detail with evidence snapshot, assigned reviewer, resolution timestamp
POST /reviews/{review_id}/approve Human approval — creates a REVIEW_APPROVED audit event, preserves the original ML decision
POST /reviews/{review_id}/reject Human rejection — same preservation guarantee
GET /exceptions?category=...&batch_id=... List exceptions, optionally filtered by category and/or batch, each with a cheap primary_root_cause summary
GET /exceptions/{exception_id} Exception detail: original immutable facts plus a full derived root-cause analysis (observed evidence, interpretation, contributing factors, investigation guidance)
GET /audit/{entity_type}/{entity_id} Full ordered audit history for a batch, decision, or review

Request/response notes

  • POST /batches accepts ledger_records and settlement_records as lists of objects (ledger_id/settlement_id, vendor_name, amount, txn_date, optional reference_id, optional description). reference_id/description accept null and normalize it to an empty string — a real bug found and fixed after submitting an actual batch through the HTTP API (see docs/frontend.md, "Important Failures").
  • POST /reviews/{id}/approve and /reject both take {"reviewer_id": "...", "comment": "..." } (comment optional). A review that no longer exists returns 404; a review that has already been resolved returns 409.
  • GET /decisions/{id} recomputes the candidate's full feature vector on demand (via the same ml.features.group_extractor.extract_group_features function the batch pipeline itself uses — no duplicated logic), so the confidence card can show evidence for every feature, not just the ones persisted as EvidenceRecord rows.
  • GET /exceptions/{id} clearly separates original, immutable, persisted facts (original_category, original_reason, and the decision fields) from derived root-cause analysis (root_cause_analysis.*) — the two are never merged into one field.
  • Batch submission is idempotent by content hash: resubmitting an identical batch returns the existing batch and creates no new decisions.

Batch processing

POST /batches runs the entire pipeline synchronously within the request: candidate generation (blocking V1 + V2) → feature extraction → LightGBM prediction + sigmoid calibration → risk-flag computation → per-candidate routing (auto-match / review / exception) → audit events → batch summary. There is no background task queue — this is appropriate at this project's data volumes but would need an async worker to scale to much larger production batches.

Decisions

A ReconciliationDecision carries the model's raw and calibrated probability, the three-way decision (HIGH_CONFIDENCE_MATCH / NEEDS_REVIEW / LIKELY_NO_MATCH), the separate workflow_state (which can advance further, e.g. to APPROVED_BY_REVIEWER, without ever changing decision itself), relationship_type (one_to_one / one_to_many / many_to_one), and a list of risk flags. All member ledger/settlement record IDs are preserved (never collapsed into a single fake ID), joined with | for structural groups.

Reviews

Only candidates routed to NEEDS_REVIEW create a ReviewTask (status OPEN). Approving or rejecting a review only ever changes workflow_state and creates a new audit event — the underlying ReconciliationDecision.decision and its probabilities are never rewritten, which is verified directly by an automated test and confirmed against real approve/reject calls over HTTP against a live PostgreSQL instance.

Exceptions

Candidates that score below the LOW threshold, or have no viable candidate, become an ExceptionRecord with a root-cause category assigned by a fixed-priority classifier (NO_CANDIDATE_FOUND > STRUCTURAL_AMBIGUITY > HIGH_COMPETITION > INSUFFICIENT_EVIDENCE > LOW_MATCH_CONFIDENCE). The detail endpoint additionally runs a deterministic root-cause engine (backend/app/workflow/exception_intelligence.py) that drills into which evidence dimension (amount/date/vendor/reference) was weakest, with precedence grounded in the model's own feature-importance findings (amount > reference > vendor > date).

Audit / evidence endpoints

GET /audit/{entity_type}/{entity_id} returns every event for a given entity (BATCH, DECISION, or REVIEW), ordered by timestamp (with id as a secondary sort key for deterministic ordering). Every event captures event_id, event_type, actor_type (SYSTEM / MODEL / HUMAN), actor_id, previous_state, new_state, a JSON payload, and a timestamp. Recording is append-only by construction: backend/app/audit.py exposes exactly one write function (record_event()), and there is no update or delete function for events anywhere in the codebase.

GET /decisions/{id}/audit is a convenience endpoint that merges a decision's own events with the events of its associated review (if any) — without it, human review outcomes would be invisible from a decision's timeline, since they're recorded under a separate REVIEW entity namespace.

No LLM agents, no chatbot layer

This API is orchestration, persistence, and human oversight around the existing learned reconciliation engine. No agent framework, no RAG, no autonomous planning was added — the model is already the AI here.

Demo flow, confirmed against real data

  1. POST /batches with a real batch of ledger/settlement records → returns a summary with auto-matched / needs-review / exception / structural-match counts (a real 130-record run against Postgres produced 72 candidates → 57 auto-matched / 1 review / 14 exceptions, 9 structural matches, 23 risk-flagged — see docs/workflow.md).
  2. GET /batches/{batch_id}/decisions → list all decisions, find a structural (one-to-many or many-to-one) one.
  3. GET /decisions/{decision_id} → shows probability, evidence, relationship type, and risk flags (a real KNOWN_LOW_GENERALIZATION flag on a genuine one-to-many case).
  4. GET /reviews?status=OPEN → the open review queue.
  5. POST /reviews/{review_id}/approve → a human resolves it; the underlying ML decision and probability are confirmed unchanged.
  6. GET /decisions/{decision_id}/audit → the merged history: e.g. MODEL_EVALUATED, REVIEW_TASK_CREATED, REVIEW_APPROVED, each a separate immutable event.