Skip to content

Evaluate and implement comprehensive API monitoring #3816

Description

@anth-volk

Problem

PolicyEngine needs functional monitoring across substantially more API behavior than the Better Stack free tier can cover. Existing host-level checks do not detect failures limited to particular route groups, and realistic calculation monitoring may require authenticated requests, response validation, and polling asynchronous operations.

Adding paid Better Stack coverage immediately would commit us to a service and recurring cost before we have compared it with an internally operated solution or other hosted products.

This issue supersedes #2155 and should preserve its required coverage while broadening the work to include option evaluation and platform design.

Objective

Evaluate and implement a cost-effective monitoring approach for PolicyEngine APIs. The selected approach may be:

  • a small custom service operated by PolicyEngine;
  • an existing self-hosted monitoring product; or
  • a paid hosted monitoring service, including Better Stack, if its cost and capabilities are the best fit.

Do not select an implementation until the alternatives have been compared against explicit requirements.

Investigation

  1. Inventory the API endpoints and complete user operations that need monitoring, including metadata, household, policy, user-profile, household-calculation, economy, and budget-window behavior.
  2. Define request frequency, acceptable response time, failure conditions, environments, and required geographic coverage.
  3. Determine requirements for:
    • public and authenticated requests;
    • safe test data and credential storage;
    • multi-request and asynchronous operations;
    • response status, content, and schema validation;
    • Slack failure and recovery notifications, including duplicate suppression;
    • execution history, latency reporting, and diagnostic evidence;
    • configuration stored in version control;
    • maintenance ownership, security updates, reliability, and data retention.
  4. Compare custom, self-hosted, and hosted approaches. Document expected recurring cost, implementation effort, operational burden, service limits, and vendor dependence.
  5. Prototype the leading candidates where documentation alone does not establish whether they satisfy the requirements.

Implementation

After recording the decision, implement the selected approach and migrate or supplement the existing checks. Monitoring should exercise representative successful operations with safe inputs rather than only checking API root responses.

Completion criteria

  • The requirements and endpoint coverage inventory are documented.
  • The alternatives and their expected costs and maintenance requirements are compared in writing.
  • The selected approach and reasons for selecting it are documented.
  • Required public, authenticated, and asynchronous API operations can be checked where applicable.
  • Failures and recoveries produce actionable Slack notifications without excessive duplicate messages.
  • Monitor definitions, credentials guidance, operating instructions, and ownership are documented.
  • Existing checks are retained, migrated, or deliberately removed with the reason recorded.
  • The route-group coverage previously tracked in Add endpoint-level external monitoring for API route groups #2155 is implemented or explicitly deferred in follow-up issues.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions