Skip to content

Workflow status is tracked in GitHub. This local document is retained as an engineering spec/evidence record.

ED2KSRV-FEAT-002 - Add production observability and restart recovery

Summary

Expose the signals required to operate the server and prove that planned or unexpected restarts return to a healthy, internally consistent state.

Acceptance Criteria

  • [ ] Metrics cover active/admitted/rejected sessions, files/sources, request rates, queue saturation, errors, and latency.
  • [ ] Health and readiness have distinct meanings and readiness fails on critical saturation or incomplete startup.
  • [ ] Structured logs identify lifecycle transitions and rejection/error classes without private payload data.
  • [ ] Restart semantics document which state is ephemeral, how clients repopulate it, and expected recovery time.
  • [ ] Crash/restart and configuration-reload tests show no stale registrations, corrupted counters, or false readiness.
  • [ ] Operator guidance defines alert thresholds used by the soak and canary gates.

Notes

Metrics must expose the admission decisions implemented by ED2KSRV-BUG-002 and ED2KSRV-BUG-003.