Skip to content

Status

All systems operational

Live operational status and 90-day uptime history for every Ember service component. 99.98% aggregate uptime over the last 90 days.

All systems operationalChecked July 22, 2026, 09:41 UTC
01Components

System status

API

REST and WebSocket endpoints

Operational
99.99% uptime (90 d)

Incident ingestion

PagerDuty, Opsgenie, CloudWatch and Grafana connectors

Operational
99.97% uptime (90 d)

Model routing

Provider fan-out and per-task failover layer

Operational
99.98% uptime (90 d)

Dashboard

Web application and asset delivery

Operational
99.99% uptime (90 d)

Webhooks

Outbound delivery to Slack, Jira and custom endpoints

Operational
99.95% uptime (90 d)

Postmortem export

PDF, Confluence and Notion export pipeline

Operational
99.99% uptime (90 d)
99.98% aggregate uptime across all components over 90 days
33 min total incident time in 90 days
02Incidents

Recent incidents

Resolved incidents from the past 90 days. One incident in this period.

ResolvedSEV3July 8, 2026Duration: 33 min

Elevated webhook delivery latency

Component: Webhooks  ·  Ref: INC-2026-0708

Impact

p99 webhook delivery latency reached 4.2 s against a normal baseline of ~80 ms for 33 minutes. No webhooks were dropped or lost. Slack and Jira notifications were delayed by up to 90 s per event during the window.

Timeline (UTC)

  1. 11:14 UTCAutomated alert fired: p99 delivery latency crossed the 1 s threshold. Incident opened.
  2. 11:19 UTCEmber correlated the spike with deploy v2026.7.8, landed at 11:08 UTC. The webhook dispatcher connection pool was at saturation (200/200 connections).
  3. 11:31 UTCRoot cause confirmed: a misconfigured socket timeout in the dispatcher caused messages to be re-queued faster than the pool could drain, triggering a retry-storm.
  4. 11:47 UTCRolled back deploy v2026.7.8 (kubectl rollout undo deploy/webhook-dispatcher). Latency returned to baseline within 90 s of rollback completing.
  5. 12:05 UTCIncident stood down. All queued webhooks delivered. Postmortem drafted and circulated to the platform team.

Resolution

The misconfigured socket timeout in v2026.7.8 caused the webhook dispatcher to re-queue unacknowledged messages faster than the connection pool could handle, exhausting all 200 slots within 6 minutes of the deploy landing. Rollback resolved pool saturation immediately. Follow-up: timeout configuration is now validated in CI before merge.

No other incidents in the past 90 days.

03Maintenance

Scheduled maintenance

No maintenance scheduled

All upcoming maintenance windows are posted here at least 72 hours in advance. Subscribers receive an email and Slack notification when a window is announced, updated or cancelled. Historical windows are retained in this log.

Subscribe to status updates

Get notified the moment an incident opens, updates or resolves, and when maintenance is scheduled. Delivered by email or Slack.

Subscribe

Product releases and changes are documented in the changelog. For security vulnerabilities, contact security@emberoncall.com.