Live status / Developer tools

Buildkite is operating normally.

Buildkite is operating normally.

Signal updated 52m ago · Sep 17, 2026, 8:40 PM UTC

OperationalOfficial signal: none

Affected surface

Components

12 components
WebOfficial componentOperational
Agent APIOfficial componentOperational
REST APIOfficial componentOperational
Job QueueOfficial componentOperational
NotificationsOfficial componentOperational
Hosted AgentsOfficial componentOperational
Package RegistriesOfficial componentOperational
Test EngineOfficial componentOperational
SCM IntegrationsOfficial componentOperational
SCM ProvidersOfficial componentOperational
Third Party ServicesOfficial componentOperational
DocsOfficial componentOperational

0 of 12 affected

Last 30 days · Status timeline

What changed

41 updates
  1. Investigating

    We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon.

  2. Investigating

    We’re investigating elevated latency in background worker processing between approximately 12:39am and 1:02am UTC. Processing has since caught up and was near real time as of 1:13am UTC. We’re continuing to monitor the service and are checking remaining retries. Some builds or hosted-agent-related operations may have experienced delays during this period. We’ll provide another update as soon as we have confirmed the service is stable and have more information about the cause. We apologise for the disruption.

  3. Monitoring

    We have confirmed that all background workers have caught up as of 1:13am UTC. Hosted Agents dispatch has returned to normal. We'll continue to to monitor the service and are checking remaining retries.

  4. Resolved

    This incident has been resolved.

  5. Monitoring

    At 06:04 UTC we deployed a defective request routing change to `agent-edge.buildkite.com`, which caused some customers' agents to abruptly disconnect. At 06:09 UTC we noticed the problem, and immediately reverted the change. We began to observe recovery at 06:10 UTC, and are continuing to monitor.

  6. Monitoring

    Agent connectivity is recovered, but some agents are not reconnecting automatically after the earlier forced disconnects. We recommend restarting agent service to make agents reconnect.

  7. Resolved

    This incident has been resolved.

  8. Investigating

    We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon.

  9. Investigating

    We are currently investigating a sitewide outage.

  10. Investigating

    We have identified a fix and applied mitigations, and we are seeing signs of recovery.

  11. Monitoring

    Systems have recovered and we are seeing new builds created. For a period of about 20min new builds could not be created due to a failure of our internal DNS services. No build requests were lost during this time and we are continuing to process the backlog

  12. Resolved

    We're considering this incident resolved.

  13. Investigating

    We're seeing increased latency and error rates for a subset of our customers. We're currently investigating and will provide status updates as they become available.

  14. Investigating

    We identified issues affecting certain customers when the buildkite-agent attempts to register with Buildkite. When agents are not able to register, builds may start to back up. We've identified a potential root case and are attempting changes to fix it.

  15. Investigating

    We have rolled out a few changes to fix the underlying issues and are seeing signs of recovery. We'll keep monitoring until all systems are fully recovered and communicate once we are confident everything is stable again.

  16. Monitoring

    Recovery is underway; we'll keep monitoring until recovery is complete.

  17. Resolved

    Mitigation strategies worked as expected and now agent registration is back to normal. All systems are stable now.

  18. Monitoring

    Internal processing of uploaded test execution data has been delayed, and is currently catching up. No data has been dropped, and all data should be processed over the next hour.

  19. Monitoring

    Catch-up processing continues, most of the backlog is processed, ingestion latency is decreasing.

  20. Resolved

    Processing of uploaded test execution data has caught up and is back to normal.

  21. Identified

    We are experiencing delivery issues with email notifications due to an issue with an upstream provider.

  22. Identified

    Our on-call engineers are working on failing over to an alternative upstream email provider to restore outbound email notifications. No other Buildkite services have been impacted by this disruption.

  23. Monitoring

    We are seeing successful email deliveries and are monitoring processing of our outbound email queue.

  24. Resolved

    Our outbound email queue has finished processing all previously failed deliveries. Email notifications are processing successfully.

  25. Investigating

    We're investigating an issue with job dispatch to our Hosted Agents.

  26. Investigating

    Our engineers are continuing to investigate an issue with delayed job dispatch for Hosted Agents. Job dispatch for self-hosted agents remains unaffected.

  27. Identified

    We're beginning to see recovery with delayed job dispatch for Hosted Agents. We are currently working through the backlog of delayed jobs.

  28. Monitoring

    We are continuing to process through the backlog of delayed jobs and monitoring.

  29. Resolved

    Job dispatch for Hosted Agents is no longer experiencing delays.

  30. Investigating

    We're seeing increased latency and error rates for a subset of our customers. We're currently investigating and will provide status updates as they become available.

  31. Investigating

    Some job dispatches to hosted agents are timing out, leading to jobs starting late. We're currently investigating and will update again soon.

  32. Monitoring

    We're seeing recovery in hosted job scheduling rates. Your jobs should run, but there may be some delays as we work through any backlogs.

  33. Resolved

    Jobs running on Hosted Agents are now being dispatched promptly.

  34. Investigating

    We're seeing increased latency and error rates within our Hosted Agents. We're currently investigating and will provide status updates as they become available.

  35. Monitoring

    We have deployed mitigations and have observed recovery across Hosted Agents. We will continue monitoring.

  36. Resolved

    Jobs running on Hosted Agents are now being dispatched promptly.

  37. Investigating

    We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon.

  38. Investigating

    We are continuing to investigate this issue. We are seeing impact on the Agent API which will affect job scheduling, artifact uploads, and an increase in 5xx responses from the Agent API endpoints.

  39. Investigating

    We are continuing to investigate elevated error rates across multiple services, and are working to determine the cause.

  40. Monitoring

    We are seeing improvements across the affected services, and are seeing services return to normal functionality. We are continuing to monitor and are determining the root cause.

  41. Resolved

    We have seen full recovery for customers since 20:28 UTC. We experienced an autoscaling feedback loop which increased the number of connections to our redis cluster above its ability to respond. This had widespread impact for all of our customers with Web UI, Agent API, REST API and job queue impact between 18:43-19:21 UTC, and again between 20:02-20:28 UTC. A full post incident review will be available later this week.