Live status / Developer tools
Buildkite is operating normally.
Buildkite is operating normally.
Signal updated 52m ago · Sep 17, 2026, 8:40 PM UTC
Affected surface
Components
0 of 12 affected
Last 30 days · Status timeline
What changed
- Investigating
We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon.
- Investigating
We’re investigating elevated latency in background worker processing between approximately 12:39am and 1:02am UTC. Processing has since caught up and was near real time as of 1:13am UTC. We’re continuing to monitor the service and are checking remaining retries. Some builds or hosted-agent-related operations may have experienced delays during this period. We’ll provide another update as soon as we have confirmed the service is stable and have more information about the cause. We apologise for the disruption.
- Monitoring
We have confirmed that all background workers have caught up as of 1:13am UTC. Hosted Agents dispatch has returned to normal. We'll continue to to monitor the service and are checking remaining retries.
- Resolved
This incident has been resolved.
- Monitoring
At 06:04 UTC we deployed a defective request routing change to `agent-edge.buildkite.com`, which caused some customers' agents to abruptly disconnect. At 06:09 UTC we noticed the problem, and immediately reverted the change. We began to observe recovery at 06:10 UTC, and are continuing to monitor.
- Monitoring
Agent connectivity is recovered, but some agents are not reconnecting automatically after the earlier forced disconnects. We recommend restarting agent service to make agents reconnect.
- Resolved
This incident has been resolved.
- Investigating
We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon.
- Investigating
We are currently investigating a sitewide outage.
- Investigating
We have identified a fix and applied mitigations, and we are seeing signs of recovery.
- Monitoring
Systems have recovered and we are seeing new builds created. For a period of about 20min new builds could not be created due to a failure of our internal DNS services. No build requests were lost during this time and we are continuing to process the backlog
- Resolved
We're considering this incident resolved.
- Investigating
We're seeing increased latency and error rates for a subset of our customers. We're currently investigating and will provide status updates as they become available.
- Investigating
We identified issues affecting certain customers when the buildkite-agent attempts to register with Buildkite. When agents are not able to register, builds may start to back up. We've identified a potential root case and are attempting changes to fix it.
- Investigating
We have rolled out a few changes to fix the underlying issues and are seeing signs of recovery. We'll keep monitoring until all systems are fully recovered and communicate once we are confident everything is stable again.
- Monitoring
Recovery is underway; we'll keep monitoring until recovery is complete.
- Resolved
Mitigation strategies worked as expected and now agent registration is back to normal. All systems are stable now.
- Monitoring
Internal processing of uploaded test execution data has been delayed, and is currently catching up. No data has been dropped, and all data should be processed over the next hour.
- Monitoring
Catch-up processing continues, most of the backlog is processed, ingestion latency is decreasing.
- Resolved
Processing of uploaded test execution data has caught up and is back to normal.
- Identified
We are experiencing delivery issues with email notifications due to an issue with an upstream provider.
- Identified
Our on-call engineers are working on failing over to an alternative upstream email provider to restore outbound email notifications. No other Buildkite services have been impacted by this disruption.
- Monitoring
We are seeing successful email deliveries and are monitoring processing of our outbound email queue.
- Resolved
Our outbound email queue has finished processing all previously failed deliveries. Email notifications are processing successfully.
- Investigating
We're investigating an issue with job dispatch to our Hosted Agents.
- Investigating
Our engineers are continuing to investigate an issue with delayed job dispatch for Hosted Agents. Job dispatch for self-hosted agents remains unaffected.
- Identified
We're beginning to see recovery with delayed job dispatch for Hosted Agents. We are currently working through the backlog of delayed jobs.
- Monitoring
We are continuing to process through the backlog of delayed jobs and monitoring.
- Resolved
Job dispatch for Hosted Agents is no longer experiencing delays.
- Investigating
We're seeing increased latency and error rates for a subset of our customers. We're currently investigating and will provide status updates as they become available.
- Investigating
Some job dispatches to hosted agents are timing out, leading to jobs starting late. We're currently investigating and will update again soon.
- Monitoring
We're seeing recovery in hosted job scheduling rates. Your jobs should run, but there may be some delays as we work through any backlogs.
- Resolved
Jobs running on Hosted Agents are now being dispatched promptly.
- Investigating
We're seeing increased latency and error rates within our Hosted Agents. We're currently investigating and will provide status updates as they become available.
- Monitoring
We have deployed mitigations and have observed recovery across Hosted Agents. We will continue monitoring.
- Resolved
Jobs running on Hosted Agents are now being dispatched promptly.
- Investigating
We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon.
- Investigating
We are continuing to investigate this issue. We are seeing impact on the Agent API which will affect job scheduling, artifact uploads, and an increase in 5xx responses from the Agent API endpoints.
- Investigating
We are continuing to investigate elevated error rates across multiple services, and are working to determine the cause.
- Monitoring
We are seeing improvements across the affected services, and are seeing services return to normal functionality. We are continuing to monitor and are determining the root cause.
- Resolved
We have seen full recovery for customers since 20:28 UTC. We experienced an autoscaling feedback loop which increased the number of connections to our redis cluster above its ability to respond. This had widespread impact for all of our customers with Web UI, Agent API, REST API and job queue impact between 18:43-19:21 UTC, and again between 20:02-20:28 UTC. A full post incident review will be available later this week.