Delay on starting Machine Job Tasks

Incident Report for CircleCI

Postmortem

Summary

Since October 1, CircleCI customers have experienced multiple incidents which have caused delays and failures in customer pipeline execution. There are two (unrelated) causes, both of which the CircleCI engineering team is actively mitigating:

  • Available Cloud Computing Capacity
  • Internal Infrastructure

Available Cloud Computing Capacity

Internal Infrastructure

  • Problem: We assign customer workflows over multiple compute providers via a workflow orchestration service. This service is currently gated by the write throughput of a single database.
  • Incidents:

  • Mitigations and Resolutions:

    • By October 23, 2026, we plan to implement multiple redundant workflow orchestration service instances, each with its own database, which will enable us split the orchestration load between them.

Our Commitment

For avoidance of doubt,

  • These particular incidents are not related to ongoing outages at Github
  • We are not currently migrating our infrastructure or services
  • We are not currently under cyberattack

Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents and works to prevent future incidents.

Please reach out to our support team with any questions or concerns.

Posted Oct 08, 2026 - 17:55 UTC

Resolved

The incident has been resolved. Thank you for your patience
Posted Oct 07, 2026 - 14:06 UTC

Monitoring

We are seeing signs of recovery and monitoring the situation.
Posted Oct 07, 2026 - 13:52 UTC

Investigating

Customers may be experiencing elevated wait times for Machine Job Tasks, we are investigating the root cause and will update when we know more.
Posted Oct 07, 2026 - 13:40 UTC
This incident affected: Machine Jobs.