Unplanned Downtime

Incident Report for Arcade

Resolved

We determined that because of recently adding thousands(!) of tools, one of our production components required a longer than expected startup process, meaning its health check endpoint was unavailable before the timeout expired. As a result, orchestrator was constantly restarting nodes, leading to poor availability of this critical component. Going forward, we plan to move these tasks out of component startup so this doesn't happen again and we can continue to support even larger numbers of tools.
Posted Nov 06, 2025 - 23:54 UTC

Monitoring

Our prod services are now recovering. We are continuing to monitor and have identified the root cause of the issue.
Posted Nov 06, 2025 - 23:30 UTC

Update

We are continuing to investigate this issue.
Posted Nov 06, 2025 - 23:22 UTC

Investigating

We are currently investigating an issue with our production deployments. We are attempting to rollback to a stable version. Investigation of the issue continues.
Posted Nov 06, 2025 - 23:20 UTC
This incident affected: Arcade Cloud Infrastructure (Engine API).