blacksmith - Backend Degradation causing job slowness and caching failures – Incident details

Backend Degradation causing job slowness and caching failures

Resolved
Degraded performance
Started 3 days agoLasted 1 hour 54 minutes

Affected

Blacksmith Managed Runners

Degraded performance from 9:10 PM to 10:42 PM, Operational from 10:20 PM to 11:04 PM

US West

Degraded performance from 9:10 PM to 10:42 PM, Operational from 10:42 PM to 11:04 PM

us-west ARM

Degraded performance from 9:10 PM to 10:42 PM, Operational from 10:42 PM to 11:04 PM

us-west x86

Degraded performance from 9:10 PM to 10:42 PM, Operational from 10:42 PM to 11:04 PM

EU Central

Degraded performance from 9:10 PM to 10:20 PM, Operational from 10:20 PM to 11:04 PM

eu-central ARM

Degraded performance from 9:10 PM to 10:20 PM, Operational from 10:20 PM to 11:04 PM

Updates
  • Resolved
    UTC
    Resolved

    This incident has been resolved.

  • Update
    UTC
    Update

    This incident is resolved. Job start times and cache operations returned to normal in all regions, and we are continuing to monitor. Jobs that failed during the incident can be re-run.

  • Update
    UTC
    Update

    Caching has recovered, and job queue times across EU regions have recovered. Queue times for larger jobs (16 and 32 vcpu jobs) in US west and US east continue to remain elevated. We are continuing to monitor for any regressions.

  • Monitoring
    UTC
    Monitoring

    Our mitigation has taken effect and job start times and cache operations are improving. Customers may still see some delayed job starts and cache failures while recovery completes. We will provide an update within the next 30 minutes.

  • Identified
    UTC
    Identified

    We have applied a fix and are monitoring its effect. Customers may still see delayed job starts and failing cache operations while recovery completes. We will provide an update within the next 30 minutes.

  • Investigating
    UTC
    Investigating

    Jobs across all regions are taking longer to start and cache operations are failing. We are continuing to investigate the root cause.