blacksmith - Backend Degradation causing job slowness and caching failures – Incident details

Backend Degradation causing job slowness and caching failures

Monitoring
Degraded performance
Started about 1 hour ago

Affected

Blacksmith Managed Runners

Degraded performance from 9:10 PM to 12:00 AM, Operational from 10:20 PM to 12:00 AM

US West

Degraded performance from 9:10 PM to 12:00 AM

us-west ARM

Degraded performance from 9:10 PM to 12:00 AM

us-west x86

Degraded performance from 9:10 PM to 12:00 AM

EU Central

Degraded performance from 9:10 PM to 10:20 PM, Operational from 10:20 PM to 12:00 AM

eu-central ARM

Degraded performance from 9:10 PM to 10:20 PM, Operational from 10:20 PM to 12:00 AM

Updates
  • Update
    UTC
    Update

    Caching has recovered, and job queue times across EU regions have recovered. Queue times for larger jobs (16 and 32 vcpu jobs) in US west and US east continue to remain elevated. We are continuing to monitor for any regressions.

  • Monitoring
    UTC
    Monitoring

    Our mitigation has taken effect and job start times and cache operations are improving. Customers may still see some delayed job starts and cache failures while recovery completes. We will provide an update within the next 30 minutes.

  • Identified
    UTC
    Identified

    We have applied a fix and are monitoring its effect. Customers may still see delayed job starts and failing cache operations while recovery completes. We will provide an update within the next 30 minutes.

  • Investigating
    UTC
    Investigating

    Jobs across all regions are taking longer to start and cache operations are failing. We are continuing to investigate the root cause.