blacksmith - US West storage cluster degradation – Incident details

US West storage cluster degradation

Resolved
Degraded performance
Started 17 days agoLasted about 3 hours

Affected

Incremental Docker Builders

Degraded performance from 4:13 PM to 7:23 PM

us-west Storage Cluster

Degraded performance from 4:13 PM to 7:23 PM

Docker Container Cache

Degraded performance from 4:13 PM to 7:23 PM

us-west Storage Cluster

Degraded performance from 4:13 PM to 7:23 PM

Sticky Disks

Degraded performance from 4:13 PM to 7:23 PM

us-west Storage Cluster

Degraded performance from 4:13 PM to 7:23 PM

Updates
  • Resolved
    UTC
    Resolved

    This incident is resolved. Sticky disk storage in our us-west region came under more read load than it could serve at normal latency, which slowed and produced higher error rates for Git-cached checkouts, incremental Docker builds, and container caching. We reduced and redistributed that load, and performance has been normal since approximately 17:15 UTC. Jobs that failed during the incident can be safely re-run.

  • Monitoring
    UTC
    Monitoring

    Our us-west region storage cluster's latency and error rates have returned to baseline as of approximately 17:15 UTC, following mitigations that reduce load on the affected storage. Git-cached checkouts, incremental Docker builds, and container caching in us-west are back to normal, and we are monitoring to confirm the recovery holds while we continue to investigate the underlying cause. We will provide a final update within the next hour.

  • Update
    UTC
    Update

    We are still working to resolve elevated latency on sticky disk storage in our us-west region; other regions are not affected. Customers running jobs in us-west may continue to see slower Git-cached checkouts, incremental Docker builds, and container caching. Our investigation into the underlying cause is ongoing and we will provide another update within the next 30 minutes.

  • Update
    UTC
    Update

    Sticky disk storage in our us-west region is experiencing higher latency; our other regions are not affected. Customers running jobs in us-west may see slower Git-cached checkouts, incremental Docker builds, and container caching. We are continuing to investigate the underlying cause and will provide another update within the next 30 minutes.

  • Investigating
    UTC
    Investigating

    We are investigating reports of higher latency on sticky disk operations in our us-west region. Customers running jobs in us-west may see slower incremental Docker builds, Git caching, and container caching, so affected jobs can take longer than usual to complete.