blacksmith - US west cache failure – Incident details

US west cache failure

Resolved
Degraded performance
Started about 3 hours agoLasted about 2 hours

Affected

Actions Cache

Degraded performance from 5:00 PM to 6:06 PM, Operational from 6:06 PM to 6:32 PM

US West Cache

Degraded performance from 5:00 PM to 6:06 PM, Operational from 6:06 PM to 6:32 PM

Updates
  • Resolved
    UTC
    Resolved

    This incident is resolved. GitHub Actions cache operations for jobs in us-west have returned to normal, and any jobs that failed during the incident can be re-run.

  • Update
    UTC
    Update

    Cache errors for jobs in us-west have largely subsided. As we complete recovery, some jobs may see a one-time cache miss on their next run. We are monitoring and will provide an update within the next hour.

  • Monitoring
    UTC
    Monitoring

    We have restored the cache infrastructure in us-west, and cache failures have decreased but are not yet back to normal. Jobs in us-west may still see some cache restores and saves fail and fall back to a full install, while the earlier delayed job starts have cleared and other regions are unaffected. We are bringing additional cache capacity online to fully restore service.

  • Identified
    UTC
    Identified

    We have identified the cause of the failures, and we are implementing a fix to restore cache service. Customers with jobs in us-west may still see cache restores and saves fail and fall back to a full install, which can make those jobs run longer or fail, while jobs in other regions are not affected.

  • Investigating
    UTC
    Investigating

    We are investigating failing GitHub Actions cache operations for jobs running in us-west since approximately 16:10 UTC. Affected jobs may see cache restores and saves fail and fall back to a full install, which can make those jobs run longer or fail, while jobs in other regions are not affected.