blacksmith - Notice history

100% - uptime

eu-central Storage Cluster - Operational

100% - uptime
Aug 2026 · 99.98%Sep · 99.92%Oct · 100.0%
Aug 2026
Sep 2026
Oct 2026

us-west Storage Cluster - Operational

99% - uptime
Aug 2026 · 96.53%Sep · 99.92%Oct · 100.0%
Aug 2026
Sep 2026
Oct 2026

eu-west Storage Cluster - Operational

100% - uptime
Aug 2026 · 99.98%Sep · 99.92%Oct · 100.0%
Aug 2026
Sep 2026
Oct 2026

us-east Storage Cluster - Operational

100% - uptime
Aug 2026 · 100.0%Sep · 100.0%Oct · 100.0%
Aug 2026
Sep 2026
Oct 2026

API - Operational

100% - uptime
Aug 2026 · 100.0%Sep · 100.0%Oct · 100.0%
Aug 2026
Sep 2026
Oct 2026
100% - uptime

https://blacksmith.sh - Operational

100% - uptime
Aug 2026 · 99.98%Sep · 100.0%Oct · 100.0%
Aug 2026
Sep 2026
Oct 2026

Github → Actions - Operational

Github → API Requests - Operational

Github → Webhooks - Operational

Codesmith - Operational

100% - uptime
Aug 2026 · 99.98%Sep · 100.0%Oct · 100.0%
Aug 2026
Sep 2026
Oct 2026

Dashboard - Operational

100% - uptime
Aug 2026 · 99.57%Sep · 99.92%Oct · 100.0%
Aug 2026
Sep 2026
Oct 2026

Notice history

View current status

Sep 2026

Backend Degradation causing job slowness and caching failures
ResolvedDegraded performance1 hour 54 minutes
  • Resolved
    UTC
    Resolved

    This incident has been resolved.

  • Update
    UTC
    Update

    This incident is resolved. Job start times and cache operations returned to normal in all regions, and we are continuing to monitor. Jobs that failed during the incident can be re-run.

  • Update
    UTC
    Update

    Caching has recovered, and job queue times across EU regions have recovered. Queue times for larger jobs (16 and 32 vcpu jobs) in US west and US east continue to remain elevated. We are continuing to monitor for any regressions.

  • Monitoring
    UTC
    Monitoring

    Our mitigation has taken effect and job start times and cache operations are improving. Customers may still see some delayed job starts and cache failures while recovery completes. We will provide an update within the next 30 minutes.

  • Identified
    UTC
    Identified

    We have applied a fix and are monitoring its effect. Customers may still see delayed job starts and failing cache operations while recovery completes. We will provide an update within the next 30 minutes.

  • Investigating
    UTC
    Investigating

    Jobs across all regions are taking longer to start and cache operations are failing. We are continuing to investigate the root cause.

US west cache failure
ResolvedDegraded performance1 hour 32 minutes
  • Resolved
    UTC
    Resolved

    This incident is resolved. GitHub Actions cache operations for jobs in us-west have returned to normal, and any jobs that failed during the incident can be re-run.

  • Update
    UTC
    Update

    Cache errors for jobs in us-west have largely subsided. As we complete recovery, some jobs may see a one-time cache miss on their next run. We are monitoring and will provide an update within the next hour.

  • Monitoring
    UTC
    Monitoring

    We have restored the cache infrastructure in us-west, and cache failures have decreased but are not yet back to normal. Jobs in us-west may still see some cache restores and saves fail and fall back to a full install, while the earlier delayed job starts have cleared and other regions are unaffected. We are bringing additional cache capacity online to fully restore service.

  • Identified
    UTC
    Identified

    We have identified the cause of the failures, and we are implementing a fix to restore cache service. Customers with jobs in us-west may still see cache restores and saves fail and fall back to a full install, which can make those jobs run longer or fail, while jobs in other regions are not affected.

  • Investigating
    UTC
    Investigating

    We are investigating failing GitHub Actions cache operations for jobs running in us-west since approximately 16:10 UTC. Affected jobs may see cache restores and saves fail and fall back to a full install, which can make those jobs run longer or fail, while jobs in other regions are not affected.

GitHub connectivity degraded in us-west
ResolvedDegraded performance33 minutes
  • Resolved
    UTC
    Resolved

    Git checkout performance in us-west has remained stable since we rerouted traffic away from the congested upstream network link, and this incident is now resolved.

  • Monitoring
    UTC
    Monitoring

    After rerouting traffic, congestion on the upstream network link has subsided, and git checkout performance in us-west has returned to normal. We are monitoring and will resolve this incident once performance has continued to stay stable.

  • Investigating
    UTC
    Investigating

    We are investigating degraded network connectivity between our us-west region and GitHub. Customers running jobs in us-west may see slower-than-normal git checkouts, while jobs in other regions are not affected. Workflows using the Blacksmith checkout action with git checkout caching are less affected, and setup instructions are in the documentation below.

    Git checkout caching documentation: https://docs.blacksmith.sh/blacksmith-caching/git-checkout-caching

Aug 2026

GitHub outage affecting job failures and dashboard errors
ResolvedMajor outage6 hours 57 minutes
  • Resolved
    UTC
    Resolved
    This incident has been resolved.
  • Update
    UTC
    Update

    We are no longer seeing upstream errors, and are continuing to monitor impact of the upstream outage.

  • Update
    UTC
    Update

    We are still seeing intermitting GitHub API errors at a low rate.

  • Monitoring
    UTC
    Monitoring

    We're seeing signs of GitHub recovery. The dashboard is now loading and jobs should be running again. We are monitoring the recovery.

  • Update
    UTC
    Update

    Upstream incident has been declared: https://www.githubstatus.com/incidents/zkxwbgr0cnmx

    We are also seeing elevated error rates in jobs as they hit upstream GitHub errors. Job adoption times are also affected and are delayed.

    We are monitoring and are looking at potential mitigations.

  • Identified
    UTC
    Identified

    The Blacksmith Dashboard is unable to load. We've identified the root cause to upstream 503s being returned from GitHub on permission-check requests.

Storage degradation in us-west
ResolvedDegraded performance30 hours 33 minutes
  • Resolved
    UTC
    Resolved

    All services have been fully restored, including sticky disks, Docker container caching, and incremental Docker builders in us-west, and job queues are operating normally in all regions. A small amount of recently written cache data could not be recovered during storage repair, so some builds may run slower on their first runs while caches rebuild. This incident is now resolved, and we will publish a detailed post-incident report in the coming days.

  • Update
    UTC
    Update

    All services have been restored, including sticky disks, Docker container caching, and incremental Docker builders in us-west, and job queues are operating normally in all regions. We are continuing to monitor the stability of the recovered storage cluster as it ramps back up with traffic, and some builds may run slower on their first runs while recently written cache data rebuilds. We will post a final update once we have confirmed stability.

  • Update
    UTC
    Update

    The storage cluster for Sticky Disks, Docker Container Caching, and Incremental Docker Builders in us-west has been restored, and disks are mounting and operating normally. As part of the recovery, a small amount of recently written cache data may need to be rebuilt, so some Docker builds may run slower over their first few runs while caches rehydrate. We are monitoring closely.

  • Update
    UTC
    Update

    We are continuing to work with our compute provider to restore the storage cluster for Sticky Disks, Docker Container Caching, and Incremental Docker Builders in us-west. This requires hands-on recovery work by our provider's team and may take a few more hours to fully restore. Workflows using the affected features in us-west may see slow or stalled Docker builds in the meantime.

  • Update
    UTC
    Update

    We have restored the Github Actions Cache in us-west. We are still working with our compute provider to restore our remaining storage cluster for Sticky Disks, Docker Container Caching, and Incremental Docker Builders.

  • Update
    UTC
    Update

    We are still working with our compute provider to recover the storage clusters.

  • Update
    UTC
    Update

    We are noticing some issues with the storage cluster backing the Github Actions cache after it was restored. We are working on restoring this alongside the storage cluster backing Sticky Disks, Incremental Docker Builders, and Docker Container Caching.

  • Update
    UTC
    Update

    Our us-west compute provider has restored power and cooling in their datacenter and the majority of our capacity has returned. Github Actions caching has been reenabled in the region but we are still working to restore Sticky Disks, Docker Container caching, and Incremental Docker Builders. We will post an update once these components are restored.

  • Update
    UTC
    Update

    Job queues have fully recovered, and runners in all regions are operating normally for all runner sizes. Cache operations, including sticky disks, remain unavailable in us-west while we bring the restored storage hardware back online. We will provide another update within the next hour.

  • Monitoring
    UTC
    Monitoring

    Queue times for all runner sizes have returned to normal in all regions, and we are monitoring closely while eu-west clears the last of its backlog. Cache operations, including sticky disks, remain unavailable in us-west while we bring the restored storage hardware back online. We will provide another update within the next hour.

  • Update
    UTC
    Update

    Runner capacity in us-west continues to recover and queues in all regions are draining. Our provider has restored the access we need to begin bringing our us-west storage systems back online. Cache operations, including sticky disks, remain unavailable while that work completes, and jobs may still take longer than usual to start.

  • Update
    UTC
    Update

    Queue times are improving across all regions as capacity comes back online, but jobs may still take longer than usual to start. Cache operations in us-west, including sticky disks, remain unavailable while our provider restores the storage systems, and jobs on our largest runner sizes may see the longest delays. We will provide another update within the next hour.

  • Update
    UTC
    Update

    Queue times in us-west have returned to near-normal levels, while us-east, eu-west, and eu-central are still working through their remaining backlogs. Cache operations, including sticky disks, remain unavailable in us-west while our provider works to restore the storage systems there. We will provide another update within the next hour.

  • Update
    UTC
    Update

    Queues in us-west have come down substantially as restored capacity comes online and we continue moving customers back. Cache operations, including sticky disks, remain unavailable in us-west, and jobs in us-east may still take significantly longer than usual to start while we work through the remaining backlog. We will provide another update within the next hour.

  • Update
    UTC
    Update

    More us-west capacity has come back online and queues in the region are steadily draining as we move customers back. Cache operations, including sticky disks, remain unavailable in us-west while our provider works to restore the storage portion of the facility, and because we shifted us-west traffic to other regions earlier today, jobs elsewhere, particularly in us-east, may still take longer than usual to start. We will provide another update within the next hour.

  • Update
    UTC
    Update

    Our provider is continuing to bring us-west servers back online, and we are gradually moving customers back to us-west as capacity returns. The backlog has not yet cleared, so jobs in all regions may still take longer than usual to start. We will provide another update within the next hour.

  • Update
    UTC
    Update

    Our provider has begun powering servers back on at the us-west facility, and a portion of our us-west capacity is back online. We have started moving some customers back to us-west to spread load across regions. Jobs in all regions may still take longer than usual to start.

  • Update
    UTC
    Update

    Our provider is continuing to restore cooling at the us-west facility; temperatures have not yet reached safe levels for servers to be powered back on, and we are staged to begin restoring runner capacity as soon as they are. The us-west region remains in a major outage, and jobs in all regions may still take longer than usual to start. We will provide another update within the next hour.

  • Update
    UTC
    Update

    Our provider has partially restored cooling at the us-west facility and expects to reach safe operating temperatures within the next few hours, at which point servers will be powered back on in phases. In the meantime jobs in all regions may still take significantly longer than usual to start. We will provide another update within the next hour.

  • Update
    UTC
    Update

    Our datacenter provider has begun deploying mitigations at the affected us-west facility, and GitHub has resolved its separate webhooks incident, though some knock-on delays may persist while affected jobs requeue. 

    The us-west region remains heavily affected and with capacity reduced while we rebalance traffic, jobs in all regions may take longer than usual to start.

  • Update
    UTC
    Update

    Conditions in our us-west region have not yet improved. We're actively working with our datacenter provider on the issue, and will provide updates as they occur.  We are also manually rebalancing traffic out of the us-west region to aid with recovery.

    Github have also declared an incident affecting webhooks - https://www.githubstatus.com/incidents/k8vbzwqjkxzn which may cause some job adoption delays

  • Update
    UTC
    Update

    Conditions at our provider's us-west facility have not yet improved. We are continuing to issue mitigations to ensure that customer impact is minimal.

  • Update
    UTC
    Update

    Conditions at our provider's us-west facility have not yet improved, and we are continuing to work around the issue to keep customer impact minimal. Such mitigations may include rebalancing customers to unaffected regions.

  • Update
    UTC
    Update

    Conditions at our provider's us-west facility have not yet improved. As we rebalance customers away from us-west, other regions may see slightly longer pickup times than usual as a side effect

  • Update
    UTC
    Update

    Conditions at our provider's us-west facility have not yet improved, and we are continuing to work around the issue to keep customer impact minimal. Such mitigations may include rebalancing customers to unaffected regions.

  • Update
    UTC
    Update

    We are currently applying mitigations for affected customers in the region.

  • Identified
    UTC
    Identified

    Caching performance in US West is currently degraded. Our upstream provider has lost cooling power in their data center, leading to failure of some storage systems in US West. As a result jobs may take longer to complete. We are taking action to mitigate the issue.

  • Investigating
    UTC
    Investigating

    Our cloud provider is experiencing a thermal event in their us-west datacenter that is affecting caching in the region. Sticky Disk, Actions Cache, Docker Container Cache, and Bazel Build Caching are impacted: cache operations may be slow or unavailable which may result in some builds running slower than usual.

Previous

Aug 2026 to Oct 2026

Next