Status | Blacksmith - Notice history

100% - uptime

eu-central Storage Cluster - Operational

100% - uptime
May 2026 · 100.0%Jun · 100.0%Jul · 100.0%
May 2026
Jun 2026
Jul 2026

us-west Storage Cluster - Operational

100% - uptime
May 2026 · 100.0%Jun · 100.0%Jul · 100.0%
May 2026
Jun 2026
Jul 2026

eu-west Storage Cluster - Operational

100% - uptime
May 2026 · 100.0%Jun · 100.0%Jul · 100.0%
May 2026
Jun 2026
Jul 2026

Actions Cache - Operational

100% - uptime
May 2026 · 100.0%Jun · 100.0%Jul · 99.84%
May 2026
Jun 2026
Jul 2026

API - Operational

100% - uptime
May 2026 · 100.0%Jun · 100.0%Jul · 100.0%
May 2026
Jun 2026
Jul 2026

Website - Operational

100% - uptime
May 2026 · 100.0%Jun · 100.0%Jul · 99.94%
May 2026
Jun 2026
Jul 2026

Github → Actions - Operational

Github → API Requests - Operational

Github → Webhooks - Operational

Website - Operational

100% - uptime
May 2026 · 100.0%Jun · 100.0%Jul · 100.0%
May 2026
Jun 2026
Jul 2026

https://blacksmith.sh - Operational

100% - uptime
May 2026 · 100.0%Jun · 100.0%Jul · 100.0%
May 2026
Jun 2026
Jul 2026

Codesmith - Operational

100% - uptime
May 2026 · 100.0%Jun · 100.0%Jul · 100.0%
May 2026
Jun 2026
Jul 2026

Notice history

Jul 2026

Delays in job adoption
  • Resolved
    UTC
    Resolved
    This incident has been resolved.
  • Update
    UTC
    Update

    We are currently scanning for any missed jobs during this degradation period and ensuring that they are being dispatched.

  • Update
    UTC
    Update

    We are seeing job adoption return to normal latencies across all regions.

  • Update
    UTC
    Update

    We are still watching over the job queue draining in one of our regions (eu-central). Customers in this region will see a delay for their jobs to be picked up.

  • Update
    UTC
    Update

    Job adoption in us-west and eu-west have returned to normal. Customers in eu-central may still see delayed job starts while we clear the remaining queue backlog. We have identified the cause of the Bazel caching issue and are now implementing a fix. We will provide another update within the next hour.

  • Update
    UTC
    Update

    Cache and sticky disk operations are healthy again. Bazel caching is not yet operational but being actively investigated by our team. The remaining impact is a backlog of queued jobs that we are actively draining, so some customers may still see delayed job starts until the queue drains.

  • Monitoring
    UTC
    Monitoring

    Our primary Redis instance, which backs our control plane, hit a saturation point, leading to a feedback loop of load. The initial cause was a burst of deliveries of delayed webhooks from GitHub, and our reconciliation systems added further load to the control plane, making things worse.

    We have improved load balancing by spreading this workload across multiple Redis instances, and our control plane has fully recovered. The remaining impact is a backlog of queued jobs that we are actively draining, so some customers may still see delayed job starts until the queue clears. We will continue monitoring and update with our findings in the next 30 minutes.

  • Update
    UTC
    Update

    We have applied a change to our backend systems and metrics are showing partial improvement, though not yet back to baseline. Customers can still expect degraded cache interactions while we continue to investigate the root cause. We will provide an update within the next 30 minutes.

  • Investigating
    UTC
    Investigating

    We have rolled out fixes to our workload to bring database query latency back to baseline. The degradation has now moved to other parts of our stack. We are continuing to investigate the root cause here. Customers can still expect to see degraded cache interactions. We will provide an update within the next 30 minutes.

  • Update
    UTC
    Update

    We are continuing to investigate degradation in our control plane services. This is causing a large % of caching and stickydisk requests to fail, which in turn cause customer jobs to run much slower than their baseline. The slower runs are exacerbating queueing of new jobs that are coming in to the system. We are actively investigating and rolling out fixes to unblock customer jobs.

  • Update
    UTC
    Update

    We are continuing to investigate the delays in webhook processing on our control plane. Customer's may also see some degradation with caching and sticky disk components.

  • Identified
    UTC
    Identified

    We have applied a mitigation for the Caching and Monitors degradation. We are seeing some delays in webhook processing on our control plane which we are investigating.

Jun 2026

Github are reporting degraded performance for Actions and Webhooks

May 2026

Github reporting degraded performance for API Requests, Git Operations and other services
  • Resolved
    UTC
    Resolved

    Github have resolved this incident.

  • Monitoring
    UTC
    Monitoring
GitHub reporting degraded performance for Actions
  • Resolved
    UTC
    Resolved

    Github have resolved this incident.

  • Monitoring
    UTC
    Monitoring

    GitHub is reporting degraded performance with Actions as well as other services. This may have an impact on job adoption.
    We are monitoring this incident - https://www.githubstatus.com/incidents/gnftqj9htp0g

Delays in job adoption in eu-west
  • Resolved
    UTC
    Resolved
    This incident has been resolved.
  • Identified
    UTC
    Identified

    We are continuing to monitor this, as well as track an incident reported by with GitHub Actions which can be found here:

    https://www.githubstatus.com/incidents/g6ffrm0rfvz9

    We are continuing to re-balance in order to reduce delays.

  • Investigating
    UTC
    Investigating

    We're seeing a spike of jobs in the eu-west region leading to a temporary capacity shortage. We're looking into re-balancing to reduce the delays

May 2026 to Jul 2026

Next