blacksmith - Degraded job performance, metrics, and log ingestion in us-west – Incident details

Degraded job performance, metrics, and log ingestion in us-west

Resolved
Degraded performance
Started 20 days agoLasted about 1 hour

Affected

Blacksmith Managed Runners

Operational from 7:17 PM to 7:45 PM, Degraded performance from 7:45 PM to 8:44 PM

US West

Operational from 7:17 PM to 7:45 PM, Degraded performance from 7:45 PM to 8:44 PM

us-west ARM

Operational from 7:17 PM to 7:45 PM, Degraded performance from 7:45 PM to 8:44 PM

us-west x86

Operational from 7:17 PM to 7:45 PM, Degraded performance from 7:45 PM to 8:44 PM

Dashboard

Degraded performance from 7:17 PM to 8:44 PM

Updates
  • Resolved
    UTC
    Resolved

    This incident is resolved, with job performance and the ingestion of metrics and logs in our us-west region stable for the past 30 minutes. Jobs that failed or timed out during the incident can be safely re-run.

  • Monitoring
    UTC
    Monitoring

    Metrics and log ingestion in our us-west region is recovering, and job performance in the region has returned to normal. We are monitoring to confirm the recovery holds and are continuing to investigate the underlying cause. We will provide an update shortly.

  • Update
    UTC
    Update

    We are experiencing degraded network performance in our us-west region, affecting metrics and log ingestion as well as job performance. Jobs in us-west that upload artifacts or transfer large amounts of data may run slower than normal and in some cases hit their configured timeouts and fail. Other regions are not affected, and we are actively investigating the issue.

  • Update
    UTC
    Update

    We are continuing to investigate degraded metrics and log ingestion in our us-west region. Customers may still see metrics and logs for their jobs appear missing or delayed in the Blacksmith dashboard, while other regions remain unaffected. We will provide another update within the next 30 minutes.

  • Investigating
    UTC
    Investigating

    We are experiencing degradation in our metrics and log ingestion services in our us-west region. We are actively investigating the issue.