Lead Business Analyst – Alert Management, Observability Standards

Stellenbeschreibung:

Responsibilities

  • Provide solutions that help attain business outcomes.
  • Responsible for rationalizing and governing all system alerts to ensure they align with department priorities, operational coverage models, and service reliability goals.
  • Define alerting standards, review and approve alerts before they are routed to the 24x7 Eyes‑on‑Glass Operations team, and establish a scalable approach to cataloging alert response instructions (runbooks/playbooks).
  • Operate at the intersection of the IT Operations Command Center (OCC), engineering/application teams, platform/monitoring tool owners, and service owners, ensuring alerts are actionable, prioritized, and paired with clear response guidance.
  • Establish and maintain a department‑wide alert rationalization framework that evaluates alerts for: business/service criticality and operational priority, actionability, signal‑to‑noise, and ownership.
  • Perform regular alert reviews to ensure alert quality, correct routing, and alignment with operational coverage.
  • Lead continuous improvement efforts to reduce alert fatigue while preserving detection of true incidents and high‑impact degradation.
  • Define and enforce alerting standards.

Requirements

  • 5+ years in IT Operations, SRE, Observability, Monitoring Engineering, or Incident Management
  • Demonstrated success reducing noise and improving actionability across enterprise alerting ecosystems
  • Experience with common monitoring/observability tools (e.g., Splunk, AppDynamics, Dynatrace, Datadog, Prometheus/Grafana, Azure Monitor, CloudWatch, ServiceNow Event Mgmt or similar)
  • Strong understanding of: Incident response workflows and operational coverage models (24x7 vs. business hours)
  • CMDB/service ownership concepts and dependency mapping
  • Standard operating procedures/runbooks and knowledge management
  • Excellent stakeholder management and ability to drive standards across teams.

Hard Skills

  • IT Operations
  • SRE
  • Observability
  • Monitoring Engineering
  • Incident Management
  • alert rationalization
  • alert standards
  • incident response workflows
  • dependency mapping
  • knowledge management

Soft Skills

  • stakeholder management
  • continuous improvement
  • leadership
  • communication
  • organizational skills

Industry Keywords

  • alert fatigue
  • operational coverage models
  • business/service criticality
  • actionability
  • signal-to-noise

Tools & Technologies

  • Splunk
  • AppDynamics
  • Dynatrace
  • Datadog
  • Prometheus
  • Grafana
  • Azure Monitor
  • CloudWatch
  • ServiceNow Event Management

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    24 Jul 2026
  • Standort:

    WorkFromHome
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

    Development & IT
  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!