Senior Manager, Engineering – Observability Platform

Stellenbeschreibung:

Responsibilities

  • Lead a team of engineers focused on observability platform engineering, driving build-out of a unified observability stack used by all engineering teams at Smartsheet.
  • Own and evolve the platform's technical roadmap, consolidating multiple tooling platforms, and AI observability tooling into a coherent, scalable capability.
  • Define platform standards, contribute to architectural direction, and ensure the team operates with engineering rigor and strong operational habits.
  • Build and scale the team, hiring senior engineers and establishing effective global practices across distributed stakeholders.
  • Lead design and delivery of centralized observability infrastructure covering metrics pipelines, distributed tracing, alerting frameworks, and log analytics across Smartsheet services.
  • Drive SLO/SLA definition and tooling for platform-wide reliability visibility, partnering closely with infrastructure, platform engineering, and on‑call teams.
  • Own governance including instrumentation standards, cost optimization, and rollout of advanced capabilities such as APM, RUM, and custom dashboards.
  • Lead architecture, scaling, and operational practices for log analytics across high‑throughput production workloads.
  • Establish shared observability libraries, agents, and SDKs that reduce instrumentation burden for application engineering teams.
  • Build and maintain AI/ML observability integrations in partnership with the AI Platform team.
  • Partner with the Data & AI Platform team to integrate MLflow tracing, Inference Tables, and LLM‑as‑judge evaluation pipelines into the observability stack.
  • Develop dashboards and alerting for agentic AI workloads, including latency, token consumption, error rates, and evaluation metric drift.
  • Contribute to the AI governance and cost observability program, providing telemetry for model usage, cost attribution, and compliance reporting.
  • Serve as the primary engineering partner for platform consumers across Data & AI, Commerce, Infrastructure, and Security teams, ensuring observability needs are met across workstreams.
  • Lead complex, cross‑functional observability projects with high ambiguity, managing delivery risk, communicating clearly to senior stakeholders, and building alignment across teams.
  • Partner with delivery partners to coordinate instrumentation across platform modernization and migration workstreams.
  • Contribute to quarterly and annual platform goals, reporting on key reliability and observability metrics to engineering leadership.
  • Communicate platform status, risks, and roadmap progress to Engineering leadership and above audiences in a clear, executive‑ready format.
  • Embed on‑call culture and incident management discipline into the team, ensuring clear runbooks, fast MTTR, and post‑incident learning loops.
  • Drive cost governance for observability tooling, including spend optimization and efficient resource management.
  • Champion AI‑assisted engineering practices within the team, applying tooling and automation to reduce toil and accelerate delivery.

Requirements

  • 10+ years of software or platform engineering experience, with strong fundamentals in distributed systems, infrastructure, and backend services.
  • 3 years of engineering management experience, including direct team building, performance management, and cross‑functional delivery ownership.
  • Deep hands‑on expertise with observability tooling: Datadog (APM, metrics, logs, alerting), OpenSearch or Elasticsearch, distributed tracing (OpenTelemetry or equivalent), and SLO/SLA management at scale.
  • Proven experience operating observability platforms for high‑availability, high‑throughput production environments.
  • Experience building and scaling engineering teams in distributed or international focus.
  • Strong execution track record on complex, cross‑functional infrastructure programs with high ambiguity.
  • Clear, direct communication (written and verbal) with both technical and non‑technical audiences, including leadership and executive stakeholders.
  • Proactive risk identification and status communication without prompting.
  • Experience managing vendors, external delivery partners, and third‑party integrations in a platform context.

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    31 Jul 2026
  • Standort:

    WorkFromHome
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!