Grafana Labs

Staff Software Engineer - Databases SRE | Germany | Remote

Stellenbeschreibung:

Overview

As Staff Software Engineer - SRE at Grafana Labs, you will own production reliability for high-SLA Grafana Cloud databases and collaborate with product engineering squads to scale reliability practices. You’ll design automation, define per-tenant SLOs, and lead incident response and post-incident reviews for multi-cloud SaaS environments. This role blends customer needs, production systems, and product design in a globally distributed, open culture. You’ll influence architecture to improve observability, reduce toil, and ship robust, scalable services.

Leistungen / Benefits
  • equity
  • bonus (if applicable)
  • remote work
  • 30 days annual leave with Grafana Shutdown Days
  • global culture and onboarding support
  • open-communication culture
Verantwortungsbereiche
  • Own production reliability for high-SLA and complex customer environments
  • Design and implement automation to scale reliability practices
  • Ensure customers meet SLO targets and evolve per-tenant SLOs
  • Proactively reduce SLO burn and prevent repeat incidents
  • Lead incident response, on-call duties, and post-incident reviews
  • Contribute to design docs and code reviews; influence feature design for scalability
  • Build automation to eliminate toil and improve alert quality
  • Participate in incident response and communicate with customers via bridge calls
Zentrale Anforderungen
  • 8+ years engineering experience, 4+ in SRE/production engineering
  • Strong Kubernetes experience in AWS, GCP, or Azure
  • Familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet)
  • Experience leading teams and mentoring engineers
  • Experience operating multi-tenant systems in production
  • Experience designing and implementing SLOs
  • Proficiency in one or more programming languages (Go, Python, Java)
  • Knowledge of Linux internals, networking, cloud storage, and scaling
  • Excellent problem-solving and troubleshooting skills
  • Experience with blame-free incident response, PIRs, and post-mortems
  • Ability to collaborate with product engineering teams
  • Autonomy, transparency, and a bias toward action
  • intellectual curiosity
  • transparency
  • bias toward action
  • Go
  • Python
  • Java
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    14 Sep 2026
  • Standort:

    Berlin

    Einsatzort:

    Canada
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!