Senior Site Reliability Engineer / Platform Engineer

Stellenbeschreibung:

We are looking for a Senior SRE / Platform Engineer (m/f/d) to own and improve the cloud infrastructure behind SimScale’s browser-based simulation platform. The role spans AWS and EKS, observability, disaster recovery, security and compliance controls, multi-region architecture, elastic GPU/HPC capacity, and internal developer tooling.

Responsibilities

  • SimScale’s engineering teams run workloads directly on AWS; you will build the standards, guardrails, and self-service tooling that let them do so safely, raising reliability and security without slowing engineering velocity.
  • You will join a small, tightly knit infrastructure team supporting 50+ engineers across the company. This is a hands‑on senior individual contributor role; people management is not required, but there is a genuine path toward tech‑lead ownership as the team grows.
  • Evolve our Kubernetes platform: Evaluate and adopt technologies such as Kubernetes Gateway API and service mesh patterns, and coordinate platform evolution across 10+ engineering teams.
  • Take observability to the next level: Drive organization-wide adoption of OpenTelemetry for distributed tracing and metrics, and help teams define meaningful SLOs.
  • Shape multi-region architecture and data residency: Support our move from an EU-centered footprint toward a global, multi-cloud architecture that satisfies disaster‑recovery and data‑residency requirements.
  • Own cloud cost and efficiency at scale: Keep petabyte-scale infrastructure cost‑efficient, secure, and well‑instrumented.
  • Improve tooling: Build self‑service AWS account provisioning, guardrails and AI‑assisted automations that help engineering teams manage infrastructure safely and efficiently at scale.

Qualifications

  • Strong systems foundation: You understand Linux internals and distributed systems well enough to debug complex production behavior.
  • Software development experience: Your background is rooted in software development, and you moved into SRE from there. You write production-quality software in at least one of Python, Go, Rust, or Java.
  • Security and compliance awareness: You understand how infrastructure decisions affect access control, auditability, disaster recovery, logging, and standards such as SOC 2.
  • Production debugging depth: You can investigate complex failures, communicate clearly during incidents, and turn findings into durable improvements.
  • Hands‑on cloud and infrastructure experience: AWS (or GCP), declarative infrastructure (Terraform), gitops‑workflow (ArgoCD) and container orchestration (Kubernetes).
  • 5+ years of professional experience in SRE, platform, or infrastructure engineering.
  • Clear communication: You can explain trade‑offs to engineering teams and help others adopt better platform practices without unnecessary friction.
  • Observability and reliability experience: You have worked with OpenTelemetry, Prometheus, distributed tracing, monitoring, and meaningful SLOs/SLIs.
  • An open source portfolio or contributions.
  • Prior technical leadership experience, especially in infrastructure, reliability, or platform engineering.

Benefits

  • Mobile Working: Modern technology coupled with the widespread use of digital communication enables remote working. We embrace mobile working as it offers many possibilities for our employees.
  • Competitive Health Benefits: Whilst working in a competitive and ambitious environment, it is just as important to take care of your health. At SimScale, we’ve got you covered.
  • Discounted Gym Membership: We offer a gym membership with multiple locations to make it easier for our employees to invest in their health, well‑being, and workout regularly.
  • Tech Talks and Social Events: Whether it’s through tech talks, social events or other organized activities, we are not just working together, we like sharing, and enjoying each other’s company.
  • Flexible Working Hours: It doesn’t matter if you are an early bird or a night owl, at SimScale you have the freedom to plan your workday.
  • Learning and Development: With the right training in place, we offer our employees the opportunity to further develop their competencies and skill-sets, to ultimately become more successful and satisfied.
  • Child Care Contributions: In addition to flexible working hours, and home office possibilities to support family commitments, we also provide contributions for our mini SimScaler’s nursery or kindergarten childcare.
  • Retirement Plan: We are future-oriented, and take care of our employees! We provide employees with a retirement plan that will contribute towards a comfortable and secure future.

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    20 Jul 2026
  • Standort:

    München

    Einsatzort:

    Germany (inferred from text)
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!