Senior Site Reliability Engineer, SRE, Backend

Stellenbeschreibung:

  • Own our AWS infrastructure end to end — Lambda, ECS Fargate, SQS, SNS, EventBridge, SES, Cognito, DynamoDB, RDS Postgres, and DMS
  • Manage everything as code in Terraform, with well-designed modules, clean state management, and a solid review workflow
  • Build and maintain CI/CD pipelines with safe rollout and rollback across web, mobile backends, and infrastructure
  • Design our event-driven services for resilience: retries, dead-letter queues, idempotency, graceful degradation
  • Own our Datadog and Sentry setup — define SLOs, build dashboards, and keep alerting actionable instead of noisy
  • Lead incident response and run blameless postmortems that actually change how we build
  • Harden our security posture: IAM, secrets management, network boundaries, Cognito auth flows, and vulnerability remediation
  • Protect sensitive health data and support our GDPR and compliance requirements
  • Monitor and optimize AWS spend without compromising reliability
  • Contribute to backend development — APIs, event consumers, data pipelines, and Postgres and DynamoDB performance
  • Participate in architectural discussions and mentor engineers on operational excellence

Requirements

  • 5+ years in SRE, DevOps, platform, or backend engineering, with real production ownership
  • Deep AWS experience across serverless and containers — Lambda, ECS Fargate, and debugging both under pressure
  • Strong Terraform skills, including module design and managing state across multiple environments
  • Hands-on experience with event-driven architecture (SQS, SNS, EventBridge) and a healthy respect for its failure modes
  • Solid PostgreSQL: query tuning, indexing, connection management, and zero-downtime migrations
  • Production experience with Datadog or a comparable observability platform (Grafana, New Relic, Honeycomb)
  • Comfortable writing production backend code in (Python / Node.js / Go)
  • Genuine on-call and incident response experience — you've led an incident and written the postmortem
  • Strong security fundamentals: IAM, least privilege, secrets, network isolation, common web vulnerabilities
  • Pragmatic about complexity — you reach for the simplest thing that meets the reliability bar
  • Effective communicator who can explain a technical tradeoff without jargon

Core Competencies

Demonstrates extensive expertise in AWS infrastructure management, including serverless and containerized environments, while ensuring compliance with GDPR and security best practices. Proficient in Terraform for infrastructure as code, alongside strong backend development capabilities in Python, Node.js, or Go.

Highest-signal resume keywords

  • AWS Infrastructure Management
  • Terraform Module Design
  • Event-Driven Architecture
  • PostgreSQL Performance Optimization
  • Incident Response Leadership

ATS Optimization Keywords

Hard Skills

  • AWS Lambda
  • ECS Fargate
  • Terraform
  • PostgreSQL
  • SQS
  • SNS
  • EventBridge
  • Datadog
  • Python
  • Node.js

Soft Skills

  • Effective Communication
  • Incident Response

Industry Keywords

  • GDPR Compliance
  • Security Fundamentals
  • Operational Excellence
  • Event-Driven Services
  • Production Ownership

Tools & Technologies

  • Datadog
  • Terraform
  • AWS
  • Sentry
  • CI/CD Pipelines

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    16 Aug 2026
  • Standort:

    Berlin
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!