Mid-Level SRE

Stellenbeschreibung:

  • Prevent production incidents by identifying operational risks, points of failure, and noisy alerts before they become problems.
  • Participate in building the observability platform (logs, metrics, and tracing), contributing to the definition of processes, SLIs, SLOs, and error budgets.
  • Conduct root cause analyses and lead postmortems, documenting lessons learned and tracking action plans through to completion.
  • Resolve critical incidents through troubleshooting in AWS and on-premises production environments.
  • Identify FinOps opportunities and contribute to cloud cost predictability.
  • Collaborate with the development team on the continuous improvement of application reliability and performance.
  • Support scalability initiatives and the creation of new infrastructure, with a focus on automation.
  • Contribute to the team’s SRE maturity through practices, documentation, and incident management culture.

Requirements

  • Solid experience troubleshooting production environments and distributed systems.
  • Production experience with AWS (EC2, networking, load balancing, IAM); experience with on-premises environments is a plus.
  • Knowledge of Docker/Docker Compose, including running containers in production.
  • Experience with observability and monitoring tools (e.g., Grafana, Prometheus, Datadog, SigNoz, or similar) and OpenTelemetry (logs, metrics, and tracing).
  • Understanding of SRE practices: SLI/SLO, error budgets, incident management and resolution, and postmortems.
  • Strong knowledge of Linux, networking, and protocols (HTTP, TCP/IP, DNS).
  • Strong communication skills, autonomy, and resilience when working during critical incidents, including occasional direct interaction with customers.
  • Experience with automation (Python, Bash, Terraform, Ansible, or similar) is a plus.
  • Familiarity with DevOps practices, including CI/CD and infrastructure as code (IaC), is a plus.

Core Competencies

Demonstrates expertise in troubleshooting production environments and distributed systems, with a strong focus on AWS, observability tools, and SRE practices. Capable of driving incident management and resolution while contributing to infrastructure automation and cost predictability.

Highest-signal resume keywords

  • AWS Production Experience
  • Observability Tools Proficiency
  • SRE Practices Knowledge
  • Linux Networking Expertise
  • Automation Skills

Hard Skills

  • Troubleshooting Production Environments
  • Distributed Systems
  • Docker/Docker Compose
  • OpenTelemetry
  • SLI/SLO Understanding
  • Incident Management
  • Python
  • Bash
  • Terraform
  • Ansible

Soft Skills

  • Strong Communication Skills
  • Autonomy
  • Resilience

Industry Keywords

  • FinOps
  • Cloud Cost Predictability
  • Infrastructure as Code (IaC)
  • CI/CD
  • Incident Management Culture

Tools & Technologies

  • AWS (EC2, Networking, Load Balancing, IAM)
  • Grafana
  • Prometheus
  • Datadog
  • SigNoz

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    09 Sep 2026
  • Standort:

    Remote
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

    Development & IT
  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!