For our growing team, we are looking for an experienced Sovereign Cloud / Site Reliability Engineer (SRE) to support the secure, reliable, and continuous operation of modern cloud-native and Kubernetes platforms.
You will work with an experienced DevOps/SRE team on highly secure, business-critical platforms, taking responsibility for platform operations, automation, monitoring, incident management, security, and continuous improvement.
The role focuses particularly on Kubernetes, CI/CD, Infrastructure as Code, observability, logging, and operational excellence within a highly available Unified Observability Platform in a 24x7 operational environment.
Mandatory Requirements – Please Read Before Applying
Citizenship
You must hold valid citizenship in a country that is a full member of both the European Union (EU) and NATO .
If you hold multiple citizenships, all citizenships must be from countries that are members of both the EU and NATO.
Employment in Germany
You must:
- Be employed directly by a legal entity registered in Germany.
- Hold a German employment contract.
- Be subject exclusively to German labor law.
- Comply with all applicable German tax, social-security, and employment regulations.
- Reside in Germany.
Employment through non-German entities, including foreign subcontractors or affiliates, does not meet this requirement.
**
Security Clearance – Ü2
You must hold a valid and verifiable Ü2 security clearance in accordance with the German Security Clearance Act (Sicherheitsüberprüfungsgesetz – SÜG) and applicable preventive personnel sabotage-protection requirements.
Tasks
Kubernetes & Platform Operations
- Operate and maintain Kubernetes clusters with Gardener, including workloads, deployments, Helm charts and platform components.
- Troubleshoot availability, performance and deployment issues.
- Ensure secure, scalable, resilient and highly available platforms.
CI/CD & Automation
- Operate and optimize Jenkins and ArgoCD pipelines for automated deployments.
- Implement Infrastructure as Code (IaC) and Git-based deployment workflows.
- Develop automation and operational tools using Python, Go and/or Bash.
- Automate provisioning, health and compliance checks, alerting and reporting.
Monitoring & Observability
- Manage Prometheus, Thanos and OpenTelemetry environments, including scrape jobs, alert rules and PromQL.
- Develop and maintain Grafana dashboards.
- Continuously improve monitoring and alerting capabilities.
Logging & Log Management
- Operate and optimize Elasticsearch/OpenSearch, Logstash and Kibana.
- Monitor log ingestion, storage, performance and reliability.
- Support centralized troubleshooting, anomaly detection, security and compliance.
Integration & Operations
- Integrate observability and logging platforms with ServiceNow, PagerDuty and other enterprise tools via secure APIs.
- Handle operational requests.
- Participate in Scrum, DevOps and service‑improvement activities.
- Collaborate with internal teams, SAP, suppliers and stakeholders.
Incident & Problem Management
- Participate in a 24/7 on‑call and shift rotation, including weekends and public holidays.
- Respond to platform, monitoring, logging and deployment incidents.
- Perform Root Cause Analysis (RCA).
- Support Major Incident Management (MIM).
- Implement sustainable corrective actions.
- Continuously improve platform stability, resilience and operational processes.
Requirements
Key Technologies
Kubernetes | Gardener | Helm | Jenkins | ArgoCD | Git | IaC | Python | Go | Bash | Prometheus | PromQL | Thanos | OpenTelemetry | Grafana | Elasticsearch | OpenSearch | Logstash | Kibana | ServiceNow | PagerDuty | REST APIs
Technical Skills & Experience
- Proven experience as a Sovereign Cloud Engineer and/or Site Reliability Engineer (SRE) .
- Strong experience operating and maintaining Kubernetes clusters , ideally with Gardener.
- Hands‑on experience with Kubernetes workloads, deployments, Helm charts and platform components .
- Experience troubleshooting availability, performance and deployment issues .
- Experience with Jenkins and ArgoCD for automated CI/CD deployments.
- Solid understanding of Infrastructure as Code (IaC) and Git-based deployment workflows.
- Programming or scripting experience with Python, Go and/or Bash .
- Experience with automation of provisioning, health checks, compliance checks, alerting and reporting .
- Hands‑on experience with Prometheus, PromQL, Thanos and OpenTelemetry .
- Experience developing and maintaining Grafana dashboards and monitoring/alerting solutions.
- Experience with Elasticsearch/OpenSearch, Logstash and Kibana .
- Understanding of log ingestion, storage, performance and reliability .
- Experience supporting centralized troubleshooting, anomaly detection, security and compliance .
- Experience integrating platforms with ServiceNow, PagerDuty and other enterprise tools via secure REST APIs .
- Experience with incident management, Root Cause Analysis (RCA) and Major Incident Management (MIM) .
- Experience working in 24/7 operational environments , including on-call and shift rotations.
- Strong understanding of platform security, scalability, resilience, reliability and high availability .
- Experience working in Scrum and DevOps environments .
- Strong collaboration skills and experience working with internal teams, SAP, suppliers and other stakeholders .
Languages
- English: Required
- German: A plus
Benefits
What You Can Expect
- A technically challenging role within a highly secure and business‑critical cloud environment .
- The opportunity to work with modern cloud-native, Kubernetes and observability technologies .
- An international working environment with teams and stakeholders across different locations.
- The opportunity to contribute to automation, platform reliability and continuous improvement .
- Remote / hybrid working possibilities within Germany.
We Look Forward to Hearing from You
Are you ready to bring your expertise to a challenging cloud engineering environment?
We look forward to receiving your application and getting to know you.
#J-18808-Ljbffr