Senior Site Reliability Engineer, SRE

Stellenbeschreibung:

  • Operate, harden and extend production OpenShift / OKD / Kubernetes clusters across on-premises and hybrid environments.
  • Supporting migrations, helping modernise the underlying compute and infrastructure layer.
  • Own CI/CD processes across the full lifecycle of platform and application components.
  • Own and mature GitOps deployment practices, particularly using tools such as Argo CD.
  • Support cloud-native application delivery using tools such as Helm and Kustomize.
  • Maintain and improve core platform services including Keycloak, ingress, observability, certificate management, service mesh and container registry capabilities.
  • Build and operate observability across logs, metrics, traces, alerting, SLOs and error budgets.
  • Improve platform hardening in line with secure and regulated environment requirements.
  • Automate repeatable operational tasks using tools such as Ansible, Terraform, Helm, Kustomize, Go, Python or similar.
  • Lead incident response activity, support blameless post-mortems and drive systemic fixes.
  • Partner with networking and security teams on platform integration, segmentation, load balancing and accreditation evidence.
  • Create and maintain clear technical documentation, runbooks, design notes and operational guidance.
  • Mentor engineers and act as a senior technical authority across cloud and Kubernetes operations.
  • Participate in an on-call rota, with appropriate compensation.

Requirements

  • Strong experience running production Kubernetes environments, not just consuming or deploying into them.
  • Strong Linux fundamentals, including systems, networking, storage and performance troubleshooting.
  • Experience with Kubernetes distributions such as OKD, OpenShift, vanilla Kubernetes, Rancher, EKS, AKS or GKE.
  • Infrastructure as code experience, including Ansible plus Terraform or equivalent.
  • Experience with Helm, Kustomize or similar cloud-native deployment tooling.
  • GitOps and CI/CD experience managing full application and component lifecycles, using tools such as Argo CD, Flux, GitHub Actions or similar.
  • Observability across logs, metrics and traces, using tools such as Prometheus, Grafana, Elastic Stack, LGTM and OpenTelemetry.
  • Experience with identity and access technologies such as OIDC, SAML, SCIM or Keycloak.
  • Experience with virtualisation or infrastructure platforms such as KVM, libvirt or VMware.
  • Scripting or tooling experience using Go, Python, shell scripting or similar.
  • Strong troubleshooting, problem-solving and analytical skills.
  • Experience working in secure, regulated or enterprise-scale environments.
  • Strong written and verbal communication skills, with the ability to produce clear documentation, runbooks, post-mortems and technical guidance.
  • Eligibility to hold UK SC clearance.
  • Desirable (Not Essential)
  • Specific OpenShift or OKD experience, including operators, MachineConfig or SCCs.
  • Service mesh experience such as Istio or Linkerd.
  • Policy engine experience such as OPA, Gatekeeper or Kyverno.
  • Software supply chain security experience, including SBOMs, image signing, admission control or tools such as Sigstore.
  • Storage experience such as Ceph, Longhorn, OpenShift Data Foundation or equivalent.
  • Networking experience including BGP, VXLAN, Palo Alto or Juniper technologies.
  • AI, ML or GPU-enabled platform operations.
  • CKA, CKAD, CKS, Red Hat certifications or equivalent.
  • Active or recent UK SC clearance.
  • Recognised open-source contributions to the Kubernetes ecosystem.

Demonstrates expertise in operating and hardening Kubernetes environments, with a strong focus on CI/CD processes, GitOps practices, and cloud-native application delivery. Proficient in automation and observability tools, ensuring secure and efficient platform management.

Highest-signal resume keywords

  • Kubernetes Environment Management
  • CI/CD Processes Ownership
  • GitOps Deployment Practices
  • Infrastructure as Code (Ansible, Terraform)
  • Observability Tools (Prometheus, Grafana)

ATS Optimization Keywords

Hard Skills

  • Kubernetes
  • OpenShift
  • OKD
  • Ansible
  • Terraform
  • Helm
  • Kustomize
  • Go
  • Python
  • Linux Fundamentals

Soft Skills

  • Strong Troubleshooting Skills
  • Analytical Skills
  • Strong Written Communication
  • Strong Verbal Communication

Certifications & Qualifications

  • CKA
  • CKAD
  • CKS
  • Red Hat Certifications
  • UK SC Clearance

Industry Keywords

  • Cloud-Native
  • CI/CD
  • GitOps
  • Observability
  • Secure Environments
  • Regulated Environments
  • Enterprise-Scale
  • Identity and Access Management
  • Infrastructure as Code
  • Software Supply Chain Security

Tools & Technologies

  • Argo CD
  • Prometheus
  • Grafana
  • Elastic Stack
  • Keycloak
  • Service Mesh
  • Virtualization Platforms
  • GitHub Actions
  • OpenTelemetry
  • Networking Technologies

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    31 Jul 2026
  • Standort:

    Remote
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!