We are building a European AI Infrastructure & Platform Operations team responsible for operating large-scale AI infrastructure environments powered by NVIDIA GPUs, high-performance networking, Kubernetes, and next-generation platform technologies
The team is responsible for ensuring the availability, performance, and operational stability of critical AI infrastructure platforms deployed across multiple datacenters. Working at the intersection of infrastructure, networking, and platform operations, you will help support the environments that power modern AI workloads
This is an opportunity to work with some of the latest technologies in AI infrastructure while contributing to the evolution of AI-powered operational services through platforms such as k0rdent AI
Monitor, operate, and support production AI infrastructure platforms
Investigate and resolve infrastructure, networking, hardware, and platform-related incidents
Support NVIDIA GPU infrastructure and associated platform services
Monitor and troubleshoot Kubernetes-based environments
Investigate performance, availability, and reliability issues across infrastructure and platform components
Collaborate with engineering teams, hardware vendors, datacenter personnel, and service delivery teams to resolve technical issues
Participate in incident response, root cause analysis, and operational improvement activities
Contribute to improvements in monitoring, observability, automation, and operational processes
Maintain operational documentation, runbooks, and knowledge articles
Qualifications
Strong Linux administration and troubleshooting skills
Ability to work within a shift-based operational environment
Experience supporting production infrastructure and services
Excellent communication and collaboration skills
Experience working within structured operational and incident management processes
3+ years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or related technical roles
Good understanding of networking concepts and experience diagnosing infrastructure-related issues
Strong analytical and problem-solving skills
Working knowledge of Kubernetes in production environments
NVIDIA GPU infrastructure and accelerated computing platforms
InfiniBand networking and NVIDIA UFM
Kubernetes platform operations
AI infrastructure or HPC environments
Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry
Site Reliability Engineering (SRE) or Platform Engineering
Infrastructure automation technologies and Infrastructure-as-Code practices
Large-scale distributed systems and production platforms
#J-18808-Ljbffr
NOTE / HINWEIS:
EN: Please refer to Fuchsjobs for the source of your application
DE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung
Stelleninformationen
Veröffentlichungsdatum:
20 Jul 2026
Standort:
Berlin
Typ:
Vollzeit
Arbeitsmodell:
Vor Ort
Kategorie:
Erfahrung:
2+ years
Arbeitsverhältnis:
Angestellt
KI Suchagent
Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!