Mirantis

AI Infrastructure & Platform Operations Engineer

Stellenbeschreibung:

Responsibilities

  • We are building a European AI Infrastructure & Platform Operations team responsible for operating large-scale AI infrastructure environments powered by NVIDIA GPUs, high-performance networking, Kubernetes, and next-generation platform technologies
  • The team is responsible for ensuring the availability, performance, and operational stability of critical AI infrastructure platforms deployed across multiple datacenters. Working at the intersection of infrastructure, networking, and platform operations, you will help support the environments that power modern AI workloads
  • This is an opportunity to work with some of the latest technologies in AI infrastructure while contributing to the evolution of AI-powered operational services through platforms such as k0rdent AI
  • Monitor, operate, and support production AI infrastructure platforms
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents
  • Support NVIDIA GPU infrastructure and associated platform services
  • Monitor and troubleshoot Kubernetes-based environments
  • Investigate performance, availability, and reliability issues across infrastructure and platform components
  • Collaborate with engineering teams, hardware vendors, datacenter personnel, and service delivery teams to resolve technical issues
  • Participate in incident response, root cause analysis, and operational improvement activities
  • Contribute to improvements in monitoring, observability, automation, and operational processes
  • Maintain operational documentation, runbooks, and knowledge articles

Qualifications

  • Strong Linux administration and troubleshooting skills
  • Ability to work within a shift-based operational environment
  • Experience supporting production infrastructure and services
  • Excellent communication and collaboration skills
  • Experience working within structured operational and incident management processes
  • 3+ years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or related technical roles
  • Good understanding of networking concepts and experience diagnosing infrastructure-related issues
  • Strong analytical and problem-solving skills
  • Working knowledge of Kubernetes in production environments
  • NVIDIA GPU infrastructure and accelerated computing platforms
  • InfiniBand networking and NVIDIA UFM
  • Kubernetes platform operations
  • AI infrastructure or HPC environments
  • Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry
  • Site Reliability Engineering (SRE) or Platform Engineering
  • Infrastructure automation technologies and Infrastructure-as-Code practices
  • Large-scale distributed systems and production platforms

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    20 Jul 2026
  • Standort:

    Berlin
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!