Jobgether

Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)

Stellenbeschreibung:

This position islisted on behalfof a partnercompany, who manages allapplications and nextsteps. Our partneris looking foraSenior Technical Operations& Deployment Engineer (GPU Cloud Infrastructure)based inGermany.

This is a highly hands-oninfrastructure role focusedon deploying, commissioning, and operating GPUcloud environments acrossregional and coredatacenters.
You will turnvalidated architecturesand bills ofmaterials into production-ready infrastructure spanninghardware, networking, storage, Linux, andplatform software.
The role sitsat the intersectionof datacenter operations, GPU infrastructure, networkengineering, and cloudplatform operations.
You will workwith high-densityNVIDIA GPU systems, advanced networking,storage platforms, Kubernetes, virtualization, and observabilitytooling.
As a practicaltechnical escalation point, you will troubleshootcomplex issues acrossphysical and softwarelayers and driveincidents through resolution.
You willalso help establishdeployment standards, validationprocedures, documentation, and operational practicesfor a rapidlyevolving AI infrastructureenvironment.
The role offersbroad technical ownershipinan international, fast-moving setting wherehands-on executionand operational excellenceare essential.

Accountabilities

  • Datacenter deployment: Coordinate deployments with datacenter providers, integrators, logistics teams, vendors, and internal engineering; validate rack layouts, power, cooling, airflow, cabling, labeling, and physical readiness.
  • Rack and infrastructure commissioning: Support rack-and-stack activities for GPU and CPU servers, storage, switches, routers, firewalls, PDUs, serial/OOB systems, and supporting infrastructure.
  • Cabling and connectivity: Validate fiber and copper cabling, optics, transceivers, breakout cables, port mappings, link speeds, redundancy, and management, storage, north-south, and east-west connectivity.
  • Hardware bring-up: Commission GPU servers, storage nodes, and platform infrastructure while validating BIOS, BMC, firmware, NICs, DPUs, GPUs, NVMe, RAID/HBA, PCIe topology, NUMA, thermals, power, and hardware health.
  • Hardware validation: Execute burn-in, stress, network, storage, and acceptance testing before production handover; troubleshoot issues involving GPUs, DPUs, NICs, optics, memory, disks, firmware, and BIOS.
  • Network deployment support: Work with network engineering to validate switch configurations, routing, VLAN/VRF segmentation, BGP, ECMP, EVPN/VXLAN, OVS/OVN, VyOS, firewalls, WAF infrastructure, and customer connectivity.
  • AI networking: Support validation of RoCE/RDMA fabrics for distributed AI workloads and troubleshoot issues such as link flaps, MTU mismatches, route errors, packet loss, PFC/ECN problems, and congestion.
  • Platform installation: Install and validate Ubuntu/Linux environments, NVIDIA drivers, CUDA, OFED or inbox drivers, Docker/containerd, KVM/QEMU, platform agents, and GPU infrastructure components.
  • Cloud and Kubernetes environments: Support CloudStack, Kubernetes, KubeVirt, GPU Operator, CSI/CNI integrations, GPU passthrough, SR-IOV, BlueField DPUs, VM networking, and container networking.
  • Storage integration: Support integration and validation of StorPool, Weka, local NVMe, and other supported storage platforms.
  • Operational readiness: Execute acceptance testing, produce deployment readiness reports, maintain runbooks, and ensure infrastructure is fully operational before customer or production handover.
  • Day-2 operations: Perform controlled firmware, OS, driver, BIOS, switch, and hardware maintenance while supporting production incidents and infrastructure escalations.
  • Incident management: Investigate operational failures, perform root-cause analysis, distinguish temporary workarounds from permanent fixes, and work with engineering to eliminate recurring issues.
  • Observability: Validate telemetry and monitoring across hosts, GPUs, DPUs, switches, storage, and platform components using tools such as Zabbix, Prometheus, Grafana, Loki, DCGM/NVML, and NVIDIA NetQ or equivalents.
  • Performance validation: Establish baselines for GPU, network, storage, and host performance and support benchmarking and infrastructure validation.
  • Documentation: Maintain accurate as-built records covering rack elevations, cable maps, port mappings, serial numbers, asset records, IP allocations, changes, and operational procedures.
  • Cross-functional coordination: Partner with infrastructure, networking, storage, platform, fleet automation, observability, product engineering, sales engineering, and service delivery teams.
  • Vendor management: Coordinate with datacenter providers, system integrators, server and storage vendors, NVIDIA, and networking suppliers to resolve deployment and infrastructure issues.
  • Continuous improvement: Feed field experience back into reference architectures, BOMs, rack designs, cabling standards, deployment playbooks, validation processes, and automation.

Requirements

  • Datacenter infrastructure: Strong hands-on experience deploying and maintaining datacenter infrastructure, ideally within GPU, HPC, AI cloud, private cloud, or high-density compute environments.
  • Bare-metal deployment: Proven ability to bring servers from physical installation and bare metal through validation and production readiness.
  • GPU infrastructure: Experience with NVIDIA GPU servers, drivers, firmware, PCIe topology, hardware validation, and high-performance compute environments.
  • Next-generation AI infrastructure: Familiarity with NVL72-style rack-scale architectures, NVLink/NVSwitch domains, in-rack networking, high-density power delivery, and OEM/NVIDIA validation requirements.
  • Datacenter readiness: Ability to assess power density, cooling, rack dimensions, floor loading, containment, serviceability, maintenance access, and other physical requirements for AI infrastructure.
  • Linux: Strong Linux troubleshooting capabilities and experience managing operating systems, kernels, drivers, and hardware interfaces.
  • Networking: Practical knowledge of VLANs, VRFs, BGP, ECMP, OVS/OVN, routing, OOB management, and high-speed datacenter connectivity.
  • GPU networking: Familiarity with NVIDIA/Mellanox networking, RoCE/RDMA, SR-IOV, BlueField DPUs, and high-performance east-west infrastructure.
  • Virtualization and containers: Experience with KVM/QEMU, VFIO, PCI passthrough, Docker/containerd, Kubernetes, and/or KubeVirt.
  • Storage: Experience integrating or troubleshooting local NVMe, storage nodes, and enterprise or distributed storage platforms.
  • Automation: Familiarity with Terraform, Ansible, Bash, and/or Python for deployment, validation, configuration, or operational automation.
  • Observability: Experience with infrastructure monitoring, telemetry, logs, metrics, health checks, and performance dashboards.
  • Documentation: Strong attention to detail and discipline in producing accurate as-built documentation, runbooks, validation records, and handover materials.
  • Troubleshooting: Strong systems-thinking ability across physical infrastructure, hardware, firmware, networking, Linux, storage, and platform layers.
  • Operational mindset: Comfortable supporting production environments, deployment windows, operational escalations, and customer-impacting incidents.
  • Communication: Able to clearly explain technical issues, risks, workarounds, and permanent solutions to engineering teams, vendors, and leadership.
  • Personal qualities: Highly practical, detail-oriented, calm under pressure, autonomous, and comfortable working both inside datacenters and remotely with smart-hands teams.

Benefits

  • Attractive compensation package reflecting your expertise, experience, transferable skills, and market conditions.
  • Full-time or contract engagement, depending on the agreed arrangement.
  • Europe-based remote working environment with flexibility.
  • Opportunity to work on cutting-edge GPU cloud and AI infrastructure at significant scale.
  • Hands-on exposure to NVIDIA GPU platforms, high-density datacenter environments, RoCE/RDMA networking, Kubernetes, virtualization, storage, and advanced observability.
  • Broad cross-functional scope spanning hardware, datacenter operations, networking, storage, Linux, and cloud platforms.
  • High-impact role within a fast-growing international scale-up.
  • Strong opportunities for technical growth and career development as the infrastructure platform expands.
  • Friendly, diverse, flexible, and international working environment.
  • Inclusive workplace committed to equal opportunity and respect for all qualified candidates.

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    30 Aug 2026
  • Standort:

    Remote

    Einsatzort:

    France
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!