HPC Infrastructure Engineer – GPU Clusters

Stellenbeschreibung:

  • Operate and improve the GPU fleet end to end, including provisioning, scheduling, monitoring, upgrades, and capacity planning
  • Build automation for node health checks, automated draining and remediation, and burn-in pipelines
  • Own the infrastructure stack beneath training code, including OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, and InfiniBand/RoCE networking
  • Run and tune job scheduling with Slurm or similar systems
  • Build and maintain high-performance storage for datasets and checkpoints
  • Investigate and resolve performance problems involving stragglers, degraded links, thermal issues, and faulty GPUs
  • Evaluate rented GPU capacity by benchmarking providers, validating capacity, and enforcing SLAs
  • Perform hands-on hardware work, including racking, cabling, and diagnostics
  • Coordinate with datacenter staff and vendors
  • Maintain cluster security through access control, network isolation, and secrets management
  • Work directly with researchers to improve training throughput and researcher velocity

Requirements

  • Production experience running large-scale Linux server or GPU environments
  • Strong knowledge of NVIDIA drivers, CUDA, NCCL, and DCGM, or deep systems experience with ability to learn hardware stacks quickly
  • Experience with bare-metal environments, server hardware, and high-speed networking
  • Proficiency in Python and/or Bash automation
  • Experience with infrastructure-as-code tools such as Ansible or Terraform
  • Ability to analyze metrics, logs, and PromQL
  • Experience supporting ML training workloads from the infrastructure side (nice to have)
  • Experience evaluating and working with GPU cloud providers (nice to have)
  • Experience with parallel filesystems such as WEKA or VAST, or large-scale object storage (nice to have)
  • Experience with BMC/IPMI/Redfish automation and PXE provisioning at scale (nice to have)
  • Power and cooling awareness for dense GPU deployments (nice to have)
  • Willingness to perform datacenter trips and hands-on hardware work

Core Competencies

Demonstrates expertise in managing and optimizing GPU infrastructure, including proficiency in NVIDIA drivers, CUDA, and automation tools. Capable of performing hands-on hardware work while ensuring cluster security and high-performance storage management.

Highest-signal resume keywords

  • GPU Infrastructure Management
  • NVIDIA Drivers and CUDA Proficiency
  • Python and Bash Automation
  • Infrastructure-as-Code Tools (Ansible, Terraform)
  • Performance Analysis and Troubleshooting

ATS Optimization Keywords

Hard Skills

  • Linux Server Management
  • GPU Environment Operations
  • Job Scheduling with Slurm
  • High-Speed Networking
  • Metrics and Log Analysis
  • Bare-Metal Environments
  • Parallel Filesystems (WEKA, VAST)
  • BMC/IPMI/Redfish Automation
  • PXE Provisioning
  • Capacity Planning

Soft Skills

  • Collaboration with Researchers
  • Problem-Solving
  • Communication with Datacenter Staff

Industry Keywords

  • GPU Cloud Providers
  • ML Training Workloads
  • Cluster Security
  • Access Control
  • Network Isolation

Tools & Technologies

  • NCCL
  • DCGM
  • PromQL
  • InfiniBand/RoCE Networking
  • Automated Health Checks

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    07 Sep 2026
  • Standort:

    Remote
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!