HPC System Engineer

Stellenbeschreibung:

Responsibilities

  • Work closely with hardware and development teams to profile and analyze GPU performance at system and kernel level.
  • Evaluate and compare GPU performance across different platforms, architectures, and software stacks (e.g., CUDA, ROCm).
  • Perform acceptance testing for new GPU clusters, ensuring hardware and software meet performance, stability, and compatibility requirements for AI workloads.
  • Perform experiments across diverse GPU system configurations to assess the impact of varying interconnect strategies and system‑level optimizations on performance and scalability.

Requirements

  • Proficient in Unix/Linux, plus Python and Bash for automation.
  • Good understanding of the GPU stack: CUDA, NCCL, drivers, and relevant libraries.
  • Proven ability to troubleshoot complex system issues including hardware, software, and networking problems.
  • Familiarity with containerized environments (e.g., Docker, Kubernetes).

Hard Skills

  • GPU performance analysis
  • CUDA
  • ROCm
  • Python
  • Bash
  • NCCL
  • troubleshooting
  • system optimization
  • acceptance testingperformance evaluation

Soft Skills

  • collaboration
  • problem‑solving

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    31 Jul 2026
  • Standort:

    Remote
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!