GPU Performance Engineer

Stellenbeschreibung:

We make complex software run efficiently on any processor using a self-learning compiler and cloud-scale optimization infrastructure. Our team brings together researchers and engineers from RWTH Aachen, TU Munich, TU Darmstadt, and ETH Zurich to tackle some of the hardest problems in systems and infrastructure software. If you want to work on deeply technical challenges with real-world impact, join us and help shape the future of compute. As a GPU Performance Engineer at Daisytuner, you will operate at the intersection of GPU architecture, performance engineering, and compiler technology. You will analyze demanding machine learning and high-performance computing workloads, identify performance bottlenecks, and develop highly optimized GPU kernels. Working closely with our compiler engineers, you will translate performance insights into compiler transformations, optimization heuristics, and autotuning strategies that enable our compiler to automatically generate high-performance GPU code.

What you will be doing

  • Analyze and optimize performance-critical GPU workloads, with a strong focus on machine learning applications
  • Benchmark, profile, and hand-tune GPU kernels across modern accelerator architectures
  • Identify bottlenecks related to memory hierarchy, occupancy, instruction throughput, synchronization, and kernel launch configuration
  • Develop highly optimized CUDA kernels and GPU programming techniques for real-world workloads
  • Use profiling and performance analysis tools to understand kernel behavior and hardware utilization
  • Collaborate closely with compiler engineers to translate manual optimization techniques into automated compiler transformations and optimization strategies
  • Design reproducible benchmarking methodologies and performance evaluation pipelines
  • Evaluate optimization quality across different GPU architectures and vendors
  • Contribute to the development of performance models and heuristics for GPU optimization
  • Stay up to date with the latest GPU architectures, programming models, and optimization techniques

What you are bringing

  • Master's degree (or equivalent experience) in Computer Science, Electrical Engineering, Mathematics, or a related technical field
  • Strong C/C++ and CUDA programming skills
  • Deep understanding of modern GPU architectures and performance optimization
  • Experience optimizing machine learning, HPC, or other performance-critical GPU workloads
  • Solid understanding of GPU memory hierarchies, shared memory, caches, register usage, occupancy, warp scheduling, and instruction-level performance
  • Experience profiling GPU applications using tools such as NVIDIA Nsight Compute, Nsight Systems, CUPTI, rocProfiler, or similar
  • Ability to analyze low-level performance bottlenecks and systematically improve kernel efficiency
  • Strong Python skills for benchmarking, automation, and experimentation
  • Comfortable working in Linux environments and with Git-based workflows
  • Strong communication skills in English and/or German
  • A structured, analytical working style and willingness to take ownership in an early-stage environment

Nice to have

  • Experience with GPU programming frameworks such as CUTLASS, Triton, cuBLAS, cuDNN, ROCm, or SYCL
  • Familiarity with compiler infrastructures such as LLVM, MLIR, or similar systems
  • Experience with kernel fusion, code generation, or compiler optimizations
  • Knowledge of transformer architectures, LLM inference/training, or other modern ML workloads
  • Experience across multiple GPU vendors (NVIDIA, AMD, Intel)

What we offer

  • A small, highly technical team with direct impact on core technology
  • Competitive compensation and potential equity participation
  • The opportunity to work at the intersection of compilers, machine learning, and high-performance computing
  • Real ownership over GPU optimization strategies that directly influence the capabilities of our compiler and the performance of production workloads

We are an equal opportunity employer and welcome applications from people of all backgrounds. We value diversity and believe that different perspectives make us stronger. We do not discriminate based on gender, nationality, ethnic origin, religion, disability, age, sexual orientation, or identity.

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    05 Sep 2026
  • Standort:

    München
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

    Development & IT
  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!