NVIDIA

Senior GPU Networking Architect

Stellenbeschreibung:

Overview

As a Senior GPU Networking Architect, you’ll design and optimize GPU communication kernels for large-scale AI systems. You’ll work at the intersection of GPU compute and networking to improve kernel efficiency and reduce latency. Collaborating with cross-functional teams, you will shape GPU-aware communication strategies and evaluate new ideas through experiments and modeling. This role offers impact across AI frameworks and large infrastructure, contributing to NVIDIA’s mission to power the world’s AI workloads. You’ll thrive in a fast-paced, innovation-driven environment and help push the boundaries of GPU networking.

Leistungen / Benefits
  • competitive salaries
  • comprehensive benefits package
  • remote work options
  • career growth opportunities
Verantwortungsbereiche
  • Build, implement, and optimize GPU communication kernels for collective and point-to-point operations in large-scale AI systems
  • Enhance kernel efficiency by applying deep knowledge of GPU architectures and execution pipelines to minimize latency
  • Develop GPU-resident communication primitives and device-side APIs for fine-grained device-to-device data movement
  • Profile and tune GPU kernels end-to-end, identifying bottlenecks at compute-memory-network intersections and drive optimizations
  • Collaborate with network software, hardware, and AI framework teams to co-design communication strategies aligned with GPU execution patterns
  • Build proofs-of-concept, run experiments, and modelably evaluate new communication strategies before production
  • Contribute to evolving programming models that expose GPU-aware networking capabilities to developers
Zentrale Anforderungen
  • 5+ years of hands-on CUDA programming with optimization of non-trivial GPU kernels
  • M.Sc. or equivalent in computer science, computer engineering, or closely related field
  • Strong understanding of GPU architecture fundamentals (warp scheduling, shared memory, L2 cache, memory coalescing, occupancy, asynchronous execution)
  • Systems-level C/C++ development in performance-critical environments
  • Familiarity with GPU data movement mechanisms such as GPUDirect RDMA and GPU-initiated communication
  • Ability to read and reason about GPU performance profiles (Nsight Compute, Nsight Systems) and translate observations into optimizations
  • Strong collaboration skills in a multi-national, interdisciplinary environment
  • collaboration
  • problem-solving
  • communication
  • CUDA programming
  • GPU kernel optimization
  • GPU architecture (warp scheduling, memory hierarchy)
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    14 Sep 2026
  • Standort:

    Berlin

    Einsatzort:

    Santa Clara, CA, USA
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!