Staff Software Engineer, Machine Learning Inference Platform

Stellenbeschreibung:

Responsibilities

  • Design platform architecture for multi-tenant inference workloads across serving, orchestration, control plane, APIs, SDKs, observability, and model‑engine integration.
  • Develop robust API layers (gRPC, WebSockets, REST, etc.) and developer SDKs that abstract complex distributed inference orchestration into seamless, reliable token streams.
  • Build and harden a multi-tenant control plane to enable accurate metering, rate limiting, quotas, tenant isolation, and noisy‑neighbor fairness across the platform.
  • Optimize inference performance across the entire system stack, including the model engine layer.
  • Build observability and SLOs to gain insights into system economics, cache‑hit rates, GPU utilization, and cost accounting per model and per tenant.
  • Partner with product and infrastructure teams on model onboarding, capacity planning, external API contracts, and customer adoption.
  • Promote Engineering Excellence: maintain a high bar for engineering excellence in one’s own work while setting a culture of excellence within the team.

Requirements

  • Education: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • Experience: 7+ years of experience building and operating backend distributed systems end to end.
  • Demonstrated cross‑team technical leadership in backend distributed systems, ML infrastructure, inference serving, or high‑performance compute platforms.
  • Strong Data & ML systems fundamentals: data‑intensive distributed systems, concurrency, networking, and performance profiling.
  • Hands‑on experience running large‑scale inference services on GPUs, including KV caches, prefill/decode stages, and throughput/latency trade‑offs.
  • Direct experience with inference engines (TensorRT, vLLM, etc.) or serving frameworks (Dynamo, Triton, or equivalent).
  • Technical Skills: strong programming skills in C++, Go, Rust, or Python.
  • Familiarity with deep learning frameworks (PyTorch, etc.) and model parallelism.
  • Familiarity with GPU computing primitives such as CUDA, NCCL, NVLink, and hardware‑specific optimizations.
  • Practical understanding of high‑performance networking architectures, including InfiniBand, RoCE, and low‑latency cluster communication.
  • Communication: excellent verbal and written communication skills, with the ability to convey complex technical concepts to non‑technical stakeholders.
  • Autonomous vehicles (AV) experience is a bonus.

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    18 Aug 2026
  • Standort:

    WorkFromHome
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!