Inference Optimization — Member of Technical Staff

Stellenbeschreibung:

Back to careers


Inference Optimization — Member of Technical Staff


Berlin, Full-time, In person


Who Are We


Models should learn from what happens after deployment. We're building the loops that make this possible.


We will be the default inference provider for specialized tokens, running models that continuously adapt to each customer and workload. That means serving many continuously adapting models with frontier-level performance and attractive economics.


The Work



  • Finding novel, hardware-informed quantization methods across model architectures

  • Writing and tuning Triton and CUDA kernels

  • Finding bottlenecks across compute, memory, and communication

  • Efficiently serving LoRAs, learned memory, specialized weights, and other forms of model adaptation


This is capture-the-flag for GPU efficiency: profile the system, find something everyone else missed, and prove the gain on real workloads.


Great candidates might come from ML systems, compilers, mathematics, physics, competitive programming, security, or HPC. The common thread is strong first-principles reasoning, an obsession with efficiency, and a habit of going deep on hard problems for the fun of it.


We work together in person from our office in Berlin.


Interview process



  • Two technical interviews

  • An onsite interview at our Berlin office

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    18 Aug 2026
  • Standort:

    Berlin
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!