Jobgether

Senior Site Reliability Engineer — Token Factory (Inference Platform)

Stellenbeschreibung:

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer — Token Factory (Inference Platform) based in Germany.

This is a senior engineering role focused on the reliability, performance, and observability of a large-scale AI inference platform.
You will help operate infrastructure serving foundation models across text, vision, audio, and emerging multimodal workloads.
The role combines Kubernetes, infrastructure-as-code, observability, automation, and production incident management at significant scale.
You will optimize GPU-heavy workloads, strengthen resilience, and ensure high-throughput APIs meet demanding reliability and cost targets.
You will work closely with software engineers and infrastructure teams to build self-healing systems and robust operational processes.
The environment is fast-moving, highly technical, international, and focused on solving complex infrastructure challenges for the AI ecosystem.
This is an opportunity to have a direct impact on the infrastructure powering next-generation AI applications.

Accountabilities

  • Own the reliability, performance, and observability of the inference platform and its supporting infrastructure.
  • Design, implement, and continuously improve telemetry pipelines covering metrics, logs, and traces .
  • Build monitoring and observability solutions capable of processing large volumes of production signals and converting them into actionable insights.
  • Configure and optimize Kubernetes infrastructure for high availability, scalability, and efficient resource utilization.
  • Tune Kubernetes autoscaling mechanisms to improve the efficiency and utilization of GPU resources.
  • Develop and maintain Terraform modules and infrastructure-as-code patterns that embed resilience and reliability into new clusters and services.
  • Design and improve request-routing, retry, and failure-handling mechanisms to minimize the impact of transient infrastructure or service failures.
  • Develop automation and operational tooling to detect, isolate, and remediate incidents quickly.
  • Create, maintain, and improve runbooks for incident response and operational procedures.
  • Participate in production incident management, troubleshooting issues and restoring services within demanding reliability objectives.
  • Lead or contribute to post-mortem processes and implement corrective actions to prevent recurring incidents.
  • Define and improve reliability practices for high-throughput APIs, including alerting strategies and Service Level Objectives (SLOs) .
  • Investigate distributed-system failures and performance issues across infrastructure and application layers.
  • Optimize systems from the kernel and infrastructure layer through to the application layer .
  • Support and improve the operation of GPU-intensive inference workloads and accelerator-based infrastructure.
  • Contribute to scaling the inference platform while balancing performance, reliability, and infrastructure costs .
  • Collaborate closely with software engineers to incorporate reliability and operational excellence into product and platform development.
  • Promote automation, self-healing capabilities, and engineering practices that reduce operational overhead and improve system resilience.

Requirements:

  • Significant experience in Site Reliability Engineering, Production Engineering, DevOps, or a closely related infrastructure discipline .
  • Deep practical knowledge of Kubernetes in production environments.
  • Strong experience with Prometheus and Grafana for monitoring, metrics, dashboards, and observability.
  • Advanced experience with Terraform and infrastructure-as-code practices.
  • Strong scripting and automation skills using Python and/or Bash .
  • Solid understanding of distributed systems and the ways production backends can fail under real-world conditions.
  • Experience designing effective alerts, monitoring strategies, and SLOs for high-throughput services or APIs.
  • Strong troubleshooting and debugging skills across infrastructure, networking, operating systems, and application layers.
  • Experience designing systems for high availability, resilience, scalability, and graceful failure recovery.
  • Hands‑on experience with GPU-heavy workloads or accelerator‑based infrastructure is highly valuable.
  • Familiarity with GPU inference technologies such as vLLM, Triton, Ray , or comparable accelerator and model‑serving stacks.
  • Experience with MLOps, model hosting, AI infrastructure, or machine‑learning platforms is advantageous.
  • Strong understanding of infrastructure automation, deployment, configuration management, and operational tooling.
  • Ability to analyze complex performance and reliability problems and translate findings into practical engineering improvements.
  • Strong incident‑management and root‑cause‑analysis capabilities.
  • Ability to collaborate effectively with software engineers and other technical teams to integrate reliability into platform development.
  • Proactive mindset with a strong focus on automation, self‑healing systems, and continuous improvement.
  • Comfortable working independently, taking ownership of critical infrastructure, and operating effectively in a fast‑paced technical environment.

Benefits:

  • Competitive compensation .
  • Career growth and continuous learning opportunities .
  • Flexibility and significant ownership in your work.
  • Collaborative and innovative international working environment.
  • Opportunity to work on high-impact AI infrastructure and inference technologies .
  • Exposure to large-scale GPU infrastructure and complex distributed systems.
  • Opportunity to contribute to infrastructure supporting next-generation multimodal AI applications.
  • Diverse and highly technical international teams.
  • Inclusive workplace committed to equal employment opportunities.
  • Workplace accommodations available throughout the application process where required.
  • Employment is subject to authorization to work in the country where the position is based.

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    21 Sep 2026
  • Standort:

    Remote

    Einsatzort:

    France
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!