We’re looking for a Senior AI Reliability Engineer (Platform) to help us accomplish our mission to improve and extend lives by learning from the experience of every person with cancer. Are you ready to be the next changemaker in cancer care?
A model may not crash, but it may silently degrade, become less accurate, respond inconsistently, produce poor outputs, or create business risk in ways that are hard to detect without the right evaluation and observability patterns.
You are comfortable operating in ambiguous spaces where the right answer is not always obvious, and you are motivated by turning emerging AI capabilities into production-ready systems that teams can actually trust. You understand that AI systems fail differently from traditional software.
You’re a senior technical practitioner with experience working across data science, machine learning, software engineering, platform engineering, or reliability engineering. This role is focused on that production behaviour and system health, not on pure model research or training.
Experience with modern cloud and ML infrastructure, including AWS, containers, Kubernetes, CI/CD, data pipelines, workflow orchestration, versioning, and distributed compute platforms. Strong understanding of AI reliability and observability, including logging, tracing, monitoring, drift detection, statistical analysis, uncertainty, alerting, and production system health.
5+ years of experience in platform engineering, SRE, machine learning, MLOps or a related technical field, with strong Python skills and experience building production-quality systems.
Experience designing experiments, evaluation frameworks, statistical analyses, and quality metrics for ML or AI systems, with familiarity in LLMs, RAG, AI agents, prompt evaluation, and model behaviour. Fluent in English. Strong communication and collaboration skills, with the ability to explain AI behaviour and tradeoffs to technical and non-technical stakeholders and thrive in a fast-moving, ambiguous environment with a pragmatic, enablement-focused mindset.
Knowledge of agentic and multi-agent systems, including orchestration, state management, tool execution, governance, reliability, human-in-the-loop controls, and selecting the appropriate level of AI autonomy for a given problem.
Experience with LLM evaluation, red-teaming, adversarial testing, AI safety, RAG evaluation, retrieval quality measurement, embedding drift, or AI observability and model monitoring.
Hands-on experience with observability and data/ML platforms such as Datadog, Splunk, OpenTelemetry, Databricks, Spark, Airflow, dbt, Ray, SageMaker, GitLab CI/CD, or similar technologies. Experience working in healthcare, life sciences, or other regulated, privacy-sensitive environments.
#J-18808-LjbffrVeröffentlichungsdatum:
24 Jul 2026Standort:
BerlinTyp:
VollzeitArbeitsmodell:
Vor OrtKategorie:
Erfahrung:
2+ yearsArbeitsverhältnis:
Angestellt
Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!