Senior Systems Software Engineer, Observability and Telemetry Platform

Stellenbeschreibung:

  • Design, implement, and support operational and reliability aspects of a large-scale Observability & Telemetry collection platform, focusing on performance at scale, real-time monitoring, logging, and alerting
  • Engage in and improve the full lifecycle of services, from inception and design through deployment, operation, and refinement
  • Support services before launch through system design consulting, development of software tools, platforms, and frameworks, capacity management, and launch reviews
  • Maintain live services by measuring and monitoring availability, latency, and overall system health
  • Scale systems sustainably through automation and evolve systems by driving changes that improve reliability and velocity
  • Practice sustainable incident response and blameless postmortems
  • Participate in an on-call rotation to support production systems

Requirements

  • BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics), or equivalent experience
  • 5+ years of experience with Infrastructure automation, distributed systems design, and designing and developing tools for running large-scale private or public cloud systems in production
  • 5+ years of experience delivering foundational infrastructure and observability platforms
  • Experience with one or more of: Python, Go, Perl, or Ruby
  • In-depth knowledge of Linux, networking, and containers
  • Experience in using or running large private and public cloud systems based on Kubernetes, OpenStack, and Docker
  • Experience running Grafana, OpenTelemetry, Prometheus, and similar observability-focused tools

Core Competencies

Demonstrates expertise in designing and implementing large-scale Observability and Telemetry platforms, with a strong focus on performance, reliability, and automation. Proficient in managing cloud systems and utilizing observability tools to ensure system health and operational excellence.

Highest-signal resume keywords

  • Infrastructure Automation
  • Distributed Systems Design
  • Python Programming
  • Kubernetes Management
  • Observability Tools Experience

Hard Skills

  • Infrastructure Automation
  • Distributed Systems Design
  • Python
  • Go
  • Perl
  • Ruby
  • Linux
  • Networking
  • Containers
  • Cloud Systems

Soft Skills

  • Incident Response
  • Collaboration

Certifications & Qualifications

  • BS Degree in Computer Science

Industry Keywords

  • Observability
  • Telemetry
  • Performance Monitoring
  • System Health
  • Automation

Tools & Technologies

  • Kubernetes
  • OpenStack
  • Docker
  • Grafana
  • OpenTelemetry
  • Prometheus

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    16 Sep 2026
  • Standort:

    WorkFromHome
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!