Customer Site Reliability Engineer - OpenShift Managed Cloud Services (Spoken Japanese, Kuberne[...]

Stellenbeschreibung:

Red Hat is looking for a Customer Site Reliability Engineer (CSRE) to join our OpenShift Managed Cloud Services (MCS) team. The CSRE plays a crucial role in ensuring the availability, reliability, and performance of critical services at scale, independently managing complex systems and solving intricate problems that significantly impact service quality and stability.

A CSRE has a customer‑first mindset and will act as a technical lead for customer escalations, applying expert troubleshooting to ensure timely and effective resolutions that maintain trust and confidence. Leveraging extensive experience in software and systems engineering, the CSRE automates operations, reduces toil, and drives continuous improvement across the service lifecycle. The role is autonomous, demonstrating strong judgment and decision‑making while managing non‑routine assignments.

Collaboration is essential; you will partner with Technical Account Managers, Services, Fleet SRE, DevOps, and infrastructure teams to address customer‑specific and fleet‑wide issues, ensuring the stability and functionality of our cloud‑based systems.

As a champion of Knowledge‑Centered Support (KCS), you will document resolutions, root causes, and best practices to enrich the knowledge base and promote self‑service solutions, while mentoring team members to foster a collaborative and continuous‑learning culture.

This role is ideal for a highly skilled and motivated individual who thrives in a fast‑paced, collaborative environment and is passionate about driving reliability, scalability, and customer satisfaction.

What you will do

  • Manage large‑scale, distributed systems, focusing on minimizing downtime and improving system resilience.
  • Maintain customer trust and confidence by ensuring stability and functionality of services.
  • Drive continuous enhancement of processes, tools, and methodologies to support the evolving needs of the service.
  • Lead the development of code and automation scripts to optimize the scalability, reliability, and performance of services.
  • Lead and participate in high‑priority customer escalations, adopting a customer‑first mindset.
  • Coordinate and execute complex incident response procedures, ensuring timely resolution and thorough post‑mortems.
  • Collaborate with cross‑functional teams to enhance system robustness.
  • Demonstrate a proactive mindset to help preempt escalations and ensure reliable operations.
  • Document resolutions, root causes, and best practices to enrich the knowledge base and promote self‑service solutions.
  • Mentor and coach team members, fostering a culture of continuous learning, knowledge sharing and collaboration.
  • Participate in on‑call rotation and provide leadership during critical incidents.
  • Collaborate on strategic AI and automation projects designed to increase the efficiency of fleet operations and troubleshooting, ultimately delivering a better product experience for customers.
  • Given the customer‑facing nature of this SRE role, exceptional communication skills are essential. You must demonstrate the ability to articulate complex technical solutions and lead critical incident calls with confidence, even in high‑pressure environments.

What you will bring

  • Advanced experience with OpenShift/Kubernetes container platform support or administration.
  • Proficiency with container‑based technologies on Linux.
  • Experience managing Linux‑based systems in a public cloud such as AWS, Azure, or GCP.
  • Advanced experience with enterprise systems monitoring; knowledge of Prometheus is preferred.
  • Advanced experience with enterprise configuration management such as Ansible, Terraform.
  • Software engineering experience using object‑oriented languages; GoLang is preferred.
  • Superior communication skills and experience working directly with and presenting to customers.
  • Ability to quickly learn new technologies and follow industry trends.
  • Demonstrated ability to quickly and accurately troubleshoot systems issues.
  • Solid understanding of standard TCP/IP networking and common protocols.
  • Fluent in English, and additional languages such as Japanese, Chinese, Korean, or Spanish are an advantage.

Equal Opportunity Policy (EEO)

Red Hat is proud to be an equal opportunity workplace and an affirmative action employer. We review applications for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, citizenship, age, veteran status, genetic information, physical or mental disability, medical condition, marital status, or any other basis prohibited by law.

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    24 Jul 2026
  • Standort:

    Remote

    Einsatzort:

    Texas and California
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!