Software Platform Support Engineer – GPU Cloud

Stellenbeschreibung:

  • Coordinate with multiple internal teams to provide Tier 1 support for complex cloud platforms
  • Define and improve operational workflows, including runbooks, escalation paths, and support processes
  • Triage and investigate root causes of customer issues and elevate as needed
  • File bugs and report issues while working closely with the Site Reliability team
  • Build tooling to improve the customer support process and visibility
  • Develop a deep understanding of user workloads and use cases
  • Partner with multiple internal teams to provide feedback to engineering teams and develop solutions
  • Participate in an on-call rotation to support production systems

Requirements

  • BS/MS degree in Computer science or related areas (or equivalent experience)
  • 5+ yrs of experience with supporting distributed software systems, supporting end-user software platforms, and experience with Linux
  • Experience with Kubernetes, AWS, Azure, OCI, and GCP
  • Background of Infrastructure, Networking, Storage, and DevOps scripting/tooling
  • Understanding of data storage technologies (databases, file, block, blob)
  • Customer Service/Support Experience
  • Willingness to work up and down the stack as well as across multiple teams
  • Strong skills in troubleshooting and Communication
  • Experience with MLOps workflows or ML infrastructure
  • Familiarity with GPU workloads or distributed training systems
  • SLURM or HPC previous experience
  • Strong drive to work with internal customers and make them successful
  • A drive to improve process with strong organizational skills

Core Competencies

Demonstrates expertise in supporting distributed software systems and cloud platforms, with a strong focus on troubleshooting, customer service, and operational workflow improvement. Proficient in collaborating across teams to enhance user experience and support processes.

Highest-signal resume keywords

  • 5+ Years Experience Supporting Distributed Software Systems
  • Experience With Kubernetes, AWS, Azure, OCI, And GCP
  • Strong Skills In Troubleshooting And Communication
  • Customer Service/Support Experience
  • Understanding Of Data Storage Technologies

Hard Skills

  • Linux
  • DevOps Scripting/Tooling
  • MLOps Workflows
  • GPU Workloads
  • SLURM
  • HPC
  • Operational Workflow Improvement
  • Triage And Root Cause Analysis
  • Bug Filing And Issue Reporting
  • Building Tooling For Support Processes

Soft Skills

  • Strong Organizational Skills
  • Willingness To Collaborate Across Teams
  • Drive To Improve Processes

Certifications & Qualifications

  • BS/MS Degree In Computer Science Or Related Areas

Industry Keywords

  • Cloud Platforms
  • Customer Support
  • Infrastructure
  • Networking
  • Storage

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    16 Sep 2026
  • Standort:

    WorkFromHome
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!