Senior GPU Cloud Storage Solutions Expert – SRE SME

Stellenbeschreibung:

  • Deploy and operate parallel/distributed storage systems including WEKA, VAST Data, Ceph, and DDN/Lustre
  • Design storage architectures optimized for AI workload patterns such as checkpoint I/O bursts, sequential dataset reads, and KV cache for inference
  • Implement multi-tenant storage isolation with per-tenant QoS, quotas, and access controls
  • Configure and optimize GPU Direct Storage for direct GPU-to-storage data paths
  • Deploy and manage storage networking including NFS over RDMA, NVMe-oF, high-speed storage fabrics, and Nvidia CMX for cluster-wide storage orchestration
  • Diagnose and tune storage performance using IOPS, throughput, latency profiling, fio, IOR, and mdtest
  • Own the runbook for common failure modes
  • Plan storage capacity aligned with GPU cluster growth and customer workload projections
  • Manage firmware, data migration, and disaster recovery procedures
  • Instrument storage telemetry including IO tail latency, checkpoint durations, NVMe SMART, filesystem health, and RDMA counters
  • Feed telemetry into the platform team's metrics, logs, and traces store
  • Partner with the platform team to define the storage-fault predictor, including signals, labels, and false-positive tolerances
  • Convert novel incidents into automation, progressing from SOPs to runbook-as-code and agent-executable remediation
  • Deliver observability and a baseline predictor for the top three storage-fault classes
  • Reduce storage-incident MTTR
  • Design storage for Nvidia GB200-class clusters

Requirements

  • 5+ years in enterprise or HPC storage operations, with at least 2 years supporting AI/ML workloads
  • Hands-on deployment and operations experience with at least two of: WEKA, VAST Data, Ceph, DDN/Lustre
  • Strong understanding of AI training I/O patterns: checkpoint frequency, dataset loading, shuffle buffers
  • Experience with high-performance storage networking (NFS over RDMA, NVMe-oF)
  • Knowledge of GPU Direct Storage and RDMA-based data transfer
  • Proficiency in storage performance benchmarking and tuning (fio, IOR, mdtest)
  • Experience implementing multi-tenant storage with isolation and QoS
  • Strong Linux systems knowledge (kernel tuning, filesystem internals, block device management)
  • Experience shipping an anomaly detector for storage/IO telemetry or ability to articulate the labels and features needed
  • Runbook-as-code mindset, with every SOP executable by a machine within a quarter

Core Competencies

Demonstrates expertise in deploying and managing parallel and distributed storage systems optimized for AI workloads, with a strong focus on performance tuning, multi-tenant isolation, and telemetry instrumentation. Proficient in high-performance storage networking and capable of implementing automation for storage operations.

Highest-signal resume keywords

  • WEKA Deployment
  • Ceph Operations
  • GPU Direct Storage
  • Storage Performance Benchmarking
  • Multi-Tenant Storage Isolation

ATS Optimization Keywords

Hard Skills

  • Storage Architecture Design
  • AI Workload Optimization
  • Storage Performance Tuning
  • Linux Systems Knowledge
  • Anomaly Detection for Telemetry

Soft Skills

  • Problem-Solving
  • Collaboration

Industry Keywords

  • Enterprise Storage Operations
  • HPC Storage
  • AI/ML Workloads
  • Telemetry Instrumentation
  • Disaster Recovery Procedures

Tools & Technologies

  • NFS over RDMA
  • NVMe-oF
  • Fio
  • IOR
  • Mdtest

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    16 Sep 2026
  • Standort:

    WorkFromHome
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!