Principal ML Solutions Architect - Token Factory

Stellenbeschreibung:

Your responsibilities will include:

Own the most complex, highest-stakes customer engagements from architecture through production across multiple modalities, driving measurable business value. Optimize LLM inference at the framework and hardware level and codify the resulting best practices into reusable playbooks. Lead supervised and reinforcement fine-tuning efforts to maximize model quality. Design and implement production-ready LLM solutions using Token Factory's inference services. Provide deep technical expertise in prompt engineering, RAG architectures, model selection, and cost/performance trade-offs at scale. Partner closely with product, engineering and research to surface customer needs, prototype platform features, and directly influence the roadmap. Guide customers from PoC to production with a focus on performance, reliability, and cost efficiency — and define the standards by which the team does so. Mentor Senior and mid-level Solutions Architects; raise the technical bar of the team through review, enablement, and knowledge sharing. Represent Token Factory externally through talks, blog posts, and conferences.

We expect you to have:

  • 8+ years of experience in ML/AI systems, with at least 4 years focused on LLMs and generative AI.
  • Demonstrated technical leadership: owning ambiguous, high-impact problems end to end and influencing decisions across teams and customers.
  • Expert knowledge of the LLM ecosystem: model architectures, fine-tuning approaches, and inference internals.
  • Deep, hands-on command of inference optimization: quantization, KV-cache management, batching, routing, etc.
  • Hands‑on experience with: Running LLMs in production at scale, LLM fine‑tuning including SFT/LoRA, LLM evaluation and deployment of LLM‑powered applications.
  • Strong Python programming skills and excellent communication.

Nice-to-have: Contributions to OSS inference/ML projects, published research, multimodal AI, DevOps tooling, internal tooling for ML workflows. Preferred tech stack includes Python, vLLM, TensorRT-LLM, SGLang, Transformers, OpenAI/Anthropic SDKs, Kubernetes, Docker, cloud platforms.

Key Employee Benefits and Pay Transparency outlined with salary: 208k-261k USD base.

Additional details:

  • 8+ years of experience in ML/AI systems, with at least 4 years focused on LLMs and generative AI.
  • Demonstrated technical leadership: owning ambiguous, high-impact problems end to end and influencing decisions across teams and customers.
  • Expert knowledge of the LLM ecosystem: model architectures, fine-tuning approaches, and inference internals.
  • Deep, hands‑on command of inference optimization: quantization, KV-cache management, batching, routing, etc.
  • Hands‑on experience with: Running LLMs in production at scale, LLM fine‑tuning including SFT/LoRA, LLM evaluation and deployment of LLM‑powered applications.
  • Strong Python programming skills and excellent communication.

Key Employee Benefits and Pay Transparency outlined with salary: 208k-261k USD base.

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    02 Sep 2026
  • Standort:

    WorkFromHome
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

    Development & IT
  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!