Own the most complex, highest-stakes customer engagements from architecture through production across multiple modalities, driving measurable business value. Optimize LLM inference at the framework and hardware level and codify the resulting best practices into reusable playbooks. Lead supervised and reinforcement fine-tuning efforts to maximize model quality. Design and implement production-ready LLM solutions using Token Factory's inference services. Provide deep technical expertise in prompt engineering, RAG architectures, model selection, and cost/performance trade-offs at scale. Partner closely with product, engineering and research to surface customer needs, prototype platform features, and directly influence the roadmap. Guide customers from PoC to production with a focus on performance, reliability, and cost efficiency — and define the standards by which the team does so. Mentor Senior and mid-level Solutions Architects; raise the technical bar of the team through review, enablement, and knowledge sharing. Represent Token Factory externally through talks, blog posts, and conferences.
Nice-to-have: Contributions to OSS inference/ML projects, published research, multimodal AI, DevOps tooling, internal tooling for ML workflows. Preferred tech stack includes Python, vLLM, TensorRT-LLM, SGLang, Transformers, OpenAI/Anthropic SDKs, Kubernetes, Docker, cloud platforms.
Key Employee Benefits and Pay Transparency outlined with salary: 208k-261k USD base.
Key Employee Benefits and Pay Transparency outlined with salary: 208k-261k USD base.
#J-18808-LjbffrVeröffentlichungsdatum:
02 Sep 2026Standort:
WorkFromHomeTyp:
VollzeitArbeitsmodell:
Vor OrtKategorie:
Development & ITErfahrung:
2+ yearsArbeitsverhältnis:
Angestellt
Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!