Senior Software Engineer – Applied AI

Stellenbeschreibung:

  • Design and run randomized experiments on tier routing, MCP coverage, permission configuration, repository context quality, and budget headroom
  • Own the analytical layer of the measurement program, including work classification, evaluation design, longitudinal within-unit analysis, and staggered-adoption estimates
  • Build and validate LLM-as-judge and classification pipelines with sampling, hand-labeled ground truth, precision and recall measurement, and revalidation
  • Extend AI capabilities across testing, environment and data setup, migration and modernization, code review, security remediation, and certification evidence assembly
  • Work directly with constrained teams to identify delivery bottlenecks and target capabilities accordingly
  • Build evaluations for internal AI capabilities, including golden sets, regression suites, groundedness, answer-quality scoring, cost, and latency telemetry
  • Identify effective practitioner behaviors, document and teach practices, and publish practices rather than rankings
  • Partner with the platform team on Claude Code configuration, MCP servers, gateway telemetry, and the model registry
  • Report findings to engineering leadership and finance, including what is working, delivery constraints, and the limits of each claim

Requirements

  • 7+ years spanning software engineering and quantitative analysis
  • Production experience with LLM applications: prompting, tool and function calling, context management, evaluation, and knowing where models fail in practice
  • Experimental design and causal inference, including randomized and quasi-experimental designs, difference-in-differences, instrumental variables, and hierarchical models
  • Strong Python and SQL, with a statistical stack such as pandas, statsmodels, scikit-learn, or R
  • Data collection from operational systems and APIs; robust sampling design
  • Identity resolution across systems and complex data joins
  • Exploratory analysis, distributions, cohort and time-series analysis, and reporting of coverage and limitations
  • Familiarity with the software delivery lifecycle, including code review, CI/CD, test strategy, release and change management
  • Git and GitLab at instrumentation depth, including merge request and pipeline data models, diffs and SHAs, merge, squash, rebase, cherry-pick, APIs, and hooks
  • Jira and Confluence integration experience, including REST APIs, changelogs, page version history, GitLab–Jira development panel, fields, and labels
  • Care with personnel-adjacent data and aggregate reporting
  • Communication with executive and engineering audiences
  • Preferred MCP servers and clients or comparable connector frameworks
  • Agent frameworks such as LangGraph, LangChain, Bedrock Agents, Strands, or equivalent
  • Enterprise deployment of coding assistants and telemetry
  • Server-side Git hooks, GitLab CI, and system or webhook-driven capture on a self-managed instance
  • Confluence and Jira as MCP-connected systems, including permission propagation, scoped credentials, and audit logging
  • Evaluation tooling such as Ragas, DeepEval, or Bedrock model evaluation; LLM observability such as LangFuse, Arize, or OpenTelemetry-based tracing
  • Amazon Bedrock and AWS cost and usage data
  • Engineering productivity frameworks such as DORA, DX Core 4, or SPACE
  • Program analysis, test generation, or developer tooling research
  • dbt, Airflow, Dagster, or equivalent transformation and orchestration; warehouse or lakehouse modeling
  • BI and visualization tooling; summary-table-based reporting
  • Queueing and flow analysis, including utilization, batch economics, and constraint identification

Core Competencies

Demonstrates expertise in experimental design and causal inference, with a strong focus on LLM applications and quantitative analysis. Proficient in Python and SQL, with experience in data collection, exploratory analysis, and software delivery lifecycle.

Highest-signal resume keywords

  • LLM Application Development
  • Experimental Design and Causal Inference
  • Python and SQL Proficiency
  • Data Collection and Sampling Design
  • Git and GitLab Expertise

Hard Skills

  • Experimental Design
  • Causal Inference
  • Python
  • SQL
  • Data Collection
  • Exploratory Analysis
  • LLM Applications
  • Statistical Analysis
  • Sampling Design
  • Evaluation Tooling

Soft Skills

  • Communication with Executive Audiences
  • Documentation and Teaching Practices

Industry Keywords

  • Software Engineering
  • Quantitative Analysis
  • AI Capabilities
  • Software Delivery Lifecycle
  • Engineering Productivity Frameworks

Tools & Technologies

  • Git
  • GitLab
  • Jira
  • Confluence
  • Amazon Bedrock
  • AWS
  • Ragas
  • DeepEval
  • LangChain
  • Airflow

#J-18808-Ljbffr
NOTE / HINWEIS:
EnglishEN: Please refer to Fuchsjobs for the source of your application
DeutschDE: Bitte erwähne Fuchsjobs, als Quelle Deiner Bewerbung

Stelleninformationen

  • Veröffentlichungsdatum:

    16 Sep 2026
  • Standort:

    WorkFromHome
  • Typ:

    Vollzeit
  • Arbeitsmodell:

    Vor Ort
  • Kategorie:

  • Erfahrung:

    2+ years
  • Arbeitsverhältnis:

    Angestellt

KI Suchagent

AI job search

Möchtest über ähnliche Jobs informiert werden? Dann beauftrage jetzt den Fuchsjobs KI Suchagenten!