We are seeking an experienced Senior Infrastructure/Systems Engineer to help operate, maintain, and improve our on-premise infrastructure. You will work hands-on with networking equipment, Linux servers, Kubernetes, Redis.Ensuring that our platforms are secure, scalable, and highly available.
In this role, you will be responsible for the availability, scalability, and security of our systems and networks, ensuring smooth day-to-day operations, reliable deployments, and secure access for our teams.
This is a hands-on technical role, ideal for a technically driven engineer who thrives in complex, production-grade infrastructure environments and enjoys keeping systems stable, secure, and performant.
This position offers a hybrid work mode, combining on-site collaboration with flexible remote work.
Key Responsibilities
- Administer and maintain network infrastructure, including switches, routers, VPNs, and firewalls, ensuring reliable and secure connectivity across the datacenter.
- Manage routing, switching, and secure remote access for internal teams and services.
- Perform datacenter maintenance, including hardware installation, cabling, capacity planning, and coordination with datacenter and connectivity providers.
- Operate, maintain, and harden Linux-based servers in our on-premise datacenter.
- Manage server lifecycle activities, including provisioning, patching, upgrades, and decommissioning.
- Manage and enhance Kubernetes clusters, including configuration, upgrades, scaling, and deployment automation using Helm and Docker.
- Manage, monitor, and troubleshoot RabbitMQ clusters, ensuring message delivery reliability, scalability, and fault tolerance.
- Administer and optimize Redis, Elasticsearch, and MariaDB/MySQL for performance, stability, and data integrity.
- Support and execute database migration and infrastructure modernization projects.
- Implement and maintain infrastructure-as-code practices using Terraform, Ansible, GitLab CI/CD, and Puppet.
- Maintain and improve monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, Zabbix).
- Contribute to incident response, root cause analysis, and on-call rotations, and help establish incident response strategies and procedures.
- Collaborate with development and product teams to support deployment pipelines and performance optimization.
- Ensure compliance with internal security and operational policies (e.g., GDPR, data protection).
- Evaluate and maintain third-party tools, hardware, and services.
- 5+ years of hands-on experience in network, system, or infrastructure engineering, focused on on-premise environments.
- Strong networking skills (TCP/IP, routing, switching, firewalls, VPNs, and secure access).
- Strong expertise in Linux server administration (Debian/Ubuntu preferred.
- Hands-on experience with datacenter operations and hardware maintenance.
- Solid practical knowledge of Kubernetes, Docker, and Helm, including cluster management, deployments, upgrades, and troubleshooting.
- Deep understanding of RabbitMQ, including clustering, high availability, performance tuning, and troubleshooting.
- Experience with infrastructure automation and configuration management using Terraform, Ansible, GitLab CI/CD, and Puppet.
- Experience with monitoring and observability tools (Prometheus, Grafana, Zabbix).
- Excellent problem-solving and analytical skills, with attention to performance, reliability, and maintainability.
Bonus Points / Nice to have
- Experience managing RabbitMQ, Redis, Elasticsearch, and MySQL/MariaDB at scale in production.
- Experience with high-availability and distributed on-prem systems.
- Familiarity with Rancher for Kubernetes cluster management, and Foreman for provisioning.
- Knowledge of service meshes (e.g., Istio).
- Familiarity with container security and system hardening best practices.
- Knowledge of disaster recovery and business continuity planning for on-prem environments.
Operating Systems
- Ubuntu
- MariaDB
- Elasticsearch
- Redis
Web & Proxy Services
Orchestration & Automation
- Terraform
- Ansible, AWX
- Rancher
- Puppet
Monitoring & Observability
What We Offer
- Everything you need to do a great job (MacBook, etc.).
- Free weekly German classes to help you settle into Berlin life.
- Wellpass gym access and many other perks.
- Flexible hours with hybrid working between our great offices and home.
- A friendly, diverse group of colleagues from all different nationalities, genders, and orientations!
#J-18808-Ljbffr