The DEKRA Arbeit Group, with its more than 140 locations throughout Europe, is one of the most successful and innovative umbrella organizations in the area of personnel services. We are a highly professional and ambitious company driven by talented and dedicated people who are committed in delivering all our services.
For our Client, we are looking for a:
Senior Site Reliability Engineer (m/f)
Hybrid
Location: Zagreb or Split
This role focuses on Linux-based systems, Kubernetes platforms, CI/CD pipelines, cloud infrastructure, and observability stacks supporting distributed applications at scale. It is intended for someone who has worked in a product company and supported large-scale, user-facing production systems.
You will work closely with engineering teams to improve system reliability, scalability, deployment workflows, and developer experience through robust infrastructure, automation, and internal platform tooling.
Key Responsibilities
- Design, build, and operate cloud-native infrastructure platforms
- Deploy and maintain production Kubernetes clusters
- Operate and troubleshoot large-scale Linux environments
- Design and maintain CI/CD pipelines and automated deployment workflows
- Implement GitOps-based deployment models
- Design and implement observability and monitoring systems
- Build automation for infrastructure provisioning and operations
- Improve reliability, scalability, and performance of distributed systems
- Investigate and resolve production incidents and performance issues
- Work closely with developers to improve deployment workflows and system reliability
- Contribute to infrastructure tooling and internal platform development
Candidate Requirements
- Strong Linux systems expertise
- Deep understanding of Kubernetes architecture and operations
- Experience operating production container platforms
- Experience supporting scalable systems with high availability requirements
- Experience working on a product/platform with a large user base or high traffic
- Experience with observability tools such as Prometheus, Grafana, OpenTelemetry, logging systems, or similar
- Experience building and maintaining CI/CD pipelines
- Experience with GitHub Actions or similar CI/CD tools
- Experience with GitOps workflows and tools such as ArgoCD, Flux, or similar
- Experience operating systems in cloud environments such as AWS, GCP, Azure, or similar
- Experience with Infrastructure as Code, preferably Terraform or similar
- Ability to read and write code used for automation, infrastructure, and tooling
- Programming and Automation (preferably using Go, Python, or TypeScript/JavaScript)
Bonus Points for
- Experience building or improving internal developer platforms
- Experience with service mesh or advanced Kubernetes networking
- Experience running multi-tenant infrastructure
- Experience with infrastructure security and hardening
- Experience operating high-scale container workloads
- Experience defining or improving SLOs, SLIs, incident response processes, or reliability metrics
Our Client Offers
- Opportunity to work on a large-scale product with real production complexity
- Work with modern cloud-native technologies and engineering practices
- High level of ownership and technical impact
- Competitive salary aligned with experience and seniority
- Employee stock options
- Bonus scheme
- Multisport card