
// Open Role at Sporty Group
Join Sporty Group as a Site Reliability Engineer to enhance cloud infrastructure, optimize Kubernetes platforms with GitOps, and ensure robust, high-availability systems for global deployments. Leverage your SRE and DevOps expertise in a remote-first environment.
What you’ll be doing
What you’ll bring
Our stack
What’s in it for you
Interview process
If you’re interested, we encourage you to apply! Every application is reviewed by a member of our team (AI is not used in our recruitment process), and we aim to respond within 48 hours.
- 3+ years DevOps / SRE / platform engineering experience - Must be based in Europe - Experience independently leading the planning and deployment of a project - Experienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and production - Strong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and Helm being highly valued - Experience with Infrastructure-as-Code, particularly Terraform - Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus - Hands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and OpenTelemetry - Experience with real user monitoring (RUM), with familiarity in Grafana Faro or OpenTelemetry SDK instrumentation being a plus - Proven on-call and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actions - Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patterns - Experience defining SLIs and SLOs and using them to inform reliability work - Familiarity with service mesh concepts is a plus, as we are actively evaluating Cilium-based service mesh in non-production environments - Solid networking knowledge, especially the TCP / IP stack and HTTP protocol - Experience handling high HTTP request volumes and designing systems for high availability and high traffic environments - A strong understanding of cache, including CDN, HTTP cache, Redis / Memcached - Excellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous