Senior Site Reliability Engineer
Clera
Job Description
<h3>About the Role</h3><p style="min-height:1.5em">We are a well-funded AI/ML company operating at the intersection of geospatial intelligence and infrastructure analytics. Our engineering team is distributed across Europe and North America, and we're looking for a <strong>Senior Site Reliability Engineer</strong> to take ownership of our cloud infrastructure and elevate our DevOps and reliability practices.</p><p style="min-height:1.5em">In this role, you'll evolve our <strong>Google Cloud Platform (GCP)</strong> infrastructure, mature our observability platform, drive incident management processes, and partner closely with Product & Engineering teams to ship reliable, high-quality software. You'll be a key voice in championing SLOs, error budgets, and DORA metrics across the organisation.</p><h3>What You'll Do</h3><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Design, evolve, and scale our cloud infrastructure on GCP.</p></li><li><p style="min-height:1.5em">Build tooling and automation that promote team autonomy and reduce toil.</p></li><li><p style="min-height:1.5em">Advance our observability platform, improving mean time to recovery (MTTR) and system visibility.</p></li><li><p style="min-height:1.5em">Build transparency into infrastructure costs and drive cost optimisation initiatives.</p></li><li><p style="min-height:1.5em">Champion reliability best practices including SLOs/SLIs, error budgets, and post-incident reviews.</p></li><li><p style="min-height:1.5em">Lead on-call rotations and incident management, fostering a blameless culture.</p></li><li><p style="min-height:1.5em">Help engineering teams leverage GCP effectively and govern usage at scale.</p></li></ul><h3>What We're Looking For</h3><p style="min-height:1.5em"><strong>Required:</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">3+ years of Site Reliability Engineering or production SRE experience.</p></li><li><p style="min-height:1.5em">Strong proficiency with <strong>Google Cloud Platform (GCP)</strong>, including cost optimisation and governance.</p></li></ul><p style="min-height:1.5em"><strong>Nice to Have / Additional Skills:</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Hands-on experience with <strong>Kubernetes</strong> for cluster and workload management.</p></li><li><p style="min-height:1.5em">Infrastructure as Code experience — <strong>Terraform</strong>, Deployment Manager, or similar.</p></li><li><p style="min-height:1.5em">Scripting and automation skills in <strong>Python</strong>, <strong>Bash</strong>, or <strong>Go</strong>.</p></li><li><p style="min-height:1.5em">Strong observability stack experience: <strong>Prometheus</strong>, <strong>Grafana</strong>, <strong>OpenTelemetry</strong>, logging, and tracing.</p></li><li><p style="min-height:1.5em">Proven ability to define and implement SLOs/SLIs and error budgets.</p></li><li><p style="min-height:1.5em">Experience with incident management, post-inc
Skills