Senior Site Reliability Engineer
name
Job Description
<p>We are looking for a <strong>Senior Site Reliability Engineer</strong> with Cloud platform experience. This individual will be part of a team responsible for operating and maintaining production clusters and developing our observability solutions; they will collaborate with team members to develop automation strategies, monitoring & alerting, and ensuring overall platform reliability. Your goal will be to become an integral part of the team, making every challenge of the platform – your own challenge, and solving them accordingly.</p><h3>Responsibilities</h3><ul>
<li>Ensure platform reliability and availability across production and pre-production environments through proactive monitoring, alerting, and automation.</li>
<li>First response for incidents, contribute to problem management and root cause analysis.</li>
<li>Supporting the development team's effort towards reliability, creating a solid reliability culture within the development lifecycle.</li>
<li>Develop troubleshooting documentation for production support resources.</li>
<li>Collaborate with Engineering teams to develop optimised and productive runbooks, operational documentation and automation of operational tasks.</li>
<li>Collaborate with development and cloud engineering teams to embed reliability and performance into the software delivery lifecycle.</li>
<li>Design, implement, and evolve observability solutions (metrics, logs, traces, dashboards) using tools such as Prometheus, Grafana, and ELK.</li>
<li>Participate in on-call rotations and continuously improve alert quality and response processes.</li>
<li>Champion a culture of reliability, performance, and continuous improvement across teams.</li>
</ul><h3>Requirements</h3><ul>
<li>Bachelor's Degree or MS in Engineering or equivalent.</li>
<li>Experience in operating at least one container orchestration cluster (Kubernetes, Docker Swarm).</li>
<li>Experience developing or maintaining software for production services at scale.</li>
<li>Experience with ELK.</li>
<li>Experience with AWS.</li>
<li>Experience with Grafana/Prometheus stack.</li>
<li>Strong scripting skills (Bash, Python or Go).</li>
<li>Excellent communication skills.</li>
<li>Thinking out of the box and anticipating challenges. It is imperative we are not simply reactive; we must expect challenges and question technologies, procedures and thinking already in place. You will be expected to constantly review and challenge at all levels.</li>
<li>Versatility. We work with agile/lean methods. We'd much rather iterate and learn than assume we know all the answers.</li>
<li>Being a team player. You don't (always) work in isolation and are excited by the thought of using your team whilst involving product, experience design, engineering, and more in the process.</li>
</ul><h3>Will be considered as a plus:</h3><ul>
<li> Telephony knowledge (SIP, VoIP);</li>
<li> Experience in Linux Administration (RedHat, CentOS, AL);</li>
<li> Working knowledge in Configur