JobHunter AI
This position is no longer accepting applications. See similar jobs below.
Senior Site Reliability Engineer
name
Location
Remote
Work Mode
Remote
Type
Full-Time
Sector
Education
First Seen
2026-08-10
Source
himalayas
Remote Education IT MEAL Administration Deadline Unclear Remote
Job Description
<p>We are looking for a <strong>Senior Site Reliability Engineer</strong> with Cloud platform experience. This individual will be part of a team responsible for operating and maintaining production clusters and developing our observability solutions; they will collaborate with team members to develop automation strategies, monitoring &amp; alerting, and ensuring overall platform reliability. Your goal will be to become an integral part of the team, making every challenge of the platform – your own challenge, and solving them accordingly.</p><h3>Responsibilities</h3><ul> <li>Ensure platform reliability and availability across production and pre-production environments through proactive monitoring, alerting, and automation.</li> <li>First response for incidents, contribute to problem management and root cause analysis.</li> <li>Supporting the development team's effort towards reliability, creating a solid reliability culture within the development lifecycle.</li> <li>Develop troubleshooting documentation for production support resources.</li> <li>Collaborate with Engineering teams to develop optimised and productive runbooks, operational documentation and automation of operational tasks.</li> <li>Collaborate with development and cloud engineering teams to embed reliability and performance into the software delivery lifecycle.</li> <li>Design, implement, and evolve observability solutions (metrics, logs, traces, dashboards) using tools such as Prometheus, Grafana, and ELK.</li> <li>Participate in on-call rotations and continuously improve alert quality and response processes.</li> <li>Champion a culture of reliability, performance, and continuous improvement across teams.</li> </ul><h3>Requirements</h3><ul> <li>Bachelor's Degree or MS in Engineering or equivalent.</li> <li>Experience in operating at least one container orchestration cluster (Kubernetes, Docker Swarm).</li> <li>Experience developing or maintaining software for production services at scale.</li> <li>Experience with ELK.</li> <li>Experience with AWS.</li> <li>Experience with Grafana/Prometheus stack.</li> <li>Strong scripting skills (Bash, Python or Go).</li> <li>Excellent communication skills.</li> <li>Thinking out of the box and anticipating challenges. It is imperative we are not simply reactive; we must expect challenges and question technologies, procedures and thinking already in place. You will be expected to constantly review and challenge at all levels.</li> <li>Versatility. We work with agile/lean methods. We'd much rather iterate and learn than assume we know all the answers.</li> <li>Being a team player. You don't (always) work in isolation and are excited by the thought of using your team whilst involving product, experience design, engineering, and more in the process.</li> </ul><h3>Will be considered as a plus:</h3><ul> <li> Telephony knowledge (SIP, VoIP);</li> <li> Experience in Linux Administration (RedHat, CentOS, AL);</li> <li> Working knowledge in Configur