£450 per day
Outside Spy
Remote (London, UK)
My client are looking for a skilled Contract Site Reliability Engineer (SRE) to help improve the reliability, availability, and operational performance of a cloud platform. You'll work closely with engineering and operations teams to ensure critical services remain highly available while driving automation and operational excellence. Responsibilities of the Site Reliability Engineer Maintain and improve platform reliability, availability, and performance. Monitor production systems and proactively identify potential issues. Respond to incidents, perform root cause analysis, and implement permanent fixes. Automate operational tasks to reduce manual effort and improve consistency. Develop and enhance observability using metrics, logging, and tracing. Improve deployment processes and support CI/CD pipelines. Work with development teams to enhance application resilience and operational readiness. Contribute to disaster recovery, capacity planning, and scalability...