Site Reliability Engineer - Contract - Remote - Outside IR35 at SixteenFifty, London, 6 Months, £450 per day

£450 per day

Contract Description

My client are looking for a skilled Contract Site Reliability Engineer (SRE) to help improve the reliability, availability, and operational performance of a cloud platform. You'll work closely with engineering and operations teams to ensure critical services remain highly available while driving automation and operational excellence.

Responsibilities of the Site Reliability Engineer

  • Maintain and improve platform reliability, availability, and performance.
  • Monitor production systems and proactively identify potential issues.
  • Respond to incidents, perform root cause analysis, and implement permanent fixes.
  • Automate operational tasks to reduce manual effort and improve consistency.
  • Develop and enhance observability using metrics, logging, and tracing.
  • Improve deployment processes and support CI/CD pipelines.
  • Work with development teams to enhance application resilience and operational readiness.
  • Contribute to disaster recovery, capacity planning, and scalability initiatives.
  • Produce and maintain operational documentation and runbooks.

Required Skills of the Site Reliability Engineer

  • Commercial experience as a Site Reliability Engineer, DevOps Engineer, or Platform Engineer.
  • Strong Linux systems administration skills.
  • Experience with public cloud platforms (AWS, Azure, or Google Cloud).
  • Hands-on experience with Kubernetes and container technologies.
  • Strong understanding of Infrastructure as Code (Terraform or similar).
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, or CloudWatch.
  • Experience supporting CI/CD pipelines and automation.
  • Scripting skills using Bash, Python, or Go.
  • Good understanding of networking, security, and distributed systems.
  • Excellent troubleshooting and communication skills.

Desirable

  • Experience operating production Kubernetes clusters.
  • Knowledge of service level indicators (SLIs), service level objectives (SLOs), and error budgets.
  • Experience with incident management and post-incident reviews.
  • Familiarity with GitOps practices and modern deployment strategies.

If you're an experienced SRE looking for your next contract opportunity and enjoy building reliable, scalable cloud platforms, we'd love to hear from you.