Leave us your email address and we'll send you all the new jobs according to your preferences.

Site Reliability Engineer (SRE)

Posted 41 minutes 26 seconds ago by Spencer Rose Ltd

Contract
Not Specified
Other
Lancashire, Manchester, United Kingdom, M21 0
Job Description

Spencer Rose is urgently seeking a number of Site Reliability Engineer's (SREs) to work on an exciting transformation programme for our Tier-1 banking client located in Manchester.

Start: ASAP
Duration: 6-12 month contract (strong possibility for extensions)
Rate: £400-420 per day Inside IR35
Location: Manchester - hybrid 2 days a week in the office

ROLE PURPOSE Improve the reliability, scalability, security and operational excellence of cloud-hosted services on Google Cloud Platform. This is a hands-on engineering role spanning SRE practices, production Kubernetes, infrastructure automation, observability, CI/CD, incident response and continuous service improvement.

About the role

The Senior Site Reliability Engineer will work with Cloud Platform, Software Engineering, Product, Security and Service teams to design and operate resilient services. The role will use software engineering and automation to reduce operational toil, improve service health and turn incident learning into measurable reliability improvements.

Key responsibilities

  • Lead reliability, availability, scalability and performance improvements across GCP-hosted applications, platforms and shared services.
  • Define and operate service level indicators, service level objectives and error-budget practices that connect technical health to customer impact.
  • Design actionable observability using Dynatrace, including instrumentation, dashboards, distributed tracing, service health views and SLO-based alerting.
  • Build modular, reusable and maintainable Terraform code for secure cloud infrastructure, platform services and environment provisioning.
  • Administer production Kubernetes environments, covering cluster life cycle, workload deployment, capacity, upgrades, networking and platform troubleshooting.
  • Automate operational activities and repetitive support work using Python, Groovy, Bash or PowerShell to reduce toil and improve consistency.
  • Build and enhance CI/CD pipelines using Jenkins, Azure DevOps, GitHub Actions or equivalent tooling, with automated quality and deployment controls.
  • Lead incident response, complex troubleshooting, root-cause analysis and post-incident reviews; ensure corrective actions are tracked to completion.
  • Embed security, resilience, monitoring and supportability into platform and application designs from the outset.
  • Coach engineers, contribute reusable standards and patterns, and promote SRE and operational excellence across engineering communities.

Interviews for these roles will take place w/c 5th October, if you are available and have the above experience APPLY NOW

Email this Job