Staff Site Reliability Engineer
Posted 1 day 3 hours ago by United States Digital Space LLC
The company is building data pipelines to power the modern data stack for thousands of companies. We're looking for a high-performance, experienced engineer to be a part of a team of Site Reliability Engineers. You will be working closely with engineering teams, product managers, as well as support, sales engineers and solution architects to build the future of the company Data Platform Reliability.
This is a full-time position based out of our Dublin office. Our hybrid work model offers a blend of remote flexibility and in-person collaboration, including two days in the office each week to connect and build as a team.
Technologies You'll Use- Cloud Service Providers (CSPs): AWS, Azure and Google Cloud
- Kubernetes: Managed Kubernetes services (EKS, AKS and GKE)
- Continuous Integration tools: GitHub Actions, Buildkite
- Continuous Delivery: ArgoCD
- Databases: Postgres and all the major databases
- Languages: Go, Java
- Scripts: Typescript, Python, Shell
- IaC: Terraform and Pulumi
- RESTful API: FastAPI
- Cloud networking: Privatelinks in Azure and AWS, Private Service Connect in GCP and site to site VPN tunnels in all 3 major cloud service providers
As a member of the Site Reliability Engineering team, you will take ownership over the overall performance and reliability of the company's infrastructure, the robustness of the deployment pipeline, as well as timely and effective incident response and resolution. You will take responsibility for the growth and stability of the company's infrastructure, and be a key player driving effective incident response and overall issue avoidance.
- Responsible for ongoing reliability and robustness of the company's production infrastructure by monitoring availability, capacity, and throughput.
- Evolve systems by adding reliability into our product roadmap.
- Coordinate the re prioritise or fix critical bugs for support or sales requirements as needed.
- Make recommendations to production infrastructure by interfacing with engineering to ensure 100% availability.
- Ensure scalable artifacts deployment to all environments by automation scripts.