About GitLab
GitLab is a comprehensive DevOps platform that enables organizations to collaborate on all stages of the software development lifecycle, from planning and creating to securing and deploying. They prioritize a remote-first work culture and strong company values.
About the Position
Introduction
Site Reliability Engineers (SREs) at GitLab ensure all user-facing services and production systems run smoothly. This role combines pragmatic operational skills with software craftsmanship, applying sound engineering principles, operational discipline, and robust automation to our environments and codebase.
In this position, you will join the Tenant Services, Geo team. Geo is a GitLab feature for data replication to a warm-standby, used for data migrations and disaster recovery. The Tenant Services, Geo team supports GitLab Dedicated customer migrations and Geo-related escalations within GitLab Dedicated (excluding FedRAMP environments). While prior Geo or Gitaly experience is not mandatory, familiarity with disaster recovery technologies or GitLab is beneficial; we will provide support for learning our specific stack.
Your team's responsibilities cover the entire Geo operational surface, including pre- and post-cutover data hygiene, migration execution, and non-migration Geo escalations. You will collaborate closely with the core Geo team, Dedicated migrations, and Support to develop a reliable, low-risk cutover model for Dedicated migrations, enhancing tooling, automation, and observability to make migrations faster, safer, and more predictable.
Responsibilities
- Execute end-to-end Dedicated Geo migrations and cutovers, covering planning, pre-cutover validation, execution, and post-cutover verification and cleanup.
- Participate in the team’s shift and weekend rotation for Dedicated cutovers across EMEA and US time zones, and join the SaaS Site Reliability Engineering (SRE) on-call rotation to address incidents affecting GitLab.com availability.
- Operate and enhance the Geo operational surface for Dedicated, including:
- Environment preparation and data hygiene checks before migrations.
- Execution of replication, validation, and cutover procedures.
- Handling Geo-related escalations from Support and internal partners.
- Design, build, and maintain automation, tooling, and runbooks to ensure migrations, cutovers, and Geo escalations are repeatable and efficient.
- Manage our infrastructure using tools like Ansible, Chef, Terraform, GitLab CI/CD, and Kubernetes, contributing improvements to GitLab’s product and infrastructure where relevant.
- Develop and maintain monitoring, alerting, and dashboards that:
- Detect symptoms early, not just outages.
- Track migration and cutover success rates, duration, rollback frequency, and critical SLOs.
- Collaborate closely with:
- The core Geo team to improve Geo features and operability.
- Dedicated migrations and Support teams on migration planning, customer communications, and escalation management.
- Other Infrastructure teams on capacity planning, disaster recovery, and reliability enhancements.
- Contribute to readiness reviews, incident reviews, and root cause analyses, applying lessons learned to improve automation, processes, or products.
- Document all actions, including runbooks, architectural decisions, and post-incident reviews, to establish repeatable practices and automation.
- Proactively identify and reduce operational overhead by automating repetitive tasks and streamlining migration workflows.
Requirements
- Experience operating highly-available distributed systems at scale, preferably in a SaaS environment with customer-facing SLAs.
- Hands-on experience with at least one major cloud provider (e.g., Google Cloud Platform or Amazon Web Services), including networking, storage, and managed services.
- Experience with Kubernetes and its ecosystem (e.g., Helm), including deploying and troubleshooting workloads.
- Experience with infrastructure as code and configuration management tools such as Terraform, Ansible, or Chef.
- Strong programming skills in at least one general-purpose language (preferably Go or Ruby) and proficiency with scripting (e.g., Shell, Python).
- Experience with observability systems (e.g., Prometheus, Grafana, logging stacks) and using metrics and logs to troubleshoot performance and reliability issues.
- Practical exposure to data replication, backup/restore, or migration scenarios (e.g., database replication, storage replication, or Geo-like technologies) where data integrity and downtime risk must be carefully managed.
- Comfort participating in an on-call rotation, investigating incidents across the stack, and driving follow-through on corrective actions.
- Ability to engage directly with enterprise customers during migrations and incidents, including on live calls and through clear written updates.
- Ability to clearly define problems, propose options, and think beyond immediate fixes to improve systems and processes over time.
- Ability to be a “manager of one”: self-directed, organized, and able to drive work to completion in a remote, asynchronous environment.
- Strong written and verbal communication skills, with a bias toward clear, asynchronous documentation and collaboration.
- Alignment with our company values and a commitment to working in accordance with those values.
Nice to Have
- Experience working with disaster recovery technologies.
- Experience with managed/hosted environments similar to GitLab Dedicated, including regulated or compliance-sensitive customers (e.g., SOC2, ISO).
- Prior work on large-scale data migrations or cutovers where customer data integrity, performance, and downtime risk had to be carefully balanced.
- Hands-on experience designing and operating database replication, backup/restore, and cutover workflows (for example, PostgreSQL or cloud-managed equivalents such as AWS RDS), including planning and executing low-risk migrations for large datasets.
- Experience with multi-tenant architectures, sharding, or routing strategies in high-traffic SaaS platforms.
- Familiarity with GitLab (self-managed or SaaS), and/or contributions to open source projects.
How to Apply
Apply Now
Your data is only shared with GitLab
Location
Remote, India
Type
Full-Time
Keywords
Similar Roles
Explore comparable positions
Senior Site Reliability Engineer, Tenant Services: Geo
GitLab