MENU
  • Remote Jobs
  • Companies
  • Go Premium
  • Job Alerts
  • Post a Job
  • Log in
  • Sign up
Working Nomads logo Working Nomads
  • Remote Jobs
  • Companies
  • Post Jobs
  • Go Premium
  • Get Free Job Alerts
  • Log in

Senior Engineering Manager - Site Reliability

Ditto

Full-time
USA
devops
java
python
program management
customer experience
Apply for this position

About the role

Ditto is at an inflection point. As we scale to meet the growing demands of our enterprise customers, we need experienced SRE Leads to drive and mature our Site Reliability Engineering practice.

This is a unique opportunity to play a leading role in shaping enterprise-grade reliability, observability and incident management to ensure Ditto's systems meet the high standards our customers expect.

As a Senior Engineering Manager of Site Reliability Engineering, you will lead a multi-layered team of SREs, including other SRE managers, to shape and scale reliability practices across our platform. You will drive strategy, execution, and people development across regions while embedding a culture that values high availability, resiliency, and operational excellence.

As a Senior Engineering Manager, you will:

  • Lead and scale a globally distributed SRE organization, including managers and ICs, setting the long-term vision and execution plan for reliability at scale

  • Develop engineering leaders and senior talent, coaching on both technical depth and leadership maturity to create a high-trust, high-performance organization

  • Drive adoption of SRE best practices, including:

    • Embedding SREs in product teams to influence design and early detection of failure modes

    • Defining production-readiness checklists and launch gates tied to SLOs

    • Championing error budgets as a shared accountability mechanism between product and reliability

  • Establish and evolve an incident management practice, including:

    • Clear roles (Incident Commander, Scribe, Subject Matter Experts, CX & affected customer communication)

    • Blameless postmortems with systemic and meaningful remediations

    • Active tracking of incident themes and reliability KPIs, and reporting to senior leadership

  • Lead the architecture and execution of observability systems that offer real-time visibility into system health and customer experience

  • Partner with platform, infrastructure, and security teams to build scalable, self-service reliability tooling (e.g. circuit breakers, automated rollback, chaos testing frameworks)

  • Guide teams to define, implement, and iterate on SLIs, SLOs, and SLAs that are meaningful to end user experience

  • Establish best-in-class documentation and operational hygiene, including runbooks, architectural decision records (ADRs), and deep operational reviews

  • Model on-call excellence, including burnout prevention, clear handoffs, and leveraging automation and toil elimination

  • Lead strategic programs to transform engineering culture toward reliability, such as:

    • Annual 'Reliability Weeks', engineering health reviews

    • Incentivizing reliability work, such as inclusion in promotion criteria and roadmap planning

    • Designing systems to hold engineering teams accountable for the reliability of their respective systems

  • Design talent acquisition strategies, hiring criteria, and interview modules to build a team of exceptional talent

  • Design and implement a highly effective SRE org structure, including geo located teams, internal leadership and management lines, and integration/partnership points with other team

  • Play a central role in the transformation of Ditto’s engineering culture into a culture that prioritizes reliability and resilience of our mission critical software. Communicate & articulate this mission across the entire company in all hands, presentations, working sessions, and via enactment of strategic objectives

What you’ll bring

  • 8+ years of experience in Site Reliability Engineering or related operational engineering roles

  • 4+ years in engineering leadership, including managing other managers, with a track record of scaling high-performing teams

  • Demonstrated experience leading cultural transformation in engineering organizations, such as:

    • Shifting from reactive firefighting to proactive reliability investment

    • Introducing SLOs and reducing incidents through continuous improvement programs

    • Empower and ultimately require teams to own their service health end-to-end

  • Expertise in cloud-native platforms (Kubernetes, Istio, etc.) and modern IaaS tooling (Terraform, Helm)

  • Strong background in observability and alerting strategy, using tools like Prometheus, Datadog, or OpenTelemetry

  • Hands-on experience with at least two major cloud providers (AWS, GCP, Azure)

  • Previous programing experience in one or more of the following languages (Go, Rust, Java, Python)

  • Excellent communicator and cross-functional partner, capable of influencing product and exec teams on engineering trade-offs

  • Adept at project and program management across multiple priorities, including balancing feature delivery with operational improvements

  • Exposure to chaos engineering, load testing, or resiliency modeling

Nice to have

  • Experience building multi-tenant SaaS platforms with high uptime requirements

  • Familiarity with compliance-driven reliability environments (e.g., regulated industries)

  • Experience scaling DevOps platforms and internal tooling ecosystems

  • Passion for knowledge sharing through internal tech talks, RFCs, or public speaking

Apply for this position
Bookmark Report

About the job

Full-time
USA
Posted 9 hours ago
devops
java
python
program management
customer experience

Apply for this position

Bookmark
Report
Enhancv advertisement

30,000+
REMOTE JOBS

Unlock access to our database and
kickstart your remote career
Join Premium

Senior Engineering Manager - Site Reliability

Ditto

About the role

Ditto is at an inflection point. As we scale to meet the growing demands of our enterprise customers, we need experienced SRE Leads to drive and mature our Site Reliability Engineering practice.

This is a unique opportunity to play a leading role in shaping enterprise-grade reliability, observability and incident management to ensure Ditto's systems meet the high standards our customers expect.

As a Senior Engineering Manager of Site Reliability Engineering, you will lead a multi-layered team of SREs, including other SRE managers, to shape and scale reliability practices across our platform. You will drive strategy, execution, and people development across regions while embedding a culture that values high availability, resiliency, and operational excellence.

As a Senior Engineering Manager, you will:

  • Lead and scale a globally distributed SRE organization, including managers and ICs, setting the long-term vision and execution plan for reliability at scale

  • Develop engineering leaders and senior talent, coaching on both technical depth and leadership maturity to create a high-trust, high-performance organization

  • Drive adoption of SRE best practices, including:

    • Embedding SREs in product teams to influence design and early detection of failure modes

    • Defining production-readiness checklists and launch gates tied to SLOs

    • Championing error budgets as a shared accountability mechanism between product and reliability

  • Establish and evolve an incident management practice, including:

    • Clear roles (Incident Commander, Scribe, Subject Matter Experts, CX & affected customer communication)

    • Blameless postmortems with systemic and meaningful remediations

    • Active tracking of incident themes and reliability KPIs, and reporting to senior leadership

  • Lead the architecture and execution of observability systems that offer real-time visibility into system health and customer experience

  • Partner with platform, infrastructure, and security teams to build scalable, self-service reliability tooling (e.g. circuit breakers, automated rollback, chaos testing frameworks)

  • Guide teams to define, implement, and iterate on SLIs, SLOs, and SLAs that are meaningful to end user experience

  • Establish best-in-class documentation and operational hygiene, including runbooks, architectural decision records (ADRs), and deep operational reviews

  • Model on-call excellence, including burnout prevention, clear handoffs, and leveraging automation and toil elimination

  • Lead strategic programs to transform engineering culture toward reliability, such as:

    • Annual 'Reliability Weeks', engineering health reviews

    • Incentivizing reliability work, such as inclusion in promotion criteria and roadmap planning

    • Designing systems to hold engineering teams accountable for the reliability of their respective systems

  • Design talent acquisition strategies, hiring criteria, and interview modules to build a team of exceptional talent

  • Design and implement a highly effective SRE org structure, including geo located teams, internal leadership and management lines, and integration/partnership points with other team

  • Play a central role in the transformation of Ditto’s engineering culture into a culture that prioritizes reliability and resilience of our mission critical software. Communicate & articulate this mission across the entire company in all hands, presentations, working sessions, and via enactment of strategic objectives

What you’ll bring

  • 8+ years of experience in Site Reliability Engineering or related operational engineering roles

  • 4+ years in engineering leadership, including managing other managers, with a track record of scaling high-performing teams

  • Demonstrated experience leading cultural transformation in engineering organizations, such as:

    • Shifting from reactive firefighting to proactive reliability investment

    • Introducing SLOs and reducing incidents through continuous improvement programs

    • Empower and ultimately require teams to own their service health end-to-end

  • Expertise in cloud-native platforms (Kubernetes, Istio, etc.) and modern IaaS tooling (Terraform, Helm)

  • Strong background in observability and alerting strategy, using tools like Prometheus, Datadog, or OpenTelemetry

  • Hands-on experience with at least two major cloud providers (AWS, GCP, Azure)

  • Previous programing experience in one or more of the following languages (Go, Rust, Java, Python)

  • Excellent communicator and cross-functional partner, capable of influencing product and exec teams on engineering trade-offs

  • Adept at project and program management across multiple priorities, including balancing feature delivery with operational improvements

  • Exposure to chaos engineering, load testing, or resiliency modeling

Nice to have

  • Experience building multi-tenant SaaS platforms with high uptime requirements

  • Familiarity with compliance-driven reliability environments (e.g., regulated industries)

  • Experience scaling DevOps platforms and internal tooling ecosystems

  • Passion for knowledge sharing through internal tech talks, RFCs, or public speaking

Working Nomads

Post Jobs
Premium Subscription
Sponsorship
Free Job Alerts

Job Skills
API
FAQ
Privacy policy
Terms and conditions
Contact us
About us

Jobs by Category

Remote Administration jobs
Remote Consulting jobs
Remote Customer Success jobs
Remote Development jobs
Remote Design jobs
Remote Education jobs
Remote Finance jobs
Remote Legal jobs
Remote Healthcare jobs
Remote Human Resources jobs
Remote Management jobs
Remote Marketing jobs
Remote Sales jobs
Remote System Administration jobs
Remote Writing jobs

Jobs by Position Type

Remote Full-time jobs
Remote Part-time jobs
Remote Contract jobs

Jobs by Region

Remote jobs Anywhere
Remote jobs North America
Remote jobs Latin America
Remote jobs Europe
Remote jobs Middle East
Remote jobs Africa
Remote jobs APAC

Jobs by Skill

Remote Accounting jobs
Remote Assistant jobs
Remote Copywriting jobs
Remote Cyber Security jobs
Remote Data Analyst jobs
Remote Data Entry jobs
Remote English jobs
Remote Spanish jobs
Remote Project Management jobs
Remote QA jobs
Remote SEO jobs

Jobs by Country

Remote jobs Australia
Remote jobs Argentina
Remote jobs Brazil
Remote jobs Canada
Remote jobs Colombia
Remote jobs France
Remote jobs Germany
Remote jobs Ireland
Remote jobs India
Remote jobs Japan
Remote jobs Mexico
Remote jobs Netherlands
Remote jobs New Zealand
Remote jobs Philippines
Remote jobs Poland
Remote jobs Portugal
Remote jobs Singapore
Remote jobs Spain
Remote jobs UK
Remote jobs USA


Working Nomads curates remote digital jobs from around the web.

© 2025 Working Nomads.