Engineering Manager(SevOps)

Posted 4d ago

Pay not disclosedGurgaonOn-site · Job
docker
kubernetes
azure
gcp
terraform
cloud
security
data science
leadership
stakeholder management
global infrastructure

dunnhumby is the global leader in Customer Data Science, partnering with the world’s most ambitious retailers and brands to put the customer at the heart of every decision. We combine deep insight, advanced technology, and close collaboration to help our clients grow, innovate, and deliver measurable value for their customers. dunnhumby employs nearly 2,500 experts in offices throughout Europe, Asia, Africa, and the Americas working for transformative, iconic brands such as Tesco, Coca-Cola, Nestlé, Unilever and Metro. We are looking for a highly motivated Engineering Manager to lead and evolve our Service Operations function into a modern, observability-led, engineering-focused capability. This role is responsible for ensuring operational excellence across production platforms through proactive monitoring, incident management, service reliability engineering (SRE) practices, automation, and continuous service improvement. The Engineering Manager will lead a team of SevOps and Observability specialists, partnering closely with Product, Engineering, Platform, Infrastructure and Support teams to ensure services are resilient, recoverable, scalable, and aligned with business objectives. The role will drive the transition from traditional operational processes to a telemetrydriven, automation-first operating model. What you’ll need: • 15+ years of experience in Engineering, with 7+ years in platform engineering/DevOps/SRE leadership roles. • Proven success leading large-scale platform transformations in cloud-native environments (preferably GCP and Azure). • Hands-on and strategic experience with Kubernetes, CI/CD, GitOps, Terraform, Crossplane, Docker, Infrastructure-as-Code, and multi-tenant platform design. • Deep expertise in platform observability, developer self-service, golden paths, and IDPs such as Backstage. • Advanced understanding of DevSecOps, compliance automation, and security bydesign principles. • Proven experience operating large-scale production environments in cloud and hybrid infrastructure. • Demonstrated success in defining and driving engineering OKRs, metrics based decision making, and cost accountability. • Strong ability to balance technical depth with cross-functional influence, managing senior stakeholders and C-level engagement. • A builder’s mindset with a focus on automation, resilience, scalability, and simplification. • Excellent communication and stakeholder management skills; able to collaborate effectively with teams in India and internationally, and to balance ambition, feasibility and risk. Key Responsibilities: Service Reliability & Operational Excellence • Own the day-to-day operational reliability and availability of business-critical services.Establish and mature SRE practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and reliability reporting. • Drive proactive identification and mitigation of operational risks through telemetry, observability, and data-driven insights. Lead the continual improvement of Incident, Problem, Change, and Major Incide

Found on the web. You apply on the company's own site