Manager Site Reliability Engineering
Posted on: December 09, 2025 | Job ID: #LI-IU
Company Name: Veeam Software | Team: Product Development
Division: Site Reliability Engineering (SRE) / Engineering Mgmt.
Location: Pune, MH, India | Employment Type: Full-Time
Schedule: Hybrid (Office-Centric) | Experience Level: 07+ Years
Discover Pune
Pune, Maharashtra offers an excellent blend of professional opportunities and quality of life. As one of India's leading IT hubs, the city provides world-class infrastructure, vibrant culture, and a thriving tech community. Explore the location and discover what's nearby as you consider joining our team in this dynamic city.
About Veeam Software
Veeam is the number one global market leader in data resilience, empowering businesses with complete control over their data whenever and wherever they need it. Our comprehensive platform delivers cutting-edge solutions across five critical pillars: data backup, data recovery, data portability, data security, and data intelligence.
Headquartered in Seattle, Washington, Veeam protects over 550,000 customers worldwide who trust us to keep their businesses running without interruption. We partner with some of the world's biggest brands, providing them with the reliability and resilience they need in today's data-driven economy. At Veeam, we're not just building software—we're shaping the future of data resilience while creating an environment where our people can grow, learn, and make real impact. Join us as we move forward together, going fearlessly forward into the future.
About This Role
Veeam is strategically expanding its global Site Reliability Engineering organization to support the rapidly growing Veeam Data Cloud platform. As an SRE Manager, you will report directly to our Global Director of SRE and play a pivotal role in building and leading a high-performing team that operates at the critical intersection of product innovation, platform engineering, and security.
Your team will partner closely with product teams, platform engineers, and security engineering to design and build systems that are reliable, scalable, and observable from the ground up. This isn't about bolting on reliability after the fact—you'll be embedding it into the DNA of our services from day one.
You will collaborate extensively with peer engineering leaders across the organization to integrate reliability principles into service roadmaps. As a key voice in global SRE planning, you'll represent your team in the design and delivery of cross-cutting reliability initiatives that span all Veeam Data Cloud services, ensuring consistency and excellence across our entire platform.
A core aspect of your leadership will be driving the adoption of modern SRE principles throughout the organization. This includes establishing and operationalizing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. You'll champion toil reduction strategies, ensuring your team spends time on high-value engineering work rather than repetitive operational tasks. You'll also foster a blameless learning culture where incidents become opportunities for growth and systemic improvement.
You will establish and operate a healthy, sustainable, daytime follow-the-sun on-call model in partnership with our SRE teams in other global regions. This approach ensures 24/7 coverage while maintaining work-life balance and preventing burnout among team members.
Your team will be deeply technical, making direct code contributions that lead to measurable improvements in the overall operability, reliability, resilience, and security of the codebases you support. This is hands-on leadership that combines people management with technical excellence.
What You'll Do
People & Team Leadership
Your primary responsibility is building and nurturing a world-class SRE team. You will lead the full talent lifecycle, from hiring and onboarding new team members to developing their careers and managing performance. You'll identify high-potential talent, conduct thorough interviews, and build a diverse team with complementary skills.
Creating the right culture is paramount. You will foster a psychologically safe environment where team members feel empowered to take calculated risks, speak up about problems, and learn from failures without fear of blame. This blameless culture prioritizes learning over finger-pointing and emphasizes proactive engineering over reactive firefighting. You'll model these behaviors yourself and hold your team accountable to these values.
Operational sustainability is critical. You will ensure sustainable operational coverage by carefully monitoring on-call health metrics, workload distribution, and team burnout indicators. You'll implement rotation schedules that are fair and sustainable, ensuring no team member is overburdened.
A key metric you'll track is toil—the repetitive, manual, automatable work that drains engineering time. You will actively track and cap toil so that your engineers spend the majority of their time (ideally 50% or more) on project work that reduces future toil through automation, better tooling, and systemic improvements. You'll celebrate toil reduction as a key team objective.
You will conduct regular one-on-one meetings, provide thoughtful feedback, identify career growth opportunities, and advocate for your team members' advancement within the organization. Your leadership will directly impact their professional development and job satisfaction.
Reliability Strategy & Governance
You will establish and operationalize SLIs, SLOs, and error budgets in close partnership with service owners and product teams. This means working collaboratively to define what reliability means for each service, setting measurable objectives, and creating error budgets that balance feature velocity with system stability.
You'll run regular reliability reviews where you examine SLO compliance, incident trends, and reliability investments. These reviews will hold teams accountable to outcomes while providing constructive support to improve reliability posture.
You will define organization-wide reliability standards, create comprehensive runbooks that enable anyone to respond to incidents effectively, develop readiness checklists for new service launches, and establish intelligent alerting patterns including SLO-based alerting that reduces noise and focuses on what truly matters.
Crucially, you'll partner with product managers and engineering managers to position reliability work as an enabler rather than a gate. Reliability investments should accelerate feature delivery by reducing incidents, improving deployment confidence, and creating stable foundations for innovation. You'll build these partnerships through influence, collaboration, and demonstrating clear business value.
Operations & Incident Excellence
When incidents occur—and they will—your team must be ready. You will ensure incident response readiness through regular training, clear escalation paths, well-maintained runbooks, and effective tooling. During major incidents, you will lead or coordinate response efforts, ensuring clear communication, rapid mitigation, and thorough documentation.
After every significant incident, you will drive fast, high-quality postmortems that identify root causes and, more importantly, systemic improvements to prevent recurrence. These postmortems will be blameless, focusing on what went wrong with systems and processes rather than who made mistakes.
You will measure and improve key reliability metrics including Mean Time To Recovery (MTTR), change failure rate, SLO compliance posture, and repeat-incident reduction. These metrics will guide your team's priorities and demonstrate progress over time.
Importantly, you'll publish learning broadly across the organization. Incident reports, reliability patterns, and hard-won lessons will be shared so that all teams can benefit from your team's experiences and expertise.
Engineering & Automation
Your team will take a software-first approach to reliability. You will lead investments in observability infrastructure, ensuring teams have the telemetry and insights they need to understand system behavior. This includes instrumentation, logging, metrics, tracing, and dashboards.
You'll drive deployment safety improvements through techniques like canary deployments, blue-green deployments, feature flags, and automated rollback mechanisms. These patterns enable teams to deploy confidently and frequently while minimizing risk.
You will champion resilience testing and chaos engineering practices, proactively identifying weaknesses before they cause customer impact. You'll also establish self-service guardrails that make it easy for teams to do the right thing by default.
Beyond service reliability, you'll drive platform improvements that scale operations and enhance developer experience. This includes advancing Infrastructure as Code practices using tools like Terraform or Pulumi, optimizing CI/CD pipelines with GitHub Actions, ArgoCD, or Azure DevOps, and enhancing Kubernetes operations.
You'll also lead the development of internal tools that multiply team effectiveness—whether that's automation frameworks, deployment tooling, incident management systems, or observability platforms.
What You'll Bring
Required Qualifications
Experience: You bring seven or more years of hands-on experience in Software Engineering, Platform Engineering, or Reliability Engineering, with at least two years in engineering management roles where you've directly managed and developed engineers.
Delivery Track Record: You have demonstrable experience leading engineering teams to predictably deliver outcomes. You know how to set goals, break down complex work, unblock teams, and ship results on schedule.
Cross-Functional Leadership: You have experience leading cross-functional initiatives collaboratively with peers through influence rather than authority. You know how to build consensus, navigate organisational complexity, and drive alignment across teams.
Cloud & Infrastructure Expertise: You have deep, hands-on experience with public cloud platforms, with Azure strongly preferred. You understand cloud architecture patterns, cost optimisation, security, and scalability.
Kubernetes Mastery: You have production experience with Kubernetes, understanding its architecture, deployment patterns, networking, storage, security, and operational considerations.
Infrastructure as Code: You're proficient with IaC tools like Terraform or Pulumi, using them to manage infrastructure declaratively, repeatably, and at scale.
CI/CD Knowledge: You have hands-on experience with modern CI/CD systems including GitHub Actions, ArgoCD, and Azure DevOps. You understand pipeline design, artifact management, deployment strategies, and release automation.
Observability Skills: You have practical experience with observability stacks including OpenTelemetry for instrumentation, Elastic for logging and search, Datadog for monitoring and APM, Prometheus for metrics collection, and Grafana for visualization and dashboards.
Coding Background: You have a solid coding background and have personally improved service reliability through code. You can review pull requests, contribute to codebases, and guide technical decisions.
Incident Management: You have hands-on experience managing incidents, running postmortems, and implementing systemic fixes. You understand incident command systems and effective crisis communication.
Communication Excellence: You excel at cross-geographical communication, able to collaborate effectively with distributed teams across time zones, cultures, and working styles.
On-Call Willingness: You're willing to participate in an on-call rotation, typically during daytime hours including weekends and holidays, demonstrating that you're in the trenches with your team.
Preferred Qualifications
SLO/Error Budget Leadership: You have demonstrated success leading SLO adoption programs and error-budget frameworks for cloud services. You've seen these practices transform how organizations think about reliability.
Follow-the-Sun Experience: You have experience operating multi-region, follow-the-sun on-call models where teams across different geographies collaborate to provide continuous coverage.
Resilience Testing Background: You have background in chaos engineering, resilience testing, performance testing, and release validation. You've used these techniques to proactively find and fix issues.
Team Building Track Record: You have a proven track record building or scaling SRE teams from the ground up and influencing organization-wide reliability standards and practices.
Compliance Knowledge: You have familiarity with compliance frameworks common to SaaS businesses such as SOC 2, ISO 27001, GDPR, or industry-specific regulations.
What You'll Get - Comprehensive Benefits Package
Time Off & Work-Life Balance
Veeam values your wellbeing and recognizes that rest and personal time are essential for sustained performance. You'll receive 18 paid vacation days annually to recharge and spend time with loved ones. Additionally, we provide three global VeeaMe Days specifically dedicated to self-care, mental health, and personal wellbeing—use them however you need. We also offer 24 volunteer hours annually, enabling you to give back to causes you care about in your community.
Health & Insurance Coverage
Your health and security matter to us. We provide comprehensive private medical coverage for you and up to four dependents, ensuring your family has access to quality healthcare when needed. We also provide life insurance, accident insurance, and disability insurance with enhanced coverage options, protecting you and your family against unforeseen circumstances.
Wellbeing Support
We offer a wellbeing allowance specifically for physical and mental wellness activities—whether that's gym memberships, fitness classes, meditation apps, therapy sessions, or other wellness investments. Additionally, you'll have free access to confidential counseling and coaching services through our Employee Assistance Programme (EAP), providing professional support whenever you need it.
Practical Daily Benefits
We provide meal benefits, fuel allowances, and transportation benefits based on your specific work arrangement and location, making your daily work life more convenient and affordable. For eligible employees, we offer daycare reimbursement to help with childcare costs and a safe cab facility to ensure secure transportation.
Professional Development & Learning
Your growth is our priority. We invest heavily in professional training and education opportunities. You'll have access to courses, workshops, and internal meetups where you can learn from colleagues and industry experts. We provide unlimited access to premium online learning platforms including LinkedIn Learning for professional skills, Athena for leadership development, and O'Reilly for technical education. Through our MentorLab program, you'll be matched with experienced mentors who can guide your career development and help you achieve your professional goals.
Equal Opportunity Employment
Veeam Software is proud to be an equal opportunity employer. We are committed to creating an inclusive environment where everyone can thrive regardless of race, color, religion, gender, age, national origin, citizenship status, disability, veteran status, or any other classification protected by federal, state, or local law. We believe diversity makes us stronger and more innovative.
All information you provide will be kept strictly confidential and processed in accordance with our Recruiting Privacy Notice, which outlines how we handle personal data during recruitment.
Application Process
By applying for this position, you consent to the processing of your personal data in accordance with our Recruiting Privacy Notice. By submitting your application, you acknowledge that the information provided in your job application and any supporting documents is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification of information may result in disqualification from consideration for employment or, if discovered after employment begins, termination of employment.
Ready to Lead Reliability Engineering?
This is your opportunity to shape the future of reliability engineering at a global leader in data resilience. You'll build and lead an exceptional team, work with cutting-edge technologies, drive meaningful organisational change, and make a real impact for customers worldwide.
Save this role. Apply now. Go fearlessly forward with Veeam

