infrastructure
Posted 5 days agoStaff Technical Program Manager, Site Reliability Engineering
at MongoDB
Atlanta, United StatesOn-site
Responsibilities
- Drive Program Planning & Execution – Define program scope, milestones, and success criteria with SRE engineers and leaders.
- Manage dependencies across platform teams, keep work clearly tracked in Jira, and deliver on time
- Strengthen Production Reliability – Lead change management and launch readiness programs.
- Lead Cross-Functional Coordination – Align SRE with Security, Compliance, Cloud platform, and other engineering teams.
- Coordinate cross-team incident response, ensure clear follow-through, and build trust as the go-to driver of complex, multi-team efforts
- Build Scalable Systems & Processes – Design lightweight frameworks and communication patterns that help SRE deliver reliably at scale.
- Proven track record leading large-scale, cross-team platform initiatives through ambiguity and change
Requirements
- As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products.
- Strong knowledge of production change management, software development lifecycle, and reliability metrics (SLOs, SLIs)
- experience with Kubernetes, cloud networking, or observability stacks (metrics, logs, tracing, alerting) Prior
- experience working with or alongside SRE teams
- Familiarity with MongoDB Atlas or other modern cloud database platforms Why This Role
- And you'll think big—scaling the platform and practices that power MongoDB's next phase of growth.
- You'll be in the critical path of high-impact projects, working with executive stakeholders, and supported to grow in your career. About MongoDB
- We have redefined the data platform for the AI era, enabling builders to create, transform, and disrupt industries with software.
- MongoDB’s unified data platform, the most widely available, globally distributed data platform on the market, helps organizations modernize legacy workloads, embrace innovation, and unleash AI.
- Our cloud-native platform, MongoDB Atlas, is the only globally distributed, multi-cloud data platform and is available across AWS, Google Cloud, and Microsoft Azure.
- With offices worldwide and over 67,000 customers, including 75% of the Fortune 100 and AI-native startups, relying on MongoDB for their most important applications, we’re powering the next era of software.
- Our compass at MongoDB is our Leadership Commitment, guiding how and why we make decisions, show up for each other, and win. It’s what makes us MongoDB.
- Learn more about what it’s like to work at MongoDB , and help us make an impact on the world!
Experience
- 8+ years in technical program management, engineering management, or a comparable technical role partnering with software engineering teams
Benefits
- From employee affinity groups, to fertility assistance and a generous parental leave policy , we value our employees’ wellbeing and want to support them along every step of their professional and personal journeys.
- MongoDB’s base salary range for this role is posted below.
- Compensation at the time of offer is unique to each candidate and based on a variety of factors such as skill set, experience, qualifications, and work location.
- Salary is one part of MongoDB’s total compensation and benefits package. Other
- benefits for eligible employees may include: equity, participation in the employee stock purchase program, flexible paid time off, 20 weeks fully-paid gender-neutral parental leave, fertility and adoption assistance, 401(k) plan, mental health counseling, access to transgender-inclusive health insurance coverage, and health
- benefits offerings. Please note, the base salary range listed below and the
- MongoDB’s base salary range for this role in the U.S. is: $126,000 — $248,000 USD
Additional details
- You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA teams.
- Success in this role means smoother launches, clearer roadmaps, stronger reliability metrics and an SRE organization that's better-equipped to deliver predictability at scale.
- This role can be based remotely on the East Coast What You'll Do
- Partner with SREs and product teams to define and operationalize SLOs/SLIs, and use incident data, metrics, and capacity signals to drive prioritization and continuous improvement
- Work yourself out of the "hero" role by leaving teams better-equipped to execute independently Requirements
- Skilled at shaping roadmaps and managing dependencies
- Able to query and interpret metrics, logs, or other data sources to inform decisions and communicate risk
- Excellent communicator—clear, concise, and calm—across engineers, cross-functional partners, and executives
- Low-ego, highly collaborative, and motivated by ownership of hard problems end to end Nice to Have
- Background in large-scale cloud infrastructure or platform engineering