infrastructure
Posted Nov 13, 2025Site Reliability Engineer
at Gamma
San Francisco, United StatesOn-site
Requirements
- WHAT YOU'LL DO - Own the reliability, availability, and performance of Gamma's production systems across our AWS infrastructure - Build observability infrastructure from the ground up: metrics, logging, tracing, and alerting that give the team genuine visibility into system health before users feel the impact - Design and ship automation that reduces toil, makes deployments safer, and gets us back on our feet faster when things go wrong - Lead incident response and blameless post-mortems, then follow
- experience with infrastructure-as-code (Terraform, CloudFormation) and end-to-end observability solutions - Track record of making systems meaningfully more reliable through automation, smarter monitoring, and architectural improvements - Deep understanding of networking, distributed systems, containerization (Docker, Kubernetes), and database performance at scale - Sharp incident management instincts and the debugging skills to navigate complex production failures -
- Experience scaling SaaS products to millions of users, or background with Kafka, chaos engineering, or service mesh technologies (Nice to have) - AWS certifications, or
Benefits
- experience with security and compliance frameworks like SOC 2 or ISO 27001 (Nice to have) COMPENSATION RANGE: The base salary for this full-time position, which spans multiple internal levels depending on qualifications, ranges between $230K - $310K plus
- benefits & equity. Final offer amounts are determined by multiple factors, including but not limited to experience and expertise in the
Additional details
- ABOUT THE ROLE Gamma's infrastructure needs to be rock-solid for millions of daily users while enabling our engineering teams to ship fast.
- You'll own the operational health of our full backend platform, building automation and tooling that improves reliability and partnering with engineering to design systems that are observable, resilient, and easy to operate.
- Your work directly impacts every Gamma user's experience.
- This is a high-impact role where you'll balance reliability with velocity, knowing when to move fast and when to prioritize stability.