infrastructure
Posted Feb 16Staff Software Engineer I - SRE
at Confluent
In India, IndiaRemote
Responsibilities
- Analyze systemic failure patterns and design improvements that prevent incident recurrence
- Define and maintain SLO/SLA frameworks; use error budgets to guide reliability investments
- Build tooling and automation to reduce incident response toil and scale team impact
- Own Rootly configuration, workflows, and integrations with PagerDuty, Jira, Confluence, and Slack
- Analyze reliability data to identify systemic improvements; build dashboards that drive action
- Design scalable reliability standards that reduce reactive workload over time. - Incident Management Program (~25% of role)
- Own standards, practices, and continuous improvement of incident response
- Develop and deliver training programs for engineering teams at all levels
- Coach teams through post-mortems and on developing actionable corrective actions - Customer Root Cause Analysis (CRCA)
- Edit and review customer-facing incident documents to ensure quality and clarity
- Drive turnaround SLAs while maintaining technical accuracy
- Ensure clear explanation of what happened, why, and how we'll prevent recurrence - Cross-Team Leadership
Requirements
- ABOUT THE ROLE: Confluent Cloud processes millions of events per second across AWS, GCP, and Azure.
- Explore AI-assisted approaches to documentation quality and incident analysis
- experience with at least one of AWS, GCP, or Azure
- - Deep expertise with incident management tooling (Rootly, PagerDuty, or similar platforms) - Strong understanding of distributed systems and failure modes at scale—Kafka/event streaming expertise preferred, or demonstrated rapid mastery of complex systems - Deep
- experience with observability: metrics, logging, tracing—ability to diagnose complex issues
- Kubernetes and container orchestration experience
- Understanding of CI/CD pipelines and release processes
- Familiarity with SLO/SLA frameworks. - Track record as a trusted advisor across engineering organizations
- Strong written communication (design docs, one-pagers, runbooks)
- Experience with async collaboration across time zones - Large company
Experience
- Be the expert who teams proactively engage for guidance WHAT YOU WILL BRING: - 10+ years in SRE, incident management, or reliability engineering Cloud
Additional details
- We’re rewriting how data moves and what the world can do with it.
- Our platform puts information in motion, streaming in near real-time so companies can react faster, build smarter, and deliver experiences as dynamic as the world around them.
- It takes a certain kind of person to join this team.
- Those who ask hard questions, give honest feedback, and show up for each other.
- Just smart, curious humans pushing toward something bigger, together. One Confluent. One Team.
- When incidents happen in a multi-cloud streaming platform, they happen at scale—data in motion, exactly-once semantics, and cascading failure modes that require deep systems thinking.
- This role combines hands-on technical work with strategic program ownership.
- You'll spend roughly 75% of your time on engineering: building automation, improving tooling, analyzing systemic failure patterns, and designing reliability improvements.
- The remaining 25% is teaching and coordination: coaching teams through post-mortems, training incident commanders, and evolving our incident response practices.
- You'll be part of a global team with follow-the-sun coverage, with clean handoffs that keep everyone working sustainable hours.