jobloom

JobLoom finds jobs directly from company career sites before many job boards, then routes you into detailed role pages like this one.

infrastructure

Posted Feb 16

Staff Software Engineer I - SRE

at Confluent

In India, IndiaRemote

Responsibilities

  • Analyze systemic failure patterns and design improvements that prevent incident recurrence
  • Define and maintain SLO/SLA frameworks; use error budgets to guide reliability investments
  • Build tooling and automation to reduce incident response toil and scale team impact
  • Own Rootly configuration, workflows, and integrations with PagerDuty, Jira, Confluence, and Slack
  • Analyze reliability data to identify systemic improvements; build dashboards that drive action
  • Design scalable reliability standards that reduce reactive workload over time. - Incident Management Program (~25% of role)
  • Own standards, practices, and continuous improvement of incident response
  • Develop and deliver training programs for engineering teams at all levels
  • Coach teams through post-mortems and on developing actionable corrective actions - Customer Root Cause Analysis (CRCA)
  • Edit and review customer-facing incident documents to ensure quality and clarity
  • Drive turnaround SLAs while maintaining technical accuracy
  • Ensure clear explanation of what happened, why, and how we'll prevent recurrence - Cross-Team Leadership

Requirements

  • ABOUT THE ROLE: Confluent Cloud processes millions of events per second across AWS, GCP, and Azure.
  • Explore AI-assisted approaches to documentation quality and incident analysis
  • experience with at least one of AWS, GCP, or Azure
  • - Deep expertise with incident management tooling (Rootly, PagerDuty, or similar platforms) - Strong understanding of distributed systems and failure modes at scale—Kafka/event streaming expertise preferred, or demonstrated rapid mastery of complex systems - Deep
  • experience with observability: metrics, logging, tracing—ability to diagnose complex issues
  • Kubernetes and container orchestration experience
  • Understanding of CI/CD pipelines and release processes
  • Familiarity with SLO/SLA frameworks. - Track record as a trusted advisor across engineering organizations
  • Strong written communication (design docs, one-pagers, runbooks)
  • Experience with async collaboration across time zones - Large company

Experience

  • Be the expert who teams proactively engage for guidance WHAT YOU WILL BRING: - 10+ years in SRE, incident management, or reliability engineering Cloud

Additional details

  • We’re rewriting how data moves and what the world can do with it.
  • Our platform puts information in motion, streaming in near real-time so companies can react faster, build smarter, and deliver experiences as dynamic as the world around them.
  • It takes a certain kind of person to join this team.
  • Those who ask hard questions, give honest feedback, and show up for each other.
  • Just smart, curious humans pushing toward something bigger, together. One Confluent. One Team.
  • When incidents happen in a multi-cloud streaming platform, they happen at scale—data in motion, exactly-once semantics, and cascading failure modes that require deep systems thinking.
  • This role combines hands-on technical work with strategic program ownership.
  • You'll spend roughly 75% of your time on engineering: building automation, improving tooling, analyzing systemic failure patterns, and designing reliability improvements.
  • The remaining 25% is teaching and coordination: coaching teams through post-mortems, training incident commanders, and evolving our incident response practices.
  • You'll be part of a global team with follow-the-sun coverage, with clean handoffs that keep everyone working sustainable hours.

Find more real-time jobs on JobLoom.