jobloom

JobLoom finds jobs directly from company career sites before many job boards, then routes you into detailed role pages like this one.

operations

Posted 4 days ago

Senior Manager, Platform Operations

at Collibra

East Coast Usa, United StatesRemote

Responsibilities

  • Own 24x7 incident response, SaaS customer operations, and change management in order to deliver consistent platform reliability for every enterprise customer
  • Build succession and development plans across the team's two technical domains in order to create space for the team's strengths to surface and grow
  • Build working relationships with both technical leads and the full team across the two operational domains
  • Identify at least one workflow suited for AI-driven automation and begin scoping it
  • Build working partnerships across engineering, product, security, finance, and support organizations
  • Demonstrate measurable improvement in at least one operational health metric, such as MTTR or change failure rate, in at least one domain

Requirements

  • Advance AI-powered automation across incident remediation, release delivery, and internal tooling in order to reduce manual toil and continuously modernize how the team operates
  • experience in engineering, with at least 3+ years in a leadership or management role, overseeing incident response, release, or deployment operations for customer-facing SaaS production environments •
  • Experience managing 24x7 production operations for customer-facing systems, including on-call rotation and escalation models •
  • Experience operating within a FedRAMP or comparable regulated compliance environment, including continuous monitoring and audit cycles •
  • Experience with cloud infrastructure at enterprise scale, including AWS, AWS GovCloud and GCP. Azure experience is a plus. •
  • Experience managing distributed infrastructure fleets, virtual machines, containers, or equivalent, supporting production SaaS environments
  • Demonstrated proficiency in leveraging AI tools (e.g., Claude, Gemini, ChatGPT, Copilot) to solve real-world business challenges, drive measurable outcomes, or streamline workflows
  • A bachelor's degree or equivalent related working experience is required
  • Apply infrastructure automation and observability tooling such as Ansible, Terraform, Python, Kubernetes, and platforms like Datadog or Prometheus/Grafana Measures of success
  • Have implemented at least one AI-driven improvement to incident remediation, release delivery, or internal tooling

Experience

  • Partner across engineering, product, security, finance, and support organizations in order to align priorities, resolve cross-team dependencies, and maintain compliance in a regulated environment You have 7+ years of

Benefits

  • Shape a documented point of view on how the team's operating model evolves for an AI-first future Compensation for this role
  • The standard base salary range for this position is 168,000 - 210,000 per year.
  • This position is not eligible for additional commission-based compensation.
  • Salary offers are based on a combination of factors, including, but not limited to, experience, skills, and location. In addition to base salary, we offer a competitive total rewards package, including bonus potential, equity for eligible roles, a Flex Fund monthly stipend, pension/401k plans, and more.
  • These flexible offerings sit on a foundation of competitive compensation, health coverage, and time off.
  • We create inclusion and belonging through how we onboard, meet, connect, engage, and communicate. Learn more about diversity, equity, and inclusion at Collibra.

Additional details

  • You'll manage the team that keeps the Collibra Platform running and current for every enterprise customer, around the clock, spanning incident response, change management, release execution, and the health of the fleet that powers it.
  • You'll report to the Senior Director of Reliability Engineering & Operations and oversee the team directly, partnering closely with two technical leads who anchor deep expertise across the team's two domains.
  • This is a highly visible, business-critical function.
  • When this team is at its best, customers don't notice it.
  • Their environments are healthy, current, and secure.
  • We're looking for someone who treats that responsibility as a craft, not just a job.
  • Our customers are our true north. Every alert answered, every release shipped, and every patch applied is in direct service of the enterprise customers running on this platform.
  • Direct deployment and fleet operations across thousands of virtual machines underpinning the Collibra Platform, including weekly release delivery, in order to keep customer environments healthy, current, and secure
  • Because this role supports the US government, it is required that this candidate be a US citizen who resides on US soil You are able to
  • Balance hands-on technical depth with people leadership, staying credible in an incident or a release window while building growth paths for the team around you

Find more real-time jobs on JobLoom.