jobloom

JobLoom finds jobs directly from company career sites before many job boards, then routes you into detailed role pages like this one.

other

Posted Jun 3

Senior Incident Manager

at Lambdalabs

United StatesHybrid

Responsibilities

  • Incident Management Operations - Own the incident response lifecycle including: - Assisting Technical Triage - Escalation - Coordination - Resolution Post-incident review - Ensure timely and accurate communication with internal stakeholders and leadership.
  • - Conduct analysis on incidents and identify patterns / trends for improvement in response and systems reliability.
  • Identify systemic reliability gaps and implement corrective actions. - Track incident metrics including MTTR, MTTD, and incident recurrence rates.
  • Operational Excellence - Improve incident response processes, escalation paths, and tooling by working with technical support and engineering teams.. - Contribute to runbooks, operational standards, and reliability frameworks. - Support implementation of automation and observability improvements.

Requirements

  • Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers.
  • Our customers range from AI researchers to enterprises and hyperscalers.
  • If you'd like to build the world's best AI cloud, join us. We are seeking a Senior Incident Manager to lead critical incident response across our AI data center infrastructure.
  • Role Overview The Senior Incident Manager is responsible for leading the end-to-end lifecycle of operational incidents impacting AI infrastructure and data center services.
  • This role requires deep operational expertise in high-availability infrastructure, large-scale GPU clusters, networking, and cloud platforms, along with strong leadership and communication skills.
  • What You’ll Do Incident Leadership - Lead the response to critical (SEV-1 / SEV-2) incidents impacting AI infrastructure, GPU clusters, networking, storage, and data center operations.
  • experience in incident management, site reliability engineering, or infrastructure operations -
  • Experience managing incidents in large-scale distributed infrastructure environments - Strong understanding of: - Data center operations - GPU compute clusters Networking and storage infrastructure - Cloud or hybrid infrastructure platforms - Proven ability to lead high-pressure incident response situations -
  • Experience with incident management frameworks (ITIL, SRE, or equivalent) - Excellent communication and stakeholder management skills -
  • Experience with incident tracking and monitoring tools such as: - PagerDuty - ServiceNow - Jira - Datadog - Prometheus / Grafana Nice to Have -
  • Experience operating AI or HPC infrastructure - Background in SRE, infrastructure engineering, or data center operations - Familiarity with high-density GPU environments (NVIDIA clusters, InfiniBand networks) -
  • Experience with hyperscale or colocation data center environments - Knowledge of automation and incident response tooling - Knowledge of and
  • experience with Incident command system (ICS) -
  • Experience in leading and developing incident command from stractch Key Competencies - Incident Command & Leadership - Operational Decision Making - Cross-Team Coordination - Root Cause Analysis - Crisis Communication - Infrastructure Reliability What Success Looks Like in This Role - Reduced Mean Time to Resolution (MTTR) for critical incidents - Improved cross-team incident coordination - High-quality post-incident reviews and corrective actions - Increased infrastructure reliability and operational
  • About Lambda - Founded in 2012, with 500+ employees, and growing fast - Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove - We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG - Our

Benefits

  • However, a salary higher or lower than this range may be appropriate for a candidate whose

Additional details

  • Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence.
  • This role is responsible for coordinating rapid resolution of service-impacting events, improving operational resilience, and driving incident management best practices across infrastructure, networking, platform engineering, and data center operations.
  • This individual acts as the central command point during major incidents, ensuring rapid triage, cross-team coordination, effective communication, and structured post-incident analysis.
  • - Serve as the Incident Commander during major outages, coordinating engineering, networking, facilities, and vendor teams.
  • - Act as the liaison between leadership and external teams during incidents / post-incidents to provide updates and status summaries.
  • - Establish clear incident timelines, triage actions, and resolution plans.
  • - Maintain incident response documentation and operational playbooks.
  • - Work in an On-Call Rotation to respond to, lead, and coordinate incidents Cross-Functional Coordination - Work closely with: - Data center operations - Infrastructure engineering & operations - Network engineering - Platform reliability engineering - Security operations - Hardware and facility vendors - Drive alignment during outages involving multiple infrastructure layers.
  • Post-Incident Analysis & Continuous Improvement - Lead post-incident reviews (PIRs) and root cause analysis.
  • Communication & Reporting - Provide executive-level incident summaries and reports. - Deliver clear, concise updates during active incidents. - Maintain incident dashboards and operational health reporting. You - 8+ years

Find more real-time jobs on JobLoom.