other
Posted Jun 3Senior Incident Manager
at Lambdalabs
United StatesHybrid
Responsibilities
- Incident Management Operations - Own the incident response lifecycle including: - Assisting Technical Triage - Escalation - Coordination - Resolution Post-incident review - Ensure timely and accurate communication with internal stakeholders and leadership.
- - Conduct analysis on incidents and identify patterns / trends for improvement in response and systems reliability.
- Identify systemic reliability gaps and implement corrective actions. - Track incident metrics including MTTR, MTTD, and incident recurrence rates.
- Operational Excellence - Improve incident response processes, escalation paths, and tooling by working with technical support and engineering teams.. - Contribute to runbooks, operational standards, and reliability frameworks. - Support implementation of automation and observability improvements.
Requirements
- Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers.
- Our customers range from AI researchers to enterprises and hyperscalers.
- If you'd like to build the world's best AI cloud, join us. We are seeking a Senior Incident Manager to lead critical incident response across our AI data center infrastructure.
- Role Overview The Senior Incident Manager is responsible for leading the end-to-end lifecycle of operational incidents impacting AI infrastructure and data center services.
- This role requires deep operational expertise in high-availability infrastructure, large-scale GPU clusters, networking, and cloud platforms, along with strong leadership and communication skills.
- What You’ll Do Incident Leadership - Lead the response to critical (SEV-1 / SEV-2) incidents impacting AI infrastructure, GPU clusters, networking, storage, and data center operations.
- experience in incident management, site reliability engineering, or infrastructure operations -
- Experience managing incidents in large-scale distributed infrastructure environments - Strong understanding of: - Data center operations - GPU compute clusters Networking and storage infrastructure - Cloud or hybrid infrastructure platforms - Proven ability to lead high-pressure incident response situations -
- Experience with incident management frameworks (ITIL, SRE, or equivalent) - Excellent communication and stakeholder management skills -
- Experience with incident tracking and monitoring tools such as: - PagerDuty - ServiceNow - Jira - Datadog - Prometheus / Grafana Nice to Have -
- Experience operating AI or HPC infrastructure - Background in SRE, infrastructure engineering, or data center operations - Familiarity with high-density GPU environments (NVIDIA clusters, InfiniBand networks) -
- Experience with hyperscale or colocation data center environments - Knowledge of automation and incident response tooling - Knowledge of and
- experience with Incident command system (ICS) -
- Experience in leading and developing incident command from stractch Key Competencies - Incident Command & Leadership - Operational Decision Making - Cross-Team Coordination - Root Cause Analysis - Crisis Communication - Infrastructure Reliability What Success Looks Like in This Role - Reduced Mean Time to Resolution (MTTR) for critical incidents - Improved cross-team incident coordination - High-quality post-incident reviews and corrective actions - Increased infrastructure reliability and operational
- About Lambda - Founded in 2012, with 500+ employees, and growing fast - Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove - We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG - Our
Benefits
- However, a salary higher or lower than this range may be appropriate for a candidate whose
Additional details
- Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence.
- This role is responsible for coordinating rapid resolution of service-impacting events, improving operational resilience, and driving incident management best practices across infrastructure, networking, platform engineering, and data center operations.
- This individual acts as the central command point during major incidents, ensuring rapid triage, cross-team coordination, effective communication, and structured post-incident analysis.
- - Serve as the Incident Commander during major outages, coordinating engineering, networking, facilities, and vendor teams.
- - Act as the liaison between leadership and external teams during incidents / post-incidents to provide updates and status summaries.
- - Establish clear incident timelines, triage actions, and resolution plans.
- - Maintain incident response documentation and operational playbooks.
- - Work in an On-Call Rotation to respond to, lead, and coordinate incidents Cross-Functional Coordination - Work closely with: - Data center operations - Infrastructure engineering & operations - Network engineering - Platform reliability engineering - Security operations - Hardware and facility vendors - Drive alignment during outages involving multiple infrastructure layers.
- Post-Incident Analysis & Continuous Improvement - Lead post-incident reviews (PIRs) and root cause analysis.
- Communication & Reporting - Provide executive-level incident summaries and reports. - Deliver clear, concise updates during active incidents. - Maintain incident dashboards and operational health reporting. You - 8+ years