engineering
Posted 5 days agoSenior Software Engineer, Reliability
at Klaviyo
Dublin, IrelandOn-site
You are nearing today's limit. Upgrade for unlimited access.
Responsibilities
- Build and operate foundational, security-critical services with a strong emphasis on availability, scalability, latency, and fault tolerance
- Design, implement, and evolve systems using SRE best practices
- Define and refine SLIs, SLOs, and error budgets to guide engineering decisions
- Improve observability, alerting, and incident response to reduce mean time to detection and recovery
- Perform quantitative analysis to understand system behavior, capacity constraints, and scaling limits
- Identify systemic risks and reliability bottlenecks and drive long-term, preventative solutions
- Collaborate closely with product, platform, and security engineers to influence architecture early and ship reliable systems
- Mentor and pair with other engineers, helping raise the bar for reliability, operational maturity, and engineering excellence Who You Are:
Requirements
- You write and maintain production-quality code (e.g. Python, Go, or similar) to build internal platforms, automate operations, and improve system reliability
- experience operating containerized workloads and platforms (e.g. Kubernetes) in production, including deployment strategies, scaling behavior, and service networking
- You apply SRE concepts such as SLIs, SLOs, error budgets, and burn-rate–based alerting to guide engineering decisions and operational response You have hands-on
- experience with infrastructure as code and declarative configuration (e.g. Terraform, Kubernetes manifests, policy-as-code)
- You’ve already experimented with AI in work or personal projects, and you’re excited to dive in and learn fast.
- You’re hungry to responsibly explore new AI tools and workflows, finding ways to make your work smarter and more efficient. Nice to Have: •
- Experience supporting security-critical platforms or building internal security tooling
- Familiarity with identity, access management, secrets management, or policy enforcement systems •
- Experience operating systems at scale in cloud environments (AWS preferred)
- Background in resilience testing, fault injection, or chaos engineering
- A strong comprehension of algorithms and data structures at scale Tech Stack:
- Klaviyo’s platform is primarily built with Python and React and runs on AWS. Engineers join us from a wide range of technical backgrounds and are supported in learning our stack.
- Python / Django / FastAPI
- MySQL / Redis / Memcached
- RabbitMQ / Celery / Apache Kafka / Apache Pulsar
- AWS / Terraform / Kubernetes
Benefits
- Our salary range reflects the cost of labour in the country where the job post is advertised.
- The base salary offered for this position is determined by several factors, including the applicant’s job-related skills, relevant experience, education or training, and work location.
- In addition to base salary, our total compensation package may include participation in the company’s annual cash bonus plan, variable compensation (OTE) for sales and customer success roles, equity, sign-on payments, and a comprehensive range of health, welfare, and wellbeing
- Your recruiter can provide more details about the specific salary/OTE range for your preferred location during the hiring process.
- Base Pay Range in Local Currency: €92.000 — €138.000 EUR
Contact
- Want to learn more about life at Klaviyo? Visit klaviyo.com/careers to see how we empower creators to own their own destiny.
Additional details
- At Klaviyo, we value the unique backgrounds, experiences and perspectives each Klaviyo (we call ourselves Klaviyos) brings to our workplace each and every day.
- We believe everyone deserves a fair shot at success and appreciate the experiences each person brings beyond the traditional job requirements.
- If you’re a close but not exact match with the description, we hope you’ll still consider applying.
- As a Senior Software Engineer, Reliability, you’ll ensure Klaviyo’s critical platforms are reliable, scalable, and sustainable while enabling rapid product development.
- We treat reliability as a core product feature and use software engineering to solve complex systems and operational challenges.
- Our work spans security, infrastructure, and software development, requiring us to understand systems and engineering.
- We build complex, foundational solutions that must be extremely reliable, secure, and performant at global scale.
- Our charter is to build and operate foundational services and infrastructure, define clear reliability objectives, reduce operational toil through automation, and continuously improve systems based on real production learnings.
- The work is highly visible and directly impacts how Klaviyos build software and how customers experience Klaviyo every day.
- As a Senior Software Engineer, Reliability, you will build and operate the platforms, systems, and services that underpin Klaviyo’s reliability and operational excellence. You will: