Build and scale a strong culture of operational excellence by defining standards and coaching teams to own reliability and availability.
Drive mature DevOps/SRE practices, including incident response and PIRs, on-call readiness, runbooks, alerting, observability, and release/change management.
Establish reliability frameworks such as SLIs/SLOs and error budgets, and use them to guide prioritization and engineering trade-offs.
Guide teams in the design, development, evolution, and operation of large-scale, distributed cloud systems.
Influence product and system direction through design reviews, architectural discussions, and cross-team collaboration.
Requirements
There are more than 20M users of Grafana, the open source visualization tool, around the globe, monitoring everything from beehives to climate change in the Alps.
Grafana Labs also helps more than 3,000 companies -- including Bloomberg, JPMorgan Chase, and eBay -- manage their observability strategies with the Grafana LGTM Stack, which can be run fully managed with Grafana Cloud or self-managed with the Grafana Enterprise Stack , both featuring scalable metrics ( Grafana Mimir ), logs ( Grafana Loki ), and traces ( Grafana Tempo ).
Grafana Cloud k6 is built around the OSS k6 and targeted at users looking to run performance tests at scale.
You can use modern AI coding assistants as part of your daily workflow (your choice of tools, within security guidelines), backed by a company-funded usage budget so you can iterate quickly without unnecessary friction.
We encourage pragmatic AI-assisted development: faster prototyping, test generation, refactors, documentation, and incident follow-ups—always paired with strong code review and quality standards.
You’ll also have access to frontier models (e.g., GPT-Codex 5/3, Claude Opus 4.6, Gemini 3 Pro). Requirements: Strong
experience with DevOps/SRE practices, including operating and evolving production systems at scale
Strong programming background in a modern language (Python and Go are primary, but prior experience is not required) •
Strong understanding of reliability engineering concepts (e.g. incident management, observability, and failure modes) •
Experience with test automation, including performance and functional testing
Ability to influence engineering practices through clear technical communication, reviews, and collaboration
Strong interpersonal skills and ability to work effectively across teams
Familiarity with modern software engineering processes and delivery practices
Experience with containerized and cloud-native systems (Docker, Kubernetes, AWS)
Familiarity with observability tooling and platforms (e.g. the Grafana stack) •
Experience working with Python, Go, JavaScript and/or Jsonnet •
Experience defining or applying SLIs/SLOs, error budgets, or reliability metrics Interest in, or
experience with, building testing frameworks or developer tooling
Grafana Labs may utilize AI tools in its recruitment process to assist in matching information provided in CVs to job postings.
Benefits
Self-driven and comfortable operating with a high degree of autonomy and ambiguity Bonus Points For: •
Compensation & Rewards:
In the US, the Base compensation range for this role is $ 174,986 - $ 209,983 .
Actual compensation may vary based on level, experience, and skillset as assessed throughout the interview process.
*Compensation ranges are country specific. If you are applying for this role from a different location than listed above, your recruiter will discuss your specific market’s defined pay range &
Balance is Key - We operate a global annual leave policy of 30 days per annum. 3 days of your annual leave entitlement are reserved for Grafana Shutdown Days to allow the team to really disconnect. *We will comply with local legislation where applicable.
Equal Opportunity Employer: We will recruit, train, compensate and promote regardless of race, religion, color, national origin, gender, disability, age, veteran status, and all the other fascinating characteristics that make us different and unique.
Additional details
Grafana Labs is a remote-first, open-source powerhouse.
The instantly recognizable dashboards have been spotted everywhere from a NASA launch and Minecraft HQ to Wimbledon and the Tour de France.
We’re scaling fast and staying true to what makes us different: an open-source legacy, a global collaborative culture, and a passion for meaningful work.
Our team thrives in an innovation-driven environment where transparency, autonomy, and trust fuel everything we do.
You may not meet every requirement, and that’s okay. If this role excites you, we’d love you to raise your hand for what could be a truly career-defining opportunity.
This is a remote opportunity, and we would be interested in applicants in United States time zones. The Opportunity
We are the team behind Grafana k6 , Grafana Cloud k6 , and Grafana Cloud Synthetics , used by teams globally to ensure resilient, high-performing systems. This opportunity is with the Grafana Cloud k6 squad, who build and operate our performance testing product.
Our enterprise and SaaS offerings allow customers to load test their systems by running distributed tests from 15+ regions worldwide, using hundreds of thousands of virtual users sending millions of requests per second.
We ingest huge volumes of data generated by k6, which can be used to view, correlate and analyze metrics from each test.
k6 is a product used by other engineers, and as such, we are looking for people enthusiastic about building high-quality tools they would want to use themselves.