Principal Site Reliability Engineer

Remote: 
Full Remote
Contract: 
Work from: 

Offer summary

Qualifications:

10+ years of experience in a DevOps or SRE role, Extensive experience with enterprise scale continuous delivery environments, Development experience with JavaScript/Node.js/TypeScript in a Linux/Mac environment, Knowledge of cloud platforms, preferably AWS, and container orchestration technologies..

Key responsibilities:

  • Chart the future of Cribl’s observability and reliability systems and practices
  • Engage with Product and Engineering teams to improve service delivery and reliability
  • Measure and monitor all production systems for availability and overall health
  • Identify and drive down toil with creative innovation and automation.

Cribl logo
Cribl Scaleup https://cribl.io/
201 - 500 Employees
See all jobs

Job description

Cribl does differently. 

What does that mean? It means we are a serious company that doesn’t take itself too seriously; and we’re looking for people who love to get stuff done, and laugh a bit along the way. We’re growing rapidly - looking for collaborative, curious, and motivated team members who are passionate about putting customers first. As a remote-first company we believe in empowering our employees to do their best work, wherever they are. 

As the data engine for IT and Security many of the biggest names in the most demanding industries trust Cribl to solve their most pressing data needs. Ready to do the best work of your career? Join the herd and unlock your opportunity.

Why you’ll love this role:

Cribl Inc is seeking a Principal Site Reliability Engineer to join our mission to unlock the value of all observability data. Cribl provides users a new level of observability, intelligence and control over their real-time data. You will join a team of software engineers and SREs who are committed to shipping only high-quality software and enjoying all the goat gifs the internet has to offer. This role is remote and you will be part of the engineering organization where you will contribute in our efforts to envision, create, deploy, test, and ship Cribl products. 

We are looking for  Site Reliability Engineers at all levels at Cribl, who enjoy being in the thick of it. Our SREs are involved from conception, design, development and testing, all the way through to production

If reliability is your passion, if you have always had strong opinions on how to make things better and if you have the desire to build consensus around ideas, let's talk!

As An Active Member Of Our Team, You Will...

  • Chart the future of Cribl’s observability and reliability systems and practices
  • Conceptualize and direct the evolution of our reliability metrics, programs and process based on the state of the art and industry best practices
  • Engage with Product and Engineering teams to improve service delivery and reliability across the entire software lifecycle
  • Measure and monitor all production systems with an eye towards availability, latency and overall system health
  • Uncover risks and seek out the sources of errors and instability in our production systems.
  • Advocate engineering-wide improvements in reliability, observability and promote antifragility
  • Identify and drive down toil with creative innovation and automation
  • Participate in on-call 

If You Got It, We Want It:

  • Extensive experience with enterprise scale continuous delivery environments
  • 10+ years of experience in a DevOps or SRE role
  • Development with JavaScript/Node.js/TypeScript in a Linux/Mac environment
  • Experience with IaC tools like Terraform (preferred) or similar
  • Experience with sustainable incident response in a blameless environment
  • Knowledge of cloud platforms (prefer AWS) and container + orchestration technologies
  • Experience with APM and Observability and related tools such as, New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry etc.
  • Deep understanding of SRE practices, such as SLOs, Error Budgets, PRRs, Problem Management
  • Comfortable with a high level of autonomy and working with a distributed team

Preferred Qualifications

  • Knowledge of Cloud and application security
  • Strong knowledge of cloud design patterns for scale, data management, resiliency, etc.
  • Passion for software quality and craft

 

Salary Range  ($240,000 - $400,000)

The salary for this role is dependent on geographic location. The salary offered within the range described will be based on the individual candidate’s job-related knowledge, skills, and experience.  In addition to a competitive salary, Cribl also offers a generous benefits package which includes health, dental, vision, short-term disability, and life insurance, paid holidays and paid time off, a fertility treatment benefit, 401(k), equity, and eligibility for a discretionary company-wide bonus.

#LI-EL1

#Remote

Bring Your Whole Self
Diversity drives innovation, enables better decisions to support our customers, and inspires change for the better. We’re building a culture where differences are valued and welcomed, and we work together to bring out the best in each other. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or any other applicable legally protected characteristics in the location in which the candidate is applying.

Interested in joining the Cribl herd? Learn more about the smartest, funniest, most passionate goats you’ll ever meet at cribl.io/about-us

Required profile

Experience

Spoken language(s):
English
Check out the description to know which languages are mandatory.

Other Skills

  • Collaboration

Site Reliability Engineer (SRE) Related jobs