SRE Manager

Number of employees

170

San Francisco, CA, United States

Posted on: 2024-12-06

Category: energy

Apply now

Please let Crusoe know you found this job on Work in Green. This will help us grow!

Employment type:

Full time

Experience required:

Senior

Salary

Salary not provided

About the company:

Crusoe exists to bring energy to ideas. We are the pioneers of clean computing infrastructure that reduces both the costs and the environmental impact of the world’s expanding digital economy. By unlocking stranded sources of energy to power cloud and data center services, we are creating the climate-aligned future of compute-intensive innovation that reduces rather than adds to emissions. The world’s appetite for computation, energy, and progress will never stop growing. Crusoe is here to bring energy to ideas in ways that are aligned with the needs of our climate.


Crusoe is building the World’s Favorite AI-first Cloud infrastructure company. We’re pioneering vertically integrated,  purpose-built AI infrastructure solutions trusted by Fortune 500 companies to power their most advanced AI applications. Crusoe is redefining AI cloud infrastructure, with a mission to align the future of computing with the future of the climate. Our AI platform is recognized as the "gold standard" for reliability and performance. Our data centers are optimized for AI workloads and are powered by clean, renewable energy.

Be part of the AI revolution with sustainable technology at Crusoe. Here, you'll drive meaningful innovation, make a tangible impact, and join a team that’s setting the pace for responsible, transformative cloud infrastructure.

About This Role:

As the SRE Manager, you will lead the creation and operation of a 24/7 Site Reliability Engineering team. Your primary goal is to ensure continuous availability and optimal performance of our cloud infrastructure, providing customers with uninterrupted access to their GPUs. You will design and implement advanced alerting and monitoring systems, manage incident response, and drive system improvements. Collaborating with remote teams across time zones, you will prioritize projects and streamline workflows to achieve rapid results. This role offers the opportunity to significantly impact the reliability of our cutting-edge cloud services and drive the success of our team.

A Day in the Life:

As a Site Reliability Engineering Manager at Crusoe Energy Systems, your day is a blend of people management and operational oversight. Your morning starts with one-on-one meetings and team stand-ups, focusing on guidance, support, and aligning daily goals. You'll spend about 40% of your time on team development, strategic planning, and fostering a collaborative environment.

The remaining 60% is dedicated to operational tasks, such as reviewing performance metrics, overseeing incident responses, and driving automation projects. You ensure high SLIs and SLOs while resolving technical issues and optimizing processes. By day's end, you review project progress and plan the next steps, maintaining a high-performing, customer-centric SRE organization.

You Will Thrive In This Role If:

  • You have at least 3 years of experience with building and managing a 24/7 technical support team in a cloud operations environment.

  • You have a strong background in Linux, containerization technologies, and Kubernetes. You understand virtualization and cloud computing concepts.

  • You have worked with Prometheus, Victoria Metrics, exporters, against bare-metal endpoints

  • You have some experience with Infrastructure as it relates to Data Center Operations.

  • You’re interested in playing a key role in talent acquisition and retention. This includes diligent performance management and coaching/developing your team according to their individual needs.

  • You’ve developed training programs for new hires and ongoing professional development opportunities for your team members.

  • You like the idea of serving as a technical escalation point and ensuring the highest quality of support. You have experience with Implementing quality assurance measures.

  • You have supported, monitored, and handled Service Level Agreements (SLAs) for a variety of categories that enable an end customer

  • You have used technologies such as RabbitMQ, Kafka, Temporal, NATs

  • You can produce solid solutions in Golang or Python

  • You’re strategic about tracking and reporting KPIs, with a focus on team performance and customer satisfaction. You’ve played a big part in the strategic planning for a team’s growth and scalability.

  • You like the idea of working with other departments to align on technical escalations, live incidents, customer needs, and feedback.

  • Leadership & Communication: Demonstrated leadership ability and excellent communication skills.

  • Problem-Solving & Adaptability: Robust problem-solving skills and adaptability in a fast-paced environment.

  • Project Management: Experience with project management tools and methodologies.

  • Embody the Company values

Benefits: 

  • Hybrid work schedule

  • Competitive Paid Time Off

  • Industry competitive pay

  • Retirement benefits

  • Healthcare benefits including Medical, Dental, and Vision

  • Short and Long-Term Disability Insurance

  • Life Insurance

  • Paid Parental Leave

  • Subscription to Calm App

Compensation Range

Compensation will be paid as salary. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant’s education, experience, knowledge, skills, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Similar climate jobs

These are some of our top picks for great climate jobs on Work in Green.

View all jobs
Fluence logo
United States
Arcadia logo
India
Fluence logo
United States
Number of employees

1010

Full time
Energy
Crusoe logo
United States
Number of employees

170

Full time
Energy
Fluence logo
United States
Number of employees

1010

Full time
Energy
Fluence logo
United States
Number of employees

1010

Full time
Energy

280 Energy jobs at Crusoe

Crusoe is hiring Service Desk 3,Mechanical Quality Assurance Specialist,CNC Operator, and more.

View all jobs at Crusoe
Crusoe logo
United States
Number of employees

170

Full time
Energy
Crusoe logo
United States
Crusoe logo
USA
Number of employees

170

Full time
Energy
Crusoe logo
USA
Number of employees

170

Full time
Energy
Crusoe logo
USA
Number of employees

170

Full time
Energy
Crusoe logo
United States
Number of employees

170

Full time
Energy