Member of Technical Staff - AI Cloud Infrastructure

Number of employees

Bay Area, United States

Posted on: 2026-08-05

Category: energy

Ready to make this your next chapter?

Let Emerald AI know you found them on WorkInGreen. It helps more companies post climate jobs here.

Employment type:

Full time

Remote?

Yes

Experience required:

Intermediate

Salary

Salary not provided

About the company:

Emerald AI is a green technology company operating at the intersection of artificial intelligence infrastructure and power grid management. The company develops software and platform solutions that make AI data centers flexible electricity consumers, enabling them to dynamically adjust their power demand in response to grid conditions. This approach transforms data centers from passive power users into active grid partners, helping utilities and grid operators manage stress on existing infrastructure without requiring costly new transmission or generation capacity. Emerald AI's flagship product, the Conductor platform, serves as an intelligent operating layer between power grids and data centers, coordinating real-time power flexibility while guaranteeing AI workload performance. Emerald AI's mission is centered on enabling the continued scaling of AI infrastructure without triggering grid crises, blackouts, or prohibitive energy cost increases for consumers. The company estimates that intelligent demand flexibility can unlock up to 100 GW of latent capacity on existing grid infrastructure, dramatically accelerating AI deployment timelines while also supporting grid reliability and clean energy integration. By reducing the need to build redundant generation and transmission assets, Emerald AI contributes directly to climate and sustainability goals in the energy sector. The company has achieved significant recognition and commercial traction, being named a 2026 World Economic Forum Technology Pioneer and a 2026 TIME100 Most Influential Company. Emerald AI has announced a collaboration with NVIDIA on its DSX Flex deployment with Silicon Valley Power, and is working within ERCOT's new framework for grid-responsive AI infrastructure in Texas. These milestones position Emerald AI as a pioneering force in the emerging field of grid-interactive AI infrastructure.

About Emerald AI

We’re at a pivotal moment for AI and energy. Demand for compute is skyrocketing, but power constraints are becoming a critical bottleneck. Emerald AI sits at the intersection of these two worlds, enabling AI data centers to scale without overwhelming the grid.

Our Emerald Conductor software platform makes data centers flexible and responsive, allowing them to adjust power usage dynamically. This unlocks massive AI growth without major new infrastructure, while also strengthening the grid and supporting the expansion of renewable energy.

We’re a team of experts across AI, cloud, software, and energy—on a mission to scale AI sustainably. We’re backed by leading investors and partners including Radical Ventures and NVIDIA.

Learn more about our vision, team, and backers at https://www.emeraldai.co/.

About the Role

Emerald AI is building the world's first power flexible managed cloud infrastructure. We are hiring a senior infrastructure engineer to architect and stand up our managed cloud services from end to end. The work covers the platform, the control plane, and the customer experience that together make up a managed AI cloud.

The right person has done this before. They have built or served as a core early engineer on a managed cloud or AI platform, whether at a GPU cloud, an internal machine learning platform run at scale, a hyperscaler AI service, or a HPC research computing center operated as a service. This is a role for an architect who still builds. You will make the major design decisions and then implement them yourself.

Key Responsibilities

  • Architect our managed services from 0→1. Define the productization of GPU capacity, encompassing isolation boundaries, tenant models, provisioning flows, and service catalogs that scale across diverse providers.

  • Engineer the platform core. Build robust control-plane services, self-service customer interfaces, and automated lifecycle systems, including usage metering integrated with billing infrastructure.

  • Onboard and vet infrastructure partners. Conduct deep technical assessments of bare-metal GPU vendors, evaluating fabric quality, network isolation, and economics to automate the path from handoff to active tenant.

  • Design end-to-end multi-tenancy. Implement rigorous isolation across compute, storage, and networking (InfiniBand/VLANs), ensuring secure boundaries, QoS, and encryption even when customers possess root access.

  • Drive workload orchestration. Manage Kubernetes and Slurm environments for large-scale training and inference, overseeing node health, driver fleets, and kernel management across heterogeneous clouds.

  • Lead high-performance storage strategy. Deploy and integrate parallel storage solutions like Lustre, VAST, or Weka, leveraging your deep experience with these systems to ensure they fold cleanly into our provisioning model.

  • Ensure operational excellence. Define SLOs, observability standards, and incident response protocols that bridge our internal standards with underlying provider SLAs to deliver a reliable, sellable product.

Minimum requirements

  • At least 7+ years of experience in infrastructure or platform engineering, including the architecture and launch of a managed cloud or AI platform that reached production users.

  • Strong experience with Kubernetes and Slurm and offering them as managed service

  • Production experience deploying or operating Lustre or a comparable parallel filesystem such as GPFS, Weka, VAST, or BeeGFS, with a solid understanding of parallel filesystem architecture, tuning, and failure modes.

  • A strong grasp of cloud service fundamentals, including control planes, tenancy and isolation models, APIs, quota and metering systems, and the operational discipline of running a service that customers pay for.

  • Deep Linux systems knowledge, mature infrastructure as code practice with tools such as Terraform and Ansible, and solid programming ability in Python or Go.

  • Familiarity with GPU infrastructure, including high performance networking with InfiniBand, RoCE, and RDMA, and the GPU software stack.

Preferred requirements

  • Prior time at a GPU cloud, a hyperscaler AI service, or an HPC center that delivers compute and storage as a service, especially one built on rented or colocated capacity.

  • Familiarity with NVIDIA reference architectures such as SuperPOD, along with GPUDirect Storage, NCCL debugging, and DCGM.

  • Experience with Lustre multitenancy features such as nodemap, fileset mounts, and Kerberos, or with service provider deployments of VAST or Weka.

  • Experience negotiating with and integrating multiple infrastructure vendors, together with a practice of designing for portability between them.

  • Experience running object storage at scale with systems such as S3, Ceph, or MinIO, including the design of data tiering.

  • Experience building billing, metering, or FinOps pipelines for services that charge by usage.

What We Offer

  • Make an impact. Solve the AI power bottleneck and shape how data centers scale sustainably.

  • Join a world-class team of AI, cloud, software, and energy experts in a collaborative, low-ego environment.

  • Build from 0→1. Influence strategy, GTM, org design, and customer/investor engagement from day one.

  • Competitive pay + equity. Stock options let you share in the value you help create.

  • Comprehensive benefits, including medical, dental, vision, and 401(k) matching.

  • Flexible location. Work from D.C., Boston, or the Bay Area, with 2 WFH days/week.

  • Backed by top investors, including Radical Ventures and NVIDIA.

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, disability, and any other protected ground of discrimination under applicable human rights legislation.

Emerald AI strives to respect the dignity and ‎independence of people with disabilities and is committed to giving them the same ‎‎opportunity to succeed as all other employees.

Inclusiveness is core to our culture at Emerald AI, and we strive to ensure you get the most from your interview experience. Emerald AI makes reasonable accommodations for applicants with disabilities. If a reasonable accommodation is needed to participate in the job application or interview process, please reach out to the Talent team.

Get job alerts

Receive new climate job opportunities matching your preferences.

Not quite the right fit? Keep looking.

More climate roles that match your skills and values.

View all jobs
Neara logo
United States
Number of employees

150

Dollar sign

$130,000.00 - $190,000.00

Full time
Energy
Electric Hydrogen logo
United States
Number of employees

360

Full time
Energy
Electric Hydrogen logo
United States
Number of employees

360

Full time
Energy
Neara logo
United States
Number of employees

150

Dollar sign

-

Full time
Energy
WeaveGrid logo
United States
Number of employees

86

Full time
Energy
Crusoe logo
United States
Number of employees

780

Full time
Energy

11 Energy jobs at Emerald AI

Emerald AI is hiring Member of Technical Staff - Research,Member of Technical Staff - AI Cloud Infrastructure,Senior Product Manager, AI Cloud Services, and more.

Hiring at Emerald AI? Feature these jobs →

View all jobs at Emerald AI
Emerald AI logo
United States
Emerald AI logo
United States
Emerald AI logo
United States
Emerald AI logo
United States
Number of employees

Full time
Energy
Emerald AI logo
United States
Number of employees

Full time
Energy