header alt image test
Microloft Logo

Microloft

Senior Technical Program Manager, Cluster Operations & Quota Management

Posted 4 Hours Ago
Be an Early Applicant
In-Office
Mountain View, CA, USA
143K-304K Annually
Senior level
In-Office
Mountain View, CA, USA
143K-304K Annually
Senior level
Lead end-to-end execution of quota and cluster operations for MAI's GPU fleet: prepare and sequence changes, run cluster moves and cutovers, coordinate provisioning and access, validate usable capacity, remove blockers, and continuously improve tooling and operating model to increase speed, reliability, and automation.
The summary above was generated by AI
Overview

At Microsoft AI, compute is the foundation everything else is built on: every frontier training run, every eval, and every inference workload depends on our GPU fleet. GPUs are our scarcest and most valuable resource. Leadership sets how that fleet is allocated; this role owns the entire system that makes those allocation decisions real - moving, provisioning, and validating quota across a constrained pool as fast and as cleanly as possible.


We are looking for a Technical Program Manager with deep, hands-on experience in cluster operations and compute capacity management to own quota and cluster operations execution end to end at MAI. You will own the entire system that turns those allocation decisions into usable capacity - reliably, at speed, and at growing scale. This is a high-agency, service-oriented role for someone who is relentless on detail and follow-through, never lets anything drop, and likes turning a fragmented, high-stakes process into a clean, scalable machine.


What You'll Do

As the TPM for Cluster Operations, you own the end-to-end system that gets quota and cluster changes implemented across MAI's GPU fleet - not just the execution of any single change, but the machinery that makes every change fast, clean, and repeatable. When an allocation is decided, you own getting it implemented: preparing and sequencing the changes ahead of time, running cluster moves and cycle cutovers cleanly, coordinating with infrastructure, platform, HPC, and vendor teams to provision capacity, and validating that quota is genuinely usable rather than merely configured. You chase every dependency down, surface and clear blockers, and keep quota flowing to where it matters most.

Owning the system means continuously rebuilding it. You champion the tooling squads and leadership use to see quota, utilization, and idle capacity - partnering with the engineering teams that build it so it reflects reality and so new clusters and tenants onboard smoothly as the fleet grows. And you work proactively: spotting cross-functional dependencies before they bite, streamlining the handoffs between teams, and convening the right working groups across research, infrastructure, platform, and vendor organizations to fix the operating model at its root, not just the symptom in front of you. Every cycle should be faster, cleaner, and more automated than the last, and you are the person accountable for making that true.


Responsibilities
  • Own the end-to-end operating system for quota and cluster execution: the process, playbooks, tooling agenda, and cross-company coordination that turn allocation decisions into usable capacity
  • Execute approved quota and capacity allocations across MAI's squads end to end: prepare and sequence changes in advance, run cluster moves and cycle cutovers cleanly, and drive each one to completion across every partner team.
  • Own the full chain to usable capacity, not just configured quota: coordinate provisioning and environment readiness with infrastructure, platform, HPC, and vendor teams, and validate identity, access, and utilization so researchers are productive from day one.
  • Relentlessly chase dependencies, surface blockers early, and unblock them, keeping a crisp, auditable status on every change so nothing is dropped or misreported.


Qualifications

Qualifications / requirements

We know that strong candidates come from many different backgrounds and experiences. If you are excited about the role and believe you could make an impact, we encourage you to apply, even if you do not meet every qualification listed below.

  • Significant experience in technical program management within infrastructure, platform engineering, or other compute-intensive environments.
  • A record of proactively improving operating models: spotting process failures, convening the right people across organizations, and driving automation and streamlining without waiting to be asked.
  • High agency and a service-oriented mindset, comfortable relentlessly chasing and pushing across teams, without formal authority, to get things done fast.

Technical Program Management IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar Jobs at Microloft

4 Hours Ago
In-Office
Mountain View, CA, USA
143K-304K Annually
Senior level
143K-304K Annually
Senior level
Automation
Lead signal and power integrity (SIPI) architecture and implementation for compute and AI SoCs and platforms. Design, model, and simulate end-to-end power delivery (DC/AC/transient), validate I/O performance, drive chiplet and advanced packaging solutions, and collaborate across silicon, package, motherboard, thermal, and platform teams to deliver production-ready HPC processor systems.
Top Skills: Ac SimulationAdvanced PackagingChiplet ArchitectureDc SimulationDecoupling Capacitor SolutionsFoundry TechnologiesHpc PlatformsIp DesignMotherboard/Pcb DesignOsat TechnologiesPower Integrity ModelingSignal And Power Integrity (Sipi)Silicon NodesSoc DesignSubstrate TechnologiesTransient SimulationVoltage Regulator Design
4 Hours Ago
In-Office
Mountain View, CA, USA
102K-261K Annually
Senior level
102K-261K Annually
Senior level
Automation
Lead applied research projects that advance AI for software developers (code completion, embeddings, agentic systems). Build datasets from public and internal code, train and manage large-scale ML experiments and models, collaborate across Microsoft Research, product teams, and GitHub, and transition research into product through experimentation and evaluation.
Top Skills: Code EmbeddingsCodellmCopilot CliGithub CopilotLlmRagVisual StudioVs Code
4 Hours Ago
In-Office
Mountain View, CA, USA
120K-304K Annually
Mid level
120K-304K Annually
Mid level
Automation
Build and ship scalable, user-centric AI web experiences (Copilot) using front-end systems. Collaborate with researchers, PMs, and designers to design, build, test, and deploy performant React/TypeScript applications while ensuring safety, reliability, and strong UX.
Top Skills: ReactTypescript

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account