header alt image test
Microloft Logo

Microloft

Principal Software Engineer

Reposted Yesterday
Be an Early Applicant
Remote
Hiring Remotely in United States
143K-304K Annually
Senior level
Remote
Hiring Remotely in United States
143K-304K Annually
Senior level
Lead architecture and technical direction for Azure cloud infrastructure, Kubernetes platforms, storage systems, and machine learning infrastructure. Design scalable distributed systems for model training, inference, data processing, and platform services. Improve reliability, security, resource efficiency, and operational performance while supporting MLOps and production AI workloads. Partner with engineering and product leaders, establish technical standards, guide cross-team architecture decisions, and mentor engineers.
The summary above was generated by AI
Overview
Shape the future of cloud-native compute at global scale by joining as a Principal Software Engineer driving the next generation of Azure Kubernetes Service and Azure Storage platforms. You will help deliver reliable, secure, and high‑performance infrastructure that powers critical workloads for customers around the world, including modern artificial intelligence scenarios. In this role, you will lead the architecture and evolution of large-scale infrastructure spanning distributed systems, machine learning infrastructure, model serving, training platforms, observability, and platform engineering. You will partner closely with engineering and product leaders to define technical direction, set high engineering standards, and ensure operational excellence for planet-scale services. You will mentor engineers across teams, elevate engineering practices, and help translate complex technical challenges into durable, customer-centric solutions that advance our cloud platform.
 
At Microsoft, our mission to empower every person and every organization on the planet to achieve more guides how we partner with customers to deliver trusted, impactful solutions. With a growth‑mindset culture, we innovate responsibly and measure success by shared progress, people, teams, and customers. Join us to do meaningful work that changes the world and helps shape what’s next for everyone.

Responsibilities
  • Define technical direction for cloud infrastructure and machine learning platforms across multiple engineering teams.
  • Design and review distributed systems that support model training, model inference, data processing, and platform services.
  • Work with partner teams to align architecture, reliability, security, scalability, and operational requirements.
  • Improve platform efficiency, including compute utilization, resource management, and service performance.
  • Support Machine Learning Operations (MLOps) practices for model development, deployment, monitoring, and lifecycle management.
  • Provide technical guidance for Kubernetes-based platforms and Artificial Intelligence (AI) workloads running in production environments.
  • Contribute to long-term platform planning, technical standards, and engineering best practices across the organization. 

Qualifications

Required Qualifications:

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, Go, C, C++, C#, Java, JavaScript, or Python 
    • OR equivalent experience.

Other Requirements:

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:
    • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.
Preferred Qualifications
  • Master's Degree in Computer Science, Computer Engineering, or a related technical field AND 8+ years of technical engineering experience developing software in Go, C, C++, C#, Java, JavaScript, or Python
    • OR Bachelor's Degree in Computer Science, Computer Engineering, or a related technical field AND 12+ years of technical engineering experience developing software in Go, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent practical experience.
    • Deep experience with Kubernetes, containers, cloud platforms, networking, storage systems, site reliability engineering, and large-scale distributed systems.
  • Experience designing and operating machine learning platforms, large-scale model training environments, graphics processing unit infrastructure, distributed batch scheduling systems, machine learning operations frameworks, KubeRay, and Kueue.
  • Strong understanding of modern large language model and foundation model ecosystems, including training, inference, model serving, and observability.
  • Experience building and scaling enterprise artificial intelligence infrastructure and platforms that support production workloads.
  • Demonstrated ability to lead architecture decisions and drive technical strategy across cross-functional engineering teams.
#azurecorejobs

Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar Jobs at Microloft

38 Minutes Ago
Remote
United States
166K-331K Annually
Senior level
166K-331K Annually
Senior level
Automation
Design, develop, test, and support highly scalable Azure Storage services and distributed systems. Drive architectural decisions, technical direction, reliability, scalability, performance, security, observability, and operational excellence at hyperscale. Provide technical leadership across projects, support highly available services used by millions of users, and apply AIOps for incident detection, root-cause analysis, and mitigation.
Top Skills: AiopsAzure StorageCC#C++Cloud ServicesDistributed SystemsJavaJavaScriptAzurePython
Yesterday
Remote
United States
143K-304K Annually
Senior level
143K-304K Annually
Senior level
Automation
Define architecture and technical strategy for Azure Specialized cloud platforms. Design highly available distributed services, improve scalability, reliability, observability, security, diagnostics, and operational efficiency. Establish service health frameworks, SLOs, metrics, and governance while reducing incidents and customer impact. Partner across engineering, networking, security, product, and operations teams, and mentor engineers to advance technical excellence.
Top Skills: Artificial IntelligenceAzureAzure Vmware SolutionCC#C++Cloud InfrastructureDistributed SystemsJavaJavaScriptNetworkingPythonTelemetry
5 Days Ago
Remote
United States
143K-304K Annually
Senior level
143K-304K Annually
Senior level
Automation
Define technical strategy for enabling AI models and hardware generations on Microsoft's Maia accelerators. Lead cross-stack investigations across models, serving systems, infrastructure, and hardware; improve performance, scalability, reliability, validation, and automation; and align teams on durable platform capabilities and long-term technical investments.
Top Skills: Ai AcceleratorsAi Model ArchitecturesAi-Assisted Engineering WorkflowsAutomationC++Inference SystemsPythonServing Stacks

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account