header alt image test
Microloft Logo

Microloft

Software Engineer II

Reposted One Month Ago
Be an Early Applicant
Remote
Hiring Remotely in United States
102K-219K Annually
Junior
Remote
Hiring Remotely in United States
102K-219K Annually
Junior
Design, build, and operate telemetry and data pipelines to monitor flagship supercomputers at scale. Troubleshoot and mitigate incidents, improve observability and runbooks, author postmortems, and implement systemic solutions to improve availability, reliability, and performance for HPC and AI infrastructure.
The summary above was generated by AI
Overview

Microsoft Azure High Performance Computing & AI Engineering (HPC & AI Eng) team is responsible for managing the core platform & fleet of AI & High Performance Computing products that customers use to run their most performant and demanding workloads. The AI Customer Experience (AICE) engineering team within the HPC & AI Eng. team is on the frontlines managing the flagship supercomputers and infrastructure used by top tier AI customers that enable breakthroughs such as ChatGPT and are highlighted in Top500, MLPerf and Graph500 rankings.

We run lean, obsess about customer experience and use evidence-based approach to decision making. We have live-site first, metrics-driven culture that prevents us from accumulating debt and necessity to put out fires on daily basis. You will be in a position that carries a ton of responsibility and provides opportunities to directly impact customers satisfaction.

As a Supercomputing Software Engineer on the AICE team, you will design & develop capabilities needed to monitor & efficiently operate across the infrastructure & fleet of supercomputers at scale. To enable first to know of critical incidents impacting customer capacity, you will create end to end data pipelines that process & synthesize large volume of telemetry, log files and other data sources to create actionable alerts.

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.


Responsibilities
  • Contribute to improving key metrics such as Job Mean Time to Interrupt, Nodes in Service, Mean Time to Resolve on flagship supercomputers.
  • Manages operations of supercomputers by responding quickly to mitigate issues.
  • Implements systemic solutions and mitigations to more complex issues impacting performance or functionality of supercomputers
  • Reviews and writes incident postmortem and presents insights that drive changes to reduce or eliminate incidents.
  • Independently improves troubleshooting guides (TSGs), wikis, tests, and telemetry, adding comprehensive observability and monitoring capabilities.
  • Proactively seeks new knowledge and adapts to new trends, technical solutions, and patterns that will improve the availability, reliability, efficiency, observability, and performance of supercomputers while also driving consistency in monitoring and operations at scale. 

Qualifications

Required Qualifications:

  • Bachelor's Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent experience.

Other Requirements:

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings: 
    • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.

Preferred Qualifications:             

  • Bachelor's Degree in Computer Science
    • OR related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, OR Python
    • OR Master's Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent experience.

Software Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar Jobs at Microloft

Yesterday
In-Office or Remote
MD, USA
102K-219K Annually
Junior
102K-219K Annually
Junior
Automation
Develops, tests, deploys, and operates Azure Data software in highly secured, air-gapped environments. Responsibilities include coding, architecture, automation, secure development, incident response, telemetry, reliability, feature experimentation, and live-site on-call support. The engineer collaborates with partner teams, contributes to design and deployment plans, applies security and compliance controls, and supports Microsoft Azure services for government and regulated-industry customers.
Top Skills: Artificial IntelligenceAutomated TestingAzure DataCC#C++Deployment AutomationFeature FlagsFlightingGenerative AiJavaJavaScriptLoggingAzurePythonSecurity MonitoringTelemetry
6 Days Ago
In-Office or Remote
MD, USA
102K-219K Annually
Junior
102K-219K Annually
Junior
Automation
Develop scalable, secure, and reliable Azure Data services for high-volume usage billing. Responsibilities include designing features, writing code, improving test coverage, automating deployment, integrating telemetry, maintaining live services, responding to incidents, troubleshooting distributed systems, and supporting compliance and security requirements.
Top Skills: AzureAzure DataCC#C++Distributed SystemsIntegration TestingJavaJavaScriptPythonTelemetryWindows
6 Days Ago
In-Office or Remote
MD, USA
102K-219K Annually
Junior
102K-219K Annually
Junior
Automation
Designs, develops, deploys, and operates secure Teams Phone services for sovereign, air-gapped, and compliance-sensitive cloud environments. Responsibilities include writing reliable code, collaborating with partner teams, building scalable distributed services, participating in on-call rotations, troubleshooting production incidents, improving telemetry and automation, and supporting CI/CD and operational excellence. The role requires active Top Secret/SCI clearance with polygraph and U.S. citizenship verification.
Top Skills: AnsibleAzure DevopsCC#C++Ci/CdCloud ComputingDistributed SystemsGithub ActionsJavaJavaScriptObservabilityPstn TelephonyPython

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account