header alt image test
Microloft Logo

Microloft

Network Engineer II

Reposted 3 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United States
102K-219K Annually
Junior
Remote
Hiring Remotely in United States
102K-219K Annually
Junior
Design, build, test, and operate high-performance, low-latency networking for large-scale AI/HPC workloads in Azure. Troubleshoot incidents, validate devices and firmware, automate deployments, monitor telemetry, and support datacenter technicians.
The summary above was generated by AI
Overview

The HPC/AI (High performance Computing and Artificial Intelligence) team is on a mission to build the next-generation distributed AI supercomputer, enabling breakthroughs in artificial intelligence by delivering unmatched computational power, scalability and reliability. We design and develop cutting-edge infrastructure that supports high-performance AI model training at scale, laying the foundation for innovations that redefine what AI can achieve.

We are seeking passionate and innovative engineers to design, build and manage cutting-edge networking infrastructure that powers large-scale AI training. This role focuses on developing next-generation networking capabilities to ensure high performance, low latency, and minimal jitter for distributed AI workloads. You will play a critical role in enabling state-of-the-art AI systems to achieve their full potential.
As a Network Engineer on the HPC/AI team, you will play a pivotal role in shaping and managing the next-generation networking infrastructure for AI training and inference in Azure Cloud. This is a unique opportunity to work at the intersection of two of the hottest fields in technology: AI and high-performance computing. With the explosive growth of generative AI and the increasing demand for large-scale, low-latency systems, this area is at the forefront of innovation and impact. You will work across diverse network architectures and cutting-edge processor and accelerator technologies, driving the design and delivery of a comprehensive, end-to-end solution with a relentless focus on performance, scalability, and observability. If you’re passionate about groundbreaking technology, large-scale systems, and AI infrastructure, join us to build the platform that will power the future of AI supercomputing!
Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.


Responsibilities
  • Demonstrates some knowledge of data — knows what data is needed, knows how to find new or missing data, and can describe defects and their relevance to product and service targets. Identifies patterns and trends in data and interprets them to inform decisions related to products and/or services.
  • Collaborates with teams across the organization to support and manage safe and secure network deployments.
  • Works with machine-readable definitions to manage deployments.
  • Supports the management of incidents by applying technical knowledge to diagnose and triage issues with a commitment to maintaining the quality of products and services. Takes notes during incidents and participates in postmortem and root cause analysis processes.
  • Performs testing and validation of network devices, firmware, and configurations. Defines and implements test cases with existing automation tools, and exposes test coverage gaps.
  • Triages, troubleshoots, and repairs live site issues by applying an understanding of network components and features (e.g., device operating systems) as well as problem management tools (e.g., root cause analysis, trend analysis, postmortems), to discover and drive solutions with minimal or no disruption to customers. Actively participates in on-call/DRI duties to troubleshoot and may actively resolve incidents in production.
  • Monitors network telemetry and performs analyses to identify patterns that reveal errors and unexpected problems. Makes suggestions on improvements to monitoring based on observations and experience.
  • Provides instructions to datacenter or network site staff/technicians on how to securely repair, replace, and maintain physical network hardware and components deployed in production. Identifies gaps and inefficiencies in processes related to securely installing and deploying new hardware and components and provides instructions to address gaps. 

Qualifications

Required Qualifications:

  • Master's Degree in Electrical Engineering, Optical Engineering, Computer Science, Information Technology, or related field AND 1+ year(s) technical experience in network design, development, and automation OR Bachelor's Degree in Electrical Engineering, Optical Engineering, Computer Science, Information Technology, or related field AND 2+ years technical experience in network design, development, and automation OR equivalent experience.
 

Other Requirements:

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings: 
    • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.
 
Preferred Qualifications:
  • Doctorate Degree in Electrical Engineering, Optical Engineering, Computer Science, Information Technology, or related field OR Master's Degree in Electrical Engineering, Optical Engineering, Computer Science, Information Technology, or related field AND 3+ years technical experience in network design, development, and automation OR Bachelor's Degree in Electrical Engineering, Optical Engineering, Computer Science, Information Technology, or related field AND 5+ years technical experience in network design, development, and automation OR equivalent experience.


Cloud Network Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar Jobs at Microloft

5 Hours Ago
Remote
United States
120K-304K Annually
Mid level
120K-304K Annually
Mid level
Automation
Develop and evaluate multimodal foundation model architectures, design datasets and pipelines, run large-scale experiments, improve training/deployment efficiency, collaborate with infrastructure and product teams, and ensure responsible, safety-aligned AI development.
Top Skills: AzureCC#C++JavaJavaScriptNumpyPandasPythonPyTorchTensorFlow
5 Hours Ago
Remote
United States
120K-261K Annually
Senior level
120K-261K Annually
Senior level
Automation
Design, implement, and debug low-level software for Azure HPC/AI VMs, focusing on hardware/software interactions, device virtualization, and GPU workload performance. Drive architecture, collaborate with partners, own reliability as an on-call DRI, and develop operational playbooks to improve stability and performance.
Top Skills: AzureCC#C++Distributed SystemsGpuHigh Performance ComputingJavaJavaScriptMachine Learning MiddlewareOperating SystemsPythonVirtual MachinesVirtualization Technologies
5 Hours Ago
Remote
United States
120K-304K Annually
Mid level
120K-304K Annually
Mid level
Automation
Design and implement algorithms, model architectures, data mixtures, and scaling laws for large-scale pre-training. Run and oversee flagship training experiments on distributed infrastructure, collaborate with infrastructure/data/post-training teams, and iterate using rigorous, data-driven ablations to advance foundation models.
Top Skills: CC#C++JavaJavaScriptLarge-Scale Distributed SystemsPython

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account