header alt image test
Microloft Logo

Microloft

Member of Technical Staff - Data Flywheel Infra, Frontier Models

Posted 11 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United States
143K-331K Annually
Senior level
Remote
Hiring Remotely in United States
143K-331K Annually
Senior level
Build scalable infrastructure that transforms first- and third-party data, model signals, evaluations, and synthetic data into governed training datasets for frontier AI models. Responsibilities include data ingestion, curation, provenance, licensing, privacy, access control, policy enforcement, quality measurement, synthetic data generation, failure mining, and evaluation-to-training feedback loops. The role requires Python, SQL, and distributed data-processing experience using Spark, Flink, or Ray.
The summary above was generated by AI
Overview
We are looking for a Data Flywheel Infrastructure Engineer to build the infrastructure that continuously turns 1P data, 3P data, model signals, evaluation results, and synthetic data into high-quality training data for frontier LLM and multimodal models.
This role owns the systems connecting:
Data Acquisition → Governance & Compliance → Curation → Training → Evaluation → Failure Mining → Data Improvement
A critical part of the role is enabling aggressive data iteration while ensuring that every dataset is secure, policy-compliant, rights-aware, traceable, and auditable.
 
Starting January 26, 2026, MAI employees are expected to work from a designated Microsoft office at least four days a week if they live within 50 miles (U.S.) or 25 miles (non-U.S., country-specific) of that location. This expectation is subject to local law and may vary by jurisdiction.
 
This role is part of Microsoft AI's Superintelligence Team. The MAIST is a startup-like team inside Microsoft AI, created to push the boundaries of AI toward Humanist Superintelligence—ultra-capable systems that remain controllable, safety-aligned, and anchored to human values. Our mission is to create AI that amplifies human potential while ensuring humanity remains firmly in control. We aim to deliver breakthroughs that benefit society—advancing science, education, and global well-being.
 
We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models! 

Responsibilities

Build 1P & 3P Data Flywheel Infrastructure
Build scalable systems for ingesting, processing, curating, versioning, and serving first-party and third-party data for pre-training and post-training. Connect model failures, evaluations, and product signals back into targeted data acquisition, generation, and improvement workflows.

Own Data Governance, Security & Compliance Infrastructure
Build governance and policy enforcement directly into the data platform, including:

    • Data provenance and lineage
    • Usage rights, licensing, and consent metadata
    • PII / sensitive-data detection and protection
    • Access control and data isolation
    • Retention and deletion enforcement
    • Geographic and regulatory restrictions
    • Dataset approval and audit workflows
    • Training eligibility and purpose-based usage controls

Build Policy-Aware Data Acquisition & Curation Systems
Develop automated pipelines for 1P and 3P data ingestion, classification, filtering, deduplication, quality scoring, semantic enrichment, and dataset construction.
Make governance policies machine-enforceable so that data can automatically be included, excluded, quarantined, or restricted based on its origin, license, sensitivity, consent, geography, and intended model use.

Build Evaluation-to-Data Feedback Loops
Convert model evaluations and real-world failure signals into actionable data tasks through failure clustering, hard-example mining, long-tail discovery, capability-gap detection, and targeted dataset generation.
Enable rapid iteration from:
Model Failure → Data Gap → Data Intervention → Training → Evaluation

Build Synthetic & AI-Native Data Pipelines
Use LLMs, VLMs, and Agents to automate data generation, labeling, filtering, quality validation, enrichment, and transformation.
Maintain clear provenance between human-created, first-party, third-party, model-generated, and derived data, and enforce appropriate policies across each category.

Build Data Quality, Attribution & Observability
Develop metrics and infrastructure to measure dataset quality, coverage, diversity, contamination, duplication, policy compliance, and contribution to model capability improvements.
Enable researchers to understand which data improves which capabilities and under what governance constraints.



Qualifications

Required

  • Master's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 4+ years experience in business analytics, data science, software development, data modeling, or data engineering OR Bachelor's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 6+ years experience in business analytics, data science, software development, data modeling, or data engineering OR equivalent experience.  
  • Software Engineering experience using Python,SQL, Spark/Flink/Ray
    Preferred
  • Experience building AI training-data governance platforms, including provenance, licensing/rights metadata, consent management, PII handling, policy enforcement, or auditable lineage.
  • Experience managing third-party datasets, data partnerships, licensed content, or externally sourced data with complex contractual and usage restrictions.
  • Experience building privacy- and security-aware systems for first-party product or user data, including isolation, access controls, retention/deletion, and purpose limitation.
  • Experience with data clean rooms, privacy-preserving processing, de-identification, confidential computing, or secure data collaboration.
  • Experience building evaluation → failure mining → data generation → training feedback loops.
  • Experience with synthetic data, model graders, reward signals, hard-example mining, active learning, or data-mixture optimization.
  • Experience with multimodal or agentic datasets including text, image, video, audio, web, GUI, tool-use, or interaction trajectories.
  • Understanding of Modern LLM training workflows including Pre-training, SFT, RL/post-training, evaluation, and synthetic data. 
  • Strong understanding of data governance, security, privacy, provenance, access control, and data lifecycle management.

Data Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay

Data Engineering IC6 - The typical base pay range for this role across the U.S. is USD $165,600 - $296,400 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $220,800 - $331,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar Jobs at Microloft

Yesterday
Remote
United States
143K-331K Annually
Senior level
143K-331K Annually
Senior level
Automation
Lead and grow a team building core VM and container platform capabilities for Azure. Drive technical strategy, execution, design, release planning, incident resolution, and cross-team stakeholder collaboration to deliver secure, scalable, high-performance virtualization, container, and confidential computing solutions.
Top Skills: AzureCC#C++Confidential ComputingContainersDevice DriversDistributed SystemsEdgeFirmwareJavaJavaScriptOperating SystemsPythonVirtual MachinesVirtualizationWindows
Yesterday
Remote
United States
86K-223K Annually
Junior
86K-223K Annually
Junior
Automation
Lead delivery for Customer Success Contracts, build relationships with customer stakeholders, drive account planning and consumption, orchestrate cross-functional teams to remove blockers, enable cloud adoption, manage escalations and renewals, and recommend process improvements to achieve customer outcomes.
Top Skills: Cloud SolutionsItilMicrosoft CloudMipPmiProsciUnified Support
Yesterday
Remote
United States
86K-223K Annually
Junior
86K-223K Annually
Junior
Automation
Drive technical pre-sales for cloud and AI application solutions: run demos, PoCs, hackathons, and architecture workshops; design secure, scalable cloud-native architectures; unblock technical issues; and evangelize Microsoft AI Foundry, Azure AI, and Responsible AI to accelerate customer adoption and deployments.
Top Skills: .NetAgentic AiAi Assisted Dev ToolsAi FoundryAPIsAzure AiC#C++ContainerizationCopilot StudioEvent-DrivenFoundry SdkGen Ai OpsGitJavaJupyterMicroservicesMonitoringNode.jsOrchestratorPycharmPythonResponsible AiSemantic KernelSentinelVs Code

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account