xAI logo

Software Engineer - Data

xAI

RemoteFull timeMid levelPosted today
Apply with JobAssist

About the role

  • At xAI, we are building AI systems that push the frontier of human knowledge and scientific discovery. High-quality data is fundamental to every stage of that mission

  • Our Data team is responsible for ensuring that the models are trained on the right data, in the right form, at the right quality, across every phase of the training lifecycle

  • This includes partnering closely with acquisition teams to identify where valuable data can be sourced, determining what data is needed to improve model performance, and building the production pipelines and systems that transform raw inputs into high-quality training data at scale

  • We work at the intersection of data, infrastructure, and machine learning to ensure our models train effectively and reliably

  • As a Data Engineer / AI Engineer on xAI’s Data team, you will be responsible for developing the systems, processes, and production code that power data acquisition, preparation, quality evaluation, and delivery for model training

  • You will work closely with acquisition teams, ML engineers, and software engineers to identify data needs, build scalable data pipelines, and continuously improve the quality of the data that shapes model behavior

  • Analyze the performance and impact of data used throughout the model training lifecycle

  • Investigate anomalous model behavior and rigorously identify the data issues that drive poor downstream performance

  • Design, build, and improve the data cleaning, transformation, and quality-control steps required to produce high-quality training data

  • Research, evaluate, and develop frontier methods for improving data quality and effectiveness in AI model development

  • Apply statistical techniques and empirical analysis to make informed, data-driven decisions about dataset quality and model outcomes

  • Partner across teams to identify where data needs exist and define the highest-impact opportunities for new data acquisition and improvement

  • Build and maintain production-grade data pipelines, tooling, and software systems that ingest, process, validate, and deliver data for training

  • Develop metrics, evaluation frameworks, and monitoring systems to assess how data quality influences model behavior at scale

  • Fuse data from multiple sources into reliable, usable datasets for research and production model training

  • Create shared datasets, tooling, and internal data products that enable other teams to analyze, debug, and improve model performance

Benefits

  • Health and wellness: Comprehensive health insurance including medical, dental, vision, and disability coverage

  • Life and family: Life and AD&D insurance and fertility benefits to ensure our team’s well-being and peace of mind

  • Flexible vacation: We work hard but avoid burn out. Take time off when you need it

  • Visa sponsorship: We support international talent with visa sponsorship to join our team

  • 401(k) plan: Retirement savings plan to secure your financial future

Qualifications

  • The ideal candidate combines strong software engineering fundamentals and excellent coding practices with deep intuition for statistics, neural networks, and how data quality influences training outcomes

  • Bachelor’s degree in computer science, data science, physics, mathematics, or a STEM discipline

  • 1+ years of data/software engineering experience (internship experience is applicable)

  • Experience in implementing or analyzing language models or neural networks

  • Professional experience in analytics, data science, machine learning, or data engineering

  • Experience building and operating production data pipelines for neural network or large-scale machine learning workloads

  • Experience working with Parquet or similar columnar storage formats in large-scale data systems

  • Strong experience with Python and the broader ecosystem of libraries and tools used in modern machine learning and data development

  • Experience developing predictive models and machine learning pipelines, including clustering, forecasting, anomaly detection, or related techniques

  • Familiarity with Kubernetes and distributed production environments

  • Experience working with very large-scale datasets, including terabyte- to petabyte-scale data systems

  • Ability to operate effectively in a dynamic environment with evolving priorities, changing requirements, and fast-moving technical challenges

  • Strong statistical intuition and the ability to use quantitative analysis to guide technical and product decision, including familiarity of scaling ladder design studies

  • Demonstrated ability to take ownership of ambiguous problems, drive projects independently, and develop new expertise where needed

Millions of jobs, with real people getting hired every day

20,000+
New jobs added daily
7,000,000+
Verified job listings
500,000+
Tailored applications submitted
FAQ

Questions, answered

Click "Apply with JobAssist" – we tailor your resume and application to this role and submit it for your approval.

Yes. This role at xAI was screened before publishing – we confirmed the employer before listing it.

The employer didn't disclose a salary range for this listing. JobAssist shows pay whenever it's available.

This position can be done from anywhere, with no in-office requirement.

Yes – every application is tailored from your profile and this job's requirements, and you can review and edit before it's sent.