
Software Engineer (Kernels/CUDA (C++))
xAI
RemoteFull timeMid level$180k – $440kPosted today
Apply with JobAssistAbout the role
- We are building one of the world’s largest AI supercomputers from the ground up
- As part of the Compute Infrastructure team, you will own both the raw GPU supercomputer and the platform layer that runs on top of it
- You will work across the full stack — from low-level GPU kernel optimizations and Linux kernel internals to massive-scale orchestration and virtualization — to make training and inference at xAI as fast, reliable, and scalable as possible
- This is a broad, high-impact role that combines hardcore supercompute and compute infrastructure work
- Your contributions will directly accelerate Grok’s training speed and overall AI progress
- Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads
- Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance
- Work on Linux kernel internals, scheduling, memory management, and resource isolation at cluster scale
- Build custom container orchestration, virtualization layers (KVM, Firecracker, etc.), and distributed systems that go beyond standard Kubernetes
- Profile, debug, and eliminate bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operations
- Create and maintain infrastructure-as-code, automation, and tools that keep the entire supercomputer reliable and efficient
- Collaborate closely with AI research teams to deliver production-grade performance and scalability
Benefits
- Health and wellness: Comprehensive health insurance including medical, dental, vision, and disability coverage
- Life and family: Life and AD&D insurance and fertility benefits to ensure our team’s well-being and peace of mind
- Flexible vacation: We work hard but avoid burn out. Take time off when you need it
- Visa sponsorship: We support international talent with visa sponsorship to join our team
- 401(k) plan: Retirement savings plan to secure your financial future- Ability to reason from first principles and optimize for both memory-bound and compute-bound scenarios
- Track record of building or running high-performance infrastructure for AI workloads (training or inference platforms)
- Strong experience with large-scale GPU clusters or distributed compute infrastructure at production scale
- Experience building and operating high performance exabyte scale storage systems
- Experience with Linux kernel internals, scheduling, virtualization, or large-scale orchestration
- Hands-on work with GPU kernel optimization (CUTLASS, custom kernels, Nsight profiling)
- Deep low-level systems programming (C/C++ or Rust)
Millions of jobs, with real people getting hired every day
20,000+
New jobs added daily7,000,000+
Verified job listings500,000+
Tailored applications submittedFAQ
Questions, answered
Click "Apply with JobAssist" – we tailor your resume and application to this role and submit it for your approval.
Yes. This role at xAI was screened before publishing – we confirmed the employer before listing it.
The employer didn't disclose a salary range for this listing. JobAssist shows pay whenever it's available.
This position can be done from anywhere, with no in-office requirement.
Yes – every application is tailored from your profile and this job's requirements, and you can review and edit before it's sent.