Crusoe Energy Systems logo

Staff Hardware Systems Engineer

Crusoe Energy Systems

RemoteFull timeMid level$215k – $260kPosted today
Apply with JobAssist

About the role

  • We are seeking a Staff Hardware Systems Engineer to strengthen Crusoe’s Hardware Systems Engineering team and close critical skill gaps in debugging, validation, performance evaluation and production support of high-performance compute systems
  • In this role, you will participate in the full hardware lifecycle – from prototype bring-up to large-scale production while driving automation, deep issue resolution, and reliability across Crusoe Cloud’s GPU- and CPU-based infrastructure
  • You will be collaborating with hardware, software, infrastructure, and vendor engineering teams while working across platform bring-up, validation and performance characterization
  • Your work will directly impact Crusoe’s ability to deploy and operate sustainable, AI-first compute systems with world-class performance and reliability
  • Drive the end-to-end lifecycle of next-generation compute platforms, including evaluation, bring-up, validation, deployment, and production readiness
  • Define and execute performance characterization and validation strategies for CPU, GPU, and accelerated computing platforms
  • Conduct in-depth workload characterization studies across training and inference - dense, MoE, long-context, and multimodal models to understand compute, memory, communication, and I/O behavior on target platforms
  • Translate workload and platform insights into cluster-level tuning and configuration recommendations: topology, parallelism strategy, scheduling, power, and software stack settings to maximize delivered performance and efficiency
  • Build and maintain workload performance profiles and reference configurations that guide how clusters are deployed, tuned, and scaled for specific model families and workload classes
  • Analyze system and workload performance, identify bottlenecks, and work across hardware and software layers to drive improvements
  • Lead complex system-level debugging across compute, memory, storage, networking, accelerators, and platform firmware
  • Partner with vendors and internal engineering teams on prototyping, qualification, NPI, and production readiness of new technologies
  • Collaborate across hardware, firmware, networking, software, infrastructure, reliability, and operations teams to resolve complex platform issues
  • Use data and system-level insights to influence platform architecture, technology selection, hardware roadmaps, and long-term infrastructure strategy

Benefits

  • Health & wellbeing: Comprehensive health benefits designed to support your overall wellness
  • Time away: Paid time off for vacations, family bonding, and unexpected needs
  • 401(k) match: Build your financial future with our 401(k) matching program
  • Mental wellness: Resources and support for your emotional wellbeing and navigating life’s challenges- Hands-on experience with system bring-up, validation, performance characterization, and root-cause analysis of complex hardware/software issues
  • Excellent technical communication skills and experience collaborating with internal engineering teams, customers, and external technology partners
  • Hands-on experience with large-scale GPU or accelerated computing infrastructure for AI/ML or HPC workloads
  • Experience with workload benchmarking, performance profiling, and system performance optimization across hardware and software layers
  • Experience working across multiple engineering disciplines, including hardware, firmware, software, networking, and infrastructure teams
  • Experience developing automation, testing, diagnostics, or data-analysis frameworks using Python, Shell, or similar languages
  • Strong analytical and problem-solving skills with the ability to operate effectively in ambiguous and rapidly evolving environments
  • Strong understanding of modern server and accelerator architectures, including CPU, GPU, memory, storage, networking, and high-speed interconnects such as PCIe, InfiniBand, or NVLink
  • Hands-on experience with distributed training and/or inference workloads at scale, including parallelism strategies and performance tuning across the hardware/software stack
  • Ability to analyze system behavior using telemetry, benchmarks, profiling tools, and other quantitative data
  • Bachelor’s or Master’s degree in Electrical Engineering, Computer Engineering, Computer Science, or equivalent experience
  • 8+ years of experience in hardware systems engineering, platform engineering, performance engineering, ML systems engineering, infrastructure engineering, or related areas
  • Experience influencing hardware or system configuration decisions based on workload performance data (e.g: HW/SW co-design, platform tuning studies)
  • Deep experience with RDMA, RoCE, CXL, NVLink or fabric-level performance analysis
  • Experience with inference serving frameworks, training frameworks, or ML compiler/runtime stacks
  • Familiarity with both x86 and ARM-based server platforms
  • Experience building observability, diagnostics, or fleet-level performance and reliability systems
  • Experience introducing new compute technologies into production cloud or large-scale datacenter environments
  • Understanding of infrastructure efficiency, power, cooling, performance-per-dollar, or total cost of ownership considerations
  • Background in sustainable or energy-efficient hardware design practices
  • Advanced certifications or coursework in AI/HPC hardware systems

Millions of jobs, with real people getting hired every day

20,000+
New jobs added daily
7,000,000+
Verified job listings
500,000+
Tailored applications submitted
FAQ

Questions, answered

Click "Apply with JobAssist" – we tailor your resume and application to this role and submit it for your approval.

Yes. This role at Crusoe Energy Systems was screened before publishing – we confirmed the employer before listing it.

The employer didn't disclose a salary range for this listing. JobAssist shows pay whenever it's available.

This position can be done from anywhere, with no in-office requirement.

Yes – every application is tailored from your profile and this job's requirements, and you can review and edit before it's sent.