Cohere logo

Software Engineer (GPU Infrastructure, High Performance Computing)

Cohere

RemoteFull timeMid levelPosted today
Apply with JobAssist

About the role

  • The internal infrastructure team is responsible for building world-class infrastructure and tools used to train, evaluate and serve Cohere’s foundational models
  • By joining our team, you will work in close collaboration with AI researchers to support their AI workload needs on the cutting edge, with a strong focus on stability, scalability, and observability
  • You will be responsible for building and operating superclusters across multiple clouds
  • Your work will directly accelerate the development of industry-leading AI models that power Cohere’s platform North
  • Build and scale ML-optimized HPC infrastructure: Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads
  • Optimize for AI/ML training: Collaborate with cloud providers to fine-tune infrastructure for cost efficiency, reliability, and performance, leveraging technologies like RDMA, NCCL, and high-speed interconnects
  • Troubleshoot and resolve complex issues: Proactively identify and resolve infrastructure bottlenecks, performance degradation, and system failures to ensure minimal disruption to AI/ML workflows
  • Enable researchers with self-service tools: Design intuitive interfaces and workflows that allow researchers to monitor, debug, and optimize their training jobs independently
  • Drive innovation in ML infrastructure: Work closely with AI researchers to understand emerging needs (e.g., JAX, PyTorch, distributed training) and translate them into robust, scalable infrastructure solutions
  • Champion best practices: Advocate for observability, automation, and infrastructure-as-code (IaC) across the organization, ensuring systems are maintainable and resilient
  • Mentorship and collaboration: Share expertise through code reviews, documentation, and cross-team collaboration, fostering a culture of knowledge transfer and engineering excellence

Benefits

  • Six weeks’ paid vacation
  • Equity / stock options
  • RRSP, 401(k), and Pension Scheme contributions
  • Coverage for 100% of your insurance premiums across health, dental, vision, and travel
  • Additional coverage for accessing mental health providers/services
  • Six months of fully paid parental leave, including adoption and surrogacy
  • Financial support for egg freezing and IVF in Canada and the UK
  • A monthly fitness and wellness allowance
  • Globally dispersed company that supports a remote work culture
  • A $2,000 annual education benefit for professional development
  • A weekly stipend for meals when working remotely and catered lunch when working from one of our global offices
  • A monthly arts and culture allowance
  • A monthly quality time allowance- Self-directed problem-solving: The ability to identify bottlenecks, propose solutions, and drive impact in a fast-paced environment
  • Low-level systems knowledge: Familiarity with Linux internals, RDMA networking, and performance optimization for ML workloads
  • Kubernetes at scale: Proven ability to deploy, manage, and troubleshoot cloud-native Kubernetes clusters for AI workloads
  • Research collaboration experience: A track record of working closely with AI researchers or ML engineers to solve infrastructure challenges
  • Deep expertise in ML/HPC infrastructure: Experience with GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and high-performance computing (HPC) environments
  • Strong programming skills: Proficiency in Python (for ML tooling) and Go (for systems engineering), with a preference for open-source contributions over reinventing solutions
  • If some of the above doesn’t line up perfectly with your experience, we still encourage you to apply!

Millions of jobs, with real people getting hired every day

20,000+
New jobs added daily
7,000,000+
Verified job listings
500,000+
Tailored applications submitted
FAQ

Questions, answered

Click "Apply with JobAssist" – we tailor your resume and application to this role and submit it for your approval.

Yes. This role at Cohere was screened before publishing – we confirmed the employer before listing it.

The employer didn't disclose a salary range for this listing. JobAssist shows pay whenever it's available.

This position can be done from anywhere, with no in-office requirement.

Yes – every application is tailored from your profile and this job's requirements, and you can review and edit before it's sent.