RemoteFull timeMid level$181k – $226kPosted today
Apply with JobAssistAbout the role
- Scale works with the industry’s leading AI labs to provide high quality data and accelerate progress in GenAI research
- We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation
- This role is on the evaluation pod within the GenAI Research Organization and will focus on building benchmarks and diagnosing model failure modes in both text and multimodal modalities
- In this role, you will develop rigorous evaluations and diagnostic methods that reveal where frontier models fail and why
- You will collaborate with researchers and engineers to define best practices in evaluation-driven AI development
- You will also partner with top foundation model labs to translate failure analysis into technical and strategic input on the next generation of generative AI models
- Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents. You’ll identify everything from capability gaps and reasoning errors to robustness and alignment issues, all focusing on RCA
- Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities
- Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to the data and training interventions that address them
- Publish research findings in top-tier AI conferences
Benefits
- Health & Wellbeing: Our holistic approach to supporting Scaliens includes comprehensive health coverage, dental and vision insurance, mental healthcare services, and more. PTO policies and accommodating schedules ensure you’ll get time off when you need it to relax and recharge. Note that our offerings may vary by region as we strive to respond to the unique needs of Scaliens around the globe.
- Personal & Career Growth: Continuously learn and grow through annual learning & development stipend, attending leadership breakfasts, manager training, speaker series, and joining an ERG.
- Building Scale Community: We welcome guests to our offices, and you can expect to see Scalien families and friends around. Join local happy hours, and accept invites to game nights, book clubs, and many other employee-led community events.
- Parental Support: Balancing work and family is essential, and Scale understands the importance of having adequate leave policies in place to promote a healthy home and work life.- Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development
- Excellent written and verbal communication skills
- Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals
- Ph.D. or Master’s degree in Computer Science, Machine Learning, AI, or a related field
- Previous experience in a customer facing role
- Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning
Millions of jobs, with real people getting hired every day
20,000+
New jobs added daily7,000,000+
Verified job listings500,000+
Tailored applications submittedFAQ
Questions, answered
Click "Apply with JobAssist" – we tailor your resume and application to this role and submit it for your approval.
Yes. This role at Scale AI was screened before publishing – we confirmed the employer before listing it.
The employer didn't disclose a salary range for this listing. JobAssist shows pay whenever it's available.
This position can be done from anywhere, with no in-office requirement.
Yes – every application is tailored from your profile and this job's requirements, and you can review and edit before it's sent.
