RemoteFull timeMid level$207k – $259kPosted today
Apply with JobAssistAbout the role
- We are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role, you will be responsible for the reliability, scalability, performance, and security of our core systems and services
- You will leverage your extensive expertise in various technologies to design, implement, and maintain robust infrastructure and automation solutions
- Implement and maintain the infrastructure and pipeline required for an internal LLM-powered chat service, potentially leveraging platforms like OpenRouter or similar alternatives
- Implement and maintain highly available, scalable, and secure cloud-native infrastructure on Amazon Elastic Kubernetes Service (EKS)
- Develop and implement comprehensive observability strategies, including monitoring, logging, and alerting, to ensure the health and performance of our systems
- Architect and optimize data pipelines to ensure efficient and reliable data flow across various platforms
- Drive the continuous improvement of our CI/CD pipelines, promoting best practices for automated testing, deployment, and release management
- Champion cloud-first strategies, leveraging the full capabilities of cloud platforms for infrastructure, services, and operations
- Implement and enforce robust security practices across our infrastructure, applications, and data
- Design and maintain Docker-based containerization solutions for our applications
- Develop and maintain automation scripts and tools using Python, Bash, and PowerShell
- Collaborate with development teams to ensure reliability is built into the software development lifecycle from inception
- Troubleshoot complex production issues across various layers of the stack, identifying root causes and implementing preventative measures
- Participate in on-call rotations to support production systems- Proven track record in designing and implementing robust data pipelines (e.g., Kafka, Airflow, Spark)
- Ability to work independently and as part of a highly collaborative team
- Expert-level knowledge of cloud platforms (AWS preferred), including infrastructure-as-code principles
- Strong background in CI/CD methodologies and tools (e.g., Jenkins, GitLab CI, ArgoCD)
- Comprehensive understanding of security best practices for cloud environments, applications, and data
- Solid understanding of networking concepts, distributed systems, and operating systems
- 12+ years of experience in Site Reliability Engineering, DevOps, or a similar role with a strong focus on operational excellence
- Excellent problem-solving, analytical, and communication skills
- Advanced scripting and programming skills in Python, Bash, and PowerShell
- Proficiency in Docker for containerization and orchestration
- Extensive experience with observability tools and practices, including Prometheus, Grafana, ELK stack, or similar
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience
- Deep expertise in Amazon EKS, including cluster provisioning, management, and troubleshooting
- Successful candidates must be able to demonstrate U.S. citizenship, permanent residency, or status as a protected individual to satisfy ITAR, contractual, and/or regulatory requirements
- Certifications in AWS, Kubernetes, or other relevant technologies
- Experience with other Kubernetes distributions or cloud providers
- Familiarity with compliance frameworks (e.g., SOC 2, HIPAA, GDPR)
Millions of jobs, with real people getting hired every day
20,000+
New jobs added daily7,000,000+
Verified job listings500,000+
Tailored applications submittedFAQ
Questions, answered
Click "Apply with JobAssist" – we tailor your resume and application to this role and submit it for your approval.
Yes. This role at Archer was screened before publishing – we confirmed the employer before listing it.
The employer didn't disclose a salary range for this listing. JobAssist shows pay whenever it's available.
This position can be done from anywhere, with no in-office requirement.
Yes – every application is tailored from your profile and this job's requirements, and you can review and edit before it's sent.
