About the role
Production Support / SRE, LINUX/UNIX Administration, SQL, Automation (Python/bash/Perl/Ruby/Shell), Messaging (MQ, CPS, XML, FIX), CI/CD (Git, Artifactory, Jenkins, Docker), Data Streaming (SPARK, Kafka), Observability and Monitoring (Grafana, Splunk, Dynatrace, AppDynamics), Configuration/Release Management Tools (Puppet, Ansible, Chef, GitHub)
12+ Months Contract to Hire
Location: Hybrid - 3 days a week in office is mandatory
Duration: 12+ months (Contract to Hire)
Interview Process:
- 1st: Zoom
- 2nd: Onsite
Description:
- Bachelor's Degree: Yes
- Industry Background: Plus
- Years of Experience: 2 - 5 years
- Shifts: Morning: 8 am - 5:00 pm, Evening: 12:30 pm - 8:00 am, Weekend: On Call (Remote)
Must:
- Linux and Unix Hands-on Experience
- Comfortable troubleshooting at OS Level
- Production Support Experience
- Incident Handling, Debugging live systems, and on-call experience
- Scripting/Automation experience within Python
- Not a Software Engineer but must be able to automate repetitive tasks
- Communication skills MUST BE THERE!
- Will be working hands-on with Dev Teams and the business
- ServiceNow Experience from a ticketing perspective
- Knowledge of ITIL Principles
Plus:
- Good understanding of Java, GO, C++, Scala, etc.
- Grafana and Snowflake
- Any Cloud Experience
- Agentic AI background or knowledge
Must Haves:
- Good hands-on Linux experience and SQL Knowledge for Python (more from a reading code perspective opposed to scripting)
This is a pipeline requisition to be used for all open consultant positions for Reliability & Production Engineering (RPE) in Utah. Please submit candidates who satisfy requirements for either Production Support or System Reliability Engineering (SRE) roles.
Detailed Job Descriptions:
- Production Support Analyst
As Production Support Analyst, your responsibilities will include, but not be limited to:
- Monitoring for and resolving issues across the entire tech stack: hardware, software, application, and network.
- A majority of your time will be devoted to production support activities.
- Working closely with engineering/development teams to address repetitive issues, reduce operational effort, and the likelihood of future service disruptions.
- Partnering with business users and other technology teams to manage significant events such as business continuity/disaster recovery tests, IPOs, stock splits, and major infrastructure changes.
- Defining and refining standard operating procedures (SOP) for everything from monitoring to troubleshooting complex code and infrastructure issues.
- Identifying and driving opportunities to improve platform supportability through automation.
- Advocating for reliability priorities in application design reviews and operational readiness exercises for new and existing services.
- Participating in weekend and off-hours on-call rotation.
- Collaborating and striving to understand business users' needs and problems.
Qualifications External
What skills and experience do I need? You should apply if you have at least a Bachelor's degree in Computer Science or other technical discipline(s), plus hands-on experience with any combination of the following:
- 3-5+ years practical experience in production systems support or application development.
- Hands-on experience managing systems in a large-scale distributed Unix/Linux environment is essential.
- Effective communicator who is comfortable speaking in front of both internal/external groups as well as business clients.
- Demonstrated ability to troubleshoot problems and debug to conclusively identify root causes.
- Knowledge of ITIL Principles. ITIL certification is a plus.
- Knowledge of Unix/Linux operating system level concepts such as processes, memory allocation, and networking, with an understanding of how applications are affected by these, and ability to debug and troubleshoot accordingly.
- Automation-related experience is particularly valued, using scripting languages such as Python, bash, Perl, and/or Ruby.
- Higher-level compiled languages such as C++, C#, JAVA, Scala, and Go are a big plus.
- Working ability to interact with message transport platforms and protocols (MQ, CPS, XML, FIX) and distributed database technologies (DB2, Sybase, Mongo, GreenPlum, Postgres, KDB).
- Autosys scheduling and batch processing concepts.
- Experience with source code and binary repositories, build tools, and CI/CD (Git, Artifactory, Jenkins, Docker) etc., and data streaming technologies like Spark, Kafka etc.
- Hands-on experience on enterprise tools set such as Grafana, Splunk, Dynatrace, AppDynamics, etc.
- Awareness of, and ability to reason through modern software & systems architectures, including load-balancing, queueing, caching, distributed systems failure modes generally, microservices, etc.
- System Reliability Analyst
As a System Reliability Analyst, your responsibilities will include, but not be limited to:
- Working closely with engineering/development teams to design, build, optimize, and maintain systems.
- Troubleshooting issues across the entire technology stack: hardware, software, application, and network.
- Aggressively targeting toil and operational risk, and deploying solutions to reduce these.
- Broadening infrastructure and application observability.
- Proactively identifying and addressing active or potential risks to system reliability.
- Advocating for reliability priorities in application design reviews and operational readiness exercises for new and existing services.
Qualifications External
What skills and experience do I need? You should apply if you have at least a Bachelor's degree in Computer Science or other technical discipline(s), plus hands-on experience with any combination of the following:
- 3-5+ years practical experience in production systems support or application development.
- Hands-on experience managing systems in a large-scale distributed Unix/Linux environment is essential.
- Automation-related experience is required, using scripting languages such as Python, bash, Perl, and/or Ruby. Higher-level compiled languages such as C++, C#, JAVA, Scala, and Go are a big plus.
- Deep knowledge of and hands-on experience applying the principles of System/Site Reliability Engineering (SRE).
- Practical experience designing and instrumenting SLO/SLI dashboards is particularly valuable.
- Hands-on experience on enterprise tools such as AppDynamics, Grafana, Splunk, Dynatrace.
- Experience with Puppet, Ansible, Chef, GitHub or any automation/configuration/release management tools.
- Awareness of, and ability to reason through modern software and systems architectures, including load-balancing, databases, queueing, caching, distributed systems failure modes, microservices, Cloud, etc.
- Working ability to interact with message transport platforms and protocols (MQ, CPS, XML, FIX) and distributed database technologies (DB2, Sybase, Mongo, GreenPlum, Postgres, KDB).
- Autosys scheduling and batch processing concepts.
- Deep understanding of infrastructure and operating system concepts such as processes, memory allocation, and networking, with an understanding of how applications are affected by the above, and ability to debug and troubleshoot accordingly.
For applications and inquiries, contact: hirings@openkyber.com
Millions of jobs, with real people getting hired every day
Questions, answered
Click "Apply with JobAssist" – we tailor your resume and application to this role and submit it for your approval.
Yes. This role at Openkyber was screened before publishing – we confirmed the employer before listing it.
The employer didn't disclose a salary range for this listing. JobAssist shows pay whenever it's available.
This position can be done from anywhere, with no in-office requirement.
Yes – every application is tailored from your profile and this job's requirements, and you can review and edit before it's sent.