About the role
Location: Tanjong Pagar
Working Hours: Office Hours
Salary: up $5,800 monthly basic
Monitoring, Observability & Service Assurance
- Manage and operate centralized monitoring dashboards and observability platforms across applications, databases, infrastructure (compute and storage), and network environments (on-premises and cloud) supporting 24/7 services.
- Continuously track system health using metrics, logs, and alerts to proactively identify anomalies, performance degradation, and potential issues.
Alerting, Triage & Coordination
- Respond to alerts and anomalies by conducting initial triage, impact assessment, and cross-system event correlation.
- Coordinate and escalate issues to the appropriate teams (Application, Cloud/Infrastructure, Network, Database) to ensure timely resolution in line with SLAs.
Monitoring Strategy, Design & Governance
- Define and implement monitoring strategies, including frameworks, alert thresholds, escalation policies, and observability standards.
- Collaborate with engineering teams to onboard systems into monitoring platforms and establish meaningful metrics and alerts.
- Continuously review and refine monitoring frameworks to reduce noise and improve the signal-to-noise ratio.
Service Health & Reporting
- Develop and maintain real-time service health dashboards for operational monitoring and reporting.
- Track and analyze system, network, and cloud availability, performance trends, and recurring incident patterns.
- Support reporting needs for service management, senior leadership, and key stakeholders.
Incident Support & Service Reliability
- Support major incident management by providing system visibility, diagnostics, and cross-team coordination.
- Identify recurring issues and reliability gaps, driving improvements in system stability, monitoring coverage, and response times.
FinOps, Cost Monitoring & Governance
- Monitor and manage cloud and infrastructure costs, including AWS usage (compute, storage, data transfer), as well as network and connectivity expenses.
- Implement cost allocation, tagging strategies, and budget monitoring with alerting mechanisms.
- Analyze cost drivers and identify optimization opportunities.
Cost Reporting & Optimisation
- Prepare and present cost reports, dashboards, forecasts, and trend analyses.
- Partner with engineering teams to optimize resource utilization, recommending rightsizing and cost-saving initiatives.
- Ensure a balanced approach between cost efficiency, performance, and reliability.
Cross-Team Coordination
- Serve as the central coordination point across Application, Cloud/Infrastructure, Network, and Database teams to align monitoring insights with operational actions.
Continuous Improvement & SRE Evolution
- Drive initiatives to enhance observability, expand monitoring coverage, and automate alerting and response workflows.
- Contribute to the adoption of Site Reliability Engineering (SRE) practices, including SLIs, SLOs, and error budgets.
Documentation
- Maintain documentation for monitoring architecture and dashboards, alerting rules and escalation procedures, cost governance models, and reports.
Operational Support
- Participate in major incident response and critical service monitoring.
- Provide after-hours support, including weekends and public holidays, as required.
Requirements:
- Training in Computer Science, Information Technology, Engineering, or a related field.
- 3 to 5 years of experience in IT operations, system monitoring, NOC, service assurance, or cloud/infrastructure operations.
- Hands-on experience with monitoring and observability platforms.
- Experience working in hybrid environments (on-premises and AWS cloud).
- Strong understanding of system and network monitoring concepts, as well as application and infrastructure health metrics.
- Proficiency with tools such as CloudWatch, Grafana, Prometheus, Splunk, ELK Stack, or similar platforms.
- Ability to analyze and interpret logs, metrics, and alerts effectively.
- Experience with AWS Cost Explorer, budgeting, and tagging strategies, with a good understanding of cloud cost structures and optimization techniques.
- Familiarity with ITIL processes (Incident, Problem, and Change Management), service level management, and observability practices.
- AWS certifications (Associate level or above) or AWS FinOps Certified Practitioner are preferred.
- Familiarity with ITIL processes (Incident, Problem, Change Management), Service Level Management and Observability principles
- AWS Certification (Associate level or above) or AWS FinOps Certified Practitioner preferred
- Strong analytical thinking and problem-solving abilities.
Interested applicants kindly click "Apply now"
We regret that only shortlisted applicants will be notified.
Michelle Lim Yan Ling | R1985041
RecruitFirst | EA13C6342
Millions of jobs, with real people getting hired every day
Questions, answered
Click "Apply with JobAssist" – we tailor your resume and application to this role and submit it for your approval.
Yes. This role at RecruitFirst Pte. Ltd was screened before publishing – we confirmed the employer before listing it.
The employer didn't disclose a salary range for this listing. JobAssist shows pay whenever it's available.
This position can be done from anywhere, with no in-office requirement.
Yes – every application is tailored from your profile and this job's requirements, and you can review and edit before it's sent.
