Jobs Career Advice Post Job
X

Send this job to a friend

X

Did you notice an error or suspect this job is scam? Tell us.

  • Posted: Sep 20, 2026
    Deadline: Not specified
    • @gmail.com
    • @yahoo.com
    • @outlook.com
  • Imagine a world where people live healthier, more enhanced and protected lives… A world in which each organisation is a powerful influencer and responsible corporate citizen, committed to being a force for social good. As a leading innovator in healthcare, wellness, insurance, investments, financial and life planning, Discovery works ceaselessly to...

     

    Observability Engineer

    Job Summary

    • As part of the Discovery Central Services - Technology Services, the Observability Engineer role at Discovery Limited is responsible for the implementation, delivery, support, maintenance, and continuous enhancement of observability, AIOps, and IT Operations Management (ITOM) capabilities that enable the performance, availability, scalability, reliability, and operational visibility of technology services. Working closely with senior engineers, specialists, and operational teams, the role ensures effective monitoring and management of infrastructure, applications, networks, and services through the configuration, optimisation, and integration of observability platforms and ITOM solutions.
    • The role is accountable for designing, maintaining, and improving telemetry pipelines, supporting cloud-native and containerised environments, and delivering AIOps capabilities that leverage analytics, automation, and intelligent event correlation to improve operational efficiency and reduce service disruption. This includes enhancing monitoring coverage, alert effectiveness, automated remediation, service mapping, discovery, and integration of observability capabilities using various software technologies.
    • The Observability Engineer collaborates across development, infrastructure, platform, and operations teams to support incident management, problem management, and continuous service improvement initiatives. The position requires strong technical expertise, analytical thinking, and a proactive approach to identifying risks, optimising system performance, and improving service reliability through observability data and operational insights.
    • The role plays a key part in advancing observability, AIOps, and ITOM maturity across the organisation, enabling data-driven decision-making, improving operational resilience, and fostering a culture of reliability, automation, and continuous improvement throughout the technology landscape.

    Key Responsibilities

    • Lead the delivery, implementation, support, configuration, and continuous improvement of Observability, AIOps, and ITOM platforms and capabilities to enhance service performance, reliability, and operational visibility.
    • Design, implement, and maintain scalable telemetry pipelines that collect, process, and analyse metrics, logs, traces, and events across infrastructure, applications, networks, and services.
    • Configure, optimise, and manage observability and monitoring solutions, including technologies such as Dynatrace, Prometheus, Grafana, ELK Stack, Datadog, OpenTelemetry, and related platforms.
    • Analyse telemetry and operational data to proactively identify trends, anomalies, performance bottlenecks, and reliability risks, leveraging insights to support AIOps-driven operations, service optimisation, and continuous improvement initiatives.
    • Define, implement, and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to measure service health, guide engineering priorities, and support reliability objectives.
    • Support and enhance ITOM capabilities, including Discovery, Service Mapping, Event Management, and automation initiatives to improve operational efficiency and reduce service disruption.
    • Collaborate with development, infrastructure, platform, and operations teams to improve monitoring coverage, alert quality, system resilience, and end-to-end operational visibility across cloud, hybrid, and containerised environments.
    • Participate in incident, problem, and change management activities by providing observability insights, performing root cause analysis, supporting post-incident reviews, and driving preventative improvements.
    • Develop and implement automation, scripts, and operational workflows that streamline observability processes, reduce manual effort, and improve service reliability and operational effectiveness.
    • Maintain observability standards, dashboards, documentation, and best practices while providing observability, AIOps, and ITOM expertise by promoting knowledge sharing, operational excellence, and continuous service improvement.

    Stakeholder Engagement:

    • Internal Stakeholders: Development, Infrastructure & Operation teams (across the Discovery Group)
    • External Stakeholders: Vendor support teams for observability tools, Third-party service providers for cloud and monitoring solutions

    Skills

    • Agent Deployment & configuration
    • Dynatrace Alert handling
    • Data Integrity
    • Data Collection
    • Monitoring
    • Dashboard Development / Dashboards
    • Infrastructure Automation
    • Incident and Problem Tracking
    • CI/CD
    • Infrastructure as Code
    • Amazon Web Services / AWS
    • Microsoft Azure
    • Kubernetes
    • Docker
    • Big Data

    Qualifications

    • Bachelor’s degree in Computer Science, Information Technology, or related field

    Certifications

    • Certified Cloud Practitioner or equivalent (AWS, Azure or GCP)
    • ITIL Foundation Certification
    • Certified Kubernetes Administrator (CKA)
    • Observability Platform Certification

    Work Experience

    • 5+ years' experience in IT Operations, Infrastructure, Monitoring, Observability, SRE, or a related technology field.
    • 3+ years' hands-on experience with observability platforms such as Dynatrace, Prometheus, Grafana, ELK, OpenTelemetry, Splunk, or similar.
    • Experience implementing and supporting Observability, AIOps, and ITOM capabilities, including Event Management, Discovery, and Service Mapping.
    • Experience monitoring enterprise infrastructure, applications, cloud platforms, and containerised environments.
    • Strong experience in incident management, problem management, root cause analysis, and service reliability improvement.
    • Practical experience with automation and scripting using Python, PowerShell, Bash, DQL, SQL or APIs.
    • Practical experience with networking technologies
    • Experience integrating observability solutions within ITSM platforms such as ServiceNow.
    • Proven ability to collaborate across development, infrastructure, and operations teams to improve service performance, availability, and operational efficiency.

    Preferred

    • Experience with Dynatrace and ServiceNow ITOM.
    • Knowledge of Kubernetes, cloud-native technologies (Azure, AWS, GCP), and SRE practices.

    Check how your CV aligns with this job

    Method of Application

    Interested and qualified? Go to Discovery Limited on careers.discovery.co.za to apply

    Build your CV for free. Download in different templates.

  • Get new ICT / Computer jobs like this on Telegram.Subscribe on Telegram
  • Send your application

    Back To Home

Career Advice

View All Career Advice
 

Subscribe to Job Alert

 

Join our happy subscribers

 
 
Send your application through

GmailGmail YahoomailYahoomail