Location
London - Stockley Park with hybrid working options where applicable.
Hours
Permanent Position, Monday to Friday, 9am to 5pm. Occasional unsociable hours including weekends or on-call rotations may be required to support live operations and critical systems. Occasional travel may be required depending on project and client needs.
Salary
Negotiable
About the Role
IMG is seeking a Site Reliability Engineer to design, build, operate, and continuously improve resilient, secure, and highly available platforms supporting digital, cloud, and broadcast-adjacent services. This role combines strong infrastructure and software engineering skills with an operational mindset to embed reliability engineering practices across live, business-critical environments. You will improve service reliability, observability, incident response, automation, and disaster recovery readiness while collaborating closely with engineering, operations, and project stakeholders. Key responsibilities include designing and maintaining scalable infrastructure across on-premises and cloud environments, enhancing observability and automation, defining SLIs and SLOs, leading incident response and root cause analysis, supporting disaster recovery planning, enforcing security best practices, optimizing system capacity and cost, producing technical documentation, and supporting live event workflows where reliability and rapid response are essential. This role plays a crucial part in driving higher automation, clearer operational ownership, and stronger resilience for live and client-facing workflows through tested failover and recovery approaches.
Experience
- Proven experience as a Site Reliability Engineer, DevOps Engineer, Platform Engineer, or similar role.
- Strong knowledge of Linux and operating system fundamentals.
- Hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud.
- Experience with containerisation and orchestration technologies like Docker and Kubernetes.
- Strong experience with CI/CD tooling and modern software delivery practices.
- Hands-on experience with Infrastructure as Code tools such as Terraform or CloudFormation.
- Experience with monitoring, logging, alerting tooling, and designing actionable observability solutions.
- Solid understanding of networking, security, system architecture, and distributed systems principles.
- Strong scripting or programming skills in Python, Bash, or similar languages.
- Experience working in high-availability, live production, or business-critical operational environments.
- Strong troubleshooting skills, calm decision-making under pressure, and a continuous improvement mindset.
- Excellent communication and collaboration skills with technical and non-technical stakeholders.
About you
Proactive and ownership driven, methodical, analytical, and detail oriented. Comfortable operating in fast-moving, high-pressure environments. Pragmatic in balancing engineering excellence with operational needs. Collaborative, service oriented, and committed to raising reliability standards across teams. Success at IMG is driven by core competencies including business acumen, operational excellence, innovation mindset, and leadership & collaboration.
Qualifications
While specific formal qualifications are not detailed, the role requires strong technical expertise and relevant professional experience in site reliability, cloud infrastructure, and software engineering disciplines. Desirable experience includes supporting media, broadcast, streaming, or live event platforms; familiarity with incident management and postmortem practices; resilience engineering and disaster recovery testing; exposure to event-driven or low-latency systems; and understanding compliance and operational risk management in client-facing environments.
IMG




















