Skip to Main Content
Location icon
Central London

AI Inference Engineer

Fuse Energy
Office & Professional
Office & Professional
Negotiable
Company logo image
Description

Location
Central London

Hours
Full Time

Salary
Competitive, commensurate with experience

About the Role
Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast. Combining first-principles thinking with cutting-edge technology, Fuse is building a radically better energy system. Having raised $210M from top-tier investors, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI.

We are seeking a Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. This role owns the inference serving layer above GPU/CUDA engineering, focusing on how models are served, scaled, and delivered against committed performance targets.

The successful candidate will define Fuse's inference serving strategy and architecture from first principles, design and build the serving stack including request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads. You will own model-level optimisation strategies such as quantisation, distillation, and speculative decoding, partnering closely with CUDA/GPU engineers.

You will make core software architecture decisions on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents), translate throughput, latency, and uptime commitments into technical specifications and capacity plans, and act as the direct technical owner of inference performance and reliability. Collaboration with CUDA and GPU engineering teams to integrate custom kernels and hardware performance work into the serving layer is essential. You will also set standards, tooling, and benchmarks as this function grows.

Requirements

Experience
4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience.
Deep, hands-on experience with inference serving frameworks and optimisation techniques including batching, KV-cache management, quantisation, and speculative decoding.
Strong systems thinking with the ability to reason about the full path from incoming request to served response across a large cluster.
Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system.
Proven track record of making high-stakes architecture decisions and owning the outcomes.
Ability to operate without a playbook in a founding role shaping a new, early-stage function.

About you
Innovative and self-driven with a passion for building scalable AI infrastructure.
Collaborative mindset with strong communication skills to work across engineering teams.
Comfortable with ambiguity and excited by the challenge of defining new architecture and systems.
Interest in renewable energy, sustainability, or energy markets is a plus.

Qualifications
Experience with Triton or custom ML inference/training frameworks is desirable.
Experience with autoscaling or capacity planning for large-scale inference workloads.
Exposure to multi-tenant serving or SLA-driven infrastructure.
Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system is advantageous.
Familiarity with Kubernetes or Slurm for cluster orchestration is a plus.

Expiry date: 21/08/2026
AI Inference Engineer
Company:
Fuse Energy
Job Type:
Full-time
Location:
Central London