Location
Central London
Hours
Full Time
Salary
Competitive salary
About the Role
Feeld is building a world where everyone is more intimately connected to each other and themselves. We are an inclusive, human-centred product with a distributed engineering team. We are hiring a Staff Reliability Engineer (Full Stack) to raise the reliability and operability of our production systems across backend and mobile integration. This hands-on individual contributor role focuses on improving how we detect, respond to, and prevent production incidents, while strengthening engineering practices such as documentation, runbooks, and quality bars to help teams move quickly without compromising stability. Reporting to the Head of Platform Engineering and working within the Platform team, you will collaborate closely with backend engineers and engineering leadership. You will influence multiple squads to improve production ownership, reliability, and backend-to-mobile integration patterns. This is a remote-first, async-friendly, high-trust environment where clear written communication and consistent operational practices are essential.
What Success Looks Like
Within your first year, you will have:
- Reduced incident frequency and/or impact through concrete reliability improvements such as better alerting, safer deploy patterns, guardrails, and playbooks.
- Made incident response more effective with clear ownership, faster mean time to recovery (MTTR), and improved post-incident follow-through.
- Delivered improvements to backend/mobile integration that reduce breakages and production risk.
- Established or materially improved documentation and operational standards that are consistently used by engineers.
Key Responsibilities
- Own reliability outcomes across critical backend services and their integration with mobile clients (React Native).
- Lead technical problem-solving during incidents by coordinating response, diagnosing root causes, communicating status, and driving resolution.
- Build and evolve monitoring and observability tools including dashboards, alerts, tracing, and logging to enable fast detection and diagnosis.
- Drive blameless post-incident reviews and ensure learnings translate into durable fixes such as technical changes, runbooks, automation, and process updates.
- Improve engineering safety and quality through guardrails, safer migrations, feature-flag practices, rollout strategies, and resilience patterns.
- Partner with product, design, QA, and engineering teams early to align delivery plans with operational risk and reliability needs.
- Strengthen documentation and onboarding materials including architecture notes, runbooks, service ownership documents, and “how we work” guides.
- Mentor engineers through pairing, code reviews, incident shadowing, and pragmatic coaching on production ownership.
Experience
- Significant experience building and operating production backend systems at scale, including debugging distributed systems and performance issues.
- Strong TypeScript/Node.js (or equivalent) backend experience with comfort working across services and APIs.
- Proven incident response leadership including on-call participation, triage, mitigation, and root-cause analysis with follow-through.
- Solid observability skills with practical experience in logging, metrics, tracing, and converting signals into actionable alerts and dashboards.
- Experience collaborating with mobile teams and understanding mobile-backend integration concerns such as API compatibility, releases, and feature flags.
- Demonstrated Staff-level individual contributor leadership through design reviews, technical direction, documentation, and cross-team alignment.
About You
- You are proactive and pragmatic with a strong sense of ownership and a passion for improving system reliability.
- You communicate clearly and effectively in writing and enjoy working in a remote-first, asynchronous environment.
- You are collaborative and able to influence cross-functional teams through technical expertise and mentorship.
- You thrive in a high-growth environment where prioritization and pragmatic trade-offs are essential.
Qualifications
Nice-to-haves include:
- React Native experience and/or strong understanding of mobile architecture patterns and release constraints.
- AWS or similar cloud experience, familiarity with infrastructure-as-code, CI/CD, and production tooling.
- Experience designing reliability programs such as SLOs, error budgets, and incident processes.
- Experience with PostgreSQL, Redis, and performance tuning in high-traffic systems.

