About the role
Scale AI is a leading data infrastructure company supporting the world's most advanced AI systems. For a decade, they've worked with major institutions like Meta, Ernst & Young, and U.S. government agencies to develop the high-quality datasets and technologies powering frontier AI applications. Their Enterprise Engineering team owns the complete lifecycle of Scale's products in live customer environments, from initial deployment through ongoing operational excellence.
This is a chance to lead infrastructure strategy and execution for a fast-growing organization. As Engineering Manager for Infrastructure, you'll oversee a team building the backend systems that power AI agents, evaluation tooling, and enterprise SaaS products. You'll define the infrastructure roadmap, ensure production reliability at scale, and work cross-functionally with product, security, and GTM to translate customer deployment challenges into platform improvements.
Your responsibilities will span
- Leading and mentoring an infrastructure engineering team while driving technical delivery and shaping team culture
- Defining the infrastructure roadmap in alignment with business priorities
- Designing and implementing scalable, secure, reliable systems with clear SLAs and SLOs for uptime, performance, and developer experience
- Building and optimizing backend services for AI-driven applications, focusing on agents, evaluation tooling, and automation
- Collaborating with product, security, and leadership to align on goals
- Staying close to customer deployment challenges to build sustainable playbooks and automation
You'll need at least 5 years of infrastructure experience, including 2+ years managing platform or infrastructure teams. Hands-on proficiency with cloud platforms (AWS, GCP, Azure, or OCI) and Kubernetes on on-premises infrastructure is essential. You should have strong expertise in CI/CD systems like CircleCI or GitHub Actions, infrastructure-as-code tools like Terraform, and deep knowledge of monitoring, alerting, and incident response. Real-world network engineering experience and familiarity with modern developer platforms to improve engineering velocity round out the core requirements.
This is a full-time, remote position based in London, UK.
Pay, location & hours
Salary not listed. Based in London, UK.
About Anthropic
6 open roles in this building · Company page → · See it on the map