About the role
Scale AI is building the infrastructure and data systems that power AI applications for the world's most critical decisions. They partner with industry leaders and government agencies to develop reliable AI systems, and they're growing their engineering teams to accelerate this work.
The Applications Platform Engineering team at Scale is looking for an Infrastructure Software Engineer to help design, deploy, and operate their platform across multiple cloud providers and on-premises environments. You'll own both the deployment layer and observability tooling that powers Scale's infrastructure, working closely with internal teams and forward-deployed engineers to understand real-world needs and shape the technical roadmap.
What you'll do
- Build and expand deployment pipelines that work consistently across AWS, Azure, GCP, OCI and on-premises infrastructure, ensuring security and reproducibility at every step
- Develop comprehensive observability and monitoring for distributed platform infrastructure so internal and customer teams can operate deployments efficiently
- Work directly with internal engineering teams to understand how they use the platform, troubleshoot issues, and build tooling that solves their problems
- Own infrastructure projects end-to-end, from architectural design through implementation and production deployment, in cross-functional settings
- Lead incident response and production issue resolution, performing root cause analysis and implementing lasting preventive fixes
- Help evolve the platform's deployment and observability roadmap, balancing immediate operational needs with long-term architectural improvements
What you'll bring
- 5+ years of hands-on experience building and deploying enterprise and public sector infrastructure across multiple cloud platforms and on-premises
- Deep expertise in infrastructure architecture, networking, VPNs, load balancers, and firewall design
- Strong proficiency with Kubernetes, Terraform, Docker and other infrastructure-as-code and container orchestration tools
- Solid understanding of CI/CD pipelines and software delivery using GitHub Actions or CircleCI
- Experience with GitOps patterns and a demonstrated track record of ensuring consistent, repeatable deployments
- Strong debugging skills and comfort navigating the tradeoffs between performance and security in production systems
- Ability to work comfortably with ambiguity and switch between reactive incident work and proactive product development
Nice to have
- Experience as a founder or early-stage engineer at an infrastructure-focused startup, owning a product end-to-end
- Background operating secure workloads in multi-tenant or untrusted environments like FaaS platforms or CI sandboxes
- Open-source contributions to systems or developer tools projects
- History of on-call incident response responsibilities
Position is full-time and based in London, UK.
Pay, location & hours
Salary not listed. Based in London, UK.
About Anthropic
6 open roles in this building · Company page → · See it on the map