About the role
Scale AI is building the infrastructure and data platforms that power AI development at massive scale. The company works at the intersection of AI research, engineering, and operations, helping teams train and deploy frontier models efficiently.
The AI Research & Engineering team is looking for a Research Engineer to work on the distributed systems that power reinforcement learning at scale. This is infrastructure work at the frontier: you'll own pieces of the system that trains Claude and similar models across thousands of accelerators and hosts. The role spans scheduling, data movement, fault tolerance, storage, networking, and observability. You'll move between layers as needed, solving whatever is currently the bottleneck.
What you'll do
- Design and operate distributed systems running RL training, sampling, and environment execution concurrently across heterogeneous clusters
- Identify system bottlenecks in scheduling, data movement, storage, networking, or coordination and build solutions that remove them
- Build fault tolerance throughout the stack: failure detection, isolation, and recovery that keep long-running jobs progressing without manual intervention
- Implement resource management and autoscaling so compute allocation shifts as a run's needs change
- Create observability systems that surface what's happening in a run, why throughput dropped, or why results diverged from expected
- Design automation and operational interfaces that let engineers and tools safely diagnose and adjust runs
- Collaborate with researchers and performance engineers to ensure system changes don't introduce nondeterminism or compromise training correctness
- Conduct incident reviews and write design documents that prevent classes of failure from recurring
What you'll bring
- Strong software engineering in Python and at least one systems language like Rust, C++, or Go
- Proven experience designing, building, and operating large-scale distributed systems in production
- Deep knowledge of distributed systems fundamentals: consistency, coordination, consensus, failure modes, and recovery strategies
- Ability to reason quantitatively about throughput, latency, and costs across compute, memory, storage, and networking
- Experience debugging complex failures across many hosts and services, including ones you can't reproduce locally
- Clear written communication, especially design documents and incident writeups
- Generalist mindset: comfort reasoning from first principles about unfamiliar systems and picking the most impactful problem over the most comfortable one
Nice to have
- Experience running ML training or inference infrastructure at scale
- Hands-on work with Kubernetes, container orchestration, or sandboxed code execution
- Background with schedulers, autoscalers, or resource management systems
- Knowledge of high-performance networking, RDMA, or collective communication libraries
- Experience with async Python frameworks like Trio or asyncio
- Exposure to reinforcement learning or large language model training
What we offer
- Annual compensation of 500,000 to 850,000 USD
- Full-time position with flexibility across San Francisco, New York City, and Seattle offices
Pay, location & hours
Salary not listed. Based in San Francisco, CA, New York City, NY, Seattle, WA.
About Anthropic
5 open roles in this building · Company page → · See it on the map
Together AI · San Francisco