About the role
Together AI is a leading AI infrastructure platform serving hundreds of enterprises with compute and model capabilities at scale. As the company expands its cluster footprint, this role will establish and manage the technical qualification process that ensures every new compute provider meets Together's standards before taking on customer workloads.
You'll own the end-to-end qualification pipeline for new compute capacity. That means coordinating evaluations across multiple providers simultaneously, partnering with infrastructure, network, data center, and SRE teams through validation phases, and ultimately delivering clear go or no-go recommendations to leadership. You'll review provider specifications and test results for accuracy, conduct your own first-pass analysis to spot inconsistencies and performance claim issues, and maintain the standards and templates that define what passes across compute, networking, storage, power, cooling, and operations.
The work requires you to be both technical and organized. You'll need to read detailed spec sheets and hardware performance data, identify potential bottlenecks in cluster architecture before they affect customer training or inference, and flag gaps that need engineering diligence. You'll also build an auditable record of every evaluation outcome to inform sourcing decisions and help the function scale as Together grows.
What you'll bring
- 5+ years in technical program management, infrastructure program management, or technical operations, preferably with hardware, data center, or large-scale compute experience.
- Track record running multiple complex cross-functional workstreams on deadline with strong stakeholder management and organizational discipline.
- Working technical knowledge of data center infrastructure: server and GPU hardware, high-performance networking like InfiniBand or Ethernet fabrics, storage, power and cooling fundamentals. Enough depth to question what you read in a spec.
- Hands-on comfort with data analysis. You can write Python or SQL scripts to independently compare, validate, and analyze provider specifications and test results.
- Clear written and verbal communication. You translate dense technical detail into actionable recommendations for both engineers and executives.
- Willingness to travel to provider sites and data centers as needed. This is not an engineering management role.
Nice to have
- Experience qualifying, commissioning, or accepting GPU clusters or HPC infrastructure against performance and reliability standards.
- Familiarity with AI training and inference infrastructure, cluster topologies, bring-up, and acceptance testing.
- Background working directly with hardware vendors, colocation providers, or cloud capacity providers.
What we offer
- US base salary of 200-250K plus equity and benefits, adjusted by location, level, and experience.
- Health insurance and flexibility around remote work. The primary office is in San Francisco or New York City.
Pay, location & hours
Salary not listed. Based in San Francisco or NYC.
About Together AI

3 open roles in this building · Company page → · See it on the map