About the role
Tether is a blockchain infrastructure company that powers global financial innovation through reserve-backed stablecoins, energy solutions, AI infrastructure, and educational technology. Operating across more than 180 countries, Tether serves hundreds of millions of users and processes trillions of dollars in transactions annually. The company has established itself as a trusted bridge between traditional finance and decentralized systems, enabling seamless digital asset transfers across blockchain networks.
Tether Data division is at the forefront of AI infrastructure development, building the next generation of GPU compute platforms that reduce operational costs while expanding access to machine learning capabilities worldwide. The Data team operates Cosmic AC, a sophisticated orchestration system that manages GPU resources, inference endpoints, and computational workloads at scale across global infrastructure.
You will join as Technical Lead for GPU Infrastructure in a hands-on leadership role overseeing the evolution of Cosmic AC from a managed-cluster platform to a full bare-metal GPU stack. You will lead approximately twelve distributed engineers across backend, frontend, DevOps, QA and documentation roles spanning Europe and India. This is a fixed-scope role with clear delivery targets in the first six months, balancing direct technical contributions with team leadership and vendor partnerships.
What you'll do
- Design and maintain platform architecture end to end, from high-level strategy through detailed implementation, managing proposals and design reviews with a live baseline architecture document
- Lead a geographically distributed engineering team, providing technical direction, code and design review, release governance, individual performance feedback, and career development guidance
- Build and operate a managed Slurm scheduling layer for research workloads, handling controller setup, accounting systems, partition configuration, node onboarding, CUDA driver management, health detection, and autohealing procedures
- Architect Kubernetes control plane deployment on bare-metal infrastructure, including cluster bootstrap, NVIDIA GPU integration, virtual machine-based GPU isolation using KubeVirt and VFIO, and day-two operations like upgrades and recovery
- Design managed inference serving architecture supporting multi-GPU and multi-node parallel processing, autoscaling policies, request routing, and endpoint reliability for production workloads
- Establish observability systems across all layers including metrics, logging, alerting, SLOs, incident response processes, and sustainable on-call rotation practices
- Serve as primary technical contact with infrastructure partners and vendors, translating requirements into specifications, managing escalations, and informing capacity planning decisions
What you'll bring
- Eight or more years of hands-on infrastructure engineering experience, with at least three years leading teams that build and operate platforms depended on by other engineering organizations
- Demonstrated production experience running Slurm at scale: direct experience with slurmctld and slurmdbd, partition and QoS configuration, accounting systems, job control scripting, node health monitoring, and managing upgrades with live workloads
- Deep knowledge of Kubernetes architecture, cluster operations, and the NVIDIA GPU Operator ecosystem
- Strong capability with backend technologies including Node.js, or demonstrated ability to lead teams using these tools
- Excellent written and verbal English communication skills, essential for coordinating across distributed teams and external partners
- Bachelor's or Master's degree in computer science or engineering, or equivalent professional infrastructure engineering experience
- Comfort with infrastructure-as-code approaches and vendor relationship management
Nice to have
- Experience with confidential computing technologies and secure workload isolation
- Familiarity with KubeVirt, VFIO, or other GPU virtualization approaches
- Knowledge of AI and machine learning workflow requirements
- Previous experience in blockchain or fintech infrastructure environments
Pay, location & hours
Salary not listed. Fully remote, open to applicants in job.
About Tether
19 open roles in this building · Company page → · See it on the map
Gemini · Remote