TopWeb3JobsTopWeb3Jobs
Engineering district · Plot MB-B2-41

Principal Software Development Engineer (Microservices)

zscaler ✓ Verified · Santa Clara, California, USA · Full-time
⌘ Engineering📍 On-site
Salary not listed

About the role

Zscaler, a NASDAQ-listed security platform company, operates the world's largest in-line cloud security infrastructure, protecting thousands of organizations through its Zero Trust Exchange. The platform spans over 200 public data centers globally and thousands of private edge locations, delivering SASE-based security at massive scale. As enterprises increasingly adopt AI-native architectures, Zscaler is at the forefront of security transformation in the AI era.

The Zero Trust Exchange platform serves as the technical backbone for modern cloud security. Zscaler's engineering teams are modernizing this infrastructure to handle unprecedented scale, complexity, and performance demands. The company views the intersection of human expertise and AI capabilities as essential to solving today's most demanding security challenges.

We are seeking a Principal Software Development Engineer specializing in microservices to lead the evolution of our Zscaler Internet Access product's control plane. You will own the technical direction of a critical platform serving thousands of customers globally, working from our Santa Clara, California office on a hybrid basis requiring at least three days per week on-site. You will report to the Senior Director of Software Development Engineering within the Engineering ZIA Core department.

What you'll do

  • Shape the technical roadmap to modernize the ZIA control plane through API-first, domain-aligned microservices architecture with explicit service ownership and multi-region deployment patterns including active-active configurations and disaster recovery strategies.
  • Design and operate event-driven services across the full stack with emphasis on low latency, implementing resilience patterns, progressive delivery techniques such as blue-green and canary deployments, GitOps automation with Argo CD or Flux, and secure rollout and rollback capabilities.
  • Champion engineering standards across the organization through code review processes, comprehensive testing strategies, shared patterns, API governance including versioning and contract management, event schema governance with registries, and deep collaboration with Product, SRE, Security, Data, and Architecture teams.
  • Establish and maintain service reliability standards, defining and tracking SLIs and SLOs for p95 and p99 latencies and availability targets, implementing SLO-driven alerting and error budgets, managing incidents and postmortems, and deploying advanced observability using OpenTelemetry, Jaeger or Tempo, and structured logging approaches.
  • Integrate security throughout microservices design using OAuth2 and OIDC protocols, mutual TLS authentication, service-to-service authorization, and secrets management via KMS or Vault, alongside supply chain security practices including SBOM generation and SLSA compliance.

What you'll bring

  • Twelve or more years of software development experience with at least four years actively building and operating production microservices at significant scale, with demonstrated proficiency in Go programming.
  • Mastery of distributed systems and event messaging patterns, preferably with Kafka but also comfortable with SQS or RabbitMQ, along with deep understanding of CQRS and Event Sourcing where appropriate.
  • Production-grade Kubernetes expertise in AWS environments using EKS, strong Docker fundamentals, proficiency with Infrastructure-as-Code practices, and hands-on experience implementing GitOps workflows.
  • Advanced database design and optimization across SQL systems such as PostgreSQL, MySQL, Amazon RDS, and Aurora alongside NoSQL platforms including Cassandra, MongoDB, and DynamoDB, with capability to manage multi-region tradeoffs and optimize costs across AWS.
  • Comprehensive reliability and security background encompassing SLI and SLO definition and tracking, error budget management, metrics collection using Prometheus, distributed tracing with OpenTelemetry, alert strategy design, incident response, postmortem analysis, service mesh technologies such as Istio or Linkerd, mTLS implementation, OAuth2 and OIDC protocols, and secrets management systems.
  • Foundational understanding of AI and machine learning technologies with demonstrated ability to leverage, secure, or position AI-driven solutions within your technical domain to deliver measurable outcomes.

Nice to have

  • Experience with chaos engineering, advanced load testing, fault injection techniques, and disaster recovery execution paired with multi-region architecture design expertise and leadership of migration efforts using strangler patterns and domain decomposition strategies.
  • Proficiency with event contract governance using Avro or Protobuf with schema registries, dead letter queue processing strategies, and data replay tooling.
  • Open source contributions, published technical writing, conference presentations, or patent record demonstrating thought leadership.

What they offer

  • Competitive base salary commensurate with role, level, and relevant experience, plus commission, bonus, and equity opportunities where applicable.
  • Comprehensive benefits package supporting professional and personal well-being.

Pay, location & hours

Salary not listed. Based in Santa Clara, California, USA.

About zscaler

11 open roles in this building · Company page → · See it on the map

Apply ↗

More roles to explore

Salary not listed
zscaler
Apply ↗

☆ Save this job

We'll e-mail you this role so you can come back to it. No account needed.

Report this job

Reports go to the TopWeb3Jobs team. Scam reports are checked first.