TopWeb3JobsTopWeb3Jobs
AI & ML district

Senior/Staff Software Engineer, Search & Retrieval Infrastructure

Pinecone · New York City (Hybrid) · Full-time
◎ AI & ML🏢 Hybrid
$190K–270K / year

About the role

Pinecone is the leading AI knowledge platform, trusted by over 10,000 customers and 1 million developers worldwide. The company provides vector database, retrieval, and marketplace solutions that power intelligent applications capable of understanding and acting on complex information. Pinecone's platform enables organizations to connect knowledge from diverse data sources directly to large language models and AI agents, making artificial intelligence genuinely knowledgeable rather than merely responsive.

The R&D team at Pinecone is building the next generation of knowledge retrieval systems designed specifically for the AI era. You'll join a group focused on creating infrastructure that allows customers to synthesize structured and unstructured data into high-quality, scalable retrieval capabilities for enterprise AI applications. This is a high-impact position with significant ownership across system design, performance optimization, and reliability.

What you'll do

  • Design and build scalable platform components for semantic search, hybrid retrieval, metadata-aware search, and LLM-integrated generation
  • Create optimized indexing pipelines that process both structured databases and unstructured data at scale
  • Develop backend services supporting retrieval orchestration, knowledge graph construction, and multi-tier search strategies
  • Build and refine evaluation and observability frameworks to continuously improve retrieval quality and system insights
  • Design intuitive APIs for both human developers and autonomous agents to consume retrieval capabilities
  • Optimize end-to-end latency, throughput, and cost across large-scale inference and retrieval workloads
  • Establish technical direction around system reliability and security across the platform

What you'll bring

  • Six or more years shipping production backend systems handling high throughput and low latency at scale, with proven architectural thinking beyond writing individual features
  • Hands-on experience building high-throughput data pipelines that work with messy unstructured data and rigid structured schemas
  • Direct experience or deep theoretical knowledge in semantic search, vector databases, hybrid retrieval approaches, or traditional search platforms like Elasticsearch and OpenSearch
  • Understanding of Retrieval-Augmented Generation patterns, embedding workflows, query planning, and how metadata filtering impacts LLM reasoning
  • Expert proficiency in at least one systems language such as Go, Rust, C++, Java, or Python
  • Practical experience with Kubernetes, cloud-native architectures, observability systems, and infrastructure-as-code tools like Terraform or Pulumi
  • User-focused product thinking that extends to designing clean APIs for both developers and agentic systems
  • Comfort operating in high-growth environments where you own problems end-to-end rather than executing isolated tickets

Nice to have

  • Experience designing and operating multi-tenant SaaS systems
  • Hands-on work with retrieval evaluation frameworks and measuring search quality
  • Background with query planning, agentic reasoning loops, or teaching systems to decompose complex problems

What they offer

  • Medical, dental, vision, and mental health coverage
  • 401(k) plan with company matching
  • Equity awards
  • Flexible time off and paid parental leave
  • Annual company retreat
  • Remote work equipment stipend

Pay, location & hours

$190K–270K / year. Based in New York City (Hybrid).

About Pinecone

5 open roles in this building · Company page → · See it on the map

Apply ↗

More roles to explore

$190K–270K / year
Pinecone
Apply ↗

☆ Save this job

We'll e-mail you this role so you can come back to it. No account needed.

Report this job

Reports go to the TopWeb3Jobs team. Scam reports are checked first.