About the role
Pinecone is the leading AI knowledge platform, trusted by over 10,000 customers and 1 million developers worldwide. The company provides vector database, retrieval, and marketplace solutions that power intelligent applications capable of understanding and acting on complex information. Pinecone's platform enables organizations to connect knowledge from diverse data sources directly to large language models and AI agents, making artificial intelligence genuinely knowledgeable rather than merely responsive.
The R&D team at Pinecone is building the next generation of knowledge retrieval systems designed specifically for the AI era. You'll join a group focused on creating infrastructure that allows customers to synthesize structured and unstructured data into high-quality, scalable retrieval capabilities for enterprise AI applications. This is a high-impact position with significant ownership across system design, performance optimization, and reliability.
What you'll do
- Design and build scalable platform components for semantic search, hybrid retrieval, metadata-aware search, and LLM-integrated generation
- Create optimized indexing pipelines that process both structured databases and unstructured data at scale
- Develop backend services supporting retrieval orchestration, knowledge graph construction, and multi-tier search strategies
- Build and refine evaluation and observability frameworks to continuously improve retrieval quality and system insights
- Design intuitive APIs for both human developers and autonomous agents to consume retrieval capabilities
- Optimize end-to-end latency, throughput, and cost across large-scale inference and retrieval workloads
- Establish technical direction around system reliability and security across the platform
What you'll bring
- Six or more years shipping production backend systems handling high throughput and low latency at scale, with proven architectural thinking beyond writing individual features
- Hands-on experience building high-throughput data pipelines that work with messy unstructured data and rigid structured schemas
- Direct experience or deep theoretical knowledge in semantic search, vector databases, hybrid retrieval approaches, or traditional search platforms like Elasticsearch and OpenSearch
- Understanding of Retrieval-Augmented Generation patterns, embedding workflows, query planning, and how metadata filtering impacts LLM reasoning
- Expert proficiency in at least one systems language such as Go, Rust, C++, Java, or Python
- Practical experience with Kubernetes, cloud-native architectures, observability systems, and infrastructure-as-code tools like Terraform or Pulumi
- User-focused product thinking that extends to designing clean APIs for both developers and agentic systems
- Comfort operating in high-growth environments where you own problems end-to-end rather than executing isolated tickets
Nice to have
- Experience designing and operating multi-tenant SaaS systems
- Hands-on work with retrieval evaluation frameworks and measuring search quality
- Background with query planning, agentic reasoning loops, or teaching systems to decompose complex problems
What they offer
- Medical, dental, vision, and mental health coverage
- 401(k) plan with company matching
- Equity awards
- Flexible time off and paid parental leave
- Annual company retreat
- Remote work equipment stipend
Pay, location & hours
$190K–270K / year. Based in New York City (Hybrid).
About Pinecone
5 open roles in this building · Company page → · See it on the map
Databricks · New York