Engineering Manager - Inference Job at Perplexity (San Francisco)

Engineering Manager - Inference

Perplexity

Location:
United States , San Francisco

Category:
IT - Software Development

Contract Type:
Not provided

Salary:

300000.00 - 385000.00 USD / Year

Save Job

Apply Position

Job Description:

We are looking for an Inference Engineering Manager to lead our AI Inference team. This is a unique opportunity to build and scale the infrastructure that powers Perplexity's products and APIs, serving millions of users with state-of-the-art AI capabilities. You will own the technical direction and execution of our inference systems while building and leading a world-class team of inference engineers. Our current stack includes Python, PyTorch, Rust, C++, and Kubernetes. You will help architect and scale the large-scale deployment of machine learning models behind Perplexity's Comet, Sonar, Search, Deep Research products.

Job Responsibility:

Lead and grow a high-performing team of AI inference engineers
Develop APIs for AI inference used by both internal and external customers
Architect and scale our inference infrastructure for reliability and efficiency
Benchmark and eliminate bottlenecks throughout our inference stack
Drive large sparse/MoE model inference at rack scale, including sharding strategies for massive models
Push the frontier with building inference systems to support sparse attention, disaggregated pre-fill/decoding serving, etc.
Improve the reliability and observability of our systems and lead incident response
Own technical decisions around batching, throughput, latency, and GPU utilization
Partner with ML research teams on model optimization and deployment
Recruit, mentor, and develop engineering talent
Establish team processes, engineering standards, and operational excellence

Requirements:

5+ years of engineering experience with 2+ years in a technical leadership or management role
Deep experience with ML systems and inference frameworks (PyTorch, TensorFlow, ONNX, TensorRT, vLLM)
Strong understanding of LLM architecture: Multi-Head Attention, Multi/Grouped-Query Attention, and common layers
Experience with inference optimizations: batching, quantization, kernel fusion, FlashAttention
Familiarity with GPU characteristics, roofline models, and performance analysis
Experience deploying reliable, distributed, real-time systems at scale
Track record of building and leading high-performing engineering teams
Experience with parallelism strategies: tensor parallelism, pipeline parallelism, expert parallelism
Strong technical communication and cross-functional collaboration skills

Nice to have:

Experience with CUDA, Triton, or custom kernel development
Background in training infrastructure and RL workloads
Experience with Kubernetes and container orchestration at scale
Published work or contributions to inference optimization research

What we offer:

Equity
Health
Dental
Vision
Retirement
Fitness
Commuter and dependent care accounts

Additional Information:

Job Posted:
February 21, 2026

Employment Type:

Fulltime

Perplexity - All Job Offers

Job Link Share:

Engineering Manager - Inference

Perplexity

Location:
United States , San Francisco

Category:
IT - Software Development

Contract Type:
Not provided

Salary:

Job Description:

Job Responsibility:

Requirements:

Nice to have:

Additional Information:

Job Posted:
February 21, 2026

Looking for more opportunities? Search for other job offers that match your skills and interests.

Similar Jobs for Engineering Manager - Inference

Engineering Manager - Machine Learning

Engineering Manager - Machine Learning Infrastructure

Senior Machine Learning Engineering Manager, Gen AI

Staff Product Manager, Managed Inference

Staff Product Manager, Managed Inference

Director of Engineering, Cloud Availability

Director of Engineering, Cloud Availability

Engineering Manager, GenAI Platform

Engineering Manager - Inference

Perplexity

Location:United States , San Francisco

Category:IT - Software Development

Contract Type:Not provided

Salary:

Job Description:

Job Responsibility:

Requirements:

Nice to have:

Additional Information:

Job Posted:February 21, 2026

Looking for more opportunities? Search for other job offers that match your skills and interests.

Similar Jobs for Engineering Manager - Inference

Engineering Manager - Machine Learning

Engineering Manager - Machine Learning Infrastructure

Senior Machine Learning Engineering Manager, Gen AI

Staff Product Manager, Managed Inference

Staff Product Manager, Managed Inference

Director of Engineering, Cloud Availability

Director of Engineering, Cloud Availability

Engineering Manager, GenAI Platform

Location:
United States , San Francisco

Category:
IT - Software Development

Contract Type:
Not provided

Job Posted:
February 21, 2026