招聘
Required Skills
Python
PyTorch
Rust
C++
Leadership
System Design
ABOUT THE ROLE:
We are looking for an Inference Engineering Manager to lead our AI Inference team. This is a unique opportunity to build and scale the infrastructure that powers Perplexity's products and APIs, serving millions of users with state-of-the-art AI capabilities.
You will own the technical direction and execution of our inference systems while building and leading a world-class team of inference engineers. Our current stack includes Python, Py Torch, Rust, C++, and Kubernetes. You will help architect and scale the large-scale deployment of machine learning models behind Perplexity's Comet, Sonar, Search, Deep Research products.
WHY PERPLEXITY?
-
Build SOTA systems that are the fastest in the industry with cutting-edge technology
-
High-impact work on a smaller team with significant ownership and autonomy
-
Opportunity to build 0-to-1 infrastructure from scratch rather than maintaining legacy systems
-
Work on the full spectrum: reducing cost, scaling traffic, and pushing the boundaries of inference
-
Direct influence on technical roadmap and team culture at a rapidly growing company
RESPONSIBILITIES:
-
Lead and grow a high-performing team of AI inference engineers
-
Develop APIs for AI inference used by both internal and external customers
-
Architect and scale our inference infrastructure for reliability and efficiency
-
Benchmark and eliminate bottlenecks throughout our inference stack
-
Drive large sparse/MoE model inference at rack scale, including sharding strategies for massive models
-
Push the frontier with building inference systems to support sparse attention, disaggregated pre-fill/decoding serving, etc.
-
Improve the reliability and observability of our systems and lead incident response
-
Own technical decisions around batching, throughput, latency, and GPU utilization
-
Partner with ML research teams on model optimization and deployment
-
Recruit, mentor, and develop engineering talent
-
Establish team processes, engineering standards, and operational excellence
QUALIFICATIONS:
-
5+ years of engineering experience with 2+ years in a technical leadership or management role
-
Deep experience with ML systems and inference frameworks (Py Torch, Tensor Flow, ONNX, TensorRT, vLLM)
-
Strong understanding of LLM architecture: Multi-Head Attention, Multi/Grouped-Query Attention, and common layers
-
Experience with inference optimizations: batching, quantization, kernel fusion, Flash Attention
-
Familiarity with GPU characteristics, roofline models, and performance analysis
-
Experience deploying reliable, distributed, real-time systems at scale
-
Track record of building and leading high-performing engineering teams
-
Experience with parallelism strategies: tensor parallelism, pipeline parallelism, expert parallelism
-
Strong technical communication and cross-functional collaboration skills
NICE TO HAVE:
-
Experience with CUDA, Triton, or custom kernel development
-
Background in training infrastructure and RL workloads
-
Experience with Kubernetes and container orchestration at scale
-
Published work or contributions to inference optimization research
Total Views
0
Apply Clicks
0
Mock Applicants
0
Scraps
0
Similar Jobs

Senior Software Engineer
Neon · New York City

Staff Software Engineer (Backend)
Faraday Future · Gardena, California, United States

Staff Software Engineer, Backend (Affiliate Platform)
Faraday Future · El Segundo, California, United States

Data Engineer / Sr Data Engineer
Hartford · 4 Locations

Sr. Staff Data Engineer (Tech Lead) - Hybrid
Hartford · 4 Locations
About Perplexity AI

Perplexity AI
Series BPerplexity AI, Inc., or simply Perplexity, is an American privately held software company offering a web search engine that processes user queries and synthesizes responses.
51-200
Employees
San Francisco
Headquarters
$1B
Valuation
Reviews
4.0
1 reviews
Work Life Balance
3.0
Compensation
3.0
Culture
3.0
Career
3.5
Management
3.0
70%
Recommend to a Friend
Pros
Helpful tool for research and analysis
Useful for job application preparation
Effective for complex marketing challenges
Cons
Limited feedback provided
No specific criticisms mentioned
Insufficient detail on potential drawbacks
Salary Ranges
28 data points
Senior/L5
Senior/L5 · Data Scientist
0 reports
$791,025
total / year
Base
-
Stock
-
Bonus
-
$672,171
$909,879
Interview Experience
1 interviews
Difficulty
4.0
/ 5
Duration
14-28 weeks
Experience
Positive 0%
Neutral 0%
Negative 100%
Interview Process
1
Application Review
2
HR Screen
3
Take-home Marketing Challenge
4
Hiring Manager Interview
5
Panel Interview
6
Offer
Common Questions
Digital Marketing Strategy
Campaign Performance Analysis
Behavioral/STAR
Technical Marketing Knowledge
Case Study
News & Buzz
After lawsuit, one of the biggest Amazon customers, Perplexity, signs $750 million deal with Microsoft, says 'AWS remains' - MSN
Source: MSN
News
·
5w ago
Perplexity signs $750 million AI cloud deal with Microsoft - The American Bazaar
Source: The American Bazaar
News
·
5w ago
Perplexity strikes Microsoft AI cloud deal amid Amazon legal fight - Cryptopolitan
Source: Cryptopolitan
News
·
5w ago
Perplexity Inks Microsoft AI Cloud Deal Amid Dispute With Amazon - Bloomberg
Source: Bloomberg
News
·
5w ago