Job Description
Senior Product Manager – AI Inference Performance
NVIDIA is seeking an exceptional Senior Product Manager to own the products that enable customers to extract maximum performance from AI models running on NVIDIA hardware. Every inference deployment—from a single-GPU workstation to multi-thousand-GPU data centers—depends on latency, efficiency, and cost per token. This role turns deep optimization techniques into broadly adoptable products that make NVIDIA the clear choice for AI inference.
You will set the strategy across the full inference stack: model representation, memory and state management, request scheduling and serving, and token generation. The techniques evolve rapidly; you will evaluate which matter most, decide what to build, what to adopt from the ecosystem, and what to retire. The focus is on platforms that generalize across model families, deployment sizes, and customer needs rather than one-off solutions.
Key Responsibilities
- Own the inference performance roadmap – Define direction for how models are represented, how KV caches and state are managed, how requests are scheduled, and how tokens are generated at scale.
- Build platforms, not point solutions – Deliver capabilities with sane defaults, easy adoption, and extensibility that work for single-GPU users through hyperscale deployments.
- Agentic and multi-turn workloads – Develop performance strategies for long-running agent sessions, tool-call stalls, unpredictable output lengths, cross-turn cache reuse, request prioritization, and efficient idle-time handling.
- Framework and ecosystem strategy – Ensure optimizations land effectively across TensorRT-LLM, vLLM, SGLang, NVIDIA Dynamo, and related serving stacks. Partner with open-source communities and internal engineering teams.
- Benchmarking and credible performance claims – Own methodology, key metrics (TTFT, inter-token latency, throughput per GPU, cost per million tokens), and guardrails so published numbers remain trustworthy and reproducible.
- Day-to-day product ownership – Drive release readiness, quality bars, regression tracking, customer blocking issues, and the continuous feedback loop from production deployments back into the roadmap.
You will operate with high autonomy inside a small, high-leverage Product Management organization that shapes NVIDIA’s Deep Learning and Generative AI strategy. Success requires forming a clear point of view from data and customer conversations and driving it to shipped outcomes.
Required Qualifications
- 12+ years in product management at a technology company, or equivalent experience as a founder, engineering lead, or technical product owner.
- Deep expertise in AI inference optimization techniques including KV caching and reuse, quantization, speculative decoding, and disaggregated serving—and how each impacts accuracy, latency, and cost.
- Familiarity with major inference frameworks: TensorRT-LLM, vLLM, SGLang, NVIDIA Dynamo, and the broader serving and orchestration ecosystem.
- Proven ability to take ambiguous problem spaces, define strategy, and ship results independently.
- Operational experience running live products: release management, quality and regression rigor, customer issues, and support processes.
- Skill translating low-level technical capability into clear business value (lower TCO, faster response times, higher GPU utilization) for both engineers and executives.
- BS, MS, or PhD in Computer Science, Computer Engineering, or a related field (or equivalent practical experience).
Preferred / Stand-Out Qualifications
- Hands-on engineering experience with LLM inference performance—profiling, kernel-level or serving-level optimization, or building serving stacks.
- Open-source contributions or product leadership in vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or Dynamo.
- Production-scale experience with capacity planning, autoscaling, SLA management, or stateful multi-turn / agentic applications.
- A demonstrated habit of reading current research and converting insights into concrete roadmap decisions.
This role sits at the intersection of cutting-edge AI research, large-scale systems engineering, and customer-facing product strategy. You will influence how the world’s most important AI workloads run efficiently and cost-effectively on NVIDIA platforms.