Post-Training Platform Infrastructure Engineer Job at AMD (San Jose)

Job Description

We are looking for a systems-minded engineer who lives at the intersection of large-scale model inference, distributed systems, and performance optimization. This role focuses on post-training and inference infrastructure, with particular emphasis on P/D disaggregation, KV cache lifecycle management, and efficient offloading mechanisms across both inference and reinforcement learning (RL) systems.

Job Responsibility

Research and deeply understand modern LLM inference frameworks
Analyze and compare inference execution paths to identify performance bottlenecks and inefficiencies
Develop and implement infrastructure-level features to improve inference latency, throughput, and memory efficiency
Optimize KV cache management and offloading strategies
Enhance scalability across multi-GPU and multi-node deployments
Apply the same research-driven approach to RL frameworks
Study post-training and RL systems
Debug performance and correctness issues in distributed RL pipelines
Optimize inference, rollout efficiency, and memory usage during training
Collaborate with research and applied ML teams
Translate model-level requirements into infrastructure capabilities
Validate performance gains with benchmarks and real workloads
Document findings, architectural insights, and best practices to guide future system design

Requirements

Strong background in systems engineering, distributed systems, or ML infrastructure
Hands-on experience with GPU-accelerated workloads and memory-constrained systems
Solid understanding of: LLM inference workflows (prefill vs decode)
Attention mechanisms and KV cache behavior
Multi-process / multi-GPU execution models
Proficiency in Python and C++ (or similar systems languages)
Experience debugging performance issues using profiling tools (GPU, CPU, memory)
Ability to read, understand, and modify complex open-source codebases
Strong analytical skills and comfort working in research-heavy, ambiguous problem spaces
Bachelor's or master's degree in computer science, computer engineering, electrical engineering, or equivalent

Nice to have

Direct experience with LLM inference frameworks or serving stacks
Familiarity with: GPU memory hierarchies (HBM, pinned memory, NUMA considerations)
KV cache compression, paging, or eviction strategies
Storage-backed offloading (NVMe, object stores, distributed file system)
Experience with distributed RL or post-training pipelines
Knowledge of scheduling systems, async execution, or actor-based runtimes
Contributions to open-source ML or systems projects
Experience designing benchmarking suites or performance evaluation frameworks

AMD - All Job Offers

Select Country

Post-Training Platform Infrastructure Engineer

Job Description

Job Responsibility

Requirements

Nice to have

Looking for more opportunities?

Post-Training Platform Infrastructure Engineer

Software Engineer, AI Platform

Applied Research - Forward-Deployed

Technical Program Manager - Infrastructure

Engineering Director, AI Solutions and Automation (ASA)-AI Product Acceleration

AI Researcher

Head of Enterprise Sales

Applied Research - RL & Agents

Member of Technical Staff, MLE

Our AI answers in your language