Principal Software Engineer
About the role
### About the Company
NVIDIA is a technology company known for its work in accelerated computing, artificial intelligence, GPUs, and high-performance computing. Its engineering teams develop infrastructure that supports demanding AI and machine learning workloads at large scale.
The company is expanding teams focused on generative AI infrastructure and systems engineering. This role offers an opportunity to work on technologies supporting large language model inference across distributed, heterogeneous computing environments.
### About the Role
NVIDIA is seeking a **Principal Software Engineer** to define the vision and technical roadmap for memory management within large-scale LLM inference and storage systems.
The position will work closely with NVIDIA Dynamo, a high-throughput, low-latency inference framework designed for serving generative AI and reasoning models across multi-node distributed environments. Dynamo uses Rust for performance and Python for extensibility and coordinates GPU shards, request routing, and shared KV-cache management across heterogeneous clusters.
As LLM workloads increasingly exceed the memory capacity of individual GPUs, the role will focus on building infrastructure that efficiently manages data across multiple memory and storage tiers.
### Key Responsibilities
- Design and evolve a unified memory layer spanning GPU memory, pinned host memory, RDMA-accessible memory, SSDs, and remote file, object, or cloud storage.
- Develop systems that support large-scale LLM inference with high throughput and low latency.
- Architect integrations with LLM serving engines such as vLLM, SGLang, and TensorRT-LLM.
- Develop solutions for KV-cache offloading, reuse, sharing, and remote access.
- Design interfaces and protocols supporting disaggregated prefill and peer-to-peer KV-cache sharing.
- Build multi-tier KV-cache storage across GPU memory, CPU memory, local disks, and remote memory.
- Work with GPU architecture, networking, and platform teams on GPUDirect, RDMA, NVLink, and related technologies.
- Optimize KV-cache access and sharing across heterogeneous and disaggregated accelerator environments.
- Profile and optimize systems across CPU, GPU, memory, and networking components.
- Use performance metrics to guide architectural decisions and validate improvements in time-to-first-token (TTFT) and throughput.
- Mentor senior and junior engineers and establish technical direction for memory and storage subsystems.
- Lead cross-functional technical initiatives involving research, product, platform, and customer teams.
- Represent the team in internal technical reviews, open-source communities, conferences, and customer-facing technical discussions.
### Required Qualifications
- Master's degree, PhD, or equivalent professional experience.
- 15+ years of experience building large-scale distributed systems, high-performance storage, or ML infrastructure.
- Strong programming experience with C/C++ and Python.
- Demonstrated experience delivering production-scale services.
- Deep knowledge of memory hierarchies, including GPU HBM, host DRAM, SSD, and remote or object storage.
- Experience designing multi-tier systems optimized for performance and cost efficiency.
- Experience with distributed caching or key-value systems designed for low latency and high concurrency.
- Hands-on experience with networked I/O and technologies such as RDMA, NVMe-oF, or NVLink.
- Understanding of disaggregated and aggregated architectures for AI clusters.
- Strong systems profiling and optimization skills across CPU, GPU, memory, and network resources.
- Ability to use quantitative metrics to evaluate system performance and architectural improvements.
- Excellent communication and technical leadership skills.
- Experience leading initiatives involving research, product, engineering, and customer-facing teams.
### Preferred Qualifications
Candidates may stand out with experience in:
- Open-source LLM serving or systems projects.
- KV-cache optimization, compression, streaming, or reuse.
- Unified memory or storage layers spanning GPU, host, SSD, and cloud storage.
- Enterprise-scale or hyperscale infrastructure.
- Memory-disaggregated architectures.
- RDMA- or NVLink-based data planes.
- KV-cache or CDN-style systems for machine learning.
- Research publications or patents involving LLM systems, distributed memory, storage, or high-performance networking.
### Compensation and Benefits
The stated base salary range for this position is **$272,000 to $425,500 USD**, with the actual base salary determined according to factors such as location, experience, and compensation for comparable positions. The position also includes eligibility for equity and a comprehensive benefits package.
### Work Environment
This is a **full-time remote role** with listed locations in Santa Clara, California; Washington; and Massachusetts. Candidates should review NVIDIA's current employment requirements and the specific application details for the location applicable to them.
### Equal Opportunity
NVIDIA is committed to maintaining a diverse and inclusive workplace. The company states that it does not discriminate in hiring or promotion based on race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability, or other characteristics protected by law.
**Important:** The supplied listing contains a date inconsistency: it says the position was posted **August 31, 2026**, while also stating that applications will be accepted at least until **December 26, 2025**. Applicants should verify the current application deadline directly with the employer before applying.
Responsibilities
- Design and evolve a unified memory layer spanning GPU memory, pinned host memory, RDMA-accessible memory, SSDs, and remote file, object, or cloud storage.
- Develop systems that support large-scale LLM inference with high throughput and low latency.
- Architect integrations with LLM serving engines such as vLLM, SGLang, and TensorRT-LLM.
- Develop solutions for KV-cache offloading, reuse, sharing, and remote access.
- Design interfaces and protocols supporting disaggregated prefill and peer-to-peer KV-cache sharing.
- Build multi-tier KV-cache storage across GPU memory, CPU memory, local disks, and remote memory.
- Work with GPU architecture, networking, and platform teams on GPUDirect, RDMA, NVLink, and related technologies.
- Optimize KV-cache access and sharing across heterogeneous and disaggregated accelerator environments.
Qualifications
- Master's degree, PhD, or equivalent professional experience.
- 15+ years of experience building large-scale distributed systems, high-performance storage, or ML infrastructure.
- Strong programming experience with C/C++ and Python.
- Demonstrated experience delivering production-scale services.
- Deep knowledge of memory hierarchies, including GPU HBM, host DRAM, SSD, and remote or object storage.
- Experience designing multi-tier systems optimized for performance and cost efficiency.
- Experience with distributed caching or key-value systems designed for low latency and high concurrency.
- Hands-on experience with networked I/O and technologies such as RDMA, NVMe-oF, or NVLink.
Benefits
- Eligibility for equity.
- Comprehensive benefits package.
Skills mentioned
About NVIDIA
Torentify Jobs is a modern job platform helping candidates find verified job opportunities quickly and easily. From remote jobs to full-time roles, we connect talent with the right companies. Our goal is to simplify job search and help recruiters hire faster with quality candidates.
H-1B sponsorship history
Historical employer filing data was found for NVIDIA. The employer record includes 6,598 historical certified applications. This is employer-level history, not a guarantee that this role currently offers sponsorship.