ByteDance
ByteDance

Research Scientist Graduate (DPU & AI Infra) - 2026 Start (PhD)

RoleMachine Learning
LevelEntry
LocationSan Jose, Canada, United States
WorkOn-site
TypeRegular
PostedToday
Apply now

About the role

About the Team:

The Byte Dance DPU (Data Processing Unit) team is building the foundational computing infrastructure for Byte Dance and Volcano Engine Public Cloud. Our mission is to advance the architecture, development, and research of next-generation software-hardware technologies across compute, networking, and storage for cloud and AI computing.

Our technology stack spans:

  • Cloud virtualization & hypervisors
  • High-performance user-space network protocols (DPDK, RDMA, etc.)
  • High-speed interconnect and virtual switching
  • Distributed storage acceleration
  • GPU virtualization and scheduling for AI/ML workloads

We work at the intersection of software systems, distributed infrastructure, and custom hardware acceleration, shaping the next wave of cloud-scale computing.

We are looking for talented individuals to join our team in 2026. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Launch your career where inspiration is infinite at Byte Dance.

Successful candidates must be able to commit to an onboarding date by end of year 2026. Please state your availability and graduation date clearly in your resume.

Responsibilities:

  • Design and develop DPU network software with a focus on high performance, low latency, and reliability.
  • Collaborate with hardware teams to build software-hardware co-design solutions for networking and storage acceleration.
  • Explore AI/ML infrastructure acceleration, leveraging DPUs, GPUs, and custom hardware to optimize distributed training and inference.
  • Drive end-to-end performance optimization, from OS kernels and drivers to user-space runtime systems.
  • Contribute to architecture design, technical proposals, and long-term research directions.

Requirements:

  • Minimum Qualifications

  • Individuals who are completing or have recently completed a PhD degree in related fields with research training and publications.

  • 2+ years of relevant industry experience (exception for Ph.D. with strong background).

  • Proficiency in C/C++ development and debugging.

  • Strong Linux systems development experience.

  • Solid understanding of compute, network architecture, and operating systems.

  • Background in at least one of: software-hardware co-design, distributed systems, high-performance networking, or AI/ML systems.

  • Preferred Qualifications

  • Experience with software-hardware co-design (networking, storage, or distributed compute).

  • Hands-on experience with network virtualization (OVS, SR-IOV, eBPF).

  • Familiarity with DPDK and high-performance user-space networking.

  • Bonus points for hardware acceleration experience, FPGA/ASIC/GPU/CUDA

  • Bonus points for experience with NCCL Collectives along with AI communication patterns and parallelization techniques

  • Proven experience designing and building AI/ML infrastructure related but not limited to inference kv cache system, data preprocessing system.

For Byte Dance:

By submitting an application for this role, you accept and agree to our global applicant privacy policy, which may be accessed here: https://jobs.bytedance.com/en/legal/privacy

Required skills

Machine learning

Model evaluation

Data workflows

About ByteDance

San Jose

Headquarters