ByteDance
ByteDance

Research Engineer — Training Performance & ML Compilation (Torch Compile) - Seed Infra

职能机器学习
级别中级
地点Seattle, WA, United States
方式现场办公
类型Regular
发布今天
立即申请

职位介绍

About the Team:

The Seed Infrastructures team oversees the distributed training, reinforcement learning framework, high-performance inference, and heterogeneous hardware compilation technologies for AI foundation models.

Responsibilities:

  • Optimize training performance for large-scale foundation models through compiler-level techniques, including graph optimization, operator fusion, and kernel generation.
  • Develop and extend ML compilation capabilities based on the Py Torch compilation stack (e.g. FX, Dynamo, Inductor) to improve training efficiency across heterogeneous GPU platforms.
  • Design and optimize high-performance GPU kernels for training workloads.
  • Conduct performance profiling and analysis of large-scale training jobs; identify and resolve bottlenecks in collaboration with research and infrastructure teams.

Requirements:

  • Minimum Qualification(s)

  • Bachelor's degree or above in Computer Science, Electrical Engineering, or a related field.

  • Strong proficiency in C/C++ and Python; solid foundations in algorithms, data structures, and systems programming.

  • Hands-on experience in training-side performance optimization for deep learning workloads.

  • Hands-on experience writing and optimizing GPU kernels (e.g., CUDA, Triton).

  • Experience with the Py Torch compilation stack, meeting at least one of the following: direct experience using, debugging, or extending Inductor or FX; proficiency in Triton kernel development; or solid experience with Py Torch computation graph work (graph optimization, graph capture, operator fusion).

  • Preferred Qualification(s)

  • Experience with Torch Dynamo or bytecode-level program transformation.

  • Experience with Triton compiler internals or other ML compiler backends (e.g., MLIR, LLVM).

  • Contributions to related open-source projects (e.g., Py Torch, Triton, Flash Attention).

  • Publications in relevant venues (e.g., MLSys, OSDI, ASPLOS).

必备技能

Machine learning

Model evaluation

Data workflows

关于ByteDance

Seattle

总部位置