ByteDance
ByteDance

Research Scientist - LLM Training System as a Service - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

RoleMachine Learning
LevelMid Level
LocationSan Jose, Canada, United States
WorkOn-site
TypeRegular
PostedToday
Apply now

About the role

We are looking for talented individuals to join our team in 2027. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Launch your career where inspiration is infinite at our Company.

Successful candidates must be able to commit to an onboarding date by end of year 2027. Please state your availability and graduation date clearly in your resume.

Team Introduction:
AML-Ark combines system engineering and the art of machine learning to develop and maintain massively distributed ML training and Inference system/services around the world, providing high-performance, highly reliable, scalable systems for LLM/AIGC/AGI.

Topic Content:
With the evolution from large language models (LLMs) to AI Agents, the training paradigm is undergoing a fundamental shift. Traditional distributed training frameworks like Megatron-LM are designed around relatively static parallelism strategies, whereas Agent training introduces more dynamic patterns, including external tool interactions, multi-step reasoning, and iterative self-improvement.

In this context, tightly coupled system design can limit flexibility and efficiency. To better support these emerging workloads, we aim to build a robust architecture that cleanly separates “logical control” from “compute execution,” enabling more scalable and adaptable training workflows.

Responsibilities:

  • Responsible for developing and optimizing LLM training & inference & Reinforcement Learning framework.
  • Working closely with model researchers to scale LLM training & Reinforcement Learning to the next level.
  • Responsible for GPU and CUDA Performance optimization to create an industry-leading high-performance LLM training and inference and RL engine.

Requirements:

Minimum Qualifications:

  • Currently pursuing a PhD in computer science, automation, electronics engineering or a related technical discipline
  • Proficient in algorithms and data structures, familiar with Python
  • Understand the basic principles of deep learning algorithms, be familiar with the basic architecture of neural networks and understand deep learning training frameworks such as Pytorch.

Preferred Qualifications:

  • Proficient in GPU high-performance computing optimization technology on CUDA, in-depth understanding of computer architecture, familiar with parallel computing optimization, memory access optimization, low-bit computing, etc.
  • Familiar with FSDP, Deepspeed, JAX SPMD, Megatron-LM, Verl, TensorRT-LLM, ORCA, VLLM, SGLang, etc.
  • Knowledge of LLM models, experience in accelerating LLM model optimization is preferred.

Required skills

Machine learning

Model evaluation

Data workflows

About ByteDance

San Jose

Headquarters