ByteDance
ByteDance

Student Researcher (LLM Post Training – Agent & Reinforcement Learning) - 2026 Start (PhD)

RoleMachine Learning
LevelMid Level
LocationSan Jose, Canada, United States
WorkOn-site
TypeInternship
PostedToday
Apply now

About the role

About the team
The Seed LLM Post Training team is responsible for researching cutting-edge posttrain technologies and providing core posttrain capabilities for unified multimodal large models. The team's goal is to research and explore next-generation advanced technologies such as SFT, RM, RL, and self-learning during the posttrain phase, while significantly optimizing and improving key areas including reasoning, coding, agent, and omni model.

Responsibilities:

  • Explore large-scale models and optimize systems.
  • Data construction, instruction tuning, preference alignment, and model optimization.
  • Improving relevant model capabilities, such as reasoning, code, math etc.
  • In-depth research and exploration of future use cases.

Requirements:

Minimum Qualifications:

  • Currently pursuing a PhD in Computer Science, AI, or a related field.
  • Research experience in reinforcement learning, sequential decision-making, or agent behavior.
  • First-author publications in accredited ML/AI conferences (e.g., NeurIPS, ICLR, ICML).
  • Solid programming and experimentation skills, including with RL or LLM frameworks.

Preferred Qualifications:

  • Experience with LLM agents, tool use, or prompt-based control.
  • Familiarity with environments such as Web Arena, ALFWorld, or programmatic reasoning tasks.
  • Understanding of RL techniques such as reward shaping, memory augmentation, or curriculum learning.

As a condition of employment, all successful candidates must be able to establish authorization to work in the United States. For this position, the Company does not provide sponsorship or any immigration-related benefits.

Required skills

Machine learning

Model evaluation

Data workflows

About ByteDance

San Jose

Headquarters