
ByteDance
Research Scientist Graduate (Seed Multimodal Interaction and World Model) - 2027 Start
RoleMachine Learning
LevelMid Level
LocationSan Jose, Canada, United States
WorkOn-site
TypeRegular
PostedToday
About the role
About the team
The Seed Multimodal Interaction and World Model team is dedicated to developing models that have human-level multimodal understanding and interaction capabilities. The team is working to advance the exploration and development of multimodal assistant products.
Responsibilities:
- Develop multimodal foundation models integrating vision, language, audio, and environment signals.
- Design and optimize world models for reasoning, planning, and interaction.
- Build training pipelines including data curation, alignment, and reinforcement learning.
- Improve agent capabilities such as perception, memory, decision-making, and tool use.
- Explore next-generation interaction paradigms between humans and intelligent systems.
Requirements:
Minimum Qualifications:
- Individuals who are completing or have recently completed a Bachelor's in Computer Science, Electrical Engineering, Electrical and Computer Engineering, Physics, Mathematics, or a related discipline.
- Excellent coding ability, data structures, and fundamental algorithm skills, proficient in C/C++ or Python, etc.
- Demonstrated interest or project experience in relevant areas.
Preferred Qualifications:
- Experience in multimodal learning, reinforcement learning, or agent systems through internships is preferred.
- Strong problem-solving and collaboration skills.
Required skills
Machine learning
Model evaluation
Data workflows
About ByteDance
San Jose
Headquarters