ByteDance
ByteDance

Research Scientist in Vision Foundation Model - Seed - Graduates - 2027 Start (PhD)

RoleMachine Learning
LevelMid Level
LocationSan Jose, Canada, United States
WorkOn-site
TypeRegular
PostedToday
Apply now

About the role

About the team
The Seed Vision team focuses on foundational models for visual generation, developing multimodal generative models, and carrying out leading research and application development to solve fundamental computer vision challenges in GenAI.

Responsibilities:

  • Develop and scale vision foundation models across image and video modalities.
  • Design data pipelines, pre-training strategies, and post-training methods for vision tasks.
  • Improve core capabilities such as perception, reasoning, and multimodal understanding.
  • Optimize model architectures, training efficiency, and evaluation frameworks.
  • Explore real-world applications of vision models in multimodal systems.

Requirements:

Minimum Qualifications:

  • Currently pursuing a PhD in computer science, mathematics, engineering, or a related field, with an expected graduation date in 2027 and the ability to commit to an onboarding date by the end of 2027.
  • Excellent coding ability, data structures, and fundamental algorithm skills, proficient in C/C++ or Python, etc.
  • Experience with computer vision, multimodal learning, or large-scale model training.
  • Strong understanding of deep learning architectures and training methodologies.

Preferred Qualifications:

  • Proven research experience through impactful papers or projects.
  • Strong problem-solving and communication skills.

Required skills

Machine learning

Model evaluation

Data workflows

About ByteDance

San Jose

Headquarters