ByteDance
ByteDance

Research Scientist - ByteBrain - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

RoleMachine Learning
LevelMid Level
LocationSan Jose, Canada, United States
WorkOn-site
TypeRegular
PostedToday
Apply now

About the role

Byte Brain is Byte Dance’s AI for Infrastructure (AI4Infra) platform, dedicated to improving the efficiency, reliability, and intelligence of large-scale infrastructure systems through AI and machine learning. Byte Brain supports a wide range of infrastructure domains, including AI data center supply chains, AIOps, Operations Research and Agent Ops, powering infrastructure optimization at massive scale.

Why This Role Is Unique:

This role sits at the intersection of: Operations Research × AIOps × AI for Infra
You will have the opportunity to solve some of the most challenging optimization problems behind large-scale AI datacenters while pioneering the next generation of AI-powered decision-making systems, where LLMs, and optimization algorithms work together to improve efficiency, resource utilization, and operational intelligence across Byte Dance's global infrastructure.

Responsibilities:

  • Design and develop AI, machine learning, and optimization algorithms to improve the efficiency, reliability, and performance of large-scale infrastructure systems and AI supply chain. Areas may include AIOps, operations research, software engineering, Agent Ops, and system optimization.
  • Drive the deployment, scaling, and continuous improvement of algorithms in production environments, supporting large-scale services.
  • Identify optimization opportunities and emerging challenges from real-world infrastructure scenarios, translating them into impactful research and engineering solutions.
  • Conduct cutting-edge research and publish high-quality papers in top-tier conferences and journals.

Topic Content:
With the large-scale adoption of LLMs and AI agents, traditional cloud-native infrastructure can no longer meet the ultra-high performance and elasticity requirements of AI workloads. This topic conducts systematic research across the entire AI infrastructure stack:

  1. Network and Observability: Research intelligent fault localization and root cause analysis for large-scale AI clusters, combined with intelligent tuning of time-series databases to improve cluster stability.
  2. Storage Systems: Develop serverless high-performance elastic file systems and storage acceleration architectures specifically for AI scenarios, explore hardware-software co-optimization for DPU, and overcome AI storage performance bottlenecks.
  3. Data Center Power Scheduling: Research GPU/CPU/MEM heterogeneous collaborative scheduling technologies, build a heterogeneous power orchestration system for AI agents, and address scheduling challenges including heterogenous workloads and state dependencies.
  4. Vector Retrieval: Optimize core vector retrieval technologies for LLM-powered applications, building a cloud-native distributed vector index engine to meet ultra-large-scale vector retrieval demands with low latency and low cost.
  5. Intelligence and Agent Architecture: Explore automatic infrastructure optimization based on AI Agent workflows, build a self-evolvable business agent framework, and enable full-stack intelligent optimization through AI for Infra.

This topic aims to build a next-generation AI-native infrastructure to support the deployment of LLMs and AI agents, improve resource utilization, reduce costs, support elastic scaling, and drive the technological evolution of AI infrastructure.

Requirements:

  • Minimum Qualifications
  • Individuals who are completing or have recently completed a PhD degree in Computer Science or a related discipline.
  • Proven research track record with multiple publications in top-tier conferences or journals related to AI, machine learning, operations research, systems, or related fields.
  • Deep expertise in AI, machine learning, and/or operations research, with hands-on experience in large-scale data analysis and algorithm development.
  • Strong coding, implementation, and problem-solving skills, with the ability to bridge research and production systems.
  • Excellent communication and cross-functional collaboration skills.

Preferred Qualifications:

Industry experience applying AI and optimization techniques to real-world infrastructure challenges, such as:

  • AI data center supply chain optimization
  • AI Ops and intelligent operations
  • Software engineering productivity optimization
  • Operations research and resource scheduling
  • System tuning and performance optimization
  • Large-scale infrastructure management and automation

About ByteDance

San Jose

Headquarters