
Research Scientist - ByteBrain - Global Frontier Tech Recruitment Program - 2027 Start (PhD)
About the role
Byte Brain is Byte Dance’s AI for Infrastructure (AI4Infra) platform, dedicated to improving the efficiency, reliability, and intelligence of large-scale infrastructure systems through AI and machine learning. Byte Brain supports a wide range of infrastructure domains, including AI data center supply chains, AIOps, Operations Research and Agent Ops, powering infrastructure optimization at massive scale.
Why This Role Is Unique:
This role sits at the intersection of: Operations Research × AIOps × AI for Infra
You will have the opportunity to solve some of the most challenging optimization problems behind large-scale AI datacenters while pioneering the next generation of AI-powered decision-making systems, where LLMs, and optimization algorithms work together to improve efficiency, resource utilization, and operational intelligence across Byte Dance's global infrastructure.
Responsibilities:
- Design and develop AI, machine learning, and optimization algorithms to improve the efficiency, reliability, and performance of large-scale infrastructure systems and AI supply chain. Areas may include AIOps, operations research, software engineering, Agent Ops, and system optimization.
- Drive the deployment, scaling, and continuous improvement of algorithms in production environments, supporting large-scale services.
- Identify optimization opportunities and emerging challenges from real-world infrastructure scenarios, translating them into impactful research and engineering solutions.
- Conduct cutting-edge research and publish high-quality papers in top-tier conferences and journals.
Topic Content:
With the large-scale adoption of LLMs and AI agents, traditional cloud-native infrastructure can no longer meet the ultra-high performance and elasticity requirements of AI workloads. This topic conducts systematic research across the entire AI infrastructure stack:
- Network and Observability: Research intelligent fault localization and root cause analysis for large-scale AI clusters, combined with intelligent tuning of time-series databases to improve cluster stability.
- Storage Systems: Develop serverless high-performance elastic file systems and storage acceleration architectures specifically for AI scenarios, explore hardware-software co-optimization for DPU, and overcome AI storage performance bottlenecks.
- Data Center Power Scheduling: Research GPU/CPU/MEM heterogeneous collaborative scheduling technologies, build a heterogeneous power orchestration system for AI agents, and address scheduling challenges including heterogenous workloads and state dependencies.
- Vector Retrieval: Optimize core vector retrieval technologies for LLM-powered applications, building a cloud-native distributed vector index engine to meet ultra-large-scale vector retrieval demands with low latency and low cost.
- Intelligence and Agent Architecture: Explore automatic infrastructure optimization based on AI Agent workflows, build a self-evolvable business agent framework, and enable full-stack intelligent optimization through AI for Infra.
This topic aims to build a next-generation AI-native infrastructure to support the deployment of LLMs and AI agents, improve resource utilization, reduce costs, support elastic scaling, and drive the technological evolution of AI infrastructure.
Requirements:
- Minimum Qualifications
- Individuals who are completing or have recently completed a PhD degree in Computer Science or a related discipline.
- Proven research track record with multiple publications in top-tier conferences or journals related to AI, machine learning, operations research, systems, or related fields.
- Deep expertise in AI, machine learning, and/or operations research, with hands-on experience in large-scale data analysis and algorithm development.
- Strong coding, implementation, and problem-solving skills, with the ability to bridge research and production systems.
- Excellent communication and cross-functional collaboration skills.
Preferred Qualifications:
Industry experience applying AI and optimization techniques to real-world infrastructure challenges, such as:
- AI data center supply chain optimization
- AI Ops and intelligent operations
- Software engineering productivity optimization
- Operations research and resource scheduling
- System tuning and performance optimization
- Large-scale infrastructure management and automation
About ByteDance
San Jose
Headquarters