
ByteDance
Site Reliability Engineer - AI Application
RoleDevops
LevelMid Level
LocationSingapore
WorkOn-site
TypeRegular
PostedToday
About the role
About the team
We are an AI-driven search and recommendation team focused on building innovative, scalable products for global users.
Responsibilities:
- Ensure the reliability and normal operation of multiple core systems related to Viking Team's Big data and online services, while focusing on system capacity planning and stability assurance;
- Enhance system visibility by monitoring the availability and performance metrics of system components, helping development teams quickly locate faults, and especially ensuring operation of critical links such as AI search/vector databases;
- Improve the reliability, scalability, and Performance optimization of services to ensure the achievement of the core system SLA;
- Participated in the design and implementation of the automation platform, ensuring the rapid iteration and efficient operation and maintenance of large-scale online Viking clusters and AI search-related clusters;
- Combining with the usage scenarios of AI Search/Viking business, in-depth optimization of service governance practices, including but not limited to analysis of performance bottlenecks in key AI Search/Viking links, business problem location and troubleshooting, promoting the transformation and upgrading of the system's high-availability architecture, and those familiar with Viking-related technologies are preferred to participate in core optimization work.
Requirements:
Minimum Qualifications:
- Bachelor's degree or above, majoring in computer-related fields, with more than five years of relevant work experience;
- Has a solid foundation in computer software knowledge, and understands the relevant principles of Linux operating systems, storage, network IO, etc.
- Familiar with at least one programming language (such as Python/Go/Java/Shell/Ansible), with moderate development capabilities, and placing more emphasis on operations and maintenance practices and problem-solving abilities;
- Understand at least one type of knowledge related to cloud infrastructure such as AWS/Volcano Engine/Aliyun/GCP; those with experience in computing/distributed systems are preferred (e.g., Nginx/Kubernetes/Docker/Open Stack/Hadoop/Spark/Flink, etc.);
Preferred Qualifications:
- Familiar with algorithmic thinking, good data structure and system design capabilities
- Have certain understanding of AI Cloud, large model-related Search Suggestion, and Recommender system.
Required skills
Cloud infrastructure
Reliability engineering
Automation
About ByteDance
Singapore
Headquarters