ByteDance
ByteDance

Research Engineer, LLM/VLM Inference Optimization (Kernel & Compiler) - Seed Infra

직무머신러닝
경력중급
위치Seattle, WA, United States
근무오피스 출근
고용Regular
게시오늘
지원하기

포지션 소개

About the Team:

The Seed Infrastructures team oversees the distributed training, reinforcement learning framework, high-performance inference, and heterogeneous hardware compilation technologies for AI foundation models.

Responsibilities:

  • Design, implement, and optimize high-performance GPU kernels for large-scale LLM/VLM inference workloads, including attention, GEMM, and other compute- and memory-intensive operators.
  • Develop and tune inference kernels in CUDA and Triton, and drive end-to-end performance optimization of production inference systems at scale.
  • Conduct in-depth performance analysis and profiling to identify bottlenecks across the inference stack, from kernel level to serving level.
  • Collaborate with research and infrastructure teams to land kernel- and compiler-level optimizations in production inference systems.

Requirements:

Minimum Qualifications:

  • Bachelor's degree or above in Computer Science, Electrical Engineering, or a related field.
  • Strong proficiency in C/C++ and Python; solid foundations in algorithms, data structures, and systems programming.
  • Hands-on experience in LLM/VLM inference optimization with demonstrated impact on latency, throughput, or serving cost.
  • Hands-on experience writing and optimizing GPU kernels in CUDA and/or Triton.
  • Deep understanding of GPU architecture (memory hierarchy, occupancy, instruction throughput) with solid optimization experience.

Preferred Qualifications:

  • Experience with ML compiler internals (e.g., Triton, MLIR, LLVM).
  • Contributions to related open-source projects (e.g., Triton, vLLM, SGLang, Flash Attention, CUTLASS).
  • Publications in relevant venues (e.g., MLSys, OSDI, ASPLOS).

필수 스킬

Machine learning

Model evaluation

Data workflows

ByteDance 소개

Seattle

본사 위치