Data AML is ByteDance's Machine Learning mid-platform, providing training and inference systems for recommendation/advertising for businesses such as Douyin, Jinri Toutiao, and Xigua Video. It provides powerful Machine Learning computing power for internal business units within the company and conducts research on some general and innovative algorithms for issues in these businesses.
We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.
Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.
Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.
Job Description
- Responsible for the computational performance optimization of ByteDance's recommendation mid-platform models, conduct in-depth tuning for inference/training bottlenecks in business scenarios, and improve computing utilization.
- Lead the design and development of high-performance kernel libraries, covering general-purpose and business-customized kernels, including general computation and communication parallelism, to ensure the ultimate performance of kernels.
- Deeply cultivate model compilation optimization technology, and based on directions such as graph optimization, kernel fusion, computation scheduling, and code generation, build and improve the mid-platform model compilation system.
- Collaborate with the business and algorithm teams to identify performance issues, provide full stack performance analysis, bottleneck diagnosis, and optimization solutions, consolidate general-purpose performance optimization components, toolchains, and platform capabilities, and empower multiple internal business.
Minimum Qualifications
- Individuals who are completing or have recently completed a Bachelor's/ Master's degree in Software Development, Computer Science, Computer Engineering, or a related technical discipline, or a related discipline.
- Familiar with mainstream model compilation stacks (such as TVM, MLIR, XLA, etc.), with relevant experience in development, and optimization;
- Proficient in C/C++ development, familiar with assembly, CPU/GPU architecture, and cache mechanism, with practical experience in high-performance kernel development;
- Familiar with the underlying principles of deep learning frameworks (such as TensorFlow, PyTorch, OneFlow, etc.), understand the computational graph structure, inference/training execution process, and have experience in implementing model graph optimization and compilation optimization;
- Possess strong abilities in independent thinking, problem decomposition, performance troubleshooting, and practical optimization, and be able to independently overcome complex performance bottlenecks;
Preferred Qualifications
- Have experience in joint hardware and software design, and possess experience in heterogeneous computing projects;
- Have experience in contributing to or developing open-source deep learning kernel libraries, compilers, or inference engines;
- Have in-depth research experience on the underlying architecture and mechanisms of at least one machine learning framework (TensorFlow / PyTorch / MxNet or other self-developed frameworks).