拓达教育 Touchdown Education 拓达求职支持办公室Career Support Office touchdown.org.cn ↗
JobFinder为留学生筛选身份匹配的岗位

MLSys Backend Engineer Graduate (AML Ark)

TikTok · 新加坡 · 收录于 8月15日
CSO 可内推 本科硕士 软件工程
评分
92
发布日期
收录于 8月15日 (2周前)
内推
有 CSO 内推渠道
公司规模
10k+
行业
互联网

岗位描述

Team Introduction

Volcano Ark is Volcano Engine's one-stop foundation model and Agent platform for enterprises and developers. Our team is building the next-generation infrastructure for general-purpose Agents capable of handling complex tasks—from Agent runtimes and execution engines, to platform-level resource models, lifecycle management, and public APIs, as well as evaluation and observability systems.

We are looking for talented individuals to join our team in 2027. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Launch your career where inspiration is infinite at ByteDance.

Successful candidates must be able to commit to an onboarding date by end of year 2027. Please state your availability and graduation date clearly in your resume.

Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to ByteDance and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.

Responsibilities

-Design and develop resource scheduling systems for machine learning platforms, supporting model training, evaluation, and inference workloads across domains such as NLP, Computer Vision (CV), and Speech.

-Optimize the orchestration and scheduling of heterogeneous computing resources, including GPUs, CPUs, and other specialized accelerators, to maximize the utilization of dedicated, opportunistic (spot), co-located, and multi-cloud resources.

-Develop scheduling solutions that optimize the allocation of compute resources, RDMA high-speed networking, and storage resources, enabling large-scale distributed clusters to achieve maximum performance and efficiency.

-Design and implement workload scheduling across multiple data centers, regions, and cloud environments, ensuring balanced resource utilization and efficient placement of both online services and offline batch jobs.

Minimum Qualifications

-Individuals who are completing or have recently completed a Bachelor's/ Master's degree in computing or a related discipline.

-Proficient in one or more programming languages such as Go, Python, or Shell scripting in a Linux environment.

-Strong understanding of the Kubernetes architecture and ecosystem, with hands-on experience in container technologies such as Docker, Containerd, Kata Containers, and Podman. Extensive experience in developing and operating machine learning systems is preferred.

-Solid understanding of distributed systems principles, with experience designing, developing, and maintaining large-scale distributed systems.

Preferred Qualifications

-Familiar with at least one ML framework: TensorFlow / PyTorch

-Hands on with AI Infrastructure,HW/SW Co-Design,High Performance Computing,ML Hardware Architecture (GPU、Accelerators、Networking)

加入我的投递 →