About the team
The Seed Multimodal Interaction and World Model team is dedicated to developing models that have human-level multimodal understanding and interaction capabilities. The team is working to advance the exploration and development of multimodal assistant products.
Responsibilities
- Develop multimodal foundation models integrating vision, language, audio, and environment signals.
- Design and optimize world models for reasoning, planning, and interaction.
- Build training pipelines including data curation, alignment, and reinforcement learning.
- Improve agent capabilities such as perception, memory, decision-making, and tool use.
- Explore next-generation interaction paradigms between humans and intelligent systems.
Minimum Qualifications
- Individuals who are completing or have recently completed a Bachelor's in Computer Science, Electrical Engineering, Electrical and Computer Engineering, Physics, Mathematics, or a related discipline.
- Excellent coding ability, data structures, and fundamental algorithm skills, proficient in C/C++ or Python, etc.
- Demonstrated interest or project experience in relevant areas.
Preferred Qualifications
- Experience in multimodal learning, reinforcement learning, or agent systems through internships is preferred.
- Strong problem-solving and collaboration skills.