ByteDance Seed · San Jose, CA · Research Intern · May 2026–Present
Core contributor to the multimodal RFT/RL recipe for the audio understanding component of the Seed (Doubao) flagship model, including data-mixture design, advantage aggregation, and process-reward optimization with PPO/VAPO. Improved audio-text understanding by 5.1% on average across internal and public benchmarks (MMAU, MMSU), outperforming previous SOTA Gemini 3.1 Pro in this domain.
Built an agentic audio framework (thinking-with-audio) for training and evaluation from scratch for long-form speaker diarization, reducing DER from 55.7% to 11.3% on the Seed Omni model.
