Speech & Omni Models
Speech, vision, video and audio intelligence across understanding and generation.
I am currently a Staff Researcher/Tech Lead at Alibaba Group (ATH), where I lead the Speech and Omni LLM Applied Research team. Previously, I was a researcher at the Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), focusing on multilingual and multimodal large language models. I earned my Ph.D. in Machine Learning from Dublin City University's ML-Labs in 2023, following a Bachelor of Engineering from Northeastern University in China in 2018.
My research spans three main lines: speech and omni models, represented by Marco-Voice for expressive voice cloning and emotional speech synthesis, CVQA for culturally diverse multilingual visual question answering, and GPT4Video for unified video understanding and generation; LLM agents and reliability, including Trust No Tool on agents operating under untrusted tool feedback, Crayotter for traceable long-form video editing, and CoQuIR for quality-aware code retrieval; and multilingual models, including Marco-MoE for efficient sparse multilingual modeling, Marco-LLM for cross-lingual enhancement, and CulturALL for grounded multilingual and multicultural evaluation across 14 languages and 51 regions. My papers have been accepted at top-tier conferences, including NeurIPS, ICML, COLM, ACL, EMNLP, ICASSP, CVPR, and ACM-MM.
My work has received both research and personal recognition. CoQuIR was selected for an ACL 2026 SAC Highlights Award, GPT4Video was nominated for the Best Paper Award at ACM-MM 2024, and I won two championships and two runner-up prizes at the IWSLT 2025 speech translation competition. My open-source projects and contributions have collectively earned 4k+ GitHub stars. I have also received personal honors, including the German DAAD AInet Fellowship, the 2023 Young AI Role Model of the Year award at the Irish AI Awards, and an SFI PhD Scholarship. Individual projects have also received external coverage: CVQA was featured by MBZUAI News, Microsoft Research, and a Microsoft Research podcast; Marco-Voice was covered by Slator; and Marco-LLM was reported by Bloomberg, CNBC, and the South China Morning Post. My broader work and career have also been featured by RTÉ and Irish Tech News, including a podcast interview on LLMs. Before joining Alibaba Group, I held research and visiting positions at Tencent AI Lab, the National Institute of Informatics (NII), and IBM Research-China.
Speech, vision, video and audio intelligence across understanding and generation.
Traceable multi-agent workflows, tool-feedback defense and robust evaluation.
Efficient model adaptation and culturally grounded intelligence across languages.

Dublin City University · ML-Labs


Alibaba Group · ATH · Speech and Omni LLM Applied Research

Mohamed bin Zayed University of Artificial Intelligence · Multilingual and Multimodal LLMs




🎬 Crayotter, a major initiative on long-horizon video-editing agents, now brings together an open-source project and two papers, with nearly 200 GitHub stars. Two articles featured our work: 从“一句成片”到“长轨推演”:探究多模态智能体在长视频编辑中的应用 and Crayotter:用群体相对偏好反向传播,训练长视频编辑智能体.
🏆 CoQuIR was selected for an ACL 2026 SAC Highlights Award.
🎉 Marco-MoE was accepted to COLM 2026 and featured by 24 AI.
🎉 Spurious Rewards Paradox was accepted to ICML 2026.
🎉 ElasticFormer was accepted to CVPR 2026.
🎉 Marco-Voice, LongSpeech and MECap-R1 were accepted to ICASSP 2026.
🎤 Invited industry expert talk on Marco Models at ACM Multimedia Asia 2025.
📰 Marco-Voice was featured by Slator for its unified approach to voice cloning and emotional speech synthesis.
🏆 Our team secured two championships and two runner-up prizes at IWSLT 2025.
🎉 Four papers on multilingual LLMs and hallucination detection were accepted to ACL 2025.
🎙️ CVQA was featured in a Microsoft Research podcast on culturally aware and linguistically diverse multimodal evaluation.
* denotes corresponding or equal contribution. See Google Scholar and DBLP for the complete record.
arXiv preprint, 2026.
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics.
Conference on Language Modeling.
IEEE/CVF Conference on Computer Vision and Pattern Recognition.
IEEE International Conference on Acoustics, Speech, and Signal Processing.
IEEE International Conference on Acoustics, Speech, and Signal Processing.
Proceedings of the 32nd ACM International Conference on Multimedia.
NeurIPS Datasets and Benchmarks Track.