Reinforcement Learning
Reward modeling, policy optimization, and post-training for agents that plan and act.面向会规划、会行动的智能体,研究奖励建模、策略优化与后训练。
AI researcher / builder AI 研究者 / 构建者
I am a Year-1 Ph.D. student at MMLab, CUHK. My recent work spans reinforcement learning, multimodal generation, LLM/Agent, and Physical AI. 我是香港中文大学 MMLab 一年级博士生。近期研究覆盖 Reinforcement Learning、Multimodal Generation、LLM/Agent 与 Physical AI。
Research Focus研究主线
Reward modeling, policy optimization, and post-training for agents that plan and act.面向会规划、会行动的智能体,研究奖励建模、策略优化与后训练。
Unified generation across language, vision, audio, and video.面向语言、视觉、音频、视频的统一生成。
Foundation models, agent reasoning, tool use, and autonomous workflows.面向基础模型、智能体推理、工具使用与自主工作流。
World-action modeling, memory, intent, and control for embodied agents.面向具身智能体的世界-动作建模、记忆、意图与控制。
News近况
Beginning my Ph.D. studies at the Multimedia Laboratory.在 Multimedia Laboratory 开始博士阶段学习。
Physical intelligence series spanning MindVLA-U1, Streaming Intent, DIAL, MindSim, and MindLabel.覆盖 MindVLA-U1、Streaming Intent、DIAL、MindSim 与 MindLabel 的物理智能系列。 Project hub.
Publications论文
arXiv
arXiv
AAAI
AAAI
ICML
arXiv
arXiv
ACL
CVPR
CVPR
AAAI
AAAI
CVPR
CVPR
AAAI
Arxiv
Arxiv
Arxiv
Arxiv
Arxiv
Arxiv Preprint
Arxiv
Arxiv
ICML
Arxiv
arXiv
arXiv
arXiv
AAAI
AAAI
Arxiv
ACL
Arxiv
Blog博客
A compact batch of VLA, world-action model, and embodied world-model papers.一组关于 VLA、世界-动作模型与具身世界模型的阅读索引。
A note on reward, branch-native score objects, and update interfaces.关于奖励、分支原生 score object 与更新接口的笔记。
A short note on keeping research traces public before they become polished work.为什么把尚未打磨成论文的研究痕迹公开留下来。
Contact联系
For project ideas, research collaboration, or joining related efforts, reach out by email.如果你想交流项目想法、研究合作或参与相关工作,可以通过邮件联系我。
jeix782@gmail.comAcademic Service学术服务