About Me 🌲
Hi, I am a third-year Ph.D. student at the University of New South Wales (UNSW Sydney), supervised by Dr. Dong Gong. Before that, I received my MPhil degree from the Australian Institute for Machine Learning (AIML), University of Adelaide, supervised by A/Prof. Qi Wu and Dr. Yuankai Qi. I obtained my Bachelor’s degree from Beijing Jiaotong University, where I was advised by Prof. Runmin Cong.
My research focuses on continual learning of foundation models, particularly Dynamic MoE for learning when, where, and how much capacity to add, and which experiences should drive learning. More broadly, I ask: when a trained model encounters new experience, what should change, where should it be stored, and what should drive the update?
Building on this question, I am currently exploring how models can learn from experience through updates to their parameters, recurrent states, and representations of the world:
- Continual and self-improving learning — how a pretrained model can continue acquiring capabilities through post-training, test-time learning, and interaction, while deciding what to update and which experience should drive that update.
- Recurrent sequence modeling — how a fixed-size hidden state can compress and preserve long histories, what information should be written, retained, or overwritten, and whether the state-update rule itself can be learned rather than manually designed.
- World models and interactive simulation — how generative models can learn compact, persistent representations of physical or digital environments from interaction, enabling agents to predict, act, and continue learning over long horizons.
News 🔥
Aug 2026CoDyRA is accepted to EMNLP 2026 Findings. Congratulations to Jeff!
Forgetting in continual post-training grows with LoRA rank — a sparsity penalty on per-component importance keeps each update at the rank its task needs.Aug 2026Released QwenReasoning, an open-source evaluation harness for test-time reasoning.
Soft, entropy-gated and CoT decoding for Qwen3 / Qwen3.5 / Qwen3.8, on HF and vLLM.May 2026Recognized as an ICML 2026 Gold Reviewer!May 2026MoRAM is accepted to ICML 2026. Congratulations to Jeff!
No router — each expert is a rank-1 key–value memory that activates on its own key, so capacity grows as memory atoms instead of coarse experts a router must tell apart.Feb 2026DyMoE is accepted to CVPR 2026. Thanks to all collaborators!
Frozen experts still forget: old and ambiguous tokens in new-task data teach little but drift onto the new experts — DyMoE filters them out token by token.Jun 2025SAME is accepted to ICCV 2025. Congratulations to Gengze!
One agent for seven navigation tasks — experts routed by the agent's state, not by the instruction's granularity.Dec 2024New preprint! Check CoDyRA on arXiv!
Forgetting grows with update rank, so each LoRA learns only the rank it needs — rank minimization as an implicit forgetting regularizer.Dec 2024New preprint! Check MambaCL on arXiv!
Meta-trained so continual learning happens at test time in the forward pass — Mamba's constant-size state is what gets updated, no gradient steps and no growing KV cache.Oct 2024Recognized as an ACM MM 2024 Outstanding Reviewer!Dec 2023WebVLN is accepted to AAAI 2024. Thanks to all collaborators!
Vision-and-language navigation on shopping websites.Jul 2023MG-VLN is accepted to ACM MM 2023. Thanks to all collaborators!
Backtracking to passed correct locations via first-person video grounding for navigation.
Research 🌴
CVPRCVPR 2026
Ambiguous tokens, scored similarly by the old and the new experts, cause most of the drift and contribute least to learning the new task. DyMoE therefore filters tokens before assignment, restricting reassignment to the ambiguous ones and to the router. This reduces forgetting while preserving new-task performance.
PreprintPreprint
We study that update rule as a learning rule: what the state should write, and how that writing should be trained. The rule is meta-learned rather than hand-designed, and training is regularized on what the state keeps. Across Mamba, Longhorn and Gated DeltaNet, as well as RWKV-4 and -7, rules closer to the delta rule learn better from a non-stationary stream.
ICMLICML 2026
EMNLPEMNLP 2026 Findings
ProjectOpen-source project, 2026
ACM MMACM MM 2023
AAAIAAAI 2024
ICCVICCV 2025
