About Me 🌲

My research focuses on continual learning of foundation models, particularly Dynamic MoE for learning when, where, and how much capacity to add, and which experiences should drive learning. More broadly, I ask: when a trained model encounters new experience, what should change, where should it be stored, and what should drive the update?

Building on this question, I am currently exploring how models can learn from experience through updates to their parameters, recurrent states, and representations of the world:

  • Continual and self-improving learning — how a pretrained model can continue acquiring capabilities through post-training, test-time learning, and interaction, while deciding what to update and which experience should drive that update.
  • Recurrent sequence modeling — how a fixed-size hidden state can compress and preserve long histories, what information should be written, retained, or overwritten, and whether the state-update rule itself can be learned rather than manually designed.
  • World models and interactive simulation — how generative models can learn compact, persistent representations of physical or digital environments from interaction, enabling agents to predict, act, and continue learning over long horizons.

News 🔥

  • Aug 2026    CoDyRA is accepted to EMNLP 2026 Findings. Congratulations to Jeff!
    Forgetting in continual post-training grows with LoRA rank — a sparsity penalty on per-component importance keeps each update at the rank its task needs.
  • Aug 2026    Released QwenReasoning, an open-source evaluation harness for test-time reasoning.
    Soft, entropy-gated and CoT decoding for Qwen3 / Qwen3.5 / Qwen3.8, on HF and vLLM.
  • May 2026    Recognized as an ICML 2026 Gold Reviewer!
  • May 2026    MoRAM is accepted to ICML 2026. Congratulations to Jeff!
    No router — each expert is a rank-1 key–value memory that activates on its own key, so capacity grows as memory atoms instead of coarse experts a router must tell apart.
  • Feb 2026    DyMoE is accepted to CVPR 2026. Thanks to all collaborators!
    Frozen experts still forget: old and ambiguous tokens in new-task data teach little but drift onto the new experts — DyMoE filters them out token by token.
  • Jun 2025    SAME is accepted to ICCV 2025. Congratulations to Gengze!
    One agent for seven navigation tasks — experts routed by the agent's state, not by the instruction's granularity.
  • Dec 2024    New preprint! Check CoDyRA on arXiv!
    Forgetting grows with update rank, so each LoRA learns only the rank it needs — rank minimization as an implicit forgetting regularizer.
  • Dec 2024    New preprint! Check MambaCL on arXiv!
    Meta-trained so continual learning happens at test time in the forward pass — Mamba's constant-size state is what gets updated, no gradient steps and no growing KV cache.
  • Oct 2024    Recognized as an ACM MM 2024 Outstanding Reviewer!
  • Dec 2023    WebVLN is accepted to AAAI 2024. Thanks to all collaborators!
    Vision-and-language navigation on shopping websites.
  • Jul 2023    MG-VLN is accepted to ACM MM 2023. Thanks to all collaborators!
    Backtracking to passed correct locations via first-person video grounding for navigation.
Zoomed image

Research 🌴

Continual Learning
CVPR
On Token’s Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language Models
Chongyang Zhao, Mingsong Li, Haodong Lu, Dong Gong
CVPR 2026
TL;DR Paper Project Page Code
In an LVLM whose MoE grows task by task, freezing the old experts and their routers does not prevent forgetting, because of routing drift, in which training on new-task data reassigns old-task tokens to the new experts.
Ambiguous tokens, scored similarly by the old and the new experts, cause most of the drift and contribute least to learning the new task. DyMoE therefore filters tokens before assignment, restricting reassignment to the ambiguous ones and to the router. This reduces forgetting while preserving new-task performance.
Preprint
Learning Mamba as a Continual Learner: Meta-learning Selective State Space Models for Continual Learning
Chongyang Zhao, Dong Gong
Preprint
TL;DR Paper
A recurrent model replaces the unbounded KV cache of attention with a fixed-size state, so at each step its update rule has to decide what to keep and what to overwrite.
We study that update rule as a learning rule: what the state should write, and how that writing should be trained. The rule is meta-learned rather than hand-designed, and training is regularized on what the state keeps. Across Mamba, Longhorn and Gated DeltaNet, as well as RWKV-4 and -7, rules closer to the delta rule learn better from a non-stationary stream.
ICML
Little By Little: Continual Learning via Incremental Mixture of Rank-1 Associative Memory Experts
Haodong Lu, Chongyang Zhao, Jason Xue, Lina Yao, Kristen Moore, Dong Gong
ICML 2026
TL;DR Paper Project Page Code
Taking a whole LoRA module as one expert is a coarse unit of growth, since each new expert overlaps with the existing ones and routing becomes less reliable as they accumulate. Read as a linear associative memory, an expert decomposes into rank-1 key–value pairs, each responding only to its own key. Capacity then grows one memory at a time, and no router is required.
EMNLP
Take Only What You Need: Rank Minimization as an Implicit Forgetting Regularizer in Continual Learning
Haodong Lu, Chongyang Zhao, Jason Xue, Lina Yao, Kristen Moore, Dong Gong
EMNLP 2026 Findings
TL;DR Paper Code
The best rank shifts with the module and the task, and a bound explains why: forgetting scales with both the rank of an update, how many directions it uses, and its norm, how far it moves along each. One sparsity penalty on per-component importance presses both at once, leaving every update at the rank its task needs and mergeable straight back into the backbone — no task ID, no replay.
Reasoning
Project
QwenReasoning: Test-Time Reasoning Methods for Qwen3, Qwen3.5 and Qwen3.8
Chongyang Zhao
Open-source project, 2026
TL;DR Code
Training-free decoding that skips the argmax: the probability-weighted embedding goes back into the model, so the next step reads the full distribution rather than one sampled token. Soft, entropy-gated and CoT variants share an entry point, on HF and a vLLM fork, over five math and science benchmarks.
Embodied Navigation
ACM MM
Mind the Gap: Improving Success Rate of Vision-and-Language Navigation by Revisiting Oracle Success Routes
Chongyang Zhao, Yuankai Qi, Qi Wu
ACM MM 2023
TL;DR Paper
Across four VLN agents on R2R and REVERIE, oracle success runs up to 9% above success — the agent reaches the target and walks on. The stop is therefore recoverable from a trajectory an off-the-shelf agent already produced: read the route as a first-person video and ground the instruction in it to locate the frame to stop at.
AAAI
WebVLN: Vision-and-Language Navigation on Websites
Qi Chen*, Dileepa Pitawela*, Chongyang Zhao*, Gengze Zhou, Hsiang-Ting Chen, Qi Wu
AAAI 2024
TL;DR Paper Code
One of the early benchmarks for language-guided web agents, moving navigation from 3D houses to tree-structured web pages and replacing turn-by-turn instructions with questions to answer. The agent reads the raw HTML as well as the rendering, which exposes text and structure the page never displays. I contributed to the task formulation and was responsible for constructing the WebVLN-v1 benchmark and adapting embodied navigation methods to the web.
ICCV
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
Gengze Zhou, Yicong Hong, Zun Wang, Chongyang Zhao, Mohit Bansal, Qi Wu
ICCV 2025
TL;DR Paper Code
Category-level search and step-by-step instruction following are built as separate tasks, though both read an instruction, look around, and pick a move. SAME routes its experts on the agent's navigation state rather than on the instruction's granularity, covering seven language-guided navigation tasks with one agent.

Services

Conference Reviewer

ICML '24 '25 '26 Gold Reviewer NeurIPS '24 '25 '26 ICLR '25 '26 '27
CVPR '24 '25 '26 ICCV '25 ECCV '24 '26
ACM MM '24 Outstanding Reviewer AAAI '25 '27 ICDM '24

Journal Reviewer

IJCV IEEE-TNNLS IEEE-TCSVT Neurocomputing CAAI-TIT