RESEARCH OUTPUT

Publications

Research in brain-inspired learning, language models, multimodal learning, and intelligent decision-making.

Published and accepted work, alongside manuscripts under review.

Global credit assignment via dynamical criticality ICML 2026

Global Credit Assignment via Dynamical Criticality

First author | Peking University

ICML 2026 | Accepted

This ICML 2026 paper studies temporal credit assignment in recurrent, convolutional recurrent, and spiking systems through criticality-driven online local learning with constant activation memory.

Research & contributions
  • Studies temporal credit assignment beyond BPTT through a criticality-driven local learning rule.
  • Uses long-range spatial inertia and temporal correlations to approximate global credit propagation with a locally computable teaching signal.
  • Derives a closed-form update requiring constant activation memory and O(H) auxiliary state, and validates accuracy, stability, and scalability on RNN, ConvRNN, and SNN benchmarks.
Survey of decision-making large language models for wireless communication IEEE COMST 2025

Decision-Making Large Language Model for Wireless Communication: A Comprehensive Survey on Key Techniques

Ning Yang, Mingrui Fan, Wentao Wang, Haijun Zhang

Third author (advisor is first author) | Institute of Automation, Chinese Academy of Sciences

IEEE COMST | Published

This survey organizes decision-making LLMs for wireless communication across data construction, architecture adaptation, reasoning control, multi-agent coordination, and open challenges.

Research & contributions
  • Surveys LLM-enabled decision-making for wireless communication across data, modeling, reasoning, inference, learning-based control, and multi-agent settings.
  • I independently contributed the sections on data generation and augmentation, architecture adaptation, and open challenges.
  • The paper was published in IEEE Communications Surveys & Tutorials.
Cooperative edge caching with large language models IEEE TMC

Cooperative Edge Caching with Large Language Model in Wireless Networks

Ning Yang, Wentao Wang, Lingtao Ouyang, Haijun Zhang

Second author (student first; advisor is first author) | Institute of Automation, Chinese Academy of Sciences

IEEE TMC | Accepted

We formulate cooperative multi-base-station edge caching as an LLM-native sequential decision problem, build an SFT+GRPO training pipeline, and design an opportunity-aware reward.

Research & contributions
  • Builds a multi-BS, multi-user cooperative edge-caching environment and formulates it as an LLM-native sequential decision problem.
  • Implements a two-stage SFT+GRPO training pipeline with TRL, Unsloth, and QLoRA under a strictly validated text-to-action interface.
  • Achieves strong generalization across user scale, base-station count, and request distributions.
MC-TRCM architecture from Figure 2: source tokens, FiLM-conditioned Transformer fusion, and recursive prediction ICONIP 2026

MC-TRCM: Observation-Aware Recursive Fusion for Incomplete Mobile and Wearable Mental-Health Feature Views

First author | Dalian University of Technology

ICONIP 2026 | Accepted

Observation-aware recursive fusion for asynchronous and incomplete mobile and wearable feature views, combining shared-private representations with task-conditioned multi-task prediction.

Research & contributions
  • Encodes each feature source independently and explicitly models missingness.
  • Combines shared-private latent representations, task-conditioned fusion, and recursive prediction.
  • Achieves leading or competitive results on six tasks across DepreST-CAT and PSYCHE-D.
LinearARD for long-context RoPE restoration NeurIPS 2026

LinearARD: Linear-Memory Attention Distillation for RoPE Restoration

Co-first author (advisor listed first) | Institute of Automation, Chinese Academy of Sciences

NeurIPS 2026 | Accepted

LinearARD aligns row-wise dense self-relations between native-RoPE teachers and RoPE-scaled students, restoring long-context performance with an exact linear-memory distillation objective.

Research & contributions
  • Proposes a restoration-distillation framework that aligns row-wise dense self-relations between a native-RoPE teacher and a RoPE-scaled student.
  • Derives an exact linear-memory KL kernel, avoiding explicit construction of full relation matrices on long sequences.
  • On LLaMA2-7B scaled from 4K to 32K, the method recovers 98.3% of short-context performance using only 4.25M training tokens.
  • Responsible for writing the full manuscript.
Cluster-aware attention model for pickup and delivery problems MAJOR REVISION

Cluster-Aware Attention-Based Deep Reinforcement Learning for Pickup and Delivery Problems

Wentao Wang, Lifeng Han, Guangyu Zou

First author | Dalian University of Technology

Applied Intelligence | Major Revision

CAADRL combines global self-attention, intra-cluster attention, and a dynamic dual-decoder to exploit clustered structure in pickup and delivery problems while reducing inference latency.

Research & contributions
  • Proposes CAADRL, a cluster-aware deep reinforcement learning framework for pickup and delivery problems with clustered structure.
  • Integrates global self-attention, intra-cluster attention, and a dynamic dual-decoder with a learnable gate.
  • Achieves leading performance on clustered benchmarks while reducing inference latency relative to neural collaborative-search baselines.

← Back to homepage