RESEARCH OUTPUT
Publications
Research in brain-inspired learning, language models, multimodal learning, and intelligent decision-making.
ICML 2026
First author | Peking University
ICML 2026 | Accepted
This ICML 2026 paper studies temporal credit assignment in recurrent, convolutional recurrent, and spiking systems through criticality-driven online local learning with constant activation memory.
Research & contributions
- Studies temporal credit assignment beyond BPTT through a criticality-driven local learning rule.
- Uses long-range spatial inertia and temporal correlations to approximate global credit propagation with a locally computable teaching signal.
- Derives a closed-form update requiring constant activation memory and O(H) auxiliary state, and validates accuracy, stability, and scalability on RNN, ConvRNN, and SNN benchmarks.
IEEE COMST 2025
Ning Yang, Mingrui Fan, Wentao Wang, Haijun Zhang
Third author (advisor is first author) | Institute of Automation, Chinese Academy of Sciences
IEEE COMST | Published
This survey organizes decision-making LLMs for wireless communication across data construction, architecture adaptation, reasoning control, multi-agent coordination, and open challenges.
Research & contributions
- Surveys LLM-enabled decision-making for wireless communication across data, modeling, reasoning, inference, learning-based control, and multi-agent settings.
- I independently contributed the sections on data generation and augmentation, architecture adaptation, and open challenges.
- The paper was published in IEEE Communications Surveys & Tutorials.
IEEE TMC
Ning Yang, Wentao Wang, Lingtao Ouyang, Haijun Zhang
Second author (student first; advisor is first author) | Institute of Automation, Chinese Academy of Sciences
IEEE TMC | Accepted
We formulate cooperative multi-base-station edge caching as an LLM-native sequential decision problem, build an SFT+GRPO training pipeline, and design an opportunity-aware reward.
Research & contributions
- Builds a multi-BS, multi-user cooperative edge-caching environment and formulates it as an LLM-native sequential decision problem.
- Implements a two-stage SFT+GRPO training pipeline with TRL, Unsloth, and QLoRA under a strictly validated text-to-action interface.
- Achieves strong generalization across user scale, base-station count, and request distributions.
ICONIP 2026
First author | Dalian University of Technology
ICONIP 2026 | Accepted
Observation-aware recursive fusion for asynchronous and incomplete mobile and wearable feature views, combining shared-private representations with task-conditioned multi-task prediction.
Research & contributions
- Encodes each feature source independently and explicitly models missingness.
- Combines shared-private latent representations, task-conditioned fusion, and recursive prediction.
- Achieves leading or competitive results on six tasks across DepreST-CAT and PSYCHE-D.
NeurIPS 2026
Co-first author (advisor listed first) | Institute of Automation, Chinese Academy of Sciences
NeurIPS 2026 | Accepted
LinearARD aligns row-wise dense self-relations between native-RoPE teachers and RoPE-scaled students, restoring long-context performance with an exact linear-memory distillation objective.
Research & contributions
- Proposes a restoration-distillation framework that aligns row-wise dense self-relations between a native-RoPE teacher and a RoPE-scaled student.
- Derives an exact linear-memory KL kernel, avoiding explicit construction of full relation matrices on long sequences.
- On LLaMA2-7B scaled from 4K to 32K, the method recovers 98.3% of short-context performance using only 4.25M training tokens.
- Responsible for writing the full manuscript.
MAJOR REVISION
Wentao Wang, Lifeng Han, Guangyu Zou
First author | Dalian University of Technology
Applied Intelligence | Major Revision
CAADRL combines global self-attention, intra-cluster attention, and a dynamic dual-decoder to exploit clustered structure in pickup and delivery problems while reducing inference latency.
Research & contributions
- Proposes CAADRL, a cluster-aware deep reinforcement learning framework for pickup and delivery problems with clustered structure.
- Integrates global self-attention, intra-cluster attention, and a dynamic dual-decoder with a learnable gate.
- Achieves leading performance on clustered benchmarks while reducing inference latency relative to neural collaborative-search baselines.
← Back to homepage