About
I am a senior undergraduate in Computer Science at Shanghai Jiao Tong University (ACM Honors Class, Zhiyuan College). I am currently a research intern at HuLab, University of Wisconsin–Madison, advised by Prof. Junjie Hu. At SJTU, I work in the Advanced Network Laboratory, advised by Prof. Shengzhong Liu and Prof. Fan Wu. Previously, I was a research intern at Audiocc, advised by Prof. Chenda Li and Prof. Yanmin Qian. Friends call me Hugo.
My research centers on test-time computation and efficient reasoning in large language models: RL-based post-training, inference-time reasoning control, and test-time training and adaptation. More broadly, I want to understand the principles of computation and inference that lead to efficient, generalizable intelligence across modalities.
I am always open to collaboration and academic discussion. Feel free to reach me at hugo0713@sjtu.edu.cn or qhu99@wisc.edu, or on WeChat (ID: HUGO--2025).
I am applying to PhD programs for Fall 2027.
Research interests
- Test-time computation and test-time training
- Efficient reasoning and RL-based post-training of LLMs
- Long-context modeling with fast-weight memory

Universal Test-Time Training
* Equal contribution
COLM 2026 Workshop on Efficient Reasoning SpotlightUnder review at ICLR 2027
Shares test-time-training fast weights across the whole depth stack instead of keeping one private state per layer, which turns the model into a single recurrence over (chunk, layer) and improves long-context retrieval at equal state size.

Test-Time Generation
Zefan Cai,Qinzhe Hu,Yuchen Zhu,Hao Tan,Tianyuan Zhang,Sai Bi,Haoyi Qiu,Yicong Hong,Haozhe Zhao,Jason Kuen,Haoliang Wang,Ruiyi Zhang,Ryan A. Rossi,Wanrong Zhu,Molei Tao,Wen Xiao,Junjie Hu,Jiuxiang Gu Under review at ICLR 2027
Reads a fast-weight memory by generating from it: a key-conditioned denoising readout samples a value instead of regressing to the conditional mean, so a compressed memory can recall competing values instead of their average, improving long-context retrieval and novel view synthesis at the same state size.

Last Layer KV Cache is All You Need
Zefan Cai,Sai Bi,Lin Zhang,Yicong Hong,Yuchen Zhu,Lu Li,Haozhe Zhao,Hanwen Jiang,Qinzhe Hu,Zeyi Huang,Shawn Lin,Cheng Luo,Chongjian GE,Rui-Jie Zhu,Xun Huang,Jiuxiang Gu,Junjie Hu,Hao Tan Under review at ICLR 2027
Lets every layer read only the last layer's KV cache of earlier chunks, turning the Transformer into a chunk recurrence across depth and time: long-term KV storage shrinks toward 1/L while RULER retrieval stays the best among the compared methods at 124M and 760M.

TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation
INTERSPEECH 2026 Oral
A sparse Mixture-of-Experts framework that routes experts over time frames and mel bands, raising the capacity of speech-separation models with almost no extra inference cost (+3.8 dB SDR over BSRNN on Libri2Mix at a comparable 4.1 GMACs/s).

SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning
ICML 2026
GRPO-based efficient reasoning with progressive chain-of-thought length calibration: it estimates the optimal reasoning length for each question and adapts the length-reward coefficient, cutting reasoning tokens by up to 52% while improving accuracy on benchmarks such as AIME25.
Experience
Research intern, advised by Prof. Junjie HuMadison, WI, USA · on-site since Jul 2026 Three projects on long-context memory, all under review at ICLR 2027. Universal Test-Time Training (co-first author): shared fast-weight memory across time and depth—one state pool for the whole stack instead of a private state per layer; at matched state and active compute, uTTT-MoE gains 2.6 / 2.1 RULER points over layer-private TTT-MoE at 124M / 760M. Test-Time Generation (second author): generate each memory read from fresh noise, guided by the query, so a compressed memory can recover alternative values instead of their average. LastKV: let every layer read a shared last-layer KV history, feeding deep past representations back into future computation and cutting long-term KV storage toward 1/L of a standard Transformer.
Reinforcement learning and post-training for efficient LLM reasoning, including chain-of-thought compression and on-policy distillation.
Co-built JEDI, an AI participant that joins Jitsi meetings and decides when to speak; owned its Kubernetes-native data layer (Ingress → PostgREST → PostgreSQL).
Making speech-separation models cheaper to run without shrinking their capacity, via sparse MoE routing over temporal frames and mel-frequency bands.
Education

上海交通大学 · ACM Honors Class, Zhiyuan College
B.S. in Computer ScienceShanghai, China

新加坡国立大学 · School of Computing · Summer Workshop 2025
Cloud Computing track, grade A+Singapore
First Prize of the course for the group project
JEDI 
湖南师大附中
High schoolChangsha, Hunan, China
Honors
- Zhiyuan Honor Scholarship, awarded to the top 5% of students at SJTU
Service
- Reviewer: COLM 2026, ICLR 2027
Latest notes
Course notes and research notes. English translations of the Chinese originals.
All posts →