CV

Dairu Liu

ironmt00@gmail.com
+86-175-0102-4327
Beijing, China

Research Interests

My primary research goal is to leverage large-scale data to build generalizable embodied agents capable of performing diverse real world tasks. I am particularly interested in Robot Learning (scaling up robot data, efficient robot learning algorithms, humanoid whole-body control, loco-manipulation) and Multimodal Reasoning (agentic perception, cross-modal understanding, long-horizon decision making).

Education

  • B.Eng. in Software Engineering
    Sep. 2023 - Jun. 2027 (expected)
    Nankai University, Tianjin, China
    GPA: 3.75; Top 15%
  • The Affiliated High School of Peking University, Beijing, China

Research Experience

  • Research Intern
    Dec. 2025 - Present
    GALBOT | IIIS, Tsinghua University
    Advisor: Prof. Li Yi
    • Built a large-scale humanoid motion tracking benchmark and proposed a preference-aligned reward model for whole-body control evaluation, addressing the misalignment between kinematic metrics and human perception. [1]
  • Research Intern
    Apr. 2024 - Present
    THUNLP | AIR, Tsinghua University
    Advisor: Prof. Peng Li, Prof. Yang Liu
    • Proposed Visual Abstract Thinking, a novel multimodal reasoning paradigm that enhances MLLMs via visual abstraction rather than verbose verbal rationales. [4]
    • Developed a 4D escape-room benchmark for evaluating time awareness and cross-modal active perception in Omni models. [5]
    • Exploring self-evolving agents that generalize from weak to strong supervision.

Skills

Robot Platforms

  • Unitree G1

Programming

  • C++
  • Python
  • MATLAB
  • LaTeX

Publications

  • HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
    2026
    ECCV 2026
    Dairu Liu*, Zekun Qi*, Jiayu Zeng*, Yu Guan, Chenghuai Lin, Xuchuan Chen, Xinqiang Yu, Wenyao Zhang, He Wang†, Li Yi†
    • Built a standardized humanoid motion tracking benchmark with 150 hours of newly captured optical mocap from 24 professional performers, comprising 24K clips organized into 4 motion families for fine-grained diagnosis.
    • Proposed HumanScore, a preference-aligned reward model trained on 3K pairwise human comparisons; achieves 0.8834 alignment rate with human judgments, outperforming MPJPE (0.8164), and validated on Unitree G1 real-robot deployment.
  • Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking
    2026
    CVPR 2026
    Zekun Qi*, Xuchuan Chen*, Dairu Liu*, Chenghuai Lin*, Yunrui Lian, Zhikai Zhang, Wenyao Zhang, Yu Guan, Jilong Wang, Xinqiang Yu, He Wang†, Li Yi†
    • Introduced Humanoid-GPT, a GPT style causal Transformer trained on an approximately 2B-frame retargeted motion corpus, distilling hundreds of RL motion experts via DAgger into a single general motion tracker with strong zero-shot generalization on Unitree G1.
    • Proposed Harmonic Motion Embedding (HME) for diversity-aware, distribution-balanced sampling; characterized scaling trends for data and model size and deployed with real-time inference (ONNX/TensorRT, <1.5 ms on RTX 4090).
  • LIMMT: Less is More for Motion Tracking
    2026
    ICML 2026
    Yu Guan*, Zekun Qi*, Chenghuai Lin, Xuchuan Chen, Dairu Liu, Wenyao Zhang, Jilong Wang, Xinqiang Yu, He Wang†, Li Yi†
    • First data centric study for humanoid motion tracking: General Quality Selection (GQS) scores motion along physics feasibility, diversity (HME embeddings), and complexity, then selects compact subsets via complexity-weighted farthest-point sampling.
    • Showed a strong less-is-more effect: training on approximately 3% of AMASS can outperform the full corpus across trackers (e.g., Any2Track, TWIST2); validated cross-dataset transfer and real-world Unitree G1 deployment without fine-tuning.
  • Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction
    2026
    COLM 2026 Under Review; NeurIPS 2025 Workshop
    Dairu Liu*, Ziyue Wang*, Minyuan Ruan*, Fuwen Luo, Chi Chen, Peng Li†, Yang Liu†
    • Proposed Visual Abstract Thinking (VAT), a reasoning paradigm that replaces verbose verbal CoT with visual abstractions, focusing models on semantic, structural, and geometric cues.
    • VAT achieves +2.21% average gain over the GPT-5 baseline, exceeding CoT (+0.96%); uses only 0.68x CoT runtime and 0.1x the tokens of Visual Sketchpad on GPT-4o. RL-trained VAT further gains +4.52% on Qwen-2.5-VL-3B.
  • Evaluating Time Awareness and Cross-modal Active Perception of Large Models via 4D Escape Room Task
    2026
    EMNLP 2026 Under Review
    Yurui Dong*, Ziyue Wang*, Dairu Liu*, Shuyun Lu*, Xuechen Liu*, Fuwen Luo, Peng Li†, Yang Liu†
    • Developed EscapeCraft-4D, a customizable 4D escape-room benchmark (66 scenes, 1,032 objects) incorporating trigger-based audio, misleading auditory distractors, and temporally transient clues for evaluating time-aware multimodal reasoning.
    • Introduced 4D specific process level metrics beyond final success rate: AMR (Audio Misleading Rate) and TCSS (Transient Clue Seeing Speed), revealing significant modality bias and failures in time-conditioned decision making across Omni models.

Honours and Awards

  • National Scholarship
    Sep. 2024
    Top 0.2%; Undergraduate, 10000CNY
  • Second Prize, NOIP 2018
    Dec. 2018
    National Olympiad in Informatics in Provinces
  • Second Prize, CSP-S 2019
    Dec. 2019
    CCF Certified Software Professional

Leadership and Services

  • Vice President
    Sep. 2025 - Present
    Quantitative Finance & Investment Society, Nankai University
  • Member
    2024 - 2025
    University Ski Team, Nankai University
  • Academic Department Vice-Head
    Sep. 2024 - Sep. 2025
    Psychology Association, Nankai University

Languages

  • Mandarin Chinese
    Native
  • English
    IELTS 7.0 (L 7.5, R 8.0, S 6.0, W 6.0); CET-6 608

Interests

  • Skiing