Xinming Wei

Xinming WeiLucas

Ph.D. Candidate in Computer Science/Peking University

I am a final-year Ph.D. student in Computer Science at Peking University, advised by Prof. Guojie Luo at the Center for Energy-Efficient Computing and Applications (CECA). I received my B.S. in School of EECS, Peking University in 2022.

I build machine learning systems, optimizing the full stack — from heterogeneous hardware to large-scale distributed systems — for efficient and scalable LLM training and inference. My recent focus is rack-scale systems (aka “超节点”) that carry production deployment of trillion-parameter MoE models, alongside large-scale RL systems and efficient multimodal models. Earlier, I worked on parallel acceleration and efficient algorithms for design automation.

Xinming Wei

Experience

Dots Infra, RedNote

Aug 2025 – Jun 2026

Research Intern, LLM Training Infrastructure

RedStar Intern

Systems Research Group, Microsoft Research Asia

Aug 2023 – Nov 2023

Research Intern, LLM Reasoning and Autoformalization

Stars of Tomorrow (Excellent Intern)Mentors: Xian Zhang and Fan Yang

Center for Energy-Efficient Computing and Applications, Peking University

Sep 2021 – Present

Research Assistant, Computer Architecture and Systems

News

All news
  1. Awarded the Tencent Hunyuan Ph.D. Fellowship (中国电子学会-腾讯博士生科研激励计划·混元大模型专项)

  2. We open-source UltraEP, a production-ready expert load balancing library

    Our real-time expert load balancer for large-scale MoE training and inference, running in production with near-optimal balancing, is now public. Read more

Selected Publications

Full list

UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing

Xinming Wei, Chao Jin, Tuo Dai, Yinmin Zhong, Shan Yu, Chengxu Yang, Bingyang Wu, Zili Zhang, Jing Mai, Qianchao Zhu, Zhouyang Li, Yuliang Liu, Guojie Luo

Preprint 2026

Expert load imbalance caps MoE training and inference throughput under large-scale expert parallelism. UltraEP is the first balancer to act on exact post-gating load in real time — replicating hot experts and rerouting tokens inside every layer of every microbatch — holding critical-path cost under 300 µs and reaching 94.3% of the force-balanced ideal throughput.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning

Chao Jin, Xinming Wei, Yinmin Zhong, C. Yang, B. Wu, R. Zhu, Z. Zhang, Yuliang Liu, Xin Jin

Preprint 2026

MoE RL shows significant expert load imbalance: its data is domain-skewed and load distribution swings hard. ReLibra leverages the replay of routing decisions already observed during rollout to plan balancing ahead of time for the RL training phase.

Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC

Xinming Wei, Jiahao Zhang, Haoran Li, J. Chen, H. Guan, R. Qu, M. Li, X. Chen, Guojie Luo

Preprint 2025

Agentic workloads put latency-critical turns and throughput-bound background work on the same heterogeneous SoC. Agent.xpu unleashes NPU and GPU with phase disaggregation and elastic tensor chunking, adds kernel-level preemption with slack-aware piggybacking, and dispatches kernels with bandwidth awareness, keeping interaction responsive without leaving accelerators idle.

AceRoute: Adaptive Compute-Efficient FPGA Routing with Pluggable Intra-Connection Bidirectional Exploration

Xinming Wei, Ziyun Zhang, Sunan Zou, Kaiwen Sun, Jiahao Zhang, Jiaxi Zhang, Ping Fan, Guojie Luo

ICCAD 2024

FPGA routing is essentially node-disjoint path finding on a DAG, where congested regions blow the search space up exponentially. AceRoute pairs adaptive bidirectional A* with fine-grained intra-connection parallelism and a lock-free multi-threaded runtime for substantial speedups.

Rethinking IC Layout Vulnerability: Simulation-Based Hardware Trojan Threat Assessment with High Fidelity

Xinming Wei, Jiaxi Zhang, Guojie Luo

IEEE S&P 2024

Outsourced fabrication lets a foundry slip hardware Trojans into a finalized layout. SiliconCritic reproduces black-box, foundry-level attacks and post-fabrication analysis at design time, quantifying insertion difficulty from the resulting shifts in timing and power side channels.

GDSII-Guard: ECO Anti-Trojan Optimization with Exploratory Timing-Security Trade-Offs

Xinming Wei, Jiaxi Zhang, Guojie Luo

DAC 2023

Hardening a layout against Trojan insertion normally costs timing. GDSII-Guard frames it as ECO-style incremental edits over a connected-graph layout representation, using greedy cell shift, clustering and wire scaling to walk the security–performance trade-off.