Portrait of Jiaxuan Chen is unavailable

Jiaxuan Chen

PhD - McGill University
Supervisor
Co-supervisor
Research Topics
Computer Systems
Energy Systems
Large Language Models (LLM)

Publications

DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs
Jianshu She
Rajat Ghosh
Karan Gupta
Qirong Ho
Xue Liu
LLM serving systems are provisioned for peak load to meet strict latency targets, leaving substantial GPU compute idle whenever traffic fall… (see more)s below peak. We present DeltaServe, a host-agnostic co-serving design that converts this idle inference capacity into LoRA fine-tuning throughput while preserving inference service-level objectives (SLOs). DeltaServe integrates with existing inference engines through a compact hook interface that requires only multi-LoRA batching support. It exploits the shared execution structure of inference prefill and LoRA fine-tuning forward passes, and uses an SLO-aware scheduler to admit and execute fine-tuning only when sufficient inference headroom is available. The scheduler is driven by a CUDA-graph-aware latency model calibrated offline and refined online. We integrate DeltaServe with vLLM, SGLang, and S-LoRA. On a production trace from Company X, DeltaServe on vLLM delivers 2.9x higher fine-tuning throughput than LLMStation at 100% inference SLO compliance, versus 85% for LLMStation. It also achieves 39% higher fine-tuning throughput than a baseline running vLLM+torchtune, using no additional hardware and maintaining full SLO compliance.
The Cost of Expertise: Understanding MoE Decode Performance
Sami Abuzakuk
Anne-Marie Kermarrec
Rafael Pires
Ramya Prabhu
Martijn de Vos