首页 > AI前沿 > From Overloaded to Guaranteed: High-Throughput Multi-SLO Enforcement for LoRA-Assisted On-Premise LLM Deployment

From Overloaded to Guaranteed: High-Throughput Multi-SLO Enforcement for LoRA-Assisted On-Premise LLM Deployment

arXiv自然语言 2026-10-04 13:04 3 阅读 查看原文

As Large Language Models (LLMs) become essential in privacy-sensitive sectors like hospitals and government agencies, the on-premise LLM servers offer a cost-effective and secure alternative to public cloud services.

However, these resource-constrained servers struggle to guarantee heterogeneous Service Level Objectives (SLOs) when serving multiple LoRA-adapted services simultaneously.

Existing serving frameworks suffer from severe SLO violations due to the computational overhead of LoRA layers and the rigid nature of batch scheduling.

To address this, we propose HALO, a scheduling method tailored for LoRA-assisted on-premise LLM deployment.

HALO introduces two key innovations: a spatial multiplexing strategy that overlaps Base and LoRA computations by partitioning GPU Streaming Multiprocessors (SMs), and an SLO-aware scheduler that decouples request execution based on "request-level slack."

By prioritizing urgent tasks and utilizing idle budget for traffic shaping, HALO significantly mitigates resource contention.

Our evaluation demonstrates that HALO minimizes SLO violations while improving throughput compared to state-of-the-art baselines.