Poster
in
Workshop: Methods and Opportunities at Small Scale (MOSS)

TinyServe: Query-Aware Cache Selection for Efficient LLM Inference

Dong Liu ⋅ Yanxuan Yu

Keywords: Efficient serving LLMs system behavior Cache management small LLMs Token selection

Project Page [ OpenReview]

Abstract

Serving large language models (LLMs) efficiently remains challenging due to the high memory and latency overhead of key-value (KV) cache access during autoregressive decoding. We present \textbf{TinyServe}, a lightweight and extensible runtime system for deploying tiny LLMs (e.g., TinyLLaMA, GPT2-345M) with support for structured KV sparsity, plugin-based token selection, and hardware-efficient attention kernels. Unlike prior simulation frameworks, TinyServe executes real-time decoding with configurable sparsity strategies and fine-grained instrumentation.To reduce decoding cost, we introduce a \textit{query-aware page selection} mechanism that leverages bounding-box metadata to estimate attention relevance between the query and KV cache blocks. This enables selective KV loading with minimal overhead and no model modifications. Our fused CUDA kernel integrates page scoring, sparse memory access, and masked attention in a single pass.Experiments show that TinyServe achieves up to \textbf{3.4×} speedup and over \textbf{2×} memory savings with negligible accuracy drop. Additional analysis of cache reuse, page hit rate, and multi-GPU scaling confirms its practicality as a system-level testbed for LLM inference research on resource-constrained hardware.

Chat is not available.