[1] Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., Kaiser Ł., andPolosukhin I., 2017. Attention is all you need.Advances in Neural Information Processing Systems, 30. [2] Kwon W., Li Z., Zhuang S., Sheng Y., Zheng L., Yu C.H., Gonzalez J., Zhang H., andStoica I., 2023. Efficient memory management for large language model serving with paged attention. InProceedings of the 29th Symposium on Operating Systems Principles, pp. 611-626. [3] Yang H., Zhang R., Huang M., Wang W., Tang Y., Li Y., Liu Y., andZhang D., 2025. Kvshare: an LLM service system with efficient and effective multi-tenant kv cache reuse.Arxiv Preprint Arxiv:2503.16525. [4] Aminabadi R.Y., Rajbhandari S., Awan A.A., Li C., Li D., Zheng E., Ruwase O., Smith S., Zhang M., Rasley J., andHe Y., 2022. Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale. InSC22: International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1-15. [5] Sheng Y., Zheng L., Yuan B., Li Z., Ryabinin M., Chen B., Liang P., Ré C., Stoica I., andZhang C., 2023. Flexgen: high-throughput generative inference of large language models with a single gpu. InInternational Conference on Machine Learning, pp. 31094-31116. [6] Walia E.,2002. Operating system concepts. Khanna Publishing House. [7] Chu K., Shen Z., Cheng S.R., Xiang D., Liu Z., andZhang W., 2025. Mcam: efficient LLM inference with multi-tier kv cache management. In2025 IEEE 45th International Conference on Distributed Computing Systems (ICDCS), pp. 571-581. [8] Qiu S., Hu Y., Wang X., Zhu W., Yan J., Chen H., Xu K., Chen K., andZhang Y., 2026. Tutti: making SSD-backed KV cache practical for long-context LLM serving.Arxiv Preprint Arxiv:2605.03375. [9] Chen T., Xu B., Zhang C., andGuestrin C., 2016. Training deep nets with sublinear memory cost.Arxiv Preprint Arxiv:1604.06174. [10] Liang X., Yao L., Wu S., Li Y., andXu Y., 2023. Care: A cost-aware eviction strategy for improving throughput in cloud environments. In2023 IEEE 29th International Conference on Parallel and Distributed Systems (ICPADS), pp. 2269-2276. [11] Child R., Gray S., Radford A., andSutskever I., 2019. Generating long sequences with sparse transformers.Arxiv Preprint Arxiv:1904.10509. [12] Beltagy I., Peters M.E., andCohan A., 2020. Longformer: the long-document transformer.Arxiv Preprint Arxiv:2004.05150. [13] Chu K., Lin Z., Xiang D., Shen Z., Su J., Chu C., Yang Y., Zhang W., Wu W., andZhang W., 2025. Selective KV-cache sharing to mitigate timing side-channels in LLM inference.Arxiv Preprint Arxiv:2508.08438. [14] Yoon D., Min Y., Kim H., Noh S.H., andKim J., 2025. TraCT: disaggregated LLM serving with CXL shared memory KV cache at rack-scale.Arxiv Preprint Arxiv:2512.18194. [15] Touvron H., Lavril T., Izacard G., Martinet X., Lachaux M.A., Lacroix T., Rozière B., Goyal N., Hambro E., Azhar F., andRodriguez A., 2023. Llama: open and efficient foundation language models.Arxiv Preprint Arxiv:2302.13971. [16] Bai Y., Lv X., Zhang J., Lyu H., Tang J., Huang Z., Du Z., Liu X., Zeng A., Hou L., andDong Y., 2024. Longbench: A bilingual, multitask benchmark for long context understanding. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3119-3137. [17] Liu T., Xu C., andMcAuley J., 2024. Repobench: benchmarking repository-level code auto-completion systems. InInternational Conference on Learning Representations, 2024, pp. 47832-47850. [18] Huang L., Cao S., Parulian N., Ji H., andWang L., 2021. Efficient attentions for long document summarization. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 1419-1436. [19] Zaheer M., Guruganesh G., Dubey K.A., Ainslie J., Alberti C., Ontanon S., Pham P., Ravula A., Wang Q., Yang L., andAhmed A., 2020. Big bird: transformers for longer sequences.Advances in Neural Information Processing Systems, 33, pp. 17283-17297. [20] Heo B., Park S., Han D., andYun S., 2024. Rotary position embedding for vision transformer. InEuropean Conference on Computer Vision, pp. 289-305. [21] Gu A., andDao T., 2023. Mamba: linear-time sequence modeling with selective state spaces.Arxiv Preprint Arxiv:2312.00752. |