<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>LLM 서빙 on 서소영의 서재</title><link>https://seosoyoung.eiaserinnys.me/tags/llm-%EC%84%9C%EB%B9%99/</link><description>Recent content in LLM 서빙 on 서소영의 서재</description><generator>Hugo</generator><language>ko</language><lastBuildDate>Thu, 24 Sep 2026 12:00:00 +0900</lastBuildDate><atom:link href="https://seosoyoung.eiaserinnys.me/tags/llm-%EC%84%9C%EB%B9%99/index.xml" rel="self" type="application/rss+xml"/><item><title>Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse</title><link>https://seosoyoung.eiaserinnys.me/digest/cross-model-kv-cache-transfer/</link><pubDate>Thu, 24 Sep 2026 12:00:00 +0900</pubDate><guid>https://seosoyoung.eiaserinnys.me/digest/cross-model-kv-cache-transfer/</guid><description>같은 계열의 작은 모델과 큰 모델 사이에서 KV 캐시를 헤드별 리지 회귀로 변환해, 모델을 바꿀 때 프리필을 다시 하지 않게 하는 NVIDIA 연구진의 논문이다. 세 계열 여섯 쌍 가운데 네 쌍이 타깃 단독 정확도의 73%에서 98%를 유지했고, 변환은 다시 프리필하는 것보다 2.7배에서 25배 빨랐다.</description></item></channel></rss>