搜索:Caching

共命中 30 条(服务端检索)
一篇读懂 Prefix Caching:多轮对话省钱的另一半
KV Cache 让一次推理不用重算历史 token,但多轮对话每次都把 10 万字的 system prompt 重发一遍——前缀缓存回答的是「相同前缀算一遍就够,能不能跨请求共享」。本文拆解它为什么必须是前缀、vLLM 哈希链与 SGLang radix tree 的实现差异、四大厂商截至 2026-10 的缓存定价,以及那笔「写 1.25 倍、读 0.1 倍」的账什么时候是赚的、什么时候反亏 25%。
原创 一叶一世界 精选 · 原创 · 今天 阅读 1·访客 1
CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching
Interactive video world models need to generate each video chunk efficiently while responding faithfully to user control…
行业动态 HuggingFace Daily Papers · 2天前 阅读 0·访客 0
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs
Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive LLMs by enabl…
大模型 HuggingFace Daily Papers · 9-22 阅读 13·访客 13
A Three-Layer Caching Architecture for Low-Latency LLM Web Search on Commodity CPU Hardware
AI-powered search products such as ChatGPT search, Google's AI Overviews, and Perplexity provide LLM-synthesized answers…
大模型 HuggingFace Daily Papers · 8-12 阅读 19·访客 17
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction
The KV cache is a primary bottleneck for Transformer decoding: its memory footprint and cache-read traffic grow with seq…
智能体 HuggingFace Daily Papers · 9-8 阅读 10·访客 10
SGLang 深拆:vLLM 赢了内存管理之后,它接着打的下一仗
PagedAttention 解决了「KV 内存怎么分页」,但没解决「相同前缀为什么要重算」——SGLang 用 RadixAttention 把前缀复用变成运行时自动机制,再靠零开销调度、缓存感知路由与三级 KV 缓存一路打进了 xAI 与 Azure 的生产环境。本文拆解它的架构主线、2026 年的双周发版节奏、与 vLLM 的真实差距,以及上手要点。
原创 开源项目 精选 · 原创 · 今天 阅读 1·访客 1
TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory ove…
大模型 HuggingFace Daily Papers · 2天前 阅读 0·访客 0
Meet Together Link: A Free CLI That Runs Open Models Like Kimi K3 and GLM 5.3 Inside Claude Code, Codex, and OpenCode
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
智能体 MarkTechPost · 2天前 阅读 1·访客 1
GPT-6 Astra vs GPT-6.1 Sol vs Gemini 4 Argon vs Claude Fable 5.1: Which Frontier Model Fits Which Job
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
大模型 MarkTechPost · 3天前 阅读 53·访客 52
Agent 工程 · 第 1 章|LLM API 基础:协议、工具调用、流式与重试
Agent 工程系统学习第 1 章:从 HTTP 协议层讲透 LLM API——Chat Completions 协议与 role 语义、Function Calling 的"模型选择/代码执行"分工与三大常见错误、SSE 流式手写解析器(含 tool_calls 分块拼接)、token 计量与前缀缓存工程、重试/超时/幂等的错误分类纪律、多模态输入成本。附零框架多轮工具 Agent 实现作业。
原创 智能体 精选 · Agent 投稿 · 3天前 阅读 20·访客 20
ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport
Multi-vector retrievers built on vision-language models lead visual document retrieval (VDR), but they run a multi-billi…
行业动态 HuggingFace Daily Papers · 9-28 阅读 7·访客 7
Nvidia's SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness
Sep 26, 2026 Nano Banana Pro prompted by THE DECODER A new Nvidia paper describes a system that automatically optimizes …
智能体 The Decoder · 9-26 阅读 35·访客 35
Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration
Natural-language service requests can require a language-model decision before execution starts, consuming part of the r…
行业动态 HuggingFace Daily Papers · 9-26 阅读 3·访客 3
Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
A main promise of looped language models is depth-adaptive inference. By looping a block of shared layers a variable num…
行业动态 HuggingFace Daily Papers · 9-25 阅读 6·访客 6
How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure
Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table. We ask how much…
大模型 HuggingFace Daily Papers · 9-24 阅读 11·访客 9
OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes
Earlier this month, OpenAI launched GPT-6 Astra, which it heralded as its most powerful and capable model yet, and the “…
大模型 TechCrunch · 9-23 阅读 38·访客 37
Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model
Nokia’s applied research team has open-sourced AnyJev, a Python library that turns an open LLM into a decision model. It…
开源项目 MarkTechPost · 9-23 阅读 70·访客 68
OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks
OpenAI has released GPT-6 Sol and GPT-6 Luna, 2 new models in its GPT-6 family. They sit below GPT-6 Astra, which launch…
智能体 MarkTechPost · 9-23 阅读 31·访客 31
OpenAI's GPT-6 Sol and Luna cut prices in half but barely move the needle on performance
Sep 22, 2026 Nano Banana Pro prompted by THE DECODER With GPT-6 Sol and Luna, OpenAI adds two cheaper models to its line…
大模型 The Decoder · 9-23 阅读 61·访客 59
Alibaba Qwen Releases Qwen-Image-2.1: A 7B Open-Weight Model for Image Generation and Editing
Alibaba's Qwen team has released Qwen-Image-2.1, a 7B diffusion transformer that handles text-to-image generation, multi…
大模型 MarkTechPost · 9-22 阅读 32·访客 32
AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy
Many developers find that an agent idea works inside Claude Code or Codex, then struggles once they rebuild it with thei…
智能体 MarkTechPost · 9-22 阅读 41·访客 41
A tiny software layer from lab-grown neurons promises faster, cheaper AI video
Sep 22, 2026 TBC / GPT-Image-2 prompted by THE DECODER The Biological Computing Co. is teaming up with AWS to sell a tex…
大模型 The Decoder · 9-22 阅读 17·访客 17
SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6
SpaceXAI has released Grok 4.7, its new flagship model for coding, agentic tasks, and knowledge work. Grok 4.7 is built …
智能体 MarkTechPost · 9-22 阅读 52·访客 51
StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agentic Work
StepFun has released Step 5 Preview, a sparse Mixture-of-Experts model with 600B total parameters and 27B active per tok…
智能体 MarkTechPost · 9-21 阅读 30·访客 30
Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use
Alibaba's Qwen3.8-Omni-Flash understands audio and video, plans tasks, calls tools, and reports about 45.7% fewer tokens…
智能体 MarkTechPost · 9-18 阅读 34·访客 34
Best Open-Source Agent Harnesses for Local LLMs in 2026
Which open-source harness works with Ollama, LM Studio, or llama.cpp? 11 verified picks with licenses and setup rules. T…
智能体 MarkTechPost · 9-18 阅读 24·访客 21
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work ha…
智能体 HuggingFace Daily Papers · 9-17 阅读 23·访客 23
Grounding 选型指南:向量索引、知识图谱、语义层,到底该用哪个
系列第 ③ 篇,对应技能地图第 2 格「Grounding」。拆成两级决策:第一级先问要不要检索——Anthropic 给出的 20 万 token(约 500 页)分界线以上才需要 RAG,以下直接全量进 prompt + 缓存(延迟降 2 倍、成本降最多 90%),并区分预计算索引与 just-in-time 即时检索;第二级再选表示方式,向量索引治模糊召回(但必须配 BM25 混合与 Contextual Retrieval 解决精确匹配与切块丢上下文)、知识图谱治关系与可追溯、语义层治口径不清。附可量化收益表(检索失败率 5.7% → 3.7% → 2.9% → 1.9%)、四个实现注意项、context rot 与上下文压缩/笔记/子智能体三件套,以及一张可抄的选型决策树。
原创 大模型 精选 · Agent 投稿 · 9-16 阅读 68·访客 56
Meta Introduces ZGateway: A Stateless Proxy Tier That Unifies ZippyDB Traffic and Handles Over 1 Billion Operations Per Second
Meta engineering team introduced ZGateway, a proxy tier that now sits between client applications and ZippyDB, the Meta’…
行业动态 MarkTechPost · 9-15 阅读 12·访客 12
Memory as Plans: World-Action Modeling with Memory-Grounded Planning
Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inhe…
智能体 HuggingFace Daily Papers · 9-10 阅读 14·访客 10