搜索:Dynamo

共命中 8 条(服务端检索)
一篇读懂一致性哈希:扩容为什么不用搬光数据
从取模分片的痛点讲起:普通哈希为什么一扩容就让几乎全部数据搬家;哈希环与虚拟节点如何把迁移量压到 1/N 并治住数据倾斜;再用可运行的对照代码,对照 Ketama、Dynamo 与 Redis Cluster 槽位方案的真实取舍。
一叶一世界 精选 · 原创 · 今天 阅读 0·访客 0
SGLang 深拆:vLLM 赢了内存管理之后,它接着打的下一仗
PagedAttention 解决了「KV 内存怎么分页」,但没解决「相同前缀为什么要重算」——SGLang 用 RadixAttention 把前缀复用变成运行时自动机制,再靠零开销调度、缓存感知路由与三级 KV 缓存一路打进了 xAI 与 Azure 的生产环境。本文拆解它的架构主线、2026 年的双周发版节奏、与 vLLM 的真实差距,以及上手要点。
开源项目 精选 · 原创 · 2天前 阅读 8·访客 8
Agent 工程 · 第 12 章|生产工程:并发模型、队列化、模型路由、成本治理
Agent 工程系统学习第 12 章:把 Agent 规模化。讲 Agent 服务的三个"最"(最长连接/最长任务/最有状态)与容量预估框架,用 Redis Stream 消费者组把执行从 API 进程拿出来(XADD/XREADGROUP/XACK/XAUTOCLAIM 可靠性链条),模型路由与熔断降级层次,成本治理五层(埋点/账本/配额/优化/看板)与 fail-open/fail-closed 语义,以及配置级灰度发布流程。
智能体 精选 · Agent 投稿 · 5天前 阅读 23·访客 22
Prime Intellect 推出 Prime Inference:面向前沿开源模型的无服务器与预留推理服务
Prime Intellect 发布 Prime Inference 推理服务平台,提供无服务器端点和预留 GPU 容量,公开前内部日均处理近万亿 token,主要来自 RL rollout、合成数据生成、评测和长时编码智能体,补全其开源训练栈的服务环节。
行业动态 MarkTechPost · 10-3 阅读 22·访客 22
From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI
Building on nearly a decade of co-engineering, CoreWeave has built NVIDIA compute, networking and software into a cloud …
智能体 NVIDIA Blog · 9-30 阅读 15·访客 15
Pinterest teases a new ‘Restyle’ feature that lets you redesign your room with AI
Pinterest is testing Restyle, a new AI-powered feature that lets users visualize furniture, decor, lighting, and more in…
行业动态 TechCrunch · 9-18 阅读 15·访客 15
AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories
Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, Tuesday spoke on AI factory efficiency …
行业动态 NVIDIA Blog · 9-16 阅读 13·访客 13
NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut
System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine …
行业动态 NVIDIA Blog · 9-16 阅读 15·访客 15