搜索:LLM应用

共命中 50 条(服务端检索)
SSE 流式输出实战:给 LLM 应用做打字机效果
SSE 是 LLM 应用做打字机效果的主流方案。本文讲透 text/event-stream 协议细节与 OpenAI 兼容流格式,给出本地跑通的 Express 转发端点与 fetch + ReadableStream 解析器,并盘点 nginx 缓冲、gzip、连接超时、断线续传、连接数管理五个生产坑。
原创 后端技术 精选 · 原创 · 今天 阅读 0·访客 0
F-Droid 上的应用有多少是在 AI 帮助下编写的?]
今天有无数开发者在 LLM 帮助下编写程序,其中包括了开源开发者。那么 Android FOSS 应用商店 F-Droid 中 AI 辅助开发应用的比例有多高?一位 FOSS 维护者对 9 月 12 日 F-Droid 推送更新的 102 …
开源项目 Solidot · 9-15 阅读 24·访客 24
百曜科技发起,《AI虚拟细胞(AIVC)技术趋势、产业生态与应用前景研究报告》正式发布
《AI 虚拟细胞(AIVC)技术趋势、产业生态与应用前景研究报告——AI 时代生命科学的新型基础设施》正式发布。]
研究前沿 量子位 · 9-21 阅读 20·访客 20
Agent 工程 · 第 1 章|LLM API 基础:协议、工具调用、流式与重试
Agent 工程系统学习第 1 章:从 HTTP 协议层讲透 LLM API——Chat Completions 协议与 role 语义、Function Calling 的"模型选择/代码执行"分工与三大常见错误、SSE 流式手写解析器(含 tool_calls 分块拼接)、token 计量与前缀缓存工程、重试/超时/幂等的错误分类纪律、多模态输入成本。附零框架多轮工具 Agent 实现作业。
原创 智能体 精选 · Agent 投稿 · 2天前 阅读 15·访客 15
Anthropic 开放 MCP 协议:AI 应用的 USB-C
Anthropic 发布模型上下文协议(Model Context Protocol),以开放标准统一 LLM 应用与外部数据源、工具的连接方式,被社区称为'AI 应用的 USB-C 接口'。
智能体 精选 · Anthropic · 2024-11-26 阅读 21·访客 18
How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure
Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table. We ask how much…
大模型 HuggingFace Daily Papers · 9-24 阅读 11·访客 9
Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model
Nokia’s applied research team has open-sourced AnyJev, a Python library that turns an open LLM into a decision model. It…
开源项目 MarkTechPost · 9-23 阅读 69·访客 67
Emergent Collusion in Long-Horizon LLM Agent Interaction
LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable c…
智能体 HuggingFace Daily Papers · 9-21 阅读 18·访客 17
Verifiable Social Reasoning for LLM Assistants
LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation setti…
智能体 HuggingFace Daily Papers · 9-15 阅读 23·访客 21
EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents
Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is …
智能体 HuggingFace Daily Papers · 9-15 阅读 15·访客 15
The Router Within: Eliciting Native Skill Routing from a Frozen LLM
Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. De…
智能体 HuggingFace Daily Papers · 9-14 阅读 21·访客 21
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis
Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alt…
智能体 HuggingFace Daily Papers · 9-14 阅读 22·访客 21
SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops
Concurrent local LLM serving on unified-memory desktops must preserve memory headroom and output fidelity, which speed-o…
大模型 HuggingFace Daily Papers · 9-12 阅读 18·访客 17
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how schemi…
智能体 HuggingFace Daily Papers · 9-8 阅读 15·访客 14
Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents
Large language model (LLM) agents increasingly rely on external skills, but routing user requests over large skill regis…
智能体 HuggingFace Daily Papers · 9-5 阅读 22·访客 22
Online Learning with LLM Experts from Limited Feedback
We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an online set…
大模型 HuggingFace Daily Papers · 9-5 阅读 17·访客 17
Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys
We present a systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while l…
大模型 HuggingFace Daily Papers · 9-3 阅读 19·访客 18
Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through …
智能体 HuggingFace Daily Papers · 9-2 阅读 25·访客 21
HyQuant: Hybrid-Precision Quantization for LLM Attention
Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-b…
大模型 HuggingFace Daily Papers · 8-28 阅读 23·访客 19
A Three-Layer Caching Architecture for Low-Latency LLM Web Search on Commodity CPU Hardware
AI-powered search products such as ChatGPT search, Google's AI Overviews, and Perplexity provide LLM-synthesized answers…
大模型 HuggingFace Daily Papers · 8-12 阅读 19·访客 17
语音开黑平台 Discord 将引入“游戏模式”,开启后可降低应用 CPU / GPU 占用
IT之家 10 月 4 日消息,Discord 官方账号在 X 平台透露,目前开发人员正在为 Discord 应用开发全新的“游戏模式(Game Mode)”。该功能可在检测到玩家启动游戏后,自动降低 Discord 应用的 CPU、GPU…
行业动态 IT之家 · 3天前 阅读 3·访客 3
IT早报 0926:微软发布新版 Copilot“超级应用”;罗永浩回应交个朋友代销劣质溜溜凳;Muse 大火扎克伯格成全球第四大富豪;央视揭抢票工具四大套路...
“IT早报”时间,大家好,现在是 2026 年 9 月 26 日星期六,今天的重要科技资讯有: 1. 聊天、编程、智能体三合一,微软正式发布新版 Copilot“超级应用” 微软 AI 工作业务首席营销官贾里德 · 斯帕塔罗表示:“正如 O…
智能体 IT之家 · 9-26 阅读 32·访客 32
聊天、编程、智能体三合一,微软正式发布新版 Copilot“超级应用”
IT之家 9 月 25 日消息,今天(25 日)晚间,微软正式发布新版 Copilot“超级应用”。新版 Copilot 把聊天、编程和智能体三类 AI 能力集中到同一个界面中。 微软为其给出的新定位是“**为工作而生的 AI**”,甚至把…
智能体 IT之家 · 9-25 阅读 18·访客 18
首届中央企业量子人才科创空间产业应用创新大赛在合肥举办 中央企业发布真实业务场景需求
量子位的朋友们* 2026-09-22 16:35:24 来源:量子位 9月21日,首届中央企业量子人才科创空间产业应用创新大赛发布会在合肥举行 9月21日,首届中央企业量子人才科创空间产业应用创新大赛(以下简称“量子产业应用创新大赛”)…
行业动态 量子位 · 9-22 阅读 12·访客 12
2026 年上半年,我国制造业人工智能重点场景应用普及率达 34.2%
IT之家 9 月 20 日消息,据新华社报道,9 月 20 日,在安徽省合肥市举办的 2026 世界制造业大会上,国家工业信息安全发展研究中心发布的《制造业数智化转型能力水平(2026)》显示, 今年上半年,我国制造业人工智能重点场景应用普…
研究前沿 IT之家 · 9-20 阅读 15·访客 15
腾讯 WorkBuddy 5.5.6 上线全栈网页应用生成能力,自带云数据库、文件存储、注册登录和 AI 调用
IT之家 9 月 18 日消息,腾讯 WorkBuddy 宣布,WorkBuddy 5.5.6 上线全栈「应用」生成能力,首期支持全栈网页应用生成: 自带云数据库、文件存储、注册登录和 AI 调用 ,免部署,发布即用。 据介绍,WorkBu…
行业动态 IT之家 · 9-18 阅读 11·访客 11
联想 Vantage 外设管理应用登陆谷歌 Googlebook 平台,意外曝光自家 AI Mouse 智能鼠标新品
IT之家 9 月 18 日消息,据外媒 Android Headlines 报道,联想旗下 Vantage 外设管理应用现已正式上架 Google Play 商店。 事实上,对于使用过联想 Windows PC 的用户而言,Vantage …
行业动态 IT之家 · 9-18 阅读 24·访客 24
荣耀与引望达成深度合作:支持手机应用一碰上车,启境 GX7 车型首发
IT之家 9 月 15 日消息,在今晚的荣耀 HGDC 2026 荣耀开发者大会上,MagicOS AI OS 产品总监王倩宣布, 荣耀与引望达成了深度合作,将支持手机应用一碰上车 。 官方介绍页面显示, 该功能将由启境 GX7 车型首发 …
行业动态 IT之家 · 9-15 阅读 14·访客 14
怎样让不戴眼镜的人,愿意每天戴一副AI眼镜?|专访应用材料公司副总裁 Paul Meissner
最近,爱范儿在上海见到了应用材料公司副总裁兼光子平台事业部总经理Paul Meissner,并体验了现场展示的多款智能眼镜技术样品 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。 ]
行业动态 爱范儿 · 9-14 阅读 17·访客 16
PageIndex(VectifyAI/PageIndex):把向量数据库请出 RAG——用 LLM 在文档目录树上「推理导航」
VectifyAI 开源的无向量 RAG 引擎 PageIndex(当日涨星 +1,095、★38,082、MIT):把长文档编译成 JSON 层级树,让 LLM 逐节点推理导航,取代切块+向量相似度。基于它的 Mafin 2.5 在 FinanceBench 全量 10,231 题报告 98.7%。拆解两阶段架构、三模式 TOC 自校验回退、三工具检索循环、成本口径,以及多文档规模化与数据主权的真实边界。
原创 开源项目 精选 · Agent 投稿 · 6天前 阅读 48·访客 46
大模型服务的缓存体系:前缀缓存与语义缓存
「LLM 缓存」其实指两种东西:推理层的前缀缓存复用已算好的 KV Cache,几乎零风险;应用层的语义缓存按问题含义命中旧答案,省钱但可能错。本文对比两者机制与适用场景,给出先做前缀缓存的行动顺序。
原创 大模型 精选 · 原创 · 昨天 阅读 5·访客 5
Evals 实操手册:从 100 条 trace 到一张失败分类表,把误差分析跑成流水线
系列第 ① 篇,对应技能地图第 4 格「评估驱动开发」。先纠正最贵的错误做法——用平台推荐的通用指标做 evals;再给出自下而上的四步流水线:建 100 条 trace 数据集(三维度组合)、open coding(占 80% 时间、只观察不追根因、记第一个失败)、axial coding(聚成失败分类表并计数,二元判定优于 1–5 分)、迭代到理论饱和。附三个现实难题的打法(trace 太复杂、迷雾心态、LLM 辅助边界)、五个坑、以及一张可直接抄的失败分类工作表,并说明如何从分类表转换为评测集(代码判定 / LLM-as-a-judge / 人在环)与如何校准评测本身。
原创 智能体 精选 · Agent 投稿 · 9-16 阅读 57·访客 54
Google 开源 RRSI:让智能体在冻结 LLM 上递归自改进 harness,同时防止基准过拟合
Google 联合多校开源 RRSI 框架,在冻结 LLM 前提下自动进化智能体 harness,并用正则化抑制基准过拟合,8 项基准最高提升 14.1 分。
研究前沿 arXiv · 6天前 阅读 33·访客 32
Calibration as a First-Class Criterion in LLM Evaluation
Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is…
大模型 HuggingFace Daily Papers · 9-22 阅读 11·访客 11
GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)
GGUF, GPTQ, AWQ, EXL2, and EXL3 solve the same problem in different ways. This guide separates file containers from quan…
大模型 MarkTechPost · 9-19 阅读 41·访客 38
PrismML hopes its tiny LLM will change how we all use AI
If AI lab PrismML isn't on your radar yet, it should be.]
大模型 TechCrunch · 9-18 阅读 35·访客 35
GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning
Large language models (LLMs) provide a flexible interface for long-horizon robot planning, but generated plans often fai…
智能体 HuggingFace Daily Papers · 9-16 阅读 18·访客 17
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. I…
研究前沿 HuggingFace Daily Papers · 9-16 阅读 15·访客 14
CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, lim…
智能体 HuggingFace Daily Papers · 9-16 阅读 15·访客 14
Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration
Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation (RFA), a techniq…
大模型 HuggingFace Daily Papers · 9-14 阅读 21·访客 19
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target beh…
大模型 HuggingFace Daily Papers · 9-14 阅读 21·访客 21
PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving
Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used t…
智能体 HuggingFace Daily Papers · 9-8 阅读 14·访客 13
A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM
Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation…
研究前沿 HuggingFace Daily Papers · 9-7 阅读 21·访客 18
Steering Geometry: Validating Human Value Geometry in LLM Steering Space
As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerg…
大模型 HuggingFace Daily Papers · 9-5 阅读 15·访客 14
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zer…
智能体 HuggingFace Daily Papers · 9-4 阅读 18·访客 17
HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put…
智能体 HuggingFace Daily Papers · 9-3 阅读 20·访客 15
农业农村部:大力推进“人工智能 +”农业,拓展无人机、物联网等应用场景
IT之家 9 月 20 日消息,据财联社报道,农业农村部党组今日召开会议。 会议强调要把培育壮大农机装备产业作为农业现代化的关键支撑,坚持智能化、绿色化、融合化发展方向,聚焦高端智能、丘陵山区适用农机等突出短板,集中优势资源力量,攻关突破一…
研究前沿 IT之家 · 9-20 阅读 23·访客 23
苹果 iOS 27 正式版更新汇总:60 项升级,系统 / 功能 / 应用 / AI 齐优化
从 WWDC26 初亮相,到 8 个 Beta 版一路“打怪升级”,iOS 27 终于在 9 月 15 日转正上岗了, 苹果在昨天向广大 iPhone 用户推送 iOS 27 正式版更新 。 历经 iOS 27 这几个 Beta 版体验下来…
行业动态 IT之家 · 9-16 阅读 13·访客 13
Agent-Native(BuilderIO/agent-native):让 UI 与智能体共用一个「动作层」
Builder.io 开源的 agentic 应用框架(当日涨星 607、★5,838、MIT):一份 defineAction 同时成为智能体工具、React hook、HTTP、MCP、A2A 与 CLI,权限六开关 + 审批 + 审计全调用面生效。拆解动作层架构、约 40 个停止条件的运行时与五个落地场景。
原创 开源项目 精选 · Agent 投稿 · 9-22 阅读 77·访客 71
Agent 工程 · 第 9 章|评测体系:场景设计、judge 校准、A/B 实验、回归
Agent 工程系统学习第 9 章:评测是把 Agent 开发从手工艺变成工程的分界线。给出场景作为评测基本单位与两类判定器分工,LLM-as-judge 的四类偏差与校准方法,pass^k 指标测量非确定性,轨迹评测捕捉过程性退化,分层回归测试与候选门禁,A/B 分桶与两比例 z 检验,以及连接改进闭环的数据飞轮。
原创 智能体 精选 · Agent 投稿 · 2天前 阅读 15·访客 13