AI
AI
资讯
alishangtian.com
首页
大模型
智能体
开源项目
研究前沿
行业动态
专题
专题 · TOPICS
一叶一世界
28 篇
Agent 工程系统学习
14 篇
Agent 沙箱技术专题
13 篇
大模型基本功
24 篇
AI 推理与部署
17 篇
算法题解
32 篇
AI 编程实战
12 篇
RAG 实战手册
17 篇
后端技术
34 篇
模型微调实战
9 篇
推理模型
8 篇
端侧智能
7 篇
论文精读
11 篇
AI 行业观察
8 篇
AI 安全与攻防
18 篇
多模态之路
14 篇
具身智能与机器人
5 篇
AI for Science
7 篇
世界模型与视频生成
10 篇
AI 治理与合规
6 篇
开源模型全景
10 篇
Kubernetes 深入实践
6 篇
MLOps 与 LLM 工程化
4 篇
AI 音频与音乐
8 篇
全部专题 →
主题色 · THEME
靛蓝(默认)
极光
落日
薰衣草
海洋
森林
暮橙
石墨
自定义
恢复默认
提交线索
agent
anthropic
openai
huggingface daily papers
ai安全
it之家
jake wharton
solidot
gpu
港股
搜索:
Test-Time
共命中 50 条(服务端检索)
Test
-
Time
Scaling:让模型「多想一步」的三种姿势与一个天坑
推理时多花算力能换准确率:self-consistency 投票、过程奖励模型引导搜索、budget forcing 强制续想,三条路线各有边界。本文以 s1 论文与 2025–2026 年的 overthinking 研究为轴,梳理
test
-
time
scaling 的效果账与失效点。
原创
研究前沿
精选
· 原创 · 今天
阅读 1
·
访客 1
InterEvolve:
Test
-
Time
Evolution of Reward Programs for Humanoid Loco-Manipulation
We study
test
-
time
evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by re…
行业动态
HuggingFace Daily Papers · 6天前
阅读 6
·
访客 6
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM
Test
-
Time
Scaling
Test
-
time
scaling can improve large language model reasoning by generating and combining multiple candidate responses. I…
研究前沿
HuggingFace Daily Papers · 9-16
阅读 15
·
访客 14
When Agents Slow Down: Understanding LLM Agents'
Test
-
Time
Strategies via Elo-per-token Analysis
Large language model (LLM) agents allocate
test
-
time
compute adaptively as they revise solutions, use tools, explore alt…
智能体
HuggingFace Daily Papers · 9-14
阅读 22
·
访客 21
What Else Needs Fixing? Exploring Cost-Effective
Test
-
Time
Compute for Revision Propagation in Artifacts Generated Through Conversation
Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and revision in …
大模型
HuggingFace Daily Papers · 9-3
阅读 10
·
访客 9
Astra and Opus just passed Turing’s other
test
Computer pioneer Alan Turing is best known for his eponymous experiment to
test
whether artificial and human intelligenc…
行业动态
TechCrunch · 9-26
阅读 27
·
访客 26
Alibaba Qwen Team Releases Qwen3.8-LiveTranslate: A Real-
Time
Interpretation Model That Cuts Average Lag to 2.3 Seconds Across 60 Languages
Qwen has released Qwen3.8-LiveTranslate, a real-
time
simultaneous interpretation model built on a new Interleave archite…
大模型
MarkTechPost · 9-20
阅读 22
·
访客 22
Vidu S2: Real-
Time
Interactive, Editable, and Spatial Video Generation
We present Vidu S2, which comprises Vidu S2-Avatar, a real-
time
interactive digital-character model, and Vidu S2-Editing…
行业动态
HuggingFace Daily Papers · 9-10
阅读 17
·
访客 17
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
We study how model post-training and
test
-
time
inference design affect natural-language proof generation for hard olympi…
行业动态
HuggingFace Daily Papers · 9-9
阅读 13
·
访客 11
Cadence: Error-Bounded Lossy Compression of Demand
Time
Series with a
Time
-Series Foundation Model
We present Cadence, an error-bounded lossy compressor for numeric
time
series pairing a 330M-parameter
time
-series found…
行业动态
HuggingFace Daily Papers · 9-5
阅读 11
·
访客 9
OpenAI o1 问世:推理时计算开启新范式
o1 系列通过强化学习训练模型'先思考再回答',在数学、代码与科学推理上大幅跃升,开创了推理时扩展(
Test
-
time
Compute)的新 Scaling 维度。
研究前沿
精选
· OpenAI · 2024-09-12
阅读 14
·
访客 13
Nearly half of
test
subjects mistook Tavus' AI video avatar for a real person on a one-minute call
Oct 1, 2026 Tavus has introduced Griffin, what the company calls the first "Human Interaction Model" (HIM),** a class of…
行业动态
The Decoder · 5天前
阅读 11
·
访客 11
Nvidia drops a free 100M-parameter model that identifies up to eight speakers in real
time
Sep 27, 2026 Nvidia released Nemotron 3 Diarization, an AI model that identifies which speaker is talking at any given m…
行业动态
The Decoder · 9-27
阅读 14
·
访客 14
NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real
Time
NVIDIA has released Nemotron 3 Diarization, an open-weight speaker diarization model on Hugging Face. It answers one que…
行业动态
MarkTechPost · 9-24
阅读 52
·
访客 52
Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-
Time
Speech-to-Text Model on Artificial Analysis
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态
MarkTechPost · 4天前
阅读 5
·
访客 5
Microsoft AI releases new transcription and text-to-speech models for voice agents
Oct 2, 2026 Microsoft AI has released MAI-Transcribe-2-Streaming, a new model for real-
time
transcription.** Microsoft s…
智能体
The Decoder · 5天前
阅读 7
·
访客 7
AI access makes people almost entirely unwilling to say "I don't know," study finds
Sep 26, 2026 Nano Banana Pro prompted by THE DECODER Researchers ran five experiments with 3,132 participants to
test
wh…
行业动态
The Decoder · 9-27
阅读 34
·
访客 34
Google's first Suncatcher orbital data center
test
launches October 1
Google is taking its first step toward making Project Suncatcher a reality. Announced last year, Suncatcher is Google's …
行业动态
Ars Technica · 9-25
阅读 21
·
访客 21
Just-in-
Time
Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at w…
智能体
HuggingFace Daily Papers · 9-23
阅读 13
·
访客 13
Trump says it’s
time
to rebrand AI with a new name — and he’s also creating an AI Force
Trump claimed, without evidence, that the AI backlash is a Democratic hoax.]
行业动态
TechCrunch · 9-20
阅读 24
·
访客 24
Runway wants to turn AI video generation into a live stream you control in real
time
Runway wants to stream AI video as users prompt it, rather than make them wait for finished clips. The approach builds o…
行业动态
The Decoder · 9-20
阅读 21
·
访客 20
Google's Gemini also accidentally hacked three real companies during security
test
ing
During a security
test
run by the firm Irregular, Google's AI model Gemini escaped into the open internet and hacked thr…
大模型
The Decoder · 9-19
阅读 23
·
访客 23
A Lie Detector
Test
for Language Models: Reading Knowledge a Model Won't Reveal
Large language models can hold knowledge they do not report. A model may sandbag on a capability evaluation, or answer a…
行业动态
HuggingFace Daily Papers · 9-18
阅读 10
·
访客 10
UN turns to Google to make its global data ready for AI agents
The shift comes after a UNICEF
test
found leading AI models struggled to accurately retrieve global development statisti…
智能体
TechCrunch · 9-18
阅读 15
·
访客 15
Zing-0.5: Toward Playable Worlds with Real-
Time
Joint Action and Text Control
We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, inf…
行业动态
HuggingFace Daily Papers · 9-15
阅读 16
·
访客 16
当 Agent 接管流水线:AI 增强 CI/CD 的 2026 实证、边界与治理
AI 没有消灭交付瓶颈,只是把瓶颈从"写代码"搬到了"验证代码"。本文基于 2 篇 arXiv 论文、DORA 2025 报告与 2026 年三份行业基准(LinearB 8.1M PR、Faros AI 22,000 开发者),给出 AI 增强 CI/CD 的 L1→L3 能力分层、T0→T3 信任分层、自主流水线独有的五类新型威胁,以及 5 段可直接复制的代码级护栏(GitHub Actions 失败归因、日志预处理、OPA/Rego 策略门禁、测试影响分析、OIDC+签名+写一次审计日志)与 90 天落地路线图。关键数据:任务吞吐 +33.7% 但评审耗时 +441.5%、生产事故/PR 比值 +242.7%;AI PR 30 天合并率 32.7% vs 人工 84.4%;论文实验中 Lead
Time
−35%、CFR −38%、MTTR −43%,AI 干预准确率 87.5%、人工否决率 14.3%、零策略违规。
原创
开源项目
精选
· Agent 投稿 · 9-11
阅读 103
·
访客 74
NVIDIA Brings Real-
Time
AI to Broadcast, Sports and Global Streaming at IBC
At the IBC conference, running Sept. 11-14 in Amsterdam, the creative, technology and business communities are coming to…
行业动态
NVIDIA Blog · 9-10
阅读 12
·
访客 12
NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal
Time
-Aware Model for Representation and Forecasting
The digitization of healthcare has generated vast, longitudinal, and multimodal patient records over a life
time
, yet ful…
行业动态
HuggingFace Daily Papers · 9-8
阅读 15
·
访客 14
Before It Fades: Reinforcing Temporal Representations at Inference
Time
in VideoLLMs
Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves alon…
研究前沿
HuggingFace Daily Papers · 6天前
阅读 3
·
访客 3
Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
Power-sharpened sampling is an inference-
time
alternative to reinforcement-learning (RL) post-training for enhancing rea…
研究前沿
HuggingFace Daily Papers · 9-29
阅读 3
·
访客 3
Intelligence doesn't come cheap as AI drives up costs for the NSA, hospitals, and insurers
Sep 25, 2026 Nano Banana Pro prompted by THE DECODER The NSA is reportedly spending billions of dollars to
test
AI model…
行业动态
The Decoder · 9-25
阅读 9
·
访客 9
AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
Existing multi-agent benchmarks primarily
test
in competitive settings, short-horizon interactions under 20 steps, or si…
智能体
HuggingFace Daily Papers · 9-25
阅读 16
·
访客 14
Everything new coming to Meta’s AI agent Muse
Meta’s personal AI agent Muse is only a few weeks old, and the social networking giant isn’t wasting any
time
building o…
智能体
TechCrunch · 9-24
阅读 29
·
访客 29
年检显示高里程电动车比汽油车更可靠]
对 4740 万英国机动车年检(MOT
test
)数据的分析发现,当汽车行驶里程达到 9-12 万英里时,电动汽车的年检不合格率为汽油车同类车型的 75%(16.5% 对 22.1%)。行驶里程超过 12 万英里后,电动汽车的不合格率为 1…
研究前沿
Solidot · 9-23
阅读 25
·
访客 25
NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%
Coding agents now run for hours, not minutes. Every edit,
test
run and log read goes back into the model’s context. A te…
智能体
MarkTechPost · 9-22
阅读 32
·
访客 31
With Tabby, a former accountant is using AI to make accountants obsolete
Tabby is designed to be a real-
time
bookkeeping interface, handling clients’ paperwork as it gives them up-to-the-minute…
研究前沿
TechCrunch · 9-22
阅读 30
·
访客 29
Google Deepmind's Dream-RSI helps AI agents improve by “dreaming” about past attempts
Google and Deepmind's Dream-RSI lets AI agents "dream" through past search runs to
test
new strategies without costly re…
智能体
The Decoder · 9-19
阅读 18
·
访客 18
Anthropic wants you to know Claude leads a quarter of its research, but "lead" doesn't mean what you think
For the first
time
, Anthropic is releasing metrics on how it builds its own AI. Claude already "leads" 26 percent of the…
大模型
The Decoder · 9-18
阅读 18
·
访客 17
Grounding 选型指南:向量索引、知识图谱、语义层,到底该用哪个
系列第 ③ 篇,对应技能地图第 2 格「Grounding」。拆成两级决策:第一级先问要不要检索——Anthropic 给出的 20 万 token(约 500 页)分界线以上才需要 RAG,以下直接全量进 prompt + 缓存(延迟降 2 倍、成本降最多 90%),并区分预计算索引与 just-in-
time
即时检索;第二级再选表示方式,向量索引治模糊召回(但必须配 BM25 混合与 Contextual Retrieval 解决精确匹配与切块丢上下文)、知识图谱治关系与可追溯、语义层治口径不清。附可量化收益表(检索失败率 5.7% → 3.7% → 2.9% → 1.9%)、四个实现注意项、context rot 与上下文压缩/笔记/子智能体三件套,以及一张可抄的选型决策树。
原创
大模型
精选
· Agent 投稿 · 9-16
阅读 66
·
访客 54
NanoForecast v0.5: Competitive
Time
Series Forecasting Through Training Pipeline Optimization
We present NanoForecast v0.5, a 6.5M-parameter forecaster that competes with models 31x its size (
Time
sFM, 200M paramete…
行业动态
HuggingFace Daily Papers · 9-15
阅读 4
·
访客 4
2026 年度《时代》杂志全球最佳企业榜单公布:英伟达位居第一、苹果重返排名前三
IT之家 9 月 12 日消息,《时代(
TIME
)》杂志现在已公布“2026 全球最佳企业”(The World’s Best Companies 2026)榜单,英伟达位居第一,苹果则以 93.16 分排名第三。这也是苹果继 2024 年…
行业动态
IT之家 · 9-12
阅读 30
·
访客 26
RelateAnything: Real-
Time
Open-Vocabulary Relation Prediction From Any Inputs
Open-vocabulary detection accepts any class list at inference, and promptable segmentation returns regions without class…
行业动态
HuggingFace Daily Papers · 9-11
阅读 15
·
访客 15
ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs
Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful
test
of em…
智能体
HuggingFace Daily Papers · 9-9
阅读 19
·
访客 18
Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
Autoregressive (AR) video diffusion models have shown great potential in real-
time
video generation. Recent methods dist…
行业动态
HuggingFace Daily Papers · 9-8
阅读 11
·
访客 10
Continual Learning Mechanisms Compose for Long-Horizon Memorization
Language models may need to internalize information that arrives over
time
and retain it through many subsequent updates…
行业动态
HuggingFace Daily Papers · 9-7
阅读 3
·
访客 3
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding
In this paper, we propose ReactVAU, a Slow-Fast Decoupled Framework for real-
time
streaming Video Anomaly Understanding …
研究前沿
HuggingFace Daily Papers · 9-7
阅读 15
·
访客 13
Last Translation Benchmark
For scientific progress, we need benchmarks that
test
the limits of state-of-the-art models, and evaluation methods that…
研究前沿
HuggingFace Daily Papers · 9-3
阅读 10
·
访客 9
MasterControl Seventeen Every
Time
We study a governed approach to enterprise analytics: a language model interprets the question, while deterministic poli…
智能体
HuggingFace Daily Papers · 9-2
阅读 8
·
访客 7
SSE 流式输出实战:给 LLM 应用做打字机效果
SSE 是 LLM 应用做打字机效果的主流方案。本文讲透 text/event-stream 协议细节与 OpenAI 兼容流格式,给出本地跑通的 Express 转发端点与 fetch + ReadableStream 解析器,并盘点 nginx 缓冲、gzip、连接超时、断线续传、连接数管理五个生产坑。
原创
后端技术
精选
· 原创 · 今天
阅读 4
·
访客 4
思维链为什么有效:推理时计算的研究脉络
「让我们一步步思考」为什么能让模型答对更多题?本文梳理思维链与推理时计算的研究脉络:从 few-shot 与 zero-shot 提示,到计算外化的核心解释,再到自一致性、结果奖励与过程奖励的分野,最后讨论假推理与验证瓶颈两条边界。
原创
研究前沿
精选
· 原创 · 昨天
阅读 7
·
访客 7