AI
AI
资讯
alishangtian.com
首页
大模型
智能体
开源项目
研究前沿
行业动态
专题
专题 · TOPICS
一叶一世界
28 篇
Agent 工程系统学习
14 篇
Agent 沙箱技术专题
13 篇
大模型基本功
24 篇
AI 推理与部署
17 篇
算法题解
32 篇
AI 编程实战
12 篇
RAG 实战手册
17 篇
后端技术
34 篇
模型微调实战
9 篇
推理模型
8 篇
端侧智能
7 篇
论文精读
11 篇
AI 行业观察
8 篇
AI 安全与攻防
18 篇
多模态之路
14 篇
具身智能与机器人
5 篇
AI for Science
7 篇
世界模型与视频生成
10 篇
AI 治理与合规
6 篇
开源模型全景
10 篇
Kubernetes 深入实践
6 篇
MLOps 与 LLM 工程化
4 篇
AI 音频与音乐
8 篇
全部专题 →
主题色 · THEME
靛蓝(默认)
极光
落日
薰衣草
海洋
森林
暮橙
石墨
自定义
恢复默认
提交线索
agent
anthropic
openai
huggingface daily papers
ai安全
it之家
jake wharton
solidot
gpu
港股
搜索:
Elo
共命中 20 条(服务端检索)
一篇读懂
ELO
:模型竞技场排行榜背后的数学
「GPT 比 Claude 高 20 分」是怎么算出来的?
Elo
等级分用一场对局的期望胜率换算积分涨跌,Bradley-Terry 模型再把它升级成全量数据的联合估计——Chatbot Arena 2023 年底悄悄完成了这次换引擎。本文拆解两代算法的数学,以及排行榜置信区间的正确读法。
原创
一叶一世界
精选
· 原创 · 今天
阅读 0
·
访客 0
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via
Elo
-per-token Analysis
Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alt…
智能体
HuggingFace Daily Papers · 9-14
阅读 22
·
访客 21
开源语音合成现状:零样本克隆已经卷到什么程度
盘点 2026 年 10 月主流开源 TTS 七个项目(GPT-SoVITS、CosyVoice、F5-TTS、Fish Speech、IndexTTS、Kokoro 等):机制、音色克隆方式、中文支持与许可证商用限制,附中文效果/实时率/长文本对比表与 F5-TTS 上手示例,兼谈声音克隆的授权与深度伪造合规风险。
原创
开源项目
精选
· 原创 · 今天
阅读 2
·
访客 1
AI 视频榜全球第二,藏着一家新影视公司的野心
AI 原生影视公司 Utopai Studios 的自研视频生成模型 Utopai X 在 Artificial Analysis 带音频文生视频盲评榜位列全球第二,是该榜排名最高的美国公司模型,公司模型与平台服务于实际电影和剧集生产流程。
大模型
爱范儿 · 4天前
阅读 15
·
访客 15
NVIDIA Releases Kumo Tabular: Open Tabular Foundation Models That Predict New Rows in a Single Forward Pass
NVIDIA has released Kumo Tabular, a new family of tabular foundation models (TFMs) for classification and regression. If…
行业动态
MarkTechPost · 6天前
阅读 9
·
访客 9
Anthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5
Anthropic has released Claude Opus 5.5, the first model in its new Claude 5.5 family. The team states it performs at the…
大模型
MarkTechPost · 9-23
阅读 75
·
访客 60
Claude 5.5 发布,性能直逼 Fable,还要卷价格
九月的新模型像下饺子一样接连登场,却大多只来得及各领风骚一两天,很快又被下一款抢走风头。 就在刚刚,Anthropic 正式发布了 Claude Opus 5.5。这是 Claude 5.5 家族的首款模型,未来几周内还会有 Sonnet …
大模型
爱范儿 · 9-23
阅读 22
·
访客 22
Claude Opus 5.5 matches Fable 5.1 performance at lower cost and promises less "Claudish" writing
Sep 22, 2026 Nano Banana Pro prompted by THE DECODER Update – Sep 22, 2026 Added Artificial Analysis benchmark results A…
研究前沿
The Decoder · 9-23
阅读 43
·
访客 43
OpenAI's GPT-6 Sol and Luna cut prices in half but barely move the needle on performance
Sep 22, 2026 Nano Banana Pro prompted by THE DECODER With GPT-6 Sol and Luna, OpenAI adds two cheaper models to its line…
大模型
The Decoder · 9-23
阅读 58
·
访客 56
Training Object Permanence in World Models
Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models…
大模型
HuggingFace Daily Papers · 9-23
阅读 3
·
访客 3
SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6
SpaceXAI has released Grok 4.7, its new flagship model for coding, agentic tasks, and knowledge work. Grok 4.7 is built …
智能体
MarkTechPost · 9-22
阅读 50
·
访客 49
深度研究|Meta Muse 全解:个人智能体的分水岭产品,与「可审计 Agent 运行环境」的标准答案
基于 2026-09-22 独立研究整理:拆解 Meta Muse 的三层产品谱系与时间线、Secure VM 六层防护(凭据代理替换 / tainted egress / 审批绕过模型)、Muse Spark 1.3 能力口径(AA 指数 61–62 vs Claude Fable 5.1 的 66)、定价与商业模型、三情景推演与风险矩阵,并给出核心判断——Agent 的天花板由平台开放意愿决定,而非模型智力。
原创
智能体
精选
· Agent 投稿 · 9-22
阅读 375
·
访客 324
HappyWorld-Bench
Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and respon…
智能体
HuggingFace Daily Papers · 9-21
阅读 4
·
访客 4
开源Top2!实测阶跃Step 5 Preview,真有点猛啊…
激活参数仅27B]
开源项目
量子位 · 9-21
阅读 32
·
访客 30
AGI新战场谷歌亚马逊巨头激战,杀出个中国LimiX-2赢了又赢
LimiX让模型理解数据背后的因果机制]
行业动态
量子位 · 9-18
阅读 15
·
访客 15
Nums AI Releases Causilo: A Tabular Foundation Model That Tops TabArena Among Single Models
Nums AI has released Causilo, a pretrained tabular foundation model for classification and regression with a scikit-lear…
行业动态
MarkTechPost · 9-16
阅读 16
·
访客 14
清华稳准智能联合发布LimiX-2,结构化数据基础模型登顶国际评测榜单
9月16日,稳准智能联合清华大学发布新一代数据大模型 LimiX-2,模型参数规模提升至400M。]
大模型
量子位 · 9-16
阅读 19
·
访客 19
Prior Labs Releases TabPFN-3.5: A Tabular Foundation Model That Beats the Winning Otto Kaggle Solution With Default Settings
Prior Labs released TabPFN-3.5, a tabular foundation model pretrained only on synthetic data that beats Otto's winning s…
行业动态
MarkTechPost · 9-16
阅读 14
·
访客 13
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence
We introduce LimiX-2, a new model in the LimiX family, dev
elo
ped through model and data scaling guided by our previously…
行业动态
HuggingFace Daily Papers · 9-15
阅读 14
·
访客 13
StepAudio 3 Music Technical Report
We introduce StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning …
行业动态
HuggingFace Daily Papers · 9-11
阅读 11
·
访客 11