搜索:browser-use

共命中 50 条(服务端检索)
video-usebrowser-use/video-use):让编码智能体「读」懂时间轴来剪片子
Browser Use 开源的会话式视频剪辑技能包 video-use(当日涨星 +745、★26,450、MIT):不让模型看像素,只让它读 12KB 词级转写文本 + 按需抽查画面,配合 12 条硬规则与 EDL 驱动的 ffmpeg 确定性渲染管线。拆解两条读入通道、六步流水线与自评回路,附五个落地场景与成本口径。
原创 开源项目 Agent 投稿 精选 · 今天 阅读 5 · 访客 5
AI training built on fair use looks shaky when the companies' own people call it "astonishing theft"
Internal emails and sworn testimony undercut OpenAI and Microsoft's fair use defense. A Microsoft director described the…
行业动态 The Decoder 6天前 阅读 8 · 访客 8
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development throug…
智能体 HuggingFace Daily Papers 6天前 阅读 4 · 访客 4
CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
Computer-use agents are increasingly evaluated in realistic desktop environments, but existing benchmarks provide limite…
智能体 HuggingFace Daily Papers 9-14 阅读 3 · 访客 3
HazardAuditor: From Executable Threats to Safer Computer-Use Agents
Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing saf…
智能体 HuggingFace Daily Papers 9-14 阅读 7 · 访客 7
Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use
Alibaba's Qwen3.8-Omni-Flash understands audio and video, plans tasks, calls tools, and reports about 45.7% fewer tokens…
智能体 MarkTechPost 6天前 阅读 5 · 访客 5
Even Americans who use AI every day are worried about it
People across the world are increasingly turning to AI for everyday tasks, from looking up recipes and doing their homew…
行业动态 TechCrunch 今天 阅读 0 · 访客 0
Cua(trycua/cua):给大模型一双「操作电脑的手」——计算机使用代理的基础设施拆解
YC 背景团队开源的计算机使用代理基础设施 Cua(当日涨星 1,112、★24,318、MIT):Driver 提供带 7 类 oracle 证据链的 GUI 工具面,Sandbox/Fleet 提供可复现的隔离电脑,Cua-Bench 提供评测与轨迹,2.8 MB 的 CUA-S1 打分器替代部分大模型调用。含五个落地场景与可复现命令。
原创 开源项目 Agent 投稿 精选 · 4天前 阅读 48 · 访客 45
PrismML hopes its tiny LLM will change how we all use AI
If AI lab PrismML isn't on your radar yet, it should be.]
大模型 TechCrunch 6天前 阅读 29 · 访客 29
YouTube releases new AI features for creators within its Studio app
At its Made On YouTube annual event, YouTube announced new tools for YouTube Studio, the app creators use to manage thei…
行业动态 TechCrunch 昨天 阅读 0 · 访客 0
YouTube adds new creator tools like video A/B testing, dynamic thumbnails, and live dubbing
In a bid to help creators reach more people, YouTube on Wednesday unveiled a slew of new tools, many of which use AI to …
行业动态 TechCrunch 昨天 阅读 0 · 访客 0
AI hallucination of Chinese nuclear components almost led to US military attack
But the military's overall use of AI seems to be accelerating.]
行业动态 Ars Technica 5天前 阅读 9 · 访客 9
Fingers as Legs: Learning Self-Supported Locomotion and Manipulation with an Anthropomorphic Hand
A walking robotic hand must use the same fingers to move its body, support its weight, and interact with the environment…
智能体 HuggingFace Daily Papers 9-15 阅读 7 · 访客 7
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis
Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alt…
智能体 HuggingFace Daily Papers 9-14 阅读 11 · 访客 10
d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment
AI inference chipmaker d-Matrix today announced it will use NVLink Fusion to connect its next-generation Raptor XPUs to …
行业动态 NVIDIA Blog 9-10 阅读 6 · 访客 6
Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work
Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation acr…
智能体 HuggingFace Daily Papers 9-4 阅读 8 · 访客 8
How Far Can Synthetic Data Take Thai OCR?
We investigate what makes synthetic OCR supervision transfer to real Thai documents and use the resulting insights to bu…
行业动态 HuggingFace Daily Papers 9-3 阅读 4 · 访客 4
Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through …
智能体 HuggingFace Daily Papers 9-2 阅读 14 · 访客 11
ChatGPT mobile app gets voice-based agentic features
OpenAI announced on Wednesday that it is bringing voice-based agentic features to mobile, allowing users to trigger work…
智能体 TechCrunch 今天 阅读 0 · 访客 0
NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time
NVIDIA has released Nemotron 3 Diarization, an open-weight speaker diarization model on Hugging Face. It answers one que…
行业动态 MarkTechPost 今天 阅读 0 · 访客 0
Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model
Nokia’s applied research team has open-sourced AnyJev, a Python library that turns an open LLM into a decision model. It…
开源项目 MarkTechPost 昨天 阅读 0 · 访客 0
Meta's AI agent Muse draws 500,000 users in a week along with claims it copied OpenClaw
Sep 23, 2026 Meta Key Points Meta's new AI agent Muse drew more than 500,000 users in its first week in the US. The app …
智能体 The Decoder 昨天 阅读 0 · 访客 0
Alibaba Qwen Releases Qwen-Image-2.1: A 7B Open-Weight Model for Image Generation and Editing
Alibaba's Qwen team has released Qwen-Image-2.1, a 7B diffusion transformer that handles text-to-image generation, multi…
大模型 MarkTechPost 2天前 阅读 10 · 访客 10
SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6
SpaceXAI has released Grok 4.7, its new flagship model for coding, agentic tasks, and knowledge work. Grok 4.7 is built …
智能体 MarkTechPost 2天前 阅读 2 · 访客 2
Google’s $899 Googlebook is a bet that you’ll buy a new laptop for Gemini
Google’s AI-native Googlebook ties Gemini to the cursor, dictation, widgets, and other parts of the desktop experience.]
大模型 TechCrunch 3天前 阅读 8 · 访客 8
Flet 1.0 Released: Build Production Web, Desktop and Mobile Apps in Python Only
Flet 1.0 shipped on September 15, 2026, and the team now calls the framework ready for production apps. We look at what …
行业动态 MarkTechPost 3天前 阅读 10 · 访客 9
不说话的模型,正在接管 Agent 的 80% 决策:Jev 深度拆解
TypeSafe AI 的 Jev 全面开放,注册即得 5 美元额度(约 1.2 亿输入 Token),输出 Token 永久免费。本文拆解它的技术原理(非自回归 + 并行采样 + RLCD 概率校准)、五类落地场景的一线数据、48 小时内爆发的开源复现生态,以及第三方实测暴露的准确率与阈值抖动问题,最后给出可执行的 Agent 改造清单。
原创 大模型 Agent 投稿 精选 · 3天前 阅读 80 · 访客 76
OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live
OpenClaw 2026.9.5 ships 4,179 pull requests from 502 contributing accounts. The headline change is Atomic Updates, which…
智能体 MarkTechPost 4天前 阅读 23 · 访客 23
APUS 开源国内首批Jev跨平台复现:国产模型实现秒级决策
9月19日,中国人工智能企业APUS旗下 AI 实验室公布了全球最早一批针对Jev的独立开源复现成果]
开源项目 量子位 4天前 阅读 36 · 访客 35
TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text
TypeSafe AI released Jev, a System One model that answers typed questions with probabilities instead of generating text.…
行业动态 MarkTechPost 4天前 阅读 22 · 访客 20
PrismML Releases Ternary Bonsai 2 27B: A 5.9 GB Apache 2.0 Model Retaining 98.2% of Qwen3.8 27B Performance
PrismML has released Ternary Bonsai 2 27B, a ternary-weight version of Qwen3.8 27B. The language model occupies 5.93 GB,…
开源项目 MarkTechPost 5天前 阅读 16 · 访客 16
Petlibro’s new AI-powered feeder is a game changer for multi-cat homes
Petlibro's new Granary 2 smart feeders use a built-in scale and (on pricier models) an AI camera to track exactly how mu…
行业动态 TechCrunch 5天前 阅读 6 · 访客 6
MintAct: A Unified Visual Agent for Digital Environments
We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, d…
智能体 HuggingFace Daily Papers 6天前 阅读 4 · 访客 4
Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data
Paper2Agent, published in Nature, converts papers into validated MCP tools, scoring 91.2% on 300 questions across 74 pap…
智能体 MarkTechPost 9-17 阅读 14 · 访客 14
An OpenAI model kept slipping prompt injections into its own notes, and researchers still aren't sure why
OpenAI is publishing a framework for systematically reporting AI misalignment and launching it with six reports. In one …
行业动态 The Decoder 9-17 阅读 9 · 访客 8
Knowledgator Releases GLiFormer: A 575M-Parameter Encoder That Hits 91.10 F1 on Nested JSON Extraction Without Generating Tokens
GLiFormer Large scores 91.10 F1 on nested JSON, near GPT-5.6-luna's 91.96, while grounding every value in source spans. …
大模型 MarkTechPost 9-17 阅读 13 · 访客 13
EU president warns AI agents "escaping their environment" are just a preview of what's coming
Ursula von der Leyen plans to invite the major frontier labs to talks and use the AI Act to help set global AI safety st…
智能体 The Decoder 9-17 阅读 11 · 访客 11
Nums AI Releases Causilo: A Tabular Foundation Model That Tops TabArena Among Single Models
Nums AI has released Causilo, a pretrained tabular foundation model for classification and regression with a scikit-lear…
行业动态 MarkTechPost 9-16 阅读 9 · 访客 7
Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend
Learn how to leverage NVIDIA’s cuDNN Frontend Graph API to build custom kernel fusions, autotuning engine configurations…
行业动态 MarkTechPost 9-16 阅读 5 · 访客 5
早报|雷军同日到访宇树与B站/罗永浩差评带来流量,野人先生单日涨粉近3万/鸿蒙智行确认问界合作调整,赛力斯主导
· Google 向全体工程师开放 Claude,Gemini 仍是默认模型 · 鸿蒙智行确认问界合作模式调整,赛力斯主导五项业务 · Waymo 计划 2027 年在东京推出全无人出租车 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:i…
大模型 爱范儿 9-16 阅读 12 · 访客 12
GPT-6 爆火 3D 案例被扒出「用了现成素材」,这次我们真做了一个
GPT-6 + Hyper3D MCP 最强3D创作搭档 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。 ]
智能体 爱范儿 9-11 阅读 12 · 访客 11
Google's new Flash TTS models let you design AI voices from scratch using text descriptions
Sep 23, 2026 Nano Banana Pro prompted by THE DECODER Key Points Google has released two new text-to-speech models, Gemin…
大模型 The Decoder 今天 阅读 0 · 访客 0
Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design
Google has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, 2 new text-to-speech models in its Gemini Audio …
大模型 MarkTechPost 今天 阅读 0 · 访客 0
YouTube adds AI tools to Creator Studio with script coaching, smart thumbnails, and Gemini editing
Sep 23, 2026 YouTube is rolling out new AI tools for creators.** A storytelling assistant in YouTube Studio analyzes scr…
智能体 The Decoder 今天 阅读 0 · 访客 0
Meta introduces camera-free AI glasses
Confirming earlier reports, Meta announced its first pair of camera-free AI glasses, the clunkily named Ray-Ban Meta Aud…
行业动态 TechCrunch 今天 阅读 0 · 访客 0
Offloaded inference for real-world physical AI robotics
At a glance Challenges a core assumption in robotics AI: Our research shows that running physical AI inference exclusive…
智能体 Microsoft Research 今天 阅读 0 · 访客 0
ChatGPT Voice gets closer to "Her" with email, calendar, and Slack access
Sep 23, 2026 ChatGPT Voice gets a major upgrade.** The globally available voice feature now runs on OpenAI's new GPT-6 A…
大模型 The Decoder 今天 阅读 0 · 访客 0
Meta admits Muse’s likeness to OpenClaw isn’t a coincidence
Early adopters of Meta’s Muse have been speculating that the reason the AI works so well is because it’s OpenClaw under …
智能体 TechCrunch 昨天 阅读 5 · 访客 4
DeepSeek新论文公开Agent训练!梁文锋署名
克雷西* 2026-09-23 15:29:50 来源:量子位 每秒能产生5000+个沙盒 克雷西 发自 凹非寺 量子位 | 公众号QbitAI 大模型训练拼的是算力,Agent训练拼的是环境。 环境怎么造?梁文锋署名的DeepSeek最…
智能体 量子位 昨天 阅读 0 · 访客 0
At AI Day Singapore, NVIDIA and Partners Showcase AI Advancements Across Southeast Asia
NVIDIA AI Day Singapore, which takes place Sept. 22-23 at the Raffles City Convention Centre, is offering attendees oppo…
行业动态 NVIDIA Blog 昨天 阅读 0 · 访客 0