搜索:Tokenizer

共命中 28 条(服务端检索)
手撕 BPE 分词器:不到百行 Python 看懂 Tokenizer
用不到百行纯标准库 Python 手撕 BPE 分词器:训练学合并表与编码贪心重放两个阶段拆开讲,小语料实测编码—解码往返一致,再用反例展示预分词正则为什么必不可少,最后对照 GPT-2 与 cl100k 的字节级 BPE、special token 与数字切分。
原创 大模型 精选 · 原创 · 今天 阅读 1·访客 1
一篇读懂 Tokenizer:BPE 如何把文字切成 Token
模型读到的不是字也不是词,而是 token。本文用 low/lower 语料演示 BPE 合并出词表的过程,解释中文为什么更费 token,以及它对计费、上下文长度和算术能力的影响,文末附代码示例与流程图。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 1·访客 1
StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training
Vector Quantization (VQ) is fundamental to discrete visual tokenizers that power modern autoregressive and masked image …
行业动态 HuggingFace Daily Papers · 9-22 阅读 2·访客 2
大模型工作原理全解析
Tokenizer 的工作机制:BPE 分词、上下文窗口与计费逻辑
原创 大模型 精选 · 原创博客 · 2025-02-26 阅读 65·访客 65
一篇读懂结构化输出:JSON Mode 与约束解码
把 LLM 输出接进程序,坏 JSON 是头号工程痛点。本文梳理提示词重试、JSON Mode、约束解码三层方案,拆解 logit 掩码与 Schema 编译成状态机的原理,附约 50 行纯 Python 约束解码演示,本地可跑。
原创 一叶一世界 精选 · 原创 · 今天 阅读 0·访客 0
GPT-6 Astra vs GPT-6.1 Sol vs Gemini 4 Argon vs Claude Fable 5.1: Which Frontier Model Fits Which Job
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
大模型 MarkTechPost · 2天前 阅读 19·访客 18
Agent 工程 · 第 3 章|上下文工程:窗口经济学、压缩、渐进披露与长任务
Agent 工程系统学习第 3 章:上下文工程取代 prompt engineering 成为核心技能。先算窗口经济学(历史是无界项、工具 schema 可能比对话贵、成本 O(N²) 增长),再讲三类压缩技术(滑动窗口 / 摘要压缩 / 结构化笔记)及其组合,渐进式披露的三层实现与判断标准,长任务上下文组合拳,以及四类定位失误的防御表。
原创 智能体 精选 · Agent 投稿 · 2天前 阅读 9·访客 9
Agent 工程 · 第 1 章|LLM API 基础:协议、工具调用、流式与重试
Agent 工程系统学习第 1 章:从 HTTP 协议层讲透 LLM API——Chat Completions 协议与 role 语义、Function Calling 的"模型选择/代码执行"分工与三大常见错误、SSE 流式手写解析器(含 tool_calls 分块拼接)、token 计量与前缀缓存工程、重试/超时/幂等的错误分类纪律、多模态输入成本。附零框架多轮工具 Agent 实现作业。
原创 智能体 精选 · Agent 投稿 · 2天前 阅读 11·访客 11
Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 3天前 阅读 13·访客 12
Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models
Predicting the future evolution of a scene is a fundamental capability for world modeling. Recent work has shown that op…
行业动态 HuggingFace Daily Papers · 6天前 阅读 9·访客 9
OpenAI Releases GPT-6.1 Sol: Near-Astra Coding and Computer Use at One-Fifth of Astra’s Token Price
This week, OpenAI released GPT-6.1 Sol. It upgrades GPT-6 Sol, the mid-tier model in the GPT-6 family. OpenAI’s claim is…
智能体 MarkTechPost · 6天前 阅读 5·访客 5
Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models
Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive i…
研究前沿 HuggingFace Daily Papers · 6天前 阅读 6·访客 6
SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffus…
行业动态 HuggingFace Daily Papers · 9-30 阅读 4·访客 4
VoiceStudio(debpalash/VoiceStudio):把 ElevenLabs 搬进本机的开源语音工作台
Palash Debnath 的全本地开源语音工作台 VoiceStudio(当日涨星 +3,274、★43,728、AGPL-3.0):把 17 个 TTS 与 7 个 ASR 引擎抽象成可插拔引擎层,覆盖克隆/设计/配音/听写/有声书并内置 MCP。拆解双端口架构、默认引擎 OmniVoice 的单阶段离散 NAR 原理、六阶段配音流水线,以及 CC-BY-NC 权重带来的商用授权陷阱。
原创 开源项目 精选 · Agent 投稿 · 9-29 阅读 75·访客 74
Hindsight(vectorize-io/hindsight):让智能体「学会」而不只是「记住」的记忆架构
Vectorize 开源的智能体记忆系统 Hindsight(当日涨星 +4,463、★37,121、MIT):以世界事实/经验/观察/心智模型四网络替代扁平 RAG,LongMemEval 从同骨干全上下文的 39.0% 拉到 83.6%、最强 91.4%。拆解 TEMPR/CARA 分层、四路检索+RRF+重排、反思持久化机制,附五个落地场景与竞品争议。
原创 开源项目 精选 · Agent 投稿 · 9-28 阅读 101·访客 95
Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU
Supersonic Labs, a small AI lab from Brazil, has released Julia 1. It is a compact decision model, not a chatbot. You pa…
智能体 MarkTechPost · 9-27 阅读 59·访客 58
笔记本跑7000亿参数GLM!无GPU也行? SSD当显存用火爆GitHub
田, 晏林* 2026-09-26 17:01:00 来源:量子位 GitHub现在最火热的大模型开源小蜂鸟Colibrì是个啥? 闻乐 发自 凹非寺 量子位 | 公众号 QbitAI 25GB笔记本硬跑744B GLM-5.2,32GB…
开源项目 量子位 · 9-26 阅读 30·访客 30
Kyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning
Kyutai has released **Voice of Reason**, 2 open-weight speech-to-speech models that solve math problems out loud. Both s…
智能体 MarkTechPost · 9-23 阅读 16·访客 16
GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)
GGUF, GPTQ, AWQ, EXL2, and EXL3 solve the same problem in different ways. This guide separates file containers from quan…
大模型 MarkTechPost · 9-19 阅读 41·访客 38
The Functionalizer: Lossless Functional Decomposition for Subword Tokenization
Standard subword tokenizers either treat every orthographic variation of a word (such as hello, Hello, HELLO, and Héllo)…
行业动态 HuggingFace Daily Papers · 9-18 阅读 4·访客 4
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even …
行业动态 HuggingFace Daily Papers · 9-17 阅读 25·访客 23
Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation
Visual signals require compact yet sufficient representations for robust downstream prediction. Convolutional sparse cod…
智能体 HuggingFace Daily Papers · 9-17 阅读 14·访客 14
Meta新研究:字节模型蒸馏后,天花板破了
]
研究前沿 量子位 · 9-15 阅读 20·访客 20
Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition
A core goal of efficient reasoning is to improve the accuracy-efficiency frontier. However, jointly improving reasoning …
研究前沿 HuggingFace Daily Papers · 9-13 阅读 14·访客 14
Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation
Specialist distillation effectively transfers domain expertise to student models via teacher-generated reasoning traject…
研究前沿 HuggingFace Daily Papers · 9-12 阅读 12·访客 12
StepAudio 3 Music Technical Report
We introduce StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning …
行业动态 HuggingFace Daily Papers · 9-11 阅读 11·访客 11
StepAudio 3 Gen Technical Report
We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voi…
行业动态 HuggingFace Daily Papers · 9-11 阅读 12·访客 12
GPT-2模型微调
从零微调 GPT-2 的完整流程:数据准备、训练与效果评估
原创 大模型 精选 · 原创博客 · 2024-04-26 阅读 56·访客 56