输入框那个模型名,你大概从没换过。这一章你拿自己的三个真问题跑一次对照,做出一张属于你的三行路由表 —— 哪类活走 Haiku,哪类走 Sonnet,哪类值得花 Opus,然后把结论钉进工作区默认。That model name in the input box — you probably never change it. Here you run three of your own real questions side by side and build a three-row routing table of your own: which work goes to Haiku, which to Sonnet, which earns Opus — then pin the answer as a workspace default.
输入框右上角那个模型名,是你每次对话都在用、却几乎从没动过的开关。默认那一档不一定是最划算的那一档 —— 但「哪一档划算」这件事,别人替你答不了,只能你拿自己的活跑一遍。That model name in the corner of the input box is the switch you use every conversation and almost never touch. The default isn't always the most economical setting — and no one can tell you which one is, because the answer comes out of your own work.
— I
当前阵容,与三种代价The Current Lineup, and Three Kinds of Cost.
Sonnet 是默认主力,Opus 贵在难题,Haiku 快在量大。Sonnet is the default workhorse, Opus is worth its price on hard problems, Haiku is fast at volume.三个日常主力,价格差到 5 倍。截至 2026-08:1注 1Note 1Claude API Docs · Models overview —— 当前阵容:Claude Opus 5($5/$25 每百万输入/输出 token,1M 上下文,128K 输出)、Claude Sonnet 5($3/$15,1M;至 2026-08-31 入门价 $2/$10)、Claude Haiku 4.5($1/$5,200K,64K 输出)。Opus 4.8 / 4.7 / 4.6、Sonnet 4.6 / 4.5 等已列入 legacy。截至 2026-08。Claude API Docs · Models overview — the current lineup: Claude Opus 5 ($5/$25 per Mtok in/out, 1M context, 128K output), Claude Sonnet 5 ($3/$15, 1M; introductory $2/$10 through 2026-08-31), Claude Haiku 4.5 ($1/$5, 200K, 64K output). Opus 4.8 / 4.7 / 4.6 and Sonnet 4.6 / 4.5 are now listed as legacy. As of 2026-08.Three everyday workhorses, a 5x price spread. As of 2026-08:1注 1Note 1Claude API Docs · Models overview —— 当前阵容:Claude Opus 5($5/$25 每百万输入/输出 token,1M 上下文,128K 输出)、Claude Sonnet 5($3/$15,1M;至 2026-08-31 入门价 $2/$10)、Claude Haiku 4.5($1/$5,200K,64K 输出)。Opus 4.8 / 4.7 / 4.6、Sonnet 4.6 / 4.5 等已列入 legacy。截至 2026-08。Claude API Docs · Models overview — the current lineup: Claude Opus 5 ($5/$25 per Mtok in/out, 1M context, 128K output), Claude Sonnet 5 ($3/$15, 1M; introductory $2/$10 through 2026-08-31), Claude Haiku 4.5 ($1/$5, 200K, 64K output). Opus 4.8 / 4.7 / 4.6 and Sonnet 4.6 / 4.5 are now listed as legacy. As of 2026-08.
模型Model
每百万 token(入/出)Per Mtok (in/out)
上下文Context
什么时候选它When to pick it
Claude Opus 5
$5 / $25
1M
最难的推理、长代码、要一次想透的题Hardest reasoning, long code, problems to think through in one pass
Claude Sonnet 5
$3 / $15
1M
日常主力:大多数活它就够(至 8 月底入门价 $2/$10)The daily workhorse: enough for most work (intro $2/$10 through Aug 31)
Claude Haiku 4.5
$1 / $5
200K
量大、要快、不烧脑的轻活High-volume, fast, low-difficulty work
对只用网页的人,价格不体现为账单,而是三种代价 —— 而且只有第一种你能直接感觉到:For web-only users, price doesn't show up as a bill but as three costs — and only the first is one you feel directly:
额度消耗速度How fast quota drains
同一件活,Opus 比 Sonnet 更快撞到限额。这是你唯一能直接观察到的代价,也是第 11 章那本账要记的东西。On the same task, Opus hits the limit sooner than Sonnet. It's the only cost you can observe directly, and it's exactly what the ledger in Chapter 11 records.
等待时间Waiting
越强的模型想得越久。一个五分钟的深度回答和一个八秒的够用回答,在「我正等着用」的场景里不是同一件商品。A stronger model thinks longer. A five-minute deep answer and an eight-second good-enough answer are not the same product when you're standing there waiting for it.
你的注意力Your attention
被低估的一种。弱模型答得快,但你要多审两遍 —— 省下的额度换成了你的时间,这笔账常常是亏的。The underrated one. A weaker model answers fast, but you re-read it twice — quota saved, your time spent. That trade is often a loss.
所以那句常听的「贵的模型买的不是上限,是下限」有一半是对的:你多花的额度,换的是它在难题上更少翻车 —— 但只有当你真的在做难题时才成立。So the familiar line — 'the pricier model buys not a higher ceiling but a higher floor' — is half right: the extra quota buys fewer faceplants on hard problems, and only holds when you're actually doing hard problems.
— II
思考已经不是一个开关Thinking Isn't a Switch Anymore.
模型自己决定想多久,你不用再替它按。The model decides how long to think; you no longer flip a switch for it.当前主力模型都用「自适应思考」—— 按任务难度自动决定要不要多想一会儿、想多久,旧版那个手动的「延长思考」token 预算,在当代模型上已经移除。3注 3Note 3Claude API Docs · Models overview —— 当前模型使用自适应思考(模型自行决定想多久),手动的「延长思考」token 预算已在当代模型上移除;effort 参数在 Claude Opus 5 与 Sonnet 5 上于 Claude API 与 Claude Code 默认 high。截至 2026-08。Claude API Docs · Models overview — current models use adaptive thinking (the model decides how long to think); the manual extended-thinking token budget is removed on current models. The effort parameter defaults to high on Claude Opus 5 and Sonnet 5 across the Claude API and Claude Code. As of 2026-08.如果你从来没动过那个开关,什么都不用改。The current workhorses all use adaptive thinking — deciding on their own whether and how long to reason harder based on the task. The old manual extended-thinking token budget has been removed on current models.3注 3Note 3Claude API Docs · Models overview —— 当前模型使用自适应思考(模型自行决定想多久),手动的「延长思考」token 预算已在当代模型上移除;effort 参数在 Claude Opus 5 与 Sonnet 5 上于 Claude API 与 Claude Code 默认 high。截至 2026-08。Claude API Docs · Models overview — current models use adaptive thinking (the model decides how long to think); the manual extended-thinking token budget is removed on current models. The effort parameter defaults to high on Claude Opus 5 and Sonnet 5 across the Claude API and Claude Code. As of 2026-08. If you never touched that switch, nothing changes for you.代价和以前一样:想得越多,推理越稳,额度消耗也越快。所以同一段提示,复杂题比简单题更耗额度 —— 因为它在后台多想了。这意味着你真正在调的从来不是「开不开思考」,而是「我这件活配不配得上它想那么久」。那正是这一章要你做的路由表。The cost equation is unchanged: more thinking means steadier reasoning and faster quota consumption. So the same prompt uses more quota on a hard problem than an easy one — because it thought harder behind the scenes. Which means the dial you're actually turning was never 'thinking on or off,' it's 'does this task deserve that much thinking.' That dial is the routing table.
— III
动手:跑一次三行对照Do It: Run the Three-Row Comparison.
十五分钟,换掉你以后几百次对话的默认选项。Fifteen minutes, in exchange for the default on your next several hundred conversations.
01
从近一周的对话里挑三个真问题Pull three real questions from your past week
一个简单的(查个说法、翻译一段、格式转换),一个中等的(写一段东西、整理一份材料),一个真正烧脑的(多步推理、长文档综合、难调的 bug)。必须是你真问过的 —— 编的题会得到编的结论。One easy (look something up, translate a passage, convert a format), one medium (write something, tidy up some material), one genuinely hard (multi-step reasoning, long-document synthesis, a gnarly bug). They must be questions you actually asked — invented ones give you invented conclusions.
02
每个问题在两个模型上各跑一遍Run each question on two models
简单题跑 Haiku 和 Sonnet,中等题跑 Sonnet 和 Opus,烧脑题跑 Sonnet 和 Opus。同一句提示词,不要改。每次记三样:答案够不够用、等了多久、你有没有再追问。Run the easy one on Haiku and Sonnet, the medium and hard ones on Sonnet and Opus. Same prompt, unchanged. Record three things each time: was the answer usable, how long you waited, and whether you had to follow up.
03
只问一个问题:便宜那档够不够用Ask one question only: was the cheaper one enough
不是「哪个更好」——贵的那个几乎总是略好一点,那个比较没有意义。要问的是:便宜那档的答案,我能不能直接用?能,这行就归它。Not 'which is better' — the pricier one is almost always a bit better, and that comparison tells you nothing. Ask instead: could I have used the cheaper one's answer as-is? If yes, that row belongs to it.
04
写成三行,钉进工作区Write the three rows and pin them
按「这类活 → 走这个模型 → 因为」写三行。然后把结论钉进你常用的工作区默认(第 4 章会建那个 Project)—— 光记在脑子里,两周后你又回到默认那档了。Three rows, each shaped 'this kind of work → this model → because.' Then pin the conclusion as the default of the workspace you use most (Chapter 4 builds that Project) — keep it in your head only and you'll be back on the default tier in two weeks.
嫌自己判断费劲,可以让它先替你定档、再动手 —— 这一句同时也是个长期习惯:把「该用哪档」变成它每次开工前的第一步。If judging it yourself feels like work, have it pick the tier before it starts — and keep the habit: make 'which tier is this' the first thing it does on every job.
提示词Prompt让 Claude 替你定档Let Claude pick the tier
先别急着答。看一眼下面这个任务,告诉我:
1. 它该用 Haiku、Sonnet 还是 Opus?一句话说为什么(按难度 / 步数 / 有没有唯一正确答案)。
2. 如果我现在用的档比你建议的高,说清高在哪儿浪费了。
说完你的判断,再按它直接做。
任务:<把你的任务贴这里>。Don't answer yet. Look at the task below and tell me:
1. Should it run on Haiku, Sonnet, or Opus? One line on why (by difficulty / number of steps / whether there's a single correct answer).
2. If the tier I'm on is higher than the one you'd pick, say specifically what the extra is being wasted on.
State your call, then do the task by it.
Task: <paste your task here>.
先别急着答。看一眼下面这个任务,告诉我:
1. 它该用 Haiku、Sonnet 还是 Opus?一句话说为什么(按难度 / 步数 / 有没有唯一正确答案)。
2. 如果我现在用的档比你建议的高,说清高在哪儿浪费了。
说完你的判断,再按它直接做。
任务:把这段 300 字的产品公告翻成英文,保留所有数字和产品名,语气正式。Don't answer yet. Look at the task below and tell me:
1. Should it run on Haiku, Sonnet, or Opus? One line on why (by difficulty / number of steps / whether there's a single correct answer).
2. If the tier I'm on is higher than the one you'd pick, say specifically what the extra is being wasted on.
State your call, then do the task by it.
Task: translate this 300-word product announcement into English, keep every number and product name, formal tone.
错了要紧 · 做一遍Matters · onceOpus。一次性的高风险判断 —— 合同条款、架构选型、要发出去的对外声明。这格最值得花。Opus. A one-off, high-stakes call — a contract clause, an architecture choice, a public statement. This is the cell worth paying for.
错了要紧 · 做很多遍Matters · many timesOpus 打样,Sonnet 量产。先用 Opus 把一件做到位、连同判断标准一起沉淀成 Skill(第 5 章),之后交给 Sonnet 照着做。Prototype on Opus, produce on Sonnet. Get one done properly on Opus, distill it with its acceptance criteria into a Skill (Chapter 5), then let Sonnet run the rest against it.
错了不要紧 · 做一遍Doesn't matter · onceSonnet,别想了。这格占你日常的大头,而它的正确答案就是「用默认那档,别在选模型上花时间」。Sonnet, no deliberation. This cell is the bulk of your day, and its right answer is 'use the default and don't spend time picking a model.'
错了不要紧 · 做很多遍Doesn't matter · many timesHaiku。批量分类、格式转换、初筛。这格是唯一一个「换到更便宜的档」能立刻见效的地方。Haiku. Bulk classification, format conversion, first-pass triage. This is the one cell where dropping to a cheaper tier pays off immediately.
两个轴就够:这件活错了要不要紧(纵),和你要做多少遍(横)。跑完对照之后,把你的三行填进对应的格子 —— 大多数人会发现自己长期站在右上角,用最贵的档做量大又不要紧的活。Two axes are enough: how much it matters if this is wrong (vertical) and how many times you'll do it (horizontal). After the comparison, drop your three rows into the matching cell — most people find they've been camped in the top right, running the priciest tier on high-volume work that doesn't matter.
— IV
换个域,路由表长什么样What the Table Looks Like in Other Fields.
Haiku:把英文材料批量翻成中文、生成填空卡片、把一节笔记压成三句话。Sonnet:讲解一个概念、出练习题、批改我的答案。Opus:留给「我卡住了、而且不知道自己卡在哪」的那种题 —— 它更擅长指出你问错了问题。这一栏的经验是:批改用 Sonnet 就够,但「为什么我这个思路是错的」值得上 Opus。Haiku: bulk-translate source material, generate cloze cards, compress a section of notes into three sentences. Sonnet: explain a concept, produce practice problems, grade my answers. Opus: reserve it for 'I'm stuck and I don't know where' — it's better at telling you that you asked the wrong question. The lesson from this field: grading is a Sonnet job, but 'why is my approach wrong' earns Opus.
Haiku:抽取元数据、按关键词初筛一批文献、统一引用格式。Sonnet:读单篇、写摘要、做对比表。Opus:跨十几个来源的综合,尤其是证据互相打架、需要有人指出「这两篇的分歧其实在样本选择上」的时候 —— 便宜的档在这里会把分歧抹平,而抹平了你看不出来。这一栏是全书里最不该省钱的地方。Haiku: extract metadata, keyword-screen a batch of papers, normalize citation formats. Sonnet: read one paper, write the abstract, build a comparison table. Opus: synthesis across a dozen sources, especially when the evidence conflicts and someone needs to point out that 'these two disagree because of how they sampled' — cheaper tiers flatten the disagreement here, and a flattened disagreement is invisible to you. This is the field where saving money costs the most.
Haiku:写测试桩、批量改名、生成样板代码、把报错翻译成人话。Sonnet:日常改功能、写单测、code review 第一遍。Opus:调那种「日志看着一切正常但结果就是不对」的 bug,和要动多个模块的重构。判断标准很土但很准:你自己看着 diff 心里发虚的改动,就该上 Opus。Haiku: write test stubs, bulk rename, generate boilerplate, translate a stack trace into plain language. Sonnet: everyday feature work, unit tests, the first pass of code review. Opus: the bug where 'the logs look fine and the result is still wrong,' and refactors that touch several modules. The test is crude and accurate: if reading the diff makes you uneasy, it was an Opus job.
Haiku:把一堆邮件分类、把会议纪要拆成待办、统一表格格式。Sonnet:写周报、起草对外邮件、把数据整理成一页纸。Opus:定价、取舍、要向上汇报的判断 —— 凡是「这件事我要为它背书」的,别省那点额度。反过来也成立:日报周报这类每周都写、错了改一句就好的东西,用 Opus 是纯浪费。Haiku: sort a pile of email, split meeting notes into action items, normalize table formats. Sonnet: weekly updates, drafting external email, turning data into a one-pager. Opus: pricing, trade-offs, judgments you'll present upward — anything you'll personally stand behind is not the place to save quota. The converse holds too: weekly updates you write every week and fix with one edit are pure waste on Opus.
四张示例路由表。别照抄 —— 抄下来的表和你自己跑出来的表,唯一的区别是你会不会真的照着切。Four sample routing tables. Don't copy them — the only difference between a borrowed table and one you ran yourself is whether you'll actually switch.
— V
边界:三个容易想歪的地方Boundaries: Three Easy Misreads.
1M 上下文不是越满越好A 1M window isn't better full
真需要长上下文的场景很少:一次读完整本书、整个大代码库。什么时候别:把零散资料一股脑全贴进对话 —— 那种该分段,或者交给 Project 的知识库去检索(第 4 章),让它只取相关的几页,而不是每次都重读全部。Genuinely needing the long window is rare: a whole book, a large codebase read at once. When not: dumping scattered material into the chat — that should be chunked, or handed to a Project's knowledge base for retrieval (Chapter 4) so it pulls only the relevant pages instead of re-reading everything each time.
窗口大小不等于装得下多少字Window size isn't how much text fits
当代模型用的新 tokenizer,同一段文字比早期模型产生约 30% 更多 token —— 这是变化量(相对 Opus 4.7 之前),不是水平值,具体幅度随内容而定。4注 4Note 4Claude API Docs · Models overview —— Claude Opus 4.7 起引入的新 tokenizer 沿用至今:同一段文字产生约 30% 更多 token(相对 Opus 4.7 之前的模型;具体幅度随内容而定)。这是变化量,不是水平值。截至 2026-08。Claude API Docs · Models overview — the tokenizer introduced with Claude Opus 4.7 is still in use: the same text produces roughly 30% more tokens than on models predating Opus 4.7 (the exact increase depends on the content). That figure is a change, not a level. As of 2026-08.日常问答感觉不到,整本书、整个代码库这种场景可能比你预期更早碰到边界。Current models use a newer tokenizer that produces roughly 30% more tokens for the same text than earlier ones — that's a change (relative to models predating Opus 4.7), not a level, and the exact figure depends on the content.4注 4Note 4Claude API Docs · Models overview —— Claude Opus 4.7 起引入的新 tokenizer 沿用至今:同一段文字产生约 30% 更多 token(相对 Opus 4.7 之前的模型;具体幅度随内容而定)。这是变化量,不是水平值。截至 2026-08。Claude API Docs · Models overview — the tokenizer introduced with Claude Opus 4.7 is still in use: the same text produces roughly 30% more tokens than on models predating Opus 4.7 (the exact increase depends on the content). That figure is a change, not a level. As of 2026-08. You won't notice in everyday Q&A; on a whole book or a full codebase you may hit the edge sooner than expected.
换模型本身不花钱,但会打断上下文Switching models is free, but it breaks the thread
在同一段对话里切档,前面聊的它还看得见,但它的思路是新起的 —— 复杂问题聊到一半才想起来该升档,不如换个新对话把问题重述一遍,反而更快。Switching mid-conversation keeps the history visible, but the reasoning restarts. On a complex problem, realizing halfway through that you should have gone up a tier is usually better fixed by restating the question in a fresh chat than by flipping the switch in place.
— VI
收束 · 本章验收Sign-off.
挑了三个自己真问过的问题:一个简单、一个中等、一个烧脑。Picked three questions you actually asked: one easy, one medium, one hard.
每个问题在两个模型上各跑了一遍,同一句提示词没改过。Ran each on two models with the same, unmodified prompt.
记下了三样:够不够用、等了多久、有没有追问。Recorded three things: usable or not, how long you waited, whether you followed up.
写出了三行路由表,每行都带一个「因为」。Wrote three routing rows, each with a 'because'.
至少有一行的结论是「降档就够」—— 如果三行全是「上 Opus」,多半是你没敢让便宜的那档真的交卷。At least one row concludes 'the cheaper tier was enough.' If all three say 'go Opus,' you probably never let the cheaper one actually turn in its work.
下一章把这张表钉住 —— Project 的默认模型、固定文件和一段自定义指令,是同一个工作区里的三件事。The next chapter pins this table down — a Project's default model, its fixed files, and its custom instruction are three parts of one workspace.
默认选项替你选了模型,
没替你省钱.
The default picks your model,
not your bill.
Aklman Library
— 讨论Discussion
讨论Discussion.
评论区初始化中…Initializing comments…
01 / 01
没有匹配结果No matches.
换个关键词,或按 Esc 回到页面Try another keyword, or press Esc to return